Deskripsi Pekerjaan
Snaphunt is seeking a highly skilled Reliability Engineer to join our innovative engineering team in Singapore. In this critical role, you will be responsible for ensuring the stability, availability, and performance of our mission-critical systems and infrastructure. We are looking for a proactive technical leader who can drive continuous improvement and minimize downtime through robust engineering practices.
As a Reliability Engineer, you will bridge the gap between software development, operations, and quality assurance to build resilient software solutions. You will leverage data-driven methodologies to identify potential failure points, perform deep root cause analysis, and implement long-term solutions that significantly enhance system reliability. If you are passionate about maintaining high standards of technical excellence and want to make a tangible impact on our platform's uptime, we encourage you to apply.
Tanggung Jawab
- Design and implement comprehensive strategies to maximize system uptime and availability.
- Conduct thorough root cause analysis (RCA) for incidents, outages, and system anomalies.
- Develop and maintain key reliability metrics and dashboards to monitor system health in real-time.
- Collaborate closely with cross-functional teams to integrate reliability best practices into the software development lifecycle (SDLC).
- Perform load testing, stress testing, and capacity planning to anticipate and mitigate potential bottlenecks.
- Automate reliability checks and implement Site Reliability Engineering (SRE) principles to improve operational efficiency.
Kualifikasi
- Bachelor’s degree in Computer Science, Systems Engineering, or a related technical field.
- Minimum of 3-5 years of professional experience as a Reliability Engineer, SRE, or Systems Engineer.
- Strong understanding of distributed systems, cloud architecture (AWS, Azure, or GCP), and containerization (Docker, Kubernetes).
- Proficiency in scripting and programming languages such as Python, Go, or Bash.
- Excellent analytical and problem-solving skills with a meticulous attention to detail.
- Familiarity with modern monitoring and logging tools (e.g., Datadog, Prometheus, ELK Stack).