Deskripsi Pekerjaan
HAVI is seeking a passionate and experienced Site Reliability Engineer (Azure) to join our dynamic team in Kuala Lumpur. As a key member of our infrastructure and operations group, you will focus on designing and implementing automation solutions that enhance system reliability, scalability, and performance while reducing operational toil. You will work closely with development teams to build and maintain robust, cloud-native systems on Microsoft Azure, ensuring high availability and seamless user experiences. This role offers a unique opportunity to drive reliability engineering best practices, automate operational workflows, and contribute to the continuous improvement of our digital platforms. If you are a problem-solver with a deep understanding of Azure services and a commitment to operational excellence, we want to hear from you.
Tanggung Jawab
- Design, develop, and maintain automation scripts and tools to automate deployment, monitoring, and incident response processes on Azure infrastructure.
- Collaborate with software engineering teams to ensure applications are built with reliability, scalability, and performance in mind from the ground up.
- Implement and manage monitoring, alerting, and logging solutions to proactively identify and resolve system issues.
- Conduct post-incident reviews and implement preventive measures to reduce mean time to recovery (MTTR).
- Manage and optimize Azure resources (compute, storage, networking) to control costs while meeting performance and availability targets.
- Participate in on-call rotations to provide operational support for critical systems and respond to incidents.
- Drive the adoption of infrastructure as code (IaC) practices using tools such as Terraform or ARM templates.
- Create and maintain comprehensive documentation for system architecture, runbooks, and reliability processes.
Kualifikasi
- Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent practical experience.
- 5+ years of experience in Site Reliability Engineering, DevOps, or a similar role with a strong focus on Azure.
- Proven expertise in Azure services (Azure Kubernetes Service, Azure DevOps, Azure Monitor, Azure Functions, etc.).
- Strong proficiency in scripting languages such as Python, PowerShell, or Bash.
- Hands-on experience with infrastructure as code tools (Terraform, ARM templates) and CI/CD pipelines.
- Deep understanding of containerization technologies (Docker, Kubernetes) and microservices architecture.
- Excellent troubleshooting and analytical skills, with a proactive approach to system reliability and performance optimization.
- Relevant Azure certifications (e.g., Azure Administrator, Azure DevOps Engineer, or Azure Solutions Architect) are highly desirable.