Deskripsi Pekerjaan
Join Atos as an L3 Site Reliability Engineer and become a cornerstone of our digital infrastructure. This role empowers you to drive proactive system reliability through SLI/SLO management, automation implementation, and performance optimization. You'll tackle complex challenges head-on, ensuring our services meet the highest standards of availability and efficiency. Collaborate with cross-functional teams to architect resilient systems, automate repetitive tasks, and innovate solutions that scale with our growing demands. At Atos, you'll thrive in a dynamic environment where your expertise directly impacts millions of users worldwide. This position offers opportunities to master cutting-edge technologies while contributing to mission-critical projects that shape the future of IT.
Tanggung Jawab
- Monitor and maintain system reliability, performance, and scalability across cloud and on-premise environments
- Implement automation tools and scripts to reduce manual intervention and improve operational efficiency
- Define, track, and optimize Service Level Indicators (SLIs) and Objectives (SLOs)
- Troubleshoot and resolve complex system incidents with root cause analysis
- Collaborate with development teams to implement DevOps best practices
- Design and maintain monitoring dashboards for real-time system health
- Participate in on-call rotation and incident response protocols
Kualifikasi
- Bachelor's degree in Computer Science, Engineering, or related field
- 3+ years of experience in site reliability engineering or IT operations
- Proficiency in scripting languages (Python, Bash, PowerShell)
- Strong knowledge of cloud platforms (AWS, Azure, or GCP)
- Experience with monitoring tools (Prometheus, Grafana, Datadog)
- Familiarity with CI/CD pipelines and infrastructure-as-code (Terraform, Ansible)
- Excellent problem-solving skills with ability to troubleshoot complex distributed systems