Deskripsi Pekerjaan
Atos is a global leader in digital transformation, managing critical systems for clients across every industry. We are currently seeking a dynamic and experienced L3 AWS Cloud Operations Engineer (CloudOps – Level 3) to join our premier Cloud Center of Excellence in Taguig City, Metro Manila.
In this senior role, you will act as the final escalation point and technical authority for our clients' mission-critical AWS environments. Your number one priority will be guaranteeing exceptional high availability, robust security, and optimal performance. You will own the most complex technical incidents, perform deep-dive root cause analyses, and design permanent, automated solutions that prevent recurrence.
This is a high-impact engineering role where you will drive the evolution of our cloud operations. You will work closely with world-class DevOps, Security, and Architecture teams to optimize multi-account AWS organizations, enhance system observability, and implement cutting-edge automation using Infrastructure as Code (IaC). As a trusted advisor, you will also mentor L1 and L2 engineers, elevating the technical expertise of the entire team.
Atos holds premier tier partner status with AWS, providing you with access to the latest technologies, extensive training resources, and a clear path for career advancement. Our tech stack heavily utilizes the Kubernetes ecosystem (Amazon EKS, Helm), serverless architectures (Lambda, Step Functions), and modern observability tools (Grafana, Prometheus, Datadog). We offer a competitive compensation package, hybrid work arrangements, and the opportunity to solve complex challenges for global clients. If you are passionate about cloud technology and want to make a significant impact within a Fortune 500 company, we encourage you to apply.
Tanggung Jawab
- Manage and optimize complex, multi-account AWS cloud infrastructure (EC2, S3, VPC, RDS, ELB, Lambda, EKS) to ensure high availability, scalability, and cost efficiency.
- Function as the L3 escalation point for critical infrastructure incidents, providing expert-level troubleshooting, root cause analysis, and the implementation of permanent fixes.
- Develop, maintain, and champion Infrastructure as Code (IaC) using Terraform and AWS CloudFormation to automate provisioning and configuration management.
- Design, implement, and refine comprehensive monitoring, alerting, and observability strategies (CloudWatch, Grafana, Prometheus, Datadog) to enable proactive system management.
- Lead performance tuning, cost optimization (FinOps), and capacity planning initiatives across the entire AWS landscape.
- Automate operational workflows and tasks using scripting languages (Python, Bash) and CI/CD pipelines (Jenkins, GitLab CI, AWS CodePipeline).
- Enforce cloud security best practices and compliance standards (IAM Policies, Security Groups, KMS, AWS WAF, GuardDuty) in collaboration with the Security team.
- Mentor and provide technical guidance to L1 and L2 Cloud Operations engineers, fostering a culture of continuous learning and knowledge sharing.
Kualifikasi
- Bachelor's degree in Computer Science, Information Technology, or a related quantitative discipline; equivalent practical experience is highly valued.
- Minimum of 5 years of experience in Cloud Operations, DevOps, or Site Reliability Engineering, with at least 3 years focused specifically on Amazon Web Services.
- AWS Professional Level Certification (Solutions Architect Professional or DevOps Engineer Professional) is mandatory.
- Expert-level proficiency in administering Linux and Windows Server operating systems in a cloud environment.
- Deep understanding of cloud networking concepts (DNS, TCP/IP, BGP, VPN, VPC Peering, Transit Gateway, Application/Network Load Balancers).
- Strong proficiency in one or more scripting/programming languages (Python, Bash, PowerShell, Go) with a strong focus on automation.
- Extensive hands-on experience with containerization and orchestration (Docker, Kubernetes / Amazon EKS) and related ecosystem tools (Helm, Istio).
- Excellent analytical and problem-solving skills with a proven track record of resolving complex technical incidents under pressure.