Deskripsi Pekerjaan
Are you a seasoned SRE professional looking to make an impact on global financial infrastructure? The London Stock Exchange Group (LSEG) is seeking a visionary Technical Lead for Site Reliability Engineering to join our team in Taguig City. In this pivotal role, you will be at the forefront of driving reliability, observability, and security across our massive AWS cloud platforms.
As a Technical Lead, you won't just be maintaining systems; you will be an architect of change. You will champion SRE best practices, mentor engineering teams, and implement cutting-edge automation to ensure our platforms meet the rigorous demands of the global financial market. If you are passionate about building resilient systems and thrive in a high-stakes, collaborative environment, we want to hear from you.
Tanggung Jawab
- Lead the design and implementation of highly available, scalable, and secure AWS infrastructure.
- Drive the adoption of SRE principles, including error budgets, incident response, and capacity planning.
- Architect advanced observability solutions to proactively detect and resolve platform latency and performance bottlenecks.
- Collaborate with cross-functional development teams to integrate CI/CD pipelines and automated testing into the deployment lifecycle.
- Lead post-incident reviews and implement blameless engineering practices to prevent recurrence of systemic issues.
- Mentor junior engineers and foster a culture of operational excellence and technical innovation.
- Ensure strict adherence to financial security standards and compliance protocols across all cloud environments.
Kualifikasi
- Bachelor’s degree in Computer Science, Engineering, or a related technical field.
- Minimum of 7+ years of experience in SRE, DevOps, or Software Engineering roles.
- Expert-level proficiency in AWS cloud services (EKS, RDS, Lambda, VPC, IAM).
- Proven experience with Infrastructure as Code (IaC) tools such as Terraform or CloudFormation.
- Strong proficiency in scripting/programming languages such as Python, Go, or Bash.
- Deep understanding of observability tools (e.g., Datadog, Prometheus, Grafana, or Splunk).
- Expertise in managing large-scale distributed systems in a high-traffic or regulated industry.
- Strong leadership and communication skills, with the ability to influence technical roadmaps.