Deskripsi Pekerjaan
Are you ready to be the operational backbone of next-generation AI solutions? TELUS Digital is seeking a highly skilled SRE DevOps Engineer to join our elite AI Automation team in Metro Manila. In this high-impact role, you will be the bridge between cutting-edge GenAI research and production-grade reliability, ensuring our AI/ML solutions perform seamlessly at scale.
You will be responsible for designing, deploying, and maintaining the infrastructure that powers our high-stakes GenAI models. We are looking for an engineer who thrives on automation, site reliability engineering principles, and cloud-native architecture. If you are passionate about observability, latency optimization, and robust CI/CD pipelines for AI workloads, we want to hear from you.
Tanggung Jawab
- Design, build, and maintain scalable infrastructure to support complex GenAI and LLM-based applications.
- Implement and optimize CI/CD pipelines to facilitate rapid, reliable deployment of AI/ML models.
- Establish robust observability, monitoring, and alerting frameworks to ensure 99.9% uptime for AI services.
- Collaborate closely with AI/ML Developers to troubleshoot production issues and optimize model inference latency.
- Manage cloud resources (AWS/GCP/Azure) with a focus on cost-efficiency and security compliance.
- Develop automation scripts and tools to reduce toil and improve operational efficiency across the environment.
- Participate in on-call rotations to resolve critical incidents and maintain system integrity.
Kualifikasi
- Bachelor’s degree in Computer Science, Engineering, or a related technical field.
- Minimum 4+ years of experience in DevOps, Site Reliability Engineering, or a similar platform-focused role.
- Proven expertise in cloud infrastructure management (AWS, GCP, or Azure).
- Strong proficiency in Infrastructure as Code (Terraform, CloudFormation, or Pulumi).
- Deep understanding of containerization and orchestration technologies like Docker and Kubernetes.
- Solid experience with monitoring and logging stacks (Prometheus, Grafana, ELK, or Datadog).
- Scripting and automation skills using Python, Go, or Bash.
- Prior experience supporting AI/ML production environments is highly preferred.