Deskripsi Pekerjaan
Are you ready to shape the future of AI-driven search and recommendation systems? ByteDance is seeking a talented Site Reliability Engineer (SRE) to join our innovative team. We are dedicated to building scalable, reliable, and high-performance products that serve millions of global users. As an SRE in this role, you will bridge the gap between development and operations, ensuring our AI infrastructure remains robust and efficient. You will work in a fast-paced environment leveraging cutting-edge technologies to deploy, maintain, and scale machine learning models that drive our core business.
In this position, you will be responsible for the reliability and performance of our AI services. You will collaborate closely with machine learning engineers to optimize model serving latency and cost, while maintaining the highest standards of system availability. If you have a passion for automation, system architecture, and AI technologies, we want to hear from you. Join us in revolutionizing how users discover content through intelligent recommendation algorithms.
Tanggung Jawab
- Design, build, and maintain highly available, scalable, and efficient AI infrastructure and services.
- Implement and manage CI/CD pipelines specifically tailored for deploying machine learning models to production.
- Monitor system health, performance metrics, and error rates to ensure 99.99% uptime for critical AI applications.
- Collaborate with ML engineers to optimize model serving performance and resource utilization.
- Conduct thorough incident management and root cause analysis to drive continuous improvement in system resilience.
- Automate operational tasks and deployment processes to increase engineering efficiency and reduce manual toil.
Kualifikasi
- Bachelor’s or Master’s degree in Computer Science, Software Engineering, or a related technical field.
- 3+ years of professional experience as a Site Reliability Engineer, DevOps Engineer, or SDET.
- Strong proficiency in programming languages such as Python, Go, or Java.
- Deep understanding of containerization (Docker) and orchestration tools like Kubernetes.
- Experience with cloud platforms (AWS, GCP, or Azure) and their specific AI/ML services.
- Familiarity with AI/ML frameworks (TensorFlow, PyTorch) and serving platforms (Triton, TorchServe).