Deskripsi Pekerjaan
Are you passionate about building and scaling some of the most complex, high-traffic systems in the world? ByteDance is seeking a high-caliber Site Reliability Engineer to join our Traffic Platform team in Singapore. As we continue to scale our global platforms, our infrastructure needs to be resilient, performant, and highly automated.
In this role, you will be at the heart of our engineering operations, working on massively distributed systems that handle billions of requests daily. You will design, develop, and maintain the traffic management layers that ensure our services remain stable and responsive under intense pressure. You will leverage your expertise in networking, load balancing, and observability to optimize our infrastructure while championing SRE best practices across the engineering organization.
Tanggung Jawab
- Design and maintain highly available and scalable traffic management infrastructure, including load balancers, proxies, and global traffic routing systems.
- Automate operational tasks, deployments, and infrastructure management to minimize manual intervention and reduce toil.
- Proactively monitor system performance and identify bottlenecks, implementing robust observability solutions.
- Lead incident response for critical traffic-related issues, conducting thorough post-mortems and driving long-term architectural improvements.
- Collaborate with cross-functional teams to integrate new services into the platform while maintaining strict uptime SLAs.
- Develop and implement strategies for capacity planning to ensure our infrastructure keeps pace with rapid user growth.
- Contribute to the culture of reliability by mentoring team members and improving engineering standards.
Kualifikasi
- Bachelor’s or Master’s degree in Computer Science, Software Engineering, or a related technical field.
- 3+ years of experience in SRE, DevOps, or Software Engineering roles focused on large-scale distributed systems.
- Strong proficiency in programming languages such as Go, Python, or C++.
- Deep understanding of network protocols (TCP/IP, HTTP/HTTPS, DNS, TLS) and traffic management concepts.
- Hands-on experience with containerization and orchestration technologies like Kubernetes and Docker.
- Proven track record of managing large-scale infrastructure on public or private cloud platforms.
- Excellent analytical, troubleshooting, and communication skills in a fast-paced environment.