Beranda Loker Detail
B
Information & Communication Technology 🏢 Full Time ⭐️ Terverifikasi

Site Reliability Engineer, Edge Services - Traffic Infrastructure

ByteDance
Singapore
Estimasi Gaji
SGD 144.000 – SGD 216.000
Live Update
10 Mei 2026
Batas Akhir
10 Mei 2027

Deskripsi Pekerjaan

ByteDance is a global technology company operating a range of content platforms that inspire creativity and enrich life. Our suite of products includes TikTok, Douyin, Lark, and CapCut. We are looking for a highly skilled Site Reliability Engineer, Edge Services - Traffic Infrastructure to join our world-class team in Singapore.

In this role, you will be a key member of the Content Distribution Networks (CDN) team, responsible for ensuring the reliability, performance, and efficiency of our globally distributed edge services platform. You will work on a hybrid platform that seamlessly integrates commercial CDN vendors with ByteDance's proprietary infrastructure, serving billions of users worldwide with low latency and high availability.

You will tackle complex challenges in massive-scale distributed systems, traffic engineering, and real-time incident management. Your work will directly influence the architecture of our hybrid edge platform, balancing cost, performance, and resilience across hundreds of Points of Presence (PoPs) worldwide. From debugging kernel-level network anomalies to architecting multi-region failover systems, this role offers a diverse and deeply technical landscape.

As an SRE on this team, you will design and build automation frameworks, capacity planning models, and global traffic steering policies. Our data-driven culture empowers engineers to take ownership of their services from architecture to operation. By joining ByteDance, you will work alongside some of the brightest minds in the industry on infrastructure that powers billions of daily active users.

If you are passionate about building robust systems, automating operational workflows, and diving deep into traffic infrastructure, we want to hear from you. Join us to push the boundaries of what's possible in edge computing and content delivery.

Tanggung Jawab

  • Design, build, and maintain highly reliable and scalable edge services and traffic infrastructure to support ByteDance's global products.
  • Develop and operate a hybrid CDN platform, analyzing performance metrics and making data-driven decisions to optimize content delivery across commercial and proprietary vendors.
  • Lead incident response and root cause analysis for production outages affecting traffic, ensuring swift resolution and implementing preventative measures.
  • Automate operational tasks, including deployment, monitoring, capacity planning, and failover processes, to improve efficiency and reduce manual intervention.
  • Collaborate with software engineering teams to define Service Level Objectives (SLOs) and Service Level Agreements (SLAs) for edge services, and build dashboards to track reliability.
  • Conduct performance and stress testing of traffic systems, identifying bottlenecks and implementing optimizations for latency, throughput, and cost efficiency.
  • Participate in an on-call rotation to provide 24/7 support for critical production incidents, championing a culture of blameless post-mortems and continuous improvement.
  • Drive automation and tooling improvements for traffic routing, DNS management, and global load balancing to ensure resilient edge operations.

Kualifikasi

  • Experience: Minimum 5 years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering, with a strong focus on edge services, CDNs, or traffic management.
  • Technical Skills: Deep understanding of CDN architectures, DNS, HTTP/HTTPS protocols, TLS, and global load balancing (GLB) techniques.
  • Programming: Proficiency in at least one programming language such as Go, Python, or C++ for infrastructure automation and tooling.
  • Systems Knowledge: Strong experience with Linux systems internals, networking (TCP/IP, BGP), and performance tuning in a high-traffic environment.
  • Cloud/Orchestration: Hands-on experience with containerization (Docker, Kubernetes) and cloud platforms (AWS, GCP, Azure) or large-scale bare-metal environments.
  • Monitoring: Experience building observability solutions using tools like Prometheus, Grafana, Elasticsearch, and distributed tracing.
  • Problem Solving: Exceptional analytical and troubleshooting skills, with the ability to diagnose complex issues across a distributed stack.
  • Collaboration: Excellent communication skills and experience working in a fast-paced, cross-functional environment with global teams.

Keahlian yang Dibutuhkan

CDN Site Reliability Engineering Infrastructure DevOps Cloud Computing Linux Kubernetes Docker Python Go TCP/IP DNS HTTP Traffic Engineering Edge Computing Incident Management Automation Monitoring Prometheus Grafana AWS GCP Azure Singapore ByteDance

Siap Mengambil Tantangan Ini?

Pastikan resume Anda sudah siap. Kirimkan lamaran Anda sekarang sebelum tanggal deadline.

Lamar Sekarang

Lowongan Terkait

Rekomendasi pekerjaan serupa untuk Anda

Lihat Semua