Deskripsi Pekerjaan
Join ByteDance's global data center operations team and play a pivotal role in supporting our hypergrowth trajectory. As a Data Center Monitoring System Engineer, you'll architect and maintain the monitoring infrastructure that powers our hyperscale data centers – the backbone serving billions of users worldwide. This role combines deep technical expertise with strategic impact, ensuring mission-critical infrastructure maintains optimal performance, security, and scalability in our rapidly evolving ecosystem. You'll collaborate with cross-functional teams to implement cutting-edge monitoring solutions, drive automation initiatives, and establish industry best practices for data center resilience. ByteDance offers a dynamic environment where innovation thrives, and your work directly influences the reliability of our global platforms.
Our data center operation team is the engine behind our technological expansion, managing complex infrastructure that demands proactive monitoring and rapid response capabilities. In this role, you'll be at the forefront of detecting and resolving potential issues before they impact user experience, while continuously optimizing monitoring systems to support our exponential growth. This position offers unparalleled exposure to large-scale distributed systems and the opportunity to shape the future of data center monitoring at one of the world's leading technology companies.
Tanggung Jawab
- Design, implement, and maintain comprehensive monitoring systems for hyperscale data center infrastructure
- Develop automation scripts and tools for real-time monitoring, alerting, and incident response
- Proactively identify and resolve performance bottlenecks, security vulnerabilities, and system failures
- Collaborate with infrastructure and network teams to define monitoring requirements and SLAs
- Create and maintain documentation for monitoring systems, runbooks, and operational procedures
- Lead continuous improvement initiatives for monitoring dashboards and visualization tools
- Ensure 24/7 system availability through on-call rotation and rapid incident resolution
- Conduct capacity planning and forecasting to support future infrastructure scaling
Kualifikasi
- Bachelor's degree in Computer Science, Engineering, or related technical field
- 3+ years of experience in data center operations or system administration
- Expertise with monitoring tools such as Nagios, Zabbix, Prometheus, or Grafana
- Strong proficiency in scripting languages (Python, Bash, PowerShell)
- Deep understanding of network protocols (TCP/IP, DNS, HTTP) and virtualization technologies
- Experience with cloud platforms (AWS, Azure, GCP) and container orchestration
- Ability to troubleshoot complex distributed systems and root cause analysis
- Strong communication skills with ability to collaborate across global teams