Deskripsi Pekerjaan
At Google, we build the infrastructure that powers the internet. Our data centers are the heartbeat of our global operations, and we are committed to making them the most reliable and efficient in the world.
We are seeking an exceptional Program Manager II, Data Center Incidents and Availability to join our team in Singapore. In this critical role, you will lead the charge in managing incidents that impact our global data center fleet. You will drive end-to-end incident management, from immediate response to deep-dive root cause analysis, and you will collaborate with world-class engineering and operations teams to launch proactive solutions that prevent future disruptions.
This is not just a reactive role; you will be a strategic architect of reliability, shaping how Google anticipates and mitigates risks across millions of servers. You will define the standards for incident response, build robust playbooks, and ensure that lessons learned translate into systemic improvements. Your work will directly impact the availability of Google's core products, from Search and Cloud to YouTube and Maps, used by billions of people globally.
As a leader on the team, you will navigate complex technical challenges, communicate with executive stakeholders during high-pressure events, and mentor teams on best practices in availability and disaster recovery.
The ideal candidate possesses a rare blend of strategic thinking and technical depth. You will partner with engineering teams to define Service Level Objectives (SLOs) and Error Budgets, ensuring our infrastructure meets the demanding needs of our users. Your insights will directly influence the design of future data center architectures, embedding resilience from the ground up.
Google offers a dynamic, collaborative environment where you will work alongside some of the brightest minds in the industry. We provide a world-class benefits package and a culture that values diversity, inclusion, and the continuous pursuit of excellence. If you are ready to take ownership of reliability at a global scale and make a tangible impact on billions of users, we invite you to apply.
Tanggung Jawab
- Lead and coordinate global incident response for critical data center availability events, ensuring timely communication and effective resolution.
- Drive comprehensive post-incident root cause analysis and implement preventative action plans to enhance fleet reliability.
- Develop, maintain, and iterate upon incident management frameworks, runbooks, and escalation procedures.
- Collaborate cross-functionally with Engineering, Site Reliability Engineering (SRE), and Data Center Operations teams to identify and mitigate systemic risks.
- Analyze incident trends and operational metrics to design and launch proactive reliability improvements.
- Provide clear, concise communication to leadership and stakeholders regarding incident status, impact, and long-term mitigation strategies.
- Champion 'availability by design' and disaster recovery best practices across the organization.
- Mentor team members and contribute to a culture of operational excellence and continuous improvement.
Kualifikasi
- Minimum qualifications:
- Bachelor's degree in Computer Science, Engineering, a related technical field, or equivalent practical experience.
- Experience in program or project management within large-scale infrastructure, data center, or cloud computing environments.
- Experience in incident management, crisis management, or site reliability engineering.
- Demonstrated ability to lead cross-functional teams and influence without direct authority.
- Preferred qualifications:
- Experience managing large-scale, high-severity outages or availability incidents.
- Technical knowledge of data center infrastructure components (power, cooling, networking, server hardware).
- Experience developing automation or software tools to improve incident detection or resolution.
- Excellent written and verbal communication skills in English, with experience presenting to executive leadership.