Deskripsi Pekerjaan
Join our dynamic team at Cambridge University Press & Assessment as a Site Reliability Engineer and play a crucial role in ensuring the reliability, performance, and scalability of our English Technology platforms, applications, services, and websites. You will be at the forefront of designing, implementing, and maintaining robust systems that support our global educational mission.
As a Site Reliability Engineer, you will bridge the gap between development and operations by implementing automation, monitoring, and best practices to enhance system reliability. You will work with cutting-edge technologies to build resilient infrastructure, troubleshoot complex issues, and continuously improve our systems' performance. Your expertise will directly impact the educational experience of millions of users worldwide.
We offer a collaborative environment where innovation is encouraged, and your contributions make a tangible difference. If you are passionate about building and maintaining high-availability systems and want to be part of an organization dedicated to excellence in education, we invite you to apply.
Tanggung Jawab
- Design, implement, and maintain robust, scalable, and reliable systems for English Technology platforms and applications
- Develop automation tools and scripts to streamline deployment, monitoring, and maintenance processes
- Monitor system performance, identify bottlenecks, and implement solutions to enhance reliability
- Collaborate with development teams to ensure smooth integration of new features and services
- Participate in on-call rotation and respond to incidents with urgency and precision
- Document system architecture, processes, and best practices for knowledge sharing
- Continuously evaluate and implement new technologies and methodologies to improve system reliability
Kualifikasi
- Bachelor's degree in Computer Science, Engineering, or related field
- 3+ years of experience in site reliability engineering, DevOps, or similar role
- Strong proficiency in scripting languages (Python, Bash, etc.) and automation tools
- Experience with cloud platforms (AWS, Azure, or GCP) and containerization technologies
- Knowledge of monitoring tools (Prometheus, Grafana, etc.) and logging systems
- Familiarity with CI/CD pipelines and infrastructure as code (Terraform, Ansible, etc.)
- Excellent problem-solving skills and ability to work under pressure
- Strong communication skills and ability to collaborate effectively with cross-functional teams