Deskripsi Pekerjaan
Join Arvion Services as a Site Reliability Engineer and play a pivotal role in building and maintaining the robust infrastructure that powers our global tech platform. We're seeking a passionate professional to architect scalable, fault-tolerant systems that ensure maximum uptime and performance for our BPO operations. In this dynamic role, you'll leverage cutting-edge technologies to automate processes, optimize system reliability, and drive continuous improvement in our cloud-native environment. Your expertise will directly impact millions of users worldwide while working with a collaborative team dedicated to engineering excellence.
This position offers the opportunity to work at the intersection of DevOps, cloud infrastructure, and automation. You'll tackle complex challenges in monitoring, incident response, and capacity planning while implementing best practices for observability and resilience. If you thrive in fast-paced environments and are passionate about creating systems that just work, we invite you to contribute to our mission of delivering exceptional digital solutions.
Tanggung Jawab
- Design, implement, and maintain scalable infrastructure solutions for cloud-based BPO platforms
- Develop and execute automation strategies for deployment, monitoring, and incident response
- Ensure system reliability through proactive monitoring, alerting, and capacity planning
- Collaborate with development teams to integrate reliability requirements into the SDLC
- Lead incident response and root cause analysis for production issues
- Optimize system performance and cost efficiency across cloud environments
- Document infrastructure architecture and operational procedures
- Maintain security compliance and best practices across all systems
Kualifikasi
- Bachelor's degree in Computer Science, Engineering, or related technical field
- 3+ years of experience in Site Reliability Engineering or DevOps roles
- Expertise in cloud platforms (AWS/Azure/GCP) and container orchestration
- Proficiency in automation tools (Ansible, Terraform, Kubernetes)
- Strong scripting skills (Python, Bash, Go)
- Experience with monitoring tools (Prometheus, Grafana, ELK stack)
- Knowledge of CI/CD pipelines and infrastructure-as-code principles
- Ability to work effectively in cross-functional teams