Deskripsi Pekerjaan
Are you a seasoned Site Reliability Engineer passionate about building rock-solid Linux infrastructure? KMC Solutions is seeking a Senior SRE to join our high-performing team. In this role, you will be the backbone of our platform operations, ensuring 99.99% availability, scalability, and performance for our large-scale data center environments.
We are looking for a hands-on engineer who thrives on automating manual tasks, debugging complex system-level bottlenecks, and implementing site reliability best practices. You will work in a fully remote setup, collaborating with global teams to optimize our Linux-based stacks. If you are a proponent of Infrastructure as Code (IaC) and have an obsession with system stability, we want to hear from you.
Tanggung Jawab
- Design, deploy, and manage scalable Linux-based infrastructure to support mission-critical applications.
- Automate operational tasks using scripting languages (Python, Bash, or Go) to minimize toil.
- Proactively monitor system performance, identify bottlenecks, and implement long-term optimization solutions.
- Collaborate with development teams to integrate CI/CD pipelines and improve deployment reliability.
- Lead incident response efforts, conduct root cause analysis (RCA), and implement preventative measures.
- Manage configuration management tools (Ansible, Puppet, or SaltStack) across multi-server environments.
- Ensure security compliance and hardening of Linux server environments.
Kualifikasi
- Bachelor’s degree in Computer Science, Engineering, or equivalent professional experience.
- 5+ years of experience in Linux Systems Administration and Site Reliability Engineering.
- Deep understanding of Linux kernel, networking (TCP/IP, DNS, Load Balancing), and storage subsystems.
- Expertise in container orchestration tools such as Kubernetes, Docker, or OpenShift.
- Proficiency in Infrastructure as Code (IaC) tools like Terraform or CloudFormation.
- Strong background in monitoring and observability tools (Prometheus, Grafana, ELK stack).
- Excellent problem-solving skills with a mindset focused on automation and proactive system health.