Deskripsi pekerjaan DevOps Engineer PT Karisma Zona Kreatifku (KAZOKKU)
DevOps / Site Reliability Engineer (SRE)
About the Role
Join our Managed Services Team and play a critical role in supporting enterprise clients with large-scale infrastructure. You will be responsible for ensuring the stability, performance, and resilience of diverse environments, ranging from traditional VM-based setups to modern containerized workloads. This role focuses on proactive maintenance, rapid disaster recovery, and end-to-end application support.
Key Responsibilities:
* Infrastructure Management: Maintain and optimize Linux-based environments (RHEL, Rocky Linux, Ubuntu Server).
* Lifecycle Operations: Manage VM provisioning, snapshots, and full-lifecycle troubleshooting across hybrid cloud environments.
* System Reliability: Diagnose and resolve complex Linux issues, including boot failures (GRUB, dracut, initramfs) and storage bottlenecks.
* Disaster Recovery: Execute and refine DR strategies, managing failover flows while meeting RPO/RTO targets.
* Container Orchestration: Deploy and troubleshoot workloads on Kubernetes or Docker Swarm.
* Collaboration: Work closely with the team to follow CI/CD workflows and ensure seamless application request flows from DNS to Database.
Requirements:
Technical Core
* Linux Mastery: Strong experience in production Linux administration and deep knowledge of system logs (journalctl, dmesg).
* Infrastructure & Virtualization: Proficiency in hypervisor platforms (On-premise/Public Cloud) and storage management.
* Networking: Solid understanding of OS-level networking (IP, Routing, DNS, VLAN, Firewall).
* Modern Stack: Familiarity with Kubernetes/Docker and basic CI/CD pipeline execution.
* Troubleshooting: Expert-level ability to analyze end-to-end application request flows (DNS -> Proxy -> App -> DB).
Soft Skills
* Communication: Strong team awareness with a positive and proactive attitude.
* Problem Solver: Exceptional analytical skills and a drive to take on new technical challenges.
* Ownership: A high sense of responsibility and attention to detail.
Nice to Have
* Automation: Scripting skills in Python or Ansible.
* Data Protection: Experience with enterprise backup tools like Acronis or Veeam.
* Monitoring: Hands-on experience with Prometheus, Grafana, or Zabbix.
* Windows Administration: Basic Windows Server management (User MGMT, WinRM, Event Logs).
* Database Ops: Basic backup/restore operations for MySQL or PostgreSQL.
* Quality Assurance: Experience in application integrity checks and testing.