Job description for Site Reliability Engineer (SRE) at PT Tricada Intronik
RESPONSIBILITIES :
- Maintain and improve the availability, reliability, performance, and scalability of production infrastructure and services.
- Monitor infrastructure, applications, and services using observability and monitoring platforms.
- Handle incidents, perform troubleshooting, and conduct root cause analysis to prevent recurring problems.
- Automate repetitive operational activities to reduce manual effort and operational risk.
- Work with Infrastructure, DevOps, Security, and application teams to improve overall service reliability.
- Support capacity planning, performance tuning, and continuous service improvement.
- Monitor production infrastructure and services to ensure agreed levels of availability, reliability, performance, and capacity.
- Define and maintain appropriate SLIs, SLOs, monitoring thresholds, dashboards, and alerting mechanisms.
- Respond to infrastructure and service incidents and coordinate troubleshooting until service restoration.
- Perform root cause analysis (RCA) and identify corrective and preventive actions for recurring incidents.
- Implement and maintain operational automation using scripting and configuration management tools.
- Perform performance analysis, tuning, and capacity planning for servers, VMs, containers, networks, and supporting platforms.
- Collaborate with DevOps and application teams to improve deployment reliability and production readiness.
- Support infrastructure and application changes by evaluating operational risk, monitoring impact, and validating service health.
- Maintain monitoring, logging, observability, and alerting standards across environments.
- Participate in continuous improvement activities based on incidents, operational metrics, performance trends, and capacity utilization.
- Ensure infrastructure operations comply with applicable security, access control, and operational standards.
- Support backup, recovery, resiliency, and disaster recovery validation where required.
- Participate in production support and escalation for critical services when required.
REQUIREMENTS :
- Bachelor’s degree in Computer Science, Information Technology, Computer Engineering, or related discipline.
- Minimum 3 years experience in infrastructure operations, system engineering, DevOps, Site Reliability Engineering, or similar roles.
- Experience supporting production or mission-critical systems is preferred.
- Strong Linux system administration and troubleshooting skills.
- Good understanding of networking fundamentals including TCP/IP, DNS, routing, firewall, load balancing, and HTTP/HTTPS.
- Experience with virtualization platforms such as VMware, Proxmox, OpenStack, or KVM.
- Hands-on experience with containers and orchestration platforms such as Docker and Kubernetes.
- Experience with monitoring and observability tools such as Prometheus, Grafana, Opentelemetry, OpenSearch/ELK, or equivalent.
- Understanding of metrics, logs, traces, alerting, and dashboards.
- Proficient in scripting and automation using Bash, Python, and/or Ansible.
- Familiarity with Infrastructure as Code such as Terraform is a plus.
- Understanding of SLI, SLO, SLA, availability, error rate, latency, and service health indicators.
- Strong troubleshooting capability across OS, network, infrastructure, middleware, container, and application layers.
- Understanding of incident, problem, change, and capacity management concepts.
- Familiarity with ITIL practices is a plus.
- Experience with Kafka, databases, storage systems, or distributed systems is a plus.
- Good documentation and communication skills.
- Able to collaborate effectively with Infrastructure, DevOps, Security, developers, and product/project teams.





