Deskripsi pekerjaan Server Engineer - (GPU Server) PT Aktualisasi Gratia Talenta Indonesia
Job Description:
- Design, deploy, configure, and maintain server infrastructure, particularly GPU-based environments supporting AI and Machine Learning workloads.
- Manage and maintain GPU Servers, including hardware installation, configuration, monitoring, troubleshooting, and performance optimization.
- Support the implementation and operation of AI Infrastructure, including compute, storage, networking, and supporting infrastructure components.
- Install, configure, and maintain server operating systems, drivers, firmware, and system-level software required for GPU and AI workloads.
- Configure and optimize NVIDIA GPU environments, including GPU drivers, CUDA, and related NVIDIA software components.
- Perform server health checks, capacity planning, performance monitoring, and preventive maintenance.
- Troubleshoot hardware and software issues involving servers, GPUs, storage, networking, operating systems, and system components.
- Monitor GPU utilization, temperature, memory usage, system performance, and overall infrastructure availability.
- Support the deployment and configuration of AI/ML computing environments for development, testing, and production workloads.
- Work closely with AI/ML Engineers, DevOps Engineers, Network Engineers, System Administrators, and other technical teams to ensure infrastructure readiness.
- Implement infrastructure security, access control, backup, monitoring, and availability best practices.
- Perform server installation, rack-and-stack activities, cabling, hardware replacement, and infrastructure upgrades when required.
- Maintain technical documentation covering server configurations, infrastructure architecture, operational procedures, and troubleshooting activities.
- Conduct incident investigation, root cause analysis, and corrective actions to maintain system reliability.
- Support infrastructure migration, expansion, and technology refresh projects.
- Ensure infrastructure complies with operational standards, security requirements, and project specifications.
Requirements:
- Minimum D3 in Computer Science, Information Technology, Computer Engineering, Electrical Engineering, or a related field.
- Hands-on experience as a Server Engineer, System Engineer, Infrastructure Engineer, Data Center Engineer, or similar role.
- Proven experience working on GPU Server projects and AI Infrastructure.
- Strong understanding of server hardware, including CPU, RAM, GPU, RAID, storage, power supply, and server components.
- Experience with NVIDIA GPU Servers and NVIDIA GPU technologies.
- Familiarity with CUDA, NVIDIA GPU Drivers, NVIDIA GPU monitoring, and GPU computing environments.
- Good knowledge of Linux/Windows Server operating systems.
- Understanding of server virtualization, storage, networking, backup, monitoring, and high-availability concepts.
- Experience with data center infrastructure and server deployment is highly preferred.
- Strong troubleshooting and problem-solving skills for both hardware and software infrastructure.


