
AI Infrastructure Engineer
Genesis Networks Pte Ltd
Posted 18 hours ago
1. Compute & Cluster Management
Architect, configure, and maintain high-density multi-GPU compute clusters (e.g., NVIDIA HGX/DGX architectures).
Implement and manage container orchestration platforms (Kubernetes, Slurm, or Ray) optimized for AI/ML distributed workloads.
Monitor GPU health, telemetry, utilization, and thermals; minimize idle compute time and prevent single-node bottlenecks.
2. High-Performance Networking & Storage
Design and optimize low-latency, lossless network fabrics supporting distributed training (InfiniBand, RoCE v2, NVLink, spine-leaf topologies).
Configure and scale high-throughput parallel file systems and object storage (e.g., Lustre, GPFS/IBM Spectrum Scale, Ceph, MinIO, NVMe-oF) to feed high-speed data pipelines.
3. Automation & Infrastructure as Code (IaC)
Build and manage automated deployment pipelines using Terraform, Ansible, Helm, or Pulumi.
Maintain standard golden images, Linux OS tuning (kernel parameters, NUMA node binding, GPU drivers, CUDA/cuDNN libraries), and firmware updates.
4. Operations, Observability & Performance
Set up end-to-end monitoring, alerting, and metrics dashboards (Prometheus, Grafana, DCGM exporter, NVIDIA System Management Interface).
Partner with AI/ML engineering teams to diagnose network bottlenecks, NCCL communication latency, and I/O wait states during distributed training jobs.
Lead incident response, root-cause analysis (RCA), and disaster recovery plans for mission-critical AI environments.
Qualifications & Requirements
Technical Competencies
Operating Systems: Deep expertise in Linux systems administration, kernel tuning, and shell scripting (Bash/Python).
Accelerated Compute: Strong understanding of GPU hardware architectures, CUDA runtimes, and PCIe/NVLink topologies.
Orchestration & Workload Scheduling: Hands-on experience with Kubernetes (GPU operator, device plugins) and/or HPC schedulers (Slurm, Run:ai, Ray).
High-Speed Networking: Proven experience with RDMA (RoCE v2 / InfiniBand), PFC (Priority Flow Control), and ECN configurations.
Storage Systems: Familiarity with high-IOPS, low-latency shared storage architectures for AI datasets and model checkpoints.
Automation: Proficiency in Infrastructure as Code (Terraform) and configuration management (Ansible).
Experience & Education
Bachelor’s Degree in Computer Science, Information Technology, Computer Engineering, or equivalent practical experience.
3–6+ years of hands-on experience in infrastructure engineering, high-performance computing (HPC), DevOps, or cloud infrastructure.
Relevant certifications are a plus (e.g., CKA/CKAD, NVIDIA Certified Associate/Professional, AWS/Azure/GCP Solutions Architect).
About Genesis Networks Pte Ltd
Founded in the year 2001, Genesis Networks is a leading IT Service Provider in the Information and Communications Technology (ICT) sector. We are dedicated to building and managing infocomm network infrastructure, implementing business solutions and delivering high value professional services for corporate organizations. This is done so that our customers have an integrated and trusted platform to conduct their businesses.
Product & Services
Genesis is a unique breed of the fast emerging ICT player to provide end-to-end consulting and implementation services to corporate customers. The 5 main areas of focus and specialization are: IT Disaster Recovery and Business Continuity Planning, Managed Services (End-to-end solutions focusing on Enterprise Storage, Security and Infrastructure), IT Manpower Outsourcing, Web Applications (Focusing on Fast Enterprise Search), Security Audit and Consulting. We provide coverage in Asia Pacific. Our key market segments are enterprise corporations and MNCs in this region. We also facilitate companies who wish to set up branch offices overseas, in their IT infrastructure needs and requirements.