Sobre esta vaga de HPC Engineer (AL-FNC260723 003/01) na Xcellink Pte Ltd
We are looking for HPC Engineers to design, deploy and operate high-performance computing infrastructure supporting research, simulation and data-intensive workloads.
You will manage compute clusters, parallel storage, high-speed interconnects and job scheduling environments.
The role requires deep knowledge of HPC hardware and software alongside strong operational skills for cluster administration, performance optimisation and user support.
Key Responsibilities
- Design, deploy, and manage HPC compute clusters and GPU infrastructure
- Administer job scheduling platforms such as Slurm, PBS Pro, or LSF
- Maintain and optimise parallel file systems including Lustre, BeeGFS, or IBM Spectrum Scale (GPFS)
- Manage high-speed interconnect technologies such as InfiniBand and Omni-Path
- Deploy and support scientific computing software, MPI frameworks, compilers, and research applications
- Monitor cluster health, utilisation, job performance, and storage capacity
- Troubleshoot system, application, and infrastructure-related issues
- Implement automation and configuration management using tools such as Ansible
- Collaborate with researchers, faculty members, IT teams, and technology vendors to support research workloads
- Maintain technical documentation, capacity planning records, and operational procedures
- Support security, compliance, and governance requirements within the HPC environment
Requirements
For HPC Engineer (L2)
- Degree in Computer Science, Engineering, Information Technology, or a related discipline
- 2 to 5 years of experience in Linux systems administration, infrastructure operations, or HPC environments
- Experience with cluster management, job schedulers, or large-scale Linux platforms
- Knowledge of storage, networking, and performance tuning concepts
- Familiarity with automation and scripting tools
For Senior HPC Engineer (L3)
- 5 to 10 years of experience supporting HPC, scientific computing, or large-scale distributed infrastructure
- Strong expertise in HPC cluster architecture and operations
- Experience supporting GPU environments and accelerator technologies
- Hands-on experience with parallel file systems and high-speed interconnect fabrics
- Proven track record in performance optimisation, capacity planning, and enterprise-scale operations
Technical Skills
Experience in several of the following areas:
- Linux Administration (Red Hat / Rocky Linux / CentOS)
- Slurm, PBS Pro, LSF
- NVIDIA GPU Platforms
- MPI (OpenMPI, Intel MPI)
- InfiniBand / Omni-Path
- GPFS / Spectrum Scale, Lustre, BeeGFS
- Ansible, Infrastructure Automation
- Singularity / Apptainer Containers
- Spack, Environment Modules, Lmod
- Performance Monitoring and Capacity Management