À propos de ce poste Senior HPC Cloud Engineer chez Accenture Federal Services
Accenture Federal Services is seeking a Senior Cloud Engineer / HPC specialist to join our team and support our client at Hill AFB in Utah. This senior individual-contributor role owns the design and operation of high-performance compute infrastructure for mission workloads, from cluster architecture through GPU-accelerated job scheduling. You will independently make cluster sizing and architecture decisions, mentor engineers, and partner with data platform and AI/ML teams to ensure infrastructure meets mission demand.
Success in This Role Means:
- HPC clusters are right-sized and cost-efficient, scaling compute and GPU resources to workload demand
- Job scheduling (Slurm, PBS, AWS Batch/ParallelCluster) is reliable and self-service, minimizing manual intervention
- GPU-accelerated workloads run efficiently, with proactive CUDA/cuDNN tuning and resource allocation
- Data pipeline teams (Kafka, Airflow, Spark, EMR) have a stable, well-documented compute foundation
- Security and accreditation requirements for classified HPC workloads are met without impeding mission delivery
What you’ll do:
- Design, size, tune, and operate HPC clusters for compute-intensive workloads
- Own HPC job scheduling infrastructure (Slurm, PBS, or equivalent)
- Architect/manage AWS ParallelCluster and/or AWS Batch for cloud-based HPC provisioning
- Design for parallel computing paradigms (MPI, shared-memory, distributed task orchestration)
- Configure industry GPU hardware/software for accelerated workloads
- Tune CUDA, cuDNN, and GPU libraries for scientific computing/data processing
- Optimize scheduling/resource allocation across CPU/GPU node pools
- Integrate HPC compute with data pipelines (Kafka, Airflow, Spark) and AWS data services (EMR, Redshift, Glue)
- Ensure HPC infrastructure meets security controls/accreditation for classified environments
- Mentor engineers building HPC/GPU-compute familiarity
- Document cluster architecture decisions, runbooks, and operational procedures
What you’ll need:
- 5 years’ cloud engineering/infrastructure experience, including 2 years focused on HPC cluster design/operations
- Experience with HPC job scheduling systems (Slurm, PBS, or equivalent)
- Proficiency with AWS ParallelCluster and/or AWS Batch for cloud-based HPC provisioning
- Understanding of parallel computing paradigms (e.g.: MPI, shared-memory, distributed task orchestration)
- Working knowledge of DevOps practices (e.g.: CI/CD, infrastructure-as-code, GitOps)
- Scripting proficiency in Python, Bash, or PowerShell
Bonus Points if you have:
- Bachelor’s in Computer Science, Computational Science, Engineering, or related field (certifications considered in lieu)
- Hands-on experience with industry GPU hardware/software ecosystems
- Familiarity with CUDA, cuDNN, and GPU-accelerated libraries
- Experience designing/operating data pipelines (Kafka, Airflow, Spark)
- Familiarity with AWS data services (EMR, Redshift, Glue)
- AWS certifications (Solutions Architect or Advanced Networking)
Clearance:
- Must have an active Secret clearance; Top secret preferred
Work Environment & Culture Fit
- Comfortable in fast-paced, dynamic environments with evolving compute demands
- Strong Agile framework familiarity (sprints, standups, retrospectives)
- Solution ownership mindset—responsible for cluster efficiency/reliability
- Fail-fast, fail-forward mentality—iterates on cluster tuning
- Growth-oriented—invested in mentoring engineers
Who Thrives in This Role
- Extreme ownership mindset—accountable for cluster efficiency/workload reliability
- Comfortable with ambiguity—translates vague mission needs into concrete architectures
- Detail-oriented on performance—tunes resource allocation to avoid over-provisioning
- Strong collaborator—partners with data engineering and AI/ML teams
- Self-directed—proactively identifies scaling/tuning opportunities
Why This Role Matters
Directly determines whether mission workloads—simulation, modeling, AI/ML pipelines—have the capacity and performance needed. Work is visible at every level, from data engineers and AI/ML teams to mission stakeholders relying on timely simulation and analysis results.
As required by local law, Accenture Federal Services provides reasonable ranges of compensation for hired roles based on labor costs in the states of California, Colorado, Connecticut, Hawaii, Illinois, Maine, Maryland, Massachusetts, Minnesota, New Jersey, New York, Ohio, Vermont, Virginia, Washington, and the District of Columbia. The base pay range for this position in these locations is shown below. Compensation for roles at Accenture Federal Services varies depending on a wide array of factors, including but not limited to office location, role, skill set, and level of experience. Accenture Federal Services offers a wide variety of benefits. You can find more information on benefits here. We accept applications on an on-going basis and there is no fixed deadline to apply.