Sobre este puesto de Systems Engineer en Prime System Solutions
Position summary
The Systems Engineer builds, hardens, and operates the Linux server fleet and core infrastructure services that underpin GPU One (GPUaaS). This role keeps the operating system, networking, and storage layers healthy, secure, and automated so that customer and platform workloads run reliably at scale.
Key responsibilities
· Administer, patch, and harden a large fleet of Linux servers (RHEL, Rocky, or Ubuntu) across bare-metal and cloud environments
· Troubleshoot complex OS, kernel, performance, and hardware issues down to root cause
· Design, configure, and operate TCP/IP networking including routing, VLANs, DNS, DHCP, firewalls, and bonding
· Deploy and manage NFS storage and other shared file systems for high-throughput workloads
· Automate provisioning, configuration, and remediation using Ansible and infrastructure-as-code
· Build and maintain monitoring, alerting, and dashboards with Grafana and related observability tooling
· Manage the operational workflow through ticketing systems such as ServiceNow (SNOW) and JIRA
· Own capacity planning, OS lifecycle, and standardized system build and image management
· Partner with Platform Engineering and SRE on reliability, security, and rollout of new services
· Document standards, runbooks, and procedures, and drive continuous reduction of operational toil
Requirements
Required qualifications
· 5+ years in Linux systems administration or infrastructure engineering at scale
· Strong Linux internals and troubleshooting skills across OS, kernel, storage, and performance
· Solid TCP/IP networking fundamentals and hands-on network troubleshooting
· Hands-on experience operating NFS and other shared storage in production
· Proven experience with configuration management using Ansible
· Experience with monitoring tools such as Grafana and ticketing tools such as ServiceNow and JIRA
· Scripting proficiency in Bash and Python for automation
· Bachelor's degree in computer science, engineering, or equivalent experience
Preferred qualifications
· Experience operating GPU servers, HPC, or AI infrastructure
· Familiarity with Kubernetes, Slurm, or other cluster schedulers
· Exposure to storage technologies such as Lustre, GPFS/Spectrum Scale, or Ceph
· Knowledge of InfiniBand or RDMA and high-performance networking
· Familiarity or working knowledge of using AI coding tools such as Claude, OpenAI, or others to accelerate automation and troubleshooting
· Relevant certifications (RHCSA/RHCE, CCNA) or cloud provider certifications