Über diese Hardware Engineer Stelle bei AION
About aion
Aion is the enterprise AI platform, a full-stack solution for building, fine-tuning, and deploying AI at scale. Whether an organization is modernizing internal operations, launching AI-powered products, or transforming customer experiences, Aion takes them from concept to production on a single, unified platform.
We work differently than most AI companies: our teams deploy alongside our customers, turning production-ready AI into real business outcomes in weeks, not quarters.
We’re a fast-growing, VC-backed startup led by founders with a track record of successful exits. With teams across the US, UK, and India, we’re building the next generation of enterprise AI and we’re looking for exceptional people to help us scale.
Who You Are
You're a hardware engineer passionate about building the infrastructure that powers large-scale AI systems. You understand the complexities of modern compute platforms including servers, GPUs, networking, storage, and high-performance computing environments that enable enterprise AI workloads.
You have experience designing, integrating, validating, and optimizing hardware platforms for performance, scalability, and reliability. You're comfortable working across compute infrastructure, hardware bring-up, system diagnostics, firmware interactions, and data center deployments.
You're product-minded and understand how hardware decisions directly impact AI performance, infrastructure efficiency, and customer experience. You enjoy solving complex systems challenges and collaborating with software, platform, and infrastructure teams to build production-ready AI infrastructure.
What You'll Do
Hardware Platform Design & Integration
- Design and build hardware platforms optimized for AI training and inference workloads.
- Evaluate, integrate, and validate servers, GPUs, networking equipment, storage systems, and accelerator hardware.
- Develop scalable hardware architectures supporting enterprise AI deployments across cloud, hybrid, and on-premises environments.
- Collaborate with software and platform engineering teams to optimize hardware and software integration.
System Performance & Optimization
- Optimize compute, memory, storage, networking, and GPU performance for AI workloads.
- Benchmark and analyze system performance under production-scale AI deployments.
- Identify hardware bottlenecks and implement improvements to maximize throughput, latency, and infrastructure efficiency.
- Validate hardware compatibility across multiple AI frameworks and deployment environments.
Infrastructure Reliability
- Build reliable and fault-tolerant hardware infrastructure capable of supporting mission-critical AI workloads.
- Develop hardware validation, diagnostics, stress testing, and failure analysis processes.
- Support hardware lifecycle management including provisioning, upgrades, maintenance, and replacement strategies.
- Collaborate with vendors and partners to evaluate emerging hardware technologies.
Deployment & Operations
- Support hardware deployment across enterprise customer environments and internal infrastructure.
- Develop automation and operational procedures for hardware provisioning, monitoring, and troubleshooting.
- Ensure infrastructure meets enterprise standards for availability, security, and operational excellence.
- Work closely with customer-facing teams to resolve deployment and infrastructure challenges.
Engineering Excellence
- Establish best practices for hardware validation, documentation, testing, and operational readiness.
- Conduct technical reviews and contribute to infrastructure architecture decisions.
- Collaborate across hardware, software, platform, and AI engineering teams to continuously improve system performance and reliability.
Requirements
Technical Skills & Experience
- 4+ years of experience in hardware engineering, systems engineering, or data center infrastructure.
- Strong understanding of server architecture, CPUs, GPUs, memory, storage, networking, and hardware components.
- Experience with AI infrastructure including NVIDIA GPUs, AMD accelerators, or other high-performance computing platforms.
- Knowledge of PCIe, NVLink, InfiniBand, Ethernet, storage architectures, and high-speed interconnects.
- Experience with Linux systems, hardware diagnostics, firmware updates, and system bring-up.
- Familiarity with rack-scale deployments, data center operations, and hardware lifecycle management.
- Understanding of thermal design, power management, and hardware reliability engineering.
- Experience with infrastructure monitoring and hardware observability tools.
- Knowledge of automation and scripting using Python, Bash, or similar languages is preferred.
- Familiarity with Kubernetes, virtualization, or cloud infrastructure is a plus.
- Experience supporting AI infrastructure for training or inference workloads is highly desirable.
Bonus/ Good to Have
- AI Infrastructure: Experience deploying and optimizing infrastructure for large language models, deep learning, or distributed AI workloads.
- Hardware Validation: Experience in hardware qualification, manufacturing validation, stress testing, and performance benchmarking.
- Enterprise Infrastructure: Experience supporting enterprise data centers, on-premises deployments, or hybrid cloud infrastructure.
- Platform Engineering: Exposure to infrastructure automation, provisioning systems, Infrastructure-as-Code, or hardware management platforms.
Benefits
Preferred Attributes:
- Founder-level ownership and bias for action.
- Strong strategic thinking and ability to connect technical decisions to business impact.
- Excellent communication and mentoring skills.
- Thrives in ambiguity, fast-paced environments, and early-stage startup culture.
Why Join aion?
- Work directly with high-pedigree founders shaping technical and product strategy.
- Build infrastructure powering the future of AI compute globally.
- Significant ownership and impact with equity reflective of your contributions.
- Competitive compensation, flexible work options, and wellness benefits