Über diese Virtualization Engineer Stelle bei AION
About aion
Aion is the enterprise AI platform, a full-stack solution for building, fine-tuning, and deploying AI at scale. Whether an organization is modernizing internal operations, launching AI-powered products, or transforming customer experiences, Aion takes them from concept to production on a single, unified platform.
We work differently than most AI companies: our teams deploy alongside our customers, turning production-ready AI into real business outcomes in weeks, not quarters.
We’re a fast-growing, VC-backed startup led by founders with a track record of successful exits. With teams across the US, UK, and India, we’re building the next generation of enterprise AI and we’re looking for exceptional people to help us scale.
Who You Are
You're a systems engineer with deep expertise in virtualization, cloud infrastructure, and compute resource management. You understand how modern virtualization technologies power scalable, secure, and high-performance AI infrastructure across on-premises, hybrid, and cloud environments.
You have experience building virtualization platforms that maximize infrastructure efficiency while providing strong isolation, performance, and operational reliability. You're comfortable working across Linux systems, hypervisors, containers, Kubernetes, networking, and virtualization technologies.
You think beyond infrastructure. You understand how virtualization directly impacts AI model performance, developer productivity, platform scalability, and enterprise deployments. You're excited to solve complex systems problems while building the foundation that powers enterprise AI workloads.
You're a team player who enjoys working across infrastructure, platform engineering, product, and customer success teams to deliver production-ready virtualization solutions for customers globally.
What You'll Do
Virtualization Platform & Infrastructure
- Design and build Aion's virtualization platform for running AI workloads across cloud, hybrid, and on-premises environments.
- Architect secure, scalable, and high-performance virtualization infrastructure supporting enterprise AI deployments.
- Build virtualization services that optimize compute utilization while maintaining workload isolation and reliability.
- Develop reusable platform components that simplify provisioning, lifecycle management, and infrastructure automation.
Compute & Resource Management
- Design intelligent compute scheduling and resource allocation strategies for CPU, memory, storage, networking, and GPU resources.
- Implement virtualization technologies that improve infrastructure utilization while ensuring predictable performance.
- Optimize workload placement, scaling, and resource balancing across distributed clusters.
- Build automation for provisioning, migration, recovery, and lifecycle management of virtualized environments.
Performance & Systems Engineering
- Optimize virtualization stacks for low-latency, high-throughput, and efficient resource utilization.
- Debug complex infrastructure issues across Linux systems, networking, storage, hypervisors, and Kubernetes environments.
- Improve system performance through profiling, benchmarking, and infrastructure tuning.
- Design highly available infrastructure with redundancy, fault tolerance, and disaster recovery capabilities.
Enterprise Platform & Security
- Build secure multi-tenant virtualization environments for enterprise customers.
- Implement networking, storage, identity, and access controls across virtualized infrastructure.
- Design infrastructure supporting customer VPC deployments, air-gapped environments, and on-premises installations.
- Ensure compliance with enterprise security, governance, and operational best practices.
Observability & Engineering Excellence
- Build comprehensive monitoring for virtualization infrastructure, resource utilization, and system health.
- Implement telemetry for compute performance, storage, networking, GPU utilization, and infrastructure reliability.
- Conduct code reviews and establish engineering best practices for infrastructure automation and platform reliability.
- Collaborate closely with AI platform, orchestration, inference, and DevOps teams to continuously improve infrastructure capabilities.
Requirements
Technical Skills & Experience
- 4+ years of experience building virtualization platforms, cloud infrastructure, or systems software.
- Strong understanding of Linux internals, operating systems, and virtualization technologies.
- Experience with hypervisors such as KVM, QEMU, VMware, Xen, or similar platforms.
- Strong proficiency in Golang is preferred. Python, Rust, or C++ experience is a plus.
- Experience with Kubernetes, Docker, container runtimes, and cloud-native infrastructure.
- Knowledge of networking concepts including virtual networking, overlays, load balancing, and service meshes.
- Experience with infrastructure automation using Terraform, Ansible, or Infrastructure-as-Code tools.
- Understanding of distributed storage systems, filesystems, and storage virtualization.
- Experience building highly available distributed infrastructure platforms.
- Familiarity with observability tools such as Prometheus, Grafana, OpenTelemetry, and distributed logging.
- Understanding of authentication, authorization, secrets management, and enterprise security best practices.
- Experience supporting GPU workloads or AI infrastructure is highly desirable.
Bonus/ Good to Have
- Cloud & Hybrid Infrastructure: Experience building platforms across AWS, Azure, GCP, OpenStack, VMware, or hybrid cloud environments.
- Systems Programming: Experience developing low-level systems software, Linux kernel modules, virtualization services, device drivers, or storage systems.
- Enterprise Infrastructure: Experience designing infrastructure for customer VPC deployments, on-premises environments, or regulated enterprise workloads.
- Platform Engineering: Experience building Internal Developer Platforms (IDPs), infrastructure control planes, Kubernetes Operators, or cluster lifecycle management systems.
Benefits
Preferred Attributes:
- Founder-level ownership and bias for action.
- Strong strategic thinking and ability to connect technical decisions to business impact.
- Excellent communication and mentoring skills.
- Thrives in ambiguity, fast-paced environments, and early-stage startup culture.
Why Join aion?
- Work directly with high-pedigree founders shaping technical and product strategy.
- Build infrastructure powering the future of AI compute globally.
- Significant ownership and impact with equity reflective of your contributions.
- Competitive compensation, flexible work options, and wellness benefits