Über diese Network Engineer Stelle bei AION
About aion
Aion is the enterprise AI platform, a full-stack solution for building, fine-tuning, and deploying AI at scale. Whether an organization is modernizing internal operations, launching AI-powered products, or transforming customer experiences, Aion takes them from concept to production on a single, unified platform.
We work differently than most AI companies: our teams deploy alongside our customers, turning production-ready AI into real business outcomes in weeks, not quarters.
We’re a fast-growing, VC-backed startup led by founders with a track record of successful exits. With teams across the US, UK, and India, we’re building the next generation of enterprise AI and we’re looking for exceptional people to help us scale.
Who You Are
You're an experienced Network Engineer with a strong foundation in designing, implementing, and operating highly available, secure, and scalable network infrastructure. You understand that networking is the backbone of distributed AI systems and have experience building resilient environments that support high-performance computing, cloud-native applications, and large-scale enterprise deployments.
You're comfortable working across cloud and on-premises environments, troubleshooting complex networking issues, and automating network operations. You think beyond connectivity. You optimize for performance, security, observability, and operational excellence.
You're product-minded and collaborative. Whether you're designing multi-region network architecture, optimizing traffic flow for AI workloads, supporting customer deployments, or working closely with platform and infrastructure engineers, you take ownership and focus on delivering reliable infrastructure that enables developers and customers to succeed.
What You'll Do
Network Architecture & Infrastructure
- Design, build, and maintain scalable, secure, and highly available network infrastructure across cloud and on-premises environments.
- Architect enterprise-grade networking solutions supporting AI infrastructure, Kubernetes clusters, GPU workloads, and customer deployments.
- Design network topologies, routing strategies, segmentation, and redundancy to ensure high availability and fault tolerance.
- Drive network architecture decisions while balancing scalability, security, performance, and operational simplicity.
Cloud & Hybrid Networking
- Build and manage networking across AWS, Azure, or GCP environments.
- Design secure hybrid connectivity between cloud infrastructure, customer VPCs, and on-premises environments using VPNs, Direct Connect, ExpressRoute, or similar technologies.
- Implement secure multi-region and multi-cloud networking architectures.
- Support enterprise customer deployments in private cloud and on-premises environments.
Network Operations & Reliability
- Monitor, troubleshoot, and optimize network performance across production environments.
- Investigate latency, packet loss, bandwidth utilization, and connectivity issues impacting AI workloads.
- Build resilient networking with automated failover, redundancy, and disaster recovery strategies.
- Participate in incident response, root cause analysis, and continuous infrastructure improvements.
Network Security
- Implement secure networking practices including firewalls, network segmentation, Zero Trust principles, IDS/IPS, and DDoS protection.
- Configure and maintain load balancers, reverse proxies, WAFs, and secure ingress and egress policies.
- Partner with security teams to ensure compliance with enterprise security standards and customer requirements.
- Support authentication, access control, certificate management, and secure service communication.
Automation & Observability
- Automate network provisioning and configuration using Infrastructure-as-Code and network automation tools.
- Build monitoring dashboards for network health, latency, throughput, utilization, and availability.
- Improve observability through metrics, logging, tracing, and proactive alerting.
- Document network architecture, operational procedures, and best practices.
Requirements
Technical Skills & Experience
- 4+ years of experience designing and managing enterprise or cloud network infrastructure.
- Strong understanding of TCP/IP, routing protocols (BGP, OSPF), VLANs, VPNs, DNS, DHCP, NAT, and load balancing.
- Experience with cloud networking across AWS, Azure, or Google Cloud Platform.
- Hands-on experience with Kubernetes networking, CNI plugins, ingress controllers, and service meshes is highly desirable.
- Experience with firewalls, network security appliances, VPN gateways, and network access control.
- Familiarity with Linux networking, packet analysis, and troubleshooting tools such as tcpdump, Wireshark, iperf, and traceroute.
- Experience with Infrastructure-as-Code tools such as Terraform and Ansible.
- Knowledge of container networking using Docker and Kubernetes.
- Experience with observability tools including Prometheus, Grafana, OpenTelemetry, or equivalent monitoring platforms.
- Familiarity with CI/CD pipelines and infrastructure automation.
- Understanding of high-availability networking, disaster recovery, and business continuity planning.
- Strong troubleshooting and debugging skills in distributed production environments.
Bonus/ Good to Have
- AI & High-Performance Infrastructure:
- Experience supporting networking for AI/ML infrastructure, GPU clusters, or high-performance computing (HPC) environments.
- Familiarity with InfiniBand, RDMA, RoCE, or high-speed data center networking.
- Cloud & Platform Engineering:
- Experience supporting Kubernetes production clusters and cloud-native networking.
- Exposure to service meshes such as Istio or Linkerd.
- Experience with SDN technologies and software-defined networking platforms.
- Automation:
- Experience developing automation using Python, Go, or Bash.
- Knowledge of GitOps workflows and infrastructure lifecycle management.
Benefits
Preferred Attributes:
- Founder-level ownership and bias for action.
- Strong strategic thinking and ability to connect technical decisions to business impact.
- Excellent communication and mentoring skills.
- Thrives in ambiguity, fast-paced environments, and early-stage startup culture.
Why Join aion?
- Work directly with high-pedigree founders shaping technical and product strategy.
- Build infrastructure powering the future of AI compute globally.
- Significant ownership and impact with equity reflective of your contributions.
- Competitive compensation, flexible work options, and wellness benefits