AI Compiler Engineer

EnCharge AI · U.S., Canada, Germany, Norway, India

Software Canada Germany Norway Remote - US Posted Apr 17, 2026

EnCharge AI is a leader in advanced AI hardware and software systems for edge-to-cloud computing. EnCharge’s robust and scalable next-generation in-memory computing technology provides orders-of-magnitude higher compute efficiency and density compared to today’s best-in-class solutions. The high-performance architecture is coupled with seamless software integration and will enable the immense potential of AI to be accessible in power, energy, and space constrained applications. EnCharge AI launched in 2022 and is led by veteran technologists with backgrounds in semiconductor design and AI systems.

About the Role

EnCharge AI is seeking a highly skilled and experienced AI Compiler Engineer to spearhead the efforts in developing and optimizing graph compilers tailored to cutting-edge AI and ML workloads. You will collaborate with hardware architects, and AI researchers to enhance performance, optimize computation graphs, and enable efficient model deployment on EnCharge’s Inference Accelerators.

Responsibilities

Architect, design, and implement optimizations for AI model execution on graph compilers to improve performance, reduce latency, and maximize hardware utilization.
Work closely with ML researchers, hardware engineers, and software developers to design and deploy AI models, understanding and addressing hardware-specific challenges.
Work on performance optimizations for neural network models, such as layer fusion, operator fusion, and graph-level transformations.
Develop compiler optimizations and passes that convert high-level AI models (e.g., from TensorFlow, PyTorch) into intermediate representations (IR).
Implement parsing, semantic analysis, and IR generation for deep learning frameworks.
Research and integrate the latest advancements in compiler design, ML model optimizations, and hardware acceleration into graph compilers.
Provide leadership, mentorship, and technical guidance to a team of engineers focused on graph compiler optimizations.

Qualifications

Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or related field (Ph.D. preferred).
3+ years in compiler development, with a strong focus on AI or ML graph compilers.
Proficiency in AI graph compiler frameworks (e.g., MLIR, Torch-FX)
Solid background in hardware architectures (e.g., GPUs, TPUs, ASICs) and optimization techniques such as fusion, quantization, and tiling.
Familiarity with neural networks operators and code generation.
Strong understanding of intermediate representations, code parsing, and semantic analysis in compiler design.
Proficiency in C++, Python, or other programming languages commonly used in compiler development.
Open-source contributions to AI software frameworks and libraries is a plus
Demonstrated experience leading and mentoring engineering teams with successful project delivery.

EnchargeAI is an equal employment opportunity employer in the United States.

Ready to apply?

Apply to EnCharge AI

EN

EnCharge AI

View all jobs →

EN

AI Research Engineer

EnCharge AI · Canada, Germany, Norway, United States

Apply now

Software Canada Germany Norway Remote - US Posted Mar 18, 2026

EnCharge AI is a leader in advanced AI hardware and software systems for edge-to-cloud computing. EnCharge’s robust and scalable next-generation in-memory computing technology provides orders-of-magnitude higher compute efficiency and density compared to today’s best-in-class solutions. The high-performance architecture is coupled with seamless software integration and will enable the immense potential of AI to be accessible in power, energy, and space constrained applications. EnCharge AI launched in 2022 and is led by veteran technologists with backgrounds in semiconductor design and AI systems.

About the Role

EnCharge AI is looking for an experienced AI Research Engineer to optimize deep learning models for deployment on edge AI platforms. You will work on model compression, quantization strategies, and efficient inference techniques to improve the performance of AI workloads.

Responsibilities

Research and develop quantization-aware training (QAT) and post-training quantization (PTQ) techniques for deep learning models.
Implement low-bit precision optimizations (e.g., INT8, BF16).
Design and optimize efficient inference algorithms for AI workloads, focusing on latency, memory footprint, and power efficiency.
Work with frameworks such as PyTorch, ONNX Runtime, and TVM to deploy optimized models.
Analyze accuracy trade-offs and develop calibration techniques to mitigate precision loss in quantized models.
Collaborate with hardware engineers to optimize model execution for edge devices, and NPUs.
Contribute to research on knowledge distillation, sparsity, pruning, and model compression techniques.
Benchmark performance across different hardware and software stacks.
Stay updated with the latest advancements in AI efficiency, model compression, and hardware acceleration.

Qualifications

Master’s or Ph.D. in Computer Science, Electrical Engineering, or a related field.
Strong expertise in deep learning, model optimization, and numerical precision analysis.
Hands-on experience with model quantization techniques (QAT, PTQ, mixed precision).
Proficiency in Python, C++, CUDA, or OpenCL for performance optimization.
Experience with AI frameworks: PyTorch, TensorFlow, ONNX Runtime, TVM, TensorRT, or OpenVINO.
Understanding of low-level hardware acceleration (e.g., SIMD, AVX, Tensor Cores, VNNI).
Familiarity with compiler optimizations for ML workloads (e.g., XLA, MLIR, LLVM).

EnchargeAI is an equal employment opportunity employer in the United States.

Ready to apply?

Apply to EnCharge AI

EN

EnCharge AI

View all jobs →

EN

LLM Inference Deployment Engineer

EnCharge AI · U.S., Canada, Germany, Norway

Apply now

Software Canada Germany Norway Remote - US Posted Jan 27, 2026

EnCharge AI is a leader in advanced AI hardware and software systems for edge-to-cloud computing. EnCharge’s robust and scalable next-generation in-memory computing technology provides orders-of-magnitude higher compute efficiency and density compared to today’s best-in-class solutions. The high-performance architecture is coupled with seamless software integration and will enable the immense potential of AI to be accessible in power, energy, and space constrained applications. EnCharge AI launched in 2022 and is led by veteran technologists with backgrounds in semiconductor design and AI systems.

About the Role

EnCharge AI is seeking an LLM Inference Deployment Engineer to optimize, deploy, and scale large language models (LLMs) for high-performance inference on its energy efficient AI accelerators. You will work at the intersection of AI frameworks, model optimization, and runtime execution to ensure efficient model execution and low-latency AI inference.

Responsibilities

Deploy and optimize LLMs (GPT, LLaMA, Mistral, Falcon, etc.) post-training from libraries like HuggingFace
Utilize inference runtimes such as ONNX Runtime, vLLM for efficient execution.
Optimize batching, caching, and tensor parallelism to improve LLM scalability in real-time applications.
Develop and maintain high-performance inference pipelines using Docker, Kubernetes, and other inference servers.

Qualifications

Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or related field.
Experience in LLM inference deployment, model optimization, and runtime engineering.
Strong expertise in LLM inference frameworks (PyTorch, ONNX Runtime, vLLM, TensorRT-LLM, DeepSpeed).
In-depth knowledge of the Python programming language for model integration and performance tuning.
Strong understanding of high-level model representations and experience implementing framework-level optimizations for Generative AI use cases
Experience with containerized AI deployments (Docker, Kubernetes, Triton Inference Server, TensorFlow Serving, TorchServe).
Strong knowledge of LLM memory optimization strategies for long-context applications.
Experience with real-time LLM applications (chatbots, code generation, retrieval-augmented generation).

EnchargeAI is an equal employment opportunity employer in the United States.

Ready to apply?

Apply to EnCharge AI

EN

EnCharge AI

View all jobs →

PyTorch Jobs in Norway.

AI Compiler Engineer

AI Research Engineer

LLM Inference Deployment Engineer