Jobs Companies OPSWAT Senior AI Engineer

À propos de ce poste Senior AI Engineer chez OPSWAT

OPSWAT · Sur site · Ho Chi Minh City, Ho Chi Minh City, Vietnam

OPSWAT, a global leader in IT, OT, and ICS critical infrastructure cybersecurity, delivers an end-to-end platform that gives public and private sector organizations and enterprises the critical advantage needed to protect their complex networks, secure their devices, and ensure compliance. Over the last 20 years our commitment to innovative technology has earned the trust of more than 1,700 organizations, governments, and institutions globally, solidifying our role in protecting the world’s critical infrastructure and securing our way of life.

The Position

OPSWAT is committed to harnessing the power of AI to drive meaningful advancements across all areas of our business. We are currently seeking a highly skilled Senior AI Engineer to lead the deployment and optimization of Large Language Models for efficient, high performance inference on GPU hardware.

In this role, you will own the infrastructure layer that turns our AI models into fast, cost effective, production grade services. You will work closely with the AI/ML Team, the Enterprise & Data Team, and Infrastructure to design local LLM serving environments, maximize GPU utilization, and apply model compression techniques such as quantization from FP32 down to FP8/FP4. This is a chance to work at the deep intersection of models, serving software, and hardware within a global cybersecurity company, delivering the performance that powers OPSWAT's AI products at scale.

What You Will Do

  • Local LLM Deployment & Hardware Setup: Design and build local LLM serving environments on GPU hardware, selecting GPU configurations based on VRAM, memory bandwidth, and workload. Install and maintain the full inference stack (NVIDIA drivers, CUDA, cuDNN, inference engines) across single and multi GPU deployments, including tensor and pipeline parallelism.
  • Inference Optimization & Model Compression: Optimize LLM models for efficient GPU serving, balancing latency, throughput, memory, and cost. Apply quantization from FP32 to FP16/BF16, FP8, and FP4 using PTQ methods (GPTQ, AWQ, SmoothQuant) and QAT. Design mixed precision strategies and run calibration to minimize quality loss. Apply complementary techniques including pruning, distillation, and sparsity.
  • Serving Engine & Performance Engineering: Deploy and tune high throughput serving engines (vLLM, TensorRT-LLM, TGI), optimizing KV cache (PagedAttention), continuous batching, and speculative decoding. Leverage optimized kernels (FlashAttention) and FP8 tensor cores. Build benchmarking for TTFT, inter token latency, throughput, GPU utilization, and cost per million tokens.
  • Production Serving Infrastructure: Deploy quantized models with autoscaling, load balancing, and observability. Establish quality regression gates and run A/B tests of quantized vs full precision models on real traffic.
  • Research & Propose Innovative Solutions: exploring and implementing novel inference optimization and model compression techniques.

What We Need From You

  • Bachelor's degree in Computer Science, Software Engineering, or a related field.
  • 5+ years of experience in software engineering, with a focus on ML infrastructure, LLM inference, or model optimization.
  • Hands on experience deploying and serving LLMs on GPU hardware in production.
  • Strong understanding of quantization and model compression (FP8/FP4/INT8/INT4, GPTQ, AWQ, SmoothQuant, QAT).
  • Experience with high throughput inference engines (vLLM, TensorRT-LLM, TGI, llama.cpp).
  • Solid understanding of GPU architecture, CUDA, and the memory bandwidth bound nature of LLM inference.
  • Familiarity with KV cache, continuous batching, PagedAttention, and speculative decoding.
  • Solid understanding of software development methodologies (Agile, Scrum, etc.).
  • Good communication, collaboration, and problem solving skills.
  • Ability to work independently and as part of a team in a fast paced environment.
  • Proficiency in Python; familiarity with C++/CUDA is a strong plus.

It Would Be Nice if You Have

  • Experience writing or tuning custom CUDA / Triton kernels.
  • Experience with multi GPU and distributed inference.
  • Certifications in AWS or Azure architecture.
  • Experience with CI/CD and cloud production deployment (Azure, AWS, or GCP), including GPU backed instances.

OPSWAT is an equal opportunity employer. We celebrate diversity and are committed to providing an environment where equal employment opportunities are extended to all employees and applicants, free of discrimination and harassment of any type. All employment decisions are based on individual qualifications, job requirements, and business needs without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other category protected by federal, state, or local laws.

Recruiting Agencies: we do not accept unsolicited resumes from third party agencies for any of our open positions. To submit resumes for our jobs, there must be a recruiting contract approved by our legal team and endorsed by both parties. We are currently not accepting additional 3rd party agencies at this time.

 

Prêt à postuler chez OPSWAT ?
Postuler chez OPSWAT

À propos de OPSWAT

OPSWAT, a global leader in IT, OT, and ICS critical infrastructure cybersecurity, delivers an end-to-end platform that gives public and private sector organizations and enterprises the critical advantage needed to protect their complex networks, secure their devices, and ensure compliance. Over the last 20 years our commitment to innovative technology has earned the trust of more than 1,700 organizations, governments, and institutions globally, solidifying our role in protecting the world’s critical infrastructure and securing our way of life.

Voir tous les emplois chez OPSWAT →

Emplois similaires

OPSWAT
AI Security Engineer
OPSWAT
⚡ Postuler tôt Ho Chi Minh City, Ho Chi Minh... Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 6 j
OPSWAT
Senior Full-stack Engineer (.NET, AI SDLC)
OPSWAT
⚡ Postuler tôt Ho Chi Minh City, Ho Chi Minh... Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 1 sem.
OPSWAT
Machine Learning Engineer
OPSWAT
⚡ Postuler tôt Ho Chi Minh City, Ho Chi Minh... Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 1 sem.
KNOREX
MLOps Engineer
KNOREX
⚡ Postuler tôt Ho Chi Minh City, Ho Chi Minh,... Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 1 mois
CoverGo
Senior Full Stack Engineer, AI Builder
CoverGo
⚡ Postuler tôt Ho Chi Minh City, Ho Chi Minh,... Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 1 mois
AX
Machine Learning Engineer II
Axon
⚡ Postuler tôt Ho Chi Minh City, Ho Chi Minh... Hybride
● Nouveau 👁 Vu ✓ Postulé il y a 1 mois
NVIDIA
Senior Software Engineer, Metropolis Vision AI
NVIDIA
⚡ Postuler tôt Vietnam, Ho Chi Minh City Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 1 mois
Substance | Level Up by Substance
Agentic AI Engineer (Fractional)
Substance | Level Up by Substance
⚡ Postuler tôt Bengaluru, Karnataka, India Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 3 mois
NVIDIA
Senior Deep Learning Engineer - AI for Wireless Systems
NVIDIA
⚡ Postuler tôt Vietnam, Hanoi Sur site ⚠ 6 mois+
● Nouveau 👁 Vu ✓ Postulé il y a 7 mois

Inscrivez-vous pour des suggestions adaptées aux emplois que vous ouvrez et aux recherches que vous enregistrez.

Plus d’emplois chez OPSWAT

Voir tous les emplois chez OPSWAT →

Postuler maintenant
🤖

Doucement — un instant

JobsRadar a été conçu pour de vraies personnes qui traversent une période difficile dans leur recherche d’emploi — pas pour des requêtes automatisées. Vous cliquez beaucoup trop vite et vous êtes maintenant temporairement bloqué.

Revenez plus tard. Si vous cherchez réellement un emploi, nous sommes de votre côté — agissez simplement comme un être humain.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Prenez une longueur d’avance dans votre recherche d’emploi.

Rejoignez notre canal Telegram pour ce qui vous aide à décrocher le poste — références salariales, le pouls hebdomadaire du marché et les annonces de nouveautés. Pas de spam, que du signal.

Rejoindre le canal — c’est gratuit