Jobs Companies NVIDIA Senior LLM Agents Architect

About this Senior LLM Agents Architect role at NVIDIA

NVIDIA · Onsite · Israel, Yokneam

We don't just build the hardware and software that powers the AI revolution — we are building the AI that designs the next generation of both. Our team sits at the intersection of inference software and GPU architecture, creating autonomous LLM-driven systems that reason about hardware, write high-performance CUDA, and automate the complex loops of architectural simulation, analysis, and optimization.

We are looking for a senior LLM Agents Architect to work hands-on with hardware architects, verification engineers, GPU performance experts, and software developers to build end-to-end agent flows that drive significant improvements in kernel optimization, architectural exploration, and developer efficiency.

What you'll be doing:

  • Design and build agentic AI systems that generate, analyze, and optimize GPU compute kernels — targeting speed-of-light performance on NVIDIA hardware.

  • Collaborate with GPU architects and performance engineers to encode domain expertise — memory hierarchy trade-offs, occupancy tuning, instruction-level reasoning — into agent workflows that rival hand-tuned optimization.

  • Build automated performance forensics agents capable of ingesting large-scale simulation traces and Nsight profiler data to identify bottlenecks and propose architectural or software mitigations.

  • Partner with HW architects to develop agentic flows for GPU architectural studies — enabling rapid what-if analysis across micro-architecture configurations such as cache sizing, memory controller design, and compute unit scaling.

  • Explore agentic approaches to HW/SW co-design challenges, including replacing or augmenting graph-compiler functionality (e.g., TorchInductor) with LLM-driven optimization and code-generation pipelines.

  • Rapidly prototype and thoughtfully productize; integrate with internal services, utilize GPU capabilities, remove bottlenecks, and deliver fitting solutions.

  • Set up evaluation backbone using offline golden sets and online telemetry for confident iterations, cost control, and safe improvements.

  • Mentor and improve teams through insights in agent orchestration, prompting, RAG, observability, crafting documentation and playbooks for NVIDIA's teams.

What we need to see:

  • 8+ years in applied ML/AI or large-scale systems, with 2+ years crafting agentic or LLM-powered applications in production environments.

  • B.Sc in Computer Science / Electrical Engineering.

  • Solid grounding in computer architecture: memory hierarchies, parallelism models, pipelining, and cache behavior. Specific familiarity with NVIDIA GPU architecture — streaming multiprocessors, warp scheduling, shared/global memory model, and occupancy reasoning — is essential.

  • Hands-on CUDA programming experience: writing, profiling, and optimizing GPU kernels — not just calling into CUDA-accelerated libraries. Comfortable with tools such as Nsight Compute, Nsight Systems, or equivalent profiling workflows.

  • Proven ownership of at least one end-to-end agentic system or LLM application: requirements, architecture, implementation, evaluation, and incremental hardening in production — not just experience with off-the-shelf frameworks.

  • Strong software engineering skills in Python and one systems language (C++ preferred).

  • Proficient in tool use, RAG pipelines, and model adaptation techniques for building agentic systems.

  • Demonstrated ability to collaborate with HW/SW domain experts and translate their heuristics into deterministic tools, constraints, and evaluation metrics.

  • Excellence in communication and facilitation: aligning diverse collaborators, documenting decisions/assumptions, and influencing without authority.

  • Track record of building observability for AI systems: dataset/version management, offline test suites, online telemetry, guardrails/safety checks, and rollback plans.

Ways to stand out from the crowd:

  • Familiarity with the PyTorch compilation and lowering stack (torch.compile, TorchDynamo, TorchInductor, Triton, down to PTX), and with GPU graph compilers, kernel fusion strategies, or auto-tuning frameworks.

  • Background in performance engineering for HPC or GPU-accelerated workloads, including experience with performance modeling or hardware simulators.

  • Familiarity with distributed processing, multi-GPU workloads, and networking (e.g., NVLink, InfiniBand).

  • Familiarity with frontier agentic coding tools (e.g., Claude Code, Codex, Cursor) — understanding their underlying architecture: tool orchestration, context management, and autonomous task execution patterns.

  • Hands-on experience building a domain-specific coding agent — whether on top of frontier agentic harnesses (e.g., Claude Code, Codex SDK) or lower-level agent frameworks (e.g., LangChain/LangGraph deep agents, CrewAI). Comfortable with the design choices that make a coding agent useful in practice: task scoping, tool and context curation, evaluation, and failure recovery

Widely considered to be one of the technology world's most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/

Ready to apply to NVIDIA?
Apply to NVIDIA

About NVIDIA

NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry. Learn more about NVIDIA .

See all jobs at NVIDIA →

Similar jobs

NVIDIA
Senior Software Architect, GPU Networking
NVIDIA
⚡ Apply early Israel, Tel Aviv Onsite
● New 👁 Seen ✓ Applied 4d ago
NVIDIA
Senior RDMA and Networking Architect
NVIDIA
⚡ Apply early Israel, Yokneam Onsite
● New 👁 Seen ✓ Applied 4d ago
NVIDIA
Senior System Design Test Architect
NVIDIA
⚡ Apply early Israel, Yokneam Onsite
● New 👁 Seen ✓ Applied 4d ago
Guidewire
Guidewire Technical Architect – ClaimCenter / Jutro
Guidewire
⚡ Apply early India - Bangalore Onsite
● New 👁 Seen ✓ Applied 58m ago
Cisco
Mechanical Engineering Technical Architect | Rack Server, Solid modelling, Structural Simulation, System Thermal Design | 15+ years
Cisco
⚡ Apply early Bangalore, India Onsite
● New 👁 Seen ✓ Applied 3h ago
Redwood Materials
ERP Solutions Architect
Redwood Materials
⚡ Apply early McCarran, NV Onsite
● New 👁 Seen ✓ Applied 9h ago
TraceLink, Inc
Software Architect I
TraceLink, Inc
⚡ Apply early APAC - India - Pune Onsite
● New 👁 Seen ✓ Applied 10h ago
TraceLink, Inc
Solutions Consultant (Supply Chain Business Process Architect)
TraceLink, Inc
⚡ Apply early EMEA - Spain - Remote · location restricted €89,451–€95,666
● New 👁 Seen ✓ Applied 10h ago
Airbus
Solution Architect
Airbus
⚡ Apply early Bangalore Area Onsite
● New 👁 Seen ✓ Applied 11h ago

Sign up for suggestions tailored to the jobs you open and the searches you save.

More jobs at NVIDIA

See all jobs at NVIDIA →

Apply now
🤖

Whoa — hold up

JobsRadar was built for real people having a rough time in their job search — not for automated requests. You're clicking way too fast and you're now temporarily blocked.

Come back later. If you're genuinely job hunting, we've got your back — just act like a human.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Get an edge on your job hunt.

Join our Telegram channel for the stuff that helps you land the role — salary benchmarks, the weekly market pulse, and new-feature drops. No spam, just signal.

Join the channel — it's free