Jobs Companies Arcade Applied AI Engineer

About this Applied AI Engineer role at Arcade

Arcade · Onsite · San Francisco, CA

Everyone's building AI agents, but almost nobody gets them to production.

Building an impressive demo is easy. Building an AI agent that can securely take action inside enterprise systems is hard. The moment an agent accesses customer data, executes a workflow, or makes changes on behalf of a user, authorization, governance, and trust become the real engineering challenge.

Arcade is the MCP runtime that gives agents the power to do both seamlessly. We connect agents to the systems they act in, then give each one a permission slip and a paper trail - proof of what it's allowed to do, and a record of what it did. That's what makes AI safe to turn loose: real actions, on real systems, already shipping inside Fortune 100 companies.

The Revolution Needs You

Every AI app needs agentic tools that let AI models take real actions. Without tools, AI can only chat. With tools, AI can actually do things. We're building the definitive tools catalog, actions platform, and governance model that will unlock AI's true potential. Think Zapier for AI Actions. Think Auth0 for AI. Think really big.

Why This Is The Opportunity of a Lifetime

  • Traction: Real deployments with Fortune-100 customers like Morgan Stanley and Open Table

  • Founder-Market Fit: Our CEO previously founded Stormpath (acquired by Okta), where he created the first Authentication API for developers. He's done this before - and this time the market is 10x bigger. Our CTO led the vector database team at Redis, shipped 100+ LLM applications, and is a contributor to LangChain and LlamaIndex. He knows this space better than anyone.

  • Dream Team: We've assembled authentication, integrations, distributed systems, and AI experts from Okta, Redis, Microsoft, Splunk, Ngrok, Google, Airbyte, Disney, and HPE who've built and founded multiple successful developer platforms.

  • Perfect Timing: Every enterprise is racing to put agents in production - almost none get there. The problem isn't better models, it's proving which agent can take which action, on behalf of which user, against which system. That's us.

  • Massive Market : We're building critical infrastructure for the biggest technological shift of our generation. Every AI app will need what we're building.

  • Backed By The Best: Our Series A round is led by SYN Ventures, with strategic investment from Morgan Stanley and Wipro. Our earlier investors have also backed Databricks, Clickhouse, MongoDB, Perplexity, Cohere, ScaleAI, Confluent, Elastic, and Firebase. They see what we see - this is going to be huge.

The Challenge

You'll report to the Engineering Manager for Tools and Growth. The Tools team owns Arcade's tool catalog — thousands of tools across many services, growing faster than any human can review by hand. The next leap in agent quality lives inside this team's work, and you'll be the applied-AI seat that pushes it forward.

Three real problems define the role.

Agentic tools vs. deterministic tools. Most tools today are deterministic: call X API with Y arguments, get Z result. That model breaks down for entire classes of agent work — research a topic, summarize a thread, decide which of three accounts to act on. Agentic tools, the ones that internally reason, plan, or call models are the answer, but the design space is wide open. When is agentic better than deterministic? How do you make an agentic tool fast, reliable, and debuggable? You'll set the bar for what these look like at Arcade.

Agents that build tools. The toolkit catalog is too big for hand-crafting to scale. We need agent harnesses that can take a vendor's API and produce a high-quality toolkit — design, code, eval, docs with a human in the loop only where the human is actually needed. There's early work on this already. You'll take it from a prototype into the production pipeline that produces the next thousand tools.

Workflows that compose tools. Individual tools solve narrow problems. Real customer outcomes: "close the quarter," "triage the inbox," "stand up the integration" need many tools, chained, with the right control flow. We need to figure out what the right primitive looks like above the tool layer, and you'll lead that design.

The most honest thing we can say about this work: most of the problems you'll be solving didn't exist three months ago. There's no prior art. There's no known solution. If that's the part of the job that makes you nervous, this isn't the right role. If that's the part that makes you lean in, it is.


We do real experiments. We form hypotheses. We publish learnings. Research is part of the job. But the role is built around shipping. If you want to spend six months proving an idea in a notebook before anything reaches a customer, this isn't the right role. If you want to ship the experiment and the writeup in the same quarter, it is.

What You'll Do

  • Design and ship agentic tools that go beyond deterministic API wrappers — and define the patterns the rest of the Tools team will use to build more.

  • Build the agent harness that automates tool creation — take a vendor's API, produce a high-quality toolkit end-to-end, keep humans in the loop only where humans add real value.

  • Design workflows that compose tools into higher-level abstractions customers can actually point at outcomes ("triage this inbox," "close out this account") rather than individual API calls.

  • Bring applied-ML rigor to tool design — evals, model-aware iteration, retrieval, tool description tuning, response shaping. Make decisions defensible with data.

  • Run model-aware experiments across Claude, GPT, Gemini — agentic tool behavior diverges across models in ways nobody else is studying, and we should.

  • Set the technical bar for what "good tool-building" looks like as the team scales — your patterns get inherited by every toolkit author after you.

  • Contribute back to the MCP and agent ecosystem where the conversation about agentic tools is forming.

Required Skills

  • 5+ years software engineering experience, with at least 2 years shipping production ML or applied-AI systems. Formal title matters less than the work.

  • Strong Python.

  • LLM application depth — prompting, retrieval, tool use, agent design. You've built non-trivial agent systems and know where the rough edges are.

  • Experience designing or composing multi-tool / multi-agent workflows that produced real outcomes.

  • You've built evals at scale — not "I ran a benchmark once," but a measurement system real engineering decisions were made against.

  • Statistics fluency — significance, confidence intervals, A/B test design. You can defend whether a small delta is real or noise.

  • Comfort across multiple frontier models and reasoning about their behavioral differences.

  • A do-er, not a researcher-in-residence. You'd rather ship a working v0.5 next week than a polished v2.0 next quarter.

  • Comfort with ambiguity — early team, narrow charter that will expand. You make good decisions with incomplete data.

  • An insatiable desire to ship.

Bonus Points

  • You've built agents that build software (codegen agents, harness-style systems, meta-agents).

  • Prior work on tool-use specifically — BFCL, τ-bench, ToolBench, MCP eval work, or equivalent.

  • MCP ecosystem familiarity — extra bonus if you've filed an issue against the spec.

  • You've worked on agent frameworks (LangChain, CrewAI, AutoGen, Mastra) and have opinions about where they get tool use and workflow composition wrong.

  • Prior experience at an API platform, integrations-heavy product, or developer tools company.

Join The Movement

We're not just building a product - we're leading a movement to transform AI from just chatbots to agents that can take actions against real systems. This is your chance to be at the forefront of that revolution.

If you want to look back in 5 years and say, "I helped build that", then we want to talk to you. Ready to make AI actually useful? Apply Now

Compensation and Benefits

This role offers a competitive salary, equity, and benefits. Compensation is aligned with the range below and determined based on a candidate's background, experience, and performance.

Salary Range

$179,000-240,000 USD



Ready to apply to Arcade?
Apply to Arcade

How this AI Engineer salary compares

This role pays $209,500/yrin line with the typical range for AI Engineer roles.

$169,415 median $235,000 $312,323

Typical range $201,264–$277,990/yr, from 195 comparable AI Engineer listings on JobsRadar (pay annualized to USD). See AI Engineer salary insights →

Similar jobs

Sign up for suggestions tailored to the jobs you open and the searches you save.

More jobs at Arcade

See all jobs at Arcade →

Apply now
🤖

Whoa — hold up

JobsRadar was built for real people having a rough time in their job search — not for automated requests. You're clicking way too fast and you're now temporarily blocked.

Come back later. If you're genuinely job hunting, we've got your back — just act like a human.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Get an edge on your job hunt.

Join our Telegram channel for the stuff that helps you land the role — salary benchmarks, the weekly market pulse, and new-feature drops. No spam, just signal.

Join the channel — it's free