Sobre este puesto de Senior Applied Scientist en Caseware
We are building the agentic platform that powers Caseware's next generation of audit and accounting products, and applied science is how we measure and improve quality across our agents. This is a role for someone who thinks in experiments, has built evaluations for LLM-based systems, and wants to build the tools that other teams and our customers rely on to ship their own agents.
You will design and run experiments that turn ambiguous quality questions into measurable results, and own the evaluation methodology that internal teams and customers depend on. You will build the AI-based capabilities behind our end-to-end agent builder, including the synthetic data generation and eval builders that let teams and customers create, evaluate, and deploy their own agents. You will also work on something genuinely novel: self-learning, self-improving systems that get better both offline and online. A significant part of that is agentic memory: deciding what is worth remembering and validating the criteria that let agents compound useful knowledge on behalf of customers over time, while rigorously upholding the legal and contractual obligations owed to them and their clients. At the staff level, you will also influence technical direction and mentor others.
📍 Location: This is a fully remote position located in Colombia.
Contact
Maira Russo - Senior Talent Acquisition Partner
The domain: financial audit
This role sits squarely in the financial audit domain, and that domain matters more than most people realize. Independent audit is one of the quiet foundations of the global economy. When investors, lenders, regulators, and the public can trust that a company's financial statements are accurate, capital can flow, markets can function, and organizations can be held accountable. That trust rests on the quality of audit and assurance work, which makes the tools that audit professionals use genuinely consequential.
You do not need prior knowledge of accounting or auditing to succeed here. What you will gain is deep, durable expertise in how audits are performed and where AI can make them faster, more reliable, and more insightful. You will learn the domain on the job alongside experienced domain experts, and that fluency will make you a stronger applied scientist. If you already bring that background, even better.
What you'll be doing:
-
Design and run experiments that measure and improve the quality of LLM-based applications and agents, turning ambiguous quality questions into measurable, reproducible results.
-
Build out and own the evaluation methodology in practice that internal teams and customers depend on: benchmarks, regression suites, scoring methods, acceptance thresholds across task success, faithfulness, safety, latency, and cost, and the scientific foundations for comparing offline and online performance, partnering with platform engineering on the automated infrastructure that runs those comparisons and alerts on drift or regression, and wiring the signals into release gates.
-
Build AI-based tools and capabilities that power the end-to-end agent builder developer experience used by other Caseware teams and by our customers.
-
Build the synthetic data generation and eval-builder capabilities that let internal teams and customers create, evaluate, and deploy their own agents.
-
Turn domain procedures into task-plus-verifier structures, drawing on deep intuition for how LLM-based systems fail.
-
Contribute to applied science on agentic memory: use statistical methods grounded in audit domain knowledge to surface which experiences are worth remembering, and define and validate the promotion and demotion criteria that let agents compound that knowledge over time, partnering with platform engineering on the infrastructure that executes it.
-
Advance self-learning, self-improving systems that get better both offline and online as they are used.
-
Help ensure memory promotion and demotion criteria uphold the legal and contractual obligations owed to customers and their clients, partnering with Security, Legal, and Domain SMEs.
-
Design experiments and evaluations that demonstrate agent quality improves the more customers use the platform.
-
Translate advances in GenAI (models, retrieval, agent frameworks, evaluation techniques) into practical, maintainable capabilities.
-
At the staff level, influence technical direction through RFCs and design reviews, and mentor other scientists and engineers.
What you'll bring:
-
6+ years (senior) to 8+ years (staff) of professional experience in applied science, machine learning, research, or data-intensive engineering. Staff-level candidates bring demonstrated impact beyond a single team.
-
2+ years working on production GenAI or LLM-based systems (applications, agents, or the tooling and evaluation around them).
-
Proven experience designing and building evaluation frameworks for ML or LLM systems: metrics, benchmarks, scoring approaches, and rigorous experiment design.
-
A strong experimentation mindset. You form clear hypotheses, design sound experiments, and draw defensible conclusions from noisy, real-world data.
-
Ability to build and operate your own production tooling and services, ideally on AWS.
-
Strong understanding of GenAI system tradeoffs including quality, latency, cost, reliability, and safety.
-
A strong foundation in probability and statistics: experimental design, significance testing, and reasoning under uncertainty.
-
Strong English language communication and collaboration skills.
-
Comfortable operating in fast-moving environments with ambiguity and evolving requirements.
Nice to have
-
PhD or MS in a quantitative field (Computer Science, Statistics, Machine Learning, or similar) preferred, though equivalent industry experience is welcome.
-
Prior machine learning experience (classical ML, model training, or MLOps).
-
Experience with AI guardrails, governance, or safety mechanisms.
-
Experience with distributed, SaaS, cloud-native, multi-tenant platforms at scale.
-
Familiarity with Infrastructure as Code (CDK, CloudFormation, or Terraform).
-
Experience operating in regulated or compliance-heavy domains.
-
Familiarity with financial audit, accounting, or assurance workflows.
Tech stack you'll be working with
-
Backend & Platform: TypeScript, NestJS, Python
-
Cloud & Infrastructure: AWS EKS, AWS Lambda, AWS Bedrock, AWS AgentCore
-
Search & Retrieval: AWS OpenSearch and S3 Vectors
-
Document & Data Processing: AWS Textract, DynamoDB, S3
-
AI Evaluation & Observability: LangFuse, LangSmith, LangChain, LangGraph
-
AI-Assisted Development: GitHub Copilot, Claude Code, Devin
-
Developer Tooling: GitHub, GitHub Actions, Nx Monorepo