We use cookies. Find out more about it here. By continuing to browse this site you are agreeing to our use of cookies.
#alert
Back to search results
New

Senior AI Agent Quality Engineer

MultiPlan
$140,000 - $160,000
401(k)
United States, Virginia, McLean
7900 Tysons One Place (Show on map)
Oct 01, 2026

Why This Role Matters

Claritev is seeking a Senior AI Operations Engineer to serve as the quality gatekeeper for the agentic AI systems powering the next generation of healthcare products. In this highly visible role, you will help ensure AI agents perform reliably, safely, and consistently across healthcare workflows and business-critical processes.

Working at the intersection of AI engineering, quality assurance, and operations, you will design the evaluation frameworks, benchmarks, and automated testing pipelines used to validate agent behavior before and after release. Your work will help identify regressions, hallucinations, unsafe actions, and workflow failures before they impact customers or production environments.

This role offers a unique opportunity for a QA automation, MLOps, backend, or software engineer looking to specialize in the emerging field of AI agent quality, evaluation, and reliability. You will partner closely with AI engineers, data scientists, product leaders, and operations teams to improve agent performance, observability, compliance, and trustworthiness at scale.

What You'll Do

  • Design, build, and maintain automated test suites and evaluation pipelines for agentic AI systems across single-turn, multi-turn, and end-to-end workflow scenarios.

  • Develop and curate golden datasets, simulation environments, and test scenarios that reflect real-world healthcare workflows and edge cases.

  • Execute automated benchmarking and regression testing across model, prompt, and tool changes with measurable quality gates integrated into CI/CD pipelines.

  • Measure and report agent quality metrics including task completion rates, grounding accuracy, hallucination rates, response quality, latency, cost, and safety compliance.

  • Implement LLM-as-judge and rubric-based evaluation frameworks and validate automated assessments against human review.

  • Triage, reproduce, and root-cause agent failures through analysis of LLM interactions, retrieval workflows, tool invocations, and orchestration paths.

  • Perform adversarial testing, prompt-injection assessments, boundary-condition testing, and guardrail validation.

  • Monitor production agent performance, investigate quality incidents and model drift, and continuously improve evaluation coverage.

  • Create dashboards, runbooks, and quality reporting frameworks used by engineering and product teams.

  • Validate PHI, PII, auditability, and human-in-the-loop controls to support HIPAA and data governance requirements.

  • Collaborate with AI engineers and data scientists to improve agent reliability, observability, and testability by design.

What You Will Bring

Required Qualifications

  • Bachelor's degree in Computer Science, Engineering, Data Science, a quantitative discipline, or related field.

  • Master's degree preferred.

  • 5+ years of experience in software engineering, QA/test automation, ML engineering, or a related technical discipline.

  • 2+ years of experience testing, operating, or supporting production AI or machine learning systems.

  • 1+ years of hands-on experience with Generative AI, LLMs, Retrieval-Augmented Generation (RAG), or agentic AI systems.

  • Proven track record building test automation or evaluation capabilities that improved product or model quality.

  • Strong Python programming skills and experience with automation frameworks such as pytest.

  • Experience with AI evaluation concepts including benchmark datasets, rubric-based scoring, LLM-as-judge methodologies, and regression detection.

  • Familiarity with AI observability and evaluation platforms such as LangSmith, Langfuse, Arize Phoenix, Ragas, DeepEval, promptfoo, or similar tools.

  • Working knowledge of agent frameworks including LangGraph, LangChain, AutoGen, CrewAI, and associated orchestration patterns.

  • Understanding of RAG architectures, embeddings, vector search, and retrieval quality measurement.

  • Experience integrating automated testing and quality controls into CI/CD pipelines.

  • Familiarity with cloud environments, containers, and distributed-system debugging.

  • Knowledge of statistical concepts used to interpret evaluation outcomes.

  • Strong analytical, troubleshooting, communication, and documentation skills.

Preferred Qualifications

  • Experience with Oracle Cloud Infrastructure (OCI), including Generative AI, Data Science, Database, and Observability services.

  • Healthcare, insurance, payment integrity, claims, or other regulated-industry experience.

  • Experience testing systems that process PHI or PII within HIPAA-regulated environments.

  • Exposure to public AI benchmarks such as SWE-bench, GAIA, AgentBench, or tau-bench.

  • Experience with AI red-teaming, safety evaluation, and security testing.

  • Experience with Kubernetes, Terraform, Site Reliability Engineering practices, or infrastructure automation.

  • Previous QA or SDET experience supporting ML, AI, or data-intensive applications.

Compensation

The salary range for this position is $140,000 - $160,000. Actual compensation is based on experience, skills, education, work location, and internal equity. This position may also be eligible for incentive compensation, health insurance, 401(k), and bonus opportunities.

#LI-MC2

Why Claritev?

Healthcare is complex. We help make it clearer.

At Claritev, you'll do work that matters. Together, we're helping make healthcare more transparent and affordable for all through the power of data, technology, and expertise. We offer meaningful opportunities to grow your career, collaborate with talented colleagues, and make an impact on the clients and communities we serve. If you're looking for purpose, growth, and a team that succeeds together, you'll find it here.

What Guides Us

At Claritev, innovation, agility, and a focus on results drive our success. We embrace bold thinking, work as one team, take ownership, and strive for excellence in everything we do - creating meaningful impact for our clients, communities, and each other.

Applied = 0

(web-9db6c7984-5xhmq)