← All roles

AgentOps

Keeps agents and model systems running in production — evals, observability, cost and latency control, failure handling.

Also posted as: LLMOps Engineer · AI Reliability Engineer · Evals Engineer · AI Platform Engineer

Definition

What this role actually is

AgentOps is the operations layer under everything else on this list. The work is measuring whether a non-deterministic system is still doing its job: eval suites, tracing, regression detection on model upgrades, spend and latency budgets, and a plan for the run that goes wrong at 3am.

It is the difference between an agent that demoed well in March and one that is still trusted in September. Almost nobody holds this job yet, which is why most agent projects quietly stop being used.

The line in the sand

What it is NOT

Not an Agent Engineer.

Agent engineers build the system. AgentOps keeps it honest over time. The overlap is real; the accountability is different.

Not general DevOps.

Uptime is the easy part here. The hard part is detecting quality degradation in a system that never returns an error code.

Not a dashboard screenshot.

A monitoring panel with no decision attached proves nothing. Evidence is a caught regression, a cost cut, or a failure diagnosed and fixed.

Output

What they actually ship

  • Eval suites that catch quality regressions before users do
  • Tracing and observability that make a failed run diagnosable after the fact
  • Cost and latency budgets, with real numbers before and after
  • Model-upgrade regression testing and rollback paths
  • Written incident accounts — what failed, why, and what changed

Stack

Tool stack

LangSmithBraintrustLangfuseArize PhoenixOpenTelemetryHeliconeGrafanapytestRagasSentry

Depth

The proof standard

Tier 1

Core

Owns evals and observability for a model system in production. Can show a regression the suite caught, state a cost or latency improvement with numbers, and describe a real incident end to end.

Tier 2

Working

Built an eval suite or tracing setup for a live system. Catches obvious failures; measurement may be partial.

Tier 3

Light

Built evals or monitoring for prototypes and personal projects. Not yet load-bearing in production.

Depth is what the builder claims. The badge is awarded separately, from the evidence — read the standard every badge is judged against.

Evidence

What a strong portfolio piece looks like

An eval suite or trace setup plus a written account of one real failure and the fix. The incident write-up is the whole role; the tooling is replaceable.

Market

Market snapshot

Barely titled and rising fast, hidden inside Platform Engineer, MLOps and Senior AI Engineer postings. Teams typically create the role only after an agent has failed expensively in production, which means demand appears suddenly and pays well when it does.

Last updated: 2026-09-04

Do you match this?

List your work

Hiring this role?

See who matches