AgentOps
Keeps agents and model systems running in production — evals, observability, cost and latency control, failure handling.
Also posted as: LLMOps Engineer · AI Reliability Engineer · Evals Engineer · AI Platform Engineer
Definition
What this role actually is
AgentOps is the operations layer under everything else on this list. The work is measuring whether a non-deterministic system is still doing its job: eval suites, tracing, regression detection on model upgrades, spend and latency budgets, and a plan for the run that goes wrong at 3am.
It is the difference between an agent that demoed well in March and one that is still trusted in September. Almost nobody holds this job yet, which is why most agent projects quietly stop being used.
The line in the sand
What it is NOT
Not an Agent Engineer.
Agent engineers build the system. AgentOps keeps it honest over time. The overlap is real; the accountability is different.
Not general DevOps.
Uptime is the easy part here. The hard part is detecting quality degradation in a system that never returns an error code.
Not a dashboard screenshot.
A monitoring panel with no decision attached proves nothing. Evidence is a caught regression, a cost cut, or a failure diagnosed and fixed.
Output
What they actually ship
- Eval suites that catch quality regressions before users do
- Tracing and observability that make a failed run diagnosable after the fact
- Cost and latency budgets, with real numbers before and after
- Model-upgrade regression testing and rollback paths
- Written incident accounts — what failed, why, and what changed
Stack
Tool stack
Depth
The proof standard
Tier 1
Core
Owns evals and observability for a model system in production. Can show a regression the suite caught, state a cost or latency improvement with numbers, and describe a real incident end to end.
Tier 2
Working
Built an eval suite or tracing setup for a live system. Catches obvious failures; measurement may be partial.
Tier 3
Light
Built evals or monitoring for prototypes and personal projects. Not yet load-bearing in production.
Depth is what the builder claims. The badge is awarded separately, from the evidence — read the standard every badge is judged against.
Evidence
What a strong portfolio piece looks like
An eval suite or trace setup plus a written account of one real failure and the fix. The incident write-up is the whole role; the tooling is replaceable.
Market
Market snapshot
Barely titled and rising fast, hidden inside Platform Engineer, MLOps and Senior AI Engineer postings. Teams typically create the role only after an agent has failed expensively in production, which means demand appears suddenly and pays well when it does.
Last updated: 2026-09-04