Measure. An evaluation framework — criteria agreed with you, scored datasets, judge models, versioned reports — and a regression suite wired to your CI, so a model or prompt change cannot silently degrade production. Where the problem is made-up claims, a per-claim fact-checking component.
Optimise. Then the loop: error analysis on the failing cases, one change at a time (prompt, retrieval, data, fine-tuning), re-measure, keep only what moves the number, and record what did not. This is how performance improves on a schedule instead of by luck. The same loop is how I took an essay-scoring model to state of the art before it served a million learners.
Run unattended. Where the goal is autonomy: MCP servers over your data and tools, a harness with permission rules, and a scheduled runner with locking, timeouts, bounded retries, cost caps and one aggregated alert per run.
The rule behind it. Verification runs in a separate process with a different model, because a system reviewing its own work misses its own mistakes.
In technical terms. Eval harnesses, LLM-as-judge scoring, hallucination detection, error-analysis-driven optimisation; MCP servers that expose your data and APIs as agent tools; the agent harness — scheduling, retries, permissions, human-in-the-loop — so agents run unattended.
See: Catching an AI's made-up claims before the user does · State-of-the-art AI model research, put into production · Twenty specialist agents, two human sign-offs