Model assurance
Usage counts and average accuracy are vanity in production. Override rate, correction rate, failed validations and refusal accuracy are what move before an incident.
Read article · 7 min
Model assurance
Numeric models drift in ways a chart catches. When the output is code, prose or a derivation, the drift is qualitative and silent. Here is how to surface it.
Read article · 8 min
Model assurance
Grading one model with another feels independent. If they share training and assumptions, they share blind spots, and the check passes on the same mistake.
Read article · 7 min
Model assurance
LLM-as-judge is the default and often the wrong one. Match the validator to the output: code where you can check it, a human for judgment, a model only for the open-ended.
Read article · 7 min
Model assurance
Averaged accuracy is meaningless when a single rare error carries the whole consequence. Set the bar on the tail, and build the eval set to find it.
Read article · 8 min
Model assurance
Success criteria, a representative task set, tool-use correctness and cost per resolved task — and where to set a go/no-go bar before you ship.
Read article · 8 min
Model assurance
Grounding rate, refusal accuracy, recall at k and tool-call accuracy each predict a kind of incident. Overall accuracy predicts none of them.
Read article · 7 min
Model assurance
Aggregates rise while things break. A diff between two runs is an afternoon of work and it changes what a team argues about.
Read article · 6 min
Applied ML
Tokens per call and average latency both flatter the system. The number a business can act on is what one completed piece of work costs, at the tail.
Read article · 6 min
Model assurance
Policy pass rate, over refusal and sensitive data in output. A system can score well on every other line of a benchmark and still be the one that leaks.
Read article · 7 min
Model assurance
A system that answers differently on the third attempt is a support problem. A suite run once per case can never show it.
Read article · 6 min
Model assurance
Real users paraphrase, misspell, pad with irrelevant detail and paste in text carrying instructions. Robustness is the measurement of what that does to your answers.
Read article · 7 min
Applied ML
Schema validity, required fields and refusal precision. The measurements that decide whether the system downstream survives contact with your agent.
Read article · 6 min
Applied ML
Tool selection accuracy is the easy measurement and the least informative. What decides production behaviour is what happens after a call times out.
Read article · 7 min
Applied ML
If the passage holding the answer never reaches the prompt, no model recovers it. Measure retrieval separately or you will spend months tuning the wrong component.
Read article · 7 min
Model assurance
Hallucination is not a vibe. It is the share of claims no retrieved passage supports, plus what the system does when the answer is not there at all.
Read article · 8 min
Model assurance
An agent retrieves, decides, calls tools and writes. Each part fails on its own, and a single accuracy figure is the average that hides which one did.
Read article · 7 min
Applied ML
A contradiction needs both halves in context at once. Retrieval ranks
passages by similarity to a query, which is not the same thing, and no model
detects a disagreement it has only seen one side of.
Read article · 9 min
Applied ML
If the passage holding the answer never reaches the prompt, no model can
recover it. Most disappointing document question answering systems are failing at
chunking and search, not at generation.
Read article · 9 min
Model assurance
Without a fixed set of cases and a definition of correct, every change is an
opinion and the team argues from anecdotes. The harness is what turns it into a
measurement.
Read article · 8 min