Field Note · Applied AI and system ownership
What Should You Evaluate Before You Trust an AI System?
An AI evaluation should answer more than whether the output looks good. It should tell us whether the system can change a consequential decision or action without hiding the work, risk, and correction burden it creates.
A thinking frame by Andrew Moss
The questions I get
Usually some version of these:
- How do we know whether this AI output is good enough?
- What should an eval include for expert work?
- When can we reduce human review?
What a lot of people seem to think
If the model sounds fluent, returns information faster, and the average output looks good, the system is ready.
How I look at it
Fluency and speed are weak proxies. I want the expert to define the purpose, must-have facts, prohibited errors, source rules, judgment criteria, and escalation conditions. Then I want to know whether the result actually changes a decision, responsible action, or valuable capability. Every verified correction should improve the next test rather than disappear into a chat.
Why the decision matters
The cost is rarely confined to the line item.
If the sequence is wrong
The system can make fast, polished, repeatable mistakes that users stop noticing.
If the sequence is right
Quality is explicit, high-consequence failures are tested, human review matches risk, and corrections improve the operating system.
How reversible is it?
Low when output reaches clients, changes rights or money, influences a consequential decision, or becomes trusted memory.
The short answer
Evaluate the decision, the action, and the correction loop.
Use representative cases, including ugly exceptions. Score source fidelity, consequential error, usefulness, total review work, and the action the output supports. Preserve verified corrections with their reason and scope. Trust should increase only where repeated evidence earns it.
The action testFaster information is not the outcome.
The relevant chain is source → output → expert review → decision → responsible action → measured result → verified correction. A system that improves only the output may leave the business exactly where it started.
Move fromFluent, fast output→Move towardEvidence that changes a consequential action
The order I would use
Take the right steps in the right order.
- 01
Define the real purpose
Name the user, decision, action, and consequence the system is meant to improve.
- 02
Write the expert standard
Specify must-have facts, source requirements, prohibited errors, judgment criteria, uncertainty, and escalation.
- 03
Build representative cases
Include normal work, sparse evidence, conflicting sources, ugly exceptions, and cases where the correct action is to stop.
- 04
Measure total work and action delta
Count preparation, review, correction, integration, and follow-through; then identify what decision or action actually changed.
- 05
Store verified corrections
Record the correction, source, reviewer, reason, scope, permission, and outcome so the next use can improve responsibly.
- 06
Earn narrower trust
Increase autonomy only for the cases and consequences the evidence supports.
Questions worth answering
Before the next irreversible move:
- What decision or responsible action changes if the output is better?
- Which error would be costly even if the average score is high?
- How much human review and correction does the result still require?
- What verified correction should make the next use better?
- Where must the system stop and escalate?
What not to do
Do not grade style while the decision remains untested.
Do not reward fluency, speed, or average accuracy alone. Do not use only clean examples. Do not let corrections vanish or become universal without provenance. Do not expand authority beyond the cases the evidence covers.
Keep the perspective
Evaluation converts speed into earned trust.
The point is not to make the model look competent. It is to protect the decision, make expert judgment explicit, and create a correction layer that improves the system without obscuring human responsibility.
Independent sources
Useful primary material
These sources support the public frame. They do not replace the private facts or the accountable professional.
Common follow-up questions
How many test cases are enough?
Start with a small set that covers normal work, important variations, known failures, and high-consequence edges. Expand it whenever a material new failure appears.
Can AI grade AI output?
It can assist when the rubric is explicit and its grading is itself checked. High-consequence decisions still need appropriate human review and source verification.