FIND JOBS IN AI GUIDE
What is a rubric? How to write accurate, atomic criteria
A rubric is a structured way to judge whether a piece of work meets a standard. In AI evaluation, it helps reviewers assess answers consistently instead of relying on an overall impression. A useful rubric is grounded in the task, tests observable evidence and makes each decision clear.

Why do rubrics matter?
If two reviewers read the same answer and reach different scores, the problem may be an unclear standard rather than the answer itself. A rubric tells them which facts, behaviours or omissions matter and how to recognise them.
Rubrics can also make feedback more useful. Instead of “not good enough”, a reviewer can identify exactly which criterion was missed.
What do atomic and self-contained mean?
An atomic criterion asks one question. If an answer can satisfy half of the sentence but fail the other half, split the sentence into two criteria.
A self-contained criterion carries enough context to be judged without guessing the intended answer. It names the expected fact or behaviour, including the relevant task conditions where necessary.
Start with evidence, then write the criterion
Read the prompt and reference material first. Decide what a correct answer must include and what would be an actual error. Write a short pass condition for each distinct item, then check whether a different reviewer could use it without extra explanation.
Bad versus better criterion
Vague: “The answer is accurate, helpful and complete.” Better: “The answer states that the UK minimum wage rate depends on the worker’s age.” A second criterion can check whether the answer gives the relevant rate for the date in the prompt. Each item tests one claim and can be judged on its own.
Atomic means one decision per criterion
Do not combine several requirements with “and” if a response could meet one but miss another. Split “identifies the issue and explains the exception” into two criteria when each is independently important. This makes partial credit and error analysis clearer.
Atomic does not mean reducing a judgment to meaningless fragments. A criterion should still test a useful fact or behaviour.
Self-contained means the grader has enough context
Write the expected answer or evidence into the criterion where practical. “Mentions the correct date” is weak if the grader has to search elsewhere to learn what date is correct. “States that the deadline is 30 June 2026” is easier to apply, assuming that date is established by the task source.
Avoid private shorthand, unexplained acronyms and references such as “as above”. If external evidence is needed, identify the authoritative source and the relevant passage.
A repeatable drafting method
First, read the original task and authoritative reference material. List the facts or behaviours an ideal answer needs. Draft one criterion for each distinct item. For each, specify an observable pass condition, then test it against one answer that should pass and one that should fail.
Finally, ask another reviewer to score the same examples without your help. If their decisions differ, tighten the criterion or split it. Check that the full set covers the task without double-counting.
What to avoid
Do not reward confident wording when the fact is wrong. Do not make a criterion depend on information unavailable to the person being graded. Avoid criteria that penalise harmless wording differences when the meaning is equivalent. Keep mandatory safety or eligibility checks separate from style preferences.
Questions to ask before you start
- Does this criterion test just one observable thing?
- Could a reviewer judge it without reading the rubric writer’s mind?
- Does it say what evidence counts as meeting it?
- Is it grounded in the actual task and reference material?
- Could two criteria award credit for the same fact?
Further reading: OpenAI HealthBench paper on self-contained, objective rubric criteria ↗