AI Rubric: Don't Just Ask AI for an Answer — Ask What Good Looks Like

· By Peter Lowe

Category: Governance

Woman holding a checklist clipboard while marking and improving a document, brand illustration

Before you accept an AI answer, define what a good one contains. A short AI rubric does more for quality than any prompt trick.

*Part 6 of 6 in a series on the habits that make AI useful. Previously: [AI Adoption Is a Management Challenge, Not a Technology Project](/insights/ai-adoption-management-challenge/).* Most people use AI as a vending machine. An AI rubric changes that. Question in, answer out, take it or leave it. There's a better pattern, and it costs one extra step: before you accept the answer, ask what a good answer would have to contain — then have the work marked against it. We call it a rubric. It's the same thing a teacher uses to mark an essay, or an interview panel uses to score candidates: a short list of criteria, agreed before anyone starts, so that judgement is consistent rather than a matter of mood. ## Why a rubric helps The reason AI output often feels plausible but disappointing is that "good" was never defined. You knew it when you saw it, and you only saw it after it was wrong. A rubric forces the definition up front. For a customer email that might be: answers the actual question, no more than 150 words, no jargon, acknowledges the delay, offers a specific next step, doesn't promise anything we can't deliver. Now there's something to check against. You can ask the model to assess its own draft against each criterion, say where it falls short, and produce an improved version. It's remarkably good at this — far better at critiquing work against stated criteria than at guessing what you wanted from a one-line request. Three steps, and you can do it in any chat tool today: 1. Ask what criteria a good version of this should meet, and edit the list until you agree with it. 2. Ask for the draft. 3. Ask it to score the draft against each criterion, explain the gaps, and rewrite. ## Where this stops being a chat trick The interesting bit happens when you keep the rubric. Once the criteria live somewhere other than a chat window, they become a shared standard — and everyone's output gets more consistent, including the humans. We do this in two places on this site, and both are running in production rather than in a slide. **The AI Starting Point Assessment.** Answer twelve questions and you get a scored report on where your business actually is with AI. Behind it sits a fixed rubric: what each answer means, how the dimensions are weighted, what separates a strong position from a weak one. That's what makes it a diagnostic rather than a horoscope. The scoring is deliberately deterministic — the same answers produce the same result, every time, within a couple of percent. That reliability isn't a bonus feature. It's the whole reason the output is worth reading, and it comes from the rubric, not the model. **Our own SEO work.** When the system proposes a change to a page, it doesn't just make it. The change gets evaluated against criteria — is it accurate, does it match how people actually search, does it improve the page for a reader rather than a crawler, is it consistent with how we write? Poor suggestions get rejected before they reach me. Good ones get applied. What was learned gets kept, so the next batch of suggestions is better than the last. That last part is what makes it self-improving, and it isn't magic: it's just refusing to throw away the judgement. Most people re-explain their standards from scratch in every conversation. Writing them down once and reusing them is the entire trick. ## Two things we got wrong Worth saying, since the useful version of this always comes with scar tissue. The first rubric we wrote was too long. Fourteen criteria, and everything scored somewhere in the middle. A rubric with too many criteria stops discriminating between good and bad. Five or six sharp ones beat fifteen woolly ones. The second mistake was writing criteria we couldn't check. "Engaging" isn't a criterion. "Answers the question in the first two sentences" is. If you can't tell whether something passed, neither can anything else. We also keep a human in the loop on anything that gets published or sent. The rubric raises the floor and filters out the obviously weak; it doesn't take responsibility. Nothing does. That's still yours — as we've said about [meaningful human oversight](/insights/ai-human-oversight-meaningful/). ## Try it on something small this week Take a piece of work your team produces regularly and inconsistently. A quote, a project update, a job advert, a proposal summary. Write six criteria for what good looks like. Ask a colleague whether they agree — that conversation alone is usually worth the exercise, because it's rarely as agreed as everyone assumed. Then run your next three drafts against it. You'll get better output from AI. You'll also, quietly, have documented a standard that used to exist only in someone's head — which is [the problem we keep coming back to](/insights/process-documentation-out-of-one-head/). ## Where this series lands Six articles, one argument: nobody is naturally good at AI. They're curious, they define problems clearly, they expect the first attempt to fail, they invest time before they see returns, and they work somewhere that lets them. The rubric is the smallest, most practical version of all of that. It's clear thinking, written down, applied repeatedly. If you'd like a scored, honest read on where your business currently stands, the [AI Starting Point Assessment](/ai-starting-point/) takes about ten minutes and uses exactly the approach described here. ## FAQs ### What is an AI rubric? A short list of criteria that defines what a good output looks like, agreed before the work starts. The AI produces a draft, then assesses it against those criteria and improves it. It's the same idea as a marking scheme. ### How many criteria should a rubric have? Five or six that you can genuinely check. Long rubrics stop discriminating between strong and weak work, and vague criteria like "engaging" can't be assessed by anyone, human or otherwise. ### Can AI reliably judge its own work? Against explicit criteria, it's surprisingly effective — much better than it is at guessing your intent from a short request. It's far less reliable at judging factual accuracy, so facts still need a human or a source check. ### Does this work in ordinary tools like ChatGPT or Copilot? Yes. Ask for the criteria, agree them, ask for the draft, then ask for a scored critique and a rewrite. Save the criteria somewhere reusable so you're not reinventing the standard each time. ### How is this different from just writing a better prompt? A prompt describes the task. A rubric describes the standard, and can be reused, shared and improved across a team. It also gives you something to check the output against, which a prompt doesn't. ### Do we still need a person to check the output? Yes. A rubric raises the average quality and filters out weak work; it doesn't carry accountability. Anything published, sent to a customer, or affecting money, people or legal position still needs a named human reviewer.

This article was written by Peter Lowe. The ideas and opinions are his own; AI was used to assist with drafting and editing.