Skip to main content

For the people who know the work

You don’t need an engineer to start this.

The hard part of an AI project is the judgment it gets built from, and that is already yours. You can map how the work really runs and set the standard the thing gets held to, then send colleagues a link when you want their read on it.

Compass
Prism
Workbench
Caliper

Your expertise usually arrives as somebody else’s reading of it

Someone asks for a document. It goes over the wall to whoever is building, and what comes back is their reading of it.

  1. 01

    You write the document

    Someone asks for the process in writing. You describe the ninety percent, because the exceptions are the part that takes an hour to explain and a career to learn.

  2. 02

    Someone else interprets it

    People who've never done the job read it and turn it into rules. The judgment calls you didn't think to write down get filled in with reasonable-sounding guesses.

  3. 03

    You see it when it's wrong

    The first you hear of it is a demo, or a complaint. By then the guess is built, and correcting it looks like you're being difficult rather than useful.

None of that is anyone being careless. It’s what happens when judgment has to survive a round trip through prose. The alternative is to skip the document and answer the questions directly.

Any of these is somewhere you could start on Monday

Pick whichever matches what you’re trying to work out. None of them need preparation, and each one has a link you can send when a colleague should weigh in.

A Compass interview in progress from the guest's side — the conversation mid-flow, with a follow-up question being asked
01

Map how the work actually runs

Walk through the job the way you'd explain it to a new hire, and it asks follow-ups when something sounds important. What comes out is the map everyone argues over when they decide what to build. Send the same interview to the people whose corner of it you don't see.

Twenty minutes, from your phone if you want.

A study from the respondent's side: one option at a time, the scale, and the box asking why they answered that way
02

Put the options in front of people

Line up several versions of an idea and ask the people who'd live with it which ones hold up, and why. Finding out an option is wrong here costs an afternoon. Finding out after launch costs a quarter.

Answers come back without anyone scheduling a call.

A design session from the guest's side: the draft answering in conversation, with the reaction controls beside it
03

Get a first draft and argue with it

Describe what it should do and get a working version back, then talk to it the way a real user would. Your reactions land on the thing itself rather than in a meeting nobody minuted, and the ones that are really standards become part of what it gets scored against.

As long as you feel like poking at it.

The guest rating surface as a rater sees it: one case, the AI's answer, pass/fail on each criterion, and the box for what it should have said
04

Judge what it produced

Read what it did with an actual piece of work and mark whether it's right. Where it's wrong, say what should have happened instead — that sentence is worth more than the score beside it. Hand the same set to colleagues when one opinion isn't enough.

Ten minutes. Stop whenever you like.

The spec contributor surface: a question with the ideal answer being written, and other people's answers to the same case alongside
05

Write down what a good answer is

For the cases that matter most, write what a good answer looks like in your own words. That becomes the bar every later version gets measured against, and colleagues can add theirs beside it.

As many as you feel like.

What you say becomes a test the AI keeps having to pass

This is the part worth knowing before you spend the ten minutes. What you say doesn’t get read once and filed. Each case you judge becomes a test, and every later version has to pass it.

Models get swapped. Prompts get rewritten by people who weren’t there. Your answer is what notices.

  1. In March

    You said a double charge before payroll is urgent, not normal.

  2. Every release since

    That case runs against each new version, with 46 others.

  3. This morning

    A new model routed it to Normal. The build stopped.

  4. Before anyone noticed

    It was fixed, because your answer was still the standard.

Why this beats writing it down

The exceptions survive

The rules of thumb that never make it into an SOP are exactly what this asks you about, because they're where the AI gets it wrong.

You stop being the bottleneck

Right now every hard call routes through you. Write the calls down as cases and your judgment can answer them while you're on holiday.

You get a say in the standard

Your answers set the bar before there's anything to demo. By the time a project reaches review, changing anything is expensive.

What you’re probably wondering

Am I training my replacement?
You're setting the bar it has to clear, on the cases you decide matter. A team that asks first is a team that knows the judgment is worth more than the throughput — and the standard you set is the one it keeps being measured against.
Does everyone on my team need an account?
No. Anyone you want an answer from gets a link, and that's the whole thing — no signup, no password, nothing installed, and nobody added to a tool they'll have to remember the name of.
Do I have to learn the software?
No. One question at a time, in plain language. If you can answer an email you can do this, and there's no session where someone trains you first.
How long does it take?
Ten to twenty minutes, and you can stop partway. What you've already answered still counts.
What happens to what I say?
It stays attached to you and to the case it was about, so nobody can round it off into an average later. Your colleagues can see who said what, and disagree in the open.
What if I don't have time?
Five answers beat none. One case where you say what should have happened instead is worth more than a page of general guidance.

If nobody has asked you yet

That’s the more common problem. Most of this gets built from a guess about how the work is done, and everyone acts surprised when it’s confidently wrong in the places that matter. If that’s happening to your area, send this page back to whoever is building it.