The questions are the whole product
This list costs nothing and needs no agents. Run it on yourself with a notebook.
Assembling the situation
- When you say “X”, what specific problem is that supposed to keep you from missing?
- Who besides you pays for how this turns out?
- What can't be put at risk under any scenario?
- What decision should this make possible that you can't make today?
Checking your own wording
- “More mature”, “better”, “more systematic” are labels, not findings. What would someone observe that makes them true?
- Are you looking for a system that already exists, or a mandate to build one?
Checking a person who just appeared in your story
- What claim or action becomes impossible if this person, or this role, doesn't exist?
Checking a branch you're considering
- What are you deliberately building if you stay where you are?
- What does staying postpone?
- What observable change would show this can become what you're hoping for?
- Do you have a mandate to change it?
Bets and signals
- Which actions strengthen more than one branch at once?
- What event weakens each branch, rather than only confirming it?
Checking a concession, asked the moment somebody agrees
- Which of the things you named at the start did you just give up?
The card the session has to end with
Prose evaporates. End the session in fixed fields instead, and the fields do the pushing.
ALREADY DONE carries the load. It takes only what happened while you were sitting there. If nothing goes in it, that's a result, not a gap to fill.
If you'd rather be run than run it
The same thing packaged as a Claude skill: decision-pressure-test. Clone it, drop SKILL.md in, bring a fork you haven't resolved. Forty to sixty minutes, one question at a time and never any answer options. It writes your concession condition down word for word, speaks for the people your decision changes who aren't in the room, checks you every time you agree, and ends with the card above.
It needs a real fork: a date, a cost, somebody besides you who pays. The mechanism runs on you not wanting to give something up, so a practice question gets a practice card back.
Everything below is how I found this out and why half of what I built didn't work. If you came for the tool, you have it.
An answer sorts, a question extracts
An answer sorts information that's already in the room and hands it back better arranged. Extraction builds knowledge that was never in the input, out of remarks, hesitations, constraints and collisions — including the knowledge that your question was the wrong one.
The apparatus isn't mine. Organized collective analysis with declared positions and rules for conceding runs through Georgy Shchedrovitsky, and I name him once so I'm not passing off someone else's method as mine. What I added is a count: an agreement counts only when the person can point at the condition they wrote down beforehand and say what satisfied it. Everything else is logged as uncounted, rounds where nothing moved included.
Positions hold, roles fold
A seat is one party with one stake and one thing it won't trade. Whoever holds it names that stake out loud before the first round, and every time the seat agrees it has to say what it just lost. A seat can be filled, declared empty, or refused entry under a rule.
Give agents roles instead — the skeptic and the finance one — and you get costumes draped over one opinion. Seven of mine converged in the first round, and the run only recovered because the facilitator treated that agreement as fake.
The agents were weaker without a person in the room
The run with a live person in it had a control: the same question through a single strong call from a frozen packet. Two blind judges, origin stripped, forced choice. Both picked the run with the person.
Two judges is not a sample and the count carries nothing on its own, so the part worth reporting is the other one. Both independently named the same missing construct: a branch held at "not activated", opening conditions written in advance, one of them another person's freely given consent, and a ban on treating the launch itself as the test.
The advantage was asking, not thinking
Part of the gap is unglamorous. That branch had a live person to question, the control had a frozen packet, and the fact that closed the question wasn't in the packet. A question pulled it out on the third round. Part of the advantage is informational, not procedural.
The expensive setup doesn't reason better. It elicits better.
The rest is a hypothesis, not a result. Eliciting is a procedure, which means it's comparable and copyable — and it's also the part I never separated from the information advantage. Nothing I ran holds the access fixed and varies only the asking. The ablation is simple to state: give both branches the same live person to question, and let only one of them ask. Until that runs, eliciting is what I think won, not what I showed won.
Thirty-eight to one, and every judge was a model
One full run took 39 model calls and roughly 2.19 million tokens. The same packet through a single call took 58,000. Thirty-eight to one. And the single call's answer was good.
Here's the boundary every number in this note sits inside. Every judge in every run was another instance of a language model, not a person. Producer, judge and scorer come from one model family. Human preference was never measured.
Inside that, the judges called every card. Twelve out of twelve, even after I rewrote them all into one identical format. The tell was a single field: the one naming what got given up and who gave it up. Cheap answers cancel impersonally. The expensive one has somebody paying.
That makes twelve out of twelve a smaller number than it looks. What the judges detected was one formal feature that only one branch ever produces. Build the comparison that way and a hundred percent is the only score available. It reports nothing about whether either decision was any good.
The control I built to kill it beat me
Twenty-eight worker agents produced nine candidate solutions. Critics killed eight of them with stated reasons, and the survivor went further than my expensive setup on the narrowest point in the problem. Worse for me: all four branches I ran — the cheapest single call included — independently landed on the same bottleneck.
Most of the effect was the form, not the argument
That's the finding that cost me the most. Require any answer to come back in the card's fields — seats filled and seats declared empty, entry refused under a named rule, what got withdrawn when somebody agreed. Do that and a single call reproduces nearly the whole result at a thirty-eighth of the price.
The shape of the output was doing the work, not the argument.
One thing the cheap version never did, though. It never took anything away.
Seven seats, and an output made of removals
The material was dull on purpose: whether to add a long-horizon planning module to a personal system of text files. Three prices went in — the quarter's bet, a week-long probe, or skip it.
Seven seats opened on that packet, each a separate agent blind to the others. Four refused the question as posed. The long horizon isn't empty for lack of a tool, they said. It's occupied by phrasing that can't be failed. The facilitator froze that premise and banned the word "emptiness" for a round. Four different things came back where one had been.
The change happened a level above the fork, in what counted as the problem. The horizon looked empty because nothing in it could ever turn out wrong. No date. Nobody to say "failed". The seat holding the quarter dropped one of his own live commitments on that ground.
A criterion whose only judge has never returned a guilty verdict isn't a criterion. It's self-soothing.
Two rounds later the facilitator invented a seat mid-run and filled it with the person who'd be asked to hold the deadline. Three rounds of construction had been standing on a yes nobody had gone and asked for. That seat never conceded, and its own line is the one I kept: you've already counted my refusal as your damage. Its rule wanted a name, a term, and an answer to one question: what will you do if I refuse. The first two existed. The third came back "I don't know". A round later another seat abolished itself, admitting that what it defended had been accumulating without it all along.
Three price tags went in. Nothing new came out. Entry refused under a named rule, with the condition for reviewing it written down. A file created two rounds earlier deleted. A drifting commitment taken over by a named holder, with failure dates in a registry the same day. The week that was asked for handed back by the seat that asked for it.
Advice leaves the decision where it was and adds a line on top. Here the output was made of removals, and they'd already happened by the time the run stopped.
Then I pointed it at a real meeting and it fell apart
I took an hour-long recording of an ordinary working meeting, where nobody had handed out roles, and reconstructed its structure with the same apparatus. Then a skeptic read the source in full and attacked the reconstruction. It didn't survive.
Six seats the reconstruction called empty, and not one held up against the source. Three unrelated frames fit the same quotes equally well, one of them a plain org chart. Seven of nine "positions" turned out to be job titles.
The skeptic's strongest line is the one I kept. Decision rules, the price of a concession, the boundaries between rounds exist only where somebody wrote them down before the first round. In live speech nobody ever does. A test with one possible outcome isn't a test. The one thing that survived came from reading the source, not from the apparatus: concessions in live speech are tonal, not typed.
The run that decides this hasn't happened yet
The next test is one sentence long and I haven't run it. Give the single call as many rounds of clarifying questions as the expensive version gets, put somebody on the other side whose job is to break whatever comes back, set a date, compare.
Two more things go in with it, written down now so I can't quietly drop them later.
A factorial run: format without questions, questions without format, both, and the agent version. Without that split I can't say which half does the work, and today I can't.
A delayed outcome: thirty days after the session, one check — did the decision hold or did it get rolled back. That's the only quantity in any of this that doesn't depend on the shape of the answer.
The second one is there because of what's missing from everything above. Every number in this note measures how full the card came back. Decision quality was never measured. Not once.
Here's the falsifier in advance. If the gap closes, this apparatus is a way of interviewing yourself and the cheap version interviews just as well. The rerun is designed and the result goes up in August whichever way it lands.
The list at the top is free. Run it and tell me which question did nothing.
Get the next plate in your inbox.
One strategic memo or weekend-buildable AI product idea per week. No noise. Confirm your email to activate the subscription.
Working on AI transformation?
If you are building, managing, or advising AI transformation inside a company, I'd love to compare notes.
Where does it break for you — tools, workflows, org design, or decision-making?
