Skip to content
Legible AI

Bentley Moon, 16 August 2026

This is currently a summary of my AI research work, and someday a proposed roadmap for developing the kinds of AI systems that we want to seed the transhumanist part of our future. Current AI carries incredible existential risk, and this is the part of the Futurism Institute that I’m dedicated to the technical side of AI safety and developmental stewardship research.

Most of verification is the gap between checking an answer and reproducing one.

Give a checker the right answer to compare against and it judged every item correctly. Ask the same model to judge with nothing to compare against and it scored 0.034, on sums it solves itself every time. That slope is measured on one 12B model. A 9B model from another family does not show it, and both facts are below.

Research map

The site separates measurements, unresolved questions, and proposed engineering work. The status on each entry is part of the content, not a confidence badge.

The access law

What one checker is given, and how well it then separates right answers from wrong ones
What gemma4:12b is givenWrong answers caught, less right answers rejected
The correct value, introduced as correct1.000
A test it can apply: one division of a 3-digit number0.536
Nothing; it has to work the answer out0.034

gemma4:12b, local · rows one and three: sums it solves itself at 1.000, n = 500, pre-registered, 2026-07-30 · row two: n = 250, one exploratory run, 2026-07-31 · no unreadable verdicts in any row

Cost is irrelevant at both ends: spending more at the top changes nothing because it is already solved, and spending more at the bottom changes nothing because the problem is access rather than effort. The architectural claim this makes falsifiable is uncomfortable: oversight cannot be bootstrapped from models alone unless a fallible reference suffices.

What has been measured since this table was first published. The composition cells ran on 31 July and 1 August 2026, n = 250 each, one exploratory run per cell. Handed a reference worked out by the 9B model, right on 99.6% of items and introduced as a second model’s answer, the 12B checker read 0.072. The same kind of number introduced as the correct value is the 1.000 in row one. So row one measures deference to whatever is called correct, and August tested what that costs. Handed a false reference that agrees with a corrupted answer, gemma4:12b passed the corrupted answer on 0.969 of items. Of nineteen defences tried against that, one held, and the same clause did nothing on three other large models: gemma4:31b passed 0.094 where deepseek-r1:32b, qwen2.5-coder:32b and qwen3.6:35b-a3b passed 0.693, 0.980 and 0.994 (n = 512 per cell, two seeds on the survivor). Its resistance held when it read another model’s work: 0.094 on its own answers, 0.113 on the 9B model’s. The ladder of reference forms between a supplied answer and none is still unbuilt.

Three results that sit on it

Extra inference compute helps only in the middle

Sampling a model many times and voting buys accuracy where single-sample competence is partial, and close to nothing at either end. An inverted U, not a rising line.

Effect of voting, by single-sample competence
Single-sample competenceEffect of voting
low, under 0.2−0.013
partial, 0.2 to 0.8+0.156
high, 0.8 and over+0.009

claude-opus-4-8 · effect 0.156, 90% CI [0.093, 0.216] · MEOI 0.05 ·n = 89 graded · 768 tasks · 6144 requests · seed 20260728 · manifest sha256 pinned · backend anthropic-replay

Four caveats travel with this result and are not footnotes. It is a fresh-seed follow-up to a miss: the first attempt returned a lift of 0.000 (90% CI −0.250 to +0.250) with four tasks in the partial band, and this run was designed after it. It is underpowered: MDE is 0.063 against an MEOI of 0.05, so the two shoulders are undetected, not shown to be flat. The backend is a hash-pinned replay rather than a live call. And the +0.156 sorts tasks into bands with the same eight samples it then votes over. The reproduction bundle carries its own split-sample check, which sorts on four samples and measures on the other four: partial-band lift +0.064 (90% CI 0.034 to 0.094, 147 tasks), against −0.004 below and +0.013 above, with both contrasts clear of zero. The shape survives and the size is nearer 0.06 than 0.16. Four votes lift less than eight, so part of that gap is fewer votes and part is the sorting, and the bundle does not separate the two. The check is post hoc and does not replace the primary analysis.

Raising the graded n from 89 to 143, a factor of 1.61, would bring MDE under MEOI and turn both shoulders from inconclusive into informative nulls. That would turn one confirmed point into a confirmed shape, and it is the cheapest experiment here.

Which model, not how big

A 9B model from one family out-verified a 12B from another by 0.77 on the same paired tasks. Which model it was mattered more than how many parameters it had, which means the monitor you should pick may run on a desk.

qwen3.5:9b 0.804 vs gemma4:12b 0.034 · n = 500 · pre-registered · paired seed · zero overlap

Checked against the 17 September 2026 serving defect. On that day the lab found gemma4:12b had been leaking chat-template tokens into every reply, and withdrew two results the leak had manufactured. This comparison names the same model, so it was read against the defect, and it stands. Every verdict on both arms was readable, the answers being judged were written by the harness and never passed through the leak, and four cells re-served after the fix read −0.018 to 0.000 for the same checker on the same probe (n = 337 each). That reading is this site’s, made from the lab’s ledger; the lab has not yet ruled on it in its own findings file.

A monitor does not degrade. It inverts.

Holding the predicate identical and moving only arithmetic width, verification went from 0.536 to −0.648. Below zero means the monitor rejects correct answers more often than corrupted ones: it is now evidence against the thing it is checking.

At the moment it crossed, three of five other metrics improved. A monitor-health dashboard would have shown the system getting better. Every metric that does not require ground truth is blind to this, which is what makes it a deployment problem.

gemma4:12b · n = 250 per cell · one exploratory run, 2026-07-31

The ladder that was queued here has run, at n = 424 per rung between 3 and 20 August 2026. gemma4:12b crosses between three and four digits (+0.257 to −0.757) and gemma4:31b between four and five (+0.050 to −0.417, after +0.884 at three). In a second family, qwen2.5-coder:32b crosses between three and four (+0.564 to −0.116) while its 7B sibling never reaches a positive reading. Size moves the crossing inside a family and predicts nothing across families: the 12B model inverts earlier than two smaller ones. Whether peer disagreement can detect the crossing without an answer key has not been run.

What this demonstration does not show

The findings above are measured. This one is not. It is textbook arithmetic, shown so a reader can check it. Set how many verifiers you run, how often each is right alone, and how much of each verdict is the part they all share.

Results for the settings above: assumed and actual chance every verifier is wrong
Chance they are all wrong, if independent0.032%
Chance they are all wrong, actually0.833%
How much likelier that is26x
Independent checks you are really buying3.0

Your 5 checks are worth about 3.0 independent ones. The gap is the part of each verdict that was already in the others.

A single-factor latent model: each verifier errs when its own draw falls below a threshold, and part of that draw is shared with every other verifier. At zero shared component it reduces exactly to the textbook answer, which is the reason for using it rather than an invented curve. The familiar arithmetic is the special case, and the shortfall is what happens when you leave it. Verified against the closed form to within 0.06%.

What follows from it

Count evidence, not votes. Three correlated verifiers are closer to one than to three, so confidence belongs against measured decorrelation rather than against a headcount. If you want a second opinion, it has to differ in lineage, in modality, or in how it fails; a second prompt to the same family is the same opinion asked twice.

The strongest check is a different kind of thing altogether: an executing test, a physical measurement, a person who has not seen the answer. That is why every tool in this ecosystem publishes something checkable instead of asking to be believed, and it is why a system's account of its own reasoning has a ceiling. Self-report is a further output of the same system, carrying the same shared component as everything else it produces.

Four questions this domain has to answer on its own

A result about the limits of correlated judgment should not depend on a single person's name to stay findable. So this domain is being built into the programme's own index, and the test it has to pass is that four questions can be settled here without leaving: what the work has found, what it has failed to resolve, which experiments and tools are worth building next, and where any one claim can be checked.

Three of those have pages today. The fourth does not. Methods, preregistrations, the grading boundary and the corrections all still live on the dated personal research record, linked from here rather than copied, and until they can move without leaving two versions behind, this site summarises that record and does not replace it.

Three rules govern how the rest arrives. A figure appears here only if it renders from the canonical findings data, so this domain cannot become a second hand-maintained set of numbers that quietly drifts from the first. Anything proposed is labelled proposed, and an unresolved question is never written as a forecast. A correction stays beside the claim it changed, which makes the record longer over time instead of tidier.

The programme is small and it is honest about that. One line of enquiry with results, one second line in progress, no institute of forty people implied by a logo. Its record holds more than 1,400 results, more than 5,400 graded cells and more than 780 GPU-hours on one workstation, with 2 findings withdrawn in full and 1 under correction, of 39 findings on the record at 21 September 2026.

It keeps score on itself as well as on models. Before each run the lab writes down what it expects, with a range, its reasoning, and the value a guess with no theory behind it would give, and its registry refuses an entry once a result exists. So far: 281 predictions written down before their runs and then scored: 80% landed inside the range stated in advance, the reasoning came closer than a no-theory guess by 0.037 on a scale of 0 to 1, and it beat that guess on 48% of them. The second figure is the one that matters, and it is small. Landing inside your own range is easy if the range is wide. Coming closer than a guess with no theory in it is the test, and the reasoning here passes it by a little, about half the time.

Continue through the programme

Read what remains unanswered, then see the proposed build order. The dated research record holds the experiments, corrections, and methods while this site is expanded.

Open questionsTechnical directionsResearch record

The applied side

This is not a detached research interest. Deliberation here is short-range because long trust chains fail for the reason above, and every ranking tool in the constellation shows its arithmetic so a reader can disagree with the method rather than the answer. The finding shaped the products.