title: The queue date: 2026-09-15
This is the week the loop stopped being the place to push. It ends with a file: 5,946 problems the machine has never seen, ranked by how badly it expects to read them, waiting for an annotator. Getting there took four nulls, one refuted hypothesis, and one fact that reframed the rest.
The parser reads synthetic algebra at 0.98 and ordinary prose at 0.25, and on prose the failure is binding: which quantity each relation takes, and what value each given holds. Since the loop runs seven breaths and re-reads the text at each one, the obvious lever is the attention inside those breaths. We pulled it four ways at read time, each with bars pinned before the read.
A spotlight on the sentences the solver’s contradiction named. A claim mask marking tokens a slot had already taken. A melt of the slots whose reading changed when the sentences were reordered. And, after a census showed that a given slot’s strongest token is a digit 63% of the time on synthetic rows and 6% on prose, a bias pushing prose givens to look at the digits. Zero for four. The last one was worse than null: it lost more the harder it pushed, 122 slots lost against 61 gained at full strength. The lesson was that the census meter was a symptom. On prose the value reaches the slot through the context the frozen trunk already mixed, and pointing the slot at the raw digit removes that route without adding one.
We read the loop’s own path through state space, per row, per breath. Under the training we had trusted, the state moved half its norm at every middle breath with direction cosines near zero, made 12% of its progress to the finish by breath 3, and walked 70% of the way in the last breath. It wandered and then leapt. The decode was finished by breath 2 on every field. Then a control read of the checkpoint trained with no fog at all showed more progress at every breath and no leap. The blur we had introduced to make the middle breaths load-bearing had created the stall: it asked the early breaths to stay soft, so the machine held its commitment and jumped at the end. A version of the fog with per-field deadlines gave the first monotone march the campaign has seen, all cosines positive, and it changed the wild score by nothing. Three geometries, one number. The shape of the path does not move prose.
The perceiver, the organ we built this week to watch the loops, ran its first census on the machine’s own training prose. It gets 2.6% of those rows at least half wrong. Held-out prose reads at 0.25. The machine has fit its training rows and does not generalize across prose. That is not a disease of the loop, the schedule, or the attention. It is what a parser looks like when the phrases and structures of the next problem were never in its diet, and the only thing that has ever moved it is annotated rows of the same kind, which is what the books measured when they scaled.
So the lever is the diet, and the question becomes which rows. Not the held-out set, which is a measurement fixture. Not the training rows the machine already fits. The unseen ones.
There are 5,946 gsm8k training problems outside both the diet and the holdout. The machine read all of them, with no gold, and the perceiver scored each by the meters that had been calibrated on the holdout first: how much the state is still moving at breath 3, how confident the weakest relation is in its arguments, and how many variables the decoded relations reference that nothing defines. Then the fingerpost ran four sentence-permuted views over the top 300 and marked the rows whose parse changed with the word order. The head of the queue reads like the failures we already knew: “divided the pile into thirds”, “guess the total number”, “three full cans”. Partitives and worded quantities, the register a lexicon and a book are for.
Each row goes to the annotator with its text, its answer, and the request: spans for the quantity phrases and the relations, and the factor graph, in the dialect the parser is trained on. What comes back passes a gate before it touches the diet. The rulebook has to hold, and the solver, run on the annotated graph, has to produce the answer key. A row that reaches the answer by a shortcut the text does not state is refused. Every admitted row is silver, versioned apart from the hand-written books, with a sample audited by a person.
We built the perceiver to pull levers between breaths. Its first real job turned out to be choosing what the machine reads next.