Artificial Intelligence | AI systems | abstraction | representation learning | evaluation | distillation | program synthesis
Abstraction First: Is This a Real Quotient or Just a Re-encoding?

Once you think you've found an abstraction, prove it preserves what you need.
Opening claim: "A better fit may mean your coordinate system is worse."
The industry-recognizable problem this article owns: compression mistaken for abstraction.
The recognizable failure
AI systems compress constantly, and every compression is offered as a simplification:
a distilled model that matches the teacher on the benchmark distribution;
a memory summary that "captures the essentials" of a long context;
a program-synthesis IR that lowers every expression in the corpus;
a benchmark-specific scaffold that lifts the score;
a "world model" that predicts the next observation well.
Each is smaller, faster, cleaner — and each carries the same hidden question: did the distinction you will actually need survive the reduction? A high score on the compressed representation is not evidence of preservation. The compressed form can agree with the original on everything you happened to measure and still diverge on the case that matters, because you measured on the axis the compression preserved rather than the axis it destroyed.
The conventional response
Add capacity back. Blend the compressed representation with the raw one. Ensembles, residual connections, retrieval augmentation, more features. If the number is good enough, ship it.
Why that response is insufficient
Every one of those responses operates inside the representation that produced the failure. More features compensating for a wrong coordinate system is the failure mode itself, not its repair — each addition is a stabilizer bolted to a moving camera. The question that escapes the loop is not "how do I fit better here?" but "did the move to this representation preserve the property I'm reasoning about — and could I tell if it didn't?"
The transform
Abstraction-first is one demand placed on every proposed representational collapse, stated before any trust is extended:
Name what must be preserved before trusting the compression, then falsify.
Compression is not correctness. Reducing to a canonical form is legitimate only if what mattered survived the reduction. A representation that is smaller but loses the distinction you were reasoning about is not a simplification — it is a re-encoding that has quietly discarded the thing you cared about.
This article is the entry point for a series about representational change, so it carries the two invariants every member of that series obeys, stated once here and inherited thereafter:
Declared authority. The axis, the canonical form, the classification — each cites a provenanced source, never a constant or a rendered string. A resolution that bottoms out in a hardcoded value has bypassed the move.
Proven preservation. Every collapse — a reduction, a lift, a reframe — carries a witness that what mattered survived, or a counterexample where it did not. Compression without a preservation witness is re-encoding, not simplification.
Minimal formal shape
Not every abstraction preserves the same way, and honesty requires saying how strongly a given collapse is claimed to hold. Every proposed abstraction is classified into exactly one of six buckets, from strongest to weakest:
Exact equivalence — full preservation, smaller representation.
Exact under assumptions — correct only if stated constraints hold.
Exact on a subclass — correct for bounded or structured instances.
Parameterized tractability — efficient when a parameter stays small.
Approximation / heuristic — useful, but does not preserve optimality.
Intuition only — no proof obligation met.
The buckets are a closed vocabulary, not free prose, and the justification obligation scales with the claim: exact on a subclass has to name the subclass; exact under assumptions has to state the assumptions. A stronger bucket is unrepresentable without the backing that would let a reader falsify it.
A preservation claim is not a boolean a function returns and forgets. It is a witness: a content-addressed record binding the proposed abstraction to the evidence for or against it. When preservation fails, the failure is a counterexample bundle — the concrete violating instances, with replay evidence. A falsifier that cannot show you where it broke is not a falsifier, only an opinion.
A concrete experiment
One case is an anecdote. This thesis now has three domains, and the last two are the additions that make it a replication.
Domain 1 — prime gaps: the re-encoding that fit at 99.2%
I spent two days modeling the transition law of consecutive prime gaps in raw gap coordinates. The layered model was, in hindsight, a documentary of the trap:
Order | What it captures | Features | Cumulative |
|---|---|---|---|
0th: sieve prior | each prime's exclusion pattern | — | 30% |
1st: row-local tilt | scale-varying mean reversion | 10 | 94% |
2nd: pairwise cross-terms | beats between two primes' rhythms | 168 | 98% |
3rd: triple cross-terms | compound beats of three primes | 40 | 99.2% |
218 features total, 99.2% of the serial dependence mapped — and every added feature was compensating for the same coordinate error. The prediction reality check said so directly: the full model predicted the correct next gap 18% of the time against a 16.5% marginal; all that structure was worth 0.34 bits out of ~4. The switch to slot-rank coordinates ("the k-th legal opening," wall-attached instead of camera-attached) collapsed the model to a one-parameter geometric law, r ≈ 0.705, matching the full transition matrix to within a NLL difference of 0.001 across three wheel sizes and six scale bands. The parameter turned out to be a 1930s density result (Cramér's). All 218 features were projection artifacts — accurate re-encodings of a simple object viewed through the wrong lens. The full narrative is the companion case study; the point here is the shape: high fit, wrong coordinates, quotient discovered by noticing what the fit couldn't buy.
Domain 2 — MH-2c: the same error shape with a mechanized falsifier
A transfer mechanism for cross-domain planning showed 5 of 6 beneficial BFS ratios — by every within-model measure, successful transfer. Execution replay then showed every plan from that path failing on the real kernels. The benefit was faster search in an approximate world model that did not preserve feasibility: faster hallucination, not transfer. The preservation test was the replay itself, and it supplied the honest retroactive classification — the combined-evidence method was demoted to the approximation / heuristic bucket (a scratch-world model preserved feasibility exactly, which is what isolated the failure to the combination step). Same error shape as the prime-gap feature trap, unrelated domain, and the falsifier was mechanical rather than narrative.
Domain 3 — the cube encoding probe: re-encoding does not help, as a recorded verdict
The strongest version, because the experiment was designed to discriminate. A domain-induction system was overfitting on Rubik's-cube state, and there were two candidate diagnoses: the state encoding was wrong (a Python-side problem) or the grammar of available operators was wrong (a native-side problem). So two genuinely different re-encodings of the same state — a colour-landing encoding and a slot-permutation encoding, each field recording which source slot fed it — were each driven through the governed native bridge and each replayed on a held-out transition.
Neither generalized: 10 of 24 fields off for one, 12 of 24 for the other. The recorded verdict is GRAMMAR_GAP — re-encoding does not help, because the gap is the grammar's lack of a relabel/read-from-field operator, not the state encoding. The verdict was locked by a verdict-asserting test (verdict == "GRAMMAR_GAP", both encodings fail) and a native-line witness that the emitted operator set contains only additive ops. That test has since been retired from Sterling's live suite, so the verdict now stands on its recorded manifest rather than on a running assertion. The verdict name was chosen deliberately: it records the gap as belonging to the grammar so that a future grammar with a relabel operator could legitimately change it.
Three domains, one thesis: when the representation is the problem, improving the fit inside it is re-encoding, and the only thing that distinguishes a real quotient is a preservation witness at the point of collapse.
The falsification condition
What would show abstraction-first is just another patch? Three conditions, and where each one stands:
The check never says no. A preservation check that has never failed over a meaningful population is not a check. This one has said no in all three domains above: to the 218-feature model, to the combined-evidence transfer, and to both cube encodings.
The witness is computed downstream. A preservation record derived after the fact, from the compressed representation's own outputs, grades its own homework; the property has to be recorded at the transform's own edge, where the collapse happens. Here the record is mixed. In each domain the "no" came from outside the compressed representation — the actual next gap, the real kernels, a held-out transition — so none of those verdicts was self-graded. But each was reached after the fact rather than recorded at the edge, and the one witness that does record at the edge is domain-specific and advisory (see the non-claims below).
In-distribution agreement stands in for held-out transfer. Agreement on the axis the compression preserved is vacuous. The test that counts is held out: a transition the abstraction did not see, a kernel the planner did not search, a scale band outside the fit. Domains 2 and 3 were decided by exactly such tests. A claim made with this method that skips one should be read as a re-encoding until shown otherwise.
AI-system implications
Each compression this article opened with owes the same two things: a named property, and a test that could show it was lost.
Distilled models: a student that matches its teacher on the benchmark distribution is, at best, exact on a subclass — and the subclass is the benchmark distribution. Name that bucket and that subclass, and test where teacher and student could come apart, not where they were fit to agree.
Memory summaries: a summary that preserves next-token plausibility has not been shown to preserve the distinctions later reasoning will turn on. Declare what has to survive — the commitments, constraints, and open questions — before compressing, and check the summary against that list. The obligation sits with the summary, not with the downstream task that fails quietly.
Program-synthesis IRs: an IR that lowers every expression in the corpus has shown coverage, not preservation of the structure search depends on. The cube probe is this failure in miniature: both encodings represented the state, neither generalized, and the gap was in the operator grammar. Test whether the operators can express a held-out transition, not whether the corpus lowers.
Benchmark scaffolds: a scaffold that lifts a score has passed an in-distribution agreement test. Held-out transfer is the test that counts, and a benchmark that cannot express it will keep rewarding re-encodings.
World models and agents: MH-2c's planner showed beneficial search ratios in 5 of 6 cases inside its model, and every plan from that path failed on the real kernels. A plan that succeeds in the model is a verdict about the model, not standing to act in the world: an observation, a verdict, or a successful computation is not itself authority for the inference someone wants to draw from it. Replay against the real system before calling model-side success transfer.
The non-claim
No generator. Abstraction-first answers one question — given a proposed abstraction, does it preserve? — and not the other: which abstraction should be tried next? In all three domains, a person found the coordinate switch. The separation is deliberate. A preservation check that starts proposing candidates has taken on a different job with its own evidence burden, and it becomes a search that invents abstractions and grades its own homework.
No general-purpose witness. An earlier implementation of the preservation check was retired and deleted; the requirement stands, the artifact does not. The closest working shape is the coordinate-retention witness. It declares a binding that must survive a projection — a
(role, source-token-index)coordinate — records it at the projection's edge, and returns a typed verdict (RETAINED,DROPPED, orUNDETERMINED) with an inventory of what was dropped. It is a census of one projection's drops, not a proof of a quotient, and its recorded boundary holds: necessary but not sufficient, post-run only, and advisory. It flags; it never admits.No learning from past verdicts, and no self-repair. The series this article opens makes two further non-claims. Nothing here compiles a history of adjudicated verdicts into a policy for where to look next; that learner is deliberately undesigned. And nothing amends its own failure taxonomy or invents new coordinates: detecting that a representation is insufficient is documented, repairing it autonomously is not, and no mechanism for it is reserved.
No "verified by a passing suite." The sources below cite the recorded claim, the code, and, where one still exists, the test that asserts the verdict. Test counts in Sterling's lane ledger are static counts of test functions on disk, not records that any test ran.
The open problem
Two things would have to transfer, in order.
The first is the witness. For abstraction-first to be a general operator rather than a discipline applied by hand, the witness has to generalize: from one projection's census to arbitrary transforms, and from a post-run diagnostic to a precondition checked before trust is extended — still necessary rather than sufficient, and still not a license. Nothing described here does that yet.
The second is the generator. Abstraction-first can tell you, mechanically and with counterexamples, that the representation you trusted does not preserve what you need. It cannot tell you which representation to try next, and letting the checker propose would erase the boundary that makes it a checker. A generator would need proposals that carry their own preservation obligations, failures recorded as counterexamples rather than discarded, and a check it does not control. Nothing in Sterling claims that exists. It is the next problem, and this article is the standard it will be held to.
Source material
Sterling is a private repository; its paths are given to pin each claim to a specific file or commit, not as links.
Domain 1 — prime gaps: The Prime Knows Where It Is Because It Knows Where It Isn't and github.com/darianrosebrook/prime-phase; every claim is checkable by re-running the committed scripts.
Domain 2 — MH-2c: Sterling design records 0023 and 0077; commit
b72811adb(the execution replay);docs/theory/abstraction-first-audits/000-origin-and-throughline.md, 2026-03-20 entry.Domain 3 — cube encoding probe:
proving-grounds/dips-cube-encoding-probe— verdictGRAMMAR_GAP, recorded in the probe'smanifest.yaml. The verdict-asserting test was added in commit3fc3937768and has since been retired from the live suite.Coordinate-retention witness:
proving-grounds/coordinate-retention— typed verdicts, the advisory boundary, and a live-run drop inventory.Doctrine:
docs/theory/abstraction-first.md(falsifier, not generator; the six buckets; the retired implementation) anddocs/theory/collapse-family.md("verdict is not standing"; the series' two non-claims).Ledger caveat:
proving-grounds/scripts/inventory.py— the ledger'slocal_testsfield is a static count of test functions, not a record of runs.
