Case study

Can You Induce a Domain You Do Not Already Understand?

Prime Phase as a worked example of domain induction and program synthesis: what was induced, what was synthesized, how it was checked, and the one step no program in the loop could do.

10 min read

A known answer, an inducer who did not know it, and the probes on the record.

I spent two days on consecutive prime gaps with no background in number theory. I started from a geometric hunch, wrote one analysis program after another — 36 by the end, counting the falsification work that came later — and ended at a one-parameter law that turned out to be Harald Cramér's random model from the 1930s. The full story — the helicoid, the moiré, the 99.2% model that predicted almost nothing — is in The Prime Knows Where It Is Because It Knows Where It Isn't.

This piece is about what kind of work that was. I think it is a small, complete example of two things people in AI are trying to get machines to do: domain induction, working out how an unfamiliar domain should be represented from observations alone, and program synthesis, producing programs that meet a specification. Here a person did both by hand, and there is a known answer at the end to grade the result against. That last part is what makes it useful.

Why a known answer matters

Domain induction is hard to evaluate. When a system — or a person — proposes a new way to represent a domain, you usually cannot tell whether it found the right representation or merely a workable one, because nobody knows the right one.

Prime Phase avoids that problem. For the observable I was studying, the right representation is known: treat each number that survives the small-prime sieve as independently prime, with a probability set by the prime number theorem. I did not know that when I started. I did not know what the Cramér model was, or a singular series, or an Euler product; I learned each name when the work forced it on me. So the episode has the two ingredients a fair test needs: an inducer with no access to the answer, and an answer to grade against afterwards.

That changes what counts as success. Arriving at a century-old model is not a disappointment here; it is the pass condition. It means the representation I induced is the one in which the known model is exact for what I was measuring.

What was induced: a representation, not a law

The distinction that matters is between learning a function inside a representation and inducing the representation itself.

For most of the two days I did the first. I modeled the transition law — the next gap given the current gap — in raw gap coordinates, and kept improving it: a sieve prior, then a row-local tilt, then interaction terms between pairs and triples of small primes' rhythms. By the end the model had 218 features and recovered 99.2% of the serial dependence, out of sample.

It predicted the correct next gap 18% of the time. Ignoring the current gap entirely gave 16.5%. All that structure was worth 0.34 bits out of roughly 4. A better fit in the wrong representation had bought almost nothing.

What changed the outcome was a change of representation. Conditioning on board state — where the current prime sits in the cycle of a small-prime wheel — showed the same gap leading to different next gaps on different boards. Measuring position as slot rank, "the k-th candidate the wheel leaves legal," instead of raw distance made those differences disappear. One shared matrix matched 48 board-specific matrices (NLL 2.044 against 2.031), and about 90% of the apparent board dependence turned out to be an artifact of the coordinates.

What came out is a short description of the domain:

integers
   ↓ remove multiples of the small primes (a wheel of size W)
wheel-legal candidates
   ↓ index each candidate by its rank among the legal openings
slot rank
   ↓ each legal candidate is the next prime with probability q
geometric waiting time, q = (W/φ(W)) / ln x

Every arrow in that diagram is a distinction I did not start with. Each one was forced by an observation and kept only after it survived a test:

Distinction

What forced it

What it survived

Measurement apparatus vs. signal

The wrapped-phase spectrum looked structured

A random sequence with the same density produced the same spectrum, matching the prediction from prime density alone (correlation above 0.99): the signal was my transform

Admissibility vs. probability

The small-prime sieve says which gap pairs are possible

Built into a transition model, it explained only about 30% of the anticorrelation: legal is not likely

State-relative vs. state-invariant coordinates

The same gap predicted different next gaps on different wheel boards

One slot-rank matrix matched 48 per-board matrices; 90% of the board dependence vanished

Dependence vs. projection

Rows of the slot matrix were nearly identical

A one-parameter geometric law matched the full matrix within ΔNLL 0.004 at wheels of 30, 210 and 2310

Fitted parameter vs. derived quantity

The geometric parameter changed with wheel and scale

q = (W/φ(W))/ln x predicted it across 24 wheel-and-scale measurements: correlation 1.0000, best scale factor 0.9987

Residual vs. measurement bound

A residual of about 0.3% remained, with a sign pattern

It shrank as my max-gap cap rose and vanished without the cap

That table is the induction. None of these distinctions is new to number theory. What I did not have, and had to recover from observations, was the vocabulary — and the right vocabulary is what made the law simple.

What was synthesized: two kinds of program

"Program synthesis" can sound like a stretch for a person writing scripts, so here is the precise sense in which it applies.

Instruments. Each hypothesis became a program whose job was to tell that hypothesis apart from a specific alternative. The specification was never just "compute X"; it was "produce a number that would come out differently if I were wrong." The repository keeps those programs, so the claim can be checked against them:

Hypothesis

Program

Alternative it had to beat

Outcome

The wrapped spectrum is prime structure

phase_velocity_test.py

Random sequences with the same density

Same spectrum: it was density

Gaps carry prime-specific dependence

dechirped_residual.py, dechirped_v2.py, dependence_isolation.py

Random sequences, then a mod-30 wheel-filtered null with a matched histogram; shuffles

Lag-1 anticorrelation of about −0.057 survives

Small-prime interference explains it

gap_model_validation.py, ipf_residual.py

Held-out blocks and block bootstrap

About 30%; the rest needed something else

Pairwise and triple beats close the gap

moire_test.py, moire_triples.py

Blocked cross-validation

97.9%, then 99.2%, out of sample

The fitted model predicts primes

predict_primes.py

A model that ignores the current gap

18% against 16.5%

The law depends on board state

tetris_model.py, comoving_frame.py, comoving_ladder.py

48 per-board matrices; per-wheel holdout

One slot-rank law, geometric at every wheel

The hazard is just density

renewal_test.py

A parameter-free prediction over 4 wheels × 6 scale bands

Correlation 1.0000

The leftover 0.3% is real

density_form_falsification.py

Next-order density forms, a power-law exponent, no gap cap

Residual gone; exponent 1.0000

These programs were not implementations of mathematics I already knew. They were how I found out which mathematics I needed. I did not decide to study the singular series and then write code for it. A probe built around overlapping modular rhythms — fizzbuzz, more or less — showed those rhythms doing real work, and only then did I go looking for the formula that computes them. The program came first and the name second.

The model. The other program is the one the instruments were evaluating: a generator for the next gap. Its size is the clearest record of the episode.

Representation

What had to be specified

What it bought

Raw gaps, sieve prior

Small-prime admissibility, calibrated to observed frequencies

About 30% of the anticorrelation

+ row-local tilt

10 fitted features

94%

+ pairwise interactions

168 more

98%

+ triple interactions

40 more, 218 in total

99.2%; next-gap accuracy 18% against 16.5%

Slot rank, shared matrix

One matrix for all 48 boards at W = 210

Matched the per-board matrices, and beat them out of sample

Slot rank, geometric law

One parameter, r ≈ 0.705 at W = 210

Matched the full matrix within ΔNLL 0.004

Slot rank, derived hazard

No fitted parameters: q = (W/φ(W))/ln x

Predicted r across wheels and scales (correlation 1.0000)

Read from the bottom up, that table is a synthesis result. The target program went from 218 fitted features to none without losing accuracy. It did not get shorter because I searched harder: searching harder in raw coordinates is exactly what produced 218 features. It got shorter because one new primitive entered the language — rank among legal candidates — and in that language the program is nearly trivial.

That is the part of program synthesis that is hardest to automate: not searching a fixed language for a program, but changing the language so that a short program exists. The trajectory also shows the signal that should have triggered the change sooner. While I was adding features, the model kept growing and its predictive value stayed flat. That combination is computable. In hindsight it was telling me the representation was wrong well before I listened.

The checker was a program too

This is the part I would point a synthesis researcher to first.

With the derived hazard in place, a residual of about 0.3% remained, with a sign that flipped between small and large wheels. It looked like the most interesting thing in the project: the place where real structure beyond the simple model might live.

In synthesize-and-verify terms, it was a counterexample to the model. The response was to try the cheapest explanations before believing it:

  • A missing term in the density estimate. Two standard next-order corrections both fit worse than the plain form, at every wheel and every cap tested.

  • My own measurement. The fit capped gaps at 80, which cuts off the geometric tail. Raising the cap shrank the residual steadily. Removing the cap took it to 0.08%, flat across wheels, for primes up to 8 million, and to 0.00% for primes up to 1 billion.

  • A hidden second parameter. If the true relationship were q_emp = β·q_pred^α with α ≠ 1, a linear fit could average it away. At 1 billion, α = 1.0000 ± 0.0000.

The counterexample came from the instrument, not the model. In a loop where programs check programs, the checker is part of what has to be falsified: an arbitrary bound in a measurement script produced the most interesting-looking result of the whole project.

Which parts were mine

It would be easy to overstate this, so here is the division as honestly as I can draw it.

What I supplied, and could not have specified in advance: the frames. A helicoid that wrapped the primes onto a spiral. Fizzbuzz, as overlapping modular rhythms. Moiré, as the beat between two of those rhythms. A Tetris board, with the wheel state as the board the next prime lands on. Most of these were wrong about the specific mechanism and right about the kind of structure involved. The slot-rank coordinate came from looking at board-state results and seeing that the board, not the law, was moving. No program in the loop proposed it.

What was mechanical once a frame existed: writing the probe, choosing its alternative, holding data out, reading the number. Tedious, but specifiable.

What could have been mechanized but was not: the alarms. "The model grows and predicts no better" and "the residual moves when my measurement bound moves" are both checks a program can run continuously. I noticed each of them late, by hand.

What nobody mechanized: proposing the next coordinate. A check can tell you a representation is failing — the 218-feature model failed it plainly. It cannot tell you which representation to try instead. In Abstraction First I called this the gap between a falsifier and a generator. Prime Phase is a clean trace of a person acting as the generator, with the rest of the loop on the record.

What this does not show

  • No new mathematics. The endpoint has been known since the 1930s, and a number theorist would have reached it much faster.

  • One observable. Consecutive prime gaps, primes up to 10⁹, wheels up to 2310. Real structure beyond the independence model exists in the literature; this observable at this depth does not show it.

  • One inducer, one domain, one episode. It is a worked example, not a benchmark, and on its own it says nothing about how often this process succeeds.

  • No automated system did any of it. Program synthesis here means a person turning hypotheses into programs. The claim is about the shape of the work, not about a tool.

  • Not an argument for ignorance. Not knowing the field kept me off the standard path. It also cost me much of two days rediscovering standard results.

A shape for testing induction

If you want to know whether a system can induce a domain, Prime Phase suggests a setup:

  1. Pick a domain whose right representation is known for some observable, and keep it from the inducer.

  2. Record every probe as a program, together with the alternative it was written to rule out.

  3. Track the size of the model against what it predicts, and treat "growing but not improving" as evidence against the current representation.

  4. Grade the end state on whether the known model becomes exact in the induced representation, not on whether anything new was found.

  5. Treat the measurement programs as falsifiable too.

What this setup can grade but not supply is where the next representation comes from. That is the open problem I keep returning to, and the next pieces in this series take it up from different directions.

Source material