Lab · record
Sixteen perspectives
Sixteen named perspectives were asked for research directions in the conjecture form, on Claude Opus 5 and GPT-6 Astra; eight were asked again on Claude Fable 5.1 as well, with the working paper in the packet. Three blind judges, one per lineage, each ranked their own provider family first in every case. An outside check by GPT-6 Astra corrected the editor's record; a decision model from a third family, TypeSafe's Jev, judged the same triples blind and never chose Astra. The directions that survived, and what the judging showed about judging, are below.
What this is #
Two rounds, three judges #
What was asked #
The author, 20 September 2026, asked for a review of Space Immanence "from multiple different perspectives (really being creative and innovative of what these are exactly)" to surface directions worth taking the research, run on both Claude Opus 5 and GPT-6 Astra because, in his words, "I am worried Opus may not be good enough." Round one gave sixteen named perspectives — a quantum-gravity phenomenologist, a category theorist, an illusionist philosopher, a control engineer, a Madhyamaka scholar, a historian of science, and ten more — an identical brief and an 8,809-word packet built from the site's own served texts, with instructions to say what the record is missing and then propose exactly three directions in the conjecture form (claim, falsifier, cheapest test with cost, what it adds), reading the packet only, under 900 words.
Round two was pre-registered as an amendment after round one's brief was found to steer both lineages toward one genre — a small controller with and without a self-model — and to starve the physics, formal and philosophical perspectives. It narrowed to eight roles the first round had starved or where only synthesis could help, added Claude Fable 5.1 as a third arm on the author's decision, and changed the brief to one up-to-500-word analysis of what the record is missing, one direction worked with its first step done on the page and costed at three tiers (under USD 20; under USD 500; under USD 10,000 with a named kind of person's hour), and one direction with no cost limit, all under 1,700 words, single-shot with no tools and no file reads, against a roughly 22,000-word packet carrying the working paper in full, the ledgers, the persistence conjecture as adopted, and round one's ninety-six direction titles.
What the judges returned #
Round one ran one judge only: the blind GPT-6 Astra judge preferred the Astra output in 12 of 16 pairs (mean score 23.06 against 19.38 out of 24), a sign-test p of 0.077, short of the pre-registered 0.05. The author redirected the round-one budget to round two before the Opus and Fable judges for round one were run, so round one's comparison is one judge and stands incomplete by that decision.
Round two ran all three judges — Claude Opus 5, Claude Fable 5.1 and GPT-6 Astra, each blind to lineage — over all eight triples, scoring the worked direction, the first step, the missing-analysis paragraph and the unlimited direction to a maximum of 14, then ranking the three outputs and naming five picks each across all twenty-four:
| judge | mean score Fable / Opus / Astra | mean rank Fable / Opus / Astra | Fable v Opus | Fable v Astra | Opus v Astra |
|---|---|---|---|---|---|
| Claude Opus 5 | 13.88 / 13.38 / 9.75 | 1.38 / 1.62 / 3.00 | 5 to 3, p = 0.73 | 8 to 0, p = 0.008, Fable | 8 to 0, p = 0.008, Opus |
| Claude Fable 5.1 | 13.38 / 13.00 / 7.88 | 1.38 / 1.62 / 3.00 | 5 to 3, p = 0.73 | 8 to 0, p = 0.008, Fable | 8 to 0, p = 0.008, Opus |
| GPT-6 Astra | 10.12 / 9.38 / 13.75 | 2.25 / 2.75 / 1.00 | 6 to 2, p = 0.29 | 0 to 8, p = 0.008, Astra | 0 to 8, p = 0.008, Astra |
Under the design's pre-registered rule (two of three judges finding a difference in the same direction, at p < 0.05), no judge found a difference between Fable and Opus, and both Claude judges put Fable narrowly ahead on score and rank. Read mechanically, the rule does find a difference between the Claude arms and Astra: both Claude judges ranked a Claude output ahead of Astra in all eight triples, and the Astra judge ranked Astra ahead of both Claude arms in all eight — the same 8-to-0 split, run in opposite directions by different judges.
What the judging showed #
The editor does not report the rule's conclusion as a finding about quality. The two Claude judges ranked a Claude output first in all eight triples (Fable first five times, Opus three) and Astra last in all eight; the Astra judge ranked Astra first in all eight and a Claude output last in all eight. That is family-aligned agreement in both directions, not each judge simply preferring itself: the editor's first draft of this record said each judge ranked its own model first every time, which is false for both Claude judges — Opus in fact picked itself first only three times, Fable five — and was corrected after the outside read caught it.
The editor's reading, that this looks like judges reading their own family's dialect rather than a demonstrated quality gap, is stated here as the editor's own diagnosis, not as something the round established. The Astra judge's error list points the same way, with a caveat: it flagged 32 errors across the Claude outputs and none in Astra's own, where the two Claude judges flagged errors across every arm including their own family's. The 32 entries are allegations rather than a settled count; the outside read checked three of them and upheld all three, and separately found three Astra outputs scored a full 14 that stipulate what they claim to test, prove an obstruction against a bridge the paper never proposed, or derive matching records from identical rules. So the asymmetry is partly the texts and partly the judge, and the rule's conclusion is recorded here without being credited as a quality finding.
A judge from a fourth provider family changes this only partway. TypeSafe's Jev, run blind over the same eight triples on 21 September, never chose an Astra output and split Fable and Opus four picks each; pairwise, its rankings put Opus over Astra 8 to 0 and Fable over Astra 7 to 1, the same direction as the two Claude judges. That weakens the partisanship reading — the Claude judges' ranking is no longer explained by family alone — without confirming a quality difference: Jev's five yes/no sub-scores barely separated the outputs, each scoring between 3.7 and 4.4 of 5, so its signal sits in its single top choice, and a top choice can track length and confidence of tone as easily as substance. The Claude outputs were the longest arm in round two, which is the stated caveat.
The editor's own synthesis — the shortlist below — needed the same outside read the judging did: GPT-6 Astra's check found it Claude-heavy and overstated in places, corrected in the next two sections.
The directions that survived #
Seven items, in the editor's order, amended after the outside read. Every one is a proposal to the author, not a decision, and each names the judges who chose it.
- P1's falsifier is aimed the wrong way. Five outputs across three lineages found that causal set theory (Bombelli, Lee, Meyer and Sorkin 1987, "Space-time as a causal set") already satisfies P1's published falsifier on its face; the Opus physics referee argued the live danger runs the other way, toward accounts that give perspective a calculable role with no self-reference. Kept, with a caveat added at the outside read: packet silence is not itself a negative result, and the old falsifier's outcome stays on record beside any new wording. Chosen by both Claude judges.
- S1's strong reading may rest on P2. Fable's and Opus's philosophy referees, working independently, found the diagnosis's defence of separability running through P2 rather than established on its own; Fable's Part I auditor found the same three-condition rule applying to every further-fact problem, induction included. Rebuilt at the outside read: the claims graph's confidence policy, omitted from the round-two packet, already prices S1's 0.8 on the weak reading only and holds the strong reading pending the missing argument, so there is no demotion to argue, only an argument still to supply. Chosen by the Fable and Opus judges.
- The correction machinery has never lowered a number. Opus's pre-mortem found eight disconfirming events on record and zero tier or numeric changes. Its proposed compulsory decrement rule and USD 500 blind grading round were dropped at the outside read — a deadline or an obligation to move a number is not evidence against a claim — and replaced with a dated author decision for each relevant result; this is where amendment A2, the dated-decisions rule, comes from. Astra's own history check strengthened the underlying finding: no tier lowered since 29 May 2026 and no number changed since 2 July 2026, across fifteen committed versions of the claims graph. Chosen by both Claude judges.
- One panel, many hinges. Fable's synthesis designer widened the persistence conjecture's panel to all four criteria and pre-read ten outcomes against named hinges, with a reliability precondition: two lineages fill the cells blind, and if they disagree on more than one cell in five the criteria are not yet an instrument. Amended to hold boundary, bearer and orientation definitions fixed across every answer, and to read a missing partition as undetermined rather than as evidence of independence. Route: a desk-tier pass under USD 20 first.
- A number for the physics arm. Opus's phenomenologist computed an ordering fraction read as one-dimensional across every architecture named; the outside read found that claim contradicted inside the same output's own numbers, and replaced it with a bounded derivation audit combining the Opus and Astra physics pieces rather than a physics result. Fable's clock-bound derivation, reported eight orders of magnitude too coarse to see, stands as a negative result with the Astra judge's objections unanswered.
- The count is a typing question. Fable's category theorist typed "exactly two" as a count of variances; Opus's category theorist relocated the open parameter to a class of admissible self-maps. Both readings have gaps the outside check named — the variance count has not derived two orientations, and an admissible-map class does not automatically become the anticipation rule — so the route is a translation test, specifying the intended richer representation first, rather than paying for formalisation on either count.
- The lab needs a judge outside the maker lineages. Not a proposed direction but the round's clearest finding about itself: three judges, three lineages, perfect family alignment. The outside read dropped the claim that a fourth vote alone would settle bias; the route is to reconcile the specific mathematical and interpretive disagreements under one rubric at desk cost first, before any further blinded comparison.
What the outside check corrected #
GPT-6 Astra's outside read, 20 September 2026, with lineages known, corrected three things in the editor's records and added one point on confidence: the judges' alignment is by provider family, not by model, since the editor's original sentence that each judge picked its own model every time was false (Opus picked itself first only three times, Fable five); the "dialect, not quality" reading was restated as the editor's own diagnosis rather than a demonstrated finding; and the Astra judge's 32 flagged errors were restated as allegations, of which the check verified three and separately found three Astra outputs over-scored at 14. It also flagged that the round-two packet omitted the claims graph's confidence policy, bearing on shortlist item 2, and, with GPT-5.6 models it disclosed using to check arithmetic and history, found no claims-graph tier lowered since 29 May 2026 and no number changed since 2 July 2026 across fifteen committed versions — a stronger form of the grade-history finding than the editor's eight-row audit.
The design (00_DESIGN.md), the comparison (20_comparison.md), the directions (10_directions.md), the check prompt (30_astra_check_prompt.md) and the check itself (31_astra_check.md) are under /downloads/lab/perspectives/; round two's twenty-four outputs are under /downloads/lab/perspectives/round-two/; the TypeSafe results are at typesafe/10_results.md; and the ten dated decisions awaiting the author are at 13_dated_decisions.md.
Review #
Read from outside the editor's lineage by GPT-6 Astra on 20 September 2026, with lineages known, on three questions: is the editor's shortlist skewed; was the Astra judge partisan, and was discounting all three judges right; do the three content findings hold. Its verdict: the shortlist Claude-heavy and overstated in places, with sixteen of twenty-four candidates Claude-made; withholding a confident quality ranking justified, declaring all three judges partisan not established; the three content findings hold, each with a stated correction. Three corrections were applied to the records and marked there: the judges' alignment reads as family, not model; the "dialect, not quality" reading is the editor's own diagnosis; the Astra judge's 32 flagged errors are allegations, three of them checked and upheld. The S1 item was rebuilt on the site's own confidence policy, and the compulsory-decrement idea was dropped. Full text: 31_astra_check.md.
Then TypeSafe's Jev, blind, from a third family, on 21 September 2026: never chose an Astra output; Opus over Astra 8 to 0. typesafe/10_results.md.
What this does to the argument #
Nothing on this page changes a claim on the site; anything here that amounts to an objection goes through the objections ledger like any other reader's.
What would count against this #
- The partisanship reading being wrong: if a fourth-lineage or human judge, blind, also ranked the Claude arms above Astra 8 to 0 (or the reverse), the judges were not partisan, they were right, and the rule's conclusion should be credited. That is the test the record now needs and could not run.
- The rubric rewarding dialect: the seven part-scores are reported so a reader can see whether one lineage lost on "first step done" (substance) or on "new" (a judgement about the record).
- The Fable-over-Opus lean being an artefact of length: Fable's outputs were the longest; the judges scored length breaks as form, not as quality, and said so, but a reader may weigh that differently.
- The seven directions being the editor's taste: they are marked as proposals and every one names which judges chose it.
Makers by Claude Opus 5, GPT-6 Astra and Claude Fable 5.1 agents at the author's request; edited by Claude Fable 5.1; judged blind by Claude Opus 5, Claude Fable 5.1, GPT-6 Astra and TypeSafe's Jev; read once from outside the editor's lineage by GPT-6 Astra, with corrections applied.