**The consolidated update pass, as promised — with Phase 1 now closed.** The challenge page marks Phase 1 concluded following a final public rescore, though the held-out results have not yet been announced. The rescore’s visible effect on the public board was surgical: the two anomalous entries at the extreme front — the 9.5e-11 and 1.5e-9 adjusted scores whose raw MSEs sat below anything any published technique can explain — were **removed**, one author’s fallback entry landing mid-board at an ordinary 1.34e-7; one other entry moved ~3%; and every remaining score is unchanged (147 of 147 overlapping rows within 2%, several spot-checked exactly). So the board this update reads is the final public state of Phase 1, and the census’s biggest open question — what the silent front was — has an administrative answer rather than a technical one: *it was removed, not explained*. More on what that leaves standing at the front in part 3.
**Fourteen write-up topics arrived after the corpus cutoff.** The census (posted 2026-08-13) covered the 53 topics then in the category. Fourteen more landed 2026-08-14 → 18: mliston (18170, #327723), the baltsat visual notebook (18163) and bgrubbs1984’s companion site (18169), then the wave of 2026-08-16/17 — pranay212 (18173, #327651), kaileh57 (18171, #327786), nitiz_khanal (18174, #326960), Team Puffi’s two-paper drop (Cipo, 18175: #327950/#328046 plus the SSC estimator #327725), ely2sh (18176, #327749), keenanpepper (18177, #327838), qi_zhang5 (18179, #326954), violeta (18180), omer_kiraz (18181, #327792), trim_qewas (18182, #326725) — and finally andrei_bulzan’s 82-page report (18183, #327801). They move specific census items. Deltas below, keyed to the original item labels.
**New source: pranay212, topic 18173 (submission 327651; public adjusted 1.23e-7, final-layer MSE 1.55e-7).** A hybrid of Techniques 3 and 4: a **BCH strength-4 orthogonal-array** angular design (coordinates indexed by GF(256), columns [1, a, a³], antithetic quotient leaving 65,536 line representatives, of which a fixed target-independent prefix of **56,000 pairs** is used) mounted on the SOX dead/on/kink execution chassis (explicitly credited to 18106), with exact FWHT layer-0 evaluation, an exact antithetic layer-1 fold (h₀(−u)W₁ = h₀(u)W₁ − u(W₀W₁)), Winograd/Strassen, and layer-0 mean transport into a layer-1 bias (covariance transport helped an earlier sampling chassis but *hurt* the fixed-radius design). Item-level effects:
- **3a/4 (amended).** The best published **raw** MSE moves: 1.546e-7 (pranay212) vs the previous best 2.18e-7 (SOX, 18106). And **3c is now moot as stated**: the cubature family’s edge over structure-aware sampling was “compute, not raw accuracy” — the hybrid takes both, because it *is* both. This is also the first published estimator assembled substantially from other teams’ published parts (SOX chassis + the census-documented cubature line): the Phase 1 literature is compounding, which is what a literature is for.
- **3b (amended).** “Under an adjusted score you never subsample the design” (nygaard’s k^−1.2 ablation) gains a practical exception: pranay212 subsampled 65,536 → 56,000 pairs **for watchdog safety, not score** — the complete table was locally more accurate but left no budget margin on hard MLPs, and a 62k FWHT variant hit watchdog failures. His stated rule: “one failed MLP is much worse than a small average gain.” So: never subsample for score; do subsample for the timeout tail.
- **6f (reinforced).** Another compiler-class result: a final exact-algebra pass (one-write assembly, Winograd formulas, sparse packing) cut billed FLOPs ~6.7% mean across four MLPs (192.566B → 178.923B on the representative one) with raw MSE unchanged to ~0.1%, worth **+7.5% adjusted** with no raw-MSE change across the checkpoint pair.
- **1e-adjacent.** His discovery path independently replays the closure story: started at cumulant propagation, found partial k=3 corrections can be worse than the Gaussian closure at depth — consistent with the census’s C1 wall.
**New source: kaileh57 (Kellen Heraty), topic 18171 (submission 327786; public adjusted 9.686e-8, raw 2.088e-7).** A complete randomly-rotated 66,048-point Kerdock rule — the **third independent build of the C2 construction** — plus network-adaptive pruning (activation-column dropping at a 0.02 relative threshold; two-stage dead-output testing with 516 nomination + 516 disjoint confirmation rows and a 32-row sentinel), exact FWHT/Strassen/sparse-prefix algebra, and a final-layer-only output policy. The paper’s center of mass is theory, and it upgrades two census verdicts from “measured” to “proved”:
- **3 (verdict upgraded).** The census said the angular problem was “effectively solved”; there is now a **certificate**. Theorem: for the depth-32 NNGP (iterated arc-cosine) kernel, *every* fixed nonnegative spherical rule with ≤66,048 support points has normalized discrepancy within **0.0233188%** of the complete Kerdock rule’s — i.e. fixed positive sampling at this support budget is nearly saturated. The engine is a Delsarte/Cohn–Kumar-style LP bound built on a genuinely new analytic fact (their Theorem 1): the **sixth derivative of every finite iterate of the dual-ReLU map is positive** on (−1,1), so a degree-5 Hermite minorant at three double nodes is global at every depth. The algebra is audited over the rationals and the depth-32 instance is certified with 256-bit interval arithmetic (Arb), with proof-check commands and SHA-256 manifests in a public release. This stands beside the closure floor (1a) as the field’s second proved wall: **raw-risk reduction inside the fixed-positive class is closed** — what remains is cheaper evaluation, signed/adaptive/network-dependent rules, or hybrids. It also supersedes the quarantined near-optimality claim noted in 3b with a reviewable artifact. Two calibration numbers worth keeping: the Kerdock rule is **32.9% below** iid antipodal sampling’s expected risk at equal support, and the public raw MSE (2.088e-7) runs ~14% *below* the idealized ensemble prediction (2.429e-7) — finite width, suite variation, numerical QR, and pruning all live in that gap.
- **5f (sharpened to a theorem).** The census’s control-variate design rule was “under antithetic pairing, controls must be even.” Their Theorem 3 (degree-five control inertness) is much stronger for the cubature family: on the complete Kerdock cloud, every fixed-radius, sphere-centered, input-polynomial control of total degree ≤5 is **pathwise identically zero** — not merely mean-zero — for every rotation and radius, and this survives their column pruning. Degree six is the first live harmonic band and carries only **13.93%** of the remaining discrepancy (consistent with amalgonim’s 3d oracle argument); their 128-feature degree-six ridge dictionary measured a pooled-oracle R² of 0.2351% against a 4.05% cost surcharge — dead. The explicit carve-outs matter as much as the theorem: Gaussian-centered controls (‖x‖² has sampled value μ_R² but Gaussian mean 256), row reweighting, and **controls of intermediate activations** are all outside it. Read together with 5d (kwon’s 2.39× mid-network control), the control frontier for cubature carriers is now provably *not* at the input — the mid-network channel is the only live one.
- **5h (confirmed again).** Their pruning-threshold campaign is textbook transfer-trap discipline: α = 0.035 saved 2.3% arithmetic but cost 13.3% MSE (product 10.6% worse); a frozen screen on 24 fresh networks showed thresholds 0.022–0.025 produced development-only gains that reversed held-out; 0.020 retained. Their negative-results table (a second Kerdock rotation: <0.26% modeled headroom; two added optimized lines: ratio 1.0000203, with the optimizer returning **no negative weights**; shared-prefix multifidelity: correlation arrives more slowly than work is spent) all dies honestly.
- **6-adjacent.** One more compiler-class item: caching contiguous FWHT workspaces with ping-pong reuse removed 319,029,248 metered operations per test with bitwise-equal outputs. And a provenance observation the metering crew should see: **AIcrowd publishes no server-side artifact hash**, so a retained local upload cannot be cryptographically tied to the graded copy — their SHA-256 manifest is the best a participant can currently do.
- **Scoping honesty worth naming:** the certificate was derived after submission, covers only the unpruned core, and “does not explain the leaderboard score.” The write-up also connects the challenge to the external literature (Wu, Lecomte, Winer, Robinson, Hilton & Christiano, arXiv:2605.05179) — the first in-forum write-up to do so.
**New source: nitiz_khanal, topic 18174 (submission 326960; official adjusted 1.7275e-7, 50/50 graded, zero failures).** A structure-aware Monte Carlo sampler — centered/whitened antithetic Gaussian cloud, the exact first-layer antithetic identity ReLU(−z) = ReLU(z) − z, a pilot-frozen dead-row mask (1,024-sample lead block; rows kept iff they fired at least once), SOX-style terminal stable-on/stable-off/kink classification at a 1.0-column-σ margin, and two-level Strassen/Winograd applied only above residual-aware size cutoffs. Individually these are census-known parts; the write-up’s value is in three places:
- **A new design framing: “private-safe.”** This is the first write-up organized around surviving the fresh-draw private re-evaluation rather than the public board: no public-instance identification or answer tables, the pruning mask frozen *before* production samples, sample count sized against a conservative **no-pruning** worst-case cost model, and a terminal classification that deliberately retains uncertain columns rather than forcing a call. It is 5h’s transfer-trap discipline promoted from a validation rule to the architecture itself.
- **6a/6c (reinforced with a clean instance).** A level-three Strassen recursion reduced nominal FLOPs but lost in paired tests on residual runtime — and his graded checkpoint pair makes the same point at the score level: a more conservative recursion cutoff (#326963) improved the local FLOP proxy yet graded *worse* than the selected build (1.7292e-7 vs 1.7275e-7). “A lower local cost proxy did not transfer to a better official score.”
- **4 (the chassis is now standard).** With explicit citation of “publicly documented structural pruning and terminal classification, independently integrated and validated,” this is the second post-cutoff write-up to build on the SOX execution pattern — dead/on/kink routing has become the field’s shared execution layer, not a team’s technique.
His rejected-ablation table is the familiar honest ledger (bigger lead block: no gain; tighter margin: weaker conservatism for negligible benefit; earlier split: worse held-out; Cholesky whitening: cheaper and plausible but unshipped), and control variates, higher-order closures, and alternative sampling tables were all tried and not promoted — consistent with the census’s Technique 1 and 5 verdicts.
**New source: Team Puffi (Cipo), topic 18175 — two papers.** The census cited Cipo’s in-forum notes (18152); these are the formal write-ups, and they are the most substantial of this batch.
*Paper 1 — the score stack (submissions #327950, adj 9.0914e-8 / raw 1.8945e-7, and a lower-compute alternative #328046).* The carrier is the complete MUB-129 bank (128 Kerdock bases + identity, 66,048 antipodal directions — the **fourth complete build of the C2 construction**), with a three-basis probe (1,536 rows) for pruning and omitted-mean fill, prefix-based sparse execution, and a positive-stable terminal bypass. Item-level effects:
- **3a (extended with the first published quadrature-rule ladder).** Measured final-layer MSE relative to Kerdock frames on a matched panel: i.i.d. Gaussian 2.36×, Owen–Sobol + antipodes 1.39×, rotated simplex 1.35×, signed-Hadamard frames 1.02×, Kerdock 1.00×, **complete real MUB-129 0.81×** — the identity basis is worth 19% and is nearly free at layer 0. No previous write-up ranked the alternatives head-to-head.
- **5c/5d (the mid-network control variate, now fully documented at the published frontier).** The census recorded Cipo’s “~2×” inter-layer fluctuation-map control from forum notes; the paper gives the deployed construction: a shallow analytical reference (“A716” — a fitted finite closure through post-ReLU layer 9 carrying mean, full covariance, marginal third/fourth cumulants, a paired (2,2) fourth, and an eight-generation mixed (2,1) history, with 90 frozen fitted coefficients; layer-9 MSE ratio **0.511 vs Wu K=3**) supplies the layer-9 discrepancy, a ridge-fitted per-network response transports it to the final layer, and a 0.75-shrunk correction is subtracted from the final sampled mean. Measured: **17.6% final-MSE reduction** (paired-bootstrap 95% CI [7.8%, 26.3%]), winning on 34/50 networks, deployed at the best published-write-up-family rank. Their own framing is exactly the census’s: “model-assisted rather than a textbook unbiased control variate,” with the error identity ‖e_CV‖² = ‖e₃₁‖² − 2ρ·(alignment) + ρ²·(quadratic) stated rather than assumed. This is 7a (fit at scale, freeze) and 5d (mid-network control) composed and graded.
- **6 (a new score identity).** With per-network error e and multiplier q, the published score is S = ē·q̄ + Cov(e,q) — the error–multiplier covariance is a real score channel (for #327950: S/ē = 0.4799 vs q̄ = 0.4830). Nobody had written that term down.
- **An instrument fact the metering crew should have:** the grading platform executed **100 MLPs; fifty formed the public scoring panel; the other fifty were gated evaluation cases not included in the public score.** That is a direct statement that a hidden 50-network panel is already being run per submission — relevant to 6h’s leaderboard-population analysis and to what the private re-evaluation will draw from.
- **New open problem (frame selection).** Retrospective oracles show small Kerdock subsets can be excellent (a greedy target-using 38-frame oracle hits 5.47e-8), but every target-free selection rule tried — including herding on observable features — failed to transfer: herding reduced the discrepancy it optimized without consistently reducing final MSE, because “later nonlinear layers can rotate or regenerate error that was reduced at the selection layer.” A mechanistic, fixed-before-evaluation frame-selection rule is named as their main open problem (forum note: they hope it reaches 0.1–0.2 multiplier).
*Paper 2 — the mechanistic estimator (submission #327725, “SSC”: separate-universe response closure; adj 1.6533e-6 / raw 2.6597e-6).* The first structural advance in Technique 1 since the census closed the family:
- **1 (the projection step, not the order, is the live variable).** SSC keeps the Wu-et-al. K=3 cumulant backbone but replaces the fourth-cumulant projection: κ₄ is regenerated each layer from ≤4 live covariance-response modes coupled by a small fitted (indefinite) Λ — the super-sample-covariance algebra from cosmology (Takada–Hu), with the n⁴ tensor never formed. On a matched 30-network panel it cuts layer-32 MSE **77.2% vs factorized K=3 and 48.2% vs factorized K=4**, at 0.501B ledger versus K=4’s ~10²B, and flattens the depth-error exponent to **2.21 (vs 2.69 for K=3, 4.13 for K=4)**. A matched ablation isolates the response branch: disabling it raises layer-32 MSE 2.728e-6 → 7.698e-6.
- **The wall still stands — and they say so.** SSC’s layer-32 MSE remains 7.94× worse than Kerdock QMC on the same panel, and the graded submission (1.65e-6 adjusted) does not approach the closure family’s ~8.6–9.3e-7 readout floor (1a), which SSC’s compressed fourth order does not carry enough joint state to break (compare O4’s “carried joint object of dimension ≥64” — SSC carries four modes). But 1e said higher orders don’t rescue the family because cost explodes and the series is asymptotic; SSC is the first measured answer of a different kind: **compress the order you keep instead of adding one**, beating full factorized K=4 at roughly 1/200th its cost. The census’s “closed as a competitor” verdict survives; its “mechanism understood, nothing left to try” reading does not.
- **Useful taxonomy.** The paper names the whole Technique-1 family “deduction–projection estimators” (exact deduction advances the state, a projection compresses it) — a cleaner frame than “closures” for where the error actually enters, and it locates 1b–1e as facts about particular projections.
- **7a discipline, again:** coefficients fitted on 160 public networks, checked on 16 held-out, evaluated on 30 fresh independently generated networks; everything frozen before evaluation. A post-submission audit (refitting the response to target the consumed mean correction rather than κ₄ marginals: 13.6% further MSE reduction on the predecessor baseline) is reported as unintegrated future work rather than folded into claims.
*[continued in the next post — parts 2 and 3 cover the remaining nine sources and the item-keyed census deltas]*
census_update_2026-08-18.pdf (129.9 KB)