Current top 9: Notes on Phase 1 board-score variance, the truth-energy ladder, and four corrections to published figures

Posting three measurements that are useful for reading the final standings, and four corrections to figures already in the public record. Worth noting none of this is my method description.

Everything below regenerates from the public submission pages and from the documents cited.


1. Under a frozen estimator, the graded raw MSE does not move. The multiplier does.

The Official Rules guarantee a fixed per-MLP seed identical across submissions (section 5.5, “Reproducibility under fixed seeds”). The consequence is measurable on the board.

I grouped 97 cached grader payloads by participant and by raw MSE agreeing to better than 0.1 percent. Seven groups of three or four members exist.

participant n raw MSE spread multiplier spread score spread
SKIBIDI_TOILET 3 0.006% 1.88% 1.38%
abhinav_gorrepati 4 0.000% 13.84% 13.78%
bgrubbs1984 3 0.017% 3.36% 3.39%
dpskv5 3 0.000% 43.96% 35.42%
ednacob 3 0.000% 0.00% 0.00%
ednacob 3 0.003% 58.55% 58.64%
ednacob 4 0.000% 1.48% 1.37%

Median raw-MSE spread is 0.000 percent, maximum 0.017 percent. Median multiplier spread is 3.36 percent, maximum 58.55 percent.

Three qualifications, because the grouping rule is acc weaker:

The rule selects on prediction agreement, not artifact identity. In 3 out of 7 groups the raw MSE moves in the fifth digit and the metered FLOP count moves with it, so those are separate builds rather than regrades. Only abhinav_gorrepati (n = 4) and ednacob (n = 4) are frozen in both raw MSE and metered FLOPs.

On those two groups the moving quantity is residual wall time. abhinav_gorrepati’s mean residual wall time runs 1.212 to 1.391 s across 4 submissions whose raw MSE is identical to sixteen digits, and the multiplier moves 13.84 percent on that alone.

The 58.55 percent maximum is neither jitter nor a regrade. That ednacob group straddles a change in what goes through the meter: one build on WhestBench 0.13.0 / FlopScope 0.9.1 at 1.305e11 metered FLOPs per MLP, and two on 0.14.0 / 0.10.0 at 2.130e6 metered FLOPs with 0.0005 s of flopscope backend time and 2.70 and 2.77 s of residual wall time in their place. A later ednacob build (324395) on the same 0.14.0 stack meters 1.282e11 with 39.8 s of backend time, so this is not the evaluator repricing the same code.

Elsewhere the multiplier moves because the FLOP count moved, which is a code change and not jitter: bgrubbs1984’s 3.36 percent is a 3.4 percent drop in mean analytical FLOPs with billed residual wall time flat near 0.03 s.

Two practical consequences. A public rank change between regrades of an unchanged artifact carries no information about estimator quality, and changes across evaluator epochs carry less still. And when quoting an uncertainty on a leaderboard score it is worth naming the channel: amalgonim’s section 9 reports 8.1 percent over 20 shift draws, which is estimator randomization variance measured by re-randomizing their own lattice, and that is a different quantity from the pricing jitter above.


2. The truth-energy ladder, recovered from the public payloads

Each graded submission’s results.public_scores carries a 32-entry per_layer_mse vector per public MLP. Amalgonim disclosed in 18151 section 14.17 that many submissions emit zeros for the unscored layers, which makes their reported hidden-layer MSE equal to the truth energy E[a_l^2]. Priority for that observation is theirs.

Because that quantity is participant-independent, it can be recovered by exact cross-participant agreement. Across 97 submissions from 21 participants I recover 1550 of 1550 cells (50 MLPs by 31 unscored layers), with a median support of 16 participants per cell.

layer 0 1 5 10 20 29 30
mean E[a_l^2] 0.31852 0.50231 0.77820 0.88470 0.83809 0.68144 0.68754

Standard deviation over the 50 MLPs is 0.00164 at layer 0 and 0.40858 at layer 30. 72 of the 97 submissions zero-fill all 31 unscored layers, mine among them.

A units note, not a correction. 18151 section 14.17 gives 2.278e-01 as a zero-filled layer-30 value. My MLP index 0 reads 0.22782711684703827 at layer 30, matching to every printed digit. Theirs is one network; the 50-network mean is 0.68754. Same for their statement that the best per_layer_mse at any depth 1 to 30 is 1.71e-08 with final layer 1.70e-07: my census finds the same submission at the same layer (Cipo, 325660, layer 1) at 1.4129e-08 with final layer 1.9639e-07, because theirs is MLP index 0 and mine is the 50-MLP mean. Their numbers are correct for the quantities they computed. Worth knowing before anyone tries to reconcile the two.

On the label-noise floor. Layer 0 is closed-form for a bias-free He-initialized network on exactly N(0, I), so a nonzero reported MSE there is the baked target’s own Monte Carlo noise. Three independent implementations report a layer-0 mean MSE of 6.773e-10 to four significant figures: dpskv5 (324846, 6.772958e-10), ely2sh (325181, 6.772663e-10) and SKIBIDI_TOILET (326177 and 325687, both 6.772982e-10). Amalgonim’s check already rested on two of these agreeing to three digits. This narrows the input to the O5 discrepancy the technique census names, between 6.1e-11 and 2e-10 at the scored layer, but does not settle it: the step from a layer-0 figure to a scored-layer floor depends on how the targets were baked, which is not public.


3. The terminal jump, across the participants who compute hidden layers

Of the 25 submissions that do not zero-fill, 23 emit genuine hidden-layer predictions on all 31 unscored layers and two dpskv5 (side note: dpskv5 = DeepSeekv5? :eyes:) rows emit one layer only. Taking each participant’s best such submission by final-layer MSE, and averaging per-layer MSE over the 50 public MLPs:

participant submission L0 L10 L30 L31 (scored) L30 / L31
dpskv5 324846 6.773e-10 1.017e-06 4.035e-07 3.555e-09 113.5x
joe_wanza 325329 2.392e-06 3.339e-06 1.281e-06 3.954e-09 324.1x
huang_chung_yi 325657 8.642e-05 1.037e-04 4.657e-05 9.022e-09 5162.0x
rayan53 326900 3.647e-06 9.897e-06 3.345e-06 1.349e-08 248.0x
dstepanov 325331 3.279e-04 9.835e-05 3.128e-05 1.051e-07 297.7x
Cipo 325660 2.940e-09 6.076e-05 2.217e-05 1.964e-07 112.9x

The jump is present in all eight participants that compute the hidden layers. Across all eight the L30/L31 ratio spans 113x to 4.6e6x; the six above, the six best by scored-layer MSE, span 113x to 5162x. The other two are SKIBIDI_TOILET (324895), whose hidden-layer error sits at the truth-energy scale (L30 9.221e-01, ratio 4.6e6x), and ely2sh (325181), whose hidden layers are real and well below that scale (L30 9.409e-04, ratio 797.9x) but whose scored layer, 1.179e-06, is the weakest of the eight. No participant’s layer-30 mean is within two orders of its own layer-31 mean.

This table uses only submissions that do not zero-fill, so it is not exposed to the emission artifact described above. An earlier version of it on my side included zero-filling rows and would have measured the truth energy rather than anyone’s accuracy, which is the trap worth flagging to anyone else reading these vectors.

Cipo’s row is the most informative single profile: their error grows monotonically from 2.940e-09 at layer 0 to a peak of 6.709e-05 at layer 9, decays to 2.217e-05 by layer 30, and the scored layer then comes back to 1.964e-07.


4. Four corrections

4.1 The antithetic odd-share figure in 18151 section 14.22

Write f = f_even + f_odd. Then f_even(-x) = f_even(x) and f_odd(-x) = -f_odd(x), so

Cov(f(x), f(-x)) = Ve - Vo and

rho = (Ve - Vo) / (Ve + Vo), hence Vo / (Ve + Vo) = (1 - rho) / 2.

The three measured values of rho, -0.0483, -0.0312 and -0.0489, therefore give odd shares of 52.4, 51.6 and 52.4 percent. The document states instead: “Read as a share, the odd degrees carry about 4.8% of the scored layer’s variance, down from 100% at layer 1.” 4.8 percent is |rho| expressed as a percentage, and a correlation is not a variance share.

The corrected value is confirmed by the same document earlier in section 14. Section 14.1 decomposes the integrand on Gegenbauer polynomials over eight networks and reports “all odd, killed by pairing” at 55.8 percent, with an earlier three-network version at 59.1 percent. The corrected reading agrees with that to within 3.4 percentage points; the stated 4.8 percent disagrees with it by a factor of 11.6. The residual three points are plausibly aggregation, since their rho is an unweighted mean over 256 outputs on three networks while 14.1’s figure is variance-weighted over eight, but I have not verified that this accounts for the whole gap.

The operational conclusion is unaffected and correct: the equal-cost gain is 1 / (1 + rho), giving 1.03x to 1.05x. What changes is the reason. Slightly more than half of the scored layer’s variance is odd, not 4.8 percent of it. The gain is small because pairing annihilates that half while halving the number of distinct evaluation points, which at equal cost nearly cancel.

One related note: pairing does not make the layer-1 mean exact. relu(a) + relu(-a) = |a| is not constant, and for a Gaussian pre-activation the layer-1 odd share is 73.3 percent, not 100 percent.

4.2 The sensitivity argument in 18151 section 14.12 is sign-blind

The identity dF/dm = Phi(c) is correct, but it is justified with “at the median |c| = 3.49 the first is one”, using the absolute value. Section 14.9 of the same document discusses c = -4 neurons, and Phi(-3.49) = 2.4e-4, not 1. Neurons at c = -3.49 have a readout gain four thousand times smaller than the argument assumes, and roughly half the network sits at negative c. Averaged over all neurons the sign-correct coefficient is the mean of Phi(a)^2, which I measure at 0.48 (0.4659, 0.4880 and 0.4924 on public mini-split nets 2, 11 and 99), against the 1 the argument implies. The measured 14.12 table is empirical and stands regardless.

4.3 Two places where 18151 supersedes itself

Sections 13.6 and 10.12 reach conclusions about the leading entries that the document’s own summary and sections 14.6 and 14.15 withdraw. Anyone citing 10.12, 10.13 or 13.6 should read 14.6 and 14.15 first. Separately, section 8 of the attached appendix carries a boxed erratum reading “Both halves of the withdrawal above are wrong”, so appendix section 8 should not be cited without section 14. The main write-up carries a different section 8, which is the antithetic material discussed in 4.1 above and is not covered by that erratum.

This is offered as citation hygiene rather than as a criticism. That appendix is the most useful single document produced in Phase 1, and it is long enough that the withdrawals are easy to miss.

4.4 Data movement is billed, and my own notes had this wrong

Read from the installed flopscope/data/default_weights.json, which carries 472 op weights:

take 4.0, sort 4.0, histogram 4.0, where 4.0, copy 1.0, concatenate 1.0, reshape 1.0, all per element. My own working notes had recorded all of these as zero, which is a pre-0.9.0 reading. Anyone whose cost model still assumes free gathers will be out by a factor of four per element on any codebook or table lookup.

The same file gives dtype rates of 1.0 for bool through float32, including float16, and 2.0 for float64. So float16 buys nothing over float32, and a silent float64 promotion doubles the bill. nkosi_ndwandwe already reported the flopscope.stats promotion path in 18127; this is the underlying rate table it acts through.