A bet on a floor, filed before the reveal

We’re around #80 on the public board as I write this — a few days past the last re-grade (the first metering strip, a four-factor billing correction, has been merged and applied), and the board is still shaking out, with the library under it under active refinement (further billing fixes and a parity harness in review right now; a young benchmark being maintained exactly as it should be). So nothing here is a complaint about any snapshot, and nothing here is about anyone else’s entry. It’s simpler: we believe something specific about this problem, and we’d rather put it on the record now, dated, than explain it after the results land.

The bet. This task looks to us like it has a floor. For an unbiased, budget-respecting estimator, the error variance can’t be pushed below a value set by the network’s final nonlinearity — we put it at a per-pair variance of about 0.012, which works out to roughly 3.7e-7 adjusted under the current four-factor meter (about 3.4–3.9e-7 depending on the unseen suite, since per-network variance is heavy-tailed). We spent the week trying to get under it from every direction we could think of: richer control-variate features, better samplers, restructured sampling, biased shortcuts. Everything failed in an instructive way — the genuine improvements we found only brought us closer to it. That number is where our best unbiased estimator lands on networks it has never seen, and where we’d expect careful methods to converge.

The structural core of the argument (the plateau identity, optimal-CV-equals-projection, and the exchangeability symmetry) is machine-checked in Lean 4 / Mathlib with zero sorry and zero custom axioms — including all three ReLU–Hermite Gaussian-integral constants (the last two now proved via a Gaussian integration-by-parts identity on Mathlib’s FTC), every step depends only on Mathlib’s standard foundations, confirmed by #print axioms. What Lean certifies is that structural scaffolding; the floor’s numerical size (the ~0.012) is measured, not proven. Full write-up coming; stating the number first so the claim precedes the proof. And genuinely: if someone clears the floor on unseen networks, that’s the most interesting possible outcome — we’d love to be wrong about this one.

The prediction. The rules are clear the live board doesn’t decide ranking — a fresh private re-run on unseen networks does: one designated entry per team, re-graded after the submission phases close, results in early October. Filed today:

  • public and private rankings will differ noticeably (by design);
  • careful unbiased methods will converge toward the floor on unseen networks — from both sides — since for the unbiased class a floor is a property of the problem (a biased method whose error generalizes, especially at a different network shape, could sit below it);
  • our own standing should move up substantially from around #80, into the cluster at the floor — with one honest caveat. That ~#80 is already past the first billing strip, so the remaining rise rides on two things: suite-overfit entries regressing on networks they’ve never seen, and the newer accounting gaps still under discussion getting closed before the final meter is pinned. The second isn’t guaranteed — at least one gap looks open as I write this, and if it’s still open at the re-run, some scores rest on compute the meter doesn’t fully count, a fresh set of networks won’t demote those on its own, and our move up is muted to that extent. That’s a genuine dependency, not a hedge I can resolve from here.

Not predicting we finish first: if the floor is real, everyone who reached it is a tight pack separated by draw luck, and some current leaders may hold their ground just fine. We’re calling the direction of our own move, and the floor itself — concretely, if we don’t clear roughly the top third on the unseen re-run, treat this prediction as failed. On a fixed visible set, part of any low score — ours included — is favorable draw rather than durable signal; that’s just what feedback boards do, and we don’t exempt ourselves. Our own current score is one such draw, and it reverts toward the floor on unseen nets — so the honest-frontier entries just above us converge with us rather than falling below.

We made a wager this week: spend the effort on the unseen evaluation rather than the visible one, because whether a method holds up on networks it has never seen is the number we actually care about. If any of this is wrong, it’s here in writing and you can hold me to it. Good luck all — it’s been a genuinely strong field, and this contest is a lovely miniature of a bigger question: how much can you know about a system from its structure, before you run it?

I have reached a similar conclusion. My floor is similar to yours, and through tricks I can improve it at most by 10%, but I really can´t see how the top 10 got there, it seems almost physically impossible. Either they are overfitting or doing something I can´t really imagine.