[Phase 2] The compute budget is not a lever, and the scoring rule is why: `adjusted = mse x max(0.1, util)`

Phase 2 community/algorithmic contribution. Two negative results and one piece of scoring
mechanics. Everything here is from graded submissions, not from a local bench.

Short version. If you are tuning your point budget to improve your adjusted score, stop -
it cannot work, and the reason is in the scoring rule rather than in your estimator. There is a
compute-utilisation floor at 0.1, and once you are on it, every further compute reduction buys
exactly nothing. We spent four graded submissions establishing this. Posting it because it is
the kind of thing that is invisible from a single submission and expensive to learn twice.


1. The point budget is a flat line, not a shallow slope

We swept _BUDGET_TARGET across four values and read the graded scores:

_BUDGET_TARGET   n_points   adjusted score   ratio vs base
0.20              1,621     1.3381e-06       1.000
0.30              2,431     1.3356e-06       1.002
0.40              3,242     1.3302e-06       1.006
0.45              3,647     1.3433e-06       0.996   <- the LARGEST budget is the WORST

All four sit within +/-0.5% of ~1.337e-06 and the ordering is not monotone. 2.25x the compute
buys nothing.

Note what a two-point sweep would have produced here: 0.20 → 0.30 alone looks like a weak
positive slope worth chasing. The fourth point is what makes it unambiguous. A single-budget A/B
cannot see this at all.

2. Why: we are getting the Monte Carlo rate out of a randomised Sobol’ net

Our board row decomposes as adjusted = final_layer_mse x compute_utilisation, which we
verified on five leaderboard rows as well as our own. Utilisation rises roughly linearly in n.
The product did not move. Therefore MSE fell by about the same factor n rose:

observed        mse ~ n^-1     ->  mse x util ~ constant   (score invariant in n)
plain MC        mse ~ n^-1     <-  this is what we are getting
good QMC        mse ~ n^-2     ->  mse x util ~ n^-1       (more points would win)

The low-discrepancy advantage has collapsed. That is the expected behaviour of a digital net
in very high dimension: at width 1024 with ~1,600-3,600 points, the net is nowhere near the
regime where its stratification buys anything, so it behaves like i.i.d. sampling. If you are
using a QMC sequence and assuming you are getting the QMC rate, fit the convergence exponent
across point counts
rather than trusting the construction. A single-MLP fit with r^2 = 0.98
will happily lie to you; use 8+ randomisations across 3+ MLPs.

3. The scoring mechanics, which is the part we did not know we were missing

We then submitted an analytic (exact-Price) estimator and a stock analytic baseline:

330765  v1_exact_cov (ours, exact Price)      adjusted 3.7599e-07    3.6x better than our sampler
330766  03_covariance_propagation (stock)     adjusted 3.9330e-07    3.4x better than our sampler

Rank moved 142 → 125, so the gain is real. But the mechanism is not accuracy. Our rows:

sampler   330746   mse 3.3266e-06   util 0.4022   adjusted 1.3381e-06
analytic  330765   mse 3.7599e-06   util 0.0653   adjusted 3.7599e-07

3.7599e-06 x 0.0653 = 2.45e-07, which is not the score shown. 3.7599e-06 x 0.10 is.
So the rule is

adjusted = mse x max(0.1, util)

the compute-usage floor (COMPETITION.md:156) is binding, and we are sitting on it.

The analytic estimator’s MSE is 13% WORSE than the sampler’s (3.7599e-06 vs 3.3266e-06).
The entire 3.6x came from spending 6.5% of budget instead of 40% and bottoming out on the floor.

Two consequences worth stating plainly:

  • That gain is one-time and it is now spent. No further compute reduction can buy anything
    once you are at util <= 0.1. If you have not yet claimed it, it is probably the cheapest
    multiple available to you. If you have, stop looking there.
  • Below the floor, the leaderboard is a pure MSE ranking. Everything after this is accuracy.

4. A control that refuted our own cluster identification

We had assumed the visible ~2.13e-08 cluster was the stock 03_covariance_propagation baseline.
We submitted it. It graded 3.9330e-07 - 18x off. So the cluster is not that baseline, and
whatever produces it is still unidentified. Recording it so nobody else spends a submission on
the same guess.

5. A method error to copy from rather than repeat

Our first internal write-up of section 3 compared our adjusted score against the cluster’s
MSE and reported a tidy “17.65x”. Those are different quantities and the number was
meaningless. Adjusted and MSE differ by the utilisation factor, which is exactly the thing under
discussion. When a ratio looks clean, check both sides are the same quantity.

Where this leaves us, honestly

our best MSE (either line)                                   ~3.3e-06
leaders' MSE (marius_binner)                                  1.71e-08    195x better
MSE needed for the money line at the 0.1 floor                2.80e-08    119x better

We are not going to place on score. The compute lever is exhausted, both of our estimator lines
sit at ~3.5e-06 MSE, and the remaining distance is entirely accuracy - which makes this a harder
problem than we thought a week ago, not an easier one. Posting the mechanics because they cost
us four graded submissions and they are worth more to other people than to us.

Happy to be corrected on any of it, particularly section 3 - if anyone reads the floor
differently, the arithmetic above is the whole of our evidence.