We reproduced your plateau arithmetic and it holds: for pure Monte Carlo, mse = V/n and C = c·n cancel, so the score sits at V·c/B regardless of n, and your 3.7e-7 follows from that at the competition shape.
Where we’d differ is the inference. That bounds Monte Carlo estimators; it doesn’t bound analytic methods (no n to cancel) or control-variate hybrids (cost c_a + c_mc·n, so the cancellation breaks). A score below the plateau is therefore not by itself evidence of misaccounting. Finding methods that beat naive sampling is what the challenge is for.
We will also be reviewing submissions that intentionally route computation around flopscope’s accounting to obtain unmetered compute, such submissions will be disqualified based on final review by ARC.