Not sure if this could help …
If you keep track of the recent flopscope changes, it seems that the top rankers might have found quite a few FLOP arbitrages that haven’t been priced in the public scorer yet. So this might simply be that some submissions uses way more samples but don’t get properly billed. I think the organizers are actively working on fairer accounting.
https://github.com/AIcrowd/flopscope/pulls?q=is%3Apr+is%3Aclosed