See https://arc-whest-public50-unusually-easy.netlify.app/
It’s kind of baffling because, AFAIK, the organizers have not informed us of any selection process that would make the public 50 lack this tail of high-variance MLPs. And yet, if it were really sampled from the same distribution as the huggingface datasets, there would only be a 0.2% chance of it being this easy. Which seems to strongly imply there was some kind of filtering that was used… perhaps rejection sampling?
So a question on my mind is, will the private 50 also lack this tail of difficult / high-variance MLPs, or will it contain the tail?