Hello @jiaxun_cui
Can you elaborate more on this?
distribution of result seems to be biased to be always worse when submitting
Do you mean the agent doesn’t learn as good as it does locally on your end? Or do you mean that you are getting lower score than expected during the rollouts?
You should be able to replicate the evaluation setup using the config from FAQ: Round 1 evaluations configuration