I found another weird behavior of the env.reset() function: According to the openai gym specifications, env.reset() should return an initial observation of a new episode. However, if we call env.reset() after an episode has ended, it returns the last observation of that previous episode instead of an initial observation of the next episode. Furthermore, the data format of the observations returned by env.reset() is different compared to the observations returned by env.step(action). Is there anyone else with the same problems or is there a misunderstanding on my side? Thx 