Using a trained agent in RLlib

Your approach seems correct in principle … not sure why the trainer cannot restore from checkpoint. You could compare with the example provided.