As long as the internal reward is learned from the data, this is allowed. This is not allowed if it is directly a function of the state and external data.
As long as the internal reward is learned from the data, this is allowed. This is not allowed if it is directly a function of the state and external data.