Theory-of-Mind Self-Imagination for Altruistic Reinforcement-Learning Agents
A reinforcement-learning framework in which an agent's self-imagination module estimates state values from random reward feedback and a Theory-of-Mind perspective-taking mechanism re-evaluates those imagined states from other agents' viewpoints, together forming an intrinsic altruistic motivation. Added to a sparse task reward, this motivation drives the agent to prioritize rescuing others while avoiding irreversible negative side effects.
2501.00320
A reinforcement-learning framework in which an agent's self-imagination module estimates state values from random reward feedback and a Theory-of-Mind perspective-taking mechanism re-evaluates those …