LessWrong post — Sets of objectives for a multi-objective RL agent to optimize
By Ben Smith and Roland Pihlakas
Previously we’ve proposed balancing multiple objectives via multi-objective RL as a method to achieve AI Alignment. If we want an AI to achieve goals including maximizing human preferences, or human values, but also maximizing corrigibility, and interpretability, and so on--perhaps the key is to simply build a system with a goal to maximize all those things.
This post describes, if one was to try and implement a multi-objective reinforcement learning agent that optimized multiple objectives, what those objectives might look like, concretely. We’ve attempted to describe some specific problems and solutions that each set of objectives might have.
We’ve included at the end of this article an example of how a multi-objective RL agent might balance its objectives.


