AISC V Project proposal — Model structure and useful invariants for combining pluralistic positive and negative consequentialism in parametric ML (while avoiding trivial pathologies / degenerate state
By Roland Pihlakas
How to represent the goal systems with multiple values in order to reduce the Goodhart-like behaviour and specification gaming problems. Among other subtopics this includes combining multiple positive utility maximisation goals with multiple negative utility minimisation goals - in such a way that all these goals of an AI still get the desired relatively coherent equal treatment. The negative utility minimisation part is useful for task-based/low impact aspects, but also for whitelisting, explainability, and human accountability aspects.
In other words I want to enable “common sense” and to avoid using single-dimensional measures of success. In some cases the measures may be pluralistic on the surface, but in practice one of them would start dominating over all the others, therefore still turning the system into a single-dimensional one.


