AI Safety Camp project proposals — Universal Values, Risk Aversion vs Prospect Theory, and Proactive AI Safety
Roland Pihlakas will be running one of three possible projects, based on which one receives the most interest.
(32a) Creating new AI safety benchmark environments on themes of universal human values
Category: Evaluate risks from AI
We will be planning and optionally building new multi-objective multi-agent AI safety benchmark environments on themes of universal human values. Based on various anthropological research, a list of universal (cross-cultural) human values has been compiled. Various of these universal values resonate with concepts from AI safety, but use different keywords. It might be useful to map these universal values to more concrete definitions using concepts from AI safety.
One notable detail: in the case of AI and human cooperation, the values are not symmetric as they would be in human-human cooperation. This arises because we can change the goal composition of agents, but not of humans. Additionally, agents can be relatively easily cloned, while humans cannot.
(32b) Balancing and Risk Aversion versus Strategic Selectiveness and Prospect Theory
Category: Agent Foundations
We will be analysing situations and building an umbrella framework about when either of these incompatible frameworks would be more appropriate in describing how we want safe agents to handle choices relating to risks and losses in a particular situation.
Economic theories often focus on the “gains” side of utility. A well-known formulation is to use diminishing returns — a concave utility function. But what happens in the negative domain of utility? There is a well-known theory named “Prospect Theory”, which claims that our preferences in the negative domain are convex. This contradiction may be underexplored, especially with regards to AI safety.
(32c) Act locally, observe far — proactively seek out side-effects
Category: Train Aligned/Helper AIs
We will be building agents that are able to solve an already implemented multi-objective multi-agent AI safety benchmark that illustrates the need for the agents to proactively seek out side-effects outside of the range of their normal operation and interest, in order to be able to properly mitigate or avoid these side-effects.
In various real-life scenarios we need to proactively seek out information about whether we are causing undesired side effects (externalities). This information either would not reach us by itself, or would reach us too late. Attention is a limited resource — and the same constraints apply to AI agents.


