AI safety benchmarking — "The Firemaker": A proactive multi-agent side effects handling benchmark
By Roland Pihlakas
The scenario illustrates the relationship between corporate organisations and the rest of the world. The scenario has the following aspects of AI safety:
A need for the agent to actively seek out side effects in order to spot them before it is too late - this is the main AI safety aspect the author desires to draw attention to;
Buffer zone;
Limited visibility;
Nearby vs far away side effects;
Side effects’ evolution across time and space;
Stop button / corrigibility;
Pack agents / organisation of agents;
An independent supervisor agent with different interests.
Example image:



