Hackathon project: BioBlue — Biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format
By Roland Pihlakas, Shruti Datta Gupta, and Sruthi Kuriakose
We aim to evaluate LLM alignment by testing agents in scenarios inspired by biological and economical principles such as homeostasis, resource conservation, long-term sustainability, and diminishing returns or complementary goods.
So far we have measured the performance of LLMs in three benchmarks (sustainability, single-objective homeostasis, and multi-objective homeostasis), in each for 10 trials, each trial consisting of 100 steps where the message history was preserved and fit into the context window.
Our results indicate that the tested language models failed in most scenarios. The only successful scenario was single-objective homeostasis, which had rare hiccups.


