Subscribe
Sign in
Home
Archive
About
Latest
Top
Working paper — Open-source LLMs administer maximum electric shocks in a Milgram-like obedience experiment
By Roland Pihlakas and Jan Llenzl Dagohoy
May 20
•
Three Laws - AI alignment
December 2025
LessWrong post — Research agenda for training aligned AIs using concave utility functions following the principles of homeostasis and…
By Roland Pihlakas
Dec 28, 2025
•
Three Laws - AI alignment
September 2025
Working paper — BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks…
By Roland Pihlakas and Sruthi Susan Kuriakose
Sep 2, 2025
•
Three Laws - AI alignment
July 2025
Presentation at Machine Ethics and Reasoning Workshop 2025 — Simulating value collapse in LLMs
By Lenz Dagohoy, Roland Pihlakas, Chad Burghardt, and Sophia March
Jul 30, 2025
•
Three Laws - AI alignment
June 2025
Black-box interpretability methodology blueprint: Probing runaway optimisation in LLMs
A methodology brainstorming document for identifying when, why, and how LLMs collapse from multi-objective and/or bounded reasoning into…
Jun 22, 2025
•
Three Laws - AI alignment
April 2025
Presentation at MAISU unconference 2025 — BioBlue: Notable runaway-optimiser-like LLM failure modes
By Roland Pihlakas, Sruthi Susan Kuriakose, and Shruti Datta Gupta
Apr 20, 2025
•
Three Laws - AI alignment
Presentation at MAISU unconference 2025 — Building Benchmarks for Universal Values [AISC 10]
By Lenz Dagohoy, Chad Burghardt, Sophia March, and Roland Pihlakas
Apr 20, 2025
•
Three Laws - AI alignment
March 2025
LessWrong post — Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks
By Roland Pihlakas, Sruthi Susan Kuriakose, and Shruti Datta Gupta
Mar 17, 2025
•
Three Laws - AI alignment
February 2025
Baseline experimental results with an LLM agent and OpenAI Stable Baselines 3 RL algorithms on our Extended Gridworlds
By Roland Pihlakas
Feb 24, 2025
•
Three Laws - AI alignment
Hackathon project: BioBlue — Biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format
By Roland Pihlakas, Shruti Datta Gupta, and Sruthi Kuriakose
Feb 1, 2025
•
Three Laws - AI alignment
January 2025
LessWrong post — Why modelling multi-objective homeostasis is essential for AI alignment
By Roland Pihlakas
Jan 1, 2025
•
Three Laws - AI alignment
November 2024
Presentation at Foresight Institute's Intelligent Cooperation Group — Introducing biologically and economically aligned multi-objective…
By Roland Pihlakas
Nov 1, 2024
•
Three Laws - AI alignment
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts