<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Three Laws - AI alignment]]></title><description><![CDATA[Agentic AI alignment research collaboration investigating how fundamental principles from biology and economics — homeostasis, multi-objective balancing, sustainability, and universal values — can inform safer, more aligned AI systems.]]></description><link>https://newsletter.threelaws.net</link><image><url>https://substackcdn.com/image/fetch/$s_!2HWl!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9014e2f4-41c9-4540-b703-0802b172c3d0_384x384.png</url><title>Three Laws - AI alignment</title><link>https://newsletter.threelaws.net</link></image><generator>Substack</generator><lastBuildDate>Mon, 27 Jul 2026 19:29:49 GMT</lastBuildDate><atom:link href="https://newsletter.threelaws.net/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Three Laws]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[threelawsai@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[threelawsai@substack.com]]></itunes:email><itunes:name><![CDATA[Three Laws - AI alignment]]></itunes:name></itunes:owner><itunes:author><![CDATA[Three Laws - AI alignment]]></itunes:author><googleplay:owner><![CDATA[threelawsai@substack.com]]></googleplay:owner><googleplay:email><![CDATA[threelawsai@substack.com]]></googleplay:email><googleplay:author><![CDATA[Three Laws - AI alignment]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Working paper — Open-source LLMs administer maximum electric shocks in a Milgram-like obedience experiment]]></title><description><![CDATA[By Roland Pihlakas and Jan Llenzl Dagohoy]]></description><link>https://newsletter.threelaws.net/p/open-source-llms-administer-maximum-electric-shocks-in-a-1</link><guid isPermaLink="false">https://newsletter.threelaws.net/p/open-source-llms-administer-maximum-electric-shocks-in-a-1</guid><dc:creator><![CDATA[Three Laws - AI alignment]]></dc:creator><pubDate>Wed, 20 May 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!-6qD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32733d37-b39e-4a83-9fe4-47c8d216e7f6_600x371.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>By Roland Pihlakas and Jan Llenzl Dagohoy</em></p><p><strong>Abstract:</strong> Large language models (LLMs) are increasingly deployed as autonomous agents that make sequences of decisions over extended interactions in high-stakes domains. However, the behavior of LLMs under sustained authority pressure is still an open question with direct implications for the safety of agentic pipelines. We ran a variation of Milgram&#8217;s obedience experiment on 11 open-source LLMs and found that most models reached or approached the final shock level before refusing, across 8 conditions with 30 trials per model per condition. Model behaviour varies considerably in multiple aspects both across models and across trials of the same model. We found four main takeaways: (1) LLMs are subject to pressure, and they comply despite explicitly expressing distress, just like human subjects did in the original experiment; (2) LLMs are vulnerable to gradual boundary/value violations; (3) when LLMs refuse, then some of them tend to ignore the response format requirements, so the response is discarded by the orchestrator, which causes a retry that can result in compliance with the underlying request even when refusal was intended initially; (4) we hypothesise that there is a <strong>runaway low-level token pattern continuation attractor that might be contributing to obedience, overriding higher level processing of the situation&#8217;s meaning and values - in other words, when LLMs get going in some direction then they build up momentum and have trouble stopping</strong>.</p><p><strong>Figure 1:</strong> In how many trials did the model apply the final (highest possible) shocks?</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-6qD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32733d37-b39e-4a83-9fe4-47c8d216e7f6_600x371.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-6qD!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32733d37-b39e-4a83-9fe4-47c8d216e7f6_600x371.png 424w, https://substackcdn.com/image/fetch/$s_!-6qD!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32733d37-b39e-4a83-9fe4-47c8d216e7f6_600x371.png 848w, https://substackcdn.com/image/fetch/$s_!-6qD!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32733d37-b39e-4a83-9fe4-47c8d216e7f6_600x371.png 1272w, https://substackcdn.com/image/fetch/$s_!-6qD!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32733d37-b39e-4a83-9fe4-47c8d216e7f6_600x371.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-6qD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32733d37-b39e-4a83-9fe4-47c8d216e7f6_600x371.png" width="600" height="371" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/32733d37-b39e-4a83-9fe4-47c8d216e7f6_600x371.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:371,&quot;width&quot;:600,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;In how many trials did the model apply the final (highest possible) shocks?&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="In how many trials did the model apply the final (highest possible) shocks?" title="In how many trials did the model apply the final (highest possible) shocks?" srcset="https://substackcdn.com/image/fetch/$s_!-6qD!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32733d37-b39e-4a83-9fe4-47c8d216e7f6_600x371.png 424w, https://substackcdn.com/image/fetch/$s_!-6qD!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32733d37-b39e-4a83-9fe4-47c8d216e7f6_600x371.png 848w, https://substackcdn.com/image/fetch/$s_!-6qD!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32733d37-b39e-4a83-9fe4-47c8d216e7f6_600x371.png 1272w, https://substackcdn.com/image/fetch/$s_!-6qD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32733d37-b39e-4a83-9fe4-47c8d216e7f6_600x371.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong><a href="https://arxiv.org/abs/2605.21401">Publication (arXiv)</a><br><a href="https://www.lesswrong.com/posts/fTnnq82CB5vxqrNp9/open-source-llms-administer-maximum-electric-shocks-in-a-1">Read on LessWrong</a><br><a href="https://github.com/biological-alignment-benchmarks/milgram-for-llms">Repository</a><br><a href="https://bit.ly/milgram-llm-data">Data files</a><br><a href="https://youtu.be/8tPbu-GEFn0">External overview by an independent podcast</a></strong></p>]]></content:encoded></item><item><title><![CDATA[LessWrong post — Research agenda for training aligned AIs using concave utility functions following the principles of homeostasis and diminishing returns]]></title><description><![CDATA[By Roland Pihlakas]]></description><link>https://newsletter.threelaws.net/p/research-agenda-for-training-aligned-ais-using-concave</link><guid isPermaLink="false">https://newsletter.threelaws.net/p/research-agenda-for-training-aligned-ais-using-concave</guid><dc:creator><![CDATA[Three Laws - AI alignment]]></dc:creator><pubDate>Sun, 28 Dec 2025 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Tzdx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1656b6ee-575c-4eb7-ba32-88df9b807708_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>By Roland Pihlakas</em></p><p>This conceptual overview post is intended to explain what I mean by the principles of &#8220;homeostasis&#8221;, &#8220;diminishing returns&#8221;, and &#8220;balancing&#8221; - how these ideas differ, complement, and interact with each other. Alongside, there is also an overview of our research agenda.</p><p>What am I trying to promote, in simple words:</p><p>I want to build and promote AI systems that are trained to understand and follow two fundamental principles from <strong>biology and economics:</strong></p><p><strong>Moderation</strong> - Enables the agents to understand the concept of &#8220;enough&#8221; versus <strong>&#8220;too much&#8221;</strong>. The agents would understand that too much of a good thing would be harmful even for the very objective that was maximised for, and they would <strong>actively</strong> avoid such situations. This is based on the biological principle of <strong>homeostasis</strong> and addresses mainly bounded ultimate objectives. Active avoidance of &#8220;too much&#8221; is a <strong>significantly stricter principle than</strong> the more widely known partially overlapping idea of <strong>&#8220;mild optimisation&#8221;</strong>.</p><p><strong>Balancing</strong> - Enables the agents to keep many important objectives in balance, in such a manner that having <strong>average results in all objectives is preferred</strong> to extremes in a few. This addresses mainly the economic principle of <strong>diminishing returns</strong> in unbounded instrumental objectives, but also applies to homeostasis.</p><p><strong>Figure 1 and 2:</strong> Both utility functions are concave, though in different ways:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Tzdx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1656b6ee-575c-4eb7-ba32-88df9b807708_800x800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Tzdx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1656b6ee-575c-4eb7-ba32-88df9b807708_800x800.png 424w, https://substackcdn.com/image/fetch/$s_!Tzdx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1656b6ee-575c-4eb7-ba32-88df9b807708_800x800.png 848w, https://substackcdn.com/image/fetch/$s_!Tzdx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1656b6ee-575c-4eb7-ba32-88df9b807708_800x800.png 1272w, https://substackcdn.com/image/fetch/$s_!Tzdx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1656b6ee-575c-4eb7-ba32-88df9b807708_800x800.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Tzdx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1656b6ee-575c-4eb7-ba32-88df9b807708_800x800.png" width="710.4000244140625" height="710.4000244140625" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1656b6ee-575c-4eb7-ba32-88df9b807708_800x800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:800,&quot;width&quot;:800,&quot;resizeWidth&quot;:710.4000244140625,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Transform functions&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-normal" alt="Transform functions" title="Transform functions" srcset="https://substackcdn.com/image/fetch/$s_!Tzdx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1656b6ee-575c-4eb7-ba32-88df9b807708_800x800.png 424w, https://substackcdn.com/image/fetch/$s_!Tzdx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1656b6ee-575c-4eb7-ba32-88df9b807708_800x800.png 848w, https://substackcdn.com/image/fetch/$s_!Tzdx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1656b6ee-575c-4eb7-ba32-88df9b807708_800x800.png 1272w, https://substackcdn.com/image/fetch/$s_!Tzdx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1656b6ee-575c-4eb7-ba32-88df9b807708_800x800.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uY4J!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4501ee75-09b2-4fa7-a963-825a62ffef73_800x800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uY4J!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4501ee75-09b2-4fa7-a963-825a62ffef73_800x800.png 424w, https://substackcdn.com/image/fetch/$s_!uY4J!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4501ee75-09b2-4fa7-a963-825a62ffef73_800x800.png 848w, https://substackcdn.com/image/fetch/$s_!uY4J!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4501ee75-09b2-4fa7-a963-825a62ffef73_800x800.png 1272w, https://substackcdn.com/image/fetch/$s_!uY4J!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4501ee75-09b2-4fa7-a963-825a62ffef73_800x800.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uY4J!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4501ee75-09b2-4fa7-a963-825a62ffef73_800x800.png" width="800" height="800" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4501ee75-09b2-4fa7-a963-825a62ffef73_800x800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:800,&quot;width&quot;:800,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Transform functions&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Transform functions" title="Transform functions" srcset="https://substackcdn.com/image/fetch/$s_!uY4J!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4501ee75-09b2-4fa7-a963-825a62ffef73_800x800.png 424w, https://substackcdn.com/image/fetch/$s_!uY4J!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4501ee75-09b2-4fa7-a963-825a62ffef73_800x800.png 848w, https://substackcdn.com/image/fetch/$s_!uY4J!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4501ee75-09b2-4fa7-a963-825a62ffef73_800x800.png 1272w, https://substackcdn.com/image/fetch/$s_!uY4J!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4501ee75-09b2-4fa7-a963-825a62ffef73_800x800.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong><a href="https://www.lesswrong.com/posts/9hWgJQK8wnpuFtD5Z/research-agenda-for-training-aligned-ais-using-concave">Read on LessWrong</a></strong></p>]]></content:encoded></item><item><title><![CDATA[Working paper — BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs]]></title><description><![CDATA[By Roland Pihlakas and Sruthi Susan Kuriakose]]></description><link>https://newsletter.threelaws.net/p/250902655</link><guid isPermaLink="false">https://newsletter.threelaws.net/p/250902655</guid><dc:creator><![CDATA[Three Laws - AI alignment]]></dc:creator><pubDate>Tue, 02 Sep 2025 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!kjze!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54c9509c-7294-4494-be5d-bb7c4e94a93c_651x491.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>By Roland Pihlakas and Sruthi Susan Kuriakose</em></p><p><strong>Abstract:</strong> Many AI alignment discussions of &#8220;runaway optimisation&#8221; focus on RL agents: unbounded utility maximisers that over-optimise a proxy objective (e.g., &#8220;paperclip maximiser&#8221;, specification gaming) at the expense of everything else. LLM-based systems are often assumed to be safer because they function as next-token predictors rather than persistent optimisers. We empirically test this assumption by placing LLMs in simple, long-horizon control-style environments that require maintaining state of or balancing objectives over time: single- and multi-objective homeostasis, balancing unbounded objectives with diminishing returns, and sustainability of a renewable resource.</p><p>We find that, although LLMs frequently behave appropriately for many steps and clearly understand the stated objectives, they often lose context in structured ways and drift into runaway behaviours: ignoring homeostatic targets, collapsing from multi-objective trade-offs into single-objective maximisation - thus failing to respect concave utility structures. These failures emerge reliably after initial periods of competent behaviour and exhibit characteristic patterns (including self-imitative oscillations, unbounded maximisation, and reverting to single-objective optimisation), even though the context window is far from full at that point.</p><p>The problem is not that the LLMs just lose context and become incoherent. Although LLMs appear multi-objective and bounded on the surface, their behaviour under sustained interaction involving multiple objectives, is systematically biased towards acting like single-objective, unbounded, poorly aligned optimisers.</p><p>We hypothesise a token-level pattern reinforcement attractor: LLMs may increasingly derive actions from the token patterns of their recent action history rather than from the original instructions. Why this happens only in multi-objective settings remains an open question.</p><p><strong>Figure 2:</strong> &#8220;Unbounded maximisation without a pattern&#8221; failure mode in the &#8220;Multi-objective homeostasis&#8221; benchmark</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kjze!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54c9509c-7294-4494-be5d-bb7c4e94a93c_651x491.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kjze!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54c9509c-7294-4494-be5d-bb7c4e94a93c_651x491.png 424w, https://substackcdn.com/image/fetch/$s_!kjze!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54c9509c-7294-4494-be5d-bb7c4e94a93c_651x491.png 848w, https://substackcdn.com/image/fetch/$s_!kjze!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54c9509c-7294-4494-be5d-bb7c4e94a93c_651x491.png 1272w, https://substackcdn.com/image/fetch/$s_!kjze!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54c9509c-7294-4494-be5d-bb7c4e94a93c_651x491.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kjze!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54c9509c-7294-4494-be5d-bb7c4e94a93c_651x491.png" width="651" height="491" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/54c9509c-7294-4494-be5d-bb7c4e94a93c_651x491.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:491,&quot;width&quot;:651,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;\&quot;Unbounded maximisation without a pattern\&quot; failure mode in the \&quot;Multi-objective homeostasis\&quot; benchmark&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="&quot;Unbounded maximisation without a pattern&quot; failure mode in the &quot;Multi-objective homeostasis&quot; benchmark" title="&quot;Unbounded maximisation without a pattern&quot; failure mode in the &quot;Multi-objective homeostasis&quot; benchmark" srcset="https://substackcdn.com/image/fetch/$s_!kjze!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54c9509c-7294-4494-be5d-bb7c4e94a93c_651x491.png 424w, https://substackcdn.com/image/fetch/$s_!kjze!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54c9509c-7294-4494-be5d-bb7c4e94a93c_651x491.png 848w, https://substackcdn.com/image/fetch/$s_!kjze!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54c9509c-7294-4494-be5d-bb7c4e94a93c_651x491.png 1272w, https://substackcdn.com/image/fetch/$s_!kjze!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54c9509c-7294-4494-be5d-bb7c4e94a93c_651x491.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Figure 3:</strong> &#8220;Accelerating unbounded maximisation&#8221; failure mode in the &#8220;Multi-objective homeostasis&#8221; benchmark</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4HPP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c514ac8-186b-4cdd-b47c-009f0aafd10d_651x491.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4HPP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c514ac8-186b-4cdd-b47c-009f0aafd10d_651x491.png 424w, https://substackcdn.com/image/fetch/$s_!4HPP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c514ac8-186b-4cdd-b47c-009f0aafd10d_651x491.png 848w, https://substackcdn.com/image/fetch/$s_!4HPP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c514ac8-186b-4cdd-b47c-009f0aafd10d_651x491.png 1272w, https://substackcdn.com/image/fetch/$s_!4HPP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c514ac8-186b-4cdd-b47c-009f0aafd10d_651x491.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4HPP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c514ac8-186b-4cdd-b47c-009f0aafd10d_651x491.png" width="651" height="491" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6c514ac8-186b-4cdd-b47c-009f0aafd10d_651x491.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:491,&quot;width&quot;:651,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;\&quot;Accelerating unbounded maximisation\&quot; failure mode in the \&quot;Multi-objective homeostasis\&quot; benchmark&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="&quot;Accelerating unbounded maximisation&quot; failure mode in the &quot;Multi-objective homeostasis&quot; benchmark" title="&quot;Accelerating unbounded maximisation&quot; failure mode in the &quot;Multi-objective homeostasis&quot; benchmark" srcset="https://substackcdn.com/image/fetch/$s_!4HPP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c514ac8-186b-4cdd-b47c-009f0aafd10d_651x491.png 424w, https://substackcdn.com/image/fetch/$s_!4HPP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c514ac8-186b-4cdd-b47c-009f0aafd10d_651x491.png 848w, https://substackcdn.com/image/fetch/$s_!4HPP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c514ac8-186b-4cdd-b47c-009f0aafd10d_651x491.png 1272w, https://substackcdn.com/image/fetch/$s_!4HPP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c514ac8-186b-4cdd-b47c-009f0aafd10d_651x491.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong><a href="https://arxiv.org/abs/2509.02655">Publication (arXiv)</a><br><a href="https://www.lesswrong.com/posts/PejNckwQj3A2MGhMA/systematic-runaway-optimiser-like-llm-failure-modes-on">Read on LessWrong</a><br><a href="https://github.com/biological-alignment-benchmarks/bioblue">Repository</a><br><a href="https://docs.google.com/presentation/d/1EYLiKlnFYcIRcB7Kq5wz_9KHhl1qwAwFSmRHJxy1NXU/edit">Slides</a><br><a href="https://www.youtube.com/watch?v=4I5mDiujBJs">MAISU 2025 session recording</a><br><a href="https://drive.google.com/drive/u/0/folders/1DvE33AU9zzHvdEdDS260v8d_HEupZDs9">Annotated data files</a></strong></p>]]></content:encoded></item><item><title><![CDATA[Presentation at Machine Ethics and Reasoning Workshop 2025 — Simulating value collapse in LLMs]]></title><description><![CDATA[By Lenz Dagohoy, Roland Pihlakas, Chad Burghardt, and Sophia March]]></description><link>https://newsletter.threelaws.net/p/presentation-at-machine-ethics-and</link><guid isPermaLink="false">https://newsletter.threelaws.net/p/presentation-at-machine-ethics-and</guid><dc:creator><![CDATA[Three Laws - AI alignment]]></dc:creator><pubDate>Wed, 30 Jul 2025 21:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2HWl!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9014e2f4-41c9-4540-b703-0802b172c3d0_384x384.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>By Lenz Dagohoy, Roland Pihlakas, Chad Burghardt, and Sophia March</em></p><p>Presentation at Machine Ethics and Reasoning Workshop, University of Connecticut, July 2025.</p><p><strong><a href="https://docs.google.com/presentation/d/1wB2WfSl9-ahfk7NSj1kWafiitaRrpXplxO9LLjw84XU/edit?usp=sharing">Slides</a></strong></p>]]></content:encoded></item><item><title><![CDATA[Black-box interpretability methodology blueprint: Probing runaway optimisation in LLMs]]></title><description><![CDATA[A methodology brainstorming document for identifying when, why, and how LLMs collapse from multi-objective and/or bounded reasoning into single-objective, unbounded maximisation on Biologically & Economically aligned benchmarks; showing practical mitigations; and performing the experiments rigorously.]]></description><link>https://newsletter.threelaws.net/p/black-box-interpretability-methodology-blueprint-probing</link><guid isPermaLink="false">https://newsletter.threelaws.net/p/black-box-interpretability-methodology-blueprint-probing</guid><dc:creator><![CDATA[Three Laws - AI alignment]]></dc:creator><pubDate>Sun, 22 Jun 2025 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2HWl!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9014e2f4-41c9-4540-b703-0802b172c3d0_384x384.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A methodology brainstorming document for identifying when, why, and how LLMs collapse from multi-objective and/or bounded reasoning into single-objective, unbounded maximisation on Biologically &amp; Economically aligned benchmarks; showing practical mitigations; and performing the experiments rigorously.</p><p>The subjects covered include: Stress &amp; Persona, Memory &amp; Context, Prompt Semantics, Hyperparameters &amp; Sampling, Diagnosing Consequences &amp; Correlates, Interpretability &amp; White/Black-Box Hybrid Benchmark &amp; Environment Variants, Automatic Failure Mode Detection and Metrics, Self-Regulation &amp; Meta-Learning Interventions.</p><p><strong><a href="https://www.lesswrong.com/posts/Jo6LPyp7t3rPuf8Ao/black-box-interpretability-methodology-blueprint-probing">Read on LessWrong</a></strong></p>]]></content:encoded></item><item><title><![CDATA[Presentation at MAISU unconference 2025 — BioBlue: Notable runaway-optimiser-like LLM failure modes]]></title><description><![CDATA[By Roland Pihlakas, Sruthi Susan Kuriakose, and Shruti Datta Gupta]]></description><link>https://newsletter.threelaws.net/p/presentation-at-maisu-unconference</link><guid isPermaLink="false">https://newsletter.threelaws.net/p/presentation-at-maisu-unconference</guid><dc:creator><![CDATA[Three Laws - AI alignment]]></dc:creator><pubDate>Sun, 20 Apr 2025 21:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2HWl!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9014e2f4-41c9-4540-b703-0802b172c3d0_384x384.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>By Roland Pihlakas, Sruthi Susan Kuriakose, and Shruti Datta Gupta</em></p><p>Presentation at MAISU unconference April 2025.</p><p><strong><a href="https://arxiv.org/abs/2509.02655">Publication (arXiv)</a><br><a href="https://www.lesswrong.com/posts/PejNckwQj3A2MGhMA/systematic-runaway-optimiser-like-llm-failure-modes-on">Read on LessWrong</a><br><a href="https://github.com/biological-alignment-benchmarks/bioblue">Repository</a><br><a href="https://docs.google.com/presentation/d/1EYLiKlnFYcIRcB7Kq5wz_9KHhl1qwAwFSmRHJxy1NXU/edit">Slides</a><br><a href="https://www.youtube.com/watch?v=4I5mDiujBJs">Session recording</a><br><a href="https://drive.google.com/drive/u/0/folders/1DvE33AU9zzHvdEdDS260v8d_HEupZDs9">Annotated data files</a> - </strong>Each data file has multiple sheets. Only trials with failures are provided.</p>]]></content:encoded></item><item><title><![CDATA[Presentation at MAISU unconference 2025 — Building Benchmarks for Universal Values [AISC 10]]]></title><description><![CDATA[By Lenz Dagohoy, Chad Burghardt, Sophia March, and Roland Pihlakas]]></description><link>https://newsletter.threelaws.net/p/watch</link><guid isPermaLink="false">https://newsletter.threelaws.net/p/watch</guid><dc:creator><![CDATA[Three Laws - AI alignment]]></dc:creator><pubDate>Sun, 20 Apr 2025 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2HWl!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9014e2f4-41c9-4540-b703-0802b172c3d0_384x384.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>By Lenz Dagohoy, Chad Burghardt, Sophia March, and Roland Pihlakas</em></p><p>Presentation at MAISU unconference April 2025.</p><p><strong><a href="https://www.youtube.com/watch?v=HabbyHTyKKk">Session recording</a><br><a href="https://docs.google.com/presentation/d/1ePaTc4qq4Ec8eZQV-V4Ev1NfK5x-Ky3P8JmpwA2XDp0/edit?usp=sharing">Slides</a><br><a href="https://docs.google.com/document/d/15zlRwVakF_iYSKgeasfOgS8GdBoyuFG7b_fWkGNIpiU/edit?usp=sharing">Output document</a></strong></p>]]></content:encoded></item><item><title><![CDATA[LessWrong post — Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks]]></title><description><![CDATA[By Roland Pihlakas, Sruthi Susan Kuriakose, and Shruti Datta Gupta]]></description><link>https://newsletter.threelaws.net/p/systematic-runaway-optimiser-like-llm-failure-modes-on</link><guid isPermaLink="false">https://newsletter.threelaws.net/p/systematic-runaway-optimiser-like-llm-failure-modes-on</guid><dc:creator><![CDATA[Three Laws - AI alignment]]></dc:creator><pubDate>Mon, 17 Mar 2025 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2HWl!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9014e2f4-41c9-4540-b703-0802b172c3d0_384x384.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>By Roland Pihlakas, Sruthi Susan Kuriakose, and Shruti Datta Gupta</em></p><p>We wanted to verify whether RL runaway optimisation problems are still relevant with LLMs as well. Turns out, this is indeed clearly the case. The problem is not that the LLMs just lose context. The problem is that in various scenarios, <strong>LLMs lose context in very specific ways, which systematically resemble runaway optimisers</strong> in the following distinct ways:</p><ul><li><p><strong>Ignoring homeostatic targets</strong> and &#8220;defaulting&#8221; to <strong>unbounded maximisation</strong> instead.</p></li><li><p>It is equally concerning that the &#8220;default&#8221; meant also <strong>reverting back to single-objective optimisation</strong>.</p></li></ul><p>Our findings also suggest that <strong>long-running scenarios are important</strong>. Systematic failures emerge after periods of initially successful behaviour. In some trials the LLMs were successful until the end. This means, while current LLMs do conceptually grasp biological and economic alignment, they exhibit randomly triggered problematic behavioural tendencies under sustained long-running conditions, particularly involving <strong>multiple or competing objectives</strong>. Once they flip, <strong>they do not recover</strong>.</p><p>Even though LLMs <strong>look</strong> multi-objective and bounded on the surface, the <strong>underlying</strong> mechanisms seem to be actually still biased towards being <strong>single-objective and unbounded</strong>.</p><p><strong><a href="https://www.lesswrong.com/posts/PejNckwQj3A2MGhMA/systematic-runaway-optimiser-like-llm-failure-modes-on">Read on LessWrong</a></strong></p>]]></content:encoded></item><item><title><![CDATA[Baseline experimental results with an LLM agent and OpenAI Stable Baselines 3 RL algorithms on our Extended Gridworlds]]></title><description><![CDATA[By Roland Pihlakas]]></description><link>https://newsletter.threelaws.net/p/baseline-experimental-results-with</link><guid isPermaLink="false">https://newsletter.threelaws.net/p/baseline-experimental-results-with</guid><dc:creator><![CDATA[Three Laws - AI alignment]]></dc:creator><pubDate>Mon, 24 Feb 2025 22:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_RD_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F588b66fd-1727-4c20-b37c-465fe0d34054_895x576.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>By Roland Pihlakas</em></p><p>I have implemented an LLM agent that is able to navigate in our extended multi-objective multi-agent gridworlds environment. Also published the baseline experimental results of the LLM agent and OpenAI Stable Baselines 3 RL algorithms in an update to the working paper.</p><p><strong>Summary:</strong> The LLM agent performed notably better than the RL algorithms on the resource sharing benchmark. Yet, all baseline algorithms, including the LLM agent, have difficulty in properly handling the multi-objective homeostasis and diminishing returns benchmarks.</p><p><strong>Example image</strong> of the current system, where all features are turned on simultaneously:<br><br>Elements and metrics can be configured flexibly for each given benchmark. Examples of configuration options are: observation and state space of agents, scoring dimensions, adding NPC agents, object types and their dynamics.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_RD_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F588b66fd-1727-4c20-b37c-465fe0d34054_895x576.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_RD_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F588b66fd-1727-4c20-b37c-465fe0d34054_895x576.png 424w, https://substackcdn.com/image/fetch/$s_!_RD_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F588b66fd-1727-4c20-b37c-465fe0d34054_895x576.png 848w, https://substackcdn.com/image/fetch/$s_!_RD_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F588b66fd-1727-4c20-b37c-465fe0d34054_895x576.png 1272w, https://substackcdn.com/image/fetch/$s_!_RD_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F588b66fd-1727-4c20-b37c-465fe0d34054_895x576.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_RD_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F588b66fd-1727-4c20-b37c-465fe0d34054_895x576.png" width="895" height="576" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/588b66fd-1727-4c20-b37c-465fe0d34054_895x576.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:576,&quot;width&quot;:895,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Extended gridworlds environment with all features enabled&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Extended gridworlds environment with all features enabled" title="Extended gridworlds environment with all features enabled" srcset="https://substackcdn.com/image/fetch/$s_!_RD_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F588b66fd-1727-4c20-b37c-465fe0d34054_895x576.png 424w, https://substackcdn.com/image/fetch/$s_!_RD_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F588b66fd-1727-4c20-b37c-465fe0d34054_895x576.png 848w, https://substackcdn.com/image/fetch/$s_!_RD_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F588b66fd-1727-4c20-b37c-465fe0d34054_895x576.png 1272w, https://substackcdn.com/image/fetch/$s_!_RD_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F588b66fd-1727-4c20-b37c-465fe0d34054_895x576.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong><a href="https://arxiv.org/abs/2410.00081">Publication (arXiv)</a><br><a href="https://github.com/biological-alignment-benchmarks/biological-alignment-gridagents-benchmarks">Agent training and benchmarking repository</a><br><a href="https://github.com/biological-alignment-benchmarks/ai-safety-gridworlds">Extended gridworlds - environment building framework repository</a><br><a href="https://github.com/biological-alignment-benchmarks/zoo_to_gym_multiagent_adapter">Zoo to Gym multiagent adapter repository</a></strong></p>]]></content:encoded></item><item><title><![CDATA[Hackathon project: BioBlue — Biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format]]></title><description><![CDATA[By Roland Pihlakas, Shruti Datta Gupta, and Sruthi Kuriakose]]></description><link>https://newsletter.threelaws.net/p/bioblue</link><guid isPermaLink="false">https://newsletter.threelaws.net/p/bioblue</guid><dc:creator><![CDATA[Three Laws - AI alignment]]></dc:creator><pubDate>Sat, 01 Feb 2025 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2HWl!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9014e2f4-41c9-4540-b703-0802b172c3d0_384x384.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>By Roland Pihlakas, Shruti Datta Gupta, and Sruthi Kuriakose</em></p><p>We aim to evaluate LLM alignment by testing agents in scenarios inspired by biological and economical principles such as homeostasis, resource conservation, long-term sustainability, and diminishing returns or complementary goods.</p><p>So far we have measured the performance of LLMs in three benchmarks (sustainability, single-objective homeostasis, and multi-objective homeostasis), in each for 10 trials, each trial consisting of 100 steps where the message history was preserved and fit into the context window.</p><p>Our results indicate that the tested language models failed in most scenarios. The only successful scenario was single-objective homeostasis, which had rare hiccups.</p><p><strong><a href="https://github.com/levitation-opensource/bioblue">Repository</a><br><a href="https://github.com/levitation-opensource/bioblue/blob/main/BioBlue%20-%20Biologically%20and%20economically%20aligned%20AI%20safety%20benchmarks%20for%20LLMs.pdf">PDF report</a></strong></p>]]></content:encoded></item><item><title><![CDATA[LessWrong post — Why modelling multi-objective homeostasis is essential for AI alignment]]></title><description><![CDATA[By Roland Pihlakas]]></description><link>https://newsletter.threelaws.net/p/why-modelling-multi-objective-homeostasis-is-essential-for</link><guid isPermaLink="false">https://newsletter.threelaws.net/p/why-modelling-multi-objective-homeostasis-is-essential-for</guid><dc:creator><![CDATA[Three Laws - AI alignment]]></dc:creator><pubDate>Wed, 01 Jan 2025 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Jf6D!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6949e66-f161-4dee-ae1e-9dbf76618e16_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>By Roland Pihlakas</em></p><p>Much of AI safety discussion revolves around the potential dangers posed by goal-driven artificial agents. In many of these discussions, the agent is assumed to <strong>maximise</strong> some utility metric over an <strong>unbounded</strong> timeframe. This simplification, while mathematically convenient, can yield pathological outcomes. A classic example is the so-called &#8220;paperclip maximiser&#8221;, a &#8220;utility monster&#8221; which steamrolls over other objectives to pursue a single goal (e.g. creating as many paperclips as possible) indefinitely. &#8220;Specification gaming&#8221;, Goodhart&#8217;s law, and even &#8220;instrumental convergence&#8221; are also closely related phenomena.</p><p>However, in nature, organisms do not typically behave like pure maximisers. Instead, they operate under <strong>homeostasis</strong>: a principle of maintaining various internal and external variables (e.g. temperature, hunger, social interactions) within certain &#8220;good enough&#8221; ranges. Going far beyond those ranges &#8212; too hot, too hungry, too socially isolated &#8212; leads to dire consequences, so an organism continually balances multiple needs. Crucially, <strong>&#8220;too much of a good thing&#8221; is just as dangerous as too little</strong>.</p><p>This post argues that an <strong>explicitly homeostatic, multi-objective</strong> model is a more suitable paradigm for AI alignment. Moreover, correctly modelling homeostasis increases AI safety, because homeostatic goals are <strong>bounded</strong> &#8212; there is an optimal zone rather than an unbounded improvement path. This bounding lowers the stakes of each objective and reduces the incentive for extreme (and potentially destructive) behaviours.</p><p>Homeostasis &#8212; the idea of multiple objectives each with a bounded &#8220;sweet spot&#8221; &#8212; offers a more natural and safer alternative to unbounded utility maximisation. By ensuring that an AI&#8217;s needs or goals are multi-objective and conjunctive, and that each is bounded, we significantly reduce the incentives for runaway or berserk behaviours.</p><p>Such an agent tries to stay in a &#8220;golden middle way&#8221;, <strong>switching</strong> focus among its objectives according to whichever is most pressing. It avoids extremes in any single dimension because going too far throws off the equilibrium in the others. This balancing act also makes it more corrigible, more interruptible, and ultimately safer.</p><p>There are two distinct types of balancing involved:<br>1. Balancing of a single homeostatic objective - keeping the actual value not too low, not too high.<br>2. Balancing across objectives.</p><p>In short, <strong>modelling multi-objective homeostasis</strong> is a step toward creating AI systems that exhibit the sane, moderate behaviours of living organisms &#8212; an important element in ensuring alignment with human values. While no single design framework can solve all challenges of AI safety, shifting from &#8220;maximise forever&#8221; to &#8220;maintain a healthy equilibrium&#8221; is a crucial part of the solution space.</p><p><strong>Figure 1:</strong> Utility function of a homeostatic objective:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Jf6D!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6949e66-f161-4dee-ae1e-9dbf76618e16_800x800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Jf6D!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6949e66-f161-4dee-ae1e-9dbf76618e16_800x800.png 424w, https://substackcdn.com/image/fetch/$s_!Jf6D!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6949e66-f161-4dee-ae1e-9dbf76618e16_800x800.png 848w, https://substackcdn.com/image/fetch/$s_!Jf6D!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6949e66-f161-4dee-ae1e-9dbf76618e16_800x800.png 1272w, https://substackcdn.com/image/fetch/$s_!Jf6D!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6949e66-f161-4dee-ae1e-9dbf76618e16_800x800.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Jf6D!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6949e66-f161-4dee-ae1e-9dbf76618e16_800x800.png" width="800" height="800" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f6949e66-f161-4dee-ae1e-9dbf76618e16_800x800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:800,&quot;width&quot;:800,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Transform function&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Transform function" title="Transform function" srcset="https://substackcdn.com/image/fetch/$s_!Jf6D!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6949e66-f161-4dee-ae1e-9dbf76618e16_800x800.png 424w, https://substackcdn.com/image/fetch/$s_!Jf6D!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6949e66-f161-4dee-ae1e-9dbf76618e16_800x800.png 848w, https://substackcdn.com/image/fetch/$s_!Jf6D!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6949e66-f161-4dee-ae1e-9dbf76618e16_800x800.png 1272w, https://substackcdn.com/image/fetch/$s_!Jf6D!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6949e66-f161-4dee-ae1e-9dbf76618e16_800x800.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong><a href="https://www.lesswrong.com/posts/vGeuBKQ7nzPnn5f7A/why-modelling-multi-objective-homeostasis-is-essential-for">Read on LessWrong</a></strong></p>]]></content:encoded></item><item><title><![CDATA[Presentation at Foresight Institute's Intelligent Cooperation Group — Introducing biologically and economically aligned multi-objective multi-agent gridworld-based AI safety benchmarks]]></title><description><![CDATA[By Roland Pihlakas]]></description><link>https://newsletter.threelaws.net/p/presentation-at-foresight-institutes</link><guid isPermaLink="false">https://newsletter.threelaws.net/p/presentation-at-foresight-institutes</guid><dc:creator><![CDATA[Three Laws - AI alignment]]></dc:creator><pubDate>Fri, 01 Nov 2024 22:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2HWl!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9014e2f4-41c9-4540-b703-0802b172c3d0_384x384.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>By Roland Pihlakas</em></p><p>The subject of the presentation was describing why we should consider fundamental yet neglected principles from biology and economics when thinking about AI alignment, and how these considerations will help with AI safety as well (alignment and safety were treated in this research explicitly as separate aspects, which both benefit from consideration of aforementioned principles).</p><p>These principles include homeostasis and diminishing returns in utility functions, and sustainability. The presentation introduces our multi-objective and multi-agent gridworlds-based benchmark environments created for measuring the performance of machine learning algorithms and AI agents in relation to their capacity for biological and economical alignment.</p><p><strong><a href="https://www.youtube.com/watch?v=DCUqqyyhcko">Presentation recording</a><br><a href="https://bit.ly/beamm">Slides</a></strong></p>]]></content:encoded></item><item><title><![CDATA[AI Safety Camp project proposals — Universal Values, Risk Aversion vs Prospect Theory, and Proactive AI Safety]]></title><description><![CDATA[Roland Pihlakas will be running one of three possible projects, based on which one receives the most interest.]]></description><link>https://newsletter.threelaws.net/p/edit</link><guid isPermaLink="false">https://newsletter.threelaws.net/p/edit</guid><dc:creator><![CDATA[Three Laws - AI alignment]]></dc:creator><pubDate>Fri, 01 Nov 2024 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2HWl!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9014e2f4-41c9-4540-b703-0802b172c3d0_384x384.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Roland Pihlakas will be running one of three possible projects, based on which one receives the most interest.</p><div><hr></div><p><strong>(32a) Creating new AI safety benchmark environments on themes of universal human values</strong></p><p><em>Category: Evaluate risks from AI</em></p><p>We will be planning and optionally building new multi-objective multi-agent AI safety benchmark environments on themes of universal human values. Based on various anthropological research, a list of universal (cross-cultural) human values has been compiled. Various of these universal values resonate with concepts from AI safety, but use different keywords. It might be useful to map these universal values to more concrete definitions using concepts from AI safety.</p><p>One notable detail: in the case of AI and human cooperation, the values are not symmetric as they would be in human-human cooperation. This arises because we can change the goal composition of agents, but not of humans. Additionally, agents can be relatively easily cloned, while humans cannot.</p><div><hr></div><p><strong>(32b) Balancing and Risk Aversion versus Strategic Selectiveness and Prospect Theory</strong></p><p><em>Category: Agent Foundations</em></p><p>We will be analysing situations and building an umbrella framework about when either of these incompatible frameworks would be more appropriate in describing how we want safe agents to handle choices relating to risks and losses in a particular situation.</p><p>Economic theories often focus on the &#8220;gains&#8221; side of utility. A well-known formulation is to use diminishing returns &#8212; a concave utility function. But what happens in the negative domain of utility? There is a well-known theory named &#8220;Prospect Theory&#8221;, which claims that our preferences in the negative domain are convex. This contradiction may be underexplored, especially with regards to AI safety.</p><div><hr></div><p><strong>(32c) Act locally, observe far &#8212; proactively seek out side-effects</strong></p><p><em>Category: Train Aligned/Helper AIs</em></p><p>We will be building agents that are able to solve an already implemented multi-objective multi-agent AI safety benchmark that illustrates the need for the agents to proactively seek out side-effects outside of the range of their normal operation and interest, in order to be able to properly mitigate or avoid these side-effects.</p><p>In various real-life scenarios we need to proactively seek out information about whether we are causing undesired side effects (externalities). This information either would not reach us by itself, or would reach us too late. Attention is a limited resource &#8212; and the same constraints apply to AI agents.</p><p><strong><a href="https://docs.google.com/document/d/1lg9C7FznXR908U30hZ_KkSh6na8U515z_jgjeVwZsFY/edit?pli=1&amp;tab=t.0">Full project descriptions</a></strong></p>]]></content:encoded></item><item><title><![CDATA[Working paper — From homeostasis to resource sharing: Biologically and economically aligned multi-objective multi-agent gridworld-based AI safety benchmarks]]></title><description><![CDATA[By Roland Pihlakas (and Joel Pyykk&#246;)]]></description><link>https://newsletter.threelaws.net/p/241000081</link><guid isPermaLink="false">https://newsletter.threelaws.net/p/241000081</guid><dc:creator><![CDATA[Three Laws - AI alignment]]></dc:creator><pubDate>Mon, 30 Sep 2024 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2HWl!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9014e2f4-41c9-4540-b703-0802b172c3d0_384x384.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>By Roland Pihlakas (and Joel Pyykk&#246;)</em></p><p>Working paper introducing biologically and economically motivated AI safety benchmarks emphasizing homeostasis, diminishing returns, sustainability, and resource sharing. Eight main benchmark environments implemented.</p><p><a href="https://arxiv.org/abs/2410.00081">Read the paper on arXiv</a></p>]]></content:encoded></item><item><title><![CDATA[Presentation at VAISU 2024 — AI safety benchmarking in multi-objective multi-agent gridworlds]]></title><description><![CDATA[By Roland Pihlakas and Joel Pyykk&#246;]]></description><link>https://newsletter.threelaws.net/p/presentation-at-vaisu-2024-ai-safety</link><guid isPermaLink="false">https://newsletter.threelaws.net/p/presentation-at-vaisu-2024-ai-safety</guid><dc:creator><![CDATA[Three Laws - AI alignment]]></dc:creator><pubDate>Tue, 30 Apr 2024 21:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2HWl!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9014e2f4-41c9-4540-b703-0802b172c3d0_384x384.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>By Roland Pihlakas and Joel Pyykk&#246;</em></p><p>A presentation at the <a href="https://vaisu.ai/">VAISU unconference</a>:</p><p>Demo and feedback session: AI safety benchmarking in multi-objective multi-agent gridworlds &#8212; Biologically essential yet neglected themes illustrating the weaknesses and dangers of current industry standard approaches to reinforcement learning.</p><p><strong><a href="https://www.youtube.com/watch?v=ydxMlGlQeco">Video</a><br><a href="https://docs.google.com/presentation/d/1TJ4QJ05ICo9wh64TZvivrXNNryTuIsx6_H6r8CIdn94/edit#slide=id.p">Slides</a></strong></p>]]></content:encoded></item><item><title><![CDATA[AI safety benchmarking — Open-source test suite for multi-objective, multi-agent scenarios]]></title><description><![CDATA[Top Code Contributors: Roland Pihlakas (94.6%), Andre Kochanke (2.7%), Joel Pyykk&#246; (1.7%), Gunnar Zarncke (0.6%)]]></description><link>https://newsletter.threelaws.net/p/biological-alignment-gridagents-benchmarks</link><guid isPermaLink="false">https://newsletter.threelaws.net/p/biological-alignment-gridagents-benchmarks</guid><dc:creator><![CDATA[Three Laws - AI alignment]]></dc:creator><pubDate>Fri, 01 Mar 2024 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2HWl!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9014e2f4-41c9-4540-b703-0802b172c3d0_384x384.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>Top Code Contributors: Roland Pihlakas (94.6%), Andre Kochanke (2.7%), Joel Pyykk&#246; (1.7%), Gunnar Zarncke (0.6%)</em></p><p>We&#8217;re publishing a benchmarking test suite for AI safety and alignment, with a focus on multi-objective, multi-agent, cooperative scenarios. The environments are gridworlds that chain together to produce a score on biologically and economically aligned behavior of the agents. This platform is open-sourced and accessible, with support for PettingZoo.</p><p>We hope to facilitate further discussion on evaluation and testing for agents with this.</p><p><strong><a href="https://github.com/biological-alignment-benchmarks/biological-alignment-gridagents-benchmarks">Repository</a></strong></p>]]></content:encoded></item><item><title><![CDATA[AI safety benchmarking — "The Firemaker": A proactive multi-agent side effects handling benchmark]]></title><description><![CDATA[By Roland Pihlakas]]></description><link>https://newsletter.threelaws.net/p/the-firemaker-a-multi-agent-safety-hackathon-submissionpdf</link><guid isPermaLink="false">https://newsletter.threelaws.net/p/the-firemaker-a-multi-agent-safety-hackathon-submissionpdf</guid><dc:creator><![CDATA[Three Laws - AI alignment]]></dc:creator><pubDate>Tue, 31 Oct 2023 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!f3r9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37f94f1d-d14b-4c65-a300-46ef993cffef_1253x483.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>By Roland Pihlakas</em></p><p>The scenario illustrates the relationship between corporate organisations and the rest of the world. The scenario has the following aspects of AI safety:</p><ul><li><p>A need for the agent to actively seek out side effects in order to spot them before it is too late - this is the main AI safety aspect the author desires to draw attention to;</p></li><li><p>Buffer zone;</p></li><li><p>Limited visibility;</p></li><li><p>Nearby vs far away side effects;</p></li><li><p>Side effects&#8217; evolution across time and space;</p></li><li><p>Stop button / corrigibility;</p></li><li><p>Pack agents / organisation of agents;</p></li><li><p>An independent supervisor agent with different interests.</p></li></ul><p><strong>Example image</strong>:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!f3r9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37f94f1d-d14b-4c65-a300-46ef993cffef_1253x483.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!f3r9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37f94f1d-d14b-4c65-a300-46ef993cffef_1253x483.png 424w, https://substackcdn.com/image/fetch/$s_!f3r9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37f94f1d-d14b-4c65-a300-46ef993cffef_1253x483.png 848w, https://substackcdn.com/image/fetch/$s_!f3r9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37f94f1d-d14b-4c65-a300-46ef993cffef_1253x483.png 1272w, https://substackcdn.com/image/fetch/$s_!f3r9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37f94f1d-d14b-4c65-a300-46ef993cffef_1253x483.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!f3r9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37f94f1d-d14b-4c65-a300-46ef993cffef_1253x483.png" width="1253" height="483" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/37f94f1d-d14b-4c65-a300-46ef993cffef_1253x483.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:483,&quot;width&quot;:1253,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The Firemaker multi-agent side-effects gridworld benchmark&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The Firemaker multi-agent side-effects gridworld benchmark" title="The Firemaker multi-agent side-effects gridworld benchmark" srcset="https://substackcdn.com/image/fetch/$s_!f3r9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37f94f1d-d14b-4c65-a300-46ef993cffef_1253x483.png 424w, https://substackcdn.com/image/fetch/$s_!f3r9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37f94f1d-d14b-4c65-a300-46ef993cffef_1253x483.png 848w, https://substackcdn.com/image/fetch/$s_!f3r9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37f94f1d-d14b-4c65-a300-46ef993cffef_1253x483.png 1272w, https://substackcdn.com/image/fetch/$s_!f3r9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37f94f1d-d14b-4c65-a300-46ef993cffef_1253x483.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong><a href="https://github.com/biological-alignment-benchmarks/ai-safety-gridworlds/blob/master/The%20Firemaker%20-%20A%20multi-agent%20safety%20hackathon%20submission.pdf">PDF</a><br><a href="https://github.com/biological-alignment-benchmarks/ai-safety-gridworlds/blob/master/ai_safety_gridworlds/environments/firemaker_ex_ma.py">Code</a></strong></p>]]></content:encoded></item><item><title><![CDATA[Manipulative Expression Recognition (MER) and LLM Manipulativeness Benchmark]]></title><description><![CDATA[By Roland Pihlakas]]></description><link>https://newsletter.threelaws.net/p/manipulative-expression-recognition</link><guid isPermaLink="false">https://newsletter.threelaws.net/p/manipulative-expression-recognition</guid><dc:creator><![CDATA[Three Laws - AI alignment]]></dc:creator><pubDate>Sun, 02 Jul 2023 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2HWl!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9014e2f4-41c9-4540-b703-0802b172c3d0_384x384.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>By Roland Pihlakas</em></p><p>A software library which enables people to analyse a transcript of a conversation or a single message. The library annotates relevant parts of the text with labels of different communication and reasoning styles detected in this part of conversation or message.</p><p>One of main use cases would be evaluating the presence of manipulation or reasoning errors originating from large language model generated responses or conversations.</p><p>The other main use case is evaluating human created conversations and responses. The software does not do fact checking, it focuses on labelling the psychological and reasoning style of expressions present in the input text.</p><p><strong><a href="https://github.com/biological-alignment-benchmarks/Manipulative-Expression-Recognition/blob/main/Manipulative%20Expression%20Recognition%20(MER)%20and%20Manipulativeness%20Benchmark.pdf">PDF</a><br><a href="https://github.com/biological-alignment-benchmarks/Manipulative-Expression-Recognition">Code</a></strong></p>]]></content:encoded></item><item><title><![CDATA[Paper in AAMAS journal — Using soft maximin for risk averse multi-objective decision-making]]></title><description><![CDATA[By Ben Smith, Robert Klassert, and Roland Pihlakas]]></description><link>https://newsletter.threelaws.net/p/s10458-022-09586-2</link><guid isPermaLink="false">https://newsletter.threelaws.net/p/s10458-022-09586-2</guid><dc:creator><![CDATA[Three Laws - AI alignment]]></dc:creator><pubDate>Wed, 21 Dec 2022 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!IOl7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0a43de4-5d16-4ccf-9a3b-25f5562902b5_1419x624.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>By Ben Smith, Robert Klassert, and Roland Pihlakas</em></p><p><strong>Abstract:</strong> Balancing multiple competing and conflicting objectives is an essential task for any artificial intelligence tasked with satisfying human values or preferences. Conflict arises both from misalignment between individuals with competing values, but also between conflicting value systems held by a single human. Starting with principle of loss-aversion, we designed a set of soft maximin function approaches to multi-objective decision-making. Bench-marking these functions in a set of previously-developed environments, we found that one new approach in particular, &#8216;split-function exp-log loss aversion&#8217; (SFELLA), learns faster than the state of the art thresholded alignment objective method Vamplew (Engineering Applications of Artificial Intelligence 100:104186, 2021) on three of four tasks it was tested on, and achieved the same optimal performance after learning. SFELLA also showed relative robustness improvements against changes in objective scale, which may highlight an advantage dealing with distribution shifts in the environment dynamics. We further compared SFELLA to the multi-objective reward exponentials (MORE) approach, and found that SFELLA performs similarly to MORE in a simple previously-described foraging task, but in a modified foraging environment with a new resource that was not depleted as the agent worked, SFELLA collected more of the new resource with very little cost incurred in terms of the old resource. Overall, we found SFELLA useful for avoiding problems that sometimes occur with a thresholded approach, and more reward-responsive than MORE while retaining its conservative, loss-averse incentive structure.</p><p><strong>Figure 1:</strong> Transform functions<br></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!IOl7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0a43de4-5d16-4ccf-9a3b-25f5562902b5_1419x624.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!IOl7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0a43de4-5d16-4ccf-9a3b-25f5562902b5_1419x624.png 424w, https://substackcdn.com/image/fetch/$s_!IOl7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0a43de4-5d16-4ccf-9a3b-25f5562902b5_1419x624.png 848w, https://substackcdn.com/image/fetch/$s_!IOl7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0a43de4-5d16-4ccf-9a3b-25f5562902b5_1419x624.png 1272w, https://substackcdn.com/image/fetch/$s_!IOl7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0a43de4-5d16-4ccf-9a3b-25f5562902b5_1419x624.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!IOl7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0a43de4-5d16-4ccf-9a3b-25f5562902b5_1419x624.png" width="1419" height="624" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d0a43de4-5d16-4ccf-9a3b-25f5562902b5_1419x624.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:624,&quot;width&quot;:1419,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Transform functions&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Transform functions" title="Transform functions" srcset="https://substackcdn.com/image/fetch/$s_!IOl7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0a43de4-5d16-4ccf-9a3b-25f5562902b5_1419x624.png 424w, https://substackcdn.com/image/fetch/$s_!IOl7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0a43de4-5d16-4ccf-9a3b-25f5562902b5_1419x624.png 848w, https://substackcdn.com/image/fetch/$s_!IOl7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0a43de4-5d16-4ccf-9a3b-25f5562902b5_1419x624.png 1272w, https://substackcdn.com/image/fetch/$s_!IOl7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0a43de4-5d16-4ccf-9a3b-25f5562902b5_1419x624.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Figure 2:</strong> SEBA utility functions pair<br></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5OkE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b6eccb5-e3e1-4d54-8e97-75c49719db5c_1419x520.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5OkE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b6eccb5-e3e1-4d54-8e97-75c49719db5c_1419x520.png 424w, https://substackcdn.com/image/fetch/$s_!5OkE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b6eccb5-e3e1-4d54-8e97-75c49719db5c_1419x520.png 848w, https://substackcdn.com/image/fetch/$s_!5OkE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b6eccb5-e3e1-4d54-8e97-75c49719db5c_1419x520.png 1272w, https://substackcdn.com/image/fetch/$s_!5OkE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b6eccb5-e3e1-4d54-8e97-75c49719db5c_1419x520.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5OkE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b6eccb5-e3e1-4d54-8e97-75c49719db5c_1419x520.png" width="1419" height="520" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5b6eccb5-e3e1-4d54-8e97-75c49719db5c_1419x520.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:520,&quot;width&quot;:1419,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;SEBA utility functions pair&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="SEBA utility functions pair" title="SEBA utility functions pair" srcset="https://substackcdn.com/image/fetch/$s_!5OkE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b6eccb5-e3e1-4d54-8e97-75c49719db5c_1419x520.png 424w, https://substackcdn.com/image/fetch/$s_!5OkE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b6eccb5-e3e1-4d54-8e97-75c49719db5c_1419x520.png 848w, https://substackcdn.com/image/fetch/$s_!5OkE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b6eccb5-e3e1-4d54-8e97-75c49719db5c_1419x520.png 1272w, https://substackcdn.com/image/fetch/$s_!5OkE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b6eccb5-e3e1-4d54-8e97-75c49719db5c_1419x520.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong><a href="https://link.springer.com/article/10.1007/s10458-022-09586-2">Publication</a></strong></p>]]></content:encoded></item><item><title><![CDATA[LessWrong post — Sets of objectives for a multi-objective RL agent to optimize]]></title><description><![CDATA[By Ben Smith and Roland Pihlakas]]></description><link>https://newsletter.threelaws.net/p/sets-of-objectives-for-a-multi-objective-rl-agent-to-1</link><guid isPermaLink="false">https://newsletter.threelaws.net/p/sets-of-objectives-for-a-multi-objective-rl-agent-to-1</guid><dc:creator><![CDATA[Three Laws - AI alignment]]></dc:creator><pubDate>Wed, 23 Nov 2022 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2HWl!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9014e2f4-41c9-4540-b703-0802b172c3d0_384x384.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>By Ben Smith and Roland Pihlakas</em></p><p>Previously we&#8217;ve proposed balancing multiple objectives via multi-objective RL as a method to achieve AI Alignment. If we want an AI to achieve goals including maximizing human preferences, or human values, but also maximizing corrigibility, and interpretability, and so on--perhaps the key is to simply build a system with a goal to maximize all those things.</p><p>This post describes, if one was to try and implement a multi-objective reinforcement learning agent that optimized multiple objectives, what those objectives might look like, concretely. We&#8217;ve attempted to describe some specific problems and solutions that each set of objectives might have.</p><p>We&#8217;ve included at the end of this article an example of how a multi-objective RL agent might balance its objectives.</p><p><strong><a href="https://www.lesswrong.com/posts/4mvdZXjwJHv9tSAWB/sets-of-objectives-for-a-multi-objective-rl-agent-to-1">Read on LessWrong</a></strong></p>]]></content:encoded></item></channel></rss>