Self-deception and negligence: Fundamental limits to computation due to limitations of attention-like processes (Definition of self-deception in the context of AI safety)
This document about self deception is mainly not a solution to a problem. It is a description of a (in my opinion) very serious problem that has not received (any?) attention so far. Towards the end I venture to suggest some partial solutions though.
I believe the problem could be formalised and put into code to be used as a demonstration of the danger.
The main point is that the danger is not somewhere far away requiring some very advanced AI, but rather it is more like a law of nature that starts manifesting beginning from rather simple systems without any need for self-reflection and self-modification capabilities etc. So instead of the notion that danger springs from some special capabilities of intelligent systems, I want to point out that some other special capabilities of intelligent systems would be needed to somehow evade the danger.


