The Origins of Consciousness — Know Thyself: A Physical Theory of Self-Modelling

Consciousness is overrated. We spend large fraction of our lives in willing forfeit of it: too much consciousness will kill you. It is not involved in every action that we take. Each time we act in the world, new motor programs are generated unconsciously depending on the setting. If we had to be consciousness of every action and sensation, every minute of our waking lives, we would also likely die. There is just too much at stake in surviving a single day to hand it all over to conscious control. 

Hyper vigilance is the name given to a state of “too much consciousness”. Hyper vigilance can cause increased anxiety, elevated heart rate, and challenges in social interactions. These symptoms can  impact your life, making it difficult to relax or focus on daily tasks. Thinking becomes difficult, and perception is altered, leading  to physical and mental fatigue, decreased fine motor skills, and impaired decision-making. 

Consciousness is costly. The need to sleep is the ‘smoking gun’ for consciousness. All animals must  spend a good part of every day recovering from it. That cost could only be born if it conferred a considerable advantage in survival rates. It seems unlikely that this advantage would accrue to humans alone. Animals need to sleep too, indicating that they bear the cost of conscious.  It cannot be exclusive to humans but a widespread evolutionary adaptation.

What is the necessary hardware required for a living thing to be conscious? It seems likely that a central nervous system is essential.  Simple organisms (e.g., bacteria, plants) lack the necessary neural complexity to be conscious. Insects and small animals may have minimal forms of awareness but not full consciousness. Humans, primates, cetaceans, and some birds exhibit strong evidence of consciousness. But what precisely does a central nervous system do, and are there other hardware enablers for consciousness?

All animals are learning machines. They consume energy to continuously reconfigure brains to predict the ever changing sensations that follow actions. In so doing they necessarily lower their internal entropy paying a price by dissipating heat into the environment.   Contemporary engineered learning machines do the same but much less efficiently. They consume far more energy, and dissipate far more heat, than is necessary for learning. 

All animals are learning machines, but are all animals conscious? Humans claim they are, but struggle to explain what this means, and are unsure if other animals are conscious. If animals are not consciousness, learning is not sufficient for consciousness.

A minimal model for a learning machine requires actuators, that change the local environment, and sensors that respond to changes in the environment.  A small subset  of those changes are due to internally generated actions. A learning machine must predict what sensations follow a given action. To do so they must distinguish internally generated changes in the local environment from those that are not. Evolution will drive this process to become ever more efficient. 

Self-awareness is a necessary precondition for learning in embedded learning machines.  If self-awareness is necessary for consciousness, then learning is necessary for consciousness, but is self-awareness sufficient for consciousness? Are  all embedded learning machines conscious? 

The simplest actuator for a learning machine is the ability to move. Llinas has made a strong case for motility as a necessary precondition for the evolution of a central nervous system in animals.    Motility is not essential for robots, but it would appear to be a highly desirable feature. It certainly makes it far more difficult to engineer a robot if it is required to move rather than remain stationary. 

The minimal requirement for an autonomous motile learning machine is an internal gyroscope and internal clock. Of course, a motile learning machine need not be autonomous. A drone equipped with sensors and actuators can learn a lot about the world, and effect changes in it,  even if its motion is controlled remotely by a controller.  Modern warfare has made this very apparent.

Here is a very simple example. Imagine an autonomous learning machine — an agent — confined to move on the circumference of a disc. In addition to an internal clock and a gyroscope, the agent can send and receive pulses of light. An internal actuator enables it to rotate the direction of emitted pulse —the pointer. I will assume that the boundary of the disk is perfectly reflecting so that a pulse of light sent into the interior will eventually be received back by the agent, if sent at rational multiples of pi.  The gyroscope estimates the direction of its pointer as a function of the ticks of the internal clock. An example of possible configurations is shown below.

The gyroscope is a simple sensor that measures angular accelerations of the agent’s head direction. When combined with a clock, it enables the agent to keep a record of head direction.  The agent can only `know’ its internal states indexed by ticks of the internal clock. 

I will assume that the agent emits a pulse immediately after receiving a pulse: if a pulse is received at a clock count of m, a pulse is emitted at the same clock count m.  The head direction at each step determines the clock count of the next emission. We can label the head direction `cause’ and the clock count for the next emission as ‘effect’. 

Each emission event corresponds to a distinct internal state labelled by a setting of its internal gyroscope, labelled a, and a reading of the clock at pulse emission, n. Our protocol means that a pulse received at clock count of  is a pulse returning from the previous emission when the head direction was a.   The only thing the agent has access to are these two things. An internal state is an ordered pair of numbers  S=(a, n) where  n is the clock count for the next emission, and  a is the record of head direction at the previous emission.  The first entry of the pair is a cause while the second entry is the effect, see figure below.

We now equip the agent with an internal learning machine. It works like this. For the first emission the agent elects a head  setting, a,  at random and sets the clock counter to n=0.   Simultaneously, the learning machine is sent the head direction, a.  It then quickly tries to predict when the next light pulse will be received by generating an integer n. For simplicity, I will assume there are only four settings for a,  designated (1,2,3,4). These are labelled with angles given by

, but the agent does not know this. It only knows that there are four different headings, as recorded by its internal gyroscope.  At each emission step one of these four headings is chosen at random. 

Given our external god-like view of the agent’s world, we can easily give a relation between head angle and time taken for a pulse to return to the agent. If we assume that space is Euclidean inside the disk, the time taken is then given by 

where I have used an arbitrary scale for time units. The shortest period occurs for a pulse sent along a diagonal and the longest for a pulse sent at 45 degrees to the diagonal. 

The agent does not know this function. It simply generates a list of ordered pairs made up of a label for the head angle and number of ticks of the clock until the pulse returns. The data can be displayed like the table below. 

Heading Count
141
250
456
354
354
354
354
456
354
456



I will assume that the agent has machine that enables it to learn this function.   This can be almost anything, for example a physical neural network or a decision tree. I will assume that is a machine that is described by a learning algorithm (nearest neighbour) and that it has a very large number of examples to train on. In the figure below I plot an example of the predictions made by this machine once it has been trained using 1000 training pairs.

The machine can use its sensors and actuators and learn to predict the return time of a pulse for a given head direction, working only with its internal states. If such agents were inclined to speculation, they might propose that the world they inhabit is a boundary to a two dimensional Minkowski disk. Or they might not. In any case, once the agent has leaned this relationship it can control when it will receive a pulse emitted at any head direction. This is the point of learning: once learned, the relationship can be used to intervene effectively on the world.

Once the agent has learned a good approximation to the ground truth function, it can begin to classify received pulses that it did not originate as spurious background pulses. This is the origin of self awareness. It is an entirely physical process and intrinsic to a physical learning machines. Even very simple learning machines would easily develop this ability.

Without self awareness, a robot equiped with a learning machine would be hardly viable and possibly dangerous.  I conclude that embedded learning machines would be driven by evolution to become self aware. 

How do we get from self awareness to consciousness? Before I can answer that, I need to consider how stable communities of  learning agents arise. Agents must not only learn to distinguish self generated sensations from extrinsic sensations, they must be able to distinguish sensations that arise from the actions of other agents of the same kind. I will address this in the next post.