POMDPs, filtering, and curiosity
Bayes filtering with partial observability
The structure of a partially observable Markov decision process, or POMDP (Kaelbling, Littman, and Cassandra 1998) is a tuple \langle \mathcal{S}, \mathcal{A}, \mathcal{O}, \tau, \rho, \omega \rangle, where the elements \mathcal{S}, \mathcal{A}, \tau, \rho describe a classical MDP:
| \mathcal{S} | The set of all states of the environment. |
| \mathcal{A} | The set of all actions available to an agent (implicitly constrained by the state, which can be implemented in the agent’s policy). |
| \tau(s'; a, s) | The state transition function, giving the conditional probability of transition to state s' given previous state s and action a, i.e., \tau(s'; a, s) = Pr(s' \mid a, s). |
| \rho | The reward function. |
| \mathcal{O} | A set of observations that the agent can experience. |
| \omega(o; a, s') | The observation function, the probability of observing o after taking action a and ending up in state s': \omega(o; a, s') = Pr(o \mid a, s'). If O is conditionally independent of A given S', then \omega(o; s') = Pr(o \mid s'). |
Since states are only partially observed, the agent maintains a belief about which state the environment is in. Under it uncertain beliefs, the agent must find an action policy that maximizes the long term return of discounted rewards. I.e., the agent wants to find a way of behaving that results in getting what it wants going forward. There are three major aspects of POMDP that I will discuss in the context of intrinsically motivated behavior: (1) belief maintenance, (2) action policies, and (3) values and rewards.
Belief maintenance
A belief state – denoted bel – is a probability distribution over S. The belief state is updated each time the agent takes an action a, transitions to a new state s', and gets an observation o:
bel(s' ; o,a) = \frac{\omega(o; a, s')\sum_{s \in \mathcal{S}}\tau(s';a,s)bel_0(s)}{Pr(o|a)}
It is common to define \eta = Pr(o|a)^{-1} and write
The normalization term is the summation over the values of S': Pr(o|a) = \sum_{s' \in \mathcal{S}} \omega(o; a, s')\sum_{s \in \mathcal{S}}\tau(s';a,s)bel_0(s)
bel(s'; o,a) = \eta\Big[{\omega(o; a, s')\sum_{s \in \mathcal{S}}\tau(s';a,s)bel_0(s)}\Big] \tag{1}
This updating procedure is also known in robotics as Bayes filtering (Thrun 2002). In this field, this procedure is broken down into two steps:
- Prediction, in which the robot uses its knowledge about the world to form a beleif about the current state.
- Correction, in which the robot accounts for what it is observing to arrive at a better estimate.
The prediction step may be written as \overline{bel}(s'; a) = \sum_{s \in \mathcal{S}}\tau(s'; a, s)bel_0(s) \tag{2} The correction step then becomes bel(s'; o, a) = \eta \omega(o; s') \overline{bel}(s'; a) \tag{3}
The two step Bayes filtering map nicely to a common description of Bayesian posterior inference in terms of a product between the prior and the likelihood. Here, the prior probability of some state s' is computed from the distribution over previous states and the state transition function. It is telling us the probability of s' without considering any direct evidence for that state. The likelihood of s' is evaluated against the observation (also called measurement) function \omega. This term “fine-tunes” (i.e., corrects) the prior probability by weighing it agains what is being observed.
As long as we specify \omega and \tau, we know how to describe belief updating (but not necessarily how to compute it). Specifications of these functions may reflect our assumptions about what the agent knows or should know about its environment (\tau) and what it can observe (\omega).
We shall assume that agents represent \tau internally as an internal transition model approximating how states and actions map to future states. Whereas the true state transitions are constrained solely by reality, the internal model predicts state transitions based on uncertain beliefs about states and the knowledge on how various elements of the world state interact, i.e., “how the world works”. Abstractly, we can think of \tau as having a particular parametric form that constrains how s' gets assigned a probability measure based on a and s. We can symbolize the parameters of \tau with \theta. It is natural to also assume that agents may be more or less certain about the particular values of these parameters, and accordingly maintain a belief about them. Similarly, we can imagine different individuals hosting different internal observation models \omega.
Variational inference and free energy minimization
The formulation of belief updating as Bayes filtering lets us examine it through the Free Energy Principle (FEP) framework. Specifically, the FEP framework allows us to formulate the reward function (that we have ignored so far) with normative constraints. The two frameworks are complementary. For example, Bayes filtering can be viewed as elaborating on the computation of the prior that FEP treats as a given. We can view Equation 3 as a familiar Beyesian inference considered in the FEP framework, where \overline{bel}(s') is the prior, \omega(o; s') is the likelihood, and \eta is the infamous uncomputable evidence term, which makes the posterior computation intractable.
The FEP is based on Variational Inference (VI) by which the true posterior is approximated with a variatonal (or recognition) model. In Bayes filtering, this subjective model corresponds to bel(s'). Thus, we can define a normative VI objective for a Bayes-filtering agent as the KL divergence between bel(s') and the true posterior Pr(s'|o):
\begin{align*} D_{KL} & \Big[bel(s'; o, a) \parallel Pr(s'|o) \Big] := \mathbb{E}_{bel} \Big[ -\ln bel(s'; o, a) - \ln Pr(s'|o) \Big] \\ = & \mathbb{E}_{bel} \Big[ -\ln bel(s'; o, a) - \ln Pr(s', o) \Big] - \ln Pr(o) \end{align*} \tag{4}
This quantity is minimized (to 0) as bel(s') approaches Pr(s'| o). However, since Pr(s'|o) is not available, the agent can minimize free energy instead:
F = \mathbb{E}_{bel} [ -\ln bel(s') - \ln Pr(s', o)]
As bel(s'; o, a) approaches the true Pr(s' | o), F shrinks towards \ln Pr(o).
Yanagisawa and Honda (2025) add that when bel(s'; o, a) \equiv Pr(s'|o), the \ln Pr(o) term decomposes into two:
-\ln p(o) = BS + PX
where BS denotes Bayesian surprise (Itti and Baldi 2009) which corresponds to the average subjective (log) likelihood ratio between posterior and prior:
The average is subjective, because the expectation distribution is bel.
BS = D_{KL} \Big [ bel(s'; o, a) \equiv Pr(s'| o) \parallel Pr(s') \Big ] \tag{5}
and PX denotes observational perplexity which reflects the average subjective (Shannon’s) surprise of the observation o, given s':
Yanagisawa and Honda (2025) refer to PX as “uncertainty”, but I will reserve the term for Shannon’s entropy. I think ‘observational perplexity’ is fitting because it computes the surprise (opposite of likelihood) of the given observation across different states, weighting it by subjective state probability. In other words, it tells us: given the observation-informed belief about the current state (s'), how surprising is the observation? I think it is fair to say that observing something unlikely can be perplexing.
PX = -\mathbb{E}_{bel(s'; o, a) \equiv Pr(s'| o)} \Big[ \ln Pr(o | s') \Big].
The convergence of bel(s'; o,a) towards the true posterior (however it happens) signifies a type of short term learning – the learning of the current hidden state from a given observation. This type of learning is called recognition in Yanagisawa and Honda (2025), who distinguish it from the adoption of bel(s'; o,a) as the prior for the next time point; a kind of belief shift. When Pr(s') from Equation 5 becomes bel(s'; o, a), the BS term becomes 0 and F reduces to PX.
The FEP tells us that to arrive at an accurate representation of the current state, the agent can minimize F. One part of F (D_{KL}[\overline{bel}_a \parallel bel_o]) is reduced by processing the observation and “recognizing” the current hidden state. The other part (D_{KL}[bel_o \parallel overline{bel}_a]) is reduced by adopting the recognized state as a prior for future predictions.
Goal pursuit
So far, we went over two intricately related frameworks that describe inference. However, curiosity is a driver of behavior – it determines preferences for various actions. Suppose an agent acts according to a policy pi that maps actions to probability, which reflects motivation – the relative strength or tendency to behave a particular way. Humans seem to have rather flexible policies that depend on the desired (i.e., goal) state. We can express this as follows:
\pi(a; s, s_*) = \tau(s_*; s, a)
i.e., the probability of selecting action a in state s is equal to the transition probability of arriving at a goal state s_* from s via a.