Free energy minimization in variational inference
For a long time, I’ve been trying to understand the formal framework of the Free Energy Principle (FEP), particularly its place in my mental model of curiosity and information-seeking. Part I of this post will quickly recap the logic behind the FEP. Then, in part II, I will discuss how the FEP relates to curiosity theories.
Part 1: Free Energy Principle
Motivation behind free energy
The FEP implies that organisms seek to minimize free energy (F) when deciding on how to act in the environment. At a high level, the principle holds that organisms should seek to minimize surprise that they expect to experience in their predicted futures (technically, the average surprise of future states). Before diving into how this principle is stated mathematically, it is worth mentioning from the get go that it rests on multiple assumptions that make the principle less obvious if violated. For example, Friston (Friston 2009) often explicitly invokes the ergodic assumption, which roughly means that in the long run and under an invariant behavioral policy, the agent is expected to converge on the same distribution of visited states regardless of where it started. There might be other theoretical and mathematical assumptions in the foundation of the FEP, but the one I want to make clear here is the assumption that the brain approximates the ideal Bayesian inference in a specific way – via variational inference (VI).
VI is an approximate Bayesian inference method that addresses the intractability problem of computing the posterior over hidden states. The agent is assumed to operate in partially observable conditions, meaning that it never knows with certainty the true state that the environment (including the agent itself) is in. An idealized Bayesian agent would compute the probability of being in state s given an observation o using the Bayes rule:
p(s \mid o) = \frac{p(o|s)p(s)}{p(o)}.
However, the term p(o), which denotes the marginal probability of observing o across all possible states, is often difficult to compute. VI’s approach is to sacrifice accuracy for tractability by “taking p(o) out of the equation”. We assume that while the agent does not know the true posterior p(s|o), it adopts a “working model” q(s; o) to approximate it. This working model, sometimes called recognition or variational density, is a function of the state s just like p(s|o). However, it does not need to abide by the Bayes rule and compute the marginal q(o), since q(s; o) is not defined as a conditional density. In this model, o can be constant parameter or even be effectively taken out of the equation. For convenience, I’ll use a subscript notation and write q_o(s) \equiv q(s; o).
When s is multidimensional, p(o) = \int_{s'} ds' ~ p(o|s')p(s'), or p(o) = \sum_{s'} ~ p(o|s')p(s') can become difficult to compute, as the number of p(o|s')p(s') evaluations grows exponentially. E.g., if we have n binary features, computing p(o) may require 2^n evaluations. If the n features are continuous, but discretized into m values, p(o) may require up to m^n evaluations. This problem is called the curse of dimensionality.
The goal of VI is to find a model q^* that has minimal Kullback-Liebler divergence to the true but unknown posterior:
q_o^* = \displaystyle \underset{q_o}{\operatorname{arg\,min}} ~ D_{KL}\big[q_o(s)\parallel p(s | o)\big]
This objective makes sense since D_{KL}(Q \parallel P) measures the dissimilarity (or distance) between probability distributions Q and P. If Q=P, D_{KL}(Q \parallel P)=0. D_{KL} is defined as the expected log-likelihood ratio between two models, so in VI we can write:
D_{KL}\big[q_o(s) \parallel p(s | o)\big] = \sum_s q_o(s) \ln \frac{q_o(s)}{p(s|o)}
In the discrete case; if s is continuous, we replace the summation operator with \int_s ds. More generally, we can write it as an expectation on s under q:
D_{KL}\big[q_o(s) \parallel p(s | o)\big] = \mathbb{E}_{q_o} \Bigg[ \ln \frac{q_o(s)}{p(s|o)} \Bigg]
How do we minimize the objective while avoiding computing p(s|o)? The trick is to leverage (1) the Bayes rule, by which p(s|o) = p(o|s)p(s)p(o)^{-1}, and (2) logarithmic identities to turn multiplication/division inside a logarithm into addition/subtraction. Thus:
\begin{align*} D_{KL}\big[q_o(s) \parallel p(s | o)\big] & = \mathbb{E}_{q_o} \Bigg[ \ln \frac{q_o(s)}{p(s|o)} \Bigg] \\ & = \mathbb{E}_{q_o} \Big[ \ln q_o(s) - \ln p(s|o) \Big] \\ & = \mathbb{E}_{q_o} \Big[ \ln q_o(s) - \ln \frac{ p(o|s)p(s) }{p(o)} \Big] \\ & = \mathbb{E}_{q_o} \Big[ \ln q_o(s) - \big(\ln [p(o|s)p(s)] - \ln p(o) \big)\Big] \\ & = \mathbb{E}_{q_o} \Big[ \ln q_o(s) - \ln [p(o|s)p(s)] + \ln p(o) \Big] \\ & = \mathbb{E}_{q_o} \Big[ \ln q_o(s) - \ln [p(o|s)p(s)] \Big] + \ln p(o) , \end{align*} where in the last step, we note that p(o) is a marginal distribution that does not depend on s, so it can be taken out of the expectation. The final step towards F is to notice that p(o|s)p(s) = p(s,o), giving us the following identity:
D_{KL}\big[ q_o(s) \parallel p(s | o) \big] = \mathbb{E}_{q_o} \Big[ \ln q_o(s) - \ln p(s, o) \Big] + \ln p(o). \tag{1}
It might be tempting to write \mathbb{E}_{q_o} \Big[ \ln q_o(s) - \ln p(s, o) \Big] as D_{KL}\big[ q(s) \parallel p(s,o) \big] but D_{KL}(Q \parallel P) is only defined when Q and P are distributions over the same variable. Here, the reference distribution is over a joint variable (s,o), not just s.
The expectation term on the right-hand side is what’s called variational free energy, or just free energy. Formally,
F := \mathbb{E}_{q_o} \Big[ \ln q_o(s) - \ln p(s, o) \Big] \tag{2} which allows us to write:
F = D_{KL}\big[q_o(s) \parallel p(s | o)\big] - \ln p(o). \tag{3}
Negative free energy is sometimes referred to as evidence lower bound (or ELBO), because the marginal (log) evidence term \ln p(o) cannot go below -F. To see this clearly, rearrange the above equation a bit: \ln p(o) = D_{KL}\big[q_o(s) \parallel p(s | o)\big] - F and think of what happens as q_o approaches p(\cdot|o). The KL-divergence term shrinks to 0, leaving only -F. I.e. \ln p(o) \ge -F
There are probably good reasons for starting with the definition of F like in Equation 3, but I feel like substituting F into Equation 1 provides a clearer motivation for the FEP:
D_{KL}\big[ q_o(s) \parallel p(s | o) \big] = F + \ln p(o). An agent wants to minimize F because it wants to have a mental model (q) as close to the true model p(s|o) as possible. Computing D_{KL}\big[ q_o(s) \parallel p(s | o) \big] directly is difficult because computing p(o), which shows up on the right hand side, is difficult. However, computing F is possible, as it requires only \ln q_o(s) and \ln p(o|s)p(s).
Beyond variational inference
We can decompose further the VI-based definition of F from Equation 2. As mentioned before, the joint density can be factorized into p(s|o)p(o). By rearranging some terms we get to a further decomposition of F into two interpretable terms (Yanagisawa and Honda (2025)):
\begin{align*} F & = \mathbb{E}_{q_o} \Big[ \ln q_o(s) - \ln p(s, o) \Big] \\ & = \mathbb{E}_{q_o} \Big[ \ln q_o(s) - \ln [p(o|s)p(s)] \Big] \\ & = \mathbb{E}_{q_o} \Big[ \ln q_o(s) - [ \ln p(o|s) + \ln p(s) ] \Big] \\ & = \mathbb{E}_{q_o} \Big[ \ln q_o(s) - \ln p(s) - \ln p(o|s) \Big] \\ & = D_{KL}\big[ q_o(s) \parallel p(s)\big] - \mathbb{E}_{q_o} \big[ \ln p(o|s) \big] \\ \end{align*} \tag{4}
The new KL-divergence term is distinct from the VI objective. It tells us how close the mental-model inference about the state after observing o is to the prior (not the posterior!) belief about the state. When q_o(s) \equiv p(s|o):
- F becomes -\ln p(o) (i.e., we reach the lower bound of the evidence term denoting the marginal probability of observation in the generative model);
- D_{KL}\big[ q_o(s) \parallel p(s)\big] becomes equivalent to idealized Bayesian surprise (Itti and Baldi 2009), i.e., D_{KL}\big[ p(s|o) \parallel p(s)\big];
- -\mathbb{E}_{p(s|o)} \big[ \ln p(o|s) \big] becomes -\mathbb{E}_{p} \big[ \ln p(s|o) \big] = -\sum_s p(s|o) \ln p(o|s), or the value of Shannon’s surprise of the observation (p(o|s)) expected if s was distributed according to p(s|o). It seems appropriate to call this term “residual uncertainty”, as it denotes the average surprise (just like entropy) after accounting for uncertainty over the hidden state conditioned on the observation.
Interestingly, (F. Poli et al. 2020) refer to D_{KL}\big[ p(s|o) \parallel p(s)\big] as learning progress.
To finalize Part 1, let’s write out the big equation combining all the main theoretical quantities:
\underbrace{D_{KL}\big[ q_o(s) \parallel p(s | o) \big]}_{\substack{\text{Objective function(al) or} \\ \text{what agents want to minimize}}} = \underbrace{D_{KL}\big[ q_o(s) \parallel p(s)\big]}_{\substack{\text{Distance b/w observation-informed} \\ \text{subjective belief about } s \text{ and prior}}} - \underbrace{\mathbb{E}_{q_o} \big[ \ln p(o|s) \big]}_{\substack{\text{Given current subjective belief} \\ \text{about } s, \text{ how surprising is } o?}} + \underbrace{\ln p(o)}_{\substack{\text{Overall (log-)likelihood of} \\ o \text{ across all possible states}}} Note that -\ln p(o) Overall (log-)likelihood is constant with respect to q_o. It offsets which value of F (first two terms on the right) is optimal, but has no influence on which q_o optimizes the objective.
Part 2: Free Energy and Curiosity Theories
I still need to think how the FEP framework relates to various formal theories of curiosity, such as Berlyne/Loewenstein’s operationalization of uncertainty (or information gap) as Shannon’s entropy Golman and Loewenstein (2018), Kidd’s operationalization of complexity as Shannon’s surprise (Kidd, Piantadosi, and Aslin 2012), Itti and Baldi’s Bayesian surprise (Itti and Baldi 2009), or Poli’s mapping of information theoretic quantities to psychological constructs surrounding curiosity (Francesco Poli et al. 2024), etc. There are definitely connections.
If you want to work with me on these, feel free to reach out (alexandr.ten at uni-tuebingen.de).