Information gain in sequential variational inference

(a)
Prior
Likelihood
Posterior
(b)
p(o)
−ln p(o)
KLD
BS
(c)
KLD
BS
IG = KLD + BS
Prior belief μ₀ fixed at 0
Observation
off

How to read these plots

The paper treats learning from a single observation as free energy minimization carried out in two sequential steps, each releasing its own information gain. Step 1 — recognition: before the observation, the agent's best guess about the hidden state s is just the prior p(s). Recognizing the observation updates this guess to the true posterior p(s|o); the free energy released by that update equals KLD = D_KL[p(s) ‖ p(s|o)] — the first information gain, measuring how much was learned by recognizing this particular observation right now. Step 2 — prior update (learning): that posterior is then adopted as the new prior, ready for the next observation. The free energy this second step is expected to release is bounded by BS = D_KL[p(s|o) ‖ p(s)] — the second information gain, measuring the novelty of the observation relative to what was already believed. Because KL divergence isn't symmetric, KLD and BS compare the same two distributions in opposite directions and genuinely differ; the paper sums them into a single total, IG = KLD + BS, and proposes that as the valence signal behind epistemic emotions like curiosity and interest across the Wundt curve.

(a) Prior, likelihood, posterior

  • Prior — belief about the hidden state s before observing anything. The app fixes its mean at η = 0 and lets you set its spread σ₀.p(s) = N(s; η, σ₀²)Paper: §3.1 (Gaussian model). Not touched by the §3.2 noise correction.
  • Likelihood — the observation model read as a function of s: how well each candidate state explains the observed value o. With uniform noise switched on, a small floor ε is added so that no state is ever ruled out entirely.p(o|s) = N(o; s, σ²)pε(o|s) = α · [N(o; s, σ²) + ε], α = 1/(1+ε)Paper: §3.1 (Gaussian); Eq. 16 (§3.2, uniform noise). ε = 10⁻³ throughout.
  • Posterior — prior and likelihood combined via Bayes' rule: a pure Gaussian under §3.1, or a mixture of the Gaussian posterior and the prior once uniform noise is on. Its weights depend on how expected the observation was.1/σpost² = 1/σ₀² + 1/σ², μpost = σpost² · (η/σ₀² + o/σ²)pε(s|o) = wpost · N(μpost, σpost²) + wpri · N(η, σ₀²)wpost = e(δ)/(e(δ)+ε), wpri = ε/(e(δ)+ε)Paper: §3.1 (Gaussian posterior); Eq. 18 (§3.2, mixture posterior).
  • δ bracket — the gap between the prior's and likelihood's centers: the same number swept along the x-axis in (b) and (c).δ = o − ηPaper: written η − o. The sign is flipped here so δ ≥ 0; only δ² enters any formula, so nothing changes.

All three densities sit on the same state axis, so their overlap is visible directly: a posterior that lands close to the likelihood and far from the prior means the observation dominated belief; one that barely moved off the prior means the reverse. The δ bracket above the curves makes that gap explicit — everything plotted in (b) and (c) is this one number, δ, run through different functions of the fixed σ₀ and σ.

(b) Evidence and divergences vs. δ

  • p(o) — marginal likelihood or evidence: how expected the observation was, averaged over the prior's uncertainty. Dashed line, left axis; bounded in [0, 1].e(δ) = exp(−δ² / 2(σ₀²+σ²))pε(o) = α · (e(δ) + ε)Paper: Eq. 13 (Gaussian evidence); Eq. 17 (§3.2, with uniform noise).
  • −ln p(o) — surprise, the negative log of the evidence and the model's raw prediction-error signal. Solid line, right axis. Same color as p(o) because it is the same quantity, just log-transformed.F = −ln p(o) = δ² / 2(σ₀²+σ²) (Gaussian, up to a constant)F = −ln[α · (e(δ) + ε)] (with uniform noise)Paper: the free-energy term F released by recognition; it becomes the x-axis of (c).
  • KLD — recognition information gain from step 1: how far the belief moves from the prior to the posterior. Right axis.KLD = DKL[ p(s) ‖ p(s|o) ]Paper: Eq. 14. Closed form under §3.1; with uniform noise, integrated numerically against pε(s|o) (Appendix Eq. 26–29 have no closed form).
  • BS — anticipated learning information gain from step 2: the novelty of the posterior relative to the prior. Right axis.BS = DKL[ p(s|o) ‖ p(s) ]Paper: Eq. 15. Same divergence as KLD in the opposite direction; likewise integrated numerically under §3.2.

All four curves are functions of the single number δ, swept from 0 to 20. With η fixed at 0 and o restricted to ≥ 0, δ = o − η never goes negative, so this range has no redundant mirror image. p(o) is bounded in [0,1] (left axis); the three divergence-related quantities are unbounded and can grow without limit under the pure-Gaussian model — exactly the problem §3.2's uniform-noise correction exists to fix. Under §3.1 (pure Gaussian), KLD and BS have the closed forms above; switch on "+ uniform noise" and they're instead computed by numerically integrating the mixture posterior p_ε(s|o) against the prior (Simpson's rule over s), since the paper's own Appendix (Eq. 26–29) has no closed form for that case. The dashed line and dots mark where the current sliders' δ sits on each curve.

(c) Information gains vs. surprise

  • KLD — the same recognition information gain as in (b), re-plotted against surprise F instead of δ. Under §3.1 it is an exact straight line in F.KLD = (σ₀²/σ²) · F + ½[ ln(σ²/(σ₀²+σ²)) + (σ₀²+σ²)/σ² − 1 ]Paper: Eq. 14, with F from Eq. 13. The uniform-noise version (§3.2) is numerical.
  • BS — the same anticipated learning information gain as in (b), against surprise. Also an exact straight line in F under §3.1, with a different slope.BS = (σ₀²/(σ₀²+σ²)) · F + ½[ ln((σ₀²+σ²)/σ²) + σ²/(σ₀²+σ²) − 1 ]Paper: Eq. 15, with F from Eq. 13. The uniform-noise version (§3.2) is numerical.
  • IG = KLD + BS — total information gain: the paper's proposed valence signal for epistemic emotions such as curiosity and interest. Under uniform noise it becomes the finite, inverted-U-shaped arousal curve, peaking at SIG.IG = KLD + BSPaper: Eq. 12; the inverted-U with uniform noise is the paper's Figure 4 (§3.2).

δ and F are related by a fixed, monotonic substitution (F = δ²/2(σ₀²+σ²)), so re-plotting against F rather than δ doesn't change which values pair up — it only changes the horizontal spacing. Under pure Gaussian (§3.1), that substitution makes KLD and BS come out exactly linear in F, with slopes σ₀²/σ² and σ₀²/(σ₀²+σ²) respectively — so IG is a third straight line, their sum, and nothing here has a peak yet. Switching on the uniform-noise correction (§3.2) is what bends all three into the finite, inverted-U-shaped curves from the paper's Figure 4: IG rises, peaks at the optimal arousal level S_IG, then falls back toward zero as δ grows large enough that the observation stops being informative about the state at all.