The paper treats learning from a single observation as free energy minimization carried out in two sequential steps, each releasing its own information gain. Step 1 — recognition: before the observation, the agent's best guess about the hidden state s is just the prior p(s). Recognizing the observation updates this guess to the true posterior p(s|o); the free energy released by that update equals KLD = D_KL[p(s) ‖ p(s|o)] — the first information gain, measuring how much was learned by recognizing this particular observation right now. Step 2 — prior update (learning): that posterior is then adopted as the new prior, ready for the next observation. The free energy this second step is expected to release is bounded by BS = D_KL[p(s|o) ‖ p(s)] — the second information gain, measuring the novelty of the observation relative to what was already believed. Because KL divergence isn't symmetric, KLD and BS compare the same two distributions in opposite directions and genuinely differ; the paper sums them into a single total, IG = KLD + BS, and proposes that as the valence signal behind epistemic emotions like curiosity and interest across the Wundt curve.
All three densities sit on the same state axis, so their overlap is visible directly: a posterior that lands close to the likelihood and far from the prior means the observation dominated belief; one that barely moved off the prior means the reverse. The δ bracket above the curves makes that gap explicit — everything plotted in (b) and (c) is this one number, δ, run through different functions of the fixed σ₀ and σ.
All four curves are functions of the single number δ, swept from 0 to 20. With η fixed at 0 and o restricted to ≥ 0, δ = o − η never goes negative, so this range has no redundant mirror image. p(o) is bounded in [0,1] (left axis); the three divergence-related quantities are unbounded and can grow without limit under the pure-Gaussian model — exactly the problem §3.2's uniform-noise correction exists to fix. Under §3.1 (pure Gaussian), KLD and BS have the closed forms above; switch on "+ uniform noise" and they're instead computed by numerically integrating the mixture posterior p_ε(s|o) against the prior (Simpson's rule over s), since the paper's own Appendix (Eq. 26–29) has no closed form for that case. The dashed line and dots mark where the current sliders' δ sits on each curve.
δ and F are related by a fixed, monotonic substitution (F = δ²/2(σ₀²+σ²)), so re-plotting against F rather than δ doesn't change which values pair up — it only changes the horizontal spacing. Under pure Gaussian (§3.1), that substitution makes KLD and BS come out exactly linear in F, with slopes σ₀²/σ² and σ₀²/(σ₀²+σ²) respectively — so IG is a third straight line, their sum, and nothing here has a peak yet. Switching on the uniform-noise correction (§3.2) is what bends all three into the finite, inverted-U-shaped curves from the paper's Figure 4: IG rises, peaks at the optimal arousal level S_IG, then falls back toward zero as δ grows large enough that the observation stops being informative about the state at all.