skip to primary navigationskip to content

Department of Computer Science and Technology

Courses 2026–27

 

Course pages 2026–27 (working draft)

Information, Energy and Intelligence

Principal lecturer: Prof Neil Lawrence
Taken by: MPhil ACS, Part III
Code: L172
Term: Michaelmas
Hours: 16 (8 x 2hrs lectures)
Class limit: max. 10 students
Prerequisites: Students should have an undergraduate-level understanding of probability and statistics, including probability distributions, expectation, discrete and continuous random variables, and Bayes theorem. Knowledge of linear algebra, including matrix operations, eigenvalues and eigenvectors, is required. Familiarity with basic multivariate calculus, including partial derivatives and Lagrange multipliers, is helpful. No prior knowledge of physics or thermodynamics is assumed, as all thermodynamic concepts are developed from first principles during the module.
timetable

Aims

A century ago, many patents were submitted proposing perpetual motion. We know why that was impossible: the second law of thermodynamics. Today, billions are being invested in promises of superintelligence. But what are the limits?  This module builds mathematical machinery needed to answer that question rigorously. Entropy appears in three apparently separate traditions — thermodynamics (Boltzmann, Gibbs), information theory (Shannon), and Bayesian inference (Jaynes) — and turns out to be the same mathematical object viewed from different operational assumptions. Information geometry (Amari) provides the unifying geometric language.  

Syllabus

1 | **Motivation; Boltzmann distribution; free energy decomposition** — The perpetual motion framing. Partition function Z; internal energy U; entropy S; Helmholtz free energy F = U − TS. Why "available energy" is what a system can do work with. | Lecture + motivating examples; formative LLM exercise set

 2 | **Shannon entropy and its equivalence to Boltzmann entropy** — H = −Σ pᵢ log pᵢ; Boltzmann S = kH. Three framings (information, thermodynamic, Bayesian) introduced as distinct operational perspectives on the same quantity. Partition function as generating function. | Lecture + derivation; formative LLM exercise set

  3 | **Maxwell's demon; Landauer's principle; intelligence and entropy** — The demon as an intelligent agent that appears to violate the second law. Landauer (1961): erasing one bit costs k_B T ln 2. The first rigorous link between decision-making and thermodynamic cost. | Lecture + thought-experiment workshop; formative LLM exercise set |  

4 | **Jaynes' maximum entropy principle; exponential family; synthesis** — MaxEnt with Lagrange multipliers; recovers canonical ensemble and Gaussian under appropriate constraints. Exponential family p(x|θ) = exp(θ·T(x) − A(θ)) as the MaxEnt family. Synthesis: comparing all three perspectives on entropy. | Lecture; formative LLM exercise set |  

 5 | **Information geometry: Fisher metric and dually flat geometry** — Manifold of probability distributions; Fisher information matrix as Riemannian metric; e-flat and m-flat coordinate systems; Pythagorean theorem for KL divergence. | Lecture + worked geometric examples; formative LLM exercise set | 

 6 | **Information geometry: MaxEnt as projection; natural gradient** — MaxEnt as m-projection onto a constraint surface. Natural gradient F⁻¹∇L: removes arbitrary dependence on parameterisation. Comparison with vanilla gradient descent. | Lecture + coding demonstration; formative LLM exercise set |  

7 | **Multi-information; I + H = C; von Neumann entropy** — Watanabe (1960): multi-information I = Σhᵢ − H ≥ 0; conservation law I + H = C. Classical limit I = C forces H = 0 (determinism). Quantum resolution: pure entangled state has S(ρ) = 0 but positive marginal entropies. | Lecture; formative LLM exercise set |  

8 | **Probability transport; Schrödinger bridges; limits on intelligence** — Intelligent agency abstracted as transport of probability mass. Optimal transport (Wasserstein distance). Schrödinger bridge as maximum-entropy interpolation between distributions. Closing the loop: evaluating superintelligence claims using the inaccessible game. | Lecture + synthesis discussion; in-class Quiz 4 | 

Please note: Each session is followed by a formative (unassessed) exploration exercise in which students explore the session's central concept from three disciplinary perspectives — thermodynamic, information-theoretic, and Bayesian — and write a short critical synthesis.  

 

This is the principal mechanism for building the multi-perspective fluency required; it is not separately timetabled but constitutes a significant proportion of the self-study hours. 

Objectives:

  • LO1 | Decompose the Boltzmann distribution into contributions from internal energy, entropy, and Helmholtz free energy, and interpret each term's physical and informational meaning

  • LO2 | Derive Shannon entropy as a measure of uncertainty and prove its formal equivalence to thermodynamic (Gibbs) entropy in equilibrium statistical mechanics

  • LO3 | Derive the canonical ensemble and use the partition function to compute mean energy, entropy, and free energy for simple systems

  • LO4 | Explain Maxwell's demon thought experiment, identify where the apparent second-law violation arises, and apply Landauer's principle to show that information erasure restores thermodynamic consistency

  • LO5 | Apply Jaynes' maximum entropy principle, using Lagrange multipliers, to derive the least-committal probability distribution consistent with a set of moment constraints

  • LO6 | Identify the exponential family as the MaxEnt family and explain why the canonical ensemble, Gaussian, and Bernoulli distributions all belong to it

  • LO7 | Compare and contrast the information-theoretic, thermodynamic, and Bayesian perspectives on entropy, articulating the distinct operational assumptions each makes and what each illuminates or obscures

  • LO8 | Describe the manifold of probability distributions as a Riemannian space, define the Fisher information matrix as its metric, and explain the dually flat geometry of exponential families including the Pythagorean theorem for KL divergence  

  • LO9 | Apply information geometry to interpret maximum entropy inference as a projection onto a constraint manifold, and explain why natural gradient descent is the geometrically correct gradient for statistical models

  • LO10 | Define multi-information I, state the conservation law I + H = C, and explain the analogy between this structure and the kinetic/potential energy trade-off in classical mechanics

| LO11 | Explain why the classical limit I = C requires H = 0, argue that sustaining high marginal entropies alongside I = C forces a move beyond classical probability, and show that von Neumann entropy S(ρ) = −Tr(ρ log ρ) satisfies S = 0 for a pure entangled state while its marginals carry positive entropy | Analyse | 7 |  

| LO12 | Articulate how the movement of probability mass between distributions provides an abstraction of intelligent agency — analogous to Shannon's use of probability to abstract a communication code — and sketch the role of optimal transport and Schrödinger bridges in formalising this abstraction | Evaluate | 8 |  

| LO13 | Evaluate claims about the capabilities of intelligent systems using Landauer's principle, the perpetual motion analogy, and the information-theoretic constraints derived from the inaccessible game framework | Evaluate | 8 | 

Recommended Reading

Primary references (all covered in lectures):*

  •  Shannon, C.E. (1948). "A Mathematical Theory of Communication." *Bell System Technical Journal*, 27, 379–423. 
  •  Jaynes, E.T. (1957). "Information Theory and Statistical Mechanics." *Physical Review*, 106(4), 620–630.  
  • Landauer, R. (1961). "Irreversibility and Heat Generation in the Computing Process." *IBM Journal of Research and Development*, 5(3), 183–191. 
  •  Watanabe, S. (1960). "Information Theoretical Analysis of Multivariate Correlation." *IBM Journal of Research and Development*, 4(1), 66–82. 
  •  Amari, S. & Nagaoka, H. (2000). *Methods of Information Geometry*. AMS/Oxford University Press. *(Chapters 1–3)*  
  • Lawrence, N.D. (2025). "The Inaccessible Game." *(Course notes; distributed via Moodle)*  *Supporting textbooks:*  
  •  Cover, T.M. & Thomas, J.A. (2006). *Elements of Information Theory* (2nd ed.). Wiley. *(Background reference for Weeks 1–4)*  
  • MacKay, D.J.C. (2003). *Information Theory, Inference, and Learning Algorithms*. Cambridge University Press. *(Freely available online; relevant chapters indicated per session)* 

Assessment

The module is assessed as four take-home worksheets and four short in-class Moodle quizzes.

  1. Worksheets (60% total, 15% each):  Each worksheet consists of a short Python notebook and a written reflection (300–500 words).  
  • Worksheet 1: Thermodynamics and Shannon entropy | End of Week 2 | Boltzmann sampling; empirical vs analytic entropy; free energy decomposition | LO1, LO2, LO3
  • Worksheet 2: Maxwell's demon, MaxEnt, and the exponential family | End of Week 4 | MaxEnt with Lagrange multipliers; verify canonical ensemble and Gaussian; Landauer's principle | LO4, LO5, LO6, LO7
  • Worksheet 3: Information geometry | End of Week 6 | Fisher information matrix; one step of natural gradient descent vs vanilla gradient descent | LO8, LO9
  • Worksheet 4: Multi-information, von Neumann entropy, and limits on intelligence | End of Week 8 | Multi-information; I + H = C conservation; Schrödinger bridge sketch; evaluate superintelligence claims | LO10, LO11, LO12, LO13 

  1. In-class Moodle quizzes (40% total, 10% each):
  •   Each quiz is administered at the start of the relevant lecture (10 minutes, 8–10 MCQ).