Control¶
Steady-state LQR action selection: the action-side dual of the Kalman filter. The Agent builds one of these for you when you give it a goal; you rarely touch it directly.
Internal — not part of the public API
LQRController is not exported from cpomdp and carries no stability promise; the Agent constructs it for you. It's documented here for the architecture it illustrates — LQR as the fixed-sensor reduction of active inference 1 (ADR-003). Build agents with StateGoal, not this directly.
LQRController ¶
LQRController(
model: LinearGaussianModel,
*,
goal_precision: ArrayLike,
effort_penalty: ArrayLike,
tol: float = 1e-12,
max_iter: int = 1000,
)
Steady-state LQR action selection — the action-side dual of the filter.
Where the Kalman filter front-loads perception (solve the estimation Riccati
once for the steady-state gain K∞, then mean += K∞·prediction_error),
this front-loads action: solve the dual control Riccati once for L∞,
then action = -L∞·(mean − goal). Both gains are data-independent, both are
computed at construction, and together they are LQG (see RESEARCH.md).
The load-bearing claim (ADR-003) is that LQR is active inference here, not a substitute for it. For a fixed linear-Gaussian sensor the covariance recursion is control-independent, so Expected Free Energy's epistemic term is identical for every action and drops out of the argmin; EFE-minimising selection reduces to its pragmatic term, and the pragmatic term under a Gaussian preference is a quadratic cost whose optimum is exactly LQR. The epistemic term only re-enters once sensing depends on the state or action — out of scope for v0.1.
The two cost matrices are named for the preference they encode, not by LQR's
traditional Q/R — those letters already mean the noise covariances on
the model (dynamics_noise/observation_noise), the exact collision ADR-003
warns about. The names are the same across the whole library: an Agent
hands these straight through to its controller.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model
|
LinearGaussianModel
|
The linear-Gaussian model to act in. Must carry a |
required |
goal_precision
|
ArrayLike
|
How sharply the agent prefers the goal, an |
required |
effort_penalty
|
ArrayLike
|
How much action costs, a |
required |
tol
|
float
|
Absolute tolerance on successive cost-to-go iterates; convergence is declared when they stop moving by more than this. |
1e-12
|
max_iter
|
int
|
Iteration cap before the Riccati recursion is declared to have failed to converge. |
1000
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If the model has no |
RuntimeError
|
If the control Riccati does not converge within |
Source code in src/cpomdp/control.py
action ¶
The action that drives the estimated state toward goal.
One matrix-vector product, -L∞·(mean − goal) — all the work was
front-loaded into L∞ at construction, so there is no optimisation in
the loop. The mean − goal shift turns the regulator (which drives its
state to zero) into a controller that drives the state to goal.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
mean
|
ArrayLike
|
The current belief mean — the best estimate of the state,
shape |
required |
goal
|
ArrayLike
|
The state to steer toward, shape |
required |
Returns:
| Type | Description |
|---|---|
Float64[Array, p]
|
The action, shape |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Source code in src/cpomdp/control.py
The finite-horizon schedule¶
In plain terms
A controller steers a system toward a goal. LQRController builds one by solving a
puzzle that asks how hard to push right now, given where the state is and given
forever to reach the goal. It answers by repeating one calculation until the answer
stops changing. That final answer is the steady-state gain, and the controller keeps
only it.
The intermediate answers were never junk. After one repetition the answer is how hard
to push with one step left. After five, with five steps left. finite_horizon_lqr
runs the same calculation a chosen number of times, H, and keeps every answer. The
result is a schedule: the right push with H steps left, then with H − 1, down to
the last one.
The expected-free-energy planner looks a fixed number of steps ahead, applies the
first action and plans again. A planner like that is using the "H steps left" rule
at every step. The "forever" rule is close to it when H is large and visibly
different when H is small. Held against the forever rule, the planner shows a gap
that fades as H grows and looks exactly like a bug. first_gain is the matching
rule, so no such gap is manufactured.
One convention to know. The cost is charged on the state each action arrives at, and nothing is charged after the last action. The planner's pragmatic term does the same accounting, so the two line up without adjustment.
finite_horizon_lqr ¶
finite_horizon_lqr(
model: LinearGaussianModel,
*,
goal_precision: ArrayLike,
effort_penalty: ArrayLike,
horizon: int,
) -> FiniteHorizonLQR
Run the control Riccati recursion backward over horizon steps.
The same Bellman step LQRController iterates to a fixed point, run a declared
number of times from a zero terminal cost and with every gain kept. With j
steps remaining and W = goal_precision + P_{j−1} the cost of what remains::
L_j = (effort_penalty + Bᵀ W B)⁻¹ (Bᵀ W A)
P_j = Aᵀ W A − (Aᵀ W B) L_j
starting from P_0 = 0. At horizon = 1 the gain is the one-step regulator
(effort_penalty + Bᵀ Q B)⁻¹ Bᵀ Q A. The Control page of the API reference
opens the same account in plain terms.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model
|
LinearGaussianModel
|
The linear-Gaussian model to act in. Must carry a |
required |
goal_precision
|
ArrayLike
|
The stage cost on the state, an |
required |
effort_penalty
|
ArrayLike
|
The stage cost on the action, a |
required |
horizon
|
int
|
|
required |
Returns:
| Type | Description |
|---|---|
FiniteHorizonLQR
|
The schedule, with |
FiniteHorizonLQR
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Source code in src/cpomdp/control.py
FiniteHorizonLQR
dataclass
¶
FiniteHorizonLQR(
model: LinearGaussianModel,
goal_precision: Float64[Array, "n n"],
effort_penalty: Float64[Array, "p p"],
gains: Float64[Array, "H p n"],
cost_to_go: Float64[Array, "H+1 n n"],
)
The gain schedule and cost-to-go of an H-step regulator.
The stage cost is charged on the state each action arrives at, and nothing is
charged after the last one: the terminal cost is zero. That is the sum a
receding-horizon planner scores over its lookahead, so first_gain is the gain
such a planner applies at every step, and it is what a comparison against one
has to use. LQRController.gain is its limit as H grows and differs from it
at every finite H, by an amount that shrinks with H and reads as an error
when the horizons are not matched.
The schedule carries the model and the two costs it was built from, so anything
priced against it reads the same Q, R and noise the recursion did.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model
|
LinearGaussianModel
|
The model the schedule regulates. |
required |
goal_precision
|
Float64[Array, 'n n']
|
The stage cost on the state it was built with, |
required |
effort_penalty
|
Float64[Array, 'p p']
|
The stage cost on the action it was built with, |
required |
gains
|
Float64[Array, 'H p n']
|
|
required |
cost_to_go
|
Float64[Array, 'H+1 n n']
|
|
required |
The control bracket¶
In plain terms
Two costs bound every controller that follows a plan. The lower one is what the plan costs when the state is known exactly at every step. The upper one is the best any controller can do when it has to work the state out from readings. The gap between them is the price of not knowing: what the sensor fails to deliver. That gap is the number to report. Either end on its own is a cost in units nothing calibrates.
Under fixed noise the upper end is settled. Take the gains the full-information plan would use and apply them to the filtered estimate instead of the state. That is the certainty-equivalent controller, and the separation principle says nothing that infers the state does better. Its cost is written down two ways that share nothing but the two Riccati recursions: once by propagating how the state and the estimate spread together, once as the floor plus what the estimate's error costs at each step. The two agreeing to machine precision is the signature of the fixed-noise regime, and it is what makes the upper end a closed form rather than a measurement.
An agent's cost is then read as a position inside the bracket. Zero means it did exactly as well as certainty equivalence. One means it did as well as knowing the state. An agent that can move its own sensor to where the readings are sharper can sit above zero, and that is the effect the bracket exists to measure. The bar on the agent's cost, divided by the width, is the floor below which its position cannot be told from zero.
full_information_cost ¶
full_information_cost(
schedule: FiniteHorizonLQR,
*,
goal: ArrayLike,
prior: Belief | None = None,
) -> float
J_lower: the expected cost of the plan when the state is observed exactly.
The floor of the control bracket. No controller that has to infer the state from
readings can do better, since the state itself is the most it could know. With
d the prior mean's offset from the goal, Σ₀ the prior covariance and
Q_w the process noise::
J_lower = dᵀ P_H d + tr(P_H Σ₀) + Σ_{j=0}^{H−1} tr((Q + P_j) Q_w)
The first two terms are the terminal quadratic averaged over where the state starts. The sum is what the process noise adds at each step, priced at the cost of everything that remains from there.
The goal has to be a state the dynamics hold at zero action, A · goal = goal,
since the plan regulates the offset from it and the offset only obeys the same
dynamics when the goal does not drift.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
schedule
|
FiniteHorizonLQR
|
The plan, with the model and costs it was built from. |
required |
goal
|
ArrayLike
|
The state the plan steers toward, shape |
required |
prior
|
Belief | None
|
Where the state starts, as a Gaussian. Defaults to the model's prior. |
None
|
Returns:
| Type | Description |
|---|---|
float
|
The expected cost in the units of |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the model's process noise depends on the state, if |
Source code in src/cpomdp/control.py
kalman_schedule ¶
kalman_schedule(
model: LinearGaussianModel,
horizon: int,
prior: Belief | None = None,
) -> KalmanSchedule
The exact filter's gain and covariance at each of horizon readings.
The covariance recursion the per-step KalmanBackend runs, run ahead of time
and kept, since under fixed noise it never sees a reading.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model
|
LinearGaussianModel
|
The model the filter runs under, with fixed noise. |
required |
horizon
|
int
|
How many readings, at least one. |
required |
prior
|
Belief | None
|
The belief before the first reading. Defaults to the model's prior. |
None
|
Returns:
| Type | Description |
|---|---|
KalmanSchedule
|
The gains and covariances, indexed as the plan indexes its steps. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Source code in src/cpomdp/control.py
KalmanSchedule
dataclass
¶
The exact filter's gains and error covariances over H readings.
Both are data-independent, so they can be written down before a reading exists.
The indexing follows the plan's: action k is chosen from the belief after
k readings, and reading k + 1 arrives after it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
gains
|
Float64[Array, 'H n m']
|
|
required |
covariances
|
Float64[Array, 'H+1 n n']
|
|
required |
closed_loop_cost ¶
closed_loop_cost(
schedule: FiniteHorizonLQR,
filter_gains: ArrayLike,
*,
goal: ArrayLike,
controller_gains: ArrayLike | None = None,
prior: Belief | None = None,
) -> float
The expected cost of acting on a linear estimate, in closed form.
Action k is −L_k · (estimate − goal), chosen from the estimate after k
readings; the first is chosen from the prior mean. The state then moves, is
charged where it lands, and is read, and the estimate folds the reading in with
K_{k+1}. Every map is linear and every disturbance Gaussian, so the joint
second moment of (state, estimate) about the goal propagates exactly and the
cost is a sum of traces against it. Nothing is sampled.
Any data-independent gain sequence prices this way: the exact filter's, a frozen
one's, or a degraded one's. L_k defaults to the plan's own schedule, and can be
replaced to price a gain the plan did not choose, as the steady-state gain applied
at every step of a short plan.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
schedule
|
FiniteHorizonLQR
|
The plan, with the model and costs it was built from. |
required |
filter_gains
|
ArrayLike
|
|
required |
goal
|
ArrayLike
|
The state steered toward, shape |
required |
controller_gains
|
ArrayLike | None
|
|
None
|
prior
|
Belief | None
|
Where the state starts and what the estimate starts as. Defaults to the model's prior. |
None
|
Returns:
| Type | Description |
|---|---|
float
|
The expected cost in the units of |
Raises:
| Type | Description |
|---|---|
ValueError
|
If either noise depends on the state, if |
Source code in src/cpomdp/control.py
352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 | |
CertaintyEquivalentController
dataclass
¶
Acts on the estimate as if it were the state: −L_k · (mean − goal).
The plan's gains are the ones a controller that saw the state would use. Applying
them to a filtered mean instead is certainty equivalence, and under fixed noise
it is optimal: the separation principle says no controller that has to infer the
state does better. expected_cost is J_CE, priced by closed_loop_cost
with the exact filter's gains.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
schedule
|
FiniteHorizonLQR
|
The plan, with the model and costs it was built from. |
required |
goal
|
ArrayLike
|
The state steered toward, shape |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Source code in src/cpomdp/control.py
action ¶
The action at step of the plan from the current estimate.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
step
|
int
|
Which step of the plan this is, from |
required |
mean
|
ArrayLike
|
The belief mean after |
required |
Returns:
| Type | Description |
|---|---|
Float64[Array, p]
|
|
Source code in src/cpomdp/control.py
expected_cost ¶
J_CE: this controller's expected cost with the exact filter.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
prior
|
Belief | None
|
Where the state starts and what the filter starts from. Defaults to the model's prior. |
None
|
Source code in src/cpomdp/control.py
optimal_cost ¶
optimal_cost(
schedule: FiniteHorizonLQR,
*,
goal: ArrayLike,
prior: Belief | None = None,
) -> float
J*: the least any controller that infers the state from readings can pay.
The separation principle prices it without propagating a state. The plan's gains
are optimal whatever the estimate, and acting on an estimate instead of the state
costs, at each step, the estimate's error charged at what an action error costs
from there. With Σ_k the exact filter's error covariance of the estimate action
k is chosen from, and W_k = Q + P_{H−k−1} what remains after the step::
J* = J_lower + Σ_{k=0}^{H−1} tr(L_kᵀ (R + Bᵀ W_k B) L_k Σ_k)
closed_loop_cost with the exact filter's gains reaches the same number by
propagating the joint second moment of state and estimate. The two share no
arithmetic past the two Riccati recursions, so their agreement to machine
precision is the fixed-noise signature: certainty equivalence is optimal, and no
use of the readings beats it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
schedule
|
FiniteHorizonLQR
|
The plan, with the model and costs it was built from. |
required |
goal
|
ArrayLike
|
The state the plan steers toward, shape |
required |
prior
|
Belief | None
|
Where the state starts and what the filter starts from. Defaults to the model's prior. |
None
|
Returns:
| Type | Description |
|---|---|
float
|
The expected cost in the units of |
Raises:
| Type | Description |
|---|---|
ValueError
|
If either noise depends on the state, if |
Source code in src/cpomdp/control.py
ControlBracket
dataclass
¶
The two costs every controller under one plan sits between.
The floor is the plan with the state observed exactly. The ceiling is the best any controller that has to infer the state from readings can do, which under fixed noise is the certainty-equivalent controller with the exact filter. Their difference is what the readings fail to deliver: the price of inference under this plan. That width is the object to report. Either end alone is a number in cost units that nothing calibrates.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
floor
|
float
|
|
required |
ceiling
|
float
|
|
required |
control_bracket ¶
control_bracket(
schedule: FiniteHorizonLQR,
*,
goal: ArrayLike,
prior: Belief | None = None,
) -> ControlBracket
The floor and ceiling of a plan, both in closed form.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
schedule
|
FiniteHorizonLQR
|
The plan, with the model and costs it was built from. |
required |
goal
|
ArrayLike
|
The state the plan steers toward, shape |
required |
prior
|
Belief | None
|
Where the state starts and what the filter starts from. Defaults to the model's prior. |
None
|
Raises:
| Type | Description |
|---|---|
ValueError
|
Whatever |
Source code in src/cpomdp/control.py
control_efficiency ¶
η_ctrl: how much of the price of inference an agent recovers.
(ceiling − agent_cost) / width, with the agent's bar scaled by the width. Zero
for a certainty-equivalent controller under fixed noise. Positive when an agent
knows more than the fixed-noise filter can, as one that steers its own sensor
does. Negative when it pays more than certainty equivalence would. Every term is
scored under the plan's own model, so the number is within-model and needs no
reference.
Both ends of the bracket are closed forms, so the result's bar is the agent's
alone, divided by the width. That bar is the floor below which η_ctrl cannot
be told from zero, and a claim of zero is a claim to it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
bracket
|
ControlBracket
|
The plan's floor and ceiling. |
required |
agent_cost
|
Bounded
|
The agent's expected cost under the same plan, with its bar. |
required |
Returns:
| Type | Description |
|---|---|
Bounded
|
|
Bounded
|
the width, since the agent's cost enters negated; |
Bounded
|
divided by the width. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the bracket's width is not positive, since there is then no price of inference to normalise by. |
Source code in src/cpomdp/control.py
-
Magnus T. Koudahl, Wouter M. Kouw, and Bert de Vries. On epistemics in expected free energy for linear Gaussian state space models. Entropy, 23(12):1565, 2021. URL: https://doi.org/10.3390/e23121565, doi:10.3390/e23121565. ↩