TD-MTP · Proposition 2

Termination and Truncation in the Multi-Step Target

TD-MTP's multi-step critic target handles the two ways an episode can end inside a sampled window with two separate gates. This page shows what each gate does to the target and how large the bias is when a truncation is treated as a termination.

01

Windows that cross an episode boundary

A multi-step target is computed on a window of consecutive transitions. A sequential replay buffer stores episodes back to back and samples fixed-length windows at any offset, so some windows cross an episode boundary. An episodic buffer never returns such a window, which is why n-step implementations built on one do not need to handle this case.

Three sampled windows of the same length. The middle one starts in episode A and ends in episode B.
02

Two ways an episode ends

TD-MTP builds two gates per transition from the termination flag d and the truncation flag u.

In the buffer both appear as the end of an episode. After a termination the value of the next state is zero; after a truncation it is not.
bootstrap gate ν
1 − d
closes at a termination
propagation gate χ
(1−d)(1−u)
closes at a termination or a truncation

At a termination both gates close, so the target neither bootstraps from the next state nor carries later rewards back. At a truncation only χ closes. The next state still has value and the target bootstraps from it, but the rest of the window belongs to another episode, so its rewards are not carried back.

Weight on collected rewards relative to the critic’s estimate. At 0 the target is the reference 1-step target.
Where the episode ends inside the window.
Truncation is a time limit; termination is a terminal state.
03

How the gates enter the backward recursion

The target at the front of the window is computed backwards. Each transition adds its reward to a discounted mix of the critic’s estimate and the next target, and the gates determine which terms reach the front.

04

Treating a truncation as a termination

If the time limit is recorded as a termination, ν closes at the boundary when it should stay open. The bootstrap is dropped and the target loses the value of the successor state.


Why training does not reveal the error

The loss is still a regression onto a target of roughly the right size, the gradients stay finite and the training curve still rises. The critic learns that some states are worth far less than they are, but only at positions near a boundary, which are a minority of the batch.

The error is proportional to how often windows cross a boundary. Short episodes and long windows make crossings common. An episodic buffer rules them out, so implementations built on one do not show the error.

05

Proposition 2: the target as a sum of gated terms

Proposition 2 writes the backward recursion as a single sum. Each reward term carries a product of propagation gates, so one closed gate zeroes every later reward. Each bootstrap term carries a ν, which a termination sets to zero and a truncation does not.

claims

The sum is an exact rewriting of the backward recursion: the two give the same number for every gate pattern and every λ. It is checked against 20,000 random windows at every anchor position.

claims

At λ = 0, every gate assignment reduces termwise to r + γνq, the reference 1-step target, which the λ slider shows at zero. The baseline arm of every ablation is therefore the reference agent itself rather than a reimplementation.

not

The sum is not the textbook λ-return. With gates present the mixture weights include the gate products, and (1−λχ) replaces (1−λ). The two agree only when no gate closes inside the window.

not

The proposition describes what the recursion computes and says nothing about whether multi-step targets help, which is a separate result.


Rewards and critic estimates on this page are illustrative and sized for a locomotion task; the discount is the reference agent’s γ = 0.97. All numbers are computed live from the recursion, and the repository re-checks the forward-view identity against 20,000 random windows.