TD-MTP · Proposition 2
TD-MTP's multi-step critic target handles the two ways an episode can end inside a sampled window with two separate gates. This page shows what each gate does to the target and how large the bias is when a truncation is treated as a termination.
A multi-step target is computed on a window of consecutive transitions. A sequential replay buffer stores episodes back to back and samples fixed-length windows at any offset, so some windows cross an episode boundary. An episodic buffer never returns such a window, which is why n-step implementations built on one do not need to handle this case.
TD-MTP builds two gates per transition from the termination flag d and the truncation flag u.
At a termination both gates close, so the target neither bootstraps from the next state nor carries later rewards back. At a truncation only χ closes. The next state still has value and the target bootstraps from it, but the rest of the window belongs to another episode, so its rewards are not carried back.
The target at the front of the window is computed backwards. Each transition adds its reward to a discounted mix of the critic’s estimate and the next target, and the gates determine which terms reach the front.
If the time limit is recorded as a termination, ν closes at the boundary when it should stay open. The bootstrap is dropped and the target loses the value of the successor state.
The loss is still a regression onto a target of roughly the right size, the gradients stay finite and the training curve still rises. The critic learns that some states are worth far less than they are, but only at positions near a boundary, which are a minority of the batch.
The error is proportional to how often windows cross a boundary. Short episodes and long windows make crossings common. An episodic buffer rules them out, so implementations built on one do not show the error.
Proposition 2 writes the backward recursion as a single sum. Each reward term carries a product of propagation gates, so one closed gate zeroes every later reward. Each bootstrap term carries a ν, which a termination sets to zero and a truncation does not.
The sum is an exact rewriting of the backward recursion: the two give the same number for every gate pattern and every λ. It is checked against 20,000 random windows at every anchor position.
At λ = 0, every gate assignment reduces termwise to r + γνq, the reference 1-step target, which the λ slider shows at zero. The baseline arm of every ablation is therefore the reference agent itself rather than a reimplementation.
The sum is not the textbook λ-return. With gates present the mixture weights include the gate products, and (1−λχ) replaces (1−λ). The two agree only when no gate closes inside the window.
The proposition describes what the recursion computes and says nothing about whether multi-step targets help, which is a separate result.
Rewards and critic estimates on this page are illustrative and sized for a locomotion task; the discount is the reference agent’s γ = 0.97. All numbers are computed live from the recursion, and the repository re-checks the forward-view identity against 20,000 random windows.