TD-MTP · Proposition 2
The structured planner builds a candidate by picking one control point per layer and interpolating between them. Whether a candidate can land in a given region of action space depends on which control points the graph holds, and drawing more paths does not change that.
01 · how a candidate is built
The graph holds M = 3 layers of N = 16 control points. A path picks an index independently in each layer and the interpolation matrix Φ turns those points into H = 3 actions. Every row of Φ is non-negative and sums to one, so each action is a weighted average of the points the path chose.
02 · the event
Take a box A in action space and consider sequences whose actions all lie inside it. If every control point a path selects is in A, every interpolated action is in A as well, because each action is a convex combination of those points and a box is convex.
Layers are chosen independently, so the probability of this sufficient event is the product of the per-layer fractions.
03 · interactive bound
Set how many of each layer’s sixteen points fall inside the region and how many paths the planner draws this iteration. The bound is a lower bound on the probability of at least one hit. The simulation draws that many paths from a graph with this arrangement and counts how many land inside.
04 · the case H = M = 3
With H = M = 3 the only actions produced are U1, the midpoint, and U2. If both endpoints are inside the box, the midpoint is too. If either endpoint is outside, it is itself one of the actions, so the sequence is already outside. The sufficient event is therefore also necessary, and the inequality holds with equality.
This does not hold in general. If a third layer has some weight, a path can enter the region even when the point it picks in one active layer lies outside, because a small weight on that point can be offset by the others. The product then underestimates the probability, and when any active qm is zero the bound is zero even though some paths reach the region.
05 · computing qm
The fractions count whole control points. Computing a fraction for each coordinate and multiplying them gives the wrong value, as the two-point example below shows.
06 · scope
Numbers and statements follow the manuscript’s Appendix B, checked against the 26 September draft. The simulation samples from a grid instead of replaying stored results, so the counts change on every resample. Under review at ICLR 2027; no code is released.