TD-MTP · paired 2×2 study
With a 1-step critic target, replacing the planner’s single Gaussian proposal with the tensor planner lowers final return on nine of ten task blocks. When the critic is trained on a TD(λ) target instead, with the planner, schedule and seeds unchanged, the tensor planner improves return on five of ten blocks.
Each task is trained under four conditions: with and without the tensor planner, crossed with a 1-step target and a TD(λ) target. Both planners in a block share their model initialization, and all conditions are evaluated with the same search, so differences come from the collected data and the training target.
Each row is one task, and the mark shows the planner’s effect on final return under the selected target.
Clicking a row in the chart above shows the four training runs for that task. Values are mean final return over the last three evaluations, averaged over seeds.
Two blocks, stair and pole, supply about four-fifths of the summed interaction.
The study measures an interaction on a paired design. Both planners in a block share their model initialization and evaluation search is identical across conditions, so comparisons are within-block.
The sign tests are descriptive counts for a comparison chosen after seeing the results. They are not confirmatory tests and say nothing about effect size.
The study is underpowered as a planner comparison. Cross-seed standard deviation on the planner margin is 55–77 return units, while the effects of interest are 20–40; detecting them would need 16–75 seeds per condition, and these runs have 3–5.
Tensor candidates are active only in the first half of training, so the endpoint score measures their downstream effect through the replay buffer and the learned agent rather than the search itself.
The largest block, stair, reverses when training is continued to twice the budget at a different dose of λ, and we do not have an explanation for the reversal.
The numbers on this page, including the ten block effects, both sign-test values and the four-fifths share, are recomputed from the paper’s ablation table. Effects are differences of means over the last three evaluations of each run, averaged over 3–5 seeds.