Transition histories provide context as a pendulum controller faces systems with different masses and lengths. One ID-selected recipe is held fixed across all groups.
One complete recipe is selected on ID development and held fixed across ID and four OOD groups. Each method uses three training seeds and ten 200-step episodes per seed and group.
What this setting establishes.
R7 has higher mean return than CaDM in four of five groups. Return and upright success come from the same evaluation trajectories; same-checkpoint probes and logged-action substitutions examine context separately.
Scope of the evidence
Mean return is lower in OOD c2, and advantages are not shared by every seed. Logged-action substitutions test fixed-model use; weak mass probes do not establish recovery of all physical parameters.
Full result record
Measurements and comparisons.
Reported results are kept with their own populations and conditions. Development, held-out and sealed comparisons remain separate.
01 / Reported analysis
Pendulum closed-loop control: one fixed R7 recipe
3 training seeds × 10 episodes per seed × 200 steps per episode, in each group.
Return ↑ and upright success (%) ↑
Pendulum closed-loop control: one fixed R7 recipe
Group
CaDM · return ↑
CaDM · success (%) ↑
Random relation · return ↑
Random relation · success (%) ↑
SPRII (R7) · return ↑
SPRII (R7) · success (%) ↑
ID
−326.5 ± 46.5
96.7
−321.4 ± 33.2
100.0
−323.9 ± 26.2
100.0
OOD_c0
−376.6 ± 63.4
93.3
−346.8 ± 30.6
100.0
−322.2 ± 29.4
100.0
OOD_c1
−387.0 ± 11.9
93.3
−400.1 ± 24.6
93.3
−381.0 ± 13.6
100.0
OOD_c2
−355.5 ± 33.9
90.0
−479.6 ± 126.7
76.7
−370.2 ± 51.8
90.0
OOD_c3
−1100 ± 273.5
20.0
−1185 ± 131.2
6.7
−1060 ± 384.4
30.0
BBold: best reported mean return or success in each group; ties included.
ID is separated from the four OOD groups.
Returns are mean ± sample SD across three seeds. Success requires all final 100 states within 60° of upright. One ID-selected R7 recipe is fixed across groups; mean return is lower than CaDM in OOD c2.
02 / Reported analysis
Same-checkpoint factor accessibility and logged-action context use
ID probes and a common logged-action bank, using the same checkpoints as the control study. The text reports point estimates without uncertainty.