SPRIIA study of persistent information
Value · Use · Generality

Pendulum

From dynamics context to online control.

Transition histories provide context as a pendulum controller faces systems with different masses and lengths. One ID-selected recipe is held fixed across all groups.

4 / 5 groupshigher mean return than CaDM
Motion preview
Torque demonstration · PendulumDynamics: CaDM · render geometry: OpenAI Gym · torque demonstration. Source & changes
Across independent experiences

What persists.

Mass and length within the same system.

Across distinct realizations

What changes.

Episode histories; mass and length across systems.

Experience → prediction

The task.

Transition history + current state.

→ Learned dynamics for online control.

Evaluated instance

The comparison.

CaDM interface · SPRII R7 (relation weight .003; variance target .1).

Controls. CaDM; Random relation.

Interpreting the evidence

What stays fixed in the test.

One complete recipe is selected on ID development and held fixed across ID and four OOD groups. Each method uses three training seeds and ten 200-step episodes per seed and group.

What this setting establishes.

R7 has higher mean return than CaDM in four of five groups. Return and upright success come from the same evaluation trajectories; same-checkpoint probes and logged-action substitutions examine context separately.

Scope of the evidence

Mean return is lower in OOD c2, and advantages are not shared by every seed. Logged-action substitutions test fixed-model use; weak mass probes do not establish recovery of all physical parameters.

Full result record

Measurements and comparisons.

Reported results are kept with their own populations and conditions. Development, held-out and sealed comparisons remain separate.

01 / Reported analysis

Pendulum closed-loop control: one fixed R7 recipe

3 training seeds × 10 episodes per seed × 200 steps per episode, in each group.

Return ↑ and upright success (%) ↑

Pendulum closed-loop control: one fixed R7 recipe
GroupCaDM · return ↑CaDM · success (%) ↑Random relation · return ↑Random relation · success (%) ↑SPRII (R7) · return ↑SPRII (R7) · success (%) ↑
ID−326.5 ± 46.596.7−321.4 ± 33.2100.0−323.9 ± 26.2100.0
OOD_c0−376.6 ± 63.493.3−346.8 ± 30.6100.0−322.2 ± 29.4100.0
OOD_c1−387.0 ± 11.993.3−400.1 ± 24.693.3−381.0 ± 13.6100.0
OOD_c2−355.5 ± 33.990.0−479.6 ± 126.776.7−370.2 ± 51.890.0
OOD_c3−1100 ± 273.520.0−1185 ± 131.26.7−1060 ± 384.430.0

Bold: best reported mean return or success in each group; ties included.

ID is separated from the four OOD groups.

Returns are mean ± sample SD across three seeds. Success requires all final 100 states within 60° of upright. One ID-selected R7 recipe is fixed across groups; mean return is lower than CaDM in OOD c2.

02 / Reported analysis

Same-checkpoint factor accessibility and logged-action context use

ID probes and a common logged-action bank, using the same checkpoints as the control study. The text reports point estimates without uncertainty.

Mass / length probe R² ↑; own-donor H10 MSE ↓; wrong-minus-matched MSE

Same-checkpoint factor accessibility and logged-action context use
MethodMass probe R² ↑Length probe R² ↑Own-donor H10 MSE ↓Wrong − matched MSE
CaDM0.0270.6700.01810.1132
Random relation0.0120.6830.02260.0956
SPRII (R7)0.0500.6690.01900.1018

Bold: best point estimate for each directed metric. Wrong−matched MSE is reported without ranking.

ID point estimates from the same checkpoints. Logged-action substitutions test fixed-model use, separately from online control; mass R² remains low.