Interpreting the evidenceWhat stays fixed in the test.
The task uses the same support/query access and a frozen official perception frontend; learner-specific objectives are preserved.
What this setting establishes.
On 1,994 held-out episodes, Cross improves over Structure and Random in every fitted CPC and RSSM seed.
Scope of the evidenceThis is the adapted S3/query3 prediction task. This is a held-out episode test, not a verified unseen-parameter split or the original CoPhy counterfactual protocol. The displayed pair is newly simulated in PyBullet under shared physical parameters. It is an environment demonstration, not an evaluated episode or model prediction.
Full result recordMeasurements and comparisons.
Reported results are kept with their own populations and conditions. Development, held-out and sealed comparisons remain separate.
01 / Study Table 1Held-out trajectory prediction
1,994 held-out episodes; S3/query3; 3 fitted pairs per method
Physical trajectory MSE ↓
BBold marks the Cross result within each separately evaluated learner family.
- CPCCompare methods within a learner row; this is the adapted S3/query3 held-out task.
Means ± sample SD across three jointly fitted source/reader seeds per arm. The same S3/query3 inputs are used throughout. This is a held-out episode test under the adapted protocol, not a verified unseen-parameter split.
02 / Study Table 7Collision held-out comparison — A. Matched four-arm coverage: three seeds per arm
Same 1,994 held-out episodes; 3 CPC/RSSM seeds per arm; JEPA J2: 3 Cross seeds and 1 per control
Trajectory MSE ↓; Cross error reduction ↑
BBold marks the Cross result within each three-seed learner comparison.
Means ± sample SD across three joint source/reader seeds per arm on the same 1,994 episodes. Every CPC and RSSM seed favors Cross over Structure and Random. The adapted task is not a verified unseen-parameter split.
03 / Study Table 7Collision held-out comparison — B. JEPA J2: three SPRII seeds, one fitted seed per control
Same 1,994 held-out episodes; 3 CPC/RSSM seeds per arm; JEPA J2: 3 Cross seeds and 1 per control
Trajectory MSE ↓; Cross error reduction ↑
BBold marks the three-seed JEPA J2 result.
- JEPAEach control has one fitted seed. This follow-up does not have the balanced seed coverage of CPC and RSSM.
JEPA J2 is a development-selected, matched-only retest: three Cross fits versus one fit per control. For fixed seed-0 models, the Cross-minus-Structure paired-recipient 95% interval is [−0.045874,−0.020480]; it is not a training-seed interval.
04 / Study Table 7Collision held-out comparison — C. Cross mean-error reductions
Same 1,994 held-out episodes; 3 CPC/RSSM seeds per arm; JEPA J2: 3 Cross seeds and 1 per control
Trajectory MSE ↓; Cross error reduction ↑
BBold marks reductions from the balanced three-seed CPC and RSSM comparisons.
The dashed rule separates the unequal-coverage JEPA follow-up.
Reductions compare equally weighted mean errors. CPC and RSSM have three fitted seeds per arm; JEPA J2 has three Cross fits and one per control, so its coverage is not directly equivalent.
05 / Study Table 35Full CoPhy learner × scene matrix
S3/query3 development; source100/head100; source coverage explicit
Trajectory MSE ↓; support and reader-memory probe R² where reported
BBold marks the lowest reported mean only within each scene × learner row.
Dashed rules separate scenes and the supervised CoPhyNet block.
- Align + Cross: supervised CoPhyNet objective · BallsA dash is an unevaluated arm. Development source 0 is not the held-out Collision benchmark.
The full source-0 development matrix retains favorable and adverse comparisons. SPRII uses Cross for JEPA/CPC/RSSM and Align + Cross for CoPhyNet. A dash is an unevaluated arm. These validation cohorts differ from the held-out Collision test.
06 / Study Table 36Full CoPhy learner × scene matrix
S3/query3 development; source100/head100; source coverage explicit
Trajectory MSE ↓; support and reader-memory probe R² where reported
BBold identifies Cross prediction errors, including adverse cells; it is not a winner mark.
Dashed rules separate learner × scene groups.
- JEPA / Balls · NativeSource columns remain separate. Support R² and reader-memory R² are accessibility assays, not error metrics.
Source coverage is explicit; dashes denote unevaluated cells. Support and reader-memory probes target restitution in Balls, mass in Collision, and vertical gravity in Blocktower. These development cohorts include selection examples and remain separate from the held-out test.
07 / Study Table 36RSSM Collision development replication summary
S3/query3 development; source100/head100; 3 sources
Trajectory MSE ↓; mass probe R²
BBold marks Cross prediction MSE in the three-source development comparison.
- CrossThe neighboring R² columns measure accessibility; this population is separate from the held-out test.
MSE is mean ± sample SD across three RSSM sources; support and reader-memory R² probe mass. These S3/query3 development data include selection examples and are not the held-out Collision test population.
08 / Study Table 37Frozen-source reader design across scenes
Source-50/head-100; source seed 0; full validation
Relative error reduction versus Native (%) ↑
BBold marks sign reversals across reader recipes in Balls and Blocktower.
- Main initialization, dropout 0.1, initialPositive is an error reduction versus Native. Each row is one shared recipe across scenes; the best cell per scene is not a common selected model.
Every row is one shared reader recipe across all three scenes, using the same frozen source-50 checkpoint. Positive is lower error than Native. The sign varies by scene and recipe; separately selecting each scene’s best cell would change the comparison.
09 / Study Table 39Complete follow-up outcomes
Separate configurations and populations shown row by row
MSE ↓, except FHN relative L₂ ↓
BBold identifies each follow-up configuration; heterogeneous numerical cells are deliberately not ranked.
Dashed rules separate experiments with different populations and fitted-model coverage.
All completed follow-ups from the configuration round are retained, including null and adverse results. Each row has its own population and model coverage. Paired intervals condition on the reported fitted models; the one-source development rows do not replace the original multi-source comparisons.