Full result recordMeasurements and comparisons.
Reported results are kept with their own populations and conditions. Development, held-out and sealed comparisons remain separate.
01 / Study Table 24More distinct histories through a fixed reader
S3-trained RSSM; mean of 3 retained fits
Trajectory MSE ↓
BBold compares three versus eight distinct histories through the same SPRII reader.
- SPRII (Cross)The reader was trained with S3 and is not refitted for S5/S8; this is an input substitution.
MSE averages three retained fits; Repeat averages individual-history repetitions. S3, S5 and S8 change the input history count through the same S3-trained RSSM reader. No retraining is performed, so this is not a sample-efficiency comparison.
02 / Study Table 25Reader-memory accessibility and donor physics — A. CoPhyNet reader-memory probes (R^2)
Probe: 26,558 training / 7,578 validation objects; 3 sources; separate source100 donor assays
Probe R²; same/wrong-physics distance ratio; wrong-factor-minus-matched MSE
BBold marks restitution readout under each recipe, keeping the three sources separate.
Dashed rules separate training recipes.
- SPRII (A+C) · RestitutionMass and friction remain in the table. Readout accessibility is not a direct measurement of predictor reliance.
Training-only ridge probes inspect CoPhyNet’s actual reader memory, including its learned projection, using 26,558 training and 7,578 validation objects. Sources remain separate. Strong restitution accessibility is distinct from fixed-predictor donor reliance.
03 / Study Table 25Reader-memory accessibility and donor physics — B. RSSM within/between geometry
Probe: 26,558 training / 7,578 validation objects; 3 sources; separate source100 donor assays
Probe R²; same/wrong-physics distance ratio; wrong-factor-minus-matched MSE
BBold marks the SPRII within/between ratios across all three sources.
- SPRII (Cross)This is same-physics divided by wrong-physics distance. A lower ratio is not by itself a task benefit.
The ratio divides same-physics code distance by wrong-physics code distance for each RSSM source separately. These geometry measurements do not establish task value; donor interventions provide a different test of the fitted reader.
04 / Study Table 25Reader-memory accessibility and donor physics — C. Separate source100 donor assay
Probe: 26,558 training / 7,578 validation objects; 3 sources; separate source100 donor assays
Probe R²; same/wrong-physics distance ratio; wrong-factor-minus-matched MSE
BBold marks the restitution donor penalties, the large reported factor-specific effects.
- JEPAThese are separate single-source assays. Small or negative effects for other factors do not establish physical irrelevance.
These source100 donor assays are separate single-source evaluations, not three-source replications. Each delta is wrong-factor minus matched MSE through a fixed prediction route. Small or negative effects do not establish that a factor is physically irrelevant.
05 / Study Table 35Full CoPhy learner × scene matrix
S3/query3 development; source100/head100; source coverage explicit
Trajectory MSE ↓; support and reader-memory probe R² where reported
BBold marks the lowest reported mean only within each scene × learner row.
Dashed rules separate scenes and the supervised CoPhyNet block.
- Align + Cross: supervised CoPhyNet objective · BallsA dash is an unevaluated arm. Development source 0 is not the held-out Collision benchmark.
The full source-0 development matrix retains favorable and adverse comparisons. SPRII uses Cross for JEPA/CPC/RSSM and Align + Cross for CoPhyNet. A dash is an unevaluated arm. These validation cohorts differ from the held-out Collision test.
06 / Study Table 36Full CoPhy learner × scene matrix
S3/query3 development; source100/head100; source coverage explicit
Trajectory MSE ↓; support and reader-memory probe R² where reported
BBold identifies Cross prediction errors, including adverse cells; it is not a winner mark.
Dashed rules separate learner × scene groups.
- JEPA / Balls · NativeSource columns remain separate. Support R² and reader-memory R² are accessibility assays, not error metrics.
Source coverage is explicit; dashes denote unevaluated cells. Support and reader-memory probes target restitution in Balls, mass in Collision, and vertical gravity in Blocktower. These development cohorts include selection examples and remain separate from the held-out test.
07 / Study Table 37Frozen-source reader design across scenes
Source-50/head-100; source seed 0; full validation
Relative error reduction versus Native (%) ↑
BBold marks sign reversals across reader recipes in Balls and Blocktower.
- Main initialization, dropout 0.1, initialPositive is an error reduction versus Native. Each row is one shared recipe across scenes; the best cell per scene is not a common selected model.
Every row is one shared reader recipe across all three scenes, using the same frozen source-50 checkpoint. Positive is lower error than Native. The sign varies by scene and recipe; separately selecting each scene’s best cell would change the comparison.
08 / Study Table 38Balls prediction and paired method contrasts — Absolute matched, null and wrong-donor error
2,000 development recipients; 3 fitted sources per learner
Matched/Null/Wrong MSE; comparator-minus-SPRII gap ↑
BBold marks matched-context SPRII risk within each learner family.
The dashed rule separates RSSM and CoPhyNet.
- RSSM · SPRII (C)Null and Wrong are fixed context conditions in this record; comparing their sizes does not rank context quality independently of the reader.
Matched, Null and Wrong risks share the 2,000-recipient development cohort within each learner family and average three fixed sources. RSSM uses Cross; CoPhyNet uses Align + Cross. These context conditions measure different uses of the fitted prediction route.
09 / Study Table 38Balls prediction and paired method contrasts — Adjusted paired method contrasts
2,000 development recipients; 3 fitted sources per learner
Matched/Null/Wrong MSE; comparator-minus-SPRII gap ↑
BBold marks the five prespecified comparator-minus-SPRII differences.
The dashed rule separates learner families.
- RSSM · NativeThe intervals shown are Bonferroni-adjusted across all five contrasts and condition on the fitted models.
Comparator-minus-SPRII differences are paired before averaging sources. Recipient-bootstrap intervals are Bonferroni-adjusted across all five contrasts. They condition on fitted models and exclude model-selection uncertainty.
10 / Study Table 39Complete follow-up outcomes
Separate configurations and populations shown row by row
MSE ↓, except FHN relative L₂ ↓
BBold identifies each follow-up configuration; heterogeneous numerical cells are deliberately not ranked.
Dashed rules separate experiments with different populations and fitted-model coverage.
All completed follow-ups from the configuration round are retained, including null and adverse results. Each row has its own population and model coverage. Paired intervals condition on the reported fitted models; the one-source development rows do not replace the original multi-source comparisons.