SPRIIA study of persistent information
Prediction · Generality

CoPhy · Collision

Experience before the next collision.

Related support interactions provide context for predicting how objects move after contact, across two different base learners.

CPC + RSSMheld-out prediction gains
Motion preview
Physics demonstration · initial velocity and collision · ½ speedCoPhy: Fabien Baradel et al. · reconstructed scene. Source & changes
Across independent experiences

What persists.

Matched objects under the scene-specific relation.

Across distinct realizations

What changes.

Independent support interactions.

Experience → prediction

The task.

Three support histories + three query frames.

→ The next twelve trajectory frames.

Evaluated instance

The comparison.

CPC / RSSM · Cross. Other learners appear in the extension matrix.

Controls. Native; Structure; Random.

Interpreting the evidence

What stays fixed in the test.

The task uses the same support/query access and a frozen official perception frontend; learner-specific objectives are preserved.

What this setting establishes.

On 1,994 held-out episodes, Cross improves over Structure and Random in every fitted CPC and RSSM seed.

Scope of the evidence

This is the adapted S3/query3 prediction task. This is a held-out episode test, not a verified unseen-parameter split or the original CoPhy counterfactual protocol. The displayed pair is newly simulated in PyBullet under shared physical parameters. It is an environment demonstration, not an evaluated episode or model prediction.

Full result record

Measurements and comparisons.

Reported results are kept with their own populations and conditions. Development, held-out and sealed comparisons remain separate.

01 / Study Table 1

Held-out trajectory prediction

1,994 held-out episodes; S3/query3; 3 fitted pairs per method

Physical trajectory MSE ↓

Held-out trajectory prediction
LearnerNativeStructureRandomCross
CPC.248 ± .008.223 ± .006.221 ± .006.189 ± .007
RSSM.268 ± .007.228 ± .005.222 ± .007.202 ± .007

Bold marks the Cross result within each separately evaluated learner family.

  • CPCCompare methods within a learner row; this is the adapted S3/query3 held-out task.

Means ± sample SD across three jointly fitted source/reader seeds per arm. The same S3/query3 inputs are used throughout. This is a held-out episode test under the adapted protocol, not a verified unseen-parameter split.

02 / Study Table 7

Collision held-out comparison — A. Matched four-arm coverage: three seeds per arm

Same 1,994 held-out episodes; 3 CPC/RSSM seeds per arm; JEPA J2: 3 Cross seeds and 1 per control

Trajectory MSE ↓; Cross error reduction ↑

Collision held-out comparison — A. Matched four-arm coverage: three seeds per arm
FamilyNativeStructureRandomCross
CPC0.248382 ± 0.0078220.222821 ± 0.0060050.220925 ± 0.0058140.188559 ± 0.006778
RSSM0.267842 ± 0.0073470.228131 ± 0.0048990.222168 ± 0.0067950.202082 ± 0.007051

Bold marks the Cross result within each three-seed learner comparison.

Means ± sample SD across three joint source/reader seeds per arm on the same 1,994 episodes. Every CPC and RSSM seed favors Cross over Structure and Random. The adapted task is not a verified unseen-parameter split.

03 / Study Table 7

Collision held-out comparison — B. JEPA J2: three SPRII seeds, one fitted seed per control

Same 1,994 held-out episodes; 3 CPC/RSSM seeds per arm; JEPA J2: 3 Cross seeds and 1 per control

Trajectory MSE ↓; Cross error reduction ↑

Collision held-out comparison — B. JEPA J2: three SPRII seeds, one fitted seed per control
FamilyNativeStructureRandomCross
JEPA0.247625 (–)0.214314 (–)0.261401 (–)0.184110 ± 0.005443

Bold marks the three-seed JEPA J2 result.

  • JEPAEach control has one fitted seed. This follow-up does not have the balanced seed coverage of CPC and RSSM.

JEPA J2 is a development-selected, matched-only retest: three Cross fits versus one fit per control. For fixed seed-0 models, the Cross-minus-Structure paired-recipient 95% interval is [−0.045874,−0.020480]; it is not a training-seed interval.

04 / Study Table 7

Collision held-out comparison — C. Cross mean-error reductions

Same 1,994 held-out episodes; 3 CPC/RSSM seeds per arm; JEPA J2: 3 Cross seeds and 1 per control

Trajectory MSE ↓; Cross error reduction ↑

Collision held-out comparison — C. Cross mean-error reductions
FamilyVersus StructureVersus Random
CPC15.38%14.65%
RSSM11.42%9.04%
JEPA J2 (unequal coverage)14.09%29.57%

Bold marks reductions from the balanced three-seed CPC and RSSM comparisons.

The dashed rule separates the unequal-coverage JEPA follow-up.

Reductions compare equally weighted mean errors. CPC and RSSM have three fitted seeds per arm; JEPA J2 has three Cross fits and one per control, so its coverage is not directly equivalent.

05 / Study Table 35

Full CoPhy learner × scene matrix

S3/query3 development; source100/head100; source coverage explicit

Trajectory MSE ↓; support and reader-memory probe R² where reported

Full CoPhy learner × scene matrix
GroupSceneFamilyNativeStructureSPRIIRandom
Cross: self-supervised dynamics objectivesBallsJEPA1.25631.26301.23961.5827
Cross: self-supervised dynamics objectivesBallsCPC1.43991.42391.50421.6228
Cross: self-supervised dynamics objectivesBallsRSSM1.31051.25821.12541.3155
Cross: self-supervised dynamics objectivesCollisionJEPA0.25120.22120.22970.2717
Cross: self-supervised dynamics objectivesCollisionCPC0.24480.22830.17780.2220
Cross: self-supervised dynamics objectivesCollisionRSSM0.27310.23380.19700.2192
Cross: self-supervised dynamics objectivesBlocktowerJEPA0.09310.08780.08650.0863
Cross: self-supervised dynamics objectivesBlocktowerCPC0.10300.09910.09950.1007
Cross: self-supervised dynamics objectivesBlocktowerRSSM0.09740.09350.09430.0942
Align + Cross: supervised CoPhyNet objectiveBallsCoPhyNet1.0534–1.00561.0698
Align + Cross: supervised CoPhyNet objectiveCollisionCoPhyNet0.2033–0.21840.2171
Align + Cross: supervised CoPhyNet objectiveBlocktowerCoPhyNet0.0871–0.08820.0862

Bold marks the lowest reported mean only within each scene × learner row.

Dashed rules separate scenes and the supervised CoPhyNet block.

  • Align + Cross: supervised CoPhyNet objective · BallsA dash is an unevaluated arm. Development source 0 is not the held-out Collision benchmark.

The full source-0 development matrix retains favorable and adverse comparisons. SPRII uses Cross for JEPA/CPC/RSSM and Align + Cross for CoPhyNet. A dash is an unevaluated arm. These validation cohorts differ from the held-out Collision test.

06 / Study Table 36

Full CoPhy learner × scene matrix

S3/query3 development; source100/head100; source coverage explicit

Trajectory MSE ↓; support and reader-memory probe R² where reported

Full CoPhy learner × scene matrix
GroupArmMSE src 0MSE src 1MSE src 2Support R² src 0Support R² src 1Support R² src 2Reader R² src 0Reader R² src 1Reader R² src 2
JEPA / BallsNative1.25631.3185—0.79770.7981—0.78650.7778—
JEPA / BallsStructure1.26301.2944—0.78550.7866—0.78320.7788—
JEPA / BallsCross1.23961.3054—0.91620.9185—0.90510.9116—
JEPA / BallsRandom1.58271.6436—0.67770.6918—0.68840.6823—
JEPA / CollisionNative0.25120.2444—0.44310.4366—0.56570.5619—
JEPA / CollisionStructure0.22120.2238—0.59760.6246—0.66590.6706—
JEPA / CollisionCross0.22970.2154—0.64220.6718—0.64510.6713—
JEPA / CollisionRandom0.27170.2645—0.50050.5250—0.52860.5502—
JEPA / BlocktowerNative0.09310.0895—0.10280.0902—0.04300.0467—
JEPA / BlocktowerStructure0.08780.0876—0.12520.1301—0.07140.0879—
JEPA / BlocktowerCross0.08650.0876—0.17610.1564—0.12420.0860—
JEPA / BlocktowerRandom0.08630.0853—0.16730.1678—0.12560.1154—
CPC / CollisionNative0.24480.2517—0.43720.4213—0.56040.5484—
CPC / CollisionStructure0.22830.2323—0.61260.5929—0.64760.6367—
CPC / CollisionCross0.17780.1841—0.74950.7411—0.75410.7419—
CPC / CollisionRandom0.22200.2212—0.66570.6588—0.67170.6711—
RSSM / CollisionNative0.27310.26740.27290.3340.2400.3440.4970.4940.477
RSSM / CollisionStructure0.23380.23320.22730.5640.5650.5700.6270.6420.634
RSSM / CollisionCross0.19700.19970.21120.7380.7230.7330.7360.7190.724
RSSM / CollisionRandom0.21920.21990.22840.6330.6430.6340.6780.6690.664

Bold identifies Cross prediction errors, including adverse cells; it is not a winner mark.

Dashed rules separate learner × scene groups.

  • JEPA / Balls · NativeSource columns remain separate. Support R² and reader-memory R² are accessibility assays, not error metrics.

Source coverage is explicit; dashes denote unevaluated cells. Support and reader-memory probes target restitution in Balls, mass in Collision, and vertical gravity in Blocktower. These development cohorts include selection examples and remain separate from the held-out test.

07 / Study Table 36

RSSM Collision development replication summary

S3/query3 development; source100/head100; 3 sources

Trajectory MSE ↓; mass probe R²

RSSM Collision development replication summary
ArmMSE: mean ± SDSupport R^2_SReader memory R^2_M
Native0.2711 ± 0.00320.3060.490
Structure0.2314 ± 0.00360.5660.635
Cross0.2026 ± 0.00750.7310.727
Random0.2225 ± 0.00510.6360.671

Bold marks Cross prediction MSE in the three-source development comparison.

  • CrossThe neighboring R² columns measure accessibility; this population is separate from the held-out test.

MSE is mean ± sample SD across three RSSM sources; support and reader-memory R² probe mass. These S3/query3 development data include selection examples and are not the held-out Collision test population.

08 / Study Table 37

Frozen-source reader design across scenes

Source-50/head-100; source seed 0; full validation

Relative error reduction versus Native (%) ↑

Frozen-source reader design across scenes
Reader recipeBallsCollisionBlocktower
Main initialization, dropout 0.1, initial-5.448-4.089-1.224
Alternative initialization, dropout 0+7.506-2.104-1.477
Alternative initialization, dropout 0.1+6.074-4.426+2.090
New initialization, dropout 0+0.617-7.591+0.621
New initialization, dropout 0.1, every step+6.646-2.925+1.134

Bold marks sign reversals across reader recipes in Balls and Blocktower.

  • Main initialization, dropout 0.1, initialPositive is an error reduction versus Native. Each row is one shared recipe across scenes; the best cell per scene is not a common selected model.

Every row is one shared reader recipe across all three scenes, using the same frozen source-50 checkpoint. Positive is lower error than Native. The sign varies by scene and recipe; separately selecting each scene’s best cell would change the comparison.

09 / Study Table 39

Complete follow-up outcomes

Separate configurations and populations shown row by row

MSE ↓, except FHN relative L₂ ↓

Complete follow-up outcomes
ConfigurationCoverageNumerical comparisonFinding / scope
JEPA Collision: J21,994 test episodes; 3 J2 seeds, 1/controlJ2 0.184110 ± 0.005443; Structure 0.214314Matched-only retest. Four arms and fixed-model interval: Table tab:jepa-collision-followup.
CoPhyNet Collision: Cross4,000 development recipients; seed 0Cross 0.203240; prior SPRII 0.210671Near Native; paired interval against Native includes zero.
CPC Balls: Align2,000 development recipients; seed 0Align 1.392711; prior SPRII 1.415352; Structure 1.387643Lower than prior SPRII, higher than Structure. Full contrasts are indexed in the package.
FHN: tuned SPRIIReport ICs 15/24/45; source 42; selection IC 5 excludedID K=4,H=50: SPRII 0.020132; extra-50k NOD 0.021522Paired-system interval includes zero; other eight split/horizon cells have higher error. Extra-budget control.
Poke object-only: R8 / G2400 development systems; source 0, reader 0; H=16R8: M2 0.901091; M1 0.902588; recipient-only 0.911791. M2-M1: -0.001497 [-0.011855,0.008400]. G2: M2 0.913197; M1 0.908531; recipient-only 0.908021Four-dimensional object-state error; M1 and M2 selected independently.
CoPhyNet Blocktower8,088 development recipients; seed 0SPRII 0.088180; Native/Structure 0.087114; Random 0.086230. SPRII-Random: 0.001951 [0.000668,0.003184]Paired interval versus Structure includes zero; Random has lower mean error.

Bold identifies each follow-up configuration; heterogeneous numerical cells are deliberately not ranked.

Dashed rules separate experiments with different populations and fitted-model coverage.

All completed follow-ups from the configuration round are retained, including null and adverse results. Each row has its own population and model coverage. Paired intervals condition on the reported fitted models; the one-source development rows do not replace the original multi-source comparisons.

Practical implications

What to take from this setting.

Treat Align and Cross as separate choices.

Keep the native objective, compare the components, and select on the task you want to improve.

The combined objective is not established as the best recipe for every learner and task.