SPRIIA study of persistent information
Use · Generality

CoPhy · Balls

Shared objects, different trajectories.

Ball interactions connect physical-information probes in actual reader memory with tests of whether donor context changes prediction.

Probe + interventionaccessibility and functional use
Motion preview
Physics demonstration · initial velocities and collisions · ½ speedCoPhy: Fabien Baradel et al. · reconstructed scene. Source & changes
Across independent experiences

What persists.

Matched objects under the scene-specific relation.

Across distinct realizations

What changes.

Support interactions and initial motion.

Experience → prediction

The task.

Three support histories + three query frames.

→ Ball trajectories.

Evaluated instance

The comparison.

JEPA / CPC / RSSM: Cross. CoPhyNet: Align + Cross.

Controls. Native; Random; Structure where tested.

Interpreting the evidence

What stays fixed in the test.

Probes inspect frozen reader memory. Donor interventions hold the consuming prediction route and recipient query fixed.

What this setting establishes.

This scene connects memory probes with donor-use measurements and extends the prediction comparison across learners.

Scope of the evidence

Accessibility, donor dependence and prediction error answer different questions. The displayed histories are a PyBullet reconstruction of the published scene definition; they are separate from the evaluated dataset.

Full result record

Measurements and comparisons.

Reported results are kept with their own populations and conditions. Development, held-out and sealed comparisons remain separate.

01 / Study Table 24

More distinct histories through a fixed reader

S3-trained RSSM; mean of 3 retained fits

Trajectory MSE ↓

More distinct histories through a fixed reader
RecipeRepeatS3S5S8
SPRII (Cross)1.28961.17071.15591.1467
Native1.53201.28581.24131.2065
Random1.49661.35631.33271.3164
Structure1.51721.28841.24561.2177

Bold compares three versus eight distinct histories through the same SPRII reader.

  • SPRII (Cross)The reader was trained with S3 and is not refitted for S5/S8; this is an input substitution.

MSE averages three retained fits; Repeat averages individual-history repetitions. S3, S5 and S8 change the input history count through the same S3-trained RSSM reader. No retraining is performed, so this is not a sample-efficiency comparison.

02 / Study Table 25

Reader-memory accessibility and donor physics — A. CoPhyNet reader-memory probes (R^2)

Probe: 26,558 training / 7,578 validation objects; 3 sources; separate source100 donor assays

Probe R²; same/wrong-physics distance ratio; wrong-factor-minus-matched MSE

Reader-memory accessibility and donor physics — A. CoPhyNet reader-memory probes (R^2)
RecipeFactorSource 1Source 2Source 3
SPRII (A+C)Mass0.05090.05460.0678
SPRII (A+C)Friction0.0027-0.00200.0002
SPRII (A+C)Restitution0.82970.82660.8210
NativeMass0.02830.02710.0411
NativeFriction-0.00140.00240.0009
NativeRestitution0.75810.75990.7684
RandomMass0.03060.02620.0295
RandomFriction-0.0025-0.00430.0007
RandomRestitution0.76110.76810.7598

Bold marks restitution readout under each recipe, keeping the three sources separate.

Dashed rules separate training recipes.

  • SPRII (A+C) · RestitutionMass and friction remain in the table. Readout accessibility is not a direct measurement of predictor reliance.

Training-only ridge probes inspect CoPhyNet’s actual reader memory, including its learned projection, using 26,558 training and 7,578 validation objects. Sources remain separate. Strong restitution accessibility is distinct from fixed-predictor donor reliance.

03 / Study Table 25

Reader-memory accessibility and donor physics — B. RSSM within/between geometry

Probe: 26,558 training / 7,578 validation objects; 3 sources; separate source100 donor assays

Probe R²; same/wrong-physics distance ratio; wrong-factor-minus-matched MSE

Reader-memory accessibility and donor physics — B. RSSM within/between geometry
RecipeSource 1Source 2Source 3
SPRII (Cross)1.1821.3731.059
Native3.9041.7692.493
Random2.3702.4292.334
Structure2.4792.4152.496

Bold marks the SPRII within/between ratios across all three sources.

  • SPRII (Cross)This is same-physics divided by wrong-physics distance. A lower ratio is not by itself a task benefit.

The ratio divides same-physics code distance by wrong-physics code distance for each RSSM source separately. These geometry measurements do not establish task value; donor interventions provide a different test of the fitted reader.

04 / Study Table 25

Reader-memory accessibility and donor physics — C. Separate source100 donor assay

Probe: 26,558 training / 7,578 validation objects; 3 sources; separate source100 donor assays

Probe R²; same/wrong-physics distance ratio; wrong-factor-minus-matched MSE

Reader-memory accessibility and donor physics — C. Separate source100 donor assay
LearnerMatched MSEΔ mΔ frictionΔ restitution
JEPA1.239584+0.000490+0.001173+0.514036
CoPhyNet1.005565+0.004422-0.001568+0.518382

Bold marks the restitution donor penalties, the large reported factor-specific effects.

  • JEPAThese are separate single-source assays. Small or negative effects for other factors do not establish physical irrelevance.

These source100 donor assays are separate single-source evaluations, not three-source replications. Each delta is wrong-factor minus matched MSE through a fixed prediction route. Small or negative effects do not establish that a factor is physically irrelevant.

05 / Study Table 35

Full CoPhy learner × scene matrix

S3/query3 development; source100/head100; source coverage explicit

Trajectory MSE ↓; support and reader-memory probe R² where reported

Full CoPhy learner × scene matrix
GroupSceneFamilyNativeStructureSPRIIRandom
Cross: self-supervised dynamics objectivesBallsJEPA1.25631.26301.23961.5827
Cross: self-supervised dynamics objectivesBallsCPC1.43991.42391.50421.6228
Cross: self-supervised dynamics objectivesBallsRSSM1.31051.25821.12541.3155
Cross: self-supervised dynamics objectivesCollisionJEPA0.25120.22120.22970.2717
Cross: self-supervised dynamics objectivesCollisionCPC0.24480.22830.17780.2220
Cross: self-supervised dynamics objectivesCollisionRSSM0.27310.23380.19700.2192
Cross: self-supervised dynamics objectivesBlocktowerJEPA0.09310.08780.08650.0863
Cross: self-supervised dynamics objectivesBlocktowerCPC0.10300.09910.09950.1007
Cross: self-supervised dynamics objectivesBlocktowerRSSM0.09740.09350.09430.0942
Align + Cross: supervised CoPhyNet objectiveBallsCoPhyNet1.0534–1.00561.0698
Align + Cross: supervised CoPhyNet objectiveCollisionCoPhyNet0.2033–0.21840.2171
Align + Cross: supervised CoPhyNet objectiveBlocktowerCoPhyNet0.0871–0.08820.0862

Bold marks the lowest reported mean only within each scene × learner row.

Dashed rules separate scenes and the supervised CoPhyNet block.

  • Align + Cross: supervised CoPhyNet objective · BallsA dash is an unevaluated arm. Development source 0 is not the held-out Collision benchmark.

The full source-0 development matrix retains favorable and adverse comparisons. SPRII uses Cross for JEPA/CPC/RSSM and Align + Cross for CoPhyNet. A dash is an unevaluated arm. These validation cohorts differ from the held-out Collision test.

06 / Study Table 36

Full CoPhy learner × scene matrix

S3/query3 development; source100/head100; source coverage explicit

Trajectory MSE ↓; support and reader-memory probe R² where reported

Full CoPhy learner × scene matrix
GroupArmMSE src 0MSE src 1MSE src 2Support R² src 0Support R² src 1Support R² src 2Reader R² src 0Reader R² src 1Reader R² src 2
JEPA / BallsNative1.25631.3185—0.79770.7981—0.78650.7778—
JEPA / BallsStructure1.26301.2944—0.78550.7866—0.78320.7788—
JEPA / BallsCross1.23961.3054—0.91620.9185—0.90510.9116—
JEPA / BallsRandom1.58271.6436—0.67770.6918—0.68840.6823—
JEPA / CollisionNative0.25120.2444—0.44310.4366—0.56570.5619—
JEPA / CollisionStructure0.22120.2238—0.59760.6246—0.66590.6706—
JEPA / CollisionCross0.22970.2154—0.64220.6718—0.64510.6713—
JEPA / CollisionRandom0.27170.2645—0.50050.5250—0.52860.5502—
JEPA / BlocktowerNative0.09310.0895—0.10280.0902—0.04300.0467—
JEPA / BlocktowerStructure0.08780.0876—0.12520.1301—0.07140.0879—
JEPA / BlocktowerCross0.08650.0876—0.17610.1564—0.12420.0860—
JEPA / BlocktowerRandom0.08630.0853—0.16730.1678—0.12560.1154—
CPC / CollisionNative0.24480.2517—0.43720.4213—0.56040.5484—
CPC / CollisionStructure0.22830.2323—0.61260.5929—0.64760.6367—
CPC / CollisionCross0.17780.1841—0.74950.7411—0.75410.7419—
CPC / CollisionRandom0.22200.2212—0.66570.6588—0.67170.6711—
RSSM / CollisionNative0.27310.26740.27290.3340.2400.3440.4970.4940.477
RSSM / CollisionStructure0.23380.23320.22730.5640.5650.5700.6270.6420.634
RSSM / CollisionCross0.19700.19970.21120.7380.7230.7330.7360.7190.724
RSSM / CollisionRandom0.21920.21990.22840.6330.6430.6340.6780.6690.664

Bold identifies Cross prediction errors, including adverse cells; it is not a winner mark.

Dashed rules separate learner × scene groups.

  • JEPA / Balls · NativeSource columns remain separate. Support R² and reader-memory R² are accessibility assays, not error metrics.

Source coverage is explicit; dashes denote unevaluated cells. Support and reader-memory probes target restitution in Balls, mass in Collision, and vertical gravity in Blocktower. These development cohorts include selection examples and remain separate from the held-out test.

07 / Study Table 37

Frozen-source reader design across scenes

Source-50/head-100; source seed 0; full validation

Relative error reduction versus Native (%) ↑

Frozen-source reader design across scenes
Reader recipeBallsCollisionBlocktower
Main initialization, dropout 0.1, initial-5.448-4.089-1.224
Alternative initialization, dropout 0+7.506-2.104-1.477
Alternative initialization, dropout 0.1+6.074-4.426+2.090
New initialization, dropout 0+0.617-7.591+0.621
New initialization, dropout 0.1, every step+6.646-2.925+1.134

Bold marks sign reversals across reader recipes in Balls and Blocktower.

  • Main initialization, dropout 0.1, initialPositive is an error reduction versus Native. Each row is one shared recipe across scenes; the best cell per scene is not a common selected model.

Every row is one shared reader recipe across all three scenes, using the same frozen source-50 checkpoint. Positive is lower error than Native. The sign varies by scene and recipe; separately selecting each scene’s best cell would change the comparison.

08 / Study Table 38

Balls prediction and paired method contrasts — Absolute matched, null and wrong-donor error

2,000 development recipients; 3 fitted sources per learner

Matched/Null/Wrong MSE; comparator-minus-SPRII gap ↑

Balls prediction and paired method contrasts — Absolute matched, null and wrong-donor error
FamilyArmMatchedNullWrong
RSSMSPRII (C)1.17071.89242.2957
RSSMNative1.28582.48532.2931
RSSMRandom1.35632.00412.3000
RSSMStructure1.28841.93722.3350
CoPhyNetSPRII (A+C)1.04421.79122.2277
CoPhyNetNative1.08661.77332.2445
CoPhyNetRandom1.07861.75862.2157

Bold marks matched-context SPRII risk within each learner family.

The dashed rule separates RSSM and CoPhyNet.

  • RSSM · SPRII (C)Null and Wrong are fixed context conditions in this record; comparing their sizes does not rank context quality independently of the reader.

Matched, Null and Wrong risks share the 2,000-recipient development cohort within each learner family and average three fixed sources. RSSM uses Cross; CoPhyNet uses Align + Cross. These context conditions measure different uses of the fitted prediction route.

09 / Study Table 38

Balls prediction and paired method contrasts — Adjusted paired method contrasts

2,000 development recipients; 3 fitted sources per learner

Matched/Null/Wrong MSE; comparator-minus-SPRII gap ↑

Balls prediction and paired method contrasts — Adjusted paired method contrasts
FamilyComparatorDifferenceAdjusted 95% interval
RSSMNative0.1151[0.0902,0.1395]
RSSMRandom0.1856[0.1607,0.2096]
RSSMStructure0.1177[0.0927,0.1433]
CoPhyNetNative0.0424[0.0276,0.0574]
CoPhyNetRandom0.0344[0.0193,0.0497]

Bold marks the five prespecified comparator-minus-SPRII differences.

The dashed rule separates learner families.

  • RSSM · NativeThe intervals shown are Bonferroni-adjusted across all five contrasts and condition on the fitted models.

Comparator-minus-SPRII differences are paired before averaging sources. Recipient-bootstrap intervals are Bonferroni-adjusted across all five contrasts. They condition on fitted models and exclude model-selection uncertainty.

10 / Study Table 39

Complete follow-up outcomes

Separate configurations and populations shown row by row

MSE ↓, except FHN relative L₂ ↓

Complete follow-up outcomes
ConfigurationCoverageNumerical comparisonFinding / scope
JEPA Collision: J21,994 test episodes; 3 J2 seeds, 1/controlJ2 0.184110 ± 0.005443; Structure 0.214314Matched-only retest. Four arms and fixed-model interval: Table tab:jepa-collision-followup.
CoPhyNet Collision: Cross4,000 development recipients; seed 0Cross 0.203240; prior SPRII 0.210671Near Native; paired interval against Native includes zero.
CPC Balls: Align2,000 development recipients; seed 0Align 1.392711; prior SPRII 1.415352; Structure 1.387643Lower than prior SPRII, higher than Structure. Full contrasts are indexed in the package.
FHN: tuned SPRIIReport ICs 15/24/45; source 42; selection IC 5 excludedID K=4,H=50: SPRII 0.020132; extra-50k NOD 0.021522Paired-system interval includes zero; other eight split/horizon cells have higher error. Extra-budget control.
Poke object-only: R8 / G2400 development systems; source 0, reader 0; H=16R8: M2 0.901091; M1 0.902588; recipient-only 0.911791. M2-M1: -0.001497 [-0.011855,0.008400]. G2: M2 0.913197; M1 0.908531; recipient-only 0.908021Four-dimensional object-state error; M1 and M2 selected independently.
CoPhyNet Blocktower8,088 development recipients; seed 0SPRII 0.088180; Native/Structure 0.087114; Random 0.086230. SPRII-Random: 0.001951 [0.000668,0.003184]Paired interval versus Structure includes zero; Random has lower mean error.

Bold identifies each follow-up configuration; heterogeneous numerical cells are deliberately not ranked.

Dashed rules separate experiments with different populations and fitted-model coverage.

All completed follow-ups from the configuration round are retained, including null and adverse results. Each row has its own population and model coverage. Paired intervals condition on the reported fitted models; the one-source development rows do not replace the original multi-source comparisons.