SPRIIA study of persistent information
Prediction · Formation · Use · Value

SpringWorld

Different motion. The same physical system.

A spring system changes its trajectory while mass, drag and stiffness persist. Independent histories help predict a new encounter.

22.3%lower MSE than native JEPA
Motion preview
Physics demonstration · force and releaseProject-generated demonstration · MuJoCo simulation. Source & changes
Across independent experiences

What persists.

Mass m, drag γ and stiffness k.

Across distinct realizations

What changes.

Independent interaction history: initial state and actions.

Experience → prediction

The task.

Independent history + current query and future actions.

→ Future position and velocity increments.

Evaluated instance

The comparison.

JEPA · Align + Cross. Structure, Align and Cross are separate component studies.

Controls. Native JEPA; TDS; component controls.

Interpreting the evidence

What stays fixed in the test.

For donor interventions: source, fitted reader, recipient initial state and future actions are fixed. One donor factor and the corresponding simulator target factor are varied independently.

What this setting establishes.

Sealed prediction supports a task benefit. Separate development studies trace representation geometry, donor-specific prediction and the value recovered by different readers.

Scope of the evidence

The sealed 256-system comparison and the development donor/readout studies are separate populations. A factor probe alone does not establish task benefit. The M1/M2 readers share the registered architecture but not exactly the active capacity: M1 leaves its persistent projection unused.

Full result record

Measurements and comparisons.

Reported results are kept with their own populations and conditions. Development, held-out and sealed comparisons remain separate.

01 / Study Table 1

Sealed cold-start prediction

256 sealed systems; 3 sources × 3 readers

Standardized state-increment MSE ↓

Sealed cold-start prediction
Source recipeMean ± SDSPRII reduction [95% CI]
Native JEPA.138 ± .00222.3% [13.2, 30.7]
TDS.130 ± .01317.5% [8.4, 25.9]
SPRII (Align + Cross).107 ± .004–

Bold marks SPRII’s primary mean and the two paired error reductions.

  • SPRII (Align + Cross)The absolute means and paired reductions use the same sealed population; uncertainty has the definitions shown above.

Source means ± sample SD after averaging readers. Error reductions use paired physical-system 95% intervals, conditional on the fitted models; the sealed population is separate from the development studies.

02 / Study Table 6

Sealed prediction: populations and paired contrasts

Primary: 256 systems; secondary: 36; pooled: 292; 3 sources × 3 readers

Standardized MSE ↓; reduction versus comparator ↑

Sealed prediction: populations and paired contrasts
SourcePrimary: 256 systemsSecondary: 36Pooled: 292
Native0.13824 [0.11851,0.15924]0.213710.14754
TDS0.13025 [0.11181,0.15023]0.196700.13844
Align + Cross0.10740 [0.09036,0.12668]0.127260.10985

Bold marks the SPRII mean on the prespecified primary population.

  • Align + CrossPrimary, secondary and pooled cohorts are separate; the primary result is the emphasized comparison.

Six queries per system; readers are averaged within source. Primary MSE intervals resample physical systems and condition on the fitted models. The secondary and system-weighted pooled cohorts are reported separately.

03 / Study Table 6

Sealed prediction: populations and paired contrasts

Primary: 256 systems; secondary: 36; pooled: 292; 3 sources × 3 readers

Standardized MSE ↓; reduction versus comparator ↑

Sealed prediction: populations and paired contrasts
PopulationComparatorMSE gap [95% interval]Reduction [95% interval]
PrimaryNative0.03084 [0.01735,0.04429]22.31% [13.15,30.69]
PrimaryTDS0.02284 [0.01039,0.03487]17.54% [8.40,25.87]
SecondaryNative0.08645 [0.02195,0.16640]40.45% [16.34,54.55]
SecondaryTDS0.06944 [0.01587,0.13303]35.30% [12.10,49.13]
PooledNative0.03769 [0.02285,0.05302]25.55% [16.90,33.32]
PooledTDS0.02859 [0.01560,0.04176]20.65% [12.00,28.45]

Bold marks paired contrasts on the primary 256-system population.

Dashed rules separate evaluation populations.

Each contrast is paired within its stated population. MSE gaps and percentage reductions use 95% physical-system intervals conditional on the fitted models. Primary, secondary and pooled results are distinct comparisons.

04 / Study Table 8

Adapted methods under matched history access — A. Spring: original reader

Development cold start; 100 systems / 600 cases; 3 sources × 3 readers

Standardized MSE ↓

Adapted methods under matched history access — A. Spring: original reader
ReaderSourceFresh NullMatchedFixed Wrong
OriginalAlign + Cross0.1774880.1121290.243931
OriginalTDS0.1715980.1595790.178664
OriginalCompact TDS (d=0)0.1709270.1709270.170927

Bold compares matched-context errors under the original reader.

  • Original · Compact TDS (d=0)Compact TDS has no persistent channel; its three context conditions therefore coincide.

Sources and readers are equally weighted. Fresh Null is fitted separately; Fixed Wrong replaces context through the unchanged Matched reader. The compact TDS arm has no persistent channel, so its context conditions coincide.

05 / Study Table 8

Adapted methods under matched history access — B. Spring: common adapter

Development cold start; 100 systems / 600 cases; 3 sources × 3 readers

Standardized MSE ↓

Adapted methods under matched history access — B. Spring: common adapter
ReaderSourceFresh NullMatchedFixed Wrong
Common adapterAlign + Cross0.1778370.113644—
Common adapterFCRL-style0.2767210.273478—

Bold compares matched-context errors under the common adapter.

  • Common adapter · FCRL-styleThese development results are separate from both the original-reader comparison and the sealed benchmark.

Both source families use the common downstream adapter, with three sources and three readers. These development results remain separate from the original-reader protocol and the sealed task comparison. A dash denotes an unavailable intervention.

06 / Study Table 17

Align and Cross contribute differently — A. Three-source component mean (reader 0)

Task: 100 development systems / 600 cases, reader 0; geometry: 178 systems × 4 histories; 3 sources

Cold-start MSE; frozen-source geometry and partial correlations

Align and Cross contribute differently — A. Three-source component mean (reader 0)
MeasurementNativeStructureAlignCrossAlign + Cross
Mean0.1702610.1715770.1231530.1675070.110386

Bold marks Align and Align + Cross, the two alignment-containing task recipes.

  • MeanComponent means average three sources with reader seed 0; this is a development comparison.

Means equally average three sources, each with reader seed 0 under the original 10,000-update recipe. The 100-system cold-start development population is separate from both the geometry bank and the sealed prediction benchmark.

07 / Study Table 17

Align and Cross contribute differently — B. Frozen-source geometry

Task: 100 development systems / 600 cases, reader 0; geometry: 178 systems × 4 histories; 3 sources

Cold-start MSE; frozen-source geometry and partial correlations

Align and Cross contribute differently — B. Frozen-source geometry
RecipeWithinBetweenRatioρ_mρ_γρ_kρ_k/m
Structure131.58731.6400.2400.0010.005-0.014-0.025
Align47.24790.8992.0450.3610.2240.1320.414
Cross115.70829.2150.2520.2000.1840.0560.066
Align + Cross40.58493.6322.3600.3540.1930.1410.421

Bold marks the between/within ratios of the alignment-containing recipes.

  • Align + CrossGeometry columns are distinct diagnostics; larger code separation does not by itself establish task benefit.

Frozen, training-standardized source codes use 178 systems and four histories per system; no reader is fitted for this measurement. Values equally average three sources. Ratio averages source-wise between/within ratios; partial correlations control the other factors.

08 / Study Table 18

Development reference learners and physical probes — A. Single-source reference (source 0, reader 0)

Prediction: source 0 / reader 0, 100 cold-start or 178 mixture systems; probe: 100 validation systems / 400 histories

Standardized MSE ↓; log-factor probe R² ↑

Development reference learners and physical probes — A. Single-source reference (source 0, reader 0)
GroupSourceCold startQuery mixture
Controlled JEPA objectivesNative0.17160.8511
Controlled JEPA objectivesStructure0.17110.8468
Controlled JEPA objectivesAlign0.11950.8033
Controlled JEPA objectivesCross0.16460.8443
Controlled JEPA objectivesAlign + Cross0.09200.7892
Other reference learnersCPC0.22180.9858
Other reference learnersRSSM0.16530.7882
Other reference learnersSupervised split0.18410.8864
Supervised reference encodersGRU0.09340.6328
Supervised reference encodersTransformer0.08870.6631
Supervised reference encodersTCN0.05930.5948
Supervised reference encodersDeepSets0.08920.6597

Bold marks the lowest cold-start mean within each of the controlled-JEPA and supervised-reference groups.

Dashed rules separate learner groups with different training objectives.

  • Supervised reference encoders · DeepSetsReference encoders use their own supervised training recipes; this is not one globally matched leaderboard.

Single-source, single-reader development comparisons use matched standardized MSE. Cold start has 100 systems; the mixture has 178. Supervised reference encoders retain their own training recipes, so the groups do not form one globally matched leaderboard.

09 / Study Table 18

Development reference learners and physical probes — B. Frozen-source physical probes

Prediction: source 0 / reader 0, 100 cold-start or 178 mixture systems; probe: 100 validation systems / 400 histories

Standardized MSE ↓; log-factor probe R² ↑

Development reference learners and physical probes — B. Frozen-source physical probes
Sourcemγkk/m
Native0.0076-0.0234-0.0265-0.0085
Align + Cross0.38060.15290.20230.4945

Bold marks factor accessibility from Align + Cross in this frozen-source comparison.

  • Align + CrossProbe R² measures accessibility, not fixed-predictor use or downstream value.

Training-fitted ridge probes evaluate log factors using 400 histories from 100 validation systems; hyperparameters are chosen using training data only. These frozen-source probes measure accessibility, separately from prediction error and the sealed benchmark.

10 / Study Table 20

Donor physics structures fixed predictive responses

64 development systems; 3 frozen sources; reader 0

Valley depth = off-diagonal minus diagonal h16 MSE ↑

Donor physics structures fixed predictive responses
FactorSourceValley depth95% intervalρ
MassPooled0.1820[0.1320, 0.2343]0.591
MassSeed 00.1950[0.1379, 0.2558]0.606
MassSeed 10.1612[0.1131, 0.2134]0.592
MassSeed 20.1899[0.1364, 0.2462]0.601
DragPooled0.0303[0.0216, 0.0400]0.585
DragSeed 00.0334[0.0238, 0.0439]0.600
DragSeed 10.0335[0.0229, 0.0458]0.578
DragSeed 20.0240[0.0159, 0.0327]0.542
StiffnessPooled0.0240[0.0156, 0.0338]0.672
StiffnessSeed 00.0285[0.0184, 0.0407]0.776
StiffnessSeed 10.0229[0.0143, 0.0328]0.401
StiffnessSeed 20.0205[0.0124, 0.0299]0.628

Bold marks the pooled matched-factor advantages and their paired intervals.

Dashed rules separate physical factors; the scales are not pooled.

Valley depth is off-diagonal minus diagonal h16 error. Three frozen sources and reader 0 are evaluated on 64 development systems. Paired-system 95% intervals condition on the fitted routes; physical factors are assessed separately.

11 / Study Table 20

Donor physics structures fixed predictive responses

64 development systems; 3 frozen sources; reader 0

Valley depth = off-diagonal minus diagonal h16 MSE ↑

Donor physics structures fixed predictive responses
RecipeMass: depth [95% CI]Drag: depth [95% CI]Stiffness: depth [95% CI]
Structure.01179 [.00573,.01900].00310 [.00156,.00485].00016 [-.00014,.00055]
Align.19224 [.14024,.24706].03124 [.02198,.04152].02383 [.01586,.03353]
Cross.03197 [.01837,.04807].00886 [.00508,.01344].00033 [-.00066,.00142]
Align + Cross.18205 [.13200,.23429].03033 [.02148,.03983].02395 [.01539,.03398]

Bold marks alignment-containing recipes in each factor-specific donor test.

  • Align + CrossEach recipe has its own fitted source and reader; these cells are not a causal comparison through one shared reader.

Each objective recipe has its own source and fitted reader. Values are off-diagonal-minus-diagonal h16 errors with paired-system 95% intervals on 64 development systems. Positive values favor matched donor physics; factor scales are kept separate.

12 / Study Table 27

History value grows with prediction horizon

178-system mixture; cold/moving strata; 3 sources × 3 readers

Persistent-versus-Null error reduction (%) ↑

History value grows with prediction horizon
HorizonCold: gain [95% interval]Moving: gain [95% interval]
11.95 [-8.47,11.01]1.15 [-.83,3.01]
213.39 [3.93,21.49]1.98 [-.22,4.02]
423.19 [14.37,30.35]3.23 [.91,5.77]
835.20 [26.74,41.92]11.62 [7.17,16.57]
1635.68 [28.22,42.27]29.73 [22.68,36.45]

Bold marks the shortest and longest reported horizons to make the conditional-value trend visible.

  • 1Both query conditions have zero observed recipient transitions. The h1 intervals include zero; longer-horizon results use the same fitted grid.

Persistent and Null readers are fitted separately on the same 178-system grid. Cold and moving queries both have zero observed recipient transitions. Pointwise stratified system-bootstrap 95% intervals condition on the fitted source/reader grid.

13 / Study Table 29

Same-donor readers and factor accessibility — A. Reader risk: cold-start and full-mixture populations

Reader: 100 cold-start / 178 mixture systems; probe: 178 systems / 712 actual donor histories

Standardized MSE ↓; train-fitted log-factor probe R²

Same-donor readers and factor accessibility — A. Reader risk: cold-start and full-mixture populations
PopulationNullPersistentDecodeOracle
Cold start: 100 systems / 600 cases0.1778370.1136440.1112820.061882
Full mixture: 178 systems0.8716820.8097320.8087310.765938

Bold compares direct context and decoded parameters on the cold-start population.

Dashed rules separate cold-start and full-mixture populations.

  • Cold start: 100 systems / 600 casesDecode has the lower cold-start mean here; D-Clean has the opposite interface ordering.

All arms share the common adapter and the same donor identities within each population. Cold-start Decode-minus-Persistent is −0.0023623 [−0.0045732,−0.0004762]. The 100-system cold-start and 178-system full-mixture risks remain separate.

14 / Study Table 29

Same-donor readers and factor accessibility — B. Actual-donor accessibility: 178 systems

Reader: 100 cold-start / 178 mixture systems; probe: 178 systems / 712 actual donor histories

Standardized MSE ↓; train-fitted log-factor probe R²

Same-donor readers and factor accessibility — B. Actual-donor accessibility: 178 systems
Source recipeSeedR^2_mR^2_γR^2_kR^2_k/m
Align + Cross00.39180.17540.20500.4842
Align + Cross10.31120.15370.14210.3774
Align + Cross20.35690.16180.23450.5115
FCRL-style temporal0-0.0000-0.0019-0.0180-0.0091
FCRL-style temporal10.0107-0.0073-0.0349-0.0154
FCRL-style temporal20.00300.0013-0.0134-0.0056

Bold marks the actual-donor probe values for Align + Cross.

Dashed rules separate source recipes.

  • Align + CrossThe 178-system probe bank is distinct from the 100-system cold-start risk bank.

Training-only ridge probes use 712 actual donor histories from 178 validation systems with equal system weighting. Targets are log-parameters; k/m is derived from fitted mass and stiffness coordinates. This probe population differs from the cold-start risk population.

15 / Study Table 32

Reader recovery across frozen source families

Spring Align+Cross: 3×3; Structure: 1×3; Poke G1/G2: source 2 with 2/3 readers

M1/M2 mean MSE; M1-minus-M2 gain ↑

Reader recovery across frozen source families
FamilySources × readersM_1M_2GainWeighted gain
Align + Cross3×30.6185480.5695390.0490090.050589
Structure1×30.6316070.637819-0.006213-0.006971
Poke G11×20.8098940.812698-0.002805-0.001495
Poke G21×30.7918850.798727-0.006841-0.006443

Bold marks reader gains for Spring Align + Cross and the two Poke boundary cases.

Dashed rules separate SpringWorld and PokeWorld source families.

  • Align + CrossPositive favors M₂. The shared registered architecture does not give exactly matched active capacity.
  • Poke G1Negative Poke gains favor M₁; the successful Spring route is not a universal result.

Gain is M₁−M₂; positive favors persistent conditioning. These are descriptive development means, with no interval inferred from run averages. Active capacity is not exactly matched. Poke G1/G2 use source 2 and retain negative gains under both readout variants.

16 / Study Table 34

Context amplitude and donor identity

Development checkpoints; standard and physically weighted routes

Fixed-reader MSE ↓

Context amplitude and donor identity
GroupSourceDonors=00.250.50.751
A. Standard M_2; displayed in Figure fig:v3-m1m2Align + CrossOwn0.6333050.6074330.5850970.5702880.569539
A. Standard M_2; displayed in Figure fig:v3-m1m2Align + CrossCross-system0.6333050.6338710.6449630.6700360.712522
A. Standard M_2; displayed in Figure fig:v3-m1m2StructureOwn0.6419280.6384040.6363440.6363920.637813
A. Standard M_2; displayed in Figure fig:v3-m1m2StructureCross-system0.6419280.6393700.6380590.6386250.640407
B. Physically weighted M_2,physAlign + CrossOwn0.6333050.6089850.5859110.5684460.565960
B. Physically weighted M_2,physAlign + CrossCross-system0.6333050.6337130.6434400.6671230.710223
B. Physically weighted M_2,physStructureOwn0.6419280.6387460.6364760.6361440.637209
B. Physically weighted M_2,physStructureCross-system0.6419280.6394250.6376310.6376220.638848

Bold marks the full-amplitude outcomes for own-system and cross-system context in the Align + Cross source.

Dashed rules separate source families and standard versus physically weighted routes.

  • A. Standard M_2; displayed in Figure fig:v3-m1m2 · Align + CrossAt zero amplitude the route is B₀, not M₁. The same frozen weights process all amplitudes within a row.

Own-system and whole-system cross donors pass through the same fitted weights within each source and route. At zero amplitude M₂ equals the frozen recipient-only base B₀, not M₁. Standard and physically weighted routes are distinct comparisons.

Practical implications

What to take from this setting.

Check that the right context actually matters.

Pair factor probes with Correct / Null / Wrong context interventions on the actual downstream prediction route.

Hold the predictor, current query, actions and target fixed. A zero-code ablation alone is not a factor-specific test.

When value stalls, examine the reader.

Compare direct context, decoded factors and a conditioned residual reader before deciding the representation needs to be retrained.

Development evidence; active capacity is not perfectly matched. The best interface depends on the environment.