SPRIIA study of persistent information
Formation · Use · Value

PokeWorld

What you call “shared” changes what is learned.

Change the relation between interactions and the representation shifts its emphasis between mass, drag and stiffness.

Formation → Usedoes not guarantee more task value
Motion preview
History replay · applied actions · ¼ speedProject-generated recorded histories · action overlays. Source & changes
Across independent experiences

What persists.

G1: mass m; G2: (m, γ); G3: (m, γ, k).

Across distinct realizations

What changes.

The shared-factor rule; separately, correct-pair fraction α at fixed history budget.

Experience → prediction

The task.

Independent image/action history + current query.

→ Future observation embedding.

Evaluated instance

The comparison.

JEPA · Structure, Align, Cross and Align + Cross.

Controls. Relation reliability; relation semantics; objective composition.

Interpreting the evidence

What stays fixed in the test.

The α study fixes donor/query marginals and model settings. In fixed-route use assays, predictor, decoder, current observations, actions, target and horizon remain fixed while the donor changes.

What this setting establishes.

More reliable relations strengthen organization. Adding drag to the shared relation redirects accessibility from mass toward drag; it need not preserve all previously accessible factors.

Scope of the evidence

The original acquisition bank, factorized relation bank and fresh-reader studies are different protocols. Stronger formation and measurable donor use do not imply a monotonic downstream gain.

Full result record

Measurements and comparisons.

Reported results are kept with their own populations and conditions. Development, held-out and sealed comparisons remain separate.

01 / Study Table 11

What the observed histories reveal

Rendered-history versus privileged numerical-state inputs

Factor probe R² ↑

What the observed histories reveal
Input readout (R^2)MassDragStiffness
Rendered history0.270 ± 0.0260.574 ± 0.0090.149 ± 0.034
Privileged dynamics0.7620.9370.760

Bold compares drag recoverability under the two input conditions.

The dashed rule separates rendered observations from privileged numerical-state access.

  • Privileged dynamicsPrivileged inputs do not represent what the rendered-observation model receives.

Rendered histories and privileged numerical dynamics are different input conditions. Means and source SD are retained where reported; privileged recoverability does not describe the information directly available to the rendered-observation model.

02 / Study Table 12

Relation semantics redirects factor accessibility

400 validation systems; ridge and MLP probes trained only on training systems

Frozen-code probe R² ↑; G2-minus-G1 probe MSE

Relation semantics redirects factor accessibility
ProbeTargetG1 R^2G2 R^2MSE contrast95% interval
Ridgelog m0.52710.07090.1492[0.1231,0.1758]
Ridgeγ0.30750.9249-0.7704[-0.8572,-0.6869]
MLPlog m0.52720.02730.1635[0.1303,0.1974]
MLPγ0.12910.9644-1.0422[-1.1435,-0.9438]

Bold marks the higher G1/G2 accessibility within each probe × factor pair.

The dashed rule separates linear and nonlinear probe families.

  • Ridge · log mThe mass/drag trade-off is present in both probe families; these are accessibility results.

Training-only ridge and MLP probes evaluate frozen code centroids on 400 validation systems. Error contrasts are G2 minus G1 with paired-system 95% intervals; positive favors G1. Both probe families show the mass-to-drag accessibility shift.

03 / Study Table 14

History diversity and pairing reliability — A. Independent-history budget

3 sources; validation and frozen confirmation kept separate

Within/between-system distance; ratio; partial drag correlation

History diversity and pairing reliability — A. Independent-history budget
RelationBudgetWithin ↓Between ↑Ratio ↑Partial ρ_γ ↑
CorrectR278.35 ± 4.6729.38 ± 1.780.375 ± 0.0170.095 ± 0.022
CorrectR455.66 ± 5.1743.32 ± 2.820.781 ± 0.0620.361 ± 0.036
CorrectR828.09 ± 1.93104.34 ± 2.013.729 ± 0.3130.773 ± 0.009
RandomR270.80 ± 8.7417.28 ± 2.030.244 ± 0.002-0.001 ± 0.017
RandomR483.68 ± 16.9120.39 ± 3.640.245 ± 0.008-0.010 ± 0.016
RandomR8126.87 ± 0.8934.18 ± 0.390.269 ± 0.004-0.006 ± 0.004

Bold compares the R8 endpoints under correct versus random relations.

The dashed rule separates relation conditions.

  • CorrectR2/R4/R8 vary history diversity; this differs from changing pair reliability at fixed R8.

Means ± source SD across three training seeds on validation systems. Codes use training-only standardization; partial drag correlation controls mass and stiffness. R2/R4/R8 change the number of available independent realizations.

04 / Study Table 14

History diversity and pairing reliability — B. Relation fidelity

3 sources; validation and frozen confirmation kept separate

Within/between-system distance; ratio; partial drag correlation

History diversity and pairing reliability — B. Relation fidelity
Correct-pair fraction αValidation B/W ratioValidation partial drag correlationConfirmation B/W ratioConfirmation partial drag correlation
00.272 ± 0.0060.061 ± 0.0310.260 ± 0.0060.031 ± 0.026
0.50.976 ± 0.0510.243 ± 0.0560.880 ± 0.0470.264 ± 0.063
13.749 ± 0.0580.781 ± 0.0274.391 ± 0.1560.796 ± 0.023

Bold marks the confirmation endpoints at zero and full pair reliability.

  • 1Validation and confirmation occupy separate columns. These formation measures do not guarantee reader gains.

Means ± source SD across three training seeds. Correct-pair fraction changes at fixed R8 history budget. Validation and frozen confirmation remain separate; stronger organization does not by itself establish a reader gain.

05 / Study Table 15

Factor-selective geometry under relation refinement

Validation and confirmation banks; 3 sources; split-disjoint exact factor tuples

Partial factor-distance correlations

Factor-selective geometry under relation refinement
ConditionValidation: massValidation: dragValidation: stiffnessConfirmation: massConfirmation: dragConfirmation: stiffness
Structure0.069 ± 0.0080.194 ± 0.0070.037 ± 0.0120.057 ± 0.0100.190 ± 0.0020.018 ± 0.008
G_10.419 ± 0.0060.013 ± 0.0090.103 ± 0.0040.489 ± 0.0090.017 ± 0.0060.084 ± 0.003
G_20.015 ± 0.0000.720 ± 0.009-0.013 ± 0.001-0.004 ± 0.0010.732 ± 0.012-0.012 ± 0.001
G_30.064 ± 0.0130.427 ± 0.0870.067 ± 0.0350.077 ± 0.0350.413 ± 0.0820.059 ± 0.023
Random0.030 ± 0.0070.036 ± 0.0140.054 ± 0.0080.022 ± 0.0090.013 ± 0.0070.044 ± 0.006

Bold marks G1’s mass geometry and G2’s drag geometry on the two evaluation banks.

  • G_1Sharing more factors redirects the representation; the table is not a monotonic information-accumulation ranking.

Partial factor geometry is reported as three-source mean ± sample SD. Exact factor tuples are split-disjoint, while factor levels are shared. Validation and confirmation remain separate; additional sharing need not preserve the earlier factor emphasis.

06 / Study Table 16

Objective factorial: geometry and native prediction

400 frozen confirmation systems; 3 training seeds

Partial geometry; between/within ratio; learned-embedding h16 loss

Objective factorial: geometry and native prediction
ConditionMass ↑Drag ↑Stiff. ↑Ratio ↑h16 ↓
Structure0.058 ± 0.0050.191 ± 0.0050.070 ± 0.0110.299 ± 0.0060.404 ± 0.033
Align0.012 ± 0.0090.781 ± 0.0040.016 ± 0.0054.203 ± 0.2890.386 ± 0.019
Cross0.107 ± 0.0090.229 ± 0.0090.106 ± 0.0020.362 ± 0.0090.373 ± 0.005
Align + Cross0.006 ± 0.0040.788 ± 0.0010.011 ± 0.0024.411 ± 0.1420.381 ± 0.010

Bold marks strong drag organization under Align-containing recipes and the lower native h16 loss under Cross.

  • CrossGeometry and prediction need not select the same objective. No common ranking is imposed across columns.

Three-source means ± sample SD on 400 frozen confirmation systems. The two-view batch, predictor, dimensions and training budget are held fixed. Factor geometry and native h16 prediction loss measure different effects of the objectives.

07 / Study Table 30

Fresh-reader value and Oracle component references — A. Aggregate h16 MSE

400 development systems; 3 sources × 3 readers

h16 standardized MSE; comparator-minus-Persistent contrasts

Fresh-reader value and Oracle component references — A. Aggregate h16 MSE
ArmNullPersistentShuffledOracle
MSE0.9629300.9623770.9637440.961003

Bold marks Null and Persistent, the primary context-value comparison.

  • MSEThe adjusted Null-minus-Persistent interval crosses zero. Oracle is a fitted true-parameter reader, not an optimal-risk bound.

Each arm has a separately fitted reader on the same 400-system development bank and three-source by three-reader grid. The adjusted Null-minus-Persistent interval crosses zero. Oracle is a fitted true-parameter reader.

08 / Study Table 30

Fresh-reader value and Oracle component references — B. Prespecified aggregate contrasts

400 development systems; 3 sources × 3 readers

h16 standardized MSE; comparator-minus-Persistent contrasts

Fresh-reader value and Oracle component references — B. Prespecified aggregate contrasts
ContrastDifferenceAdjusted 95% interval
Null–Persistent0.000553[-0.000419,0.001641]
Null–Oracle0.001927[0.000427,0.003615]
Shuffled–Persistent0.001367[0.000274,0.002498]

Bold keeps the Null-versus-Persistent boundary and the Shuffled-versus-Persistent contrast visible together.

  • Null–PersistentIntervals are adjusted across the three prespecified contrasts; readers in all arms are fitted separately.

Intervals adjust for the three prespecified aggregate contrasts and condition on the fitted source/reader grid. The Null-minus-Persistent interval crosses zero; the Shuffled-minus-Persistent interval is positive. Shuffled is separately trained, rather than a fixed-reader intervention.

09 / Study Table 30

Fresh-reader value and Oracle component references — C. Oracle component reference

400 development systems; 3 sources × 3 readers

h16 standardized MSE; comparator-minus-Persistent contrasts

Fresh-reader value and Oracle component references — C. Oracle component reference
MeasurementObject positionObject velocityFinger positionFinger velocity
Null–Oracle0.0028090.006168-0.0030090.001740

Bold marks object-state improvements and the opposite finger-position effect.

  • Null–OracleThese component means are not summed to reconstruct the aggregate endpoint.

Component values compare Null against the fitted Oracle reader on the same development bank. Object-state improvement is partly offset by finger-position error. These component means are not summed to reconstruct the aggregate endpoint.

10 / Study Table 31

Factor-targeted value and actual-donor probes

Expanded development grid; 3 source seeds × 3 reader seeds; 5,913 actual donor windows

Utility-versus-sensitivity slope; actual-donor probe R²

Factor-targeted value and actual-donor probes
Source seed012
Expanded-grid slope.003838.003307.001284

Bold marks the three source-specific slopes; no source is selected as best.

  • Expanded-grid slopeThe pooled 95% interval crosses zero. These directional development slopes do not establish a task-gain law.

The expanded development grid includes the pilot and uses three source and three reader seeds. The pooled slope is .002810 with a 95% system-bootstrap interval [−.004443,.009203]. Its interval crosses zero, so the directional pattern does not establish a task-gain law.

11 / Study Table 31

Factor-targeted value and actual-donor probes

Expanded development grid; 3 source seeds × 3 reader seeds; 5,913 actual donor windows

Utility-versus-sensitivity slope; actual-donor probe R²

Factor-targeted value and actual-donor probes
Recipelog mlogγlog klog(k/m)
G1.206980.124610.081871.244367
G2.070750.788914.074915.146439

Bold marks G1 mass access and G2 drag access in actual donor windows.

  • G1This donor-window protocol differs from the frozen-centroid probe comparison.

Training-fitted probes use actual donor windows and average source seeds. Their protocol differs from the frozen-centroid accessibility study. G1 favors mass readout and G2 favors drag readout; this alone does not establish targeted downstream gains.

12 / Study Table 32

Reader recovery across frozen source families

Spring Align+Cross: 3×3; Structure: 1×3; Poke G1/G2: source 2 with 2/3 readers

M1/M2 mean MSE; M1-minus-M2 gain ↑

Reader recovery across frozen source families
FamilySources × readersM_1M_2GainWeighted gain
Align + Cross3×30.6185480.5695390.0490090.050589
Structure1×30.6316070.637819-0.006213-0.006971
Poke G11×20.8098940.812698-0.002805-0.001495
Poke G21×30.7918850.798727-0.006841-0.006443

Bold marks reader gains for Spring Align + Cross and the two Poke boundary cases.

Dashed rules separate SpringWorld and PokeWorld source families.

  • Align + CrossPositive favors M₂. The shared registered architecture does not give exactly matched active capacity.
  • Poke G1Negative Poke gains favor M₁; the successful Spring route is not a universal result.

Gain is M₁−M₂; positive favors persistent conditioning. These are descriptive development means, with no interval inferred from run averages. Active capacity is not exactly matched. Poke G1/G2 use source 2 and retain negative gains under both readout variants.

13 / Study Table 33

Pair reliability versus downstream reader budget — A. Organization, accessibility and pooled reader value

400-system development bank; 3 frozen sources per α; 3 readers per source

B/W; log-drag R²; M1-minus-M2 gain in MSE ×10⁻³

Pair reliability versus downstream reader budget — A. Organization, accessibility and pooled reader value
αB/WR^2_logγ1k gain5k gain5k 95% interval
00.3200.101-2.512-1.890[-8.366,+3.116]
0.51.0420.254+0.061-2.942[-8.255,+1.948]
13.7960.797-4.469-6.291[-12.005,-0.328]

Bold contrasts strong formation at full reliability with the nonpositive 5k reader gains.

  • 1The gain is M₁−M₂; a negative value favors M₁. Better accessibility is not itself a positive downstream result.

Gains are M₁−M₂ in normalized MSE ×10⁻³. Both budgets use the same three frozen sources per α and the same 400-system bank. Intervals use paired source/system bootstrapping and are descriptive with three sources; all 5k mean gains are nonpositive.

14 / Study Table 33

Pair reliability versus downstream reader budget — B. Source-specific reader value

400-system development bank; 3 frozen sources per α; 3 readers per source

B/W; log-drag R²; M1-minus-M2 gain in MSE ×10⁻³

Pair reliability versus downstream reader budget — B. Source-specific reader value
αReader updatesSource 0Source 1Source 2
01000-3.463-3.017-1.057
0.51000+1.222-0.015-1.023
11000-3.310-4.186-5.910
05000-4.739-1.617+0.686
0.55000-2.123-3.436-3.267
15000-5.288-11.272-2.314

Bold marks all source outcomes at the larger reader budget, including the lone positive cell.

The dashed rule separates reader fitting budgets.

  • 0Three readers are averaged within each source. Cells are descriptive source outcomes, not separate significance tests.

Three readers are averaged within each source; source outcomes remain separate. Gains use normalized MSE ×10⁻³, with positive favoring M₂. The two reader budgets use the same sources and observations, rather than additional source training.

15 / Study Table 39

Complete follow-up outcomes

Separate configurations and populations shown row by row

MSE ↓, except FHN relative L₂ ↓

Complete follow-up outcomes
ConfigurationCoverageNumerical comparisonFinding / scope
JEPA Collision: J21,994 test episodes; 3 J2 seeds, 1/controlJ2 0.184110 ± 0.005443; Structure 0.214314Matched-only retest. Four arms and fixed-model interval: Table tab:jepa-collision-followup.
CoPhyNet Collision: Cross4,000 development recipients; seed 0Cross 0.203240; prior SPRII 0.210671Near Native; paired interval against Native includes zero.
CPC Balls: Align2,000 development recipients; seed 0Align 1.392711; prior SPRII 1.415352; Structure 1.387643Lower than prior SPRII, higher than Structure. Full contrasts are indexed in the package.
FHN: tuned SPRIIReport ICs 15/24/45; source 42; selection IC 5 excludedID K=4,H=50: SPRII 0.020132; extra-50k NOD 0.021522Paired-system interval includes zero; other eight split/horizon cells have higher error. Extra-budget control.
Poke object-only: R8 / G2400 development systems; source 0, reader 0; H=16R8: M2 0.901091; M1 0.902588; recipient-only 0.911791. M2-M1: -0.001497 [-0.011855,0.008400]. G2: M2 0.913197; M1 0.908531; recipient-only 0.908021Four-dimensional object-state error; M1 and M2 selected independently.
CoPhyNet Blocktower8,088 development recipients; seed 0SPRII 0.088180; Native/Structure 0.087114; Random 0.086230. SPRII-Random: 0.001951 [0.000668,0.003184]Paired interval versus Structure includes zero; Random has lower mean error.

Bold identifies each follow-up configuration; heterogeneous numerical cells are deliberately not ranked.

Dashed rules separate experiments with different populations and fitted-model coverage.

All completed follow-ups from the configuration round are retained, including null and adverse results. Each row has its own population and model coverage. Paired intervals condition on the reported fitted models; the one-source development rows do not replace the original multi-source comparisons.

16 / Study Table 5.3

Fixed-predictor use follows relation semantics

144 PokeWorld validation systems for directional contrasts; reliability assay kept separate

Dimensionless counterfactual projection contrast; wrong-drag penalty

Fixed-predictor use follows relation semantics
AssayConditionReported effect95% interval
Mass counterfactual projectionG10.0420[0.0312,0.0539]
Drag counterfactual projectionG20.1722[0.1448,0.1983]
Wrong-drag fixed-route penaltyCorrect-pair fraction 00.00033Not reported in this prose
Wrong-drag fixed-route penaltyCorrect-pair fraction 0.50.00979Not reported in this prose
Wrong-drag fixed-route penaltyCorrect-pair fraction 10.01192Not reported in this prose

Bold marks the semantic-direction effects and the reliability endpoints.

The dashed rule separates two distinct intervention assays.

  • Mass counterfactual projection · G1Projection contrasts and wrong-drag penalties have different meanings; their magnitudes are not directly ranked.

The directional projections use native observation-embedding spaces. The reliability assay is a separate fixed-route wrong-drag test; its uncertainty is not given in this prose record. The two assays measure context dependence, rather than downstream task benefit.

Practical implications

What to take from this setting.

Choose the relation for the information you need.

Inspect each task-relevant factor. A more specific relation need not preserve everything the earlier relation captured.

The change in accessibility is measured. Task-aware relation selection is a promising strategy, not a demonstrated universal gain.

Measure where persistent information should help.

Report horizon-specific and component-specific errors alongside the aggregate task score.

Longer-horizon training or reweighting outputs are candidate interventions; these analyses do not demonstrate their benefit.