Full result recordMeasurements and comparisons.
Reported results are kept with their own populations and conditions. Development, held-out and sealed comparisons remain separate.
01 / Study Table 1Sealed cold-start prediction
256 sealed systems; 3 sources × 3 readers
Standardized state-increment MSE ↓
BBold marks SPRII’s primary mean and the two paired error reductions.
- SPRII (Align + Cross)The absolute means and paired reductions use the same sealed population; uncertainty has the definitions shown above.
Source means ± sample SD after averaging readers. Error reductions use paired physical-system 95% intervals, conditional on the fitted models; the sealed population is separate from the development studies.
02 / Study Table 6Sealed prediction: populations and paired contrasts
Primary: 256 systems; secondary: 36; pooled: 292; 3 sources × 3 readers
Standardized MSE ↓; reduction versus comparator ↑
BBold marks the SPRII mean on the prespecified primary population.
- Align + CrossPrimary, secondary and pooled cohorts are separate; the primary result is the emphasized comparison.
Six queries per system; readers are averaged within source. Primary MSE intervals resample physical systems and condition on the fitted models. The secondary and system-weighted pooled cohorts are reported separately.
03 / Study Table 6Sealed prediction: populations and paired contrasts
Primary: 256 systems; secondary: 36; pooled: 292; 3 sources × 3 readers
Standardized MSE ↓; reduction versus comparator ↑
BBold marks paired contrasts on the primary 256-system population.
Dashed rules separate evaluation populations.
Each contrast is paired within its stated population. MSE gaps and percentage reductions use 95% physical-system intervals conditional on the fitted models. Primary, secondary and pooled results are distinct comparisons.
04 / Study Table 8Adapted methods under matched history access — A. Spring: original reader
Development cold start; 100 systems / 600 cases; 3 sources × 3 readers
Standardized MSE ↓
BBold compares matched-context errors under the original reader.
- Original · Compact TDS (d=0)Compact TDS has no persistent channel; its three context conditions therefore coincide.
Sources and readers are equally weighted. Fresh Null is fitted separately; Fixed Wrong replaces context through the unchanged Matched reader. The compact TDS arm has no persistent channel, so its context conditions coincide.
05 / Study Table 8Adapted methods under matched history access — B. Spring: common adapter
Development cold start; 100 systems / 600 cases; 3 sources × 3 readers
Standardized MSE ↓
BBold compares matched-context errors under the common adapter.
- Common adapter · FCRL-styleThese development results are separate from both the original-reader comparison and the sealed benchmark.
Both source families use the common downstream adapter, with three sources and three readers. These development results remain separate from the original-reader protocol and the sealed task comparison. A dash denotes an unavailable intervention.
06 / Study Table 17Align and Cross contribute differently — A. Three-source component mean (reader 0)
Task: 100 development systems / 600 cases, reader 0; geometry: 178 systems × 4 histories; 3 sources
Cold-start MSE; frozen-source geometry and partial correlations
BBold marks Align and Align + Cross, the two alignment-containing task recipes.
- MeanComponent means average three sources with reader seed 0; this is a development comparison.
Means equally average three sources, each with reader seed 0 under the original 10,000-update recipe. The 100-system cold-start development population is separate from both the geometry bank and the sealed prediction benchmark.
07 / Study Table 17Align and Cross contribute differently — B. Frozen-source geometry
Task: 100 development systems / 600 cases, reader 0; geometry: 178 systems × 4 histories; 3 sources
Cold-start MSE; frozen-source geometry and partial correlations
BBold marks the between/within ratios of the alignment-containing recipes.
- Align + CrossGeometry columns are distinct diagnostics; larger code separation does not by itself establish task benefit.
Frozen, training-standardized source codes use 178 systems and four histories per system; no reader is fitted for this measurement. Values equally average three sources. Ratio averages source-wise between/within ratios; partial correlations control the other factors.
08 / Study Table 18Development reference learners and physical probes — A. Single-source reference (source 0, reader 0)
Prediction: source 0 / reader 0, 100 cold-start or 178 mixture systems; probe: 100 validation systems / 400 histories
Standardized MSE ↓; log-factor probe R² ↑
BBold marks the lowest cold-start mean within each of the controlled-JEPA and supervised-reference groups.
Dashed rules separate learner groups with different training objectives.
- Supervised reference encoders · DeepSetsReference encoders use their own supervised training recipes; this is not one globally matched leaderboard.
Single-source, single-reader development comparisons use matched standardized MSE. Cold start has 100 systems; the mixture has 178. Supervised reference encoders retain their own training recipes, so the groups do not form one globally matched leaderboard.
09 / Study Table 18Development reference learners and physical probes — B. Frozen-source physical probes
Prediction: source 0 / reader 0, 100 cold-start or 178 mixture systems; probe: 100 validation systems / 400 histories
Standardized MSE ↓; log-factor probe R² ↑
BBold marks factor accessibility from Align + Cross in this frozen-source comparison.
- Align + CrossProbe R² measures accessibility, not fixed-predictor use or downstream value.
Training-fitted ridge probes evaluate log factors using 400 histories from 100 validation systems; hyperparameters are chosen using training data only. These frozen-source probes measure accessibility, separately from prediction error and the sealed benchmark.
10 / Study Table 20Donor physics structures fixed predictive responses
64 development systems; 3 frozen sources; reader 0
Valley depth = off-diagonal minus diagonal h16 MSE ↑
BBold marks the pooled matched-factor advantages and their paired intervals.
Dashed rules separate physical factors; the scales are not pooled.
Valley depth is off-diagonal minus diagonal h16 error. Three frozen sources and reader 0 are evaluated on 64 development systems. Paired-system 95% intervals condition on the fitted routes; physical factors are assessed separately.
11 / Study Table 20Donor physics structures fixed predictive responses
64 development systems; 3 frozen sources; reader 0
Valley depth = off-diagonal minus diagonal h16 MSE ↑
BBold marks alignment-containing recipes in each factor-specific donor test.
- Align + CrossEach recipe has its own fitted source and reader; these cells are not a causal comparison through one shared reader.
Each objective recipe has its own source and fitted reader. Values are off-diagonal-minus-diagonal h16 errors with paired-system 95% intervals on 64 development systems. Positive values favor matched donor physics; factor scales are kept separate.
12 / Study Table 27History value grows with prediction horizon
178-system mixture; cold/moving strata; 3 sources × 3 readers
Persistent-versus-Null error reduction (%) ↑
BBold marks the shortest and longest reported horizons to make the conditional-value trend visible.
- 1Both query conditions have zero observed recipient transitions. The h1 intervals include zero; longer-horizon results use the same fitted grid.
Persistent and Null readers are fitted separately on the same 178-system grid. Cold and moving queries both have zero observed recipient transitions. Pointwise stratified system-bootstrap 95% intervals condition on the fitted source/reader grid.
13 / Study Table 29Same-donor readers and factor accessibility — A. Reader risk: cold-start and full-mixture populations
Reader: 100 cold-start / 178 mixture systems; probe: 178 systems / 712 actual donor histories
Standardized MSE ↓; train-fitted log-factor probe R²
BBold compares direct context and decoded parameters on the cold-start population.
Dashed rules separate cold-start and full-mixture populations.
- Cold start: 100 systems / 600 casesDecode has the lower cold-start mean here; D-Clean has the opposite interface ordering.
All arms share the common adapter and the same donor identities within each population. Cold-start Decode-minus-Persistent is −0.0023623 [−0.0045732,−0.0004762]. The 100-system cold-start and 178-system full-mixture risks remain separate.
14 / Study Table 29Same-donor readers and factor accessibility — B. Actual-donor accessibility: 178 systems
Reader: 100 cold-start / 178 mixture systems; probe: 178 systems / 712 actual donor histories
Standardized MSE ↓; train-fitted log-factor probe R²
BBold marks the actual-donor probe values for Align + Cross.
Dashed rules separate source recipes.
- Align + CrossThe 178-system probe bank is distinct from the 100-system cold-start risk bank.
Training-only ridge probes use 712 actual donor histories from 178 validation systems with equal system weighting. Targets are log-parameters; k/m is derived from fitted mass and stiffness coordinates. This probe population differs from the cold-start risk population.
15 / Study Table 32Reader recovery across frozen source families
Spring Align+Cross: 3×3; Structure: 1×3; Poke G1/G2: source 2 with 2/3 readers
M1/M2 mean MSE; M1-minus-M2 gain ↑
BBold marks reader gains for Spring Align + Cross and the two Poke boundary cases.
Dashed rules separate SpringWorld and PokeWorld source families.
- Align + CrossPositive favors M₂. The shared registered architecture does not give exactly matched active capacity.
- Poke G1Negative Poke gains favor M₁; the successful Spring route is not a universal result.
Gain is M₁−M₂; positive favors persistent conditioning. These are descriptive development means, with no interval inferred from run averages. Active capacity is not exactly matched. Poke G1/G2 use source 2 and retain negative gains under both readout variants.
16 / Study Table 34Context amplitude and donor identity
Development checkpoints; standard and physically weighted routes
Fixed-reader MSE ↓
BBold marks the full-amplitude outcomes for own-system and cross-system context in the Align + Cross source.
Dashed rules separate source families and standard versus physically weighted routes.
- A. Standard M_2; displayed in Figure fig:v3-m1m2 · Align + CrossAt zero amplitude the route is B₀, not M₁. The same frozen weights process all amplitudes within a row.
Own-system and whole-system cross donors pass through the same fitted weights within each source and route. At zero amplitude M₂ equals the frozen recipient-only base B₀, not M₁. Standard and physically weighted routes are distinct comparisons.