Full result recordMeasurements and comparisons.
Reported results are kept with their own populations and conditions. Development, held-out and sealed comparisons remain separate.
01 / Study Table 11What the observed histories reveal
Rendered-history versus privileged numerical-state inputs
Factor probe R² ↑
BBold compares drag recoverability under the two input conditions.
The dashed rule separates rendered observations from privileged numerical-state access.
- Privileged dynamicsPrivileged inputs do not represent what the rendered-observation model receives.
Rendered histories and privileged numerical dynamics are different input conditions. Means and source SD are retained where reported; privileged recoverability does not describe the information directly available to the rendered-observation model.
02 / Study Table 12Relation semantics redirects factor accessibility
400 validation systems; ridge and MLP probes trained only on training systems
Frozen-code probe R² ↑; G2-minus-G1 probe MSE
BBold marks the higher G1/G2 accessibility within each probe × factor pair.
The dashed rule separates linear and nonlinear probe families.
- Ridge · log mThe mass/drag trade-off is present in both probe families; these are accessibility results.
Training-only ridge and MLP probes evaluate frozen code centroids on 400 validation systems. Error contrasts are G2 minus G1 with paired-system 95% intervals; positive favors G1. Both probe families show the mass-to-drag accessibility shift.
03 / Study Table 14History diversity and pairing reliability — A. Independent-history budget
3 sources; validation and frozen confirmation kept separate
Within/between-system distance; ratio; partial drag correlation
BBold compares the R8 endpoints under correct versus random relations.
The dashed rule separates relation conditions.
- CorrectR2/R4/R8 vary history diversity; this differs from changing pair reliability at fixed R8.
Means ± source SD across three training seeds on validation systems. Codes use training-only standardization; partial drag correlation controls mass and stiffness. R2/R4/R8 change the number of available independent realizations.
04 / Study Table 14History diversity and pairing reliability — B. Relation fidelity
3 sources; validation and frozen confirmation kept separate
Within/between-system distance; ratio; partial drag correlation
BBold marks the confirmation endpoints at zero and full pair reliability.
- 1Validation and confirmation occupy separate columns. These formation measures do not guarantee reader gains.
Means ± source SD across three training seeds. Correct-pair fraction changes at fixed R8 history budget. Validation and frozen confirmation remain separate; stronger organization does not by itself establish a reader gain.
05 / Study Table 15Factor-selective geometry under relation refinement
Validation and confirmation banks; 3 sources; split-disjoint exact factor tuples
Partial factor-distance correlations
BBold marks G1’s mass geometry and G2’s drag geometry on the two evaluation banks.
- G_1Sharing more factors redirects the representation; the table is not a monotonic information-accumulation ranking.
Partial factor geometry is reported as three-source mean ± sample SD. Exact factor tuples are split-disjoint, while factor levels are shared. Validation and confirmation remain separate; additional sharing need not preserve the earlier factor emphasis.
06 / Study Table 16Objective factorial: geometry and native prediction
400 frozen confirmation systems; 3 training seeds
Partial geometry; between/within ratio; learned-embedding h16 loss
BBold marks strong drag organization under Align-containing recipes and the lower native h16 loss under Cross.
- CrossGeometry and prediction need not select the same objective. No common ranking is imposed across columns.
Three-source means ± sample SD on 400 frozen confirmation systems. The two-view batch, predictor, dimensions and training budget are held fixed. Factor geometry and native h16 prediction loss measure different effects of the objectives.
07 / Study Table 30Fresh-reader value and Oracle component references — A. Aggregate h16 MSE
400 development systems; 3 sources × 3 readers
h16 standardized MSE; comparator-minus-Persistent contrasts
BBold marks Null and Persistent, the primary context-value comparison.
- MSEThe adjusted Null-minus-Persistent interval crosses zero. Oracle is a fitted true-parameter reader, not an optimal-risk bound.
Each arm has a separately fitted reader on the same 400-system development bank and three-source by three-reader grid. The adjusted Null-minus-Persistent interval crosses zero. Oracle is a fitted true-parameter reader.
08 / Study Table 30Fresh-reader value and Oracle component references — B. Prespecified aggregate contrasts
400 development systems; 3 sources × 3 readers
h16 standardized MSE; comparator-minus-Persistent contrasts
BBold keeps the Null-versus-Persistent boundary and the Shuffled-versus-Persistent contrast visible together.
- Null–PersistentIntervals are adjusted across the three prespecified contrasts; readers in all arms are fitted separately.
Intervals adjust for the three prespecified aggregate contrasts and condition on the fitted source/reader grid. The Null-minus-Persistent interval crosses zero; the Shuffled-minus-Persistent interval is positive. Shuffled is separately trained, rather than a fixed-reader intervention.
09 / Study Table 30Fresh-reader value and Oracle component references — C. Oracle component reference
400 development systems; 3 sources × 3 readers
h16 standardized MSE; comparator-minus-Persistent contrasts
BBold marks object-state improvements and the opposite finger-position effect.
- Null–OracleThese component means are not summed to reconstruct the aggregate endpoint.
Component values compare Null against the fitted Oracle reader on the same development bank. Object-state improvement is partly offset by finger-position error. These component means are not summed to reconstruct the aggregate endpoint.
10 / Study Table 31Factor-targeted value and actual-donor probes
Expanded development grid; 3 source seeds × 3 reader seeds; 5,913 actual donor windows
Utility-versus-sensitivity slope; actual-donor probe R²
BBold marks the three source-specific slopes; no source is selected as best.
- Expanded-grid slopeThe pooled 95% interval crosses zero. These directional development slopes do not establish a task-gain law.
The expanded development grid includes the pilot and uses three source and three reader seeds. The pooled slope is .002810 with a 95% system-bootstrap interval [−.004443,.009203]. Its interval crosses zero, so the directional pattern does not establish a task-gain law.
11 / Study Table 31Factor-targeted value and actual-donor probes
Expanded development grid; 3 source seeds × 3 reader seeds; 5,913 actual donor windows
Utility-versus-sensitivity slope; actual-donor probe R²
BBold marks G1 mass access and G2 drag access in actual donor windows.
- G1This donor-window protocol differs from the frozen-centroid probe comparison.
Training-fitted probes use actual donor windows and average source seeds. Their protocol differs from the frozen-centroid accessibility study. G1 favors mass readout and G2 favors drag readout; this alone does not establish targeted downstream gains.
12 / Study Table 32Reader recovery across frozen source families
Spring Align+Cross: 3×3; Structure: 1×3; Poke G1/G2: source 2 with 2/3 readers
M1/M2 mean MSE; M1-minus-M2 gain ↑
BBold marks reader gains for Spring Align + Cross and the two Poke boundary cases.
Dashed rules separate SpringWorld and PokeWorld source families.
- Align + CrossPositive favors M₂. The shared registered architecture does not give exactly matched active capacity.
- Poke G1Negative Poke gains favor M₁; the successful Spring route is not a universal result.
Gain is M₁−M₂; positive favors persistent conditioning. These are descriptive development means, with no interval inferred from run averages. Active capacity is not exactly matched. Poke G1/G2 use source 2 and retain negative gains under both readout variants.
13 / Study Table 33Pair reliability versus downstream reader budget — A. Organization, accessibility and pooled reader value
400-system development bank; 3 frozen sources per α; 3 readers per source
B/W; log-drag R²; M1-minus-M2 gain in MSE ×10⁻³
BBold contrasts strong formation at full reliability with the nonpositive 5k reader gains.
- 1The gain is M₁−M₂; a negative value favors M₁. Better accessibility is not itself a positive downstream result.
Gains are M₁−M₂ in normalized MSE ×10⁻³. Both budgets use the same three frozen sources per α and the same 400-system bank. Intervals use paired source/system bootstrapping and are descriptive with three sources; all 5k mean gains are nonpositive.
14 / Study Table 33Pair reliability versus downstream reader budget — B. Source-specific reader value
400-system development bank; 3 frozen sources per α; 3 readers per source
B/W; log-drag R²; M1-minus-M2 gain in MSE ×10⁻³
BBold marks all source outcomes at the larger reader budget, including the lone positive cell.
The dashed rule separates reader fitting budgets.
- 0Three readers are averaged within each source. Cells are descriptive source outcomes, not separate significance tests.
Three readers are averaged within each source; source outcomes remain separate. Gains use normalized MSE ×10⁻³, with positive favoring M₂. The two reader budgets use the same sources and observations, rather than additional source training.
15 / Study Table 39Complete follow-up outcomes
Separate configurations and populations shown row by row
MSE ↓, except FHN relative L₂ ↓
BBold identifies each follow-up configuration; heterogeneous numerical cells are deliberately not ranked.
Dashed rules separate experiments with different populations and fitted-model coverage.
All completed follow-ups from the configuration round are retained, including null and adverse results. Each row has its own population and model coverage. Paired intervals condition on the reported fitted models; the one-source development rows do not replace the original multi-source comparisons.
16 / Study Table 5.3Fixed-predictor use follows relation semantics
144 PokeWorld validation systems for directional contrasts; reliability assay kept separate
Dimensionless counterfactual projection contrast; wrong-drag penalty
BBold marks the semantic-direction effects and the reliability endpoints.
The dashed rule separates two distinct intervention assays.
- Mass counterfactual projection · G1Projection contrasts and wrong-drag penalties have different meanings; their magnitudes are not directly ranked.
The directional projections use native observation-embedding spaces. The reliability assay is a separate fixed-route wrong-drag test; its uncertainty is not given in this prose record. The two assays measure context dependence, rather than downstream task benefit.