Disjoint episodes, positions, actions and outcomes.
Experience → prediction
The task.
Prior episodes + current query.
→ Ego-action prediction.
Evaluated instance
The comparison.
Native multi-episode AD action learner · Align only.
Controls. Matched history slot + variance/covariance regularization.
Interpreting the evidence
What stays fixed in the test.
The grounded_coord_simple layout and partner relation are explicit. A100 and VC100 share the matched training/evaluation recipe; history budgets are analyzed separately.
What this setting establishes.
The application measures partner-behavior readout, then cooperation across history budgets. The completed three-seed record separates familiar from held-out-development partners.
Scope of the evidence
Held-out-development cooperation is not a stable positive headline: A100−VC100 is −1.111 ± 11.437 points across source seeds. The displayed native scene is not a trained-policy success demonstration.
Full result record
Measurements and comparisons.
Reported results are kept with their own populations and conditions. Development, held-out and sealed comparisons remain separate.
01 / Study Table 40
Align − H+VC: partner readout and cooperation differences
3 terminal source seeds 4200/4201/4202; familiar and held-out-development partners separated
Paired Align minus H+VC differences: negative probe-MSE changes and positive return changes are favorable.
Align − H+VC: partner readout and cooperation differences
History
Δ familiar probe MSE
Δ familiar return
Δ held-out dev. return
25%
-0.00567 ± 0.00810
1.489 ± 2.335
-2.222 ± 15.207
50%
-0.01344 ± 0.00728
0.389 ± 1.110
18.444 ± 20.659
100%
-0.02120 ± 0.00610
1.889 ± 2.175
-1.111 ± 11.437
BBold pairs full-history familiar and held-out-development return effects.
100%The familiar effect is positive in all three seeds; held-out-development effects have mixed signs. ± is sample SD, not a confidence interval.
Paired Align-minus-H+VC differences are means ± sample SD across three training seeds. Return averages episodes 6–20 per partner. Full-history familiar effects are positive in every seed; held-out-development effects have mixed signs. The two partner populations remain separate.
02 / Study Table 41
Partner behavior and cooperation across history budgets — A. Behavioral readout and representation
3 terminal source seeds 4200/4201/4202; familiar and held-out-development partners separated
Partner behavior and cooperation across history budgets — A. Behavioral readout and representation
History
Recipe
Familiar MSE ↓
Held-out dev. MSE ↓
Distance ratio
25%
VC
0.06788 ± 0.00667
0.13883 ± 0.04230
1.343 ± 0.061
25%
A
0.06220 ± 0.00974
0.09733 ± 0.01384
1.917 ± 0.071
50%
VC
0.06934 ± 0.00811
0.14091 ± 0.02324
1.408 ± 0.019
50%
A
0.05590 ± 0.00736
0.08966 ± 0.01240
2.329 ± 0.326
100%
VC
0.07516 ± 0.00791
0.12395 ± 0.02857
1.458 ± 0.036
100%
A
0.05395 ± 0.01106
0.10631 ± 0.01046
2.269 ± 0.206
100%
Random
0.07978 ± 0.01423
0.14836 ± 0.03203
1.460 ± 0.066
BBold marks Align’s behavioral-probe errors within each history budget.
Dashed rules separate budgets and the probe-only Random control.
100% · RandomRandom has probe coverage only. Familiar and held-out-development probe populations stay separate.
Means ± sample SD across three terminal checkpoints. A is Align; VC is H+VC. Familiar and held-out-development probe populations remain separate. Distance ratios use the familiar behavioral-signature bank; Random has probes only, without return evaluation.
03 / Study Table 41
Partner behavior and cooperation across history budgets — B. Cooperation returns
3 terminal source seeds 4200/4201/4202; familiar and held-out-development partners separated
Partner behavior and cooperation across history budgets — B. Cooperation returns
History
Recipe
Familiar partners
Held-out dev. partners
25%
VC
15.778 ± 1.262
-1.333 ± 4.372
25%
A
17.267 ± 1.768
-3.556 ± 14.852
50%
VC
15.200 ± 0.200
-19.556 ± 2.341
50%
A
15.589 ± 0.916
-1.111 ± 18.407
100%
VC
14.044 ± 1.500
-19.111 ± 8.572
100%
A
15.933 ± 0.742
-20.222 ± 3.672
BBold keeps both full-history policies and partner populations in view.
Dashed rules separate offline history budgets.
100% · AReturns come from frozen-policy episodes with no optimizer updates. The held-out-development population is not a stable positive generalization result.
Means ± sample SD across three training seeds. Frozen policies run without optimizer updates; return averages episodes 6–20 within each partner, then partners within each cohort. The full-history held-out-development Align-minus-H+VC difference is −1.111 ± 11.437, with mixed seed-level effects.