SPRIIA study of persistent information
Use · Generality

RH20T

A past robot episode provides task context.

Recorded multimodal interactions test how task-related context enters force, torque and motion forecasts.

Offline forecastswith donor-specificity controls
Motion preview
RH20T dataset example · drawer manipulationDemo: Hao-Shu Fang et al. · CC BY-SA 4.0 · cropped and resized. Source & changes
Across independent experiences

What persists.

Task identity.

Across distinct realizations

What changes.

Different recorded episodes.

Experience → prediction

The task.

Multimodal history + current inputs and actions.

→ Force/torque and TCP motion.

Evaluated instance

The comparison.

Supervised predictor · Cross / Align + Cross.

Controls. Structure; pairing; input-route controls.

Interpreting the evidence

What stays fixed in the test.

Query, target, horizon and actions stay fixed while self, matched and shuffled-task contexts are substituted through the fitted route.

What this setting establishes.

Forecasting and donor-specificity measurements test how task context enters a real multimodal prediction problem.

Scope of the evidence

Same task does not assert identical physical parameters. The 25-task forecasting cohort and 24-task hard-negative cohort differ. Monolithic retains lower matched absolute error on the donor-specificity bank. These are offline forecasts, not closed-loop robot gains.

Full result record

Measurements and comparisons.

Reported results are kept with their own populations and conditions. Development, held-out and sealed comparisons remain separate.

01 / Study Table 21

Offline robot forecasting across input routes — A. Self-context

25 held-out tasks; task-macro means over 3 fitted sources

Normalized force/torque MSE ↓; TCP error in cm ↓

Offline robot forecasting across input routes — A. Self-context
ConditionFT MSE h1FT MSE h4FT MSE h16TCP cm h4TCP cm h16
Native0.63461.11751.49711.96214.0632
Structure0.61991.09931.49851.94443.9645
Align–Indep0.66331.13161.50312.26184.4288
Align–Random0.62321.10641.50982.04544.2568
Cross–SameEp0.61451.08981.49481.96283.9652
Cross–Indep0.62301.09301.49741.90153.9022
Cross–Random0.63071.09611.49501.94763.9505
A+C–SameEp0.64071.12021.50032.09764.3710
A+C–Indep0.66021.12961.50462.26834.4347
A+C–Random0.62251.10051.50682.03594.2575
LD–Native0.59091.08931.48931.72963.8873
LD–Cross–Indep0.57601.06951.48801.59793.7484
LD–A+C–SameEp0.61621.11771.50452.03244.3018
LD–A+C–Random0.60901.12241.50541.86644.1913

Bold compares Structure and independent-episode Cross on self-context TCP forecasts.

The dashed rule separates full-input and LowDim models.

  • Cross–IndepOther conditions and outputs remain visible. Compare each input-eligible condition on its own target scale.

Task-macro means average three fitted sources on 25 held-out tasks. Force/torque uses normalized MSE; TCP uses cm. Self-context results are distinct from matched-donor forecasting, the 24-task specificity bank, and the additional 25-task intervention bank. These are offline forecasts.

02 / Study Table 21

Offline robot forecasting across input routes — B. Matched donor

25 held-out tasks; task-macro means over 3 fitted sources

Normalized force/torque MSE ↓; TCP error in cm ↓

Offline robot forecasting across input routes — B. Matched donor
ConditionFT MSE h1FT MSE h4FT MSE h16TCP cm h4TCP cm h16
Native–––––
Structure0.79911.22651.55122.42934.9521
Align–Indep0.66671.13461.50442.53634.6438
Align–Random0.62551.10841.50992.06224.2735
Cross–SameEp0.67091.14031.52742.25424.7451
Cross–Indep0.66121.12841.52062.13884.5837
Cross–Random0.66651.12751.51592.15774.6035
A+C–SameEp0.64981.12661.50632.40704.5776
A+C–Indep0.66331.13251.50572.57294.6695
A+C–Random0.62501.10261.50702.04464.2689
LD–Native–––––
LD–Cross–Indep0.60391.13061.53111.89284.4714
LD–A+C–SameEp0.61871.12141.50632.07774.3387
LD–A+C–Random0.61081.12401.50571.90494.2297

Bold compares Structure and independent-episode Cross on matched-donor TCP forecasts.

The dashed rule separates full-input and LowDim models.

  • NativeNative has no eligible matched-donor route. A dash denotes missing coverage, not zero error.

Task-macro means average three fitted sources on 25 held-out tasks. Dashes denote unavailable matched inputs. Force/torque MSE and TCP cm remain separate targets. This offline forecasting cohort differs from both the specificity and additional intervention banks.

03 / Study Table 22

Donor specificity and absolute prediction risk — A. Donor specificity

24-task hard-negative bank; equal task weights; 3 fitted sources

Hard-minus-matched specificity; FT MSE and TCP cm risk

Donor specificity and absolute prediction risk — A. Donor specificity
EndpointA+C specificity [95% CI]Monolithic specificity [95% CI]A+C-minus-Monolithic specificity [95% CI]
FT h4 MSE0.00260 [0.00107,0.00429]0.00030 [-0.00050,0.00134]0.00230 [0.00073,0.00404]
TCP h16 (cm)0.0891 [0.0209,0.1614]0.0176 [-0.0025,0.0414]0.0715 [-0.0042,0.1480]

Bold marks the primary force/torque specificity contrast.

The dashed rule separates the primary FT and secondary TCP endpoints.

  • FT h4 MSEMore donor specificity is not lower absolute prediction risk; the paired risk tables retain that distinction.

Specificity is hard-minus-matched risk; the final column compares A+C against Monolithic. Task-paired 95% intervals use equal task weights after averaging three sources. FT h4 is primary, TCP h16 secondary. Greater specificity does not imply lower absolute risk.

04 / Study Table 22

Donor specificity and absolute prediction risk — B. Absolute matched and hard-donor risk

24-task hard-negative bank; equal task weights; 3 fitted sources

Hard-minus-matched specificity; FT MSE and TCP cm risk

Donor specificity and absolute prediction risk — B. Absolute matched and hard-donor risk
EndpointModelMatched riskMatched 95% CIHard-donor riskHard-donor 95% CI
FT h4Align+Cross0.3343[0.2501,0.4262]0.3369[0.2518,0.4300]
FT h4Monolithic0.3205[0.2390,0.4104]0.3208[0.2391,0.4105]
TCP h16 (cm)Align+Cross4.8995[4.3548,5.4822]4.9887[4.4458,5.5615]
TCP h16 (cm)Monolithic4.3492[3.9252,4.7864]4.3668[3.9484,4.7998]

Bold marks Monolithic’s lower absolute mean risk within each endpoint and donor condition.

The dashed rule separates FT and TCP units.

  • FT h4 · MonolithicThese are descriptive mean orderings. Paired model-difference intervals are shown in the next comparison.

Absolute risks use equal task weights after averaging three sources on the 24-task bank; intervals are task-paired 95%. Monolithic retains lower mean risk under both donor conditions. The paired model differences are reported separately.

05 / Study Table 22

Donor specificity and absolute prediction risk — C. A+C minus Monolithic risk

24-task hard-negative bank; equal task weights; 3 fitted sources

Hard-minus-matched specificity; FT MSE and TCP cm risk

Donor specificity and absolute prediction risk — C. A+C minus Monolithic risk
EndpointDonorDifference95% paired interval
FT h4matched0.0138[-0.0002,0.0345]
FT h4hard0.0161[0.0028,0.0363]
TCP h16 (cm)matched0.5503[0.3699,0.7485]
TCP h16 (cm)hard0.6218[0.4302,0.8282]

Bold marks paired A+C-minus-Monolithic risk differences; positive favors Monolithic.

The dashed rule separates FT and TCP endpoints.

  • FT h4 · matchedThe primary matched-FT interval includes zero; no significance is implied by bold text.

Differences are A+C minus Monolithic; positive favors Monolithic. Intervals are paired over tasks after averaging three sources. The primary matched-FT interval includes zero. FT h4 and TCP h16 use different units and endpoint roles.

06 / Study Table 23

Action use and history-count sensitivity — A. Action and donor interventions

Separate 25-task bank; 3 fitted seeds; 10,000 task-bootstrap draws

Intervention-minus-observed errors: FT MSE; TCP cm

Action use and history-count sensitivity — A. Action and donor interventions
Intervention/modelMetricDifference95% task interval
A+C: action zeroFT h40.002074[0.000740,0.003650]
A+C: action zeroTCP h40.064386[0.045066,0.082745]
A+C: action zeroTCP h160.139929[0.076875,0.202482]
Align + CrossFT h40.009692[0.004828,0.016242]
Align + CrossTCP h40.470481[0.280141,0.694764]
Align + CrossTCP h160.338528[0.176817,0.530439]
CrossFT h40.017115[0.009042,0.028759]
CrossTCP h40.004189[-0.020586,0.029215]
CrossTCP h160.015707[-0.047204,0.073348]

Bold marks force/torque contrasts within each intervention group.

Dashed rules separate action substitution, A+C donor substitution and Cross donor substitution.

  • A+C: action zero · FT h4Pointwise task intervals are unadjusted for multiplicity; effects across FT and TCP use different units.

Action-zero contrasts subtract observed-action risk; donor contrasts subtract correct-donor risk from shuffled-donor risk. Three fitted seeds are averaged before 10,000 paired task-bootstrap draws. Intervals are pointwise without multiplicity correction; FT and TCP use different units.

07 / Study Table 23

Action use and history-count sensitivity — B. Donor-count sensitivity at fixed weights

Separate 25-task bank; 3 fitted seeds; 10,000 task-bootstrap draws

Intervention-minus-observed errors: FT MSE; TCP cm

Action use and history-count sensitivity — B. Donor-count sensitivity at fixed weights
ModelKFT h4 MSETCP h4 (cm)TCP h16 (cm)
Align + Cross10.01019 [0.00476,0.01786]0.46859 [0.27414,0.69682]0.33680 [0.17193,0.53270]
Align + Cross20.00874 [0.00325,0.01648]0.36518 [0.16746,0.60035]0.26136 [0.09066,0.46178]
Align + Cross40.00794 [0.00240,0.01566]0.29781 [0.08930,0.53826]0.21914 [0.04070,0.42493]
Cross10.01606 [0.00836,0.02703]0.00488 [-0.01979,0.03002]0.01573 [-0.04830,0.07455]
Cross20.01593 [0.00634,0.02960]0.00691 [-0.01343,0.02892]0.01599 [-0.02912,0.05778]
Cross40.01602 [0.00531,0.03114]0.00869 [-0.01117,0.02953]0.01166 [-0.02378,0.04752]
Random10.00071 [0.00017,0.00133]0.00162 [-0.00353,0.00672]0.00113 [-0.01151,0.01315]
Random20.00067 [0.00017,0.00126]0.00110 [-0.00355,0.00570]-0.00018 [-0.01234,0.01100]
Random40.00066 [0.00016,0.00122]0.00078 [-0.00368,0.00521]-0.00091 [-0.01283,0.00794]

Bold marks the K1/K4 A+C endpoints without selecting a best donor count.

Dashed rules separate fitted model families.

  • Align + CrossEach cell is random-minus-correct error at fixed K. These contrasts are not absolute risk or a retrained efficiency curve.

Every cell is random-minus-correct donor error at fixed K, not absolute risk. Donor count changes without refitting. Three fitted seeds are averaged before task bootstrapping; intervals are pointwise and unadjusted for multiplicity.