TTT-RM

Test-Time Training as Residual Memory for Robot Policies

Remembering task-relevant history beyond the current observation and explicit memory.

1University of Illinois Chicago 2Cisco Research

Abstract

Memory is essential for long-horizon robotic manipulation, where successful actions may depend on past events that are no longer recoverable from the current observation. As episodes grow longer, however, retaining the full history becomes increasingly costly, creating a fundamental scalability challenge for memory-augmented policies. Existing approaches address this challenge by either storing selected past observations in a bounded memory bank or compressing interaction history into a fixed-size parametric state through Test-Time Training (TTT). Yet these formulations do not explicitly distinguish between historical information that can already be recovered from the policy’s current context and information that must persist beyond it.

We introduce TTT-RM, which repurposes TTT as Residual Memory, using fast weights not to generically compress history but to complement a bounded memory bank by preserving task-relevant historical information that cannot be recovered from the policy’s current context. Concretely, TTT-RM learns a history decoder that reconstructs historical representations from the current observation and retrieved memory. The resulting reconstruction residual captures what this context fails to explain and serves as the learning target for TTT. The TTT slow weights are optimized against this residual target so that online fast-weight updates learn to encode complementary historical information over time. The fast-weight state is then queried to produce a residual memory representation that conditions action generation.

Extensive experiments on memory-intensive simulation benchmarks and real-world tasks show that TTT-RM consistently improves across multiple memory designs, outperforms diverse baselines, and supports sustained execution on a three-minute, eight-stage stowing task.

Explicit Memory stores selected tokens, TTT compresses full history into fast weights, and TTT-RM preserves history unexplained by the existing context.
Three ways to remember. Explicit Memory retains selected historical observations. TTT compresses history into a fixed-size fast-weight state. TTT-RM specializes that state to preserve information the current observation and retrieved memory cannot recover.

How TTT-RM Works

A complementary memory pathway that reuses the existing visual stream, with no additional encoder or VLM forward passes.

Identify what is missing

A history decoder reconstructs past representations from the current observation and retrieved memory. Its reconstruction residual defines the information this context cannot explain.

Learn residual memory

Residual supervision trains the slow write and read parameters so that online fast-weight updates preserve complementary historical information within a fixed-size state.

Recall history for control

Explicit and residual memory jointly condition action generation. At inference, the history decoder is removed; fast weights update once per policy call and reset each episode.

TTT-RM pipeline: frozen visual features feed Explicit Memory and a writer that updates TTT fast weights. A training-only history decoder supervises Residual Memory, which conditions the action expert.
Method overview. The history decoder supplies residual targets during training. At deployment, the policy reads from Explicit Memory and the updated fast-weight state to guide its next action chunk.

Simulation Results

Consistent gains across memory designs and tasks that require remembering past observations.

RoboMME — average success rate (%) across Counting, Permanence, Reference, and Imitation. Each TTT-RM variant augments the corresponding Explicit Memory (EM) baseline.
Memory configurationExplicit MemoryTTT-RMGain (pp)
TokenDrop / Expert34.5±0.437.1±1.2+2.6
FrameSamp / Expert36.7±1.038.9±0.8+2.2
TokenDrop / Modul30.9±1.138.0±0.5+7.1
FrameSamp / Modul42.8±1.945.4±1.0+2.6
RMBench — average success rate (%). M(1) tasks require a single past observation; M(n) tasks require integrating multiple past observations.
Memory complexityBaseTTTEMTTT-RM
M(1)15.6±1.711.2±0.730.5±3.336.5±4.1
M(n)8.8±0.810.3±0.810.8±0.315.8±0.3

Real-World Demonstrations

Remembering completed actions, counting objects, and sustaining an eight-stage stowing task on a mobile manipulator.

Supplementary video — real-world task rollouts and comparisons.
Mean task progress (%) over 20 trials per method per task. Progress measures completed task stages, rather than full-task success. Gains are in percentage points (pp).
TaskStrongest baselineTTT-RMGain (pp)
Pen to drawer55.565.5+10.0
Pick N pens47.073.5+26.5
Stow objects50.577.0+26.5