TR2026-142
Test-Time Attention: Can Robots Better Follow Commands?
-
- , "Test-Time Attention: Can Robots Better Follow Commands?", IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) IARL Workshop, September 2026.BibTeX TR2026-142 PDF
- @inproceedings{Liu2026sep,
- author = {Liu, Jing and Wang, Ye and Suzuki, Kei and Koike-Akino, Toshiaki},
- title = {{Test-Time Attention: Can Robots Better Follow Commands?}},
- booktitle = {IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) IARL Workshop},
- year = 2026,
- month = sep,
- url = {https://www.merl.com/publications/TR2026-142}
- }
- , "Test-Time Attention: Can Robots Better Follow Commands?", IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) IARL Workshop, September 2026.
-
MERL Contacts:
-
Research Areas:
Abstract:
Generalist Vision-Language-Action (VLA) policies promise broad manipulation capabilities, but the same breadth can weaken fidelity to a narrow deployment command. Repeated scene–object–action correlations and the large number of visual tokens relative to language tokens may cause visually suggested behaviors to compete with the operator’s current command. We ask whether an already deployed VLA can be made more command-focused by redistributing attention only at inference time, without new demonstrations or policy retraining. Using pretrained pi 0.5 policy on a R1 Pro wheeled humanoid, we propose and evaluate 7 heuristic training-free interventions that either strengthen the current subtask text or reduce the influence of less command-relevant visual information in an outfitting toolbox setting. Because task success alone can conceal interactions with objects not specified by the current command, we introduce the Command Focus Ratio (CFR), which measures interaction selectivity, and the Target Interaction Rate (TIR), which measures how often the commanded target is engaged. Empirical studies suggest that test-time attention methods can improve the command fidelity of VLA policies without additional data collection or retraining, while the remaining gap in CFR highlights the need for stronger test-time mechanisms to suppress interactions with non-target objects.



