Reflex: Enabling Fast and Predictive Vision-Language-Action Models for Reaction-Critical Manipulation

작성자

카테고리:

← 피드로
arXiv cs.AI · Yuxuan Chen, Wanruo Zhang, Xiao Li · 2026-08-17 AI

[Submitted on 14 Aug 2026]

View PDF HTML (experimental)

Abstract:Vision-Language-Action (VLA) models have recently achieved promising performance in robotic manipulation. However, existing benchmarks mainly evaluate generalization on static manipulation tasks and largely overlook dynamic interaction scenarios. To address this gap, we present ReflexBench, a benchmark for reaction-critical manipulation. ReflexBench contains six dynamic tasks and introduces an evaluation framework that decouples simulator stepping from robot control while supporting configurable latency under synchronous and asynchronous inference. Building upon ReflexBench, we propose ReflexVLA, an efficient VLA model designed for reaction-critical manipulation without large-scale robot-data pretraining. ReflexVLA enhances temporal reasoning through latent future prediction and multi-frame temporal fusion within the vision backbone, while reducing deployment latency through batched visual encoding and CUDA Graph replay. Experiments show that ReflexVLA consistently improves dynamic manipulation performance while maintaining competitive accuracy on standard static manipulation benchmarks, and real-world experiments further demonstrate its effectiveness under practical deployment conditions. Project website: this https URL

Submission history

From: Yuxuan Chen [view email]
[v1] Fri, 14 Aug 2026 15:19:04 UTC (2,510 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2608.14379