Learning to Walk With Less: A Dyna-Style Approach to Quadrupedal Locomotion

작성자

카테고리:

← 피드로
arXiv cs.AI · Francisco Affonso, Felipe Tommaselli, Jo~ao H. Al'essio, Vivian S. Medeiros, Mateus V. Gasparino, Girish Chowdhary, Marcelo Becker · 2026-08-10 AI

[Submitted on 8 Sep 2025 (v1), last revised 7 Aug 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:Traditional on-policy reinforcement learning (RL) controllers for quadrupedal locomotion often suffer from low data efficiency, requiring millions of interactions with simulated environments to achieve stable control. We integrate model-based techniques that improve sample efficiency by augmenting PPO rollouts with synthetic data in a Dyna-style framework. Our method employs a learned transition model to generate short-horizon synthetic tails for each trajectory, anchored by physics-based simulation to preserve stability. A predefined scheduling strategy gradually integrates synthetic transitions, preventing model usage during early training stages when prediction accuracy is low. Through extensive ablation studies, we analyze how varying data parameters influence PPO’s learning behavior. Finally, we validate our method in simulation on a Unitree Go1 robot, reaching convergence with substantially fewer simulation steps (19.64M vs. 27.53M) and a 12.24% reduction in wall-clock training time, without compromising policy performance or convergence. Cross-platform experiments on ANYmal and Unitree Go2 further confirm the framework’s ability to learn high-dimensional locomotion control with substantially reduced simulation experience, despite reward trade-offs on complex morphologies.

Submission history

From: Francisco Affonso [view email]
[v1] Mon, 8 Sep 2025 02:48:23 UTC (2,922 KB)
[v2] Fri, 7 Aug 2026 04:58:39 UTC (3,444 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2509.06296

코멘트

답글 남기기