← 피드로
[Submitted on 17 Jun 2026]
Abstract:This paper develops a formal account of what generalist agents must store in memory in order to act near-optimally across multiple environments and goals. It shows that when two domains share an observational bottleneck but require incompatible optimal actions, any uniformly near-optimal policy must induce distinct memory distributions at that bottleneck. The result yields a separation theorem: sufficiently successful agents cannot rely only on current state observations, but must preserve domain-relevant information in memory. The paper further shows that if an agent’s memory contains enough information to estimate values for related goals, then that memory can be used to approximately reconstruct the agent’s local transition dynamics. Together, these results characterize memory as the substrate that supports domain disambiguation, transition-model reconstruction, and planning for generalist agents.
Submission history
From: Khurram Yamin [view email]
[v1]
Wed, 17 Jun 2026 06:46:51 UTC (851 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2606.18746
답글 남기기