Policy Regret for Embedding Model Routing: Contextual Bandits with Low-Rank Experts

작성자

카테고리:

← 피드로
arXiv cs.AI · Yan Dai, Negin Golrezaei, Patrick Jaillet · 2026-09-28 AI

[Submitted on 12 Jun 2026 (v1), last revised 24 Sep 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:Modern recommendation systems increasingly rely on dynamically routing diverse queries to multiple embedding models. Despite its practical significance, this problem remains poorly understood under realistic conditions like adversarial queries, bandit feedback, and limited observability of models. We formalize embedding model routing as an adversarial contextual linear bandit with low-rank experts, where contexts are queries, actions are items, and experts are the embedding models working on low-rank latent representation spaces. We first establish that standard regret notions suffer from structural misspecification or statistical intractability, and we identify a log-quadratic policy class that is expressive enough to capture query-dependent model routing, yet structured enough to allow efficient online learning. Focusing on this log-quadratic policy optimization problem under bandit feedback — which is of independent interest — we propose a policy gradient algorithm called Hypentropy Policy Gradient (HPG). It provably adapts to the unknown low-rank structure under incomplete information and attains $\widetilde{\mathcal O}(s\sqrt{M T})$ linearized policy regret — where $s, M$, and $T$ are the intrinsic rank of the experts, the number of models, and the number of rounds — thus avoiding a curse of dimensionality. We provide computationally efficient and parameter-free implementation of HPG.

Submission history

From: Yan Dai [view email]
[v1] Fri, 12 Jun 2026 20:09:03 UTC (75 KB)
[v2] Thu, 24 Sep 2026 22:16:42 UTC (75 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2606.14929