Supervised Learning Has a Geometric Blind Spot

작성자

카테고리:

← 피드로
arXiv cs.AI · Vishal Rajput · 2026-08-07 AI

[Submitted on 23 Apr 2026 (v1), last revised 6 Aug 2026 (this version, v3)]

View PDF

Abstract:Ordinary supervised training minimises the task loss and then stops. It never pays for how far the representation moves when the input is nudged along directions that helped fit training labels—including directions that are nuisance at deployment. We call that leftover sensitivity the geometric blind spot of empirical risk minimisation. In a Gaussian linear model where the nuisance enters the label conditional and the decoder has finite Lipschitz constant, population MSE forces a floor on linearised representation drift. The same distinction predicts a failure mode of adversarial training: Jacobian magnitude can fall while clean class geometry worsens. We track that dissociation with a class-layout score and study isotropic encoder matching—penalising the squared distance between phi(x) and phi(x+delta) for Gaussian delta under a task-loss cap—when nuisance axes are unknown. On a Vision Transformer trained from scratch on CIFAR-10, projected gradient descent attains the smallest Jacobian Frobenius yet the worst clean layout score (1.353+/-0.020 over three seeds), above task-only training (1.093); isotropic matching attains the best (0.904). The drift floor is proved for the linear-Gaussian case; deep nets and cross-task orderings are protocol empirics. Design rule: report class-layout geometry beside the task score; prefer isotropic encoder matching when axes are unknown.

Submission history

From: Vishal Rajput [view email]
[v1] Thu, 23 Apr 2026 08:03:33 UTC (69 KB)
[v2] Mon, 27 Apr 2026 10:44:51 UTC (295 KB)
[v3] Thu, 6 Aug 2026 15:33:34 UTC (258 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2604.21395

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다