SURGELLM: Rethinking Multi-Task Evaluation through Task-Aware Feature Gating with Class-Balanced Normalization

작성자

카테고리:

← 피드로
arXiv cs.AI · Noor Islam S. Mohammad, Ulug Bayazit · 2026-06-24 AI

[Submitted on 23 Jun 2026]

View PDF HTML (experimental)

Abstract:Fine-tuned encoders deployed across heterogeneous NLP tasks face three compounding problems: mismatched inductive biases, class-imbalance corruption of feature statistics, and no mechanism to condition attention on external lexical knowledge. We introduce textbf{surgellm}, a unified transformer framework that addresses each with a dedicated lightweight module: a emph{surgical feature gate} (learned per-dimension sigmoid over curated lexical indicators and texttt{[CLS]}; provably degenerates to identity when features are uninformative), emph{task-conditioned prefix tokens} (quantized feature values and task identity prepended to every input), and emph{Instance-Weighted Normalization} (IWN; removes class-prior bias from gate statistics). We prove an excess-risk bound linking gate benefit to emph{surgical feature alignment}. Across four tasks, SST-2, multi-hop retrieval, LLM-prompt attribution, and authorship detection, covering 17,830 examples and eleven model variants over three seeds, the IWN variant achieves macro-F1 textbf{0.940} ($+0.036$ over the strongest non-IWN baseline; $+0.130$ on authorship detection). A random-vocabulary control ($-0.028$ avg. F1) confirms gains are lexical, not parametric. Code, vocabularies, and a $99.5%$-recovery auto-extraction recipe are released.

Submission history

From: Noor Noor S. Mohammad [view email]
[v1] Tue, 23 Jun 2026 07:47:21 UTC (161 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2606.24259

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다