← 피드로
[Submitted on 9 Jun 2026]
Abstract:Lipschitz-style individual fairness formalizes the idea that semantically similar examples should receive similar predictions, but its evaluation in multi-task learning (MTL) can be confounded by method-induced representation scales. This paper identifies threshold confounding: when the auditing tolerance is derived from each model’s own representation distances, different algorithms are compared under different semantic thresholds. A threshold-drift analysis further shows how Bias rankings can change and identifies sufficient conditions for ranking preservation.
We propose textbf{ReLiF}, a reliability-aware framework that separates evaluation-time fixed-$delta$ auditing from training-time controlled regularization. ReLiF uses a shared reference tolerance for comparable auditing and a violation-rate feedback controller to keep the Lipschitz surrogate active without letting it dominate stochastic training. This work also develops supporting analysis for threshold drift, reference-tolerance selection, and the relationship between the huberized training surrogate and its unsmoothed positive-margin counterpart.
Experiments on clinical time-series benchmarks and NYUv2 (NYU Depth V2) dense prediction show that fixed-$delta$ auditing exposes utility–fairness trade-offs that method-dependent thresholds can obscure. On NYUv2 with a ResNet50 backbone, ReLiF achieves competitive utility while substantially reducing aligned bias under shared fixed thresholds. On clinical benchmarks, ReLiF yields controlled fairness-regularized trade-offs, while fixed-$delta$ auditing reveals that task-balancing baselines can sometimes achieve lower bias and that genuine utility–fairness trade-offs persist. These results support fixed-$delta$ auditing as a semantically consistent protocol for evaluating Lipschitz fairness in MTL.
Submission history
From: Danhuai Guo [view email]
[v1]
Tue, 9 Jun 2026 09:36:54 UTC (720 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2606.10632
답글 남기기