TIF: Learning Temporal Invariance in Android Malware Detectors

작성자

카테고리:

← 피드로
arXiv cs.AI · Xinran Zheng, Shuo Yang, Edith C. H. Ngai, Suman Jana, Lorenzo Cavallaro · 2026-06-23 AI

[Submitted on 7 Feb 2025 (v1), last revised 22 Jun 2026 (this version, v3)]

View PDF HTML (experimental)

Abstract:Learning-based Android malware detectors degrade over time due to natural distribution drift caused by malware variants and new families. This paper systematically investigates the challenges classifiers trained with empirical risk minimization (ERM) face against such distribution shifts and attributes their shortcomings to their inability to learn \emph{stable} discriminative features. Invariant learning theory offers a promising solution by encouraging models to generate stable representations across environments that expose the instability of the training set. However, the lack of prior environment labels, the diversity of drift factors, and low-quality representations caused by diverse families make this task challenging. To address these issues, we propose TIF, the first temporal invariant training framework for malware detection, which aims to enhance the ability of detectors to learn stable representations across time. TIF organizes environments based on application observation dates to reveal temporal drift, integrating specialized multi-proxy contrastive learning and invariant gradient alignment to generate and align environments with high-quality, stable representations. TIF can be seamlessly integrated into any learning-based detector. Experiments on a decade-long dataset show that TIF excels, particularly in early deployment stages, addressing real-world needs and outperforming state-of-the-art methods.

Submission history

From: Xinran Zheng [view email]
[v1] Fri, 7 Feb 2025 17:17:42 UTC (1,011 KB)
[v2] Wed, 17 Sep 2025 03:12:35 UTC (1,114 KB)
[v3] Mon, 22 Jun 2026 03:24:07 UTC (1,149 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2502.05098

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다