MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding

작성자

카테고리:

← 피드로
arXiv cs.AI · Wenda Wang, Yihan Tong, Yuwei Hu, Xuchen Pan, Zhewei Wei, Yaliang Li, Bolin Ding · 2026-08-03 AI

[Submitted on 2 Jul 2026 (v1), last revised 31 Jul 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:Using molecular large language models (LLMs) as a unified framework for understanding molecular structures and functions is emerging as a new trend in tasks such as molecular design and drug discovery. However, these models struggle to fully capture the visual representation of molecular structures, limiting their potential. While existing molecular vision-language models (VLMs) show promise, they still face challenges in structural alignment and lack the necessary topological modeling for accurate molecular understanding. To address this, we propose MolSight, a graph-aware vision-language model framework designed to enhance the understanding of molecular images by VLMs. MolSight integrates a Molecular Topology Module to inject chemical-bond adjacency information into vision tokens, and a Molecular Grounding Module to align visual features with chemical symbolic semantics. Our experiments demonstrate that MolSight significantly outperforms existing VLMs, molecular LLMs, and task-specific models across multiple chemical visual understanding tasks, achieving a new level of molecular image reasoning in complex chemical scenarios.

Submission history

From: Wenda Wang [view email]
[v1] Thu, 2 Jul 2026 10:13:19 UTC (2,370 KB)
[v2] Fri, 31 Jul 2026 06:19:26 UTC (3,230 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2607.01982

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다