← 피드로
[Submitted on 12 Aug 2026]
Abstract:Large Language Models (LLMs) are increasingly used in circuit design workflows, yet their reliability on simulator-facing SPICE netlist recognition and manipulation remains poorly understood and is rarely separated from high-level design reasoning. Although netlists are textual, they encode structured circuit objects through topology and parameters. We present \textbf{NetlistBench}, a structure-verified benchmark for SPICE netlist recognition and manipulation. NetlistBench contains 2,342 cases across 24 task families, covering parameter and connectivity recognition and edits, hierarchical operations, equivalence judgment, and long-horizon compound editing. Model outputs are evaluated by a deterministic structure-aware oracle. Across six non-thinking LLMs, performance varies substantially with operation-level structural complexity. Simple local edits reach $96\%$–$100\%$ accuracy, while device addition drops to $41\%$–$83\%$ and equivalence judgment to $49\%$–$90\%$. Enabling reasoning substantially improves weaker models but does not eliminate structure-preservation failures, with performance still degrading sharply as the edit horizon increases. NetlistBench identifies netlist reliability as a distinct bottleneck for trustworthy LLM-based circuit design automation.
Submission history
From: Xiaoguang Liu [view email]
[v1]
Wed, 12 Aug 2026 15:51:52 UTC (5,254 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2608.12197
답글 남기기
댓글을 달기 위해서는 로그인해야합니다.