← 피드로
[Submitted on 18 Jul 2026 (v1), last revised 24 Jul 2026 (this version, v2)]
Abstract:Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically produce scripts that are not automatically materialized as persistent, editable platform artifacts. We call this disconnect the textit{NL2Pipeline gap}. To bridge it, we introduce textsc{DataFlow-Harness}, a platform that guides an LLM agent to construct platform-native directed acyclic graphs (DAGs) through typed, incremental mutations rather than free-form scripts. The platform combines textsc{DataFlow-Skills} for procedural guidance, a Model Context Protocol (MCP) layer that exposes the live operator registry and current pipeline state, and textsc{DataFlow-WebUI}, which synchronizes conversational authoring with a visual DAG editor. On a 12-task data-engineering benchmark, textsc{DataFlow-Harness} achieves a 93.3% observed end-to-end pass rate. Relative to Vanilla Claude Code, it reduces measured monetary cost by 72.5% and generation latency by 49.9%; its observed pass rate is within 0.9 percentage points of the Context-Aware Claude Code baseline while its cost is 42.8% lower. Per-task analysis indicates that Skills are most useful when construction depends on implicit procedural knowledge. These results show that live platform grounding can produce persistent, editable workflow artifacts with an observed reliability close to script-generation baselines and with lower measured construction cost and latency.
Submission history
From: Runming He [view email]
[v1]
Sat, 18 Jul 2026 03:30:19 UTC (4,205 KB)
[v2]
Fri, 24 Jul 2026 08:50:49 UTC (4,205 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2607.16617
답글 남기기