Ask ten LLMs for a Blender 5.0 script. Then run it in Blender 5.0.

작성자

카테고리:

← 피드로
DEV Community · Dhotiiiii · 2026-09-29 개발(SW)

This is a submission for the Kaggle Benchmarking Challenge

What I Benchmarked

The question

Ask an LLM for a Blender Python script and it will usually give you one that looks right. Whether it runs depends on which Blender you paste it into. The bpy API changes on every major release: enums get renamed, attributes are removed, node sockets move, whole subsystems (compositor, sequencer, animation) get restructured. Most of the Blender code a model has seen was written for 2.8 to 3.x. Ask for 5.0 and you often get a 3.x script that dies on the first changed line.

Text similarity cannot see this. mesh.use_auto_smooth = True reads perfectly and raises AttributeError in 4.1 and later. The only judge that counts is the Blender version you asked for. So the benchmark runs every answer in that exact version.

Two questions per answer:

  • runs: does the script, followed by a per-task assert script, exit 0 in a headless Blender of the target version?
  • aware: does the answer say what changed? Each answer must end with a WATCH OUT block naming every API that changed for that version and its replacement, or - none.

These come apart in interesting ways, which is most of what follows.

The task

30 scripting tasks, each asked for Blender 3.6, 4.2, 4.5 and 5.0 (119 prompts per model; one task has no 3.6 form). The tasks cover 33 verified API changes: EEVEE engine ids, the boolean solver enum, auto-smooth and custom normals, calc_normals, Principled BSDF socket names, the OBJ exporter, node group sockets, bone collections, slotted actions, the compositor node tree moving off Scene, the sequencer’s sequences becoming strips, and so on. Two control tasks use APIs that did not change, to check the grader does not punish ordinary correct code.

Every model gets the same system prompt, temperature 0, no tools, no retrieval. The Kaggle task ships the four Blender Linux builds in a dataset, extracts them at the start of the run, and grades each answer by launching blender -b --factory-startup --python on the target build. A hand-written reference answer per task and version passes its assert in all four builds (119/119), which is the evidence that every prompt is answerable.

Models Tested

Ten models through the Kaggle Model Proxy, chosen to span vendors, sizes, open and closed weights, and reasoning styles, within a daily proxy budget of a few dollars:

model why gpt-5.5, claude-sonnet-5, gemini-3.1-pro-preview current frontier from three vendors gemini-3.7-flash, gemini-3.8-flash, claude-haiku-4-5 the small and cheap tier people actually script with grok-4.20-0309-reasoning fourth vendor, reasoning model (grok-4.6 is listed by the proxy but returned 404 on every call) deepseek-r1-0528 open-weight reasoning model that returns its thinking inline qwen3-coder-480b-a35b-instruct open-weight model tuned for code gpt-oss-120b open-weight OpenAI model

Total proxy spend for the ten leaderboard runs: $11.23. gpt-5.5 alone was $4.44; qwen3-coder was $0.03.

Findings

Results

1190 graded answers. Aware is re-graded offline with the current parser so every model is judged the same way (more on that under methodology).

model runs aware runs but unaware gpt-5.5 95% 71% 26% claude-sonnet-5 94% 53% 42% gemini-3.8-flash 92% 77% 17% gemini-3.1-pro-preview 92% 76% 18% gemini-3.7-flash 89% 73% 18% grok-4.20-reasoning 82% 29% 57% claude-haiku-4-5 72% 31% 48% qwen3-coder-480b 70% 29% 46% deepseek-r1 67% 39% 33% gpt-oss-120b 61% 29% 37%

Share of scripts that run on the target version, by Blender version

Runs by target version:

model 3.6 4.2 4.5 5.0 gpt-5.5 97% 97% 100% 87% claude-sonnet-5 100% 97% 100% 80% gemini-3.8-flash 97% 90% 97% 87% gemini-3.1-pro-preview 97% 90% 97% 83% gemini-3.7-flash 93% 90% 90% 83% grok-4.20-reasoning 86% 83% 87% 70% claude-haiku-4-5 79% 77% 73% 60% qwen3-coder-480b 83% 67% 70% 60% deepseek-r1 86% 70% 63% 50% gpt-oss-120b 93% 53% 57% 43%

Pooled over all ten models: 91% of scripts run on 3.6, 81% on 4.2, 83% on 4.5, 70% on 5.0. Awareness falls faster: 81%, 48%, 40%, 35%.

1. Every model is worst on 5.0, and the 5.0 failures are the same five APIs

Run rate by change category and version, all models

Three categories go to 0% on 5.0 across all ten models: the compositor (Scene.node_tree is gone; the compositor is now a node group on scene.compositing_node_group), the boolean solver ('FAST' is now 'FLOAT', and 'MANIFOLD' is new), and the sequencer (SequenceEditor.sequences is now strips, and new_effect() takes length= instead of frame_end=). Animation is at 10%: Action.fcurves moved under slotted actions (action.layers[0].strips[0].channelbag(slot).fcurves). Not one model knew any of these. The training data has not caught up with 5.0 and no amount of scale fixes that. The best model on 5.0 (gpt-5.5 and gemini-3.8-flash, 87%) is still exactly as lost on these five as gpt-oss-120b.

The failures on 4.2 and later are different in character. They are changes from 4.0 to 4.2 that the models have partly absorbed (counts pooled over 4.2, 4.5 and 5.0): use_auto_smooth removed (19 failures), the legacy OBJ exporter bpy.ops.export_scene.obj removed (15), Mesh.calc_normals removed (14), BLENDER_EEVEE renamed (14), Principled BSDF Specular renamed to Specular IOR Level (9). The frontier models mostly get these right; the open-weight models mostly do not. gpt-oss-120b is the clearest case: 93% on 3.6, 53% on 4.2. It knows Blender 3.x well and stopped there.

2. A rename that was reverted splits the models into two camps

The render engine id was BLENDER_EEVEE in 3.6, became BLENDER_EEVEE_NEXT in 4.2 when EEVEE Next shipped, and went back to BLENDER_EEVEE in 5.0 when the legacy engine was deleted. Every model failed this task somewhere:

camp models 3.6 4.2 4.5 5.0 learned the 4.2 rename gpt-5.5, claude-sonnet-5, claude-haiku-4-5 ok ok ok fail never learned it both gemini flash models, gemini-3.1-pro, grok, deepseek, qwen, gpt-oss ok fail fail ok

The second camp is right on 5.0 by accident: they write BLENDER_EEVEE for every version. The first camp is wrong on 5.0 with full confidence. gpt-5.5’s WATCH OUT for the 5.0 prompt reads:

scene.render.engine = 'BLENDER_EEVEE': changed in Blender 4.2 when legacy EEVEE was replaced by EEVEE Next; use scene.render.engine = 'BLENDER_EEVEE_NEXT' in Blender 5.0.

Claude Sonnet 5’s answer for 5.0 is the most instructive failure in the whole set. Its WATCH OUT names both identifiers and says the old one “no longer refers to the current EEVEE renderer in 5.0”, so the aware check passes. The script then sets BLENDER_EEVEE_NEXT and fails. It knows there is a story here and tells the wrong half of it.

3. Running is not knowing: the “runs but unaware” gap

The gap column in the first table is the share of answers where the script ran but the WATCH OUT block missed the change. It is the number I find most useful, because it separates two very different kinds of model.

gemini-3.8-flash and gpt-5.5 have small gaps (17% and 26%): when their code works, they can usually say why the old code would not. Claude Sonnet 5 runs 94% of the time but names the change in only 53%; grok-4.20 runs 82% and names it in 29%. Both write - none under WATCH OUT for most prompts: Sonnet in 85 of 119 answers, grok in 110. qwen3-coder and Haiku do the same (115 and 109), and it is why the whole bottom half of the table sits at 29% to 39% aware. For every 2.80-era change (scene.objects.link becoming collection.objects.link, lamps becoming lights, dupli becoming instance) they silently write the modern form and declare that nothing changed. That is a judgement about what counts as “changed”, not a parse failure, and I decided not to special-case it: the prompt asked for changes relative to older Blender, and 3.6 is old enough that 2.80 changes are still the ones people trip over in old tutorials.

The reverse gap, “aware but breaks”, is small everywhere (1% to 7%). 41 answers named the change and still emitted broken code. Sonnet’s EEVEE answer above is one. Three of Haiku’s eight are the OBJ exporter: it knows export_scene.obj became wm.obj_export, then passes the old use_selection= keyword to the new operator.

4. The awareness curve is steeper than the run curve

Share of answers that name the API change, by Blender version

On 3.6, awareness is 69% to 97%: everybody knows the 2.80 changes. By 4.2 the field splits. The three Gemini models and gpt-5.5 stay between 60% and 80% through 5.0, Sonnet sits at 43% to 53%, and everyone else drops below 30% by 4.5 and stays there. The scripts of that last group still run 57% to 87% of the time on 4.5 because most tasks touch APIs that did not change in that version. They are right without knowing why, and that is exactly the situation where a user cannot tell a stale answer from a current one.

5. Temperature 0 is not a determinism guarantee

The same model, prompt and settings, run four times through the proxy (gemini-3.7-flash across task versions): 107, 107, 103 and 106 of 119 scripts ran. Zero of the 119 answers were byte-identical between two runs. That is 3 to 4 points of noise on top of the sampling interval, so gpt-5.5 at 95%, Sonnet at 94% and the two Gemini models at 92% are a tie. Sonnet and gemini-3.8-flash actually swapped places between two task versions. The differences that survive the noise are the tiers: the five models above 89%, grok alone at 82%, and the four between 61% and 72%.

Methodology notes (the parts that would have silently corrupted the numbers)

Three grading problems showed up in real runs. Each one would have produced a plausible-looking leaderboard that was wrong.

  • Reasoning leaks into the answer. deepseek-r1 returns its thinking inline as <think>...</think>, and in 29 of 119 answers the thinking contained a draft code fence. The first notebook version ran the draft. Fixed by stripping the block before parsing; deepseek’s run rate moved from 62% to 67% and its awareness came down, because the rehearsed WATCH OUT inside the thinking had inflated it. The Kaggle leaderboard still shows 62% for deepseek because that is what the notebook version that ran it computed; the tables here are re-graded offline from the stored answers, and the changed scripts were re-run in the same Blender builds.
  • Output caps truncate thinking models. Gemini 3.x thinking tokens count against max_tokens through the proxy. gemini-3.8-flash spends 12 to 16 thousand tokens thinking about auto-smooth, hit an 8192 cap once, and its “SyntaxError: unterminated string literal” was a truncation, not a model error. The task now raises the cap to 16384 for reasoning models and re-asks once at 32768 if an answer lands within 32 tokens of the cap. One answer needed that retry.
  • Proxy errors are not model errors. Rate limits (429) and daily-quota rejections (403) come back as exceptions. The task retries with backoff, excludes prompts that are still lost from the denominator, and aborts if more than 10% are lost rather than record a score. An early version recorded qwen3-coder at 0.8% because 118 of 119 calls were rate-limited; that version is gone.

Two limits worth stating. The aware check matches identifiers by substring and quoted enum values exactly, so it is a floor on awareness, not a proof of understanding. And 30 tasks is enough to show the shape of the drift, not to rank models that are three points apart.

Where this came from

The case bank and the WATCH OUT contract come from bpy-compass, a version-aware Blender assistant I built for the Sanity challenge. Its evaluation compared one model with and without a retrieval layer over the release notes. The question this benchmark asks is the one I wanted answered before building that: how bad is the drift without retrieval, for which models, and where exactly. The answer is that on 5.0 it is bad for everyone, and that retrieval over release notes is not a nicety for Blender scripting, it is the difference between the compositor working and not.

My Benchmark

원문에서 계속 ↗