π€― We’ve all been there.
You’re iterating on a prompt for an LLM. You tweak one word, the output gets better. You tweak another, it gets worse. You paste it into ChatGPT, it kind of works. You change “concisely” to “in 3 bullets”, and suddenly the summaries are perfect.
Then tomorrow comes. And you have no idea which version was the good one.
Your folder looks like this:
prompts/
βββ summarize.txt
βββ summarize_old.txt
βββ summarize_v2.txt
βββ summarize_v2_final.txt
βββ summarize_v2_final_FINAL.txt
βββ summarize_REALLY_final.txt
βββ summarize_use_this_one.txt
Enter fullscreen mode Exit fullscreen mode
And the worst part? The “good” one might be any of them. π
You can’t diff them. You can’t tell what changed between v2_final and v2_final_FINAL. You can’t roll back to the version that actually worked.
π€ Why not just use git?
Good question. I asked myself the same thing.
You can. And technically, PromptVault borrows git’s entire model β content-addressed objects, trees, commits. But in practice, prompts live next to code. They’re mixed into repos, notebooks, chat exports, and random .txt files on your desktop.
Prompts deserve their own version control that:
- π Tracks
.txt/.md/.prompt/.j2/.yamlfiles specifically - π Lives in a
.pv/folder that doesn’t collide with your code repo’s.git/ - π Is built for the iterate-and-compare workflow that prompt engineering actually is
- π« Doesn’t require a GitHub repo, doesn’t sync to the cloud, doesn’t ask for an account
So I built it. π οΈ
π Enter PromptVault
PromptVault is git, but purpose-built for prompts. It’s written in Rust, ships as a single ~2MB static binary, and runs entirely on your machine. No account, no cloud, no telemetry.
$ pv init
Initialized empty prompt vault in ./.pv
$ pv add prompts/summarize.md
added: prompts/summarize.md
$ pv commit -m "refine: summarize now reports tone + title"
[main 791151d] refine: summarize now reports tone + title
1 prompt
Enter fullscreen mode Exit fullscreen mode
Now iterate. See exactly what changed:
$ pv diff prompts/summarize.md
diff -- prompts/summarize.md
You are a precise summarizer.
-Summarize the text below in 3 concise bullets, then propose a title.
+Summarize the text below in 3 concise bullets, propose a title, and note the tone.
{{text}}
Enter fullscreen mode Exit fullscreen mode
Walk back through every iteration:
$ pv log
commit 791151dβ¦
parent 4100199β¦
Date: Mon Aug 10 10:04:33 2026 +0000
refine: summarize now reports tone + title
commit 4100199β¦
Date: Mon Aug 10 10:04:33 2026 +0000
feat: initial prompt set
Enter fullscreen mode Exit fullscreen mode
Restore any past version by hash, tag, or branch:
$ pv show 4100199 # prefix works too π
$ pv revert v1.0 # restore working tree to a tagged version βͺ
Enter fullscreen mode Exit fullscreen mode
πΏ Branches for A/B testing
This is the killer feature for prompt engineering. You want to test two variants of the same prompt? Branch it.
$ pv branch experiment
$ pv checkout experiment
# ... tweak the prompt, commit it ...
$ pv checkout main # working tree restores to main's version
$ pv checkout experiment # ...and back to experiment's version
Enter fullscreen mode Exit fullscreen mode
Switch branches and the working tree restores instantly. No copying files, no renaming, no “wait, which folder was the experiment in?” π―
When you’re done, compare them against a dataset:
$ pv ab main:summarize.md experiment:summarize.md -d cases.jsonl --show
A/B: A=main:summarize.md B=experiment:summarize.md (3 cases)
[1/3] DIFF differs
--- A vs B ---
-A You are a precise summarizer.
+B You are a concise summarizer.
hello world
--- end ---
Summary: 0 identical, 3 differing (of 3)
Enter fullscreen mode Exit fullscreen mode
Pure local. No model calls. No API keys. π
πΈ Evals without the API bill
One thing that annoyed me about existing prompt tooling: everything wants to call a model. Every eval run costs money. Every test sends data to OpenAI.
PromptVault takes a different stance: it never calls a model. It only renders templates and checks assertions.
$ pv eval summarize.md --dataset cases.jsonl --show
Eval: summarize.md (3 cases)
[1/3] PASS contains "3 concise bullets"
--- rendered prompt ---
You are a precise summarizer.
...
--- end ---
Summary: 2/2 passed (100%)
Enter fullscreen mode Exit fullscreen mode
You write a JSON Lines dataset, each line fills the prompt’s {{variables}}, and PromptVault renders + asserts. β
If you want to actually run the prompt through a model, pipe the rendered output to whatever runner you trust (curl, the OpenAI CLI, ollama, your own script).
This separates two concerns that existing tools conflate:
- Did my template render correctly? (cheap, local, deterministic) β‘
- Does the model produce good output? (expensive, non-deterministic, model-dependent) π°
PromptVault does (1). You choose how to do (2).
π The full feature list
For a “tiny git for prompts”, it ended up with more than I planned:
- πΈ Snapshots with messages (like git commits)
- π Line-level diff so you can see exactly which instruction you changed.
pv diff --statfor a one-line summary per file. - πΏ Branches & merging β
pv branch experiment,pv merge experiment(fast-forward or three-way, with conflict markers) - π·οΈ Tags & rollback β
pv tag v1.0,pv revert v1.0 - π«
.pvignoreβ gitignore-style, so drafts stay out of the vault - π Ref-to-ref diff β
pv diff v1 v2compares any two commits/tags/branches - π Ref:path access β
pv show HEAD:summarize.mdreads any file at any ref - π§° Stash / reset / clean / grep / export / stats / blame β everyday git-class utilities
- π₯οΈ TUI β
pv tuilaunches an interactive commit browser - π Shell completions β bash / zsh / fish / elvish / powershell
- βοΈ Remote sync β
pv push/pv pullsyncs the vault to any git host as a backing store - π€ Model runner (opt-in, build with
--features run) βpv runagainst OpenAI/Anthropic/Ollama. API keys live only in env vars; nothing is ever logged or stored. - βοΈ A/B testing β
pv ab main:x experiment:x -d dataset.jsonlrenders two versions against the same dataset and diffs them
π§ How it works under the hood
PromptVault is a tiny git. On pv init it creates:
.pv/
βββ HEAD β "ref: refs/heads/main"
βββ index.json β staging area (path β blob hash)
βββ objects/ β content-addressed store (SHA-256)
β βββ ab/cdefβ¦ β "<type>