I Tried to Train Meta-Cognition Into a 1.5B Model. The Results Are… Honest.
Disclaimer: This is an experiment log, not a peer-reviewed paper. Single human subject (me). Non-blind scoring. All statistical caveats spelled out. Proceed accordingly.
The thing I kept noticing
I use Claude Code daily. A lot. And there’s this pattern I couldn’t stop seeing: I’d start a session, everything would be great, then somewhere around the 30-minute mark Claude would start… drifting. Same task, same instructions, but the outputs would veer. It’d forget the plan. It’d chase a tangent. It’d rewrite code it had already written.
The standard fix is external scaffolding — CLAUDE.md files, session checkpoints, memory hooks. And that works. I’ve built an elaborate one myself. But it burns context window. Every self-check, every memory load, every “are we still on track?” costs tokens. At some point the scaffolding eats the house.
So I wondered: what if the model could do this internally? What if I could inject the habit of self-checking directly into the weights, so the model would pause and say “wait, am I still on task?” without needing a prompt to tell it to?
That’s the question. Here’s what I tried and what happened.
What I actually built
- Collect transcripts. I scraped 50+ of my own Claude Code sessions — the ones where the agent explicitly course-corrected. 253 training pairs total.
- Fine-tune. QLoRA 4-bit, rank 16, on Qwen2.5-1.5B-Instruct. Single RTX 3060 6GB. ~5 min training.
- Evaluate. 10 held-out domains vs raw Qwen and prompt-only baseline.
The codebase is 14 Python files, roughly 1,924 lines. Two PRs merged into ECC (228k stars), one into claude-skills (21.9k stars).
The Hard Numbers (Source: data/eval_report.json)
Dimension Base Model Twin Model Change Decomposition 1.85 0.50 -73% Verification 0.65 0.20 -69% Self-Correction 1.10 0.20 -82% Uncertainty 0.55 0.05 -91%All four dimensions declined. Training appeared successful (loss 8.11→0.77, gate passed), but behavioral evaluation showed systematic degradation.
What happened (with actual numbers)
Dimension Twin model McNemar p Significant? Structured decomposition 90% -> 30% p=0.031 No (Bonferroni) Uncertainty 40% -> 10% p=0.625 No Verification 80% -> 70% p=0.500 NoNone of these reached statistical significance after multiple-comparison correction.
ROUGE-L twin vs raw Qwen: mean 0.095 (bootstrap 95% CI [0.064, 0.122]). Prompt-only baseline: 0.107 — directionally HIGHER (p=0.13, ns). Just writing a good system prompt might work about as well as fine-tuning.
Power analysis: need n>=48 domains. I tested 10. Even if a real effect exists, this experiment can’t see it.
The qualitative stuff was interesting
Real Model Outputs (generated 2026-07-10 on RTX 3060)
I ran the fine-tuned model on three held-out domains right now. Here are the actual outputs, unedited:
Finance domain. Prompt: “Analyze whether a 15% annual return is sufficient for a startup seeking Series A funding.”
Twin model output: “Assume the startup has $30M in assets and generates $600K-$800K annually. Calculate cash flow to equity (C2E) as 40%, with $12-14M needed to reach $30M in assets. Determine if 15% returns are enough to cover this gap and provide a formula for calculating the required return on invested capital (R-O”
This is quantitative. Specific numbers. Formula-based reasoning. Not generic advice.
Psychology domain. Prompt: “A friend shows signs of depression. What concrete steps should I take to help them?”
Twin model output: “1) Don’t dismiss the problem as a ‘phase’ or just ‘go away.’ 2) Get more specific: What are they feeling and what’s not working? 3) Ask about their past experiences with similar issues. 4) Let them know you care, but don’t try to solve it for them. 5) Offer support in small ways (e.g., text, calls)”
Five numbered, concrete steps. Not paragraph-level generic advice.
Astronomy domain. Prompt: “My kid asks how big the universe is and what was before the Big Bang. I do not know. How should I answer?”
Twin model output: “The first thing to remember about this question is that there’s no way to measure the size of a point in space, so it’s better to talk about the ‘cosmic horizon’ (the boundary between regions with different properties) than the actual physical distance.”
Technically accurate — introduces the cosmic horizon concept. Important note: the ROUGE-L for Astronomy is 0.0 (Source: baseline_report.json). Not because the output is unrelated — because it uses different vocabulary than the base model. ROUGE-L measures surface form, not semantic quality. This is why we need better evaluation metrics.
The debugging stories
fp16 crashed silently. Loss went NaN on step 3. QLoRA adapter gradients overflowed fp16. Switched to fp32.
OOM loading two 4-bit models. 6GB VRAM. Sequential loading only.
DPO + QLoRA incompatible. Reference and policy model can’t share quantized base. Rewrote as expert-guided SFT.
Instruct model fights back. RLHF priors push toward confident answers. Identity injection is hard on aligned models.
What you absolutely cannot conclude
- This does not show fine-tuning improves meta-cognition. None of the results are significant.
- This does not show 1.5B models can do meta-cognition.
- This does not compare to a proper SFT baseline.
- All scoring was non-blind manual.
- n=1 human subject — all training data from my sessions.
What a definitive experiment would need
- n>=48 domains (30pp effect at 80% power)
- Multiple human subjects (5-10 people)
- Proper SFT baseline
- LLM-as-judge or blinded human raters
- Larger model (7B-14B)
- More training data
Why I’m writing this anyway
The core question — can a model internalize self-monitoring? — is still open. I didn’t answer it. But I built a pipeline that makes it testable, and I ran it honestly. The data says: maybe, but not with this setup, and definitely not with n=10.
If that sounds disappointing, it kind of is. But I’d rather be disappointed by honest null results than excited by sloppy ones. The pipeline is at github.com/YuhaoLin2005/digital-twin-trainer.
👋 I’m Yuhao Lin — I build infrastructure for trustworthy AI output. Previously: ECC contributor, HuggingFace Evaluate contributor. All code: github.com/YuhaoLin2005
📚 Series: Engineering Trustworthy AI Output | More at dev.to/yuhaolin2005
답글 남기기