How I Cut MCP Token Usage by 91% (and Learned a Humbling Lesson About Tokenizers)

작성자

카테고리:

← 피드로
DEV Community · MCP Token Saver · 2026-08-16 개발(SW)

MCP Token Saver

The Problem

When you add MCP servers to your AI coding agent, each one dumps its full JSON schema into context. 255 tools across all servers = 39,964 tokens. On a 128K context window, that’s 31% gone before you type a single character.

I was literally paying for JSON syntax overhead. Every API call included {"content":[{"type":"text","text":"..."}]} — 80 tokens to deliver 6 tokens of data.

The Solution: mcptoon

I built a CLI that sits between your agent and MCP servers. It does three things:

1. Schema Compression (SLIM format)

Instead of full JSON schemas, mcptoon presents tools in a compact pipe-delimited format:

# Full JSON (287 tokens per tool):
{"name":"search","description":"Search the web","inputSchema":{"type":"object","properties":{"q":{"type":"string","description":"Query"},"n":{"type":"number"}},"required":["q"]}}

# SLIM format (26 tokens):
search|q:s*|n:n

Enter fullscreen mode Exit fullscreen mode

255 tools: 39,964 → 3,511 tokens. 91% saved. Verified with tiktoken.get_encoding("cl100k_base").

2. Zero-Context Discovery

Schemas live on disk in ~/.mcptoon/config.json. Your agent runs mcptoon manifest --slim to see what’s available, then mcptoon call <server> <tool> '{"param":"value"}' --toon to execute. Only the compressed output enters context.

3. TOON Format for Results

Tool results come back as human-readable key-value pairs instead of nested JSON:

# JSON result (80 tokens):
{"content":[{"type":"text","text":"{\"name\":\"react\",\"stars\":219000}"}]}

# TOON result (12 tokens):
name: react
stars: 219000

Enter fullscreen mode Exit fullscreen mode

The Lesson

My first version replaced null with (the empty set symbol). I thought I was being clever. Then someone ran it through tiktoken: null = 1 token, = 2 tokens. I was literally increasing token count and calling it optimization.

The HN community called me out on it. Fair enough — I hadn’t measured before shipping. Now everything is tiktoken-verified. true stays true. null stays null. No unicode tricks.

Cost Impact

At GPT-4o pricing ($5/M tokens):

  • Without mcptoon: 25 requests × 40K schema tokens = 1M tokens = $5
  • With SLIM: 25 requests × 3.5K = 87.5K tokens = $0.44
  • Daily savings (100 sessions): ~$540/month

Getting Started

pip install mcptoon
mcptoon add fetch --stdio npx -y @anthropic/mcp-fetch
mcptoon manifest --slim    # see what's available, compact
mcptoon call fetch fetch '{"url":"https://example.com"}' --toon

Enter fullscreen mode Exit fullscreen mode

Works with any agent that can run shell commands. One config file for all agents — no more reconfiguring for Claude Code vs Cursor vs OpenCode.

3000 lines of Python, 309 tests, zero dependencies.

GitHub: https://github.com/activeing123/mcptoon

What’s Next

  • More MCP servers being added to the default config
  • Working on a rigorous quality benchmark (currently only have token counts, not LLM accuracy)
  • Open to feedback on the SLIM format — is there a better encoding?

What’s your biggest MCP token waste? How are others handling this?

원문에서 계속 ↗