How I Cut MCP Token Usage by 91% (and Learned a Humbling Lesson About Tokenizers)
The Problem When you add MCP servers to your AI coding agent, each one dumps its full JSON schema into context. 255 tools across all servers = 39,964 tokens. On a 128K context window, that's 31% gone before you type a single character. I was literally paying for JSON syntax overhead. Every API call included {"content":[{"type":"text","text":"..."}]} — 80 tokens to deliver 6 tokens of data. The Solution: mcptoon I built a CLI that sits between your agent and MCP servers. It does three things: 1. Schema Compression (SLIM format) Instead of full JSON schemas, mcptoon presents tools in a compact pipe-delimited format: # Full JSON (287 tokens per tool): {"name":"search","description":"Search the web","inputSchema":{"type":"object","properties":{"q":{"type":"string","description":"Query"},"n":{"type":"number"}},"required":["q"]}} # SLIM format (26 tokens): search|q:s*|n:n 255 tools: 39,964 → 3,511 tokens. 91% saved. Verified with tiktoken.get_encoding("cl100k_base"). 2. Zero-Context Discovery Schemas live on disk in ~/.mcptoon/config.json. Your agent runs mcptoon manifest --slim to see what's available, then mcptoon call '{"param":"value"}' --toon to execute. Only the compressed output enters context. 3. TOON Format for Results Tool results come back as human-readable key-value pairs instead of nested JSON: # JSON result (80 tokens): {"content":[{"type":"text","text":"{\"name\":\"react\",\"stars\":219000}"}]} # TOON result (12 tokens): name: react stars: 219000 The Lesson My first version replaced null with ∅ (the empty set symbol). I thought I was being clever. Then someone ran it through tiktoken: null = 1 token, ∅ = 2 tokens. I was literally increasing token count and calling it optimization. The HN community called me out on it. Fair enough — I hadn't measured before shipping. Now everything is tiktoken-verified. true stays true. null stays null. No unicode tricks. Cost Impact At GPT-4o pricing ($5/M tokens): Without mcptoon: 25 requests × 40K schema tokens = 1M tokens = $5 With SLIM: 25 requests × 3.5K = 87.5K tokens = $0.44 Daily savings (100 sessions): ~$540/month Getting Started pip install mcptoon mcptoon add fetch --stdio npx -y @anthropic/mcp-fetch mcptoon manifest --slim # see what's available, compact mcptoon call fetch fetch '{"url":"https://example.com"}' --toon Works with any agent that can run shell commands. One config file for all agents — no more reconfiguring for Claude Code vs Cursor vs OpenCode. 3000 lines of Python, 309 tests, zero dependencies. GitHub: https://github.com/activeing123/mcptoon What's Next More MCP servers being added to the default config Working on a rigorous quality benchmark (currently only have token counts, not LLM accuracy) Open to feedback on the SLIM format — is there a better encoding? What's your biggest MCP token waste? How are others handling this?
This is a summary aggregated from Dev.to. Read the complete article on the original site:
Read full article at Dev.to