Taming Qwen Overthinking with GBNF Grammars
Qwen 3.6 overthinks in free-form mode, wasting thousands of tokens. Constraining thinking with a GBNF grammar reduced think-token consumption by 7x without losing code quality.
Qwen 3.6 overthinks in free-form mode, wasting thousands of tokens. Constraining thinking with a GBNF grammar reduced think-token consumption by 7x without losing code quality.
Benchmarked an RTX 3090 across six power limits and found that 280W saves ~70W with less than 1% performance loss for LLM inference. Below 200W, everything collapses.
Setting up a Jekyll Chirpy blog and building a Claude Code skill that converts working sessions into publishable drafts — privacy-first, drafts only.