Nemotron 30B on 24 GB: Benchmarks and a Quantization Quirk
Downloading and benchmarking NVIDIA's Nemotron-3-Nano-30B on a single RTX 3090 — including the confusing discovery about Unsloth Dynamic quants.
Downloading and benchmarking NVIDIA's Nemotron-3-Nano-30B on a single RTX 3090 — including the confusing discovery about Unsloth Dynamic quants.
Benchmarked Qwen3.6-27B Q4_K_M on an M1 Max via llama-cpp-turboquant. The same prod config that runs on my 3090 was already optimal — but llama-cli's interactive auto-mode nearly ate my disk.
Qwen 3.6 overthinks in free-form mode, wasting thousands of tokens. Constraining thinking with a GBNF grammar reduced think-token consumption by 7x without losing code quality.
Benchmarked an RTX 3090 across six power limits and found that 280W saves ~70W with less than 1% performance loss for LLM inference. Below 200W, everything collapses.
Setting up a Jekyll Chirpy blog and building a Claude Code skill that converts working sessions into publishable drafts — privacy-first, drafts only.