John Paul Wile
I break things quickly and write about AI tech.
Full-Stack Developer & Writer · Colorado Springs, CO
Latest Posts
View all →May 27, 2026
Qwen3.6-35B-A3B-MTP at 70 tok/s on an RTX 3070
I’ve been running Qwen3.6-35B-A3B-MTP on my RTX 3070 (8GB VRAM) for a while now and previously documented getting ~55 tok/s....
May 20, 2026
Multi-Token Prediction MTP in llama.cpp How It Works and How to Use It
PR #22673 just landed in upstream llama.cpp, and it is a big deal....
May 16, 2026
The Architecture Breakthrough Nobody in Local LLM is Talking About
If you’ve been following the local LLM space this year, you’ve probably been tracking parameter counts, quantization methods, and hardware...
About Me
Full-Stack Developer & Tech Lead from Colorado Springs. I build things, break them, and write about what I learn along the way.
Learn MoreStay Updated
Get my latest thoughts on AI, development, and technology delivered to your inbox.
Subscribe on Substack