John Paul Wile

I break things quickly and write about AI tech.

Full-Stack Developer & Writer · Colorado Springs, CO

Latest Posts

View all →
May 27, 2026

Qwen3.6-35B-A3B-MTP at 70 tok/s on an RTX 3070

I’ve been running Qwen3.6-35B-A3B-MTP on my RTX 3070 (8GB VRAM) for a while now and previously documented getting ~55 tok/s....

May 20, 2026

Multi-Token Prediction MTP in llama.cpp How It Works and How to Use It

PR #22673 just landed in upstream llama.cpp, and it is a big deal....

May 16, 2026

The Architecture Breakthrough Nobody in Local LLM is Talking About

If you’ve been following the local LLM space this year, you’ve probably been tracking parameter counts, quantization methods, and hardware...

About Me

Full-Stack Developer & Tech Lead from Colorado Springs. I build things, break them, and write about what I learn along the way.

Learn More

Stay Updated

Get my latest thoughts on AI, development, and technology delivered to your inbox.

Subscribe on Substack