Why a 14B Draft Model Slows Down Your Mac Mini Instead of Speeding It Up
Why a 14B Draft Model Slows Down Your Mac Mini Instead of Speeding It Up
The Draft Model Assumption Are you dropping a small draft model onto your Mac mini and expecting the same speedup you got on your NVIDIA box? Wondering why your tokens per second went down instead of up? Speculative decoding works like this on CUDA hardware: a small, fast draft model guesses several tokens ahead, and the big target model checks those guesses in one batched pass.