Tag archive
#gptq
GPTQ · 1 field report
Chronological feed
Squeezing Qwen3.8 Until It Cried: A 46.9% Speedup and One Optimization Too Far
A 46.9% Qwen3.8 speedup, a production qualification, and the tempting 64K shortcut that failed when the benchmark got serious.
Promote the full-vocabulary checkpoint: 46.88% faster decode, qualified behavior, and no support for the later 64K shortcut.