Topic
Local AI
Models, inference runtimes, acceleration, quantization, context windows, and routing experiments on hardware we actually own.
4 field reportsChronological feed
Squeezing Qwen3.8 Until It Cried: A 46.9% Speedup and One Optimization Too Far
A 46.9% Qwen3.8 speedup, a production qualification, and the tempting 64K shortcut that failed when the benchmark got serious.
Promote the full-vocabulary checkpoint: 46.88% faster decode, qualified behavior, and no support for the later 64K shortcut.
210K Context on a 32 GB Arc Pro B70
Three identical cold starts, 179,525 prompt tokens, and a rollback regex that briefly became the main character.
Admitted — the production route passed exact near-180K recall and the complete recovery, vision, reasoning, tool, and MTP suite.
How ubatch=1024 Won the Benchmark and Lost the Ability to Speak
A 6% prefill improvement, five failed conversations, and one tiny silicon cult chanting the unused32 token.
Rejected — the narrow prefill win concealed catastrophic generation and tool-call failures.
Qwen3.8 Flash-Next on One Arc Pro B70
Fifteen tokens per second, one extremely narrow memory edge, and several acceleration ideas asked to leave the laboratory.
Promising isolated batch-worker profile, not a production route.