Tag archive
#llama-cpp
llama.cpp · 2 field reports
Chronological feed
How ubatch=1024 Won the Benchmark and Lost the Ability to Speak
A 6% prefill improvement, five failed conversations, and one tiny silicon cult chanting the unused32 token.
Rejected — the narrow prefill win concealed catastrophic generation and tool-call failures.
Qwen3.8 Flash-Next on One Arc Pro B70
Fifteen tokens per second, one extremely narrow memory edge, and several acceleration ideas asked to leave the laboratory.
Promising isolated batch-worker profile, not a production route.