Skip to content

Lab Notes6 field reports

All Lab Notes

Tag archive

#benchmarking

Benchmarking · 3 field reports

Filed under Local AI (3). All tags.

Chronological feed

  1. Local AI · 10 min read

    Squeezing Qwen3.8 Until It Cried: A 46.9% Speedup and One Optimization Too Far

    A 46.9% Qwen3.8 speedup, a production qualification, and the tempting 64K shortcut that failed when the benchmark got serious.

    Promote the full-vocabulary checkpoint: 46.88% faster decode, qualified behavior, and no support for the later 64K shortcut.

  2. Local AI · 8 min read

    How ubatch=1024 Won the Benchmark and Lost the Ability to Speak

    A 6% prefill improvement, five failed conversations, and one tiny silicon cult chanting the unused32 token.

    Rejected — the narrow prefill win concealed catastrophic generation and tool-call failures.

  3. Local AI · 9 min read

    Qwen3.8 Flash-Next on One Arc Pro B70

    Fifteen tokens per second, one extremely narrow memory edge, and several acceleration ideas asked to leave the laboratory.

    Promising isolated batch-worker profile, not a production route.