Skip to content

Lab Notes6 field reports

All Lab Notes

Tag archive

#vllm-xpu

vLLM XPU · 2 field reports

Filed under Local AI (2). All tags.

Chronological feed

  1. Local AI · 10 min read

    Squeezing Qwen3.8 Until It Cried: A 46.9% Speedup and One Optimization Too Far

    A 46.9% Qwen3.8 speedup, a production qualification, and the tempting 64K shortcut that failed when the benchmark got serious.

    Promote the full-vocabulary checkpoint: 46.88% faster decode, qualified behavior, and no support for the later 64K shortcut.

  2. Local AI · 10 min read

    210K Context on a 32 GB Arc Pro B70

    Three identical cold starts, 179,525 prompt tokens, and a rollback regex that briefly became the main character.

    Admitted — the production route passed exact near-180K recall and the complete recovery, vision, reasoning, tool, and MTP suite.