Skip to content

BetterScale

Less host waiting. More device execution. Modular graph, replay and metadata-preparation optimizations for vLLM Ascend.

Install vllm-betterscale and use betterscale.worker.Worker with your native vLLM serving command. The implementation lives directly in src/betterscale/.

Incremental DeepSeek V4 serving improvements on vLLM + vLLM-Ascend. Keep the mature engine. Measure each change. Keep improvements that survive correctness, memory and service-quality comparisons.

Install and start

See public package instructions for the version-pinned install and complete native TP8 / DP8 launch examples. Version 0.4.2 adds backing-level KV clear and a privately selected HC-pre host tiler on TP. It preserves the donor installation, the single target/draft graph pool and the 1 GiB safety reserve. See memory-patch mechanics and native build.

One Worker, model-local overrides

All current source deployments select betterscale.worker.Worker. Native model and serving configuration choose DSV4 TP/DP, Qwen owned mixed FULL (no MTP), or Qwen native MTP2. No patch-specific Worker class or private profile is required. See composition, lifecycle and migration. Unsupported configurations fail before installation; missing qualified native libraries never cause a silent fallback to a different state layout.

Qwen hybrid TP2

The no-MTP route owns a K-V GDN state pool and dynamic mixed FULL graphs keyed by token capacity plus alternating metadata bank, not request partitions. It uses wave-shared FIA planning and supports align-mode prefix caching. See the service contract and launcher. Release 0.5.1 includes the qualified Qwen native libraries and metadata-selected prefill/mixed MatmulAllReduce; pure decode keeps its native path. After installation:

python -m pip install --no-deps vllm-betterscale==0.5.1
python -m betterscale serve-qwen /models/Qwen3.8-27B --devices 0,1

Use the pinned Linux/aarch64 CANN 9.0.1 runtime; model weights remain separate. Native MTP2 remains a separate configuration with immutable convolution-weight packing; it does not activate owned GDN or mixed FULL.

The legacy betterscale.qwen_worker.Worker and MixedWorker imports now alias the same public Worker. In particular, no-MTP qwen_worker.Worker no longer selects the older native-layout single-request prefill implementation. It needs the owned route's native libraries and startup environment, even with APC off. Use a fresh process; never reuse a live state pool across implementations.

Latest matched APC-on SWE service study reports the pre-refactor qualified implementation; entry consolidation is not a new performance measurement. Historical mixed acceptance, dual-bank comparison, and GDN fusion follow-up retain their original configurations and comparator boundaries.

End-to-end service evidence

September14 HTTP acceptance compares retained programs with native donor, not one incremental patch against another. On the bounded same-host workloads, the retained TP8 path improves aggregate output throughput by 35.17%; the DP8 startup-prepared implementation now shipped in 0.3.1 improves it by 39.63%. The DP figure reuses its qualified run155; this is not a fresh wheel benchmark. DP8 0.3.0 measured −6.11% and remains in the report, not relabelled. All measured arms pass the retained32-question retrieval gate; decode latency tails do not improve universally. See all repeats, configuration and release provenance.

Pinned upstreams

Component Release Commit
vLLM v0.25.1 752a3a504485790a2e8491cacbb35c137339ad34
vLLM-Ascend v0.25.1rc1 9bf964cb4b87c8cd0d6852c41a55b3c29711fa95

These are unmodified release pins, not a claim that a new build reproduces our previously installed donor packages byte-for-byte. The Ascend pin is a release candidate, not a stable release. Submodule gitlinks are authoritative.

git clone --recurse-submodules https://github.com/vLLM-HUST/BetterScale.git
cd BetterScale
git submodule status

No build, package installation or accelerator job happens on checkout. Upstreams retain their own licenses and build instructions. Before comparing a fresh build with historical measurements, record the full software environment and establish an unchanged baseline on the same workload and hardware.

Working style

  • Keep upstream pins unchanged while testing an optimization.
  • Carry small, separately reversible patches in patches/; avoid a long-lived monolithic engine fork. Use native extension points where they fit.
  • Compare unchanged and patched runs using the same inputs and settings. Report output throughput, TTFT, output gaps, memory and correctness—not just a favorable kernel duration. Profiled runs are diagnostic, not timing baselines.
  • Package proven changes as a plugin when the actual extension boundary is clear. Core changes may remain explicit patches or become upstream contributions.
  • Keep model weights, credentials, datasets, build products and raw profiles out of Git. Publish compact evidence and compressed timeline references.

Starting evidence

Donor C32 investigation records the current cache, scheduler and FULL-graph observations. It does not establish a stable C32 throughput collapse or a measured benefit from broader FULL graph capture. The linked large artifacts remain in the original local workspace.

These starting hypotheses preceded the graph work below; they are not current completion or performance claims. Prefix-cache granularity remains outside the kept patch bundle.

Maintained serving entry and report

The kept patches now live in src/betterscale/, not just the historical prototype tree. They use vLLM's explicit worker_cls lifecycle and normal OpenAI server. No donor source files are modified, upgraded or rebuilt on startup. Each patch is a closed directory with its own install: compat_lcm, target_full, ordered_replay, qli_cpu, split_draft, cross_step. Worker owns their composition; importing a module does not install its hooks. See each module's README under src/betterscale/patches/ for its contract.

# In your existing vLLM-Ascend environment (no donor dependencies are upgraded):
python -m pip install --no-deps --no-build-isolation .

# Keep your native serving arguments; add only the worker class:
vllm serve /models/DeepSeek-V4-Flash <your-native-vllm-arguments> \
  --worker-cls betterscale.worker.Worker

Worker is the only public integration entry. There is no BetterScale CLI, environment setup, automatic plugin discovery, private profile or mandatory artifact directory. The package does not select Python/CANN, set HCCL/allocator variables, repair library paths, or change network settings. Without an explicit KV byte budget, it sizes KV from actual execution residency and physical headroom. Start with a working donor environment. A source install with --no-build-isolation needs existing setuptools>=77.0.3. Prefer the published distribution: pip install --no-deps vllm-betterscale==0.5.1.

The original TP admission remains bounded: TP8/EP/DSACP/K5, four seats,4128 token budget, max length<=524288, target FULL, native scheduler, native prefix caching supported. This entry change does not qualify arbitrary layouts or shapes. Manual KV bytes and bind address remain native settings; there is no patch-imposed minimum of four full-length resident requests. See the runbook for the complete native example and source-compatibility boundary.

To remove the patch, stop the service and return to your original native worker and command. Do not hot-unpatch a live process. Some pinned K5/TP8+SP combinations need the patch's LCM fix even to start; removing it is not necessarily a runnable same-configuration baseline. The historical isolated controls remain evidence, not a second production entry. Shared hosts still require lease/admission.

Matched K5 studies isolate20–21% draft-graph and another10–12% stable-receipt cycle improvement. Split context/query improves observed short mixed waves; large prefill has no general speedup claim. Both historical arms pass the retained 32-item OpenCompass retrieval gate. These are not proven maximum-concurrency, KV-capacity or stable end-to-end throughput improvements.

The Markdown report is canonical. Regenerate its self-contained HTML with python docs/render_report.py in a separate documentation environment containing Markdown==3.8.2; do not add documentation dependencies to the donor environment. The SVG figures are editable vector sources under docs/figures/.

DP8 stable decode continuation

The same Worker also has a separately gated TP1/DP8/EP8 path: two seats per rank,1026 local token budget,context<=524288,FULL target,DSpark K5 with native eager draft,DSACP off,native prefix caching supported. It combines native DSA FULL target with owned input slots and captured device preparation/metadata. It does not enable the rejected all-mode worker extension. Automatic physical KV sizing is shared with the TP entry; only TP installs the finite split-draft catalog. See the ownership protocol and native launch parameters. Prototype matched-cycle evidence and the packaged-worker acceptance are reported separately; a faster step is not by itself an end-to-end throughput claim.

License

BetterScale is open source under the Apache License 2.0. See third-party notices for the upstream execution paths adapted by the patches. The pinned upstream repositories retain their own licenses.

Qwen35 default: resident State + balanced decode attention (development source)

Qwen3.5-35B-A3B's default serve-qwen entry uses BetterScale-owned StateTensor lanes with balanced decode attention, the qualified native model, MTP2, asynchronous scheduler and FULL graphs. --runtime auto selects this route from model config; Qwen27 is unchanged. The source, launch configuration and complete pinned donor adaptation are packaged under src/betterscale; no experiment controller or external livemodule import is required. See the deployment instructions.

The qualified E16/R20 configuration retains262144 context and pooled regular-attention pages. The combined C16/900s SWE point measured442.98 output tokens/s/chip, P90 decode62.91 tokens/s/user and TTFT P95588.66ms. This is one observation, not a paired incremental-speedup or SWE answer-quality certification. Native graph capture remains in use. Arbitrary historical checkpoints and EOS/stop rollback are not implemented. The earlier standalone live root remains research code, not the35B serving entry. Published PyPI0.5.1 predates this integration.

About

Incremental DeepSeek V4 improvements on pinned vLLM and vLLM-Ascend.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages