Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions .github/workflows/metal-wheel.yml
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@ name: Apple Silicon Metal wheel
qwentts_ref:
description: qwentts.cpp revision (must match the binding's verified ABI)
type: string
default: 7df559a8ca25f66fee02970514ebe5f01dee9055
default: 6fae92914045cd83364d2845ceaa0f7969727319
workflow_call:
inputs:
pypi:
Expand All @@ -16,7 +16,7 @@ name: Apple Silicon Metal wheel
default: false
qwentts_ref:
type: string
default: 7df559a8ca25f66fee02970514ebe5f01dee9055
default: 6fae92914045cd83364d2845ceaa0f7969727319

permissions:
contents: read
Expand All @@ -26,7 +26,7 @@ jobs:
runs-on: macos-14
env:
MACOSX_DEPLOYMENT_TARGET: "14.0"
QWENTTS_REF: ${{ inputs.qwentts_ref || '7df559a8ca25f66fee02970514ebe5f01dee9055' }}
QWENTTS_REF: ${{ inputs.qwentts_ref || '6fae92914045cd83364d2845ceaa0f7969727319' }}
steps:
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
with:
Expand Down
6 changes: 3 additions & 3 deletions .github/workflows/publish-hf-wheels.yml
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ name: Publish Hugging Face Wheels
qwentts_ref:
description: qwentts.cpp commit, tag, or branch to build
required: true
default: 7df559a8ca25f66fee02970514ebe5f01dee9055
default: 6fae92914045cd83364d2845ceaa0f7969727319
hf_repo_id:
description: Hugging Face dataset repo for wheel hosting
required: true
Expand All @@ -17,13 +17,13 @@ permissions:

env:
HF_WHEEL_REPO_ID: ${{ inputs.hf_repo_id || vars.HF_WHEEL_REPO_ID || 'andito/qwentts-cpp-python-wheels' }}
QWENTTS_REF: ${{ inputs.qwentts_ref || '7df559a8ca25f66fee02970514ebe5f01dee9055' }}
QWENTTS_REF: ${{ inputs.qwentts_ref || '6fae92914045cd83364d2845ceaa0f7969727319' }}

jobs:
build-metal-wheel:
uses: ./.github/workflows/metal-wheel.yml
with:
qwentts_ref: ${{ inputs.qwentts_ref || '7df559a8ca25f66fee02970514ebe5f01dee9055' }}
qwentts_ref: ${{ inputs.qwentts_ref || '6fae92914045cd83364d2845ceaa0f7969727319' }}

build-wheels:
name: Linux ${{ matrix.arch }} CUDA ${{ matrix.cuda_version }} HF wheel
Expand Down
6 changes: 3 additions & 3 deletions .github/workflows/publish.yml
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ name: Publish
qwentts_ref:
description: qwentts.cpp commit, tag, or branch to build
required: true
default: 7df559a8ca25f66fee02970514ebe5f01dee9055
default: 6fae92914045cd83364d2845ceaa0f7969727319

permissions:
contents: read
Expand All @@ -19,7 +19,7 @@ jobs:
uses: ./.github/workflows/metal-wheel.yml
with:
pypi: true
qwentts_ref: ${{ inputs.qwentts_ref || '7df559a8ca25f66fee02970514ebe5f01dee9055' }}
qwentts_ref: ${{ inputs.qwentts_ref || '6fae92914045cd83364d2845ceaa0f7969727319' }}

build-wheels:
name: Linux ${{ matrix.arch }} CUDA 12.8 wheel
Expand All @@ -46,7 +46,7 @@ jobs:
QWENTTS_CPP_BACKEND: cuda
QWENTTS_CPP_BUILD_JOBS: ${{ matrix.build_jobs }}
QWENTTS_CPP_WHEEL_BUILD_TAG: 1cu128
QWENTTS_REF: ${{ inputs.qwentts_ref || '7df559a8ca25f66fee02970514ebe5f01dee9055' }}
QWENTTS_REF: ${{ inputs.qwentts_ref || '6fae92914045cd83364d2845ceaa0f7969727319' }}
CUDA_ARCHITECTURES: ${{ matrix.cuda_architectures }}
steps:
- uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
Expand Down
8 changes: 4 additions & 4 deletions .github/workflows/wheels.yml
Original file line number Diff line number Diff line change
Expand Up @@ -6,13 +6,13 @@ name: qwentts-cpp-python wheels
qwentts_ref:
description: qwentts.cpp commit, tag, or branch to build
required: true
default: 7df559a8ca25f66fee02970514ebe5f01dee9055
default: 6fae92914045cd83364d2845ceaa0f7969727319

jobs:
macos-metal:
uses: ./.github/workflows/metal-wheel.yml
with:
qwentts_ref: ${{ inputs.qwentts_ref || '7df559a8ca25f66fee02970514ebe5f01dee9055' }}
qwentts_ref: ${{ inputs.qwentts_ref || '6fae92914045cd83364d2845ceaa0f7969727319' }}

linux-cuda:
name: Linux ${{ matrix.arch }} CUDA ${{ matrix.cuda_version }} wheel
Expand Down Expand Up @@ -91,7 +91,7 @@ jobs:
QWENTTS_CPP_BACKEND: cuda
QWENTTS_CPP_BUILD_JOBS: ${{ matrix.build_jobs }}
QWENTTS_CPP_WHEEL_BUILD_TAG: ${{ matrix.wheel_build_tag }}
QWENTTS_REF: ${{ inputs.qwentts_ref || '7df559a8ca25f66fee02970514ebe5f01dee9055' }}
QWENTTS_REF: ${{ inputs.qwentts_ref || '6fae92914045cd83364d2845ceaa0f7969727319' }}
CUDA_ARCHITECTURES: ${{ matrix.cuda_architectures }}
steps:
- uses: actions/checkout@v4
Expand Down Expand Up @@ -175,7 +175,7 @@ jobs:
QWENTTS_CPP_BACKEND: cpu
QWENTTS_CPP_BUILD_JOBS: "2"
QWENTTS_CPP_WHEEL_BUILD_TAG: 1cpu
QWENTTS_REF: ${{ inputs.qwentts_ref || '7df559a8ca25f66fee02970514ebe5f01dee9055' }}
QWENTTS_REF: ${{ inputs.qwentts_ref || '6fae92914045cd83364d2845ceaa0f7969727319' }}
steps:
- uses: actions/checkout@v4

Expand Down
27 changes: 21 additions & 6 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -143,8 +143,23 @@ python -m twine check --strict wheelhouse/*.whl
### Native ABI compatibility

The CI wheel build defaults to qwentts.cpp
`7df559a8ca25f66fee02970514ebe5f01dee9055`, which retains ABI v2 and includes
the latest static-graph, streaming-decode, and widened voice-route changes.
`6fae92914045cd83364d2845ceaa0f7969727319` (September 28, 2026), which requires
ABI v5 and includes the upstream batched compute worker, codec memory
controls, language/model queries, and latest ggml fork update.

ABI v5 changes the initialization and synthesis struct layouts; libraries
built from the previous ABI v2 pin must be rebuilt. Python's `do_sample=False`
and `subtalker_do_sample=False` now select greedy decoding by setting the
corresponding native temperature to zero. Buffered codec chunk size is set
once with `QwenTTS(..., codec_chunk_sec=24.0)` or
`QwenTTS.from_pretrained(..., codec_chunk_sec=24.0)`, rather than on
`synthesize()`. Zero selects the native default. The codec derives its own
left context; `codec_left_context_sec` is no longer accepted. Streaming's
`codec_chunk_sec` still controls Python output packet size independently.
The wrapper continues to serialize requests on each context (native
`max_batch=1`). `tts.language_names()` lists the model's languages (synthesis
also accepts `lang="auto"`), and `tts.model_type()` reports `base`,
`custom_voice`, or `voice_design`.

The loader verifies this native revision before calling functions that write
ctypes parameter buffers. Upstream does not expose an ABI-version or struct-size
Expand Down Expand Up @@ -229,15 +244,15 @@ limit flushes any remaining audio as a short packet, including utterances
shorter than the requested first packet. Cancellation or errors discard the
unfinished packet; closing the iterator requests native cancellation.

Packet assembly happens in Python, using the existing verified ABI v2 library.
Packet assembly happens in Python, using the verified ABI v5 library.
The native decoder still emits its fixed 1→2→4→8-frame ramp, then 8-frame
chunks. Consequently, first packets of 1, 2, 4, and 8 frames become available
after native output has reached 1, 3, 7, and 15 frames respectively (or earlier
at end-of-speech). Packet boundaries preserve every PCM sample but do not
change native decode scheduling. Later packets may become available together
when a native callback spans several packet boundaries; this is not a timed
playback scheduler. `codec_left_context_sec` is ignored by the stateful native
stream. Buffered `synthesize()` retains its native codec chunking behavior.
playback scheduler. Buffered `synthesize()` uses the native codec chunk size
configured when the context is created; the native codec derives its left context.

`last_stream_profile` keeps `first_callback_*` and `callback_count` for raw
native callbacks. `first_packet_ready_ms`, `first_packet_audio_s`, and
Expand Down Expand Up @@ -281,7 +296,7 @@ the larger first packet is assembled by the binding.

## Cached voice references

qwentts.cpp ABI v2 can skip reference WAV encoding for Base voice cloning by
qwentts.cpp can skip reference WAV encoding for Base voice cloning by
passing precomputed latents:

- `.spk`: raw float32 speaker embedding from `qwen-codec --talker`
Expand Down
62 changes: 12 additions & 50 deletions scripts/build_native.py
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,6 @@
import argparse
from contextlib import contextmanager
import os
import re
import shutil
import shlex
import subprocess
Expand Down Expand Up @@ -32,7 +31,7 @@ def metal_shader_compatibility(source: Path, enabled: bool):
if not enabled:
yield
return
shader = source / "ggml/src/ggml-metal/ggml-metal.metal"
shader = source / "ggml/src/ggml-metal/kernels/unary.metal"
original = shader.read_text()
old = "dst_ptr[i0] = (T) args.val;"
new = "dst_ptr[i0] = (T) ((TC) args.val);"
Expand Down Expand Up @@ -83,59 +82,22 @@ def native_logging_compatibility(source: Path):

@contextmanager
def native_diagnostic_compatibility(source: Path):
"""Route the pinned source's remaining direct diagnostics through qt_log.
"""Keep a missing tokenizer visible at error log level.

These headers predate qwentts.cpp's callback API. Keep severity when
converting them, and restore the checkout once the wheel is built.
Upstream routes diagnostics through qt_log now, but reports this load
failure at INFO. Restore the checkout after building.
"""
expected = {
"audio-io.h": 10, "audio-resample.h": 2, "bpe.h": 9,
"code-predictor-forward.h": 5, "code-predictor-weights.h": 4,
"convnext-block.h": 3, "dac-decoder-v2.h": 2,
"encoder-downsample.h": 2, "encoder-transformer.h": 3,
"gguf-weights.h": 9, "graph-arena.h": 1, "kv-cache.h": 3,
"prompt-builder.h": 13, "quantizer-decode.h": 4,
"quantizer-encode.h": 2, "rvq-file.h": 7, "seanet-encoder.h": 3,
"speaker-encoder-extract.h": 6, "speaker-encoder-weights.h": 3,
"talker-forward.h": 7, "talker-weights.h": 3,
"tokenizer-transformer.h": 3, "wav.h": 7, "weight-ctx.h": 2,
}
pattern = re.compile(r"fprintf\(stderr,\s*(.*?)\);", re.DOTALL)
originals = {}
path = source / "src/bpe.h"
original = path.read_text()
old = 'qt_log(QT_LOG_INFO, "[BPE] Tokenizer not found in %s", gguf_path);'
new = 'qt_log(QT_LOG_ERROR, "[BPE] Tokenizer not found in %s", gguf_path);'
if original.count(old) != 1:
raise SystemExit("Native tokenizer diagnostic changed; review the logging compatibility fix")
try:
for name, count in expected.items():
path = source / "src" / name
original = path.read_text()
if len(pattern.findall(original)) != count or not original.startswith("#pragma once\n"):
raise SystemExit(f"Native diagnostics changed in {name}; review the logging compatibility fix")

def replace(match):
args = match.group(1)
format_match = re.search(r'"((?:[^"\\]|\\.)*)"', args)
if format_match is None:
raise SystemExit(f"Native diagnostic format changed in {name}")
message = format_match.group(1).lower()
if "warning" in message or "no spk_enc." in message:
level = "QT_LOG_WARN"
elif any(word in message for word in (
"fatal", "failed", "cannot", "oom", "unsupported",
"not a valid", "not found", "no audio data", "unknown format",
)):
level = "QT_LOG_ERROR"
else:
level = "QT_LOG_INFO"
# qt_log and the Python trampoline each add the line ending.
args = args.replace(r'\n"', '"')
return f"qt_log({level}, {args});"

transformed = pattern.sub(replace, original)
transformed = transformed.replace("#pragma once\n", '#pragma once\n#include "qt-error.h"\n', 1)
originals[path] = original
path.write_text(transformed)
path.write_text(original.replace(old, new))
yield
finally:
for path, original in originals.items():
path.write_text(original)
path.write_text(original)


def find_first(root: Path, patterns: list[str]) -> Path | None:
Expand Down
Loading
Loading