Serve the pinned upstream GGML server on vulkan/cuda/cpu #4
No reviewers
Labels
No labels
bug
documentation
duplicate
enhancement
good first issue
help wanted
invalid
question
wontfix
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
nsaspy/Qwen3-TTS_server!4
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "feature/ggml-backends"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Replaces the PyTorch/ROCm container backends with the published qwentts.cpp images pinned to 71ad93d591a2811f35db77e27c02acba091c9e9b.
Validated live on RX 5500 XT (gfx1012/RDNA1): Vulkan init, 896MB KV-cache allocation (the whisper.cpp#3611 crash path), GPU synthesis (~2.5s), and realtime PCM streaming all pass with zero kernel faults.
Tests: 33 passed, flake check green.
Pull request closed