> 한국어: READMEko.md > > Ppaso (빠소) = short for ppareunsori (빠른소리, "fast voice" in Korean). Real-time Korean TTS for edge NPU devices.
Source du modèle
Extrait de la source
한국어: READMEko.md > > Ppaso (빠소) = short for ppareunsori (빠른소리, "fast voice" in Korean). Real-time Korean TTS for edge NPU devices.
Sources
1 sourceVérifié 6 sept.
Artefacts du modèle
4 artefactsExtraits de sources
2 extraitsonnx/variance.onnx
onnx · 79,2 KB · SHA-256 378e1b0af2fa…5f33 · Hugging Face
--- license: apache-2.0 language: - ko library_name: pytorch tags: - text-to-speech - tts - korean - edge - npu - rk3576 - onnx - rknn - lightweight pipeline_tag: text-to-speech --- # Ppaso-TTS — Korean lightweight TTS (RK3576 NPU friendly) > 한국어: [README_ko.md](./README_ko.md) > > **Ppaso (빠소)** = short for *ppareunsori* (빠른소리, "fast voice" in Korean). Real-time Korean TTS for edge NPU devices. ## What it is for When pairing an LLM with TTS on an edge device (e.g. NanoPi R76S) for a Korean voice assistant, **TTS inference often becomes the latency bottleneck**. Ppaso-TTS is a lightweight TTS designed to remove that bottleneck by running on the NPU. Designed for **speed and small footprint over naturalness**. Single voice, Korean only. ## Highlights - **Korean only**, single female voice. Speed and lightweightness prioritized over quality (robotic timbre included). - **~20× real-time on NPU** — ~305 ms for a 5.94-second utterance on RK3576 (RTF 0.052). - **21 MB on-device model** (RKNN) — fits comfortably on RAM-constrained edge devices. - **Two backends**: RKNN NPU (RK3576 / RK3588) / ONNX CPU (anywhere). - **Any-length input** — built-in chunking handles long paragraphs without truncation. - **Streaming mode** — feed LLM tokens as they arrive; emit wav as soon as a sentence completes (voice assistant friendly). - **🆕 Lipsync animation** — companion module `ppaso_lipsync` generates 2D mouth + 1D eye-blink animation curves directly from TTS output. One curve drives two render routes: 2D sprite (atlas) and 3D blendshape (VRM). Pure numpy, edge-friendly. See [demo video](#lipsync-animation) below. - **Custom vocabulary focus** — training corpus is self-synthesized via a CosyVoice teacher (not an external dataset). Naturalness is sacrificed, but the data distribution for domain-specific vocabulary (industrial safety, technical terms, etc.) can be controlled. ## What's New (v8 — 2026-05-30) v8 is a from-scratch retrain on a refined training corpus. No architecture / config change. - **Cleaner pronunciation for liaison, aspiration and morpheme-boundary patterns** (*맑→말금*, *흙이→흘기*, *좋고 / 좋다* aspiration). The G2P lexicon was rebuilt with explicit phonology rules; v8 is the first acoustic model trained on the new label distribution. - **Stabilized word endings** — short utterances (e.g. *이 약 좋다.*) occasionally lost their final syllable in the teacher-s...
Source context: 76 downloads · 3 likes · Pipeline text-to-speech · Library pytorch · Repo akamotaco/ppaso-tts-v1