longcat-audiodit-license
Modellquelle
Quellenauszug
longcat-audiodit-license
Quellen
1 QuelleVerifiziert 22. Sept.
Modellartefakte
2 Artefaktemodel-00001-of-00002.safetensors
safetensors · 4,66 GB · SHA-256 f385720e2a56…c2c0 · Hugging Face
Herunterladenmodel-00002-of-00002.safetensors
safetensors · 650 MB · SHA-256 150e876b5247…a46a · Hugging Face
HerunterladenQuellenauszüge
2 Auszüge--- license: other license_name: longcat-audiodit-license base_model: meituan-longcat/LongCat-AudioDiT-1B tags: - audio - text-to-speech - tts - environmental-tts - flow-matching - dit library_name: transformers pipeline_tag: text-to-speech --- # LongCat-AudioDiT Env-TTS — 10000-step (independent noise / room-RIR / mic-IR augmentation) Fine-tune of [meituan-longcat/LongCat-AudioDiT-1B](https://huggingface.co/meituan-longcat/LongCat-AudioDiT-1B) for the **three-stream env-tts task**: given a reference environment audio, a reference speaker audio, and three text streams (env caption / speaker caption / target speech text), generate target speech that places the target text in the referenced environment with the referenced speaker timbre. > **The "rir" variant — full acoustic-scene augmentation.** This checkpoint is > trained with THREE INDEPENDENT Bernoulli augmentation dimensions (p=0.5 each): > additive **noise**, **room RIR** reverb, and **microphone-IR colouration**, > applied in the chain *RIR → mic IR → noise* (noise stays un-coloured). > The spk reference gets its own independent draw; **env and target share one > realization** (same noise clip + same RIR + same mic IR + same SNR/wet — one > acoustic scene), so the model learns env-consistent generation. Compare with > [-ablation](https://huggingface.co/ChristianYang/LongCat-AudioDiT-Env-TTS-1B-ablation) > (no augmentation) and > [-augment](https://huggingface.co/ChristianYang/LongCat-AudioDiT-Env-TTS-1B-augment) > (spk-only noise+RIR) to isolate the effect. ## Differences from the base model The transformer adds **six learnable boundary tokens** (three latent-space, three text-space): ``` latent sequence : [<boe> z_env <bos> z_spk <bon> z_target] text sequence : [<boe_t> env_text_emb <bos_t> spk_text_emb <bon_t> target_text_emb] ``` `encode_multistream_text(env, spk, target, drop_env_text=…, drop_spk_text=…, drop_target_text=…)` is the new entry-point. `AudioDiTModel.forward(...)` also accepts a pre-assembled `prompt_latent` (replaces `prompt_audio`) so the inference path can feed the boundary-tokenized three-stream prompt directly. ## Training summary | Field | Value | |---|---| | Steps | 10000 (~1.7 epochs of the 379k-row train split) | | Effective batch | 16 × grad_accum 4 × 1 GPU = **64 rows / step** | | Learning rate | cosine 5e-5 (warmup 250) | | AdamW | β₁=0.9, β₂=0.999, wd=0.0...
Source context: 7 downloads · 0 likes · Pipeline text-to-speech · Library transformers · Repo humanify/LongCat-AudioDiT-Env-TTS-1B-rir