This is realtime on-device speech-to-speech AI model. It used a series to tools to achieve that. It uses a combination of voice activity detection, speech recognition, language models, and text-to-speech synthesis to...
Fuente del modelo
Extracto de la fuente
This is realtime on-device speech-to-speech AI model. It used a series to tools to achieve that. It uses a combination of voice activity detection, speech recognition, language models, and text-to-speech synthesis to...
Fuentes
1 fuenteVerificado 5 sept
Artefactos del modelo
33 artefactosExtractos de fuentes
2 extractosdata/voices/af_aoede.pt
pt · 511 KB · SHA-256 c03bd1a4c371…1cc4 · Hugging Face
--- license: mit language: - zho - eng - fra - spa - por - deu - ita - rus - jpn - kor - vie - tha - ara base_model: - hexgrad/Kokoro-82M - Qwen/Qwen2.5-0.5B-Instruct - openai/whisper-tiny.en - pyannote/segmentation-3.0 pipeline_tag: audio-to-audio tags: - speech-to-speech, - conversational-ai, - voice-chat, - realtime, - on-device - cpu --- # On Device Speech to Speech Conversational AI  This is realtime on-device speech-to-speech AI model. It used a series to tools to achieve that. It uses a combination of voice activity detection, speech recognition, language models, and text-to-speech synthesis to create a seamless and responsive conversational AI experience. The system is designed to run on-device, ensuring low latency and minimal data usage. <h2 style="color: yellow;">HOW TO RUN IT</h2> 1. **Prerequisites:** - Install Python 3.8+ (tested with 3.12) - Install [eSpeak NG](https://github.com/espeak-ng/espeak-ng/releases/tag/1.52.0) (required for voice synthesis) - Install Ollama from https://ollama.ai/ 2. **Setup:** - Clone the repository `git clone https://github.com/asiff00/On-Device-Speech-to-Speech-Conversational-AI.git` - Run `git lfs pull` to download the models and voices - Copy `.env.template` to `.env` - Add your HuggingFace token to `.env` - Tweak other parameters in `.env`, if needed [Optional] - Install requirements: `pip install -r requirements.txt` - Add any missing packages if not already installed `pip install <package_name>` 4. **Run Ollama:** - Start Ollama service - Run: `ollama run qwen2.5:0.5b-instruct-q8_0` or any other model of your choice 5. **Start Application:** - Run: `python speech_to_speech.py` - Wait for initialization (models loading) - Start talking when you see "Voice Chat Bot Ready" - Long press `Ctrl+C` to stop the application </details> We basically put a few models together to work in a multi-threaded architecture, where each component operates independently but is integrated through a queue management system to ensure performance and responsiveness. ## The flow works as follows: Loop (VAD -> Whisper -> LM -> TextChunker -> TTS) To achieve that we use: - **Voice Activity Detection**: Pyannote:pyannote/segmentation-3.0 - **Speech Recognition**: Whisper:whisper-tiny.en (OpenAI) - **Langua...
Source context: 0 downloads · 98 likes · Pipeline audio-to-audio · Repo asif00/Kokoro-Conversational