Hardware: CPU: Ryzen 7 7735HS\ GPU: RX 7600S (8GB VRAM)\ RAM: 16GB DDR5 4800 MTS (may be used for loading the model)\ SWAP: 8 GB (not used)\ Modell: https://huggingface.co/Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice flash-attn wasn't used Installation modification: Uninstalled all cuda related deependencies and installed for all pytorch related packages the rocm version Example - Cpu \ Generation time: 657.36s\ Audio duration: 88.64s\ Real-time factor: 7.42x\ text: cpu\ Speaker: Ryan\ Language: English\ Instruct: read the text in a calm and soothing voice, with a slight emphasis on key points, and maintain a steady pace throughout the narration. Example - Raid\ Generation time: 2315.90s\ Audio duration: 655.28s\ Real-time factor: 3.53x\ text: cpu\ Speaker: Ryan\ Language: English\ Instruct: read the text in a calm and soothing voice, with a slight emphasis on key points, and maintain a steady pace throughout the narration. Example - [Huggingface](https://huggingface.co/blog/security-incident-july-2026)\ Generation time: 1407.29s\ Audio duration: 390.40s\ Real-time factor: 3.60x\ text: huggingface\ Speaker: Aiden\ Language: English\ Instruct: read the text in a calm and soothing voice, with a slight emphasis on key points, and maintain a steady pace throughout the narration. Example - [Qwen](https://dev.to/czmilo/qwen3-tts-the-complete-2026-guide-to-open-source-voice-cloning-and-ai-speech-generation-1in6) Generation time: 386.22s Audio duration: 52.88s Real-time factor: 7.30x Speaker: Aiden\ Language: qwen\ Instruct: read the text in a calm and soothing voice, with a slight emphasis on key points, and maintain a steady pace throughout the narration. Example - [Offener Brief: Microsoft, Nvidia & Co. fordern offenes KI-Ökosystem](https://www.heise.de/news/Offener-Brief-Microsoft-Nvidia-Co-fordern-offenes-KI-Oekosystem-11378076.html)\ (offically not supported/unclear which speaker supports the language)\ Generation time: 903.47s\ Audio duration: 184.40s\ Real-time factor: 4.90x\ Speaker: Aiden\ Language: German\ Instruct: read the text in a calm and soothing voice, with a slight emphasis on key points, and maintain a steady pace throughout the narration. VRAM Usage is most of the time below 4GB. Notice the modell loading needs some time so I assume that the warmup time explain the fluctuating results between the real time factors