Files

59 lines
2.3 KiB
Markdown

Hardware:
CPU: Ryzen 7 7735HS\
GPU: RX 7600S (8GB VRAM)\
RAM: 16GB DDR5 4800 MTS (may be used for loading the model)\
SWAP: 8 GB (not used)\
Modell: https://huggingface.co/Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice
flash-attn wasn't used
Installation modification:
Uninstalled all cuda related deependencies and installed for all pytorch related packages the rocm version
Example - Cpu \
Generation time: 657.36s\
Audio duration: 88.64s\
Real-time factor: 7.42x\
text: cpu\
Speaker: Ryan\
Language: English\
Instruct: read the text in a calm and soothing voice, with a slight emphasis on key points, and maintain a steady pace throughout the narration.
Example - Raid\
Generation time: 2315.90s\
Audio duration: 655.28s\
Real-time factor: 3.53x\
text: cpu\
Speaker: Ryan\
Language: English\
Instruct: read the text in a calm and soothing voice, with a slight emphasis on key points, and maintain a steady pace throughout the narration.
Example - [Huggingface](https://huggingface.co/blog/security-incident-july-2026)\
Generation time: 1407.29s\
Audio duration: 390.40s\
Real-time factor: 3.60x\
text: huggingface\
Speaker: Aiden\
Language: English\
Instruct: read the text in a calm and soothing voice, with a slight emphasis on key points, and maintain a steady pace throughout the narration.
Example - [Qwen](https://dev.to/czmilo/qwen3-tts-the-complete-2026-guide-to-open-source-voice-cloning-and-ai-speech-generation-1in6)
Generation time: 386.22s
Audio duration: 52.88s
Real-time factor: 7.30x
Speaker: Aiden\
Language: qwen\
Instruct: read the text in a calm and soothing voice, with a slight emphasis on key points, and maintain a steady pace throughout the narration.
Example - [Offener Brief: Microsoft, Nvidia & Co. fordern offenes KI-Ökosystem](https://www.heise.de/news/Offener-Brief-Microsoft-Nvidia-Co-fordern-offenes-KI-Oekosystem-11378076.html)\
(offically not supported/unclear which speaker supports the language)\
Generation time: 903.47s\
Audio duration: 184.40s\
Real-time factor: 4.90x\
Speaker: Aiden\
Language: German\
Instruct: read the text in a calm and soothing voice, with a slight emphasis on key points, and maintain a steady pace throughout the narration.
VRAM Usage is most of the time below 4GB.
Notice the modell loading needs some time so I assume that the warmup time explain the fluctuating results between the real time factors