59 lines
2.3 KiB
Markdown
59 lines
2.3 KiB
Markdown
Hardware:
|
|
CPU: Ryzen 7 7735HS\
|
|
GPU: RX 7600S (8GB VRAM)\
|
|
RAM: 16GB DDR5 4800 MTS (may be used for loading the model)\
|
|
SWAP: 8 GB (not used)\
|
|
Modell: https://huggingface.co/Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice
|
|
|
|
flash-attn wasn't used
|
|
|
|
Installation modification:
|
|
Uninstalled all cuda related deependencies and installed for all pytorch related packages the rocm version
|
|
|
|
Example - Cpu \
|
|
Generation time: 657.36s\
|
|
Audio duration: 88.64s\
|
|
Real-time factor: 7.42x\
|
|
text: cpu\
|
|
Speaker: Ryan\
|
|
Language: English\
|
|
Instruct: read the text in a calm and soothing voice, with a slight emphasis on key points, and maintain a steady pace throughout the narration.
|
|
|
|
Example - Raid\
|
|
Generation time: 2315.90s\
|
|
Audio duration: 655.28s\
|
|
Real-time factor: 3.53x\
|
|
text: cpu\
|
|
Speaker: Ryan\
|
|
Language: English\
|
|
Instruct: read the text in a calm and soothing voice, with a slight emphasis on key points, and maintain a steady pace throughout the narration.
|
|
|
|
Example - [Huggingface](https://huggingface.co/blog/security-incident-july-2026)\
|
|
Generation time: 1407.29s\
|
|
Audio duration: 390.40s\
|
|
Real-time factor: 3.60x\
|
|
text: huggingface\
|
|
Speaker: Aiden\
|
|
Language: English\
|
|
Instruct: read the text in a calm and soothing voice, with a slight emphasis on key points, and maintain a steady pace throughout the narration.
|
|
|
|
Example - [Qwen](https://dev.to/czmilo/qwen3-tts-the-complete-2026-guide-to-open-source-voice-cloning-and-ai-speech-generation-1in6)
|
|
Generation time: 386.22s
|
|
Audio duration: 52.88s
|
|
Real-time factor: 7.30x
|
|
Speaker: Aiden\
|
|
Language: qwen\
|
|
Instruct: read the text in a calm and soothing voice, with a slight emphasis on key points, and maintain a steady pace throughout the narration.
|
|
|
|
Example - [Offener Brief: Microsoft, Nvidia & Co. fordern offenes KI-Ökosystem](https://www.heise.de/news/Offener-Brief-Microsoft-Nvidia-Co-fordern-offenes-KI-Oekosystem-11378076.html)\
|
|
(offically not supported/unclear which speaker supports the language)\
|
|
Generation time: 903.47s\
|
|
Audio duration: 184.40s\
|
|
Real-time factor: 4.90x\
|
|
Speaker: Aiden\
|
|
Language: German\
|
|
Instruct: read the text in a calm and soothing voice, with a slight emphasis on key points, and maintain a steady pace throughout the narration.
|
|
|
|
VRAM Usage is most of the time below 4GB.
|
|
|
|
Notice the modell loading needs some time so I assume that the warmup time explain the fluctuating results between the real time factors |