2026-07-26 01:11:44 +02:00
2026-07-26 01:11:44 +02:00
2026-07-26 01:11:44 +02:00
2026-07-26 01:11:44 +02:00
2026-07-26 01:11:44 +02:00
2026-07-26 01:11:44 +02:00
2026-07-26 01:11:44 +02:00
2026-07-26 01:11:44 +02:00
2026-07-26 14:39:45 +02:00
2026-07-26 14:39:45 +02:00
2026-07-26 14:39:45 +02:00
2026-07-26 14:39:45 +02:00
2026-07-26 01:11:44 +02:00
2026-07-26 14:41:13 +02:00
2026-07-26 14:39:45 +02:00

Hardware: CPU: Ryzen 7 7735HS
GPU: RX 7600S (8GB VRAM)
RAM: 16GB DDR5 4800 MTS (may be used for loading the model)
SWAP: 8 GB (not used)
Modell: https://huggingface.co/Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice

flash-attn wasn't used

Installation modification: Uninstalled all cuda related deependencies and installed for all pytorch related packages the rocm version

Example - Cpu
Generation time: 657.36s
Audio duration: 88.64s
Real-time factor: 7.42x
text: cpu
Speaker: Ryan
Language: English
Instruct: read the text in a calm and soothing voice, with a slight emphasis on key points, and maintain a steady pace throughout the narration.

Example - Raid
Generation time: 2315.90s
Audio duration: 655.28s
Real-time factor: 3.53x
text: cpu
Speaker: Ryan
Language: English
Instruct: read the text in a calm and soothing voice, with a slight emphasis on key points, and maintain a steady pace throughout the narration.

Example - Huggingface
Generation time: 1407.29s
Audio duration: 390.40s
Real-time factor: 3.60x
text: huggingface
Speaker: Aiden
Language: English
Instruct: read the text in a calm and soothing voice, with a slight emphasis on key points, and maintain a steady pace throughout the narration.

Example - Qwen Generation time: 386.22s Audio duration: 52.88s Real-time factor: 7.30x Speaker: Aiden
Language: qwen
Instruct: read the text in a calm and soothing voice, with a slight emphasis on key points, and maintain a steady pace throughout the narration.

Example - Offener Brief: Microsoft, Nvidia & Co. fordern offenes KI-Ökosystem
(offically not supported/unclear which speaker supports the language)
Generation time: 903.47s
Audio duration: 184.40s
Real-time factor: 4.90x
Speaker: Aiden
Language: German
Instruct: read the text in a calm and soothing voice, with a slight emphasis on key points, and maintain a steady pace throughout the narration.

VRAM Usage is most of the time below 4GB.

Notice the modell loading needs some time so I assume that the warmup time explain the fluctuating results between the real time factors

S
Description
No description provided
Readme
44 MiB
Languages
Python 100%