Text to speech node
Purpose: Convert text into playable speech audio.
Add Text to speech node.
- 01Voice model: Click to switch voice model
- 02Voice: Select character voice
- 03Listening: Play the current sound sample
- 04Text box: enter lines or narration
- 05Tags: Add mood, sound effects or pauses
Choose a sound.
Select the model and clickrun。
- 01
Ark-Audio-3.5B: Suitable for delicate emotion and tone control - 02
Ark-Audio-2B: Suitable for natural multilingual speech - 03
MiniMax-Speech-2.8:Support system timbre and emotion control
Enter the text you want to read, such as lines or narration.
Add sound effect tags and set pauses as needed.
Adjust speed, pitch and volume as needed. The three sliders sit on the parameter bar at the top of the node, with ranges of 0.5x–2x, −12 to +12 semitones and 10%–200%.
- DescriptionClick the play button above the node to listen to the generated results.
Three common ways to wire this node:
The reading content can be passed in from the upstream text node or LLM node connection. It is suitable for the model to write the narration first and then read it directly.
Connect an audio clip to the node and any speech model that supports voice cloning will use that clip as the voice reference; the node shows “Connected audio” and a preview button. Remove the connection and the voice picker comes back. Models that rely on system voices do not take a reference clip and keep using the selected voice.
The generated speech is audio output, and when connected to the audio track of the Video composition node, narration or lines can be added to the screen.