FlowWall

Text to speech node

Purpose: Convert text into playable speech audio.

01

Add Text to speech node.

Text to speech node parameters
  1. 01Voice model: Click to switch voice model
  2. 02Voice: Select character voice
  3. 03Listening: Play the current sound sample
  4. 04Text box: enter lines or narration
  5. 05Tags: Add mood, sound effects or pauses
02

Choose a sound.

03

Select the model and clickrun

Select a voice model
  1. 01Ark-Audio-3.5B: Suitable for delicate emotion and tone control
  2. 02Ark-Audio-2B: Suitable for natural multilingual speech
  3. 03MiniMax-Speech-2.8:Support system timbre and emotion control
04

Enter the text you want to read, such as lines or narration.

05

Add sound effect tags and set pauses as needed.

Open the mood and sound effects tag menu
MiniMax label usage example
Ark-Audio tag usage example
06

Adjust speed, pitch and volume as needed. The three sliders sit on the parameter bar at the top of the node, with ranges of 0.5x–2x, −12 to +12 semitones and 10%–200%.

Listen to the speech generation results
  1. DescriptionClick the play button above the node to listen to the generated results.

Three common ways to wire this node:

The reading content can be passed in from the upstream text node or LLM node connection. It is suitable for the model to write the narration first and then read it directly.

Connect an audio clip to the node and any speech model that supports voice cloning will use that clip as the voice reference; the node shows “Connected audio” and a preview button. Remove the connection and the voice picker comes back. Models that rely on system voices do not take a reference clip and keep using the selected voice.

The generated speech is audio output, and when connected to the audio track of the Video composition node, narration or lines can be added to the screen.