Audio generation node
Generate an entire audio piece from a description. The soundtrack, multi-character dialogue, and environmental sound effects can be generated separately, or they can be written in the same description and generated at once, and the output is a mixed audio track. The difference from the Text to speech node is that Text to speech reads a piece of text, while this node performs the entire sound scene.
| What do you want | How to write the prompt | Typical scenario |
|---|---|---|
| Just the human voice | Write down the character, lines and tone, and indicate that no music is required | Narration, audiobooks, interviews, ASMR |
| Just sound effects or ambient sounds | Write around the source of the sound, its distance, its rhythm, and when it comes in and when it exits, not the characters. | The sound of rain, streets, machinery, footsteps and doors opening |
| Just the soundtrack | Write about style, instrumentation, tempo and mood, and indicate that no vocals are required | The opening music and the atmosphere pave the way |
| Three things together | Write the character and timbre first, then the background sounds, and finally how to perform each sentence in order of occurrence. | Film and television clips, two-person podcast, live broadcast, news scene |
Add Audio generation node. The currently used model is marked in the upper right corner of the node. There is currently only one, so switching is not provided.
- 1The currently used model is marked in the upper right corner of the node. There is currently only one, so switching is not provided.
- 2Reference timbre: Click the plus sign to select from the timbre library, up to 3 items; the added ones are displayed in a row, stop and listen, and remove the cross in the upper right corner.
- 3Tips: Write characters, lines, tone, and background sounds; enter @ to quote the reference sounds above. The lower right corner is the real-time word count, with an upper limit of 3,000.
- 4Command and Polish: Insert timestamps or timbre references; the Magic Wand turns creative briefs into performable scripts.
- 5Run: Stop the mouse first, and it will display the approximate number of points spent this time and the maximum number of seconds.
When you need to specify who will speak, prepare a reference tone first. Click above the prompt word box + Open the sound library and select one, up to 3; stop the mouse + It will show how many items have been added, and the button will turn gray when it is full. The sounds that have been added are displayed as a row of avatars. You can stop and listen to them. The cross in the upper right corner can be removed. If you want to use your own voice as a reference, just connect the audio node to the audio input on the left.
- 01Added reference sounds: displayed as avatars, and the cross in the upper right corner can be removed.
- 02Plus sign: Open the sound library, which uses the same list as the Text to speech node.
- 03Search: Filter by name or language, for example, only look at English sounds.
- 04Tone entry: Click to add reference timbre. The selected one will be checked. Click again to cancel.
Write prompt words. When there is a lot of content, it is most stable to write it in three paragraphs: "Characters and timbres → Background sounds → Performances in the order of occurrence"; if you only need one of the sounds, write the unnecessary part directly, such as "no music". input @ You can quote the reference timbre above to specify which voice should be used to pronounce a certain line. There is a real-time word count in the lower right corner, and the text coming in from the upstream is also included in the 3,000 words.
When you need to click on a time point, click below the prompt word boxCommandButton to insert timestamp: after filling in the start and end seconds, a form like this will be inserted at the cursor [2.0s:5.0s] The mark is used to align the placement of lines and sound effects. Audio quotes can also be inserted in the same menu.
- 01Add timestamp: fill in the start and end seconds, and the end time must be later than the start time.
- 02Audio reference: Insert a reference sound into the prompt word, which is equivalent to manually entering @.
When the prompt word is not written smoothly, click the magic wand button next to itPolish prompt, the model will make it more complete and directly replace the main text - the creative brief will be made up into an actable script, and the already written lines will only be made up for the performance instructions. Stopping the mouse on the button will first display the approximate number of points it will cost this time. After polishing, the button becomes undoable, and you can click it again to restore it to the original.
- 01Polish prompt: Click once to directly rewrite the text without popping up another window to confirm.
- 02Hover prompt: The estimated points for this polishing are given first. The evaluation itself does not cost points.
Adjust speech rate, pitch, and volume as desired, the defaults are 1x, 0, and 100% respectively.
Clickrun. The maximum length of a single audio piece is 120 seconds, and you can listen to it directly on the node after it is generated.
- DescriptionThe run button will first press the estimated points for up to 120 seconds, and after the film is released, the calculation will be based on the actual duration and the difference will be refunded.
Three common ways to wire this node:
Connect the text node or LLM node upstream, and the connected text will be combined with the node's own prompt words. It is suitable for the model to write the script first, and then directly give it to it to act.
Connect the upstream audio node or the sound output port of the main node. Every time a connected sound is connected, it will automatically become a reference sound, which also occupies 3 places.
The audio track most commonly connected to the Video composition node downstream; if connected to the Subtitles node, this audio can be directly identified as a Subtitles with a timeline.