Audio guide

Make speech, music and sound

Create speech, music and sound assets while keeping their recipes and timings.

Treat every line, bed and effect as its own audio asset. This keeps the prompt, voice or settings attached to the media and makes timing decisions inspectable later.

Choose the audio job

Read list_models for the current audio entries and their settings.

  • Speech turns text into a voiced asset. Pick a voice exposed for the paying account.
  • Music creates a track from musical direction and duration.
  • Sound effects create a short sound from a concrete description of source, action, space and perspective.

Quote the configured request with estimate_credits. Prices may depend on duration or use a flat track price; the live catalog carries the exact calculation.

Write a useful prompt

For speech, write the words first, then use the available controls for delivery. Split long copy where the edit needs independent timing or emphasis.

For music, name structure and movement: opening texture, pulse, density, change and ending. For sound, describe what produces it, where the listener is and how the sound evolves. Avoid mood-only prompts that leave the physical event unspecified.

Inspect timing

Generated speech can carry word timings. Use get_audio_word_timings to read the current timing record. Use align_audio when ready audio needs a new or corrected alignment; it changes timing data, not the audio bytes. Alignment is a priced action. Uploaded audio and audio not generated as speech must supply the exact text; generated speech may use its recorded prompt.

Upload existing audio with upload_asset and declare its origin when known. A lip-sync video can use a ready image and audio asset in the reference slots named by its catalog entry.

Name assets for their role, star the approved results and keep rejected results as ordinary unstarred assets. The browser timeline is currently unavailable, so export or download the ready audio for the final cut.