By Mirelo Team |
Add sound effects to a video without replacing its dialogue in three steps:
You get a finished mix and a separate sound effects track. When speech is detected, you also get an isolated dialogue track that you can mix yourself. Speech separation can change the audio signal, so keep your source recording if you need the original track unchanged.
Read the Preserve speech API guide or explore Mirelo SFX.
A video with good dialogue and weak sound still feels unfinished. This often happens with AI-generated video: the voice works, but footsteps, doors, traffic, and room sound are thin, generic, or missing. Replacing the entire soundtrack would also discard the voice.
Speech preservation lets you keep the voice from the source clip while generating new sound for the scene. It works with supported clips from AI video generators, avatar tools, ads, interviews, tutorials, and vlogs. Developers can enable it in the API.
Mirelo SFX generates sound effects from video. It uses the footage to place sounds on visible actions, such as a footstep, a door closing, or a splash. You do not need to describe and place every sound on a timeline; the video provides the context.
Speech preservation brings the same video-driven generation to clips that already have dialogue.
The dialogue is taken from your clip, not synthesized or rerecorded. Isolation and mixing may affect its sound, so the returned track is not a bit-for-bit copy of the original audio.
AI video creators. Add sound to a generated clip whose voice already works without generating that voice again.
AI video platforms. Offer an audio step after video generation. Send a finished clip to the Mirelo API and return the mix or the separate tracks to your users.
Avatar, ad, and explainer platforms. Add product sounds, room tone, and ambience around spoken content.
Editors, agencies, and studios. Build sound design around a recorded interview or voiceover. Keep the source recording when an exact, approved vocal track is required.
Upload the clip to Mirelo and generate effects from the picture. If the clip has speech you want to keep, turn on Preserve speech in the API. To try it without writing code, use the API playground in Mirelo Studio.
Turn on Preserve speech. Mirelo isolates the voice already in the soundtrack, generates effects that match the video, and returns a mix plus separate tracks when speech is detected.
Yes. Mirelo does not regenerate the voice. It separates detected speech from the source soundtrack and generates new effects around it. Keep the original source if exact audio fidelity matters.
When speech is detected, Mirelo returns an isolated dialogue track. Separation may leave artifacts or change the sound of the recording.
Yes. When speech is detected, you receive separate dialogue and sound effects tracks alongside the ready-made mix. You can adjust them in your video editor.
Mirelo does not generate new speech. The returned voice comes from the source soundtrack, but isolating it can change the audio signal.
Mirelo ducks the generated effects while speech is present and balances the layers in its finished mix. Review the result before publishing, especially when the source dialogue is quiet or noisy.
Yes. The video guides generation, and you can add a short prompt to steer the result or request several versions.
You still receive sound effects and a mix. If no speech is detected, there is no dialogue track and you are charged the base sound effects price. For a video without an audio track, send the soundtrack as a separate audio input to use Preserve speech. Otherwise, use standard video-to-sound effects.
With Preserve speech on SFX 1.6, each request covers up to 60 seconds of video. For a longer video, send one request per section and set start_offset_ms to where each section starts. Limits can change, so read max_with_preserve_speech from GET /v3/models before sending a job. The Preserve speech guide explains the limit.
Add Mirelo as an audio step after your video model. Send each finished clip to the API with Preserve speech enabled when it has a soundtrack you want to retain. Read the integration guide.
Mirelo uses the finished video as input rather than requiring a particular video model. The file still needs to meet the API's supported input and duration limits.
Yes. The result includes a sound effects track and a mix. If speech is detected, the isolated dialogue is available separately, so users can adjust the balance or generate another effects take.
Sound effects are priced per second of generated audio, per variant. When speech is found, speech separation is charged once per second of the requested duration. Preflight gives the maximum credit cost and estimated processing time before a job starts. See usage rates.
Yes. Jobs run in the background and can run in parallel within the limits listed in the API docs.
It means keeping speech from an existing soundtrack while creating new sound around it. Mirelo isolates detected speech rather than generating a replacement voice.
Text-to-SFX creates a sound from a written description that you place in the edit. Video-to-SFX uses the footage to generate sound timed to visible action. Mirelo offers both.
See the Preserve speech guide in the Mirelo API documentation.