
By Florian Wenzel |
AI video generation has been moving at light speed. The visuals are stunning — but there is something missing.
Sound.
Most AI video providers generate videos without sound. No matter how cinematic, a visual without sound feels unfinished. That changes today. Meet Mirelo SFX, the first sound effects model designed specifically for AI-generated video.
Available now on fal.ai.
Sound generated by Mirelo SFX.
Before, adding sounds to AI videos was a tedious process. Now, Mirelo SFX lets you easily create perfectly synced cinematic sound with one click.
Typically, AI-generated videos don't come with sound. Mirelo SFX unmutes them with one click.
By integrating Mirelo SFX through fal.ai, video-gen platforms can now deliver complete experiences by default:
Real-time generation — 10 seconds of video = 10 seconds of compute.
Multiple variations — 2-8 outputs per clip so users can pick their favorite.
The uncanny valley is finally crossed. In listening tests, users preferred Mirelo SFX's output 67-77% of the time. This isn't a nice-to-have anymore — it's a model that can be chained by default, knowing it will enhance every video, not degrade it.
In blind pairwise comparison tests, independent raters* compared generations from Mirelo SFX and three of the most popular existing video-to-sfx alternatives (PixVerse, MMAudio V2 and ThinkSound). In average, Mirelo SFX achieved a total win rate of 67.4% (excluding ties) and 77% (including ties), clearly outperforming all three alternatives.

Competitors often struggle with unwanted artifacts like music or speech, particularly on synthetic videos (see, e.g., examples 1 and 2 below). Mirelo SFX also excels at creating realistic effects across a multitude of scenes when alternative models default to unrelated or unnatural sounds (see examples 2 and 3 below).
* We recruited 20 external listeners from a crowd-sourcing platform via mabyduck. The raters were pre-screened by mabyduck to ensure they listen closely to the given audio. Raters performed pairwise comparisons between different audio generations for a given video, with the option to vote for a "tie". 740 pairwise comparisons were collected in total.