Why I built Revori
In short
Revori turns one song upload into beat-synced lyric videos, visualizers, cover art and Spotify Canvas loops. It exists because professional editors know nothing about beats or lyrics, consumer editors sync to the waveform rather than to words, and the existing lyric-video SaaS category mostly ships weak transcription and weak presets. The AI does the mechanical work of transcription and beat detection; the taste decisions stay with the person. Editing, previewing and saving are always free, and credits are spent only on analysis and final renders.
The first lyric video I made took four hours.
It was a Sunday. I had a song I liked and footage I liked and wanted to put them together. I spent four hours in Premiere nudging keyframes a frame at a time, exported, watched it back, and the second verse had already drifted. Start over.
There was no profound lesson in it. I was annoyed enough to build the thing instead.
What the existing tools each get wrong
Three categories, three different failures.
Professional editors. Premiere, DaVinci Resolve, Final Cut. Total control and zero help. They do not know what a beat is. They do not know what a lyric is. Every sync is manual, and manual sync across a three-minute song is several hundred decisions you have to get individually right and then keep right through every re-edit.
Consumer editors. CapCut, InShot, the native TikTok editor. Fast and genuinely good-looking, and they do have beat features. But those features read the audio waveform, not the vocal. They will cut your clip on the drum hits. They cannot put the word "heartbreak" on the syllable where it is actually sung, because nothing in the pipeline ever transcribed it.
Lyric-video SaaS. The category that shows up in YouTube pre-roll. In principle these understand both beats and lyrics. In practice the transcription is where they fall down, and transcription is the part everything else depends on: a lyric video built on a bad transcript is unusable no matter how good the presets are.
None of these is wrong for its own purpose. None of them was the thing I wanted.
The thing that made it buildable
Two capabilities got good enough at roughly the same time.
Speech recognition now returns word-level timestamps reliable enough to build on, particularly when you run several models and reconcile them rather than trusting one. Beat detection got a genuine step change with neural detectors: we run beat_this, from ISMIR 2024, with a multi-band librosa detector as the fallback path.
So the mechanical half of the job, the half that took me four hours, is now machine work. What is left is the half that was always the point: which typeface, which footage, which colour, where the video should hold still.
That split is the whole design. The AI does the part with a correct answer. The person does the part that does not have one.
What actually happens when you upload
Upload a song. The vocal gets separated from the mix, the beat grid gets detected, and the lyrics get transcribed with a timestamp on every word. You review the transcript before you build anything, because no transcription is perfect and the review step is cheaper than discovering an error at export.
Then you pick a preset, pick clips from a library of 1,100 or more royalty-free videos or upload your own, and export a 1080p MP4. Default aspect is 9:16 because the destination is usually TikTok or Reels, with 1:1 and 16:9 available.
Analysis runs once per song. Reuse that same song for a lyric video, a visualizer, cover art and a Spotify Canvas without paying for the analysis again.
What it costs, plainly
Editing, previewing and saving are all free, with unlimited projects on every tier including the free one.
Credits are spent on two things only: 1 credit per second of audio analysed, and 50 credits for a standard 1080p render. The free tier starts with 500 credits and does not ask for a card.
Free-tier exports carry a small "Made with Revori" mark in the bottom corner. Paid exports do not. I would rather say that here than have you find it after you have built something.
What Revori is not
Revori is not a platform, a social network or an ecosystem. It does not collaborate with your team, match you with session musicians, or distribute your track to 400 streaming services. Each of those is a different product, and trying to be all of them at once is the standard route to a tool that does nothing well.
It takes your audio and gives back something you would actually post.
Questions people ask about this
- What is Revori?
- Revori is a browser-based tool that turns a single song upload into beat-synced lyric videos, audio visualizers, cover art, Spotify Canvas loops and lyric cards. It transcribes the lyrics with word-level timing and detects the beat grid automatically, then the user makes the visual decisions and exports a 1080p MP4 sized for TikTok, Reels, YouTube or Spotify.
- How is Revori different from CapCut or Premiere for lyric videos?
- Premiere and DaVinci Resolve give full control and no assistance: they have no concept of a beat or a lyric, so every word is placed by hand. CapCut and similar consumer editors can cut to the audio waveform but cannot place individual words, because they do not transcribe the vocal. Revori aligns each word to the vocal performance and marks the beat grid separately, so the words and the cuts are handled by different systems that each know what they are looking at.
- Is Revori free to use?
- Editing, previewing and saving are always free and unlimited on every plan, including the free tier, which starts with 500 credits and no card. Credits are only spent on AI work: 1 credit per second of audio analysis, and 50 credits for a standard 1080p render. Free-tier exports carry a small "Made with Revori" mark in the corner; paid tiers export without it.