<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>Notes from building Revori</title>
  <subtitle>Field notes on beat detection, lyric timing, transcription and the craft of beat-synced video.</subtitle>
  <link rel="alternate" type="text/html" href="https://www.revori.app/blog" />
  <link rel="self" type="application/atom+xml" href="https://www.revori.app/blog/feed.xml" />
  <id>https://www.revori.app/blog</id>
  <updated>2026-09-05T00:00:00Z</updated>
  <author><name>Revori</name></author>
  <entry>
    <title>What an audio visualizer is for, and which style to pick</title>
    <link rel="alternate" type="text/html" href="https://www.revori.app/blog/audio-visualizer-for-tiktok" />
    <id>https://www.revori.app/blog/audio-visualizer-for-tiktok</id>
    <published>2026-09-05T00:00:00Z</published>
    <updated>2026-09-05T00:00:00Z</updated>
    <category term="Craft" />
    <summary>The format for a track with nothing to show: no footage, no lyrics worth reading, or an instrumental. What makes one read on a phone, and the five styles.</summary>
    <content type="html"><![CDATA[<article><section class=" py-20 md:py-24 lg:py-32  px-6 lg:px-10 "><div class="max-w-[40rem] mx-auto"><header class="mb-14"><div class="font-mono text-xs uppercase tracking-widest text-text-muted flex flex-wrap items-center gap-x-3 gap-y-1 mb-6"><span>Craft</span><span aria-hidden="true">·</span><time dateTime="2026-09-05">September 5, 2026</time><span aria-hidden="true">·</span><span>4 min<!-- --> read</span></div><h1 class="font-display text-[2.5rem] md:text-[4.5rem] font-normal text-white leading-[0.98] tracking-tight">What an audio visualizer is for, and which style to pick</h1></header><hr class="border-t border-white/[0.06] mb-14"/><div data-answer-summary="true" class="mb-14 rounded-2xl border border-white/[0.07] bg-surface-0 p-6 md:p-8"><p class="eyebrow mb-4">In short</p><p class="text-white text-lg md:text-xl leading-relaxed max-w-[42rem]">An audio visualizer is the format for a track with nothing to show: no footage, an instrumental, or a passage where the lyrics are not the point. It gives the listener motion that follows the music, so the post is not a static cover with a play button. It reads on a phone when the motion is driven by the audio rather than looped, when it sits on a backdrop that means something (the cover, a photo, one colour), and when the style matches the energy of the track. Revori has five styles, Circles, Orbit, Nebula Drift, Bars and Waves, on any backdrop, exported 9:16, 1:1 or 16:9 at 1080p for 40 credits, with no transcription involved.</p></div><div class="prose-editorial"><p>Every release has a section with nothing to show. An intro with no vocal, a beat you made before the vocal exists, a drop, a producer tag, an instrumental B-side. Posting a static cover over it is what most people do, and the feed treats a still image with audio exactly the way it looks: as something to scroll past.</p>
<p>A visualizer is the format for that section. It is not a lesser lyric video. It is the right answer to a different question.</p>
<h2 id="when-a-visualizer-is-the-right-format">When a visualizer is the right format</h2>
<p>Three situations, and they cover nearly every case.</p>
<p><strong>The track has no words in the section you are posting.</strong> Instrumentals, beats, the sixteen bars before the vocal comes in. A lyric video of silence is a blank frame.</p>
<p><strong>The words are not the point.</strong> A vocal chop, an ad-lib run, a heavily processed hook where the texture is the hook. Putting readable text over it draws attention to the one thing you did not want examined.</p>
<p><strong>You have no footage and no time to find any.</strong> A lyric video over the wrong clip is worse than a visualizer over the right colour. The footage question can be skipped entirely here, because the motion comes from the sound.</p>
<p>If none of those apply and the section has a line people will sing back, make the <a href="https://www.revori.app/blog/how-to-make-a-lyric-video">lyric video</a> instead. The two formats often end up posted for the same track, cut from different sections.</p>
<h2 id="what-makes-one-read-on-a-phone">What makes one read on a phone</h2>
<p>Four things separate a visualizer that holds attention from one that reads as a screensaver.</p>
<p><strong>The motion is driven by the audio, not looped.</strong> A loop of bars bouncing to nothing is recognisable within two seconds. A viewer cannot say what is wrong, only that the picture and the sound are not the same thing. Every one of Revori&#x27;s five styles reads the audio frame by frame and draws that, so a kick moves the picture when the kick happens and a breakdown lets it settle.</p>
<p><strong>The backdrop means something.</strong> The visualizer is a layer, and what it sits on is most of the frame. The cover art is the usual choice and a good one: the listener sees the image they will later find in the app. A photo of you, a single flat colour from the artwork, a slow gradient. What does not work is a stock image chosen because it was there; the motion draws the eye and then the eye lands on nothing.</p>
<p><strong>One style, chosen for the energy.</strong> A dense, bright style over a slow track is exhausting; a drifting particle field under a fast one looks asleep. Pick the style for the energy of the section, then leave it. Switching styles inside one post reads as indecision.</p>
<p><strong>Room to breathe at phone size.</strong> A visualizer is watched in peripheral vision most of the time. Motion that fills the frame edge to edge on a desktop monitor becomes noise on a phone, and the outer band is under the platform&#x27;s own overlays anyway. Preview at phone size before exporting, the same rule as for <a href="https://www.revori.app/blog/anatomy-of-a-lyric-video">any lyric video</a>.</p>
<h2 id="the-five-styles">The five styles</h2>
<p>Each one draws the audio differently, which is the whole basis for choosing.</p>
<p><strong>Circles.</strong> A ring of bars around a centre point, each bar driven by its own slice of the spectrum, the whole ring breathing with the bass. The centre is reserved for an image, usually the cover. Suits anything with a clear pulse, and it is the default because it frames artwork.</p>
<p><strong>Orbit.</strong> A single orb that breathes with the low end and sends a ring outward on every beat of the tempo grid. It reads the kick more than anything else, so it suits tracks whose pulse is the point: minimal, techno, a beat where the drum is the hook.</p>
<p><strong>Nebula Drift.</strong> A few hundred particles on a slow orbit, each tied to a band of the spectrum; the bass pulls the whole field inward and the highs make parts of it sparkle. Reads as texture rather than as a meter, so it suits tracks where you do not want a literal readout of the beat: downtempo, vocal-led passages, anything atmospheric.</p>
<p><strong>Bars.</strong> The classic equaliser, weighted toward the low and mid bands so it moves musically rather than flickering. The most literal readout, and the right choice for high-energy sections where the listener wants to see the drop coming. Also the most legible at small sizes.</p>
<p><strong>Waves.</strong> A mirrored wave across the centre of the frame, drawn from the mid and high bands rather than the kick, with a stroke that thickens on energy. It moves with whatever is out front, a vocal chop, a guitar line, a lead synth, and it is the one that looks most like the sound drawn.</p>
<p>If you cannot decide, Circles with the cover in the centre is the safe answer for a release and Bars is the safe answer for a drop.</p>
<h2 id="what-it-costs-and-what-comes-out">What it costs and what comes out</h2>
<p>A visualizer render is 40 credits, less than a lyric video render, because there is no lyric layer to draw. It needs the song upload but not the transcription: the visualizer flow skips lyrics entirely, so there is no transcript to correct and nothing to sync. The same upload is reused if you later make a lyric video or a Canvas from the track.</p>
<p>Output is 1080p MP4 in 9:16, 1:1 or 16:9, each a separate render because the motion is composed for the frame. Previewing and changing the style or backdrop is free and unlimited, so the right workflow is to try three styles against the section and export the one that survives the phone check.</p>
<aside style="margin-top:3rem;padding:1.5rem;border:1px solid rgba(255,255,255,0.08);border-radius:12px;background:rgba(43,146,245,0.06);font-size:0.875rem;opacity:0.8"><p>Five styles, your backdrop, no transcription. Free to preview. <a href="https://www.revori.app" style="color:#9CC7FF">revori.app</a></p></aside></div><section class="mt-20" aria-labelledby="post-faq-heading"><p id="post-faq-heading" class="eyebrow mb-6">Questions people ask about this</p><hr class="border-t border-white/[0.06] "/><dl class="mt-2"><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">Do I need lyrics to make an audio visualizer?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">No. A visualizer reacts to the audio itself, so it works for instrumentals, beats, ambient tracks and anything where the words are not the point. In Revori the visualizer flow skips transcription entirely; you upload the track, pick a style and a backdrop, and export. If the hook is a line people will sing back, a lyric video is the better format for that section.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">What size should an audio visualizer be for TikTok?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">9:16 vertical at 1080x1920 for TikTok, Instagram Reels and YouTube Shorts. Use 1:1 for an Instagram feed post and 16:9 for YouTube. Render each shape separately rather than cropping one into another, because the motion is composed for the frame it was rendered in and a crop cuts the outer bars or the orbit off the edge.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">Is a visualizer better than a lyric video?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Neither is better; they answer different situations. A lyric video carries a hook the listener will read and sing, so it wins when the words are the reason to stop scrolling. A visualizer wins when the sound is the reason: an instrumental, a producer tag, a drop, a passage where reading would get in the way. Many artists post both for the same track, cut from different sections.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">Can I put my cover art behind the visualizer?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Yes. The backdrop is yours: a photo such as the cover, a video clip, a single colour or a gradient. The visualizer draws over it. A cover as the backdrop is the most common choice for a release post because the listener sees the art they will later recognise in the app.</dd></div></dl></section></div></section></article>]]></content>
  </entry>
  <entry>
    <title>Why the words in a lyric video look early or late</title>
    <link rel="alternate" type="text/html" href="https://www.revori.app/blog/lyrics-out-of-sync" />
    <id>https://www.revori.app/blog/lyrics-out-of-sync</id>
    <published>2026-09-05T00:00:00Z</published>
    <updated>2026-09-05T00:00:00Z</updated>
    <category term="Engineering" />
    <summary>Four separate causes, from a wrong transcript to the 50 ms lead every highlight needs, and the fix for each. Most complaints are one of two of them.</summary>
    <content type="html"><![CDATA[<article><section class=" py-20 md:py-24 lg:py-32  px-6 lg:px-10 "><div class="max-w-[40rem] mx-auto"><header class="mb-14"><div class="font-mono text-xs uppercase tracking-widest text-text-muted flex flex-wrap items-center gap-x-3 gap-y-1 mb-6"><span>Engineering</span><span aria-hidden="true">·</span><time dateTime="2026-09-05">September 5, 2026</time><span aria-hidden="true">·</span><span>8 min<!-- --> read</span></div><h1 class="font-display text-[2.5rem] md:text-[4.5rem] font-normal text-white leading-[0.98] tracking-tight">Why the words in a lyric video look early or late</h1></header><hr class="border-t border-white/[0.06] mb-14"/><div data-answer-summary="true" class="mb-14 rounded-2xl border border-white/[0.07] bg-surface-0 p-6 md:p-8"><p class="eyebrow mb-4">In short</p><p class="text-white text-lg md:text-xl leading-relaxed max-w-[42rem]">When the words in a lyric video look off, it is one of four things, and they need different fixes. A wrong transcript: a word that was never sung has no true time, so fix the text first. An alignment miss on one word or one phrase, usually a held note or a quiet passage, fixed by nudging or tapping that word. A whole track that is consistently early or late, fixed with one global offset. And the perceptual effect where a highlight fired exactly on the vocal onset still reads as late, which the renderer handles by firing every highlight 50 ms early. Tell them apart by checking whether the drift is one word, one section, or everything.</p></div><div class="prose-editorial"><p>&quot;The lyrics are out of sync&quot; is one sentence that describes four different problems. They look similar in a preview and they have nothing to do with each other underneath, which is why the first fix people reach for, dragging everything a bit earlier, usually makes one of them better and another worse.</p>
<p>Sort out which one you have before touching anything.</p>
<h2 id="which-of-the-four-is-it">Which of the four is it</h2>
<p>Play the section and watch for the shape of the error.</p>
<p><strong>One word is wrong, the rest are fine.</strong> Alignment miss. Skip to the second section.</p>
<p><strong>A whole phrase or a whole section drifts, then it recovers.</strong> Usually a transcript problem upstream of alignment, sometimes a passage the aligner could not hear. First and second sections.</p>
<p><strong>Everything is off by the same amount, from the first line to the last.</strong> Global offset. Third section.</p>
<p><strong>Everything is technically on time and it still reads late.</strong> The perceptual one. Fourth section, and it is already handled for you.</p>
<h2 id="the-transcript-is-wrong">The transcript is wrong</h2>
<p>A word that was never sung has no correct time. If the transcript says &quot;fire&quot; where the singer sang &quot;higher&quot;, the aligner will find the best place for &quot;fire&quot; and it will be near the right spot, but the highlight will land on the wrong syllable count and the next few words shuffle to make room. It looks like a timing error. It is a text error.</p>
<p>This is why the <a href="https://www.revori.app/blog/where-ai-lyric-transcription-fails">transcription accuracy post</a> says to read the transcript against the song before doing anything else. On real uploads the text is the bottleneck, not the alignment.</p>
<p>The fix is in the text, then a realignment. In Lyric Studio, correct the words, or paste your own lyric sheet if the transcript is wrong in more than a few places. Then run Realign: the text stays as you wrote it, the timings are replaced with a fresh alignment against the audio, and undo restores the old timings if it made something worse. Realigning after a batch of edits is cheaper than nudging twenty words by hand, and more accurate, because the aligner sees the corrected text as a whole.</p>
<h2 id="one-word-or-one-phrase-is-off">One word or one phrase is off</h2>
<p>The transcript is right and one word still lands wrong. This is the alignment&#x27;s own error, and it has a shape.</p>
<p>Alignment pins each word to its acoustic onset by matching the text against the separated vocal. It is good at consonant attacks in a clear passage and worse at three things: a held vowel, where the &quot;onset&quot; is a slow swell with no edge to find; a quiet or sparse passage, where the vocal barely rises above the stem&#x27;s noise floor; and stacked harmonies or heavy processing, where several voices start the word at slightly different times.</p>
<p>Since September 2026 the onsets come from a CTC aligner with a second opinion. Other timing estimators vote on each word, and where two of them agree with each other and disagree with the aligner by more than 200 ms, the word takes the vote. Where the estimators disagree with each other, the word is marked with a lower timing confidence and Lyric Studio shows a Timing chip on it. That chip is the list of words to listen to first.</p>
<p>On the JamendoLyrics benchmark, given the correct text, the median word is 47 ms from its true onset and 95 percent of words are within 300 ms. The remaining 5 percent are the ones you will notice, and on a three minute song that is a couple of dozen words. Nobody has an automatic system that removes that tail on sung vocals, so the design target is a tight bulk and a cheap manual fix.</p>
<p>Three manual fixes, from cheapest up:</p>
<p><strong>Nudge.</strong> Open the word&#x27;s popover. The arrows move it 50 ms earlier or later per click, 200 ms with Shift. Two clicks fixes most misses.</p>
<p><strong>Tap.</strong> Play the song, and press T at the moment you hear the word. Its onset moves to that moment minus 120 ms, which is roughly how long a person takes to press a key after hearing something, and the word is marked as verified so the Timing chip and the issues list stop flagging it. Tapping works because your ear is exactly the instrument the aligner lacks on a held note.</p>
<p><strong>Drag.</strong> In the Studio timeline, drag the word&#x27;s pill along the waveform. Useful when the miss is large and you can see the vocal energy the word should sit on.</p>
<p>None of these touch the beat grid, and none should. Onset snapping, moving each word to the nearest detected vocal onset, was in the pipeline for six months on the theory that it was the accuracy lever. Measured against human labelled onsets it was making timing worse, and <a href="https://www.revori.app/blog/onset-snapping-made-it-worse">deleting it improved the median</a>. The lesson generalises: a correction that cannot verify its target moves words away from the truth about as often as toward it.</p>
<h2 id="the-whole-track-is-early-or-late">The whole track is early or late</h2>
<p>Every word is off by about the same amount, first line to last. Two common causes.</p>
<p>The audio changed. If the analysis ran on one file and the project now plays a different master, a version with a different amount of leading silence, or a re-encode with a different amount of padding at the start, every timestamp is shifted by that difference. The words are right relative to the audio they were aligned to and wrong relative to the audio you are hearing.</p>
<p>Or it is taste. Some people want the highlight noticeably ahead of the voice, in the way karaoke systems often run, and some want it dead on. Neither is wrong.</p>
<p>Either way the fix is one control. In the Style panel, the Sync card has a Lyric offset slider from 2,000 ms early to 2,000 ms late in 50 ms steps. It shifts every word together and applies to the preview and the export identically, because both are rendered by the same composition. Set it, scrub three or four places in the song to confirm the shift is constant, and stop there.</p>
<p>If the drift is not constant, if the first verse is fine and the last chorus is late, the offset slider is the wrong tool. That is a section of misaligned words, and the fixes in the previous section apply to that section.</p>
<h2 id="it-is-on-time-and-still-reads-late">It is on time and still reads late</h2>
<p>Now the one you cannot fix by moving anything, because nothing is wrong.</p>
<p>A visual event and a sound register as simultaneous only when the visual arrives slightly first. Everyone learned this from watching music played: the stick is visibly moving before the snare sounds. So a word highlight that fires at the exact millisecond the vocal onset was measured reads as lagging, and beat-driven motion fired exactly on the beat reads as chasing the beat.</p>
<p>The renderer handles this with three leads, each measured against a musical event and each capped so it cannot overrun the previous one in a fast passage:</p>
<table><thead><tr><th>Lead</th><th>Value</th><th>Applies to</th></tr></thead><tbody><tr><td>Word highlight</td><td>50 ms early, capped at half the gap to the previous word</td><td>The active-word highlight or wipe</td></tr><tr><td>Line entrance</td><td>Up to 80 ms early, scaled to the gap before the line</td><td>A line appearing before its first word</td></tr><tr><td>Beat anticipation</td><td>3 frames, about 100 ms at 30 fps</td><td>Beat-driven pulses and motion</td></tr></tbody></table>
<p>There is also a release rule: the highlight normally holds until the next word starts, but if the next word is more than 250 ms away it releases at the word&#x27;s end instead of lingering across a silence.</p>
<p>Until September 2026 there was a fourth number. The highlight led by 70 ms and the whole clock was then shifted 80 ms the other way, on the theory that a listener recognises a word at its vowel rather than its first consonant. The theory is fine. The arithmetic was not: the two nearly cancelled, and the aligner&#x27;s onsets at the time were themselves about 27 ms late on average, so the net effect was a highlight roughly 10 ms behind an onset that was already behind. The clock shift was removed, the aligner&#x27;s median signed error is now 0 ms on full songs, and one lead of 50 ms is the whole story. The reasoning for the model behind this is in <a href="https://www.revori.app/blog/how-beat-sync-works">how beat sync actually works</a>.</p>
<p>If words still read late to you after all of that, you are in the previous section: it is taste, and the offset slider is the control for it.</p>
<h2 id="the-order-to-check-in">The order to check in</h2>
<ol>
<li>Read the transcript against the song. Fix the text, then Realign.</li>
<li>Look for Timing chips. Nudge or tap those words.</li>
<li>Scrub the song in four places. If the error is the same everywhere, set the Lyric offset once.</li>
<li>If it still reads late and every word is measurably on its onset, that is the 50 ms question, and it is a preference. The offset slider is the answer to it too.</li>
</ol>
<p>The first step fixes more lyric videos than the other three together, which is why it is first and why it is the one most people skip.</p>
<aside style="margin-top:3rem;padding:1.5rem;border:1px solid rgba(255,255,255,0.08);border-radius:12px;background:rgba(43,146,245,0.06);font-size:0.875rem;opacity:0.8"><p>Nudge, tap, realign and the offset slider are all free and unlimited on every plan. <a href="https://www.revori.app" style="color:#9CC7FF">revori.app</a></p></aside></div><section class="mt-20" aria-labelledby="post-faq-heading"><p id="post-faq-heading" class="eyebrow mb-6">Questions people ask about this</p><hr class="border-t border-white/[0.06] "/><dl class="mt-2"><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">Why do lyrics look late even when the timestamps are right?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Because a visual event reads as simultaneous with a sound only when the visual arrives slightly first. A highlight that fires at the exact millisecond the singer starts the word looks late to almost everyone. Revori fires every word highlight 50 ms before the vocal onset, capped at half the gap to the previous word so fast passages stay in order, and fires beat-driven motion three frames ahead of the beat.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">How do I fix one word that is off?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Open the word in Lyric Studio and nudge it: the popover moves it 50 ms per click, or 200 ms with Shift held. Or press T while the song plays at the moment you hear the word, which sets its onset to that moment minus a 120 ms reaction allowance and marks the word as verified. Words where the timing estimators disagreed carry a Timing chip, so start with those.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">How do I fix a whole video that is early or late?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Use the Lyric offset slider in the Sync card of the Style panel. It runs from 2,000 ms early to 2,000 ms late in 50 ms steps and shifts every word together. The preview and the export share one composition, so what you see after adjusting it is what renders. If only one section drifts, this is the wrong tool; fix the words in that section instead.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">Does Revori move lyrics onto the beat to fix timing?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">No. Word timing comes from forced alignment against the separated vocal and is never snapped to the beat grid. Singers push and drag against the beat on purpose, and a benchmark of onset snapping showed it made timing worse, not better. The grid is used for clip cuts, section markers and animation pulses, and left out of word placement.</dd></div></dl></section></div></section></article>]]></content>
  </entry>
  <entry>
    <title>Lyric video, visualizer, Canvas or lyric card: which one to make</title>
    <link rel="alternate" type="text/html" href="https://www.revori.app/blog/which-format-for-which-platform" />
    <id>https://www.revori.app/blog/which-format-for-which-platform</id>
    <published>2026-09-05T00:00:00Z</published>
    <updated>2026-09-05T00:00:00Z</updated>
    <category term="Guides" />
    <summary>Seven outputs from one song upload, and the one question that decides between them for each post. With the shape, the platform and the credit cost of each.</summary>
    <content type="html"><![CDATA[<article><section class=" py-20 md:py-24 lg:py-32  px-6 lg:px-10 "><div class="max-w-[40rem] mx-auto"><header class="mb-14"><div class="font-mono text-xs uppercase tracking-widest text-text-muted flex flex-wrap items-center gap-x-3 gap-y-1 mb-6"><span>Guides</span><span aria-hidden="true">·</span><time dateTime="2026-09-05">September 5, 2026</time><span aria-hidden="true">·</span><span>6 min<!-- --> read</span></div><h1 class="font-display text-[2.5rem] md:text-[4.5rem] font-normal text-white leading-[0.98] tracking-tight">Lyric video, visualizer, Canvas or lyric card: which one to make</h1></header><hr class="border-t border-white/[0.06] mb-14"/><div data-answer-summary="true" class="mb-14 rounded-2xl border border-white/[0.07] bg-surface-0 p-6 md:p-8"><p class="eyebrow mb-4">In short</p><p class="text-white text-lg md:text-xl leading-relaxed max-w-[42rem]">Each format answers one question about the post, and the question is per post, not per song. A lyric video is for a hook people will read and sing back, vertical for TikTok, Reels and Shorts. A music video is the 16:9 asset for YouTube. AutoCut is footage cut on the beat with no words. A visualizer is for a section with nothing to show. A Spotify Canvas is the silent three to eight second loop behind the track in the Spotify app, and it cannot be synced to the music. A lyric card is a still image with one line for feeds and stories. Cover art is the square the release ships with. One song upload and one analysis feed all of them.</p></div><div class="prose-editorial"><p>One song upload can become seven things in Revori, and the most common mistake is picking the format by habit: the artist who always makes a lyric video, the producer who always posts a visualizer. The formats are not interchangeable. Each one answers a specific question about a specific post, and the right way to choose is to ask that question for each post you plan, not once for the song.</p>
<p>Here is the whole set in one place, then the question each one answers.</p>
<table><thead><tr><th>Format</th><th>What it is</th><th>Shape</th><th>Where it goes</th><th>Credits</th></tr></thead><tbody><tr><td>Lyric video</td><td>Words timed to the vocal, over footage</td><td>9:16, 1:1 or 16:9, 1080p; 4K on Studio</td><td>TikTok, Reels, Shorts; Instagram feed; YouTube</td><td>50 per render, 100 for 4K</td></tr><tr><td>Music video</td><td>Footage or photos cut to the song, lyrics optional</td><td>16:9</td><td>YouTube</td><td>50</td></tr><tr><td>AutoCut</td><td>Footage cut on the beat, no words</td><td>9:16, 1:1 or 16:9</td><td>TikTok, Reels, Shorts</td><td>50</td></tr><tr><td>Visualizer</td><td>Audio-reactive motion over your own backdrop</td><td>9:16, 1:1 or 16:9</td><td>TikTok, Reels, YouTube</td><td>40</td></tr><tr><td>Spotify Canvas</td><td>Silent loop, 3 to 8 seconds</td><td>9:16, 1080x1920</td><td>Behind the track in the Spotify app</td><td>25, Starter and up</td></tr><tr><td>Lyric card</td><td>A still image carrying one line</td><td>1:1, 4:5, 9:16 or 16:9</td><td>Feed posts, stories, X</td><td>5</td></tr><tr><td>Cover art</td><td>The release image</td><td>Square, 3000x3000 for Spotify and Apple Music</td><td>The distributor, every profile</td><td>From 18</td></tr></tbody></table>
<p>Analysis runs once per song at 1 credit per second of audio and is shared across all of them. Editing, previewing and saving are free on every plan. The credits above are for renders only.</p>
<h2 id="start-from-the-post-not-the-song">Start from the post, not the song</h2>
<p>The question that sorts every format is: what is the listener going to do with this post? Read along and sing back. Watch a picture move to a sound. Glance at something in the corner of an app while the track plays. Screenshot a line. Recognise the release in a grid of other releases.</p>
<p>Those are different jobs, and the same song needs a different format for each of them, usually cut from a different section.</p>
<h2 id="a-hook-people-will-sing-back-lyric-video">A hook people will sing back: lyric video</h2>
<p>The words are the reason to stop scrolling. A chorus, a line that carries the song, a bridge everyone quotes. The lyric video shows those words timed to the vocal so the viewer reads and sings at once, over footage that sets the mood without competing with the text.</p>
<p>Vertical 9:16 for TikTok, Reels and Shorts, from the 20 to 30 seconds that carry the hook. A 1:1 cut for the Instagram feed and a 16:9 cut for YouTube are separate renders rather than crops, because a crop moves the text out of the safe area. The order of work that saves the most time is in <a href="https://www.revori.app/blog/how-to-make-a-lyric-video">how to make a lyric video</a>.</p>
<p>Not for: instrumentals, heavily processed vocals where the words are texture, or a section with no lyric worth reading. Those are visualizer territory.</p>
<h2 id="the-full-song-on-youtube-music-video">The full song on YouTube: music video</h2>
<p>YouTube is the one platform where a full-length, landscape video is the native format rather than a repost. The music video module is 16:9, built from your footage or a set of photos cut to the song&#x27;s structure, with lyrics as an option rather than the point. It is the asset that goes on the channel, gets linked from the streaming profile and lives as long as the song does.</p>
<p>Not for: vertical short-form. Post the lyric video there instead of a crop of this.</p>
<h2 id="footage-but-no-words-autocut">Footage but no words: AutoCut</h2>
<p>You have clips, the song has a beat, and the words are either absent or not the point. AutoCut places footage against the beat grid so every cut lands on a musical boundary, with no lyric layer at all. It is the lyric video&#x27;s editor minus the lyrics, which is why it costs the same and exports the same shapes.</p>
<p>Not for: a section where the listener should be reading. If they should, that is the lyric video.</p>
<h2 id="nothing-to-show-visualizer">Nothing to show: visualizer</h2>
<p>No footage, an instrumental, a producer tag, the sixteen bars before the vocal comes in. A visualizer draws the audio itself, frame by frame, over a backdrop you choose: the cover, a photo, a video clip, one colour, a gradient. It skips transcription entirely, so there is no transcript to correct. Five styles, each drawing the sound differently, are described in <a href="https://www.revori.app/blog/audio-visualizer-for-tiktok">the visualizer guide</a>.</p>
<p>Not for: a hook with words. Motion over a line people want to read makes the line harder to read.</p>
<h2 id="in-the-spotify-app-canvas">In the Spotify app: Canvas</h2>
<p>The short vertical loop that plays behind the track in the Spotify mobile app, replacing the static cover while the song plays. Three to eight seconds, silent, looping without a visible seam. It does not need the song upload at all: a Canvas is built from a clip trimmed to a loop or a generated shot, and the crossfade at the loop point is applied for you.</p>
<p>The one thing it cannot do is follow the music. Canvas is not tied to playback position, so a loop timed to the chorus will land somewhere else for a listener who skipped in from a playlist. The full requirements and the loop-point craft are in <a href="https://www.revori.app/blog/spotify-canvas-requirements">Spotify Canvas requirements</a>.</p>
<p>Not for: anything with a beat you want hit. Also not available on the free plan; Canvas renders start at Starter.</p>
<h2 id="a-line-for-the-feed-lyric-card">A line for the feed: lyric card</h2>
<p>One line of the song on a still image. It is the cheapest format by a distance, at 5 credits, and it does a job none of the video formats do: it gets screenshotted, reposted to stories, quoted, and it loads instantly in a feed that autoplays everything else. Four shapes: square for the feed, 4:5 portrait for the feed with more height, 9:16 for stories, 16:9 for X and headers.</p>
<p>Not for: replacing the lyric video. A card is the line; the video is the line arriving on time.</p>
<h2 id="the-release-itself-cover-art">The release itself: cover art</h2>
<p>The square image the release ships with, 3000x3000 for Spotify and Apple Music, plus a 1:1 and a 9:16 for profiles and stories. Generated from a prompt on one of three engines, priced by engine at 18, 60 or 90 credits per image, with the most expensive being the one to use once you have found the direction on the cheap one. It does not need a song upload.</p>
<p>Not for: anything that moves. It is the one still the platforms require.</p>
<h2 id="a-release-week-in-formats">A release week, in formats</h2>
<p>The formats stop competing once they are placed in time.</p>
<p>Weeks before, the cover art, because the distributor wants it first.</p>
<p>On release day, at the moment the track goes live, the Spotify Canvas, since it appears in the app immediately. Then the lyric video of the hook, vertical, because it is the post that gets shared and sung.</p>
<p>In the days after, a lyric card of the line people are quoting, a second lyric video from a different section if the first one travelled, and a visualizer for the intro or the instrumental version.</p>
<p>When the channel needs it, the 16:9 music video, or AutoCut over footage from the shoot.</p>
<p>Every one of those comes from the same upload and the same analysis. What changes is the section of the song and the question the post is answering.</p>
<aside style="margin-top:3rem;padding:1.5rem;border:1px solid rgba(255,255,255,0.08);border-radius:12px;background:rgba(43,146,245,0.06);font-size:0.875rem;opacity:0.8"><p>One upload, every format. Free to build and preview on any plan. <a href="https://www.revori.app" style="color:#9CC7FF">revori.app</a></p></aside></div><section class="mt-20" aria-labelledby="post-faq-heading"><p id="post-faq-heading" class="eyebrow mb-6">Questions people ask about this</p><hr class="border-t border-white/[0.06] "/><dl class="mt-2"><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">Do I need a separate upload for each format?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">No. One song upload is analysed once, at 1 credit per second of audio, and every format built from that song reuses the vocal separation, beat grid and transcript. A lyric video, a visualizer, a lyric card and a Canvas from the same track share the one analysis; you pay per render after that, and rendering is the only other thing that costs credits.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">Can a Spotify Canvas be synced to the beat?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">No, and not because of the tooling. A Canvas is not tied to playback position: a listener who starts the track from a playlist or a skip sees the loop from an arbitrary point, so any motion timed to a musical moment lands on a different moment for every listener. Make the loop join cleanly and let it be ambient. Timing to the music is what the lyric video and the visualizer are for.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">Which format should be ready on release day?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Cover art, because the distributor needs it weeks ahead, and the Spotify Canvas, because it appears the moment the track goes live. A lyric video of the hook is the first thing to post, since it is the one that gets sung back and shared. The visualizer, the lyric cards and any full-length video are the following days and weeks, cut from different sections of the same upload.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">Can I try every format on the free plan?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Mostly. The free tier has 500 lifetime credits and three exports, with editing, previewing and saving free and unlimited, so you can build everything and export three times. Spotify Canvas renders need the Starter plan or above. Free-tier video exports carry a small &quot;Made with Revori&quot; mark; paid tiers export without it.</dd></div></dl></section></div></section></article>]]></content>
  </entry>
  <entry>
    <title>The best fonts for lyric videos, and the four rules behind the list</title>
    <link rel="alternate" type="text/html" href="https://www.revori.app/blog/best-fonts-for-lyric-videos" />
    <id>https://www.revori.app/blog/best-fonts-for-lyric-videos</id>
    <published>2026-08-27T00:00:00Z</published>
    <updated>2026-08-27T00:00:00Z</updated>
    <category term="Craft" />
    <summary>Type in a lyric video is read at speed, at small size, over moving footage, after compression. That narrows the field a lot faster than genre does.</summary>
    <content type="html"><![CDATA[<article><section class=" py-20 md:py-24 lg:py-32  px-6 lg:px-10 "><div class="max-w-[40rem] mx-auto"><header class="mb-14"><div class="font-mono text-xs uppercase tracking-widest text-text-muted flex flex-wrap items-center gap-x-3 gap-y-1 mb-6"><span>Craft</span><span aria-hidden="true">·</span><time dateTime="2026-08-27">August 27, 2026</time><span aria-hidden="true">·</span><span>5 min<!-- --> read</span></div><h1 class="font-display text-[2.5rem] md:text-[4.5rem] font-normal text-white leading-[0.98] tracking-tight">The best fonts for lyric videos, and the four rules behind the list</h1></header><hr class="border-t border-white/[0.06] mb-14"/><div data-answer-summary="true" class="mb-14 rounded-2xl border border-white/[0.07] bg-surface-0 p-6 md:p-8"><p class="eyebrow mb-4">In short</p><p class="text-white text-lg md:text-xl leading-relaxed max-w-[42rem]">The best fonts for lyric videos are heavy, wide-countered display faces with large x-heights, because video type is read in under a second, at phone size, over moving footage, after platform compression. Anton and Bebas Neue for hip-hop and high-impact work, Poppins for pop, Space Grotesk for electronic, Playfair Display for cinematic and indie. Avoid thin weights, delicate serifs, tightly-spaced geometrics and anything below about 500 weight, because compression eats thin strokes and moving footage destroys low-contrast text. Genre matters, but it is the second filter, not the first.</p></div><div class="prose-editorial"><p>Most font advice for lyric videos is a genre lookup table. Genre matters, but it is the second filter. The first is that video type has to survive conditions print and web typography never face, and that constraint eliminates most of the field before genre gets a vote.</p>
<h2 id="the-four-conditions-type-has-to-survive">The four conditions type has to survive</h2>
<p><strong>It is read in under a second.</strong> A line is on screen for the length of a phrase. There is no re-reading, no scanning back. Legibility at a glance is worth more than legibility on inspection, and those are not the same property. Faces with distinctive overall word-shapes beat faces with beautifully drawn individual letters.</p>
<p><strong>It is small.</strong> A phone held at arm&#x27;s length. Whatever the canvas resolution says, the delivered experience is a few inches of glass. Counters, the enclosed spaces in a, e, o and g, are the first thing to close up, and once they close the letters stop being distinguishable from each other.</p>
<p><strong>It sits over moving footage.</strong> Contrast is not fixed. A word that reads cleanly over a dark frame disappears over a bright one two seconds later. This is what kills thin weights: a hairline stroke over busy footage has nothing to hold it.</p>
<p><strong>It goes through compression.</strong> Every platform re-encodes. Compression is hardest on high-frequency detail, which is exactly what thin strokes, fine serifs and tight letterspacing are. Type that looked crisp in your editor comes back softened.</p>
<p>Those four conditions add up to one rule that does most of the work: <strong>heavy, open, and large-x-height, or it does not survive the trip.</strong></p>
<h2 id="what-that-rules-out">What that rules out</h2>
<p>Thin and light weights, below roughly 500. They look sophisticated on a still and vanish in motion.</p>
<p>Delicate serifs with fine hairlines. Didone faces in particular. Bodoni is beautiful and useless here.</p>
<p>Tightly spaced geometrics. Futura and its descendants have small x-heights and closed counters relative to their cap height, so at video scale they read smaller than they measure.</p>
<p>Anything with a script or handwriting model, unless the whole video is built around it. Connected letterforms are slow to parse and slow is the one thing you cannot afford.</p>
<h2 id="the-list-by-genre">The list, by genre</h2>
<p>These are the faces we ship as preset defaults, which means they have been through real renders over real footage rather than chosen off a specimen sheet.</p>
<table><thead><tr><th>Genre</th><th>Face</th><th>Weight</th><th>Why</th></tr></thead><tbody><tr><td>Hip-hop, trap, drill</td><td>Anton</td><td>400</td><td>Extremely heavy at a single weight, tight vertical rhythm, enormous x-height. Holds at 3-word lines over anything</td></tr><tr><td>High-impact, any genre</td><td>Bebas Neue</td><td>400</td><td>Condensed caps, fits more per line than Anton without losing weight. Single weight by design</td></tr><tr><td>Pop, K-pop</td><td>Poppins</td><td>800</td><td>Geometric but with a generous x-height and circular counters that stay open. Reads friendly rather than aggressive</td></tr><tr><td>Electronic, house, techno</td><td>Space Grotesk</td><td>700</td><td>Strong idiosyncratic letterforms, mechanical without being cold. The quirks survive compression because they are structural, not fine</td></tr><tr><td>Cinematic, marquee</td><td>Playfair Display</td><td>700</td><td>A high-contrast serif that works only at its heaviest weight and at large size. Use it big or not at all</td></tr><tr><td>Indie, editorial, folk</td><td>A mixed-case grotesque</td><td>400</td><td>Quiet, sparse, 3 to 4 words a line. The restraint is the effect</td></tr><tr><td>Minimal, one word at a time</td><td>A neutral grotesque</td><td>400</td><td>When one word fills the frame, the face should get out of the way</td></tr></tbody></table>
<p>The last two rows are deliberately generic because our own presets use licensed commercial faces there, and recommending a face you would have to buy is not useful advice for a first video. Inter, Work Sans and Archivo cover that ground for free and get you most of the way.</p>
<p>Oswald deserves a mention as a middle point between Anton and Bebas: condensed, multiple weights available, less extreme than either.</p>
<h2 id="weight-matters-more-than-family">Weight matters more than family</h2>
<p>If you take one thing from this: picking the right weight of an adequate face beats picking the wrong weight of a perfect one.</p>
<p>Most of the failures we see are not a mismatched family. They are Poppins at 400 where it needed 800, or Playfair at regular where it needed bold. The face is fine. It is set too light to survive the four conditions above.</p>
<p>Anton and Bebas Neue both ship a single weight, which is not a limitation here. It is why they are hard to misuse.</p>
<h2 id="uppercase-or-lowercase">Uppercase or lowercase</h2>
<p>Uppercase is faster to recognise as a shape and slower to read as words, because word-shape recognition relies on ascenders and descenders that caps do not have. That trade is correct for three-word lines and wrong for anything longer.</p>
<p>All-lowercase reads as deliberate and informal. It suits hyperpop, bedroom pop and indie electronic, and it looks like an accident in a trap video.</p>
<p>Mixed case is the right default for anything that wants to read as considered rather than loud.</p>
<h2 id="one-face-per-video">One face per video</h2>
<p>Two typefaces in a 30-second vertical clip reads as indecision, not as hierarchy. If you need emphasis, you have weight, size, colour and position available before you need a second family.</p>
<p>The exception is a genuinely different element, such as an artist name or a release date sitting in a corner. That is not part of the lyric, so it can be set separately.</p>
<h2 id="the-practical-test">The practical test</h2>
<p>Set your line. Export a still. Look at it on your phone at arm&#x27;s length, over the brightest frame in your footage.</p>
<p>If you read it without effort the face is fine, and if you find yourself working at it the type is set too light. That test has been more reliable for us than any amount of specimen comparison, because it reproduces all four conditions at once.</p>
<aside style="margin-top:3rem;padding:1.5rem;border:1px solid rgba(255,255,255,0.08);border-radius:12px;background:rgba(43,146,245,0.06);font-size:0.875rem;opacity:0.8"><p>Revori’s font previews render your actual lyric, not a specimen line, so you judge the face on the words you are using. <a href="https://www.revori.app" style="color:#9CC7FF">revori.app</a></p></aside></div><section class="mt-20" aria-labelledby="post-faq-heading"><p id="post-faq-heading" class="eyebrow mb-6">Questions people ask about this</p><hr class="border-t border-white/[0.06] "/><dl class="mt-2"><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">What font is used in most TikTok lyric videos?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Heavy condensed sans faces dominate, with Anton and Bebas Neue the two most common, and Poppins in its heaviest weights close behind for pop. All three are free through Google Fonts. They win because they hold up under the specific constraints of the format: large x-height, wide counters that stay open at small size, and enough stroke weight to survive platform compression.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">What size should lyric video text be?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Large enough that a line fills most of the frame width at 3 to 6 words. On a 1080x1920 vertical canvas that typically means 80 to 140 pixels of cap height. The practical test is not a number: view it at phone size and see whether you read it or squint at it. Text that is comfortable on a desktop monitor is routinely too small in a feed.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">Should lyric video text be uppercase or lowercase?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Uppercase for impact-driven genres, mixed case for anything that wants to read as considered. Uppercase is faster to recognise as a shape at a glance but slower to read as words, which is the right trade for three-word lines and the wrong one for longer phrases. All-lowercase reads as deliberate and informal and suits hyperpop, bedroom pop and indie electronic.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">Do I need to pay for fonts to make a lyric video?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">No. Google Fonts covers the genuinely useful range for this format, is free for commercial use, and includes Anton, Bebas Neue, Poppins, Space Grotesk, Playfair Display and Oswald. Paid display faces buy you distinctiveness rather than quality, which matters for an established visual identity and very little for a first video.</dd></div></dl></section></div></section></article>]]></content>
  </entry>
  <entry>
    <title>How to make a lyric video for TikTok and Reels</title>
    <link rel="alternate" type="text/html" href="https://www.revori.app/blog/how-to-make-a-lyric-video" />
    <id>https://www.revori.app/blog/how-to-make-a-lyric-video</id>
    <published>2026-08-27T00:00:00Z</published>
    <updated>2026-08-27T00:00:00Z</updated>
    <category term="Guides" />
    <summary>Eight steps, in the order that saves the most work. Most of the effort goes into two of them, and they are not the two people expect.</summary>
    <content type="html"><![CDATA[<article><section class=" py-20 md:py-24 lg:py-32  px-6 lg:px-10 "><div class="max-w-[40rem] mx-auto"><header class="mb-14"><div class="font-mono text-xs uppercase tracking-widest text-text-muted flex flex-wrap items-center gap-x-3 gap-y-1 mb-6"><span>Guides</span><span aria-hidden="true">·</span><time dateTime="2026-08-27">August 27, 2026</time><span aria-hidden="true">·</span><span>5 min<!-- --> read</span></div><h1 class="font-display text-[2.5rem] md:text-[4.5rem] font-normal text-white leading-[0.98] tracking-tight">How to make a lyric video for TikTok and Reels</h1></header><hr class="border-t border-white/[0.06] mb-14"/><div data-answer-summary="true" class="mb-14 rounded-2xl border border-white/[0.07] bg-surface-0 p-6 md:p-8"><p class="eyebrow mb-4">In short</p><p class="text-white text-lg md:text-xl leading-relaxed max-w-[42rem]">Pick the 20 to 30 seconds of the song you actually want to post before you do anything else, because a full-length lyric video is the wrong deliverable for TikTok and Reels. Upload the audio and let transcription and beat detection run. Correct the transcript first, since every later decision depends on it. Choose typography and colour before you choose footage. Place clips so their edges land on the beat grid. Then cut your motion back to two or three moments in the whole clip. Preview at phone size, not desktop size, and export 9:16 at 1080p. The two steps that decide the result are correcting the transcript and cutting motion back, and both are the ones people skip.</p></div><div class="prose-editorial"><p>Most lyric video tutorials describe the software. This one describes the order, because the order is where the time actually goes. Two of these eight steps decide whether the result is any good, and they are the two that feel like admin.</p>
<h2 id="step-1">1. Choose the section before you choose anything else</h2>
<p>Decide which 20 to 30 seconds you are posting.</p>
<p>Vertical short-form is a clip format. A three-minute lyric video posted to TikTok is not a lyric video for TikTok, it is a YouTube asset in the wrong place. The section that earns a replay is the one a listener already recognises, which is almost always the hook and rarely the opening bars.</p>
<p>Doing this first means every later decision is scoped to half a minute of material rather than to a whole song, which is roughly an eightfold reduction in work.</p>
<h2 id="step-2">2. Upload and let the analysis run</h2>
<p>Upload the audio. Three things happen: the vocal gets separated from the mix, the beat grid gets detected across the whole track, and the lyrics get transcribed with a start time on every word.</p>
<p>These are genuinely separate systems answering separate questions, which is why a tool that only detects beats cannot place words. If you want the detail, it is in <a href="https://www.revori.app/blog/how-beat-sync-works">how beat sync actually works</a>.</p>
<p>Analysis costs 1 credit per second of audio in Revori and runs once per song. Reuse the same track for a visualizer or a Canvas later and you do not pay for it again.</p>
<h2 id="step-3">3. Correct the transcript. Before anything.</h2>
<p>This is the first of the two steps that matter and the one most people click past.</p>
<p>Sung vocals are hard. Layered harmonies, ad-libs, heavy processing, a producer&#x27;s reverb tail, and words that are half swallowed by design. Any system will make errors, and if you are working in a language it handles less well it will make more of them.</p>
<p>The reason to fix them now rather than later is structural. The text is the foundation. Line breaks are computed from it. Timing hangs off it. Layout, sizing and emphasis are all downstream. Correct a word after you have styled the video and you have invalidated a chunk of the work sitting on top of it.</p>
<p>Read it once against the song. It takes a minute on a 30-second section. If the whole transcript is wrong, paste your own lyric sheet instead of correcting word by word.</p>
<h2 id="step-4">4. Decide the look before you pick footage</h2>
<p>Typography and colour, in that order, before you open the clip library.</p>
<p>One typeface, chosen for the genre. Heavy condensed display for hip-hop. Geometric sans for electronic. Rounded humanist for pop. A serif with air for indie. There is a longer version with the reasoning in <a href="https://www.revori.app/blog/best-fonts-for-lyric-videos">the font guide</a>.</p>
<p>Then two or three related tones, drawn from the mood of the song rather than from a swatch grid. Desaturated blue-grey for melancholy. High-saturation complements for club. Amber and teal for heartbreak. Red on black for anger.</p>
<p>Deciding these before the footage is a constraint, and the constraint is the point. Picking clips first and then trying to make type work over them is how you end up with a video where every element is fine and nothing agrees.</p>
<h2 id="step-5">5. Place footage against the beat grid</h2>
<p>Now the clips. The rule is that clip edges land on beats, ideally downbeats, so that cuts happen on musical boundaries instead of at arbitrary timestamps.</p>
<p>This is what the beat grid is genuinely for. It is not used to move the words, which stay pinned to the vocal performance where the singer actually put them.</p>
<p>One clip track is enough. If a section needs two things happening at once, it usually needs one better thing.</p>
<h2 id="step-6">6. Cut the motion back</h2>
<p>The second step that decides the result.</p>
<p>Motion is a budget. Every animated word spends some of it. Animate every word and it is all spent, and the viewer&#x27;s eye has nothing to land on, because everything moving is indistinguishable from nothing moving.</p>
<p>Two or three large motion moments across a 30-second clip. Put them where the song has structure: the chorus entry, a key word, a section turning over. Between them, let the text sit still and let the music carry the rhythm.</p>
<p>If your tool makes it one click to give every word a different effect, that click is working against you.</p>
<h2 id="step-7">7. Preview at phone size</h2>
<p>Look at it in a phone-shaped viewport before you export.</p>
<p>Text that reads comfortably on a 27-inch monitor is routinely too small in a feed, and the failure is invisible at desktop scale. Three to six words on screen at a time is the working range for vertical. Anything denser gets skimmed instead of read.</p>
<p>Check the safe areas too. Platform chrome, captions and the interaction rail all cover parts of the frame that look empty in your editor.</p>
<h2 id="step-8">8. Export at the right ratio</h2>
<p>9:16 at 1080p for TikTok, Reels and Shorts. 1:1 for an Instagram feed post. 16:9 for YouTube proper.</p>
<p>Render each ratio you need as its own export rather than cropping one into another. A crop moves text out of the safe area and generally clips it. In Revori each ratio is a full separate render at 50 credits, because it is a full separate render.</p>
<h2 id="where-the-time-actually-goes">Where the time actually goes</h2>
<table><thead><tr><th>Step</th><th>Feels like</th><th>Actually costs</th></tr></thead><tbody><tr><td>Choosing the section</td><td>30 seconds</td><td>30 seconds, and saves hours</td></tr><tr><td>Analysis</td><td>Waiting</td><td>A few minutes, unattended</td></tr><tr><td>Correcting the transcript</td><td>Admin</td><td>One minute, and decides everything after it</td></tr><tr><td>Type and colour</td><td>The fun part</td><td>Ten minutes if you decide before you browse</td></tr><tr><td>Footage</td><td>The fun part</td><td>Expands to fill whatever time you allow</td></tr><tr><td>Cutting motion back</td><td>Undoing your own work</td><td>Five minutes, and the largest single quality lever</td></tr><tr><td>Phone preview</td><td>Optional</td><td>One minute, catches the most common failure</td></tr><tr><td>Export</td><td>Done</td><td>Minutes, unattended</td></tr></tbody></table>
<p>The pattern is that the two cheapest steps in wall-clock time are the two that determine the outcome, and both of them feel like they are getting in the way of the work rather than being the work.</p>
<aside style="margin-top:3rem;padding:1.5rem;border:1px solid rgba(255,255,255,0.08);border-radius:12px;background:rgba(43,146,245,0.06);font-size:0.875rem;opacity:0.8"><p>All eight steps in one place, free to try with 500 credits and no card. <a href="https://www.revori.app" style="color:#9CC7FF">revori.app</a></p></aside></div><section class="mt-20" aria-labelledby="post-faq-heading"><p id="post-faq-heading" class="eyebrow mb-6">Questions people ask about this</p><hr class="border-t border-white/[0.06] "/><dl class="mt-2"><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">How long should a lyric video be for TikTok?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Around 20 to 30 seconds, built around the hook. Vertical short-form is a clip format rather than a full-song format, and the section a listener recognises is what earns a replay. Post the full-length version to YouTube if you want one, and treat the vertical cut as a separate deliverable rather than a trim of it.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">What aspect ratio should a lyric video be?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">9:16 vertical at 1080x1920 for TikTok, Instagram Reels and YouTube Shorts. 1:1 for an Instagram feed post and 16:9 for YouTube proper. Render each ratio separately rather than cropping one file into another, because a crop moves text out of the safe area and usually clips it.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">Do I need to correct the AI transcription?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Yes, always read it against the song first. No transcription system is perfect on sung vocals, and every subsequent decision, including line breaks, timing and layout, is built on the text. A correction made after you have styled the video invalidates the work stacked on top of it, so the review costs a minute up front and saves considerably more later.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">Can I make a lyric video for free?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Yes. Revori’s free tier includes 500 credits with no card required, which covers analysing a few minutes of audio and rendering several exports, and editing, previewing and saving are unlimited and free on every tier. Free-tier exports carry a small &quot;Made with Revori&quot; mark; paid tiers export without it.</dd></div></dl></section></div></section></article>]]></content>
  </entry>
  <entry>
    <title>Spotify Canvas requirements, and why most loops have a visible seam</title>
    <link rel="alternate" type="text/html" href="https://www.revori.app/blog/spotify-canvas-requirements" />
    <id>https://www.revori.app/blog/spotify-canvas-requirements</id>
    <published>2026-08-27T00:00:00Z</published>
    <updated>2026-08-27T00:00:00Z</updated>
    <category term="Guides" />
    <summary>3 to 8 seconds, vertical, silent, looping. That much is easy to look up. The part nobody writes down is what makes a loop read as continuous.</summary>
    <content type="html"><![CDATA[<article><section class=" py-20 md:py-24 lg:py-32  px-6 lg:px-10 "><div class="max-w-[40rem] mx-auto"><header class="mb-14"><div class="font-mono text-xs uppercase tracking-widest text-text-muted flex flex-wrap items-center gap-x-3 gap-y-1 mb-6"><span>Guides</span><span aria-hidden="true">·</span><time dateTime="2026-08-27">August 27, 2026</time><span aria-hidden="true">·</span><span>4 min<!-- --> read</span></div><h1 class="font-display text-[2.5rem] md:text-[4.5rem] font-normal text-white leading-[0.98] tracking-tight">Spotify Canvas requirements, and why most loops have a visible seam</h1></header><hr class="border-t border-white/[0.06] mb-14"/><div data-answer-summary="true" class="mb-14 rounded-2xl border border-white/[0.07] bg-surface-0 p-6 md:p-8"><p class="eyebrow mb-4">In short</p><p class="text-white text-lg md:text-xl leading-relaxed max-w-[42rem]">A Spotify Canvas is a silent, vertical, looping video between 3 and 8 seconds long, in 9:16 aspect ratio. Deliver it at 1080x1920 so it holds up on large phones. It must have no audio, and it repeats indefinitely while the track plays, so the last frame sits directly against the first. Almost every amateur Canvas fails at that join: the shot ends somewhere different from where it started and the repeat reads as a restart. The fixes are to use one continuous camera move with no cuts, choose motion that returns to its origin, and cross-fade the tail back into the head over a second or two.</p></div><div class="prose-editorial"><p>Canvas is the short vertical video that plays behind a track in the Spotify mobile app. The format itself is simple enough to state in one line, and most articles about it stop there. The interesting problem is not the spec, it is the loop.</p>
<h2 id="the-spec">The spec</h2>
<table><thead><tr><th>Property</th><th>Value</th></tr></thead><tbody><tr><td>Aspect ratio</td><td>9:16, vertical</td></tr><tr><td>Delivery resolution</td><td>1080x1920</td></tr><tr><td>Duration</td><td>3 to 8 seconds</td></tr><tr><td>Audio</td><td>None. Any audio track is discarded</td></tr><tr><td>Playback</td><td>Repeats continuously for the length of the track</td></tr><tr><td>Content</td><td>No explicit imagery, no trademarks you do not own, no URLs, handles or calls to action in frame</td></tr></tbody></table>
<p>Spotify for Artists is the authority on the content rules and on anything about upload and eligibility, and those do change. Treat the table as the shape of the format rather than as a substitute for checking before a release.</p>
<p>The two constraints that actually shape the work are the ones people skip past: it is silent, and it never stops.</p>
<h2 id="silent-changes-what-you-can-do">Silent changes what you can do</h2>
<p>You cannot hit a beat. Not because the tooling will not let you, but because it would not mean anything: a Canvas is not synchronised to playback position, so a listener who starts the track from a playlist sees the loop at an arbitrary offset. Any motion timed to a musical moment will land on a different moment for every listener.</p>
<p>So the visual has to work at any phase. That rules out the whole vocabulary of the beat-synced lyric video, which is a different format solving a different problem. A Canvas wants continuous motion, not punctuated motion.</p>
<h2 id="looping-is-the-entire-craft-problem">Looping is the entire craft problem</h2>
<p>A Canvas repeats with no gap and no fade to black. On a four-minute track, a 4-second loop plays about sixty times. The listener is not watching it closely, but they are watching it in peripheral vision for four minutes, and the join is the thing they will notice.</p>
<p>Three rules cover most of it.</p>
<p><strong>One continuous shot, no cuts.</strong> A cut inside a 5-second loop reads as an edit, and once there is one edit in there the loop point becomes just another edit, which is worse. Keep it to a single move.</p>
<p><strong>Motion that returns to where it started.</strong> A slow orbit, a rotation, a drift that comes back, smoke or water or light that has no fixed position to return to. What fails is directional motion: a push-in that ends closer than it began has nowhere to go but a hard cut back out.</p>
<p><strong>Cross-fade the tail into the head.</strong> Even a well-chosen shot rarely ends exactly where it started. Overlapping the final second or two back over the opening frames hides the residual difference. We cap that overlap at two seconds, which is generous on a 3-second loop and about right on an 8-second one. Past that you are fading more than you are showing.</p>
<h2 id="what-to-actually-put-in-it">What to actually put in it</h2>
<p>The honest answer is that a Canvas is atmosphere, not information. It sits behind the now-playing screen while someone listens. It is not a place to explain anything, and text in it competes with the track title and artist name that Spotify draws over the top.</p>
<p>What works: a single subject with slow ambient motion. Texture, light, water, fabric, smoke. A crop of the cover art with life in it. The artist, in one continuous move, doing very little.</p>
<p>What does not: anything with a beginning and an end, dense text, rapid cuts, or a visual that only makes sense the first time you see it.</p>
<h2 id="making-one-without-shooting-one">Making one without shooting one</h2>
<p>If you have the footage, any editor that can export 9:16 at 1080x1920 with the audio stripped will do the job. The work is in the choosing and the loop point, not the export.</p>
<p>If you do not, the two routes are a library clip trimmed to a loop, or a generated shot. Both are in Revori: pick a start point in a clip, set a loop length between 3 and 8 seconds, and the crossfade is applied for you, or generate a one-shot vertical loop from a prompt. A Canvas render costs 25 credits. The song analysis you already paid for on a lyric video is reused, so a Canvas from a track you have already uploaded costs nothing beyond the render.</p>
<h2 id="the-short-version">The short version</h2>
<p>Vertical, 1080x1920, 3 to 8 seconds, silent, looping. After that, the join is where the remaining effort should go, because it is the part a listener sees sixty times in one song.</p>
<aside style="margin-top:3rem;padding:1.5rem;border:1px solid rgba(255,255,255,0.08);border-radius:12px;background:rgba(43,146,245,0.06);font-size:0.875rem;opacity:0.8"><p>Revori builds Canvas loops from the same song upload as your lyric video. <a href="https://www.revori.app" style="color:#9CC7FF">revori.app</a></p></aside></div><section class="mt-20" aria-labelledby="post-faq-heading"><p id="post-faq-heading" class="eyebrow mb-6">Questions people ask about this</p><hr class="border-t border-white/[0.06] "/><dl class="mt-2"><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">What are the Spotify Canvas requirements?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">A Canvas is a vertical 9:16 video with no audio, between 3 and 8 seconds long, that loops continuously while the track plays. Deliver it at 1080x1920. Spotify also applies content rules: no explicit imagery, no third-party trademarks or logos you do not own, and no calls to action, URLs or social handles burned into the frame. Spotify for Artists is the authority on the current rules and they do change, so check there before a release.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">What size and dimensions should a Spotify Canvas be?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Vertical 9:16. 1080x1920 is the practical delivery size: it is a standard phone resolution, it survives re-encoding, and it looks correct on large displays where lower-resolution uploads go soft. Anything wider than 9:16 gets cropped, and cropping a horizontal shot to vertical almost always cuts the subject.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">How long should a Spotify Canvas be?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Between 3 and 8 seconds, and shorter is usually better. A Canvas repeats for the entire length of the track, so a listener on a four-minute song sees a 4-second loop around sixty times. A short loop with one clear idea survives that repetition. A longer one with narrative in it becomes obvious and then irritating.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">Can a Spotify Canvas have sound?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">No. Canvas is silent by design, because the track is the audio. Any audio track in the file is discarded. This also means the visual cannot rely on hitting a beat, since a listener can start the song at any point and the loop is not synchronised to playback position.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">Why does my Spotify Canvas look like it restarts instead of looping?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Because the last frame does not match the first. A loop repeats with no gap, so any difference in position, brightness or subject placement between the tail and the head reads as a jump. Fix it by shooting or choosing one continuous move that returns to where it began, avoiding cuts entirely, and cross-fading the final second or two back into the opening frames.</dd></div></dl></section></div></section></article>]]></content>
  </entry>
  <entry>
    <title>How accurate is AI lyric transcription, really</title>
    <link rel="alternate" type="text/html" href="https://www.revori.app/blog/where-ai-lyric-transcription-fails" />
    <id>https://www.revori.app/blog/where-ai-lyric-transcription-fails</id>
    <published>2026-08-27T00:00:00Z</published>
    <updated>2026-08-27T00:00:00Z</updated>
    <category term="Engineering" />
    <summary>Our measured numbers against a human-labelled benchmark, the five things that reliably break automatic lyric timing, and what to do about each one.</summary>
    <content type="html"><![CDATA[<article><section class=" py-20 md:py-24 lg:py-32  px-6 lg:px-10 "><div class="max-w-[40rem] mx-auto"><header class="mb-14"><div class="font-mono text-xs uppercase tracking-widest text-text-muted flex flex-wrap items-center gap-x-3 gap-y-1 mb-6"><span>Engineering</span><span aria-hidden="true">·</span><time dateTime="2026-08-27">August 27, 2026</time><span aria-hidden="true">·</span><span>5 min<!-- --> read</span></div><h1 class="font-display text-[2.5rem] md:text-[4.5rem] font-normal text-white leading-[0.98] tracking-tight">How accurate is AI lyric transcription, really</h1></header><hr class="border-t border-white/[0.06] mb-14"/><div data-answer-summary="true" class="mb-14 rounded-2xl border border-white/[0.07] bg-surface-0 p-6 md:p-8"><p class="eyebrow mb-4">In short</p><p class="text-white text-lg md:text-xl leading-relaxed max-w-[42rem]">Measured against a human-annotated benchmark with correct lyrics supplied, our alignment reaches 97ms mean absolute error, 47ms median, and 95% of words within the 300ms tolerance the field treats as correct. That is the bulk of a song landing tight. It is also the optimistic figure, because it assumes the words themselves are right. On real uploads the transcription is the bottleneck, not the timing: a wrong word is a timing error by construction. Five things break it reliably, which are dense layered harmony, heavy vocal processing, ad-libs over the lead, quiet or sparse passages, and languages outside the dominant training data. The workable product is tight bulk plus a cheap manual fix, not perfection.</p></div><div class="prose-editorial"><p>Every tool in this category implies that you upload a song and get correct, timed lyrics. It is worth stating plainly what the actual ceiling is, because the gap between the claim and the reality is where people lose an afternoon.</p>
<h2 id="the-numbers">The numbers</h2>
<p>Measured on JamendoLyrics, which is 20 English songs with word onsets labelled by hand and released under MIT:</p>
<table><thead><tr><th>Metric</th><th>Our aligner</th><th>What it means</th></tr></thead><tbody><tr><td>Mean absolute error</td><td>97ms</td><td>Average distance from the human-labelled word start</td></tr><tr><td>Median error</td><td>47ms</td><td>Half of all words land inside this</td></tr><tr><td>Within 50ms</td><td>52%</td><td>Roughly a frame and a half at 30fps</td></tr><tr><td>Within 300ms</td><td>95%</td><td>The tolerance the field generally treats as correct</td></tr></tbody></table>
<p>For context, published systems as of 2024 sat around 200ms mean error. The bulk of a song lands tight.</p>
<p>Now the caveat that matters more than any of those numbers: <strong>the benchmark hands the aligner the correct lyrics.</strong> It measures alignment in isolation, deliberately, because that is the only way to test an alignment change without transcription noise swamping the result. It proves the aligner. It does not prove the product.</p>
<p>On a real upload, the text comes from speech recognition, and a wrong word is a timing error by construction. There is no timestamp for a word that was never sung.</p>
<h2 id="why-sung-vocals-are-hard">Why sung vocals are hard</h2>
<p>Speech recognition is built on assumptions that singing deliberately violates.</p>
<p>Pitch is supposed to be incidental to meaning. In singing it is the point, and it moves across a range speech never uses. Vowels are supposed to last a certain duration. In singing they are held for whole bars. Consonants are supposed to be crisp boundary markers. Singers soften them constantly because a hard consonant interrupts tone.</p>
<p>On top of that, the vocal shares frequency space with everything else in the mix. We separate the vocal stem before doing anything else, which helps a great deal and is not free: separation itself introduces artefacts, and it is imperfect on dense material.</p>
<h2 id="the-five-things-that-actually-break-it">The five things that actually break it</h2>
<p><strong>Layered harmony.</strong> Three vocal takes stacked in thirds are three simultaneously valid transcriptions of the same line, at slightly different times. The model has to pick one, and there is no principled basis for the choice.</p>
<p><strong>Heavy processing.</strong> Autotune, doubling, heavy reverb and delay all smear the acoustic boundaries the aligner reads. A long reverb tail in particular makes a word appear to end considerably later than it was sung.</p>
<p><strong>Ad-libs over the lead.</strong> A background vocal saying something different underneath the main line is, acoustically, another voice in the same stem. Ad-libs get interleaved into the main transcript in the wrong place more often than any other single failure.</p>
<p><strong>Quiet and sparse passages.</strong> An intro sung against near-silence gives energy-based detection almost nothing to work with. Counter-intuitively these are often harder than dense sections rather than easier.</p>
<p><strong>Languages outside the dominant training data.</strong> Accuracy tracks training data availability closely. Code-switching mid-line is the worst case, because the switch has to be detected before it can be transcribed across.</p>
<h2 id="the-tail">The tail</h2>
<p>Averages hide the shape of the failure, and the shape matters more than the average.</p>
<p>Our remaining error is not spread evenly. Most songs sit near the median. A small number fail badly, mis-seating entire phrases rather than drifting by a few tens of milliseconds. On our benchmark the worst track has a mean error of 466ms and a 90th percentile over 1.3 seconds while the median song is fine.</p>
<p>This is why an average is a misleading way to describe a system like this. Ninety-five percent of words within 300ms and one song in twenty that needs real intervention are both true at the same time, and the second one is what a user experiences on the day it happens to them.</p>
<h2 id="what-honest-looks-like-in-a-product">What honest looks like in a product</h2>
<p>Given all of the above, the design question is not how to reach perfection. It is what to do about the cases that will not reach it.</p>
<p>Our answers, in order of how much work they save:</p>
<p><strong>Show the transcript before anything is built on it.</strong> The review step exists because a correction is cheap before styling and expensive after. Every downstream decision, from line breaks to layout, sits on the text.</p>
<p><strong>Accept a pasted lyric sheet.</strong> If you have the real lyrics, giving them to the system removes the transcription problem entirely and leaves only the alignment problem, which is the one with the 97ms number attached to it. It is the single most useful thing a user can do here, and it is why the feature exists.</p>
<p><strong>Make re-matching one action.</strong> Sometimes the wrong song was matched. Correcting that line by line is absurd, so it is one button.</p>
<p><strong>Keep manual editing free and unlimited.</strong> Retiming a line, nudging a word, or shifting the whole track with a global offset costs nothing and is not rationed. A system with a known error tail cannot also put its correction tools behind a paywall.</p>
<h2 id="what-we-would-tell-you-before-you-upload">What we would tell you before you upload</h2>
<p>If your track is a single lead vocal, moderately processed, in a well-supported language, expect the transcript to need a couple of corrections and the timing to be publishable as-is.</p>
<p>If it is dense, layered, heavily processed, or in a smaller language, expect to paste your lyrics and expect to nudge a section or two.</p>
<p>Either way, listen to it once before exporting. We keep an internal rule that the objective benchmark is for iterating and the human ear is the final gate, and it applies to your track for the same reason it applies to ours: the metric cannot hear.</p>
<aside style="margin-top:3rem;padding:1.5rem;border:1px solid rgba(255,255,255,0.08);border-radius:12px;background:rgba(123,213,255,0.05);font-size:0.875rem;opacity:0.8"><p>The methodology behind these figures is in <a href="https://www.revori.app/blog/onset-snapping-made-it-worse" style="color:#7BD5FF">the benchmark writeup</a>. Try it on your own track at <a href="https://www.revori.app" style="color:#7BD5FF">revori.app</a>.</p></aside></div><section class="mt-20" aria-labelledby="post-faq-heading"><p id="post-faq-heading" class="eyebrow mb-6">Questions people ask about this</p><hr class="border-t border-white/[0.06] "/><dl class="mt-2"><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">How accurate is automatic lyric transcription and timing?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">On a human-annotated benchmark with correct lyric text supplied, our aligner reaches 97ms mean absolute onset error, 47ms median, and places 95% of words within 300ms of the labelled onset, which is the tolerance the field generally treats as correct. Published systems as of 2024 sat around 200ms mean error. Those figures measure alignment only. End-to-end accuracy on a real upload is lower, because it also depends on the transcription being right.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">Why does AI get song lyrics wrong?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Sung vocals violate most assumptions speech recognition is built on. Pitch is deliberately varied, vowels are held far beyond speech duration, consonants are softened for tone, and the vocal is mixed with instruments that occupy the same frequency range. Layered harmonies present several simultaneous versions of the same line, and heavy processing such as autotune, doubling and reverb smears the acoustic boundaries the model relies on.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">Can AI transcribe lyrics in languages other than English?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Yes, though accuracy drops noticeably outside the languages with the most training data. English, Spanish and the other major European languages perform best. Smaller languages, heavy regional accents and code-switching within a single line are all harder, and code-switching is the hardest of the three because the model has to detect the switch before it can transcribe across it.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">What should I do when the transcription is wrong?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Correct it before styling anything, because everything downstream is built on the text. For scattered errors, edit the individual words. If the transcript is broadly wrong, paste your own lyric sheet and let the system align that text instead of its own guess, which removes the transcription problem entirely and leaves only the alignment one. If the wrong song was matched, re-match it rather than correcting line by line.</dd></div></dl></section></div></section></article>]]></content>
  </entry>
  <entry>
    <title>We deleted the stage that was supposed to fix our lyric timing</title>
    <link rel="alternate" type="text/html" href="https://www.revori.app/blog/onset-snapping-made-it-worse" />
    <id>https://www.revori.app/blog/onset-snapping-made-it-worse</id>
    <published>2026-08-04T00:00:00Z</published>
    <updated>2026-08-27T00:00:00Z</updated>
    <category term="Engineering" />
    <summary>Onset snapping was the load-bearing accuracy lever, according to every source we read. We finally measured it against ground truth. It was making timing worse.</summary>
    <content type="html"><![CDATA[<article><section class=" py-20 md:py-24 lg:py-32  px-6 lg:px-10 "><div class="max-w-[40rem] mx-auto"><header class="mb-14"><div class="font-mono text-xs uppercase tracking-widest text-text-muted flex flex-wrap items-center gap-x-3 gap-y-1 mb-6"><span>Engineering</span><span aria-hidden="true">·</span><time dateTime="2026-08-04">August 4, 2026</time><span aria-hidden="true">·</span><span>Updated <time dateTime="2026-08-27">August 27, 2026</time></span><span aria-hidden="true">·</span><span>7 min<!-- --> read</span></div><h1 class="font-display text-[2.5rem] md:text-[4.5rem] font-normal text-white leading-[0.98] tracking-tight">We deleted the stage that was supposed to fix our lyric timing</h1></header><hr class="border-t border-white/[0.06] mb-14"/><div data-answer-summary="true" class="mb-14 rounded-2xl border border-white/[0.07] bg-surface-0 p-6 md:p-8"><p class="eyebrow mb-4">In short</p><p class="text-white text-lg md:text-xl leading-relaxed max-w-[42rem]">We ran a spectral onset-snapping stage on top of forced alignment for six months because the literature and our own architecture doc both described it as the main accuracy lever. Measured against the JamendoLyrics human-annotated benchmark, it was making timing worse: mean absolute error 113ms and 35% of words within 50ms of truth, against 98ms and 51% with the stage simply turned off. Snapping moved words onto vibrato peaks, breath and mid-consonant transients as often as onto real word starts. We replaced it with a bounded refinement that can only sharpen an edge within 60ms of where the aligner already put it, for a final 97ms mean error, 47ms median and 52% within 50ms.</p></div><div class="prose-editorial"><p>For about six months, Revori&#x27;s lyric alignment ended with a stage called onset snapping. Every piece of research we read described it as the load-bearing part: the thing that turns approximate word timings into tight ones.</p>
<p>We measured it against human-labelled ground truth. It was making the timing worse. We deleted it, and the median word landed 31% closer to where it belongs.</p>
<p>This is what we did, what the numbers were, and why we couldn&#x27;t see it for six months.</p>
<h2 id="what-the-stage-did">What the stage did</h2>
<p>Forced alignment gives you a timestamp for every word. Ours comes from wav2vec2 running on a demucs-separated vocal stem. The timestamps are good but slightly late, because an aligner tends to mark where a phoneme is clearly present rather than where it began.</p>
<p>The standard fix is to snap. You run a spectral onset detector (superflux, in our case) over the vocal stem, get a list of acoustic events, and move each word onset to the nearest one within some window. Ours used plus or minus 120ms. The logic is clean: the aligner knows <em>which</em> word, the onset detector knows <em>when</em> something started, so combine them.</p>
<p>Everyone does this. It is in the papers. It was in our own architecture doc, described in those words: the load-bearing lever.</p>
<h2 id="why-we-couldnt-tell-it-was-broken">Why we couldn&#x27;t tell it was broken</h2>
<p>Because lyric timing is judged by ear, and the ear is a terrible A/B instrument.</p>
<p>When you watch a lyric video, you are not measuring milliseconds. You are forming an overall impression, and that impression is polluted by the font, the animation curve, the footage cut, the loudness of the kick, and whether you already know what the next word is. Change the snapping and re-watch, and you will have an opinion. It will not be reliable. We had opinions in both directions for months.</p>
<p>Worse, the failure was not uniform. Snapping helped some words and hurt others. Averaged across a song, it felt like noise. Averaged across a catalogue with a metric, it was not noise at all.</p>
<p>So the actual fix was not an algorithm. It was building something that could tell us we were wrong.</p>
<h2 id="the-benchmark">The benchmark</h2>
<p>We used JamendoLyrics: 20 English songs, 1520 words, with manually labelled word onsets, released under MIT. Human ground truth, not another model&#x27;s output.</p>
<p>Two design decisions made it useful rather than merely existent.</p>
<p><strong>Score a window, not a whole song.</strong> Real uploads are roughly 30 second clips, not four minute masters. Scoring full songs adds a long tail of instrumental traversal that no user ever sees, and runs about eight times slower. We take the densest singing window of about 35 seconds per song, which is both faster and closer to what production actually does.</p>
<p><strong>Split prep from eval.</strong> The expensive part (demucs, then forced alignment) does not change when you vary the post-processing. So we run it once per song and cache the raw pre-snap alignment. Each variant then re-applies only its own post-processing step against that identical cached alignment.</p>
<p>That second decision is what makes the comparison trustworthy. Every variant scores against a byte-identical alignment, so the only thing varying is the step under test. It is a perfectly controlled A/B rather than two separate pipeline runs that happen to differ in one setting.</p>
<p>We report mean absolute onset error (AAE), median error, and PCO@50: the percentage of words landing within 50ms of truth, which is roughly a frame and a half at 30fps.</p>
<h2 id="the-result">The result</h2>
<table><thead><tr><th>mode</th><th>what it does</th><th>AAE</th><th>median</th><th>PCO@50</th></tr></thead><tbody><tr><td><code>snap</code> (legacy)</td><td>snap onset to nearest superflux onset, ±120ms</td><td>113ms</td><td>68ms</td><td>35%</td></tr><tr><td><code>off</code></td><td>trust the debiased aligner onset</td><td>98ms</td><td>48ms</td><td>51%</td></tr><tr><td><strong><code>refine</code></strong></td><td>nudge onset to the foot of its own energy rise, ±60ms</td><td><strong>97ms</strong></td><td><strong>47ms</strong></td><td><strong>52%</strong></td></tr></tbody></table>
<p>Doing nothing beat the load-bearing lever. Turning snapping off dropped median error from 68ms to 48ms and took words landing within 50ms from 35% to 51%.</p>
<p>That is the whole finding. The most sophisticated stage in the chain was worse than deleting it.</p>
<h2 id="why-snapping-lost">Why snapping lost</h2>
<p>Superflux onsets are not word onsets. They fire on any spectral flux increase, which in a sung vocal means vibrato peaks, consonant mid-points, breath, and the attack of a note that started three syllables ago.</p>
<p>So when you snap an already-decent onset to the nearest detected event, you move it away from truth roughly as often as toward it. The wins and losses cancel, and the losses are uglier, because a word yanked 90ms onto a vibrato peak reads as broken in a way that a word 40ms late does not.</p>
<p>We also tested the obvious response, which is to loosen detection so there are more onsets to choose from. That made it worse. More candidates means a closer wrong answer is always available.</p>
<h2 id="what-replaced-it">What replaced it</h2>
<p>The version we ship is called refine, and it is deliberately timid.</p>
<p>It anchors at the aligner&#x27;s own onset and walks backward to the foot of <em>that word&#x27;s</em> energy rise, bounded to plus or minus 60ms. It cannot jump to a different transient, because it never leaves the neighbourhood of the onset it started from. It only sharpens a leading edge that is already in the right place.</p>
<p>It also halves the systematic lateness. Mean bias went from +63ms to +24ms, which matters because a consistent lateness is exactly what reads as &quot;not quite on the beat&quot; even when nothing is obviously wrong.</p>
<p>The gain over simply switching snapping off is small: 48ms to 47ms median, 51% to 52% within tolerance. We kept it because the bias improvement is real and it costs almost nothing. But the honest summary is that most of the win came from deleting a stage, not from adding a better one.</p>
<p>It ships behind <code>LYRIC_ONSET_REFINE</code>, which takes <code>refine</code>, <code>off</code>, or <code>snap</code>, so the old behaviour is one environment variable away if we are wrong again.</p>
<h2 id="what-this-benchmark-does-not-prove">What this benchmark does not prove</h2>
<p>It uses perfect lyrics. Ground truth text is handed to the aligner, so the score isolates alignment quality from transcription quality.</p>
<p>Real uploads do not work that way. They depend on our ASR ensemble and the fusion step that reconciles it, and errors there produce timing failures that this benchmark cannot see. So the benchmark proves the aligner. It does not prove the product. A click track and a person listening remain the final gate before anything ships.</p>
<p>There is also a tail we have not fixed. Most of the remaining error sits in a handful of outlier songs rather than spread evenly. One track in the set sits at 466ms mean error with a p90 of 1309ms while the median song is fine. That tail is a different problem from the one this post is about, and it is the next thing worth paying for.</p>
<h2 id="if-you-are-building-something-similar">If you are building something similar</h2>
<p>Three things we would tell ourselves six months earlier.</p>
<p><strong>A stage everyone agrees is important is exactly the one to measure first.</strong> Consensus is why nobody checks. Ours had been described as load-bearing in our own documentation, which made it the last thing we suspected.</p>
<p><strong>Build the controlled comparison before the improvement.</strong> The prep/eval cache took an afternoon. It is the only reason we could run seven variants and believe the ordering. Every hour spent on it bought back weeks of arguing about whether a change felt better.</p>
<p><strong>Prefer bounded corrections to unbounded ones.</strong> Snapping could move a word anywhere within 120ms. Refine can only sharpen an edge within 60ms of where it already was. When your correction step cannot verify its own target, the bound is the thing keeping it honest.</p>
<p>The benchmark harness lives in <code>scripts/bench/</code> and the ground truth data is a fetch script away, so any future alignment change has to beat 97 / 47 / 52% before it ships.</p>
<p>Revori is at <a href="https://www.revori.app">revori.app</a> if you want to hear what this sounds like on a real track.</p></div><section class="mt-20" aria-labelledby="post-faq-heading"><p id="post-faq-heading" class="eyebrow mb-6">Questions people ask about this</p><hr class="border-t border-white/[0.06] "/><dl class="mt-2"><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">Does snapping word timings to detected onsets improve lyric alignment?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">In our measurements, no. Snapping forced-alignment word onsets to the nearest spectral onset within 120ms increased mean absolute error from 98ms to 113ms and cut the share of words landing within 50ms of ground truth from 51% to 35%. Spectral onset detectors fire on vibrato peaks, breath and consonant mid-points as well as word starts, so the snap target is wrong often enough to cancel out the cases where it is right.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">What is a good benchmark for lyric alignment accuracy?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">JamendoLyrics, from Stoller et al., is the practical choice: 20 English songs with manually labelled word onsets, released under MIT. Score mean absolute onset error, median error and percentage of correct onsets within a 50ms tolerance, which is roughly a frame and a half at 30fps. Cache the expensive vocal separation and alignment steps once per song so that post-processing variants are compared against a byte-identical alignment.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">How accurate is automatic lyric timing?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">On a human-annotated benchmark with correct lyrics supplied, our aligner reaches 97ms mean absolute error, 47ms median, and 52% of words within 50ms of the labelled onset. That figure isolates alignment quality: it assumes perfect lyric text, so it measures the aligner rather than the end-to-end product, where transcription errors introduce timing failures the benchmark cannot see.</dd></div></dl></section></div></section></article>]]></content>
  </entry>
  <entry>
    <title>Why I built Revori</title>
    <link rel="alternate" type="text/html" href="https://www.revori.app/blog/why-revori" />
    <id>https://www.revori.app/blog/why-revori</id>
    <published>2026-04-18T00:00:00Z</published>
    <updated>2026-08-27T00:00:00Z</updated>
    <category term="Product" />
    <summary>Four hours in Premiere for one lyric video, and the second verse still drifted. What the existing tools get wrong, and the narrow thing Revori does instead.</summary>
    <content type="html"><![CDATA[<article><section class=" py-20 md:py-24 lg:py-32  px-6 lg:px-10 "><div class="max-w-[40rem] mx-auto"><header class="mb-14"><div class="font-mono text-xs uppercase tracking-widest text-text-muted flex flex-wrap items-center gap-x-3 gap-y-1 mb-6"><span>Product</span><span aria-hidden="true">·</span><time dateTime="2026-04-18">April 18, 2026</time><span aria-hidden="true">·</span><span>Updated <time dateTime="2026-08-27">August 27, 2026</time></span><span aria-hidden="true">·</span><span>3 min<!-- --> read</span></div><h1 class="font-display text-[2.5rem] md:text-[4.5rem] font-normal text-white leading-[0.98] tracking-tight">Why I built Revori</h1></header><hr class="border-t border-white/[0.06] mb-14"/><div data-answer-summary="true" class="mb-14 rounded-2xl border border-white/[0.07] bg-surface-0 p-6 md:p-8"><p class="eyebrow mb-4">In short</p><p class="text-white text-lg md:text-xl leading-relaxed max-w-[42rem]">Revori turns one song upload into beat-synced lyric videos, visualizers, cover art and Spotify Canvas loops. It exists because professional editors know nothing about beats or lyrics, consumer editors sync to the waveform rather than to words, and the existing lyric-video SaaS category mostly ships weak transcription and weak presets. The AI does the mechanical work of transcription and beat detection; the taste decisions stay with the person. Editing, previewing and saving are always free, and credits are spent only on analysis and final renders.</p></div><div class="prose-editorial"><p>The first lyric video I made took four hours.</p>
<p>It was a Sunday. I had a song I liked and footage I liked and wanted to put them together. I spent four hours in Premiere nudging keyframes a frame at a time, exported, watched it back, and the second verse had already drifted. Start over.</p>
<p>There was no profound lesson in it. I was annoyed enough to build the thing instead.</p>
<h2 id="what-the-existing-tools-each-get-wrong">What the existing tools each get wrong</h2>
<p>Three categories, three different failures.</p>
<p><strong>Professional editors.</strong> Premiere, DaVinci Resolve, Final Cut. Total control and zero help. They do not know what a beat is. They do not know what a lyric is. Every sync is manual, and manual sync across a three-minute song is several hundred decisions you have to get individually right and then keep right through every re-edit.</p>
<p><strong>Consumer editors.</strong> CapCut, InShot, the native TikTok editor. Fast and genuinely good-looking, and they do have beat features. But those features read the audio waveform, not the vocal. They will cut your clip on the drum hits. They cannot put the word &quot;heartbreak&quot; on the syllable where it is actually sung, because nothing in the pipeline ever transcribed it.</p>
<p><strong>Lyric-video SaaS.</strong> The category that shows up in YouTube pre-roll. In principle these understand both beats and lyrics. In practice the transcription is where they fall down, and transcription is the part everything else depends on: a lyric video built on a bad transcript is unusable no matter how good the presets are.</p>
<p>None of these is wrong for its own purpose. None of them was the thing I wanted.</p>
<h2 id="the-thing-that-made-it-buildable">The thing that made it buildable</h2>
<p>Two capabilities got good enough at roughly the same time.</p>
<p>Speech recognition now returns word-level timestamps reliable enough to build on, particularly when you run several models and reconcile them rather than trusting one. Beat detection got a genuine step change with neural detectors: we run <code>beat_this</code>, from ISMIR 2024, with a multi-band librosa detector as the fallback path.</p>
<p>So the mechanical half of the job, the half that took me four hours, is now machine work. What is left is the half that was always the point: which typeface, which footage, which colour, where the video should hold still.</p>
<p>That split is the whole design. The AI does the part with a correct answer. The person does the part that does not have one.</p>
<h2 id="what-actually-happens-when-you-upload">What actually happens when you upload</h2>
<p>Upload a song. The vocal gets separated from the mix, the beat grid gets detected, and the lyrics get transcribed with a timestamp on every word. You review the transcript before you build anything, because no transcription is perfect and the review step is cheaper than discovering an error at export.</p>
<p>Then you pick a preset, pick clips from a library of 1,100 or more royalty-free videos or upload your own, and export a 1080p MP4. Default aspect is 9:16 because the destination is usually TikTok or Reels, with 1:1 and 16:9 available.</p>
<p>Analysis runs once per song. Reuse that same song for a lyric video, a visualizer, cover art and a Spotify Canvas without paying for the analysis again.</p>
<h2 id="what-it-costs-plainly">What it costs, plainly</h2>
<p>Editing, previewing and saving are all free, with unlimited projects on every tier including the free one.</p>
<p>Credits are spent on two things only: 1 credit per second of audio analysed, and 50 credits for a standard 1080p render. The free tier starts with 500 credits and does not ask for a card.</p>
<p>Free-tier exports carry a small &quot;Made with Revori&quot; mark in the bottom corner. Paid exports do not. I would rather say that here than have you find it after you have built something.</p>
<h2 id="what-revori-is-not">What Revori is not</h2>
<p>Revori is not a platform, a social network or an ecosystem. It does not collaborate with your team, match you with session musicians, or distribute your track to 400 streaming services. Each of those is a different product, and trying to be all of them at once is the standard route to a tool that does nothing well.</p>
<p>It takes your audio and gives back something you would actually post.</p>
<aside style="margin-top:3rem;padding:1.5rem;border:1px solid rgba(255,255,255,0.08);border-radius:12px;background:rgba(43,146,245,0.06);font-size:0.875rem;opacity:0.8"><p>Live at <a href="https://www.revori.app" style="color:#9CC7FF">revori.app</a>. 500 credits free, no card required.</p></aside></div><section class="mt-20" aria-labelledby="post-faq-heading"><p id="post-faq-heading" class="eyebrow mb-6">Questions people ask about this</p><hr class="border-t border-white/[0.06] "/><dl class="mt-2"><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">What is Revori?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Revori is a browser-based tool that turns a single song upload into beat-synced lyric videos, audio visualizers, cover art, Spotify Canvas loops and lyric cards. It transcribes the lyrics with word-level timing and detects the beat grid automatically, then the user makes the visual decisions and exports a 1080p MP4 sized for TikTok, Reels, YouTube or Spotify.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">How is Revori different from CapCut or Premiere for lyric videos?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Premiere and DaVinci Resolve give full control and no assistance: they have no concept of a beat or a lyric, so every word is placed by hand. CapCut and similar consumer editors can cut to the audio waveform but cannot place individual words, because they do not transcribe the vocal. Revori aligns each word to the vocal performance and marks the beat grid separately, so the words and the cuts are handled by different systems that each know what they are looking at.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">Is Revori free to use?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Editing, previewing and saving are always free and unlimited on every plan, including the free tier, which starts with 500 credits and no card. Credits are only spent on AI work: 1 credit per second of audio analysis, and 50 credits for a standard 1080p render. Free-tier exports carry a small &quot;Made with Revori&quot; mark in the corner; paid tiers export without it.</dd></div></dl></section></div></section></article>]]></content>
  </entry>
  <entry>
    <title>How beat sync actually works in a lyric video</title>
    <link rel="alternate" type="text/html" href="https://www.revori.app/blog/how-beat-sync-works" />
    <id>https://www.revori.app/blog/how-beat-sync-works</id>
    <published>2026-04-10T00:00:00Z</published>
    <updated>2026-09-05T00:00:00Z</updated>
    <category term="Engineering" />
    <summary>Words and beats are two separate clocks, and most tools treat them as one. What each is for, and the three leads that decide whether sync reads as tight.</summary>
    <content type="html"><![CDATA[<article><section class=" py-20 md:py-24 lg:py-32  px-6 lg:px-10 "><div class="max-w-[40rem] mx-auto"><header class="mb-14"><div class="font-mono text-xs uppercase tracking-widest text-text-muted flex flex-wrap items-center gap-x-3 gap-y-1 mb-6"><span>Engineering</span><span aria-hidden="true">·</span><time dateTime="2026-04-10">April 10, 2026</time><span aria-hidden="true">·</span><span>Updated <time dateTime="2026-09-05">September 5, 2026</time></span><span aria-hidden="true">·</span><span>7 min<!-- --> read</span></div><h1 class="font-display text-[2.5rem] md:text-[4.5rem] font-normal text-white leading-[0.98] tracking-tight">How beat sync actually works in a lyric video</h1></header><hr class="border-t border-white/[0.06] mb-14"/><div data-answer-summary="true" class="mb-14 rounded-2xl border border-white/[0.07] bg-surface-0 p-6 md:p-8"><p class="eyebrow mb-4">In short</p><p class="text-white text-lg md:text-xl leading-relaxed max-w-[42rem]">Lyric timing and beat timing are two independent clocks. Word timing comes from forced alignment against the isolated vocal, and is never moved onto the beat grid. The beat grid drives clip cuts, section markers and animation pulses. Sync feels right only once you add three small leads: fire beat-driven motion about 100ms early, start the word highlight 50ms before the vocal onset, and bring a line in up to 80ms ahead of its first word. The clock itself is never shifted; until September 2026 it was, and the shift was cancelling the lead.</p></div><div class="prose-editorial"><p>If you have ever built a lyric video by hand, you know the specific frustration: the word looks like it lands on the beat, you play it back, and it is a frame off. You nudge it. Now it is a frame off the other way. Sometimes it is exactly on the beat and it still feels wrong.</p>
<p>That last case is the interesting one, and it is not a timing bug. It is a sign that you are treating one clock where there are two.</p>
<h2 id="beats-and-words-are-two-different-clocks">Beats and words are two different clocks</h2>
<p>A song gives you two independent timelines, and they are produced by completely different machinery.</p>
<p>The <strong>beat grid</strong> is a property of the whole mix. A beat detector listens to the full track and returns a list of beat times, which of them are downbeats, and the tempo. Ours is <code>beat_this</code>, a neural detector from ISMIR 2024, with a multi-band librosa detector as fallback when the model cannot load. It knows nothing about lyrics. It would return the same grid for an instrumental.</p>
<p><strong>Word timing</strong> is a property of the vocal performance. It comes from forced alignment: you hand an aligner the audio and the text, and it returns a start time for every word. Ours runs wav2vec2 over a vocal stem that demucs has separated from the mix, because aligning against a full mix means aligning against a drum kit. It knows nothing about tempo. It would return the same timings if you halved the BPM of the backing track.</p>
<p>These two lists do not line up, and they are not supposed to. A singer who lands every syllable exactly on the grid sounds like a metronome. The push and drag against the beat is the performance.</p>
<h2 id="what-we-do-not-do-move-words-onto-the-beat">What we do not do: move words onto the beat</h2>
<p>This is the part most tools get wrong, and it is worth being blunt about because we shipped the wrong version of it for six months.</p>
<p>Every word we produce carries an annotation naming its nearest beat: <code>nearest_beat_ms</code>, the beat number, whether that beat is a downbeat, and a confidence score. That annotation is used to draw markers, to colour the timeline, and to decide where a clip edge should snap.</p>
<p>It is never used to move the word.</p>
<p>We did once run a snapping stage, though it snapped to detected vocal onsets rather than to beats. It was described in our own architecture doc as the load-bearing accuracy lever. When we finally measured it against human-labelled ground truth it was making timing worse, and deleting it dropped median word error from 68ms to 48ms. The full numbers are in <a href="https://www.revori.app/blog/onset-snapping-made-it-worse">the benchmark writeup</a>.</p>
<p>The general lesson survives the specific case. A correction step that cannot verify its own target will move things away from truth about as often as toward it.</p>
<h2 id="where-the-beat-grid-is-genuinely-useful">Where the beat grid is genuinely useful</h2>
<p>Having said all that, the grid is not decoration. It carries four jobs:</p>
<p>Clip edges snap to it, so cuts land on musical boundaries instead of arbitrary timestamps. Section boundaries are labelled from it, so a chorus can look different from a verse. Beat-driven motion, the pulse or scale bump that makes a video feel like it is moving with the song, is triggered from it. And it gives you the visible ruler that makes manual editing tractable at all.</p>
<p>None of those require touching a single word timestamp.</p>
<h2 id="why-exactly-on-the-beat-reads-as-late">Why exactly on the beat reads as late</h2>
<p>Here is the effect that catches everyone.</p>
<p>A visual event and an audio event register as simultaneous only when the visual arrives slightly first. Live performance has physical latency baked into it. A drummer&#x27;s stick is visibly moving before the snare sounds. A guitarist&#x27;s pick crosses the string before the note. Watching music has trained everyone&#x27;s expectations around that lead.</p>
<p>So an animation fired at the exact beat timestamp reads as lagging, even though the arithmetic is perfect. The fix is to fire early:</p>
<pre><code class="language-js">const { fps } = useVideoConfig();           // never hardcode 30
const BEAT_ANTICIPATION_FRAMES = 3;         // ~100ms at 30fps

const beatFrame = Math.round((beatTimeMs / 1000) * fps);
const triggerFrame = beatFrame - BEAT_ANTICIPATION_FRAMES;
</code></pre>
<p>Three frames is the value we ship. Below two the lead stops being perceptible. Past five the motion detaches and starts reading as early, which is a different kind of wrong.</p>
<h2 id="the-lead-that-used-to-be-cancelled-out">The lead that used to be cancelled out</h2>
<p>The anticipation offset is well known. This one is a correction of our own, and it is worth writing down because the mistake is easy to make.</p>
<p>Alignment pins a word to its <strong>acoustic</strong> onset: the consonant attack, or the start of an energy rise. That is the correct answer to the question the aligner was asked. A common theory says it is not where a listener hears the word, because word recognition anchors on the vowel, which for a consonant-initial word lands 50 to 100ms after the acoustic start. Until September 2026 we acted on that theory: the highlight led the onset by 70ms, and the whole karaoke clock was then shifted 80ms the other way to land on the vowel.</p>
<p>Measured, the two nearly cancelled. Worse, the aligner&#x27;s onsets at the time were themselves about 27ms late on average, so the net result was a highlight roughly 10ms behind an onset that was already behind. The clock shift was removed. The aligner now takes a second opinion from other timing estimators on every word, its median signed error is 0ms on full songs, and there is one visual lead: the highlight fires 50ms before the onset, capped at half the gap to the previous word so a fast run stays in order.</p>
<p>Reading lags listening, and visual-early is the forgiving direction. Visual-late reads as off almost immediately. That asymmetry is the whole reason the lead exists and the reason it is small.</p>
<h2 id="the-three-leads-in-one-place">The three leads, in one place</h2>
<table><thead><tr><th>Lead</th><th>Value</th><th>What it applies to</th><th>Why</th></tr></thead><tbody><tr><td>Beat anticipation</td><td>3 frames, ~100ms at 30fps</td><td>Beat-driven pulses and motion</td><td>Visual must lead audio to read as simultaneous</td></tr><tr><td>Word highlight</td><td>50ms, capped at half the gap to the previous word</td><td>The active-word highlight</td><td>Reading lags listening; early is the forgiving direction</td></tr><tr><td>Line entrance</td><td>Up to 80ms, scaled to the gap</td><td>A line appearing before its first word</td><td>The eye needs to find the line before it needs to read it</td></tr></tbody></table>
<p>All three are measured against a musical event, and every one is capped relative to the surrounding gap, so none can overrun the previous word during fast passages. A fixed 80ms lead is fine at 90 BPM and catastrophic in a triplet run. There is also a release rule: the highlight holds to the next word unless the next word is more than 250ms away, in which case it lets go at the word&#x27;s end instead of lingering across the silence.</p>
<p>If the result still reads early or late to you, the Sync card in the Style panel has a per-project offset from 2,000ms early to 2,000ms late. The four different reasons lyrics can look off, and which tool fixes each, are in <a href="https://www.revori.app/blog/lyrics-out-of-sync">why the words look early or late</a>.</p>
<h2 id="what-this-adds-up-to">What this adds up to</h2>
<p>Beat sync sounds like one feature and is really a stack of small decisions, most of which are about perception rather than arithmetic. You detect the beats, separate the vocal, align the words against the vocal and leave them there, use the grid for cuts and motion, then apply three small leads, each bounded so it cannot run into its neighbour.</p>
<p>None of that is visible in the finished video. It shows up only as words landing where you expect them to.</p>
<aside style="margin-top:3rem;padding:1.5rem;border:1px solid rgba(255,255,255,0.08);border-radius:12px;background:rgba(123,213,255,0.05);font-size:0.875rem;opacity:0.8"><p>All of this runs once per upload, server side. Try it on your own track at <a href="https://www.revori.app" style="color:#7BD5FF">revori.app</a>.</p></aside></div><section class="mt-20" aria-labelledby="post-faq-heading"><p id="post-faq-heading" class="eyebrow mb-6">Questions people ask about this</p><hr class="border-t border-white/[0.06] "/><dl class="mt-2"><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">Should lyrics be snapped to the beat in a lyric video?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">No. Words should be timed to the vocal performance, not moved onto the beat grid. Singers push and drag against the beat deliberately, and forcing each word onto the nearest beat destroys that phrasing. Use the beat grid for clip cuts, section changes and animation pulses, and let word timing follow the voice.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">Why do my lyrics look late even when the timing is correct?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Because a visual event reads as simultaneous with a sound only when the visual arrives slightly first. Animation fired exactly on the beat timestamp reads as late, and a word highlight fired at the exact millisecond of the vocal onset reads as late too. The fix is a small lead: the highlight fires 50ms before the onset, capped at half the gap to the previous word, and beat motion fires three frames early.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">How far ahead of the beat should a lyric animation fire?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Around 100ms, which is 3 frames at 30fps. Below about 2 frames the anticipation stops being perceptible; past about 5 frames the motion detaches from the beat and reads as early rather than tight. Denser genres want the lower end of that range and sparse ones tolerate the higher end.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">What is the difference between beat detection and forced alignment?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Beat detection finds the rhythmic pulse of the whole mix and returns beat times, downbeats and tempo. Forced alignment takes a known lyric text plus the audio and returns a start time for each individual word. They answer different questions and run on different inputs, which is why a tool that only does beat detection cannot place words.</dd></div></dl></section></div></section></article>]]></content>
  </entry>
  <entry>
    <title>What makes a good lyric video</title>
    <link rel="alternate" type="text/html" href="https://www.revori.app/blog/anatomy-of-a-lyric-video" />
    <id>https://www.revori.app/blog/anatomy-of-a-lyric-video</id>
    <published>2026-04-03T00:00:00Z</published>
    <updated>2026-09-05T00:00:00Z</updated>
    <category term="Craft" />
    <summary>Four things separate the lyric videos that hold a viewer from the ones people scroll past. Effects are not one of them, and the fourth is the largest.</summary>
    <content type="html"><![CDATA[<article><section class=" py-20 md:py-24 lg:py-32  px-6 lg:px-10 "><div class="max-w-[40rem] mx-auto"><header class="mb-14"><div class="font-mono text-xs uppercase tracking-widest text-text-muted flex flex-wrap items-center gap-x-3 gap-y-1 mb-6"><span>Craft</span><span aria-hidden="true">·</span><time dateTime="2026-04-03">April 3, 2026</time><span aria-hidden="true">·</span><span>Updated <time dateTime="2026-09-05">September 5, 2026</time></span><span aria-hidden="true">·</span><span>4 min<!-- --> read</span></div><h1 class="font-display text-[2.5rem] md:text-[4.5rem] font-normal text-white leading-[0.98] tracking-tight">What makes a good lyric video</h1></header><hr class="border-t border-white/[0.06] mb-14"/><div data-answer-summary="true" class="mb-14 rounded-2xl border border-white/[0.07] bg-surface-0 p-6 md:p-8"><p class="eyebrow mb-4">In short</p><p class="text-white text-lg md:text-xl leading-relaxed max-w-[42rem]">A good lyric video gets four things right: typography that signals the genre, motion that fires slightly before the beat rather than on it, a colour palette chosen to match the mood before anything else is decided, and restraint, meaning two or three deliberate motion moments per song instead of an effect on every word. Effects, transitions and filters do not appear on that list. A video that gets all four right will outperform one with better effects and worse choices.</p></div><div class="prose-editorial"><p>Most lyric videos on TikTok are bad, and it is usually not a taste problem. It is a tooling problem. The tools make the decorative choices easy and the structural ones invisible, so people spend their effort on the layer that matters least.</p>
<p>Four things actually decide whether a lyric video holds someone. None of them is an effect.</p>
<h2 id="step-1">1. Does the typography match the genre?</h2>
<p>Typography is a genre signal before it is a legibility decision. A face carries a set of associations, and using the wrong one tells a viewer you are outside the music, which is a harder problem than looking plain.</p>
<p>The rough map, which is also roughly what our own presets ship with:</p>
<p>Heavy condensed display faces for hip-hop and trap. Anton, Bebas Neue, Oswald. The visual weight matches the low end.</p>
<p>Geometric sans with strong letterforms for electronic and house. Space Grotesk, Montserrat at its heaviest. Clean, synthetic, repeating.</p>
<p>Rounded humanist sans for pop. Poppins, Inter, Nunito. Approachable, which is the register pop is in.</p>
<p>A serif with air for indie and folk. Playfair Display, Cormorant Garamond, Instrument Serif. It should suggest a page rather than a billboard.</p>
<p>Italic scripts or oblique serifs for R&amp;B and soul, where the curves echo the melisma.</p>
<p>There is a longer version of this with the reasoning per genre in <a href="https://www.revori.app/blog/best-fonts-for-lyric-videos">the font guide</a>.</p>
<h2 id="step-2">2. Does the motion arrive before the beat?</h2>
<p>If your animation fires on the exact beat timestamp, it will feel late. This is not a preference. Visual events register as simultaneous with audio only when the visual arrives slightly first, which is why a drummer&#x27;s stick is visibly moving before the snare sounds.</p>
<p>The practical number is about 100ms, or three frames at 30fps. That single offset is the difference between motion that feels locked to the track and motion that feels like it is chasing it, and it is invisible until you flip it back and forth.</p>
<p>The word highlight has its own version of the same rule: it fires 50ms before the vocal onset, because reading lags listening and early is the forgiving direction, while a highlight that arrives late reads as off almost immediately. Both leads are covered in <a href="https://www.revori.app/blog/how-beat-sync-works">how beat sync actually works</a>.</p>
<h2 id="step-3">3. Does the colour follow the mood, or fight it?</h2>
<p>Pick the colour story first, before the font, before the footage. Everything downstream gets easier and the result is coherent rather than assembled.</p>
<p>Melancholy wants desaturated palettes. Pull the vibrance out, sit in blue-grey, keep text warm white and never hot.</p>
<p>Party and club want high-saturation contrast. Magenta against cyan, lime against navy. Colours that look lit by a fog machine, because they were.</p>
<p>Heartbreak sits in amber and teal, which is the cinematographer&#x27;s default for a reason and reads as film stock rather than filter.</p>
<p>Anger wants pure red and pure black. No gradients, nothing else in the frame.</p>
<p>The fastest way to a video that feels wrong is to skip this step and pick the brightest options from a preset picker. Nothing about it will look broken. It will just feel like it belongs to a different song.</p>
<h2 id="step-4">4. Do you know when to leave the screen still?</h2>
<p>This is the one nobody talks about and it is the largest of the four.</p>
<p>Watch a music video you actually like and count the frames where the text is moving. It is a small fraction of the runtime. Most of the video is type sitting still while the rhythm happens in the music instead of on the screen.</p>
<p>Motion is a budget. Every animated word spends some of it. Animate every word and you have spent all of it, and the viewer&#x27;s eye has nothing to catch on, because everything moving is the same as nothing moving.</p>
<p>Good lyric videos have two or three large motion moments per song, placed where the song already has structure: a chorus arriving, a key word, a section turning over. In between, the text sits still.</p>
<p>A tool that makes it easy to give every word a different effect is working against you. Our defaults animate the first word of a line and then hold until the next line starts. You can override that, but you have to decide to, which is the point.</p>
<h2 id="the-four-side-by-side">The four, side by side</h2>
<table><thead><tr><th>Element</th><th>Get it right</th><th>Get it wrong</th></tr></thead><tbody><tr><td>Typography</td><td>One genre-appropriate face, used throughout</td><td>A face chosen for novelty, or several per video</td></tr><tr><td>Motion timing</td><td>Fires ~100ms before the beat</td><td>Fires exactly on the beat, reads as late</td></tr><tr><td>Colour</td><td>Palette picked first, from the mood</td><td>Brightest preset, picked last</td></tr><tr><td>Restraint</td><td>Two or three motion moments per song</td><td>An effect on every word</td></tr></tbody></table>
<h2 id="the-short-version">The short version</h2>
<p>Good lyric videos are correct more than they are fancy: genre-appropriate type, motion that anticipates the beat, colour that matches the mood, and restraint. Getting those four right matters more than any effect you could add on top, and getting one of them wrong is not something an effect will cover.</p>
<aside style="margin-top:3rem;padding:1.5rem;border:1px solid rgba(255,255,255,0.08);border-radius:12px;background:rgba(43,146,245,0.06);font-size:0.875rem;opacity:0.8"><p>Our presets are built around these four principles rather than around effect menus. Try them at <a href="https://www.revori.app" style="color:#9CC7FF">revori.app</a>.</p></aside></div><section class="mt-20" aria-labelledby="post-faq-heading"><p id="post-faq-heading" class="eyebrow mb-6">Questions people ask about this</p><hr class="border-t border-white/[0.06] "/><dl class="mt-2"><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">What makes a lyric video look professional?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Consistency and restraint, more than production value. A professional-looking lyric video uses one typeface throughout, holds the text still for most of the runtime, keeps its colour palette to two or three related tones, and reserves motion for a handful of structural moments such as a chorus entry. Videos read as amateur when every line has a different animation.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">How many words should be on screen at once in a lyric video?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">Three to six for vertical formats such as TikTok and Reels, and rarely more than one line. Vertical video is read at arm’s length on a small screen while the viewer is also listening, so anything denser gets skimmed rather than read. Sparse text also leaves room to set the type large enough to survive platform compression.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">Should every word in a lyric video be animated?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">No. Animating every word spends the viewer’s attention on the words evenly, which means none of them stand out. Two or three large motion moments per song, placed at structural boundaries such as the chorus or a key line, read as intentional. Continuous animation reads as a template.</dd></div><div class="py-8 hairline first:border-t-0"><dt class="font-display text-xl md:text-2xl text-white leading-[1.2] tracking-tight mb-4">What colours work best for a lyric video?</dt><dd class="text-text-secondary leading-relaxed max-w-prose-narrow">The ones that match the song’s mood, chosen before any other visual decision. Melancholy tracks want desaturated blue-greys with warm white text. Club tracks want high-saturation complements such as magenta against cyan. Heartbreak sits well in amber and teal. Anger wants pure red on pure black with no gradient. Picking the brightest preset instead is the fastest route to a video that feels wrong without looking wrong.</dd></div></dl></section></div></section></article>]]></content>
  </entry>
</feed>
