Edit video the way you edit a document
Timelines were designed for film, where the footage is precious and the edit is the art. A product demo is neither. DemoRiff gives you a transcript instead: delete a sentence and the footage goes with it, reorder a paragraph and the cut follows, fix a word and the voice re-reads it.
So um the pricing engine runs on every keystroke, so you see the rate before you actually commit.
Trimmed 31% of words, cut 4 dead-air regions totalling 38s, and keyframed 7 zooms. The intro now reaches the first click in 6 seconds instead of 41.
- 11 min
- Median record-to-publish
- −31%
- Typical word count reduction
- < 90s
- Render for a 3-minute cut
- 0
- Timeline skills required
Text is the timeline
Every word maps to a frame range. Strike a sentence and the video loses exactly those frames, with the cut point placed on the nearest breath so it never sounds spliced.
Direction, not filters
Zooms land on clicks. Pans follow the cursor. Cuts fall on sentence boundaries. These are the choices an editor makes, applied automatically and overridable everywhere.
A co-editor that takes notes
“Tighten the intro,” “make it sound less salesy,” “add a callout when the rate appears.” Riff applies the change and shows you the diff before it commits.
Delete the sentence, not the frames
The moment your recording lands, it is aligned word-by-word against the video. Editing becomes reading. Select a rambling paragraph and press delete — the pipeline finds the nearest clean cut points, crossfades the room tone so the join is inaudible, and reflows every downstream asset. Change your mind and undo restores the frames exactly.
- Word-level alignment with sub-frame accuracy
- Automatic cut-point snapping to breaths and sentence ends
- Room-tone crossfade so joins are inaudible
- Full version history — every edit is reversible
So um the pricing engine runs on every keystroke, so you see the rate before you actually commit.
Trimmed 31% of words, cut 4 dead-air regions totalling 38s, and keyframed 7 zooms. The intro now reaches the first click in 6 seconds instead of 41.
Zooms keyframed to the moments that matter
The recorder captured which element you clicked, not just where the cursor was. That means the zoom can frame the button rather than the pixel — and stay framed while a panel animates open. Easing is tuned to feel like a camera operator rather than a slideshow transition, and every keyframe is a draggable object on the timeline if you disagree.
- Element-aware framing, not coordinate-based
- Motion-tracked pan that follows animated UI
- Cursor smoothing with configurable weight and size
- Per-keyframe override of scale, easing and hold duration
Studio Sound turns a laptop mic into a booth
Room reverb, HVAC hum, keyboard clatter and laptop fan are separated from your voice and removed, then the result is EQ'd and levelled to broadcast loudness. If you would rather not use your own audio at all, swap in one of 260 studio voices or your own clone — the script stays, the delivery changes.
- De-reverb, de-noise, de-click and de-ess
- Loudness normalised to −16 LUFS for web, −14 for social
- Multi-speaker levelling across separate tracks
- One-click swap between recorded audio and synthetic voice
Say what you want changed
Riff sits beside the transcript and takes direction in plain language. It can rewrite for a different audience, cut to a target duration, add callouts where a concept is introduced, generate b-roll for an abstract idea, or produce three alternative intros for you to pick between. Every action is shown as a reviewable diff, and nothing is applied until you accept it.
- “Cut this to 90 seconds without losing the pricing part”
- “Rewrite for a technical buyer, drop the adjectives”
- “Add a callout the first time we say rate shopping”
- “Give me three intros and let me choose”
The technical detail
- Max export
- 3840×2160 at 60fps, H.264 / H.265 / ProRes 422
- Aspect ratios
- 16:9, 9:16, 1:1, 4:5, 21:9 and custom, auto-reframed
- Caption formats
- Burned-in, SRT, VTT, and styled soft captions
- Audio
- 48kHz stereo, −16 LUFS web / −14 LUFS social
- Max source length
- 4 hours per recording, 12 GB per file
- Render time
- Typically 0.4× realtime; 3-minute cut in under 90 seconds
We used to budget two weeks and eleven thousand dollars for a launch video. Our last three launches were recorded on a Tuesday and live on Wednesday. The agency invoice line item is gone.
About ai video studio
Yes. The transcript is the primary surface, but a full timeline sits underneath it with video, audio, caption, overlay and music tracks. Anything the AI decided is a real, editable object there — zoom keyframes, cut points, callouts, b-roll placements — so you are never stuck with a choice you disagree with.
Translations are derived, not duplicated. Change the source script and every language shows as out of date with a one-click re-render. You can also pin a language so a manually corrected translation is never overwritten.
Yes. Speakers are separated into their own tracks, labelled, and levelled independently. You can cut one speaker's tangent without touching the other, and assign different voices or leave one person's original audio intact.
Pairs well with
In production at