AI Video Studio

Edit video the way you edit a document

Timelines were designed for film, where the footage is precious and the edit is the art. A product demo is neither. DemoRiff gives you a transcript instead: delete a sentence and the footage goes with it, reorder a paragraph and the cut follows, fix a word and the voice re-reads it.

STUDIO · SCRIPT−31% words
Remove fillersTightenFormal toneShorten 20%

So um the pricing engine runs on every keystroke, so you see the rate before you actually commit.

Riff

Trimmed 31% of words, cut 4 dead-air regions totalling 38s, and keyframed 7 zooms. The intro now reaches the first click in 6 seconds instead of 41.

11 min
Median record-to-publish
−31%
Typical word count reduction
< 90s
Render for a 3-minute cut
0
Timeline skills required

Text is the timeline

Every word maps to a frame range. Strike a sentence and the video loses exactly those frames, with the cut point placed on the nearest breath so it never sounds spliced.

Direction, not filters

Zooms land on clicks. Pans follow the cursor. Cuts fall on sentence boundaries. These are the choices an editor makes, applied automatically and overridable everywhere.

A co-editor that takes notes

“Tighten the intro,” “make it sound less salesy,” “add a callout when the rate appears.” Riff applies the change and shows you the diff before it commits.

Transcript editing

Delete the sentence, not the frames

The moment your recording lands, it is aligned word-by-word against the video. Editing becomes reading. Select a rambling paragraph and press delete — the pipeline finds the nearest clean cut points, crossfades the room tone so the join is inaudible, and reflows every downstream asset. Change your mind and undo restores the frames exactly.

  • Word-level alignment with sub-frame accuracy
  • Automatic cut-point snapping to breaths and sentence ends
  • Room-tone crossfade so joins are inaudible
  • Full version history — every edit is reversible
STUDIO · SCRIPT−31% words
Remove fillersTightenFormal toneShorten 20%

So um the pricing engine runs on every keystroke, so you see the rate before you actually commit.

Riff

Trimmed 31% of words, cut 4 dead-air regions totalling 38s, and keyframed 7 zooms. The intro now reaches the first click in 6 seconds instead of 41.

Automatic direction

Zooms keyframed to the moments that matter

The recorder captured which element you clicked, not just where the cursor was. That means the zoom can frame the button rather than the pixel — and stay framed while a panel animates open. Easing is tuned to feel like a camera operator rather than a slideshow transition, and every keyframe is a draggable object on the timeline if you disagree.

  • Element-aware framing, not coordinate-based
  • Motion-tracked pan that follows animated UI
  • Cursor smoothing with configurable weight and size
  • Per-keyframe override of scale, easing and hold duration
DEMO.MP4 · 3840×2160
In transit
1,284
+12%
On-time rate
96.4%
+2.1%
Avg. dwell
3.2h
−18m
At risk
23
+4
Shipment
Route
ETA
Status
SHP-4821
Rotterdam → Chicago
2 days
On time
SHP-4822
Shenzhen → Long Beach
9 days
Delayed
SHP-4823
Hamburg → Newark
4 days
On time
SHP-4824
Busan → Oakland
12 days
At risk
Meridian
The route prices instantly.
0:41 / 1:48
Audio

Studio Sound turns a laptop mic into a booth

Room reverb, HVAC hum, keyboard clatter and laptop fan are separated from your voice and removed, then the result is EQ'd and levelled to broadcast loudness. If you would rather not use your own audio at all, swap in one of 260 studio voices or your own clone — the script stays, the delivery changes.

  • De-reverb, de-noise, de-click and de-ess
  • Loudness normalised to −16 LUFS for web, −14 for social
  • Multi-speaker levelling across separate tracks
  • One-click swap between recorded audio and synthetic voice
VOICE LIBRARY · 260 VOICES
AveryWarm · US · Narration
RhysCrisp · UK · Technical
NoorBright · AU · Product
Your voiceCloned · 0:31 sampleCONSENTED
Riff co-editor

Say what you want changed

Riff sits beside the transcript and takes direction in plain language. It can rewrite for a different audience, cut to a target duration, add callouts where a concept is introduced, generate b-roll for an abstract idea, or produce three alternative intros for you to pick between. Every action is shown as a reviewable diff, and nothing is applied until you accept it.

  • “Cut this to 90 seconds without losing the pricing part”
  • “Rewrite for a technical buyer, drop the adjectives”
  • “Add a callout the first time we say rate shopping”
  • “Give me three intros and let me choose”
RIFF AGENT
Make a 90-second demo of the new rate-shopping feature for logistics buyers.
Reading your changelog and the Linear issue. Drafting a 5-beat outline…
Recording against staging.meridian.app with the demo tenant. 11 interactions captured.
Cut to 1:34, 7 zooms, Avery voice, Meridian brand kit. Ready for review.
rate-shopping-demo.mp4 · 1:34 · awaiting your review
Specifications

The technical detail

Max export
3840×2160 at 60fps, H.264 / H.265 / ProRes 422
Aspect ratios
16:9, 9:16, 1:1, 4:5, 21:9 and custom, auto-reframed
Caption formats
Burned-in, SRT, VTT, and styled soft captions
Audio
48kHz stereo, −16 LUFS web / −14 LUFS social
Max source length
4 hours per recording, 12 GB per file
Render time
Typically 0.4× realtime; 3-minute cut in under 90 seconds
We used to budget two weeks and eleven thousand dollars for a launch video. Our last three launches were recorded on a Tuesday and live on Wednesday. The agency invoice line item is gone.
14 days → 1Launch video turnaround
PRPriya RamanVP Product Marketing · Northwind
Questions

About ai video studio

Yes. The transcript is the primary surface, but a full timeline sits underneath it with video, audio, caption, overlay and music tracks. Anything the AI decided is a real, editable object there — zoom keyframes, cut points, callouts, b-roll placements — so you are never stuck with a choice you disagree with.

Translations are derived, not duplicated. Change the source script and every language shows as out of date with a one-click re-render. You can also pin a language so a manually corrected translation is never overwritten.

Yes. Speakers are separated into their own tracks, labelled, and levelled independently. You can cut one speaker's tangent without touching the other, and assign different voices or leave one person's original audio intact.