Record your screen once, talking as you go. VideoPolisher removes the dead time, rewrites and re-voices the narration, keeps the voice in step with what is on screen, and hands back a YouTube-ready video plus vertical shorts, captions, title and tags.
You approve every word of the script before anything is rendered. No re-recording, no timeline editing.
Long pauses while the page loads, the sentence you started twice, the click you had to redo. Cutting it by hand takes longer than recording it did.
A better script means a second take, a third take, and a room that is quiet enough. The screen part was fine. Only the voice needed work.
YouTube wants 16:9. Shorts, TikTok and Reels want 9:16 with captions. LinkedIn wants a one-minute highlight. That is three more edits of the same thing.
Subtitles, a title that people search for, a description, tags, a thumbnail, an intro and an outro. Each one small, together an afternoon.
Pick what you want, see what will run, approve the script, download. No timeline, no keyframes.
Drop a screen recording with your voice. Pick a recipe such as Polished demo or Demo + vertical shorts, or tick the outputs yourself. Mark cuts, an intro or outro to keep, or a region to crop if you like. We show exactly what will run and what it costs before you commit.
We transcribe what you said and rewrite it into clean narration that follows the order of events on screen. You read it, edit any line, and approve. Nothing is voiced or rendered until you do. Want to keep your own words? Approve as is.
A natural voice reads the approved script. The video is re-timed so the screen keeps pace with the narration, dead time disappears, and every format you picked renders. Download, or publish straight to YouTube.
Tick the outputs you want. Each one is a real file, versioned, re-made only when something it depends on changed.
Clean narration, dead time removed, ready for YouTube. With an aligned transcript.
9:16 with an AI camera that follows the cursor and the action, captions burned in.
Up to five 30–60 second highlights picked from the narration, vertical or widescreen.
Cinematic camera moves on the widescreen video, so a full-screen app reads on a phone.
Selective karaoke-style subtitles that highlight the words that matter, plus an SRT file.
Search-first YouTube metadata written from the script, with a generated thumbnail.
Keep the branded intro you already have, or generate one from the script, in every format.
Just the rewritten narration, for when you want to record it yourself.
Every step is a separate piece of work. Change one setting and only the steps that depend on it run again.
Normalises the frame rate, scales or letterboxes to a standard size, crops to a region you draw, and removes dead time. Cuts are detected by AI, or you mark them on a timeline and they are used exactly as drawn.
Transcribes your voice with word timings and picks the key frames where the screen changed. Both feed the script and the sync.
Rewrites the transcript into narration that follows the on-screen order, keeps your technical terms, and marks each sentence to the moment it belongs to. Keeps the spoken language or translates to another. You edit and approve.
Natural voices from OpenAI or ElevenLabs, with pronunciation overrides for acronyms. The video is re-timed scene by scene so the narration lands on the action, speeding through dead stretches and never cutting a scene in half. Prefer the original length? Pauses are inserted instead.
The polished 16:9 video, a transcript aligned to the final cut, SRT captions, YouTube metadata and a thumbnail. GPU encoding when available.
Reframes to 1080×1920 with a camera that follows the cursor or interprets the scene, stays inside the content, avoids black bars, and respects the safe zone under Shorts and TikTok overlays.
An AI-driven camera on the 16:9 output: zooms toward what the narration talks about, pans smoothly, snaps to edges, and holds still when nothing moves.
Karaoke-style captions on the phrases that matter, in a font that fits the language. Then 30–60 second highlights with their own titles, for Shorts and for LinkedIn.
Mark an intro or outro to keep and it is passed through untouched and re-attached in every format. Or generate a branded intro and recap outro from the script. Splice other clips in at chosen times, loudness-matched.
Detects and tracks the webcam bubble in a recording, covers the original with a card or logo so the camera can move freely, and places a clean still or your own headshot where you want it.
Connect YouTube once and publish any output with its metadata. Titles, tags and descriptions are checked against YouTube's limits before upload.
Every output remembers what it was made from. Change a voice, a caption style or one line of script, and only the steps downstream of that change run again. Your approved script is never regenerated behind your back.
The same engine, pointed at audio and at your website.
Credits per minute of output. Free credits when you sign up, no card needed to try it.
No. The screen recording is the source of truth. Only the narration is rewritten and re-voiced, and the video is re-timed to match it. If you would rather use your own voice, take the script and record it yourself.
It rewrites for clarity and keeps the order of what happens on screen. It does not invent steps. You read the whole script and approve or edit every line before it is voiced.
Transcription detects the spoken language. Narration keeps that language by default or translates to another, with captions in a font that fits the script.
Screen recordings in common formats such as MP4 and MOV, with your voice on the track. For podcasts, an audio file, a script, or an article URL. For a website tour, the site's address.
Text-to-speech voices from OpenAI and ElevenLabs, chosen per project. Pronunciations of product names and acronyms can be pinned.
Uploads are stored in our own cloud storage in the US and processed by AI providers to transcribe, write and narrate. Your files are yours, exportable and deletable on request. The privacy policy lists every provider we use.