How to use AI to write video scripts without losing your voice

The difference between a usable AI script and unusable filler is almost entirely what you paste in before you ask. Here is the input, the failure modes, and the edit.

8 minute read

This is a rewriting job, not a writing job

AI has moved into video production fast. Wistia's 2025 State of Video report, based on platform data plus a survey of more than 1,300 professionals, found 41% of brands using AI in video creation, up from 18% the year before. But the useful framing is narrow: a model is good at reorganising material you supply and bad at inventing material you have not. That single distinction decides whether your script comes back usable. Ask for a script about a topic and you get the average of every script ever written about it. Give it eight minutes of you talking through the same topic and ask it to find the argument, and you get your own thinking, ordered. This post covers that workflow; the underlying craft of structure and pacing is a separate guide on this site.

2025 report41%2024 report18%
Share of brands using AI in video creation, from Wistia's State of Video reports. The 2025 edition drew on platform data from 2024 plus a survey of 1,300-plus professionals run between 16 November 2024 and 10 January 2025. These are self-reported figures from Wistia's own business customers, so they track brands rather than independent creators, and 'using AI' covers everything from captions to generation.

Feed it a transcript of you talking, not a topic

The highest-leverage move in the whole workflow costs about five minutes. Before you prompt anything, record yourself talking through the subject as if a friend asked about it over coffee. Do not perform, do not edit, ramble. Run it through any transcription tool, paste the transcript in, and ask the model to pull out the argument you were actually making, the three strongest specifics, and the order those ideas should go in. What comes back is recognisably yours because every idea in it originated with you. The model is doing sorting, which is the thing it is genuinely reliable at. Compare that with a blank-prompt script and the gap is not subtle: one has your examples and your phrasing, the other has neither.

The other three things worth pasting in, and where to put them

Add your last three or four scripts as voice samples, the real constraint rather than a vague one, and who the viewer is in one concrete sentence. Real constraint means the runtime you actually need, the platform, and whether it will be watched muted. Position matters too: Nelson Liu and colleagues, in work published in Transactions of the ACL, showed that models retrieve information best when it sits at the start or the end of the input and measurably worse when it is buried in the middle. If your prompt is long, put the instruction at the top, the reference material in the body, and repeat the instruction at the bottom. It looks redundant. It is not.

Failure one: every number and every name

Assume any statistic, date, study, price or attributed quote in the draft is wrong until you have opened the source page yourself. This is structural rather than a bug in one product. In a September 2025 paper, OpenAI researchers Adam Kalai, Ofir Nachum, Santosh Vempala and Edwin Zhang argued that hallucination persists because standard training and evaluation reward guessing over admitting uncertainty — a model that says it does not know scores worse on the benchmarks the field optimises for than one that produces a confident plausible answer. So the confidence in the draft carries no information about correctness. The working rule for a script is blunt: if you cannot verify a number on the publisher's own page in under a minute, cut the sentence rather than soften it.

Failure two: the hook, the last line, and anything that has to be funny

These are the three places the model's default register costs you the most, and they are also the places it reaches for it hardest. You will get the you-will-not-believe opening, the nobody-is-talking-about-this opening, and a closing line that thanks people for watching. Every one of those is recognisable as filler within a second, which on short-form is most of the attention you were given. Jokes fail for a related reason: humour depends on a specific shared reference, and the model only has the general one. Write the first line and the last line yourself, then use the model the other way round — hand it your line and ask for fifteen variants that are shorter, more concrete, or lead with the number. Choosing between options is fast and low-risk. Accepting one is not.

Failure three: it cannot tell you how long the script runs

Ask for a 60-second script and you will get something that reads for 60 seconds at a pace nobody speaks at. Ask how long a draft runs and you will get a words-per-minute calculation dressed up as an estimate. Speaking pace varies enormously between people and between formats, and the model has no idea what yours is. The only reliable check is the obvious one: read the draft out loud, at the pace you would use on camera, with a timer running. Do this once and you will learn your own rate, which is worth more than any rule of thumb. It also surfaces the second problem, which is that AI drafts are written to be read rather than said, and clauses that look fine on screen fall apart in your mouth.

Keeping your voice is a measurable problem, not a vibe

There is evidence that heavy AI use flattens output toward a common middle. A working paper by Chaoran Liu, Tong Wang and S. Alex Yang, posted to SSRN in July 2025, used Italy's temporary ChatGPT ban as a natural experiment on Instagram marketing posts by Milan restaurants. During the ban, similarity between accounts fell — around 15% on lexical measures and 12% on syntactic ones — and average likes rose about 3.5%, while posting frequency and post length dropped. It is a working paper about restaurants, not creators, and the engagement effect is small. But the direction lines up with what you can hear in AI scripts. The practical counter is to supply samples of your own writing as the style reference, and to keep a short banned-phrase list for the tells you hate.

The edit, and what happens to the script afterwards

Edit in one pass with a clear target: cut the connective tissue, restore the specifics, and shorten every sentence you would not say. AI drafts run long on transitions and short on detail, so the pass is mostly deletion plus putting back the name, the number and the moment the model smoothed away. Then stop treating it as a script. Compress it to a beat sheet and speak from that — reading an AI draft verbatim is the fastest way to sound like one. On disclosure: YouTube's policy lists production assistance such as generating a script, outline, title or thumbnail among the things that do not require the altered-or-synthetic content label, which is reserved for realistic content that could mislead. Once you have recorded, the mechanical work of pulling clips and burning captions is where automation like FrameOS genuinely saves hours.

FAQ

What is the best prompt for writing a video script with AI?

There is no magic wording. The input matters far more: paste a transcript of yourself talking through the topic, two or three of your previous scripts as voice samples, the actual runtime and platform, and one sentence on who the viewer is. Then ask it to find the argument and order the beats rather than to write a script. Put the instruction at the top of the prompt and repeat it at the bottom.

Will AI-written scripts hurt my YouTube channel?

Using AI to help write a script does not require disclosure. YouTube lists production assistance, including generating an outline, script, title or thumbnail, among the uses that do not need the altered-or-synthetic content label, which is for realistic content that could mislead viewers. What does cause problems is publishing unedited, generic output at volume, which YouTube's monetisation rules treat separately.

How do I stop AI scripts sounding generic?

Give it your own material to work from and edit aggressively. Supply samples of your writing as a style reference, write the first and last lines yourself, and put back the specific names, numbers and moments the model smoothed out. A working paper by Liu, Wang and Yang found Instagram marketing posts became measurably less similar to each other during Italy's ChatGPT ban, which is the homogenising effect you are editing against.

Can AI tell me how long my script will run?

No, and its estimates are words-per-minute arithmetic rather than measurement. Speaking pace differs wildly between people and formats. Read the draft aloud at your real on-camera pace with a timer. That gives you both an accurate runtime and an early warning about sentences that read fine but do not survive being said.

Sources

Related pages