How to use AI to write video scripts without losing your voice
The difference between a usable AI script and unusable filler is almost entirely what you paste in before you ask. Here is the input, the failure modes, and the edit.
8 minute read
This is a rewriting job, not a writing job
AI has moved into video production fast. Wistia's 2025 State of Video report, based on platform data plus a survey of more than 1,300 professionals, found 41% of brands using AI in video creation, up from 18% the year before. But the useful framing is narrow: a model is good at reorganising material you supply and bad at inventing material you have not. That single distinction decides whether your script comes back usable. Ask for a script about a topic and you get the average of every script ever written about it. Give it eight minutes of you talking through the same topic and ask it to find the argument, and you get your own thinking, ordered. This post covers that workflow; the underlying craft of structure and pacing is a separate guide on this site.
Feed it a transcript of you talking, not a topic
The highest-leverage move in the whole workflow costs about five minutes. Before you prompt anything, record yourself talking through the subject as if a friend asked about it over coffee. Do not perform, do not edit, ramble. Run it through any transcription tool, paste the transcript in, and ask the model to pull out the argument you were actually making, the three strongest specifics, and the order those ideas should go in. What comes back is recognisably yours because every idea in it originated with you. The model is doing sorting, which is the thing it is genuinely reliable at. Compare that with a blank-prompt script and the gap is not subtle: one has your examples and your phrasing, the other has neither.
The other three things worth pasting in, and where to put them
Add your last three or four scripts as voice samples, the real constraint rather than a vague one, and who the viewer is in one concrete sentence. Real constraint means the runtime you actually need, the platform, and whether it will be watched muted. Position matters too: Nelson Liu and colleagues, in work published in Transactions of the ACL, showed that models retrieve information best when it sits at the start or the end of the input and measurably worse when it is buried in the middle. If your prompt is long, put the instruction at the top, the reference material in the body, and repeat the instruction at the bottom. It looks redundant. It is not.
Failure one: every number and every name
Assume any statistic, date, study, price or attributed quote in the draft is wrong until you have opened the source page yourself. This is structural rather than a bug in one product. In a September 2025 paper, OpenAI researchers Adam Kalai, Ofir Nachum, Santosh Vempala and Edwin Zhang argued that hallucination persists because standard training and evaluation reward guessing over admitting uncertainty — a model that says it does not know scores worse on the benchmarks the field optimises for than one that produces a confident plausible answer. So the confidence in the draft carries no information about correctness. The working rule for a script is blunt: if you cannot verify a number on the publisher's own page in under a minute, cut the sentence rather than soften it.
Failure two: the hook, the last line, and anything that has to be funny
These are the three places the model's default register costs you the most, and they are also the places it reaches for it hardest. You will get the you-will-not-believe opening, the nobody-is-talking-about-this opening, and a closing line that thanks people for watching. Every one of those is recognisable as filler within a second, which on short-form is most of the attention you were given. Jokes fail for a related reason: humour depends on a specific shared reference, and the model only has the general one. Write the first line and the last line yourself, then use the model the other way round — hand it your line and ask for fifteen variants that are shorter, more concrete, or lead with the number. Choosing between options is fast and low-risk. Accepting one is not.
Failure three: it cannot tell you how long the script runs
Ask for a 60-second script and you will get something that reads for 60 seconds at a pace nobody speaks at. Ask how long a draft runs and you will get a words-per-minute calculation dressed up as an estimate. Speaking pace varies enormously between people and between formats, and the model has no idea what yours is. The only reliable check is the obvious one: read the draft out loud, at the pace you would use on camera, with a timer running. Do this once and you will learn your own rate, which is worth more than any rule of thumb. It also surfaces the second problem, which is that AI drafts are written to be read rather than said, and clauses that look fine on screen fall apart in your mouth.
Keeping your voice is a measurable problem, not a vibe
There is evidence that heavy AI use flattens output toward a common middle. A working paper by Chaoran Liu, Tong Wang and S. Alex Yang, posted to SSRN in July 2025, used Italy's temporary ChatGPT ban as a natural experiment on Instagram marketing posts by Milan restaurants. During the ban, similarity between accounts fell — around 15% on lexical measures and 12% on syntactic ones — and average likes rose about 3.5%, while posting frequency and post length dropped. It is a working paper about restaurants, not creators, and the engagement effect is small. But the direction lines up with what you can hear in AI scripts. The practical counter is to supply samples of your own writing as the style reference, and to keep a short banned-phrase list for the tells you hate.
The edit, and what happens to the script afterwards
Edit in one pass with a clear target: cut the connective tissue, restore the specifics, and shorten every sentence you would not say. AI drafts run long on transitions and short on detail, so the pass is mostly deletion plus putting back the name, the number and the moment the model smoothed away. Then stop treating it as a script. Compress it to a beat sheet and speak from that — reading an AI draft verbatim is the fastest way to sound like one. On disclosure: YouTube's policy lists production assistance such as generating a script, outline, title or thumbnail among the things that do not require the altered-or-synthetic content label, which is reserved for realistic content that could mislead. Once you have recorded, the mechanical work of pulling clips and burning captions is where automation like FrameOS genuinely saves hours.
FAQ
What is the best prompt for writing a video script with AI?
There is no magic wording. The input matters far more: paste a transcript of yourself talking through the topic, two or three of your previous scripts as voice samples, the actual runtime and platform, and one sentence on who the viewer is. Then ask it to find the argument and order the beats rather than to write a script. Put the instruction at the top of the prompt and repeat it at the bottom.
Will AI-written scripts hurt my YouTube channel?
Using AI to help write a script does not require disclosure. YouTube lists production assistance, including generating an outline, script, title or thumbnail, among the uses that do not need the altered-or-synthetic content label, which is for realistic content that could mislead viewers. What does cause problems is publishing unedited, generic output at volume, which YouTube's monetisation rules treat separately.
How do I stop AI scripts sounding generic?
Give it your own material to work from and edit aggressively. Supply samples of your writing as a style reference, write the first and last lines yourself, and put back the specific names, numbers and moments the model smoothed out. A working paper by Liu, Wang and Yang found Instagram marketing posts became measurably less similar to each other during Italy's ChatGPT ban, which is the homogenising effect you are editing against.
Can AI tell me how long my script will run?
No, and its estimates are words-per-minute arithmetic rather than measurement. Speaking pace differs wildly between people and formats. Read the draft aloud at your real on-camera pace with a timer. That gives you both an accurate runtime and an early warning about sentences that read fine but do not survive being said.
Sources
- Wistia's 2025 State of Video Report Shows Use of AI in Video Production More than Doubled Over Last Year · PR Newswire / Wistia
- Why Language Models Hallucinate · arXiv / OpenAI
- Lost in the Middle: How Language Models Use Long Contexts · arXiv
- Generative AI and Content Homogenization: The Case of Digital Marketing · SSRN
- Disclosing use of altered or synthetic content - YouTube Help · YouTube Help
Related pages
Keep reading
More AI →The best AI tools for social media, by the job you need done
Most AI tool lists are directories. This one is organised by job, because the useful question is not which tool is best overall, it is which step of your week you want to stop doing by hand.
9 minute readThe best AI transcription tools, judged on measured accuracy
On clean read speech almost every engine looks excellent. On meetings, accents and crosstalk the same models lose several points of accuracy — and the published numbers show exactly where.
8 minute readThe best AI writing tools for creators, and what each one is actually for
Most roundups list forty tools that all do the same thing. The useful split is between general assistants, the editing layer, and tools that write from material you already have.
8 minute readTurn one long video into a week of posts
300 credits for 3 days · no card.
Start free