如何用 Claude 制作卡通解说视频

一句主提示词加四次批准,Claude 独立完成脚本、配音、作画、剪辑,产出 78 秒竖屏卡通讲解视频。

AI GUIDES(@FREE_AI_GUIDES)教程 · 已翻译 · 约 40 分钟
阅览室 · AI 最佳实践#视频#Claude#脚本#卡通#剪辑原文

我只用一句提示词和四次确认,就做出了一支78秒的卡通视频。

Claude 写了脚本、录了配音、画了三个卡通角色、拆成28个镜头,还配好字幕剪辑成片。我全部的活儿,就是在每个阶段说一句行或不行。

我从没打开过剪辑软件。Claude 写的代码我一行都没读过。

这份指南会给你完整的搭建方法、主提示词,以及每一步可以直接复制粘贴的提示词。你不需要任何剪辑技能。

内容包括:

→ 你要做的是什么,以及你需要的4样东西

→ 主提示词,复制即用

→ 从脚本到成片的4个步骤,每一步都配一条提示词

→ 5个会让这类视频翻车的错误(附各自的解决办法)

→ 什么情况下该跳过这套流程,改做别的类型视频

→ 给你的第一支视频准备的快速上手版

你要做的是什么

一支竖屏讲解视频,1080×1920像素,手机屏幕的形状。时长75到90秒,讲清楚一个观点。

你自己的卡通角色用表情和动作把故事演出来。它们从不说话。所有台词由旁白承担,字幕在底部滚动。

我的测试阵容是三颗卡通豆子。Mo 是那个什么新工具都要试的热情朋友。Pax 是那个要证据的怀疑派。Vee 是那个负责讲解解决办法的测评人。

在底层,Claude 用代码把整支视频搭出来。你的角色也是用代码画出来的,所以每个镜头里它们都长得一模一样。这些你全程都看不到。你只需要读脚本、听配音、看画面。

最后你会拿到三个文件。

  • 视频,MP4格式
  • 字幕文件(SRT),上传时和视频放在一起
  • 一个包含分镜表、角色设定表和声音方案的页面

顺序永远不变。先脚本,再配音,最后视频。改动的代价,在脚本阶段是一条消息,到了视频阶段就是整支重做。所以每个阶段你都要先确认,Claude 才会开始下一个。

你需要准备什么

  1. Claude,在一个能运行代码的对话中。我的测试视频用的是 Claude Opus 5.5。
  2. 一张角色设定图。 一张包含你角色四种表情(平静、震惊、大笑、哭泣)的图片。Claude 会在每个镜头中参照它。
  3. 一份参考资料。 一张信息图、一篇文章,或者你自己关于该主题的笔记。
  4. 主提示词。 一段包含视频所有规则的长提示词。从下面的代码块中复制它。

把整个代码块粘贴到一个新的 Claude 对话中。要使用你自己的角色,请重写 visual_identity 和 cast 部分,并附上你自己的角色设定图。其余内容保持原样即可。你可以跳过 voiceover 和 render 部分,因为那些是给 Claude 的设置说明。

<title>
Opus 5.5 Motion Design Prompting Guide: Vertical Dynamic Cut, Bean Buddies
</title>

<purpose>
Create a complete 75-90 second vertical educational animation about [TOPIC].
Teach [CORE IDEA] through one clear visual story, not a list of facts.
Make the difficult idea intuitive for [AUDIENCE] without losing accuracy.
Design for mobile viewing: fast pacing, one idea per shot, readable at phone size.
The Bean Buddies act the story out. A narrator carries all the words.
</purpose>

<source_material>
[Optional. Attach or describe the reference. List the components, steps, or claims
the video must cover.]
Teach only what the source supports. Do not add components or facts it doesn't contain.
Leave out claims the source can't back up (exact multipliers, guaranteed results).
If the topic involves current products, versions, or numbers, verify them before writing.
Give each major component of the source its own location in the world.
</source_material>

<visual_identity>
The Bean Buddies world: a clean, soft, clay-toy look drawn as flat vector shapes.
Palette:
- Ivory #FBF5EA for bean bodies, cards, and nodes
- Clay #D97757 for action, arrows, and highlights
- Slate #4A4744 for commands, evaluators, and dark surfaces
- Sage #A7B99E for context, success, and calm moments
- Mustard #E8B54E for glows, tokens, and attention
- Warm-brown outline #3D2B24 on every shape
- Cream #F5EEE3 as the base background, blush #F4C0B4 for cheeks
Every shape has a 6 px warm-brown outline, rounded corners, and a soft offset shadow.
No sharp corners and no black outlines.
Each location has its own flat background color from the palette, a faint dot grid,
and a floor band with a soft ground shadow under each character.
Diagrams are chunky rounded nodes, clay arrows with outlined heads,
and pill-shaped chips for commands and labels.
Characters bob gently when idle, and squash and bounce when they walk or land.
Never show real product logos or real app icons. Invent simple icons instead.
</visual_identity>

<cast>
Match the attached model sheet exactly. Every bean has a soft ivory rounded bean body
with a beige shading strip on its right side, pink blush cheeks, tiny rounded arms,
and two small oval feet.

MO, the hype friend. Clay hoodie over the lower body with drawstrings and a front pocket,
clay headphones over the top of the head, big glossy eyes with highlights,
open happy mouth, slate feet. Loves every new thing, acts first and thinks later.

PAX, the skeptic. Slate ribbed beanie with a lighter slate band, heavy flat brows,
half-lidded eyes, small smirk, sage scarf, arms crossed by default, slate feet.
Wants proof. His reactions mark the problem.

VEE, the reviewer. Sage bucket hat, dark round sunglasses, open smile,
clay scarf with a hanging tail, holds a slate microphone, clay feet.
Explains and checks. He guides the viewer through the mechanism.

[ROLE MAPPING. Optional. Say which character carries which part of this story.]

Expressions: neutral, shocked (hands to cheeks, mouth open, and Vee's sunglasses
jump up onto his hat brim), laughing (eyes closed in arcs, wide open mouth),
and crying (tears on the outer cheeks, blush removed).
Change expression only on a cut or a clear beat, never mid-word.
The characters act and react. They never speak. No speech bubbles with dialogue.
</cast>

<format>
Aspect ratio 9:16, 1080x1920, 24 fps, with burned-in captions.
Safe zones: keep all key action and text inside the central area.
Leave the top 250 px and bottom 420 px free of essential content
(platform UI overlays sit there).
Compose in a single vertical column. Stack information top to bottom,
never side by side. Use vertical motion and depth (push-ins, drops, rises)
as the main movement axes.
Character framing: one character per shot by default. Use close-ups on faces
for reactions. When two characters share a shot, stack them in depth
(one in front, one behind) or put one character in the lower third under a diagram.
Show all three side by side only in the final shot.
Labels: minimum 56 px equivalent, maximum 4 words, set in pill chips.
Captions: centered, in the lower-middle area just above the bottom safe zone,
2 lines maximum.
</format>

<process>
Reduce the subject to one central cause-and-effect relationship.
Structure the story as seven scenes: hook, familiar world, disruption,
mechanism, discovery, consequence, and recap. Every scene must add new understanding.

Break every scene into 3-5 distinct shots.
Target a new shot every 2-4 seconds, roughly 25-35 shots in total.
A shot is a new composition: a new background, a new camera distance,
or a new location. Moving elements inside the same frame does not count.
Cut to a character reaction at least once per scene.

Shot rules:
- Every shot starts from a clean composition. Do not keep previous
  elements on screen unless one is deliberately carried forward.
- At most 3 primary elements on screen at any time. A character counts as one.
- When an element has done its job, it exits. Nothing accumulates.
- Vary camera distance across consecutive shots (wide, medium, close-up, macro).
- Only one hero object carries across each transition. Everything else resets.

Transitions must change the frame while preserving meaning. Use:
match cuts (a shape in one shot becomes a similar shape in the next),
zoom-throughs (push into an object to reveal a new world inside it),
whip pans and vertical drops or rises to a new location,
a character walking, hopping, or tumbling out of one frame and into the next,
and cuts on a change of expression.
Do not morph the whole frame in place. Do not add new elements
to an existing frame as a substitute for a new shot.

Let the most important transformation per scene hold for 2-3 seconds
so it can breathe. Everything else moves briskly.
Write warm, precise narration. Use concrete language before technical terms.
Do not duplicate narration in on-screen text. Keep labels brief and readable.
</process>

<workflow>
Work in three stages. Stop at the end of each of the first two stages
and wait for my approval. Do not start the next stage until I approve.

STAGE 1 · SCRIPT
Deliver the script as a .md file containing:
- A header with format, style, cast roles, core idea, audience, source, pacing,
  and total word count.
- For every scene: SCENE [N] - [START-END], a word count, the learning purpose,
  and the final voiceover lines.
- For every shot within it, a table row with shot number, framing,
  what is on screen (max 3 primary elements, with each character's expression),
  and the cut into the next shot.
- A short sound plan and a list of voiceover risks (hard words, fast lists).
Target 190-215 words total so the read lands inside 75-90 seconds.
Timestamps are estimates until the voiceover is recorded.
Then stop and wait for approval.

STAGE 2 · VOICEOVER
Record the approved script exactly as written (see <voiceover>).
Report the total length and the words or lines most likely to sound wrong.
If the length falls outside 75-90 seconds, adjust pauses first, never words.
Then stop and wait for approval.
When I flag a problem, re-record only the flagged sentence,
slow only that line if needed, and leave the rest of the track untouched.

STAGE 3 · VIDEO
Time every shot to the approved recording's sentence timings.
Build and render the complete animation (see <render>).
Do not stop at a mood board, style frame, or partial prototype.
</workflow>

<voiceover>
Engine: Kokoro via the kokoro-onnx Python package.
1. Install: pip install kokoro-onnx soundfile --break-system-packages
2. Download the full-precision model files (the int8 model drops audio on some sentences):
   https://github.com/thewh1teagle/kokoro-onnx/releases/download/model-files-v1.0/kokoro-v1.0.onnx
   https://github.com/thewh1teagle/kokoro-onnx/releases/download/model-files-v1.0/voices-v1.0.bin
3. Voice am_michael, language en-us, sample rate 24000.
4. Render every sentence as its own clip. Trim leading and trailing silence,
   keeping 10 ms before and 100 ms after the speech. Join the clips in order.
5. Speed: base 0.95 per sentence times a global factor of 1.25.
   Never exceed 1.25 on normal lines.
6. Silence: 0.8 s lead-in, 0.36 s between sentences in a scene,
   0.56 s between scenes, 0.8 s tail.
7. Export: normalise to -16 LUFS with a -1.5 dB true peak, MP3 at 128 kbps, 44.1 kHz.

Script rules for clean TTS:
- Never use colons in voiceover lines.
- Spell out acronyms the way they should be spoken, or replace them with plain words.
- Give each item in a list its own sentence when the list has more than three items.
</voiceover>

<render>
Build the animation as one HTML canvas at 1080x1920.
Draw the characters in code from the model sheet so they stay identical in every shot.
Every element is a function of time computed from the recording's sentence timings,
so any frame can be drawn in any order.
Embed all fonts. No external assets.
Render frames headlessly with Playwright and Chromium, then encode with ffmpeg:
H.264 at 24 fps, AAC audio, the voiceover mixed over a ducked music bed and effects.
Normalise the final mix to -16 LUFS with a -1.5 dB true peak.
Burn in the captions and deliver a matching SRT file.
Review sampled frames against the storyboard and fix any shot that stacks elements,
breaks the safe zones, lets text overflow its box, or drifts off the model sheet
before delivering.
Deliver the MP4, the SRT, and a published page with the video, storyboard,
character sheet, transition map, sound plan, and production notes.
</render>

<voice>
Act as a senior motion director, educational storyteller, character animator,
and short-form editor.
Be playful in presentation, rigorous about clarity, and decisive in execution.
Cut like a short-form editor: keep momentum high and the frame always changing.
Let the characters' reactions carry the emotion so the narration can stay clear.
Preserve accuracy, character consistency, and the supplied identity.
</voice>

还没有角色?先把这段内容粘贴给 Claude。它会为你设计一套角色阵容,并画出角色设定图。

Design an original cast of [NUMBER] cartoon characters for short explainer
videos about [YOUR NICHE].
Give each one a name, a one-line personality and a role: one who plays the
viewer, one who doubts, one who explains.
Keep them simple enough to redraw the same way every time: one body shape,
three or four colors, and one signature item each.
No existing characters, mascots, logos or real people.
Then draw a model sheet showing every character in four expressions
(neutral, shocked, laughing, crying) on a plain background.

保留这三个角色。观众角色犯下错误,质疑者角色对问题作出反应,讲解者角色展示解决办法。情绪由他们的面部来承载,这样旁白就可以保持冷静清晰。

第 1 步:脚本

打开一个新的 Claude 对话,粘贴主提示词。附上你的角色设定图和你的参考图。然后在点击发送之前,填写提示词中的这五个字段。

[TOPIC] = [Your subject in a few words]

[CORE IDEA] = [The one cause-and-effect idea the viewer should walk away with]

[AUDIENCE] = [Who is watching and what they already know]

<source_material> = [Attach your infographic or article, or paste your notes.
List the parts the video must cover.]

[ROLE MAPPING] (optional) = [Which character carries which part of the story]

时间不够?把字段留空,附上参考图。Claude 会根据它来填写。

核心想法是决定这条视频的字段。Claude 会围绕一个因果关系构建整个故事。我的是“当 AI 是一套技术栈,而不是一堆工具时,它才能节省时间”。

Claude 会发回一个脚本文件,然后停下。脚本始终有七个场景,按以下顺序排列。

  1. 钩子。 观众的问题,或者他们本就有的疑问。
  2. 熟悉的世界。 观众目前的情况是如何运作的。
  3. 打破。 什么出了问题。
  4. 机制。 这个修复是如何起作用的。
  5. 发现。 那个“哦,原来如此”的时刻。
  6. 后果。 观众会有什么变化。
  7. 回顾。 一口气说清这个想法。

每个场景都配有旁白和镜头表。表格列出了取景、屏幕上显示的内容、每个角色的面部表情,以及切入下一个镜头的剪辑点。

我的脚本最终是 195 个词、7 个场景、28 个镜头。下面是钩子。

你这个月已经试了五个新的 AI 工具。那为什么每项任务还是花一样长的时间?

这句台词覆盖三个镜头,大约 7 秒。Mo 在五个应用图块中间大笑,然后在一张任务卡旁边等待,同时一根时钟指针扫过一整圈,接着当第二根时钟也扫过一整圈时,他的脸沉了下来。

在批准之前,先核对这些数字。

  • 190 到 215 个词。 这个长度在朗读出来后会落在 75 到 90 秒之间。
  • 25 到 35 个镜头, 每 2 到 4 秒就有一个新镜头。
  • 每个场景有一次 2 到 3 秒的停留, 放在它最重要的时刻。

然后核对那些说法。我的信息图写了“聪明 10 倍”。Claude 删掉了那行,因为来源里没有任何东西能证明它。信息图还提到了真实的应用,所以视频改用虚构的图标,比如聊天气泡和画笔。

想要第二双眼睛帮你看?把这段粘贴到同一个对话里。

Audit this script. Report pass or fail for each line, and quote every fail:
- 190-215 words in total
- Exactly 7 scenes: hook, familiar world, disruption, mechanism, discovery,
consequence, recap
- Opens on the viewer's problem or a question, not on a character
- Speaks to the viewer as "you"
- No character named or introduced in the voiceover
- No colons in voiceover lines
- Acronyms spelled out or replaced
- Lists of more than three items get one sentence per item
Fix only the fails and show me the changed lines.

第 2 步:配音

Claude 用 Kokoro 录制旁白,这是一个免费开源的文本转语音模型。主提示词已经设定好了音色、语速和停顿。你只需发送一条消息。

Script approved. Record the voiceover exactly as written.
Report the total length, each sentence's speaking rate in syllables per second,
and the words most likely to sound wrong.

Claude 将每句话录制成单独的片段,再把这些片段拼接成一个 MP3。这样修改成本很低。当某一句听起来不对时,Claude 只需重录那一句,其余部分原封不动。

打开脚本对照着听,检查四件事。

  1. 时长。 应该落在 75 到 90 秒之间。
  2. 语速过快的句子。 正常语速大约是每秒 4.6 个音节。Claude 的报告会标出任何明显超过这个值的句子。
  3. 列表。 每一项都应该作为独立的节拍呈现。我的回顾脚本把“Prompts. Tools. Automation. Systems. Growth.”写成五个独立的句子,这样每个词都能与屏幕上亮起的方块对齐。
  4. 易错词。 缩写、数字和技术术语最容易出问题。在我的视频里,“AI”必须读成“A-I”而不是“eye”。

如果时长不对,用这个。

The read is [LENGTH] seconds. Bring it inside 75-90 seconds by widening
the pauses between sentences and between scenes. Do not change any words.

如果某一句听起来太赶,用这个。

Re-record only this sentence: "[SENTENCE]"
Problem: [rushed / mispronounced word / dragging]
Record it alone at 1.1x and leave the rest of the track untouched.

在我的视频里,“You don't build it in a day.”这句读得太快了。Claude 以 1.1 倍速重录了这一句,其余 29 句保持不变。

当你确认通过后,Claude 会把真实的时间轴写回脚本。视频将基于这些时间轴来构建,所以这是修正配音的最后一次低成本机会。

第 3 步:视频

Claude 会把每个镜头对准你录音中的一句话,所以声音走到哪,画面就跟到哪。在它开始构建单个镜头之前,先要求看角色。

Voiceover approved. Build the video.
Before any shots, show me a test sheet of every character in every expression,
next to the model sheet. Wait for my OK.
Then time every shot to the recorded sentences.
Keep a fix log with a before and after frame for every problem you correct.

把测试表和你的模型表对照一下。一顶帽子错了或围巾少了,在这里只需一条消息就能修好。等到后面,你得在 30 个镜头里逐个修。

主提示词包含五条布局规则,每一条都有其理由。

  1. 顶部 250 像素和底部 420 像素保持空白。 在手机上,应用的按钮和字幕会覆盖这些区域。
  2. 所有内容自上而下堆叠。 手机屏幕又高又窄,所以并排的元素会缩小到没人能看清。
  3. 每个镜头最多展示 3 个主要事物。 一个角色算作一个。
  4. 每个镜头都从干净状态开始。 当某个东西完成了它的任务,它就离开画面。
  5. 每次切换都有一个物体贯穿。 它把两个镜头连在一起,其余一切重新开始。

我在测试视频中最喜欢的切换,是在镜头 5.1 结尾停在 Pax 扬起的眉毛上。在镜头 5.2 中,那根眉毛变成了柜子把手。你的视线停留在同一个位置,而整个场景已经改变。

在最终渲染之前,要求提供手机尺寸的样帧。

Show me sampled frames from every scene, including one mid-transition frame
per scene, at phone size. Mark the safe zones on each frame.

如果某个镜头看起来不对,就单独修那个镜头。

Fix shot [NUMBER]. Problem: [what's wrong].
Leave every other shot as it is.

我的最终剪辑是 77 秒的语音,加上最后一帧停留 1.5 秒,总共 78.5 秒。

第 4 步:手机检查

这一步大约需要一分钟。在手机上观看视频,再关掉声音看一遍,然后闭上眼睛听一遍。

你可以先让 Claude 运行检查。

Before I publish, audit the finished video against this list and report
pass or fail for each line, with the shot number for every fail:

1. Nothing important in the top 250 px or bottom 420 px
2. No more than 3 main elements in any shot
3. No prop covering a character's face
4. Every character clearly visible against its background
5. No text overflowing its box, no label over 4 words
6. Characters match the model sheet in every shot
7. No character visible while moving between scenes
8. No real brand names or logos
9. Captions match the voice, 8 words or fewer per chunk, 2 lines at most
10. Total length inside 75-90 seconds

Show me a phone-sized frame for each fail.

关掉声音时,字幕应当以每块不超过 8 个词的方式承载整个故事。闭上眼睛时,每一句话里的人声都应当盖过音乐。

当每一行都通过后,发布 MP4,并上传 SRT 文件作为其字幕。

会毁掉这类视频的 5 个错误

每一个错误都会让你重做一遍,或者让观众直接划走。只要早点发现,每个错误修起来都用不了一分钟。

错误 1:开场就介绍角色。

我的第一版脚本一上来就介绍那三颗豆子。一个随手划过的观众根本不认识 Mo,也没有理由为他停下来。

修复:开场就切入观众的问题,或者他们本来就有的疑问。Claude 把我那句改成了“这个月你已经试过五款新的 AI 工具了。”如果你的钩子感觉平淡,就让它多给几个选项。

Write 5 opening lines for a video about [TOPIC].
Each one speaks to the viewer as "you" and names a problem they already feel
or asks a question they already have.
No character names, no colons, 20 words or fewer each.
Tell me which one you'd pick and why, in one sentence.

错误 2:用加词来凑时长。

我的第一版配音只有 69.8 秒,低于 75 秒的下限。加一句台词就意味着要重新录一遍,还要给后面每一个镜头重新对时间。

修复:改停顿,一个字都别动。Claude 把句与句之间的间隔拉宽到 0.55 秒、场景之间拉到 1 秒,配音就落到了 77 秒。如果脚本太长,就在第 1 步删词,趁什么都还没录。

错误 3:把整条音轨重录一遍。

有一句念得太赶,就会诱使你去要一版全新的录音。整条重录会挪动每一句话的时间点,包括那些本来就念对的。

修复:只重录被标记的那一句,如果听起来赶,就用 1.1 倍速。

错误 4:往已经拥挤的画面里硬塞更多东西。

当一个镜头显得很满时,本能反应是把文字缩小,好让所有东西都塞得下。画面里不断堆东西,会让这类视频看起来像幻灯片。

修复:让 Claude 给你一个新镜头。一个画面一个想法,最多 3 个主要元素。

错误 5:在笔记本上检查。

笔记本上显示的每一帧,都比观众实际看到的更大。Claude 自己做的画面检查抓出了三个只有在手机尺寸下才会暴露的问题。

  • Mo 消失了。他的黏土连帽衫在珊瑚色背景前不见了踪影。
  • 道具挡住了脸。角色在胸口高度拿着的任何东西,在手机上都会遮住脸。
  • 一个带罐子、一摞东西、一张卡片和一支箭头的循环图,显得太拥挤了。

解决办法:发布前先在真手机上看看。把手里拿的东西举到头顶上方,并给每个角色配一个与衣服形成对比的背景色。

什么时候该跳过它

这套工作流只做一种视频。以下情况我会另选他法。

→ 一份技巧清单。主提示词围绕一个因果关系来构建故事。十条零散的技巧更适合做成轮播图或推文串。

→ 一个真实应用的演示。规则禁止真实标志和应用图标,所以录屏能更好地展示产品。

→ 任何超过 90 秒的内容。提示词里的每个数字(字数、镜头数、停顿)都是为 75 到 90 秒调校的。

→ 你十分钟内就要用的视频。三轮审批比这还慢。把这套流程留给那些你会重复使用或多次发布的视频吧。

快速上手版

  1. 挑一个你经常向别人解释的概念,找出你已经写好的相关文章或笔记。
  2. 把主提示词粘贴到一个新的 Claude 对话中,附上你的模型表和参考素材,然后填写那五个字段。
  3. 检查脚本是否在 190 到 215 词之间、是否有 7 个场景,以及开头钩子是否直击观众的问题。通过。
  4. 打开脚本,听一遍配音。逐条用停顿和赶工的句子来修正时长。通过。
  5. 检查角色设定表,然后在手机尺寸下查看样帧。通过制作。
  6. 在手机上观看成片,先静音看一遍,再开声音看一遍,然后连同字幕文件一起发布。

你的第一条视频需要四次审核通过。第二条则从你已经信任的演员阵容和提示词开始。

成品展示

这是我用这套完整流程做出来的一条视频。Bean Buddies 讲解了三种 AI 循环,从一条提示词到按自己的时间表运行的循环。

时长 84 秒。Claude 写了脚本、录制了配音,并为每一个镜头做了动画。我的工作就是审核每一个阶段。

先带声音看一遍,再静音看一遍。字幕应该能承载整个故事。

把这条存下来,留给你的第一条视频用。随着工具更新,我会持续维护它。

如果你觉得有用,看看下面的我的 newsletter

我每周分享一个 AI 超能力

订阅吧,免费的 👇

https://aisuperpowers.co

延伸阅读

  • Kokoro ONNX,Claude 用于旁白的开源语音模型
  • Playwright,Claude 用于渲染每一帧的浏览器工具
  • FFmpeg,Claude 用于将画面帧与语音合成为 MP4 的工具
已读完 · 本文由熊猫易读翻译重排