1实现原理 · 为什么它能做到
第一条不可谈判规则:一次只揭示一条消息、并且有节奏——先打字、再弹泡、再留阅读间隙,绝不第 0 帧倾倒整段对话。
1. **One message reveals at a time, on a rhythm.** The retention comes from the drip: a beat of typing, then a bubble pops in, then a pause to read. Never dump the whole thread at frame 0 — pace it like a real conversation (a typing indicator before received replies, a short read-gap after each bubble).
第二条不可谈判规则:chrome 必须在半秒内被认出是哪个 app(颜色/气泡/尾巴/回执/头部全对齐),不许混搭。
If a viewer can't tell which app it is in the first half-second, the illusion breaks.
数据模型优先:整个视频由一条 Message[] 驱动,剧情写成数据,渲染器只负责播放。
Everything is one array of messages. Author the story as data; the renderer just plays it.
时间线由数组游标推导:走一遍消息累加 cursor,产出 typingStartMs / typingLenMs / enterMs,所有弹簧、音效、滚动都读它。
const durationMs = cursor + 1000; // tail pad so the last bubble is readable
时长交给 calculateMetadata 从游标算,禁止手写帧数。
Compute `durationInFrames` from `durationMs` in `calculateMetadata` so a longer script makes a longer video automatically.
打字指示器只出现在收到的消息前,且用正弦呼吸、帧驱动(不用 CSS 动画计时器)。
Only received messages type; sent messages just pop (you don't watch yourself type).
自动滚动用弹簧逼近目标位移,让最新气泡落在视野下三分之一,而不是跳变。
Animate the scroll offset toward the target with a spring each time a message lands — it should glide, not jump.
每条消息排一个音效(发/收不同),且必须排进帧时间线而非事件回调。
Never trigger sound off a timer — schedule it on the frame timeline so it survives headless render.
竖屏安全区按四段留白定义(顶部状态栏 ~120px、头部 ~150px、底部输入+平台 UI ~360px、侧边留 gutter)。
| Bottom input + UI | ~360px | Fake compose bar AND the platform's caption/CTA/audio UI both crowd the bottom |
验证回路:先出打字帧/中段帧/末帧三张静帧,判据包括「打字帧左侧只有三点、没有真气泡」与「最新气泡没被输入框遮住」。
At a typing frame the indicator (not the bubble) shows on the received side; the bubble pops only after.
2核心能力
3外部依赖
| 类型 | 依赖 |
|---|---|
| cli | npx remotion still / render / compositions |
| package | @remotion/bundler + @remotion/renderer(批产路线:bundle 一次、renderMedia 多次) |
| cli | scripts/contact-sheet.sh + scripts/probe-mp4.sh(仓库根共享验证工具箱,SKILL.md 明列) |
| cli | scripts/seek-shot.sh(Light tier:驱动单文件 HTML 的 `?t=N` harness 截图) |
| network | (可选,reference 建议)Remotion Lambda 并行渲染上百条片 |
4风险提醒 风险提醒:蓝色 · 知晓即可
- 外部数据直投屏面,zod 只校验结构不校验内容 — CSV/JSON(含 LLM 产出)里的文本、联系人名会原样渲染与用于文件名;schema 只约束 from/status 等枚举。批量生产前应自行过滤文本(长度、敏感词、Markdown 残留),否则错误内容会被渲染 N 次。
- 音效能力依赖用户自备资产,缺失时静默退化 — send.mp3 / receive.mp3 / 字体 / 头像均不在仓内(见 second_pass discrepancy 条);缺资产时输出清单里的「One pop/ding per message」无法达成,而渲染不会报错。
- reference 批产脚本引用了未随包分发的模块 — `import { csvToThreads } from "./csvToThreads"` 在仓库中不存在,照抄即报错;需按 §3 的 CSV 列约定自行实现映射。README 与 SKILL.md 未标注这段是示意代码。
- 验收仍是人读静帧,节奏类问题最依赖主观判断 — 「滴灌节奏对不对」「打字时长是否过长」没有自动断言;probe-mp4 只验成片规格,contact-sheet 只拼图。文本长度变化会让自动滚动与安全区同时受影响,人工抽样可能漏检。
- 批量放大误差与成本 — 批产路线(一次 bundle、N 次 renderMedia;上百条建议 Lambda)意味着一个布局/时序 bug 会被复制 N 次,且渲染与云资源成本同比例放大;skill 的对策只有「先验一个」这条纪律。
- 题材自带的伦理与平台风险(非代码层面) — 「伪造聊天界面」用于 fake-text 剧情,可能触及平台对虚假信息/冒充的规定与当事人权益;skill 不做任何披露或水印机制。本报告只记录事实,不评价用途。
5第二遍独立确认
- [ok] 「一次一条 + 打字只出现在收到的消息」是否真在代码里成立(防第一遍只读散文) — reference §5/§9 的组件把每条消息包在 `<Sequence from={f(m.typingStartMs)}>` 内,仅当 `m.typingLenMs > 0`(由 `m.from === "them" ? (m.typingMs ?? 800) : 0` 决定)才渲染 `<Typing />`,随后 `<Sequence from={f(m.enterMs - m.typingStartMs)}>` 弹泡;SKILL.md 也明写 "Only received messages type; sent messages just pop"。声明与实现一致。
- [ok] 数据驱动时长(calculateMetadata)是否真能自动出正确长度 — SKILL.md 给 `durationMs = cursor + 1000` 并写 `Compute durationInFrames from durationMs in calculateMetadata`;reference §9 的 Composition 把 `durationInFrames={300}` 标为默认值并由 `calculateMetadata` 覆盖为 `Math.ceil((durationMs / 1000) * 30)`,§4 再次强调 "never hand-set the duration"。三处自洽。
- [discrepancy] reference 示例 import 的模块是否随包分发(第一遍未察觉) — data-driven.md §5 的批渲染脚本写着 `import { csvToThreads } from "./csvToThreads";`,但仓库内并不存在 csvToThreads 文件(全仓仅 .md 与 .sh/.git)。属示意代码:照抄会直接报模块不存在,读者需按 §3 的 CSV 列约定自写映射。同型问题在 CSV 示例里也可见(示例行 `2,Alex,me,no,,400,read` 省略了线程首行)。对 pin commit 有效的真实出入,非虚假能力声明。
- [discrepancy] 音效能力落地是否需要用户自备资产 — SKILL.md 与 reference §8 都用 `staticFile(m.from === "me" ? "send.mp3" : "receive.mp3")`,但仓库内不含任何 mp3;data-driven.md 只在 Tips 里提醒 "Keep `send.mp3` / `receive.mp3` / fonts in `public/`"。结论:音效同步是能力,但前置资产(发/收提示音、字体、头像)必须由用户提供——SKILL.md 的输出清单「One pop/ding per message」在缺资产时会静默退化。
- [ok] 「轻量预览(`?t=N`)」与仓库 Light tier 脚本的契约是否真的对齐 — bubble-recipes.md §10 的 HTML 里实现了 `const t = new URLSearchParams(location.search).get("t"); if (t !== null) { tl.pause(); tl.seek(parseFloat(t)); }`,与 scripts/README.md 描述的 seek-shot.sh 契约(页面 load 时 `tl.pause(); tl.seek(t)`)逐字对应。三包中只有本 skill 真能用上 seek-shot.sh。
- [unlocatable] 「one of the highest-retention faceless short-form formats」是否有据 — SKILL.md 开头断言本格式为最高留存的 faceless 形式之一,仓库内无数据、无引用、无来源标注;本次为只读侦查,不验证该断言,仅记为文档声明。
- [discrepancy] 包间一致性:同为 2026-06-22 发布的三个 iart 包,skill 目录结构是否统一 — tiktok 包四个 skill 各带人读 README(1377–1494 字节);本仓库的 text-message-animation 与 youtube 包的两个 skill 的 README 均为 **0 字节**。功能不受影响(README 非 skill 加载所必需),但包间文档完整度不一致,属可核查的结构出入。
6结论
fdd9560a3ce85697…3a800e1e9b