1实现原理 · 为什么它能做到
唯一一条防坑铁律:条形/波形高度必须由当前帧推导,绝不使用实时 AnalyserNode + rAF 循环。
**Drive every bar height from the current frame, never from a real-time analyser loop.**
给出机制解释:实时分析读的是「此刻在播什么」,而渲染器乱序/变速绘制帧,于是波形会失步或冻住。
Live `AnalyserNode` + `requestAnimationFrame` reads "what is playing right now" — but a video renderer paints frames out of order and faster/slower than real time, so the wave desyncs or freezes.
两条实现路线各自绑定 API:频谱 bars 用 visualizeAudio,平滑示波 用 visualizeAudioWaveform;numberOfSamples 必须是 2 的幂。
`numberOfSamples` must be a power of two (16/32/64); use 16–32 for a chunky branded look, 64+ for a detailed spectrum.
居中均衡器要把前半段镜像到两侧,让低频落在中间。
For a centered equalizer, take the first N bars and **mirror** them around the middle so the bass sits in the center.
长音频别整段载入:先裁到 ≤90s,或改用 useWindowedAudioData(按 HTTP range 只取当前帧附近)。
For anything long, trim first or use `useWindowedAudioData()`, which fetches only the audio around the current frame via HTTP range requests.
版式固定为五层(背景 / 封面+标题 / 波形 / 字幕 / 进度条),进度条是 frame/durationInFrames 的纯函数。
const progress = useCurrentFrame() / durationInFrames; // 0 → 1
字幕机制不在本 skill 实现,明确交给 caption-animation(那是另一个仓库的 skill);本 skill 只负责把字幕轨放进版式并共用同一音频时钟。
The mechanics of word-by-word reveal, active-word highlighting, and SRT/JSON timing belong to the **caption-animation** skill
一稿多幅:同一组件按画幅注册多个 Composition,重新取景而不是加黑边。
Design once inside the center safe zone, then reframe — don't letterbox.
HEAVY tier 的验收核心是「离线烘焙」:渲染期不做音频分析,样本要么烤进 props 要么 useAudioData 提前解码。
never analyse audio at render time; headless render has no realtime audio clock.
验证回路针对「保真 + 事故」两件事:静帧的波形高度必须匹配该时刻音频幅度、字幕同步,同时排查封面缺失/字幕出界/条溢出。
This is HEAVY tier: the deliverable is a real **MP4** (audio muxed in), and the waveform must be *data-driven* — its shape at any frame must match the audio at that time.
2核心能力
3外部依赖
| 类型 | 依赖 |
|---|---|
| package | @remotion/media-utils(useAudioData / useWindowedAudioData / visualizeAudio / visualizeAudioWaveform) |
| cli | npx remotion still / render / compositions |
| cli | scripts/contact-sheet.sh + scripts/probe-mp4.sh(仓库根共享验证工具箱,SKILL.md 明列) |
| network | (条件式)useWindowedAudioData 对远程音频走 HTTP range 请求;本地 staticFile 时无网络 |
| cli | 跨包委派:caption-animation(位于另一仓库 iart-ai/tiktok-video-skills) |
4风险提醒 风险提醒:蓝色 · 知晓即可
- 字幕能力跨仓库依赖未在文档里说明 — 本 skill 把字幕交给另一个仓库的 caption-animation;只装 youtube 包的用户会发现「字幕层」无处可来,而 SKILL.md、plugin.json 与 README 都没提这一前置(见 second_pass discrepancy 条)。
- 同名参数 windowInSeconds 两处语义不同,文档未区分 — useWindowedAudioData 的「抓取窗口」与 visualizeAudioWaveform 的「分析窗口」同名;配错会分别导致「抓太长/太短」或「幅度曲线失真」(见 discrepancy 条)。
- 保真验收依赖人眼比对音频幅度 — 「静帧波形是否匹配该时刻音频」由 agent 看 PNG 判断,没有自动断言(probe-mp4 只验规格);高频细节(如 64+ 采样的高频条)用肉眼很难判定是否真的对上了。
- 长音频/高分辨率渲染成本高 — ≤90s 的 1080p 双画幅渲染已是可观成本;若用户把整集播客直接渲染(未裁剪),内存与时间都会爆炸——skill 给了建议但无强制校验。
- 素材与版权由用户自负 — 音乐/播客音频与封面图的版权、以及「把他人节目切片做 audiogram」的合理使用边界,skill 不做任何提示或校验。
5第二遍独立确认
- [ok] 「禁止实时 analyser」是否在 reference 的两条实现路线里都成立(防第一遍只信开头那句) — Remotion 路线:Bars 组件用 `useAudioData(staticFile(src))` + `visualizeAudio({ audioData, frame, fps, numberOfSamples, smoothing: true })`,值是 frame 的纯函数。canvas 路线:reference 明写 `No Remotion. Decode the file once to a per-frame amplitude array (RMS per frame), then a deterministic loop draws each frame`,并在末尾注明实时 AnalyserNode 只能用于实时预览——`// but never use it to drive a deterministic/offline render.`。两条路线都守住了铁律,且都点名了不得使用实时分析。
- [discrepancy] windowInSeconds 参数在两处的语义是否一致(第一遍漏检) — SKILL.md 在「长文件」小节把 windowInSeconds 用在 `useWindowedAudioData(staticFile("full-episode.mp3"), fps, /* windowInSeconds */ 10)`——语义是「HTTP range 抓取窗口」;而 reference 的参数速查表写 `| windowInSeconds | visualizeAudioWaveform | small (≈1/fps) tracks per-frame amplitude |`——语义是「逐帧幅度分析窗口」,且 §参数示例里 `windowInSeconds: 1 / fps` 出现在 visualizeAudioWaveform 调用中。两个 API 确实各有同名参数,但文档未区分说明,读者极易把「抓取窗口」与「分析窗口」混为一谈(设成 1/fps 去抓长音频,或设成 10 去分析都会出问题)。属对 pin commit 有效的表述出入,非能力虚假。
- [unlocatable] 「80% 静音观看」「加字幕显著提升留存」是否有据 — SKILL.md 写 "the captions carry the message for the 80% who watch on mute" 与 "subtitled video holds attention far longer",仓库内无引用来源与数据文件;只读侦查不验证,仅记为文档声明。
- [discrepancy] 跨包字幕委派会不会造成「只装本包即能力缺失」 — SKILL.md 把词级字幕机制交给 `caption-animation`,但该 skill 位于**另一个仓库** iart-ai/tiktok-video-skills(本包目录树内不存在)。youtube 包的 .claude-plugin/plugin.json 只声明本包两个 skill,README 也未提示这一跨包依赖。结论:只安装 youtube 包时,本 skill 的「captions carry the message」这一层需要用户自行补装另一包,SKILL.md 未说明——属真实的安装面缺口(不影响档位)。
- [ok] HEAVY tier 的「离线烘焙」要求是否与 reference 的实现一致 — SKILL.md 要求「Bake the amplitude/waveform samples into props (or decode once via useAudioData) offline — never analyse audio at render time」,reference 的两条路线恰好各对应一半:Remotion 路线在一次解码后按帧取值,canvas 路线预先把整轨解码为每帧 RMS 数组。要求与实现一致,不是空头约束。
- [ok] 一稿多幅与真实音频长度对齐是否落实 — reference §Full audiogram composition 末尾给出 1x1 与 9x16 两个 `<Composition>` 注册示例,并写 `Set durationInFrames from the actual clip length: Math.round(audioData.durationInSeconds * fps), or read it ahead of time with getAudioDurationInSeconds() and feed it via calculateMetadata.`;批量命令 `npx remotion render Audiogram-1x1 …` / `Audiogram-9x16 …` 亦给出。
6结论
77694b043f16eba0…f3df381d65