首页 / 全部技能 / 内容创作 / camera-3d-captions
内容创作 · heygen-com/hyperframes-community-skills

camera-3d-captions

> Captions living in 3D space around a talking head: a camera flies between the speaker and the words (whip-in, parallax truck, push), caption groups sit at different depths so camera moves pull them apart, hero words hide BEHIND the speaker through an alpha matte, a ring of words wraps round the speaker and turns in front of them, focus racks between depths, and text steps at 15 fps with ghost motion blur. Works in any type style (editorial serif, bold sans) or hand-drawn (p5.brush write-on). Trigger on: "3D text", "3D camera captions", "text around / behind the speaker", "camera moves through the text", "depth captions on my avatar video". Covers prep (portion, padded plate, matte, word clock, font metrics), the shot grammar, depth rules, QC, and limits.

风险提醒:黄色 · 留意使用AI 侦查报告
作者 heygen-comGitHub heygen-com/hyperframes-community-skills ↗Stars 156许可 Apache-2.0(仓库根 LICENSE)commit ba7a0bb6d3
agent 宿主通常会约束 skill 执行权限;风险提醒为 AI 侦查观点,不构成质量或安全保证。第三方 skill 仅作拆解与展示,安装使用风险自负,版权归原作者。

1实现原理 · 为什么它能做到

实现路径是「把 AE 的做法重建成逐帧确定性绘制器」:一个解析时钟同时画 plate、matte 与每个字形,整个合成不使用 CSS 动画(因为渲染器要能任意 seek)。

skills/camera-3d-captions/SKILL.md
One clock paints plate, matte and every glyph. Nothing is a CSS animation.
注:这里在做什么:这条决定了后面所有规则(不能有 repeat:-1、不能有 CSS transition、时间必须可解析求值)。

相机不是真 3D,而是一个解析投影:每个镜头有 cz(沿轴推进)与 px/py(屏幕平移),所有图层活在「参考相机下的参考屏幕坐标」,投影式 screen = C + (ref − C)·D0/(D0 − cz) + (px,py);文字按 15 fps 步进并把相机也一起采样,叠 7 重影运动模糊。

skills/camera-3d-captions/assets/kit/cam3d.js
as a deterministic per-frame painter for HyperFrames. One analytic clock: CAM3D.render(t) writes every style.
注:同文件紧接的模型说明:'Model: each shot has a camera {cz (dolly toward the scene), px/py (pan/tilt as screen shift)}; every layer lives in "ref-screen" coordinates',实现里 project()/unproject() 是这一式的正反解。

「AE 感」不靠硬件能力,靠三条硬编码的时间纪律:文字 POSTERIZE 到 15 fps(连它看到的相机一起采样)、7 重影运动模糊、plate 与 matte 每帧都动。

skills/camera-3d-captions/assets/kit/cam3d.js
Text is posterized (camera AND own animation sampled at 15 fps) with 7-ghost motion blur; the plate/matte move every frame with a directional Gaussian from camera speed.
注:这里在做什么:把「Posterize Time 15 + Motion Blur」翻译成常数 POST = 15、K = 7(见同文件 var W = 1440, H = 1080, CX = 720, CY = 540, FPS = 30, POST = 15, K = 7;)。

时间轴本身**不**锁死在模板里:分镜由模型按台词现场选 4–6 拍,每个分句拆一组;锁死的是绘制契约——15 fps 网格、字幕提前 0.2 s、进场的 ENTRY 位移与深度影子/DOF 公式。

skills/camera-3d-captions/SKILL.md
Pick 4–6 beats. Split each spoken clause into its own group.
注:提前量公式同文件:'Captions lead the voice by 0.2 s, snapped to the 15-fps grid: `fr(t) = 2·round((t − 0.2)·15)`.'——即镜头表(camA/camB/wipe)是实测数据表,而分镜顺序由内容决定。

素材管线是一次性 prep:切出 portion(音频源)、反射补边的 plate-tall(防甩镜/推拉露边)、同帧 alpha matte(person.webm),然后**逐项比对三者**的尺寸/帧率/帧数,不一致直接 exit 1——这是「时间轴锁死」的真实落点。

skills/camera-3d-captions/scripts/prep-take.sh
[ "$tn" -eq "$pn" ] && [ "$mn" -eq "$pn" ] || fail "frame counts differ: portion $pn, plate $tn, matte $mn"
注:参数与默认补边在同文件 usage 注释:prep-take.sh <take.mp4> <start_s> <dur_s> <project/assets> [padTop=480] [padBottom=240] [padSide=240];matte 走 npx -y "$HF" remove-background ... --quality best(本地推理)。

字体度量外置成可生成的数据表:font-metrics.py 用 fontTools 读出每个字形的 advance(em)与 line-height 1 的基线,变量字体先在给定的轴实例化;引擎据此自己排每个字形(CSS 关掉 kerning),所以变量轴必须同时在 CSS 里钉死,否则浏览器按字号走 opsz,宽度漂移。

skills/camera-3d-captions/scripts/font-metrics.py
Writes `window.CAM3D_METRICS = {adv: {key: {char: em}}, base: {key: em}}`.
注:SKILL.md 的对应纪律:'A variable font must get its axes passed here AND pinned in CSS (`font-variation-settings`, `font-optical-sizing: none`). Otherwise the browser's opsz follows the size and words overlap.'

镜头运动曲线不是拟合出来的,是从教程片里逐帧量出来的表:tables.js 存 camA(19 帧甩入+爬行)、camB(拉出/回稳/推进)、wipe 前缘 x 与进场剩余比例,按 30 fps 帧索引、按 H/1080 缩放。

skills/camera-3d-captions/assets/kit/tables.js
/* tables.js — measured motion tables (skill camera-3d-captions). Frame-indexed at 30 fps; px values are for a 1080-tall frame (scale by H/1080);
注:同文件注释还写了宽度适配:'on a wider frame add (W - 1440) / 2 to wipe.x so the band still crosses mid-frame at ref f40'——与 SKILL.md 的 Build 段一致。

手绘风格是跨 skill 复用:借已安装的 p5-paint-animation 引擎把每个词渲成 write-on 精灵图(build-sprites.py 用 subprocess 调该引擎的 render-anim.mjs,读完帧目录后再删掉),因此该风格的前置是另一个 skill 的 setup。

skills/camera-3d-captions/SKILL.md
**Hand-drawn style (optional):** needs the `p5-paint-animation` skill installed and set up (its setup downloads pinned puppeteer, Chrome for Testing, p5 and p5.brush). `build-sprites.py` runs it headlessly; frames are written to that skill's `out/` and removed afterwards.
注:实现见 scripts/hand/build-sprites.py:subprocess.run(['node','render-anim.mjs', sk, out, ...], cwd=engine),引擎目录由 --engine 或环境变量 P5_ENGINE 指定。

2核心能力

01相机在说话人与文字之间飞行(甩入/横移/推进),文字分深度层被视差拉开
02英雄词借 alpha matte 藏在说话人身后(中间字母被头挡住,仍保证 ≥60% 可读)
03环绕环:词在说话人前方的水平圆环上转(本体在环上、按切线旋转并按深度缩放),下弧自动翻转成正读
04深度规则:每组各自一个 D,影子按近远缩放,失焦 σ 随 |D − focus| 线性变化(平板字封顶 1.2 px,环上字形吃满模糊)
05身体擦除转场:用说话人自己的 matte 变黑做带子横扫(A 镜取带左、B 镜取带右),两侧同时拉出
06手绘 write-on 体(p5.brush 笔迹自己写出来;环上整词落位、按切线旋转并按深度缩放)
07QC 三件套:帧内越界检测(贴进预览页跑)、字形间隙检测、以及 prep 阶段的尺寸/帧率/帧数对齐自检
08渲染输出:逐拍抽帧复核 + `render --crf 12` 出片;可选 finish(曲线/SAT 1.12/grain 每帧重掷)

3外部依赖

类型依赖
cliHyperFrames CLI(钉死 0.8.62,经 npx 从 npm 取)
cliffmpeg / ffprobe(切分、反射补边、探帧率与帧数、最后编码)
packagePython3 + fonttools/brotli/numpy/pillow(字体度量与手绘精灵图)
network抠像模型首次运行下载(人物分割权重 ~170 MB,本地推理)
networkGSAP 3.14.2(合成在预览/渲染时由页面加载;参考片 HTML 也引用同一 CDN)
package跨 skill 依赖:p5-paint-animation 的渲染引擎(仅手绘风格需要;用 render-anim.mjs 生成精灵图)
cli素材来源:用户自备 take 视频、自备字体(需自行确认嵌入许可)、以及用户自己本地跑的词级转录器

4风险提醒 风险提醒:黄色 · 留意使用

风险提醒:黄色 · 留意使用
  • 执行期联网不可免(每命令 npx 取 CLI) — 每条命令都经 npx [email protected] 从 npm registry 拉取;离线环境无法执行。版本已钉死,复现性可控,但需要网络与 npm 可达。
  • 首次抠像下载 ~170 MB 权重,属第三方供应链面 — remove-background 由 CLI 从自身分发拉取人物分割模型到 ~/.cache/hyperframes/ 后本地推理;权重来源不在本 repo 审计范围内(与 embedded-captions 同型)。
  • 渲染非完全离线 — 合成页在预览/渲染时从 cdn.jsdelivr.net 加载 [email protected];断网或 CDN 被拦时渲染会失败。
  • 参考片不能照文档直接跑起来 — worked-film.html 除『媒体不含』外还引用未随包的 assets/o2o-data.js 与 Inter 字体文件(见 second_pass 的 discrepancy);读者须自行桥接 window.O2O 并自备字体。SKILL.md 未提示这一点。
  • 跨 skill 依赖会把执行面扩到另一个 skill 的目录 — 手绘风格需要 p5-paint-animation 已 setup(该 setup 会拉 puppeteer/PyPI 之外的 Chrome for Testing 等),且 build-sprites.py 会在那个 skill 的 out/ 写并删除临时帧。按判级口径不改档,但用户要知道执行范围超出本目录。
  • 素材合法性自负 — take 与字体都由用户提供;SKILL.md 只提醒『Confirm the licence allows embedding』(字体),没有对 take 的授权做任何校验。
风险提醒:黄色,留意使用。全目录无凭证接触:没有 env/keychain/cookie 读取、没有 API key 拼接、没有外部生成服务,素材只用用户自己的 take 与字体,SKILL.md 明写 no credentials, no paid calls, nothing uploaded。判黄而非蓝/绿的依据是执行期确有可预期但不免费的联网:① 每条命令都经 npx 从 registry.npmjs.org 取 [email protected](版本钉死,可复现);② 首次抠像会下载 ~170 MB 人物分割权重到 ~/.cache/hyperframes/(一次性供应链面);③ 合成页在预览/渲染时从 cdn.jsdelivr.net 加载 [email protected],渲染因此非完全离线。这与同批判黄的 embedded-captions(首跑拉第三方包 + 本仓模型权重 + jsdelivr gsap)同型;未发现 TLS 降级、反爬绕过或任意代码执行面(子进程调用的都是固定二进制与随包脚本),故不升橙。手绘风格会调用另一个 skill 的引擎并在其 out/ 写临时帧——按「判级对象=skill 自身目录」的口径不抬档,已记入 risks。

5第二遍独立确认

  • [ok] skill 自身是否读凭据(黄/橙分界的关键) — 全目录 process.env / os.environ 只命中 build-sprites.py 的 P5_ENGINE(p5 引擎目录路径);无 API key、token、cookie、keychain、.env 读取。因此按范式的橙档第一条不成立,落黄。
  • [ok] 「外部资源」逐条回查调用点 — [email protected] 在 SKILL.md 与 prep-take.sh(HF="[email protected]")两处一致;ffmpeg/ffprobe 在 prep-take.sh 的 ffmpeg/probe() 内可验;fontTools 在 font-metrics.py import 段;GSAP CDN 在 worked-film.html 与 SKILL.md 的 Render time 条;p5 引擎调用在 build-sprites.py 的 subprocess.run(cwd=engine)。
  • [ok] prep 阶段的三者对齐自检是否真会失败退出(不是装饰性文字) — prep-take.sh 定义 fail(){ echo ... >&2; exit 1; },并对 plate 尺寸、matte 尺寸、三者的帧率与帧数各写了一条比较分支,任一不符即 exit 1;成功路径打印 'prep-take: plate and matte aligned with the portion ($pn frames @ $pr)'。
  • [ok] 手绘风格是否真的复用 p5-paint-animation 引擎,且临时帧会被清理 — build-sprites.py 用 subprocess.run(['node','render-anim.mjs', ...], cwd=engine) 渲染临时 sketch(写在 tempfile 目录,落在自己的 temp 而非引擎内),读引擎 out/_frames_<name>.anim/ 的 PNG 后调用 shutil.rmtree(fdir, ignore_errors=True);与 SKILL.md『frames are written to that skill's out/ and removed afterwards』一致。
  • [discrepancy] worked-film.html 能否按 SKILL 的 Build 步骤原样打开(参考片的可运行性声明) — SKILL.md 只说 'its media is not included',但参考片的 <head>/<body> 还引用了 assets/o2o-data.js(注释写 'window.O2O here = assets/kit/tables.js')与 assets/Inter-Variable.woff2、assets/InterTight-900.woff2 等字体文件;o2o-data.js 与这两个 woff2 均不在包内(全量清单已确认)。按 SKILL 的『Copy assets/kit/{cam3d.js,tables.js,finish.js} and metrics.js into the project's assets/』操作仍缺 o2o-data.js,参考片不能原样运行——读者须自行把 tables.js 改名/桥接成 window.O2O 并自带字体。该出入对 pin commit 有效。
  • [ok] 15 fps 网格与 0.2 s 提前量是否真在实现里(而不是只写在文档) — cam3d.js 顶部 var W = 1440, H = 1080, CX = 720, CY = 540, FPS = 30, POST = 15, K = 7;,并有 post(t) = Math.floor(t * POST + 1e-4) / POST;SKILL.md 给出 fr(t) = 2·round((t − 0.2)·15)(×2 即 30 fps 渲染下的 15 fps 映射),两处自洽。
  • [unlocatable] 『带 grade 与多重影渲染约 2× 成本』的量化声明 — SKILL.md 的 Limits 段写 'Rendering with the grade and many ghost layers costs about 2× a plain render.',包内没有基准脚本或计时数据可复现该倍数;本次为只读侦查且不得运行渲染,故仅作为文档声明记录,不作为已核事实使用。

6结论

  • 把 AE 的观感拆成可复现的确定性机制:解析相机 + 15 fps posterize + 7 重影 + 实测运动表,而不是靠手调关键帧
  • 素材管线自带一致性闸门:plate 与 matte 必须与 portion 逐项对齐,否则直接失败退出(避免整片合成错位)
  • 文字排版不靠浏览器排版:advance/基线由 font-metrics.py 外置成表,引擎自己摆每个字形,并强制变量轴在 CSS 钉死
  • QC 直接读浏览器里的真实 DOM 与 canvas 度量(越界词、跨词碰撞),而不是估算
  • 前置条件与代价写得诚实(Node 22+、四个 Python 包、字体许可、发丝处的 matte 抖动、说话人占满画面的限制)
  • 适合:适合:锁定机位的说话人素材(数字人或有真人),要做 5–20 s 的「文字在 3D 空间里围着人表演」的段落——相机在人与字之间飞、字幕分深度被视差拉开、英雄词躲在头后面、词环绕人转、景深在层间切换、文字以 15 fps 步进带重影。也适合把 AE 的 Postorize Time/3D camera/DOF 手感搬进 HyperFrames 的可复现管线,并有纪律地做帧内越界与字形碰撞 QC。
    不适合:不适合:① 普通字幕(应走 embedded-captions 一类);② 移动/手持镜头或极端特写(matte 与 plate 变换假设锁机位,且头部后面没有余量);③ 抠像模型分不出主体的素材;④ 需要实时/交互页面而非离线成片的场景;⑤ 完全离线且无 npm 的环境;⑥ 不愿接受首次 170 MB 模型下载与渲染期 CDN 依赖的用户。
    安装 agent 直装可复制
    ① 本站镜像 更新 2026-09-28
    方式 A · 人下载镜像包下载 camera-3d-captions.tar.gz
    sha256: 70cf48e51fa89f4e…
    方式 B · JSON 格式安装指南,复制给 agent
    安装指南
    agent 读 JSON 指南后会自动从本站下载安装,无需更多说明。
    ② 上游 GitHub · 原始来源
    能访问 GitHub?直接去上游安装(实时版,可能已更新)GitHub 原始 ↗
    本页镜像锁定 commit ba7a0bb6d3;上游为实时仓库。
    来源信息 GitHub 原始
    作者 / 仓库heygen-com / heygen-com/hyperframes-community-skills
    Stars156
    最近推送2026-09-24
    本 skill commitba7a0bb6d3
    许可Apache-2.0(仓库根 LICENSE)
    本站信息
    收录日期2026-09-06
    分类内容创作
    侦查报告AI 侦查 · 2 遍 · 2026-09-06
    本站镜像与 GitHub 原始是不同来源:本站锁定 commit 快照经 /r2 分发;GitHub 为实时上游,内容可能已更新。
    同分类邻近