1实现原理 · 为什么它能做到
核心主张:用静态截图 + 代码动效替代录屏,因为录屏有真光标抖动、改一步要重录等结构性问题。
| Real cursor drifts, overshoots, jitters | Cursor follows an eased path to an exact target |
输入必须以 2×(DPR 2)抓取,2× 放大才不糊。
Capture inputs at twice the display size (retina / `deviceScaleFactor: 2`) so a 2× zoom stays sharp.
光标用 smoothstep 缓动(非线性),涟漪与画面状态变化必须在同一 clickAt 帧发生。
The cursor uses **smoothstep** so it accelerates out of rest and decelerates into the target (a linear cursor reads as robotic). Fire the ripple *and* the screen's state change on the same `clickAt` frame so cause and effect line up.
缩放到焦点用 transform-origin 指向目标像素,上限约 2×,保持 2–4s 再回撤。
Cap zoom around 2× (over 3× loses context and disorients). Hold the zoomed state 2–4s — long enough to read the detail — then ease back out before moving to the next screen.
同一时刻只允许一种运动:要么镜头(zoom),要么光标。
Move the camera *or* the cursor, rarely both at once — two simultaneous motions split attention.
每屏都要装进 browser/device 外框并填真实 URL,否则截图读起来像 bug 报告。
A raw screenshot reads as a bug report; a framed one reads as product. Wrap every screen in a browser or device chrome on a branded background.
坐标工作流被写死:所有光标/缩放目标用 display space 量测(2× 图上读到就除以 2)。
Because the cursor `to` and the zoom target are screen-pixel coordinates, measure them once against the screenshot at display size (1280×800 here, even though the PNG is 2560×1600).
验证环重点是「正确的那张图有没有加载、光标是否落在正确像素」。
The demo ingests user assets — screenshots/designs, optional logo — so verification is mostly *did the right image actually load and land in frame*.
注释层是数据而非硬编码:spotlight、箭头 callout、typed caption 轨道、四种屏间转场全部帧驱动。
const captions = [ { text: "Open any report", in: 30, out: 110 },
2核心能力
3外部依赖
| 类型 | 依赖 |
|---|---|
| package | remotion(AbsoluteFill/Sequence/Img/staticFile/useCurrentFrame/interpolate/spring) |
| package | playwright(含首次下载 Chromium 二进制) |
| network | 用户目标应用 URL(抓图时由无头浏览器访问;示例为占位域名) |
| cli | npx remotion still / render(首帧预览与编码) |
| cli | ffmpeg / ffprobe(仓库共享 contact-sheet.sh、probe-mp4.sh) |
| network | iart.ai(reference 尾部推广链接,具 utm 参数;仅读者点击才请求) |
4风险提醒 风险提醒:橙色 · 评估后使用
- 抓图需访问真实页面 — 无头浏览器会打开使用者目标 origin(示例 URL 为占位,需替换);若目标站有反自动化或需登录,抓图会失败或被拒。
- 示例代码有结构缺陷 — Callout 把 HTML 元素包在 g 里;照抄需自行修正,否则渲染报错或结构无效。
- Playwright + Chromium 下载面 — 首次抓图会从 npm 与浏览器下载源取包与 Chromium 二进制,体积与供应链面都在工具链一侧。
- 注入页面样式的能力 — 脚本用 page.addStyleTag 注入固定 CSS(隐藏光标与 cookie 横幅)——内容受限、不改逻辑,但属向第三方页面注入的能力。
- 商业引流位 — reference 尾部固定 iart.ai 推广段与 utm 链接。
5第二遍独立确认
- [discrepancy] 抓图脚本的目标 URL 是否可直接使用 — reference 的 SHOTS 列表硬编码 https://app.example.com/dashboard 等占位 URL(含 /reports/42?export=open 这类路径),原样运行必然失败;SKILL.md 未提示必须先替换成真实或本地 URL。
- [discrepancy] Callout 组件的可渲染性 — annotation-and-transitions.md 的 Callout 把 HTML div 包在 SVG 元素 g 里(g 只在 SVG 上下文有效),作为 React 组件直接放入 HTML 层会得到无效结构;该片段需改写才能用。
- [ok] 光标/涟漪与 clickAt 的因果同步是否有实现支撑 — screenshot-demo.md 的 Cursor 用 since = frame - clickAt 同时驱动 ripple 与 press 缩放,并把「ripple 与状态变化不同帧」列入失败模式,与 SKILL.md 的因果同帧要求一致。
- [ok] 2× 抓取与 ≤2× 缩放口径 — SD 的 deviceScaleFactor: 2(→2560×1600)与 SKILL.md/reference 的 zoom scale 2(个别 1.8)自洽。
- [ok] 多比例安全区口径 — AT 的 16:9/9:16 表(caption 在下方 80–120px;9:16 保留中心 80% 宽、避开底部 18%)与 SKILL.md 的 Keep captions in the lower third 一致。
- [unlocatable] 需登录的目标页如何处理 — skill 未提供任何会话/凭证注入方式(无 storageState、无 cookie 导入);「需要登录的应用怎么抓图」源码内无解答,记 unlocatable(也正因此不存在凭证读取面)。
6结论
1e7d345b4981c5db…2bd0dd19c9