首页 / 全部技能 / 内容创作 / photo-slideshow
内容创作 · iart-ai/ecommerce-video-skills

photo-slideshow

This skill should be used when the user asks to "make a photo slideshow", "turn photos into a video", "create a photo montage video", "make a picture slideshow with music", "build a memories/recap video from photos", "add Ken Burns pan/zoom to photos", "sync slide changes to the music beat", or "render a video from a folder of images". Covers per-photo Ken Burns motion, transitions, beat-synced timing, mixed aspect-ratio handling, captions/dates, intro/outro cards, and templating a photo folder into a finished video.

风险提醒:橙色 · 评估后使用AI 侦查报告
作者 iart-aiGitHub iart-ai/ecommerce-video-skills ↗Stars 8许可 MIT(仓库根 LICENSE:「MIT License」/「Copyright (c) 2026 iart.ai」;.claude-plugin/plugin.json 与 README「## License」同声明 MIT)commit 2bd0dd19c9
agent 宿主通常会约束 skill 执行权限;风险提醒为 AI 侦查观点,不构成质量或安全保证。第三方 skill 仅作拆解与展示,安装使用风险自负,版权归原作者。

1实现原理 · 为什么它能做到

唯一铁律:没有两张照片可以用同一种运动,且这种差异必须是由索引播种的确定性伪随机。

skills/photo-slideshow/SKILL.md
**No two slides may move the same way.** The instant-giveaway of an amateur slideshow is every photo doing the identical slow zoom-in. Vary direction, zoom in vs. out, and start scale per photo — but make it **deterministic** (seeded by index), so re-renders are identical and a frame-based renderer stays stable.
注:「变化」解决观感,「确定性」解决帧渲染器可能乱序渲染。

播种机制是 mulberry32 + 索引播种:同一张照片永远同一运动,相邻照片不同。

skills/photo-slideshow/references/ken-burns.md
`mulberry32` is a tiny, fast, well-distributed 32-bit PRNG. Seed it with the photo index so photo 0 always gets the same move, but adjacent photos differ.
注:reference 给出 8 种运动词表与「回滚重复」保险。

Ken Burns 只动 transform,且起始 scale ≥1.05,保证平移不露白边。

skills/photo-slideshow/SKILL.md
Always start scale ≥ `1.05` so a pan never exposes an empty edge.
注:KB 说明动 width/height 会触发布局、卡顿甚至闪烁。

混合比例照片用 blurred-pad(模糊底 + contain 前景原图),绝不拉伸。

skills/photo-slideshow/references/aspect-and-layout.md
Blurred-pad wins for mixed folders: it fills the frame, never crops the subject, and the soft blur reads as intentional depth rather than empty bars.
注:另有像素判据:长宽比差异 <15% 用 cover,否则 blurred-pad。

配乐先行:离线测一次 beat 烘进 props,绝不在渲染时逐帧做音频分析。

skills/photo-slideshow/references/timing-and-audio.md
The rule: **lock the music first, derive slide boundaries from it.** Detect beats once, offline, and bake them into props — audio analysis per frame is non-deterministic and slow.
注:换片取每 2 或 4 拍,并提供 BPM 网格退路。

文件夹 → manifest:自然排序、EXIF 拍摄时间、文件名提取 caption,全部派生自数据。

skills/photo-slideshow/references/templating-folder.md
// natural sort so photo2 < photo10 (default lexical sort breaks this)
注:并说明不用文件 mtime:拷贝后不可靠。

图片必须加载门控(staticFile + delayRender/continueRender),否则某帧会空白弹出。

skills/photo-slideshow/SKILL.md
- Photos loaded via `staticFile()`, gated with `delayRender`/`continueRender` so each image (and any caption font) is present before its frame renders — otherwise a slide pops in blank.
注:该要求在 reference 实现中并未照做(见 second_pass)。

验证环重点是「每张照片是否真的加载并正确重构图」,而不是画面美感。

skills/photo-slideshow/SKILL.md
A slideshow ingests a whole folder of user photos — the verify pass is mostly *did every photo actually load and reframe without distortion*.
注:多相册批量时先验一个代表文件夹。

转场只选一种(默认 0.4–0.6s 交叉溶解),且要重叠而不是硬切,除非切在拍点。

skills/photo-slideshow/SKILL.md
Pick **one** transition and keep it consistent — mixing wipes, spins and cubes screams "template."
注:reference 用 TransitionSeries + fade(),15 帧(0.5s@30fps)。

2核心能力

01播种式多样化 Ken Burns(运动词表 + mulberry32 + 主题锚点)
02主题感知锚点(按人脸/焦点设 transform-origin,避免把人摇出画面)
03三种适配策略与自动判据(blurred-pad / cover / contain)
04beat-synced 换片(librosa 离线测拍 + 每 2/4 拍换片 + 落点动态)
05无音频分析的 BPM 网格退路(60/bpm × beatsPerSlide × fps)
06数据驱动 Slideshow(照片+节拍+音乐全为 props,时序由数据推导)
07纯 FFmpeg 退路(zoompan 逐张 + xfade 串联 + 合成音轨)
08多比例安全区(16:9 / 9:16 / 1:1,一次构图多次渲染)

3外部依赖

类型依赖
packageremotion(frame/interpolate/Easing/Img/Audio/Sequence)
package@remotion/transitions(TransitionSeries + fade)
packagesharp + exifr(读图像尺寸与 EXIF DateTimeOriginal)
packagelibrosa + soundfile(pip 安装;离线测 beat 与 onset)
cliffmpeg(blurred-pad、zoompan、xfade 与仓库共享 contact-sheet.sh)
networkiart.ai(reference 尾部推广链接,具 utm 参数;仅读者点击才请求)

4风险提醒 风险提醒:橙色 · 评估后使用

风险提醒:橙色 · 评估后使用
  • reference 与 SKILL.md 的加载门控口径不一致 — 照抄 reference 会绕过 staticFile/delayRender,仍可能出现空白帧与相对路径问题。
  • 文档引用的助手脚本未随包 — build-manifest.mjs 与 beats.py 只有内联源码,仓库无对应文件;照抄命令行会找不到文件。
  • 依赖面最宽且跨生态 — npm(remotion/sharp/exifr)+ pip(librosa/soundfile)+ ffmpeg 三套供应链,任一环节版本不兼容都会卡住产线。
  • 不可信二进制输入通道 — 照片与音频由本地解析库处理,恶意构造文件属解析器漏洞面;skill 无沙箱或校验建议。
  • 隐私未讨论 — EXIF 日期会被上屏、manifest 保留文件路径,全文未提 EXIF/GPS 与输出元数据策略。
  • 商业引流位 — reference 尾部固定 iart.ai 推广段与 utm 链接。
风险提醒:橙色,评估后使用。触发项是分级范式「依赖第三方插件、镜像、远程包」,且是本批 9 个 skill 中依赖面最宽的一个:remotion 与 @remotion/transitions(npm)、sharp 与 exifr(npm,manifest 构建)、librosa 与 soundfile(pip),外加 ffmpeg/ffprobe;均需从公共仓库获取执行。同仓 `remotion-video` 已按同口径判橙,此处与之一致。**明确不成立的项**:无 process.env / API key / token / cookie / 钥匙串读取(token 扫描零实质命中);不把照片与音乐上传任何外部服务(解析与渲染全在本地,素材不出网);无 TLS 降级、无 `--no-verify`、无沙箱关闭、无反自动化绕过。触发条件:仅在构建 manifest、测拍与渲染视频时发生。

5第二遍独立确认

  • [discrepancy] 图片加载门控口径是否实现一致 — SKILL.md 要求 staticFile() + delayRender/continueRender 门控,但 templating-folder.md 的 Slideshow 直接把 manifest 里的文件系统路径(join(dir, f))交给 Img/KenBurnsImage,既未用 staticFile 也无 delayRender——照抄 reference 会拿不到 Remotion public/ 之外的资源,且仍可能出现空白帧。
  • [discrepancy] 被引用的「打包脚本」是否真的存在 — reference 以 `scripts/build-manifest.mjs` 文件名给出源码、SKILL.md 也写 `node scripts/build-manifest.mjs ./photos`,但仓库根 scripts/ 只有 seek-shot.sh / contact-sheet.sh / probe-mp4.sh,没有 build-manifest.mjs;`python beats.py` 同理(脚本内联在文档里)。使用前必须自行落盘。
  • [ok] 起始 scale 与运动确定性口径 — SKILL.md 的 start scale ≥ 1.05、seeded by index 与 ken-burns.md 的 `const base = 1.05;`、`mulberry32(index * 2654435761)` 完全一致。
  • [ok] 适配策略与判据是否自洽 — aspect-and-layout.md 的三策略表与 fitStrategy 阈值(<0.15 → cover)一致;SKILL.md 的 never stretch 与 blurred-pad 默认选择一致。
  • [ok] 换片节奏的量化声明 — timing-and-audio.md 的每张 3–6s、低于约 2.5s 会显仓促、每 2/4 拍换片,与 SKILL.md 的 photos hold ≥1 musical unit, not every beat 一致。
  • [unlocatable] 隐私面:照片元数据是否被提示处理 — manifest 读取 EXIF DateTimeOriginal 并把日期上屏,但全文未讨论照片 GPS/相机信息的处理与输出视频的元数据策略;源码内定位不到任何隐私处理声明,记 unlocatable(能力缺口,不作为档位依据)。

6结论

  • 把「每张照片运动都不同」变成可实现机制(索引播种 + 运动词表),同时保住帧渲染确定性
  • 混合比例照片有默认最优解与判据,杜绝拉伸变形
  • 配乐先行的工程化:离线测一次、烘进 props,避免逐帧音频分析的非确定性
  • manifest 把自然排序/EXIF/文件名 caption 这些真实相册的坑都处理了
  • 给无 Node 环境留了完整 FFmpeg 产线,可用性边界清楚
  • 适合:把一整个相册文件夹变成成片的场景:婚礼/旅行/年度回顾/纪念视频,混合竖横照片、需要 Ken Burns 动感、音乐节拍换片与日期字幕,且要有可复用的模板化管线。
    不适合:只需单张图片做静态卡片;需要旁白叙事驱动时长(用 presentation-video);不愿处理依赖安装与脚本落盘的环境。
    安装 agent 直装可复制
    ① 本站镜像 更新 2026-09-28
    方式 A · 人下载镜像包下载 photo-slideshow.tar.gz
    sha256: 58ffd9961bdf764e…
    方式 B · JSON 格式安装指南,复制给 agent
    安装指南
    agent 读 JSON 指南后会自动从本站下载安装,无需更多说明。
    ② 上游 GitHub · 原始来源
    能访问 GitHub?直接去上游安装(实时版,可能已更新)GitHub 原始 ↗
    本页镜像锁定 commit 2bd0dd19c9;上游为实时仓库。
    来源信息 GitHub 原始
    作者 / 仓库iart-ai / iart-ai/ecommerce-video-skills
    Stars8
    最近推送2026-06-22
    本 skill commit2bd0dd19c9
    许可MIT(仓库根 LICENSE:「MIT License」/「Copyright (c) 2026 iart.ai」;.claude-plugin/plugin.json 与 README「## License」同声明 MIT)
    本站信息
    收录日期2026-09-06
    分类内容创作
    侦查报告AI 侦查 · 2 遍 · 2026-09-06
    本站镜像与 GitHub 原始是不同来源:本站锁定 commit 快照经 /r2 分发;GitHub 为实时上游,内容可能已更新。
    同分类邻近