Watch a video
This skill prepares a video for analysis. It pulls a transcript and extracts a small number of frames for visual context. No remote transcription API is used. Captions come straight from the platform when available; videos without captions are transcribed locally with whisper.cpp (~140MB model, runs on CPU). If both fail, the script falls back to vision-only analysis from frames.
When to use
Trigger when the user shares a video URL or file path and asks anything about i
[Description truncada. Veja o README completo no GitHub.]