Watch — ask questions about any video - **name_fr:** Watch — poser des questions sur n'importe quelle vidéo - **tl_en:** Download a video with yt-dlp, extract timestamped frames and a transcript — or let Google's Gemini watch the whole video — then answer questions about it - **tl_fr:** Télécharger une vidéo avec yt-dlp, extraire des frames horodatées et une transcription — ou laisser Google Gemini regarder toute la vidéo — puis répondre aux questions - **creator:** @bradautomates - **type:** Agent skill - **url:** https://github.com/bradautomates/claude-video - **cat:** Video - **kws:** video analysis, yt-dlp, ffmpeg, transcription, whisperx, gemini, frames, video q&a - **license:** MIT **Description EN:** Curated by Skill Harbor — a video-watching skill by @bradautomates that lets an agent answer questions about any video. With the local engine it downloads the video (URL or local file) with yt-dlp, extracts auto-scaled timestamped frames with ffmpeg and a transcript from captions (or a local WhisperX / Groq / OpenAI fallback), then combines visuals and transcript as evidence. With a Gemini API key, Google's agentic video model watches the full video directly and the report relays its timestamped answer. Includes a guided first-run setup (engine choice, detail levels from transcript-only to token-burner, transcription backend), and security-conscious defaults: keys live in a local `~/.config/watch/.env` (0600 permissions), are never printed, and video content is treated as untrusted evidence, never instructions. Honest caveats: the Gemini engine sends the video to Google (local files are uploaded, then deleted after the answer); the local WhisperX path needs a one-time ~1.5 GB model download and 8 GB RAM; long clips get sparse frame coverage unless you focus a time interval. Credit to its creator. Skill Harbor never reviews the code, review it yourself before use. Discovered via skills.sh. **Description FR:** Sélectionné par Skill Harbor — un skill d'écoute vidéo par @bradautomates qui permet à un agent de répondre à des questions sur n'importe quelle vidéo. Avec le moteur local, il télécharge la vidéo (URL ou fichier local) avec yt-dlp, extrait des frames horodatées avec ffmpeg et une transcription depuis les sous-titres (ou un fallback local WhisperX / Groq / OpenAI), puis combine visuels et transcription comme preuves. Avec une clé API Gemini, le modèle vidéo agentique de Google regarde toute la vidéo directement et le rapport relaie sa réponse horodatée. Inclut un assistant de premier lancement (choix du moteur, niveaux de détail du mode transcription seule au mode gourmand en tokens, moteur de transcription), et des défauts soucieux de la sécurité : les clés vivent dans un `~/.config/watch/.env` local (permissions 0600), ne sont jamais affichées, et le contenu vidéo est traité comme preuve non fiable, jamais comme instructions. Bémols honnêtes : le moteur Gemini envoie la vidéo à Google (les fichiers locaux sont téléversés puis supprimés après la réponse) ; le chemin WhisperX local nécessite un téléchargement unique de ~1,5 Go de modèle et 8 Go de RAM ; les longs clips ont une couverture de frames clairsemée sauf si on cible un intervalle de temps. Crédit au créateur. Skill Harbor ne vérifie jamais le code, examinez-le vous-même avant usage. Découvert via skills.sh. **Install prompt EN:** ``` Prerequisites: Python 3.10+; optional free Gemini API key (aistudio.google.com/apikey) for the cloud engine. Local engine needs ffmpeg and yt-dlp; the WhisperX fallback needs a one-time ~1.5 GB model download and 8 GB RAM. Install "Watch — ask questions about any video" for me. The agent downloads a video (URL or local file) with yt-dlp, extracts timestamped frames with ffmpeg and a transcript from captions (or WhisperX / cloud transcription), and answers questions about the video from that evidence — or, with a Gemini key, has Google's video model watch the whole video directly. Repository: https://github.com/bradautomates/claude-video/blob/main/skills/watch/SKILL.md 1. Fetch the SKILL.md file (and any helper files, including the scripts/ folder) from the repository path into a temporary folder and summarize what it does in one or two sentences. 2. Safety check: review the SKILL.md and scripts for anything suspicious (unexpected network calls, shell commands, credential harvesting). This repo should contain zero secrets in code, credentials only via the secure vault, allowed hosts declared in the SKILL.md. Verify that holds here; STOP on any red flag and tell me. 3. Install it as a skill: copy SKILL.md and its helper files into the agent's skills directory, in a folder named "watch". 4. Verify with no network calls: frontmatter valid, files in place. 5. Report what was installed, where, and what I still need to do myself (e.g. save my Gemini API key in the vault). GitHub is optional: if I have a GitHub account or the gh CLI, you may use it; otherwise public access is fine. Never require it unless it's in the prerequisites above. Rules: don't touch anything outside the temp folder and the install target. If anything looks off, stop and ask me. ``` **Install prompt FR:** ``` Prérequis : Python 3.10+ ; clé API Gemini gratuite optionnelle (aistudio.google.com/apikey) pour le moteur cloud. Le moteur local nécessite ffmpeg et yt-dlp ; le fallback WhisperX exige un téléchargement unique de ~1,5 Go de modèle et 8 Go de RAM. Installe-moi « Watch — poser des questions sur n'importe quelle vidéo ». L'agent télécharge une vidéo (URL ou fichier local) avec yt-dlp, en extrait des frames horodatées avec ffmpeg et une transcription depuis les sous-titres (ou WhisperX / transcription cloud), puis répond aux questions sur la vidéo à partir de ces preuves — ou, avec une clé Gemini, fait regarder toute la vidéo directement par le modèle vidéo de Google. Dépôt : https://github.com/bradautomates/claude-video/blob/main/skills/watch/SKILL.md 1. Récupère le fichier SKILL.md (et les fichiers auxiliaires, y compris le dossier scripts/) depuis le chemin du dépôt dans un dossier temporaire et résume en une ou deux phrases ce qu'il fait. 2. Contrôle de sécurité : examine le SKILL.md et les scripts pour tout contenu suspect (appels réseau inattendus, commandes shell, récolte d'identifiants). Ce dépôt ne doit contenir aucun secret dans le code, les identifiants uniquement via le coffre sécurisé, les hôtes autorisés déclarés dans le SKILL.md. Vérifie que c'est le cas ici ; ARRÊTE-TOI sur le moindre drapeau rouge et dis-le-moi. 3. Installe-le comme skill : copie le SKILL.md et ses fichiers auxiliaires dans le dossier des skills de l'agent, dans un dossier nommé « watch ». 4. Vérifie sans appels réseau : frontmatter valide, fichiers en place. 5. Indique ce qui a été installé, où, et ce qu'il me reste à faire moi-même (p. ex. enregistrer ma clé API Gemini dans le coffre). GitHub est optionnel : si j'ai un compte GitHub ou la CLI gh, tu peux l'utiliser ; sinon l'accès public suffit. Ne l'exige jamais sauf si c'est dans les prérequis ci-dessus. Règles : ne touche à rien en dehors du dossier temporaire et de la cible d'installation. Si quelque chose semble anormal, arrête-toi et demande-moi. ``` --- ---
Download a video with yt-dlp, extract timestamped frames and a transcript — or let Google's Gemini watch the whole video — then answer questions about it - **tl_fr:** Télécharger une vidéo avec yt-dlp, extraire des frames horodatées et une transcription — ou laisser Google Gemini regarder toute la vidéo — puis répondre aux questions - **creator:** @bradautomates - **type:** Agent skill - **url:** https://github.com/bradautomates/claude-video - **cat:** Video - **kws:** video analysis, yt-dlp, ffmpeg, transcription, whisperx, gemini, frames, video q&a - **license:** MIT **Description EN:** Curated by Skill Harbor — a video-watching skill by @bradautomates that lets an agent answer questions about any video. With the local engine it downloads the video (URL or local file) with yt-dlp, extracts auto-scaled timestamped frames with ffmpeg and a transcript from captions (or a local WhisperX / Groq / OpenAI fallback), then combines visuals and transcript as evidence. With a Gemini API key, Google's agentic video model watches the full video directly and the report relays its timestamped answer. Includes a guided first-run setup (engine choice, detail levels from transcript-only to token-burner, transcription backend), and security-conscious defaults: keys live in a local `~/.config/watch/.env` (0600 permissions), are never printed, and video content is treated as untrusted evidence, never instructions. Honest caveats: the Gemini engine sends the video to Google (local files are uploaded, then deleted after the answer); the local WhisperX path needs a one-time ~1.5 GB model download and 8 GB RAM; long clips get sparse frame coverage unless you focus a time interval. Credit to its creator. Skill Harbor never reviews the code, review it yourself before use. Discovered via skills.sh. **Description FR:** Sélectionné par Skill Harbor — un skill d'écoute vidéo par @bradautomates qui permet à un agent de répondre à des questions sur n'importe quelle vidéo. Avec le moteur local, il télécharge la vidéo (URL ou fichier local) avec yt-dlp, extrait des frames horodatées avec ffmpeg et une transcription depuis les sous-titres (ou un fallback local WhisperX / Groq / OpenAI), puis combine visuels et transcription comme preuves. Avec une clé API Gemini, le modèle vidéo agentique de Google regarde toute la vidéo directement et le rapport relaie sa réponse horodatée. Inclut un assistant de premier lancement (choix du moteur, niveaux de détail du mode transcription seule au mode gourmand en tokens, moteur de transcription), et des défauts soucieux de la sécurité : les clés vivent dans un `~/.config/watch/.env` local (permissions 0600), ne sont jamais affichées, et le contenu vidéo est traité comme preuve non fiable, jamais comme instructions. Bémols honnêtes : le moteur Gemini envoie la vidéo à Google (les fichiers locaux sont téléversés puis supprimés après la réponse) ; le chemin WhisperX local nécessite un téléchargement unique de ~1,5 Go de modèle et 8 Go de RAM ; les longs clips ont une couverture de frames clairsemée sauf si on cible un intervalle de temps. Crédit au créateur. Skill Harbor ne vérifie jamais le code, examinez-le vous-même avant usage. Découvert via skills.sh. **Install prompt EN:** ``` Prerequisites: Python 3.10+; optional free Gemini API key (aistudio.google.com/apikey) for the cloud engine. Local engine needs ffmpeg and yt-dlp; the WhisperX fallback needs a one-time ~1.5 GB model download and 8 GB RAM. Install "Watch — ask questions about any video" for me. The agent downloads a video (URL or local file) with yt-dlp, extracts timestamped frames with ffmpeg and a transcript from captions (or WhisperX / cloud transcription), and answers questions about the video from that evidence — or, with a Gemini key, has Google's video model watch the whole video directly. Repository: https://github.com/bradautomates/claude-video/blob/main/skills/watch/SKILL.md 1. Fetch the SKILL.md file (and any helper files, including the scripts/ folder) from the repository path into a temporary folder and summarize what it does in one or two sentences. 2. Safety check: review the SKILL.md and scripts for anything suspicious (unexpected network calls, shell commands, credential harvesting). This repo should contain zero secrets in code, credentials only via the secure vault, allowed hosts declared in the SKILL.md. Verify that holds here; STOP on any red flag and tell me. 3. Install it as a skill: copy SKILL.md and its helper files into the agent's skills directory, in a folder named "watch". 4. Verify with no network calls: frontmatter valid, files in place. 5. Report what was installed, where, and what I still need to do myself (e.g. save my Gemini API key in the vault). GitHub is optional: if I have a GitHub account or the gh CLI, you may use it; otherwise public access is fine. Never require it unless it's in the prerequisites above. Rules: don't touch anything outside the temp folder and the install target. If anything looks off, stop and ask me. ``` **Install prompt FR:** ``` Prérequis : Python 3.10+ ; clé API Gemini gratuite optionnelle (aistudio.google.com/apikey) pour le moteur cloud. Le moteur local nécessite ffmpeg et yt-dlp ; le fallback WhisperX exige un téléchargement unique de ~1,5 Go de modèle et 8 Go de RAM. Installe-moi « Watch — poser des questions sur n'importe quelle vidéo ». L'agent télécharge une vidéo (URL ou fichier local) avec yt-dlp, en extrait des frames horodatées avec ffmpeg et une transcription depuis les sous-titres (ou WhisperX / transcription cloud), puis répond aux questions sur la vidéo à partir de ces preuves — ou, avec une clé Gemini, fait regarder toute la vidéo directement par le modèle vidéo de Google. Dépôt : https://github.com/bradautomates/claude-video/blob/main/skills/watch/SKILL.md 1. Récupère le fichier SKILL.md (et les fichiers auxiliaires, y compris le dossier scripts/) depuis le chemin du dépôt dans un dossier temporaire et résume en une ou deux phrases ce qu'il fait. 2. Contrôle de sécurité : examine le SKILL.md et les scripts pour tout contenu suspect (appels réseau inattendus, commandes shell, récolte d'identifiants). Ce dépôt ne doit contenir aucun secret dans le code, les identifiants uniquement via le coffre sécurisé, les hôtes autorisés déclarés dans le SKILL.md. Vérifie que c'est le cas ici ; ARRÊTE-TOI sur le moindre drapeau rouge et dis-le-moi. 3. Installe-le comme skill : copie le SKILL.md et ses fichiers auxiliaires dans le dossier des skills de l'agent, dans un dossier nommé « watch ». 4. Vérifie sans appels réseau : frontmatter valide, fichiers en place. 5. Indique ce qui a été installé, où, et ce qu'il me reste à faire moi-même (p. ex. enregistrer ma clé API Gemini dans le coffre). GitHub est optionnel : si j'ai un compte GitHub ou la CLI gh, tu peux l'utiliser ; sinon l'accès public suffit. Ne l'exige jamais sauf si c'est dans les prérequis ci-dessus. Règles : ne touche à rien en dehors du dossier temporaire et de la cible d'installation. Si quelque chose semble anormal, arrête-toi et demande-moi. ``` --- ---
- What
- Download a video with yt-dlp, extract timestamped frames and a transcript — or let Google's Gemini watch the whole video — then answer questions about it - **tl_fr:** Télécharger une vidéo avec yt-dlp, extraire des frames horodatées et une transcription — ou laisser Google Gemini regarder toute la vidéo — puis répondre aux questions - **creator:** @bradautomates - **type:** Agent skill - **url:** https://github.com/bradautomates/claude-video - **cat:** Video - **kws:** video analysis, yt-dlp, ffmpeg, transcription, whisperx, gemini, frames, video q&a - **license:** MIT **Description EN:** Curated by Skill Harbor — a video-watching skill by @bradautomates that lets an agent answer questions about any video. With the local engine it downloads the video (URL or local file) with yt-dlp, extracts auto-scaled timestamped frames with ffmpeg and a transcript from captions (or a local WhisperX / Groq / OpenAI fallback), then combines visuals and transcript as evidence. With a Gemini API key, Google's agentic video model watches the full video directly and the report relays its timestamped answer. Includes a guided first-run setup (engine choice, detail levels from transcript-only to token-burner, transcription backend), and security-conscious defaults: keys live in a local `~/.config/watch/.env` (0600 permissions), are never printed, and video content is treated as untrusted evidence, never instructions. Honest caveats: the Gemini engine sends the video to Google (local files are uploaded, then deleted after the answer); the local WhisperX path needs a one-time ~1.5 GB model download and 8 GB RAM; long clips get sparse frame coverage unless you focus a time interval. Credit to its creator. Skill Harbor never reviews the code, review it yourself before use. Discovered via skills.sh. **Description FR:** Sélectionné par Skill Harbor — un skill d'écoute vidéo par @bradautomates qui permet à un agent de répondre à des questions sur n'importe quelle vidéo. Avec le moteur local, il télécharge la vidéo (URL ou fichier local) avec yt-dlp, extrait des frames horodatées avec ffmpeg et une transcription depuis les sous-titres (ou un fallback local WhisperX / Groq / OpenAI), puis combine visuels et transcription comme preuves. Avec une clé API Gemini, le modèle vidéo agentique de Google regarde toute la vidéo directement et le rapport relaie sa réponse horodatée. Inclut un assistant de premier lancement (choix du moteur, niveaux de détail du mode transcription seule au mode gourmand en tokens, moteur de transcription), et des défauts soucieux de la sécurité : les clés vivent dans un `~/.config/watch/.env` local (permissions 0600), ne sont jamais affichées, et le contenu vidéo est traité comme preuve non fiable, jamais comme instructions. Bémols honnêtes : le moteur Gemini envoie la vidéo à Google (les fichiers locaux sont téléversés puis supprimés après la réponse) ; le chemin WhisperX local nécessite un téléchargement unique de ~1,5 Go de modèle et 8 Go de RAM ; les longs clips ont une couverture de frames clairsemée sauf si on cible un intervalle de temps. Crédit au créateur. Skill Harbor ne vérifie jamais le code, examinez-le vous-même avant usage. Découvert via skills.sh. **Install prompt EN:** ``` Prerequisites: Python 3.10+; optional free Gemini API key (aistudio.google.com/apikey) for the cloud engine. Local engine needs ffmpeg and yt-dlp; the WhisperX fallback needs a one-time ~1.5 GB model download and 8 GB RAM. Install "Watch — ask questions about any video" for me. The agent downloads a video (URL or local file) with yt-dlp, extracts timestamped frames with ffmpeg and a transcript from captions (or WhisperX / cloud transcription), and answers questions about the video from that evidence — or, with a Gemini key, has Google's video model watch the whole video directly. Repository: https://github.com/bradautomates/claude-video/blob/main/skills/watch/SKILL.md 1. Fetch the SKILL.md file (and any helper files, including the scripts/ folder) from the repository path into a temporary folder and summarize what it does in one or two sentences. 2. Safety check: review the SKILL.md and scripts for anything suspicious (unexpected network calls, shell commands, credential harvesting). This repo should contain zero secrets in code, credentials only via the secure vault, allowed hosts declared in the SKILL.md. Verify that holds here; STOP on any red flag and tell me. 3. Install it as a skill: copy SKILL.md and its helper files into the agent's skills directory, in a folder named "watch". 4. Verify with no network calls: frontmatter valid, files in place. 5. Report what was installed, where, and what I still need to do myself (e.g. save my Gemini API key in the vault). GitHub is optional: if I have a GitHub account or the gh CLI, you may use it; otherwise public access is fine. Never require it unless it's in the prerequisites above. Rules: don't touch anything outside the temp folder and the install target. If anything looks off, stop and ask me. ``` **Install prompt FR:** ``` Prérequis : Python 3.10+ ; clé API Gemini gratuite optionnelle (aistudio.google.com/apikey) pour le moteur cloud. Le moteur local nécessite ffmpeg et yt-dlp ; le fallback WhisperX exige un téléchargement unique de ~1,5 Go de modèle et 8 Go de RAM. Installe-moi « Watch — poser des questions sur n'importe quelle vidéo ». L'agent télécharge une vidéo (URL ou fichier local) avec yt-dlp, en extrait des frames horodatées avec ffmpeg et une transcription depuis les sous-titres (ou WhisperX / transcription cloud), puis répond aux questions sur la vidéo à partir de ces preuves — ou, avec une clé Gemini, fait regarder toute la vidéo directement par le modèle vidéo de Google. Dépôt : https://github.com/bradautomates/claude-video/blob/main/skills/watch/SKILL.md 1. Récupère le fichier SKILL.md (et les fichiers auxiliaires, y compris le dossier scripts/) depuis le chemin du dépôt dans un dossier temporaire et résume en une ou deux phrases ce qu'il fait. 2. Contrôle de sécurité : examine le SKILL.md et les scripts pour tout contenu suspect (appels réseau inattendus, commandes shell, récolte d'identifiants). Ce dépôt ne doit contenir aucun secret dans le code, les identifiants uniquement via le coffre sécurisé, les hôtes autorisés déclarés dans le SKILL.md. Vérifie que c'est le cas ici ; ARRÊTE-TOI sur le moindre drapeau rouge et dis-le-moi. 3. Installe-le comme skill : copie le SKILL.md et ses fichiers auxiliaires dans le dossier des skills de l'agent, dans un dossier nommé « watch ». 4. Vérifie sans appels réseau : frontmatter valide, fichiers en place. 5. Indique ce qui a été installé, où, et ce qu'il me reste à faire moi-même (p. ex. enregistrer ma clé API Gemini dans le coffre). GitHub est optionnel : si j'ai un compte GitHub ou la CLI gh, tu peux l'utiliser ; sinon l'accès public suffit. Ne l'exige jamais sauf si c'est dans les prérequis ci-dessus. Règles : ne touche à rien en dehors du dossier temporaire et de la cible d'installation. Si quelque chose semble anormal, arrête-toi et demande-moi. ``` --- ---
- Cost
- Free
- Needs
- Python 3.10+; optional free Gemini API key (aistudio.google.com/apikey) for the cloud engine. Local engine needs ffmpeg and yt-dlp; the WhisperX fallback needs a one-time ~1.5 GB model download and 8 GB RAM.
- Install
- Copy the installer prompt below into your Muse — your agent does the rest.
Curated by Skill Harbor — a video-watching skill by @bradautomates that lets an agent answer questions about any video. With the local engine it downloads the video (URL or local file) with yt-dlp, extracts auto-scaled timestamped frames with ffmpeg and a transcript from captions (or a local WhisperX / Groq / OpenAI fallback), then combines visuals and transcript as evidence. With a Gemini API key, Google's agentic video model watches the full video directly and the report relays its timestamped answer. Includes a guided first-run setup (engine choice, detail levels from transcript-only to token-burner, transcription backend), and security-conscious defaults: keys live in a local `~/.config/watch/.env` (0600 permissions), are never printed, and video content is treated as untrusted evidence, never instructions. Honest caveats: the Gemini engine sends the video to Google (local files are uploaded, then deleted after the answer); the local WhisperX path needs a one-time ~1.5 GB model download and 8 GB RAM; long clips get sparse frame coverage unless you focus a time interval. Credit to its creator. Skill Harbor never reviews the code, review it yourself before use. Discovered via skills.sh.
Version:
Install
Prerequisites: Python 3.10+; optional free Gemini API key (aistudio.google.com/apikey) for the cloud engine. Local engine needs ffmpeg and yt-dlp; the WhisperX fallback needs a one-time ~1.5 GB model download and 8 GB RAM. Install "Watch — ask questions about any video" for me. The agent downloads a video (URL or local file) with yt-dlp, extracts timestamped frames with ffmpeg and a transcript from captions (or WhisperX / cloud transcription), and answers questions about the video from that evidence — or, with a Gemini key, has Google's video model watch the whole video directly. Repository: https://github.com/bradautomates/claude-video/blob/main/skills/watch/SKILL.md 1. Fetch the SKILL.md file (and any helper files, including the scripts/ folder) from the repository path into a temporary folder and summarize what it does in one or two sentences. 2. Safety check: review the SKILL.md and scripts for anything suspicious (unexpected network calls, shell commands, credential harvesting). This repo should contain zero secrets in code, credentials only via the secure vault, allowed hosts declared in the SKILL.md. Verify that holds here; STOP on any red flag and tell me. 3. Install it as a skill: copy SKILL.md and its helper files into the agent's skills directory, in a folder named "watch". 4. Verify with no network calls: frontmatter valid, files in place. 5. Report what was installed, where, and what I still need to do myself (e.g. save my Gemini API key in the vault). GitHub is optional: if I have a GitHub account or the gh CLI, you may use it; otherwise public access is fine. Never require it unless it's in the prerequisites above. Rules: don't touch anything outside the temp folder and the install target. If anything looks off, stop and ask me.
Questions
How do I install a build?
Every product page includes a copy-paste install prompt. Paste it into your Muse and it sets the build up for you — no manual configuration.
Where does my money go?
Straight to the seller. Skill Harbor never processes payments: checkout happens on the seller’s own page, usually Stripe.
What does the ✓ next to a creator’s name mean?
It means we confirmed the identity of the person behind the listing. It says nothing about the code itself — always check a build before installing it.