# YouTube Transcript Agent-vaardigheid

> Download YouTube-videotranscripties, ondertitels en coverafbeeldingen via URL of video-ID. Ondersteunt meerdere talen, vertaling, hoofdstukken en sprekeridentificatie om de toegankelijkheid van content te verbeteren. Begin gratis binnen enkele seconden.

- Canonical: https://nanoskill.ai/nl/skills/youtube-transcript
- Markdown: https://nanoskill.ai/nl/skills/youtube-transcript.md
- Author: JimLiu
- Published: 2026-06-03T05:58:31.787Z
- Updated: 2026-07-15T13:36:15.347Z
- Language: nl
- Source type: github
- Popularity signal: 1040

## Sources

- https://github.com/JimLiu/baoyu-skills

## Install

```shell
npx skills add https://github.com/JimLiu/baoyu-skills/tree/main/skills/baoyu-youtube-transcript
```

## About

De YouTube Transcript Downloader-vaardigheid biedt een robuuste oplossing voor het extraheren van uitgebreide informatie uit YouTube-video's. Met deze vaardigheid kunnen gebruikers moeiteloos volledige YouTube-transcripties, ondertitels en zelfs coverafbeeldingen downloaden door simpelweg een video-URL of -ID op te geven. Het is ontworpen voor contentmakers, onderzoekers en iedereen die gesproken content naar tekst moet omzetten, met functies zoals ondersteuning voor meerdere talen, vertaalmogelijkheden en geavanceerde structureringsopties.

Door gebruik te maken van directe toegang tot YouTube's InnerTube API en een slimme terugvaloptie naar \`yt-dlp\`, zorgt de vaardigheid voor betrouwbare en efficiënte gegevensopvraging zonder dat persoonlijke API-sleutels nodig zijn. Het biedt flexibele uitvoerformaten, waaronder Markdown voor gedetailleerde analyse met tijdstempels en hoofdstukmarkeringen, en SRT voor standaard ondertitelintegratie. Bovendien ondersteunt het segmentatie van hoofdstukken uit videobeschrijvingen en biedt het een workflow voor AI-gestuurde sprekeridentificatie, wat leidt tot zeer georganiseerde en toegeschreven tekst.

Met intelligente caching minimaliseert de vaardigheid overbodige netwerkverzoeken, waardoor latere bewerkingen op dezelfde video uitzonderlijk snel zijn. Of u nu video-inhoud moet analyseren, toegankelijke ondertitels wilt maken, materiaal wilt vertalen voor een wereldwijd publiek of gewoon videometadata en thumbnails wilt extraheren, deze vaardigheid stroomlijnt het proces en biedt een krachtige tool voor het beheer van YouTube-content.

## Key features

- **Directe YouTube-toegang**: Geeft rechtstreeks toegang tot de Binnenbuis-API van YouTube voor snelle transcriptieophaal, schakelt automatisch over op \`yt-dlp\` als de directe API wordt geblokkeerd, waardoor betrouwbare toegang zonder API-sleutels wordt gegarandeerd.
- **Meertalige ondersteuning en vertaling**: Specificeer voorkeurstalen voor transcripties en vertaal ze naar een doeltaal, waardoor content toegankelijk wordt voor een wereldwijd publiek.
- **Hoofdstuksegmentatie en sprekeridentificatie**: Segmenteert transcripties automatisch op basis van videohoofdstukken en ondersteunt AI-nabewerking voor sprekeridentificatie, wat gestructureerde en toegeschreven tekst oplevert.
- **Flexibele uitvoerformaten**: Genereert transcripties in Markdown met tijdstempels of SRT-ondertitelbestanden, geschikt voor diverse toepassingen van inhoudsanalyse tot videospelers.
- **Slimme caching voor efficiëntie**: Cachet ruwe videogegevens, metadata en gesegmenteerde transcripties, waardoor snelle herformattering mogelijk is en netwerkoproepen bij volgende verzoeken voor dezelfde video worden verminderd.

## Use cases

- **Genereer YouTube-transcripties voor inhoudsanalyse**: Contentmakers en onderzoekers kunnen snel volledige YouTube-transcripties verkrijgen, inclusief tijdstempels en hoofdstukmarkeringen, om video-inhoud te analyseren, belangrijke informatie te extraheren of gesproken inhoud om te zetten in tekstuele artikelen.
- **Maak ondertitelbestanden voor toegankelijkheid**: Video-editors en toegankelijkheidsspecialisten kunnen SRT-ondertitelbestanden genereren van YouTube-video's, zodat inhoud toegankelijk is voor slechthorende doelgroepen of voor degenen die de inhoud liever zonder geluid consumeren.
- **Vertaal video-inhoud voor wereldwijd bereik**: Marketeers en docenten kunnen YouTube-transcripties vertalen naar meerdere talen, waardoor het bereik van hun video-inhoud wordt uitgebreid naar niet-moedertaalsprekers en de wereldwijde betrokkenheid wordt verbeterd.
- **Metadata en omslagafbeeldingen extraheren**: Gebruikers kunnen eenvoudig videometadata en omslagafbeeldingen van hoge kwaliteit extraheren, handig voor catalogisering, promotie op sociale media of het maken van visuele middelen met betrekking tot de video-inhoud.

## Result preview

Ontdek een professioneel videoanalyserapport mogelijk gemaakt door deze vaardigheid.

![A presentation cover slide featuring a TED Talk analysis titled 'This Is How Kids Should Be Learning with AI'. The slide uses a bold red background with large white headline text centered prominently across the page. Beneath the title, the speaker and event are identified as 'Priya Lakhani | TEDNext 2025', accompanied by the subtitle 'TED Talk Analysis & Knowledge Report'. The lower portion of the slide contains a thumbnail image from the TED presentation showing the speaker on stage alongside the text 'AI Isn't a Shortcut to Learning'. A teal horizontal accent bar runs along the bottom edge, creating visual contrast. The overall design resembles a professional report or presentation cover summarizing key insights from a TED Talk about artificial intelligence and education.](https://file.nanoskill.ai/youtube-transcript-outcome1.png)

![A report page titled '\[ Executive Summary \]' from a TED Talk analysis on artificial intelligence and education. The page summarizes key arguments presented by education entrepreneur Priya Lakhani at TEDNext 2025, discussing the impact of AI in classrooms, the risks of students using AI to avoid learning, and the importance of productive struggle in effective education. Several paragraphs explain how neuroscience-informed AI systems can support deeper learning when designed to challenge rather than replace student effort. A section titled 'Key Statistics at a Glance' highlights major findings using four large color-coded statistic cards. The statistics include 20% of UK students leaving secondary school without adequate reading and writing skills, 74% of teachers considering leaving the profession within three years due to workload, one in five students using AI to complete all homework assignments, and over 40 billion data points collected on how children learn. The layout uses a clean report style with a red section header, structured text blocks, and colorful infographic-style data highlights to emphasize key educational challenges and opportunities.](https://file.nanoskill.ai/youtube-transcript-outcome2.png)

![A report page titled '\[ The Four Pillars of Learning \]' that presents four evidence-based learning principles supported by cognitive science and educational research. The page is organized into four color-coded sections: Retrieval Practice, Spacing, Generation, and Reflection. Each section includes a concise explanation of the learning principle and a highlighted 'Classroom Applications' box containing practical implementation examples. Retrieval Practice emphasizes recalling information from memory through activities such as flashcards, self-testing, and closed-book quizzes. Spacing focuses on distributing learning over time using spaced repetition and interleaved study sessions. Generation encourages learners to produce answers themselves through prediction, fill-in-the-blank exercises, and open-ended questioning. Reflection highlights structured self-assessment and feedback processes that help students evaluate progress, identify learning goals, and address knowledge gaps. The page uses a clean educational report layout with colored headers, explanatory text, and application callout boxes to summarize effective learning strategies for classrooms and AI-supported education.](https://file.nanoskill.ai/youtube-transcript-outcome3.png)

![A report page from a TED Talk analysis featuring two major sections: 'Neuroscience: The London Taxi Study' and 'AI in Education: Well-Designed vs. Poorly Used.' The first section explains a neuroscience study of London black cab drivers who memorize thousands of city streets to pass 'The Knowledge' exam. It describes how brain scans revealed that experienced drivers developed a larger hippocampus, demonstrating that sustained mental effort and navigation challenges can physically strengthen the brain. A highlighted key insight box emphasizes that mental effort is essential for durable learning, expertise development, and human creativity rather than being a flaw in the learning process. The second section presents a side-by-side comparison of educational AI usage. The left panel, labeled 'AI Well-Designed,' lists benefits such as identifying learning patterns, predicting forgetting, encouraging answer generation, providing targeted feedback, personalizing instruction, supporting teachers, and reducing workload. The right panel, labeled 'AI Poorly Used,' outlines risks including students using AI to complete all work, avoiding genuine learning, replacing thinking with shortcuts, creating false confidence, confusing fluency with understanding, and reducing productive struggle. The layout uses contrasting green and red comparison panels, educational report styling, and structured visual hierarchy to highlight the difference between AI that supports learning and AI that undermines it.](https://file.nanoskill.ai/youtube-transcript-outcome4.png)

## Result walkthrough

### Installeren

Voeg de YouTube Transcript Agent-vaardigheid toe aan uw AI-agent.

![A screenshot showing the installation process of the 'baoyu-youtube-transcript' AI agent skill from a GitHub repository using an NPX command. At the top, a blue command banner displays the installation command for adding the skill from the GitHub repository. Below, a conversational interface explains the installation workflow, including handling an interactive installation prompt, detecting that the process did not complete automatically, rerunning the command with a '-y' flag to bypass prompts, and explicitly targeting the Hermes Agent environment. The log then describes creating a symbolic link from the agent skills directory to the Hermes skills directory so the skill can be recognized by the Hermes Agent. A final confirmation states that the 'baoyu-youtube-transcript' skill was successfully installed, linked, and is available for use through a skill invocation command. The interface uses a clean chat-style layout with gray response panels, dark blue command highlighting, and monospaced code snippets for file paths and terminal commands.](https://file.nanoskill.ai/youtube-transcript-install.png)

### Content verstrekken

Deel een YouTube-link en specificeer het gewenste uitvoerformaat.

![A screenshot showing a prompt and response workflow for a YouTube transcript analysis skill. The top section contains a dark blue prompt panel describing an AI-powered content research and knowledge extraction system designed to analyze YouTube videos and transform them into structured knowledge assets. The prompt outlines requirements such as extracting complete transcripts, metadata, chapter structures, timestamps, speaker distinctions, key topics, major insights, statistics, examples, and actionable takeaways. It also lists supported output formats including executive summaries, study notes, blog articles, social media content, newsletter drafts, and research reports. Additional instructions emphasize transcript quality evaluation, improved readability, information hierarchy, redundancy removal, and generating a polished final deliverable. The requested output is a visually engaging PDF report with chapter navigation, highlighted takeaways, and presentation-ready layouts. Beneath the prompt, a gray response panel shows the AI acknowledging that the skill has been loaded successfully but explaining that a YouTube URL is still required before transcript extraction and report generation can begin. The interface uses a clean chat-style layout with large rounded panels, white text on a dark blue background, and structured instructional formatting.](https://file.nanoskill.ai/youtube-transcript-task.png)

### Inzichten extraheren

Genereer gestructureerde transcripties, samenvattingen, belangrijkste punten en herbruikbare content.

![A presentation cover slide for a TED Talk analysis and knowledge report. The slide features a bold red background with a teal accent bar running along the bottom edge. Centered at the top is the large white title 'This Is How Kids Should Be Learning with AI.' Below the title, the speaker attribution reads 'Priya Lakhani | TEDNext 2025,' followed by the subtitle 'TED Talk Analysis & Knowledge Report' in italicized white text. In the center of the slide is a thumbnail image from the TED Talk showing Priya Lakhani on stage. The thumbnail includes the prominent message 'AI Isn't a Shortcut to Learning' alongside the TED logo. The overall design resembles a professional research report cover, using strong typography, high contrast colors, and a clean layout to introduce an educational analysis focused on artificial intelligence, learning science, and the future of education.](https://file.nanoskill.ai/youtube-transcript-outcome.png)

## Skill definition

# YouTube Transcriptie

Downloadt transcripties (ondertitels/bijschriften) van YouTube-video's. Werkt met zowel handmatig gemaakte als automatisch gegenereerde transcripties. Geen API-sleutel of browser nodig — gebruikt direct de InnerTube API van YouTube en schakelt automatisch over naar `yt-dlp` wanneer YouTube het directe API-pad blokkeert.

Haalt video-metadata en omslagafbeelding op bij de eerste uitvoering, slaat ruwe gegevens op in cache voor snelle herformattering.

## Scriptmap

Scripts in de `scripts/` submap. `{baseDir}` = het mappad van deze SKILL.md. Bepaal `${BUN_X}` runtime: als `bun` is geïnstalleerd → `bun`; als `npx` beschikbaar is → `npx -y bun`; anders stel voor om bun te installeren. Vervang `{baseDir}` en `${BUN_X}` door werkelijke waarden.

| Script | Doel |
|--------|---------|
| `scripts/main.ts` | CLI voor transcriptiedownload |

## Gebruik

```bash
# Standaard: markdown met tijdstempels (Engels)
${BUN_X} {baseDir}/scripts/main.ts <youtube-url-or-id>

# Specificeer talen (prioriteitsvolgorde)
${BUN_X} {baseDir}/scripts/main.ts <url> --languages zh,en,ja

# Zonder tijdstempels
${BUN_X} {baseDir}/scripts/main.ts <url> --no-timestamps

# Met hoofdstuksegmentatie
${BUN_X} {baseDir}/scripts/main.ts <url> --chapters

# Met sprekeridentificatie (vereist AI-nabewerking)
${BUN_X} {baseDir}/scripts/main.ts <url> --speakers

# SRT-ondertitelbestand
${BUN_X} {baseDir}/scripts/main.ts <url> --format srt

# Transcriptie vertalen
${BUN_X} {baseDir}/scripts/main.ts <url> --translate zh-Hans

# Toon beschikbare transcripties
${BUN_X} {baseDir}/scripts/main.ts <url> --list

# Forceer opnieuw ophalen (negeer cache)
${BUN_X} {baseDir}/scripts/main.ts <url> --refresh
```

## Opties

| Optie | Beschrijving | Standaard |
|--------|-------------|---------|
| `<url-or-id>` | YouTube-URL of video-ID (meerdere toegestaan) | Vereist |
| `--languages <codes>` | Taalcodes, door komma's gescheiden, in prioriteitsvolgorde | `en` |
| `--format <fmt>` | Uitvoerformaat: `text`, `srt` | `text` |
| `--translate <code>` | Vertaal naar opgegeven taalcode | |
| `--list` | Toon beschikbare transcripties in plaats van op te halen | |
| `--timestamps` | Voeg `[HH:MM:SS → HH:MM:SS]` tijdstempels per alinea toe | aan |
| `--no-timestamps` | Schakel tijdstempels uit | |
| `--chapters` | Hoofdstuksegmentatie op basis van videobeschrijving | |
| `--speakers` | Ruwe transcriptie met metadata voor sprekeridentificatie | |
| `--exclude-generated` | Sla automatisch gegenereerde transcripties over | |
| `--exclude-manually-created` | Sla handmatig gemaakte transcripties over | |
| `--refresh` | Forceer opnieuw ophalen, negeer gegevens in cache | |
| `-o, --output <pad>` | Opslaan naar specifiek bestandspad | automatisch gegenereerd |
| `--output-dir <map>` | Basisuitvoermap | `youtube-transcript` |

## Optionele Omgevingsvariabelen

| Variabele | Beschrijving |
|----------|-------------|
| `YOUTUBE_TRANSCRIPT_COOKIES_FROM_BROWSER` | Doorgegeven aan `yt-dlp --cookies-from-browser` tijdens fallback, bijv. `chrome`, `safari`, `firefox`, of `chrome:Profile 1` |

## Invoerformaten

Accepteert een van de volgende als video-invoer:
- Volledige URL: `https://www.youtube.com/watch?v=dQw4w9WgXcQ`
- Verkorte URL: `https://youtu.be/dQw4w9WgXcQ`
- Insluit-URL: `https://www.youtube.com/embed/dQw4w9WgXcQ`
- Shorts-URL: `https://www.youtube.com/shorts/dQw4w9WgXcQ`
- Video-ID: `dQw4w9WgXcQ`

## Uitvoerformaten

| Formaat | Extensie | Beschrijving |
|--------|-----------|-------------|
| `text` | `.md` | Markdown met frontmatter (incl. `description`), titelkop, samenvatting, optionele inhoudsopgave/omslag/tijdstempels/hoofdstukken/sprekers |
| `srt` | `.srt` | SubRip-ondertitelformaat voor videospelers |

## Uitvoermap

```
youtube-transcript/
├── .index.json                          # Video-ID → mappadtoewijzing (voor cache-opzoekingen)
└── {channel-slug}/{title-full-slug}/
    ├── meta.json                        # Videometadata (titel, kanaal, beschrijving, duur, hoofdstukken, enz.)
    ├── transcript-raw.json              # Ruwe transcriptiefragmenten van YouTube API (gecached)
    ├── transcript-sentences.json        # In zinnen gesegmenteerd transcript (gesplitst door interpunctie, samengevoegd over fragmenten)
    ├── imgs/
    │   └── cover.jpg                    # Videominiatuur
    ├── transcript.md                    # Markdown-transcript (gegenereerd uit zinnen)
    └── transcript.srt                   # SRT-ondertiteling (gegenereerd uit ruwe fragmenten, indien --format srt)
```

- `{channel-slug}`: Kanaalnaam in kebab-case
- `{title-full-slug}`: Volledige videotitel in kebab-case

De modus `--list` voert alleen uit naar stdout (geen bestand opgeslagen).

## Caching

Bij de eerste ophaalactie slaat het script op:
- `meta.json` — videometadata, hoofdstukken, pad naar omslagafbeelding, taalinformatie
- `transcript-raw.json` — ruwe transcriptiefragmenten van YouTube API (`{ text, start, duration }[]`)
- `transcript-sentences.json` — in zinnen gesegmenteerd transcript (`{ text, start: "HH:mm:ss", end: "HH:mm:ss" }[]`), gesplitst door zinseindigende interpunctie (`.?!…。？！` enz.), tijdstempels proportioneel toegewezen op basis van karakterlengte, CJK-bewuste tekstsamenvoeging
- `imgs/cover.jpg` — videominiatuur

Latere uitvoeringen voor dezelfde video gebruiken gegevens uit de cache (geen netwerkaanroepen). Gebruik `--refresh` om opnieuw ophalen te forceren. Als een andere taal wordt aangevraagd, wordt de cache automatisch vernieuwd.

Wanneer YouTube anti-bot / geblokkeerde antwoorden retourneert op het directe InnerTube-pad, probeert het script opnieuw met alternatieve clientidentiteiten en schakelt vervolgens over naar `yt-dlp` indien beschikbaar. Als fallback nodig is maar `yt-dlp` niet beschikbaar is, moet de agent beslissen hoe hij `yt-dlp` beschikbaar maakt en doorgaan in plaats van de installatiebeslissing aan de gebruiker over te laten.

SRT-uitvoer (`--format srt`) wordt gegenereerd uit `transcript-raw.json`. Tekst/markdown-uitvoer gebruikt `transcript-sentences.json` voor natuurlijke zinsgrenzen.

## Workflow

Wanneer de gebruiker een YouTube-URL opgeeft en de transcriptie wil:

1. Voer eerst `--list` uit als de gebruiker geen taal heeft opgegeven, om beschikbare opties te tonen
2. **Gebruik altijd enkele aanhalingstekens voor de URL** bij het uitvoeren van het script — zsh behandelt `?` als een glob-wildcard, dus een niet-geciteerde YouTube-URL veroorzaakt "no matches found": gebruik `'https://www.youtube.com/watch?v=ID'`
3. Standaard: voer uit met `--chapters --speakers` voor de rijkste uitvoer (hoofdstukken + sprekeridentificatie)
3. Het script slaat automatisch gegevens in cache + uitvoerbestand op en toont het bestandspad
4. Voor de modus `--speakers`: nadat het script het ruwe bestand heeft opgeslagen, volg de onderstaande workflow voor sprekeridentificatie om na te bewerken met sprekerlabels

Wanneer de gebruiker alleen een omslagafbeelding of metadata wil, zal het uitvoeren van het script met elke optie ook `meta.json` en `imgs/cover.jpg` in de cache opslaan.

Bij het opnieuw formatteren van dezelfde video (bijv. eerst tekst, dan SRT), worden de gegevens in de cache hergebruikt — geen nieuwe ophaalactie nodig.

## Hoofdstuk- en sprekerworkflow

### Hoofdstukken (`--chapters`)

Het script ontleedt hoofdstuktijdstempels uit de videobeschrijving (bijv. `0:00 Introductie`), segmenteert de transcriptie op hoofdstukgrenzen, groepeert fragmenten in leesbare alinea's en slaat op als `.md` met een inhoudsopgave. Geen verdere verwerking nodig.

Als er geen hoofdstuktijdstempels in de beschrijving bestaan, wordt de transcriptie uitgevoerd als gegroepeerde alinea's zonder hoofdstukkoppen.

### Sprekeridentificatie (`--speakers`)

Sprekeridentificatie vereist AI-verwerking. Het script voert een ruw `.md`-bestand uit met:
- YAML-frontmatter met videometadata (titel, kanaal, datum, omslag, beschrijving, taal)
- Videobeschrijving (voor het extraheren van sprekersnamen)
- Hoofdstuklijst uit beschrijving (indien beschikbaar)
- Ruwe transcriptie in SRT-formaat (vooraf berekende start-/eindtijdstempels, token-efficiënt)

Nadat het script het ruwe bestand heeft opgeslagen, start een sub-agent (gebruik een goedkoper model zoals Sonnet voor kostenefficiëntie) om sprekeridentificatie te verwerken:

1. Lees het opgeslagen `.md`-bestand
2. Lees de promptsjabloon op `{baseDir}/prompts/speaker-transcript.md`
3. Verwerk de ruwe transcriptie volgens de prompt:
   - Identificeer sprekers met behulp van videometadata (titel → gast, kanaal → host, beschrijving → namen)
   - Detecteer sprekerwisselingen op basis van gespreksstroom, vraag-antwoordpatronen en contextuele aanwijzingen
   - Segmenteer in hoofdstukken (gebruik beschrijvingshoofdstukken indien beschikbaar, anders maak aan op basis van onderwerpverschuivingen)
   - Opmaak met `**Spreker Naam:**` labels, alineagroepering (2-4 zinnen) en `[HH:MM:SS → HH:MM:SS]` tijdstempels
4. Overschrijf het `.md`-bestand met de verwerkte transcriptie (behoud de YAML-frontmatter)

Wanneer `--speakers` wordt gebruikt, is `--chapters` impliciet — de verwerkte uitvoer bevat altijd hoofdstuksegmentatie.

## Foutgevallen

| Fout | Betekenis |
|-------|---------|
| Transcripties uitgeschakeld | Video heeft helemaal geen ondertitels |
| Geen transcriptie gevonden | Aangevraagde taal niet beschikbaar |
| Video niet beschikbaar | Video verwijderd, privé, of regiogeblokkeerd |
| IP geblokkeerd | Te veel verzoeken, probeer later opnieuw |
| Leeftijdsbeperking | Video vereist inloggen voor leeftijdsverificatie |
| bot gedetecteerd | Het script probeert opnieuw met alternatieve clients en vervolgens `yt-dlp`; als fallback-tooling ontbreekt, moet de agent dit zelf oplossen, anders als het nog steeds mislukt, probeer `YOUTUBE_TRANSCRIPT_COOKIES_FROM_BROWSER=safari` (of jouw browser) |

## FAQ

### Wat is de YouTube Transcript Downloader-vaardigheid?

De YouTube Transcript Downloader-vaardigheid is een tool waarmee je transcripties, ondertitels en omslagafbeeldingen van YouTube-video's kunt downloaden met alleen hun URL of video-ID. Het ondersteunt verschillende functies zoals meertalige ophaling, vertaling, hoofdstuksegmentatie en sprekeridentificatie.

### Hoe verkrijgt deze vaardigheid YouTube-transcripties zonder een API-sleutel?

De vaardigheid gebruikt rechtstreeks de Binnenbuis-API van YouTube om transcripties op te halen. Als directe toegang wordt geblokkeerd, schakelt het automatisch over op \`yt-dlp\` om betrouwbare transcriptieophaal te garanderen zonder dat een aparte API-sleutel nodig is.

### Kan ik transcripties krijgen in andere talen dan Engels?

Ja, je kunt een door komma's gescheiden lijst met taalcodes opgeven met de optie \`--languages\`. De vaardigheid zal proberen transcripties op te halen in de opgegeven prioriteitsvolgorde. Je kunt de transcriptie ook naar een andere taal vertalen met de optie \`--translate\`.

### Ondersteunt het sprekeridentificatie en hoofdstuksegmentatie?

Ja, de vaardigheid ondersteunt hoofdstuksegmentatie uit videobeschrijvingen met de optie \`--chapters\`. Voor sprekeridentificatie kun je de optie \`--speakers\` gebruiken, die een ruw bestand uitvoert voor AI-nabewerking om sprekers te labelen.

### Welke uitvoerformaten zijn beschikbaar voor de YouTube-transcriptie?

Je kunt de transcriptie uitvoeren in Markdown-indeling (\`.md\`), die tijdstempels, hoofdstukken en optionele sprekergegevens bevat, of als een SRT-ondertitelbestand (\`.srt\`), dat compatibel is met de meeste videospelers.

### Is er een cachingmechanisme en hoe werkt het?

Ja, de vaardigheid cachet videometadata, ruwe transcriptiegegevens en gesegmenteerde zinnen. Volgende runs voor dezelfde video gebruiken deze gecachte gegevens, waardoor de verwerking wordt versneld. Je kunt een hernieuwde ophaling forceren met de optie \`--refresh\`.
