# 유튜브 트랜스크립트 에이전트 스킬

> 유튜브 영상 트랜스크립트, 자막, 커버 이미지를 URL이나 영상 ID로 다운로드하세요. 여러 언어 지원, 번역, 챕터, 화자 식별 기능으로 콘텐츠 접근성을 높이세요. 몇 초 안에 무료로 시작하세요.

- Canonical: https://nanoskill.ai/ko/skills/youtube-transcript
- Markdown: https://nanoskill.ai/ko/skills/youtube-transcript.md
- Author: JimLiu
- Published: 2026-06-03T05:58:31.787Z
- Updated: 2026-07-15T13:36:15.347Z
- Language: ko
- Source type: github
- Popularity signal: 1040

## Sources

- https://github.com/JimLiu/baoyu-skills

## Install

```shell
npx skills add https://github.com/JimLiu/baoyu-skills/tree/main/skills/baoyu-youtube-transcript
```

## About

유튜브 트랜스크립트 다운로더 스킬은 유튜브 영상에서 포괄적인 정보를 추출하는 강력한 솔루션을 제공합니다. 이 스킬을 사용하면 영상 URL이나 ID만 제공하여 전체 유튜브 트랜스크립트, 자막, 커버 이미지까지 손쉽게 다운로드할 수 있습니다. 콘텐츠 크리에이터, 연구자, 그리고 음성 콘텐츠를 텍스트로 변환해야 하는 모든 사람을 위해 설계되었으며, 다국어 지원, 번역 기능, 고급 구조화 옵션과 같은 기능을 제공합니다.

유튜브의 InnerTube API에 직접 접근하고 \`yt-dlp\`로의 지능적인 폴백을 활용하여, 이 스킬은 개인 API 키 없이도 안정적이고 효율적인 데이터 검색을 보장합니다. 타임스탬프와 챕터 마커가 포함된 상세 분석을 위한 Markdown, 표준 자막 통합을 위한 SRT 등 유연한 출력 형식을 제공합니다. 또한 영상 설명에서 챕터 분할을 지원하고 AI 기반 화자 식별을 위한 워크플로우를 제공하여, 매우 체계적이고 속성이 부여된 텍스트를 전달합니다.

지능형 캐싱을 통해 이 스킬은 중복 네트워크 요청을 최소화하여 동일 영상에 대한 후속 작업을 매우 빠르게 수행합니다. 영상 콘텐츠를 분석하거나, 접근 가능한 자막을 만들거나, 글로벌 시청자를 위해 자료를 번역하거나, 단순히 영상 메타데이터와 썸네일을 추출해야 하는 경우에도, 이 스킬은 프로세스를 간소화하여 유튜브 콘텐츠 관리를 위한 강력한 도구를 제공합니다.

## Key features

- **직접 YouTube 액세스**: 빠른 트랜스크립트 검색을 위해 YouTube의 InnerTube API에 직접 접근하며, 직접 API가 차단된 경우 \`yt-dlp\`로 자동 폴백하여 API 키 없이 안정적으로 액세스할 수 있습니다.
- **다국어 지원 및 번역**: 트랜스크립트의 선호 언어를 지정하고 대상 언어로 번역하여 전 세계 시청자가 콘텐츠에 접근할 수 있도록 합니다.
- **챕터 분할 및 화자 식별**: 비디오 챕터별로 트랜스크립트를 자동으로 분할하고, 화자 식별을 위한 AI 후처리를 지원하여 구조화되고 귀속된 텍스트를 제공합니다.
- **유연한 출력 형식**: 타임스탬프가 포함된 마크다운 또는 SRT 자막 파일로 트랜스크립트를 생성하여 콘텐츠 분석에서 비디오 플레이어까지 다양한 용도에 적합합니다.
- **효율성을 위한 스마트 캐싱**: 원시 비디오 데이터, 메타데이터 및 분할된 트랜스크립트를 캐싱하여 빠른 재포맷을 가능하게 하고 동일한 비디오에 대한 후속 요청에서 네트워크 호출을 줄입니다.

## Use cases

- **콘텐츠 분석을 위한 YouTube 트랜스크립트 생성**: 콘텐츠 제작자와 연구자는 타임스탬프와 챕터 마커를 포함한 전체 YouTube 트랜스크립트를 신속하게 확보하여 비디오 콘텐츠를 분석하거나, 핵심 정보를 추출하거나, 음성 콘텐츠를 텍스트 기사로 재활용할 수 있습니다.
- **접근성을 위한 자막 파일 생성**: 비디오 편집자와 접근성 전문가는 YouTube 비디오에서 SRT 자막 파일을 생성하여 청각 장애 시청자나 조용히 콘텐츠를 소비하려는 사람들이 콘텐츠에 접근할 수 있도록 보장할 수 있습니다.
- **글로벌 도달을 위한 비디오 콘텐츠 번역**: 마케터와 교육자는 YouTube 트랜스크립트를 여러 언어로 번역하여 비원어민 화자에게 동영상 콘텐츠의 도달 범위를 확장하고 전 세계 참여도를 높일 수 있습니다.
- **메타데이터 및 커버 이미지 추출**: 사용자는 비디오 메타데이터와 고품질 커버 이미지를 쉽게 추출할 수 있어 카탈로그 작성, 소셜 미디어 프로모션 또는 비디오 콘텐츠와 관련된 시각 자료 제작에 유용합니다.

## Result preview

이 스킬로 생성된 전문적인 영상 분석 보고서를 살펴보세요.

![A presentation cover slide featuring a TED Talk analysis titled 'This Is How Kids Should Be Learning with AI'. The slide uses a bold red background with large white headline text centered prominently across the page. Beneath the title, the speaker and event are identified as 'Priya Lakhani | TEDNext 2025', accompanied by the subtitle 'TED Talk Analysis & Knowledge Report'. The lower portion of the slide contains a thumbnail image from the TED presentation showing the speaker on stage alongside the text 'AI Isn't a Shortcut to Learning'. A teal horizontal accent bar runs along the bottom edge, creating visual contrast. The overall design resembles a professional report or presentation cover summarizing key insights from a TED Talk about artificial intelligence and education.](https://file.nanoskill.ai/youtube-transcript-outcome1.png)

![A report page titled '\[ Executive Summary \]' from a TED Talk analysis on artificial intelligence and education. The page summarizes key arguments presented by education entrepreneur Priya Lakhani at TEDNext 2025, discussing the impact of AI in classrooms, the risks of students using AI to avoid learning, and the importance of productive struggle in effective education. Several paragraphs explain how neuroscience-informed AI systems can support deeper learning when designed to challenge rather than replace student effort. A section titled 'Key Statistics at a Glance' highlights major findings using four large color-coded statistic cards. The statistics include 20% of UK students leaving secondary school without adequate reading and writing skills, 74% of teachers considering leaving the profession within three years due to workload, one in five students using AI to complete all homework assignments, and over 40 billion data points collected on how children learn. The layout uses a clean report style with a red section header, structured text blocks, and colorful infographic-style data highlights to emphasize key educational challenges and opportunities.](https://file.nanoskill.ai/youtube-transcript-outcome2.png)

![A report page titled '\[ The Four Pillars of Learning \]' that presents four evidence-based learning principles supported by cognitive science and educational research. The page is organized into four color-coded sections: Retrieval Practice, Spacing, Generation, and Reflection. Each section includes a concise explanation of the learning principle and a highlighted 'Classroom Applications' box containing practical implementation examples. Retrieval Practice emphasizes recalling information from memory through activities such as flashcards, self-testing, and closed-book quizzes. Spacing focuses on distributing learning over time using spaced repetition and interleaved study sessions. Generation encourages learners to produce answers themselves through prediction, fill-in-the-blank exercises, and open-ended questioning. Reflection highlights structured self-assessment and feedback processes that help students evaluate progress, identify learning goals, and address knowledge gaps. The page uses a clean educational report layout with colored headers, explanatory text, and application callout boxes to summarize effective learning strategies for classrooms and AI-supported education.](https://file.nanoskill.ai/youtube-transcript-outcome3.png)

![A report page from a TED Talk analysis featuring two major sections: 'Neuroscience: The London Taxi Study' and 'AI in Education: Well-Designed vs. Poorly Used.' The first section explains a neuroscience study of London black cab drivers who memorize thousands of city streets to pass 'The Knowledge' exam. It describes how brain scans revealed that experienced drivers developed a larger hippocampus, demonstrating that sustained mental effort and navigation challenges can physically strengthen the brain. A highlighted key insight box emphasizes that mental effort is essential for durable learning, expertise development, and human creativity rather than being a flaw in the learning process. The second section presents a side-by-side comparison of educational AI usage. The left panel, labeled 'AI Well-Designed,' lists benefits such as identifying learning patterns, predicting forgetting, encouraging answer generation, providing targeted feedback, personalizing instruction, supporting teachers, and reducing workload. The right panel, labeled 'AI Poorly Used,' outlines risks including students using AI to complete all work, avoiding genuine learning, replacing thinking with shortcuts, creating false confidence, confusing fluency with understanding, and reducing productive struggle. The layout uses contrasting green and red comparison panels, educational report styling, and structured visual hierarchy to highlight the difference between AI that supports learning and AI that undermines it.](https://file.nanoskill.ai/youtube-transcript-outcome4.png)

## Result walkthrough

### 설치

AI 에이전트에 유튜브 트랜스크립트 에이전트 스킬을 추가하세요.

![A screenshot showing the installation process of the 'baoyu-youtube-transcript' AI agent skill from a GitHub repository using an NPX command. At the top, a blue command banner displays the installation command for adding the skill from the GitHub repository. Below, a conversational interface explains the installation workflow, including handling an interactive installation prompt, detecting that the process did not complete automatically, rerunning the command with a '-y' flag to bypass prompts, and explicitly targeting the Hermes Agent environment. The log then describes creating a symbolic link from the agent skills directory to the Hermes skills directory so the skill can be recognized by the Hermes Agent. A final confirmation states that the 'baoyu-youtube-transcript' skill was successfully installed, linked, and is available for use through a skill invocation command. The interface uses a clean chat-style layout with gray response panels, dark blue command highlighting, and monospaced code snippets for file paths and terminal commands.](https://file.nanoskill.ai/youtube-transcript-install.png)

### 콘텐츠 제공

유튜브 링크를 공유하고 원하는 출력 형식을 지정하세요.

![A screenshot showing a prompt and response workflow for a YouTube transcript analysis skill. The top section contains a dark blue prompt panel describing an AI-powered content research and knowledge extraction system designed to analyze YouTube videos and transform them into structured knowledge assets. The prompt outlines requirements such as extracting complete transcripts, metadata, chapter structures, timestamps, speaker distinctions, key topics, major insights, statistics, examples, and actionable takeaways. It also lists supported output formats including executive summaries, study notes, blog articles, social media content, newsletter drafts, and research reports. Additional instructions emphasize transcript quality evaluation, improved readability, information hierarchy, redundancy removal, and generating a polished final deliverable. The requested output is a visually engaging PDF report with chapter navigation, highlighted takeaways, and presentation-ready layouts. Beneath the prompt, a gray response panel shows the AI acknowledging that the skill has been loaded successfully but explaining that a YouTube URL is still required before transcript extraction and report generation can begin. The interface uses a clean chat-style layout with large rounded panels, white text on a dark blue background, and structured instructional formatting.](https://file.nanoskill.ai/youtube-transcript-task.png)

### 인사이트 추출

구조화된 트랜스크립트, 요약, 핵심 요점, 재사용 가능한 콘텐츠를 생성하세요.

![A presentation cover slide for a TED Talk analysis and knowledge report. The slide features a bold red background with a teal accent bar running along the bottom edge. Centered at the top is the large white title 'This Is How Kids Should Be Learning with AI.' Below the title, the speaker attribution reads 'Priya Lakhani | TEDNext 2025,' followed by the subtitle 'TED Talk Analysis & Knowledge Report' in italicized white text. In the center of the slide is a thumbnail image from the TED Talk showing Priya Lakhani on stage. The thumbnail includes the prominent message 'AI Isn't a Shortcut to Learning' alongside the TED logo. The overall design resembles a professional research report cover, using strong typography, high contrast colors, and a clean layout to introduce an educational analysis focused on artificial intelligence, learning science, and the future of education.](https://file.nanoskill.ai/youtube-transcript-outcome.png)

## Skill definition

# 유튜브 자막

유튜브 동영상에서 자막(캡션)을 다운로드합니다. 수동으로 생성된 자막과 자동 생성된 자막 모두 작동합니다. API 키나 브라우저가 필요 없습니다 — 유튜브의 InnerTube API를 직접 사용하며, 유튜브가 직접 API 경로를 차단할 때는 자동으로 `yt-dlp`로 대체합니다.

첫 실행 시 동영상 메타데이터와 표지 이미지를 가져오고, 빠른 재포맷을 위해 원시 데이터를 캐시합니다.

## 스크립트 디렉터리

스크립트는 `scripts/` 하위 디렉터리에 있습니다. `{baseDir}` = 이 SKILL.md의 디렉터리 경로입니다. `${BUN_X}` 런타임 해결: `bun`이 설치되어 있으면 → `bun`; `npx`가 사용 가능하면 → `npx -y bun`; 그렇지 않으면 bun 설치를 제안합니다. `{baseDir}`와 `${BUN_X}`를 실제 값으로 교체하세요.

| 스크립트 | 목적 |
|--------|---------|
| `scripts/main.ts` | 자막 다운로드 CLI |

## 사용법

```bash
# Default: markdown with timestamps (English)
${BUN_X} {baseDir}/scripts/main.ts <youtube-url-or-id>

# Specify languages (priority order)
${BUN_X} {baseDir}/scripts/main.ts <url> --languages zh,en,ja

# Without timestamps
${BUN_X} {baseDir}/scripts/main.ts <url> --no-timestamps

# With chapter segmentation
${BUN_X} {baseDir}/scripts/main.ts <url> --chapters

# With speaker identification (requires AI post-processing)
${BUN_X} {baseDir}/scripts/main.ts <url> --speakers

# SRT subtitle file
${BUN_X} {baseDir}/scripts/main.ts <url> --format srt

# Translate transcript
${BUN_X} {baseDir}/scripts/main.ts <url> --translate zh-Hans

# List available transcripts
${BUN_X} {baseDir}/scripts/main.ts <url> --list

# Force re-fetch (ignore cache)
${BUN_X} {baseDir}/scripts/main.ts <url> --refresh
```

## 옵션

| 옵션 | 설명 | 기본값 |
|--------|-------------|---------|
| `<url-or-id>` | 유튜브 URL 또는 동영상 ID (여러 개 허용) | 필수 |
| `--languages <codes>` | 쉼표로 구분된 언어 코드, 우선 순위 순서 | `en` |
| `--format <fmt>` | 출력 형식: `text`, `srt` | `text` |
| `--translate <code>` | 지정된 언어 코드로 번역 | |
| `--list` | 다운로드 대신 사용 가능한 자막 목록 표시 | |
| `--timestamps` | 문단마다 `[HH:MM:SS → HH:MM:SS]` 타임스탬프 포함 | 켜짐 |
| `--no-timestamps` | 타임스탬프 비활성화 | |
| `--chapters` | 동영상 설명에서 챕터 분할 | |
| `--speakers` | 화자 식별을 위한 메타데이터가 포함된 원시 자막 | |
| `--exclude-generated` | 자동 생성된 자막 제외 | |
| `--exclude-manually-created` | 수동 생성된 자막 제외 | |
| `--refresh` | 강제로 다시 가져오기, 캐시된 데이터 무시 | |
| `-o, --output <path>` | 특정 파일 경로에 저장 | 자동 생성 |
| `--output-dir <dir>` | 기본 출력 디렉터리 | `youtube-transcript` |

## 선택적 환경 변수

| 변수 | 설명 |
|----------|-------------|
| `YOUTUBE_TRANSCRIPT_COOKIES_FROM_BROWSER` | 대체 시 `yt-dlp --cookies-from-browser`에 전달됩니다, 예: `chrome`, `safari`, `firefox`, 또는 `chrome:Profile 1` |

## 입력 형식

동영상 입력으로 다음 중 하나를 허용합니다:
- 전체 URL: `https://www.youtube.com/watch?v=dQw4w9WgXcQ`
- 짧은 URL: `https://youtu.be/dQw4w9WgXcQ`
- 임베드 URL: `https://www.youtube.com/embed/dQw4w9WgXcQ`
- 쇼츠 URL: `https://www.youtube.com/shorts/dQw4w9WgXcQ`
- 동영상 ID: `dQw4w9WgXcQ`

## 출력 형식

| 형식 | 확장자 | 설명 |
|--------|-----------|-------------|
| `text` | `.md` | 프론트매터(`description` 포함), 제목 헤딩, 요약, 선택적 목차/표지/타임스탬프/챕터/화자 정보가 포함된 마크다운 |
| `srt` | `.srt` | 동영상 플레이어용 SubRip 자막 형식 |

## 출력 디렉터리

```
youtube-transcript/
├── .index.json                          # Video ID → directory path mapping (for cache lookup)
└── {channel-slug}/{title-full-slug}/
    ├── meta.json                        # Video metadata (title, channel, description, duration, chapters, etc.)
    ├── transcript-raw.json              # Raw transcript snippets from YouTube API (cached)
    ├── transcript-sentences.json        # Sentence-segmented transcript (split by punctuation, merged across snippets)
    ├── imgs/
    │   └── cover.jpg                    # Video thumbnail
    ├── transcript.md                    # Markdown transcript (generated from sentences)
    └── transcript.srt                   # SRT subtitle (generated from raw snippets, if --format srt)
```

- `{channel-slug}`: 케밥 케이스의 채널 이름
- `{title-full-slug}`: 케밥 케이스의 전체 동영상 제목

`--list` 모드는 표준 출력으로만 출력합니다 (파일 저장 안 함).

## 캐싱

첫 가져오기 시 스크립트는 다음을 저장합니다:
- `meta.json` — 동영상 메타데이터, 챕터, 표지 이미지 경로, 언어 정보
- `transcript-raw.json` — 유튜브 API의 원시 자막 조각들 (`{ text, start, duration }[]`)
- `transcript-sentences.json` — 문장으로 분할된 자막 (`{ text, start: "HH:mm:ss", end: "HH:mm:ss" }[]`), 문장 종결 구두점(`.?!…。？！` 등)으로 분할, 타임스탬프는 문자 길이에 비례하여 할당, CJK 인식 텍스트 병합
- `imgs/cover.jpg` — 동영상 썸네일

동일한 동영상에 대한 이후 실행은 캐시된 데이터를 사용합니다 (네트워크 호출 없음). 강제로 다시 가져오려면 `--refresh`를 사용하세요. 다른 언어가 요청되면 캐시가 자동으로 새로 고쳐집니다.

유튜브가 직접 InnerTube 경로에서 안티봇/차단 응답을 반환하면 스크립트는 대체 클라이언트 ID로 재시도한 후, 가능한 경우 `yt-dlp`로 대체합니다. 대체가 필요하지만 `yt-dlp`를 사용할 수 없는 경우, 에이전트는 사용자에게 설치 결정을 미루지 않고 직접 `yt-dlp`를 사용 가능하게 만드는 방법을 결정해야 합니다.

SRT 출력(`--format srt`)은 `transcript-raw.json`에서 생성됩니다. 텍스트/마크다운 출력은 자연스러운 문장 경계를 위해 `transcript-sentences.json`을 사용합니다.

## 워크플로우

사용자가 유튜브 URL을 제공하고 자막을 원할 때:

1. 사용자가 언어를 지정하지 않은 경우 먼저 `--list`로 실행하여 사용 가능한 옵션을 표시합니다.
2. **스크립트를 실행할 때는 항상 URL을 작은따옴표로 묶으십시오** — zsh는 `?`를 글로브 와일드카드로 처리하므로, 따옴표가 없는 유튜브 URL은 "no matches found" 오류를 발생시킵니다: `'https://www.youtube.com/watch?v=ID'`를 사용하세요.
3. 기본값: 가장 풍부한 출력을 위해 `--chapters --speakers`로 실행합니다 (챕터 + 화자 식별).
3. 스크립트는 캐시된 데이터 + 출력 파일을 자동 저장하고 파일 경로를 출력합니다.
4. `--speakers` 모드의 경우: 스크립트가 원시 파일을 저장한 후, 아래의 화자 식별 워크플로우에 따라 화자 라벨로 후처리하세요.

사용자가 표지 이미지나 메타데이터만 원하는 경우, 어떤 옵션으로 스크립트를 실행해도 `meta.json`과 `imgs/cover.jpg`가 캐시됩니다.

동일한 동영상을 재포맷할 때(예: 먼저 텍스트, 그 다음 SRT), 캐시된 데이터가 재사용됩니다 — 다시 가져올 필요 없습니다.

## 챕터 및 화자 워크플로우

### 챕터 (`--chapters`)

스크립트는 동영상 설명에서 챕터 타임스탬프(예: `0:00 Introduction`)를 파싱하여 자막을 챕터 경계로 분할하고, 조각들을 읽기 쉬운 문단으로 그룹화하여 목차와 함께 `.md`로 저장합니다. 추가 처리가 필요하지 않습니다.

설명에 챕터 타임스탬프가 없는 경우, 자막은 챕터 제목 없이 그룹화된 문단으로 출력됩니다.

### 화자 식별 (`--speakers`)

화자 식별에는 AI 처리가 필요합니다. 스크립트는 다음을 포함하는 원시 `.md` 파일을 출력합니다:
- 동영상 메타데이터(제목, 채널, 날짜, 표지, 설명, 언어)가 포함된 YAML 프론트매터
- 동영상 설명 (화자 이름 추출용)
- 설명에서 가져온 챕터 목록 (가능한 경우)
- SRT 형식의 원시 자막 (미리 계산된 시작/종료 타임스탬프, 토큰 효율적)

스크립트가 원시 파일을 저장한 후, 하위 에이전트(비용 효율성을 위해 Sonnet과 같은 저렴한 모델 사용)를 생성하여 화자 식별을 처리하세요:

1. 저장된 `.md` 파일을 읽습니다.
2. `{baseDir}/prompts/speaker-transcript.md`에 있는 프롬프트 템플릿을 읽습니다.
3. 프롬프트에 따라 원시 자막을 처리합니다:
   - 동영상 메타데이터를 사용하여 화자 식별 (제목 → 게스트, 채널 → 호스트, 설명 → 이름)
   - 대화 흐름, 질문-답변 패턴, 맥락 단서에서 화자 전환 감지
   - 챕터로 분할 (가능한 경우 설명 챕터 사용, 그렇지 않으면 주제 전환에서 생성)
   - `**화자 이름:**` 라벨, 문단 그룹화(2-4문장), `[HH:MM:SS → HH:MM:SS]` 타임스탬프로 포맷
4. 처리된 자막으로 `.md` 파일을 덮어씁니다 (YAML 프론트매터 유지)

`--speakers`가 사용되면 `--chapters`가 암시됩니다 — 처리된 출력에는 항상 챕터 분할이 포함됩니다.

## 오류 사례

| 오류 | 의미 |
|-------|---------|
| 자막 비활성화됨 | 동영상에 캡션이 전혀 없습니다 |
| 자막을 찾을 수 없음 | 요청한 언어를 사용할 수 없습니다 |
| 동영상을 사용할 수 없음 | 동영상이 삭제되었거나 비공개이거나 지역 제한이 있습니다 |
| IP 차단됨 | 요청이 너무 많습니다. 나중에 다시 시도하세요 |
| 연령 제한됨 | 연령 확인을 위해 로그인이 필요합니다 |
| 봇 감지됨 | 스크립트가 대체 클라이언트를 재시도한 후 `yt-dlp`를 시도합니다; 대체 도구가 누락된 경우 에이전트가 자체적으로 해결해야 하며, 그래도 실패하면 `YOUTUBE_TRANSCRIPT_COOKIES_FROM_BROWSER=safari`(또는 사용하는 브라우저)를 시도하세요 |

## FAQ

### YouTube 트랜스크립트 다운로더 스킬이란 무엇인가요?

YouTube 트랜스크립트 다운로더 스킬은 URL이나 비디오 ID만을 사용하여 YouTube 비디오에서 트랜스크립트, 자막, 커버 이미지를 다운로드할 수 있게 해주는 도구입니다. 다국어 검색, 번역, 챕터 분할, 화자 식별 등 다양한 기능을 지원합니다.

### 이 스킬은 API 키 없이 어떻게 YouTube 트랜스크립트를 가져오나요?

이 스킬은 YouTube의 InnerTube API를 직접 사용하여 트랜스크립트를 가져옵니다. 직접 접근이 차단되면 \`yt-dlp\`로 자동 폴백하여 별도의 API 키 없이도 안정적으로 트랜스크립트를 검색할 수 있습니다.

### 영어 이외의 언어로 트랜스크립트를 받을 수 있나요?

네, \`--languages\` 옵션을 사용하여 쉼표로 구분된 언어 코드 목록을 지정할 수 있습니다. 스킬은 지정된 우선 순위에 따라 트랜스크립트를 가져오려고 시도합니다. \`--translate\` 옵션을 사용하여 트랜스크립트를 다른 언어로 번역할 수도 있습니다.

### 화자 식별과 챕터 분할을 지원하나요?

네, 이 스킬은 \`--chapters\` 옵션을 사용하여 비디오 설명에서 챕터 분할을 지원합니다. 화자 식별을 위해서는 \`--speakers\` 옵션을 사용할 수 있으며, 이는 AI 후처리를 통해 화자를 라벨링할 수 있는 원시 파일을 출력합니다.

### YouTube 트랜스크립트에 사용할 수 있는 출력 형식은 무엇인가요?

타임스탬프, 챕터 및 선택적 화자 데이터가 포함된 마크다운(\`.md\`) 형식이나 대부분의 비디오 플레이어와 호환되는 SRT(\`.srt\`) 자막 파일로 트랜스크립트를 출력할 수 있습니다.

### 캐싱 메커니즘이 있나요? 어떻게 작동하나요?

네, 스킬은 비디오 메타데이터, 원시 트랜스크립트 데이터 및 분할된 문장을 캐싱합니다. 동일한 비디오에 대한 이후 실행에서는 이 캐시된 데이터를 사용하여 처리 속도를 높입니다. \`--refresh\` 옵션을 사용하여 다시 가져오도록 강제할 수 있습니다.
