# YouTubeトランスクリプトエージェントスキル

> URLまたは動画IDでYouTube動画のトランスクリプト、字幕、カバー画像をダウンロードします。多言語対応、翻訳、チャプター、話者識別をサポートし、コンテンツのアクセシビリティを向上させます。数秒で無料で開始できます。

- Canonical: https://nanoskill.ai/ja/skills/youtube-transcript
- Markdown: https://nanoskill.ai/ja/skills/youtube-transcript.md
- Author: JimLiu
- Published: 2026-06-03T05:58:31.787Z
- Updated: 2026-07-15T13:36:15.347Z
- Language: ja
- Source type: github
- Popularity signal: 1040

## Sources

- https://github.com/JimLiu/baoyu-skills

## Install

```shell
npx skills add https://github.com/JimLiu/baoyu-skills/tree/main/skills/baoyu-youtube-transcript
```

## About

YouTubeトランスクリプトダウンローダースキルは、YouTube動画から包括的な情報を抽出するための堅牢なソリューションを提供します。このスキルを使用すると、動画のURLまたはIDを入力するだけで、完全なYouTubeトランスクリプト、字幕、さらにはカバー画像まで簡単にダウンロードできます。コンテンツクリエイター、研究者、音声コンテンツをテキストに変換する必要があるすべてのユーザー向けに設計されており、多言語サポート、翻訳機能、高度な構造化オプションなどの機能を備えています。

YouTubeのInnerTube APIへの直接アクセスと、\`yt-dlp\`へのスマートなフォールバックを活用することで、個人のAPIキーを必要とせずに信頼性の高い効率的なデータ取得を実現します。タイムスタンプとチャプターマーカー付きの詳細な分析に適したMarkdownや、標準的な字幕統合用のSRTなど、柔軟な出力形式を提供します。さらに、動画の説明からのチャプター分割をサポートし、AIによる話者識別のワークフローを提供することで、高度に整理され、属性付けされたテキストを提供します。

インテリジェントキャッシングにより、冗長なネットワークリクエストを最小限に抑え、同じ動画に対する後続の操作を非常に高速に行えます。動画コンテンツの分析、アクセシブルな字幕の作成、グローバルな視聴者向けの素材の翻訳、動画のメタデータやサムネイルの抽出など、あらゆるニーズに対応し、このスキルはプロセスを合理化し、YouTubeコンテンツ管理のための強力なツールを提供します。

## Key features

- **YouTubeへの直接アクセス**: YouTubeのInnerTube APIに直接アクセスして高速にトランスクリプトを取得し、直接APIがブロックされた場合は自動的に\`yt-dlp\`にフォールバックすることで、APIキーなしで信頼性の高いアクセスを実現します。
- **多言語サポートと翻訳**: トランスクリプトの優先言語を指定し、ターゲット言語に翻訳することで、コンテンツを世界中の視聴者にアクセス可能にします。
- **チャプター分割と話者識別**: ビデオチャプターごとにトランスクリプトを自動分割し、話者識別のためのAI後処理をサポートし、構造化され帰属が明確なテキストを提供します。
- **柔軟な出力形式**: タイムスタンプ付きのMarkdownまたはSRT字幕ファイルでトランスクリプトを生成し、コンテンツ分析からビデオプレーヤーまでさまざまな用途に適しています。
- **効率化のためのスマートキャッシング**: 生のビデオデータ、メタデータ、分割されたトランスクリプトをキャッシュし、同じビデオに対する後続のリクエストで迅速な再フォーマットを可能にし、ネットワーク呼び出しを削減します。

## Use cases

- **コンテンツ分析のためのYouTubeトランスクリプト生成**: コンテンツクリエイターやリサーチャーは、タイムスタンプやチャプターマーカーを含む完全なYouTubeトランスクリプトを迅速に取得し、ビデオコンテンツの分析、重要な情報の抽出、または音声コンテンツのテキスト記事への再利用ができます。
- **アクセシビリティのための字幕ファイル作成**: ビデオ編集者やアクセシビリティの専門家は、YouTubeビデオからSRT字幕ファイルを生成し、聴覚障害のある視聴者や音声なしでコンテンツを消費したい人にもアクセス可能にします。
- **グローバル展開のためのビデオコンテンツ翻訳**: マーケターや教育者は、YouTubeトランスクリプトを複数言語に翻訳することで、非ネイティブスピーカーへのビデオコンテンツのリーチを拡大し、グローバルなエンゲージメントを向上させます。
- **メタデータとカバー画像の抽出**: ユーザーはビデオのメタデータと高品質のカバー画像を簡単に抽出でき、カタログ作成、ソーシャルメディアプロモーション、ビデオコンテンツに関連するビジュアルアセットの作成に役立ちます。

## Result preview

このスキルによるプロフェッショナルな動画分析レポートをご覧ください。

![A presentation cover slide featuring a TED Talk analysis titled 'This Is How Kids Should Be Learning with AI'. The slide uses a bold red background with large white headline text centered prominently across the page. Beneath the title, the speaker and event are identified as 'Priya Lakhani | TEDNext 2025', accompanied by the subtitle 'TED Talk Analysis & Knowledge Report'. The lower portion of the slide contains a thumbnail image from the TED presentation showing the speaker on stage alongside the text 'AI Isn't a Shortcut to Learning'. A teal horizontal accent bar runs along the bottom edge, creating visual contrast. The overall design resembles a professional report or presentation cover summarizing key insights from a TED Talk about artificial intelligence and education.](https://file.nanoskill.ai/youtube-transcript-outcome1.png)

![A report page titled '\[ Executive Summary \]' from a TED Talk analysis on artificial intelligence and education. The page summarizes key arguments presented by education entrepreneur Priya Lakhani at TEDNext 2025, discussing the impact of AI in classrooms, the risks of students using AI to avoid learning, and the importance of productive struggle in effective education. Several paragraphs explain how neuroscience-informed AI systems can support deeper learning when designed to challenge rather than replace student effort. A section titled 'Key Statistics at a Glance' highlights major findings using four large color-coded statistic cards. The statistics include 20% of UK students leaving secondary school without adequate reading and writing skills, 74% of teachers considering leaving the profession within three years due to workload, one in five students using AI to complete all homework assignments, and over 40 billion data points collected on how children learn. The layout uses a clean report style with a red section header, structured text blocks, and colorful infographic-style data highlights to emphasize key educational challenges and opportunities.](https://file.nanoskill.ai/youtube-transcript-outcome2.png)

![A report page titled '\[ The Four Pillars of Learning \]' that presents four evidence-based learning principles supported by cognitive science and educational research. The page is organized into four color-coded sections: Retrieval Practice, Spacing, Generation, and Reflection. Each section includes a concise explanation of the learning principle and a highlighted 'Classroom Applications' box containing practical implementation examples. Retrieval Practice emphasizes recalling information from memory through activities such as flashcards, self-testing, and closed-book quizzes. Spacing focuses on distributing learning over time using spaced repetition and interleaved study sessions. Generation encourages learners to produce answers themselves through prediction, fill-in-the-blank exercises, and open-ended questioning. Reflection highlights structured self-assessment and feedback processes that help students evaluate progress, identify learning goals, and address knowledge gaps. The page uses a clean educational report layout with colored headers, explanatory text, and application callout boxes to summarize effective learning strategies for classrooms and AI-supported education.](https://file.nanoskill.ai/youtube-transcript-outcome3.png)

![A report page from a TED Talk analysis featuring two major sections: 'Neuroscience: The London Taxi Study' and 'AI in Education: Well-Designed vs. Poorly Used.' The first section explains a neuroscience study of London black cab drivers who memorize thousands of city streets to pass 'The Knowledge' exam. It describes how brain scans revealed that experienced drivers developed a larger hippocampus, demonstrating that sustained mental effort and navigation challenges can physically strengthen the brain. A highlighted key insight box emphasizes that mental effort is essential for durable learning, expertise development, and human creativity rather than being a flaw in the learning process. The second section presents a side-by-side comparison of educational AI usage. The left panel, labeled 'AI Well-Designed,' lists benefits such as identifying learning patterns, predicting forgetting, encouraging answer generation, providing targeted feedback, personalizing instruction, supporting teachers, and reducing workload. The right panel, labeled 'AI Poorly Used,' outlines risks including students using AI to complete all work, avoiding genuine learning, replacing thinking with shortcuts, creating false confidence, confusing fluency with understanding, and reducing productive struggle. The layout uses contrasting green and red comparison panels, educational report styling, and structured visual hierarchy to highlight the difference between AI that supports learning and AI that undermines it.](https://file.nanoskill.ai/youtube-transcript-outcome4.png)

## Result walkthrough

### インストール

YouTubeトランスクリプトエージェントスキルをAIエージェントに追加します。

![A screenshot showing the installation process of the 'baoyu-youtube-transcript' AI agent skill from a GitHub repository using an NPX command. At the top, a blue command banner displays the installation command for adding the skill from the GitHub repository. Below, a conversational interface explains the installation workflow, including handling an interactive installation prompt, detecting that the process did not complete automatically, rerunning the command with a '-y' flag to bypass prompts, and explicitly targeting the Hermes Agent environment. The log then describes creating a symbolic link from the agent skills directory to the Hermes skills directory so the skill can be recognized by the Hermes Agent. A final confirmation states that the 'baoyu-youtube-transcript' skill was successfully installed, linked, and is available for use through a skill invocation command. The interface uses a clean chat-style layout with gray response panels, dark blue command highlighting, and monospaced code snippets for file paths and terminal commands.](https://file.nanoskill.ai/youtube-transcript-install.png)

### コンテンツを提供

YouTubeリンクを共有し、希望する出力形式を指定します。

![A screenshot showing a prompt and response workflow for a YouTube transcript analysis skill. The top section contains a dark blue prompt panel describing an AI-powered content research and knowledge extraction system designed to analyze YouTube videos and transform them into structured knowledge assets. The prompt outlines requirements such as extracting complete transcripts, metadata, chapter structures, timestamps, speaker distinctions, key topics, major insights, statistics, examples, and actionable takeaways. It also lists supported output formats including executive summaries, study notes, blog articles, social media content, newsletter drafts, and research reports. Additional instructions emphasize transcript quality evaluation, improved readability, information hierarchy, redundancy removal, and generating a polished final deliverable. The requested output is a visually engaging PDF report with chapter navigation, highlighted takeaways, and presentation-ready layouts. Beneath the prompt, a gray response panel shows the AI acknowledging that the skill has been loaded successfully but explaining that a YouTube URL is still required before transcript extraction and report generation can begin. The interface uses a clean chat-style layout with large rounded panels, white text on a dark blue background, and structured instructional formatting.](https://file.nanoskill.ai/youtube-transcript-task.png)

### 洞察を抽出

構造化されたトランスクリプト、要約、主要なポイント、再利用可能なコンテンツを生成します。

![A presentation cover slide for a TED Talk analysis and knowledge report. The slide features a bold red background with a teal accent bar running along the bottom edge. Centered at the top is the large white title 'This Is How Kids Should Be Learning with AI.' Below the title, the speaker attribution reads 'Priya Lakhani | TEDNext 2025,' followed by the subtitle 'TED Talk Analysis & Knowledge Report' in italicized white text. In the center of the slide is a thumbnail image from the TED Talk showing Priya Lakhani on stage. The thumbnail includes the prominent message 'AI Isn't a Shortcut to Learning' alongside the TED logo. The overall design resembles a professional research report cover, using strong typography, high contrast colors, and a clean layout to introduce an educational analysis focused on artificial intelligence, learning science, and the future of education.](https://file.nanoskill.ai/youtube-transcript-outcome.png)

## Skill definition

# YouTube トランスクリプト

YouTube動画の字幕（サブタイトル/キャプション）をダウンロードします。手動作成された字幕と自動生成された字幕の両方に対応。APIキーやブラウザは不要 — YouTubeのInnerTube APIを直接使用し、直接APIパスがブロックされた場合は自動的に`yt-dlp`にフォールバックします。

初回実行時に動画のメタデータとカバー画像を取得し、生データをキャッシュして高速な再フォーマットを可能にします。

## スクリプトディレクトリ

スクリプトは `scripts/` サブディレクトリにあります。`{baseDir}` = この SKILL.md のディレクトリパス。`${BUN_X}` ランタイムの解決方法: `bun` がインストールされている場合 → `bun`; `npx` が利用可能な場合 → `npx -y bun`; それ以外の場合は bun のインストールを提案します。`{baseDir}` と `${BUN_X}` を実際の値に置き換えてください。

| スクリプト | 目的 |
|--------|---------|
| `scripts/main.ts` | トランスクリプトダウンロードCLI |

## 使い方

```bash
# デフォルト: タイムスタンプ付きマークダウン (英語)
${BUN_X} {baseDir}/scripts/main.ts <youtube-url-or-id>

# 言語を指定 (優先順)
${BUN_X} {baseDir}/scripts/main.ts <url> --languages zh,en,ja

# タイムスタンプなし
${BUN_X} {baseDir}/scripts/main.ts <url> --no-timestamps

# チャプター分割あり
${BUN_X} {baseDir}/scripts/main.ts <url> --chapters

# 話者識別あり (AI後処理が必要)
${BUN_X} {baseDir}/scripts/main.ts <url> --speakers

# SRT字幕ファイル
${BUN_X} {baseDir}/scripts/main.ts <url> --format srt

# トランスクリプトを翻訳
${BUN_X} {baseDir}/scripts/main.ts <url> --translate zh-Hans

# 利用可能なトランスクリプトを一覧表示
${BUN_X} {baseDir}/scripts/main.ts <url> --list

# 再取得を強制 (キャッシュを無視)
${BUN_X} {baseDir}/scripts/main.ts <url> --refresh
```

## オプション

| オプション | 説明 | デフォルト |
|--------|-------------|---------|
| `<url-or-id>` | YouTubeのURLまたは動画ID (複数指定可) | 必須 |
| `--languages <codes>` | 言語コード、カンマ区切り、優先順 | `en` |
| `--format <fmt>` | 出力形式: `text`, `srt` | `text` |
| `--translate <code>` | 指定された言語コードに翻訳 | |
| `--list` | トランスクリプトの一覧表示 (取得はしない) | |
| `--timestamps` | 段落ごとに `[HH:MM:SS → HH:MM:SS]` のタイムスタンプを含める | on |
| `--no-timestamps` | タイムスタンプを無効化 | |
| `--chapters` | 動画説明からチャプター分割 | |
| `--speakers` | 話者識別のためのメタデータ付き生トランスクリプト | |
| `--exclude-generated` | 自動生成されたトランスクリプトをスキップ | |
| `--exclude-manually-created` | 手動作成されたトランスクリプトをスキップ | |
| `--refresh` | 再取得を強制、キャッシュデータを無視 | |
| `-o, --output <path>` | 特定のファイルパスに保存 | 自動生成 |
| `--output-dir <dir>` | ベース出力ディレクトリ | `youtube-transcript` |

## オプションの環境変数

| 変数 | 説明 |
|----------|-------------|
| `YOUTUBE_TRANSCRIPT_COOKIES_FROM_BROWSER` | フォールバック時に `yt-dlp --cookies-from-browser` に渡されます。例: `chrome`, `safari`, `firefox`, または `chrome:Profile 1` |

## 入力形式

動画入力として以下を受け付けます:
- 完全なURL: `https://www.youtube.com/watch?v=dQw4w9WgXcQ`
- 短縮URL: `https://youtu.be/dQw4w9WgXcQ`
- 埋め込みURL: `https://www.youtube.com/embed/dQw4w9WgXcQ`
- ショートURL: `https://www.youtube.com/shorts/dQw4w9WgXcQ`
- 動画ID: `dQw4w9WgXcQ`

## 出力形式

| 形式 | 拡張子 | 説明 |
|--------|-----------|-------------|
| `text` | `.md` | フロントマター（`description`を含む）、タイトル見出し、要約、オプションの目次/カバー/タイムスタンプ/チャプター/話者付きマークダウン |
| `srt` | `.srt` | ビデオプレーヤー用のSubRip字幕形式 |

## 出力ディレクトリ

```
youtube-transcript/
├── .index.json                          # 動画ID → ディレクトリパスのマッピング (キャッシュ検索用)
└── {channel-slug}/{title-full-slug}/
    ├── meta.json                        # 動画メタデータ (タイトル、チャンネル、説明、長さ、チャプターなど)
    ├── transcript-raw.json              # YouTube APIからの生トランスクリプトスニペット (キャッシュ)
    ├── transcript-sentences.json        # 文分割されたトランスクリプト (句読点で分割、スニペット間でマージ)
    ├── imgs/
    │   └── cover.jpg                    # 動画サムネイル
    ├── transcript.md                    # マークダウントランスクリプト (文から生成)
    └── transcript.srt                   # SRT字幕 (生スニペットから生成、--format srt時)
```

- `{channel-slug}`: チャンネル名をケバブケースで
- `{title-full-slug}`: 完全な動画タイトルをケバブケースで

`--list` モードは標準出力のみに出力します (ファイル保存なし).

## キャッシュ

初回取得時に、スクリプトは以下を保存します:
- `meta.json` — 動画メタデータ、チャプター、カバー画像パス、言語情報
- `transcript-raw.json` — YouTube APIからの生トランスクリプトスニペット (`{ text, start, duration }[]`)
- `transcript-sentences.json` — 文分割されたトランスクリプト (`{ text, start: "HH:mm:ss", end: "HH:mm:ss" }[]`)、文末の句読点 (`.?!…。？！` など) で分割、タイムスタンプは文字数に比例して割り当て、CJK対応のテキストマージ
- `imgs/cover.jpg` — 動画サムネイル

同じ動画の後続の実行ではキャッシュデータを使用します (ネットワーク呼び出しなし)。再取得を強制するには `--refresh` を使用します。異なる言語が要求された場合、キャッシュは自動的に更新されます。

YouTubeが直接InnerTubeパスでアンチボット/ブロック応答を返す場合、スクリプトは代替クライアントIDで再試行し、その後利用可能であれば `yt-dlp` にフォールバックします。フォールバックが必要だが `yt-dlp` が利用できない場合、エージェントは `yt-dlp` を利用可能にする方法を自己判断し、ユーザーにインストールの決定を押し付けるのではなく続行する必要があります。

SRT出力 (`--format srt`) は `transcript-raw.json` から生成されます。テキスト/マークダウン出力では、自然な文の区切りに `transcript-sentences.json` を使用します。

## ワークフロー

ユーザーがYouTube URLを提供し、トランスクリプトを希望する場合:

1. ユーザーが言語を指定していない場合は、最初に `--list` を実行して利用可能なオプションを表示します
2. スクリプト実行時は **常にURLをシングルクォートで囲む** — zshは `?` をグロブのワイルドカードとして扱うため、クォートなしのYouTube URLは "no matches found" を引き起こします: `'https://www.youtube.com/watch?v=ID'` を使用します
3. デフォルト: 最もリッチな出力のために `--chapters --speakers` を付けて実行します (チャプター + 話者識別)
3. スクリプトはキャッシュデータと出力ファイルを自動保存し、ファイルパスを表示します
4. `--speakers` モードの場合: スクリプトが生ファイルを保存した後、以下の話者識別ワークフローに従って話者ラベルで後処理します

ユーザーがカバー画像またはメタデータのみを希望する場合、任意のオプションでスクリプトを実行すると、`meta.json` と `imgs/cover.jpg` もキャッシュされます。

同じ動画を再フォーマットする場合 (例: 最初にテキスト、次にSRT)、キャッシュデータが再利用されます — 再取得は不要です。

## チャプターと話者ワークフロー

### チャプター (`--chapters`)

スクリプトは動画説明からチャプターのタイムスタンプ (例: `0:00 Introduction`) を解析し、チャプター境界でトランスクリプトを分割、スニペットを読みやすい段落にグループ化し、目次付きの `.md` として保存します。これ以上の処理は不要です。

説明にチャプターのタイムスタンプが存在しない場合、トランスクリプトはチャプター見出しなしのグループ化された段落として出力されます。

### 話者識別 (`--speakers`)

話者識別にはAI処理が必要です。スクリプトは以下の内容を含む生の `.md` ファイルを出力します:
- 動画メタデータ (タイトル、チャンネル、日付、カバー、説明、言語) を含むYAMLフロントマター
- 話者名抽出のための動画説明
- 説明からのチャプターリスト (利用可能な場合)
- SRT形式の生トランスクリプト (事前計算された開始/終了タイムスタンプ、トークン効率的)

スクリプトが生ファイルを保存した後、サブエージェントを起動し (コスト効率のためにSonnetなどの安価なモデルを使用)、話者識別を処理します:

1. 保存された `.md` ファイルを読み取る
2. `{baseDir}/prompts/speaker-transcript.md` のプロンプトテンプレートを読み取る
3. プロンプトに従って生トランスクリプトを処理:
   - 動画メタデータを使用して話者を識別 (タイトル → ゲスト、チャンネル → ホスト、説明 → 名前)
   - 会話の流れ、質問-回答パターン、文脈上の手がかりから話者のターンを検出
   - チャプターに分割 (利用可能な場合は説明のチャプターを使用、なければトピックの移り変わりから作成)
   - `**話者名:**` ラベル、段落グループ化 (2-4文)、`[HH:MM:SS → HH:MM:SS]` タイムスタンプでフォーマット
4. 処理済みトランスクリプトで `.md` ファイルを上書き (YAMLフロントマターは保持)

`--speakers` を使用すると、`--chapters` が暗黙的に有効になります — 処理された出力には常にチャプター分割が含まれます。

## エラーケース

| エラー | 意味 |
|-------|---------|
| トランスクリプト無効 | 動画にキャプションがまったくない |
| トランスクリプトが見つかりません | 要求された言語が利用不可 |
| 動画が利用できません | 動画が削除されている、非公開、または地域制限あり |
| IPブロック | リクエストが多すぎます、後で再試行してください |
| 年齢制限 | 動画は年齢確認のためのログインが必要です |
| bot検出 | スクリプトは代替クライアントを再試行し、次に `yt-dlp` を試します。フォールバックツールが不足している場合、エージェントが自己解決する必要があります。それでも失敗する場合は、`YOUTUBE_TRANSCRIPT_COOKIES_FROM_BROWSER=safari` (またはお使いのブラウザ) を試してください |

## FAQ

### YouTube Transcript Downloaderスキルとは何ですか？

YouTube Transcript Downloaderスキルは、YouTubeビデオのURLまたはビデオIDを使用するだけで、トランスクリプト、字幕、カバー画像をダウンロードできるツールです。多言語取得、翻訳、チャプター分割、話者識別などの機能をサポートしています。

### このスキルはAPIキーなしでどのようにYouTubeトランスクリプトを取得しますか？

スキルは直接YouTubeのInnerTube APIを使用してトランスクリプトを取得します。直接アクセスがブロックされた場合は、自動的に\`yt-dlp\`にフォールバックし、別途APIキーを必要とせずに信頼性の高いトランスクリプト取得を実現します。

### 英語以外の言語でトランスクリプトを取得できますか？

はい、\`--languages\`オプションを使用して、言語コードのカンマ区切りリストを指定できます。スキルは指定された優先順位でトランスクリプトの取得を試みます。\`--translate\`オプションを使用して、トランスクリプトを別の言語に翻訳することもできます。

### 話者識別とチャプター分割をサポートしていますか？

はい、\`--chapters\`オプションを使用して、ビデオの説明からチャプター分割をサポートします。話者識別には\`--speakers\`オプションを使用できます。これはAI後処理で話者をラベル付けするための生ファイルを出力します。

### YouTubeトランスクリプトで利用可能な出力形式は何ですか？

タイムスタンプ、チャプター、オプションの話者データを含むMarkdown（\`.md\`）形式、またはほとんどのビデオプレーヤーと互換性のあるSRT（\`.srt\`）字幕ファイルとしてトランスクリプトを出力できます。

### キャッシングメカニズムはありますか？どのように機能しますか？

はい、スキルはビデオメタデータ、生のトランスクリプトデータ、および分割された文章をキャッシュします。同じビデオの以降の実行では、このキャッシュデータを使用して処理を高速化します。\`--refresh\`オプションを使用して再取得を強制できます。
