# ทักษะตัวแทนการถอดข้อความ YouTube

> ดาวน์โหลดข้อความถอดเสียงจากวิดีโอ YouTube คำบรรยาย และภาพปกโดยใช้ URL หรือรหัสวิดีโอ รองรับหลายภาษา การแปลภาษา บท และการระบุผู้พูดเพื่อเพิ่มการเข้าถึงเนื้อหา เริ่มต้นใช้งานฟรีในไม่กี่วินาที

- Canonical: https://nanoskill.ai/th/skills/youtube-transcript
- Markdown: https://nanoskill.ai/th/skills/youtube-transcript.md
- Author: JimLiu
- Published: 2026-06-03T05:58:31.787Z
- Updated: 2026-07-15T13:36:15.347Z
- Language: th
- Source type: github
- Popularity signal: 1040

## Sources

- https://github.com/JimLiu/baoyu-skills

## Install

```shell
npx skills add https://github.com/JimLiu/baoyu-skills/tree/main/skills/baoyu-youtube-transcript
```

## About

ทักษะตัวดาวน์โหลดข้อความถอดเสียงจาก YouTube มอบโซลูชันที่แข็งแกร่งในการดึงข้อมูลที่ครอบคลุมจากวิดีโอ YouTube ทักษะนี้ช่วยให้ผู้ใช้สามารถดาวน์โหลดข้อความถอดเสียงแบบเต็ม คำบรรยาย และแม้แต่ภาพปกได้อย่างง่ายดาย เพียงแค่ระบุ URL หรือรหัสวิดีโอ ออกแบบมาสำหรับครีเอเตอร์เนื้อหา นักวิจัย และผู้ที่ต้องการแปลงเนื้อหาเสียงพูดเป็นข้อความ โดยมีฟีเจอร์ต่างๆ เช่น การรองรับหลายภาษา ความสามารถในการแปลภาษา และตัวเลือกการจัดโครงสร้างขั้นสูง

ด้วยการใช้การเข้าถึงโดยตรงไปยัง API InnerTube ของ YouTube และการสำรองข้อมูลอัจฉริยะด้วย \`yt-dlp\` ทักษะนี้รับประกันการดึงข้อมูลที่เชื่อถือได้และมีประสิทธิภาพโดยไม่ต้องใช้คีย์ API ส่วนบุคคล มีรูปแบบเอาต์พุตที่ยืดหยุ่น รวมถึง Markdown สำหรับการวิเคราะห์โดยละเอียดพร้อมการประทับเวลาและเครื่องหมายบท และ SRT สำหรับการรวมคำบรรยายมาตรฐาน นอกจากนี้ยังรองรับการแบ่งส่วนบทจากคำอธิบายวิดีโอ และให้ขั้นตอนการทำงานสำหรับการระบุผู้พูดด้วย AI ซึ่งส่งมอบข้อความที่มีการจัดระเบียบและระบุแหล่งที่มาอย่างดี

ด้วยแคชอัจฉริยะ ทักษะนี้ลดคำขอเครือข่ายที่ซ้ำซ้อน ทำให้การดำเนินการในภายหลังกับวิดีโอเดียวกันนั้นรวดเร็วอย่างมาก ไม่ว่าคุณจะต้องการวิเคราะห์เนื้อหาวิดีโอ สร้างคำบรรยายที่เข้าถึงได้ แปลเนื้อหาสำหรับผู้ชมทั่วโลก หรือเพียงแค่ดึงข้อมูลเมตาดาต้าและภาพขนาดย่อของวิดีโอ ทักษะนี้ทำให้กระบวนการคล่องตัวขึ้น เป็นเครื่องมือที่ทรงพลังสำหรับการจัดการเนื้อหา YouTube

## Key features

- **การเข้าถึง YouTube โดยตรง**: เข้าถึง YouTube's InnerTube API โดยตรงเพื่อการดึงบทบรรยายที่รวดเร็ว และเปลี่ยนไปใช้ \`yt-dlp\` โดยอัตโนมัติหาก API โดยตรงถูกบล็อก ทำให้มั่นใจได้ถึงการเข้าถึงที่เชื่อถือได้โดยไม่ต้องใช้คีย์ API
- **การรองรับหลายภาษาและการแปล**: ระบุภาษาที่ต้องการสำหรับบทบรรยายและแปลเป็นภาษาเป้าหมาย ทำให้เนื้อหาเข้าถึงได้สำหรับผู้ชมทั่วโลก
- **การแบ่งบทและการระบุผู้พูด**: แบ่งบทบรรยายตามบทของวิดีโอโดยอัตโนมัติ และรองรับการประมวลผลภายหลังด้วย AI สำหรับการระบุผู้พูด ให้ข้อความที่มีโครงสร้างและระบุแหล่งที่มา
- **รูปแบบผลลัพธ์ที่ยืดหยุ่น**: สร้างบทบรรยายในรูปแบบ Markdown พร้อมการประทับเวลา หรือไฟล์ซับไตเติ้ล SRT เหมาะสำหรับการใช้งานต่างๆ ตั้งแต่การวิเคราะห์เนื้อหาไปจนถึงเครื่องเล่นวิดีโอ
- **การแคชอัจฉริยะเพื่อประสิทธิภาพ**: แคชข้อมูลวิดีโอดิบ เมตาดาต้า และบทบรรยายที่แบ่งส่วนแล้ว ทำให้สามารถจัดรูปแบบใหม่ได้อย่างรวดเร็วและลดการเรียกเครือข่ายในคำขอครั้งต่อไปสำหรับวิดีโอเดียวกัน

## Use cases

- **สร้างบทบรรยายของ YouTube สำหรับการวิเคราะห์เนื้อหา**: ผู้สร้างเนื้อหาและนักวิจัยสามารถรับบทบรรยายฉบับเต็มของ YouTube ได้อย่างรวดเร็ว รวมถึงการประทับเวลาและเครื่องหมายบท เพื่อวิเคราะห์เนื้อหาวิดีโอ แยกข้อมูลสำคัญ หรือเปลี่ยนเนื้อหาที่พูดเป็นบทความข้อความ
- **สร้างไฟล์ซับไตเติ้ลเพื่อการเข้าถึง**: ผู้ตัดต่อวิดีโอและผู้เชี่ยวชาญด้านการเข้าถึงสามารถสร้างไฟล์ซับไตเติ้ล SRT จากวิดีโอ YouTube เพื่อให้แน่ใจว่าเนื้อหาสามารถเข้าถึงได้สำหรับผู้ชมที่มีความบกพร่องทางการได้ยินหรือผู้ที่ต้องการรับชมเนื้อหาแบบเงียบ
- **แปลเนื้อหาวิดีโอเพื่อการเข้าถึงทั่วโลก**: นักการตลาดและนักการศึกษาสามารถแปลบทบรรยายของ YouTube เป็นหลายภาษา ขยายการเข้าถึงเนื้อหาวิดีโอของพวกเขาไปยังผู้ที่ไม่ใช่เจ้าของภาษา และปรับปรุงการมีส่วนร่วมทั่วโลก
- **แยกเมตาดาต้าและภาพปก**: ผู้ใช้สามารถแยกเมตาดาต้าของวิดีโอและภาพปกคุณภาพสูงได้อย่างง่ายดาย ซึ่งมีประโยชน์สำหรับการจัดทำแคตตาล็อก การโปรโมตบนโซเชียลมีเดีย หรือการสร้างสินทรัพย์ภาพที่เกี่ยวข้องกับเนื้อหาวิดีโอ

## Result preview

สำรวจรายงานการวิเคราะห์วิดีโอระดับมืออาชีพที่ขับเคลื่อนโดยทักษะนี้

![A presentation cover slide featuring a TED Talk analysis titled 'This Is How Kids Should Be Learning with AI'. The slide uses a bold red background with large white headline text centered prominently across the page. Beneath the title, the speaker and event are identified as 'Priya Lakhani | TEDNext 2025', accompanied by the subtitle 'TED Talk Analysis & Knowledge Report'. The lower portion of the slide contains a thumbnail image from the TED presentation showing the speaker on stage alongside the text 'AI Isn't a Shortcut to Learning'. A teal horizontal accent bar runs along the bottom edge, creating visual contrast. The overall design resembles a professional report or presentation cover summarizing key insights from a TED Talk about artificial intelligence and education.](https://file.nanoskill.ai/youtube-transcript-outcome1.png)

![A report page titled '\[ Executive Summary \]' from a TED Talk analysis on artificial intelligence and education. The page summarizes key arguments presented by education entrepreneur Priya Lakhani at TEDNext 2025, discussing the impact of AI in classrooms, the risks of students using AI to avoid learning, and the importance of productive struggle in effective education. Several paragraphs explain how neuroscience-informed AI systems can support deeper learning when designed to challenge rather than replace student effort. A section titled 'Key Statistics at a Glance' highlights major findings using four large color-coded statistic cards. The statistics include 20% of UK students leaving secondary school without adequate reading and writing skills, 74% of teachers considering leaving the profession within three years due to workload, one in five students using AI to complete all homework assignments, and over 40 billion data points collected on how children learn. The layout uses a clean report style with a red section header, structured text blocks, and colorful infographic-style data highlights to emphasize key educational challenges and opportunities.](https://file.nanoskill.ai/youtube-transcript-outcome2.png)

![A report page titled '\[ The Four Pillars of Learning \]' that presents four evidence-based learning principles supported by cognitive science and educational research. The page is organized into four color-coded sections: Retrieval Practice, Spacing, Generation, and Reflection. Each section includes a concise explanation of the learning principle and a highlighted 'Classroom Applications' box containing practical implementation examples. Retrieval Practice emphasizes recalling information from memory through activities such as flashcards, self-testing, and closed-book quizzes. Spacing focuses on distributing learning over time using spaced repetition and interleaved study sessions. Generation encourages learners to produce answers themselves through prediction, fill-in-the-blank exercises, and open-ended questioning. Reflection highlights structured self-assessment and feedback processes that help students evaluate progress, identify learning goals, and address knowledge gaps. The page uses a clean educational report layout with colored headers, explanatory text, and application callout boxes to summarize effective learning strategies for classrooms and AI-supported education.](https://file.nanoskill.ai/youtube-transcript-outcome3.png)

![A report page from a TED Talk analysis featuring two major sections: 'Neuroscience: The London Taxi Study' and 'AI in Education: Well-Designed vs. Poorly Used.' The first section explains a neuroscience study of London black cab drivers who memorize thousands of city streets to pass 'The Knowledge' exam. It describes how brain scans revealed that experienced drivers developed a larger hippocampus, demonstrating that sustained mental effort and navigation challenges can physically strengthen the brain. A highlighted key insight box emphasizes that mental effort is essential for durable learning, expertise development, and human creativity rather than being a flaw in the learning process. The second section presents a side-by-side comparison of educational AI usage. The left panel, labeled 'AI Well-Designed,' lists benefits such as identifying learning patterns, predicting forgetting, encouraging answer generation, providing targeted feedback, personalizing instruction, supporting teachers, and reducing workload. The right panel, labeled 'AI Poorly Used,' outlines risks including students using AI to complete all work, avoiding genuine learning, replacing thinking with shortcuts, creating false confidence, confusing fluency with understanding, and reducing productive struggle. The layout uses contrasting green and red comparison panels, educational report styling, and structured visual hierarchy to highlight the difference between AI that supports learning and AI that undermines it.](https://file.nanoskill.ai/youtube-transcript-outcome4.png)

## Result walkthrough

### ติดตั้ง

เพิ่มทักษะตัวแทนการถอดข้อความ YouTube ให้กับตัวแทน AI ของคุณ

![A screenshot showing the installation process of the 'baoyu-youtube-transcript' AI agent skill from a GitHub repository using an NPX command. At the top, a blue command banner displays the installation command for adding the skill from the GitHub repository. Below, a conversational interface explains the installation workflow, including handling an interactive installation prompt, detecting that the process did not complete automatically, rerunning the command with a '-y' flag to bypass prompts, and explicitly targeting the Hermes Agent environment. The log then describes creating a symbolic link from the agent skills directory to the Hermes skills directory so the skill can be recognized by the Hermes Agent. A final confirmation states that the 'baoyu-youtube-transcript' skill was successfully installed, linked, and is available for use through a skill invocation command. The interface uses a clean chat-style layout with gray response panels, dark blue command highlighting, and monospaced code snippets for file paths and terminal commands.](https://file.nanoskill.ai/youtube-transcript-install.png)

### ระบุเนื้อหา

แชร์ลิงก์ YouTube และระบุรูปแบบเอาต์พุตที่คุณต้องการ

![A screenshot showing a prompt and response workflow for a YouTube transcript analysis skill. The top section contains a dark blue prompt panel describing an AI-powered content research and knowledge extraction system designed to analyze YouTube videos and transform them into structured knowledge assets. The prompt outlines requirements such as extracting complete transcripts, metadata, chapter structures, timestamps, speaker distinctions, key topics, major insights, statistics, examples, and actionable takeaways. It also lists supported output formats including executive summaries, study notes, blog articles, social media content, newsletter drafts, and research reports. Additional instructions emphasize transcript quality evaluation, improved readability, information hierarchy, redundancy removal, and generating a polished final deliverable. The requested output is a visually engaging PDF report with chapter navigation, highlighted takeaways, and presentation-ready layouts. Beneath the prompt, a gray response panel shows the AI acknowledging that the skill has been loaded successfully but explaining that a YouTube URL is still required before transcript extraction and report generation can begin. The interface uses a clean chat-style layout with large rounded panels, white text on a dark blue background, and structured instructional formatting.](https://file.nanoskill.ai/youtube-transcript-task.png)

### ดึงข้อมูลเชิงลึก

สร้างบทถอดเสียงที่มีโครงสร้าง บทสรุป ข้อสรุปสำคัญ และเนื้อหาที่นำกลับมาใช้ใหม่ได้

![A presentation cover slide for a TED Talk analysis and knowledge report. The slide features a bold red background with a teal accent bar running along the bottom edge. Centered at the top is the large white title 'This Is How Kids Should Be Learning with AI.' Below the title, the speaker attribution reads 'Priya Lakhani | TEDNext 2025,' followed by the subtitle 'TED Talk Analysis & Knowledge Report' in italicized white text. In the center of the slide is a thumbnail image from the TED Talk showing Priya Lakhani on stage. The thumbnail includes the prominent message 'AI Isn't a Shortcut to Learning' alongside the TED logo. The overall design resembles a professional research report cover, using strong typography, high contrast colors, and a clean layout to introduce an educational analysis focused on artificial intelligence, learning science, and the future of education.](https://file.nanoskill.ai/youtube-transcript-outcome.png)

## Skill definition

# ทรานสคริปต์ YouTube

ดาวน์โหลดทรานสคริปต์ (คำบรรยาย/แคปชัน) จากวิดีโอ YouTube ทำงานได้ทั้งคำบรรยายที่สร้างด้วยมือและสร้างโดยอัตโนมัติ ไม่ต้องใช้คีย์ API หรือเบราว์เซอร์ — ใช้ InnerTube API ของ YouTube โดยตรงและจะเปลี่ยนไปใช้ `yt-dlp` โดยอัตโนมัติหาก YouTube บล็อกเส้นทาง API โดยตรง

ดึงข้อมูลเมตาดาต้าของวิดีโอและภาพปกเมื่อเรียกใช้ครั้งแรก เก็บแคชข้อมูลดิบเพื่อการจัดรูปแบบใหม่ที่รวดเร็ว

## ไดเรกทอรีสคริปต์

สคริปต์ในไดเรกทอรีย่อย `scripts/` `{baseDir}` = เส้นทางไดเรกทอรีของ SKILL.md นี้ แปลงค่ารันไทม์ `${BUN_X}`: หากติดตั้ง `bun` → `bun`; หากมี `npx` → `npx -y bun`; มิฉะนั้นแนะนำให้ติดตั้ง bun แทนที่ `{baseDir}` และ `${BUN_X}` ด้วยค่าจริง

| สคริปต์ | วัตถุประสงค์ |
|--------|---------|
| `scripts/main.ts` | CLI สำหรับดาวน์โหลดทรานสคริปต์ |

## วิธีใช้

```bash
# ค่าเริ่มต้น: มาร์กดาวน์พร้อมตราประทับเวลา (อังกฤษ)
${BUN_X} {baseDir}/scripts/main.ts <youtube-url-or-id>

# ระบุภาษา (ลำดับความสำคัญ)
${BUN_X} {baseDir}/scripts/main.ts <url> --languages zh,en,ja

# ไม่มีตราประทับเวลา
${BUN_X} {baseDir}/scripts/main.ts <url> --no-timestamps

# แบ่งตามบท
${BUN_X} {baseDir}/scripts/main.ts <url> --chapters

# ระบุผู้พูด (ต้องใช้การประมวลผลภายหลังด้วย AI)
${BUN_X} {baseDir}/scripts/main.ts <url> --speakers

# ไฟล์คำบรรยาย SRT
${BUN_X} {baseDir}/scripts/main.ts <url> --format srt

# แปลทรานสคริปต์
${BUN_X} {baseDir}/scripts/main.ts <url> --translate zh-Hans

# แสดงรายการทรานสคริปต์ที่มี
${BUN_X} {baseDir}/scripts/main.ts <url> --list

# บังคับดึงใหม่ (ไม่ใช้แคช)
${BUN_X} {baseDir}/scripts/main.ts <url> --refresh
```

## ตัวเลือก

| ตัวเลือก | คำอธิบาย | ค่าเริ่มต้น |
|--------|-------------|---------|
| `<url-or-id>` | URL ของ YouTube หรือรหัสวิดีโอ (อนุญาตหลายรายการ) | จำเป็น |
| `--languages <codes>` | รหัสภาษา คั่นด้วยจุลภาค ตามลำดับความสำคัญ | `en` |
| `--format <fmt>` | รูปแบบผลลัพธ์: `text`, `srt` | `text` |
| `--translate <code>` | แปลเป็นรหัสภาษาที่ระบุ | |
| `--list` | แสดงรายการทรานสคริปต์ที่มีแทนที่จะดึง | |
| `--timestamps` | รวมตราประทับเวลา `[HH:MM:SS → HH:MM:SS]` ต่อย่อหน้า | เปิด |
| `--no-timestamps` | ปิดการใช้งานตราประทับเวลา | |
| `--chapters` | แบ่งบทจากคำอธิบายวิดีโอ | |
| `--speakers` | ทรานสคริปต์ดิบพร้อมเมตาดาต้าสำหรับระบุผู้พูด | |
| `--exclude-generated` | ข้ามทรานสคริปต์ที่สร้างโดยอัตโนมัติ | |
| `--exclude-manually-created` | ข้ามทรานสคริปต์ที่สร้างด้วยมือ | |
| `--refresh` | บังคับดึงใหม่ ไม่ใช้ข้อมูลแคช | |
| `-o, --output <path>` | บันทึกไปยังเส้นทางไฟล์ที่ระบุ | สร้างโดยอัตโนมัติ |
| `--output-dir <dir>` | ไดเรกทอรีผลลัพธ์ฐาน | `youtube-transcript` |

## ตัวแปรสภาพแวดล้อมเสริม

| ตัวแปร | คำอธิบาย |
|----------|-------------|
| `YOUTUBE_TRANSCRIPT_COOKIES_FROM_BROWSER` | ส่งผ่านไปยัง `yt-dlp --cookies-from-browser` ระหว่างการสำรอง เช่น `chrome`, `safari`, `firefox`, หรือ `chrome:Profile 1` |

## รูปแบบอินพุต

ยอมรับรูปแบบใดๆ เหล่านี้เป็นอินพุตวิดีโอ:
- URL เต็ม: `https://www.youtube.com/watch?v=dQw4w9WgXcQ`
- URL ย่อ: `https://youtu.be/dQw4w9WgXcQ`
- URL ฝัง: `https://www.youtube.com/embed/dQw4w9WgXcQ`
- URL Shorts: `https://www.youtube.com/shorts/dQw4w9WgXcQ`
- รหัสวิดีโอ: `dQw4w9WgXcQ`

## รูปแบบผลลัพธ์

| รูปแบบ | นามสกุล | คำอธิบาย |
|--------|-----------|-------------|
| `text` | `.md` | มาร์กดาวน์พร้อม frontmatter (รวม `description`), หัวเรื่อง, สรุป, TOC/ปก/ตราประทับ/บท/ผู้พูดตามตัวเลือก |
| `srt` | `.srt` | รูปแบบคำบรรยาย SubRip สำหรับโปรแกรมเล่นวิดีโอ |

## ไดเรกทอรีผลลัพธ์

```
youtube-transcript/
├── .index.json                          # การจับคู่รหัสวิดีโอ → เส้นทางไดเรกทอรี (สำหรับค้นหาแคช)
└── {channel-slug}/{title-full-slug}/
    ├── meta.json                        # เมตาดาต้าวิดีโอ (ชื่อ, ช่อง, คำอธิบาย, ระยะเวลา, บท, ฯลฯ)
    ├── transcript-raw.json              # ส่วนย่อยของทรานสคริปต์ดิบจาก YouTube API (แคช)
    ├── transcript-sentences.json        # ทรานสคริปต์ที่แบ่งเป็นประโยค (แยกตามเครื่องหมายวรรคตอน, รวมข้ามส่วนย่อย)
    ├── imgs/
    │   └── cover.jpg                    # ภาพปกขนาดย่อ
    ├── transcript.md                    # ทรานสคริปต์มาร์กดาวน์ (สร้างจากประโยค)
    └── transcript.srt                   # คำบรรยาย SRT (สร้างจากส่วนย่อยดิบ หากใช้ --format srt)
```

- `{channel-slug}`: ชื่อช่องในรูปแบบ kebab-case
- `{title-full-slug}`: ชื่อวิดีโอเต็มในรูปแบบ kebab-case

โหมด `--list` จะพิมพ์ผลลัพธ์ไปยัง stdout เท่านั้น (ไม่บันทึกไฟล์)

## การแคช

เมื่อดึงครั้งแรก สคริปต์จะบันทึก:
- `meta.json` — เมตาดาต้าวิดีโอ, บท, เส้นทางภาพปก, ข้อมูลภาษา
- `transcript-raw.json` — ส่วนย่อยของทรานสคริปต์ดิบจาก YouTube API (`{ text, start, duration }[]`)
- `transcript-sentences.json` — ทรานสคริปต์ที่แบ่งเป็นประโยค (`{ text, start: "HH:mm:ss", end: "HH:mm:ss" }[]`), แยกตามเครื่องหมายจบประโยค (`.?!…。？！` ฯลฯ), จัดสรรตราประทับเวลาตามสัดส่วนความยาวอักขระ, รวมข้อความโดยรู้จัก CJK
- `imgs/cover.jpg` — ภาพปกขนาดย่อ

การเรียกใช้ครั้งต่อไปสำหรับวิดีโอเดียวกันจะใช้ข้อมูลแคช (ไม่เรียกเครือข่าย) ใช้ `--refresh` เพื่อบังคับดึงใหม่ หากมีการร้องขอภาษาอื่น แคชจะถูกรีเฟรชโดยอัตโนมัติ

เมื่อ YouTube ตอบกลับด้วยการต่อต้านบอท / ถูกบล็อกบนเส้นทาง InnerTube โดยตรง สคริปต์จะลองใหม่ด้วยข้อมูลประจำตัวไคลเอนต์อื่นแล้วจึงเปลี่ยนไปใช้ `yt-dlp` หากมี หากจำเป็นต้องสำรองแต่ `yt-dlp` ไม่พร้อมใช้งาน เอเจนต์ควรตัดสินใจว่าจะทำให้ `yt-dlp` พร้อมใช้งานและดำเนินการต่ออย่างไรแทนที่จะผลักภาระการตัดสินใจติดตั้งให้ผู้ใช้

ผลลัพธ์ SRT (`--format srt`) สร้างจาก `transcript-raw.json` ผลลัพธ์ข้อความ/มาร์กดาวน์ใช้ `transcript-sentences.json` เพื่อขอบเขตประโยคที่เป็นธรรมชาติ

## เวิร์กโฟลว์

เมื่อผู้ใช้ระบุ URL ของ YouTube และต้องการทรานสคริปต์:

1. เรียกใช้ด้วย `--list` ก่อนหากผู้ใช้ไม่ได้ระบุภาษา เพื่อแสดงตัวเลือกที่มี
2. **ใส่ URL ในเครื่องหมายคำพูดเดี่ยวเสมอ** เมื่อเรียกใช้สคริปต์ — zsh ถือว่า `?` เป็นอักขระไวลด์การ์ดแบบ glob ดังนั้น URL YouTube ที่ไม่มีเครื่องหมายคำพูดจะทำให้เกิด "no matches found": ใช้ `'https://www.youtube.com/watch?v=ID'`
3. ค่าเริ่มต้น: เรียกใช้ด้วย `--chapters --speakers` เพื่อผลลัพธ์ที่สมบูรณ์ที่สุด (บท + การระบุผู้พูด)
3. สคริปต์จะบันทึกข้อมูลแคช + ไฟล์ผลลัพธ์โดยอัตโนมัติและพิมพ์เส้นทางไฟล์
4. สำหรับโหมด `--speakers`: หลังจากสคริปต์บันทึกไฟล์ดิบ ให้ทำตามเวิร์กโฟลว์การระบุผู้พูดด้านล่างเพื่อประมวลผลภายหลังด้วยป้ายกำกับผู้พูด

เมื่อผู้ใช้ต้องการเพียงภาพปกหรือเมตาดาต้า การเรียกใช้สคริปต์ด้วยตัวเลือกใดๆ จะแคช `meta.json` และ `imgs/cover.jpg` ด้วย

เมื่อจัดรูปแบบวิดีโอเดียวกันใหม่ (เช่น ข้อความก่อนแล้วจึง SRT) ข้อมูลที่แคชแล้วจะถูกใช้ซ้ำ — ไม่จำเป็นต้องดึงใหม่

## เวิร์กโฟลว์บทและผู้พูด

### บท (`--chapters`)

สคริปต์จะแยกวิเคราะห์ตราประทับเวลาของบทจากคำอธิบายวิดีโอ (เช่น `0:00 บทนำ`), แบ่งทรานสคริปต์ตามขอบเขตของบท, จัดกลุ่มส่วนย่อยเป็นย่อหน้าที่อ่านได้, และบันทึกเป็น `.md` พร้อมสารบัญ ไม่ต้องประมวลผลเพิ่มเติม

หากไม่มีตราประทับเวลาบทในคำอธิบาย ทรานสคริปต์จะถูกแสดงเป็นย่อหน้ากลุ่มโดยไม่มีหัวข้อบท

### การระบุผู้พูด (`--speakers`)

การระบุผู้พูดต้องใช้การประมวลผลด้วย AI สคริปต์จะสร้างไฟล์ `.md` ดิบที่ประกอบด้วย:
- frontmatter YAML พร้อมเมตาดาต้าวิดีโอ (ชื่อ, ช่อง, วันที่, ปก, คำอธิบาย, ภาษา)
- คำอธิบายวิดีโอ (สำหรับการสกัดชื่อผู้พูด)
- รายการบทจากคำอธิบาย (หากมี)
- ทรานสคริปต์ดิบในรูปแบบ SRT (ตราประทับเวลาเริ่ม/สิ้นสุดที่คำนวณล่วงหน้า, ประหยัดโทเค็น)

หลังจากสคริปต์บันทึกไฟล์ดิบ ให้สร้างซับเอเจนต์ (ใช้โมเดลที่ถูกกว่าเช่น Sonnet เพื่อประหยัดค่าใช้จ่าย) เพื่อประมวลผลการระบุผู้พูด:

1. อ่านไฟล์ `.md` ที่บันทึกไว้
2. อ่านเทมเพลตพร้อมต์ที่ `{baseDir}/prompts/speaker-transcript.md`
3. ประมวลผลทรานสคริปต์ดิบตามพร้อมต์:
   - ระบุผู้พูดโดยใช้เมตาดาต้าวิดีโอ (ชื่อ → แขก, ช่อง → โฮสต์, คำอธิบาย → ชื่อ)
   - ตรวจจับการเปลี่ยนผู้พูดจากการไหลของบทสนทนา, รูปแบบคำถาม-คำตอบ, และเบาะแสบริบท
   - แบ่งเป็นบท (ใช้บทจากคำอธิบายหากมี, มิฉะนั้นสร้างจากกะงานหัวข้อ)
   - จัดรูปแบบด้วยป้ายกำกับ `**ชื่อผู้พูด:**`, การจัดกลุ่มย่อหน้า (2-4 ประโยค), และตราประทับเวลา `[HH:MM:SS → HH:MM:SS]`
4. เขียนทับไฟล์ `.md` ด้วยทรานสคริปต์ที่ประมวลผลแล้ว (เก็บ frontmatter YAML ไว้)

เมื่อใช้ `--speakers`, `--chapters` จะถูกใช้โดยนัย — ผลลัพธ์ที่ประมวลผลจะรวมการแบ่งบทเสมอ

## กรณีข้อผิดพลาด

| ข้อผิดพลาด | ความหมาย |
|-------|---------|
| Transcripts disabled | วิดีโอไม่มีคำบรรยายเลย |
| No transcript found | ภาษาที่ร้องขอไม่มี |
| Video unavailable | วิดีโอถูกลบ, เป็นส่วนตัว, หรือถูกล็อคตามภูมิภาค |
| IP blocked | คำขอมากเกินไป, ลองใหม่ภายหลัง |
| Age restricted | วิดีโอต้องการการเข้าสู่ระบบเพื่อยืนยันอายุ |
| bot detected | สคริปต์จะลองไคลเอนต์อื่นแล้วจึง `yt-dlp`; หากเครื่องมือสำรองหายไป เอเจนต์ควรแก้ไขเอง, มิฉะนั้นหากยังล้มเหลวให้ลอง `YOUTUBE_TRANSCRIPT_COOKIES_FROM_BROWSER=safari` (หรือเบราว์เซอร์ของคุณ) |

## FAQ

### ทักษะการดาวน์โหลดบทบรรยายของ YouTube คืออะไร

ทักษะการดาวน์โหลดบทบรรยายของ YouTube เป็นเครื่องมือที่ช่วยให้คุณดาวน์โหลดบทบรรยาย ซับไตเติ้ล และภาพปกจากวิดีโอ YouTube โดยใช้เพียง URL หรือรหัสวิดีโอ รองรับคุณสมบัติต่างๆ เช่น การดึงข้อมูลหลายภาษา การแปล การแบ่งบท และการระบุผู้พูด

### ทักษะนี้ดึงบทบรรยายของ YouTube โดยไม่ต้องใช้คีย์ API ได้อย่างไร

ทักษะนี้ใช้ YouTube's InnerTube API โดยตรงเพื่อดึงบทบรรยาย หากการเข้าถึงโดยตรงถูกบล็อก จะเปลี่ยนไปใช้ \`yt-dlp\` โดยอัตโนมัติเพื่อให้แน่ใจว่าสามารถดึงบทบรรยายได้อย่างน่าเชื่อถือโดยไม่ต้องใช้คีย์ API แยกต่างหาก

### ฉันสามารถรับบทบรรยายในภาษาอื่นนอกเหนือจากภาษาอังกฤษได้หรือไม่

ได้ คุณสามารถระบุรายการรหัสภาษาที่คั่นด้วยเครื่องหมายจุลภาคโดยใช้ตัวเลือก \`--languages\` ทักษะจะพยายามดึงบทบรรยายตามลำดับความสำคัญที่ระบุ คุณยังสามารถแปลบทบรรยายเป็นภาษาอื่นโดยใช้ตัวเลือก \`--translate\`

### รองรับการระบุผู้พูดและการแบ่งบทหรือไม่

รองรับ ทักษะรองรับการแบ่งบทจากคำอธิบายวิดีโอโดยใช้ตัวเลือก \`--chapters\` สำหรับการระบุผู้พูด คุณสามารถใช้ตัวเลือก \`--speakers\` ซึ่งจะส่งออกไฟล์ดิบสำหรับการประมวลผลภายหลังด้วย AI เพื่อติดป้ายกำกับผู้พูด

### มีรูปแบบผลลัพธ์ใดบ้างสำหรับบทบรรยายของ YouTube

คุณสามารถส่งออกบทบรรยายในรูปแบบ Markdown (\`.md\`) ซึ่งรวมถึงการประทับเวลา บท และข้อมูลผู้พูดที่เป็นตัวเลือก หรือเป็นไฟล์ซับไตเติ้ล SRT (\`.srt\`) ซึ่งเข้ากันได้กับเครื่องเล่นวิดีโอส่วนใหญ่

### มีกลไกการแคชหรือไม่ และทำงานอย่างไร

มี ทักษะจะแคชข้อมูลเมตาดาต้าของวิดีโอ ข้อมูลบทบรรยายดิบ และประโยคที่แบ่งส่วนแล้ว การเรียกใช้ครั้งต่อไปสำหรับวิดีโอเดียวกันจะใช้ข้อมูลที่แคชไว้ ทำให้การประมวลผลเร็วขึ้น คุณสามารถบังคับให้ดึงข้อมูลใหม่ด้วยตัวเลือก \`--refresh\`
