ego (lite) is just a browser, ego is your personal agent across devices.
Join waitlist
YouTube transcript workflow

YouTube transcript extractor with your AI agent (free)

Give Codex or Claude Code a list of YouTube URLs, your preferred caption language, and an output format. With ego (lite), your AI agent opens the videos in visible browser Spaces and returns a source-linked transcript report instead of an anonymous block of text. Videos with no available captions stay clearly marked instead of being filled with invented copy.

Download for Mac(yes, free)

Trusted by developers from

OpenAIAnthropicGoogleMetaNVIDIACursorPerplexity
SpaceXTeslaNotionFigmaStripeNetflixAirbnb

How to get a full transcript from a YouTube video (step by step)

YouTube shows a transcript panel for videos that have captions: open a video, expand the description, and choose “Show transcript”. This workflow has your Agent do exactly that in a real browser — for one video or a list — and save what the panel displays as files, with the source and caption status recorded for every video. The same steps cover what creators search as a YouTube script extractor: on YouTube, the “script” is the caption track the panel displays.

1

Install ego (lite) and choose your browser context

Download ego (lite), then import only the Chrome context you want to authorize. Many videos show captions without signing in, but what the transcript panel offers still depends on the video, your region, and your session — an age-restricted or members-only video follows YouTube's rules, not the Agent's.

2

Send the video list and transcript brief to your AI agent

Paste one video URL, several URLs, or a playlist URL into Codex or Claude Code. Add the language and output format you want. The recording shows the batch brief on the left while ego (lite) prepares dedicated browser Spaces on the right.

3

Watch each YouTube video open in its own Space

The Spaces overview keeps every video separate, then the recording moves through the individual YouTube pages while your AI agent runs the transcript task. You can see which video is active, open that Space, or take over the browser without losing the rest of the batch.

4

Review the source-linked transcript report

The finished report shown here gives each video its own section. The source URL, video length, caption type, and timestamped dialogue stay beside the extracted text, so you can trace every excerpt back to the YouTube page it came from.

Why use ego (lite) as a YouTube transcript extractor?

A paste-URL transcript site gives you text with no history. Used as a YouTube video transcript extractor, ego (lite) gives your own Agent a visible browser, one Space per video, and output that keeps every source, language, and access status — with no tokens, sign-up, or per-video metering.

Finish browser work about 3.5× faster

In this task-time comparison, ego (lite) finished in 81.8 seconds versus 282.9 seconds for an agent browser. Faster browser execution means less waiting while your AI agent opens each YouTube video, checks its captions, and moves on to the next transcript task.

Task-time comparison for ego (lite) and an agent browser

Keep each YouTube transcript task isolated

Use a separate Space for each video or playlist item. Transcript extraction stays organized instead of mixing YouTube pages and caption states across a pile of browser tabs, and you can return to the exact Space where a run stopped.

Separate ego (lite) Spaces for isolated YouTube transcript tasks

Watch the Agent and take over when YouTube needs you

Open any Space to see which video and caption track the Agent is reading. When YouTube asks for sign-in, age verification, or consent, the Agent stops and hands you the browser — you decide whether to complete the step yourself, and nothing is bypassed.

Reuse a browser context you are authorized to use

Import the Chrome context you choose instead of pasting URLs into an unknown website. Caption availability then matches what you actually see when you open the video yourself — and YouTube's sign-in, age, region, and membership rules still decide what the session can display.

ego (lite) Chrome context import for an authorized YouTube browser session

What this transcript extractor can and cannot do

YouTube documents that a transcript is available for videos that have captions, in the video's description panel. ego (lite) reads that panel in your browser session and records what it finds — it does not manufacture transcripts or unlock restricted videos.

What the Agent can record

Everything the transcript panel shows your session.

  • Transcript text with the timestamps the panel displays
  • Caption language, offered language tracks, and YouTube's auto-generated label
  • Video title, channel, source URL, and checked-at time for every video
  • The visible video list of a playlist page, processed one video at a time
  • An explicit status when captions are missing or a video is restricted

What stays out of scope

No generated captions, no bypassed restrictions, no media files.

  • Creating a transcript for a video that has no captions, or passing a summary off as one
  • Bypassing sign-in, age verification, region blocks, members-only access, or CAPTCHA
  • Downloading the video or audio, or ripping streams in any form
  • Proofreading auto-generated captions into guaranteed-accurate text
  • Accessing private or deleted videos your session cannot open

Read YouTube's guide to video transcripts

Extract YouTube video transcripts you can still trace next month

Give your Agent the video URLs, your caption language, and a format. Every transcript comes back with its source, language, and access status attached — and every video that returned nothing says why.

Try the transcript workflow

YouTube transcript extractor FAQ

Open the video, expand the description, and choose “Show transcript” — that panel holds the full transcript for any video with captions. This workflow automates exactly that: the Agent opens each video in an ego (lite) Space, reads the whole panel top to bottom, and saves the complete text with its timestamps, so you get the full transcript as a file instead of scrolling and copying it by hand.

You give Codex or Claude Code the prompt on this page with your video URLs. The Agent opens each video in an ego (lite) browser Space, opens YouTube's own “Show transcript” panel, and saves what the panel displays — title, channel, source URL, caption language and type, checked-at time, timestamps, and the text — as Markdown, TXT, or CSV files on your machine.

YouTube's transcript panel has no download button, so most people copy the text by hand. In this workflow the Agent does the copying and saves the result as the file format you asked for — one Markdown or TXT file per video, or a single CSV across a list. The download is the transcript and its source fields, not the video or audio.

The report records that video with the status “no captions available” and moves on. The Agent does not run speech recognition, summarize the video, or fill the gap another way — a missing transcript stays visibly missing, which also tells you the difference between a failed extraction and a video that never had captions.

Yes — use it as a YouTube video playlist transcript downloader that runs as an honest loop rather than a bulk API. The Agent opens the playlist page, lists the videos it can see, and processes them one at a time in separate Spaces. Long playlists take proportionally long, and the run summary shows the status of every video, so a restriction in the middle of the list never silently swallows the rest.

Yes. The “script” most creators want is the caption track YouTube already displays, so to extract script from YouTube videos you run the same prompt. The Agent reads the transcript panel and saves the text with its title, source URL, and caption labels — it does not rewrite the transcript into a polished production script for you.

Not reliably. YouTube's automatic captions often struggle with names, accents, jargon, and unclear audio, which is why the report keeps YouTube's own “auto-generated” label attached to those tracks. Treat an auto-generated transcript as a working copy to verify against the video, not as a checked document — this workflow does not proofread it for you.

Whichever languages the transcript panel offers for that video in your session. You state a preferred language in the prompt; when the video offers it, the Agent uses it, and when it does not, the Agent records which tracks were offered, uses the default, and notes the language actually extracted. It does not translate the transcript unless you separately ask for that.

Only to the extent your own session can already open them. When YouTube asks for sign-in, age verification, or membership, the Agent stops in that Space, records the restriction as the video's status, and hands the browser to you. It never enters credentials on its own or works around a block.

ego (lite) is a free Mac app, and the Agent doing the work is your own Codex or Claude Code. There are no extraction tokens, per-video credits, or sign-up walls in this workflow — what it can extract is limited by caption availability and your session's access, not by a meter.

The transcript of someone else's video is their content, so reuse depends on permission, licensing, and the copyright rules that apply to you — quoting briefly with attribution is treated very differently from republishing a full script. This workflow keeps the source URL and checked-at time with every transcript so you can attribute accurately, but it does not give legal advice.

No. The output is text and metadata: the transcript, title, channel, URL, caption language and type, timestamps, checked-at time, and status. The workflow does not download video or audio files, and it is not a tool for saving copies of media.