ego (lite) is just a browser, ego is your personal agent across devices.
Join waitlist
Reddit scraping for posts and comments

Free Reddit scraper with your AI agent

ego (lite) is a free Reddit scraper for permitted posts, comment threads, and subreddit pages. Your AI agent works in a visible browser Space and exports CSV or Markdown with displayed post and comment data, source URLs, checked-at times, and access status.

Download for Mac(yes, free)

Trusted by developers from

OpenAIAnthropicGoogleMetaNVIDIACursorPerplexity
SpaceXTeslaNotionFigmaStripeNetflixAirbnb

How to scrape Reddit posts and comments in 5 steps

Start with a small, permitted source list and a precise output contract. ego (lite) gives your Agent the visible browser, separate Spaces, and a clear stopping point when Reddit requires login, consent, verification, or another human decision.

1

Install ego (lite) and choose an authorized browser context

Download ego (lite) for Mac, then choose only a Chrome context you are authorized to use. Public Reddit pages may load without signing in, while mature-content gates, quarantined communities, private communities, and account settings can change what the session displays. The Agent stays inside that visible access boundary.

2

Give your Agent the Reddit URLs, fields, and limits

Paste the permitted post, comment-thread, or subreddit URLs into the prompt. Name the fields you need, the sort order to preserve, the maximum number of posts and comments, and CSV or Markdown as the output. A fixed limit prevents a focused research task from turning into an open-ended crawl.

3

Watch the Agent read posts and comment threads

The Agent opens each permitted Reddit source in its own Space and records the title, body text, author as displayed, timestamp, score, comment count, and requested comments that the page visibly shows. Deleted, collapsed, unavailable, or gated content stays marked as such instead of being reconstructed.

4

Keep subreddit and thread tasks in separate Spaces

Give each subreddit, keyword set, or comment thread its own Space. Independent tasks can run side by side without mixing sort order or page state, while your normal tabs stay untouched. Open any Space to inspect progress, pause the task, or take over.

5

Review the Reddit scraper report

The final CSV or Markdown keeps one row per post or comment with its permalink, parent relationship, displayed values, checked-at time, and access status. You can filter the output without losing the link back to the page that supplied each row.

Why use ego (lite) for Reddit scraping?

A hosted Reddit scraper API returns rows from a remote service. ego (lite) is different: the coding Agent you already use works in a browser on your Mac, keeps tasks in visible Spaces, and leaves every collected post or comment tied to its source and status. There are no scraper credits or per-row charges.

Spend less time waiting on browser work

ego (lite) helps your Agent complete browser tasks faster. In the approved comparison shown here, ego (lite) finished in 81.8 seconds, compared with 282.9 seconds for an agent browser. Actual timing varies by site, workflow, and network conditions.

Approved task-time comparison between ego (lite) and an agent browser for browser data collection

Keep several Reddit research tasks separate

Assign each subreddit, search, or comment thread to its own Space. The Agent can work on independent sources side by side without mixing sort order, navigation state, or the page that supplied a row.

Parallel ego (lite) Spaces keeping independent Reddit research tasks separate

Monitor tasks and take over when Reddit needs you

Open any Space to see which permitted source the Agent is checking while other tasks continue. If one task reaches login, verification, CAPTCHA, or another human decision, take over that Space yourself and let the remaining work stay separate.

Use only the browser context you authorize

Import the Chrome context you choose. The Agent sees the same Reddit page state that session shows, and Reddit's login, community, age, and access rules still apply. It does not recover deleted content or enter gated areas on its own.

ego (lite) Chrome context import for an authorized Reddit browser session

What this Reddit scraper can and cannot do

ego (lite) records what a permitted Reddit page visibly displays in the browser context you choose. It is designed for bounded, supervised research. For applications, recurring pipelines, or collection at scale, use Reddit's official Data API and follow its terms and approval requirements.

What your Agent can record

Displayed fields from a permitted, bounded source list.

  • Post title, body text, author as displayed, subreddit, permalink, displayed timestamp, score, upvote ratio, award labels, and comment count
  • Visible comment text, author as displayed, score, timestamp, permalink, parent post, parent comment, and reply depth
  • The sort order and source URL used for each subreddit listing, search page, post, or comment thread
  • Checked-at time, access status, and limitations for every collected row

What stays out of scope

No hidden data, bypasses, engagement, or unbounded crawling.

  • Deleted or removed text, hidden scores, private or quarantined communities, private messages, moderator tools, or any page the session cannot access
  • Bypassing login, age gates, CAPTCHAs, rate limits, robots rules, community controls, or other access restrictions
  • Voting, posting, commenting, messaging, joining communities, following users, or any other engagement action
  • Open-ended subreddit crawling or production data feeds; use Reddit's official Data API for programmatic access

Read Reddit's User Agreement

Collect Reddit posts and comments you can verify later

Give your Agent a permitted source list, the exact fields you need, and a small item limit. ego (lite) keeps the browser work visible and returns a source-linked report with a clear status for every post and comment.

Try the free Reddit scraper workflow

Reddit scraper FAQ

A Reddit scraper is a tool or workflow that turns selected Reddit posts, comments, subreddit listings, or search results into structured data such as CSV or JSON. With ego (lite), your coding Agent reads a bounded list of permitted Reddit pages in a visible browser and records displayed fields together with each permalink, checked-at time, and access status.

Give Codex or Claude Code a prompt with the permitted Reddit URLs, the fields you need, and explicit limits for posts, comments, and reply depth. The Agent opens each source in its own ego (lite) Space, records what the page visibly displays, and returns CSV or Markdown. It stops when Reddit presents login, verification, a private community, a CAPTCHA, a rate limit, or another access decision.

It can record displayed post fields and the visible comments from sources you are permitted to collect. The report can preserve comment permalinks, parent-post and parent-comment relationships, reply depth, and the selected sort order. Collapsed, deleted, removed, hidden, or inaccessible comments stay marked unavailable rather than being guessed.

No. ego (lite) is a local Mac app, not a hosted Reddit scraper API, so it has no extraction endpoint, request credits, or per-row billing. If you need Reddit data inside production software or at scale, use Reddit's official Data API and comply with its OAuth, rate-limit, approval, retention, and usage requirements.

A browser can display some public Reddit pages without an API key, but technical access does not create permission to scrape. Reddit's User Agreement says scraping without prior written consent is prohibited and conditionally permits crawling only within its robots.txt parameters. Use this browser workflow only when your collection is permitted; otherwise use Reddit's official Data API or review the pages manually.

For posts: the title, body text, author as displayed, subreddit, permalink, displayed timestamp, score, upvote ratio, award labels, and comment count. For comments: visible text, author as displayed, score, timestamp, permalink, parent relationship, and reply depth. Every row also keeps source_url, checked_at, access_status, and a limitations note.

No. The workflow records only what the current page visibly displays. It does not reconstruct deleted or removed text, reveal hidden scores, identify deleted authors, expand comments the session cannot load, or access private messages and moderator data.

Not for every public page, but the visible result depends on the session. Mature-content gates, quarantined or private communities, preferences, and Reddit's own access controls can require an account or block a source. The Agent uses only the browser context you authorize and never enters credentials or joins a community on its own.

Reddit's current User Agreement prohibits scraping without prior written consent and permits crawling only according to its robots.txt parameters. Its Data API has separate terms, access requirements, limits, and approval paths. This is not legal advice: confirm that your collection and downstream use are permitted before running it, keep the task bounded, and stop at every platform control.

No. ‘Reddit scrapper’ is a common misspelling of ‘Reddit scraper.’ Both searches usually refer to software that collects Reddit posts or comments into structured data. This page uses the standard term ‘scraper’ and explains a visible, bounded workflow rather than an unobserved bulk crawler.