Reddit scrapers in 2026: five ways to get Reddit data, compared
Updated October 1, 2026
A Reddit scraper in 2026 is one of five things: a browser extension that exports the thread you have open, the official Data API, a hosted scraper such as an Apify actor, your own Python script, or copying by hand. Self-service API keys ended on 2025-11-11 and logged-out .json requests have returned 403 since 2026-05-28, so most people end up choosing between an extension and Apify, and outside an approved API app, Reddit's terms prohibit scraping without its written consent.
What changed in the last year
Tutorials on how to scrape Reddit written before these dates may no longer work.
| Date | What changed |
|---|---|
| 2025-11-11 | Reddit closed self-service sign-up for the Data API. Every new token now needs approval (r/redditdev announcement). |
| 2026-05-28 | Logged-out requests to .json URLs started returning 403. Reddit said logged-in and authenticated access was not affected (r/modnews announcement). The same day, Rule 8 was rewritten to list scraping without an authorized agreement as a violation (Rule 8). |
| 2026-08-05 | Reddit said it will gradually restrict new public API requests, move third-party apps to its Devvit platform, and asked every API app to register by 2026-09-30 (announcement). Reddit's CEO also said old Reddit may become logged-in only (post). |
| Aug–Sep 2026 | Some accounts got messages saying Reddit had detected a possible script or bot (r/redditdev). |
| Week of 2026-09-28 | New comment IDs got longer (up to 13 characters) and stopped increasing in order, which breaks code that checks ID length or sorts by ID (r/redditdev). |
We sent one logged-out .json request on 2026-09-30 to check. It came back 403.
What Reddit's terms say
Section 7 of Reddit's User Agreement (effective 2026-07-01) says "scraping the Services without Reddit's prior written consent is prohibited." Section 3 rules out commercial use of the service or its content without a written agreement. The Responsible Builder Policy (edited 2026-06-05) adds that Reddit data can't be sold or commercialized without written approval, bans inferring sensitive traits or de-anonymizing users, and says Reddit can block associated accounts and domains. Academic work has its own route, Reddit for Researchers, which requires IRB approval and a sponsoring institution.
That covers every method below except an approved API app, and it covers our own extension too. Comment Scraper reads only the thread you open, in your logged-in session, when you click, with at least 1.5 seconds between requests. That keeps it a long way from the large server-side scraping Reddit has taken to court. It does not give you Reddit's consent, and you are still bound by its User Agreement.
The five ways side by side
| Method | Works in 2026 | Cost | Risk to your Reddit account | Reddit's terms | Good for |
|---|---|---|---|---|---|
| Browser extension | Yes, while you're logged in | Free for single threads; paid plans $9.99–$12/month | Grows with volume, because requests use your session | No consent; you're bound by the User Agreement | A few threads at a time, AI analysis |
| Official Data API | Only with Reddit's approval | Free tier of 100 queries/minute; commercial use needs a contract, price not public | None | The approved route | Approved developers, companies with a contract |
| Apify actors | Yes, per their listings on 2026-09-30 | $1.19–$3.40 per 1,000 results | None, it runs on Apify's servers | Server-side scraping without consent | Search results and subreddit listings in volume |
| Python script | PRAW needs approved keys; logged-out .json returns 403 |
Free, plus your time | None with API keys; real if it uses your login | Fine only through an approved API app | Developers who already have API access |
| Copy by hand | Yes | Free, slow | Same as normal browsing | Manual collection isn't carved out | One short thread |
Browser extensions (Comment Scraper and similar)
An extension runs inside the Reddit tab you already have open and uses your logged-in session. It reads the thread the way your browser does, so it can open the collapsed "more replies" and "continue this thread" links that copy and paste leaves behind.
Comment Scraper is ours, so weigh this paragraph accordingly. It exports the Reddit or Hacker News thread you're on to CSV, JSON or Markdown, keeps reply depth, scores and permalinks, and leaves out author fields. Instead of claiming a full export, it shows three numbers: comments fetched, comments it knows it didn't fetch, and the count Reddit displays. Single-thread exports are free and need no account. The Personal plan is $12 a month or $108 a year and adds batch export of your open tabs and 60 AI-brief credits a month (pricing). It doesn't crawl subreddits, run searches or monitor keywords.
The most-installed Reddit export extension on the Chrome Web Store, Reddit Comment Scraper by Adlicio, had 5,000 users on 2026-09-30. Its Pro plan is $9.99 a month or $79.99 a year, Reddit only, up to 1,000 comments per export (pricing). Every other Reddit export extension we found had 1,000 users or fewer. General-purpose scrapers such as Thunderbit (100,000+ users, free for 6 pages a month, pricing) have Reddit templates as well. Before you rely on any of them, check whether it expands collapsed replies, keeps who-replied-to-whom and scores, tells you what it couldn't fetch, and where it stores your exports.
The risk sits with your Reddit account, because the requests go out under it. Since August 2026 Reddit has messaged some accounts about suspected scripts, and a few people who got the message said they were only running other extensions (r/redditdev). More volume means more risk. Comment Scraper only runs when you click, pauses when Reddit replies with a 429 rate-limit error, and caps how many threads you can export per day. That lowers the odds. It can't promise your account won't be flagged.
The official Reddit Data API
The Data API is the only route Reddit approves. The free tier allows 100 queries per minute per OAuth client ID, averaged over 10 minutes, and traffic without OAuth or login credentials is blocked (Data API Wiki, edited 2026-05-11).
Getting in is the hard part. Since 2025-11-11 every new token needs Reddit's approval, and people on r/redditdev report long waits or no reply at all (thread). Commercial use needs a separate contract, and Reddit's help article on data access counts subscription services, and free features that lead to a paid upgrade, as commercial. Prices aren't published. The only public reference is from 2023, when Apollo was quoted $12,000 per 50 million requests, about $0.24 per 1,000 (Ars Technica).
Reddit's 2026-08-05 announcement also moves third-party apps toward Devvit, whose rules don't allow sending users off-site or selling data (Devvit rules). If you're starting now, expect to wait.
Apify actors
Apify is a marketplace of hosted scrapers, called actors, that run on Apify's servers and hand back JSON or CSV. These were the main Reddit actors on 2026-09-30:
| Actor | Price | Users (total / monthly) |
|---|---|---|
| trudax/reddit-scraper-lite | $3.40 per 1,000 results | 44,904 / 7,462 |
| fatihtahta/reddit-scraper-search-fast | $1.19 per 1,000 results | 34,647 / 1,912 |
| harshmaur/reddit-scraper | $1.50–$2.00 per 1,000 ($0.02 per run plus $0.002 per item) | 14,552 / 3,716 |
Apify is what you want if you need search results or subreddit listings, which an extension like ours won't do. The listings recommend residential proxies, and the Reddit Scraper Lite docs say subreddit and search listings stop at around 1,000 items, while comments in a single thread aren't capped that way. Ratings vary a lot between actors, so look at recent runs before you pay.
Your Reddit account isn't involved. The terms question is different: this is server-side scraping without Reddit's consent, the pattern behind Reddit's October 2025 lawsuits against SerpApi, Oxylabs, AWMProxy and Perplexity (New York Times). On 2026-07-31 a judge in the Southern District of New York let Reddit's DMCA circumvention claim against SerpApi go ahead (Ars Technica). Reddit's Public Content Policy still says commercial use of its public content needs an agreement with Reddit.
Python scripts: PRAW and .json
There are two common approaches. PRAW is the Python library for the official API, so it needs approved API credentials, and since 2025-11-11 that means waiting on Reddit. One thing to know for big threads: replace_more() only expands 32 "load more" stubs by default, each one is a separate request, and each can return new stubs (PRAW docs).
The other approach was adding .json to a thread URL and fetching it from a script. Logged out, that has returned 403 since 2026-05-28, and people report being blocked at 30–60 requests an hour (r/webscraping). Making a script pass as your logged-in browser is scraping without an agreement under Rule 8, and the account whose session you use carries the risk. We don't recommend it.
Reddit scrapers on GitHub
A few that were active in 2026: ReddMark reads the comment elements on new Reddit pages, Validata handles both new and old Reddit markup, and RedditToMarkdown added a paste-your-own-JSON option on 2026-07-03. Check the last commit date before you depend on one. Code written before May 2026 may still assume logged-out .json works, and code written before late September may reject the new comment IDs. Either can fail without telling you.
Copying by hand
It's free, needs nothing installed, and is fine for a short thread you want to read once. It gets painful fast. Collapsed replies stay collapsed unless you click every link, the reply structure flattens when you paste, and scores and links fall off.
It also isn't carved out of the terms. Section 7 of the User Agreement covers collecting data by any means, manual included, unless the terms or a separate agreement allow it. The Reddit Pro FAQ says even Reddit Pro users can't download other people's posts and comments, by hand or automatically. The practical risk to your account is the same as reading the thread.
How to choose
Start with how many threads you need and what you'll do with them.
If you're reading a handful of threads for research, customer language or a Claude or ChatGPT session, use an extension. Single-thread exports are free in ours, and the Reddit-to-Claude guide shows the full workflow.
If you need thousands of posts from search results or whole subreddits, look at Apify, at $1.19–$3.40 per 1,000 results. Go in knowing it's server-side scraping without Reddit's consent.
If you're building a product that sells Reddit data or insights, talk to Reddit about a contract before you write code. GummySearch shut down because it couldn't reach an agreement with Reddit; our GummySearch alternatives page covers what its users moved to.
If you're an academic, apply to Reddit for Researchers. If you already hold approved API credentials, PRAW is the simplest path. If it's one short thread, copy it.
For a free Reddit scraper, your options are copying by hand, the free plan of an extension (ours needs no account for exports), open-source code from GitHub, and the free API tier if Reddit approves you.