Comment Scraper help
Updated October 1, 2026
Comment Scraper is a Chrome and Edge extension that exports the Reddit or Hacker News thread you have open to CSV, JSON or Markdown, and can write an AI brief from it. This page covers installing it, reading your first export, limits, credits, privacy and the problems people run into most.
Install
Comment Scraper works in Google Chrome and Microsoft Edge. It's the same extension in both stores: the Chrome Web Store and Microsoft Edge Add-ons. Use the install button on the home page to get the right one. There is no Firefox version yet.
After installing, pin the extension to your toolbar and click its icon to open the side panel. Everything happens there.
The extension asks for as little as it can: storage, the side panel, downloads, and access to the discussion sites it exports from. It doesn't ask for the "tabs" permission, and it doesn't collect your browsing history.
You don't need an account to export. You only need one for AI briefs and paid plans.
The first time you export, the extension shows a short note: exports use your own Reddit session, heavy use can trigger Reddit's rate limits or a message from Reddit, and you're responsible for following Reddit's terms.
Your first export
- Sign in to Reddit in the same browser. Hacker News doesn't need a sign-in.
- Open a thread on www.reddit.com or old.reddit.com. If you open a link to a single comment, the extension still exports the whole thread. The Export button isn't active on the front page, subreddit lists, user profiles or search results.
- Open the side panel, pick a format and click Export.
- Watch the progress bar and the three numbers. Keep the side panel open. Closing it pauses the export, and nothing runs in the background.
- When it stops, download the file or copy it to your clipboard.
The extension opens "more replies" and "continue this thread" links as it goes, one request at a time. Big threads take a while because of the spacing between requests (see Limits).
What the three numbers mean
We don't promise every comment in a thread. Instead you get three numbers, shown in the side panel, at the top of the file and on any brief made from it. For example: Fetched 812 · Known not fetched at least 0 · Reddit's count 845.
Fetched is the number of comments we got, including placeholders for comments marked [deleted] or [removed].
Known not fetched is the number we know we didn't get. Reddit shows a count on collapsed "more replies" links, and we add those up for any we didn't open. A "continue this thread" link doesn't say how many comments are behind it, so each one counts as 1. That's why this number reads "at least".
Reddit's count is the comment total Reddit displays for the thread.
Fetched plus known-not-fetched often won't match Reddit's count. The gap can be deleted comments, comments you aren't allowed to see, or comments we missed. We can't tell which, so we show the numbers and don't guess.
Each export also ends in one of three states:
- All known comments fetched: the extension found no collapsed links left to open.
- Partial: there are still unopened links, because of an error, a rate limit, a timeout, or because you stopped it. The reason is shown, and you can press Continue to resume.
- Limit reached: the export stopped at your plan's comment limit or request limit.
You can export what was fetched in any state.
Formats and fields
CSV is UTF-8 with a byte-order mark so Excel and Google Sheets open it without garbled characters. The thread itself is the first row, with record_type set to post. Every row also carries task_state (how the export ended) and fetched_at.
JSON comes as a nested tree or a flat array. Its header holds the full metadata: parse path, status, the three numbers, the reason if it's partial, and the fetch time.
Markdown is laid out for AI tools. Replies are indented under their parent, and each comment has an ID like [c12], its score and its age, so a model can cite specific comments. The header carries the same metadata as JSON.
Clipboard copies the Markdown plus one built-in prompt: pain points, objections and concerns, or quotes for ad copy.
All formats share the same fields (schema version 1):
| Field | What it holds |
|---|---|
schema_version |
Format version, currently 1 |
source |
The site the thread came from |
record_type |
post or comment |
thread_id |
ID of the thread |
thread_title |
Title of the thread |
subreddit |
Subreddit name, for Reddit threads |
comment_id |
ID of the comment |
parent_id |
ID of the comment it replies to; for top-level comments, the thread ID |
depth |
How deep the comment sits in the reply tree |
score |
Score as shown by Reddit |
score_hidden |
true when Reddit is hiding the score |
created_utc |
When it was posted, ISO 8601 in UTC |
edited_utc |
When it was last edited, ISO 8601 in UTC |
is_op |
true if the comment is by the person who started the thread |
body |
The comment text |
permalink |
Direct link to the comment |
comment_state |
normal, deleted or removed |
There are no author fields for posts or comments.
A few rules for empty and edge cases:
edited_utcis null when the comment was never edited.- When Reddit hides a score,
scoreis null andscore_hiddenis true. bodyis the comment's original Reddit Markdown. Usernames and links people typed inside a comment are kept as written.- Deleted and removed comments keep their place in the tree with an empty
body, so replies under them still have a parent. - If a comment's parent wasn't fetched,
parent_idis still filled in and the row getsparent_missing: true.
Treat comment_id as text. Starting the week of 2026-09-28, Reddit issues longer comment IDs (up to 13 characters) that no longer go up in order, so sort by created_utc, not by ID.
The extension normally reads Reddit's own JSON for the thread. If that fails, it reads the page instead. The file header says which path was used. On the page-reading path, body is plain text: line breaks are kept and links are written as "text (URL)".
To keep files safe to open, CSV cells that start with =, +, -, @, a tab or a carriage return get a leading apostrophe, so spreadsheets don't run them as formulas. Markdown exports escape # at the start of a line and code fences inside comments.
Limits
Plan limits, including comments per thread, threads per batch and threads per day, are on the pricing page. These rules apply on every plan.
Only one export runs at a time in your browser, and every request it makes goes through a single queue spaced at least 1.5 seconds apart. That's slower than it could be. It's deliberate, because the requests go out under your Reddit account.
There are daily caps on threads and on requests. They're counted per browser profile and reset at midnight your local time. The side panel always shows how much of today's allowance you've used. When you hit a cap, new exports won't start and a running export pauses until the next day.
If Reddit replies with a 429 (too many requests), the export pauses and retries after a wait, currently 30 seconds, then 60, then 120. If Reddit still says slow down after that, the export stops and shows a Continue button.
If a thread is bigger than your plan's per-thread limit, the extension fetches in Reddit's "Best" order until it reaches the limit, then stops with "Limit reached". The comments you get won't necessarily be the highest-scoring ones in the thread.
If the extension has to read the page instead of Reddit's JSON, it shows that it's in compatibility mode, and the per-thread limit may be lower until the normal path works again.
The extension keeps your recent exports in your browser: the last 10 on Free and the last 50 on paid plans, up to 200 MB in total. Each export is deleted 30 days after you made it, with a reminder 3 days before.
AI briefs and credits
A brief turns an export into themes, objections, requests, products people mention, a rough sentiment split and open questions. Each theme comes with up to 5 representative quotes, each with its score and a link back to the comment. The model returns only comment IDs; the extension then pulls the quote text from your export, so a brief can't contain a made-up quote. Counts and percentages are the model's estimates and are shown as "about". Only the quotes can be checked exactly.
Briefs need an account. Enter your email in the side panel and type the 6-digit code we send you.
On Free you get 3 briefs a month. Each one covers the highest-scoring fetched comments, up to 300 comments or 20,000 input tokens, with hidden scores sorted last. The top of the brief says exactly what was analysed, for example "the 300 highest-scoring of 1,240 fetched comments".
On paid plans a brief covers the whole export. It's split into chunks of up to 1,000 comments or 60,000 input tokens, whichever comes first. Each chunk is summarised, then the summaries are merged. One chunk costs 1 credit, the merge is free, and deleted comments with no text don't count.
| Thread | Chunks | Credits |
|---|---|---|
| 300 comments | 1 | 1 |
| 1,001 comments | 2 (1,000 + 1) | 2 |
| 5,000 comments | 5, then merged | 5 |
| 1,000 very long comments (over 60,000 tokens) | 2 | 2 |
A brief can have up to 10 chunks, about 10,000 comments. Bigger exports won't run as a single brief. Single comments longer than 4,000 characters are cut short in what the model sees; your export keeps the full text.
Before a brief runs, the side panel shows how many credits it will use, and nothing is sent until you confirm. You're charged only when the finished brief reaches you. If the merge step fails, nothing is charged and you can retry the merge once for free.
How credits work:
- Personal includes 60 credits a month. They arrive on the same day each month (the last day, in shorter months) and expire when the next batch arrives. Annual plans get the same monthly batches.
- The 7-day pass includes 100 credits that expire 7 days after you buy it. You can buy more than one; each has its own 7 days and 100 credits.
- Credits don't roll over. The oldest batch is used first.
- Free briefs don't use credits.
- All times are in UTC.
The first time you use a brief, a dialog explains where the comment text goes and which model provider handles it. You can turn off hosted briefs in settings at any time.
Privacy in short
Exports don't include author fields. They're stored in your browser and deleted after 30 days, and we never see them. When you run an AI brief, the comment text goes through our server to our model provider, OpenAI, and comes straight back. We don't store it. There's a button in settings to delete all local data. The full details are in the privacy policy.
Troubleshooting
Not logged in to Reddit
The extension won't export Reddit threads unless you're signed in to Reddit in that browser. Sign in, reload the thread and try again. We don't try to get past Reddit's login.
Rate limited
A 429 means Reddit wants fewer requests. The extension pauses and retries on its own. If it stops with a Continue button, wait a few minutes before pressing it. If it keeps happening, export fewer threads in a row and spread big jobs across the day.
"Reddit just changed its page"
This message means Reddit changed its page layout and the extension couldn't read the thread. Whatever was fetched is saved, marked partial and can be exported. Often we can fix it on our side without an extension update; when we can't, we ship a new version, which has to pass store review first. Check the status page for progress.
Export only partly fetched
Common causes: a rate limit, a dropped connection, a daily cap, or you closed or navigated away from the thread's tab, which pauses the export. Press Continue to resume from where it stopped.
Before resuming, the extension checks you're still signed in to the same Reddit account. If you signed out or switched accounts, the export has to start again from the beginning.
In a batch export (paid plans), a thread that fails 3 times is skipped and the others keep going. The file has a status column for each thread.
The Export button doesn't appear or stays greyed out
Check that you're on a thread page, not a subreddit list, profile or search page. Private subreddits aren't supported. Right after installing, the side panel may say it's connecting to the service; Reddit exports switch on once the extension has its settings from our server. If the extension can't reach our server for 24 hours, Reddit exports switch off until it reconnects. That's how we make sure a source can be switched off within a day if we ever have to.
FAQ
Is this allowed?
Reddit's User Agreement says scraping without Reddit's prior written consent is prohibited, and we don't have an agreement with Reddit. So we can't tell you it's allowed. What we can tell you is how the extension behaves: it reads only threads you open, in your own logged-in session, only when you click, at least 1.5 seconds between requests, and it never tries to defeat logins, CAPTCHAs or rate limits. You're responsible for following Reddit's terms. Hacker News threads are read through HN's public search API run by Algolia. None of this is legal advice; see the disclaimer.
Will my Reddit account get banned?
It can happen, and we can't promise it won't. Since August 2026 Reddit has sent some accounts a message saying it detected a possible script or bot. The extension keeps the pace slow and caps daily use to keep that risk down. If you get a message from Reddit, click "I got a Reddit warning" in the side panel and slow down. If enough people report warnings, we tighten the limits for everyone.
Why no usernames?
Exports leave out author fields for posts and comments on purpose. Most research doesn't need to know who said something, and Reddit's Responsible Builder Policy forbids de-anonymizing users. Usernames someone typed inside a comment stay, because we don't edit comment text. The is_op field still tells you when a reply came from the person who started the thread.
Where is my data?
Your exports are in your browser's extension storage (IndexedDB) and are deleted after 30 days, or sooner if you use "Delete all local data". Files you download are on your computer. For AI briefs, comment text passes through our server's memory to the model provider and isn't stored. Our database holds your email, plan, credits and usage counts, never comment text. Details are in the privacy policy.
Anything else: hello@comment-scraper.com.