How to Scrape Amazon Reviews: What Still Works Now
See how to scrape Amazon reviews in 2026 using practical no-code methods, plus what Amazon now restricts and how to keep your review data useful at scale.

Gunnar
Last updated -
Tutorials

Tired of Amazon blocking your review scrapers? Explore proven approaches using APIs, browser add-ons, and residential proxies to collect product data fast.
Amazon reviews are one of the most valuable sources of market research on the web today. After all, customers tell you, in no uncertain terms, why a product made their day or let them down, which lets businesses double down on what works.
Gathering review data at scale is another story, especially with Amazon’s latest guardrails. Amazon began restricting anonymous access to its dedicated review pages in late 2024.
Today, signed-out scrapers typically get only around 8 to 13 review samples Amazon shows on the product page while deeper review pages require a logged-in session. And that’s just one of several guardrails Amazon uses to deter scrapers.
Even so, DIY scrapers, dedicated APIs, and managed scraping tools can still collect mass public review data reliably when you configure your stack correctly. This guide explains how to scrape Amazon reviews in 2026, including which methods remain effective and what limits to expect.
Which Amazon Review Data Is Still Publicly Available?
Before choosing a scraper, it helps to separate two very different targets: the reviews Amazon shows publicly on a product page and the much deeper review archive behind its dedicated review pages.
The public product-page review sample
Signed-out visitors can still see a small set of full customer reviews on many Amazon’s product pages. A tool like ScraperAPI, for instance, retrieves publicly available reviews seamlessly through its Amazon Product API.
Scraping tools commonly report somewhere around 8 to 16 reviews per product, although there’s no fixed number. Your results can vary by Amazon Standard Identification Number (ASIN), marketplace, session, and whatever else Amazon decides to render that day.
That said, these public review samples still contain plenty of useful data, including:
Star rating
Review title and text
Review date and country
Verified Purchase badge
Helpful-vote count
Product variant, like size or color
Review ID, which becomes your deduplication key later
It’s worth noting that Amazon often leads its product pages with “Top reviews,” which Amazon describes as the ones other shoppers found most helpful, so it leans toward detailed, well-voted reviews rather than objective impartiality.
For small-scale research like finding a competitor’s biggest complaint or comparing star ratings across many ASINs, the public review sample is often good enough to work with. If, however, you need to track customer sentiment monthly or build training datasets, you’ll likely need the deeper review archive.
The full review archive
Since late 2024, Amazon’s dedicated review pages have generally required users to sign in. ScraperAPI confirmed the change at the time and subsequently made its standalone Amazon Reviews API unavailable because it doesn’t scrape behind a login.

Source: ScraperAPI
Consequently, web scraping guides that tell you to simply loop and increment a “pageNumber” parameter to pull thousands of reviews are now obsolete. Some tools can work with authenticated sessions, but that’s a different workflow with extra account, session, and policy considerations.
In short, it helps to answer a practical question first: “Is the public review sample enough for your use case, or do you need the deeper historical coverage?” Your answer determines which scraping method makes sense.
3 Ways to Scrape Amazon Reviews (and Who Each One Suits)
Once you know how much review data you need, there are three fundamental ways to collect it. The difference is mostly how much of the scraping stack you want to own yourself. Here are three primary scraping methods at a glance:
| Method | What you do | What the provider handles | Typical output | Best for |
|---|---|---|---|---|
| DIY custom scraper | Build the request, browser, parsing, retry, proxy, and validation logic | Nothing | Whatever schema you design | Data teams that want maximum control |
| Managed scraping platform | Configure URLs, workflows, fields, and schedules | Cloud infrastructure, browsers, proxies, retries, and scaling | HTML or structured datasets | Teams that want flexibility without running the infrastructure |
| Dedicated Amazon Data API | Pass an ASIN, URL, marketplace, and a few settings | Requests, parsing, normalization, and much of the scraping logic | Clean JSON or CSV | Teams that mainly want the data |
DIY custom scraping
DIY custom scraping is the most hands-on route. You can scrape Amazon using Python or Node.js with an HTTP library or browser automation tool, then add your own parsers, proxies, retries, session handling, and validation.
You have complete control over how Amazon is requested and what happens to the response. That also means that when Amazon changes something, you own the problem. Put simply, DIY gives the most flexibility and the biggest maintenance burden.
Managed scraping platforms
Managed scraping platforms like Apify and Octoparse sit in the middle space between maintenance and flexibility. They can include both visual no-code tools and more configurable cloud scrapers.
You still decide what products to collect, which fields matter, and how the job should run, but the platform handles much of the infrastructure underneath. Octoparse, for example, offers ready-made Amazon review templates and cloud runs.
Dedicated Amazon data APIs
Dedicated Amazon data APIs like Rainforest API are the least maintenance-heavy option. Instead of configuring a scraper, you give the service something like an ASIN, marketplace, or product URL and get structured data back.
For example, ScraperAPI (while a generalist scraper) has an Amazon Product API that accepts an ASIN and Amazon marketplace, then returns structured product data including the reviews Amazon shows publicly on that product page.
The appeal here is zero browser automation or selector maintenance, and very little parsing work. But the same access limit still applies. An API can simplify how you collect Amazon review data, but it can’t guarantee access to reviews Amazon no longer serves publicly.
How to Build a Clean Amazon Review Dataset

Source: Marques Thomas @ Unsplash
Whichever collection method you choose, the underlying job is much the same. You need to identify the right product, confirm Amazon returned what you expected, pull the useful review fields, and preserve enough context to make sense of them later.
1. Pair the Product ID with the Right Country
Every Amazon product page has a unique 10-character code called an ASIN (Amazon Standard Identification Number). While a product keeps the same ASIN across different countries, the customer reviews don’t always cross international borders. Amazon changes the review counts and ratings depending on the local marketplace.
In other words, an ASIN on its own isn’t enough context. The same product can sit on amazon.com, amazon.co.uk, and amazon.de, each with its own local reviews and sometimes a different lineup of variants.
Because of this, saving an ASIN like B0XXXXXXX by itself isn’t enough. To build a clean dataset (especially when analyzing competitors globally), you’ll need to pair every product ID with its specific marketplace (like amazon.com or amazon.co.uk) right from the start.
2. Make Sure Amazon Returns the Right Pages
Next comes the request itself, whether you used a dedicated API, managed scraper, or your own infrastructure. The important part is what happens after Amazon responds.
Your scraper must be able to tell the difference between a real product page and these common roadblocks:
Sign-in screens
CAPTCHA or anti-bot pages
A layout your tool doesn’t recognize
“Currently unavailable” or retired product listings
Cookie consent banners (especially on European storefronts)
Amazon’s famous 404 pages, staffed by employees’ dogs (adorable, but not review data)
Don’t rely entirely on HTTP status codes. Amazon often serves anti-bot pages while still returning a successful HTTP 200 code. Your system thinks it succeeded, but you end up saving useless data.
One vital check here is to always watch core elements like the product title or the review container. If these elements are missing, your scraper needs to mark the attempt as a failed request rather than an empty product.
3. Keep Only the Fields You Need
You don’t need to extract every piece of data just because you can. Collecting unnecessary information only slows you down and makes cleaning the data much harder. For most projects, you only need these eight key fields:
| Field Name | Example | Why You Need It |
|---|---|---|
| ASIN | B0XXXXXXXX | Links the review to the exact product. |
| Review ID | R3C7EMIUQ2O4UZ | Unique code to identify the review and spot duplicates. |
| Rating | 1 | Quick number (1–5 stars) to measure customer satisfaction. |
| Title & Text | “Battery drains…” | The actual written feedback and customer insights. |
| Date | 2024-04-04 | Tracks how customer opinions change over time. |
| Variant | 128GB, Blue | Reveals if a specific color or size has a unique defect. |
| Verified Purchase | True | Confirms if the reviewer actually bought the product. |
| Helpful Votes | 42 | Highlights which reviews other shoppers trust the most. |
Pay special attention to the Review ID. Because it never changes, it’s the most reliable way to make sure you aren’t accidentally counting the same review twice when you update your dataset later.
4. Clean and Deduplicate Your Data
Amazon primarily formats its data for human readers. For example, ratings appear as “5.0 out of 5 stars.” International Amazon sites use different languages for these same fields, which means you need to convert raw text into a single, consistent format across your entire dataset.
Next, remove duplicates using the unique review ID. When you run your scraper multiple times, check new IDs against your existing database instead of saving them as brand-new entries.
Because reviews can change over time (such as getting more “helpful” votes or being edited by the author), you should update your existing records with the fresh data rather than creating duplicate entries.
5. Save the Context Alongside the Review
Always save the marketplace (like US or UK), the product ID (ASIN), the source link, and the exact date you collected the data alongside the review text.
It might feel like extra paperwork at first, but you’ll need it later if your data looks inconsistent. For example, a product review collected from Amazon US in September can’t be directly compared to a review from Amazon UK in March without this context.
High-quality data collection is about more than just grabbing text from a website. You need to keep enough background information so the text still makes sense when you analyze it later.
Scaling Amazon Review Collection With the Right Proxy Setup
Once you move from checking a few products to monitoring hundreds or thousands, the network layer starts to matter. Proxies help mainly by giving you more control over which IPs send those requests and how the workload is distributed.
For a DIY or in-house Amazon scraper, that usually means:
Spreading requests across multiple IPs: Instead of sending every product-page request through one address, you can divide jobs across a larger pool.
Keeping an IP stable where continuity matters: Static ISP proxies keep the same address rather than rotating automatically, which can be useful for workflows where related requests should maintain a consistent network identity.
Matching Regional Context: Amazon customizes product availability, delivery estimates, pricing, and review order based on geographic location. Routing traffic through country- or state-specific proxies ensures your scraper retrieves the precise listing view local buyers see.
For teams running Amazon review data collection in-house, HypeProxies is specifically built to improve the proxy network infrastructure layer.

TrustPilot Rating: 4.8/5 ⭐ from 148 reviews
You get access to:
Dedicated static ISP proxies from real US carriers hosted on 10GBPS to 100GBPS
Unlimited bandwidth and unlimited concurrent threads on per-IP pricing, starting at $1.30 per IP.
500,000+ IPs geolocated across all 50 US states, plus Canada (that covers amazon.com and amazon.ca)
Purpose-built network infrastructure, from the hardware to the IP ranges.
Independent performance testing backs this up. In June 2025, Proxyway sent roughly 2,600 requests to Amazon through 100 HypeProxies ISP proxies and recorded a 100% success rate with a 2.08-second average response time. Explore ISP Proxies.
Key Takeaways
Amazon review scraping in 2026 starts with understanding what Amazon currently supports, then choosing the ideal collection method for your use cases. A few key things to keep in mind:
Product pages provide useful public review data, but they are only a sample of the wider review history.
Managed platforms and dedicated APIs remove much of the scraping maintenance, while DIY scrapers trade that convenience for more control.
Clean datasets need more than review text. Keep review IDs, variants, marketplaces, timestamps, and validation checks alongside the content.
Proxies help distribute larger workloads and maintain stable network identities, but they don’t get around authentication.
Running Amazon product collection at scale? HypeProxies gives you dedicated US ISP IPs, unlimited bandwidth, and stable network identities without charging by the gigabyte. Check out our ISP Proxies today.
Share on
No credit card required. Request your free trial today.
Stay in the loop
Subscribe to our newsletter for the latest updates, product news, and more.
No spam. Unsubscribe at anytime.





