How we benchmark
Scrapeway benchmarks web scraping APIs against real protected websites and publishes the results twice a month. We pay for every plan ourselves. There are no sponsors and no affiliate links.
This page covers what we measure, how a run works, which providers we cover, and what we leave out.
What we measure
Success rate. A request counts as a success only when the response contains the content we expect for that page type. We check for expected content rather than status codes, because a blocked request often returns a normal 200 carrying a challenge page, an empty shell, or a soft error. A benchmark that trusts status codes counts those as wins and reports success rates that no scraper would recognize in production.
Speed. Response time per successful request. Failed requests carry no response time, so they are not part of the average. Every figure on this page is a mean rather than a median.
Cost per 1,000 successful requests. Priced on each provider's cheapest entry plan, which averages around $40 across the services we test. The denominator is successful requests rather than requests sent, so a provider that fails often costs more per useful result. That is the figure that matters when you are budgeting a scrape, and it is why the cheapest sticker price is frequently not the cheapest option. Cost reflects what we actually spend. Some providers bill for failed requests and some do not. Where they do, those charges are included, because they are charges you would pay too.
How a run works
We benchmark a fixed set of live URLs per target, and every provider gets the same URLs. Every target uses the same number of URLs, so no target is measured on a thinner sample than another. The list and the exact request volume stay private, for the reasons below.
Each benchmark covers two weeks, and every figure we publish for a target rests on around a thousand requests per provider over that period. Spreading it across two weeks means a provider is measured through changing site conditions rather than at one favorable or unfavorable moment.
Results are published twice a month, after we validate the run and check failures for configuration or detection errors. A failure caused by our own configuration is a bug on our side, and we fix it rather than publish it as a provider's result.
Which providers we benchmark
Scrapeway benchmarks self-serve web scraping APIs: services with public per-request pricing, instant signup, and no sales call before you can send traffic. That is the set most teams are choosing between when they need to scrape a protected site this week, so that is the set we test.
The current run covers 8: Firecrawl, Scraperapi, Scrapfly, ScrapingBee, Scrapingdog, String, WebScrapingAPI, ZenRows.
Providers we do not benchmark
Bright Data and Oxylabs are proxy providers, not web scraping APIs. Their unblocking products overlap with this category, but both companies sell proxy networks through sales-led onboarding and contract pricing. That is a different product bought in a different way, so it belongs in a different comparison.
Zyte is not in the current run. We have not integrated it yet.
We publish no numbers for services we have not tested. A provider missing from our tables reflects our scope, not our assessment of it.
Suggest a provider
Want a service added? Open an issue on GitHub. We add providers that fit the category and that we can integrate cleanly.
What we test on each target
Each target has a defined page-type scope. We publish measured results only for the page types we test.
| Target | We test |
|---|---|
| amazon.com | Product pages |
| booking.com | Hotel and other public pages |
| etsy.com | Product and other public pages |
| indeed.com | Job listing pages |
| instagram.com | Public pages (profile, post) |
| linkedin.com | Public pages |
| realtor.com | Property listing and search pages |
| stockx.com | Product and other public pages |
| twitter.com | Posts and other public pages |
| walmart.com | Product pages |
| zillow.com | Property listing and search pages |
Why we do not publish the target URLs
We do not publish the exact URLs we benchmark.
Two reasons. Target URLs expire: products get delisted and listings return 404 or 410, so the set changes between runs. More importantly, a provider that knows the exact URLs in advance can treat them differently from ordinary traffic, by caching those responses ahead of a run or by cutting off our account. Both have happened.
Keeping the list private is what keeps the numbers comparable across providers.
Independence
We pay for every plan we benchmark. No provider sponsors Scrapeway, no links on this site are affiliate links, and no provider sees results before publication.