Web Scraping Statistics 2026: U.S. Market, Adoption, Costs & Trends
AI Summary: Web scraping in the U.S. in 2026 is shaped by AI demand and harder access. Bots made up 57.4% of global HTML requests in June 2026, while U.S.-origin traffic was 43.6% automated. AI referral traffic to U.S. retail sites grew 393% in Q1, slowed to 62% by July, and converted 60% better. Product pages scored 66% machine readability. No public source measures the U.S. scraping market. In July the FBI seized NetNut-linked domains, and a court let Reddit's DMCA 1201 claims against SerpApi proceed.

What does web scraping actually look like in the U.S. in 2026? This report brings together the latest data on automated traffic, AI activity, access challenges, costs, and the changing infrastructure behind web data collection.
Web scraping in the United States looks different in 2026 than it did a year ago. AI systems are sending more automated requests to retail sites, while anti-bot systems are making ordinary product pages harder to read programmatically.
I’m Dom, the founder of ScrapeBadger, and I put these numbers together to get a clearer picture of where web scraping actually stands in the U.S. right now.
The statistics below focus on observed U.S. data from January through August 2026. Where reliable U.S.-specific data isn't available, we say so rather than rely on older surveys or global forecasts.
Key Web Scraping Statistics for the U.S. in 2026

Automated requests overtook human requests globally in June 2026. Cloudflare Radar put bots at 57.4% of HTML web requests against 42.6% human, a threshold Cloudflare's CEO had publicly expected in late 2027. (Source)
U.S. traffic is less automated than the global average. Cloudflare Radar's U.S. bot-traffic view put bot traffic originating from the United States at 43.6%, against 56.4% human. That is roughly 14 points below the global figure. (Source)
AI referral traffic to U.S. retail sites grew 393% year over year in Q1 2026. Adobe measured this across more than one trillion visits to U.S. retail sites. (Source)
By July 2026 that growth had cooled to 62% year over year, but the traffic converted 60% better than non-AI sources. It was the eleventh consecutive month AI-referred traffic outperformed on conversion. (Source)
Alarum Technologies reported $11.7 million in Q1 2026 revenue, up 64% year over year. It is the only audited, filed 2026 revenue figure we found from a pure-play web data infrastructure company. (Source)
On 2 July 2026 the FBI and IRS Criminal Investigation seized hundreds of domains tied to NetNut, alongside a Google Threat Intelligence Group action against a botnet Google estimates at more than two million compromised consumer devices. (Source)
Roughly 19.8% of U.S. businesses reported using AI as of early May 2026, based on a Census Bureau survey of about 1.2 million businesses. This measures AI adoption, not web scraping. (Source)
Together, those tell a more useful story than a single global market forecast. Automated traffic now dominates the global web, while AI is reshaping U.S. retail traffic. At the same time, getting machine-readable data is becoming harder, and proxy infrastructure is facing greater legal scrutiny.
How Much of U.S. Internet Traffic Is Automated? Start with the global picture, because the U.S. number only makes sense against it. On 3 June 2026, Cloudflare Radar showed automated systems generating 57.4% of HTML web requests against 42.6% from humans.

It was the first machine majority on record, and it arrived roughly 18 months ahead of Cloudflare's own published expectation. The U.S.-specific figure comes from Cloudflare Radar's historical U.S. bot-traffic view. It put bot traffic originating from the United States at 43.6%, with 56.4% human. Cloudflare Radar's worldwide bot-traffic view also showed the U.S. was the origin of 53.5% of global bot traffic by volume (Source). Two things are worth pulling out of that:
First, the United States is meaningfully below the global average for automation, not above it. Iran sits at 81.4%, Singapore at 73.7%, Ireland at 71.1%, Germany at 45%. Small, datacenter-dense markets skew highest. The U.S. sends the largest absolute volume of bot traffic while running a comparatively human-heavy domestic mix.
Second, that 53.5% figure measures where automated requests appear to originate, not where the people or companies behind them are located. Traffic routed through U.S. infrastructure counts as U.S. traffic regardless of who initiated it. The bigger qualifier: bot traffic is not web scraping.
Cloudflare's classification covers search crawlers, AI agents, monitoring tools, automated scripts, and malicious bots. Scraping is one slice. It would be wrong to read 43.6% as a scraping figure. What the number does establish is that automated access is too large to treat as an edge case.
Cloudflare's own July 2026 report found that 52% of crawler requests were for AI training as of June 2026, up from 22% in spring 2025. Treat that one carefully. Cloudflare Radar revises its classifications retroactively, and re-running the identical 28-day window in August returned a materially lower training share. Cite it with the date you pulled it, or don't cite the precise number at all. It is global, not U.S., so we use it as context rather than as a U.S. statistic (Source).
Can We Actually Measure the U.S. Web Scraping Market in 2026?
No. Not from any transparent public source we could find. Several research firms publish global web-scraping market figures for 2026. Those are modeled estimates built from a historical base plus assumptions about growth, segmentation and pricing. They are not a tally of actual revenue earned between January and August 2026, and none of them isolates the United States.
Mordor Intelligence currently gives a global 2026 figure of $1.56 billion and identifies North America as its largest regional market. That is neither a U.S.-only figure nor an observed one, so we don't use it as either. (Source)
Published estimates also define the market differently. Some count software and services including custom development. Others count only commercial platforms and managed services, explicitly excluding scrapers that companies build and run internally. Given that in-house building is how a large share of scraping actually happens, that exclusion alone moves the number considerably. There is one hard 2026 revenue figure worth having.
Alarum Technologies, the Nasdaq- and Tel Aviv-listed parent of the proxy provider NetNut, reported $11.7 million in revenue for Q1 2026, up 64% year over year, with gross margin of 61.7%. That is a filed financial result covering January through March 2026, not a company blog post and not a model. (Source)
It doesn't size the market. One company at roughly $47 million annualized tells you very little about a sector estimated in the billions. What it does give you is a reference point: an actual, audited 2026 growth rate of 64% against forecast market CAGRs in the 17% to 18% range. Either the forecasts understate growth, or they are measuring something narrower than the businesses actually operating in this space. Worth noting that Alarum delayed its Q2 2026 results. As of September 2026 there is no published April–June figure.
What we can say with confidence
There is no sufficiently transparent public source giving an actual U.S.-only web-scraping market revenue figure for January through August 2026. That isn't a hole in the article. It's the finding. A statistics resource should distinguish between "this hasn't been measured publicly" and "we found a number on Google."
How Widely Do U.S. Businesses Use Web Scraping?
There is still no nationally representative U.S. government survey asking businesses whether they scrape. So any claim that "X% of American companies use web scraping" deserves suspicion unless the underlying survey is independently conducted with a public methodology. We looked and didn't find one.
What we can show is that demand for machine-collected web data is expanding. The U.S. Census Bureau's Business Trends and Outlook Survey is nationally representative of employer businesses, samples roughly 1.2 million of them, and collects every two weeks. Its 2026 data showed overall business AI usage running between 17% and 20% through early May 2026.
Among businesses with at least 250 employees, 37% reported using AI. In the information sector, the rate was 39.7%; finance and insurance was 33.9%; retail trade was around 14% (source). This is not a scraping statistic and shouldn't be read as one. The point isn't that a fifth of U.S. businesses are scraping. It's that the customer base for systems that consume fresh machine-readable data is getting more technically mature, and the gradient by company size is steep.
If you want to understand why U.S. AI adoption surveys disagree so much, a Federal Reserve analysis from April 2026 compares three of them and explains what each is actually asking (source).
Different Kinds of "AI Traffic" Are Not the Same Thing
image Four separate measurements get blurred together constantly. Keeping them apart is most of the work in reading this topic honestly.

AI adoption by businesses is the Census figure above. It's about firms using AI tools internally.
AI referral traffic is people arriving at a website after asking an AI assistant something. That's what Adobe measures. The visitor is human.
AI crawler traffic is bots fetching pages to train or ground a model. That's what Cloudflare measures. No human is present.
AI-assisted scraping is developers using AI to write or run extraction code. That's what practitioner surveys measure.
A rise in one does not prove a rise in another. They're related but distinct, and articles that stack them together to build a bigger number are doing something misleading.
What the U.S. Retail Data Shows
Adobe Analytics covers more than one trillion visits to U.S. retail sites, which makes it the largest U.S.-specific commercial dataset available on this question. AI-source traffic to those sites grew 393% year over year in Q1 2026. March alone was up 269%, extending momentum from the November–December 2025 holiday period when it ran 693% ahead. (Source)
Then it decelerated. May was up 138% year over year. July was up 62%. That deceleration is the part most coverage skips, and it matters. Volume is still climbing, but the growth rate has fallen sharply from a very high base across two quarters. Anyone modeling forward off the 393% figure alone will overshoot.
The value of that traffic went the other way. In July 2026, AI-referred visits converted 60% better than non-AI traffic, the eleventh consecutive month AI outperformed. Those visitors generated 53% more revenue per visit, showed 14% higher engagement, spent 59% more time on site, were 33% less likely to bounce, and added items to cart at a 28% higher rate. (Source)
Twelve months earlier, the picture was inverted. In March 2025, AI-referred traffic converted 38% worse than other channels. By March 2026, it converted 42% better.
These are AI referral statistics, not scraping statistics. They matter here because AI systems need current product information, prices, and availability to make useful recommendations, and that demand drives the market for fresh, structured web data.
Are U.S. Retail Pages Actually Machine-Readable?
Less than you'd assume, and this is where scraping and AI access converge. Adobe's April 2026 benchmark scored U.S. retail homepages at an average 75% machine readability, category pages at 74%, and individual product pages at 66%.
Product pages scoring lowest is the awkward part, since that's where price, availability, and specifications live, and it's exactly the content that tends to render client-side. (Source)
In July, Adobe reported homepage visibility of 61%. That is not a decline.
Adobe expanded its analysis to a broader set of U.S. retail sites between the two readings, so the numbers describe different samples and shouldn't be compared as a trend.

On its own, the July figure indicates that 39% of homepage content across the wider cohort isn't easily machine-readable (source). Either way, having a page online doesn't mean a machine can read what a human can see.
What Happens When Automated Clients Try to Read U.S. Retail Sites?
In August 2026, Decodo tested 10 live product pages from U.S. retailers using four access configurations. A plain self-identifying agent retrieved usable product data from 2 of 10 pages. A headless browser also managed 2. Overriding the user agent got it to 3. Correcting client hints got it to 5.
A managed setup combining a scraping API, browser rendering, and a U.S. exit retrieved product data from 8 of 10 under the study's scoring rules (source).
This is a vendor experiment, not an independent benchmark. Decodo sells the kind of managed infrastructure that scored highest, and ten pages is a small sample. Don't generalize it to U.S. retail as a whole. It's still worth reading, for two reasons:
The methodology is stated plainly. A page counted as successful only when it returned usable product information such as a product schema, price fields or an add-to-cart control. That's a stricter and more honest bar than most vendor tests use. And it surfaced something useful about measurement: one retailer returned HTTP 200 while serving a challenge page rather than content.
A successful HTTP response is not a successful extraction, which is why testing what your proxies actually return matters more than counting requests sent.
The same test examined robots.txt across 20 major U.S. retailers. Nineteen had readable files, and all 19 disallowed at least one buying-related path under the general User-agent: * group, including cart, checkout, and purchase paths. That distinction is going to matter more over time: reading a page and acting on a site are different problems.
A retailer can leave product information publicly readable while blocking automated systems from completing cart or checkout actions.
Web Scraping Infrastructure in the U.S.
There is no federal industry classification for "web scraping infrastructure," which makes a clean U.S. employment or revenue series for scraping alone impossible to produce. The closest adjacent measure is the U.S. Data Processing, Hosting, and Related Services industry. BLS reported approximately 453,300 employees in August 2026, with August employment preliminary, and 55,979 private establishments in Q1 2026, also preliminary. Average hourly earnings reached $60.25 in July 2026 (source).
These do not measure the web-scraping industry. They describe the servers, hosting and data-processing ecosystem that large-scale collection runs on. Adjacent industry numbers get presented as scraping-market statistics fairly often, so the distinction is worth preserving.
What Did Scraping Cost in the U.S. in 2026?
There is still no trustworthy figure for "the average cost of web scraping in the U.S.". Cost depends on what you're collecting, how often, how much rendering is required, how many targets you're hitting, and how hard those targets are to reach. The useful approach is inputs, not an invented average.
Proxy pricing providers use different billing units, so headline prices aren't directly comparable. At the time of writing, Decodo advertises residential proxies starting at $2 per GB, with displayed monthly tiers above that starting point depending on volume (source).
Oxylabs lists ISP proxies from $1.60 per IP on its entry monthly tier, with lower per-IP rates at volume. Its own page also documents a fair-usage condition: past 50 GB per ISP proxy, concurrent sessions drop to 10 per proxy for the rest of the billing cycle. Read the unlimited clause wherever you buy (source).
The billing model matters more than the sticker price. Bandwidth-billed residential proxies make your cost partly a function of page weight, which you don't control. IP-billed proxies with unlimited bandwidth shift the main capacity variable to IP count instead.
Why success rate beats sticker price
On protected sites, the success rate usually matters more than either. Say a workload needs 5 TB of usable transferred data. At a 40% success rate you'd attempt roughly 12.5 TB to get there. At 80%, about 6.25 TB.
The proxy price didn't change. Economics did.
What makes it more expensive
Three costs compound. Access difficulty. Rate limits, CAPTCHAs, fingerprinting and behavioral detection all require more capable infrastructure.
Rendering. JavaScript-heavy pages can need a real browser rather than a simple HTTP request.
Failure and retry volume. When a site keeps serving a scraper challenge pages instead of content, every retry burns proxy bandwidth and compute without producing usable data
The August 2026 retailer test illustrates the first point directly. On some targets, changing only the client characteristics determined whether the page returned usable product data at all. For engineering labor, the most recent national BLS wage data is still May 2025 rather than 2026, and we won't pretend otherwise in a 2026 article.
The median annual wage for U.S. software developers was $135,980 in May 2025. Treat it as the latest available benchmark, not a 2026 figure. (Source)
Build or Buy in 2026?
No independent U.S. study gives a dollar threshold where building becomes cheaper than buying. Every published comparison we found came from a company selling one side of the answer. Building tends to make sense with a small number of stable targets, minimal anti-bot protection, and an engineering team that can maintain the scraper.
Buying gets more attractive when targets change often, JavaScript rendering is required, access restrictions are significant, or the scraping infrastructure isn't a competitive advantage in itself. The retailer experiment shows why the framing matters. A scraper isn't just code that makes HTTP requests.
The access layer involves the IP, browser characteristics, client hints, rendering environment and retry behavior.
Most teams that mix approaches keep parsing logic in-house while buying the layer that needs constant maintenance against moving targets, which in practice usually means the proxy and unblocking layer.
The Infrastructure Story Nobody Expected in 2026
On 2 July 2026, the FBI and IRS Criminal Investigation seized hundreds of domains tied to NetNut, a residential proxy provider owned by the publicly traded Alarum Technologies. Google's Threat Intelligence Group acted the same day, disabling accounts used for command and control and flagging the associated SDK in Play Protect.
Google linked the network to a botnet it estimates at more than two million compromised consumer devices, largely smart TVs and streaming boxes. In a single week in June 2026, Google says it observed 316 distinct threat clusters routing traffic through suspected exit nodes on that network. Google also stated it has high confidence that a number of popular residential proxy brands were reselling or white-labeling the same infrastructure. (Source)
It was Google's second residential proxy disruption of 2026, following action against the IPIDEA network in January. Alarum disputes the allegations and stated on 3 July that neither it nor NetNut had been formally contacted by the FBI or any other authority. No criminal charges have been filed against the company or any named executive. (Source)
We sell scraping infrastructure, so take our framing with that in mind. The reason this belongs in a statistics article is narrower than it looks: it's the clearest 2026 evidence that how a proxy network sources its IPs is a real operational risk, not a compliance checkbox.
If a provider can't tell you where its addresses come from, that's now a question with a documented federal answer attached to it. The FBI published consumer guidance on residential proxy networks alongside the action. (Source)
The Biggest U.S. Web Scraping Trends in 2026
AI is now a major reason to collect fresh web data. AI referral traffic to U.S. retail grew 393% year over year in Q1 and was still up 62% in July, while converting 60% better than other channels. AI systems need current prices, specifications and availability to make recommendations.
Growth is decelerating even as volume rises. 393%, then 138%, then 62% across two quarters. The direction is up; the rate is falling fast.
Machine readability is an access problem, not just an SEO problem. A third or more of U.S. retail page content isn't cleanly readable by machines, depending on page type and cohort.
Automated traffic became the majority worldwide. 57.4% of HTML requests in June 2026, with the U.S. running below that at 43.6%. That doesn't mean scraping is the majority. It means the environment scrapers operate in is now predominantly automated.
Proxy infrastructure came under direct legal pressure. The July seizure and Google's two 2026 disruptions put IP sourcing on the table as an operational question.
One trend we deliberately aren't claiming: we found no reliable evidence for the widely discussed idea that no-code scraping tools are displacing custom-built scrapers. It gets asserted often. We couldn't find adoption or spending data behind it, so we're leaving it as discussed but not yet measured.
What Changed in U.S. Web Scraping Law in 2026?
The picture got more complicated, not simpler. Two 2026 decisions matter, and they point in different directions.
On 31 July 2026, the U.S. District Court for the Southern District of New York largely denied motions to dismiss in Reddit, Inc. v. SerpApi LLC et al., allowing Reddit's DMCA Section 1201 anti-circumvention claims to proceed. The court held that Section 1201 doesn't require a copyright holder to have authorized the specific technological measure that was circumvented, only that it broadly authorized access controls. The alleged conduct includes proxy use, IP rotation and bypassing anti-bot systems. (Source)
That is the more consequential ruling for scrapers. It suggests the legal center of gravity is shifting from the Computer Fraud and Abuse Act, which asks whether access was authorized, toward Section 1201, which asks whether you got around a technical measure.
On 4 August 2026, the Ninth Circuit decided Amazon.com Services, LLC v. Perplexity AI, Inc., vacating a preliminary injunction against Perplexity's Comet browser. The court held that on the record before it, the user rather than the AI tool "accesses" the website for CFAA purposes, describing the assistant as a tool rather than a person. (Source)
Two caveats on that one. It concerned an AI agent operating inside a user's own password-protected account at that user's direction, which is not conventional scraping of public pages. And the court explicitly declined to grant agentic AI broad immunity, noting Amazon retains other routes to regulate access, including its terms of service.
Older cases still set the framework. Van Buren v. United States (2021) narrowed the CFAA's "exceeds authorized access" clause. (Source) hiQ Labs v. LinkedIn (2022) held that scraping public pages doesn't trigger the CFAA's "without authorization" clause in the Ninth Circuit, though hiQ then lost on breach of contract. (Source)
The takeaway isn't that scraping is legal or illegal. It's that the method of access and the legal theory asserted both matter enormously. Public versus authenticated access, contracts, technical restrictions, the type of data collected and the jurisdiction all pull in different directions. This article is informational and is not legal advice.
How We Researched These Statistics
Two filters:
Date. Where a figure is described as 2026 data, the underlying observation falls between 1 January and 31 August 2026. A report published in 2026 isn't automatically 2026 data. The most widely cited practitioner survey in this field, published under a 2026 title, was fielded in December 2025, and we've treated it accordingly.
Geography. Statistics describing the U.S. rest on U.S.-specific data, not North America or the Americas. We also separated measured data from forecasts. A market report stating what it expects 2026 to be worth was not treated as actual 2026 revenue.
Where no suitable U.S. 2026 statistic existed, we either used a clearly labeled adjacent measure, such as U.S. AI retail traffic, or said the evidence wasn't available. Vendor studies are labeled as vendor studies. Company-reported revenue is labeled as company-reported. Current pricing is treated as a point-in-time commercial input, not an industry average. And where a statistic is repeated across dozens of sites but traces back to a single vendor article with no methodology, we left it out.
The Bottomline
For businesses collecting data at scale, success depends less on sending more requests and more on accessing the right data reliably, handling website restrictions, and keeping failed requests and retries under control. As scraping and AI-driven data collection evolve, those factors matter just as much as raw request volume.
Scraping shouldn't mean juggling proxies, anti-bot bypass, and rendering. Try ScrapeBadger with 1,000 free credits and start scraping.
Written by
Domas Sakavickas
Dom Sakavickas is Co-founder of ScrapeBadger, building web scraping infrastructure for developers and data teams. He writes about the web data market, tool comparisons, and business use cases for scraping. ScrapeBadger is a web scraping API platform specialising in Twitter/X, Reddit and Google data, with dedicated scrapers also covering TikTok, YouTube, LinkedIn, Amazon, eBay, Zillow and 40+ more: with built-in anti-bot bypass and an MCP server for AI agents.
Ready to get started?
Join thousands of developers using ScrapeBadger for their data needs.