Congress Introduces Bill to Expose Stealth Bots Scraping American Websites

A website can receive thousands of page requests before its owner finishes a morning coffee, yet the traffic may arrive wearing the digital equivalent of a fake moustache.

On July 23, 2026, a bipartisan House group introduced H.R. 9915, the Stealth Bot Prohibition Act, aimed at crawlers that hide who they are or pretend to be human.

The proposal would require bots to disclose their identity and purpose. It would also let the Federal Trade Commission seek penalties of up to $53,000 per violation.

Properly identified crawling would remain legal, while copyright and licensing disputes would stay under existing law, according to the full bill text.

What the Stealth Bot Prohibition Act Would Do

Stealth Bots Bill

Representatives Laurel Lee of Florida, Valerie Foushee of North Carolina, and Gus Bilirakis of Florida introduced the bill and sent it to the House Committee on Energy and Commerce.

Introduction marks the beginning of the legislative process, despite its bipartisan sponsorship. The official congressional bill record provides its current status and legislative details.

H.R. 9915 defines “bot” broadly. Crawlers, spiders, fetchers, user agents, AI agents, and similar software qualify when they retrieve, scan, index, scrape, or access an internet source.

A “stealth bot” operates without prior disclosure of its identity and purpose. An accurate user-agent string appears as one expected form of identification.

Purpose disclosure would cover search indexing, text and data mining, model training, fine-tuning, inference, or retrieval-augmented generation. The information must arrive when access is requested and in a format available to the website operator.

A person could violate the law by deploying a stealth bot in a way reasonably likely to damage, impair, or burden a website’s technical or commercial operation.

A separate clause targets intentional efforts to make a bot appear human when the activity connects to a generative AI model or service.

Why Congress Is Looking at Crawlers Now

 

View this post on Instagram

 

A post shared by Face The Nation (@facethenation)

Imperva’s 2025 Bad Bot Report estimated that bots generated 51 percent of measured web traffic during 2024, exceeding human traffic for the first time in the report’s decade-long history. Bad bots represented 37 percent.

The category covers many automated threats alongside scraping, yet the figures show how crowded the machine-driven web has become. Imperva published the figures in its annual bot report.

Publisher anxiety also reflects an altered exchange of value. Cloudflare estimated a June 2025 crawl-to-referral ratio of roughly 14 to 1 for Google, 1,700 to 1 for OpenAI, and 73,000 to 1 for Anthropic.

Native-app referrals can be difficult to attribute, according to Cloudflare, but the gap helps explain why publishers increasingly see AI crawling as a cost with little return traffic. The company presented the comparison while announcing new AI crawler controls.

A public dispute gave lawmakers a vivid example. In August 2025, Cloudflare alleged that Perplexity used undeclared crawlers, generic browser identities, rotating IP addresses, and changing networks after sites tried to block its declared crawler.

Perplexity disputed the characterization and called the report a publicity stunt, according to contemporaneous coverage. Cloudflare documented its allegations in a stealth crawler investigation.

From robots.txt Etiquette to Federal Enforcement

Website owners have relied on robots.txt since Martijn Koster proposed the Robots Exclusion Protocol in 1994. The Internet Engineering Task Force standardized it as RFC 9309 in September 2022. A site publishes instructions at /robots.txt, while a crawler identifies its product token and follows the relevant rules.

Compliance depends heavily on honest identification. The robots.txt technical standard says its rules do not function as access authorization and cannot replace authentication or other security controls. H.R. 9915 would move bot disclosure from web custom into federal law, turning a fake user-agent into potential evidence in a civil case.

Who Would Enforce the Rules?

The FTC could file a civil action in federal court, seek an injunction, or pursue a penalty capped at $53,000 for each violation. The cap would rise annually with inflation. State attorneys general and certain state officials could also seek injunctions, damages, restitution, or other relief on behalf of residents.

Website owners receive no express private right to sue under H.R. 9915. Existing claims under copyright, contract, computer-access, or state law would remain available where applicable. Civil actions would carry a six-year limitations period, and the law would take effect 180 days after enactment.

The proposal bars the FTC from issuing regulations under the section, leaving courts and enforcement actions to shape several practical boundaries.

Where Ambiguity Could Create Trouble

The phrase “each violation” carries major financial importance, yet the bill does not define the unit. One request, one crawling session, one bot, or an entire campaign could produce very different penalty totals.

Disclosure mechanics remain open as well. A user-agent can identify a crawler, but its purpose may shift between indexing, training, and live retrieval. Operators will need a machine-readable method that communicates enough detail without turning every request into a legal essay.

The phrase “commercial operation” may invite argument. Bandwidth and infrastructure expense can be measured. Lost subscription value, reduced referrals, or weaker licensing leverage are harder to tie directly to one crawler.

Offshore scrapers using residential proxies also raise questions about identity, jurisdiction, and collection.

Services such as ProxyWing show how operators can route requests through rotating 4G and 5G carrier IPs rather than a fixed server address.

Supporters and Critics See Different Risks

Website scraping

Publisher groups view disclosure as the foundation for a content-licensing market. Supporters listed by the sponsors include the News/Media Alliance, News Corp, The New York Times, Condé Nast, Vox Media, Reddit, and other companies. Their argument is practical: a publisher cannot negotiate access or block a crawler effectively when the collector is disguised. Representative Foushee’s office listed the supporters in its bill introduction announcement.

Open-web advocates worry about legitimate anonymous scraping. Re:Create argued that reporters may use automated browsers to test dynamic pricing, discrimination, or algorithmic bias without alerting the target. Its July 2026 critique of the proposal warned that broad definitions could chill public-interest investigations.

A plain reading offers some narrowing. A non-AI investigative crawler would generally need to create a reasonably likely technical or commercial burden before the first clause applies. No explicit exemption protects journalism, academic research, archiving, accessibility, or cybersecurity work. Safe harbors may become a major issue in committee debate.

What Happens Next?

As of August 4, 2026, H.R. 9915 had been introduced and referred to the House Energy and Commerce Committee. It had not passed either chamber and created no legal obligations. Website operators can still preserve server logs, document crawler identities, measure infrastructure costs, and keep robots.txt instructions current.

AI developers may need to audit first-party crawlers, data vendors, browsing agents, and retrieval services. Contracts with outside scraping providers could matter because the proposal reaches people who deploy, direct, or cause a stealth bot to be deployed.

The central idea is easy to picture: software knocking on a website’s door should give a real name and explain why it wants to enter. Workable law will require sharper answers about exemptions, evidence, penalties, and cross-border enforcement.

The debate reaches well beyond model training. It asks whether the open web can preserve anonymous public-interest research while giving publishers and businesses a practical way to identify industrial-scale extraction.

latest posts