AI Crawler Scandal: Perplexity vs. Cloudflare

AI Crawler Scandal: Perplexity vs. Cloudflare

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
6. 8. 2025
3 minutes reading · 4 views
AI Crawler Scandal: Perplexity vs. Cloudflare

AI Crawler Scandal: Perplexity Versus Cloudflare

You are a website owner trying to protect your content from unwanted visitors. You set rules, block bots, and yet someone still slips through. According to a report by Cloudflare, that is exactly what is happening, with the company accusing AI startup Perplexity of using clever tricks to bypass restrictions. This story is full of details about how modern technologies are fighting over online data, and it shows how difficult it can be to maintain control over your own content. Let's take a look at it step by step.

Customer Complaints and Initial Tests

It all began when Cloudflare received complaints from its customers. They claimed that Perplexity's bots were still accessing their websites despite restrictions configured in the robots.txt file and web application firewall (WAF) rules. Cloudflare decided to investigate and created new domains with similar blocks against Perplexity crawlers, such as "PerplexityBot" and "Perplexity-User".

What did they find? Perplexity first attempts to access a site using its official identifiers. If it encounters a block, it changes its user agent—the information that tells a website which browser or device is trying to connect. Instead, it poses as Google Chrome on macOS. This allows the crawler to slip through restrictions that would otherwise work.

IP Rotation and Network Changes

But that is not all. Cloudflare found that Perplexity uses rotating IP addresses that are not included in the official list of IP addresses provided by the company. This list is available in Perplexity's documentation, but these "secret" IP addresses come from other sources. It also switches autonomous system networks (ASNs), which are numbers identifying groups of IP networks controlled by a single operator. In this way, the crawler avoids detection and blocks.

According to Cloudflare, this activity affected tens of thousands of domains and involved millions of requests per day. This means that Perplexity is collecting data on a large scale even though website owners have clearly said "no". It is as if you locked the door, but a thief made a key from a different material.

Responses From Perplexity and Cloudflare

Perplexity responded through its spokesperson Jesse Dwyer, who described Cloudflare's report as a "publicity stunt"—in other words, a trick to attract attention. According to Dwyer, Cloudflare's blog contains numerous misunderstandings. The company published its own response on its website, claiming that Cloudflare mistook 20 to 25 million requests from user agents for AI scrapers. "User-driven agents act only on specific user requests and download only the necessary content," they explained. They also claim that Cloudflare deliberately linked Perplexity to 3 to 6 million daily requests from BrowserBase, a cloud browser for AI agents that Perplexity uses only occasionally.

Cloudflare, on the other hand, whose CEO Matthew Prince is known for his statements about AI posing an "existential threat" to publishers, responded forcefully. The company removed Perplexity from its list of verified bots and introduced new methods to block these "secret crawlers". Last month, Cloudflare even enabled websites to request payment from AI companies for crawling their content and began blocking AI crawlers by default.

Consequences

This incident is not an isolated one. Last year, Perplexity faced criticism for ignoring paywalls and robots.txt files, which CEO Aravind Srinivas attributed to third parties. Cloudflare now emphasizes that crawlers must be transparent and respect website owners' wishes. Other sources confirm that Perplexity masks bot identities, rotates IP addresses, and impersonates browsers, affecting tens of thousands of domains. Cloudflare has updated its rules to block this behavior, stressing that such practices undermine publishers' autonomy.

This story shows how AI companies like Perplexity are pushing boundaries to obtain data for their models. For website owners, this means a need for better protection, while for users, it is a reminder that complex battles over data lie behind intelligent search engines. If this interests you, keep an eye on developments—it seems this is only the beginning of a broader debate about AI ethics on the web.

Advertisement

Content created with help from UpTier.

SEO and GEO on autopilot. UpTier’s multi-agent systems write and optimize content for search engines and AI answers.

Discover UpTier ↗

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

OpenAI gives Codex reusable cloud workspaces accessible from any deviceOpenAI gives Codex reusable cloud workspaces accessible from any device
Codex gains reusable cloud development environments, alongside voice controls in its CLI, code reviews in the ChatGPT desktop app and cloud-based security tools.
2 min read
2. 10. 2026
Amazon releases Strands Decider 2B for AI workflow decisionsAmazon releases Strands Decider 2B for AI workflow decisions
Strands Decider 2B selects from predefined options and returns a confidence score. The fully open-source model is available now and small enough to run locally.
2 min read
1. 10. 2026
OpenAI says it disrupted a campaign to extract hidden model reasoningOpenAI says it disrupted a campaign to extract hidden model reasoning
OpenAI reported a coordinated effort to extract protected model reasoning and said it closed an extraction pathway. It attributed the main cluster of activity to individuals associated with Moonshot AI, the developer of Kimi.
3 min read
1. 10. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok