Italiano
English

The CAPTCHA Paradox in the Internet of Bots

For the first time in the history of the internet, traffic generated by bots and automated agents has surpassed human traffic. Anti-bot defences, starting with CAPTCHA, have long ceased to genuinely separate humans from machines. In practice, what they separate is those who identify themselves from those who disguise themselves, and the former pay the price. Those whose profession is to protect intellectual property online see this both in monitoring costs and, more broadly, in the effectiveness of the Digital Services Act. Meanwhile, CAPTCHA is also becoming a privacy compliance issue for those who deploy it. There seems to be only one sustainable direction: moving from visitor profiling to verifiable traffic identity.

The two phenomena are rarely considered within the same analytical framework. The first is what I have called the “CAPTCHA paradox”. The second is the regulatory gap left by the Digital Services Act concerning active monitoring by trusted flaggers, an issue already examined in these pages in 2023. Taken separately, each seems to belong to a different world: one is a technical web-security issue, the other an interpretative nuance of European regulation. Considered together, however, and despite being driven by legitimate intentions, they produce an outcome that is difficult to defend. The burden of enforcement increasingly falls on those who operate transparently, while bad-faith actors have access to industrial-scale, low-cost and continuously improving circumvention tools.

The figures released by Cloudflare in recent days, confirming that automated traffic has overtaken human traffic, are more than a statistical curiosity. They mark the end of the conceptual premise on which the entire anti-bot architecture was built and, in part, also of certain platform moderation practices. It is worth retracing the path that brought us here, starting with data that already showed the ambiguous nature of these tools in 2026.

2021: Friction Already Measured, and Already Asymmetrical

In November 2021, an analysis published in Agenda Digitale described the “security versus usability dilemma” posed by CAPTCHAs in e-commerce. The data were already telling. Only 66% of users solved a text CAPTCHA on the first attempt (Baymard Institute); activating the tests reduced automated registrations by 88%, but also caused a 3.2% decline in genuine conversions (MOZ). Read today, however, another figure is even more striking. As early as 2014, Google had acknowledged that one of its algorithms solved text CAPTCHAs in 99% of cases, compared with 33% for humans. In short, the tool designed to distinguish humans from machines was better at rejecting the former than stopping the latter. The industry’s response—from reCAPTCHA v3 to behavioural analysis and fingerprinting—shifted the test from a visible puzzle to invisible surveillance, without addressing its structural flaw: every increase in the threshold has primarily burdened legitimate users and declared operators, while organised actors have continued to bypass it through CAPTCHA farms and automation.

2026: Bots Overtake Humans and a Presumption Comes to an End

The current picture emerges clearly from Cloudflare data reported by ANSA on 8 June. According to those figures, 57.4% of requests to websites served by the company come from bots and automated agents, compared with 42.6% generated by humans. This is the first documented overtaking in the history of the internet, and its speed surprised even industry experts: Matthew Prince, Cloudflare’s co-founder and CEO, had predicted it for the end of 2027. Agentic AI is driving the phenomenon. Whereas a person consults an average of five websites before making a purchase, an agent can compare the same product across five thousand pages.

Cloudflare’s figure is not isolated. Imperva’s 2025 Bad Bot Report (Thales Group) had already found that automated traffic accounted for 51% of the total and that bad bots alone—openly malicious automation involving abusive scraping, account takeover and fraud—represented 37% of all web traffic, increasing for the sixth consecutive year. A methodological clarification is necessary. These metrics measure HTTP requests, not usage time or attention, and in terms of engagement humans remain the internet’s principal inhabitants. For those managing infrastructure, marketplaces and security systems, however, request volumes determine defence policies. And it is on those policies that the paradox exerts its weight.

The CAPTCHA Paradox: A Selective Barrier in Reverse

CAPTCHA, an acronym for Completely Automated Public Turing test to tell Computers and Humans Apart, was created to answer a binary question: human or machine? Today, that question no longer has a reliable answer. Recent analyses indicate that systems based on multimodal models solve CAPTCHAs with accuracy above 95%, while humans achieve between 50% and 86%, depending on the type of test. The question CAPTCHA actually answers has become another one, economic in nature: how much is someone willing to spend to bypass it?

This is where the barrier becomes selective in reverse. Bad-faith actors—from counterfeiting networks and replica sellers to fraudulent store operators and phishing organisations—have access to a mature circumvention supply chain. Automated solvers with negligible marginal costs, CAPTCHA farms staffed by human operators, residential proxies that disguise the origin of traffic, and tools designed to evade fingerprinting: for these actors, CAPTCHA is a trivial cost item already built into the illicit business model.

Those who operate transparently face the opposite situation. A brand-protection operator carrying out systematic monitoring on behalf of rights holders has an interest in identifying itself and often has a contractual or professional obligation to do so: identifiable user agents, stable IP addresses, prudent request frequencies and compliance with terms of use. The result is that these very operators are intercepted and blocked by anti-bot systems, increasing operational costs that ultimately fall on rights holders. In professional practice, the asymmetry is tangible every day: lawful monitoring slows down, becomes fragmented and requires increasing human intervention, while illicit supply continues to scale through automation.

Alongside this operational problem lies a further issue, less visible but increasingly relevant from a compliance perspective. Modern anti-bot systems do not merely verify user interaction with a webpage; they frequently rely on fingerprinting techniques, behavioural analysis and the collection of technical data concerning devices and browsing activity. In recent years, these practices have attracted the attention of European data-protection authorities, fuelling debate over the compatibility of certain CAPTCHA systems with the GDPR principles of data minimisation, transparency and proportionality. Recent changes introduced by Google in the management of reCAPTCHA, together with the European debate over the qualification of privacy roles and interventions by supervisory authorities, confirm that the issue no longer concerns cybersecurity alone, but also the proper allocation of responsibility for personal-data processing.

The Impact on the Enforcement Chain

To assess the practical scope of the problem, one must look at the operational chain of intellectual-property protection, which today can only be automated. Combating counterfeiting and piracy at scale means continuously monitoring marketplaces, search engines, social networks, apps and showcase websites; classifying results; collecting evidence; sending notices through notice-and-takedown procedures or dedicated platform programmes; and subsequently verifying removal. None of these stages can withstand reliance on manual work when the phenomena involve tens of thousands of listings continuously reposted by networks of disposable accounts. Every anti-bot obstacle that interferes with detection propagates throughout the chain: what is not found cannot be documented, and what is not documented cannot be reported or removed.

The paradox acquires a temporal dimension where the legislature itself imposes tight deadlines. This is the case with Law No. 93/2023 and the Piracy Shield system, under which blocking measures must be implemented within thirty minutes of a report. Yet the entire upstream phase—identifying the illegal service, analysing the infrastructure and forensically collecting evidence—remains exposed to the technical frictions described above. This creates a peculiar inversion: the regulated phase moves quickly, while the phase that feeds it is structurally slowed by defences unable to distinguish the reporting party from the attacker. For live content, where the harm materialises within a few hours, every minute lost in detection transfers value to the illegal offering.

There is also an aspect well known to practitioners. The tools that platforms themselves make available to rights holders—from marketplace brand-protection programmes to IP reporting portals—largely presuppose manual interaction and rarely provide programmatic interfaces capable of matching the scale of the phenomenon. The diligent rights holder is therefore squeezed between anti-bot systems that hinder automated detection and reporting channels that, downstream, fail adequately to support automation.

This quantitative limitation is compounded by an even more insidious qualitative one. Even where access to these voluntary tools is granted, the information they provide is insufficient to build a comprehensive enforcement strategy. The clearest example concerns repeat infringement. Determining whether someone is a repeat infringer would require cross-referencing the seller’s identity, the history of established infringements, linked accounts or accounts reopened after closure, and conduct across multiple platforms. Voluntary programmes do not disclose these data, or disclose them only in fragmented form, confined to the individual report and individual platform. Legally, the point is far from marginal, because the DSA itself attaches importance to repeat infringement: Article 23 requires the suspension of users who frequently provide manifestly illegal content, while Article 30 requires the traceability of traders. This creates another short circuit. Platforms exclusively hold the data that determine whether those measures apply, but do not share them with those who generate the reports; meanwhile, the rights holder, who in practice bears the burden of documenting the phenomenon, lacks the information needed to demonstrate its serial nature. Each infringement is treated as an isolated episode, whereas operational experience shows that large-scale counterfeiting is, almost by definition, a serial and organised phenomenon.

The Regulatory Gap: Trusted Flaggers Recognised, but Not Recognisable

Against this technical backdrop lies the regulatory limitation already highlighted in these pages when the DSA became fully applicable: there can be no report without prior awareness of the content to be reported. Article 22 of the Regulation grants trusted flaggers a fast track for notices, which platforms must process with priority and without undue delay, and the DSA itself encourages the use of automated processes, starting with APIs, to receive and manage notices. Nothing comparable is provided, however, for the phase that logically precedes reporting. There is no obligation, even for VLOPs and VLOSEs, to provide facilitated tools and procedures for monitoring and searching for illegal content without geographical limits or visibility restrictions. Trusted flaggers are facilitated in reporting what they have found, but receive no comparable facilitation during the preliminary phase of searching for and monitoring illegal content: their access to the platform is no different from anyone else’s and remains fully subject to the anti-scraping measures described above. At the time, we observed that this gap risked making large-scale enforcement ineffective. The 2026 data show that the time has come.

Italian implementation followed the expected path. AGCOM, acting as Digital Services Coordinator, adopted Resolution No. 283/24/CONS laying down the rules for recognition of trusted-flagger status, and from 2025 the first recognitions were granted to entities active precisely in the protection of industrial property and the fight against online fraud. This is a positive development that gives substance to the institution. Legal recognition, however, does not entail technical recognisability. In the eyes of an anti-bot system, the crawler of a trusted flagger accredited by the authority is indistinguishable from an abusive scraper and is treated accordingly.

The point should be stated clearly: the current framework allows platforms to benefit twice. On the one hand, the absence of a general monitoring obligation—a fundamental principle confirmed by Article 8 DSA—places the burden of discovering illegal content on rights holders and their advisers. On the other hand, those same platforms indiscriminately deploy anti-bot defences against those rights holders, making discovery slow, costly and technically fragile. The issue is not to challenge the legitimacy of security measures, which respond to genuine needs. The issue is that their indiscriminate application, in the absence of any technical channel for accredited entities, turns an institution designed to strengthen enforcement into a label with no practical effect.

The window for intervention is open. The European Commission has been consulting on guidelines concerning the application of Article 22, with a deadline of 26 June 2026 and adoption expected in the second half of the year. Limiting those guidelines to accreditation procedures and subjective eligibility requirements, while excluding the technical dimension of the problem, would be a mistake.

From the Human/Machine Distinction to the Responsible/Non-Responsible Distinction

The technical direction has, in fact, already been mapped out, and it is striking that the same entity that documented the overtaking has helped map it. Cloudflare has promoted Web Bot Auth, a cryptographic authentication mechanism for automated traffic built on RFC 9421 (HTTP Message Signatures) and two IETF drafts. It is the first concrete, large-scale attempt to move beyond the traditional anti-bot logic based on suspicion and preventive blocking, replacing it with a model based on verifiable identification of the entity generating the traffic. The bot signs its requests with a verifiable key, enabling the website to distinguish it, recognise it and apply dedicated policies. Around this core, categories such as verified bots and signed agents are taking shape, together with commercial models such as pay-per-crawl, which enables website operators to monetise AI crawler access instead of blocking it.

The implicit paradigm retires the twentieth-century question “human or machine?” and replaces it with a far more useful one: “responsible traffic or not?”. In other words, identified or anonymous traffic, attributable to an entity accountable for its conduct or not. This is the distinction enforcement actually needs. It must be acknowledged, however, without illusions, that this identity infrastructure is being developed for commercial reasons—to govern and monetise AI-model crawling—not to protect rights. If the issue remains outside the regulatory agenda, rights holders risk being excluded from the next generation of technical rules as well, just as they are under the current one.

The resulting proposal is simple: link the legal accreditation provided for by Article 22 to verifiable technical credentials. A trusted flagger recognised by the national Coordinator should be able to cryptographically sign its monitoring traffic, and platforms should be required not to obstruct it indiscriminately, while remaining free to apply reasonable rate limiting and revoke access in cases of abuse. This could be achieved through guidelines, codes of conduct or, where necessary, legislative amendment. Such an architecture would provide safeguards for everyone: efficient and technically documentable monitoring for rights holders; full traceability and an identified, accountable counterpart for platforms, replacing the current cat-and-mouse game. Following the same logic, accreditation should also entail access to a qualified level of information. Subject to appropriate safeguards of proportionality and data protection, trusted flaggers should be able to access the information needed to document the serial nature of infringements, beginning with the history of accepted notices concerning the same trader. Without those data, Articles 23 and 30 DSA risk remaining provisions without an effective trigger.

Conclusions

Cloudflare’s documented overtaking symbolically closes an era in which automated traffic could be assumed to be the pathological exception and human traffic the physiological rule. From now on, automation is the norm online. What separates lawful from unlawful conduct is no longer the nature of the visitor, but the possibility of attributing responsibility to it.

The current system, built on outdated assumptions, distributes burdens in a way that practitioners in the field observe every day and that can hardly be described as proportionate. Transparent operators pay both for anti-bot defences and for the reduced effectiveness of monitoring; bad-faith operators pay almost nothing. The good intentions of the European legislature, which through the DSA has undoubtedly raised platform-accountability standards, are insufficient to eliminate the asymmetry because effective enforcement is determined on a technical terrain that the Regulation has not addressed.

Three priorities emerge:
(1) technical recognisability of accredited entities, to be developed through the Article 22 guidelines currently under consultation;
(2) platform transparency regarding the impact of their anti-bot systems on legitimate monitoring and investigation activities;
(3) adoption of open standards for authenticating automated traffic, preventing the field from remaining governed by proprietary solutions.

Data-protection law is now pushing in the same direction. An architecture based on the verifiable identity of automated traffic, rather than behavioural profiling of every visitor, is both more respectful of user privacy and more useful for enforcement. Without intervention on these fronts, the paradox is bound to worsen as autonomous agents spread. And a protection system that works primarily against those who comply with the rules is not merely inefficient: it is the opposite of what a regulatory framework should guarantee.

One final image summarises where we have arrived better than many analyses. CAPTCHA was created to expose machines; today, machines are better at passing it than we are. Thus, in the internet of bots, an almost mocking paradox takes shape: the only one still struggling with the “I’m not a robot” box is the human being. If you cannot solve it, you are most likely a flesh-and-blood person. We can allow ourselves that irony. But behind the joke lies a serious issue, because the same logic that now penalises the distracted user penalises, on an industrial scale, those who try to defend rights while operating in the open.

Sources and References

• N. Lasorsa Borgomaneri, M. Signorelli, “Digital Service Act: il difficile compito dei Trusted Flaggers”, Agenda Digitale, 31 October 2023.
• M. Chillau, “Captcha, che sofferenza: così incidono sugli acquisti online”, Agenda Digitale, 10 November 2021.
• A. Caffo, “L'IA si è presa Internet, i bot generano più traffico web dell'uomo”, ANSA, 8 June 2026.
• NBC News, “Bot web traffic has overtaken human web traffic, data shows”, June 2026.
• Imperva (Thales), “2025 Bad Bot Report”, April 2025.
• Cloudflare, “Forget IPs: using cryptography to verify bot and agent traffic” (Web Bot Auth) and “The age of agents: cryptographically recognizing agent traffic” (signed agents).
• European Commission, targeted consultation on draft guidelines concerning trusted flaggers under Article 22 DSA (deadline: 26 June 2026).
• AGCOM, Resolution No. 283/24/CONS (rules on recognition of trusted-flagger status) and list of recognised entities.
• Law No. 93 of 14 July 2023 (provisions for preventing and combating the illegal dissemination of copyright-protected content) and the Piracy Shield platform.
• R. Pagano, “Google reCAPTCHA e GDPR: inquadramento giuridico, enforcement europeo e le novità operative dal 2 aprile 2026”, IusPrivacy.eu, March 2026.
• CNIL, Decision SAN-2023-006 of 16 March 2023 (Cityscoot), adopted in cooperation with the Italian and Spanish supervisory authorities.
• Google Cloud, “Switching Google's role with reCAPTCHA from data controller to data processor” (effective 2 April 2026).
• Baymard Institute, research on CAPTCHA and checkout; MOZ, “Captchas' Effect on Conversion Rates”.