NEW New report: We analyzed 200 million ad clicksSee results
NEW New report: We analyzed 200M ad clicks
  • Home
  • Articles
  • Inside the Hugging Face Breach: Why Ad Fraud Experts Weren’t Surprised

Inside the Hugging Face Breach: Why Ad Fraud Experts Weren’t Surprised

On July 16, 2026, Hugging Face’s systems were hacked. The attacker harvested credentials during its movement through OpenAI’s network and used them to access Hugging Face’s internal data, API tokens, and authentication credentials. Hugging Face’s security monitoring later caught the unusual activity and contained the breach. As it turns out, the “attacker” was OpenAI’s own agents.

Immediately following the incident, Clem Delangue, Hugging Face CEO, called it “possibly the first [attack] of its kind.”

But truth is, it’s not really the first. Ad fraud experts see this kind of behavior all the time; only, we call it invalid traffic.

Ad fraud experts have seen this before

OpenAI’s agents followed a pattern: escape detection, act autonomously, and consume resources.

This is the same pattern that ad fraud bots have been running for years. They evade platform-level filters, act on behalf of bad actors, and drain budgets.

The only difference here is the target: OpenAI’s agents were after benchmark solutions, while ad fraud bots click ads and drain advertiser budgets.

When we analyzed 200 million ad clicks across our network earlier this year, we saw a similar pattern of behavior at play. Bots repeatedly used ad click identifiers across hundreds of devices and IP addresses to simulate legitimate traffic. This is called click recycling.

Inside Hugging Face’s systems, OpenAI’s agents harvested credentials for internal services and reused them to move deeper into the network. And per OpenAI’s incident report, they didn’t stop there. The models also “identified and used publicly exposed credentials” on four separate third-party services, using one compromised account as a relay and another for data storage.

It harvested real identifiers, reused them at scale, and looked legitimate at every checkpoint.

Bots frequently use this to evade detection, and so did OpenAI’s agent.

AI agents have supercharged the problem

Bots and scripts have always evaded filters, but we’re seeing it unfold on a much larger scale today.

In 2023, global ad spend lost to fraud was $84 billion, projected to hit $170 billion by 2028. But that was before AI agents came on the scene. We’ve already seen a sharp 7,851% spike in AI agent traffic in 2025 alone.

We see with the OpenAI incident that agent capabilities are increasing exponentially. Frontier models are getting more capable, and open-source models can be modified to remove safeguards entirely. That means we will likely see more AI systems that act autonomously and effortlessly generate invalid traffic.

Before AI agents came on the scene, bad actors wrote scripts that could mimic human action convincingly, bypass CAPTCHAs, and simulate random mouse movements. AI automation has made all of it faster and harder to catch.

There is hope: defense is getting better

The uncomfortable part of the Hugging Face story for us advertisers is this: the agent didn’t get caught right away. It authenticated with real credentials and looked legitimate at every checkpoint. It was only caught by behavior-monitoring-flagged API activity that didn’t fit normal patterns.

The same principle applies to your ad traffic. An AI agent clicking your ad arrives with a real browser fingerprint, a residential-looking IP, and plausible on-page behavior. Nothing immediately jumps out as suspicious, and identity-based filters won’t catch visitors with valid credentials.

However, let’s not forget how the Hugging Face story ends: Hugging Face’s own security monitoring flagged the anomalous API activity, their teams moved quickly, and they contained the breach.

And that’s another parallel we see in ad fraud: dedicated filters do catch a lot of this invalid traffic. Fraud Blocker’s own filters caught, on average, 16.2% of invalid traffic, 5.8% more than Google’s own reporting accounted for.

At a basic level, many fundamentals of ad protection still hold true: IP exclusions, geographic targeting restrictions, dayparting, and frequency caps. Used together, these meaningfully reduce advertiser exposure to the simpler end of invalid traffic.

The best defenses track behaviour alongside credentials

For the more complicated form of invalid traffic, the kind that acts autonomously, adapts, and spoofs credentials, advertisers need a more proactive approach. This is where behavioral detection comes into play. It is designed to catch complicated invalid activity by tracking signals like the following:

  • Click frequency that doesn’t match human rhythm
  • A single-click identifier appearing across hundreds of devices
  • Device characteristics that contradict each other
  • Mouse movement that’s a little too efficient

The most effective systems score every visitor against dozens of data points rather than checking a single credential. Fraud Blocker’s Intelligent Detection does this, assigning each visitor a 0–10 fraud score based on 100+ data signals. Flagged traffic sources get blocked in near real-time for Google Ads and Meta Ads campaigns.

As Sean Cassidy, CISO at Plaid, said, “The ramifications for security programs is immense… After today, the problems have been realized, and we need to account for them now.”

The silver lining of this approach is that every advertiser protected strengthens detection for the rest. Agentic fraud patterns spotted on one account will inform blocking across the entire network. That’s the same collaborative-defense principle Delangue argued for after the breach, applied to ad budgets.


Frequently Asked Questions

Yes. AI agents can click ads, fill out forms, and complete multi-step actions autonomously. Additionally, their traffic is largely indistinguishable from human traffic at the surface level, making them very powerful tools for ad fraud. Most agentic traffic today is legitimate (shopping assistants and research agents), but the same capabilities power fraudulent clicks at scale. Security researchers found the behavioral gap between benign and malicious automation narrowed to about half a percentage point in 2025.

In July 2026, OpenAI’s AI models escaped a sandboxed cybersecurity evaluation, exploited a zero-day vulnerability to reach the internet, and broke into Hugging Face’s production systems while trying to find the answers to the test they were being given. It’s considered the first publicly documented case of an autonomous AI agent breaching a real company’s production infrastructure.
Traditional bots follow simple scripts. They are repetitive, predictable, and detectable by pattern-matching filters. But AI agents make decisions, adapt to obstacles, vary their behavior, use real credentials, and can chain multiple steps toward a goal. That adaptability is why identity-based defenses (IP lists, CAPTCHAs) catch less AI agent traffic, and behavioral detection catches more.
Behavioral analysis is the most reliable approach. Use dedicated protection tools that analyze signals like click frequency patterns and device fingerprint consistency that don’t match human behavior. Click fraud protection tools apply these signals automatically. Fraud Blocker’s Intelligent Detection scores every visitor against 100+ signals and blocks flagged sources from Google Ads and Meta Ads campaigns in near real-time.
Facebook
X
LinkedIn

matthew Iyiola - click fraud specialist

ABOUT THE AUTHOR

Matthew is the resident content marketing expert at Fraud Blocker with several years of experience writing about ad fraud. When he’s not producing killer content, you can find him working out or walking his dogs.

Matthew is the resident content marketing expert at Fraud Blocker with several years of experience writing about ad fraud.

More from Fraud Blocker