ownlife-web-logo
AnalysisMetaEthicsRegulationAugust 29, 20267 min read

The Data Scraping Double Standard: Why Individuals Get Prosecuted and Meta Gets Sued

Aaron Swartz faced 35 years in prison for downloading academic papers. Meta torrented 81 terabytes of copyrighted material to train AI models. The leg...

Sponsor

Photo by Mariia Berezovsky on Unsplash

The Data Scraping Double Standard: Why Individuals Get Prosecuted and Meta Gets Sued

Aaron Swartz faced 35 years in prison for downloading academic papers. Meta torrented 81 terabytes of copyrighted material to train AI models. The legal system treats these acts very differently, and the reasons why should concern every developer and small operator on the internet.

In 2013, Aaron Swartz — co-creator of RSS, early architect of Creative Commons, and one of the internet's most committed advocates for open knowledge — took his own life while facing federal prosecution under the Computer Fraud and Abuse Act. His alleged crime: downloading roughly 70 gigabytes of academic articles from JSTOR, a digital library, with the apparent intent of making them freely available. The charges he faced included up to 35 years in prison, a $1 million fine, and asset forfeiture.

Today, Meta is in active litigation for scraping copyrighted content at a scale that dwarfs Swartz's downloads by several orders of magnitude. The company faces civil lawsuits, not criminal prosecution. No executive is looking at prison time. The contrast is hard to ignore, and it reveals something structural about how the American legal system handles data access depending on who's doing the accessing.

What Meta Actually Did

The specifics matter here. As 404 Media reported, internal Meta emails revealed in litigation showed that the company downloaded over 81 terabytes of data by scraping Anna's Archive, an open search engine used for torrenting copyrighted books, movies, TV shows, and adult content. That's not a rogue employee downloading files on a lunch break. Strike 3 Holdings, the company behind adult content sites Blacked, Vixen, and Tushy, found that 47 IP addresses belonging to Meta were used to torrent 2,396 of its videos a total of 6,008 times between 2018 and 2025.

When sued, Meta tried to get the case dismissed by arguing that the downloads might have been the work of individual employees acting on their own, not a coordinated corporate effort. The judge wasn't buying it. As 404 Media reported, U.S. District Court Judge Eumi K. Lee found that Meta's attempt to blame the scraping on rogue employees "strains credulity," noting that Strike 3 Holdings' investigation showed coordination across Meta's IP addresses. The lawsuit is moving forward.

But "moving forward" in civil court is a fundamentally different experience than being hauled into federal criminal court. Meta's worst-case outcome is a financial settlement. Swartz was facing the destruction of his life.

The CFAA Gap

The Computer Fraud and Abuse Act, originally passed in 1986, is the federal statute that made Swartz's prosecution possible. It was designed to criminalize unauthorized access to computer systems, but its language is broad enough that prosecutors have used it to go after activities that look more like terms-of-service violations than hacking. Swartz accessed JSTOR through MIT's network, which he was authorized to use as a research fellow. The government argued he exceeded that authorization by downloading articles in bulk using a script.

The CFAA gives federal prosecutors enormous discretion. They can choose to bring criminal charges for conduct that might otherwise be a civil dispute between a platform and a user. In Swartz's case, prosecutors chose to pursue the maximum possible charges, reportedly rejecting plea deals that would have resulted in minimal jail time.

Meta's scraping, by contrast, is being handled entirely through civil litigation brought by private parties. No federal prosecutor has filed criminal charges against Meta for torrenting copyrighted material. The DOJ has shown no interest in applying the same CFAA framework that it used against Swartz to a company with a market cap in the hundreds of billions.

The structural reasons are straightforward but uncomfortable. Federal prosecutors have limited resources and tend to pursue cases they can win cleanly. Going after an individual downloading academic papers is a simple case with a clear defendant. Going after one of the world's largest technology companies, which employs armies of lawyers and lobbyists, is a multi-year war with uncertain outcomes. Prosecutorial discretion, in practice, means prosecutorial path-of-least-resistance.

The Scale Problem

A blog post by CuriousQuail, writing during Blaugust 2026, captures the frustration bluntly: Swartz downloaded about 70 gigabytes of academic articles. Meta torrented 80-plus terabytes of copyrighted material. That's roughly a thousand-to-one ratio in data volume, and the person who downloaded less is the one who faced criminal charges.

The use cases diverge sharply too. Swartz wanted to make publicly funded research freely accessible. Meta wanted training data for proprietary AI models that generate revenue. You don't have to agree with Swartz's methods to notice that the legal system came down harder on the person whose goals were non-commercial.

This isn't just historical grievance. It has direct implications for anyone building software today. If you're an independent developer who scrapes a website in violation of its terms of service, you could theoretically face CFAA prosecution. If you're a trillion-dollar company doing the same thing at industrial scale, you face civil suits that your legal team can manage as a cost of doing business.

Meta's Expanding Footprint

The scraping lawsuits arrive at a moment when Meta is investing aggressively in the AI infrastructure that this training data feeds. As we covered in our reporting on Meta's $10 billion Tulsa data center, the company is building AI-optimized facilities across the country, designed specifically for the GPU-dense compute that trains and runs large language models. The scraped data has a destination, and Meta is spending billions to build it.

Meanwhile, Meta's public messaging frames its AI efforts as broadly beneficial. In a letter published on August 10, Mark Zuckerberg laid out a vision for AI as "the path to a positive AI future," arguing for wider distribution of AI access across countries and companies. But as Rest of World reported, global AI experts pushed back hard on that framing. Chinasa T. Okolo, founder of policy incubator Technecultura, told Rest of World that "Meta has grossly overstated the economic benefits of its data centers." Others questioned who gets to decide what "everyone" means when a single company controls the infrastructure.

The gap between the public narrative and the legal record is striking. Meta positions itself as democratizing AI while facing lawsuits alleging it built that AI on torrented copyrighted material. The company's defense in the Strike 3 Holdings case — that maybe employees just happened to download thousands of porn videos from company IP addresses — suggests it's not eager to own the full scope of its data acquisition strategy.

What This Means for Developers and Small Operators

The practical lesson for anyone who isn't a Fortune 50 company is sobering. The same laws that could be used to prosecute you for scraping a website are functionally unenforced against companies with sufficient legal and lobbying resources. The CFAA doesn't distinguish between a graduate student and a corporation. But the enforcement apparatus does.

This creates a chilling effect that runs in one direction. Independent researchers, open-source developers, and small companies have to think carefully about data access in ways that large platforms simply don't. A startup that scraped training data the way Meta did would face existential legal risk. Meta faces a line item.

There's no clean fix here. The CFAA is overdue for reform, and various proposals have circulated in Congress for years without gaining enough traction to pass. Aaron's Law, a bill introduced after Swartz's death to narrow the CFAA's scope, has never been enacted. Meanwhile, the question of whether AI training on copyrighted material constitutes fair use remains unsettled, with multiple cases working through the courts.

The Swartz case and the Meta scraping lawsuits aren't perfectly analogous. One involved accessing a system in arguably unauthorized ways; the other involved downloading publicly available torrents. The legal theories differ. But the disparity in consequences — criminal prosecution and potential decades in prison versus civil litigation and potential financial settlements — reflects something real about who the legal system is designed to protect and who it's designed to punish.

Until that changes, the rules of data access on the internet will continue to mean different things depending on how many lawyers you can afford.

What's your next step?

Every journey begins with a single step. Which insight from this article will you act on first?

Sponsor