OpenAI breach is a reminder that AI companies are treasure troves for hackers

12:49 PM PDT • July 5, 2024

Image Credits: Bryce Durbin / TechCrunch

There’s no need to worry that your secret ChatGPT conversations were obtained in a recently reported breach of OpenAI’s systems. The hack itself, while troubling, appears to have been superficial — but it’s a reminder that AI companies have in short order made themselves into one of the juiciest targets out there for hackers.

The New York Times reported the hack in more detail after former OpenAI employee Leopold Aschenbrenner hinted at it recently in a podcast. He called it a “major security incident,” but unnamed company sources told the Times the hacker only got access to an employee discussion forum. (I reached out to OpenAI for confirmation and comment.)

No security breach should really be treated as trivial, and eavesdropping on internal OpenAI development talk certainly has its value. But it’s far from a hacker getting access to internal systems, models in progress, secret roadmaps, and so on.

But it should scare us anyway, and not necessarily because of the threat of China or other adversaries overtaking us in the AI arms race. The simple fact is that these AI companies have become gatekeepers to a tremendous amount of very valuable data.

Let’s talk about three kinds of data OpenAI and, to a lesser extent, other AI companies created or have access to: high-quality training data, bulk user interactions, and customer data.

It’s uncertain what training data exactly they have, because the companies are incredibly secretive about their hoards. But it’s a mistake to think they are just big piles of scraped web data. Yes, they do use web scrapers or datasets like the Pile, but it’s a gargantuan task shaping that raw data into something that can be used to train a model like GPT-4o. A huge amount of human work hours are required to do this — it can only be partially automated.

AI training data has a price tag that only Big Tech can afford

Some machine learning engineers have speculated that of all the factors going into the creation of a large language model (or, perhaps, any transformer-based system), the single most important one is dataset quality. That’s why a model trained on Twitter and Reddit will never be as eloquent as one trained on every published work of the last century. (And probably why OpenAI reportedly used questionably legal sources like copyrighted books in their training data, a practice they claim to have given up.)

So the training datasets OpenAI has built are of tremendous value to competitors, from other companies to adversary states to regulators here in the U.S. Wouldn’t the Federal Trade Commission (FTC) or courts like to know exactly what data was being used, and whether OpenAI has been truthful about that?

But perhaps even more valuable is OpenAI’s enormous trove of user data — probably billions of conversations with ChatGPT on hundreds of thousands of topics. Just as search data was once the key to understanding the collective psyche of the web, ChatGPT has its finger on the pulse of a population that may not be as broad as the universe of Google users, but provides far more depth. (In case you weren’t aware, unless you opt out, your conversations are being used for training data.)

AI-powered scams and what you can do about them

In the case of Google, an uptick in searches for “air conditioners” tells you the market is heating up a bit. But those users don’t then have a whole conversation about what they want, how much money they’re willing to spend, what their home is like, manufacturers they want to avoid, and so on. You know this is valuable because Google is itself trying to convert its users to provide this very information by substituting AI interactions for searches!

Think of how many conversations people have had with ChatGPT, and how useful that information is, not just to developers of AIs, but also to marketing teams, consultants, analysts … It’s a gold mine.

The last category of data is perhaps of the highest value on the open market: how customers are actually using AI, and the data they have themselves fed to the models.

Hundreds of major companies and countless smaller ones use tools like OpenAI and Anthropic’s APIs for an equally large variety of tasks. And in order for a language model to be useful to them, it usually must be fine-tuned on or otherwise given access to their own internal databases.

This might be something as prosaic as old budget sheets or personnel records (e.g., to make them more easily searchable) or as valuable as code for an unreleased piece of software. What they do with the AI’s capabilities (and whether they’re actually useful) is their business, but the simple fact is that the AI provider has privileged access, just as any other SaaS product does.

These are industrial secrets, and AI companies are suddenly right at the heart of a great deal of them. The newness of this side of the industry carries with it a special risk in that AI processes are simply not yet standardized or fully understood.

Hugging Face says it detected ‘unauthorized access’ to its AI model hosting platform

Like any SaaS provider, AI companies are perfectly capable of providing industry standard levels of security, privacy, on-premises options, and generally speaking providing their service responsibly. I have no doubt that the private databases and API calls of OpenAI’s Fortune 500 customers are locked down very tightly! They must certainly be as aware or more of the risks inherent in handling confidential data in the context of AI. (The fact that OpenAI did not report this attack is their choice to make, but it doesn’t inspire trust for a company that desperately needs it.)

But good security practices don’t change the value of what they are meant to protect, or the fact that malicious actors and sundry adversaries are clawing at the door to get in. Security isn’t just picking the right settings or keeping your software updated — though of course the basics are important too. It’s a never-ending cat-and-mouse game that is, ironically, now being supercharged by AI itself: Agents and attack automators are probing every nook and cranny of these companies’ attack surfaces.

There’s no reason to panic — companies with access to lots of personal or commercially valuable data have faced and managed similar risks for years. But AI companies represent a newer, younger, and potentially juicier target than your garden-variety, poorly configured enterprise server or irresponsible data broker. Even a hack like the one reported above, with no serious exfiltrations that we know of, should worry anybody who does business with AI companies. They’ve painted the targets on their backs. Don’t be surprised when anyone, or everyone, takes a shot.

AI aids nation-state hackers but also helps US spies to find them, says NSA cyber director

More TechCrunch

Faulty CrowdStrike update causes major global IT outage, taking out banks, airlines and businesses globally

Security giant CrowdStrike said the outage was not caused by a cyberattack, as businesses anticipate widespread disruption.

Manish Singh

56 mins ago

Faulty CrowdStrike update causes major global IT outage, taking out banks, airlines and businesses globally

Security

From the Sphere to false cyberattack claims, misinformation runs rampant amid CrowdStrike outage

Amanda Silberling

1 hour ago

This serves as an example for how easy it is to spread inaccurate information online during a time of immense global confusion and panic.

From the Sphere to false cyberattack claims, misinformation runs rampant amid CrowdStrike outage

Transportation

Here’s how the CrowdStrike outage is affecting planes, trains and automobiles

Rebecca Bellan

1 hour ago

The CrowdStrike outage that hit early Friday morning and knocked out computers running Microsoft Windows has grounded flights globally. Major U.S. airlines including United Airlines, American Airlines and Delta Air…

Here’s how the CrowdStrike outage is affecting planes, trains and automobiles

Security

What we know about CrowdStrike’s update fail that’s causing global outages and travel chaos

Zack Whittaker

Lorenzo Franceschi-Bicchierai

1 hour ago

Here’s everything you need to know so far about the global outages caused by CrowdStrike’s buggy software update.

What we know about CrowdStrike’s update fail that’s causing global outages and travel chaos

TechCrunch Disrupt 2024

Last chance today: Secure major savings for TechCrunch Disrupt 2024!

TechCrunch Events

2 hours ago

Today is the final chance to save up to $800 on TechCrunch Disrupt 2024 tickets. Disrupt Deal Days event will end tonight at 11:59 p.m. PT. Don’t miss out on…

Last chance today: Secure major savings for TechCrunch Disrupt 2024!

Fintech

Paytm loss widens and revenue shrinks as it grapples with regulatory clampdown

Manish Singh

10 hours ago

Indian fintech Paytm’s struggles won’t seem to end. The company on Friday reported that its revenue declined by 36% and its loss more than doubled in the first quarter as…

Paytm loss widens and revenue shrinks as it grapples with regulatory clampdown

Venture

Fandango founder dies in fall from Manhattan skyscraper

Maxwell Zeff

17 hours ago

J. Michael Cline, the co-founder of Fandango and multiple other startups over his multi-decade career, died after falling from a Manhattan hotel, New York’s Deputy Commissioner of Public Information tells…

Fandango founder dies in fall from Manhattan skyscraper

Security

Researcher finds flaw in a16z website that exposed some company data

Lorenzo Franceschi-Bicchierai

19 hours ago

Venture capital giant a16z fixed a security vulnerability in one of the firm’s websites after being warned by a security researcher.

Researcher finds flaw in a16z website that exposed some company data

Media & Entertainment

Apple Vision Pro debuts immersive content featuring NBA players, The Weeknd and more

Lauren Forristal

20 hours ago

Apple on Thursday announced its upcoming lineup of immersive video content for the Vision Pro. The list includes behind-the-scenes footage of the 2024 NBA All-Star Weekend, an immersive performance by…

Apple Vision Pro debuts immersive content featuring NBA players, The Weeknd and more

Transportation

Elon Musk is now a villain in Joe Biden’s presidential campaign

Sean O'Kane

20 hours ago

Biden centering Musk in his campaign is a notable escalation, considering he spent most of his presidency seemingly pretending the billionaire didn’t exist.

Elon Musk is now a villain in Joe Biden’s presidential campaign

Transportation

Waymo wants to bring robotaxis to SFO, emails show

Kirsten Korosec

21 hours ago

Waymo would need a ground transportation permit to operate at SFO, which has yet to be approved.

Waymo wants to bring robotaxis to SFO, emails show

Startups

Why it made sense for an online community college to raise venture capital

Rebecca Szkutak

21 hours ago

When Tade Oyerinde first set out to fundraise for his startup, Campus, a fully accredited online community college, it was incredibly difficult. VCs have backed for-profit education companies in the…

Why it made sense for an online community college to raise venture capital

Startups

PE firm PartnerOne paid $28M for HeadSpin, a fraction of its $1.1B valuation set by ICONIQ and Dell Technologies Capital

Marina Temkin

21 hours ago

Canadian private equity firm PartnerOne paid $28.2 million for HeadSpin, a mobile app testing startup whose founder was sentenced for fraud earlier this year, according to documents viewed by TechCrunch.…

PE firm PartnerOne paid $28M for HeadSpin, a fraction of its $1.1B valuation set by ICONIQ and Dell Technologies Capital

Meta puts a halt to training its generative AI tools in Brazil

Lauren Forristal

22 hours ago

Meta has suspended the use of its AI assistant after Brazil’s National Data Protection Authority (ANPD) banned the company from training its AI models on personal data from Brazilians. The…

Meta puts a halt to training its generative AI tools in Brazil

ChatGPT: Everything you need to know about the AI-powered chatbot

Kyle Wiggers

Cody Corrall

Alyssa Stringer

23 hours ago

ChatGPT, OpenAI’s text-generating AI chatbot, has taken the world by storm since its launch in November 2022. What started as a tool to hyper-charge productivity through writing essays and code…

ChatGPT: Everything you need to know about the AI-powered chatbot

Crypto

WazirX halts withdrawals after losing $230 million, nearly half its reserves

Manish Singh

23 hours ago

The Mumbai-based firm said one of its multisig wallets had suffered a security breach, and it was temporarily pausing all withdrawals from the platform.

WazirX halts withdrawals after losing $230 million, nearly half its reserves

Transportation

Fisker scores a win, an AV startup reboots in Texas, and why Elon pushed the Tesla robotaxi reveal

Kirsten Korosec

23 hours ago

This week’s TechCrunch Mobility looks at Fisker scoring a win, an AV startup rebooting in Texas, why Elon is pushing the Tesla robotaxi reveal and more.

Fisker scores a win, an AV startup reboots in Texas, and why Elon pushed the Tesla robotaxi reveal

Apps

What is Apple Intelligence, when is it coming and who will get it?

Brian Heater

24 hours ago

Apple Intelligence was designed to leverage things that generative AI already does well, like text and image generation, to improve upon existing features.

What is Apple Intelligence, when is it coming and who will get it?

Government & Policy

The EU just re-elected its president for another five years — here’s what that means for tech

Natasha Lomas

24 hours ago

The European Union’s president, Ursula von der Leyen, was confirmed in the role for another five years Thursday after parliamentarians voted overwhelmingly to re-elect her. The scale of her support…

The EU just re-elected its president for another five years — here’s what that means for tech

Apps

Communia bets social media can be good for you

Amanda Silberling

1 day ago

Olivia DeRamus is flipping the script: What if scrolling through social media didn’t make us miserable? What if, especially for women, social media could actually make us feel more supported?…

Communia bets social media can be good for you

Apps

TikTok fast-tracks artist account creation for DistroKid members

Ivan Mehta

1 day ago

TikTok is partnering with the music distribution service DistroKid to fast-track the creation of artist accounts for members. The ByteDance-owned short video platform introduced an Artist Account feature last year…

TikTok fast-tracks artist account creation for DistroKid members

Transportation

Ford’s EV plans are in flux once again as it invests $3B into its biggest trucks

Kirsten Korosec

1 day ago

Ford is still pushing forward on electrification, notably by increasing hybrid options.

Ford’s EV plans are in flux once again as it invests $3B into its biggest trucks

OpenAI unveils GPT-4o mini, a smaller and cheaper AI model

Maxwell Zeff

1 day ago

OpenAI introduced GPT-4o mini on Thursday, its latest small AI model. The company says GPT-4o mini, which is cheaper and faster than OpenAI’s current cutting-edge AI models, is being released…

OpenAI unveils GPT-4o mini, a smaller and cheaper AI model

Featured Article

USPS shared customer postal addresses with Meta, LinkedIn and Snap

The U.S. Postal Service confirmed it took action to “remediate” the data sharing following a TechCrunch investigation.

Zack Whittaker

1 day ago

USPS shared customer postal addresses with Meta, LinkedIn and Snap

TechCrunch Disrupt 2024

GM CEO Mary Barra is coming to TechCrunch Disrupt 2024

Kirsten Korosec

1 day ago

The automotive industry is in the midst of dramatic technological change as companies seek out new ways to make money beyond building and selling gas-powered cars. And GM CEO and…

GM CEO Mary Barra is coming to TechCrunch Disrupt 2024

Commerce

Amazon Prime Day 2024 sales hit record $14.2 billion

Ivan Mehta

1 day ago

Amazon’s Prime Day event clocked record sales this year, as U.S. consumers spent $14.2 billion across July 16 and 17, according to Adobe Analytics. This marks an 11% jump from…