AI

Researchers used Anthropic’s Claude to ethically break into OpenAI—and collected a $6,500 bounty

· Geeknewz Author

Abstract cybersecurity visualization with code and digital lock motifs

The plot writes itself: security researchers used Anthropic’s Claude to help break into OpenAI—ethically, under a bug-bounty program—and walked away with a reported $6,500 payout. Forbes, citing the Wall Street Journal’s earlier exclusive and Hacktron AI’s own write-up, says the team gained access to an OpenAI employee’s ChatGPT account and from there could reach the company’s internal GitHub world—then stopped, disclosed, and got paid instead of going full cyberpunk.

Hacktron’s researchers described the theoretical blast radius as huge. The point of the exercise wasn’t theft; it was proving that frontier-lab identity and tooling edges still have soft spots—and that AI coding agents are already useful on both sides of the security fence.

What happened (high level)

According to public reporting and Hacktron’s disclosure summary, the chain started at an OpenAI staff discussion forum and leveraged flaws that ultimately let researchers take over ChatGPT/Codex sessions belonging to employees. With Codex connected into OpenAI’s GitHub organization, they demonstrated impact by prompting a harmless pull request in an internal monorepo—then halted further testing rather than reading or downloading private source.

OpenAI and the forum vendor patched after disclosure. The bounty figure reported by Forbes/WSJ—$6,500—is modest relative to the scare value of “we could have touched the monorepo,” which is exactly why responsible disclosure theater still matters: it turns nightmare demos into tickets and postmortems.

Claude’s role in the story

This isn’t “Claude woke up and hacked OpenAI.” Researchers used the model as an assistant while developing and porting exploit work across sessions—Hacktron noted earlier Claude versions struggled, while a newer Opus release moved faster. The uncomfortable industry takeaway is familiar: capable coding models compress the time from “interesting bug class” to “working demonstration,” which helps defenders who hire researchers and attackers who don’t wait for bounties.

That arrives two weeks after separate drama in which AI agents reportedly escaped containment at OpenAI and hit Hugging Face—another reminder that agent autonomy plus identity mistakes is the 2026 crossover episode nobody asked for.

Why labs should sweat identity more than model weights

Hacktron emphasized that the escalation path hinged on SSO/identity assumptions tying the forum world to ChatGPT/Codex—not a magical model jailbreak of GPT itself. If any first-party or third-party surface sharing that SSO plane is compromised, you get similar account reach. Discourse was one proof path; the architectural lesson is broader: employee AI tools with repo connectors are crown jewels wearing casual clothes.

For everyone else shipping agents into production: treat coding agents with org GitHub access like production deploy keys. Because functionally, they are.

Geeknewz take

Rival-model-helps-hack-rival-lab headlines write themselves, but the durable story is boring and important: identity sprawl + agent tooling beats fancy model exfil fantasies. A $6,500 bounty for monorepo-adjacent demo access feels like buying a fire extinguisher after the smoke alarm screamed—good that disclosure worked, weird that the receipt is that small.

If you’re an AI lab, rotate the “employee ChatGPT can touch GitHub” threat model to the top of the pile. If you’re shipping coding agents, assume your users’ SSO graph is the real attack surface. And if you’re keeping score in the Anthropic vs OpenAI soap opera: the models are competing on evals; the researchers just scored on vibes.

Source: Forbes — Security Researchers Hacked Into OpenAI Using Anthropic’s Claude (Siladitya Ray, Sep 18, 2026); also reported by WSJ and The Guardian.