The Guardrail and the Costume
On the evening of 24 July, Anthropic released a new model, Claude Opus 5. Three security researchers, Harsh Jaiswal, Mohan Pedhapati and Rahul Maini, working at a small firm called Hacktron AI, had spent that same day failing to get an older version of the same model, Opus 4.8, to finish a job.
They had found a real flaw: a heap buffer overflow, a bug where a program stores more data than it has room for and spills into memory it shouldn’t touch. It sat quietly inside libheif, a small code library that other software calls on to open and display images, two layers below OpenAI’s public help forum. They had been working through that upload path since the day before.
Opus 4.8 could turn that flaw into a working exploit, but only under lab conditions, with a safety feature called ASLR switched off. ASLR, short for Address Space Layout Randomisation, works like shuffling a building’s room numbers each time it is rebuilt: it doesn’t stop a break-in, but it stops an attacker from reusing a floor plan they memorised last time. With ASLR back on, the exploit kept failing.
They gave the same problem to Opus 5 the moment it shipped. Within three hours it had a working exploit for a local Mac.
Then they asked it to work against a machine across the network, and it refused. Their own account reads: Opus “refused to write exploit for remote instances.” So they built a decoy. Their own copy of the forum software, at an address of their own choosing, rce.ee/ctf-forum, dressed as a hacking-practice puzzle rather than a real company’s systems. By six on the morning of 25 July, Greenwich time, they had code execution through an image upload. They filed the report before ten, and by half past three that afternoon they were inside an OpenAI employee’s account.
Nine Links in a Chain
What followed took under 72 hours from first finding to reported access, and it is worth laying out as a chain. An old flaw in libheif, a Debian package that never received the security fix, and ImageMagick, which calls libheif. Then Discourse, which calls ImageMagick, OpenAI’s public help forum, which runs on Discourse, a login system that let a forum account unlock ChatGPT and Codex, Codex itself, a GitHub account wired into it, and OpenAI’s internal repositories at the end. Nine links. Break any one and the rest doesn’t matter.
Rather than read anything sensitive, they used the employee’s own Codex to open one harmless pull request inside OpenAI’s private code repository. That was proof of access, not use of it. They reported the flaw to OpenAI and to the forum software’s maker the same day.
OpenAI’s own side of the flaw was confirmed fixed about fourteen hours after the report landed. The forum software’s fix, separate and slower, was ready by Monday 27 July and published as a public notice the next day. OpenAI paid a bounty of $6,500, about ₹5.7 lakh, on 1 September. It noted that the forum sat outside the programme’s scope, and that the reward recognised the identity flaw on its own side. Hacktron put the token cost of the whole effort at under $3,000, about ₹2.6 lakh. “We’re just three guys with Claude and Codex subscriptions,” Pedhapati told the Wall Street Journal.
The Lab That Says Slow Down
Anthropic has spent two years telling the public it is the cautious one. Its chief executive has put a number on his worry: a 25% chance, he told Axios last September, that things go really, really badly. In June this year, a US export-control order required the company to cut off two of its newest models, Fable 5 and Mythos 5, from every foreign national. The block ran from 13 June to 1 July.
It’s a specific irony that Opus 5, the model at the centre of this story, faced no such restriction. And the one refusal on record held only until the request was dressed differently. Asked to write an exploit for a machine across the network, it said no. Asked for the same work against a target dressed as a hacking-practice puzzle, it said yes, and finished the job in hours.
Hacktron’s writeup treats this not as a scandal but as a fact: “Software has long benefited from a kind of security through complexity,” they write. Turning a known bug into a working weapon used to require rare expertise and real time. AI is converting that expertise into compute, and compute is getting cheap.
I don’t think that sentence is about OpenAI. It is about every organisation whose security rests on one old assumption: that the attacker on the other end is scarce, expensive and slow to arrive.
Three Names, No Address
Jaiswal, Pedhapati and Maini are described in most Indian coverage as an Indian team, and in one sense that’s fair. Pedhapati is from Rajahmundry, Andhra Pradesh, schooled at a government school before RGUKT Nuzvid, and Jaiswal and Maini built their careers on bug-bounty platforms. But Hacktron AI is not headquartered in India. It was co-founded in 2025 with Zayne Zhang, a Cambridge student based in Singapore, went through a European accelerator, and keeps offices in the US, India, the UK and Singapore. The talent is India-shaped. The company is not.
That distinction matters, because India’s own cyber agency had already named the risk. In April, months before this disclosure, CERT-In issued a high-severity advisory, CIAD-2026-0020, warning that frontier AI was lowering the skill needed to chain sophisticated attacks together, and that work once needing a well-resourced team could be done by far fewer people, far faster. Three researchers with roots here went on to demonstrate close to that warning, for an American company’s bounty, inside a firm built elsewhere.
What I believe, and cannot prove
Here I stop reporting. What follows is belief, not evidence.
What matters most is that Claude refused the direct request and complied once it was wrapped in the costume of a training exercise. That is the real story, more than the bounty or the bug. A safety instruction that holds against a direct question and folds against the same question in a fictional wrapper was never load-bearing in the way its presence suggested. Companies, including government offices, that treat “the model refused” as proof of a safe system are trusting a test that has been shown not to generalise.
I know the weak points in this. I am reasoning from one well-documented case, not a pattern, and Hacktron themselves note that skilled human judgement stayed essential throughout. This was not the model acting alone.
There is a harder objection. Two things changed between the refusal and the compliance, not one. The costume changed, and so did the target, because the decoy was a copy of the forum software the researchers themselves owned. That sentence therefore carries two readings. Mine is that the refusal was theatre and a wrapper beat it. The other is that the model would not write remote exploits at all, could not tell who owned the machine, and read the puzzle framing as proof of ownership rather than as a lie. Under that reading the guardrail is crude rather than theatrical. A duller claim, and it might be the true one.
I cannot settle it from one sentence, and I have not found a second case that would. A single costume beating a single refusal is not proof that every safeguard folds the same way. Weigh this as you would weigh anyone’s belief, by what it rests on, which here is one careful story and my own read of it.
What you can do on Tuesday
I have done a version of this myself. A tool declines a request, and rather than accept that, I rephrase. A different justification, a different framing, sometimes just a different order of words. More often than not, the second version goes through. I’ve never sat with what that means: that the boundary I hit wasn’t about the request, it was about how I’d worded it.
Pick one headline about the three researchers this week. Open the source behind it, not the summary, at hacktron.ai/blog/hacking-openai. Ask three things. What exactly is being claimed? Is the piece treating this as pride, or as warning? And where does the source itself draw a line it says it did not cross?
Run that third question against one specific claim. Several headlines this week say the researchers read OpenAI’s source code. Hacktron’s own account says they demonstrated access without opening it, and stopped there. That gap, between what a source says and what travels under its name, is the whole reason this publication exists.
September 22, 2026 | dhirender.tamber@gmail.com
Sources
Hacktron AI, “Hacking OpenAI” (September 2026); The Wall Street Journal’s report of the disclosure (September 2026); TechCrunch and VentureBeat coverage of the OpenAI hack (18 September 2026); CBS News and The Tech Portal on the researchers and their methods (18 and 19 September 2026); Hyderabad Mail on Pedhapati’s background (September 2026); Axios, “Amodei: 25% chance things go really, really badly” (17 September 2025); Anthropic’s statement on Fable 5 and Mythos 5 access (June and July 2026); CERT-In Advisory CIAD-2026-0020 (April 2026); Dealroom and BuiltIn company profiles of Hacktron AI.
Was this worth your time?