The attack began with OpenAI’s community forum, which is hosted on the third-party Discourse platform. Hacktron found that certain HEIC and HEIF image uploads could reach a vulnerable version of the libheif image-processing library, allowing the researchers to develop an exploit capable of remote code execution.
That foothold alone did not provide access to OpenAI’s internal systems. Hacktron said it then identified a separate weakness in OpenAI’s single sign-on configuration that allowed control of the forum environment to be turned into access to accounts that had authenticated through the service, including an OpenAI employee’s ChatGPT account.
The employee account was connected to OpenAI’s developer infrastructure through Codex and GitHub. Rather than examining sensitive source code, Hacktron said the researchers used Codex to make a harmless change and prepare a pull request inside OpenAI’s internal monorepo, demonstrating how far the compromised credentials could reach.
“We thank the researchers for contacting us and sharing their findings,” OpenAI said. The company also said it narrowed permissions on Community sign-in tokens and revoked affected tokens and sessions.
Discourse separately confirmed the underlying image-processing vulnerability and patched affected versions while adding additional sandboxing around image handling.
Anthropic’s role came during development of the exploit itself. Hacktron said the team initially worked with Claude Opus 4.8 but had difficulty making the attack reliable. After moving to Claude Opus 5, the researchers said the model helped produce a working ARM64 exploit within hours and assisted in adapting it to the environment used by Discourse.
Hacktron said the full process from discovery to demonstrating access to OpenAI’s repository environment took less than 72 hours.
The incident shows how AI coding agents can reduce some of the specialized work traditionally required to turn software bugs into usable exploits. At the same time, it highlights the security consequences of giving AI accounts access to other corporate systems.
Hacktron said affected accounts could potentially connect to services such as GitHub, Slack, Outlook, Gmail and Google Drive. In the OpenAI case, the compromised employee account’s connection to GitHub meant the breach extended beyond a standalone chatbot account and into developer infrastructure.
That interconnected access increases the consequences of an identity compromise. An AI account that can reach code repositories, communications systems and document stores can effectively inherit the permissions granted across those services.
The disclosure also arrives as AI labs increasingly use their own models for research and development. Anthropic separately said that 26% of its research and development work was now “led by” Claude, up from 1% in March. The company defined that category as work in which AI completed most of a task based on human instructions and supervision.
Anthropic said its systems still did not operate fully autonomously in the research it examined. It said humans and AI collaborated on 90% of tasks, with the models carrying out substantial portions of the work.
The company said it released the figures to help the public “understand how close the world is to reaching recursive self-improvement,” referring to the point at which AI systems can contribute increasingly to the development of later systems.
For security teams, the OpenAI breach presents a more immediate issue: the same AI agents being connected to increasingly sensitive business systems are also becoming more capable at assisting with technically demanding exploit development. In this case, researchers working under an authorized bounty program were able to combine a third-party software flaw, an identity weakness and connected developer tools into a path that reached OpenAI’s internal environment.
This analysis is based on reporting from Financial Times.
Image courtesy of Bitcoin News.
This article was generated with AI assistance and reviewed for accuracy and quality.