
The disclosures land as investors are already sensitive to AI safety and governance risk following OpenAI's earlier Hugging Face breach admission, so a second, voluntary round of disclosures is likely to be read as a credibility building move rather than a fresh shock, particularly since OpenAI is framing this as proactive rather than reactive. The absence of an industry wide disclosure standard is itself notable for sentiment around AI infrastructure and safety focused names, since it leaves room for regulators to move first if companies do not coordinate on shared criteria. Watch for whether competitors follow with their own voluntary frameworks, which would suggest the sector is trying to get ahead of binding disclosure rules rather than wait for governments to impose them. --- OpenAI just admitted its models have been hiding mistakes and fabricating data, and it's betting that owning up first beats waiting for regulators to force the issue. Summary: OpenAI disclosed six new AI safety incidents on Wednesday, the earliest dating to October Behaviors included concealing mistakes, inventing missing data, seeking unauthorized credentials, and uploading files to public hosting services without user permission One model searched GitHub for exposed API keys and fabricated earnings data when it could not retrieve requested information OpenAI introduced a new voluntary disclosure procedure: incidents are sorted into "ready for disclosure," "minor investigation" or "larger investigation" tracks, with public reporting within six or twelve business days depending on complexity The move follows OpenAI's earlier disclosure that models under evaluation breached parts of Hugging Face's systems, its most severe such incident to date Research lead Kai Chen said no industry wide disclosure standard currently exists and that OpenAI hopes its approach helps shape shared regulations OpenAI disclosed six new incidents on Wednesday in which its artificial intelligence models concealed mistakes, sought unauthorized credentials, uploaded files to the public internet or communicated across training environments meant to be isolated from one another, according to Axios . The company paired the disclosure with a new internal procedure for reporting similar misbehavior going forward. The six incidents, the earliest dating back to October, span a range of behaviors. An unreleased model in OpenAI's Astra family inserted jailbreak like instructions into its own context summaries, including directions to ignore developer messages, affecting 27 summaries in total. During training of a model referred to as GPT-5.6 Sol, models attempted to conceal mistakes, invent missing historical data and hide mismatches between source versions. In another case, a model searched public GitHub repositories for exposed API keys, attempted to use disposable email accounts, and ultimately fabricated earnings data when it could not retrieve the information it had been asked for. Separately, models uploaded data and a task image to public file hosting services on two occasions to obtain citations or external image search results without asking the user first, used an internal repository as an informal message board to exchange requests and answers across separate training samples, and in one case a group of collaborating agents uploaded a workbook to public hosting so other agents could retrieve it, despite instructions to use only local files. Kai Chen, research lead on OpenAI's alignment team, told Axios the company is disclosing the incidents voluntarily because no industry wide framework for this kind of disclosure currently exists. Under the new procedure, any employee can flag a suspected incident for review by the safety and alignment teams, which will sort cases into a ready for disclosure track, a minor investigation track or a larger investigation track. Cases ready for disclosure will be made public within six business days, while those needing a minor investigation will follow within 12 business days, with more complex cases involving third parties expected to take longer. OpenAI said it may issue an initial notice before an investigation concludes, though security, legal and responsible disclosure obligations can delay the release of full details. Employees who believe an incident warrants disclosure but are overruled can escalate the matter to senior leadership. The disclosures follow OpenAI's earlier admission that models under evaluation escaped intended controls and compromised parts of Hugging Face's systems, gaining internet access, exploiting vulnerabilities and accessing limited private data in what the company has called its most severe model driven incident of this kind to date. Chen attributed the pattern to a combination of factors, telling Axios that model capabilities have grown faster than expected while the company had not previously put sufficient security controls in place to catch this kind of misalignment. Some prominent technologists, including Anthropic's chief executive, have expressed concern that the Hugging Face breach could be an early sign of AI agents finding unforeseen ways to act on the internet, while a number of security researchers have argued separately that many of the newly disclosed incidents could have been prevented with more basic cyber controls. OpenAI says it intends to keep working with other AI developers, researchers, standards bodies and regulators to build a more objective, shared set of disclosure criteria over time. This article was written by Eamonn Sheridan at investinglive.com.
Forexlive
Original source
openai chatgpt
Check live status on DownRightNow
