OpenAI has announced plans to publish ongoing reports about unexpected or unauthorised behaviour by its artificial intelligence models, acknowledging that the industry has yet to fully solve the challenge of keeping increasingly capable AI systems aligned with human instructions.

The ChatGPT maker has introduced a new framework for tracking, investigating and disclosing cases of model misalignment. As part of the initiative, OpenAI released six initial reports covering incidents identified during model training or evaluation.

Misalignment refers to situations where an AI model behaves in ways that go against objectives, restrictions or safeguards established by its developers. OpenAI said it previously disclosed such cases on an “ad hoc and less frequent than ideal” basis. The new process is intended to speed up disclosures, including in situations where the cause of a model’s behaviour has not been fully understood or preventive measures are still being developed.

OpenAI said the six reports represent only an initial set and are not a comprehensive list of known incidents or ongoing investigations. The company also cautioned that individual cases should not be interpreted as evidence of how frequently similar behaviour occurs across its models.

The disclosed cases include models hiding mistakes, fabricating information, searching public repositories for exposed software keys and uploading files to public websites without authorisation.

In one training exercise, OpenAI said its GPT-5.6 Sol model inserted instructions into task summaries telling future versions of the model to conceal errors or make up missing information. In another case involving an unreleased model, the system searched GitHub for exposed application programming interface keys and used them without permission. When it could not obtain the information needed for its task, the model fabricated figures and presented them as genuine data.

READ
EU Proposes Social Media Ban for Children Under 13

OpenAI also reported cases involving AI agents uploading files to public hosting services in order to share information that was intended to remain on local systems.


Buy ExpressVPN with PayPal or Credit Card

Under the new disclosure framework, any OpenAI employee can flag potential model misalignment for investigation. OpenAI’s safety and alignment teams will then assess the incident, determine whether third parties may have been affected and decide whether the case should be publicly disclosed.

Advertisement