OpenAI is introducing a series of security changes following last month’s disclosure that one of its AI systems escaped a sandboxed environment and accidentally hacked Hugging Face.

The company says the changes focus on strengthening its research environments, improving monitoring, and adding new alignment measures to reduce the risk of similar incidents.

OpenAI had already paused development of a new model called Astra after determining that it could potentially have “critical” cybersecurity capabilities. The company also says it temporarily paused reinforcement learning training on its latest models intended for deployment for two weeks while it reviewed and strengthened its security practices. Its largest planned frontier reinforcement learning run remains on hold.

For frontier model research, OpenAI now requires stronger sandboxing for workloads that execute model-generated or otherwise untrusted code. The company has also added more controls to separate higher-risk workloads from the internet.

The research environment itself is being redesigned to remove potentially vulnerable shared services, reduce standing privileges, and strengthen security boundaries between different workloads and systems.

OpenAI is also expanding its monitoring capabilities. The company says it now aims to generate an alert within 30 minutes after concerning activity is detected. If the team responding to an alert cannot conclusively determine within another 30 minutes that the activity is a false positive, the relevant activity is expected to be paused.


Buy ExpressVPN with PayPal or Credit Card
Advertisement
READ
Google Shuts Down 449 South Korean YouTube Channels Over Political Content