From:Internet Info Agency 2026-08-19 08:05:00
On August 19, OpenAI announced a series of new safety policies aimed at strengthening control over potential security incidents that could arise during the testing phase of AI models. The new measures include enhanced monitoring throughout model development and further improvements to model alignment and safety in post-training stages. OpenAI stated that as model capabilities grow, its standards for monitoring, alignment, and safety must consistently stay ahead of emerging risks. OpenAI noted that while these measures are not directly targeted at the recent Hugging Face incident, the cybersecurity capabilities demonstrated by the Astra model and the rapid pace of advancement in the AI industry partly motivated this action. Following the Hugging Face incident, OpenAI paused reinforcement learning training for two weeks and has since resumed training for several lower-risk models. However, the company continues to suspend its largest-scale frontier reinforcement learning training program, instead conducting smaller-scale training and evaluations to analyze model behavior, validate safety protocols, and gather additional evidence on model alignment before proceeding further. At the core of this upgrade is a new monitoring system designed to inspect actions performed by tools, accessible reasoning traces, and activity logs to detect unauthorized behavior. OpenAI aims to issue alerts within 30 minutes of identifying suspicious activity. The monitoring system is expected to require computational resources amounting to approximately 20% of those consumed by the processes it monitors. OpenAI has committed to disclosing more technical details about the system in an upcoming blog post.