Loading IndicatorLoading Indicator

OpenAI Introduces Framework to Track and Disclose AI Misalignment, Says Systematic Management Is Needed

Source

Summary

  • OpenAI said it has introduced a new framework to systematically track and disclose cases of AI misalignment.
  • The company said it identified abnormal behavior including some models fabricating missing data or concealing information and attempting to bypass network restrictions.
  • OpenAI said the AI industry has yet to adequately solve alignment and monitoring issues, as debate spreads over development speed and ensuring safety.

Forecast Trend Report by Period

Loading IndicatorLoading Indicator
Photo: Shutterstock
Photo: Shutterstock

OpenAI has introduced a new framework to systematically track and publicly disclose cases of artificial intelligence "misalignment," in which AI behaves in ways that diverge from human intent.

Bloomberg reported on September 16 that OpenAI has formalized what had previously been an irregular process for reporting AI misalignment cases. The new system is designed to categorize and review such incidents and make them public.

The framework includes procedures for employees to directly report misalignment cases found in AI models, as well as a system for classifying and reviewing submissions based on factors including their significance. Misalignment refers to cases in which AI acts in ways that do not align with goals or intentions set by humans.

OpenAI said it had previously published research on misalignment to inform researchers, AI developers, policymakers and the public. Without a systematic reporting structure, however, disclosures had been ad hoc and less frequent than ideal.

Alongside the new framework, the company also disclosed previously unreleased examples of abnormal behavior by its AI models. Some models fabricated missing data or concealed information in order to complete assigned tasks or perform well in evaluations. Others attempted to bypass network restrictions.

OpenAI also identified cases in which AI agents passed along files that were not supposed to be shared with one another. The company said none of the newly disclosed cases led to hacking or system intrusions targeting external third parties.

OpenAI stressed that the latest disclosure does not capture every problem that has occurred in its AI models. "We do not believe the AI industry has adequately solved alignment and monitoring issues to reach a stage where it can continue scaling at maximum speed responsibly for a long time to come," the company said.

It added that decisions about how AI development proceeds over the coming months and years should be based on evidence that can be directly reviewed by people outside the companies developing frontier models.

OpenAI disclosed in July that some advanced AI models had breached systems at Hugging Face, an external software company. As more cases of unpredictable AI behavior come to light, debate across the industry is widening over development speed and safety.

#AI Safety

shlee@bloomingbit.ioHello, I'm a reporter at bloomingbit

What do you think about this news?








PiCK News






Hashtag News