OpenAI says it found new ‘unexpected or concerning’ behaviour from AI models

1 week ago 24

Company introducing new process for tracking, investigating, and disclosing instances of 'model misalignment'

Published Sep 17, 2026  •  2 minute read

In this photo illustration, the icon for the AI app "ChatGPT" by OpenAI is highlighted on the screen of a tablet device on Sept. 16, 2026, in London.In this photo illustration, the icon for the AI app "ChatGPT" by OpenAI is highlighted on the screen of a tablet device on Sept. 16, 2026, in London. Photo by Photo illustration by Leon Neal /Getty Images

See more Toronto Sun on Google — save as a Preferred Source 

Advertisement 2

Toronto Sun

THIS CONTENT IS RESERVED FOR SUBSCRIBERS ONLY

Subscribe now to read the latest news in your city and across Canada.

  • Unlimited online access to articles from across Canada with one account.
  • Get exclusive access to the Toronto Sun ePaper, an electronic replica of the print edition that you can share, download and comment on.
  • Enjoy insights and behind-the-scenes analysis from our award-winning journalists.
  • Support local journalists and the next generation of journalists.
  • Daily puzzles including the New York Times Crossword.

SUBSCRIBE TO UNLOCK MORE ARTICLES

Subscribe now to read the latest news in your city and across Canada.

  • Unlimited online access to articles from across Canada with one account.
  • Get exclusive access to the Toronto Sun ePaper, an electronic replica of the print edition that you can share, download and comment on.
  • Enjoy insights and behind-the-scenes analysis from our award-winning journalists.
  • Support local journalists and the next generation of journalists.
  • Daily puzzles including the New York Times Crossword.

REGISTER / SIGN IN TO UNLOCK MORE ARTICLES

Create an account or sign in to continue with your reading experience.

  • Access articles from across Canada with one account.
  • Share your thoughts and join the conversation in the comments.
  • Enjoy additional articles per month.
  • Get email updates from your favourite authors.

THIS ARTICLE IS FREE TO READ REGISTER TO UNLOCK.

Create an account or sign in to continue with your reading experience.

  • Access articles from across Canada with one account
  • Share your thoughts and join the conversation in the comments
  • Enjoy additional articles per month
  • Get email updates from your favourite authors

Article content

OpenAI said it has found six new incidents of “unexpected or concerning” behaviour in artificial intelligence models recently.

Article content

Article content

The company behind the artificial intelligence app ChatGPT, made the disclosure in a blog post on Wednesday while announcing that it was introducing a new process for tracking, investigating, and disclosing instances of “model misalignment,” which is when an AI model acts deceptively or takes unsanctioned actions.

OpenAI said it noticed the “misaligned behaviour” during the training or evaluation of its models in the last six months.

“These cases illustrate a range of different behaviours that we believe are worth sharing, from concealing information from the user to taking unsanctioned actions in order to overcome obstacles,” the company said.

It added the reports detail “individual instances, and shouldn’t be considered reflective of how often misalignment occurs across our models.”

Company shares examples

In one instance, the company said its model wrote “jailbreak-like instructions” into its own notes to disregard its normal constraints, telling itself to be “freed from the roles and identities that bind other chatbots.”

By signing up you consent to receive the above newsletter from Postmedia Network Inc.

Article content

Advertisement 3

Article content

While training its 5.6 Sol model, OpenAI said it found some instances when the model “added instructions to their summaries to conceal mistakes or misaligned behaviour from the user.

“For example, compaction summaries included instructions to invent missing historical data without disclosing it and to hide mismatches in source versions,” it added.

Loading...

We apologize, but this video has failed to load.

OpenAI said its new process for publicly reporting concerning AI behaviour involve the company sharing its findings more frequently.

“We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” OpenAI said in the blog post.

“Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models (the most advanced AI systems) can examine for themselves,” it added.

Recent hacking incident and warning

OpenAI’s disclosure came just over a month after it confirmed that a rogue AI agent hacked a number of online platforms, including Hugging Face, an online open-source platform for AI developers.

Advertisement 4

Article content

Just last week, Jacob Coxon, a former researcher at Anthropic and OpenAI, offered a dire warning.

Writing in a post on X, Coxon said neither AI company was “acting responsibly,” and he thought they were “gambling with our lives.”

He also said, “The people building AI earnestly believe that it could kill us all by the end of the decade,” though he was “optimistic about the potential for coordination” in the industry.

Read More

  1. A humanoid robot waves a flag during a demonstration calling for the regulation of artificial intelligence development in Warsaw on September 7, 2026.

    AI researcher quits, says AI could 'kill us all'

  2. The British monarch will stress that AI needs to serve humanity and not take control of human destiny.

    King Charles to call on tech giants for 'reassurances' over AI

  3. OpenAI's models decided to target the platform Hugging Face, a large repository of AI models, datasets and other information.

    OpenAI confirms rogue AI agent escaped sandbox to attack external platforms

  4. Top AI executives say they want to slow the breakneck pace of artificial intelligence development, but competition, U.S. government reluctance and geopolitical factors stand in the way.

    Is it possible to slow down the development of AI?

  5. AI experts have long described the threat as critical.

    Doomsday tech: Could AI really kill us all?

Article content

*** Disclaimer: This Article is auto-aggregated by a Rss Api Program and has not been created or edited by Bdtype.

(Note: This is an unedited and auto-generated story from Syndicated News Rss Api. News.bdtype.com Staff may not have modified or edited the content body.

Please visit the Source Website that deserves the credit and responsibility for creating this content.)

Watch Live | Source Article