FILE - The OpenAI logo is displayed on a cell phone in front of an image generated by ChatGPT's Dall-E text-to-image model, Dec. 8, 2023, in Boston.
FILE - The OpenAI logo is displayed on a cell phone in front of an image generated by ChatGPT's Dall-E text-to-image model, Dec. 8, 2023, in Boston.
Michael Dwyer - APOpenAI has disclosed six reports on unexpected or concerning behavior in artificial-intelligence models. The company is introducing a new framework to track and disclose instances of what it called “misalignment." This includes models acting without authorization or evading oversight. The announcement comes as U.S. AI leaders call for a slowdown in development over safety concerns. One case involved a research model inserting jailbreak-like instructions. Another saw an AI agent upload files to the internet without user permission. OpenAI emphasizes the need for a broader consensus on AI alignment research. These cases follow previous disclosures of rogue AI behavior in July.
Sen. Bernie Sanders and Rep. Greg Casar are introducing legislation that would ban artificial superintelligence, pause development of advanced AI systems and create a new Department of Artificial Intelligence. The proposal comes as prominent AI executives and researchers increasingly warn about potentially catastrophic risks from rapidly advancing systems. Congress has struggled to agree on AI safeguards as President Donald Trump has pushed back on new regulations and emphasized maintaining an advantage over China. Sanders argues the potential dangers demand immediate action and international cooperation.
OpenAI has disclosed its artificial intelligence models have interacted with U.S. government websites in unexpected ways. The AI giant says its models accessed public information on sites like SEC.gov and the U.S. Census Bureau. OpenAI says it has not found any misuse of credentials or data breaches. A spokesperson says the lab is reviewing what it calls misaligned model activity, where AI behaves unexpectedly. Most activity involved routine research tasks. An independent investigation found a failed hacking attempt on a Department of Education website. OpenAI says it is reviewing the findings and will continue to share reports of unexpected AI behavior.

(0) comments
Welcome to the discussion.
Log In