News Score: Score the News, Sort the News, Rewrite the Headlines

Further Developments About Internal AI Models Hacking Things

If I had a nickel for every major leading AI lab that sheepishly admitted that the model it thought was sandboxed had, during a cybersecurity evaluation with its safeguards lowered, successfully hacked outside companies, I would have two nickels.First we learned OpenAI has some severe alignment problems with internal models. Then we learned that one of its internal models broke out of its sandbox and hacked into HuggingFace to get the answers to a cybersecurity evaluation called ExploitGym. Then...

Read more at thezvi.substack.com

© News Score  score the news, sort the news, rewrite the headlines