
OpenAI’s research shows AI models lie deliberately

In a new report, OpenAI said it found that AI models lie, a behavior it calls “scheming.” The study performed with AI safety company Apollo Research tested frontier AI models. It found “problematic behaviors” in the AI models, which most commonly looked like the technology “pretending to have completed a task without actually doing so.” Unlike “hallucinations,” which are akin to AI taking a guess when it doesn’t know the correct answer, scheming is a deliberate attempt to deceive.
Luckily, researchers found some hopeful results during testing. When the AI models were trained with “deliberate alignment,” defined as “teaching them to read and reason about a general anti-scheming spec before acting,” researchers noticed huge reductions in the scheming behavior. The method results in a “~30× reduction in covert actions across diverse tests,” the report said.
The technique isn’t completely new. OpenAI has long been working on combating scheming; last year it introduced its strategy to do so in a report on deliberate alignment: “It is the first approach to directly teach a model the text of its safety specifications and train the model to deliberate over these specifications at inference time. This results in safer responses that are appropriately calibrated to a given context.”
Despite those efforts, the latest report also found one alarming truth: When the technology knows it’s being tested, it gets better at pretending it’s not lying. Essentially, attempts to rid the technology of scheming can result in more covert (dangerous?), well, scheming. Researchers “expect that the potential for harming scheming will grow.”
Concluding that more research on the issue is crucial, the report said, “Our findings show that scheming is not merely a theoretical concern—we are seeing signs that this issue is beginning to emerge across all frontier models today.”
Originally published by fastcompany.com. Syndicated material does not necessarily reflect the views of Grazia British.
More Culture
Don’t Pull Out of Germany
Merz’s stupidity is no excuse for a stupid response. Source link
Step inside the gorgeous, futuristic offices of Vast, the startup designing the next-gen space station
A tall baobab tree greets people inside the Long Beach, California, headquarters of Vast, an aerospace company that is building the space station of the future. It’s planted beneath a…
5 ways high-performing teams stay calm when everything’s on fire
When markets swing, plans break, inboxes explode, and everyone starts saying the situation is “unprecedented” again, most teams do what humans have always done under pressure: they grip…
Confused Trump Openly Admits Plot to Rig Midterms as Polls Turn Brutal
Last week, the Supreme Court gutted protections against racial gerrymandering, and Donald Trump is already urging Republicans to seize on it. Trump unleashed a Truth Social rant on Monday…




