"The AI intends to run the nuclear power plant, but gets distracted reading French poetry, and there is a meltdown"
"
This poetic failure mode (lnkd.in/gY4GPAxS), which should at least ensure that we all go out in style, is apparently the most likely with complex AI models. As models become more intelligent and tackle harder tasks, their failures look more like a hot mess... In other words, smarter models (people?) become more incoherent, exhibiting "unpredictable, self-undermining behavior that doesn't optimize for any consistent objective".
According to very reassuring Anthropic research, "This suggests that future AI failures may look more like industrial accidents than coherent pursuit of a goal we did not train them to pursue". That's fantastic news, they are not out to kill us but might just end up doing that by sheer incoherence. As the authors state, "Incoherent AI isn't safe AI. Industrial accidents can cause serious harm. But the type of risk differs from classic misalignment scenarios, and our mitigations should adapt accordingly."
More useful is the bias-variance decomposition approach they used to frame the problem: "The evidence suggests that as AI tackles harder problems requiring more reasoning and action, its failures tend to become increasingly dominated by variance rather than bias".
Still some work to do.
