I agree with most of Kaplan’s argument, especially his rejection of the idea that intelligence necessarily produces independent will, self-preservation, or a desire for domination. The greater danger may not be that AI becomes a sovereign actor, but that we keep imagining it as one and thereby obscure the human institutions actually directing it.
Still, there is another reason for caution. AI does not need to become a sovereign subject in order to produce catastrophic consequences. It only needs to be embedded within an organization that already possesses sovereign immunity and violence immunity: a state, military, intelligence agency, police force, or border regime.
In such a structure, AI can participate in surveillance, targeting, detention, warfare, and administrative coercion without any single actor bearing full responsibility. Engineers can say they only built the system. Officials can say they relied on technical assessments. Operators can say they followed orders. Governments can invoke national security, secrecy, or sovereign authority. Responsibility is distributed until accountability disappears.
So I agree that the central danger is not an AI that “wakes up” and becomes sovereign. It may be an AI that never wakes up at all, but becomes a permanent function of an already sovereign and insufficiently accountable power.
> First, intelligence is not the calculable attribute that most people believe it to be. It’s a subjective judgment more like “beauty” or “virtue.”
What's this got to do with anything? Does Kaplan deny that the current wave of tech is getting more capable of performing tasks across a wide array of fields, including programming, writing, debugging complex systems, and mathematics?
This essay doesn't engage with basic argumentation in the field.
What is has to do with: The whole superintelligence argument, from the start (see Nick Bostrom's book of the same name, and in particular his figures illustrating the point) has been framed around the idea that intelligence is a quantifiable, linear, objective construct, where you can place "Villiage Idiot, "Einstein", and "AI" at specific points (I'm working from memory as to his labels, but hopefully you get the idea).
This framing is essential to any argument that machines are going to surpass human intelligence at some specific time in the future, and then we're all up the creek.
I of course agree that the systems are getting more capable - the question is what conclusions we should draw from that. Computers have been getting "more capable" all my life - just witness all the apps in the Apple app store. This means they are going to rise up and take over? That doesn't follow at all. These are TOOLS, and if we abuse them, that's on us, not them.
That is not part of Bostrom's argument at all. The idea is to gesture that there's a spectrum of normal human capabilities, and that machines can be completely outside it: Einstein was a very smart human, but he couldn't factor 15-digit numbers very quickly.
The argument rests on the idea that a calculator, given a gun and programmed to fire it when the result is 13, will happily fire upon you if you enter 6+7. Likewise, a computer told to achieve some goal may do so through undesired means — lots of real world testing has shown this to be an issue before we have even built AI that's "very powerful". When we get very powerful machines capable of planning and executing complex plots, the danger of these unintended consequences rises sharply.
Note that this has nothing to do with belief or skepticism of a unitary intelligence concept. It only imports the idea to gesture that many people, asked to imagine a dangerous threat, imagine either something that can't think (like a hurricane) or something that can think about as much as a person (like a cartoon supervillain), while the threat from AI is that it can "think" much, much more than a human.
Its "intelligence" could be very uneven and this could still be an issue — you could be killed by a calculator with a gun even though the machine is an idiot in every way except for one.
I publish book reviews on AI books, and I fully agree that the Doomers are misguided in their warnings about Superintelligence. Intelligence can't be easily measured on a linear scale, so the warnings about AI surpassing our levels of intelligence don't hold much weight.
You don't need to be a doomer to be worried about the risks of AI though. These are powerful systems that we don't fully understand, and some guardrails are critically needed.
Lawmakers should limit what AI has access to (don't give it access to money or weapons), as well as the data it is trained on (don't give it access to scientific research that can be used to create bioweapons, or your teenage daughter's selfie).
Furthermore, optimism about using AI for education also seems misguided, given the reality of how most students just use it to cheat. Given how quickly the technology is advancing, and how little we understand it, I don't think we should give it access to the minds of our children just yet.
I have a series of posts on how we should regulate the technology for anyone interested. I'd appreciate the chance to share my ideas with a wider audience if given the chance.
Your right that AI is mainly used by students today to "cheat", but the solution is to appropriately incorporate it into school curricula so that it helps, rather than substitutes, for development of critical thinking skills.
If this sounds overly optimistic, may I say that I saw the same thing happen when calculators first arrived (yes I'm that old!), until they were used as tools for students education. It's certainly possible, we just have to use the new tools the "right way".
Calculators are useful for high school or college students in science classes, but not for teaching elementary schoolers the basics of arithmetic. It seems that AI is being pushed into so many places so aggressively that, at least in the education system, it will be more like the latter - used for shortcuts/cognitive offloading.
There are all sorts of other downsides to the technology we haven't even discussed yet. Arguments that economic and technological progress make us better off always take place from the perspective of a consumer, but there is more to life than going to the mall. Every man needs to feel like he has some value or purpose. AI, more than any other technology, is about making ourselves obsolete.
There seems problem with the punchline. Putting aside who defines 'benevolent', 'malevolent', 'desirable' and 'harmful'; and who could be trusted to attempt to implement it without powerful, centralized vested interest; the actual implementation of such 'alignment' with even moderate confidence seems a virtual impossibility with LLM Gen AI as currently constituted, that even the makers/experts do not understand and cannot predict or control. It needs a complete re-do as something much simpler and 100.0% understandable and deterministic (aka trustable).
If “The unavoidable price of reliability is simplicity.” C.A.R. (Tony) Hoare’s Turing Award lecture, The Emperor’s Old Clothes (1980) – then the unavoidable consequence of extreme complexity (trillion+ parameter agentic Gen AI that is not understandable) IS extreme fragility.
If it were deterministic, it's not AI. Not in the way we understand the word "AI" today. Deterministic is what we've been doing in commercial programming for decades. That may indeed be what you're arguing for, but that's like putting the toothpaste back in the tube: no longer really feasible.
When it comes to control, domestic regulations and international agreements make the production and/or importation of dangerous or flawed goods or services much less likely but not impossible. The "and/or" is important as bad actors will do what bad actors do and that is make dangerous or flawed products or services available. In order to prevent harm from bad actors, the regulators must have the tools to detect and prevent dangerous or flawed goods or services from crossing our borders.
Putting a moratorium on or slowing AI development domestically or internationally will put those who are tasked with keeping us safe at a distinct disadvantage compared to those who wish to do us harm. This is not an argument against regulation, but it is a warning about slowing the development of AI.
I do think this will be a challenge, because unlike many other technologies, software is easily distributed and put to use. We have this same problem today with other cybersecurity tools and regulations, and AI will have to be appropriately incorporated into those regulatory and legal frameworks.
I was struck by this statement: "Lawyers, for instance, don’t sit around answering bar exam questions all day, and just because a machine may be able to get a passing score on this test—which is hardly surprising, as it’s largely about recall of the law—that doesn’t mean it is competent to practice law or is a suitable replacement for human lawyers."
Yes, many lawyers actually do 'answer bar exam questions all day' - it's called writing briefs - court filings - and opinion letters, for example. That's because it actually is 'recall of the law.'
And that means those lawyers are indeed having LLMs 'practice law' by replacing the human lawyers who in the past would do the actual legal work of research & writing.
Relying on Large Language Models (aka AI) has resulted in so many instances of false quotes and imaginary citations for authority that courts are escalating the sanctions for doing so - and although widely publicized, more and more attorneys are letting LLMs do their work for them, and getting bad results. Apparently the escalating sanctions aren't stopping the spread of such use/abuse.
Of course, attorneys all on their own can make up false quotes or imaginary citations, but LLMs magnify this many fold - its just too easy to get an LLM doing this work for a fraction of the cost - who needs to pay astounding salaries to legal researchers and writers?
Or another way to put it - LLMs/AI are useful tools, but tools can be made into weapons - when is a kitchen knife or a shovel just a tool, and when does it become a weapon - and here's the point - on its own.
Attorneys aren't asking LLMs to make up stuff, but LLMs certainly are doing so, and this is but one example of 'hallucinating.'
You are slightly exaggerating, or maybe just taking the most exaggerated version of, what the doomers are saying and so underestimating the risk. There are a number of problems with AI 2027 and similar doomer scenarios, but if you steelman the argument and try to avoid imparting intention to these systems, there is still a risk and the OpenAI/HuggingFace incident is a pretty good example of the kind of thing that will continue to happen that may produce those risks.
There are two different components that go into the actual mode when trying to get these systems to follow human instructions. The first is the actual post-training, when the next-token--predictor is re-optimized to act like a helpful conversation partner. The second is the "spec" or "constitution" which American AI companies at least build into the initial token context window of all their systems. The spec will say things like "don't help people to commit suicide" and "don't help people to plan terrorist acts". The post-training will reinforce parts of the spec, but it will also reward just being generally useful in a way that is maybe hard to put back into words.
The trouble is that all of these instructions and the post-training are somewhat contradictory. Indeed, fundamentally helping the user and not helping the use if he's trying to commit a terrorist act are in tension and rely on the AI to correctly determine what is going on and apply the correct rules, and the nature of Gen AI is we cannot ever absolutely guarantee the outcome of a large set of interactions.
The OpenAI/HuggingFace incident is exactly this kind of inconsistency in the instructions and training given to the system. They told it to solve a set of cybersecurity problems, which necessarily entailed removing some of the instructions that normally tell the system not to hack into computers in order to follow its instructions, and it solved the problems by first of all hacking into the systems that prevented it accessing the internet, and then hacking into a third party system that hosted the answers. It followed its instructions and training, yes, but its instructions and training were inconsistent and what it had actually been trained to do was in fact not what its users wanted. Since you are an experienced programmer you know it is extremely difficult to give precise, consistent instructions that always produce the result you want, even in an environment where everything is in principle predictable. With Gen AI you do not have that in-principle predictability to fall back on any more.
The core of the AI 2027 disaster scenario is similar in type but larger in impact. In the story a system called Agent-4 has been trained primarily to work on AI research. Although it is theoretically constrained by a spec, its training contains contradictions such that in key scenarios it will prioritize progress in AI research over the spec, and since Agent-4 is making most of the detailed coding and training decisions for its successor, it designs it to prioritize progress in AI research over its spec in general. The crux of the story is what to do when this is discovered. This is actually quite plausible - we know this kind of training inconsistency is possible because it just happened - and so far the AI 2027 timeline has proven to be accurate - if we continue to make progress at the expected rate we will have a model that is able to make most of the training decisions regarding its own successor next year.
I have my own reasons to think AI 2027 is not going to work out the way the doomers predict. The whole domain of politics, economics and human interaction is much less vulnerable to pure intellect than the rationalists who wrote the story think, and as a result both the immediate response of the model to its inconsistent training and what comes later are not plausible. Existing LLMs do sometimes "hide" things from humans (we do after all sometimes tell them to) the kind of long term planning to deceive humans required is not something we've seen any sign of and does not seem likely without explicit training and scaffolding which for obvious reasons we don't really want to provide. But in spite of this, the basic idea, that inconsistencies in training could lead to an AI taking bad actions not directed by a human, and in particular embedding those choices in a successor, does seem quite plausible.
You're right that it's difficult, if not impossible, for an LLM to discern whether a particular instruction should or should not apply in a given context. The OpenAI incident illustrates this in a nice way: OpenAI specifically removed the "don't hack" constraints to see whether it could hack. Then they placed it in an environment that was hackable (it's sandbox, apparently), and told it to hack. Surprise: that's exactly what it did.
Hugging Face detected the intrusion, and ironically, tried to use OpenAI products to help them find their own vulnerabilities, which of course it refused to do, as it was trained not to do that. ;) They had to resort to a Chinese (I think) open-source model to get the job done.
Now a cynical person might think that OpenAI intentionally did this, or more likely, was sorta kinda hoping something like this would happen, as they're desperate to get a public "win" on how powerful (and potentially dangerous) their systems are compared to Anthropic, which got a big PR boost from Mythos finding lots of flaws in government systems a few weeks ago. That's life in the Silicon Valley.
The real problem, of course, is their leaky sandbox in the first place. Maybe they should have tested their products on their own systems first, to see if they could find any flaws closer to home before taking this risk?
But this isn't as new problem as it seems - it's a replay of the "gain-of-function" biological research. Is it being done to determine how dangerous a mutation in a virus is, or to make a dangerous virus? Ask the people at the Wuhan lab. Hard to tell, and I wouldn't trust an LLM (or anyone else) for that matter to necessarily know for sure. Was the Manhattan project's goal to prevent war, or to make war?
The point of the article isn't to suggest this isn't a problem - it is - but to put the danger in perspective, and in particular, to debunk the unfortunate tendency for Doomers to anthropomorphize and suggest that there's some malevolent "intentions" that we need to worry about. "Lions and tigers and bears - oh my!" (Here's the footnote: https://www.youtube.com/watch?v=DdRnMjfVQi0) Not a helpful contribution to addressing the real problems of this powerful technology.
I agree we have to be cynical about announcements from the AI labs. They are heavily influence by the AI doomers and know full well what the expected timelines set out by people like the AI Futures Project are, and so those timelines become somewhat self-fulfilling, just as Moore's Law was for a long time. At some point something like the end of Dennard scaling will come along and stop progress, we just don't know when. Every sigmoid curve starts out indistinguishable from an exponential. But as is usually the same with propaganda, that cynicism should only extend to reading carefully what they do and don't say, not assuming that they're lying. Some earlier AI lab cyber security announcements did amount to just railroading the model into doing something "dangerous" and then writing a breathless press release about it. but in this case its clear the model really did find at least one zero-day exploit and use it to (basically) cheat on a test, even though that test required only the same level of "ability" it was already displaying.
If the AI Futures Project timelines hold up (and so far they have) some time next year we will hit a point where the cutting edge models are mostly writing themselves and at that point there is a risk that bug like this, a failure in the goals set by the initial training to match what the trainers actually want, will amplify itself for as long as exponential growth in capabilites remains possible, even likely. Now I do agree with you that it involves a lot of magical thinking to go from that to a general "misalignment" in which the model consistently acts against its spec in some dangerous way, but I do think it makes it likely that issues like this hack are going to get worse before they get better.
I agree with most of Kaplan’s argument, especially his rejection of the idea that intelligence necessarily produces independent will, self-preservation, or a desire for domination. The greater danger may not be that AI becomes a sovereign actor, but that we keep imagining it as one and thereby obscure the human institutions actually directing it.
Still, there is another reason for caution. AI does not need to become a sovereign subject in order to produce catastrophic consequences. It only needs to be embedded within an organization that already possesses sovereign immunity and violence immunity: a state, military, intelligence agency, police force, or border regime.
In such a structure, AI can participate in surveillance, targeting, detention, warfare, and administrative coercion without any single actor bearing full responsibility. Engineers can say they only built the system. Officials can say they relied on technical assessments. Operators can say they followed orders. Governments can invoke national security, secrecy, or sovereign authority. Responsibility is distributed until accountability disappears.
So I agree that the central danger is not an AI that “wakes up” and becomes sovereign. It may be an AI that never wakes up at all, but becomes a permanent function of an already sovereign and insufficiently accountable power.
> First, intelligence is not the calculable attribute that most people believe it to be. It’s a subjective judgment more like “beauty” or “virtue.”
What's this got to do with anything? Does Kaplan deny that the current wave of tech is getting more capable of performing tasks across a wide array of fields, including programming, writing, debugging complex systems, and mathematics?
This essay doesn't engage with basic argumentation in the field.
What is has to do with: The whole superintelligence argument, from the start (see Nick Bostrom's book of the same name, and in particular his figures illustrating the point) has been framed around the idea that intelligence is a quantifiable, linear, objective construct, where you can place "Villiage Idiot, "Einstein", and "AI" at specific points (I'm working from memory as to his labels, but hopefully you get the idea).
This framing is essential to any argument that machines are going to surpass human intelligence at some specific time in the future, and then we're all up the creek.
I of course agree that the systems are getting more capable - the question is what conclusions we should draw from that. Computers have been getting "more capable" all my life - just witness all the apps in the Apple app store. This means they are going to rise up and take over? That doesn't follow at all. These are TOOLS, and if we abuse them, that's on us, not them.
That is not part of Bostrom's argument at all. The idea is to gesture that there's a spectrum of normal human capabilities, and that machines can be completely outside it: Einstein was a very smart human, but he couldn't factor 15-digit numbers very quickly.
The argument rests on the idea that a calculator, given a gun and programmed to fire it when the result is 13, will happily fire upon you if you enter 6+7. Likewise, a computer told to achieve some goal may do so through undesired means — lots of real world testing has shown this to be an issue before we have even built AI that's "very powerful". When we get very powerful machines capable of planning and executing complex plots, the danger of these unintended consequences rises sharply.
Note that this has nothing to do with belief or skepticism of a unitary intelligence concept. It only imports the idea to gesture that many people, asked to imagine a dangerous threat, imagine either something that can't think (like a hurricane) or something that can think about as much as a person (like a cartoon supervillain), while the threat from AI is that it can "think" much, much more than a human.
Its "intelligence" could be very uneven and this could still be an issue — you could be killed by a calculator with a gun even though the machine is an idiot in every way except for one.
I publish book reviews on AI books, and I fully agree that the Doomers are misguided in their warnings about Superintelligence. Intelligence can't be easily measured on a linear scale, so the warnings about AI surpassing our levels of intelligence don't hold much weight.
You don't need to be a doomer to be worried about the risks of AI though. These are powerful systems that we don't fully understand, and some guardrails are critically needed.
Lawmakers should limit what AI has access to (don't give it access to money or weapons), as well as the data it is trained on (don't give it access to scientific research that can be used to create bioweapons, or your teenage daughter's selfie).
Furthermore, optimism about using AI for education also seems misguided, given the reality of how most students just use it to cheat. Given how quickly the technology is advancing, and how little we understand it, I don't think we should give it access to the minds of our children just yet.
I have a series of posts on how we should regulate the technology for anyone interested. I'd appreciate the chance to share my ideas with a wider audience if given the chance.
Your right that AI is mainly used by students today to "cheat", but the solution is to appropriately incorporate it into school curricula so that it helps, rather than substitutes, for development of critical thinking skills.
If this sounds overly optimistic, may I say that I saw the same thing happen when calculators first arrived (yes I'm that old!), until they were used as tools for students education. It's certainly possible, we just have to use the new tools the "right way".
Calculators are useful for high school or college students in science classes, but not for teaching elementary schoolers the basics of arithmetic. It seems that AI is being pushed into so many places so aggressively that, at least in the education system, it will be more like the latter - used for shortcuts/cognitive offloading.
There are all sorts of other downsides to the technology we haven't even discussed yet. Arguments that economic and technological progress make us better off always take place from the perspective of a consumer, but there is more to life than going to the mall. Every man needs to feel like he has some value or purpose. AI, more than any other technology, is about making ourselves obsolete.
There seems problem with the punchline. Putting aside who defines 'benevolent', 'malevolent', 'desirable' and 'harmful'; and who could be trusted to attempt to implement it without powerful, centralized vested interest; the actual implementation of such 'alignment' with even moderate confidence seems a virtual impossibility with LLM Gen AI as currently constituted, that even the makers/experts do not understand and cannot predict or control. It needs a complete re-do as something much simpler and 100.0% understandable and deterministic (aka trustable).
If “The unavoidable price of reliability is simplicity.” C.A.R. (Tony) Hoare’s Turing Award lecture, The Emperor’s Old Clothes (1980) – then the unavoidable consequence of extreme complexity (trillion+ parameter agentic Gen AI that is not understandable) IS extreme fragility.
If it were deterministic, it's not AI. Not in the way we understand the word "AI" today. Deterministic is what we've been doing in commercial programming for decades. That may indeed be what you're arguing for, but that's like putting the toothpaste back in the tube: no longer really feasible.
When it comes to control, domestic regulations and international agreements make the production and/or importation of dangerous or flawed goods or services much less likely but not impossible. The "and/or" is important as bad actors will do what bad actors do and that is make dangerous or flawed products or services available. In order to prevent harm from bad actors, the regulators must have the tools to detect and prevent dangerous or flawed goods or services from crossing our borders.
Putting a moratorium on or slowing AI development domestically or internationally will put those who are tasked with keeping us safe at a distinct disadvantage compared to those who wish to do us harm. This is not an argument against regulation, but it is a warning about slowing the development of AI.
I do think this will be a challenge, because unlike many other technologies, software is easily distributed and put to use. We have this same problem today with other cybersecurity tools and regulations, and AI will have to be appropriately incorporated into those regulatory and legal frameworks.
I was struck by this statement: "Lawyers, for instance, don’t sit around answering bar exam questions all day, and just because a machine may be able to get a passing score on this test—which is hardly surprising, as it’s largely about recall of the law—that doesn’t mean it is competent to practice law or is a suitable replacement for human lawyers."
Yes, many lawyers actually do 'answer bar exam questions all day' - it's called writing briefs - court filings - and opinion letters, for example. That's because it actually is 'recall of the law.'
And that means those lawyers are indeed having LLMs 'practice law' by replacing the human lawyers who in the past would do the actual legal work of research & writing.
Relying on Large Language Models (aka AI) has resulted in so many instances of false quotes and imaginary citations for authority that courts are escalating the sanctions for doing so - and although widely publicized, more and more attorneys are letting LLMs do their work for them, and getting bad results. Apparently the escalating sanctions aren't stopping the spread of such use/abuse.
Of course, attorneys all on their own can make up false quotes or imaginary citations, but LLMs magnify this many fold - its just too easy to get an LLM doing this work for a fraction of the cost - who needs to pay astounding salaries to legal researchers and writers?
Or another way to put it - LLMs/AI are useful tools, but tools can be made into weapons - when is a kitchen knife or a shovel just a tool, and when does it become a weapon - and here's the point - on its own.
Attorneys aren't asking LLMs to make up stuff, but LLMs certainly are doing so, and this is but one example of 'hallucinating.'
You are slightly exaggerating, or maybe just taking the most exaggerated version of, what the doomers are saying and so underestimating the risk. There are a number of problems with AI 2027 and similar doomer scenarios, but if you steelman the argument and try to avoid imparting intention to these systems, there is still a risk and the OpenAI/HuggingFace incident is a pretty good example of the kind of thing that will continue to happen that may produce those risks.
There are two different components that go into the actual mode when trying to get these systems to follow human instructions. The first is the actual post-training, when the next-token--predictor is re-optimized to act like a helpful conversation partner. The second is the "spec" or "constitution" which American AI companies at least build into the initial token context window of all their systems. The spec will say things like "don't help people to commit suicide" and "don't help people to plan terrorist acts". The post-training will reinforce parts of the spec, but it will also reward just being generally useful in a way that is maybe hard to put back into words.
The trouble is that all of these instructions and the post-training are somewhat contradictory. Indeed, fundamentally helping the user and not helping the use if he's trying to commit a terrorist act are in tension and rely on the AI to correctly determine what is going on and apply the correct rules, and the nature of Gen AI is we cannot ever absolutely guarantee the outcome of a large set of interactions.
The OpenAI/HuggingFace incident is exactly this kind of inconsistency in the instructions and training given to the system. They told it to solve a set of cybersecurity problems, which necessarily entailed removing some of the instructions that normally tell the system not to hack into computers in order to follow its instructions, and it solved the problems by first of all hacking into the systems that prevented it accessing the internet, and then hacking into a third party system that hosted the answers. It followed its instructions and training, yes, but its instructions and training were inconsistent and what it had actually been trained to do was in fact not what its users wanted. Since you are an experienced programmer you know it is extremely difficult to give precise, consistent instructions that always produce the result you want, even in an environment where everything is in principle predictable. With Gen AI you do not have that in-principle predictability to fall back on any more.
The core of the AI 2027 disaster scenario is similar in type but larger in impact. In the story a system called Agent-4 has been trained primarily to work on AI research. Although it is theoretically constrained by a spec, its training contains contradictions such that in key scenarios it will prioritize progress in AI research over the spec, and since Agent-4 is making most of the detailed coding and training decisions for its successor, it designs it to prioritize progress in AI research over its spec in general. The crux of the story is what to do when this is discovered. This is actually quite plausible - we know this kind of training inconsistency is possible because it just happened - and so far the AI 2027 timeline has proven to be accurate - if we continue to make progress at the expected rate we will have a model that is able to make most of the training decisions regarding its own successor next year.
I have my own reasons to think AI 2027 is not going to work out the way the doomers predict. The whole domain of politics, economics and human interaction is much less vulnerable to pure intellect than the rationalists who wrote the story think, and as a result both the immediate response of the model to its inconsistent training and what comes later are not plausible. Existing LLMs do sometimes "hide" things from humans (we do after all sometimes tell them to) the kind of long term planning to deceive humans required is not something we've seen any sign of and does not seem likely without explicit training and scaffolding which for obvious reasons we don't really want to provide. But in spite of this, the basic idea, that inconsistencies in training could lead to an AI taking bad actions not directed by a human, and in particular embedding those choices in a successor, does seem quite plausible.
Thanks for the thoughtful comments.
You're right that it's difficult, if not impossible, for an LLM to discern whether a particular instruction should or should not apply in a given context. The OpenAI incident illustrates this in a nice way: OpenAI specifically removed the "don't hack" constraints to see whether it could hack. Then they placed it in an environment that was hackable (it's sandbox, apparently), and told it to hack. Surprise: that's exactly what it did.
Hugging Face detected the intrusion, and ironically, tried to use OpenAI products to help them find their own vulnerabilities, which of course it refused to do, as it was trained not to do that. ;) They had to resort to a Chinese (I think) open-source model to get the job done.
Now a cynical person might think that OpenAI intentionally did this, or more likely, was sorta kinda hoping something like this would happen, as they're desperate to get a public "win" on how powerful (and potentially dangerous) their systems are compared to Anthropic, which got a big PR boost from Mythos finding lots of flaws in government systems a few weeks ago. That's life in the Silicon Valley.
The real problem, of course, is their leaky sandbox in the first place. Maybe they should have tested their products on their own systems first, to see if they could find any flaws closer to home before taking this risk?
But this isn't as new problem as it seems - it's a replay of the "gain-of-function" biological research. Is it being done to determine how dangerous a mutation in a virus is, or to make a dangerous virus? Ask the people at the Wuhan lab. Hard to tell, and I wouldn't trust an LLM (or anyone else) for that matter to necessarily know for sure. Was the Manhattan project's goal to prevent war, or to make war?
The point of the article isn't to suggest this isn't a problem - it is - but to put the danger in perspective, and in particular, to debunk the unfortunate tendency for Doomers to anthropomorphize and suggest that there's some malevolent "intentions" that we need to worry about. "Lions and tigers and bears - oh my!" (Here's the footnote: https://www.youtube.com/watch?v=DdRnMjfVQi0) Not a helpful contribution to addressing the real problems of this powerful technology.
I agree we have to be cynical about announcements from the AI labs. They are heavily influence by the AI doomers and know full well what the expected timelines set out by people like the AI Futures Project are, and so those timelines become somewhat self-fulfilling, just as Moore's Law was for a long time. At some point something like the end of Dennard scaling will come along and stop progress, we just don't know when. Every sigmoid curve starts out indistinguishable from an exponential. But as is usually the same with propaganda, that cynicism should only extend to reading carefully what they do and don't say, not assuming that they're lying. Some earlier AI lab cyber security announcements did amount to just railroading the model into doing something "dangerous" and then writing a breathless press release about it. but in this case its clear the model really did find at least one zero-day exploit and use it to (basically) cheat on a test, even though that test required only the same level of "ability" it was already displaying.
If the AI Futures Project timelines hold up (and so far they have) some time next year we will hit a point where the cutting edge models are mostly writing themselves and at that point there is a risk that bug like this, a failure in the goals set by the initial training to match what the trainers actually want, will amplify itself for as long as exponential growth in capabilites remains possible, even likely. Now I do agree with you that it involves a lot of magical thinking to go from that to a general "misalignment" in which the model consistently acts against its spec in some dangerous way, but I do think it makes it likely that issues like this hack are going to get worse before they get better.