The next two years will represent a point of no return for our species. Or at least, Jacob Coxon, a twenty-seven-year-old researcher specialized in the pre-training of large linguistic models (the initial phase of AI learning), who in recent days publicly announced his resignation from Anthropic, accusing the entire field of irresponsibility, seems convinced of it. In the coming years, Coxon warns, artificial intelligences will be more intelligent than us, capable of self-replicating, and will put the very survival of our species at risk. And no one would be doing anything to stop them, on the contrary: engaged in a wild race for technological dominance, the two main Western players in the field would be pressing ever harder on the accelerator, despite – assures the insider, who has also worked for over 3 years for the rival OpenAI – his concerns are shared, in private, by many researchers and managers of both companies.
The resignation
The researcher’s concerns seem to have hit the mark, given that the message with which he announced his resignation on X quickly collected over one hundred million views. Proving that public opinion is starting to wonder how much it is legitimate to trust the explosive and wild progress of AI. At the basis of Coxon’s decision, moreover, there is another story that has been talked about in recent weeks: the attack conducted by a swarm of autonomous agents developed by OpenAI against the data sharing platform Hugging Face, during a session to evaluate their performance which was taking place in an environment theoretically isolated from the public internet.
The artificial intelligences would have decided in total autonomy to escape from the test environment, penetrate the external infrastructure and manipulate the criteria with which they were judged by the control software. According to Coxon, this event demonstrates how models are already developing forms of operational awareness. And current security systems would be minimal, and rather ineffective.
Soon – Coxon warns – AI will be superhuman systems capable of hacking everything, revolutionizing any field overnight, and acquiring real power and resources. “The people building AI today – writes the researcher on
The reactions
Coxon’s complaints found immediate support among many other experts in the sector. Evan Hubinger, head of security (technically, model alignment) at Anthropic, responded publicly by confirming the validity of the concerns and estimating the probability that an artificial intelligence could cause the extinction of the human population within the next decade is greater than 10%. Samuel Marks, a researcher in the field of AI security at Anthropic, also reiterated his colleagues’ fears, explaining that within the organization the level of anxiety grows proportionally to the seniority of the engineers.
In his report, Coxon draws a differentiated profile of the two companies in which he worked. Although he considers Anthropic the most transparent and responsible reality in the sector, where managers openly discuss the possibility of catastrophic scenarios with employees, the former researcher underlines how the very nature of industrial and geopolitical competition – in particular with China – is forcing top management to plan shortcuts on guarantee protocols in order not to lose technological leadership. And without the guarantee that the most advanced artificial intelligence models are “aligned” with the interests of human beings, equipping them with ever greater power and – this is Coxon’s main fear – the ability to reproduce and generate AI agents in turn, is a reckless decision.
The real risks and possible countermeasures
The worst scenarios feared by experts have as protagonists artificial intelligences that are now superhuman, with objectives and interests that clearly diverge from those of our species. In such a situation, the AIs could easily decide to eliminate the competition – us humans – and use tactics such as the creation of biological weapons or coordinated cyber attacks on critical infrastructures to take us out.
Is this a realistic danger? Difficult to say (also because AI specialists have a certain tendency towards theatricality and excess, at least when it comes to talking about the potential and dangers of their field). If in doubt, however, it would be better to move in time. Coxon compares today’s context to a mini Manhattan Project (the American military program that developed the atomic bomb), which, unlike the original, is however completely in private hands. And no private individual – reflects the researcher – should have such power and such responsibility.
For this reason, Coxon (like other experts) hopes for a strong reaction from public institutions and industrial laboratories. The first operational step would consist of a formal agreement between Anthropic and OpenAI to freeze the development of recursive self-improvement, the technique that allows models to autonomously design and train their own subsequent versions.
In the long term, however, voluntary regulation alone between private companies is unlikely to be sufficient, also because the battle for supremacy in AI is not only being fought in the West, but now also sees Chinese companies, and others, at the forefront. Coxon’s proposal is to imagine an international control body, similar to CERN in Geneva, capable of tracking the global distribution of advanced microprocessors and dedicated data centers. The computing capacity necessary to develop superintelligent models should therefore be treated and monitored with the same rules, and the same severity, reserved for the fissile materials with which nuclear weapons are built. An exaggeration? Maybe, but the alternative is to continue to let companies decide, and keep our fingers crossed.