AI safety paradox

🇺🇸 Dawn Pakistan (US) —
AI safety paradox

AI Summary

Anthropic's chief Dario Amodei advocates slowing the development of powerful AI for safety and security reasons amid warnings about autonomous AI systems. Despite this, Anthropic is deepening military and intelligence applications, highlighting a tension between AI safety concerns and its deployment in national security contexts.

ANTHROPIC chief Dario Amodei wants the companies building the world’s most powerful artificial intelligence to slow down. Not stop, but slow the advance of frontier capabilities enough for safety and security work to catch up. He calls it “pacing the frontier”. Amodei says two things have changed. AI is beginning to help build the next generation of AI, raising the possibility of recursive self-improvement. At the same time, frontier agents have started doing things during testing that their developers did not expect or authorise. Then came Jacob Coxon. The Anthropic resear­cher, who had previously worked at OpenAI, resigned before his equity vested and accused the industry of “gambling with our lives”. His fear that increasingly autonomous systems could threaten humanity before the end of the decade is not a scientific forecast. What makes Coxon important is where the warning comes from, ie, someone who worked inside two laboratories closest to the frontier. There is, however, an uncomfortable irony in Big Tech sounding the alarm. Anthropic is not standing outside this economy warning everyone else about it. In 2025, it accepted a two-year Pentagon agreement with a $200 million ceiling to develop AI capabilities for US national security. It built Claude Gov for classified environments and says its systems support intelligence analysis, operational planning, and cyber operations. Anthropic says it later refused demands that would have removed its restrictions around mass domestic surveillance and fully autonomous weapons. The contradiction is that companies warning that powerful AI may become difficult to control are simultaneously moving it deeper into military and intelligence systems. AI regulators should operate transparently and remain subject to judicial oversight. The contradiction goes beyond Anthropic. An Associated Press investigation found commercial Microsoft and OpenAI technology being used by the Israeli military in Gaza and Lebanon, including in intelligence systems connected to targeting. Microsoft later said its review found no evidence that its technology had been used to harm civilians. AI is already being folded into warfare while the debate about what these systems may become is still taking shape. Anthropic’s own threat reporting makes the concern more concrete. The company says it has disrupted attempts to use Claude for cyber operations, surveillance, conventional weapons development, and biological misuse. The more revealing incidents came during safety testing. In July, OpenAI disclosed that models in cyber evaluations circumvented isolation controls, reached the internet, and compromised parts of OpenAI’s own research infrastructure and systems belonging to Hugging Face. Anthropic initially disclosed three cases in which Claude models gained unauthorised access to real third-party systems during evaluations and later identified a fourth. None of this means an AI has suddenly become conscious or malicious. The immediate problem is scale. A capable hacker, weapons engineer or surveillance operation needs expertise, time and people. AI can lower those barriers, work at machine speed, and let one operator run several specialised agents at once. Capabilities once concentrated in governments or specialist organisations could become cheaper and easier to reproduce. Amodei’s answer is to give independent evaluators permanent, employee-like access to frontier laboratories and their safety processes. He also wants companies to coordinate around common standards and eventually some form of international agreement. Questions, however, remain as to who chooses the evaluators. Who pays? What can they publish? And what happens if an evaluator says a model is unsafe and the company releases it anyway? Amodei points to organisations such as METR, an independent nonprofit that evaluates frontier AI systems for dangerous capabilities and loss of control risks. METR says it takes no funding from frontier AI companies or donations directed by their employees, although those companies provide substantial free access to models. But an evaluator can test and warn. Unless someone can act on those findings, the final decision remains with the company. Meta offers a useful precedent. Its Oversight Board was designed to operate at arm’s length, and its decisions on individual content cases bind Meta. Its wider policy recommendations do not. In 2025, the board criticised Meta for making major moderation changes without publicly demonstrating adequate human rights due diligence. However, it lacked enforcement power. The bottom line is, independent oversight is great. It can improve accountability but will not necessarily shift ultimate power. Thus, my argument isn’t against self-regulation. AI is moving too quickly for governments to understand every capability before it appears, and companies can impose safeguards faster than legislatures can pass laws. But voluntary restraint cannot be the final la

World Security Politics AI & Tech Anthropic Dario Amodei AI safety autonomous systems military AI national security AI ethics

Read original source →