The US government labels military AI entities it treats as genuine threats as supply chain risk. It normally lands on shell companies and foreign hardware vendors. But in late February 2026, it landed on a San Francisco startup whose entire brand is being the cautious one.
That startup is Anthropic, and the Pentagon is now testing AI models from OpenAI, Google, and xAI to replace its Claude across classified systems. Testing started in early March, three days after the supply-chain-risk designation, with 25 of the department’s top users. The trigger wasn’t performance. It was Anthropic’s refusal to drop two safety guardrails. That single decision is the most interesting thing happening in military AI right now, and it points to a question every AI company will eventually face: who sets the limits?
Key Takeaways
- The Pentagon labeled Anthropic a supply chain risk in late February 2026. This was the first time the designation has hit a US company, and it then began testing its rivals.
- The trigger was Anthropic’s two red lines: no mass surveillance of Americans, and no fully autonomous weapons without human oversight.
- The testing covers OpenAI, Google, and xAI, with 25 “power users,” and began roughly three days after the designation.
- The Pentagon wanted Claude usable for “all lawful purposes” and argues that existing law already covers those concerns.
- The July 2025 awards were up to $200 million each to four labs, Anthropic, OpenAI, Google DeepMind, and xAI, not an Anthropic-only deal.
- Claude is deeply embedded in the Maven Smart System, so a swap is slow. Contractors got an 180-day window.
- No replacement has been named, and Anthropic is challenging the designation in court.
Why Is the Pentagon AI Program Replacing Anthropic’s Claude?
The Pentagon is replacing Anthropic’s Claude because Anthropic refused to remove safety limits the military wanted gone, and the Defense Secretary responded by branding the company a security risk.

The designation of supply chain risk is like pulling the plug. It doesn’t just end one contract. It pressures every other contractor to drop Anthropic, too, which is how a single policy disagreement turned into a full search for new military AI models to replace Claude.
What Happened: How the Anthropic Military AI Deal Fell Apart
Here’s the sequence, stripped to dates and events.
- July 2025: Anthropic is one of four labs awarded up to $200M each for frontier AI prototyping.
- Late 2025: Claude’s government version gets embedded in classified military intelligence and the Maven Smart System.
- February 2026: The Pentagon pushes for unrestricted “all lawful purposes” use. Anthropic refuses to drop its guardrails.
- Week of Feb 23: Hegseth gives Amodei a deadline to comply or face a government blacklist.
- Late February: Trump orders all federal agencies to stop using Anthropic, and the supply-chain-risk designation is formalized.
- March 1: The Pentagon begins testing rival AI models with 25 power users.
- March 6: An internal memo signed by DoD CIO Kirsten Davies orders Anthropic AI to be removed within 180 days from systems including nuclear weapons, missile defense, and cyber warfare.
- Spring 2026: Testing continues. Anthropic challenges the designation in court, saying it could cost billions.
From the $200M deal to “supply-chain risk”
| Date | Event |
| July 2025 | Four labs win up to $200M each |
| Feb 27, 2026 | Anthropic named a supply chain risk |
| March 1, 2026 | Rival model testing begins |
| March 6, 2026 | 180-day removal order issued |
OpenAI, Google, and xAI: The AI Models the Pentagon Is Testing
Three labs are in the running, and their pitches differ less on raw capability than on how far each will go.
OpenAI moved fastest. When Anthropic walked, OpenAI signed its own Pentagon deal within days, reportedly with some safety protocols but structured around lawful government use. Its strength is the same one that makes it Claude’s closest rival commercially: coding and reasoning.
Google is the more complicated case. Google’s Pentagon agreement permits use for any lawful government purpose, paired with a non-binding statement that it still opposes mass surveillance and autonomous weapons without human oversight. A preference isn’t a contract clause. Gemini’s edge is Google’s ecosystem.
xAI’s Grok rounds out the field. A Pentagon official told CNN that xAI was on board with being in a classified setting early, which signals the most permissive posture of the three. Grok’s pitch is real-time data handling.
No winner has been named. The Pentagon has said it wants multiple models available so no single vendor sits in a favored position, which is a lesson drawn straight from the Anthropic episode.
| Model | Vendor | Claimed strength | Guardrail stance |
| GPT / Codex | OpenAI | Coding and reasoning | Lawful-use deal, some protocols |
| Gemini | Sensors, scale, robotics | Broad lawful use, non-binding limits | |
| Grok | xAI | Real-time data | Most permissive, early adopter |
| Claude (incumbent) | Anthropic | Ease of use, performance | Hard red lines, refused to drop them |
How the Pentagon Is Running the Evaluation
The test is built to avoid locking in a favorite. The mechanics, as reported:
- 25 designated “power users” run the evaluations, rather than a central committee.
- They work inside a separate digital platform, reportedly called GenAI.mil, that runs independently of the Maven Smart System where Claude lives.
- Operators feed competing AI models identical prompts, then compare outputs. The official running it said the models respond differently to the same prompts, and that varying the prompts helps get the best out of each, while declining to share interim rankings.

Here I made a few observations. First, testing on a platform separate from Maven means they are trying the models in isolation before touching the system Claude is actually wired into. That’s quietly admitting how hard the real swap will be. Second, the no single favored vendor language is the lesson from Anthropic. Third, officials said early users didn’t push back as hard as expected against trying alternatives, which undercuts the idea that Claude was irreplaceable on quality alone.
AI Safety vs. Warfighting: Why Anthropic and the Pentagon Disagree
Anthropic’s two red lines, in its own framing; no mass surveillance of Americans, and no fully autonomous weapons that target people without human oversight. Amodei says these date to “day one.” His argument on weapons isn’t pacifism. It’s that current models aren’t reliable enough to be trusted to select and kill targets on their own.
On surveillance, Anthropic’s reasoning is about scale. In its statement to the Department of War, the company argued that powerful AI can assemble scattered, individually harmless data into a full picture of a person’s life automatically, and that the law hasn’t caught up to that capability.
Now, the Pentagon’s side stated as strongly as it deserves. Officials argue that mass surveillance and autonomous weapons are already restricted by US law and military policy, so Anthropic’s contractual bans are redundant and amount to a vendor inserting itself into the chain of command. The Pentagon’s chief technology officer, Emil Michael, said the military even offered written acknowledgments of the laws that restrict those uses, and that at some level, you have to trust the military to follow them.
Here the definitions themselves are contested. Lawfare points out that there is no commonly agreed definition of lethal autonomous weapons, and that “mass surveillance” covers everything from clearly illegal bulk warrantless collection to filming people who walk up to a building. So both red lines are fuzzier than a headline makes them sound.
The Technical Reality: What Replacing Claude’s Military AI Requires
“Replace” is doing a lot of work in the headlines. Swapping out Claude isn’t like changing a setting, and this is where a lot of coverage oversells the speed.
Claude isn’t a side tool. It’s embedded in the Maven Smart System, and Anthropic was the only AI company with models on the Pentagon’s classified networks, which also reported that Claude was in use during operations against Iran.
The friction is procedural as much as technical:
- Every replacement model has to be validated and, in many cases, re-authorized for secure environments before it can run in operations.
- Contractors who built Anthropic into their systems got an 180-day window to find alternative military AI solutions.
- The removal order covers systems for nuclear weapons, missile defense, and cyber warfare, which are not places anyone rushes a migration.
Joe Saunders, CEO of RunSafe Security, put it plainly in reporting on the transition: these models are woven into accredited environments and mission-specific workflows, and each new one needs validation, often re-authorization, before it can be used operationally. In plain terms, the Pentagon can pick a new model quickly and still spend months actually switching.
The Future of National Security: What This Means for the Military AI Market
Zoom out, and this stops being an Anthropic story and becomes a precedent. Bloomberg framed it as a race to find alternatives, but the more durable signal is what it tells every other lab about the cost of drawing lines in the military AI sector.
A few things look solid, and I will flag where I am speculating.
It’s confirmed that OpenAI structured its deal around lawful government use, and xAI signaled early it would work in classified settings. The labs that stayed flexible got the contracts. Michael said he expects rivals to ship capabilities comparable to Anthropic’s every month or two, which is the Pentagon’s bet that no single lab is irreplaceable for long.

Speculative but reasonable: enterprise and government buyers are starting to ask not just which AI models a vendor uses, but what that vendor has committed to.
There’s history here, too. The closest parallel is Google’s 2018 Project Maven backlash, when employee protests pushed Google off a Pentagon drone-analysis contract. Eight years later, Google signed a broad defense deal, and the safety-first lab is the one walking away. The center of gravity in defense technology and national security work has clearly shifted.
Final Thoughts
Strip away the procurement detail, and you are left with one clean question; when a private company builds frontier AI, who decides how far it goes in war?
Anthropic answered by drawing two lines and accepting the consequences. The Pentagon answered by deciding that a vendor doesn’t get to set those limits by contract, then moving to other AI models fast enough to make the point. Both answers are defensible. Neither is obviously right.
What happens next runs through a courtroom, not a benchmark. Anthropic is challenging the supply-chain-risk designation, and that case will shape whether any future lab can hold a red line against its single biggest possible customer. For now, the most safety-restricted company in the field is being pushed out of the military’s most sensitive systems while more permissive rivals line up to take its place. Whether that is a cautionary tale or simply how military AI was always going to go is still being decided.
FAQs
Because Anthropic refused to drop guardrails that the military wanted removed. Defense Secretary Hegseth applied the designation in February 2026, reportedly the first time it has hit a US company.
The Pentagon is testing OpenAI, Google, and xAI. No winner has been named, and officials want several AI models available rather than one dominant vendor.
Yes, for now. Claude remains embedded in the Maven Smart System, but a March 2026 order set a 180-day window to remove Anthropic’s technology.
No mass surveillance of Americans, and no fully autonomous weapons that select and kill targets without human oversight. Anthropic says it has held both since its founding.
OpenAI signed its own defense deal within days of the dispute, built around lawful government use. It is one of three labs now being tested, not an automatic replacement.

