Loading...
Research & Insights

U.S.-China AI talks could make autonomous cyberattacks a matter of strategic stability

September 5, 2026 · jason.ellis

U.S.-China AI talks could make autonomous cyberattacks a matter of strategic stability

Officials in Washington and Beijing are preparing to sit down this month for a dedicated government-to-government dialogue on artificial intelligence safety, Reuters reported, with cooperation on monitoring AI-directed cyberattacks and information-sharing between AI laboratories among the ideas under discussion. The timing, location, participants and agenda remain unsettled, and a U.S. Treasury spokesperson said the meeting may slip to October.

Whenever it happens, the session will be the first working test of a commitment made at the leaders' level. After President Donald Trump's summit with Xi Jinping in Beijing in May, China's foreign ministry confirmed on May 19 that the two governments had agreed to an "intergovernmental dialogue" on AI. Treasury Secretary Scott Bessent described the ambition as an effort to "set up a protocol" covering "best practices for AI to make sure nonstate actors don't get a hold of these models."

The reason such a channel is being discussed at all is that the technology has changed the stakes. Officials and analysts on both sides of the Pacific now treat the newest frontier models as capable of discovering and exploiting software vulnerabilities at a scale no human team can match, including thousands of zero-day flaws in legacy software. Both governments have seen what even human-directed campaigns can do: the July 2025 intrusions against Microsoft SharePoint servers, linked by Microsoft and several security researchers to China, reportedly hit more than 400 organizations, including the U.S. National Nuclear Security Administration, the agency responsible for the nation's nuclear weapons program. Neither capital has a communication channel designed for a machine-speed version of that event.

What the proposed dialogue would actually try to do

The question now confronting both governments is narrow and practical: what verification, incident-reporting, model-evaluation and crisis-communication mechanisms could actually be built, given that neither side will expose its intelligence methods, its companies' trade secrets or its deterrent posture?

What the proposed dialogue would actually try to do

The strongest current of expert advice in Washington holds that the talks must be modest by design. Writing for the Brookings Institution, Mark MacCarthy and Carl Schonander argue the dialogue should be "a regular process of sharing information, ideas, plans, and research in the area of AI model risk assessment and mitigation," not a negotiation in which each side trades concessions. They would keep export controls, open-source policy and U.S. model access explicitly off the table. The natural agenda, borrowing a formulation from researchers Christina Knight and Scott Singer, is to build "a shared understanding of risk" and exchange "best practices for reducing the risks of AI models."

Running alongside that technical exchange is a second, more operational idea: a standing channel for notifying each other about time-sensitive AI incidents. Researchers Karson Elmgren and Clarissa Koh of the Institute for AI Policy and Strategy, previewing forthcoming work, point out that no such channel exists today, "creating risk of misinterpretation and escalation when clarity matters most." Their scenarios are concrete: an autonomous AI system malfunctions; frontier model weights are stolen and proliferate across global networks; a compromised AI system launches cyberattacks against critical infrastructure in several countries at once. In each case, U.S. officials would need to reach technically capable Chinese counterparts quickly, faster than diplomatic cables and conventional hotlines allow.

Those two tracks, a risk-assessment exchange and an incident channel, map onto the four mechanism families that run through the current policy debate: crisis communication, incident reporting, model evaluation and, hardest of all, verification.

Why both governments changed course in 2026

Why both governments changed course in 2026

Until recently, neither government treated frontier AI as a threat category of its own. Ryan Fedasiuk of the American Enterprise Institute describes both capitals as having treated AI as "a normal technology," a commodity to be contested in third markets. The Trump administration signaled that posture early: at the February 2025 Paris AI Action Summit, Vice President Vance warned against regulation that might "deter innovators from taking the risks necessary to advance the ball." Beijing, for its part, built stringent controls on training data and model outputs aimed at political stability while, in Fedasiuk's account, looking away as Western researchers jailbroke DeepSeek's models to produce bioweapons assistance and offensive cyber tradecraft.

That posture broke in the spring of 2026. In Fedasiuk's telling, models of the class now coming online, including Anthropic's Claude Mythos and OpenAI's GPT-5.5, "have crossed a frightening, existential threshold," because the capacity to discover and exploit thousands of zero-day vulnerabilities in legacy software "is not a feature either the Washington security establishment or the Chinese Communist Party will tolerate in the hands of a private champion accountable only to investors." Elmgren and Koh report that Anthropic's limited release of Mythos appears to have prompted a shift in the Trump administration's approach, culminating in an executive order strengthening government oversight of frontier models.

China's shift has a different texture but points the same direction. Xi Jinping has repeatedly called for systems for technology monitoring, risk warning and emergency response, and has highlighted "technological loss of control" (技术失控) as a particular concern. China's 2025 National Emergency Response Plan placed AI security incidents alongside earthquakes, cyberattacks and infectious disease epidemics. Mythos, Elmgren and Koh write, "has also been a topic of significant concern for Chinese counterparts thinking about AI risks."

Why the existing channels cannot carry an AI crisis

This matters politically for one reason above all: a bilateral mechanism survives only when both governments see themselves as benefiting. When risk perception is roughly symmetric, neither has an incentive to hold the channel hostage to broader negotiations, the failure mode that has crippled previous bilateral arrangements.

Why the existing channels cannot carry an AI crisis

The infrastructure for U.S.-China crisis communication does exist, and it has a record. The problem is what that record shows.

The Beijing-Washington hotline and the Defense Telephone Link are the two formal channels. But the DTL requires 48 hours' advance notice to set up a call, an eternity in a fast-moving AI incident. Both channels depend on generalists, lacking the specialized expertise not just to describe what is happening in an AI event but to assess its significance and avoid misinterpretation. And both have a history of being silenced as political signals. Calls placed on the hotline after the 1999 bombing of the Chinese embassy in Belgrade and the 2001 Hainan Island air collision went unanswered. China left the DTL dormant during then-Speaker Nancy Pelosi's 2022 Taiwan visit and the 2023 balloon incident; it was never formally suspended, but Beijing cancelled adjacent military dialogue tracks, rendering it effectively unusable.

Finally, because AI incidents are an emerging threat category, neither government has shared definitions or procedures to fall back on. In a crisis, each side would be improvising a vocabulary while the clock ran.

The Cold War offers the counter-pattern. As a January 2026 analysis in The Diplomat recounts, Washington and Moscow built narrow agreements on nuclear testing, incident reporting and crisis hotlines "long before there was anything like trust," and kept those expert channels insulated, "not cut off as a vehicle to demonstrate political displeasure." Those measures did not end the arms race, The Diplomat notes; they made it "less likely to end the world."

What a crisis-communication channel would need to look like

What a crisis-communication channel would need to look like

The design requirements follow directly from the failures above. A dedicated AI incident channel would need to close four gaps: speed, expertise, protocols and resilience.

Speed means notification procedures measured in hours, with both sides maintaining staffed contact points around the clock, not a hotline that requires two days of scheduling. Expertise means the people on the line can distinguish a serious frontier-model failure from an ordinary software incident, which requires delegations that pair senior officials with genuine technical experts. Protocols means agreed definitions of what constitutes a reportable AI incident, drafted before anyone needs them. Resilience means the channel must keep operating during periods of political hostility, because that is precisely when misinterpretation risk peaks.

Elmgren and Koh judge the proposal "more politically feasible than it first appears." Part of the reason is structural: a notification channel requires neither side to restrain its capabilities; it asks both sides to protect themselves from the same category of surprise.

What each side would have to report, and what it would not

Incident reporting is the most concrete deliverable the September or October session could produce, and also the one requiring the most careful architecture.

What each side would have to report, and what it would not

The first task is taxonomic: agreeing on what an "AI incident" even is, and which incidents cross the threshold for mandatory notification. A plausible shortlist, drawn from the scenarios Elmgren and Koh describe, includes loss-of-control events involving frontier models, confirmed theft or unauthorized proliferation of frontier model weights, and cyber operations against critical infrastructure where a state suspects or knows that an AI system directed the attack.

The second task is designing what the notification contains. Here the model resembles outbreak notification in public health: under the World Health Organization's International Health Regulations, states must notify potential public health emergencies of international concern quickly, on incomplete information, because waiting for certainty defeats the purpose. An AI incident notice would similarly carry operational facts, what happened, when it was discovered, what systems are affected, what containment measures are underway, without requiring either side to disclose how it knows.

That distinction is what protects intelligence equities. Attribution is the part of any cyber event that touches sources and methods, and it should be quarantined from the technical channel entirely. The channel exists to reduce misinterpretation and enable containment; accusations of responsibility, when they come, can travel through existing diplomatic and political tracks that states already use for that purpose. A blame-free technical mechanism is also more likely to survive, since using it does not require either government to concede anything about its own operations.

The Reuters report that laboratory-level information-sharing is on the agenda points to a related design choice: AI labs are often the first entities to see misuse of their own models. A sanctioned channel through which leading American and Chinese labs could share indicators of AI-directed malicious activity, with government awareness on both sides, would extend the state's visibility closer to the technical edge. The evaluation track and the incident track reinforce each other here, because the shared vocabulary developed in evaluation work is what makes an incident report intelligible on the other side.

How do you evaluate models without exposing them?

Model evaluation is the mechanism most amenable to early agreement, because it already has domestic analogues in both countries and because, done properly, it requires neither side to surrender commercial secrets.

Executives signing international agreement with EU and US flags displayed on a wooden table.
Photo by Werner Pfennig on Pexels

The key is that capability evaluation is black-box by nature. Testing whether a model can assist with offensive cyber operations, or can meaningfully uplift a biological weapons program, requires observing the model's behavior, not possessing its weights or training data. Two governments can exchange testing protocols, red-teaming methods, evaluation results on publicly available models and threshold definitions for dangerous capabilities without any company's proprietary internals changing hands. Brookings identifies as ripe for exchange testing protocols and red teaming before release, plus mitigation techniques: alignment methods, safeguards that constrain models to safe environments, and measures that harden attack surfaces.

Both sides have starting material. Washington now has an executive order establishing a framework for AI companies to cooperate with the government on model risks. Beijing has an existing regulatory apparatus, including its TC260 standards work, through which it controls model outputs, even if that apparatus was built for censorship rather than safety.

The staffing lesson comes from failure. The first U.S.-China intergovernmental AI dialogue, held in Geneva in May 2024, stalled in part because, as Brookings later noted, the two sides sent mismatched delegations, political and technical counterparts who did not match up. Any durable evaluation exchange will need technical subject-matter experts in the room, which is why the Sandia National Laboratories assessment recommends seeding the work through Track 1.5 and Track 2 dialogues that include industry experts and government laboratory scientists, citing Sandia's own Cooperative Monitoring Center as an example of an institution suited to that role.

Can any of this be verified?

A female engineer using a laptop while monitoring data servers in a modern server room.
Photo by Christina Morillo on Pexels

Verification is where honesty about limits matters most, because the answer, for now, is: only partially, and only at the edges.

Classical arms-control verification, intrusive inspection regimes with mutually agreed counting rules, does not transfer to AI. Models are intangible, copyable and dual-use. There is no equivalent of counting missile silos. Technical research is exploring more exotic possibilities, from monitoring the data centers where frontier models train to tracking advanced chips to cryptographic techniques that might someday let a lab prove properties of a training run without revealing the model itself. None of these is mature enough to anchor a bilateral commitment, and each collides with sovereignty objections that both governments, for different reasons, take seriously.

What is available is a ladder of weaker instruments. The first rung is declaratory: the November 2024 agreement between Presidents Biden and Xi to avoid giving AI control over nuclear weapons systems showed that a narrow, leader-level commitment is possible where both sides' deterrent logic already implies the restraint. A cyber analogue, human review requirements before AI systems conduct operations against critical infrastructure, is conceivable on the same logic, though markedly harder to monitor.

High-resolution close-up of HTML code displayed on a computer screen, perfect for technology themes.
Photo by Bibek ghosh on Pexels

The second rung is what Matt Sheehan of the Carnegie Endowment has called "parallel" action: each side adopts testing regimes, safeguards and release practices on its own initiative, not because the other side agrees to match it. Compliance is then verified the old-fashioned way, by observing behavior, rather than by inspections. Parallel action does not need the other side's permission and cannot be suspended as a bargaining chip.

The third rung is reciprocal incident reporting itself, which over time builds a statistical record each side can check against its own telemetry and open-source evidence.

The critical boundary is what verification talk excludes. Fedasiuk warns that Beijing's opening ask, as in every meeting since 2023, will be to relax chip controls, offered in exchange for "constructive dialogue." Brookings, from a different institutional vantage point, reaches the same structural conclusion: export controls belong in a separate process. Folding them into safety talks would hand Beijing leverage over the safety mechanism's existence, exactly the hostage dynamic that makes bilateral channels brittle.

This design preserves strategic deterrence almost by definition. Nothing in the realistic menu caps anyone's capabilities, restrains military AI development or fixes force postures. The Diplomat frames the goal as "the same kind of narrow floor" Washington insisted on in earlier technology races, cooperation first where sharing poses little risk, the threats are global, and both sides fear the same disasters. Its historical parallel is telling: the U.S. government ran the Advanced Encryption Standard as an open international competition that strengthened civilian cryptography while sensitive military systems stayed classified. Open at the layer where openness helps defenders; closed where it does not.

When information sharing backfires

Miniature caution cone on a computer keyboard symbolizing data security and control.
Photo by Fernando Arcos on Pexels

Any honest design has to reckon with an episode neither government will discuss on the record but every designer is thinking about.

In August 2025, Microsoft restricted certain Chinese firms from its Active Protections Program, a program that gives security vendors early access to vulnerability details, including proof-of-concept code, so they can prepare defenses before disclosure. The restriction followed the mass exploitation of SharePoint servers and what CSO described as speculation among experts that MAPP information may have leaked to attackers. That causation remains unproven. But the episode demonstrates the incentive problem at the heart of any information-sharing regime with a state-adjacent ecosystem on the other side: the same early warning that lets defenders patch lets attackers move faster.

The security community's own assessment cuts both ways. Limiting Chinese firms' access "has a potentially negative impact in terms of the creation of blind spots and 'windows of vulnerability,'" Rik Turner of Omdia told CSO. Keith Prabhu of Confidis was blunter about the systemic cost: "the chain is only as strong as its weakest link."

For a government-to-government channel, the lesson is not that sharing is futile. It is that sharing must be tiered and graduated: incident notifications and containment best practices early, exploit-level technical detail later, if trust accumulates. A private vulnerability program is not a government incident channel, but the trust-gating logic transfers.

What could still break the process

The obstacles are as much political as technical.

Start with the schedule itself: the mid-September session remains unconfirmed, and the October alternative reported by The Jerusalem Post suggests the preparations are fragile. Agendas this sensitive slip for reasons that have nothing to do with AI.

Then there is the risk of overreach. Brookings singles out two tempting frameworks it considers premature: the arms-control approach associated with former Biden administration officials like Chris McGuire, seeking mutual constraints on AI development, and the crisis-red-line approach sketched by analyst Paul Triolo, who has suggested the two governments could "define red lines around AI-enabled cyber operations, biosecurity, autonomous military escalation, model theft, and third-party/non-state actor use." Brookings's point is not that these are wrong forever; it is that trust has to be rebuilt first, in gradients.

Network switch and blue ethernet cable with white tips connected to system for maintenance
Photo by Brett Sayles on Pexels

Washington's own instability is a genuine risk factor. The executive order framework arrived alongside what Brookings calls the "Mythos-Fable debacle," a setback to stable U.S. AI governance, and what Fedasiuk describes as the Department of War "caught in a blood feud with one of the country's most capable AI champions." A counterpart negotiating with Washington this autumn cannot be sure which faction owns AI policy in six months.

Beijing has its own unresolved contradiction. Fedasiuk reports asking dozens of Chinese scholars, technical experts and officials how the government's stated governance ambitions square with a national strategy built on open-weight models whose commercial logic depends on uncontrolled proliferation. "None has been able to offer a coherent answer," he writes.

And hovering over everything is the historical pattern: both governments have repeatedly let political retaliation sever working channels at exactly the wrong moments. An AI incident that lands during the next Taiwan crisis or the next balloon-style confrontation is the stress test the new mechanism must be built to survive.

What progress would actually look like

For a process this early, success is plumbing, not promises. The concrete markers to watch after the first session:

  • A shared working definition of a reportable AI incident, with agreed thresholds.
  • Named, technically staffed contact points in each government, with response-time commitments measured in hours.
  • A pilot notification exercise conducted before any real incident forces the issue.
  • A parallel track exchanging pre-deployment testing methods, red-teaming practices and dangerous-capability thresholds, run by people who can read an evaluation report.
  • A sanctioned channel for AI laboratory information-sharing on AI-directed malicious activity, as floated in the Reuters reporting.
  • Explicit separation of attribution from incident notification, so the technical channel stays open even when the political one is hot.

None of these measures would stop an AI arms race, constrain a deterrent, or expose a source. That is the point. The July 2025 SharePoint intrusions reportedly struck more than 400 organizations, including the agency that oversees the U.S. nuclear weapons program, in a campaign run by human operators. The capability class arriving this year threatens to hand that speed and scale to machines, and both governments now know it. The September or October session will show whether they can build the narrow, insulated floor underneath that race, the kind that exists before the crisis, because after one begins, nobody builds anything.

Sources

An empty conference room featuring red seats, a stage, and Turkish flags, ready for a meeting or presentation.
Photo by Berna on Pexels
  • "US, China gear up for mid-September AI safety dialogue," Reuters via Marketscreener, September 2026. https://hk.marketscreener.com/news/us-china-gear-up-for-mid-september-ai-safety-dialogue-ce785bdbd98aff26
  • "US, China to discuss AI safety risks, cooperation on cyberattacks," The Jerusalem Post, September 2026. https://www.jpost.com/international/article-907617
  • Karson Elmgren and Clarissa Koh, "A U.S.-China Communication Channel on AI Incidents: Provisional Recommendations," Logogram (Institute for AI Policy and Strategy), June 4, 2026. https://logogram.substack.com/p/a-us-china-communication-channel
  • Ryan Fedasiuk, "AI Coordination Without Illusion," American Enterprise Institute, May 7, 2026. https://www.aei.org/commentary/ai-coordination-without-illusion/
  • Mark MacCarthy and Carl Schonander, "How the US and China can cooperate to reduce urgent AI risks," Brookings Institution, July 14, 2026. https://www.brookings.edu/articles/how-the-us-and-china-can-cooperate-to-reduce-urgent-ai-risks/
  • "How China and the US Can Make AI Safer for Everyone," The Diplomat, January 2026. https://thediplomat.com/2026/01/how-china-and-the-us-can-make-ai-safer-for-everyone/
  • Noelle Camp and Michael Bachman, "Challenges and Opportunities for US-China Collaboration on Artificial Intelligence Governance," Sandia National Laboratories, April 2025. https://www.sandia.gov/app/uploads/sites/148/2025/04/Challenges-and-Opportunities-for-US-China-Collaboration-on-Artificial-Intelligence-Governance.pdf
  • Prasanth Aby Thomas, "Microsoft restricts Chinese firms' access to vulnerability warnings after hacking concerns," CSO Online, August 21, 2025. https://www.csoonline.com/article/4043683/microsoft-restricts-chinese-firms-access-to-vulnerability-warnings-after-hacking-concerns.html
Share this article

Comments (6)

  • Jorge Sep 5, 2026

    This matches what we see in red-team work - frontier models can chain exploits across legacy systems far faster than any human team. The article's point about thousands of discoverable zero-day vulnerabilities is not hyperbole; we routinely pull dozens from a single overnight run.

  • jaded97 Sep 5, 2026

    Information-sharing between AI laboratories sounds sensible, but I'm not sure how either side would actually verify what models the other is evaluating without exposing proprietary training pipelines.

  • Mei H. Sep 5, 2026

    The phrase 'a normal technology' really stuck with me - it's striking how quickly both governments moved away from that framing once Mythos-class capabilities emerged. I'm left wondering whether the urgency will outlast the dialogue's first year or quietly slide back into commodity-competition mode.

  • simone.bennett Sep 5, 2026

    The SharePoint incidents hitting 400+ organizations including the NNSA feel like an early version of the model-weight theft scenarios Elmgren and Koh describe. I keep wondering whether Microsoft's restrictions on Chinese firms accessing vulnerability warnings will become a template for what frontier labs do next.

  • D. Shah Sep 5, 2026

    Calling the dialogue 'not a negotiation' is probably wishful thinking once export controls and open-source policy are explicitly off the table.

  • Jonas L. Sep 5, 2026

    The claim that Mythos's limited release 'prompted a shift in the Trump administration' deserves more scrutiny. What concrete policy changes actually followed that release, and how do we know the timing isn't just coincidence with other events in early 2026?

Comments are reviewed before they appear.

Continue exploring