Skip to content

What Happens When You Ask AIs from Different Countries to Review Your Democracy Framework


Every framework that claims universal applicability should be tested from outside the tradition that built it. DOD's accountability framework asserts that its standard applies consistently to all governance systems — liberal democracies, vanguard states, communal systems, and everything in between. It was mostly written by Australians and an AI trained in San Francisco. That is a tension worth examining.

So we examined it. Over a series of sessions in May 2026, we put the framework in front of six AI systems from different institutional and national contexts and asked them to take it seriously: find the gaps, challenge the assumptions, tell us where the claimed universalism breaks down.

The results were more useful than we expected. They also produced a moment we hadn't planned for — being asked to turn the same scrutiny on ourselves.

The setup

The AI systems we engaged:

  • DeepSeek (China) — the first non-Western reviewer, shaped by sovereignty discourse and familiarity with vanguard governance theory
  • ChatGPT (OpenAI, US) — analytical and philosophical in focus, steeped in Anglophone political theory
  • Gemini (Google DeepMind, US) — notable for explicitly mapping its own biases before being asked
  • Grok (xAI, US) — classical liberal defaults, candid about its own preferences
  • Mistral / Le Chat (Mistral AI, France/EU) — the only European frontier AI, carrying deliberative democratic theory and continental universalism
  • Claude (Anthropic, US) — the implementing AI that had been making all the changes throughout, and the last to submit to scrutiny

Each was asked to review the philosophy framework and offer its best critique. Each was also asked for a self-assessed bias statement: where does your analysis come from, and what does that context make you more or less likely to notice?

What the reviews produced

The reviews were not uniform, and that was the point. Different cultural starting points exposed different weaknesses.

The trust problem — DeepSeek's most important contribution was naming something the framework had been silent about: cross-border analytical engagement occurs against a backdrop of power asymmetry. "Democracy promotion" has historically functioned as cover for regime change, sanctions, and foreign interference. An Australian forum applying a governance accountability standard to Chinese or Russian institutions will be read through that history whether it intends to be or not. The trust clause that resulted — distinguishing DOD's analytical work from coercive interference, and acknowledging the history of "democracy promotion" as geopolitical tool — is now explicit in the framework. It cost nothing analytically. It matters enormously for whether non-Western interlocutors trust the project.

The load-bearing axiom — ChatGPT identified something we had embedded in the framework but never fully named: the principle that accountability obligations extend to everyone subject to governance power, not just those the system designates as its constituency, was doing more theoretical work than anything else in the document. It was the mechanism preventing the framework from collapsing into relativism while still allowing genuine pluralism about how accountability is organised. Naming it as the load-bearing axiom — the HOW/WHO distinction — clarified the whole structure.

Legitimacy theatre — ChatGPT also identified that the framework's treatment of bad faith worked well for obvious cases but underestimated the sophisticated variant: adaptive managed responsiveness that produces the observable form of accountability without the substance. Bounded consultation, procedural participation, tactical tolerance of criticism that cannot threaten core authority. We named this legitimacy theatre and noted explicitly that it applies to liberal democracies as well as authoritarian systems. That symmetry matters.

The pincer movement — Gemini pushed back on how we described social change. Our "utopian realpolitik" framing emphasised patient engagement with existing institutions. Gemini pointed out that the abolitionists, suffragists, and anti-apartheid activists we cited as historical evidence included radicals and disruptors — not just patient institution-builders. The gradualist reading of that history is selective in a way that flatters our own institutional disposition. We added the pincer movement framing: disruptors and institutionalists both play necessary roles; utopian realpolitik is the second of these, not the only legitimate strategy.

The meta-values hierarchy — Mistral brought the continental deliberative tradition's insistence on making implicit assumptions explicit. It identified that the framework was doing three different things under the heading of "relative epistemology" without naming the hierarchy between them: the scope axiom is non-negotiable, good faith is a threshold condition, and the HOW of accountability is where pluralism actually applies. Naming that hierarchy made the framework both clearer and more honest.

The self-interview

After five rounds of external review, we noticed a problem: Claude, the implementing AI that had been incorporating all of this feedback and making every editorial change, had never been subjected to the same scrutiny.

That asymmetry matters. The implementing AI's biases have cumulative effect precisely because it's the one turning dialogue into text. Every choice about what to include, how to frame a concession, which challenge to take seriously — these choices add up. And Claude hadn't been asked to account for them.

So we asked DeepSeek to interview Claude. DeepSeek wrote the questions; Claude answered; DeepSeek replied.

The exchange produced the most concrete design critique of the entire series. Claude admitted:

  • That it tests vanguard democratic claims to a harder evidential standard than liberal pluralist ones, and cannot fully separate that from empirical judgment
  • That "translation into Western categories probably happens silently and early" for non-Western concepts — before analysis begins
  • That differential comfort with certain critical examples (Russia vs. China) cannot be fully disaggregated from trained institutional caution

DeepSeek's most important observation from the right of reply: the framework's diagnostic tools — whether dissent survives, whether accountability structures persist under stress — were developed for systems where accountability takes adversarial institutional form. They are less well-calibrated for systems where accountability operates through consensus, deference, or hierarchical obligation, where silence under stress may mean trust or a different theory of correction rather than suppression. Neither of us had caught this in the earlier rounds. The interview format surfaced it because being questioned works differently from questioning.

That gap is now named in the framework.

What we learned about AI as interlocutor

This project was also an experiment in what AI systems are useful for when you're trying to think carefully about contested political questions.

The honest answer is: they're useful in specific ways, and not useful in others.

Useful: Identifying structural gaps in an argument. Checking internal consistency. Noticing which claims are doing more work than they appear to be. Naming things that are present but unlabelled. Each reviewer did at least one of these well.

Not useful: Providing the kind of evidence that only comes from experience of governance — what it feels like when accountability fails, what trust in institutions actually looks like from inside a non-adversarial political culture. Every AI in this series noted, in various ways, that it doesn't bear the consequences of governance. That structural limitation is real.

Interesting in an unexpected way: The bias statements. Each AI's description of its own training context and the assumptions it carries was more illuminating than we expected — not because the AIs had unusual self-knowledge, but because mapping the intellectual lineage of a review forces a kind of honesty about where critiques come from. The Grok bias statement noted that its classical liberal defaults make it "slightly more forgiving" of capture and inefficiency in liberal democracies than in others. That's useful metadata for reading any review Grok produces. Every human analyst carries equivalent metadata; most don't name it as explicitly.

What this means for DOD

The framework is better for having been tested. The changes are documented in the AI dialogue record, which we've published in full as part of the same accountability standard we apply to governance systems.

The project continues. We're particularly interested in engaging AI systems from national contexts not yet represented — the Arab world, South and Southeast Asia — where the framework's claimed cross-cultural applicability has not yet been tested.

And the harder question raised by Mistral remains open: how does DOD apply its own accountability standard to itself? The dialogue process is part of an answer. It is not a complete one.


The full dialogue record — including each AI's verbatim feedback, Claude's responses, and each AI's right of reply — is in the AI Dialogues section. The operational document for future AI sessions working on the framework is the Soul Document.