Design Notes
Design Notes are non-normative essays on ARGUS design choices, trade-offs and limits. They do not amend the protocol.
What makes an analysis auditable? Four design choices behind ARGUS, and the limit they do not overcome.
A language model can produce a competent-sounding critique of almost any document put in front of it. Whether the criteria were fixed before the analysis began or fitted to the text afterwards, whether contrary evidence was sought or quietly avoided, whether the model deferred to the author’s standing or strained to look independent: none of this is visible from the prose alone. A plausible critique and a sound one look alike on the page.
ARGUS is built around that problem. The protocol’s analytical steps are documented elsewhere on this site. These notes cover something different: the design decisions that let a reader check how an analysis was produced, rather than take its conclusions on trust.
Four of them carry most of the weight. The examples come from the Amodei case study, which was run in a deliberately awkward configuration: an Anthropic model analysing an essay by Anthropic’s chief executive. No procedure makes a model impartial in that situation. The aim was narrower. The procedure should remain checkable by someone who trusts neither the model nor the analyst.
The criteria are fixed before the analysis begins
Under the current versioning policy, each new ARGUS release states whether it is major, minor or a patch, and why. Earlier entries remain visible as published. This is why the version history is part of the audit trail rather than a changelog for the curious.
In the Amodei case, the stronger condition happened to hold as well. The test that produced the central finding, the modal constancy check, had been published in ARGUS V5.1.0 in August 2026, before the September essay existed. The test asks whether the same load-bearing proposition is stated strongly where that strength does argumentative work, then narrowed where maintaining it would carry a cost.
One risk of ad hoc textual criticism is a criterion shaped around the text already in front of the analyst. A dated public specification does not make that impossible, but it makes the chronology checkable instead of deniable.
The standard of proof is declared before substantive testing begins
The analysis states what it will accept as adequate support for the genre in hand before the substantive tests are run. In the Amodei case, that meant an identifiable or retrievable source for a supported claim, a stated causal mechanism for a forward-looking one, a commitment determinate enough for a third party to check later, and a declared base and perimeter for a quantitative claim.
The same discipline applies to the numerical controls. The figures treated as probative are declared before calculation, rather than selected afterwards according to which results prove useful.
This does not make the chosen standard correct. A reader may find it too strict or too lax. The point is that the reader can argue with the standard separately from the verdicts it produced.
The publication summary is blind to everything but the analysis
“Blind” has a specific meaning here.
Once the normative analysis is complete, the publication synthesis is produced in a separate run from that analysis alone, with its own primary summary masked. The analysed object is not visible. Neither are the external source objects, the previous conversation or the translations, and external research is disabled.
The second run is therefore unable to repair the analysis by returning to the original essay, add a fact from a source the analyst omitted, or quietly strengthen a conclusion through a fresh search. Its task is narrower: reconstruct what the text argues, state what the completed analysis finds solid, identify the principal weaknesses and what survives them, preserve the distinction between a strong claim that fails and a weaker one that holds, and introduce no finding that is absent from the supplied analysis.
Keeping the summary in the same run would leave it exposed to the same contextual emphases. Separation does not make the second model impartial. It narrows the dependency and makes it inspectable.
Corrections are declared, including those that soften the result
An early reading of the Amodei rendering, before its link annotations had been extracted and incorporated into the audit, treated the essay as essentially unsourced. That reading was wrong.
Once the link layer was examined, the source audit and the assessment of interpretive closure had to be rebuilt. The correction moved the latter from level 3 to level 2, in the essay’s favour. The analysis records that change rather than absorbing it silently.
It also records a bias risk running in the opposite direction from the obvious one. A model produced by Anthropic might defer to Anthropic’s chief executive. It might also overcorrect and become unusually severe in order to display independence. The analysis names both possibilities and states where that second tendency would be expected to appear: inflated severity findings or an unnecessarily high closure level.
There is even a concrete example of deference in the record. The analyst initially failed to check the essay’s first stated reason because a technical statement by the chief executive of the company doing the work appeared to come from the person best placed to know. The protocol eventually forced the check.
The text’s own links supplied some of the strongest objections
What these properties produce is visible in the case study.
The essay says that AI progress is now driven primarily by AI’s growing ability to build the next generation of AI. The Anthropic Institute page linked at that claim describes substantial acceleration of human engineering work, while stating that it remains genuinely unclear whether current methods could unlock the stronger capacity claimed.
The essay also places Anthropic’s incidents beside the OpenAI–Hugging Face episode as “similar, though less severe”. Anthropic’s own linked assessment says its incidents involved single model instances, that there was no coordination between agents, and that the models did not depart from their assigned exercises. Those differences matter to the comparison the essay is making.
A third example concerns the standards body proposed by Demis Hassabis. The essay invokes it as an example of industry dialogue. The proposal itself goes further: it envisages a body that begins voluntarily, can become mandatory for US deployment, and can coordinate a slowdown among frontier labs if necessary.
The important point here is not that the analyst found sources hostile to the essay. It is almost the reverse. Some of the analysis’s strongest objections were reached by following sources the essay itself chose to provide.
The same apparatus records evidence in the text’s favour
The procedure also produced credits.
The analysis records the essay’s dated 6–12 month prediction, which can later be checked against events. It credits the disclosure of Anthropic’s own failures and the links supplied to the underlying reports. It also credits a provision giving outside reviewers the right to publish findings without Anthropic’s editorial control.
The essay’s sourcing matters for the same reason. At its most charged description of the OpenAI–Hugging Face incident, it links an independent investigation that supports most of what the sentence says. The analysis records that support rather than treating charged language as a defect by default.
An apparatus that only ever produces criticism is not an apparatus. It is a verdict written in advance.
Auditability is not experimental reproducibility
The limit should be stated rather than left for someone else to point out.
This is not reproducibility in the experimental sense. Re-running ARGUS on the same text does not guarantee an identical analysis, and claiming otherwise would be false.
What the design can provide is weaker and still useful: disagreement becomes locatable.
Because the steps, evidential standard, sources and corrections are declared, two analysts who differ can identify the point at which they part company and argue there, rather than exchange general impressions of the text.
The Amodei analysis does this on its own central finding. The modal-retraction verdict is recorded at medium confidence. The competing reading is stated explicitly: perhaps “pacing” always meant the weaker requirement of allowing enough time for alignment and verification, in which case there is no retraction and only an over-compressed opening sentence. The analysis rejects that reading, gives its reasons, and labels the rejection a judgement rather than a demonstration.
That is a smaller claim than reproducibility. It is also one a reader can test.
At the level of principle, none of these requirements needs specialised infrastructure. The cost is additional inference and disciplined record-keeping.
Publish the criteria, with dates, before applying them. Declare what will count as evidence before substantive testing begins. Separate the runs. Record the corrections, above all those that weaken your own conclusion.
Without that record, an analysis asks to be believed on the strength of how it sounds.