A Desk in the Building
On 12 September the chief executive of Anthropic published an essay arguing that the industry should slow the rate at which it improves model capabilities. The heads of OpenAI, xAI and Google DeepMind all responded in public within a day, and the coverage has treated that as three rivals agreeing. Agreement on the direction is not agreement on the mechanism, and on the mechanism they do not agree. This piece does not judge whether the argument is right. It examines the first of the three proposed steps, because that step is a disclosure mechanism, and disclosure mechanisms are readable from here. Each frontier company would give ongoing, employee-like access to a team of embedded third-party evaluators, covering not just finished models but training pipelines and processes. Strip the subject matter away and the shape is familiar. This is a proposal to install an independent party inside the building, with standing access and the right to publish what it finds.

16 September 2026. Statements and filings are as of the dates stated in the text.
Interest disclosure: T-OpenAI is a tokenized loan participation right providing economic exposure linked to OpenAI's pre-IPO valuation, issued through a dedicated issuing subsidiary. T-Kalshi is issued on the same basis. T-SpaceX, issued through its own dedicated issuing subsidiary, provides economic exposure linked to SpaceX; following the SpaceX listing, the investment underlying T-SpaceX is being divested. Redemption opens only once a Redemption Start Date is announced, and as of 16 September 2026 none has been. Tessera's founder, and entities he controls, hold related exposure. This piece concerns the timing of an OpenAI listing and a company that competes with OpenAI, so Tessera's interest here is direct rather than incidental. Read what follows with that in mind.
On 12 September the chief executive of Anthropic published an essay arguing that the industry should slow the rate at which it improves model capabilities. The heads of OpenAI, xAI and Google DeepMind all responded in public within a day, and the coverage has treated that as three rivals agreeing. More care is warranted than the coverage has shown: agreement on the direction is not agreement on the mechanism, and on the mechanism they do not agree.
Almost everything written since has been about whether the argument is right. We are not qualified to judge that, and this piece does not try. What it examines is the first of the three proposed steps, because that step is a disclosure mechanism, and disclosure mechanisms are readable from here.
What Step One Actually Proposes
The proposal is not a pause. The essay is explicit that pacing "does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this." That last clause is the one to hold on to, because verification by an outside party sits inside the definition rather than being bolted onto it.
The first step is narrower and more concrete than the headline. Each frontier company commits to giving "ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes."
Read that scope again, because the second half is the unusual part. The evaluators examine not only finished models but how the models are made. "Anthropic is unilaterally committing to this step now," the essay says, and it calls on governments to require other frontier companies to match it. OpenAI's chief executive is reported to have responded that committing to independent evaluators with employee-like access "is a great idea, and we will do the same," reported in major outlets and not confirmed here against a company statement.
The head of Google DeepMind, who responded within nine hours, is reported to have called the essay's direction correct while saying the details need working through, arguing in substance for a permanent standards body rather than lab-by-lab embedded-evaluator arrangements. That is a real objection to step one specifically, from someone qualified to make it, and it is the one the coverage lost when it rounded three responses up to agreement. It returns later in this piece, because it arrives at the same place as the objection reached here from a different direction.
Strip the subject matter away and look at the shape. This is a proposal to install an independent party, inside the building, with standing access and the right to publish what it finds.
Why That Is the Interesting Part
Who Counts the Capex, published on 17 August, examined why an announced capital expenditure number is never marked against anyone. Incurred spending enters the cash flow statement and is re-examined every ninety days. Announced, non-binding spending creates no obligation, commits no period, and leaves no line item to record a shortfall if it never happens.
A claim about catastrophic risk has exactly the same shape, and bluntness is warranted about it. "This technology could be dangerous by 2030" produces no line item in any quarter between now and 2030. There is no period in which it is marked, no statement it appears on, and no mechanism that records whether it was right. It is, in the strict sense, unmarkable.
That is why safety announcements have historically been discussed rather than priced, and the reason is not cynicism on anyone's part. There is nothing to price. A claim that no future document will ever settle cannot be discounted.
Step one is an attempt to build the missing document, and precision about which part is new matters here. Third-party evaluation of frontier models is not new; outside evaluators have been running pre-deployment assessments for years, and several jurisdictions are legislating conformity assessment regimes. What is being proposed is a change of degree that becomes a change of kind: permanent, employee-level access, and a contractual right to publish findings without editorial control.
That combination reaches for the shape of an audit function, meaning a party whose job is to convert an assertion by management into a record produced by somebody else. It is the thing the announced-capex problem has always lacked, and here a company is volunteering it rather than waiting to have it imposed.
The essay does not reach for the audit comparison. It reaches for a different one, and a better-chosen one: this "has precedent in the banking industry, which sometimes involves regulatory supervisors embedded along with employees." That is not the same institution. A supervisor is placed by a regulator and paid by one; an auditor is engaged and paid by the company it examines. The distinction matters for everything that follows. This piece keeps the audit frame because the output is an audit-shaped artifact, a record made by someone other than management, while flagging that the essay's own analogy answers part of the objection that follows.
The Half That Is Missing
A record only changes behavior if somebody is obliged to read it.
This is the place to push on the proposal rather than applaud it. As described, the evaluator's leverage is contractual: a right to publish, granted by the company being evaluated, without editorial control. That is a real commitment and more than anyone else has offered. But it is voluntary, it is unilateral, and no counterparty must act on what it says.
In fairness, the essay knows this. It is why step one is put as something one company does unilaterally while calling on governments to require others to match, and why steps two and three are about coordination rather than commitment. The gap described here is one the author has already named. What this piece adds is that the gap concerns more than who else adopts the practice. It concerns what happens to a finding once it exists.
Compare how the equivalent works in a market where one half of this problem, enforceability rather than independence, was solved a century ago. An auditor's opinion matters because a filing cannot be made without one, a regulator will act on a qualified opinion, and holders have a remedy if it turns out to be wrong. The report is load-bearing because the system is built so that it cannot be ignored. The auditor being admirable has nothing to do with it.
Step one has one of those features and none of the others. An evaluator who publishes a finding into a market with no obligation to price it has produced information, not consequence.
The analogy deserves finishing honestly, because it cuts the other way too. The reason auditing is the comparison anyone reaches for is also the reason it is a warning: an auditor is engaged and paid by the company it examines, and the entire modern disclosure regime is built on top of repeated, expensive discoveries of what that does to independence. Anyone who has read an accounting scandal knows that embedding a reviewer inside the firm it reviews is the beginning of the problem rather than the end of it. If the strongest case for step one is that it looks like an audit, the strongest case against it is the same sentence.
Except that this is where the essay's own analogy is better, and conceding the point is preferable to winning it cheaply. An embedded bank supervisor is not paid by the bank. The independence problem just described belongs to the auditing relationship specifically, and the banking model the essay cites is the one institutional design that avoids it. So the objection survives only in this narrower form: as offered today, the arrangement has the auditor's funding structure and none of the supervisor's mandate. If governments do what the essay asks and require it, that changes, and it changes precisely because the relationship stops being contractual.
Which is where the DeepMind objection from the top of this piece arrives, at the same destination. A permanent standards body is less a rival to embedded evaluators than the thing that would give an embedded evaluator's finding somewhere to land. Argued from safety, it is a point about who sets the standard. Argued from disclosure, which is the only way these pages can argue it, it is a point about who is obliged to read the report. Those turn out to be the same institution.
So the honest test is not whether the evaluators get their badges. It is whether their first adverse finding shows up anywhere that anybody has to respond to, and whether it survives having been written by someone whose access depends on the company that granted it.
Both Companies Have Filed
There is one feature of this that makes it different from every previous safety commitment, and we have not seen it discussed.
Anthropic confirmed on its own newsroom, on 1 June 2026, that it had "confidentially submitted a draft registration statement on Form S-1 to the U.S. Securities and Exchange Commission for a proposed initial public offering of our common stock," in a notice published under Rule 135 of the Securities Act. Its own description of what that buys is exact: "This gives us the option to go public after the SEC completes its review." OpenAI has been reported to have filed confidentially too, and its chief executive spent the same weekend discussing when it will list.
That changes what an embedded evaluator is. A confidential draft submission is not yet a registration statement in the sense that matters legally; the liability attaches when a registration statement becomes effective, and neither of these has. But the drafting has started. And once the drafting has started, an independent party publishing adverse findings is producing more than commentary. It is producing facts that somebody's lawyers have to decide whether to put in the risk factors of a document that officers will eventually sign.
So step one operates as more than a safety mechanism. It is a mechanism that generates candidate disclosure about two companies preparing to sell securities to the public. That is the route by which an unmarkable claim could become a marked one, and it is why this part of the proposal matters more than the pacing debate it arrived inside.
Care is needed here: we have not seen either draft registration statement, they are confidential, and their contents are unknown to us. The point is structural, about what an evaluator's report becomes once a company is in registration, rather than a claim about any particular filing.
One Translation Has Already Happened
And it falls where such things always fall.
In an interview published the same day as the essay, asked whether OpenAI still feels pressure to move quickly because of its IPO plans, its chief executive said that given everything happening with safety, "right now would be an ill-advised moment to go public." Pressed on whether 2026 was off the table: "I would say not 2026."
This piece does not claim the essay caused that, and the dates would not support such a claim: the two published the same day. The reason he gives is safety generally, which is his own stated reason and not an inference drawn here.
That is a real consequence, and notice who bears it. A deferred listing is a deferred liquidity event. It falls on the people who already hold the asset: employees with vested stock, early investors, anyone whose exit was supposed to be that listing. It does not fall on the public, because the public is not in yet and cannot price a delay to an event it has no position in.
This is the same shape these pages keep describing from the other end. An IPO is an exit before it is an entry, and the corollary is that everything happening before the IPO, including a safety debate that moves the date, is borne entirely by one side of a transaction that has not happened.
One qualification belongs on that, and it runs against Tessera's own business. The public equity market will price OpenAI's safety posture for the first time on the day it can buy the equity. It does not follow that nobody outside can express a view before then. Instruments offering economic exposure linked to a private company's valuation exist, Tessera's products are among them, issued through dedicated issuing subsidiaries and not available in the US or other restricted territories, and they are disclosed at the top of this piece. What those instruments do not do is produce the thing this whole piece is about. A price formed in a thin market against no filing, no audited statement and no risk factors is a price, but it is not a mark against a document. The distinction being drawn is between a market that can form an opinion and a market that can settle one, and only the second requires the paperwork. That distinction holds regardless of which instruments exist or who is eligible to hold them.
What This Argument Does Not Claim
It holds no view on whether the pace of AI development should be slowed, or on whether the risks described are correctly stated. These pages are not qualified, and the people who are qualified do not agree with each other.
It does not predict that a multi-company agreement happens. Public agreement between chief executives is not an agreement, and a coordinated slowdown between direct competitors raises questions belonging to somebody else's expertise.
It does not call the evaluator proposal inadequate. The proposal is more than anyone has offered before and it costs the company that offered it something real. The point made here concerns what would have to be true for it to bind, rather than whether it was offered in good faith.
And it makes no price claim about anything. Market reactions have been kept out of this piece deliberately: attributing a move in any security to a weekend essay, in a week containing a central bank decision, is a claim these pages could not defend.
What Follows
Watch for the first adverse finding, and then watch what happens to it.
If an embedded evaluator publishes something unwelcome and it appears in a risk factor, this will have been the week the industry built itself an audit function. If it publishes something unwelcome and the only consequence is a news cycle, then what was announced on 12 September was a commitment to produce information, which is not a commitment to be held to it.
The difference between those two outcomes is the whole of it, and it will not be visible for months.
Sources: Dario Amodei, "We Must Pace the Frontier", darioamodei.com, re-read in full on 16 September 2026, for the statement that pacing "does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this"; the description of step one, quoted verbatim, including the naming of METR and the scope covering "not just completed AI models but training pipelines and processes"; the statement that "Anthropic is unilaterally committing to this step now" and the call on governments to require other frontier companies to match; and the banking-supervisor precedent. The fact that the chief executives of OpenAI, xAI and Google DeepMind all responded publicly within a day: reported across Fortune, TechCrunch, Forbes and others, 12 September 2026. Their responses were not identical and this piece does not treat them as such. The characterization of Demis Hassabis's response, posted just under nine hours after the essay, calling its direction correct while saying the details need working through, and favoring a permanent standards body over lab-by-lab embedded-evaluator arrangements, is taken from secondary reporting of 13 September 2026 and has not been read against his own post; it is marked as reported in the body. The essay itself names "the mechanism suggested by Demis Hassabis" as a route for its second step. OpenAI's chief executive's response, that committing to independent evaluators with employee-like access "is a great idea, and we will do the same," is taken from major-outlet reporting including CNBC, 14 September 2026, and is not confirmed here against a company statement; it is marked as reported in the body for that reason. Sam Altman's remarks on listing timing, "right now would be an ill-advised moment to go public" and "I would say not 2026," are from his interview with Fortune published 12 September 2026, read via Fortune and TechCrunch. Anthropic's confidential submission of a draft registration statement on Form S-1 and the publication of that notice under Rule 135 of the Securities Act: Anthropic's own newsroom at anthropic.com/news/confidential-draft-s1-sec, read in full 16 September 2026; the notice and the submission are both dated 1 June 2026. OpenAI's confidential filing is as reported and has not been confirmed against a company statement. Chan Ahn's piece Who Counts the Capex (17 August 2026), as published under his byline, for the announced-versus-incurred framing. No market reaction, incident detail, or figure from any company's internal reporting is used in this piece; where such material exists it has not been independently verified and is deliberately absent. Figures are as of the dates given.
This is market commentary, not investment advice. It is not a recommendation to acquire, hold or redeem any Tessera product, or to take or avoid exposure to any company mentioned.
T-Tokens are tokenized loan participation rights, not equity. High risk. DYOR. Not financial advice. Not available in the US or other restricted territories. tessera.pe/terms
