A Weekly AI Brief · Every Tuesday

The Signal.

What the AI noise is hiding. One brief, every Tuesday. We cut the hype and show you the numbers that actually matter, named for what they are.

Back Issue · No. 10

The Signal — Issue 10

A supervisory agent tier that improved nothing cost 51.5% more tokens — 43 paired products, 86 runs, one task, one link varied, in a September preprint in which no human rated anything. Tuesday, September 15, 2026.

Current issue · Issue 11 · September 22 · Issue 9 · September 8 · Issue 8 · September 1 · Issue 7 · August 26 · Issue 6 · August 19 · Issue 5 · August 11 · Issue 4 · August 5 · Issue 3 · July 28 · Issue 2 · July 21 · Issue 1 · July 14

Five agents were given a manager who could send work back for revision, with everything else held fixed, and the reports came out no better while burning half again as many tokens. The thing you add to make a capable model safe to use is itself the expense. The same shape turns up in the barriers firms name for not adopting, in the price of a retraining course, and in the one new official number. One item runs the other way, and it is where the argument stops.

This week's signals
  • A supervisory agent tier that improved nothing cost 51.5% more tokens, with one link varied: whether a Manager may reject a worker’s output and oblige a revision. (Loop-Back Authority in LLM Agent Teams, arXiv:2609.14767, a preprint submitted 13 September 2026 · 43 paired products, 86 runs of one business-intelligence task, scored by a five-model judge panel and a deterministic check · no human rated anything and no firm’s P&L was measured.)
  • 23.2% of US employer businesses reported using AI in any business function in the last two weeks, against 22.4% a cycle earlier. 0.28 and 0.36 are the errors on those two levels; the error on the change between them is neither published nor computed here. (Census Bureau, Business Trends and Outlook Survey, cycle 202618, reference period 10–23 August 2026, published 10 September 2026 · one respondent per US employer business, farms excluded, sample about 1.2 million · nine-minute self-report, nothing verified, no threshold of intensity.)
  • Cost was the least-cited of the reasons non-adopters gave — behind the work not lending itself to AI, named by about half of them, and a lack of technical skills, named by roughly a third. Retraining is far and away the commonest response among adopters. (New York Fed, Liberty Street Economics, 1 September 2026; the post states no n · two monthly non-probability surveys of New York-area executives, pools of about 200 and about 150, with about 100 replies to each in a typical month · regional, not a national statistic.)
  • Meetings take 12% of work hours and 14% of firm wage bills, and the firms with the highest pay and revenue spend the most on them — associations, not causal estimates. (Meetings, NBER Working Paper 35706, issue date September 2026 · more than 9,000 workers linked to matched employer–employee administrative data · Norway only, and meetings in general rather than AI coordination.)
  • Salesforce reports having “delivered 7 billion Agentic Work Units”, and SAS with IDC report trustworthy-AI firms “15 times more likely” to report strong or high ROI. Neither is evidence here. (Salesforce newsroom, 11 September 2026, the unit defined and counted by Salesforce · SAS with IDC, 1 September 2026, 2,699 self-reporting decision-makers in 28 countries, on an index the vendor scores.)
  • Not checked this week: Eurostat, the ONS, the OECD, the ECB, the BIS, and arXiv’s cs.AI, cs.HC and cs.SE. Gaps in this issue, not absences in the world. (bls.gov returns HTTP 403 to this machine, confirmed 15 September 2026; the arXiv API rate-limited at HTTP 429.)

Read the first signal against the fourth. Where the supervising layer could not verify anything — specification accuracy was already at ceiling — it added tokens, hedging and revision loops, and the reports written without it scored higher on Utility. Where the layer is a meeting, the firms with the most expensive hours buy the most of it. Same layer, two prices, and the difference is whether it can check anything.

The full brief
The supervisor improved nothing and burned half again as many tokens

A supervisory agent tier that improved nothing cost 51.5% more tokens — five LLM agents with their roles, prompts, tools, models and data all held fixed, and one link varied: whether a Manager agent may reject a worker’s output and oblige a revision. 43 paired products, 86 runs of one business-intelligence reporting task, each report scored by a five-model judge panel and a deterministic specification check. A September arXiv preprint: no human rated anything, and no firm’s P&L was measured. In money rather than tokens the same run costs 20.2% more end to end, $0.606 a report against $0.504.

The flat organisation scored higher, not level: Utility d = 0.42, p = 0.009, and that result survives every robustness check the paper runs. Specification accuracy sat at ceiling in both. The hierarchical reports carry 53% more hedges per thousand words, and each revision loop is associated with a 0.14-point drop in Writing Clarity. The hierarchical Writer’s first draft is statistically indistinguishable from the flat report: the gap opens inside the revision loop.

Population and screen. Forty-three paired report products from five fixed LLM agents on one task class — not firms, not people, not peer reviewed. The Manager is an LLM, the judges are LLMs, and quality is a panel’s score rather than a customer, a revenue line or a production error rate. The five judges agreed on which report of a pair was better, not on how good either was: absolute agreement between them is low, Krippendorff’s alpha 0.22 to 0.30 on Utility. The paper’s own conclusion is conditional and runs here as it states it: a supervisor pays for itself when it can verify and becomes a liability when it can only opine — and with accuracy at ceiling, there was nothing to verify. Source: Loop-Back Authority in LLM Agent Teams, arXiv:2609.14767, submitted 13 September 2026. A preprint.

A new official number landed, and it moved eight-tenths of a point

23.2% of US employer businesses reported using AI in any business function in the last two weeks, against 67.2% that did not and 9.6% that did not know; the cycle before read 22.4%. Reference period 10–23 August 2026, published 10 September 2026, release CB26-TPS.53.

That is a move of eight-tenths of a point. The standard error is 0.28 points on the new figure and 0.36 on the old; those are the errors on each level, and the error on the change between them is a third number the Census does not publish and that we have not computed, because the panels partly overlap. And no test was run.

Population and screen. One respondent per US employer business, farms excluded, from a sample of about 1.2 million businesses in six panels of roughly 200,000, each panel reporting every twelve weeks on a nine-minute self-report form. Nothing is verified and there is no threshold of intensity: one employee using a chatbot once counts the same as a firm running agents across its operations. Forward intent is 27.3%, up from 25.9%, and has run three to four points above realised use in each of the three cycles we checked. Source: U.S. Census Bureau, Business Trends and Outlook Survey, cycle 202618, release CB26-TPS.53.

The barriers firms name are not on the price list

In a Liberty Street Economics post dated 1 September 2026, outside this issue’s window, the New York Fed reported 61% of service firms and 51% of manufacturers using AI in business processes over the past six months — and, among those adopters, a median of 17% of workers at service firms and 7% at manufacturers actually using it. In that same 1 September post, cost was the least-cited of the reasons non-adopters gave — behind the work not lending itself to AI, which about half of them named, and a lack of staff with the technical skills, named by roughly a third, with data privacy and accuracy above it as well. Still in that post: retraining is far and away the commonest response among the firms using AI, and layoffs are rare — the post’s own rounded words rather than percentages, because the chart workbook those percentages come from was parsed by one seat and reopened by nobody.

The post states no n. The two surveys behind it are mailed monthly to pools of about 200 manufacturing executives in New York State and about 150 service-sector executives in the region, and about 100 replies come back to each in a typical month; the August 2026 count is not published. Regional, non-probability, executive, and not a national statistic: 61% cannot be set against the 23.2% above — different frame, geography and respondent.

Retraining has a price, and it comes from a different population. A meta-analysis of 56 randomised trials of US subsidised job training since 1973, posted 7 September 2026, one day outside this window, finds employment up 1.7 percentage points in years 3–5 and pre-tax earnings up roughly $800 a year per person offered training, at $13,598 per participant. That is not AI retraining. It is subsidised training for disadvantaged and displaced workers over five decades, the best available prior rather than a measurement of what firm-funded upskilling returns. Sources: Liberty Street Economics, 1 September 2026; arXiv:2609.07011, submitted 7 September 2026, a preprint.

Some of the layer buys something

An original survey of more than 9,000 workers, linked to matched employer–employee administrative data from Norway, puts meetings at 12% of work hours and 14% of firm wage bills, mostly planning, problem solving, information sharing and project coordination. Firms with higher pay and higher revenue devote more to meetings, not less, despite a much higher opportunity cost per hour, and meeting intensity is positively related to wage growth and to reported on-the-job learning. The limits run with the result: those are associations in observational data, not causal estimates — high-paying firms both meet more and pay more, and selection explains that as easily as knowledge transmission does — and it is one country, with a compressed wage distribution, measuring meetings in general rather than anything to do with AI.

It runs because it locates the boundary of everything above it. If the organisational layer were dead weight, the firms that can least afford an hour of it would buy the least, and they buy the most. What separates a price from a waste is whether the layer can verify anything — and the lead item’s manager had nothing to verify. Source: Meetings, NBER Working Paper 35706, Deming, Løken, Willén and Xu, issue date September 2026 — a month, not a day, so placing it in the week before this one is an inference, labelled as one.

Two numbers that will reach you, and neither is evidence

Salesforce, 11 September 2026: it has “delivered 7 billion Agentic Work Units (AWUs) across Agentforce and Slack, including 3.2 billion in Q2 alone.” An Agentic Work Unit is Salesforce’s own unit, defined and counted by Salesforce, and the page carries no denominator: deployments are given as “thousands”, the customer figures beside them are single names with no base, and there is no error rate, no cost and no P&L. It counts the work the agents did and says nothing about the work the humans did to let them.

SAS with IDC: “organizations investing in trustworthy AI measures are 15 times more likely to report strong or high ROI on their AI projects.” Released 1 September 2026 and recirculated through the trade press between 10 and 15 September, and a recirculation date is not a publication date. The methodology is stated and is not the problem — 2,699 decision-makers, 28 countries, four industries. The multiple is. It sets a group scoring highly on SAS’s own trust index against everyone else, on a self-reported ROI question, published by the company that sells the governance software.

One checked silence, and a list of gaps

One silence here is explained rather than guessed at. The New York Fed’s Liberty Street Economics carries three posts for September and none is new AI work; its own About text says it does not publish during the blackout periods surrounding Federal Open Market Committee meetings, and the September meeting opens on 15 September, the day this issue ships. That is why the item above is dated 1 September.

The rest are gaps, printed as gaps, because the alternative is to let a hole read as a finding. Eurostat, the ONS, the OECD, the ECB and the BIS were not checked at all this week — a decision taken under time pressure — so this issue does not tell you that no European statistic landed, because it does not know. bls.gov answers this machine with HTTP 403, confirmed again today, so no BLS check could be run from here. No Epoch AI result was checked. And on arXiv the API rate-limited at HTTP 429, forcing the slower listing pages: cs.AI, cs.HC and cs.SE were not swept.

Source scan and verification

Census BTOS national data file, cycle 202618. DOWNLOADED AND PARSED, 15 September 2026. Read with openpyxl from the XLSX rather than from the dashboard or a snippet: the response estimates, the response standard errors, and the collection and reference dates. 23.2 / 67.2 / 9.6, standard errors 0.28 and 0.36, reference period 10–23 August 2026 — all off the file. The published response count for the cycle is not stated in the workbook, and neither is any standard error on the change between cycles, which is why the section says that number was not computed rather than printing one. The claim check re-downloaded the same file today and confirmed it is the current cycle, but that seat has no spreadsheet tooling: the cell values above rest on one parse, not two. Cite this one by cycle and release number and never by the bare URL: the file is overwritten every cycle and the address keeps resolving to different numbers.

Census press release CB26-TPS.53. OPENED, 15 September 2026. The release number, the 1.2 million sample, the six panels of about 200,000, the nine-minute form and the exclusion of farms were read off the page body.

arXiv:2609.14767, Loop-Back Authority. ALL EIGHT PAGES OF THE PDF READ, 15 September 2026. The abstract page was opened first and the paper itself was then read page by page. Every figure used above — 43 paired products and 86 runs, d = 0.42 at p = 0.009, the 53% hedging difference per thousand words, the 0.14-point drop per revision loop, the 51.5% token ratio and the 20.2% money ratio — is off that PDF, as are the judge-agreement figures. The paper’s second between-arms quality effect, on writing clarity, is printed nowhere above: its own authors call that one real in direction and fragile under their robustness checks, and printing the single favourable test would have been the defect. The twenty-one pages of supplementary material were not opened. Preprint, not peer reviewed.

NY Fed, Liberty Street Economics, 1 September 2026. OPENED IN FULL, 15 September 2026, AS WERE BOTH SURVEY OVERVIEW PAGES — which is where the response counts come from: about 100 replies to each survey in a typical month, and the August 2026 count is not published. Its chart-data workbook was parsed by one seat and reopened by nobody, so no barrier percentage and no retraining percentage runs above: the post’s own rounded words run instead. The post states its own base in its own prose — “Among businesses that use AI” — re-read at source today. The post states no n, and the pools of about 200 and about 150 executives are the counts mailed, not the counts who reply. Nobody in this house had opened it before today: it reached our own agenda from a machine pulse, and that agenda’s “September” is 1 September.

arXiv:2609.07011, an evidence review of worker retraining. ABSTRACT PAGE OPENED, 15 September 2026. The 56 trials, the 1.7 percentage points, the roughly $800, the $13,598 and the tenfold sector-programme result were taken verbatim. Preprint, not peer reviewed.

NBER Working Paper 35706, Meetings. OPENED, 15 September 2026. The 12% and 14%, the more than 9,000 workers, Norway, and the wage-growth association are off the paper page. The day of posting was not established — only “issue date September 2026” is printed — and placing it in the previous week is an inference from the new-papers flag rather than a date read off the source. Working paper, not peer reviewed.

arXiv:2609.12976, automated insulin delivery. PDF TITLE PAGE, ABSTRACT AND INTRODUCTION READ, 15 September 2026, AND NOT USED. 1,608 adults with type 1 diabetes in four Italian clinics, 283 of whom adopted, 181 retained after matching, and an estimated rise in routine outpatient visits after adoption that the authors say remains sensitive to their trend-extrapolation assumption. No figure from it appears above. It is one disease in four clinics and its screens need more room than this issue had; an item whose limits cannot be printed in full is better left out than compressed. Preprint, not peer reviewed.

Salesforce Agentforce post, 11 September 2026. OPENED, 15 September 2026. The Agentic Work Unit clause is quoted verbatim from the live page. A foil, and evidence for nothing.

SAS and IDC press release, 1 September 2026. OPENED, 15 September 2026. The sentence quoted above is verbatim from the release, and the 2,699 decision-makers, the 28 countries and the four industries were confirmed on the page. The sub-figures circulating this week were NOT chased and are unverified here, which is why none of them appears above.

Oliver Wyman executive survey. DOWNLOADED AND READ IN FULL, 15 September 2026 — the fourteen-page key-findings PDF, including its own methodology note, which describes it as a professionally fielded, self-reported, commissioned survey and not an independent audit of any organisation. No publication date is printed on the fourteen-page PDF and none could be verified; its HTML article page was not opened. No figure from it appears in this issue.

Liberty Street Economics September index and About text. OPENED, 15 September 2026. Three posts in the month, and the blog’s own statement that it does not publish during FOMC blackout periods. The meeting date was checked separately today against the Federal Reserve’s own 2026 calendar: 15–16 September. That is what makes the silence above an explanation rather than a guess.

arXiv:2609.12482 and arXiv:2609.09533. BOTH OPENED BY THE RESEARCH SEAT, 15 September 2026, BOTH HELD. The first defines six conditions for augmentation and carries no figures at all; the second reports on more than 350,000 AI coaching conversations, which is a deployment count and not a sample with a test. No figure from either appears above. Both preprints, neither peer reviewed.

NOT OPENED. gartner.com is bot-walled from this machine and the “1 in 5” was not opened at source. A UiPath orchestration survey circulating on 9 September was status-checked and not opened, and no figure from it runs. mckinsey.com is unreachable from this machine, no McKinsey page or archived capture was opened for this issue, and no McKinsey figure is printed in it. An IDC white paper that reached us through a sweep was taken on trust and is therefore used nowhere.

NOT CHECKED, which is not the same as absent. Eurostat, the ONS, the OECD, the ECB and the BIS were not checked. bls.gov answers this machine with HTTP 403, confirmed again today. No Epoch AI result was checked. arXiv’s cs.AI, cs.HC and cs.SE were not swept, because the API rate-limited at HTTP 429 and the monthly listing pages that replaced it do not scale to those categories’ volume. None of those is a finding about the world; each is a limit on this issue.

Method, and the standing caveat on preprints. One machine sweep was run, for vendor and analyst coverage only, and its returns were treated as leads: it mis-attributed two of its own citations, one of them to a different company’s 2025 release, which is why every figure this issue relies on was opened at its own URL. Both preprints above appeared in no sweep return at all and were found by walking arXiv’s monthly listing pages by hand. Two of the sources carrying figures above are arXiv preprints, labelled as such at every appearance. They are not peer reviewed, their figures can change between versions, and neither measured a firm’s costs. They run because their populations and their tests are stated plainly enough to argue with.

Want a number on your own coordination tax? Measure it right now with the free Agentic Readiness Score, built on Align-ify™.