The 79 Percent Nobody Is Selling: What 1,642 Agent Traces Reveal About Where Reliability Budgets Go Wrong

CryptoWoo โ€ข โ€ข Prediction Markets

Forty-four point two percent. It is not a hallucination rate, a jailbreak statistic, or a parameter count dressed up as a milestone. It is the share of multi-agent system failures that a 1,642-trace annotated dataset classified as FC1 โ€” system design problems. Another 34.4 percent landed in FC2, inter-agent mismatch. Stack the two and roughly 78.6 percent of everything that went wrong was decided before a single token was generated.

Tracing the ghost in the gas receipts taught me this before agents existed. In late 2017, I spent six weeks dissecting ERC-20 contract logic for a Riyadh venture fund. The vulnerabilities I surfaced were never runtime artifacts. They were specification artifacts โ€” a missing guard, an unstated invariant, a modifier that assumed something its author never wrote down. Three projects. An estimated $4.2 million in investor losses prevented. Not one of those bugs would have been stopped by bolting a firewall onto a deployed contract.

That is the shape of the argument now circling agent reliability. And nearly every dollar is going into the firewall.

The 79 Percent Nobody Is Selling: What 1,642 Agent Traces Reveal About Where Reliability Budgets Go Wrong

Context: the taxonomy is the case

The MAST taxonomy draws from seven mainstream multi-agent frameworks and 1,642 annotated traces, sorting failures into three classes: FC1, system design problems; FC2, inter-agent mismatch; and FC3, task verification. FC1 carries 44.2 percent of the load. FC2 carries 34.4 percent. Inside FC1 the recurring culprits are step repetition, failure to recognize termination conditions, and non-compliance with task specification. Inside FC2, the pattern is reasoning-action mismatch and outright task derailment.

Read that list again and notice what is absent. There is no jailbreak in it. No prompt injection, no data exfiltration, no credential theft. Every one of those failure modes is an orchestration defect โ€” a role definition that was never written, a termination condition nobody specified, a verification step that was assumed rather than built.

The 79 Percent Nobody Is Selling: What 1,642 Agent Traces Reveal About Where Reliability Budgets Go Wrong

Now look at where the industry's protocol stack has converged. MCP for tool invocation. OWASP ACS and the NIST AI Agent Standards for framework legitimacy. OAuth 2.0 and SPIFFE/SPIRE for identity and authorization. Broadcom AgentMinder for identity and intent binding. Microsoft MXC for policy-driven OS-level isolation and sandboxed execution. MCP sits under the AAIF umbrella and, by March 2026, was reportedly logging 97 million monthly SDK downloads.

Hold that download number. I will come back to it, because it is the weakest link in this entire narrative.

One evidentiary note before we go further, and I make it because forensic habits are cheap and reputations are not: the 2026 vendor releases, the AAIF governance arrangement, and the download figures are not independently corroborated in this analysis. Treat them as scenario inputs. The taxonomy is the part with a dataset behind it.

And here is why a crypto audience should care about a failure taxonomy built on model traces rather than block space. Capital is rotating. Funds that spent 2021 pricing L1 throughput and 2023 pricing rollup sequencer revenue are now pricing agent infrastructure โ€” tool-calling protocols, agent wallets, orchestration frameworks. The same instinct that underwrote dozens of Layer 2s on the theory that more chains meant more users is now underwriting an agent stack on the theory that more agents mean more autonomy. The taxonomy above is the first real measurement of whether that theory survives contact with production.

Core: what the traces actually describe

Here is the part that should worry anyone allocating into agent infrastructure this quarter. The failure distribution and the investment distribution point in opposite directions.

Consider step repetition. An agent loops on a subtask it already completed because nothing in its role definition told it what "done" looks like. I watched an almost identical pathology in 2020, when I deployed $50,000 of ETH across Uniswap V2 and SushiSwap to measure impermanent loss against pool volume in real time. The failure there was never the swap math. It was the exit condition โ€” the point at which a rational liquidity provider stops chasing emissions and accepts the divergence loss. Every dashboard I built that summer was technically correct. The strategy around it was underspecified, and the underspecification cost more than the slippage ever did.

Now scale that pathology to a system of twelve agents. One loops. Another notices the loop and re-plans around it, burning context. A third interprets the re-plan as a new objective and drifts. Two more inherit the drift and execute against a goal nobody authorized. The ensemble does not fail because any single model is weak. It fails because the contract between the agents was never written.

FC1 failures are the Solidity reentrancy bugs of this cycle โ€” and we already know how that movie ends. In 2017 the industry response to reentrancy was not better runtime monitoring. It was checks-effects-interactions, a design pattern. We changed how contracts were authored, because no amount of post-deployment vigilance substitutes for an invariant stated at authoring time. The MAST intervention data points the same direction: improving role specifications measurably reduced FC1 failures. The underlying work does not publish effect sizes, task domains, or model versions โ€” a real gap โ€” but the direction is consistent with everything I have seen in audit work.

FC3, task verification, is the quiet class, and it is where the closed loop should live. Runtime telemetry tells you an agent failed. It rarely tells you why in a form a specification author can act on. Without that feedback path, design-time specs get written once, drift, and rot. The vendors best positioned to build that loop are precisely the runtime governance companies โ€” which is the structural problem for anyone planning an independent specification-engineering startup. Your most natural distribution partner is also your most natural acquirer.

Notice the shape of this. In Layer 2 we watched dozens of teams slice a fixed pool of users into ever-smaller fragments and call it scaling. In agent governance we are watching dozens of vendors slice a fixed pool of enterprise security budget into runtime control planes and call it reliability. Both moves are defensible individually. Both leave the underlying constraint untouched. The constraint here is not enforcement. It is authoring.

The signature is in the silent transfer โ€” the action an agent takes that no dashboard flags because it was never declared out of bounds. Runtime policy engines enforce only what someone encoded. Every unencoded prohibition is an open door with the security camera pointed at the wrong hallway.

So why is the money chasing runtime? Hunting liquidity where the charts lie taught me to read flows rather than narratives, and the flow here is unambiguous. Runtime governance maps cleanly onto line items enterprises already own: identity, authorization, sandboxing, audit evidence, observability. It is procurable. It is demonstrable in a compliance review. It produces a dashboard a CISO can show a board. Broadcom can sell AgentMinder through VMware's private-cloud channel. Microsoft can bind MXC to Windows and Azure. Both are extensions of existing platforms, which means the marginal cost of distribution is close to zero.

Specification engineering has none of that. There is no SOC 2 control called "role definition quality." There is no renewal trigger for "we wrote better task specs." It is a methodology gap wearing the costume of a product category.

I have written about the human cost of technical failure before. When Celsius froze withdrawals in June 2022, I spent that summer collecting retail investors' accounts alongside on-chain tracking of the 6,000 BTC treasury movement, because a balance sheet alone does not explain what a collapse does to a family. The same discipline applies here. Step repetition and task derailment are not merely compute waste. They are an agent that keeps trading after the risk limit breached, keeps filing after the deadline passed, keeps escalating after the incident was resolved. The cost is not measured in tokens. It is measured in the actions nobody specified as forbidden.

The 79 Percent Nobody Is Selling: What 1,642 Agent Traces Reveal About Where Reliability Budgets Go Wrong

Contrarian: the 79 percent is a correlation, not a confession

Here is where I break from the consensus this analysis is building toward, and I want to be precise about it.

Seventy-nine percent is a number from one study. It is a correlation derived from a specific sample โ€” seven frameworks, 1,642 traces, one annotation schema. Change the task domain, swap in a stronger base model, or loosen the labeling rules and the distribution moves. FC1 and FC2 may also overlap; a single root cause โ€” an ambiguous role definition that produces both a reasoning-action mismatch and a derailment โ€” could be double-counted into two buckets and inflate the total. Nobody in this narrative has published the overlap analysis. Until someone does, the honest statement is: in this sample, specification-shaped failures dominate. Not: specification-shaped failures dominate everywhere.

There is a second blind spot, and it is the one I would bet on. Better base models will silently absorb some FC1 failures. Step repetition and missed termination conditions are exactly the kind of thing that scaling and stronger instruction-following erode over time. If that happens, a meaningful chunk of the specification-engineering thesis evaporates โ€” not because the diagnosis was wrong, but because it was made at a moment when models were weaker than they were about to become. I have watched this pattern before: an entire tooling category justified by a model deficiency that two releases later was gone.

And a third: specification engineering without adversarial testing is documentation theater. If the spec lives in a wiki and no red team ever tries to hijack it through a poisoned tool description, you have produced a compliance artifact, not a reliability improvement.

There is a specific failure I keep returning to, because it is the one runtime tooling actively conceals. A sandboxed agent with a verified identity, a bound intent token, and full audit logging can still execute a task whose goal was misstated at authoring time. Every control passes. The incident report will read clean. That is not a security posture โ€” it is a compliance posture, and the two have never been the same thing. The riskiest system in production right now is not the unsafe one. It is the secure one executing a task nobody defined correctly: sandboxed, fully audited, and completely wrong.

Following the money through the validator maze, I keep landing on the same read. The investment skew is not irrational. Runtime governance is being purchased because it is purchasable, and the design-time layer is underfunded because nobody has figured out how to invoice for it. The analysis itself, self-referencing its own prior series, may be doing double duty as category creation. That does not make it wrong. It makes it promotional, and promotional arguments deserve the same forensic scrutiny as promotional charts.

Takeaway

Watch four signals over the next two quarters, and watch them as a set. First, whether MAST-style failure distributions replicate across model families โ€” if FC1 shrinks as models improve, the thesis weakens fast. Second, the actual governance structure of AAIF, because whoever controls the MCP center controls the tool-calling layer. Third, whether AgentMinder and MXC ship with published pricing and named customers, or remain launch-deck vapor. Fourth, whether any regulator accepts a role definition or a verification log as compliance evidence โ€” the moment that happens, specification engineering stops being a methodology and starts being a product.

MCP's 97 million monthly downloads are not a revenue figure. I spent three months in 2024 correlating ETF inflows with exchange reserves, and the first rule I learned was that gross flows lie unless you net them against custody. The same discipline applies here. Downloads include CI pipelines, mirror pulls, and dependency auto-resolution. Adoption is not usage, and usage is not payment.

The question worth sitting with is not whether specification engineering deserves to exist. It is whether an industry that has spent three years building fences will notice that the failures happen before anyone arrives at the gate.

Market Prices

BTC Bitcoin
$83,034.6 +0.07%
ETH Ethereum
$2,509.92 +0.77%
SOL Solana
$110.57 +0.81%
BNB BNB Chain
$751.3 +1.51%
XRP XRP Ledger
$1.41 +1.84%
DOGE Dogecoin
$0.0862 +1.89%
ADA Cardano
$0.2551 +7.41%
AVAX Avalanche
$10.53 +3.32%
DOT Polkadot
$1.26 +7.16%
LINK Chainlink
$13.14 +2.50%

Fear & Greed

64

Greed

Market Sentiment

Event Calendar

{{ๅนดไปฝ}}
18
03
unlock Sui Token Unlock

Team and early investor shares released

15
04
halving Bitcoin Halving

Block reward reduced to 3.125 BTC

28
03
unlock Arbitrum Token Unlock

92 million ARB released

30
04
upgrade Celestia Mainnet Upgrade

Improves data availability sampling efficiency

12
05
halving BCH Halving

Block reward halving event

22
03
unlock Optimism Unlock

Circulating supply increases by about 2%

10
05
upgrade Ethereum Pectra Upgrade

Raises validator limit and account abstraction

08
04
upgrade Solana Firedancer

Independent validator client goes live on mainnet

Market Cap

All โ†’
1
Bitcoin
BTC
$83,034.6
1
Ethereum
ETH
$2,509.92
1
Solana
SOL
$110.57
1
BNB Chain
BNB
$751.3
1
XRP Ledger
XRP
$1.41
1
Dogecoin
DOGE
$0.0862
1
Cardano
ADA
$0.2551
1
Avalanche
AVAX
$10.53
1
Polkadot
DOT
$1.26
1
Chainlink
LINK
$13.14

Tools

All โ†’

Altseason Index

41

Bitcoin Season

BTC Dominance Altseason

Gas Tracker

Ethereum 28 Gwei
BNB Chain 3 Gwei
Polygon 42 Gwei
Arbitrum 0.5 Gwei
Optimism 0.3 Gwei

๐Ÿ‹ Whale Tracker

๐ŸŸข
0xf63b...2098
6h ago
In
2,970 ETH
๐Ÿ”ต
0x5c70...e137
1h ago
Stake
1,487,750 DOGE
๐Ÿ”ต
0x9cfe...3d4b
3h ago
Stake
4,025.70 BTC

๐Ÿ’ก Smart Money

0x74e2...e9c2
Top DeFi Miner
+$3.4M
77%
0x4be6...73fa
Institutional Custody
+$4.4M
70%
0x9835...5b20
Top DeFi Miner
+$0.6M
91%