The bank that Alexander Hamilton founded in 1784 is now asking a question he never had to ask: can a machine be trusted to sign for commerce? We built the temple, but forgot who the god is. BNY Mellon — the oldest bank in the United States, custodian of more than fifty trillion dollars, a chapel of settlement and safekeeping — recently held an internal demo day where the phrase “agentic commerce” was spoken aloud. Agents that do not merely suggest, but execute. Agents that settle, move cash, reconcile positions, and speak to other machines as counterparties, all within the silent, sacred architecture of a system designed for human instructions.
The irony would be funny if it were not so consequential. Hamilton believed in paper, in signature, in the visible hand. The system he helped design was built on the assumption that a human being bears final responsibility for an order. A teller verifies. A compliance officer reviews. A counterparty confirms. BNY Mellon’s demo day did not announce the abolition of that assumption, but it did something more subtle and more significant: it opened the door to letting probabilistic software stand where deterministic authority once stood. We are not talking about a chatbot that summarises client emails. We were always told that Level 4 autonomy for finance was five years away. That was five years ago. The real news is not what the demo day showed. It is that a 250-year-old custodian felt compelled to hold one at all.
I have spent the better part of a decade watching this industry promise to decentralise trust, only to watch it relocate trust into code that no one reads. I audited more than forty ICO white papers during the 2017 mania. I interviewed twelve people who lost their savings in algorithmic stablecoin failures during the summer of 2020. At some point I stopped looking for villains and began looking for the seams in the architecture. BNY Mellon is not a villain. It is an institution. And institutions, like ledgers, remember everything and admit nothing. When an institution of this size starts teaching machines to act commercially, what matters is not the press release. What matters is the seam.
The only authoritative account we have of BNY Mellon’s agentic commerce ambitions comes from a single short industry brief, a handful of signal points pressed into a very thin plate of metal. We know that there was an internal demo day. We know that BNY Mellon frames its strategy around empowering employees to become AI builders. We know that the phrase “agentic commerce” was treated as something that could reshape the labour dynamics of financial services. That is almost no information at all. And yet, because the bank sits at the centre of global custody, because its systems touch nearly every large asset manager, pension fund, and sovereign wealth fund on earth, even a whisper in its corridors is worth interpreting.
So let us do the only useful thing a critic can do with scarce information: reconstruct the plausible architecture underneath the announcement, weight the available evidence, and be honest about what remains unknown.
First, a necessary piece of context for anyone who did not grow up reading bank annual reports. A custodian bank does not lend like a commercial bank or underwrite like an investment bank. It holds assets on behalf of others. It settles trades, collects dividends, processes corporate actions, manages foreign exchange, optimises collateral, and reports balances to clients. Revenue comes from fees tied to assets under custody and from transaction processing. Margins are thin; scale is everything. The competitive battle is fought over basis points of efficiency, over the cost of processing a single instruction, over the speed and accuracy of settlement. This is a business where the difference between a good year and a bad year is operational excellence expressed in decimal places.
BNY Mellon is one of the last great repositories of trust in the Western financial order, alongside State Street and a handful of others. Its size — over fifty trillion dollars in assets under custody and administration — is almost incomprehensible. That number is larger than the GDP of every country on earth except a couple. To place agentic commerce inside that number is to understand why the story matters. If an agent makes an error at a rate of 0.01 percent, and the bank is processing trillions of dollars daily, the absolute magnitude of the error can still reach catastrophic levels. This is the mathematics that keeps compliance officers awake at night.
It is also the mathematics that makes the technology route predictable. No serious custodian bank is going to train its own foundation models. The cost is prohibitive, the talent is scarce, and the core competency of a custodian is not intelligence, it is institutional memory. A far more rational path is a combination-level innovation: take the best available large language models, connect them to the bank’s proprietary knowledge base through retrieval-augmented generation, wrap them in a workflow orchestration layer equipped with function calling, and constrain them with the same risk controls that govern human traders. This is what the phrase “empowering employees to become AI builders” really implies. It is not a claim about artificial general intelligence. It is a claim about tooling. The bank is telling its staff: you will build agents that are tailored to your specific workflows rather than waiting for a vendor to deliver a finished product.
That framing deserves more scrutiny than it has received. “Empowerment” is a beautiful word. It suggests agency, creativity, upward mobility. In the context of a global custodian facing cost pressure, it can also mean something less romantic: everyday employees are being asked to construct the instruments that will absorb their own routine tasks. An agent that reconciles cash positions is not a promotion. It is the beginning of a conversation about whether the person who used to reconcile cash positions is now needed. I do not say this cynically. I say it because I have watched the same pattern repeat in the crypto industry, where every protocol that promised to “empower the community” eventually had to explain away a layoff. We traded soul for speed, and called it progress.
The strategic logic, to be fair, is sound. Custody banking’s revenues are driven by fees attached to a massive stock of assets — BNY’s fee-based income is dominant — and those fees are extremely sensitive to the operating cost required to service each relationship. If an agentic layer can reduce the human cost per transaction by twenty or thirty percent, the effect on the cost-to-income ratio is measurable. BNY Mellon’s cost-to-income ratio has hovered in a range near 65–70 percent in recent years. A hundred basis points of improvement on an expense base of tens of billions of dollars is a meaningful infusion to the bottom line. Wall Street is not going to issue a press release for that. It is going to demonstrate it. The demo day is an internal and external signal that BNY Mellon intends to be on the right side of the automation curve.
But the commercial logic is not the same as the technological reality. Let us look at what an agent actually has to do inside a custodian bank. The mundane workflows are the most agent-compatible: matching settlement instructions, reconciling positions between internal ledgers and external counterparties, monitoring collateral calls, initiating client money movements, drafting regulatory reports, responding to audit evidence requests. Each of these tasks is highly rule-based, heavily logged, and trapped in deterministic system architectures that were designed before machine learning was a discipline. The core systems of a bank like BNY Mellon are not modern microservices. They are mainframe-based transaction platforms built over decades, layered with a palimpsest of patches, middleware, and interfaces that no single engineer fully understands.
This is where the true engineering challenge lies. Connecting a probabilistic agent to a deterministic mainframe is not a prompt-engineering problem. It is an integration problem at the enterprise scale, requiring an internal data layer that governs access, identity, entitlements, audit trails, and reversibility. In my own work helping demonstrate how zero-knowledge proofs could protect AI training-data privacy at three Nordic workshops in 2024, the most difficult part was never the cryptography. It was making the cryptographic abstractions transparent to engineers who were not accustomed to thinking about data provenance. A bank faces the same dilemma but with higher stakes. Based on my experience auditing financial automation architectures, I would estimate that more than half of the total budget for any serious agentic deployment in a custodian bank goes into this plumbing: the middle layer that lets an agent invoke a settlement instruction safely, the identity layer that authenticates the agent’s permissions, the observability layer that records every chain-of-thought token so that a regulator can later reconstruct why the machine acted the way it did.
That last point is the one most outsiders miss. Agentic commerce is not a natural-language problem. It is a jurisdiction problem. Code is law, until the law breaks the code. A human trader who makes an erroneous trade can be disciplined through existing employment contracts and financial regulations. An autonomous agent that initiates an erroneous settlement occupies a strange legal vacuum. In every major jurisdiction, the liability frameworks for autonomous financial action remain underdeveloped, unlitigated, and deeply uncertain. If BNY Mellon’s agent signs a transaction that causes a loss, who is the counterparty supposed to sue? The bank? The software vendor? The employee who “built” the agent as part of the internal empowerment program? Nobody knows. The fact that nobody knows is itself a constraint on how fast this rollout can proceed.
Regulators are not blind to this. The OCC, the Federal Reserve, and the SEC have each signalled, in different accents, that automated decision-making in financial services must be explainable, auditable, and safe. The United States has no comprehensive federal AI law, but the existing risk-management framework for large financial institutions already imposes enough expectations to slow any deployment that cannot produce a transparent reason for every action. In Europe, the AI Act imposes additional obligations on high-risk AI systems, including conformity assessments and human oversight. BNY Mellon is a global bank; its deployments will have to satisfy Brussels, Washington, New York, and London simultaneously. The result is a working environment where agents are likely to be introduced with a human-in-the-loop for quite some time.
Which brings us to the question of autonomy levels. The automotive industry gave us a useful framing: Level 0 is no automation, Level 5 is full autonomy with no human oversight. Most banks talking about agentic AI are really operating between Level 2 and Level 3 — the agent proposes, a human approves, a standing instruction allows for a limited set of automatic operations within well-defined risk limits. The term “agentic commerce” conjures images of machines negotiating with each other without supervision, but the practical near-term reality is narrower. In the first phase, we will see automation of workflows with human-in-the-loop approval. In the second phase, we will see standing instruction, where agents operate within explicitly authorized bounds under the supervision of a risk system. The third phase, full machine-to-machine commerce with no human supervision, is the phase that crypto enthusiasts have been anticipating for a decade.
The irony is that the crypto industry built the rails for phase three years ago. Stablecoins, custodial wallets, smart contracts, and deterministic settlement are exactly the infrastructure that autonomous agents would need to transact with each other trustlessly. BNY Mellon already has a digital-assets division that has been experimenting with tokenized collateral and digital cash. If agentic commerce is eventually married to a stablecoin settlement layer running on a programmable blockchain, the bank — or one of its competitors — could build the first compliant bridge between the old world of custody and the new world of machine actors. That is the faint signal buried in the crypto press coverage of this event. Crypto Briefing, a news outlet with a clear Web3 audience, chose to report on an internal BNY Mellon demo day not because the event was groundbreaking but because it validated a narrative: the traditional financial machine is slowly admitting that the future is machine-payable.
A reporter can only validate narratives when the underlying story is already magnetised. The magnetic charge here is real. If an AI agent needs an identity, a balance, and a way to transfer value, it does not need a bank account in the traditional sense. It needs a public key infrastructure, a settlement token, and a legal framework. That the largest custodian bank in the United States is experimenting with agentic commerce is a signal — however faint — that the traditional industry is taking the machine-economy scenario seriously enough to build internal prototypes. Yet the source material describing this event contains no mention of blockchain, cryptoassets, or stablecoins. That absence is itself an important piece of information. BNY Mellon’s current agentic commerce experiments, at least within the frame of the internal demo day, appear to be contained within the traditional fiat-and-ledger system, not extended into the digital-asset stack. The media interpretation is a projection. It is an authentic signal lost in the noise. The ledgers BNY Mellon will let its agents touch in the near term are its own, not the public chains.
This distinction matters because it reshapes the analysis of competitive threats. Across Wall Street, the AI race has already been running for years. JPMorgan has its COiN system, which famously read hundreds of thousands of hours of legal documents in seconds, and a technology budget in the neighbourhood of seventeen billion dollars. Morgan Stanley deployed an LLM assistant for its financial advisors and became one of the largest purchasers of OpenAI’s APIs in the early GPT-4 era. Citigroup, Wells Fargo, and Capital One have all launched generative AI tools internally. Against those benchmarks, BNY Mellon’s internal demo day is not a display of leadership. It is a display of conformity, a recognition that agentic tools have become table stakes in the efficiency game. The true strategic differentiation is not the technology; it is the vertical depth. A custodian bank that can build agents specific to the arcane workflows of global custody — tri-party repo collateral, securities lending, corporate action elections — will defend its niche better than a general-purpose AI vendor offering a generic chat assistant.
Non-bank competitors are already circling. Payment companies such as Stripe and PayPal have announced agentic toolkits designed to let AI agents send and receive payments on behalf of users. The faster these non-bank rails mature, the more pressure traditional custodians feel to expose their own infrastructure through APIs that agents can call. Yet banks move slowly. The average large bank takes years to move an innovative project from demo to production. This is not necessarily a sign of failure; it is a result of the risk environment. A startup can ship an unstable agent that loses a thousand dollars. A global custodian cannot afford to ship an agent that destabilises counterparty trust in a fifty-trillion-dollar custody franchise. BNY Mellon’s agentic commerce strategy, if it follows the standard trajectory of financial innovation, will be incremental, cautious, and heavily documented.
What, then, are the risks that should keep a thoughtful observer up at night? Let me enumerate them not in the register of a hype cycle, but in the register of a security audit. First, prompt injection. In a conventional chatbot, prompt injection is an annoyance; an attacker can jailbreak the model into saying something embarrassing. In an agent connected to settlement infrastructure, prompt injection becomes a systemic vulnerability. An attacker may not need to break encryption or authenticate fraudulently. It suffices to poison the instructions that the agent processes. If an agent ingests a forged email, a manipulated PDF, or a malicious research report, and that document contains hidden instructions that are executed by the agent’s tool-calling layer, the agent could be induced to move money, alter settlement instructions, or release confidential information. The threat landscape shifts from hacking the system to hacking the model’s context. This is a well-established threat model in AI security research, and it is especially dangerous in financial settings where irreversible transactions are possible.
Second, the problem of autonomy boundaries. Every deployment of agentic AI requires a risk control framework that defines what the agent may do without approval and what it must escalate. The industry has begun to use circuit breakers — analogous to those used in algorithmic trading — that halt an agent’s operations if it approaches a position limit or a loss threshold. But circuit breakers in trading are tested daily by market conditions. Circuit breakers for agents that handle client money are more difficult to calibrate, because the frequency of targeted activity is lower and the consequences of disabled workflows are higher. If the kill switch is too aggressive, the bank will suffer operational inefficiencies that erase the gains from automation. If the kill switch is too permissive, the bank wakes up to a million-dollar error that no human reviewed in time.
Third, the data containment problem. BNY Mellon serves thousands of institutional clients, many of whom are direct competitors with each other. An AI agent that processes data for Client A and Client B must be compartmentalized with the rigidity of a Chinese wall, or the bank will violate confidentiality expectations and likely regulations. Model memory and retrieval augmentation complicate this. If the agent stores a vector embedding of a transaction pattern from Client A, and that embedding is later retrieved in service of Client B’s query, the bank has created a subtle but real form of information leakage. The problem is even more acute in the context of “employee builders.” If employees across the bank are constructing agents autonomously within their business units, absent a centrally governed platform, the institution risks a proliferation of shadow AI: an unknowable inventory of semi-autonomous systems running on unvetted prompts, unsecured tool access, and unaudited data flows. Every CISO I know who works at a large financial institution fears shadow AI more than it fears external attackers, because the attack surface is being created inadvertently from within.
My own perspective on these risks is shaped by a career spent reading the fine print of technological promises. During the DeFi summer of 2020, when I was conducting roughly three months of research at a small Copenhagen-based DAO focused on lending protocols, I watched the industry celebrate mathematical elegance while quietly ignoring the human vulnerability of its own users. I interviewed people whose lives had been upended by oracle failures in algorithmic stablecoin systems that their creators insisted were safe. The gap between the perfection of smart contract logic and the fragility of human circumstance became the central theme of my entire writing practice. When I see a bank embracing agents, I feel the same tension. The agent may be mathematically precise in its intended logic; the environment in which it operates will not be precise. Data is messy. Adversaries are creative. Humans are unpredictable. And the interface between a deterministic machine and a chaotic society is precisely where the most dangerous errors live.
This does not mean the bank should not pursue agentic commerce. It means the bank should pursue it with an epistemic humility that is rare in corporate technology announcements. Large financial institutions have a deep adversarial relationship with innovation. Historically, they absorb new technologies slowly, often a decade or more after those technologies have been proven elsewhere. RPA, or robotic process automation, has been deployed in banking for years with mixed results, a portion of implementations failing after the pilot stage precisely because the underlying processes were not as stable as the RPA tooling assumed. Agentic AI inherits all of those constraints and layers a probabilistic reasoning system on top. The failure modes are therefore not simply those of traditional automation; they include hallucinated tool calls, misleading arguments presented with false confidence, overruns in inference costs, and an immeasurable dependence on the quality of context.
Let me speak directly to the cost dimension, because it is the element that cynical observers love to ignore. Running a serious agentic workflow consumes enormous volumes of tokens. Each tool-call step, each re-versioning of the reasoning loop, each retrieval of context uses compute. For a bank with tens of thousands of employees building and calling agents multiple times an hour, inference costs can accumulate quickly into nine-figure annual estimates. These costs are not theoretical. They are already visible in enterprises that have deployed copilot-style tools at scale and reported cost overruns. BNY Mellon is not going to escape this; it is one of the most scale-sensitive businesses in banking. The commercial return on agentic AI will depend not only on reducing labour costs but also on managing the marginal cost of machine reasoning. This is a less glamorous constraint than the philosophical debate about machine autonomy, but it is the constraint that will determine pacing.
The infrastructure implication is subtle. Financial agents do not need to train large models, which means they do not require bank-owned GPU superclusters in the near term. The primary workload is inference. But financial data sovereignty requires that not all inference happens in public clouds. Certain data may not legally leave the institution’s control, which forces banks into hybrid architectures. Generic, non-sensitive reasoning can run on public cloud APIs; reasoning near the core transaction systems must be local or private. This split-brain architecture creates significant engineering complexity around data leakage. The silent infrastructure competition is not about GPUs. It is about the data middleware layer — the vector databases, the semantic indexing, the access governance, the audit pipeline, and the model-routing framework that allows a bank to send each query to the appropriate model while respecting jurisdiction and data classification. Every major bank is currently quietly rebuilding this layer, and BNY Mellon is certainly among them, without issuing a single press release about it.
There is another phenomenon happening below the public radar: the emergence of agent-to-agent communication standards as a blue-ocean opportunity. If banks, payment companies, and e-commerce platforms all deploy agents, those agents must eventually transact with each other. They need a way to authenticate themselves, a way to present permissions, and a way to settle value. An open protocol for machine actors would be immensely valuable. One can imagine a future where such a protocol is equivalent to what SWIFT became for human banks. The difference is that this new protocol would operate between autonomous software actors, and it might not require a central intermediary at all. The crypto ecosystem has been building decentralized identity frameworks, attestation networks, and machine-payable rails for years. It is not unreasonable to think that banks, as they scale their agent programs, will run into the very problems that decentralized technologies have already attempted to solve. BNY Mellon’s digital-asset division gives the bank a rare opportunity to understand both worlds. Whether it will seize that opportunity is unknown, and anyone who claims certainty is fooling themselves. Faith in the protocol is not faith in the people.
The proper frame for this entire story is probably best expressed as a question: was the demo day a genuine milestone in the history of autonomous commerce, or was it a piece of institutional theatre assembled for investor relations and internal morale? The honest answer is that the available evidence cannot distinguish between the two. We have no details about which departments presented, what level of technical maturity the demonstrations achieved, whether external clients were involved, or when the technology will reach production. Without these data points, the most intellectually defensible position is calibrated skepticism. Demo days in banks are a well-established genre. They have existed since the middle of the 2010s, when large financial firms began imitating the innovation-lab aesthetics of Silicon Valley. The founders of those internal innovation labs often complained that the labs were isolated from the operating core of the institution, that they produced elegant prototypes that never saw production because the risk, compliance, and procurement functions were not engaged early enough. A demo day whose results do not proceed to production within a reasonable period is not innovation; it is ritual.
I believe BNY Mellon wants its program to be more than ritual. The bank has an unusually long history of engagement with digital and tokenized assets relative to its peers, and its leadership has spoken publicly about the convergence of AI and distributed ledger technology. Still, wanting is not doing. The strongest signal to track over the next six to twelve months is employee and hiring data. If BNY Mellon is seriously deploying agentic commerce, its job postings for AI engineers, data engineers, and compliance technologists will accelerate. If those postings stay flat, the demo day is best understood as a flag planted for investors. The second signal to track is authorization for external pilots. A bank that moves agentic commerce from demo day to a supervised pilot with a real client will announce it in some form, because client participation is fundamentally a strategic statement. Third, watch the quarterly earnings calls. Management’s language about AI will shift from future-oriented imagery to concrete references about cost-to-income ratio, operating leverage, and transaction processing cost. That shift from metaphor to measurement is the moment when the story becomes operational.
Let us also mention the uncomfortable labour question, because for all my emphasis on technology and governance, the deepest impact of agentic commerce may be human. Custody banking employs significant middle- and back-office populations across hubs such as New York, Pittsburgh, Dublin, and Manchester. Much of the work is the reconciliation and transfer of instructions — precisely the kind of work that agents can automate quickly. The bank’s narrative about “empowering employees as AI builders” is a strategy designed to transform potential resistance into participation. It is smart organisational change management. But the long-run arithmetic is inescapable. If an AI builder creates an agent that replaces a task previously performed by forty people, the bank will not keep forty people around to build more agents. It will move the survivors into higher-value activities and let attrition handle the remainder. This is the pattern of every major automation wave in modern finance. ATM adoption did not eliminate bank tellers; it transformed the teller role and reduced the need for branch-based transactional labour. Similarly, agentic AI will not necessarily eliminate middle-office roles en masse; it will change what those roles require. The employees who can construct, supervise, and audit agents will prosper. Those who cannot will face considerable pressure.
There is a version of the future where agentic commerce becomes an engine of creative collaboration, where the bank’s 50,000 employees each build small automations that reduce toil and free their time for complex judgment that machines cannot provide. This is the tech-optimist vision, and it deserves fair statement. There is another version, more dangerous and less frequently articulated, where the centralisation of machine agency concentrates power in a handful of financial institutions that control both the models and the rails. If autonomous commerce flourishes only within the walls of regulated custodians, it may never reach the open, permissionless ecosystems that the crypto industry has spent a decade defending. The machine-payable economy would be a walled garden — efficient, compliant, and centrally owned.
As someone who built a career on the conviction that blockchain technology could encode democratic values into immutable logic, I find that centralised version disheartening. But dismissing the bank’s initiative because it does not carry a token would be a category error. BNY Mellon’s entry into agentic commerce, however conservative and incremental, nonetheless normalises the idea that non-human actors can lawfully participate in commerce. Once that idea is normalised inside the most traditional financial institution in America, the semantic ground shifts for everyone else. The demand for machine identity, machine auditability, and machine accountability will grow. And with that demand will come a new architecture of trust in the bedrock of finance.
Truth is not a token you can trade. The BNY Mellon story is not proof that agentic commerce will reshape global finance in the next twelve months. The source material contains a distinct lack of specificity; a rigorous analysis can therefore conclude only that the bank is running experiments and messaging its ambition. The analytical conclusion, based on everything public, is that BNY Mellon is not leading the AI race among its Wall Street peers, that its chosen route is likely a combination of existing models and proprietary workflow integration rather than foundational model research, and that its most important competitive advantage is not AI talent or compute budget but rather the depth of its custody vertical and its unusual openness to digital-asset experimentation.
In other words: the temple is large, the priests are cautious, and the new machine god has not yet been given a set of keys. But the debate over whether it will be given keys is over. The only debates remaining are about the locks, the oversight, and who gets to hold the master key.
The ledger remembers, but the heart forgets. We built this financial system to channel human effort into common progress. If we hand it to agents, we must remember that the ultimate beneficiary — the only being that can feel a loss or celebrate a gain — is still human. Machines do not suffer when a retirement fund loses its value. They merely settle the transaction and log the outcome. The design challenge of agentic commerce, then, is to ensure that the efficiency we gain in the speed of machines does not cost us the dignity that only humans can confer. The code will do what we write. The question is whether we will stay awake long enough to write it well, and to keep writing it better.
I have ended my monthly newsletters during volatile periods with a similar refrain, because I believe it is true: in the machine economy, kindness is a feature, not a bug. BNY Mellon would be wise to build its agents not only for speed and compliance, but for legibility and grace. The agents of the future will execute at the speed of light, but the trust that underpins them will only move at the speed of human understanding. Every great financial institution was built on that understanding. The machines should be built to honour it.
We can now draw the focus to this: over the coming quarters, watch BNY Mellon’s agentic commerce with your eyes open and your credulity in check. Follow the engineering hires. Follow the regulatory guides from Washington, especially any statement from the Fed or OCC about the acceptable parameters for autonomous financial decision-making. Follow the bank’s digital-asset division — if stablecoin settlement and agentic commerce ever appear in the same BNY Mellon announcement, then the transformative scenario described in this article moves from speculative to imminent. And most importantly, ask the question that is too rarely asked in the technology press: who is accountable? When a machine acts, who answers? The answer to that question will determine whether agentic commerce becomes a new chapter in the liberation of human effort or just another automated mechanism for the concentration of wealth and power.
Until then, the temple doors are open. The builders are inside. The silence of the deep ledger is absolute, a vast, unblinking record of everything that has ever moved beneath the surface of finance. Listen closely. For a new kind of signature is being drafted, and it is not a signature that any human will write. The question is whether we will claim it as our own — or find ourselves outsmarted by the very code we created to serve us.


