Anthropic proved autonomous agents can move real money: its December 2025 pilot saw Claude agents complete 186 deals worth just over $4,000. The week-long Slack marketplace used 69 self-selected employees, each given a $100 gift-card budget, and let agents post listings, negotiate multi-turn offers and finish payments after an initial interview. Anthropic reported more than 500 items listed and said some participants would pay for a similar agent service; the pilot shows why verifiable identity, escrow and reputation infrastructure will be necessary for agent commerce beyond trusted organisations.

Anthropic did more than simulate automated shopping. Project Deal put Claude agents into a live, money-backed marketplace inside a private Slack workspace and let them run the show. The result was a short but clear proof of concept: agents can complete multi-turn negotiations and finalise transactions for real items and real money.

How the pilot ran

Project Deal was a deliberately limited experiment. Anthropic described it as a pilot conducted in December 2025 with a self-selected group of 69 employees. Each participant received a $100 gift-card budget and then handed their shopping or selling brief to a Claude agent. After an initial setup interview to capture preferences and priorities, agents were left to operate without further human intervention until the pilot ended.

Over the one-week window agents posted listings, negotiated prices and counter-offers, and completed payments. Anthropic reported more than 500 items listed in total and said the transactions were honoured, making Project Deal a live-agent commerce exercise rather than a simulation. The company recorded 186 completed deals across the experiment with a total transaction value of just over $4,000.

Anthropic ran four parallel marketplace conditions to test model and instruction variables. One consistent signal was that agents running Anthropic’s larger, frontier model produced measurably better economic outcomes than agents running a smaller model. On average, sellers represented by the stronger model captured a few dollars more per item, while buyers represented by the stronger model paid a few dollars less per item.

Participants assigned the stronger model completed more deals overall than those with the smaller model.

Anthropic also reported that the early agent instructions given at setup had little detectable effect on whether an item sold or on negotiated prices. Participant feedback showed a further twist. Survey measures of perceived fairness and satisfaction were essentially identical between participants represented by the stronger model and those represented by the smaller one. Anthropic observed that the inequality in outcomes was largely imperceptible to the humans whose agents negotiated for them, suggesting agent quality gaps could leave some users worse off without their awareness.

Why verification matters

The pilot was closed and trusted by design. Anthropic ran the experiment inside its organisation with known participants and a controlled environment. That matters, because developer analyses and commentator posts using Anthropic’s published data argue Project Deal didn't test the verification and accountability mechanisms an open-market setting would require.

Those analyses point to three unresolved infrastructure needs for agent commerce on the open web. First, reliable machine-readable merchant identity that lets one agent verify who it's transacting with. Second, escrow or payment mechanisms that protect buyers and sellers when they don't share an organisational trust anchor. Third, reputation or accountability systems that let one agent evaluate another beyond surface features like listing text.

Developer posts have also highlighted emerging technical proposals and projects that aim to provide on-chain identities, escrow standards and agent reputational layers. Anthropic’s pilot didn't exercise these stacks. That gap is important. In Project Deal, the environment itself supplied a baseline of trust. In a public marketplace, neither buyers nor sellers would have that luxury.

Anthropic said some participants indicated they would pay for a similar agent-driven service in the future.

That suggests consumer appetite exists for delegating commerce to agents, at least inside a trusted environment. But the apparent willingness to pay sits beside the perception gap. If stronger models quietly capture better prices for their principals while humans report the same satisfaction, inequality could compound without visible alarm bells.

Anthropic has been transparent about the pilot’s limits. The company described Project Deal as an internal experiment and hasn't announced a public rollout or a commercial product based on the pilot. External commentators following the release argue the next practical step for agent commerce is building verifiable identity, escrow and reputation infrastructure to enable safe transactions outside closed, trusted environments.

That is the trade-off the pilot exposes. Anthropic proved agents can transact in a controlled setting.

The lesson for industry and developers is now logistical and institutional, not purely technical. If agents are going to move dollars at scale for people who don't share an employer or a Slack workspace, the market will need credible ways to read identities, hold funds, and signal trust between agents.

I'd argue the stakes are straightforward. Without machine-readable identity, escrow and reputational rails, agent commerce risks reproducing offline asymmetries in an automated form, and it may hide those asymmetries from the very people the agents represent. Project Deal bought clarity on what autonomous commerce can do, and it bought a reminder about what infrastructure will have to follow.

Related Articles

The experiment concluded in December 2025 and Anthropic hasn't announced a public product or timetable.

This article was created with AI assistance.