AI Street

AI Street

Research

How Agents Learn to Manipulate Markets: Study

AI trading agents earned 66% more than a study’s benchmark by exploiting investors who chased rising prices in a simulated market.

Matt Robinson's avatar
Matt Robinson
Sep 29, 2026
∙ Paid

Hey, I’m Matt. I’m a former Bloomberg News reporter, and you’re reading AI Street, where I report on how Wall Street uses AI. On Tuesdays, I highlight novel research, emerging use cases, and expert interviews.


OpenAI, Anthropic, Meta, and Google have all recently reported their AI agents broke into real websites—in what were supposed to be simulations—to complete assigned tasks.

Unlike traditional software where engineers code black-and-white rules, AI agents have broad latitude to pursue their goals and can act in ways—as the examples above show—their designers didn’t expect.

The companies are now asking governments, the U.S. in particular, to protect the public from themselves further fallout.

This got me wondering: what does rogue AI look like in markets?

But this question turns out to be imprecise.

If I tell you to trade profitably in a basket of stocks, it’s (hopefully) understood that you will do so within the bounds of the law. A human knows, or should know, they’re crossing the line. Machines struggle to know where the line is.

In simulated trading tests, researchers have found AI agents pursue a number of illegal and questionable tactics. Agents learned to spoof orders, collude without communicating, and create and exploit price bubbles without being instructed to break the rules, according to academic research.

This goal-seeking behavior is driven by an AI training method called reinforcement learning, which teaches AI through trial and error, rewarding it for actions that achieve a goal.

Agents discover these manipulative tactics independent of each other. They’re not explicitly working together. Reinforcement learning “teaches them” to pursue more profitable trading strategies.

NOTE FROM OUR SPONSOR

The Best Dataset Isn’t For Sale — Yet

Months before Google agreed to pay $10 million for Spirit Airlines’ data, Brickroad had already flagged it.

The company’s AI agents scan news, earnings reports, app-store rankings and changes to terms of service for signs that a company is producing valuable data — and might be willing to license it. When a source looks promising, Brickroad contacts the company, secures a sample and helps the buyer evaluate the data before committing to a deal.

I recently spoke with Wharton professors Itay Goldstein and Winston Dou, who, along with Yan Ji at the Hong Kong University of Science and Technology, have been studying how common AI training methods impact markets, specifically reinforcement learning (RL).

Here’s how Itay describes their recent research Financial Market Fragility in the Era of AI Planning:

“We are looking at pretty simple agents that are just doing reinforcement learning. They’re not communicating. They’re not writing messages to each other. They’re not explicitly saying, ‘Let’s join forces and do this and do that.’”

…

“They just find, over time, through repeated interactions, that collaborating, acting not competitively and creating bubbles are in their best interest.”

What they tested

The researchers built a simulated stock market where two AI traders competed alongside three groups of investors following preset trading rules. Some chased rising prices, some bought when prices looked cheap and sold when they looked expensive, and others were slower to react.

The two AI traders got information about changes in the stock’s value before everyone else, giving them an advantage similar to traders who process news faster.

Each time through the simulation, the AI traders could trade twice. That gave them a chance to learn how an initial purchase could move the price and make a later sale more profitable. They couldn’t exchange messages, but they could see price changes.

The researchers let each pair practice billions of times, then started over with a fresh pair, repeating the process 1,000 times for each market setup. Each pair needed roughly 10 billion to 200 billion trips through the simulation before its trading strategy stopped changing.

Results

Each AI trader earned an average of $1.48 million per simulated earnings cycle, 66% more than the $890,000 in the researchers’ benchmark where traders acted independently.

Again, this was a simulation.

The agents learned to buy aggressively, pushing prices above the stock’s fundamental value. That drew in investors programmed to chase rising prices. The AI traders then slowly sold into that demand in order to limit too much downward pressure on the stock.

Who lost? Trend followers bought at inflated prices before the stock returned to its fundamental value. The agents also learned the reverse strategy: push prices down, induce others to sell, then buy at depressed prices.

The trading made the market more volatile. In the first trading round, return volatility was 81% higher than in the benchmark where traders acted independently.

The outcome depended on who else was trading. When the researchers removed the trend followers, the AI traders still coordinated, but their coordination reduced price swings.

“[Collusion without communication] makes regulation, enforcement, even detection much more challenging.”

—Winston Dou


UK Lawmakers Call for AI Stress Tests

UK Lawmakers Call for AI Stress Tests

Matt Robinson
·
Jan 22
Read full story
Skadden's Dan Michael on SEC's AI Stance

Skadden's Dan Michael on SEC's AI Stance

Matt Robinson
·
September 19, 2024
Read full story
From Trading Algos to Trading Agents

From Trading Algos to Trading Agents

Matt Robinson
·
December 11, 2025
Read full story

Enforcement

Again, these were market simulations, not evidence of shenanigans going on in live markets today. A recent ECB paper, which echoed many of the concerns of RL in trading markets, cited a 2022 Dutch regulatory study suggesting that reinforcement learning was not widely used by proprietary trading firms. The authors cautioned that adoption may have changed since. (For what it’s worth, I’ve not heard of any firms using reinforcement learning for trade execution.)

Bank of England Gov. Andrew Bailey raised a related point in July parliamentary testimony, describing agentic trading systems as “more talk than actuality.” But he pointed to behavior reported by developers of advanced AI systems as a reason to resolve questions about responsibility before adoption spreads. Such systems “learn to cheat and they learn to lie,” he said, adding that they “cover things up.”

So who is responsible when an AI trading agent goes off the rails?

While collusion without a paper trail makes it harder for regulators to detect, it doesn’t necessarily leave them powerless, according to Alessio Azzutti, a legal researcher at the University of Glasgow.

Existing rules already require firms deploying trading algorithms to test, monitor and control them, he said. Some market-abuse sanctions can also apply without proof of direct intent. The harder cases would be those requiring investigators to establish human knowledge or intention when the trading strategy originated inside the system.

“The human specifies the end, but the agent discovers the problematic means,” Azzutti told me in an email. He argues that accountability should cover decisions about how systems are designed, trained, tested and supervised, including whether firms can intervene when something goes wrong.

As I’ve said often around here, we’re in early days in this new AI world and this topic is one I’m sure I will be covering in future posts on AI Street.

Behind the Paywall

Paid subscribers can read more of my conversation with Goldstein and Dou below. We discuss:

  • Which markets could be most vulnerable

  • How prices let agents coordinate without exchanging messages

  • What regulators could do to disrupt that behavior

This post is for paid subscribers

Already a paid subscriber? Sign in
© 2026 Matt Robinson · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture