AI Street

AI Street

Research

Balyasny’s AI Beat Market Odds on Merger Deals

The hedge fund says its fine-tuned GPT-4o model produced more accurate probabilities than a market-implied benchmark and a conventional machine-learning model in a held-out backtest.

Matt Robinson's avatar
Matt Robinson
Sep 16, 2026
∙ Paid

Hey, I’m Matt. I’m a former Bloomberg News reporter, and you’re reading AI Street, where I report on how Wall Street uses AI. On Tuesdays Wednesday this week, I highlight novel research, emerging use cases, and expert interviews.


Over the past few weeks, we’ve looked at how hedge funds are putting their own expertise into AI systems, rather than relying on off-the-shelf models:

D. E. Shaw, Point72 and Ares Hire to Put Investment Know-How Into AI

D. E. Shaw, Point72 and Ares Hire to Put Investment Know-How Into AI

Matt Robinson
·
Aug 25
Read full story
Bridgewater Trains AI to Think Like an Investor

Bridgewater Trains AI to Think Like an Investor

Matt Robinson
·
Jul 2
Read full story

Balyasny Asset Management joins this list. The $38 billion hedge fund fine-tuned a model on merger arbitrage deals, where investors bet on whether announced corporate deals will close, attract a higher bid, or collapse.

In M&A deals, the target company’s shares typically trade below the offer price, reflecting the risk that the deal might fall apart.

Merger-arbitrage traders buy the target’s stock after a deal is announced, sometimes while shorting the acquirer, to capture that gap. They earn it if the deal closes and take losses if it fails.

To judge those odds, analysts read through hundred-page merger agreements, assess regulatory regimes across different jurisdictions, and scrutinize shareholder bases and voting histories.

Balyasny researchers built an AI system that they say can help do that work. In fact, in the July paper outlining their research, they say they deployed it as a decision-support tool for analysts and portfolio managers.

A company spokesperson didn’t respond to requests for comment on whether the hedge fund is still using the model.



How Balyasny Trained the Model

Since there is no readily available dataset of historical M&A deals, Balyasny built one.

The researchers assembled 1,648 public-company deals announced between January 2022 and December 2025. They excluded minority-stake purchases, asset sales, SPAC and smaller deals, such as those under $1 billion. (Where certain deal details were missing, they used AI agents to fill them in and reviewed incomplete information by hand.)

They split the deals by announcement date, using 1,244 deals announced through January 2025 for training and validation. The held-out test set contained 404 deals announced from February to December 2025. Across the full dataset, the system made 4,155 forecasts at different points in a deal’s life, including 1,115 in the test set.

Next, they built 12 AI research agents, each covering one part of a deal. They run in stages:

First

  • Ticker Resolution: Matches the buyer and target to their identifiers across financial databases.

  • Deal Card: Pulls the core deal terms, including structure, key dates, regulatory requirements, termination fees and voting thresholds.

Then, the detailed research

  • Filings Analysis: Reads merger agreements, proxy statements and other SEC filings for red flags, such as litigation risks, and for provisions that protect the deal.

  • Stakeholder Ownership: Identifies major shareholders on both sides, activist involvement and likely voting dynamics.

  • Market View: Reviews expert commentary and market sentiment about the deal.

  • Current Climate: Assesses the wider backdrop, including antitrust climate, interest rates, sector trends and geopolitics.

  • Regulatory Risk: Evaluates antitrust and other approval risks across countries.

  • Financing Risk: Checks whether the buyer can pay, including debt availability, credit ratings and balance-sheet health.

  • Precedent Mergers: Finds similar past deals and pulls their completion rates, timelines and risks.

  • Timeline Events: Lays out past and upcoming milestones, such as filings, shareholder votes and court dates.

Then, a check

  • Gap Analysis: Reviews everything collected so far for missing or overlooked information.

Finally, what’s next

  • Catalyst Tracker: Flags upcoming events that could shift the deal’s odds, such as regulatory decisions, shareholder votes and financing deadlines.

Merger-Arb Specialists Review Agent Work

For the training deals, the researchers built example reports for the model to learn from. First, an AI agent wrote a post-mortem of each deal, covering how it unfolded and which early signals turned out to matter. It was the only part of the system allowed on the open web, since the outcome was already known. Then GPT-5.1 combined the agents’ research with the post-mortem to write a report as of the forecast date. It could cite only information public by that date, but it knew which facts deserved the most weight.

Veteran merger-arb specialists spent months reviewing the agents’ output on historical deals they had covered and flagging gaps. (The paper doesn’t say how many specialists were involved or give a more precise timeline, only that their feedback helped refine the agents’ responsibilities, prompts and output formats.)

The researchers also set the probabilities the model would learn from. Training on final outcomes alone makes forecasts overconfident, especially early in a deal, according to the paper.

They then fine-tuned GPT-4o on those reports and probabilities.

Finally, they tested it. For each test-set forecast date, the fine-tuned model read the agents’ research and produced five forecasts. The system used the median probabilities, while GPT-5 checked the five written reports against source documents and combined them into one.

Results

Most deals end well for shareholders. In the sample, about 14% collapsed without a better offer. So a model could look accurate just by predicting every deal would go through. To prevent that, the researchers gave each collapsed deal more weight in the score than each completed one.

The fine-tuned model beat every method they tested.

This post is for paid subscribers

Already a paid subscriber? Sign in
© 2026 Matt Robinson · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture