AI Street

AI Street

Interviews

Jump Trading’s Lucas Baker on AI Agents

In an interview, Baker explains why smarter models should test fewer ideas, organized fleets beat swarms and LLM backtests are vulnerable to leakage.

Matt Robinson's avatar
Matt Robinson
Sep 22, 2026
∙ Paid

Hey, I’m Matt. I’m a former Bloomberg News reporter, and you’re reading AI Street, where I report on how Wall Street uses AI. On Tuesdays, I highlight novel research, emerging use cases, and expert interviews.


The technical divide between quant firms and frontier AI labs has all but vanished, according to Lucas Baker, head of LLM R&D at Jump Trading.

Both work with the same basic ingredients: models, compute, data and infrastructure. Both recruit similar talent from the same pool of computer science, mathematics and physics majors. You’ll often see them sponsoring the same major machine learning conferences, like OpenAI and Jane Street did for ICML 2026.

Lucas has worked in both worlds. Prior to Jump, he spent time at Google DeepMind as a software engineer, helping build the evaluation framework for AlphaGo Zero, the system that learned the game without using data from human games.

In our conversation below, Lucas explains why he expects successful quant firms of the future to resemble “frontier labs with trading arms attached.”

We also discuss:

  • How agents are automating quantitative research—and why smarter models should test fewer—not more—ideas.

  • Why an organized “fleet” of agents is more productive than an unstructured “swarm”

  • The difficulty of separating real signals from leakage, p-hacking and spurious correlations

  • Why compute, data and infrastructure are becoming prerequisites for competing at the highest level

  • Where specialized financial models still make sense as general-purpose models become more capable

NOTE FROM OUR SPONSOR

The Best Dataset Isn’t For Sale — Yet

Months before Google agreed to pay $10 million for Spirit Airlines’ data, Brickroad had already flagged it.

The company’s AI agents scan news, earnings reports, app-store rankings and changes to terms of service for signs that a company is producing valuable data — and might be willing to license it. When a source looks promising, Brickroad contacts the company, secures a sample and helps the buyer evaluate the data before committing to a deal.


This interview has been edited for length and clarity.

Where Quant Finance Is Heading

Matt: What will quant work look like in a year or two?

Lucas: My favorite answer that I’ve heard for this is that the successful quant trading firms within two years will look like frontier labs with trading arms attached.

That is certainly the vision I’m aiming for.

Matt: People increasingly move between quantitative trading firms and frontier AI labs. How similar have the required skills become?

Lucas: The flow goes both ways, but the skill sets have converged almost completely.

This is not only because the fundamental techniques involved are similar and the base tech is the same, but also because the presence of intelligent agents allows a specialist to extend into adjacent domains faster than before.

If you are world-class at kernel optimization but you don’t know quant trading, AI can give you a faster on-ramp to draw on the adjacencies that come with the infrastructure, compute, and data of quant trading firms and help you become more broadly capable. Rather than progress being evaluated by abstract benchmarks or debates, quant trading grounds you in the numbers themselves, so you always know whether you’ve actually discovered or proven something.

Matt: How open is the AI research community compared with quantitative finance?

Lucas: I think this is particular to AI in that the techniques are generally shared quickly, and the greater advantage comes from understanding them, capitalizing on them and connecting with the researchers who know the right hyperparameters of the research process itself.

Traditionally, quant finance is a very closed and rival domain where the mere knowledge of what you’re working on can be sufficient to leak some of your alpha and reveal strategies. This is still true, but practically everybody is taking agents and attempting to use them to enhance their own systems as much as possible. We’re happy to build our capabilities internally as well as to work with external parties.

How Agents Are Changing Research

Matt: What has surprised you most about the development of AI agents?

Lucas: I did not anticipate that they would be this good this soon.

Nonetheless, if you’d asked me three years ago, I would have guessed that we would still be integrating capabilities more than seeing the smooth continuation of every ability curve at once. The generality is, I think, even more stunning than insiders would have predicted.

Matt: Firms are trying to capture their institutional knowledge and train models to think like their analysts. As everyone gains access to similar tools, does the advantage come from incorporating how a particular firm thinks?

Lucas: There’s certainly an aspect of that. I think there are different concentrations and different levels of awareness.

In terms of distinction, you can look at the fundamentals, discretionary, or quant systematic side. If you’re on the fundamentals side, the goal is to model someone’s investment process, and that includes the knowledge and the theories in their head. You want to know how they think about markets, how they would assess a given situation, and apply that whole workflow. It is possible to interview someone and get a very detailed idea of how they operate.

If you take it on the other side, one of the largest changes even in the past six months, let alone a year, would be the emergence of autoresearch. I mean that in the sense of the original Karpathy autoresearch, where you take a model, set it on a problem and say, “Here’s what you’re allowed to modify, and here’s the thing that you’re trying to optimize. Go.” You leave it running overnight and see what it has produced.

Of course, the problem is much harder than that in quant finance because you have longer-running jobs over more data. They’re more finicky, the signal-to-noise ratio is lower, and it’s quite possible to reward hack along many dimensions that don’t necessarily show up in the first report. You have to be careful to analyze the underlying factors.

It’s a much harder problem than simply setting the model to optimize a given repo. Nonetheless, it is possible, and the agents are now smart, general, nuanced and capable enough to handle it. That wasn’t necessarily the case six months ago.

To put it simply, I’d say for eight to 12 months, agents have been able to pretty reliably make numbers go up. What you want is to make meaningful numbers go up in the correct way.

Matt: Does that mean you can test more ideas?

Lucas: Not necessarily. In fact, I would say the contrary. The smarter the model gets, the fewer ideas you should be testing. This is because the smarter the model is, the shorter a path it should take to the correct answer.

This post is for paid subscribers

Already a paid subscriber? Sign in
© 2026 Matt Robinson · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture