AI Street

AI Street

Interviews

How AI Runs $200 Million in Portfolios

An interview with Alejandro Lopez-Lira on AI-generated stock portfolios and what he's learned running them.

Matt Robinson's avatar
Matt Robinson
Jul 28, 2026
∙ Paid

Hey, I’m Matt. I’m a former Bloomberg News reporter, and you’re reading AI Street, where I report on how Wall Street uses AI.


Alejandro Lopez-Lira has been running one of the largest public experiments in AI investing. Nearly 52,000 investors have allocated about $200 million on Autopilot across seven public portfolios built with AI models including ChatGPT, Claude, DeepSeek and Grok.

These portfolios aren’t responding to a simple “What stocks should I buy?” prompt. Instead, they break investing into a series of steps—macro analysis, company research, portfolio construction and verification—with different AI agents handling each stage.

So far, that approach has produced stronger results than many public AI investing experiments. The DeepSeek portfolio, for example, has gained 56% over the past year, compared with a 16% gain for the S&P 500. Lopez-Lira cautions that it can take a decade to know whether a diversified strategy is outperforming the market.

I recently caught up with Lopez-Lira, an associate professor of finance at the University of Florida, to see how his thinking on AI and investing has evolved since he became AI Street's first interviewee two years ago. Back then, Lopez-Lira argued that AI could already handle many of the tasks typically assigned to an intern. Today, he believes AI systems have reached roughly the level of a fourth-year Ph.D. student.

We discussed what's changed over the past two years, how he builds AI systems that break investing into specialized tasks, and why he believes the next wave of AI on Wall Street will be driven by systems rather than standalone models.

This interview has been edited for clarity and length.

How capable have AI systems become over the past two years?

I would say large language model systems are at the level of a fourth-year Ph.D. student in every field.

I don’t think there is any specific task that does not require human-to-human interaction or doing things in the physical world that can generally be done better by humans than by AI.

Jobs are collections of tasks, sometimes well-defined and sometimes not. That determines how much substitutability there is between humans and AI systems.

I say AI systems because LLMs by themselves haven’t changed that much. But embedded in systems such as Claude Code, it’s insane what you can do. I’ve been mostly working on autonomous systems that can run for a long time without human intervention.


ICYMI

Alejandro Lopez-Lira on AI in Financial Forecasting

Alejandro Lopez-Lira on AI in Financial Forecasting

Matt Robinson
·
July 10, 2024
Read full story
Morgan Stanley's Ex-AI Head on Scaling AI Beyond Pilots

Morgan Stanley's Ex-AI Head on Scaling AI Beyond Pilots

Matt Robinson
·
Mar 25
Read full story
Former Millennium-Backed Traders Start AI Commodities Hedge Fund

Former Millennium-Backed Traders Start AI Commodities Hedge Fund

Matt Robinson
·
Jun 10
Read full story

The Claude Portfolio

One example is your Claude portfolio. How does it differ from the ChatGPT, DeepSeek and Grok portfolios?

The other models receive a list of the best potential stocks from a large scoring system. Grok decides how to make its portfolio, GPT decides how to make its portfolio and so on.

Claude is more iterative. It can launch an agent for each stock to check that the research is sound and that there are no issues. You can also have a portfolio-optimization agent assemble the portfolio and another agent audit it to make sure there’s nothing weird.

It involves many more cycles and revisions, calls more agents and potentially trades more frequently. The other portfolios trade once a month. Claude decides when to trade, with a limit of twice per week. It trades less often than that on average.

How has it performed?

It was launched about four months ago. It was up about 10% as of yesterday, roughly the same as the S&P 500.

Four months is way too little time to tell. It will take a couple more years to evaluate.

The ChatGPT portfolio has outperformed the S&P 500, but it trails Grok and DeepSeek. What explains the difference?

I think ChatGPT is more opinionated. DeepSeek says, “These seem like the stocks with the highest expected returns. I’m just going to make a bet on that.” That seems to work.

ChatGPT says, “I’m going to consider what’s happening with the macro environment and other things.” That may not necessarily be correct. Grok is somewhere in between.

The system produces very good information, and the models use it more or less effectively. But it’s still too early to draw conclusions about any of them.

How much money is invested across the portfolios?

There is $200 million across all portfolios now.

Some experiments that let LLMs trade have produced large losses and volatile results. Why haven’t your portfolios experienced the same swings?

This post is for paid subscribers

Already a paid subscriber? Sign in
© 2026 Matt Robinson · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture