Hey, I’m Matt. I’m a former Bloomberg News reporter, and you’re reading AI Street, where I report on how Wall Street uses AI. On Tuesdays, I highlight novel research, emerging use cases, and expert interviews.
Wall Street’s Compute Arms Race Rolls on
Despite the high cost of Nvidia’s GPUs, with its most advanced chips estimated at about $55,000 apiece, trading firms can’t get enough of them.
Ron Minsky, co-head of Jane Street’s tech group, on a company podcast in May:
“These days we are in something like the range of like tens of thousands of GPUs and we will in not too long be in the range of hundreds of thousands of GPUs. And we think it’s like well justified by the business.”
Iain Dunning, head of AI at Hudson River Trading, on Bloomberg’s Odd Lots podcast in June:
“One of my greatest failures has been… predicting how many GPUs we would need,” Dunning said. “You’re constantly playing catch up.”
Álvaro Cartea, director of the Oxford-Man Institute of Quantitative Finance, said the firms would not be spending billions unless they expected the investment to pay off.
“However many billions these guys spend, it’s because they know that it’s very likely they’ll recover them,” Cartea told me. “These billions are a bet that the odds are on their side.”
A new STAC research report helps explain in part why trading firms want dedicated access to GPUs rather than relying entirely on external APIs: for some workloads, it can dramatically reduce the inference time involved in building and testing quantitative models.
NOTE FROM OUR SPONSOR
The Best Dataset Isn’t For Sale — Yet
Months before Google agreed to pay $10 million for Spirit Airlines’ data, Brickroad had already flagged it.
The company’s AI agents scan news, earnings reports, app-store rankings and changes to terms of service for signs that a company is producing valuable data — and might be willing to license it. When a source looks promising, Brickroad contacts the company, secures a sample and helps the buyer evaluate the data before committing to a deal.
Benchmarking Agents in Quant Workflows
STAC, an industry group that develops and runs technology benchmarks for financial firms, tested a workload based on a real-world quant research session supplied by a large US market maker.
A Real-World Quant Research Workload
The original session involved an AI coding agent developing and evaluating short-term trading models using Bitcoin market data.
“We worked with one of the big quant trading firms to use a real-life example from their agentic quant research sessions,” James Corcoran, STAC’s head of AI, said in an email. “So we know this is representative of what some in the industry are doing.”
It was the first time STAC had benchmarked inference infrastructure for agentic workflows, Corcoran said.
STAC replayed 300 requests from the session using DeepSeek-V4-Pro, the same model used in the original session.
It ran the requests on a dedicated Lambda cluster containing 16 Nvidia B200 GPUs, then compared the results with DeepSeek’s public API.
Faster on Dedicated GPUs, but More Expensive
When STAC ran a single 300-request research session, the dedicated cluster completed the replay in about 25 minutes. The same requests took roughly 72 to 75 minutes through DeepSeek’s API.
“With this use-case, speed is the name of the game: how quickly can an agent build and test a quant trading model,” Corcoran said. “Time to task completion is extremely important.”
STAC then ran 8, 32 and 64 sessions simultaneously. The speed advantage narrowed and reversed as more sessions shared the fixed cluster.
At 64 simultaneous sessions, the dedicated cluster took about 93 minutes per session, while the API remained around 72 to 75 minutes.
At off-peak rates, the API was also cheaper per task. The dedicated setup offered greater speed for a single session, but at a higher cost.
One caveat: The test focused on how quickly the systems generated the agent’s responses. It excluded the time needed to execute code and train trading models.
This reminded me of what Matthew Dixon, co-author of Machine Learning in Finance, told me in May:
Brute force gets you a long way in finance. Most applications aren’t Rolls-Royce engineering problems. They’re bread-and-butter: get me a number, good enough, ballpark. AI fits that very well.
More speed means firms can discard weak ideas and pursue promising ones while competitors are still waiting for the same research cycle to finish.
Coming up for paid subscribers: More of my conversation with Oxford-Man’s Cartea on why compute is becoming the new latency race, why firms want models carrying their own “genetic code,” and whether billion-dollar infrastructure costs will leave only a handful of firms able to compete.







