Discussion about this post

User's avatar
blake harper's avatar

You’re far too charitable bro.

> “The more defensible forecast is that AI will make frontier research teams much more productive and may dramatically increase the number of experiments they can plan, code, and analyze”

I’m sorry but I’m deep in this stuff and I have not seen evidence for that. Experiments are expensive, they can’t just run them Willy-nilly on low-credibility guesses. Coding an experiment is rarely the long tail in R&D. OpenAI even admitted in the GPT 5.5 system card that the model is terrible when evaluated against their own internal RSI benchmark.

Their whole IPO depends upon convincing the street that RSI is imminent because it’s the only way they maintain pricing power in a world where open source just commoditizes it all. That’s why Jack has to parrot this stuff about RSI b/c from what I hear their S-1 prep is not looking good.

For those who want an even more skeptical take that looks at the automation opportunity within each phase of the model development lifecycle, I wrote about that here: https://tailwindthinking.substack.com/p/the-ai-bubbles-favorite-fairy-tale

I know it’s fashionable to assign determinate probabilities to this stuff, but I truly don’t understand how this is anything but that. It’s not epistemically responsible to assign low probabilities to events that involve compounding uncertainties and assumptions. Better to just say “we don’t yet have reason to believe this is possible in any determinate kind of way” and leave it at that. Assigning a determinate possibility means you have some reason to believe the obstacles will be cleared — and if you do, you owe us that account of precisely how it would be for that 1 in 10 scenario.

Elvin Haley's avatar

In the beginning of 2025, the best model on the market scored 0% on [FrontierMath tier 4](https://epoch.ai/benchmarks/frontiermath-tier-4-v2). If you asked me back then when I thought the benchmark would be saturated, I'd say 2030 or so. But the benchmark is saturated already.

So the fact that there are benchmarks where AI scores essentially 0 doesn't convince me that AI won't be able to do these tasks perfectly a year or two from now. But I'm not saying they _will_ be. I guess we will see...

5 more comments...

No posts

Ready for more?