I love this video. It makes me laugh every time. I also think it captures what current models are good at and bad at better than any benchmark.
Ask an AI to make you money in the stock market, or to run your fantasy team in a sport you don’t follow super closely, or to build you a profitable business by tomorrow morning. What comes back will sound smart. It will be confident and fluent, and it will lose. The reason it loses is worth walking through, because it makes us better AI users, and it also points to where the real moats in AI startups are.
You can only delegate what you can evaluate
The prompt is the first step. After the prompt, an expert user makes hundreds of micro-corrections along the way. An expert can tell output that is 80 percent right from output that is confidently wrong. A novice sees the same fluent paragraph in both cases. Remove the evaluator and the model settles into a plausible-sounding average of everything it has read.
Is coding different?
Sure, people who cannot code ship real apps now. Andrej Karpathy named vibe coding in February 2025, and it works because it’s verifiable and compilers and tests grade the work. The app runs, or it crashes. The error message tells you what to paste back into the chat. The user can’t judge the code, and it doesn’t matter. Delegation works when something can grade the output. Your own judgment can do the grading, or something else. You need one of them in the loop.
Why the intelligence is jagged
This pattern comes straight from how the models were trained. Many of the sharpest recent gains in reasoning came from reinforcement learning on verifiable rewards. Let the model attempt a problem a million times and reward it whenever the answer checks out. Math comes with answer keys, and code gets graded by compilers and test suites. That recipe took frontier models from mediocre competition mathematicians to IMO gold medalists in about two years.
No equivalent flywheel exists for writing a great strategy memo, making a good hire, or picking a stock. Nobody has a cheap, automatic checker for those, so nobody could train against them at scale. The closest thing those domains got was training on human thumbs-up, and a rater who rewards whatever sounds right since they can’t check for correctness (they’re not experts in that particular field). Ethan Mollick and his co-authors called the result the jagged frontier back in 2023: AI that beats experts on some tasks and stumbles on neighbouring ones. Since then, the models have soared wherever cheap verification existed and remained uneven elsewhere.
Markets add one more problem. A model trained on the internet’s text is a consensus machine, and whatever everyone already believes about a stock is baked into its price. Agreeing with everyone earns you nothing.
Confident mediocrity
On tasks with no answer key, the models are usually decent. They write a passable memo, a reasonable stock thesis, a defensible fantasy lineup. And they deliver everything with the same assurance they bring to gold-medal math. I call this confident mediocrity. Confident mediocrity is exactly what loses in the stock market and in fantasy leagues, because zero-sum games grade you against players with real edge. The model cannot tell you when it is wrong, and a non-expert user cannot tell either. That leaves two blind parties at the table and no referee.
Next frontier
Labs are working on this, and some of their current research attempts to manufacture answer keys where none exist: rubric-based rewards, models grading other models, supervision of reasoning steps instead of final answers. The third AGI test from ARC tests for some of this. Labs will show us whether judgment (or taste) can be turned into a reward signal at all. The thing to watch is whether training on these manufactured keys starts producing jumps like the math results.
What this says about AI today
We have superhuman intelligence for the subset of reality that comes with an answer key. AGI is the rest of it. My personal primitive AGI benchmark comes from the video at the top. The day an AI wins me a fantasy league in lacrosse, we’re there. (cuz, you know, I don’t follow lacrosse - not sure who does).
The so what for startups
Lack of judgment or a grader in the loop produces AI slop machines. Exhibit A: content marketing generation tools ;)
Products built for experts can leave the grading to the user. A lawyer can review a draft contract. A developer can inspect generated code. An investor can challenge a stock thesis. These products can be very useful, but they still depend on expertise outside the software. And their market size is capped by the number of experts. Some expert markets are big enough; many are not. And you’ll have to look for a moat in something other than the proprietary answer key.
Sometimes the world provides the answer key. The code compiles. The payment arrives. The parcel reaches the customer. These checks make the product easier to use, although every competitor has access to the same answer key. Again, no moat there.
In other cases, the company has to build its own answer key. That might include proprietary evaluations, a domain-specific verification system, proprietary datasets or a human review process that generates better feedback over time. Or all or some of it.
That third answer is the only one the company owns.
— Raif


