Independent Research & Evaluation

Effortball publishes independent, human-verified engineering breakdowns. We never accept sponsored ratings, paid ranking boosts, or affiliate manipulation. Read our full testing methodology →

In February I picked up a flyer at an anti-AI march in London. It borrowed (whether intentionally or not) from one of the internet’s best memes.

“Step 1: Grow a digital super mind.
Step 2: ?
Step 3: ?”

The group behind it, Pause AI, finished with a simple demand: pause development until we figure out what the hell Step 2 actually is.

If the reference feels familiar, you can thank South Park. In the 1998 episode “Gnomes,” a band of tiny underpants thieves present their business plan: Phase 1 collect underpants. Phase 2 ? Phase 3 profit. The joke has outlived the episode and become a perfect shorthand for any plan that skips the hard middle part.

Right now that middle part is the entire AI industry.

The Promise and the Blank Space

Companies have completed Step 1. They have built large language models and AI agents that can write code, draft emails, summarize documents, and answer questions with impressive fluency. They promise Step 3 with equal confidence: economic transformation, massive productivity gains, and a future that looks radically different from today.

What sits between those two steps remains largely undefined.

AI optimists tend to wave at the gap. They talk about “economically transformative technology” and assume the path will reveal itself. Skeptics and activists want regulation, safety testing, or outright pauses to fill the blank. Both sides agree that something big is supposed to happen. Neither has a clear, evidence-based map of how we get there.

What the Evidence Actually Shows

Two recent pieces of research highlight how thin the middle still is.

Anthropic published analysis predicting which jobs large language models are most likely to affect. Managers, architects, and media workers appear more exposed. Groundskeepers, construction workers, and hospitality staff appear less so. The study is useful as a first pass, but it rests on matching AI strengths to task descriptions rather than measuring real workplace performance.

A sharper test came from researchers at Mercor. They gave AI agents powered by leading models from OpenAI, Anthropic, and Google DeepMind hundreds of realistic tasks drawn from the daily work of bankers, consultants, and lawyers. The agents failed most of them.

These results do not prove AI is useless. They do show that current systems still struggle when the work involves judgment, messy context, and the need to operate inside existing human processes.

Why the Gap Persists

Several forces keep Step 2 fuzzy.

First, many of the loudest voices have skin in the game. Companies that sell the technology have clear incentives to emphasize speed and potential.

Second, progress in coding tools has been genuinely fast. That success gets over-generalized. Writing software is one kind of work. Making strategic decisions, navigating office politics, or handling ambiguous client needs is another.

Third, real deployment is messy. AI systems do not land in clean laboratories. They land in organizations full of people, legacy software, imperfect data, and established habits. Sometimes the technology improves results. Sometimes it creates new friction. Redesigning workflows around AI takes time, money, and institutional courage that many companies still lack.

The result is an information vacuum. Bold claims rush in to fill it. Markets swing on single social media posts. Public conversation oscillates between breathless optimism and existential dread, often with little solid ground underneath either.

What Would Actually Help

We need fewer projections and more evidence. That requires three things that are still scarce:

Transparency from the companies building the models.
Closer collaboration between independent researchers and the businesses trying to use the tools.
Better evaluation methods that measure performance in real environments rather than artificial benchmarks.

Until those pieces improve, the industry will keep repeating the underpants gnomes’ mistake. Collect the impressive technology. Skip the hard work of figuring out how it actually creates value at scale. Promise the profit.

The technology is real. The potential is real. The missing step is also real. Most organizations are still standing in that blank space, trying to decide what to do with the underpants they have collected.

RM

Rudy Mazer@RudyMazer

Contributing Analyst & AI Researcher

Covers AI productivity tools, computer vision workflows, and tech telemetry. Specializes in hands-on performance benchmarking and deep-dive technical reviews.