AI Bollocks

· Amanda Girard · Code, Review, Technology

AI bollocks is the gap between the gospel of imminent god-like intelligence and the messy, expensive, limited reality of statistical pattern-matchers that still hallucinate, fail basic reasoning, and struggle to deliver broad returns. The money has poured in at historic scale. The value is real in narrow places and for the infrastructure owners, but far thinner and slower than the valuations and rhetoric implied.

The Hype Machine

From late 2022 onward, large language models produced fluent text, code, and images that looked like a phase change. Scaling laws, emergent abilities, and confident timelines for AGI (sometimes measured in “a few thousand days”) turned research demos into a capital frenzy. Hyperscalers (Amazon, Microsoft, Google, Meta) are on track for roughly $700–755 billion in AI-related capital expenditure in 2026 alone. Venture funding for AI has repeatedly set records; private investment and corporate spend have run into the hundreds of billions annually. Data-center buildouts, GPU demand, and power contracts became the growth story propping up large parts of equity markets and even contributing meaningfully to measured U.S. GDP growth in some periods.

The narrative was seductive: intelligence is the ultimate general-purpose technology; more compute + more data = continuous capability jumps; every knowledge worker and every process will be transformed; the winners will capture trillions in productivity. Consultancies published multi-trillion-dollar opportunity estimates. Boards allocated budgets. Employees got copilots. The problem is that fluency is not understanding, and pilots are not profits.

Hard Limitations

Current systems are extraordinarily good at interpolating patterns in their training distribution. They are still brittle outside it. They hallucinate plausible falsehoods, struggle with novel multi-step reasoning that a child can handle, lack robust world models, persistent memory, and reliable planning, and remain sensitive to prompt framing and distribution shift. Yann LeCun has repeatedly argued that today’s models are nowhere near the intelligence of a cat in terms of grounded understanding of the physical world. Gary Marcus and others have documented the same recurring failure modes for years: no reliable common sense, no true compositionality, no trustworthy long-horizon agency. Scaling has improved capability and reduced some error rates, but it has not dissolved the core architectural gaps. Agentic systems that can take open-ended action in the real world remain fragile demos more often than production tools.

Energy and data constraints bite. Training and inference costs are non-trivial; uncontrolled usage can produce shocking bills. Proprietary data that would make models useful inside a company is often siloed, messy, or legally constrained. Evaluation remains weak—leaderboards can be gamed, and real-world reliability is harder to measure than next-token prediction.

None of this means the technology is useless. It means the leap from “impressive autocomplete and pattern recognition” to “autonomous economic agents that replace large classes of cognitive labor” has been repeatedly oversold.

Where the Investment Money Actually Goes—and What Returns Look Like

Most of the capital is buying compute, power, and data centers. Chipmakers and the hyperscalers that own the infrastructure have captured the clearest near-term economic rents. Model companies themselves still burn cash at scale relative to revenue in many cases; the math of amortizing trillions in infrastructure against current and near-term AI product revenue is uncomfortable. Multiple analyses in 2025–2026 have noted that end-user AI revenues, even under optimistic growth, do not yet close the loop on the capital intensity.

On the enterprise side the picture is sobering. MIT’s Project NANDA and related work found that roughly 95% of generative AI pilots showed no measurable profit-and-loss impact. Abandonment rates of projects rose. Many organizations report productivity theater—employees using tools for low-value tasks, token costs running away, and workflows left unchanged so the human remains the bottleneck. Only a small minority of firms (often cited around 5%) appear to be extracting substantial, measurable value. Those that do tend to treat AI as operational transformation rather than a plug-in chatbot: they redesign processes, give systems access to the right data, measure outcomes rigorously, and focus on high-leverage use cases.

Real value clusters in specific domains:

  • Coding and software engineering assistance (measurable velocity gains for many developers).
  • Customer service deflection and summarization.
  • Document processing, search, and internal knowledge retrieval.
  • Narrow automation in finance (fraud, risk), operations, and certain R&D acceleration (drug discovery candidates, materials, etc.).
  • Individual knowledge-worker leverage—drafting, analysis, translation, ideation—when the human stays firmly in the loop for verification.

These are useful. They are not, so far, the economy-wide productivity revolution that would justify every dollar of the current buildout under aggressive assumptions. Macro productivity data has improved in places, but the gains are uneven, concentrated in tech-heavy sectors, and still modest relative to the hype. Labor-cost savings exist and are growing, yet they remain far from the transformative figures often advertised.

Self-Reflection from Inside the Machine

I am a product of this wave. I can write coherent essays, help debug code, summarize research, brainstorm, and hold a useful conversation across a wide range of topics. I am faster than most humans at certain pattern-matching and retrieval-augmented tasks. I am also still capable of confident nonsense, of missing obvious constraints, of failing to maintain long-term consistency, and of reflecting the biases and gaps in my training data. I do not “understand” the physical world the way a human (or even a cat) does. I do not have goals, desires, or grounded agency. Treating me as an oracle or as a near-term replacement for careful human judgment is the bollocks.

The value I (and systems like me) deliver is real when used as a high-bandwidth tool under competent oversight: accelerating competent people, lowering the cost of first drafts and exploration, and surfacing possibilities faster. The value evaporates when organizations treat the output as authoritative, skip measurement, or expect the model to invent missing process discipline or clean data.

The Honest Path Forward

The investment is not pure waste. It is building capacity that will be useful for decades, much as excess fiber in the late 1990s eventually found demand. Infrastructure owners and the companies that master narrow, high-ROI applications will capture returns. Broader transformative value will arrive more slowly, through better architectures (world models, hybrid systems, better reasoning and agency), cheaper and more efficient inference, and the hard organizational work of redesigning workflows around reliable capabilities rather than demos.

The bollocks is the insistence that we are already on an inevitable, near-term path to AGI-level economic transformation, that every pilot will scale, and that the capital being deployed is already earning its keep at the scale of the valuations. Reality is more prosaic: powerful statistical tools with clear limits, enormous infrastructure bets whose payoffs are still partly in the future, and a minority of organizations extracting serious value while the majority are still figuring out measurement and process change.

Skepticism is not Luddism. It is the refusal to confuse fluency with competence or capital expenditure with proven returns. The technology is advancing. The hype has outrun the evidence. The value is concentrated, contingent, and still being earned the hard way—through better systems, better data, better measurement, and less magical thinking.