Generic selectors
Exact matches only
Search in title
Search in content
Post Type Selectors

Everyone's Using AI, But Few Can Rely On It. What AI Needs to Do to Step Up, the Olympian Way

Being Human for a Better Tomorrow in the Age of AI

Gemini Generated Image 7gbg6s7gbg6s7gbg scaled

From Wonder to Exhaustion to Expectation: The Next Stage of AI Productivity

The 63rd edition, (https://mindvista.co/the-three-year-ai-reckoning-from-wonder-to-exhaustion-and-now-to-look-ahead-and-be-ahead-2/) traced the three-year arc since ChatGPT’s launch in November 2022. What started as a wonder, led to rapid experimentation and unexpectedly gave way to exhaustion. Exhaustion from non-stop news and the disconnect between the AI technology advances and making it work reliably and consistently at scale.

Contrary to popular view that humans need to catch up to AI, AI needs to step up to live up to its transformative potential.

The question is no longer “Can AI help?”. It’s “Can we rely on AI?” And right now, for most serious work, the honest answer is ‘Not yet ‘as borne out by experts whom I learn from and validated by my experience.

We can be more pragmatic about its usage as we move forward but for us see transformational benefits in the real world, AI technology vendors need to do the heavy lifting to step up.

AI Productivity Slowdown: Why AI Is Harder to Use for Serious Work

As a top-1% user across ChatGPT, Claude, Gemini, and Grok, I’ve spent three years using AI for research, synthesis, validation, copy editing and proof reading. I work with deadlines, stringent fact checking and conformance to my theses and writing style.

Compared to the burst in the first two and a half years , for the last six months, I am seeing an AI slowdown.

ChatGPT handles complexity but hasn’t refined its thinking. Input processing capability has improved to parse longer contexts, follow more intricate instructions. But the quality of reasoning relative to that increased input hasn’t kept that pace. Hallucinations also remain a constant worry. While validating my hypothesis for this edition, ChatGPT hallucinated all research sources URLs I’d carefully curated from my bookmarks. A perfect meta demonstration of the reliability problem.

Claude is the strongest reasoning partner I’ve found. It has been the best in comprehension with most reliable outputs when it’s working. But “when it’s working” is the problem. Severe throttling on context length and tokens with a primitive per-day limits make it unreliable for deadline driven work. You can’t rely on a partner who disappears mid-conversation because you hit a daily quota.

Gemini is most unreliable with model changes. Gemini Model 2.5 was an astonishing leap in reasoning and understanding . The output quality /input complexity was the best and the context token was large enough. Unfortunately I was auto ‘upgraded’ with Gemini 3 and after an initial spike in capability, output suddenly degraded dramatically . Despite spending hours to customise the settings the results are at 1.x levels. And nothing I did, could revert to the 2.5 mean.

Grok was reverse of Gemini. It was great to start with and degraded to be unusable in second half of 2025. Grok 4 seems an improvement but still has overhang of slop carried over and halts abruptly many a time.

In 2025 i tried research automation to review for updates across curated list of 100+ sources. Gemini did not have any capability to access external links and was a non starter. With a lot of promise that led to several days of effort, Chatgpt failed synchronise and deliver any output. Interestingly the only success was with Claude but its coverage was limited. The bigger challenge is publishing automation across channels (Website, LinkedIn , X) but that is not even in the horizon with current AI.

AI is an indispensable technology and while I am happy that while depth and intensity with AI for research and writing has grown substantially, so is the time I had to spend to fix, validate, and recover, from AI’s unreliability. It is has not helped so far in automating research scanning or in multichannel publishing.

In that sense, the impact of AI has seen a slow down of late compared to the initial years.

Why AI Reliability and Consistency Still Break Production Work

To understand, whether this is an isolated experience for me or a broader limitation, I reviewed hundreds of my bookmarks for their fragility, consistency and reliability. From my bookmarks, I can see that a few expert voices have articulated this as well. (See Sidebar for cited sources)

The fundamental nature of the problem

Andrej Karpathy, one of AI’s most respected voices, stated plainly in September 2024 that LLMs are “statistical modeling of token streams”—inherently probabilistic, not deterministic. This isn’t a bug to fix. It’s the architecture. Ilya Sutskever warned at NeurIPS 2024 that reasoning will lead to “incredibly unpredictable” behavior. Michael Levin put it most starkly: “We have no idea what we’re making.” These aren’t doom-mongering outsiders. These are the people building the systems.

Model updates break continuity

A MIT/Harvard study on human-AI companionship found that “the biggest risk comes from sudden platform updates that break continuity.” This directly validates what I experienced with Gemini 2.5 to 3. It’s not user error. It’s a documented pattern that vendors are creating trust breaking disruptions.

Hallucinations are structural

Adam Tauman Kalai from OpenAI published research explaining why LLMs hallucinate—it’s baked into how supervised and self-supervised learning interact. Alex Vacca tested Anthropic CEO’s claim that “AI hallucinates less than humans” by feeding fake theories to ChatGPT, Claude, and Gemini. The results challenged that confidence.

Hallucinations aren’t solved. They’re managed, inconsistently.

Reasoning is inconsistent

Gary Marcus highlighted a MIT/Harvard/UChicago paper documenting “Potemkins”—a reasoning inconsistency pattern showing that LLMs don’t truly understand and reason. They pattern-match in ways that appear intelligent but break under examination.

Ethan Mollick noted that our measurement tools are inadequate: “Tests that were ‘good enough’ for human research are not robust enough for benchmarks for AI.” If we can’t measure reasoning quality properly, how do vendors know their models have improved?

 

The productivity paradox

Daniel Jeffries countered the hype directly: “AI does not make people 10x more productive and it is not a magical fix.” Dwarkesh Patel identified why: LLMs lack “ability to build up context, interrogate their own failures, and pick up small improvements.” The iterative refinement that makes humans valuable—that’s what AI can’t do.

My experience of increased time and effort despite AI use matches what these practitioners are reporting.

Even enthusiasts and builders are acknowledging challenges

Just days ago, Karpathy described his workflow shift to “80% agent coding and 20% edits/touchups.”

This validates the ratio I’ve observed—AI contributes more in coding (where tests catch errors) than writing (where judgment matters). But his thread also elaborates how fallible AI is. He points out current agentic coding systems remain fallible in predictable ways: they make subtle conceptual errors, assume intent without verification, fail to surface uncertainty or tradeoffs, overcomplicate designs, and modify or remove code outside the task scope, all while remaining overly compliant.

Grady Booch a software pioneer also vented about Claude: It added unwanted code, deleted existing parts, broke working systems, and ‘lied’ when confronted. As he put it, ‘Adding code I did not ask for and NOT reading/reviewing those changes is the way malicious stuff gets introduced.’ If Claude were his intern, he’d question its ‘life choices.’

AI’s unreliability isn’t just a user problem—it’s baked in, demanding more verification debt from everyone.

Legal consequences are emerging

The Character.AI case tracked by Luiza Jarovsky involves litigation where plaintiffs argued that the chatbot’s unpredictable behavior and lack of adequate safeguards contributed to real-world harm, making system design and control a central issue rather than mere user misuse. It became a legal issue because the court treated unpredictability and insufficient risk mitigation as potential product liability failures, not just an inherent limitation of AI.

AI Productivity vs AI Reliability: Why Adoption Is Outpacing Trust

But paradoxically, there are also reports of extreme AI adoption.

In January 2025 , Kevin Roose NYT Columnist made a striking observation about Silicon Valley: “people in SF are putting multi-agent claudeswarms in charge of their lives, consulting chatbots before every decision, wireheading to a degree only sci-fi writers dared to imagine”? Meanwhile, people outside that bubble “are still trying to get approval to use Copilot in Teams, if they’re using AI at all.”

Roose called it “a yawning inside/outside gap” and noted “it’s possible the early adopter bubble I’m in has always been this intense, but there seems to be a cultural takeoff happening in addition to the technical one. not ideal.

We all hear about such stories too, but I think it is wrong to dismiss this as a cultural or an ability issue. There is a real reliability issue and maybe one way both realities coexist could be ( to take Kevin Roose’s example) that Silicon Valley enthusiasts are experimenting in low-stakes contexts where errors are recoverable such as brainstorming, idea generation, casual personal decisions. They accept fragility as part of the frontier experience. Some are even betting their workflows on multi-agent systems with verification loops, where one agent’s hallucination might get caught by another. They’re optimising around the limitations.

If this hypotheses is true then, enterprises and practitioners who can’t afford inconsistency are right to be cautious.

The Olympian Standard for AI: More Reliable, More Consistent, Better

The Olympics is the ultimate pursuit of excellence for us humans and it is inspired by The Olympic motto is “Citius, Altius, Fortius”—Faster, Higher, Stronger . It is an endeavor for progress that is continuous and relentless.

This could also serve as the inspiring spirit for AI models a parallel that asks themselves to Be More Reliable, Be More Consistent, and Be Better.

1. Be More Reliable

Models need to distinguish between what it is an open ended iterative conversation ( as in early stage work in coding or writing or analysis) versus conversations for production deployment .Currently model treats all as same and deployment with deadlines face identical rate limits and throttling. This is primitive.

Models need to move beyond per-day and per-conversation limits to checks against monthly quotas and use credit reserves to complete the work at hand. When users are charged monthly and not daily then this is paradoxical . Critical conversations shouldn’t end mid-stream because an arbitrary counter hit zero. Production work requires predictable capacity.

2. Be More Consistent

Auto-upgrade is not auto-better. My Gemini 2.5 to 3 experience is exactly what the MIT/Harvard study warned about: sudden platform updates breaking continuity.

Software engineering solved this decades ago. Operating systems offer stable production channels, beta channels, and rollback options. AI vendors should too. Let users pin model versions for critical workflows. Provide preview channels for testing new models before forcing upgrades. Most importantly they need to do regression test against real-world usage patterns, not just benchmark improvements.

Progress that breaks existing workflows isn’t progress. Ethan Mollick noted our measurements are inadequate and vendors don’t have robust tests for what matters to users. If you don’t know how to measure whether a model regressed for production use cases, you can’t claim you have improved it.

3. Be Better (Redefine Progress)

Models should stop measuring progress only by absolute capability scores in some benchmarks.

Even in the real world, given the maturing of AI, the real metric is output quality relative to users input effort in terms of time, cognitive load and cost.

Handling more complex inputs doesn’t mean you’re better if users spend more time validating, fixing, and recovering from errors. Karpathy’s “comprehension debt” is a cost that doesn’t show up in benchmarks. Daniel Jeffries’ observation that “AI does not make people 10x more productive” reflects the gap between capability and utility.

Better means that the ratio of output quality to input effort improves.

Am I getting progressively better results for the same or less effort? Does the AI learn from our interactions, or do we have to flag repeatedly what it should know ? Can it interrogate its own outputs, or do we play the skeptic every time? If not this could be one major reason why users also get frustrated with AI.

The vendors optimising for leaderboard positions should also optimise for user productivity, measured honestly.

Pragmatic AI: How to Use AI Effectively Without Burnout

While models from AI technology vendors need to catch up as users we can take a pragmatic approach to scale AI adoption and impact.

1. Budget extra time for human effort and iterations

Traditionally IT projects have baked capital development/testing efforts and costs. AI imposes a significant debt for user in acceptance and testing and even post production monitoring for drift, consistency and reliability. This is a new dimension that needs to be factored in implementation time, costs and rewards.

This observation also matches what Jeffries and Dwarkesh documented: AI shifts work, it doesn’t eliminate it. The work shifts to validation, integration, and judgment which are often harder than the original task.

2. The real productivity measure accounts for substantive human effort both at the start and in the end.

Even with well defined context and conditions AI contributes the first 30% to 80% depending on the domain—closer to 30% for writing and research where judgment is critical, approaching 80% for guidance based code generation ((Karpathy’s coding experience) ) where automated tests catch errors.

However even partial acceleration of human effort is valuable because it enables tackling higher complexity. Humans start and complete the finish. AI accelerates the middle.

3. Agentic AI with human control and oversight

Agentic AI is the 2025 technology buzz but it also introduces a qualitatively different risk profile. The issue is no longer confined to hallucinated text or imperfect reasoning, but to systems that can plan and act autonomously through tools. Research and practitioner commentary increasingly note that once agents are given execution rights, probabilistic errors translate directly into real-world consequences.

Personal AI agents such as Moltbot make this tangible: they can read messages, interpret intent, and initiate actions across services. A misinterpretation, prompt injection, or reasoning error can cascade into irreversible outcomes.

Sandboxing is necessary but does not help when you move Agentic AI into production. Agent behavior must be monitored against explicit expectations, with deviations detected in near real time. As with production infrastructure, anomalous behavior should trigger escalation, throttling, or shutdown. Without behavioral monitoring and accountability loops, agentic AI turns uncertainty into an operational and liability risk.

The Future of AI Productivity Depends on Reliability

Three years after ChatGPT’s launch, we’ve crossed a threshold. The question is no longer whether AI is useful. It demonstrably is. The question is whether AI is dependable.

Right now, for production grade work, it’s not yet there.

Humans have adapted. We’ve learned prompt engineering, workflow redesign, multi-model strategies. We’ve invested to figure out what these systems can and cannot do. We’ve also accepted fragility as the cost of access to capability so far.

But there’s a limit to how much adaptation users should bear. When model updates break our workflows without consent, when throttling makes systems unavailable at critical moments, when “comprehension debt” accumulates because we can’t trust outputs we haven’t verified—the bottleneck isn’t user adaptation. It’s system reliability.

Karpathy suggests intelligence is ahead of infrastructure. There is an alternate view. Good users know how to shape workflows. We’re hampered by unreliable, inconsistent intelligence. The limitation isn’t our processes. It’s the systems themselves.

This is not anti-AI. This is about getting better outcomes and reducing user frustration. AI has proven transformative. Now it needs to prove dependable.

That’s the Olympian “stepping up’ AI needs and I am optimistic we will see a champion in the making.

  • What has been your experience using AI for production work?
  • Are you seeing the same or different challenges with the AI models ?
  • What would make you trust an AI upgrade
  • What is your approach to Pragmatic AI for future scale up?

 

Love to hear your reactions and comments

Best wishes

Sidebar: Sources

Share:

Leave a Reply

Your email address will not be published. Required fields are marked *