AI Agents in 2026: The gap between hype, usefulness and trust

AI agent workflow showing human review, task automation and trust checkpoints

Ross Harrington

Head of Design

Published:

AI agents have become one of the most hyped areas in technology. Every new release seems to arrive with the same wave of excitement, promises of autonomy, productivity breakthroughs and systems that can finally “do things for you” rather than just respond to prompts. But has that hype actually translated into reality?


A few years ago, much of the conversation around AI agents was bold and almost immediate in tone: agents will replace entire jobs imminently. Fast forward to 2026 and the discussion has matured significantly. The focus has shifted away from headline-grabbing claims toward more grounded questions; Where do agents actually work well? Why is production deployment so difficult?, and What infrastructure, governance and safeguards are required to make them usable at scale?


This shift in perspective is important. It reflects a growing understanding that building a convincing demo is very different from building a reliable system. Recent research into production AI agents suggests this is now the central challenge: reliability, governance and human oversight matter as much as model capability.


The early vision: AI acting in the real world


One of the most famous early demonstrations of this vision came in 2018, when Google showcased Duplex, an artificial intelligence system capable of making phone calls on behalf of a user. In the demo, it successfully booked a hair appointment over the phone, complete with natural conversational pauses and human-like speech patterns.


For many people, this was a defining moment. It wasn’t just a chatbot answering questions, it was an AI acting in the real world, interacting with humans and completing tasks independently. Duplex set the tone for what “AI agents” were supposed to become: systems that could understand intent and carry out actions in place of the user.


Do AI agents actually deliver value today?


In practice, AI agents are already useful in certain environments. In work settings, they can be genuinely valuable tools. They help generate written content, support research at both industry and organisational levels, summarise large volumes of information and assist with early-stage creative thinking.


In these contexts, the structure of work helps them succeed. Tasks are often well-scoped, expectations are clearer and humans remain in the loop to review and refine outputs. Under these conditions, AI agents can enhance productivity and reduce cognitive load.


However, outside of structured work environments, the picture becomes more complicated.


The gap between AI agent capability and reliability


In everyday use, experiences with AI agents are often far more inconsistent. Not because they are useless, but because their behaviour is still not reliable enough to trust without supervision. The issues that arise are not rare edge cases, they tend to be recurring patterns that appear across different systems and use cases.


Three of my recent experiences illustrate this gap quite clearly.


1. Context and memory failures


When using ChatGPT to learn chess, the system can provide helpful explanations, suggest openings and break down basic strategy effectively. However, as the interaction continues, it often loses track of key context. It may forget which colour I am playing, lose awareness of piece positions, or even contradict explanations of the rules.


For a system capable of discussing complex strategic ideas, this kind of inconsistency is surprisingly disruptive. The core issue isn’t intelligence, it’s persistence. The inability to reliably maintain context over time limits how far the system can be trusted in ongoing tasks.


2. Voice assistants that struggle with natural conversation


Voice-based agents such as Google’s Gemini assistant demonstrate impressive capabilities. They can alert users to incoming messages, read them aloud with natural tone and rhythm and even suggest appropriate responses.


The breakdown happens in the interaction loop. When a user begins to respond, any brief pause, whether to think, breathe, or structure a sentence is often interpreted as the end of input. The system stops listening prematurely, forcing the user to restart or repeat their entire response. Instead of feeling like a fluid conversation, the experience becomes constrained and fragile, where natural human pauses conflict with machine expectations of input.


3. The chatbot loop problem


In banking and customer service contexts, AI agents often perform well within predefined boundaries. For example, when using NatWest’s Cora assistant to learn about savings pots, the system can clearly explain what a savings pot is, how to set one up and how to manage it. It can also guide users through common predefined questions effectively.


However, once the conversation moves outside those expected paths, the system tends to break down. Rather than adapting to the new query, it often falls back into repeating earlier prompts or cycling through the same set of responses. This creates a loop where the original question remains unanswered and there is no effective escalation to a human agent. The result is not just frustration, it’s a loss of trust in the system’s ability to handle anything beyond scripted interactions.


The core tension: usefulness vs trust


What emerges across these examples is a consistent pattern. AI agents are not incapable. In fact, they are often impressively capable within constrained environments. The issue is not raw intelligence, it is reliability. They are just reliable enough to be useful in many situations, but not reliable enough to be fully trusted in high-stakes or open-ended contexts. And that gap creates friction. Users can see what these systems are capable of, but they also repeatedly encounter their limitations in real-world use.


A final thought


Perhaps the most important question is not when AI agents will become capable of replacing human work, but something more subtle: when will they become reliable enough that people stop feeling the need to constantly verify their output? Because until that point is reached, AI agents will remain in an awkward middle ground, powerful, promising and increasingly present, but still not quite dependable enough to be fully trusted.


Related reading:


Ross Harrington

Head of Design

Reimagine your digital future today

Send us a message for more information about how we can help you and your business

Reimagine your digital future today

Send us a message for more information about how we can help you and your business

Reimagine your digital future today

Send us a message for more information about how we can help you and your business

Reimagine your digital future today

Send us a message for more information about how we can help you and your business

Services

Capabilities

About

Linebreak

New Icon is a Linebreak company

© Newicon Ltd. Registered in England and Wales. Company No: 05904359 | VAT: GB 993768447.

Designed and built by New Icon in Bristol, a Linebreak company.

Linebreak

New Icon is a Linebreak company

© Newicon Ltd. Registered in England and Wales. Company No: 05904359 | VAT: GB 993768447.

Designed and built by New Icon in Bristol, a Linebreak company.

Linebreak

New Icon is a Linebreak company

© Newicon Ltd. Registered in England and Wales. Company No: 05904359 | VAT: GB 993768447.

Designed and built by New Icon in Bristol, a Linebreak company.