Skip to main content

The Enterprise AI Intent Gap

The Enterprise AI Intent Gap

Every hype cycle produces its own comfortable silence, and this one has a good one: almost everybody is doing AI, and almost nobody will say out loud whether it worked. In our recent CRMKonvo with Jon Reed, co-founder of diginomica, we spent an hour poking at that silence, with Ralf Korb doing the poking alongside me. Jon is one of the analysts who actually tests the thing before writing about it, which makes him tiresome company for vendors and excellent company for buyers. The conversation did not land on whether AI works. It landed somewhere considerably more uncomfortable: most enterprises cannot say what working would look like, and they started spending anyway.

TL;DR

If you want to watch the full CRMKonvo, please go ahead here (optimized for smartphones) or here (optimized for tablets/computers).


Else, be my guest and continue to read.

Or do both …

Sixty Percent, And Nobody Is Blushing

Jon opened with numbers rather than opinion, a habit more of us should copy. McKinsey's recent work on AI measurement found that nearly eight in ten companies are using generative AI in some capacity, while around sixty percent report not seeing enterprise-wide EBIT impact from those programs. The gap between activity and impact is apparently not closing. Instead, it seems to be widening. A small group of over-performers is automating end-to-end workflows inside specific domains and getting results, and even they argue about what to measure and how to attribute the improvement.

Let that sink in for a moment. These are organizations that committed budget, headcount and executive credibility to a program, then discovered they had never agreed on a definition of success. Jon asked the obvious question: "Why would you undertake a project like this if you had no idea how you were going to measure the success of it?"

The answer is not stupidity.

It is fear.

There is "a profound fear of missing out or being left behind", and the people applying that pressure are usually the ones furthest from the technology. Executives and board members "have some of the most unrealistic ideas about AI in the entire organization". So the CIO is told to spend on something, anything, and the something arrives in the shape of a forward deployed engineer. Fine role, wrong instinct: a forward deployed engineer is a technologist. They do not know your business, and this was never an engineering exercise.

The Model Stopped Being the Product

This is the part of the story most vendor keynotes skip. Five or six years ago, scaling language models produced results that looked like emergent intelligence, and the valuations followed: if you can build truly cognitive systems, you can replace large parts of the workforce and justify almost any capital expenditure. What actually arrived is "a facsimile of intelligence". Then the scaling laws slowed, the training data ran thin, and investors got nervous.

What the labs did affects you more than the models themselves do. They wrapped the models in compound architectures with external verification, symbolic tools, deterministic cross-checks that refuse to let an agent post to the general ledger when the entry fails a rule. The industry sells this as context engineering and harness engineering. Strip the vocabulary away and the model has become the least interesting component in the stack. The architecture around it has become the product.

Jon's summary of the exercise is simple: "We're taking a tool that was not intended to be deterministic, and we're trying to see how far we can push that." That is a candid description of the state of the art, and it carries a warning no vendor slide will show you. Guardrails work in one direction only. "It's a lot easier to stop agents from doing something wrong than to know for sure that they did something right, because they don't understand what the right thing is." Blocking a bad output, better one too many than missing one, is engineering. Certifying a good one is still your problem.

Expertise Is Not a Commodity, and Nothing Is Learning

Two claims circulating on LinkedIn got taken apart, and both had it coming.

The first is that expertise has been commoditized.

It. Has. Not.

What has been commoditized is working-level knowledge across many domains, which the models absorbed during training. Working knowledge is not expertise. "It's only expertise that can identify the problems in the model output," Jon argued. That sentence should reorganize your hiring plans. If you believe the machine is the expert because it passed the bar exam, you will ship its mistakes at scale and file the result under productivity. Mathematics is an exception, because synthetic data works inside the closed confines of maths. Your industry is not maths.

The second claim is that the agent learns from your users.

It does not.

The language model is not learning while you talk to it. It is pre-trained, then trained, then frozen, and adjusting weights on the fly runs into catastrophic forgetting; Rich Sutton has a Turing Award and a working explanation of why. When a vendor says the system learns continuously, they mean that a knowledge graph or memory store are updated with your preferences. That is a storage mechanism disguised as learning. It is dangerous because it gives buyers the wrong idea of what is possible, and having wrong ideas about the possible is how budgets get burned.

The CSAT Trap: When Nothing Got Worse Counts as a Win

Now to CX, where the reasoning gets worse rather than better. We looked at companies that replaced level one service with AI assistants, reduced headcount, and reported the result as a win because their CSAT scores did not go down. Hold that up to the light. The stated ambition was to change nothing about how customers experience you while spending less on them.

"Fine, but you're not Amazon." Unless you are a behemoth or an airline, service is one of the few places where you can still out-compete companies that have more money than you. The question is not whether the bot held the line at nine in the evening. It is whether these tools let you run the best service in your industry, including at nine in the evening when your people have gone home.

Corporate intent decides that outcome. If the corporate desire is to solve issues, the technology gets designed to solve issues. If the desire is to deflect them, you have bought a deflection machine with better grammar. And the failure mode is almost never the answer itself; it is the escalation. Jon's own pharmacy routes him through a voice system that offers to help, asks him to describe the problem, then loops him back into the automation he was trying to escape: "If the automated system had answered my question, I wouldn't be asking to talk to the frigging pharmacist." Every enterprise reading this has built that loop somewhere.

Not every customer warrants the same treatment either, and pretending otherwise is not fairness, it is laziness. Your largest account should probably not be routed into the same voice system as everybody else.

Architecture Follows Intent

The best line of the hour was not Jon's own. He borrowed it from a diginomica colleague writing about the International Rescue Committee's AI operating model: architecture follows intent. Decide who you want to be, then build the thing that makes it possible. Most enterprises run that sequence backwards, buying architecture and hoping an intent turns up later.

Equifax came up as the counter-example, and the detail is the useful part. They credit their AI results not to a clever agent but to five years of cloud migration that left their data in a standard fabric, plus proprietary data the models have never seen. Nobody sensible will tell you to spend two years modernizing before touching AI. Jon did not, and neither will I. But the modernization track and the AI track run in parallel, and the sprinkle-sauce theory, the one where AI lets you skip the discipline, is "a LinkedIn feed fantasy land".

Pragmatic Playbook for Enterprise CX Buyers

Settle three things before the next AI proposal reaches your desk.

Build the evaluation suite before the program office. You have to have transparency over what your AI is doing. Define the business outcome, the baseline and the attribution method before the contract is signed, not after the pilot disappoints. Pick a problem meaningful enough to matter and contained enough that getting it wrong does not break the business. If nobody in the room can state success as a number, you are not ready to buy.

Put your pricing and your data in the contract. Any change to outcome-based or consumption-based pricing requires six months of notice so you can adjust. Moving off user-based licences to pay for tokens with no business result attached is not an advancement, it is a higher invoice. And when the vendor says their agent learns from your users, ask these two questions: how exactly does it learn from my users, and how do you protect that data? An update to the knowledge graph is not learning.

Design the escalation first, then hire someone to check the whole thing. Most customer anger at AI support is not about the answer, it is about being unable to get out. Build the route to a human before you build the deflection, and keep your most valuable accounts out of the automation entirely. Then consider the role Jon would add to the org chart: an AI ombudsperson whose job is to walk into departments, gut-check what is being built, and flag the vulnerabilities and the opportunities nobody else is positioned to see.

Architecture follows intent. Buy the intent first.

Comments

Last Year's Top 5 Popular Posts

You are only as good as your customer remembers

As you know, I am very interested in how organizations are using business applications, which problems they do address, and how they review their success. In a next instance of these customer interviews, I had the opportunity to talk with Melissa Gordon , Executive Vice President, Enterprise Solutions at Tidal Basin about their journey with Zoho. You can watch the full interview on YouTube. Tidal Basin is a government contractor that provides various services throughout the government space, including disaster response, technology and financial services, and contact centers. Tidal Basin started with Zoho CRM and was searching for a project management tool in 2019. This was prompted by mainly two drivers. First, employees were asking for tools to help them running their projects. Second, with a focus on organizational growth and bigger projects that involved more people, Tidal Basin wanted to reduce its risk exposure and increase the efficiency of project delivery. This way, the compa...

SAP Draws a Perimeter around Agentic AI and What That Means for the Rest of US

The most consequential enterprise AI governance document published this year arrived in late April with surprisingly little fanfare. SAP's updated API Policy, version 4/2026 , is a short document in plain English. The clause that is most interesting is Section 2.2.2. It restricts how autonomous and generative AI systems are permitted to interact with SAP APIs. Read literally, it has the potential to change the architecture of agentic AI projects across every SAP customer landscape. Read carefully, it is also more interesting than the lock-in headlines suggest. The policy targets a specific category of AI behavior, not AI as such. It connects to commercial mechanics that go well beyond API stability. And the literal text, in its current form, will probably not survive the next two policy revisions intact. There is a lot to unpack. I will walk through what the policy actually says, how the SAP-watching community is reading it, what the rest of the major enterprise vendors are doin...

Beyond GDPR: Is MyTerms the New Standard for Enforceable Personal Data Agreements?

The news IEEE just released standard 7012-2025 for machine readable personal privacy terms , nicknamed MyTerms. MyTerms covers interactions and agreements between individuals and service providers they interact with on a network. It defines a way for personal privacy requirements to be expressed as standard-form contractual agreements.  MyTerms is intended to replace today’s “notice and consent” pattern with a standardized, machine-readable contract handshake between an individual and a service provider. The standard considers individuals true first parties who can proffer privacy terms as contractual terms, typically through an automated agent acting on their behalf. The system relies on a neutral, non-business entity that hosts a bounded set of standard-form privacy agreements. These agreements are designed to be understandable and usable in practice by humans and by machines. They must be available in plain-language human-readable form, maintain legally meaningful wording, and a...

The Illusion of Value: Why Salesforce’s Agentic Work Unit is the New "Bad Query" of the AI Era

The News On February. 25, 2026, Salesforce announced a pricing and metrics update . During the company’s Q4 FY2026 earnings call, CEO Marc Benio ff, together with CMO Patrick Stokes , unveiled the Agentic Work Unit (AWU). Positioned as a metric to quantify the labor performed by autonomous digital systems, Salesforce defines an AWU as one discrete task accomplished by an AI agent. According to Salesforce, this discrete task represents the exact moment " raw intelligence is converted into real work ". It is not a fixed unit but measured as a processed prompt, a completed reasoning chain, or an invoked tool. Salesforce explicitly designed the AWU to move the industry conversation away from the raw consumption of Large Language Model (LLM) tokens. As Benioff noted, tokens only measure "how much an AI talks," whereas the AWU is intended to measure actual business execution. The scale of this rollout is massive. Salesforce reported that its platform has already processe...

LLM Showdown: Comparing ChatGPT, Gemini, and Grok for Automated News Research

The analyst’s day is full of research. Now, this is the age of AI and AI is here to help, isn’t it? As everyone is talking about copilots and AI agents, why not using the tools at hand to do a little research on research. NB., no one really has a good definition of an AI agent, so this might become an additional topic for research. But I digress. Imagine the following project at hand, which is not only interesting for analysts, btw, but also for a variety of roles in the corporate world. Let’s call it vendor (competitor) monitoring. The job is the following: Research reputable sites for news about a number of vendors, relating to a set of keywords. Reputable sites are high quality news sites, high quality tech publications, high quality analyst sites and, of course the news pages of the vendors in question. Limit the time frame of the search matching to the cadence of my information requirement, e.g., “yesterday” for a daily update or “last week” for a weekly update. Provide a summary ...