• AboutLearn more about Oxx, our investment strategy and how we work.
  • PortfolioDiscover the Oxx portfolio and learn about each of the companies we have partnered with.
  • TeamMeet the Oxx team and learn more about each team member.
  • News & InsightsCheck out the latest news, insights and perspectives from the Oxx team and our portfolio.
  • About
  • Portfolio
  • News & Insights
  • Team
  • Get in touch

Go-to-Market Fit

Your guide with tools, templates and examples and resources to nail the scaling journey.

Learn more

The Rising Stars of Open Source

Who is winning in the open source space? Oxx annual deep-dive into of the best upcoming OSS companies.

Learn more
  1. Oxx
  2. News & Insights
  3. The (not-so) new paradigm for engagement with AI software

The (not-so) new paradigm for engagement with AI software

minutes

|

August 18, 2026
Phil Edmondson-Jones - Partner, Oxx

Has AI fundamentally broken the way we measure software engagement? Recently, this view has become increasingly prevalent, but it’s only half right.

Software has always encompassed two styles of product. The first style; including infrastructure, monitoring and protection, was always designed to run autonomously. In the second, software was a tool operated by a human, sitting in a seat, and the usage, engagement and behaviour of said ‘seats’ was a fair enough proxy for value. AI has left the first style broadly intact. But it has collapsed the second: in a growing set of products, the tool now does more of the work the human used to do inside it, and the human’s presence stops being evidence of anything.

Over decades of reviewing B2B software companies, and now several years deep within the enterprise AI era, we have come to recognise six distinct product archetypes with associated differences in end-customer engagement and value. Using these archetypes to orient ourselves, we’ve been able to get a stronger grasp of which parts of the AI-first landscape are surprisingly consistent vs those that are genuinely different from previous eras, and what type of engagement is worth evaluating per archetype. 

The six archetypes of (AI) software products

ArchetypeExamplesThe value exchange
The Grid
Infrastructure & Flow
OpenAI / Anthropic, CoreWeave, Nebius, Together AIPay for what you use. The customer buys reliability, scale and cost predictability on someone else’s infrastructure rather than building it; the vendor carries the capital risk and sells the resulting capacity by the unit.
The Watchtower
Monitor & Vigilance
Dataminr, Chainalysis, SOCRadarContinuous coverage of a scope too large or too complex to monitor manually. The product watches, filters and escalates; the human retains the judgement and the action.
The Shield
Protection
Darktrace, Abnormal Security, Vanta, CyberSmartQuantified risk reduction. The customer is buying the absence of a bad outcome: a breach, a fine, an operational failure, so the product succeeds precisely when nothing happens and there is nothing to look at.
The Concierge
Customer-Facing Automation
Sierra, Decagon, Intercom Fin, ParloaDeflection without degradation. Customer interactions are resolved before they reach a human agent, at a lower cost per contact, without the customer experience deteriorating.
The Intelligence Layer
Insight & Analysis
Glean, AlphaSense, HebbiaDemocratisation of insight. Answers arrive at the speed of a question, without the person asking needing SQL, an analyst’s time, or any knowledge of where the underlying information sits.
The Output Engine
Final Output Generation
Cursor (co-pilot), Harvey (autopilot), Cognition (agentic)Labour efficiency. The product accelerates or replaces human output on a defined body of work, so the customer is buying units of work delivered rather than access for people to do that work themselves.

Let’s have a more detailed look at each of these in turn.

The Grid – Infrastructure & Flow

The Grid represents the foundational plumbing of the AI stack: raw compute, model access, or data throughput. Usage is high-frequency, transactional, and continuously consumed. The value exchange is simple: pay for what you use, and get reliability and scale in return.

What good looks like is a cohort utilisation on-track rate above 80% of committed annual volume by months eight or nine of a contract year, alongside a low credit idle rate; credits purchased are consumed, not left dormant at period end. What misleads is aggregate token or credit volume without normalising for cohort vintage or commitment size, and new logo growth that masks declining utilisation in the existing base.

The one genuinely AI-specific wrinkle is not an engagement question at all but a pricing one: unit price per token has deflated quickly enough that volume growth and revenue decline can occur simultaneously. Utilisation curves have to be read alongside realised price per unit, never in isolation.

The Watchtower – Monitor & Vigilance

Watchtower products watch over a defined estate: a live data stream, a bounded estate of documents, vendors, or assets, and surface anomalies, risks, or changes that require attention. The value is informing a human who then acts. Intermittent or background usage is correct behaviour here, not a warning sign, which makes MAU/DAU (monthly active users/daily active users) almost meaningless for this archetype.

The first important metric here is coverage completeness: the percentage of intended scope actively monitored versus total in scope. The second is signal precision, the true positive rate on issues surfaced, trending up over time. Rising alert or flag volume is typically a negative signal: it suggests the product is generating noise, not insight.

What AI expands is the scope of what can be watched at all: unstructured sources that rules-based systems could not parse. Coverage completeness is therefore a moving target, because the denominator grows as capability does.

The Shield – Protection

Protection products defend organisations from breaches, regulatory violations, or operational risk. The ideal user experience is near-invisibility: the bad thing doesn’t happen, and the CXO reports a clean quarter. High engagement here in time spent on platform or high frequency of alerts is a negative signal, not a positive one.

Quantified risk outcomes are what matter: incidents detected and remediated, translated into avoided cost or liability. Alert precision is a leading indicator of product quality, and a leading indicator of churn when it falls. Compliance posture trajectory, meaning aggregate risk score across the estate improving over time, is the equivalent of a utilisation curve for this archetype.

The Concierge – Customer-Facing Automation

Concierge products handle inbound customer interactions: support tickets, queries, complaints, autonomously on behalf of the business. The AI is the first, and ideally the only, point of contact. The value is deflection without degradation: containing customer interactions before they reach a human agent, while maintaining or improving customer experience.

Containment rate against a pre-AI baseline, CSAT on AI-handled interactions benchmarked against human-handled equivalents, and resolution accuracy are the metrics that define success. Raw volume of interactions handled is not. Scale without quality is a liability: fast incorrect resolutions are worse than slower correct ones.

This is the archetype where pricing has moved fastest, and for a specific reason. Because containment is measurable at the interaction level, it can be sold per resolution rather than per seat, which also means the vendor now carries the cost of every failed one.

The Intelligence Layer – Insight & Analysis

Intelligence Layer products answer questions, against structured databases or unstructured enterprise knowledge, and enable users to extract insight without manual analysis or search. The value is democratisation of insight: answers at the speed of a question, without requiring technical expertise.

This is where the break with the old world becomes real. In the BI and enterprise search tools that preceded these products, seat usage genuinely was the health metric: the human performed the analysis inside the tool, so adoption was the entire battle and depth of engagement was worth measuring. Now the tool performs the analysis, interactions are short and sparse by design, and adoption tells you nothing about whether the answers are any good.

Answer quality is the only metric that matters: accuracy rate against ground truth for structured products, or first-hit resolution rate for unstructured ones. Deflection rate, the proportion of queries resolved without escalation to an analyst, SQL resource, or colleague, is a useful secondary signal. Total query volume and session length tell you almost nothing about whether the product is working.

There is also a failure mode here with no equivalent in deterministic BI: a confidently wrong answer the user has no reason to interrogate. In this archetype, accuracy measurement is not a reporting nicety, it is the only defence.

The Output Engine – Final Output Generation

Output Engines produce knowledge work outputs on behalf of internal users, spanning AI-assisted workflows (co-pilot), fully autonomous drafting (autopilot), and multi-step agentic task execution. The AI does work that was previously done by a person.

This archetype has no clean SaaS predecessor. Document automation and RPA are the nearest analogues, and neither was measured on intervention rates or priced against units of work delivered. The value exchange is labour substitution, which is why the metrics look unlike anything in the categories above.

The critical nuance is that the right intervention benchmark depends on the autonomy level of the product. For a co-pilot, a healthy edit rate confirms the human is engaged and applying judgement; near-zero edit rates are a warning sign of rubber-stamping. For an autopilot, a low override rate and a falling error rate are the target. For agentic products, task completion rate trending up is the signal. Applying the wrong benchmark will produce a deeply misleading picture of product health, and the same raw number can be evidence of health or of failure depending on which mode you are in.

Efficiency versus baseline is the underlying metric for all three modes: measurable reduction in time-to-complete or cycle time for the target workflow versus the pre-AI equivalent.

What has actually changed

ArchetypeSaaS-era predecessorWhat has actually changed
The GridAWS, Twilio, SnowflakeLittle. Consumption metrics carry over intact
The WatchtowerDatadog, Splunk, Recorded FutureLittle. Low engagement was always correct here; scope is wider
The ShieldSymantec, Palo Alto Networks, ArcherLittle. Better detection, identical value exchange
The ConciergeZendesk Answer Bot, Ada, LivePersonAccountability. Same metrics, vendor now owns and prices them
The Intelligence LayerTableau, Looker, CoveoAdoption has stopped proxying for value; accuracy is the constraint
The Output EngineNone. Nearest analogues are UiPath and document automationEverything. Labour substitution priced by units of work delivered

Interestingly, three of the six archetypes are barely touched from the last era. The Grid runs on the consumption playbook that AWS, Twilio and Snowflake established a decade ago; the Watchtower inherits coverage and precision from monitoring and threat intelligence, where nobody ever measured daily actives; and the Shield’s near-invisibility was already the operating model of antivirus, firewalls and GRC. For all three, AI has improved capability without altering the value exchange or how to evidence it. Reaching for a fashionable new metric set here is a mistake in the opposite direction.

The Concierge sits in between. The software used to be a seat-based tool for human agents, with containment as the buyer’s operational problem, and it is now a service where the vendor owns the outcome, prices against it, and absorbs the cost of getting it wrong.

Only the Intelligence Layer and the Output Engine represent a genuine break, and they break differently. The Intelligence Layer has a predecessor but an inverted relationship to it: seat usage really was the health metric in BI and enterprise search, because the human did the analysis inside the product. Now the product does it, and the signal has flipped from adoption to accuracy. The Output Engine has no predecessor at all.

One thing, however, is new to all six. Output quality is now a probabilistic variable that can regress without a line of code changing. The precision of a Watchtower alert or the accuracy of a Shield’s detection is a property of a model that can drift, not a fixed characteristic of a codebase, which means it has to be monitored continuously rather than validated once at implementation. Even in the archetypes where nothing else has changed, that is a genuine break with the rules-based ancestors of these products.

The metrics that matter, by archetype

ArchetypeThe right metricsWhat misleads
The GridCohort utilisation on-track rate; low credit idle rate; realised price per unitAggregate token/credit volume without normalising for cohort vintage or commitment size
The WatchtowerCoverage completeness; signal precision (true positive rate)Total alert volume; MAU/DAU
The ShieldIncidents detected and remediated; alert precision; compliance posture trajectoryTime spent in platform; total alert volume; reports produced
The ConciergeContainment rate vs. pre-AI baseline; CSAT on AI-handled interactions; resolution accuracyRaw interaction volume; response speed in isolation
The Intelligence LayerAnswer accuracy against ground truth or first-hit resolution rate; deflection rateTotal query volume without quality measurement; session length
The Output EngineIntervention rate calibrated to autonomy level; efficiency vs. baseline; quality trajectoryActivity metrics (sessions, outputs) without quality measurement

What this means in practice

The common thread is not that activity has universally decoupled from value. In half of these archetypes it was never coupled in the first place. What has happened is that the set of companies where engagement is a legitimate proxy has shrunk. For those, the old question: “how intensively is the product used?” has been replaced by a better one: “is the product verifiably getting to work?”

For founders, that means instrumenting outcome and quality metrics from day one, not as a reporting exercise but as the primary mechanism for understanding whether the product works, not just whether it’s being used. These metrics also matter commercially: the shift from seat-based to outcome-based pricing is structurally enabled by the ability to evidence the outcome. Founders who can demonstrate containment rates, intervention rates, or utilisation trajectories are in a materially stronger position at renewal, expansion, and fundraise.

For operators, it means building dashboards around the right archetype from the outset: utilisation, precision, containment, coverage, intervention rates, not sessions and logins. The absence of conventional engagement signals is expected and healthy. The presence of quality signals trending in the right direction is the story.

For investors, it means diagnosing which archetype a company actually occupies before reaching for a benchmark. Get that wrong in either direction and the diligence is worthless: apply engagement metrics to a Shield and you will penalise a healthy company, apply them to an Intelligence Layer and you will underwrite a broken one. Then ask for the corresponding measures. Their absence is itself a finding: a company that cannot evidence its core value exchange may not yet fully understand it. And those that can evidence it, clearly and consistently at the customer level, are significantly better positioned to defend and grow revenue than those that cannot.

The engagement metrics that characterised the last generation of B2B software companies are not wrong, and for half of these company archetypes, very little has changed. The fallacy is in applying them to the archetypes where the software user is now an empty seat.

Phil Edmondson-Jones - Partner at Oxx
Author

Phil Edmondson-Jones

Partner

Profile

Subscribe to our newsletter


info@oxx.vc

Strandvägen 7a, 114 56 Stockholm, Sweden

19 Langham Street, London, W1W 6BP, United Kingdom

Contact

Submit your pitch

Diversity VC logo

Stay up to date

  • AboutLearn more about Oxx, our investment strategy and how we work.
  • PortfolioDiscover the Oxx portfolio and learn about each of the companies we have partnered with.
  • TeamMeet the Oxx team and learn more about each team member.
  • News & InsightsCheck out the latest news, insights and perspectives from the Oxx team and our portfolio.
  • Get in touchContact the Oxx team.
  • Go to market fitWe focus on the Go-to-Market fit stage of companies’ growth journeys. Tap in to our resources about this stage.
  • FAQFind answers to frequently asked questions.
  • LinkedIn
  • Twitter
  • Legal Notice & Privacy Policy
  • ESG Policy
  • SFDR Disclosure
  • Cookie Policy
  • Careers