The AI Revenue Attribution Challenge: Why SaaS Companies Are Getting ARR Wrong in 2025
In the first half of 2025, AI startups announced some eye-popping growth milestones: Lovable claimed $100M ARR in eight months. Cursor hit the same mark in twelve months. Cluely reported doubling their ARR to $7M in a single week.
These numbers would make any SaaS founder jealous. They'd also make seasoned investors very nervous.
Because buried in those announcements is a measurement crisis that's making Annual Recurring Revenue—the SaaS industry's most sacred metric—increasingly unreliable. VCs are starting to ask uncomfortable questions: How much of that "recurring" revenue is actually just paid pilots? What percentage converts to production? And when founders say they doubled ARR in a week, what exactly are they counting?
According to Fortune, six VCs confirmed what many in the industry suspected: founders are routinely counting pilots, one-time deals, and unactivated contracts as recurring revenue. One investor recounted a pitch where a founder claimed a customer would "probably keep paying" after a two-week pilot—and counted it as ARR.
The problem isn't just founder optimism. It's that AI has fundamentally broken the assumptions that made ARR meaningful in the first place. And most companies making billion-dollar decisions based on these numbers don't even realize it.
The measurement crisis
The ARR inflation problem
Annual Recurring Revenue was supposed to be simple. You count your subscription contracts, annualize them, and you have a clean number that represents predictable future revenue. For two decades, ARR served as the North Star for SaaS companies, investors, and operators alike. It was the foundation for valuations, hiring plans, and strategic decisions.
Then AI arrived and broke everything.
Today, founders are counting pilots as recurring revenue. One-time deals are being annualized. Unactivated contracts sit in ARR calculations. Companies with 60% annual churn rates claim they have "recurring" revenue. And according to six VCs interviewed by Fortune, this isn't happening at the margins—it's become standard practice.
The numbers tell a damning story:
- Less than 50% of AI proof-of-concepts and pilots convert to production deployments
- AI-native companies are seeing 60% of their install base churn annually
- For AI products priced under $50 per month, gross revenue retention sits at just 23%—compared to 70% for traditional B2B SaaS
The SaaS Barometer newsletter warns against counting paid pilots or proof of concepts as ARR: with a conversion rate of less than 50% from pilots to production, doing so materially inflates the true nature of ARR in AI products.
But companies are doing it anyway. Why? Because everyone else is. Because the competitive pressure to show hypergrowth is intense. And because, frankly, no one has agreed on what the rules should be anymore.
Why traditional ARR breaks down for AI
Traditional ARR made three fundamental assumptions about SaaS revenue.
Assumption 1: Predictability
In the classic SaaS model, predictability was straightforward. You sold seats. Each seat cost a fixed amount. You could forecast with confidence because next month looked a lot like this month, just with a few more customers.
AI obliterated this assumption.
Consider the "inference whales" that Business Insider reported on—customers generating $35,000 in AI usage while paying just $200 monthly for an unlimited plan. That's 175x variance between what the customer pays and what they consume. Now imagine trying to forecast next quarter's revenue when your top 10 customers could swing usage by 2-10x in either direction based on how aggressively they adopt your AI features.
Token usage doesn't follow neat patterns. A customer might run 1,000 queries one month and 50,000 the next because they rolled out your AI feature to their entire sales team. Or they might drop from 10,000 queries to zero because they decided to pause their AI experiments during a budget freeze.
The elegant simplicity of seats × price = ARR has been replaced by an unknowable equation involving usage patterns, token consumption, compute costs, and customer behavior that can shift dramatically month to month.
Assumption 2: Commitment equals recurring
The second broken assumption is that a signed contract represents commitment.
In traditional SaaS, when an enterprise signed a 12-month contract for 500 seats, they were making a real commitment. The switching costs were high. The implementation effort was significant. Annual churn rates of 5-10% were normal and predictable.
AI has created what I call "the pilot plague." Enterprises aren't buying AI features—they're experimenting with them. They sign 12-month contracts with 90-day exit clauses. They run pilots with no obligation to convert. They explore multiple competing solutions simultaneously, planning to pick one winner (maybe).
The data is stark: 70-80% of customers with these short-term exit clauses actually use them. That's not recurring revenue. That's experimental spending that happens to be packaged in a contract structure that looks like a subscription.
Luigi Mallardo, a SaaS GTM expert, has proposed a new category for this reality: ERR, or Experimental Run-rate Revenue. His argument: a pilot is not Annual Recurring Revenue, even if you signed a 12-month contract with an exit clause, even if they paid for it. Until the client has seen the value, the exit window has closed, and they have made a confident, active decision to commit, it is Experimental Run-rate Revenue.
Assumption 3: Low, predictable COGS
The third pillar of traditional SaaS was the beautiful economics of near-zero marginal cost. Once you built the software, serving additional customers cost almost nothing. Traditional SaaS companies ran gross margins of 70-85%, with COGS representing just 15-30% of revenue.
AI shattered this model.
AI-native companies are seeing COGS of 35-50% of revenue. Foundation model companies are running at 50-75% COGS. Gross margins for early-stage AI companies have compressed by nearly 10 percentage points year-over-year, driven entirely by compute costs.
Every query costs money. Every token processed burns GPU cycles. The inference whales mentioned earlier? They're not just revenue anomalies—they're cost nightmares. When a customer on a $200/month unlimited plan generates $35,000 in compute costs, you're not running a high-margin SaaS business. You're subsidizing their AI experimentation and hoping they don't tell their friends.
The implication? Even if you could accurately measure ARR, it wouldn't tell you what it used to. A company with $10M ARR and $2M in compute costs is fundamentally different from one with $10M ARR and $8M in compute costs. Yet revenue multiples treat them identically.
As Priya Saiprasad, general partner at Touring Capital, put it: the classic SaaS model is dying as we speak, and the industry needs to collectively evolve to a new set of metrics it feels comfortable measuring these companies by.
The industry's confused response
The proliferation of "new ARRs"
In the absence of industry standards, companies have started creating their own definitions. The result is a Babel tower of metrics that makes comparison nearly impossible.
AI ARR. Some companies, like Verint and Salesforce, have started breaking out "AI ARR" as a separate line item. Salesforce's Agentforce product line hit $1.4B in AI ARR with a 114% surge in Q3 2025. Verint defines AI ARR as all of the ARR derived from solutions that include AI functionality, representing the quarterly run rate value of both active and newly signed SaaS agreements.
Notice the nuance: "newly signed" contracts are included in ARR even if revenue hasn't started. This is definitional drift in real-time—what Verint is calling ARR might more accurately be described as CARR (Contracted ARR). And definitions vary wildly across companies:
- Some count only revenue from AI-specific SKUs
- Others allocate a percentage of platform revenue based on AI feature usage
- Still others include any product that "includes AI functionality"—which in 2025 could mean almost anything
CARR (Contracted ARR). To address timing issues, some companies distinguish between contracts signed and revenue recognized. CARR counts newly signed agreements before customers go live. The logic is sound—you want to track forward-looking commitments—but the risk is clear: contracts get delayed, customers don't activate, and pilots don't convert. As one VC told Fortune, founders are claiming "booked ARR" based on what customers might pay in the future rather than what they're actually paying now, even though contracts frequently have provisions that let customers opt out at any time for any reason.
UARR (Usage-based ARR). For companies with heavy usage-based pricing, annualizing recent usage has become standard practice. The challenge? Usage patterns often take 6-12 months to stabilize. Early enthusiasm can crater. Pilot programs can scale up or wind down. Seasonal businesses can have 3-5x variance between peak and trough months. L.E.K. Consulting reports that companies are adopting UARR to reflect account-level ramp patterns—but this creates a new problem: you're forecasting stability that may never arrive.
ERR (Experimental Run-rate Revenue). Perhaps the most honest of the new metrics, ERR acknowledges that pilot revenue is fundamentally different from production revenue. The problem? Almost nobody is using it, because it requires admitting that a chunk of your reported ARR isn't actually recurring.
What companies are actually doing
The data on AI monetization reveals an industry in transition:
- Over 80% of traditional SaaS companies have introduced at least one AI product
- ~50% are charging separately for AI features
- 48% have prioritized AI adoption over monetization (giving it away free)
- But: there's no consensus on measurement standards
For companies that are charging, the approaches vary wildly: pure usage-based pricing (pay per token/query/task), hybrid models (base subscription plus usage overages), tiered feature access (AI capabilities in premium tiers only), and outcome-based pricing (charge based on results delivered).
Each approach creates different attribution challenges. How do you calculate ARR when pricing is outcome-based? How do you forecast when 48% of your AI features are being given away for free as adoption plays? How do you compare your AI ARR to a competitor's when they might be using completely different definitions?
The answer is: you can't. Not reliably.
Why this matters (beyond vanity metrics)
Real business consequences
This isn't just an accounting nuance. The measurement crisis is creating real, material problems for companies, investors, and the broader SaaS ecosystem.
1. Misleading investors
AI companies are trading at premium multiples—an average of 24× revenue compared to 20× for traditional SaaS, according to Bessemer's Cloud 100 report. These valuations are predicated on high growth rates and the assumption of recurring revenue.
But what happens when investors discover that 40% of a company's ARR is actually pilot revenue with under-50% conversion rates? Or that gross margins are 35% instead of the 75% they expected? Or that the vaunted "$100M ARR achieved in 8 months" includes massive churn that will erase half that revenue in the next quarter?
We're already seeing the beginning of what Cassie Young from Primary Venture Partners calls the coming "gross retention apocalypse." AI-native companies with products priced under $250/month are seeing gross revenue retention rates of just 40-60%—compared to 85-90% for traditional B2B SaaS.
Companies that raised at inflated valuations based on inflated ARR metrics are going to face brutal down rounds when reality catches up.
2. Internal decision-making failures
When you believe your ARR is $10M but $4M of it is experimental revenue that won't convert, you make catastrophic decisions:
- You hire for $10M worth of growth, not $6M
- You invest in product features based on phantom demand
- You set sales quotas that can't be achieved
- You communicate impossible expectations to your board
How do you measure whether AI features are driving ARR when customers are on multi-product platforms, usage spans both base functionality and AI enhancements, revenue is a mix of committed minimums and usage-based overages, and some customers are in pilots while others are in full production?
Without proper data governance and clear definitions of what "counts," you're building your strategy on quicksand.
3. Customer success disasters
Perhaps the most insidious problem is what bad measurement does to customer success and retention.
When you can't distinguish between experimental and committed revenue, you can't effectively:
- Predict churn: Which customers are at risk of leaving vs. which are deeply embedded?
- Allocate CS resources: Should your CSM spend time with the "$50K ARR" customer who's actually running a 30-day pilot?
- Calculate true NRR: Net Revenue Retention becomes meaningless when your base keeps shifting
The data shows this playing out in real-time. AI-native companies are seeing wildly different retention profiles based on deal size:
- Products >$250/month: 70% GRR, 85% NRR (similar to B2B SaaS)
- Products $50-$249/month: 45% GRR, 61% NRR (15 points worse than SaaS)
- Products under $50/month: 23% GRR, 32% NRR (catastrophic)
These aren't bugs—they're features of a measurement system that treats experimental spend the same as committed spend.
The data governance dimension
The fundamental problem isn't just that we lack standard metrics. It's that most companies lack the data infrastructure to measure honestly even if they wanted to.
Proper AI revenue attribution requires answering questions like: Which specific features are customers using? How much value are AI capabilities creating vs. base product functionality? When did a customer transition from pilot to production usage? What's the actual compute cost per customer segment? How do usage patterns correlate with retention and expansion?
Customer vs. company data clarity. At Recurly, we've been developing frameworks to distinguish between customer data (belongs to the customer, limited use rights) and company data (platform operational data we can use for analytics). This distinction becomes critical for AI attribution. If I can't analyze cross-customer usage patterns because of data ownership constraints, I can't build accurate models for what "normal" AI adoption looks like. If I can't link product telemetry to revenue data, I can't measure which AI features actually drive expansion ARR vs. which ones customers use but don't value enough to pay for.
Operational data architecture. You need systems that can track feature-level usage in real-time, link usage to specific revenue streams, distinguish pilot/trial usage from production workloads, calculate per-customer unit economics, and connect product analytics to finance systems. Most companies built their data infrastructure for seat-based SaaS. Product usage was a nice-to-have for engagement metrics, not a critical input for revenue recognition. AI requires rethinking the entire architecture.
Cross-functional transparency. Honest AI revenue measurement requires finance teams seeing product usage data, product teams understanding revenue attribution, sales teams knowing real pilot conversion metrics, and leadership accepting transparent reporting of experimental vs. committed revenue. This is hard. It means sales can't sandbag by calling pilots "committed ARR." It means finance can't hide behind accounting definitions when the underlying economics tell a different story. But without this transparency, you're flying blind.
A framework for honest AI revenue measurement
The industry needs standards. While we wait for them to emerge, here's a practical framework that companies can implement now.
Principle 1: Separate experimental from real revenue
Create two distinct categories.
Experimental Run-rate Revenue (ERR):
- Paid pilots with exit clauses
- Free trials where customers are using AI features
- Proof-of-concept deployments
- Contracts within the first 90 days
- Any revenue where the customer hasn't demonstrated committed, production usage
Annual Recurring Revenue (ARR):
- Contracts past the trial/pilot exit window
- Customers who have achieved defined success criteria
- Active production workloads sustained for 90+ days
- Revenue you would actually defend in a due diligence process
Your reported metrics should look like this:
Total Reported Revenue = ARR + ERR (disclosed separately)
Conversion Metrics:
- ERR → ARR conversion rate
- Average time to conversion
- Churn rate by cohort (ERR vs. ARR customers)
Don't convert ERR to ARR until all three conditions are met: the exit window has closed, the customer achieved defined success criteria, and an active production workload has been sustained 90+ days.
This is hard discipline. Your ARR number will be lower. Your board will ask uncomfortable questions. Do it anyway. The companies that maintain this discipline now will be the ones that survive when investors start demanding it.
Principle 2: Break down AI attribution components
Stop reporting a single ARR number. Break it down:
Total ARR = Base SaaS ARR + AI Feature ARR
AI Feature ARR further segmented:
├── Committed AI ARR (contracted minimums)
├── Usage-Based AI ARR (annualized actual usage)
└── Hybrid AI ARR (base + overages)
For each segment, track revenue concentration (what % comes from top 10 customers?), serious vs. experimental spend (what % of AI ARR comes from customers spending >$250/month?), conversion metrics (what's the pilot-to-production conversion rate?), gross margin (what's the actual GM on AI features vs. base product?), and time to stability (how long until usage patterns become predictable?).
Beyond the standard SaaS metrics, add these AI-specific measures:
- AI ARR contribution %: What percentage of total ARR comes from AI features? If you're claiming AI is transformational, this should be >20% and growing. Reality check: most companies are under 10%, many under 5%.
- AI feature adoption rate: What % of customers use AI capabilities—by cohort, by tier, and by use case?
- AI unit economics: Revenue per AI-active customer, compute cost per AI-active customer, gross margin per customer segment, and CAC payback period.
- Conversion funnel metrics: % of pilots that convert to paid, % of paid that reach production usage, % of production users that expand, and time to reach each stage.
Principle 3: Build data governance for attribution
None of the above works without the right data infrastructure and governance.
Product telemetry linked to revenue. You need feature flags and usage tracking at the customer level, event streams that capture AI interactions, attribution logic that connects features to revenue, and dashboards that make this data accessible to non-technical stakeholders.
Customer data ownership clarity. You need clear policies on how you classify customers in trial/pilot vs. production stages, clear definitions for the adoption journey (experimenting → evaluating → adopting → expanding), and whether you can use consumption data to inform pricing and packaging. Consider frameworks that distinguish customer data (usage information that belongs to the customer), company data (aggregated, anonymized operational metrics), and shared data (information that benefits both parties). These distinctions matter because AI attribution often requires analyzing patterns across customers. If your data governance policies prohibit that analysis, you can't measure accurately.
Cross-functional visibility. Finance needs to see product usage data, feature adoption metrics, compute costs per customer, and conversion funnel performance. Product needs to see revenue attribution by feature, which capabilities drive expansion, cost-to-serve per feature, and customer willingness to pay. Sales needs to know real pilot conversion metrics, average time from pilot to production, signals that predict conversion, and red flags that indicate experimental rather than committed spend.
A practical implementation sequence:
- Define clear AI product boundaries: Which features count as "AI-driven" vs. enhanced base functionality?
- Build attribution logic: Create models that allocate revenue across a multi-product platform based on actual usage patterns
- Establish measurement cadences: Monthly for AI feature adoption and usage, quarterly for AI ARR contribution and conversion metrics, annually for attribution model reviews
- Create transparency mechanisms: Exec dashboards showing ARR vs. ERR, product analytics showing feature-level revenue impact, regular reviews of AI investment ROI
- Set honest targets: Base AI ARR goals on realistic conversion assumptions and clear attribution methodology—not aspirational math
The goal is to build the muscle of honest measurement before you have billions of dollars riding on these numbers.
What this means for 2026
Several forces are converging that will make the current measurement chaos untenable.
1. Regulatory pressure. The SEC is watching. Public companies claiming massive AI ARR growth will face scrutiny about how they're calculating those numbers. Expect required disclosures separating committed vs. usage-based revenue, mandatory reporting of pilot conversion rates for material revenue streams, standardized definitions emerging from accounting bodies, and private companies adopting these standards 12-18 months before going public.
2. Investor sophistication. VCs are getting burned. The companies that raised at 30× revenue multiples based on pilot ARR are struggling. Smart investors are already demanding separate disclosure of ERR vs. ARR, due diligence on pilot-to-production conversion rates, detailed analysis of AI unit economics, and proof of sustainable gross margins, not just growth rates.
3. The great winnowing. Not every company will survive this transition. The ones at risk: companies with >40% of ARR in pilots that won't convert, businesses claiming high ARR while subsidizing compute costs, and startups that raised at inflated valuations based on inflated metrics. The survivors will be companies with honest measurement frameworks built early, platforms with hybrid pricing that balances predictability and usage, businesses with data governance enabling accurate attribution, and teams that chose transparency over vanity metrics.
4. Industry standardization. Nature abhors a vacuum. Within 18-24 months, we'll see informal standards emerging from benchmark reports and investor expectations, billing and analytics vendors building attribution capabilities, industry groups codifying measurement approaches, and companies with clean metrics gaining advantages in fundraising and M&A. The companies defining these standards now will be the ones others copy.
The call to action
Remember that VC from the opening story? The one who heard "we'll probably keep paying" as justification for doubling ARR in a week?
He now asks three questions in every pitch:
- What percentage of your ARR is past the pilot exit window? If they can't answer, or if it's under 50%, that's not recurring revenue
- How do you separate AI revenue from base product revenue? If they haven't built attribution logic, they're guessing
- What's your ERR-to-ARR conversion rate? This single metric reveals whether pilots are converting or churning
These aren't gotcha questions. They're basic discipline in an age where the old metrics don't work anymore.
The SaaS industry standardized around ARR because predictability created trust. Investors knew what the number meant. Operators could benchmark against it. Boards could govern with it.
AI doesn't have to destroy that trust—but it requires us to measure honestly.
The choice isn't between growth and accuracy. It's between building on a foundation of truth versus building on fog. Companies choosing fog might hit impressive-looking milestones faster. But when the market demands substance, they'll have nothing to show.
The time to build these frameworks is now. Before your next board meeting. Before your next fundraise. Before half your ARR evaporates because it was never really recurring in the first place.