What stays scarce when intelligence is free? We think frontier labs will keep spending billions to push the frontier, open source will keep making last quarter’s frontier free, and the companies that own the data, the workflow, and the trust can build on this free intelligence to drive huge TAM expansion.
This month Moonshot released Kimi K3, a 2.8 trillion parameter model that scored within three points of the best closed model on earth. Then came the punchline. Moonshot is giving the weights away for free, in the same month Anthropic posted a 50% price increase for its Opus model.1 We think this is a watershed moment. For three years enterprises treated AI spend as the cost of looking curious, and nobody read the invoice too closely. The CFOs have now read the invoice. ROI decides which projects survive, and someone in every budget meeting is asking the obvious question. Do we really need Fable 5, a model so capable that Anthropic briefly pulled it after Washington got nervous, to draft an email, summarize a PDF, or write the first pass of an MD&A? That is a Bugatti doing the school run. A free model three points behind does the job at a rounding-error price.
The consensus take on AI and enterprise software goes like this: foundation models will eat the application layer, SaaS multiples deserve to compress, and the value will accrue to the lab oligopoly. We think the consensus is wrong.
Open-source AI models (Llama, DeepSeek, Qwen, Mistral, and the rest) are doing to intelligence what open-source Linux did to the operating system: converting it from a priced product into free infrastructure. When a critical input becomes free, the winners are the companies that own everything the free input cannot replicate: proprietary data, control of end-to-end workflows, vertical context, security, compliance, SLAs, a customer success motion already sitting inside the account, and distribution that took two decades to build. The right SaaS companies will use free intelligence to add tremendous value into the captive installed base. Open-source AI is a subsidy that turns the best of them into AI companies, with someone else funding the R&D.
We call the winners AI Thrivers: companies whose data gravity and mission-critical workflows get more valuable as the models get smarter. The label has to be earned. A Thriver must use those advantages (the system of record, the permissioned write-path, decades of proprietary exhaust) to become the agentic layer of its category, the place where AI actually does the work. Incumbents that skip that step get relegated to headless systems of record. A few will make that work as toll-collecting rails; most become data utilities that someone else’s agents read from and write to, with seat-based pricing eroding underneath them.
The argument runs on six principles.
Axiom 1 // The price of intelligence falls roughly 10x per year
a16z calls it LLMflation: for an LLM of equivalent performance, inference cost drops about 10x every year. GPT-3-class intelligence cost $60 per million tokens in November 2021; three years later, Llama 3.2 3B delivered the same benchmark performance for $0.06, a 1,000x collapse.2 Stanford’s AI Index found the same thing at the GPT-3.5 tier: from $20 per million tokens in November 2022 to $0.07 by October 2024, a 280-fold drop in under two years.3 Epoch AI, tracking six benchmarks, found prices to reach a fixed capability level falling between 9x and 900x per year, with a median of 50x. The fastest declines all came most recently.4
Cheapest available model at each intelligence level, by Artificial Analysis Intelligence Index band, redrawn from Artificial Analysis’ live trends dashboard, November 2022 to July 2026. Prices are AA’s 7:2:1 blend of cache, input, and output rates.5 Every band runs the same script: a capability level debuts expensive and collapses. The 10–20 band fell from $12 to $0.06. The 20–30 band fell from $15 to $0.04. The 40–50 band, frontier territory a year ago, fell over 60x in six months. Epoch AI measures the same pattern on public benchmarks: prices to reach a fixed capability fall 9x to 900x per year, with a median of 50x.4
The slope matters more than any single price point. Nothing an enterprise buys deflates at 10x a year. Which raises the obvious question: what stops the model vendors from simply holding price?
Axiom 2 // Open-source models are hugely capable and economical
The benchmark data backs this up. When DeepSeek released R1 in January 2025, an open-weight model matched OpenAI’s o1, the frontier reasoning model of the day, across math, coding, and science tests, at roughly 4% of the API price.6 The gap has only narrowed since. In July 2026 Moonshot released Kimi K3, at 2.8 trillion parameters the largest model any lab has disclosed and confirmed for release as a free open-weight download.1 It scored 57 on Artificial Analysis’ Intelligence Index, third of any model on earth, behind only Claude Fable 5 at 60 and GPT-5.6 Sol at 59. On Moonshot’s own benchmarks it beats Claude Opus 4.8 and GPT-5.5 across most coding and agentic tasks, at about a third of the leader’s cost per task.5,1 This is not only Moonshot’s telling: a16z’s Marc Andreessen recently called Z.ai’s GLM-5.2 the first Chinese model to “match and often beat the American big lab public AI models.”1
For years the received wisdom held that Chinese models ran eight to twelve months behind the US frontier. K3 lands within three points of the best model on earth the same week that model is current, and the strongest model you can download today, DeepSeek V4 Pro, already scores 44.5 Epoch AI, which tracks the open-closed gap formally, found the best open models trailed the frontier by about a year in 2025; by mid-2026 the lag was down to roughly four months.7 Today’s open model is last quarter’s frontier at a fraction of the price.
Intelligence vs. cost per task, redrawn from Artificial Analysis’ Intelligence Index v4.1, July 2026.5 Filled green points are open weights; white points are closed; the green ring is Moonshot’s Kimi K3, a confirmed open-weight release shipping imminently. K3 lands at 57, third of any model on earth, for about a third of the leader’s cost per task. Most shipped open models still cluster at the cheap left edge, led by DeepSeek V4 Pro at 44, but the open frontier now reaches into ground that was closed-only a year ago.
This is the answer to the question axiom 1 raised. Model vendors cannot hold price because a free alternative sits one quarter behind them on Hugging Face. CIOs now report the total cost of ownership of open and closed models converging, because the labs price against free.8 Enterprises benefit from Llama without ever running it: the free option sets the price of every closed bid.
The one place that umbrella still holds is the very top, where no open substitute exists yet. Anthropic will raise Claude Opus 4.8 by 50% in September, to $3 and $15 per million input and output tokens.1 That is a bet that the frontier premium is real and durable, and Moonshot placed its counter-bet the same month, putting an open model three points behind at a third of the price.
Open weights reach any given capability level about four months after the closed frontier does, per Epoch AI’s Capabilities Index, down from roughly a year in 2025.7 The buyer’s trade: wait one quarter, keep the capability, drop the price to a rounding error.
Axiom 3 // Most use cases don’t need a frontier model
Bugatti makes beautiful, expensive cars with little day-to-day utility. Frontier models are the same: stunning engineering, priced accordingly, and huge overkill for most jobs. The typical enterprise application needs an F-150: cheap, dependable, easy to service, and available everywhere. Open-source models are the F-150.
Now look at what enterprises actually run. Menlo’s production data shows only 16% of enterprise AI deployments are true agents; most are simple prompt-and-retrieval workflows wrapped around a single model call.9 These are marketing drafts, sales research, support ticket triage, finance reconciliation, and legal document review: high-volume, well-bounded tasks that an open model handles at F-150 cost. Coding is the one big exception where buyers pay up for the frontier.9 The pattern shows up in production choices. Airbnb runs Qwen for user-facing features, and Cursor built its in-house model on an open-source base.9 Even a16z’s Global 2000 CIOs, most of whom buy closed models, report that older, cheaper models “work well enough” for established workloads.8
Three of four quadrants default to open weights. Only deep, high-stakes reasoning clears the bar for frontier pricing, and Menlo’s production data says that is a minority of enterprise workloads today.9
One assumption deserves daylight: this axiom holds capabilities and use cases constant. If enterprises keep expanding into higher-value work that only the newest frontier models can do, the labs keep real pricing leverage at the top of the market. We accept that. It shifts which model gets the call. Ownership of the account stays where it was: frontier or open, every workload has to tap the system of record to act. The Thriver moats persist across use cases and model choices.
Axiom 4 // Margin flows to scarce complements
This is the oldest move in technology strategy: commoditize your complement. Cheap PCs made Windows valuable. Free Linux made AWS valuable. Free browsers made Google valuable. Worth noting: each of those winners was a new entrant. This cycle tests whether incumbents can run the same play, and axioms 5 and 6 say the right ones can. The commodity layer expands the market; the scarce layer captures it.
Cloud is the closest parallel, and the most instructive one. Amazon, Microsoft, and Google built magnificent businesses renting compute: global spend on cloud infrastructure reached roughly $320 billion in 2024.10 And still the larger share of value flowed upward, to the SaaS companies and enterprises that built on top of that compute. The SaaS industry alone clears $400 billion a year,11 and the operating value that businesses created by rebuilding on cloud never shows up on an infrastructure income statement. Foundation models are tracing the same arc. The labs are this cycle’s hyperscalers: enormous, necessary, capital-hungry, and selling a metered input. As token costs fall and open source spreads, we expect value to accrue the same way, up the stack to the Thrivers who build on top of the LLMs.
Intelligence is already behaving exactly like the textbook says. As unit prices collapsed 100x to 1,000x, enterprise generative AI spend went up from $1.7B in 2023 to $11.5B in 2024 to $37B in 2025.9 Jevons paradox is working: the average Global 2000 enterprise spent about $7M on LLMs in 2025 and expects roughly $11.6M in 2026, even as every token gets cheaper.8
Those aggregate numbers skew to the largest enterprises, and they hide the more important story. Falling inference costs push AI across the affordability threshold for thousands of mid-market companies that could never justify it at 2023 prices. The arithmetic is blunt: a ten-billion-token annual workload that cost $375,000 a year at GPT-4 launch prices runs for under $2,000 today at the same capability level.4 The demand side confirms the crossing. SMB adoption trackers show usage nearly doubling between 2024 and 2026, with the adoption gap between large enterprises and the mid-market narrowing from 1.8x to 1.2x.12 Yipit’s panel of roughly a thousand mid-market and enterprise companies shows adoption rates that mirror the Global 2000 survey,8 and more than a quarter of AI application spend now arrives through product-led motions that require no procurement department at all.9 Jevons has two halves: existing buyers consume more, and buyers who were priced out walk in. For a mid-market investor, the second half is the thesis. Cheap intelligence deepens the existing market and widens the addressable one at the same time.
Follow the dollars. The cloud pattern is already repeating: more than half of 2025 spend landed at the application layer. Applications out-earn the model APIs that power them by roughly 1.5 to 1.9
So the strategic question for every enterprise software company is: what, exactly, stays scarce when intelligence is free?
Axiom 5 // Code and intelligence are no longer the scarce assets
A frontier model has read the entire public internet. It has never read your customer’s claims history, ERP ledger, or clinical workflow, and it cannot sign a BAA. We see this in our own work: LLMs are at their most powerful as API calls inside highly tuned, mission-critical, data-intensive applications, because a foundation model needs context, rules, and workflows before it can produce a relevant answer. Those are the scarce inputs. Tokens are abundant. Run the inventory of what a raw model cannot supply, and notice that each item is a thing enterprise SaaS vendors have spent twenty years building:
| Moat | Why the model can’t cross it | The number |
|---|---|---|
| Proprietary data | Training corpora are public data; the valuable exhaust (transactions, tickets, claims, patient records) lives inside systems of record behind the firewall. | Epic holds ~44% of the U.S. acute-care EHR market and ~57% of hospital beds; Veeva holds an estimated ~80% of life-sciences CRM.13,14 |
| End-to-end workflows | Intelligence without permission to act is a chatbot. The durable write-paths are permissioned: regulatory sign-offs, controlled ledgers, steps with a high cost of error. Interface-centric paths (UI, navigation) are easier for agents to bypass. | ServiceNow’s AI suite hit $750M ACV entering 2026; its 2026 target was raised from $1B to $1.5B.15 |
| Vertical context | Domain logic (ICD codes, GAAP treatment, FDA process, Oil & Gas accounting) is encoded in the product, not the prompt. Vertical AI spend nearly tripled to $3.5B in 2025.9 | Healthcare alone: $1.5B, more than the next four verticals combined.9 |
| Security & compliance | SOC 2, HIPAA, FedRAMP are purchased with years and audits, not tokens. This is a literal, priced barrier to entry. | SOC 2 Type II: $20K–$150K and 6–12 months. FedRAMP Moderate: $1M–$2M+ and 12–18 months.16 |
| SLAs & trust | Enterprises buy guaranteed uptime, audit trails, and someone to sue. Models ship “as is.” | 65% of enterprises prefer incumbents when available, citing trust, integration, and procurement simplicity.8 |
| Distribution | Once anyone can build the feature, the constraint is who puts it in front of customers. The install base is the channel: AI ships as a line item on a contract that already exists, with no new procurement cycle. | Growth-stage software companies spend roughly a third of revenue on sales and marketing to build the reach an incumbent already owns.17 |
| Customer success in place | The renewal motion, the champion, the expansion path… already installed. AI features ride existing rails at near-zero CAC! | Median public SaaS gross retention ~90%; enterprise-segment NRR median ~110%+.17 |
The buy-side behavior confirms it. In 2024, enterprises split roughly 50/50 between building AI solutions and buying them. In 2025, 76% of AI use cases were purchased.9 AI deals convert to production at 47%, nearly double traditional software’s 25%.9 Enterprises are waiting for their vendors to show up with AI in the product.
Axiom 6 // Thrivers are going to Thrive
The winners will share two traits. First, high data gravity: they own data exhaust no one else has, produced automatically by operating the business. Think inspection logs, field records, transaction histories, and compliance filings built up over decades. Second, they own mission-critical workflows. That is the permissioned write-path: agents need permission to act, and the system of record is where permission lives.
We map every company we evaluate on these two dimensions. The result is the Covalent AI Thriver Matrix.18 Utilities hold valuable data with weak workflow lock-in; AI-native platforms can pull that data away from them. Specialists own niche workflows with thin data; horizontal AI players can crowd them out. Survivors score low on both counts and face displacement from below and from above. Thrivers sit in the top-right quadrant: data-rich, mission-critical, deeply embedded, and typically found in verticals where regulation and decades of workflow logic keep general-purpose AI out. The test is one question: does this company get more valuable as the models get smarter? The quadrants move. A Thriver that fails to build the agentic layer slides toward Utility: still data-rich, increasingly headless, its seat-based pricing eroding while someone else’s agents become the interface.
This is Covalent’s work. We identify the AI Thrivers, and we help incumbents build the agentic layer so they cross into the AI era in the top-right quadrant.
The Inversion
The thesis in one picture. Open-source models compress the price of intelligence toward zero at the base of the stack. Every layer above the commodity line is owned by the software vendor, and each is where pricing power now collects.
In 1999 the question was “what’s your internet strategy?” Back then the disrupters swept most incumbents aside: the value went to new entrants like Amazon and Google, who built the fresh distribution and owned the new data. In 2026 the question is “what’s your AI strategy?”, and this time the enterprise incumbents hold cards the consumer incumbents of 1999 never had: longitudinal data gravity, systems of record that take years to replace, customer and regulatory trust, and distribution that nothing routes around. The internet of 1999 was itself a new channel, and the disrupters used it to reach customers directly. AI ships with no channel of its own. It reaches the enterprise through software that is already installed. Enterprise AI is already a $37B market growing 3x a year inside a $400B+ SaaS market.9,11 The labs will keep spending billions to push the frontier, open source will keep making last quarter’s frontier free, and the companies that own the data, the workflow, and the trust will keep converting free intelligence into paid outcomes.
That is the paradox of open-source LLMs: the cheaper intelligence gets, the more the moats around it are worth.
- Constant-quality inference deflates about 10x per year. GPT-3-class intelligence fell 1,000x in three years.
- Open weights have closed to within three points of the best model on earth (Moonshot’s Kimi K3), at about a third of the cost, and set the floor price for every closed API.
- Only 16% of enterprise AI deployments are true agents. Most enterprise workloads run fine on an open-source F-150.
- The application layer already out-earns the model layer, $19B to $12.5B in 2025.
- Thrivers own the two things models cannot buy: data gravity and mission-critical workflow surface area.
- Thriver moats persist across use cases and model choices. Frontier or open, every workload has to tap the system of record to act.
// Sources & Further Reading
- Financial Times, coverage of Moonshot’s Kimi K3 launch (Jul 2026): 2.8T parameters, open-weight release, benchmark claims vs. Claude Opus 4.8 and GPT-5.5, Claude Opus 4.8 September price increase, and Marc Andreessen on GLM-5.2
- Guido Appenzeller, a16z — "Welcome to LLMflation" (Nov 2024)
- Stanford HAI — AI Index Report (2025 & 2026 editions)
- Epoch AI — "LLM inference prices have fallen rapidly but unequally across tasks" (Mar 2025); underlying model-level dataset on GitHub
- Artificial Analysis — Intelligence Index v4.1 and "GPT-5.6 benchmarks" (Jul 9, 2026); provider-published API prices (OpenAI, DeepSeek); AA Trends dashboard: inference price by Intelligence Index band (accessed Jul 2026); Kimi K3 scores and cost per task via Artificial Analysis (Jul 16, 2026)
- DeepSeek R1 launch benchmarks and pricing vs. OpenAI o1, Jan 2025 — summary
- Epoch AI — "Open models lag state-of-the-art closed models by 4 months" and "How behind are open models?" (2025–2026)
- Sarah Wang, Justin Kahl, Shangda Xu, a16z — "Leaders, gainers and unexpected winners in the Enterprise AI arms race" (Jan 2026)
- Menlo Ventures — "2025: The State of Generative AI in the Enterprise" (Dec 2025)
- Worldwide cloud infrastructure services spend, 2024: Canalys ($321B, +20%); Synergy Research ($330B)
- Gartner / Statista — worldwide SaaS spend ~$408B (2025), ~$465B forecast (2026)
- SMB and mid-market AI adoption, 2024–2026 — compiled adoption statistics; see also OECD, "AI adoption by SMEs" (Dec 2025) and JPMorganChase Institute
- KLAS data via CNBC and HealthsystemCIO (2025–2026)
- Veeva life-sciences CRM share estimates — IntuitionLabs
- ServiceNow disclosures — TNW, Fortune (2026)
- Compliance cost benchmarks — Vanta, A-LIGN, Paramify
- SaaS retention benchmarks — SaaS Capital, Benchmarkit 2025 (retention and sales & marketing spend benchmarks)
- Covalent Equity Partners — "Finding the AI Thrivers" (Insights Vol. 08)
- DeepSeek API pricing trackers, e.g. lmmarketcap.com (2026); directional, provider-published prices vary