Open-weight models pressure token prices, but compute scarcity still supports cloud vendor returns
AI summary card
Open-weight models pressure token prices, but compute scarcity still supports cloud vendor returns
Morgan Stanley believes open-weight models will intensify price competition at the model layer, but expanding inference demand, improved throughput efficiency, and incremental cloud infrastructure service revenue can allow hyperscale cloud vendors to continue earning healthy returns.
- Open-weight models can be downloaded, fine-tuned on private data, and deployed on-premises, in the cloud, or via APIs, thereby helping reduce costs and expand AI adoption.
- Model labs face downward pressure on per-token pricing, increasing the importance of model architecture efficiency, throughput, differentiation, and ecosystem development.
- Under assumptions of approximately $1.75 per million tokens and approximately 2,000–3,500 tokens per second per GPU, 1 gigawatt of owned GB300 infrastructure can still generate approximately 20%–60% ROIC.
- Compute scarcity, demand growth, hyperscalers' throughput optimization capabilities, and incremental revenue from other cloud services are the four core supports for keeping their unit economics attractive.
- The report maintains Overweight views on META, AMZN, and GOOGL, with META as the top pick.
Report interpretation
Overview
The report examines the impact of the rapid adoption of open-weight models on the unit economics and return on invested capital of the generative AI industry. The core conclusion is that lower per-token prices will create direct pressure on model providers, but will also promote adoption and growth in inference volume; hyperscale cloud vendors, leveraging scarce compute, pricing power, custom chips, software optimization, enterprise customer bases, and cloud service cross-selling, are still expected to earn healthy returns.
Core views
First, open-weight models expand model distribution channels through free downloads, fine-tuning on private data, and flexible deployment, accelerating technology diffusion. Second, price competition at the model layer will intensify, and labs must rely on higher token throughput, stronger model differentiation, and broader ecosystems to sustain revenue and returns. Third, hyperscale cloud vendors control scarce and capital-intensive compute resources and can monetize them through rational pricing and higher utilization; the volume growth brought by lower-priced models may matter more than margin percentage. Fourth, beyond API revenue, open models will also drive revenue from accelerator leasing, storage and databases required for retrieval-augmented generation, as well as security, governance, identity, and monitoring services. Fifth, the lowest cost to serve will become a key competitive factor, and cloud vendors with custom chips, first-party models, and stronger hardware-software co-optimization capabilities will have greater advantages.
Analysis framework
The report analyzes model labs and hyperscale cloud vendors by layer: for the model layer, it evaluates returns based on price per million tokens, token throughput per GPU, the share of capacity used for inference, and revenue per gigawatt; for the infrastructure layer, it estimates EBIT, NOPAT, and ROIC based on the number of GPUs, utilization, hourly rental prices, depreciation, energy, and other operating costs. It also maps assets by combining the business drivers, scenario valuations, and risk-reward profiles of META, AMZN, and GOOGL.
Methodology notes
Estimate revenue per gigawatt and return on invested capital based on token pricing, throughput, and inference capacity
The framework assumes a 1-gigawatt data center is equipped with approximately 410,256 GB300 chips, and estimates annual revenue based on the share of capacity used for inference, tokens per second per GPU, and price per million tokens, then compares it with compute leasing costs.
Estimate NOPAT and ROIC based on utilization, hourly rental price, and capital costs
The framework uses a 75% utilization rate and incorporates costs such as IT depreciation, non-IT depreciation, energy, labor, and maintenance; different rental price scenarios correspond to ROIC of approximately 23%–39%.
Compare model control, customization capabilities, deployment methods, costs, and distribution paths
Open-weight models allow users to obtain model weights, fine-tune them on private data, and independently choose deployment locations; closed-source models are typically accessible only through subscriptions or APIs.
Form target price ranges based on future-year EPS and P/E multiples
The price targets for META, AMZN, and GOOGL are primarily based on the average EPS for 2027 and 2028 and corresponding P/E multiples, with scenarios constructed using differences in revenue growth, margins, and AI monetization.
Asset mapping & comparison
Structured mapping from thesis to named assets (strengths, weaknesses, peers, risks).
- META PLATFORMS INC (US.META)Open-weight model provider, AI infrastructure investor, and digital advertising platform; the report names it as the top pick and assigns an Overweight rating.
- Strengths
- AI investment is expected to improve Reels engagement and advertising monetization, while ad measurement and attribution continue to improve; an efficiency-oriented organizational transformation, subscriptions, and click-to-message businesses provide additional upside optionality.
- Weaknesses
- High data center and AI capital expenditures increase capital intensity, and the low-price strategy for open models may pressure direct monetization capability at the model layer.
- Comparison
- Compared with pure cloud infrastructure vendors, META realizes AI returns more indirectly through advertising, user engagement, and new products, while also having distribution advantages from its open model ecosystem.
- Risks
- Reels monetization below expectations, declining user engagement, advertising targeting regulation, weak macro consumption, widening losses at Reality Labs, and poor execution in data center construction leading to lower returns.
- AMAZON COM INC (US.AMZN)Captures inference demand from open-weight models through AWS, Bedrock, accelerator leasing, and related cloud services; the report assigns an Overweight rating.
- Strengths
- AWS is in a long-term cloud adoption cycle and can achieve multilayer monetization through compute leasing, hosted APIs, storage, databases, and security services; high-margin businesses support continued investment.
- Weaknesses
- Both cloud and retail businesses require sustained high investment, and short-term profits are relatively sensitive to investment pace, fulfillment costs, and AWS growth.
- Comparison
- Compared with META, AMZN benefits more directly from cloud infrastructure usage driven by open models; compared with GOOGL, its advantages are concentrated in AWS scale, enterprise coverage, and a broad cloud service portfolio.
- Risks
- AWS revenue deceleration or margin decline, higher-than-expected investment intensity, deterioration in merchandise margins, and insufficient operating leverage in delivery and fulfillment.
- ALPHABET INC (US.GOOGL)Owns Gemini models, Google Cloud, TPU custom chips, search, and YouTube distribution channels; the report assigns an Overweight rating.
- Strengths
- Custom TPUs and model-cloud platform synergy help reduce cost to serve; efficient models such as Gemini Flash, Google Cloud backlog, and commercialization of Search and YouTube provide diversified growth sources.
- Weaknesses
- AI search and low-priced models may increase compute intensity and pressure margins while monetization of new products is not yet mature.
- Comparison
- GOOGL has attributes of both a model lab and a hyperscale cloud vendor; compared with pure model providers, it has infrastructure and distribution advantages, and compared with AMZN, it has stronger first-party model, search, and video ecosystem synergies.
- Risks
- Slowing search advertising growth, AI products cannibalizing core search, low monetization rates of new products, rising compute costs, insufficient expense discipline, and cloud business acceleration falling short of expectations.
Key data
- Baseline token price for model APIsApproximately $1.75/million tokensThe baseline assumption used in the report, which is expected to potentially decline further over the long term.
- Inference throughput assumptionApproximately 2,000–3,500 tokens/second/GPUBased on public inference benchmarks for the open-weight DeepSeek-V4-Pro 1.6T on GB300 hardware.
- Model-layer return on invested capitalApproximately 20%–60%The estimated result for 1 gigawatt of owned GB300 infrastructure.
- Number of GB300 chips per 1 gigawattApproximately 410,256Morgan Stanley estimate.
- Hyperscale cloud vendor GPU leasing returnApproximately 23%–39%Based on 75% utilization and an hourly rental price range of $7–$10.
- META price target$775Approximately 23x the average 2027 and 2028 EPS of $34/$35; bull and bear case scenarios are $1,000 and $450, respectively.
- AMZN price target$335Approximately 25x the average 2027 and 2028 EPS of $14; the bull case scenario is $410.
- GOOGL price target$400Approximately 24x the average 2027 and 2028 EPS of $15/$18; bull and bear case scenarios are $450 and $225, respectively.
Impact & implications
The main impact of open-weight models is not to eliminate returns on AI infrastructure, but to redistribute profits across the value chain: model labs face more visible pricing and gross margin pressure and must offset it through throughput gains, model innovation, and product differentiation; hyperscale cloud vendors may benefit from larger inference workloads, GPU leasing, and incremental services such as storage, databases, and security. From an investment perspective, the focus should be on compute owners with low cost to serve, custom chips, strong software scheduling capabilities, first-party models, and large volumes of enterprise requests.
Risks
- Competition from open-weight models may drive per-token prices down faster than throughput efficiency improves, compressing model-layer ROIC.
- AI demand or GPU utilization below expectations may weaken data center returns for hyperscale cloud vendors.
- Capital expenditures, energy, depreciation, and maintenance costs above assumptions may cause actual unit economics to fall below model estimates.
- Efficiency improvements in model architecture, custom chips, interconnects, and request batching may fall short of expectations.
- New AI products may cannibalize search or other high-margin businesses and increase compute burdens at lower monetization rates.
- Weak macro consumption and advertising spend, regulatory restrictions, and slower enterprise adoption may affect the revenue of related companies.
- The returns in the report are assumption-driven scenario estimates and are highly sensitive to token prices, throughput, utilization, and rental prices.
What to watch
- Changes in per-million-token prices for open-weight and near-frontier models.
- Token throughput per GPU, share of capacity used for inference, and request batching efficiency.
- The release cadence, capability improvements, and ecosystem adoption of META's upcoming Muse model.
- Progress of Google Gemini Flash and Gemini 4 in efficiency, performance, and commercialization.
- AI revenue, backlog, utilization, and margins at AWS, Google Cloud, and other cloud platforms.
- Cost and performance advantages of custom chips such as TPU and Trainium versus general-purpose GPUs.
- Incremental revenue contribution from open models to storage, databases, retrieval-augmented generation, security, and governance services.
- Capital expenditures, free cash flow, and AI investment return disclosures from META, AMZN, and GOOGL.