Anthropic and OpenAI continue to lead frontier AI competition, while Token price cuts and memory inflation become key variables for internet AI valuations
AI summary card
Anthropic and OpenAI continue to lead frontier AI competition, while Token price cuts and memory inflation become key variables for internet AI valuations
August data indicate that model capabilities and GPU demand remain strong, but inference-pricing competition is intensifying and DRAM costs are rising; investors should monitor new model launches by hyperscale cloud providers and cost pass-through.
- Claude Opus 5 leads several rankings, including intelligence, agentic intelligence, and coding-agent intelligence, with GPT-5.6 Sol close behind.
- On Vercel, DeepSeek's month-to-date August Token share is 29.7%, while Anthropic's spending share is 64.8%.
- The AI Token Price Index stood at $2.21 per million Tokens as of August 13, down roughly 6% to 9% month over month but still up 87% year over year.
- Rental prices for B200, H100, and A100 GPUs have generally remained resilient, indicating that AI infrastructure demand remains healthy.
- DDR5 spot prices rose roughly 8% month over month and 483% year over year in August, with memory-cost pressures yet to ease.
Report interpretation
Overview
This report launches the BofA Frontier AI Data Tracker for the first time, covering model-capability rankings, model usage and spending shares, Token pricing, GPU rental prices, and memory costs to assess the AI competitive landscape, returns on capital expenditures, and valuation sentiment for large internet companies.
Core views
Anthropic and OpenAI currently set the pace in frontier-model capabilities, while upcoming model launches from Meta and Google could be catalysts. Lower-cost open models have recently gained some usage share, showing that developers continue to weigh capability against cost. Token prices retreated in August after strengthening significantly through July, leaving cloud inference margins exposed to more intense price competition; GPU rental prices remain solid, indicating healthy AI infrastructure demand and server useful lives. Meanwhile, DRAM and NAND prices remain elevated, with no sign that memory-cost pressures are easing.
Analysis framework
The report combines data from Artificial Analysis, BenchLM, OpenRouter, Vercel, Silicon Data, and spot markets to assess model capabilities, developer routing usage, platform spending, inference costs, GPU supply and demand, and memory costs. Relevant platform data primarily reflect traffic routed through their gateways or routers and do not represent total market usage.
Methodology notes
Frontier-model intelligence and task cost
Model capabilities are assessed through intelligence, agentic intelligence, and coding-agent intelligence indices; task cost per unit of intelligence is used to compare the effective inference cost of achieving similar capabilities, with lower values indicating better cost efficiency.
Token usage share and spending share
OpenRouter and Vercel data can reflect developer mindshare, cost-effectiveness trade-offs, and commercial traction, but exclude substantial workloads from direct model-provider API calls, other cloud platforms, private deployments, and enterprise-owned data centers.
Blended cost per million Tokens
The Token Price Index tracks changes in average list prices among major model providers; the LLM Token Expenditure Index measures effective inference costs across a broader ecosystem using usage weights.
Standardized GPU hourly rental prices
Based on data from neoclouds, hyperscale clouds, managed providers, brokered clusters, and private rental platforms, this indicator reflects AI GPU supply and demand and compute-rental pricing.
Asset mapping & comparison
Structured mapping from thesis to named assets (strengths, weaknesses, peers, risks).
- METAAI lab progress and AI monetization across the internet advertising platform
- Strengths
- MuseSpark 1.2 ranks seventh in the Intelligence Index, indicating progress in model capabilities; the report maintains a Buy view.
- Weaknesses
- Model capabilities still trail leading Anthropic and OpenAI products, while AI investment may pressure margins.
- Comparison
- It remains in catch-up mode versus Anthropic and OpenAI; compared with Google Gemini 3.6 Flash, MuseSpark 1.2 ranks higher in the Intelligence Index.
- Risks
- Digital advertising sensitivity to macro conditions, rising AI capital expenditures, increased fixed assets, competition from AI-native platforms for users and advertising budgets, and regulatory and litigation risks.
- GOOGL / GOOGBeneficiary of AI through Gemini models, search, and cloud operations
- Strengths
- Strong AI assets, expanding cloud margins, and expected double-digit revenue growth support valuation; future Gemini releases could serve as catalysts.
- Weaknesses
- Gemini 3.6 Flash trails leading models in relevant intelligence rankings, and Vercel Token share declined in August.
- Comparison
- Both model capabilities and platform usage share face competition from Anthropic, OpenAI, and lower-cost open models.
- Risks
- Search traffic diversion to AI tools, weaker-than-expected monetization of LLMs embedded in search, EU DMA compliance pressure, and AI capital expenditures increasing and depressing free cash flow.
- AMZNBeneficiary of AWS inference and AI infrastructure demand
- Strengths
- Resilient GPU rental prices support AI infrastructure demand, and AWS has a foundation for cloud-service monetization.
- Weaknesses
- Token pricing competition and higher infrastructure costs may pressure cloud margins and free cash flow.
- Comparison
- Relative to other hyperscale cloud providers, the competitive focus is on model ecosystems, inference costs, and cloud customer workloads.
- Risks
- Competition among cloud providers and large retailers, higher AWS infrastructure costs, agentic AI diverting direct traffic and high-margin advertising revenue, and margin volatility driven by macroeconomic uncertainty.
Key data
- Model capability rankingsClaude Opus 5 ranks first, with Claude Fable 5 and GPT-5.6 Sol also among the leadersMeta MuseSpark 1.2 ranks seventh in the Artificial Analysis Intelligence Index; Gemini 3.6 Flash ranks thirteenth.
- Open-model Token usageMiMo-V2.5 at 328 trillion Tokens; DeepSeek V4 Flash at 264 trillion TokensOpenRouter month-to-date August data.
- Closed-model Token usageGPT-5.6 Luna at 86 trillion Tokens; Claude Opus 4.8 at 48 trillion TokensOpenRouter month-to-date August data.
- Vercel Token shareDeepSeek 29.7%, Anthropic 24.7%, OpenAI 16.3%, Google 5.1%Month-to-date August average; DeepSeek and OpenAI gained share month over month, while Google and Anthropic declined.
- Vercel spending shareAnthropic 64.8%, OpenAI 11.2%, Google 7.8%, Moonshot.ai 6.4%Month-to-date August average; Anthropic was flat versus July.
- AI Token Price Index$2.21/million TokensDown roughly 6% to 9% month over month as of August 13 and up 87% year over year; OpenAI price cuts and Gemini 3.6 Flash replacing older versions are potential drivers.
- LLM Token Expenditure Index$1.15/million TokensDown 27% month over month in August, versus $1.57 in July.
- GPU rental pricesB200 $5.63/hour, H100 $2.77/hour, A100 $1.65/hourIn August, B200 was down 2% month over month, H100 up 2%, and A100 flat.
- Memory pricesDDR5 approximately +8% month over month; NAND flat month over monthDDR5 was approximately +483% year over year and NAND approximately +432% year over year; cost pressures remain high.
Impact & implications
For large internet stocks, improved model capabilities can strengthen product competitiveness and the AI narrative, but Token price cuts will test the unit economics of cloud services and inference businesses. Firm GPU rental prices support expectations for AI infrastructure demand, while also implying that capital-expenditure and depreciation pressures may persist; rising memory prices could further compress margins for hardware and cloud infrastructure. Meta's improved model ranking is a positive signal, but its AI investment, advertising competition, and regulatory risks still require parallel assessment.
Risks
- Platform usage data have limited coverage and cannot be extrapolated as industry-wide model usage shares.
- If Token price cuts persist, they could erode margins for cloud inference and model services.
- Rising DRAM and NAND prices could increase server and cloud infrastructure costs.
- High AI capital expenditures, fixed-asset expansion, and changes in server depreciation assumptions could depress free cash flow.
- Model iteration, competition from open models, and regulatory changes could rapidly reshape the competitive landscape.
What to watch
- New model launches by hyperscale providers, including Meta Watermelon and Google Gemini 4.
- Token pricing adjustments by OpenAI, Anthropic, Google, and open-model providers.
- Whether Token share for lower-cost open models continues to increase on OpenRouter and Vercel.
- Whether GPU rental prices can remain resilient, and price changes among B200, H100, and A100.
- Whether DRAM and NAND spot prices begin to ease, and their impact on cloud-provider capital expenditures and gross margins.
- Earnings guidance from large internet companies on AI infrastructure investment, inference demand, and monetization progress.