Quick Summary
Covering the latest research from top Wall Street investment banks

Token usage continues to drive AI spending growth and memory prices higher, while GPU rental pricing diverges

Institution
JPMorgan
Date
20260826
Authors
Joseph Cardoso, Manmohanpreet Singh, Akanksh Chauhan
Company
Ticker
Industry
AI computing infrastructure, data center hardware, and networking equipment
Rating
BullishMedium confidenceShort-termThe report believes that continued growth in LLM token spending and strengthening DRAM and NAND prices collectively point to a persistently positive demand environment for AI infrastructure, despite mixed GPU rental pricing.
AuthorsJoseph Cardoso, Manmohanpreet Singh, Akanksh Chauhan
Research firm divisions/subsidiariesJ.P. Morgan Securities LLC(Subsidiary/Legal Entity)

AI summary card

Token usage continues to drive AI spending growth and memory prices higher, while GPU rental pricing diverges

JPMorgan's August data show that OpenRouter token usage increased 47% m/m, driving a 7% m/m rise in spending; DRAM and NAND spot prices rose significantly. GPU performance was mixed, with H100 prices recovering while B200 recorded its first m/m decline in more than six months.

AI computing infrastructureData centersLLM tokensOpenRouterGPU rentalsNvidiaDRAMNAND
  • August token usage increased 47% m/m and reached 28x y/y, while token spending grew 7% m/m and reached 12x y/y.
  • Discounts and adoption of lower-cost models drove the volume-weighted token price down 28% m/m and 56% y/y.
  • H100 rental prices rose 1.2% m/m, while B200 prices declined 1.5% m/m, marking the latter's first decrease in more than six months.
  • DDR5 16Gb spot prices rose to $50.20, up 6% m/m and 880% y/y.
  • NAND 1Tb spot prices rose to $30.50, up 14% m/m and 470% y/y.

Report interpretation

Overview

The report assesses August AI infrastructure demand using samples of LLM usage and pricing from OpenRouter, GPU rental prices from non-hyperscale cloud providers, and DRAM and NAND spot prices. It concludes that overall demand trends remain positive: token spending and memory prices continue to rise, although GPU rental pricing diverges by model.

Core views

The report first evaluates AI infrastructure demand through LLM token activity. Its tracker samples one week of data each month, covering approximately 300 LLMs on the OpenRouter platform. August token usage reaccelerated after slowing in July, increasing 47% m/m, above July's 18% but below June's 70%; it reached 28x y/y, compared with 21x and 20x in July and June, respectively. Closed-weight models accounted for 33% of total usage, up from 29% in July and unchanged from 33% in June; their usage increased 66% m/m, versus 3% and 34% in July and June, respectively, and reached 14x y/y. Open-weight model usage increased 39% m/m, compared with 26% and 97% in July and June, respectively; it reached 56x y/y, below July's 63x and June's 74x. Usage growth was accompanied by price declines. The average token price decreased 5% m/m in August, versus a 3% decline in July and a 7% increase in June; it increased 5% y/y, improving from 2% in July and -5% in June. The volume-weighted price, which better reflects the actual model mix, declined 28% m/m and 56% y/y, mainly due to a 50% price cut for GPT-5.6 Sol, an 80% price cut for GPT-5.6 Luna, and widespread adoption of low-cost models. Tencent Hy3, DeepSeek v4 Flash-0731, and GPT-5.6 Luna together contributed approximately 99% of the incremental m/m token volume. The volume-weighted price of open-weight models decreased 17% m/m, compared with increases of 36% and 12% in July and June, respectively; it still rose 60% y/y. The volume-weighted price of closed-weight models fell 37% m/m and 29% y/y, both substantially weaker than in previous months. Despite falling prices, strong usage expansion continued to drive token spending growth. Total spending increased 7% m/m in August, unchanged from July and below June's 70%; it reached 12x y/y, compared with 11x and 16x in July and June, respectively. Closed-weight models accounted for 78% of total spending, down from 80% in July and 88% in June; their spending increased 4% m/m and reached 10x y/y. Spending on open-weight models increased 16% m/m, slowing markedly from 71% in July and 120% in June, but still reached 89x y/y. This indicates that open-weight models contributed faster usage expansion, while higher-priced closed-weight models continued to account for the vast majority of spending. Model rankings further illustrate the difference between usage and revenue composition. By token usage, the top five models in August were DeepSeek v4 Flash-0731, Tencent Hy3, MiMo v2.5, GPT-5.6 Luna, and DeepSeek v4 Flash, together accounting for more than 50% of total usage. By token spending, the top five were Claude Opus 5, Kimi K3, Claude Opus 4.7, GPT-5.6 Sol, and Claude Fable 5, together accounting for 48% of total spending. There was no overlap between the two top-five groups, indicating that low-cost models dominate usage while higher-priced models continue to dominate spending. GPU rental prices diverged by model. In the non-hyperscale cloud provider market, the average August rental price for the A100 was $1.65 per GPU-hour, unchanged m/m, following increases of 0.9% and 6.4% in July and June, respectively. The H100 price was $2.73 per GPU-hour, recovering 1.2% m/m after declining 0.6% in July and rising 3.5% in June; the H100-to-A100 price ratio was 1.65x, compared with 1.64x and 1.66x in July and June, respectively. The B200 price was $5.63 per GPU-hour, down 1.5% m/m, its first m/m decline in more than six months, following increases of 7.2% and 2.7% in July and June, respectively. The B200-to-H100 price ratio declined to 2.06x, below July's 2.12x but above June's 1.97x, and substantially below the 2.58x level around the index's launch in September 2025. Memory prices provided a more consistent upward signal. The DDR5 16Gb spot price rose to $50.20 in August, up 6% m/m from $47.00 in July and 880% y/y from $5.09 in August 2025. This marked the fifth consecutive monthly increase following modest m/m declines in February and March, which the report characterized as an increase of more than 8x y/y. The NAND 1Tb spot price rose to $30.50, up 14% m/m from $26.90 in July and 470% y/y from $5.29 in August 2025. This was the first return to growth after four consecutive months of modest declines, which the report characterized as an increase of more than 4x y/y. Overall, expanding token spending and rising memory prices support a positive assessment of AI infrastructure demand, while the decline in B200 rental prices indicates that pricing momentum is not uniform across computing products.

Analysis framework

The report sequentially analyzes AI application activity, computing capacity rentals, and memory prices. It first decomposes token spending into usage and price, comparing m/m and y/y changes across all models, open-weight models, and closed-weight models. It then examines the non-hyperscale cloud computing market using rental prices and relative price ratios across GPU models. Finally, it assesses memory demand and pricing trends through DRAM and NAND spot prices, combining the three groups of indicators to evaluate the AI infrastructure demand environment.

Methodology notes

  • Industry/sector analysis frameworkVolume-price decomposition

    Decomposition of token spending into usage and price

    The report separately tracks token usage, average price, volume-weighted price, and total spending to determine whether spending growth is driven by usage expansion or changes in unit pricing.

  • Industry/sector analysis frameworkSupply-demand framework

    Supply-demand signals from GPU rentals and memory spot prices

    The report uses GPU rental prices and DRAM and NAND spot prices as indicators of supply-demand conditions in the computing and memory markets, comparing the direction and magnitude of price changes across products.

  • (Out-of-vocabulary methodology)

    OpenRouter monthly one-week snapshot tracking

    The tracker samples one week of data each month, covering approximately 300 LLMs on OpenRouter, to produce a monthly snapshot of usage and pricing trends. The sample is skewed toward developers, startups, and agentic coding traffic and excludes first-party API usage from model providers and hyperscale cloud platforms.

Key data

  • August token usage+47% m/m, 28x y/yM/m growth reaccelerated from 18% in July; y/y growth increased from 21x in July to 28x.
  • August average token price-5% m/m, +5% y/yThe price continued to decline m/m, but y/y growth improved from 2% in July.
  • August volume-weighted token price-28% m/m, -56% y/yPrimarily affected by substantial model price cuts and increased adoption of lower-cost models.
  • August token spending+7% m/m, 12x y/yUsage expansion offset the decline in the volume-weighted price.
  • Closed-weight model share of spending78%Down from 80% in July and 88% in June.
  • Average A100 rental price$1.65/GPU-hourUnchanged m/m in August.
  • Average H100 rental price$2.73/GPU-hourUp 1.2% m/m in August, with the H100/A100 price ratio at 1.65x.
  • Average B200 rental price$5.63/GPU-hourDown 1.5% m/m in August, marking the first m/m decline in more than six months; the B200/H100 price ratio was 2.06x.
  • DDR5 16Gb spot price$50.20Up 6% m/m and 880% y/y in August, marking the fifth consecutive monthly increase.
  • NAND 1Tb spot price$30.50Up 14% m/m and 470% y/y in August, ending four consecutive months of m/m declines.

Impact & implications

The report believes that rapid expansion in token usage is sufficient to continue driving spending growth despite falling prices, indicating resilient demand for AI model usage. At the same time, rising DRAM and NAND spot prices reinforce signals of improving memory demand and pricing conditions. However, the first m/m decline in B200 rental pricing and its narrowing premium over the H100 indicate that GPU pricing trends vary by product generation, meaning that overall demand improvement should not be equated with simultaneous price increases across all computing products.

Risks

  • OpenRouter data are skewed toward developers, startups, and agentic coding traffic and exclude usage from first-party APIs such as OpenAI and Anthropic, as well as hosted endpoints on hyperscale cloud platforms. Therefore, the sample is not representative of the entire LLM market.

What to watch

  • Continue monitoring whether token usage growth can offset declines in average and volume-weighted prices, allowing total token spending to keep growing.
  • Monitor changes in usage, pricing, and spending shares between open-weight and closed-weight models.
  • Track whether B200 rental prices continue to decline, as well as changes in the B200/H100 and H100/A100 price ratios.
  • Observe whether DRAM can sustain consecutive monthly increases and whether NAND can maintain growth after ending four months of declines.
Zhejiang ICP No. 2022035445-5
Disclaimer: Market data, charts, indicators, research views, and other information provided on this website are intended solely for information display, research communication, and educational reference. They should not be regarded as personalized investment advice, securities recommendations, trading instructions, solicitations, or guarantees of return. While we strive to improve the reliability of our data and content, such information may still be subject to delays, errors, incompleteness, or untimely updates due to source differences, methodological limitations, system processing, or market volatility. Users should exercise independent judgment based on their own circumstances and bear all risks and responsibilities arising from the use of this website.

Settings

Sign in to view recent logins