Humanoid robotics brain-model development and generalization: Helix 2.5 shows scene generalization, but humanoid robots remain far from generalized commercial deployment
Bernstein views Figure AI's Helix 2.5 as evidence that proprietary human-video data and integrated brain models are improving humanoid performance. The report stresses that 56% task success across unseen homes is meaningful progress, not proof of transferable zero-shot task capability or commercial readiness.
Summary
Bernstein views Figure AI's Helix 2.5 as evidence that proprietary human-video data and integrated brain models are improving humanoid performance. The report stresses that 56% task success across unseen homes is meaningful progress, not proof of transferable zero-shot task capability or commercial readiness.
- Helix 2.5 completed three trained household tasks across 30 unseen homes with a 56% overall success rate.
- Index pretraining raised zero-shot success from 9% to 56% while halving the data needed to match Helix 02 performance.
- Bernstein argues that proprietary data and model-data integration are becoming a critical competitive moat.
- The next tests are transferable skills across related tasks and an efficient path from roughly 50–60% success toward 99% reliability.
Report Interpretation
Overview
This Global Automation note assesses Figure AI's Helix 2.5 demonstration as a concrete advance in humanoid-robot scene generalization. Bernstein argues that the result reinforces the importance of scalable proprietary data and tightly integrated robotic brain models, while leaving generalized tasks and mass deployment unresolved.
Core views
Figure AI's Helix 2.5 performed living-room tidying, towel folding and bed making in 30 homes for which it had no training data, achieving a 56% overall success rate. Bernstein characterizes this as scene generalization: the robot operated in previously unseen environments and handled full-body loco-manipulation, but it had been specifically trained for all three tasks. Object variation and arrangements across the homes were also limited. Accordingly, the report distinguishes this advance from genuine zero-shot task generalization, which it believes remains a long way off. The report attributes much of the improvement to the Index dataset, Figure AI's proprietary corpus of egocentric human task videos. Holding other variables constant, Index pretraining increased overall zero-shot success from 9% to 56% and halved the data required to match Helix 02 performance. As of August 2026, Index contained 23 million videos and was adding roughly 35 minutes of human experience per second; each 1,000 hours of collected data represented 373 tasks, 1,146 manipulated objects and 116 environments. Bernstein also cites continued declines in action-prediction loss as human-to-robot pretraining data expands, interpreting this as early evidence of scaling effects. Bernstein argues that the usable training-data universe for robotic models is expanding. Egocentric human video is increasingly suitable for pretraining, while real robot data remains valuable but teleoperation is likely to lose importance because it does not scale as readily. The report also sees deeper integration between data and models: instead of relying on fixed-weight vision-language models, Helix 2.5 is trained from scratch on proprietary data. It places Figure AI alongside AgiBot's GE-ACT 2.0, Dyna Robotics' Dyna 2.0 and Physical Intelligence's π0.7 as evidence of an industry trend toward broader data modalities, data-brain integration, scaling and generalization. The crucial next milestones are more demanding than the current demonstration. Bernstein wants evidence that a robot trained on one task can perform similar tasks without further training; it highlights Physical Intelligence's π0.7 as a notable example of compositional task generalization, in which untrained skills emerge from related learned skills. The report also stresses that a 56% success rate is not commercially deployable. It will watch for a post-training and deployment data flywheel that automatically improves performance, and views an efficient move from the 50–60% pretraining success range toward 99% reliability as the next major test. Overall, Bernstein believes generalized tasks and large-scale deployment beyond a limited set of applications are not imminent, although brain-model development is advancing rapidly quarter by quarter.
Analysis framework
Bernstein evaluates the Helix 2.5 demonstration by separating performance in unseen scenes from performance on untrained tasks, then compares pretraining results and dataset scale to isolate the role of proprietary human-video data. It broadens the assessment through a cross-company comparison of robotic brain-model architectures, data composition and reported scaling evidence, before identifying commercialization tests for future releases.
Methodology notes
Data scaling and data-brain integration in robotic model development
The report treats proprietary training data as a critical input to humanoid performance, comparing dataset scale, composition and pretraining outcomes to explain why larger and richer data supply may improve model capability.
Forward EV/EBITDA valuation for covered automation companies
For the covered equities, Bernstein applies target EV/EBITDA multiples to one-year-forward EBITDA estimates, with multiples referenced to prior cycles and adjusted for secular or competitive trends.
DCF as a reference for long-term intrinsic value
Bernstein uses DCF as a long-term valuation reference and notes that cycle-based price targets may differ from DCF-implied value.
Asset mapping & comparison
Structured mapping from thesis to named assets (strengths, weaknesses, peers, risks).
- Keyence (6861.JP)Covered automation company rated Outperform.
- Comparison
- Target price JPY95,000 versus JPY76,580 closing price; 22x EV/EBITDA on JPY913.2bn one-year-forward EBITDA.
- Risks
- Weaker manufacturing-capacity utilization, weaker global automation demand and JPY appreciation.
- Cognex (CGNX)Covered automation company rated Outperform.
- Comparison
- Target price USD85.00 versus USD61.99 closing price; 36.5x EV/EBITDA on USD372.6 million one-year-forward EBITDA.
- Risks
- Weaker global automation demand, slower emerging-customer progress, or delayed Moritex integration and synergies.
- FANUC (6954.JP)Covered automation company rated Outperform.
- Comparison
- Target price JPY7,200 versus JPY5,852 closing price; 23.0x EV/EBITDA on JPY255bn one-year-forward EBITDA.
- Risks
- Weaker global automation demand, market-share loss and JPY appreciation.
- Harmonic Drive Systems (6324.JP)Covered automation company rated Outperform.
- Strengths
- More than 50% global share in strain-wave reducers.
- Weaknesses
- Competitive landscape changes are particularly relevant given its market share.
- Comparison
- Target price JPY7,800 versus JPY6,200 closing price; 45.5x EV/EBITDA on JPY16,194 million one-year-forward EBITDA.
- Risks
- Weaker robot demand, market-share loss and JPY appreciation.
- Shenzhen Inovance Technology-A (300124.CH)Covered automation company rated Outperform.
- Comparison
- Target price CNY82.00 versus CNY54.05 closing price; 24.0x EV/EBITDA on CNY8,971.6 million one-year-forward EBITDA.
- Risks
- Weaker China automation demand, slower share gains outside servomotors and VFDs, and weaker EV demand.
- Estun Automation Co Ltd (2715.HK; 002747.CH)Covered automation company rated Market-Perform.
- Strengths
- Potential for faster-than-expected margin expansion and market-share gains.
- Weaknesses
- Cloos integration and planned synergies remain important to the thesis.
- Comparison
- H-share target HKD17.12 versus HKD14.80; A-share target CNY31.60 versus CNY29.68. Both targets use 45.5x EV/EBITDA on CNY681.4 million one-year-forward EBITDA.
- Risks
- Weaker China automation demand, slower margin improvement or market-share gains, and delayed Cloos integration and synergies.
Key data
- Helix 2.5 overall success rate56%Across three household tasks in 30 unseen homes.
- Zero-shot success with Index pretraining9% to 56%Improvement holding other variables constant.
- Index dataset scale23 million videosAs of August 2026; adding approximately 35 minutes of human experience per second.
- Index data diversity per 1,000 hours373 tasks, 1,146 manipulated objects, 116 environmentsReported composition of collected data.
- Keyence target priceJPY95,000Versus JPY76,580 closing price on 18 September 2026.
- Cognex target priceUSD85.00Versus USD61.99 closing price on 18 September 2026.
- FANUC target priceJPY7,200Versus JPY5,852 closing price on 18 September 2026.
- Harmonic Drive target priceJPY7,800Versus JPY6,200 closing price on 18 September 2026.
- Inovance target priceCNY82.00Versus CNY54.05 closing price on 18 September 2026.
- Estun target pricesCNY31.60 for A-shares; HKD17.12 for H-sharesVersus CNY29.68 and HKD14.80, respectively, on 18 September 2026.
Impact & implications
The report argues that proprietary data collection, pretraining scale and proprietary model training are becoming key sources of competitive differentiation in humanoid robotics. It supports a constructive view of automation suppliers under coverage while cautioning that current humanoid capability has not yet cleared the threshold for broad commercial deployment.
Risks
- Generalized task capability and broad deployment remain unproven; Bernstein considers the current 56% success rate insufficient for commercial deployment.
- The report identifies the challenge of finding an efficient path from roughly 50–60% pretraining success to 99% reliability.
- Covered automation companies face explicit risks from weaker global or China automation demand, industrial-capex cycles, competition, trade frictions and currency movements.
What to watch
- Whether future releases demonstrate transferable skills across related tasks without additional training.
- Evidence of a post-training and deployment data flywheel that automatically improves robot performance.
- Progress from current pretraining success rates toward 99% reliability.
- Further evidence that larger human-video datasets improve robot performance on unseen tasks.