Agentic AI for dynamic factor investing: Selective LLM committees show promise for European factor rotation, but QMI remains the benchmark
J.P. Morgan finds that tightly scoped LLM agents can turn broad macro data into useful regime classifications and positive factor-rotation returns. The best two-model committee improves balance and drawdowns, although residual look-ahead bias and pessimistic model skew prevent it from displacing QMI.
Summary
J.P. Morgan finds that tightly scoped LLM agents can turn broad macro data into useful regime classifications and positive factor-rotation returns. The best two-model committee improves balance and drawdowns, although residual look-ahead bias and pessimistic model skew prevent it from displacing QMI.
- All tested LLM regime signals generated positive excess returns in the 1994-2026 European backtest.
- The GPT 5.5 and Opus 4.8 committee delivered 6.0% annualised excess return with a 0.68 Sharpe ratio and -10.6% maximum drawdown.
- The selective two-member committee matched QMI regimes 48.1% of the time versus 39.1% for the six-member committee.
- Date anonymisation reduces but does not eliminate the risk that models recall historical macro episodes.
- Both committee approaches called Expansion for September 2026, while J.P. Morgan's QMI remained at Slowdown in the US and Contraction in Europe.
Report Interpretation
Overview
This European quantitative-strategy study tests whether LLM agents can classify monthly macro regimes and rotate equity factor portfolios accordingly. J.P. Morgan concludes that the approach is encouraging as a complementary input, particularly when using newer, selectively chosen models, but not as a replacement for its QMI systematic framework.
Core views
The report addresses whether a narrowly instructed LLM can read a broad macro context pack, assign probabilities to Expansion, Slowdown, Contraction and Recovery for the following month, and translate those calls into profitable European equity factor rotation. The motivation is that individual equity styles can suffer severe drawdowns during regime shifts; rotating between cycle-specific factor portfolios may preserve upside while reducing exposure before leadership changes. J.P. Morgan uses its European QMI as both the regime benchmark and the established systematic reference for style rotation. The agent receives a generalist macro pack rather than hand-selected signals. Inputs span growth, labour, inflation, policy and financial conditions across the US, UK and Europe, including PMIs, unemployment, CPI/HICP, policy rates, yield curves, real yields, the US dollar, oil, volatility and credit spreads. Each series includes its level, changes over 1/3/6/12 months as applicable, and a 36-month rolling z-score. The four regimes are defined using both the level of activity relative to trend and whether conditions are improving or deteriorating, emphasizing that the second derivative of macro data can matter more for expected style leadership than the level alone. The report treats look-ahead bias as the central implementation risk. Because LLMs may recognize historical episodes from training data, simply using lagged and non-restated data cannot guarantee an out-of-sample backtest. In ablation tests using Opus 4.8, regime paths remained coherent when dates were removed and when both dates and raw levels were removed, suggesting the model can use the remaining macro evidence. However, a date-only input still produced plausible historical regime patterns, particularly around the dot-com episode, the Global Financial Crisis and COVID, indicating residual historical memory. Replacing familiar regime names with neutral labels left macro-informed calls broadly unchanged: agreement was 80%-83% and Cohen's kappa was 0.71-0.76 across the three macro-data variants, versus 59% agreement and 0.45 kappa when only the date remained. J.P. Morgan therefore concludes that the agent primarily reasons from macro data when it has genuine evidence, but that leakage cannot be eliminated. The study tests GPT 5.5, GPT 5.4, GPT 5.1, Claude Opus 4.8, Sonnet 4.6 and Haiku 4.5 using date-anonymised packs, low-reasoning mode and temperature of 1. Because outputs are stochastic, the researchers run each model 100 times per month and use averaged probabilities or majority-style aggregation. Bootstrap resampling of 200 observed calls per date finds that consensus dates stabilize quickly, while ambiguous dates may require about 100 runs to reach the tightest error threshold; gains beyond that point are marginal. The report identifies systematic directional bias as a first-order issue: several older models over-select Slowdown and Contraction. Haiku 4.5 spent more than 80% of the sample in negative regimes and never called Recovery, while GPT 5.5 and Opus 4.8 generated a more balanced regime mix closer to QMI. Ensembles improve results only when their members are sufficiently capable and balanced. Pooling all six models retained the shared pessimistic skew, whereas the selective GPT 5.5-plus-Opus 4.8 committee aligned more closely with QMI. Across 391 monthly observations, the two-member committee matched QMI 48.1% of the time, compared with 39.1% for the six-member group. Its Expansion precision and recall were 46.0% and 36.4%, respectively, compared with 36.1% and 11.8% for the larger committee; its Recovery precision and recall were both 47.4%, versus 33.7% and 39.5%. The trade-off is that the two-member committee changes regimes more frequently, implying higher turnover and transaction costs, whereas the six-member committee is mechanically smoother. For implementation, the report rotates among four cycle-phase equity portfolios using each model's monthly regime call, subject to a two-month confirmation rule at regime transitions to reduce turnover. In the 1994-2026 MSCI Europe excess-long backtest, every LLM strategy generated positive excess returns, and newer models generally produced stronger realized outcomes. The two-member committee produced the highest agent annualized return at 6.0%, a 0.68 Sharpe ratio, a 60.3% hit rate and the tightest maximum drawdown at -10.6%. The six-member committee returned 5.6% with a 0.72 Sharpe ratio, 60.8% hit rate and -20.1% maximum drawdown. QMI Cycle Investing remained ahead at 7.2% annualized return, a 0.74 Sharpe ratio and -17.9% maximum drawdown. On a long/short basis, QMI delivered the strongest return at 13.7% and a 1.39 Sharpe ratio, although several agent variants achieved higher Sharpe ratios through lower volatility and shallower drawdowns. J.P. Morgan cautions that stronger apparent performance from newer models may be concentrated around 2020, precisely when historical recognition could most inflate results. It also warns that Sharpe ratios can flatter older models with few regime switches: Haiku 4.5 and GPT 5.1 posted Sharpe ratios near QMI, but their returns of 3.9% and 4.4%, respectively, lagged QMI's 7.2%, suggesting lower volatility rather than better regime inference. For September 2026, four of six models called Expansion and two called Slowdown; both committees selected Expansion, with probabilities of 0.49 for the two-member group and 0.43 for the six-member group. Constructive models emphasized PMIs above 50, calm volatility and tight credit spreads, while cautious models emphasized fading UK and Eurozone momentum, sticky inflation and still-restrictive policy. The report frames this split as evidence of a difficult turning-point environment and retains QMI as the trusted reference, while viewing Agentic AI as a potentially useful supplementary input as model capability develops.
Analysis framework
J.P. Morgan defines four macro regimes, gives each LLM a fixed prompt and a date-anonymised multi-indicator macro pack, then repeats the exercise across six models and 100 runs per month. It tests leakage through data, date and label ablations; compares individual and committee calls against QMI; and backtests the resulting signals in European cycle-phase factor portfolios using a two-month confirmation rule.
Methodology notes
Cycle-based multi-factor rotation
The strategy rotates between four equity factor portfolios designed for different macro phases, using regime calls as the monthly allocation signal.
Regime-linked equity style leadership
The report links macro-cycle changes to the equity styles expected to lead, seeking to avoid drawdowns when factor leadership shifts.
Ablation testing, ensemble aggregation and bootstrap resampling
The researchers remove dates, raw levels and regime names to test leakage; aggregate model calls into committees; and use bootstrap resampling to select 100 runs per date as a practical stability standard.
Asset mapping & comparison
Structured mapping from thesis to named assets (strengths, weaknesses, peers, risks).
- MSCI Europe cycle-phase equity factor portfoliosThe backtest rotates among four regime-specific portfolios using LLM or QMI regime calls.
- Strengths
- Selective LLM committees generated positive excess returns and improved downside outcomes in the excess-long test.
- Weaknesses
- No LLM or committee exceeded QMI's full-sample Cycle Investing result.
- Comparison
- The two-member GPT 5.5 and Opus 4.8 committee outperformed individual agents and the six-model committee in annualized excess-long return and maximum drawdown.
- Risks
- Residual look-ahead bias, model pessimism, concentrated historical-period performance and higher turnover from more frequent regime switching.
Key data
- Backtest period1994-2026European factor-rotation analysis using date-anonymised macro packs
- Two-member committee QMI match rate48.1%Versus 39.1% for the six-member committee across 391 months
- Two-member committee excess-long annualized return6.0%Sharpe ratio 0.68; hit rate 60.3%; maximum drawdown -10.6%
- QMI Cycle Investing excess-long annualized return7.2%Sharpe ratio 0.74; maximum drawdown -17.9%
- Macro-informed phase-label agreement80%-83%Cohen's kappa was 0.71-0.76, versus 59% agreement and 0.45 kappa for date-only input
- September 2026 committee callExpansionTwo-member probability 0.49 and confidence 0.63; six-member probability 0.43 and confidence 0.65
Impact & implications
The report argues that LLMs can become a complementary regime signal for dynamic equity factor timing, especially through selective ensembles of newer models. It does not view them as a replacement for QMI, given persistent historical-memory risk, directional skew, turnover trade-offs and QMI's stronger full-sample performance.
Risks
- LLMs may recognize historical episodes from their training data, so date anonymisation cannot guarantee out-of-sample purity.
- Several models show a persistent pessimistic bias toward Slowdown and Contraction, which can distort committee outputs.
- The strongest newer-model backtest gains appear concentrated around 2020, when historical-recognition effects may be especially influential.
- The more accurate two-member committee switches regimes more often, implying higher turnover and transaction costs.
- Sharpe ratios may overstate older models' skill when low volatility results from infrequent regime changes rather than better inference.
What to watch
- Whether newer LLM generations continue to improve regime balance and realized factor-rotation performance.
- Whether future testing can further reduce or isolate look-ahead and historical-memory effects.
- The balance between committee accuracy and turnover costs as models are added or removed.
- European PMI momentum, inflation, policy restrictiveness, credit spreads and volatility in assessing the Expansion-versus-Slowdown debate.