Would You Trust an AI Model to Invest Real Money?
**Important:** This discussion is educational and is not financial advice. The results below are attributed to the experiment described in The Koerner Office and have not been independently audited by WhatAI.
A recent Koerner Office experiment compared investment recommendations from Claude, ChatGPT, Gemini, Grok and Perplexity. One model reportedly more than doubled its allocation while another lost over half. Does that reveal investment skill or short-term noise?
**The experiment**
Chris Koerner interviewed Brandon Doyle about an experiment that gave five AI systems ([Claude](/tool/claude), [ChatGPT](/tool/chatgpt), [Gemini](/tool/google-gemini), [Grok](/tool/grok-ai) and [Perplexity](/tool/perplexity)) $1,000 allocations to manage through recommendations.
According to the official episode description, Brandon prompted the models, made weekly trades from their recommendations and compared the results. Claude reportedly more than doubled its money, while Gemini lost more than half after increasingly risky bets. The conversation also covered leveraged exchange-traded funds, selling decisions and investor psychology.
The outcome is attention-grabbing. It is not automatically evidence that Claude is the best investing model.
**Why one profitable run is not a reliable benchmark**
Investment results combine decision quality with market conditions and chance. A risky position can produce the best return over a short period and still be a poor strategy over hundreds of trials. A conservative portfolio can lag during a sharp rally while providing better downside protection over a longer horizon.
To compare the models fairly, we would need to know:
- Whether every model received exactly the same information and prompt.
- Whether each model had access to current market data at the same time.
- How often the models were allowed to trade.
- Whether transaction costs, spreads and taxes were included.
- Whether returns were adjusted for volatility and drawdown.
- How the experiment handled hallucinated facts or stale information.
- Whether the models were allowed to revise their strategy after losses.
- How the result would change across different market periods.
Without those controls, the experiment is better understood as a provocative demonstration than a scientific ranking.
**The deeper risk: confidence without accountability**
AI models can produce a persuasive explanation for a trade. The language may be organised, detailed and confident even when the underlying assumptions are weak. Investors can mistake fluency for accuracy.
There is also an automation problem. An assistant that suggests an idea leaves a human responsible for checking it. An agent that executes the idea can create losses before anyone notices an error. The appropriate level of autonomy depends on the size of the position, the reliability of the data and the consequences of being wrong.
Financial markets are adversarial and adaptive. A pattern described in public information may already be reflected in prices. Models trained on historical material can explain what happened more easily than they can predict what will happen next.
**Where AI may still be useful for investors**
Rejecting autonomous trading does not mean AI has no value in finance. A model can help organise research, summarise filings, compare competing arguments, identify questions, monitor a watchlist and expose assumptions in a thesis.
Safer uses include:
- Summarising a company's public disclosures with links to the original documents.
- Generating a bear case against your preferred investment.
- Comparing stated management goals with later results.
- Building a checklist before a human makes a decision.
- Explaining portfolio concentration, scenario risk and possible drawdowns.
- Flagging changes that require human review rather than automatically trading them.
The model should not become the source of truth. It should help the user interrogate sources and their own reasoning.
**What a stronger AI investing test would look like**
A more informative comparison would run for a longer period, use identical information, define risk limits in advance and compare performance against a simple benchmark. It would measure maximum drawdown, volatility, turnover and risk-adjusted return, not only the final balance.
The models should also be tested across different tasks:
- Research accuracy
- Recognition of uncertainty
- Ability to correct an error
- Portfolio construction
- Risk control
- Source citation
A model that produces the highest return by accepting extreme risk may be less useful than one that consistently identifies missing information and avoids catastrophic mistakes.
**The question for the WhatAI community**
**Would you allow an AI model to execute trades with real money, or should AI remain a research assistant with a human making every final decision?**
Also:
- What would an honest Claude versus ChatGPT versus Gemini investing benchmark need to measure?
- Is source quality more important than the model itself?
- Would you trust AI more for long-term analysis or short-term trading?
- Should an investing assistant be optimised for returns, risk control or admitting uncertainty?
- Has AI ever changed your mind about an investment after showing evidence you had missed?
**Source:** [The Koerner Office, episode 317, published July 14, 2026](https://toolkit.tkopod.com/podcast/episode/6c9a)