Published on October 8, 2026
Can an AI really trade the markets?
Promises of bots that win on the markets thanks to AI are everywhere. I built a system where Claude places its own orders on a trading account, then put it up against eleven years of data.
Can an AI really trade the markets?
The idea is appealing and the pitch well rehearsed: an artificial intelligence that reads the markets day and night, makes the right calls and grows an account without human input. These promises are everywhere, pushed by sellers of courses and turnkey robots. What remained was to find out what they are worth once tested. Rather than believe them or dismiss them, I built the system, let it run on a demo account for several months and measured every transaction. The result fits in one sentence: the tool works, the promise does not.
A system, not a hacked-together script
The principle is simple. Claude trades EUR/USD on a demo account the way a disciplined operator would: it reads a briefing in the morning, watches economic announcements, applies a method fixed in advance, places the order and records every trade in a journal. The whole thing runs continuously on a server, inside a Docker container, with a session that starts on its own each morning and stops in the evening.
The first version drove the broker's web terminal, which Claude operated directly through the browser. The approach works, but the click is the fragile link: a volume field that concatenates, a session that expires, and a 0.01 lot order can turn into something else entirely. So I moved to MetaTrader 5 and its Python library, where the order goes out as controlled values, framed by safeguards that refuse anything outside the rules: demo account only, fixed size, mandatory stop, one position at a time. The goal was never speed, but reliability.
Around that execution I added a watcher that reads the ECB, Fed and SNB feeds every two minutes, an economic calendar, reproducible backtests on historical data, a risk simulation and an automatic post-mortem after each loss. The point is worth stressing, because it changes how to read what follows: the problem I ran into is not a plumbing problem.
A deliberately simple method
No martingale, no model meant to predict the future. A classic trend-following method, one I can understand and test.
| Pair | EUR/USD only |
| Signal | the 10 moving average crosses the 30 on 4-hour candles |
| Stop | 3 times the 4h ATR (about 56 pips), set at entry, never widened |
| Exit | reverse cross, stop hit, or Friday before 8pm |
| Size | 0.01 lot, one position at a time |
| Pace | 3 to 4 trades a month |
Staying simple is a deliberate choice: the more settings a method has, the more it risks fitting the past by coincidence and being worthless on the future.
A few definitions
For anyone who does not trade, a handful of words come up below and deserve a plain definition.
- Pip: the smallest price move of a currency pair. On EUR/USD it is 0.0001. It is the unit used to count gains and losses.
- Lot: the unit of position size. 0.01 lot is the smallest common size, the one I kept to limit risk.
- Spread: the gap between the buy price and the sell price. It is the cost charged on every round trip, and it goes to the broker.
- Swap: the interest paid or received for holding a position open from one day to the next.
- Moving average: the average price over the last candles. When the short average (10) rises above the longer one (30), that is the buy signal used here; below it, the sell signal.
- ATR: a measure of the typical size of recent moves, used to place the stop at a distance proportional to how agitated the market is.
- Stop: the order that closes the position automatically if price drifts too far the wrong way, to cap the loss.
What the backtests reveal
A backtest replays the method on past prices, with the real costs measured on the platform (a 2 pip spread, the overnight swap). To avoid fooling yourself, the past is split in two: a tuning period, on which the parameters are chosen, and a validation period, never looked at during tuning, which says whether the parameters hold on something new.
The first test covered nearly three years, from late 2023 to 2026: +50 CHF on tuning, +34 CHF on validation, positive every year. Enough to think you are onto something.
The next test, over eleven years of Dukascopy data (tuning from 2015 to 2025, validation on 2026), tells a different story.
| Result over eleven years | Value |
|---|---|
| Tuning 2015 to 2025 | +42 CHF over 267 trades |
| Validation 2026 | +28 CHF over 27 trades |
| Average gain per trade | about 2 pips, or 0.16 CHF |
| Winning trades | 42% |
| Worst drawdown | 59 CHF |
| Longest losing streak | 11 in a row |
The year-by-year breakdown leaves no doubt (in CHF): 2015 -2, 2016 -12, 2017 -4, 2018 -16, 2019 +57, 2020 +4, 2021 -18, 2022 +26, 2023 -35, 2024 +8, 2025 +34, 2026 +28.
The method wins clearly in the years when the market has a direction (2019, 2022, 2025) and loses the others. Over eleven years, the expectancy is roughly zero. The good numbers from the first test came from a favourable period, not from a real edge. One last check confirms it: re-choosing the parameters each year on the past only (walk-forward), the tally falls to -39 CHF over ten years, with five negative years out of ten.
The only improvement that holds across both periods is to stay out when a Fed or ECB decision, an NFP or a CPI release falls within 24 hours. It lifts tuning from +42 to +105 CHF and validation from +28 to +39 CHF. Even then, it is not about guessing direction, only about avoiding exposure at the wrong moment. Around forty other ideas (Ichimoku, multiple crossovers, volatility filters) were tested: none does better in a reliable way, and none reaches the statistical threshold that would set it apart from chance.
The weight of costs
Why does an average gain of 2 pips weigh almost nothing? Because the value of a trade comes down to an expectancy:
expectancy = (win rate x average gain) - (loss rate x average loss)
With 42% of trades winning, the gains have to clearly exceed the losses to stay positive. They do, but barely: the result falls back to about 2 pips per trade. Yet each round trip already costs 2 pips of spread. The transaction cost is therefore of the same order of magnitude as the signal itself, and the entire margin of the method sits inside a gap thinner than what the platform takes.
The same mechanism had already killed an earlier idea, scalping on 10 minutes: a loser in every variant of the backtest, even with the spread removed from the simulation. At that scale, the spread alone absorbs close to 40% of the move of a 5-minute candle. You start out losing before you have even been right.
What the probabilities say
You cannot predict the order in which gains and losses will occur. To estimate the risk, a Monte Carlo simulation therefore replays the backtest trades 10,000 times in a random order, and counts how often a normal losing streak would blow the risk budget (25 CHF).
| Horizon | Probability of being stopped | Probability of ending negative |
|---|---|---|
| 20 trades (about 9 months) | 18 to 20% | 43 to 44% |
| 30 trades (about 13 months) | 26 to 28% | 41 to 42% |
| 50 trades (about 22 months) | 36 to 37% | 38 to 39% |
In other words, with the method intact, there is about a one in four chance of being stopped within the year by sheer bad luck, and four chances in ten of ending negative. The implication matters: on a sample this small, a gain does not prove the method works, and a loss does not prove it is broken. Chance dominates the outcome, whether a human or a machine is at the wheel.
A reliable tool, an unfounded promise
What does not hold up is not the tool. Claude does exactly what is asked of it: read the context, apply a discipline without tiring, never forget a stop, keep an honest journal. On execution, a disciplined human operator would do no better.
The problem lies elsewhere. An artificial intelligence does not create an edge where none exists. The retail currency market behaves like a negative-sum game: price moves are close to a coin toss, and transaction costs tilt the balance against the operator. No model turns a signal worth 2 pips into a source of income when placing the order costs just as much. Offers of robots presented as profitable usually rest on a lucky stretch mistaken for skill, when they are not simply a sales argument.
What the experiment is really worth
Its value is not in a return. It is in the method: backtests that separate tuning from validation, an end-of-test rule written in advance so the decision is not made under the sway of emotion, a coded safeguard that refuses the dangerous order, a post-mortem after each loss to tell ordinary bad luck from a genuine mistake.
As a learning ground, the exercise is hard to beat: an autonomous agent running for months, driven first through a browser and then through a programming interface, infrastructure, market data and statistical rigour. The demo account was never meant to make me rich. It answered a more useful question: am I able to build a rigorous system and judge it honestly, even when the verdict does not suit me. That verdict, here it is: AI does not beat the market. Mostly, it helped me prove it cleanly.
A word of caution
This experiment finally sheds light on the murkiest part of the subject: the courses and robots sold as money machines. A few signals should always raise a flag.
A return presented as guaranteed, or the image of money made while you sleep, contradicts what the numbers show: once costs are deducted, retail trading behaves like a negative-sum game. A flattering backtest proves nothing either until it has held on a period never used to tune it, as my first test reminded me, brilliant over three years and worthless over eleven. A winning-trade rate put forward without the ratio between gains and losses, or the costs, means nothing: you can win eight times out of ten and still lose money.
The business model deserves a look too. A seller of courses, signals or a bot earns from the subscription, not from the market. If they truly held a method that beat the market, they would have no reason to hand it over. And when the method loses, the blame conveniently shifts to the student: not enough discipline, not enough capital.
There remains the distance between demo and live, which my test did not even cross. Real conditions add price slippage, emotions and the risk of losing money that matters. The simplest rule is probably the sturdiest: if a method really won for sure, no one would waste their time selling it.
Sources and data
The figures in this article come from the test setup itself, fed by identified and cross-checked sources:
- Historical prices: Dukascopy (hourly data since 2015), Yahoo Finance as an independent check.
- Calendar and announcement dates (rate decisions, NFP, CPI): ForexFactory, FRED, official Fed and ECB sites.
- News feeds and statements: ECB, Fed, SNB, Investing, FXStreet.
- Macro context and positioning: CME FedWatch, the CFTC Commitments of Traders report, Myfxbook.
- Execution: Swissquote demo account (CFXD web terminal, then MetaTrader 5 and its Python library).