Last spring a friend showed me his AI trading model. Four months of work, a neural network trained on three years of Bitcoin minute bars, and a backtest that printed a 94% hit rate. He went live with $8,000. Nine days later he was down 22%, and he still couldn’t explain why.
The model wasn’t the problem. Everything around the model was. Here’s the process I’d use instead, in the order I’d actually do it.
Step 1: Shrink the question until it’s almost boring
“Can AI predict the market?” isn’t a project, it’s a wish. You need a question with a yes-or-no answer and a timestamp attached.
- Too big: “Which stocks will go up?”
- Buildable: “Will SPY close above today’s open tomorrow, given the previous 20 sessions of price and volume?”
The second version gives you one instrument, a defined label, and a prediction horizon you can actually measure. It’s dull. That’s a feature.
Pick one instrument and one bar size to start
A single liquid ETF, one timeframe, a few years of history. SPY one-minute bars run to roughly 98,000 rows a year, which is plenty to learn on. Once the pipeline works end to end, adding instruments becomes a config change rather than a rewrite.
Step 2: Get data you can actually trust
Free data will get you started and then quietly cost you money. Before you write a single line of model code, check these:
- Adjusted versus raw prices. A 4:1 split looks like a 75% crash to an untrained model, and it will happily trade the phantom dip.
- Survivorship bias. A ticker list pulled today excludes every company that delisted along the way, which flatters any strategy that buys weakness.
- Timestamp alignment. Price, volume and news feeds need to share a clock and a timezone, or you are leaking future information into past rows.
- Missing bars. Halts and holidays leave holes. Decide whether you forward-fill, drop or flag them, and write the decision down.
Ten hours of cleaning for every hour of modelling is normal. It isn’t a sign you’re doing it wrong.
Step 3: Build a baseline before you train anything
Run two deliberately dumb strategies first: buy and hold, plus a 20/50 moving average crossover. Record Sharpe ratio, maximum drawdown and turnover for both. Then compute the same figures for whatever clever thing you build.
This matters because a lot of what gets sold as AI trading is a complicated way of rediscovering a trend filter. Knowing what machine learning actually adds to a portfolio, and what it doesn’t, stops you mistaking novelty for edge.
Step 4: Split your data by time, never at random
Shuffling rows into training and test buckets kills more projects than bad features do. If Tuesday’s bar lands in training and Monday’s in test, the model has effectively seen tomorrow’s newspaper.
A workable walk-forward setup on daily bars looks like this:
- Train on 24 months, test on the next 3.
- Roll forward 3 months and repeat until you run out of history.
- Stitch the out-of-sample windows into one continuous equity curve. That curve is the only one worth judging.
Expect the stitched result to look considerably worse than your first attempt. That gap is the honest one.
Step 5: Make the backtest pay real tolls
A backtest without costs is a story, not a result. Add these before you get excited:
- Commission. Even zero-fee brokers earn on the spread, so assume something.
- Slippage. One tick on a liquid name is optimistic; two is safer for anything mid-cap.
- Borrow costs if you short, plus financing if you use leverage.
Here’s the sobering arithmetic. A strategy trading four round trips a day at 1.5 basis points of cost each pays roughly 12% a year in tolls. Most signals don’t clear that bar, which is precisely why lower-frequency ideas tend to survive longer than clever ones.
Step 6: Paper trade, then size small enough to be embarrassed
Sixty days of paper trading won’t prove a strategy works, but it will break your plumbing: order routing, partial fills, what happens when the API times out at 9:31am on a payrolls Friday.
When you do go live, size at a quarter of what the backtest implies. If the model is as good as the numbers claim, a small position still compounds nicely. If it isn’t, you’ve paid tuition rather than a mortgage payment.
Step 7: Watch for decay and schedule the retrain
Models go stale. Correlations shift, volatility regimes flip, and a feature that predicted volume spikes in 2021 stops working once market structure changes beneath it. Set a monthly retrain with a validation gate, and write a hard rule that the model goes offline after a fixed drawdown. Write that rule now, because you will not be objective at the moment it triggers.
Where compute and platforms actually matter
For a personal project, a laptop and a $20-a-month cloud instance run everything above. Latency only becomes a constraint when your holding period is measured in seconds, and that is a different sport, closer to the hardware race that turned a wafer-scale processor into a category, like the story behind Cerebras and its dinner-plate-sized chip. Institutional desks pay for microseconds. You probably don’t need to.
The same holds on the software side. Enterprise AI platforms such as C3 AI’s application layer bundle data pipelines, feature stores and monitoring so teams don’t rebuild them per project. You can borrow the architecture without buying the product: keep the data layer separate from the model layer, and keep monitoring in one place.
Keep learning deliberately too. There’s a real difference between building AI skills without drowning in hype and collecting certificates that never touch a live signal.
What to do on Monday morning
Pick one liquid ETF. Pull five years of daily bars. Write a script that computes buy-and-hold and a moving average crossover, both with costs baked in. That’s a day of work, and it sets the bar every later idea has to clear.
Then, and only then, train a logistic regression on three features: five-day return, volume relative to its 20-day average, and distance from the 50-day moving average. Logistic regression, not a transformer. If three features and a linear model can’t beat the baseline after costs, twelve layers of attention won’t rescue you.
The edge in AI trading was never the algorithm. It’s the discipline to test the boring version first, and to believe the out-of-sample numbers over the in-sample story you’d rather tell.

