Crypto News

Why complex Bitcoin price models often fail to beat simple benchmarks

Why complex Bitcoin price models often fail to beat simple benchmarks

Bitcoin models struggle against simple forecasts

A May 2026 review of Bitcoin prediction research found that no forecasting model has consistently beaten the simplest possible answers at horizons of one to six months.

The analysis, written by Carlos Baquero of the University of Porto, examined hundreds of papers and selected 23 for close review based on their methods and use of genuine out-of-sample evaluation. The preprint has not yet undergone peer review.

This finding highlights a major challenge in crypto forecasting: the market changes so much over time that relationships models learn in one period often disappear in another.

Key findings on forecasting accuracy

  • A May 2026 review found no Bitcoin forecasting model consistently beat naive benchmarks at one- to six-month horizons across different market regimes.
  • The review highlights risks from non-stationarity, overfitting, and information leakage as Bitcoin's market structure changes across cycles.
  • The preprint remains unpeer-reviewed, and stronger tests would require repeated holdout windows, naive comparisons, trading costs, and transparent disclosure.

How naive models stay competitive

Naive forecasting works because financial prices tend to be persistent. If Bitcoin is trading at $100,000 today, a simple forecast of $100,100 tomorrow might have a tiny percentage error even if the model has learned almost nothing about direction or return.

Using today's price as a forecast produces nearly the same accuracy. Evaluating only a complex model gives it credit for information the market already supplied.

The benchmark becomes more demanding as the time horizon expands because Bitcoin can move violently over a month, giving a forecaster room to add value. However, the relationships a model learns decay as the market evolves.

A rule calibrated to the retail-led 2017 cycle encountered a different derivatives structure in 2021, and spot exchange-traded funds (ETFs) created another route for capital in 2024. Each era provides historical data from a version of the market that no longer exists in quite the same form.

This problem is known as non-stationarity, which appears when the relationships between variables do not stay stable enough for past observations to describe the future.

Independent study confirms similar results

Francesco Puoti, Fabrizio Pittorino, and Manuel Roveri reached a similar result in a separate study comparing statistical, machine-learning, and deep-learning forecasts.

They applied 12 different approaches to five major cryptocurrencies at one-day, seven-day, and 30-day horizons. Simple naive models consistently produced better forecasts than ARIMA, Prophet, random forests, XGBoost, LSTM networks, and N-BEATS.

The result says more about the available information than the sophistication of each method. A complex model can add value when stable patterns exist for it to learn, but it can memorize noise when those patterns are weak or temporary.

Bitcoin offers enormous quantities of data, but the number of independent market cycles it has gone through is still quite small. Millions of minute bars keep repeating observations from the same 2018 bear market or the same 2020 liquidity shock.

The danger of backtest overfitting

Many Bitcoin models look strongest once their creators have seen the entire historical period used to build them. Researchers can try different variables, lookback windows, start dates, or architectures before publishing the best result.

The winner may have discovered a durable relationship, but it also could have won a large lottery conducted on the same price history. This outcome is known as backtest overfitting.

David Bailey and his co-authors formalized the problem in their research on the probability of backtest overfitting. Trying more model variations raises the odds of finding an excellent historical result through chance. Selecting the winner and presenting its performance alone hides the number of failed attempts that made the winner possible.

A single chronological split offers little protection because a researcher can train through 2020 and evaluate the model in 2021, producing an apparently out-of-sample result that owes much of its performance to a single bull market.

Walk-forward evaluation is stronger because the model repeatedly retrains on past data and forecasts the next unseen period. Multiple non-overlapping holdout windows are stronger again because they force the same method to encounter bull markets, crashes, sideways periods, and different liquidity conditions.

Among the peer-reviewed papers Baquero examined, none evaluated the same approach across several non-overlapping holdout windows covering different regimes.

Information leakage creates false confidence

Information leakage can also lead to false confidence because a feature calculated with future data can give a model a faint view of the answer.

You get the same problem when you normalize variables across the full sample, and overlapping return windows can carry future observations across the training boundary.

The error can be subtle enough to survive peer review, especially when a complicated architecture puts several transformations between the raw data and the reported forecast.

What is still unclear

Short-horizon order flow and daily return forecasts occupy a separate field and some have produced real predictive value. However, online discussions often blend them with longer-horizon price forecasts and valuation models, although each task asks for a different answer.

A formula describing Bitcoin's historical path tells us little about tomorrow's direction, while a daily direction model says little about the price six months from now.

Why this matters for investors

The findings suggest that claims of superior forecasting ability should be treated with caution. Stronger evaluation methods are needed to separate genuine predictive insight from chance patterns in noisy data.

Sources

Comments (0)

Leave a comment
Your comment will appear publicly after submission.
No comments yet. Be the first to comment!