The Risk of Mistaking Chance Results for Skill in Financial Machine Learning—and How to Challenge Them
Evidence and scope — Failure modes and falsification criteria
The cited literature informs checks for overfitting, leakage, costs, and other reasons to question strong results. Timing rules are educational examples, not validation of actual executions or the superiority of a particular model.
Part 8 covered a design for comparing momentum, logistic regression, decision trees, and random forests under the same conditions. It did not report actual ETF performance or establish that any particular model was superior. This article takes the next step: why a result that looks good in a comparison table should not immediately be called “skill.”
A fair comparison is only the starting point. Even if one model leads during a particular period, we still need to test whether that result survives under different market periods and cost assumptions.
The ways financial machine learning can fail can be grouped into three broad areas.
| Core area | Included causes | First question to ask |
|---|---|---|
| Unstable data relationships | 1. More noise than signal<br>2. Market relationships change over time | Does the signal persist across other periods and market conditions? |
| Validation and selection errors | 3. Excessive repeated searching<br>4. Future-information leakage<br>6. Selecting only favorable results | Have choices made before seeing the results been separated from choices made afterward? |
| Economic meaning | 5. A small predictive edge disappears after costs | Does anything meaningful remain after trading costs? |
This list is not a diagnosis that declares why a particular experiment failed. Without finalized data, trade records, and execution results, we cannot know which issue actually occurred. It is better understood as an audit map for avoiding overconfidence in attractive results.
| Cause | Symptoms to watch for | How to check | Limits of the check |
|---|---|---|---|
| More noise than signal [S1] | Results shift sharply when the period or split changes | Compare against simple baselines and review results by period | Instability alone does not prove the model fitted noise |
| Market relationships change over time [S2][S3][S11][S12] | Features that worked in one period weaken in another | Record errors by period, feature changes, and results by market state | Causes such as crowding cannot be confirmed without additional evidence |
| Excessive repeated searching [S4][S5][S8] | Only the best settings remain, while changes multiply after viewing test results | Record the number of candidates, seeds, runs, and changes | Keeping records does not guarantee future performance |
| Future-information leakage [S6] | Values unavailable at the time are mixed into features, preprocessing, or labels | Audit the signal date, information-availability time, label-finalization date, and execution date | Actual data-availability timing and fill feasibility require separate checks |
| A small edge disappears after costs [S7] | Accuracy improves, but results weaken after costs | Calculate turnover alongside cost scenarios | Implicit costs are difficult to measure precisely without execution data |
| Selecting only favorable results [S9][S10] | Failed runs and unfavorable findings disappear from the record | Disclose failed runs, code, data timing, and selection rules | Replication failure and publication bias are not the same phenomenon |
1. When There Is More Noise Than Signal
Predicting stock returns is a low signal-to-noise problem. The more flexible the model, the more carefully we need to examine whether it has fitted random fluctuations in the training data. [S1]
Here, noise does not mean only erroneous data. One-off moves that have no stable relationship with the future can also look like convincing patterns to a model. If results look strong during the training period but change substantially when you shift the date boundary or training window slightly, the model may have learned more from sample fluctuations than from genuine signal.
That does not mean complex models are always bad. Rather, the more flexible the model, the more important it becomes to check whether it discovered a relationship or merely fitted noise. [S1]
A practical check is to compare a simple baseline such as momentum with an ML model using the same chronological order, and to retain performance by period and variation across splits—not just one average result. The fact that a model worked once is less persuasive than evidence that it continues to point in a similar direction when conditions change.
Still, unstable results alone do not justify concluding that the model fitted noise. Changes in market relationships, discussed next, or differences in cost and execution assumptions may also be contributing.
2. When Market Relationships Change Over Time
It helps to distinguish three related terms.
- Regime change describes the possibility that relationships shift as an unobserved state changes in a nonstationary time series. Whether one fixed rule can explain every period is itself something to test. [S3]
- Concept drift is the broader machine-learning concept in which the input distribution or the relationship between inputs and outcomes changes over time. [S2]
- Feature decay is a practical term for the weakening over time of the predictive or economic value of a feature that was once useful. Research finding weaker out-of-sample and post-publication effects for published return predictors gives us reason to examine this possibility. [S12]
For that reason, we cannot assume that momentum, volatility, or trading-volume features that worked well in one period will continue to work the same way later. This is why it is useful to record errors by period, changes in feature importance, and results across different market states. [S2][S3]
Research also suggests that when many participants use similar factors, correlated trading and liquidity problems can emerge. [S11] But to claim that crowding caused a result in an individual ETF experiment, we would need evidence such as overlapping holdings, order flow, liquidity, and market-impact data. Without those data, crowding is only a possible explanation.
Time variation is not a confirmed cause of failure. It is a reason not to place blind trust in a single fixed rule.
3. When You Search Too Many Models and Parameters
When the same data are repeatedly used for model selection, a result that happened to look appealing can be mistaken for evidence that a method is genuinely good. This is the central risk of data snooping. Searching across many strategy and parameter variations and selecting the best in-sample result can lead to backtest overfitting. [S4][S5]
Suppose you keep changing tree depth, leaf size, the number of trees and features in a random forest, buy thresholds, and cost assumptions while repeatedly checking the results. The setting left at the end is no longer a single hypothesis you intended to test from the beginning. In particular, changing candidate models, thresholds, or cost assumptions after viewing performance in the test period should be recorded as a new search, not treated as a minor adjustment. [S4][S5][S8]
Before running the experiment, it is useful to set a search budget. Keep a record of the candidate list, number of runs, random seeds, selection criteria, when the test period was first viewed, and any changes made afterward. Ideas that arise after seeing the test results should be separated into a new exploration rather than selected again using the original test period.
This process does not automatically eliminate overfitting. It does, however, prevent us from hiding the question, “How many attempts did it take to produce this best result?” A test period is a report card, not an unlimited laboratory for making the next choice.
4. Future-Information Leakage
Leakage occurs when a model uses information that could not legitimately have been available for the prediction target. The important question is not whether a value exists in a data column, but whether it was actually known at the time the investment decision was made. [S6]
For an ETF panel, each row can record at least these four points in time.
| Recorded item | Audit question |
|---|---|
signal_date | When was the signal created? |
| Feature availability time | Was this feature finalized before the signal time? |
label_end_date | When was the future label finalized? |
| Execution date | After the signal, when and at what price is the trade assumed to occur? |
Creating a signal at month-end and executing at the next trading day’s close is one example of an educational timing rule. It does not mean the desired quantity could actually have been filled at that closing price. In addition, rows whose labels have not yet been finalized should not enter training, and preprocessing such as scaling or missing-value handling should be fitted within each training period only. As an audit rule for timing consistency in this article, ETF rows from the same date can be split together by date so they are not scattered across training, validation, and test sets. This rule alone does not prove that leakage was absent. [S6]
The availability timing of adjusted prices, delisting treatment, changes in ETF constituents, and the feasibility of execution at the market close all require separate evidence. These rules are an audit design for preventing leakage, not a performance report claiming that an actual implementation was leakage-free.
5. When a Small Predictive Edge Vanishes After Costs
Portfolio trading costs include not only commissions but also spreads, market impact, and opportunity costs, while implicit costs can be difficult to measure directly. [S7] Therefore, a modest improvement in accuracy or model score does not by itself establish economic significance.
The following is an arithmetic example, not actual ETF performance.
Assumption: If the expected gross edge before trading is 10bp and round-trip costs are 14bp, the net effect after costs is
10bp - 14bp = -4bp.
When the predictive edge is small, costs can reverse the direction of the outcome. In particular, a model whose signals change frequently and produce high turnover may improve accuracy while becoming weaker after costs.
When checking this, define cost scenarios that distinguish commissions, spreads, market impact, and opportunity costs alongside turnover. A rule such as executing at the next trading day’s close after a month-end signal is also a simplifying assumption for calculation; it does not guarantee the actual fill price or fill feasibility. [S7]
Without execution data, it is more honest not to treat costs as one exact number. Checking whether results remain after costs is necessary, but that check alone cannot guarantee that future execution will be feasible.
6. Researcher Bias: Selecting Only Favorable Results
Researcher bias is easier to understand when separated into two layers. The first is selection within a study: trying many candidates and retaining only favorable results. Reusing the same data, conducting large-scale searches, and performing multiple tests all increase the risk of selecting a result that looks good by chance. [S4][S5][S8] The second is publication bias, in which favorable findings are more likely to be published in papers or reports. [S10] The two are related, but they are not the same question.
A large-scale anomaly replication study showed that many results may fail to meet earlier standards when sample construction, weighting methods, and multiple-testing criteria change. This supports making replication procedures and design choices transparent. [S9] At the same time, publication bias and out-of-sample weakening should be measured separately, and uncertainty remains about the size and interpretation of their effects. [S10]
For this reason, no single study justifies the conclusion that “most financial ML results are false,” nor that bias can safely be ignored. Instead, failed runs, excluded candidates, changed rules, data timing, and code should be retained so that a third party can check the result again under the same conditions.
Reproducibility is not a mechanism that guarantees a good result. It is a condition that makes a result possible to verify or challenge. Failed runs are not disposable byproducts; they are part of the selection process.
Falsification Conditions to Attach to Any Good Result
In financial ML, the more important question is not, “Which model looked best?” It is whether you can answer these three questions.
- Does the result persist across other periods, splits, and market conditions? [S1][S2][S3]
- Can the same conclusion be evaluated if you disclose the number of candidates, number of runs, changes made after viewing the test set, and timing rules? [S4][S5][S6][S8][S9][S10]
- Does the result retain economic meaning after considering not only commissions but also spreads, market impact, and opportunity costs? [S7]
A result that cannot yet answer these questions does not need to be labeled a failure. But it is also too early to call it an investment edge. The most accurate label is a hypothesis requiring further validation.
In your next experiment, check data timing, search history, cost assumptions, and records of failed runs before focusing on a performance chart. Part 10 will examine how LLMs can be used beyond price prediction—as aids for organizing research, assisting with coding, and checking for errors.
This article is an educational and research-oriented methodological explanation. It is not a recommendation to trade any specific ETF or financial product.
Sources
- [S1] Empirical Asset Pricing via Machine Learning | The Review of Financial Studies / Shihao Gu, Bryan Kelly, Dacheng Xiu | 2020-02-26 | https://doi.org/10.1093/rfs/hhaa009 ↩
- [S2] A Survey on Concept Drift Adaptation | Association for Computing Machinery / João Gama, Indrė Žliobaitė, Albert Bifet, Mykola Pechenizkiy, Abdelhamid Bouchachia | 2014 | https://dl.acm.org/doi/10.1145/2523813 ↩
- [S3] A New Approach to the Economic Analysis of Nonstationary Time Series and the Business Cycle | Econometric Society / James D. Hamilton | 1989-03-01 | https://www.jstor.org/stable/1912559 ↩
- [S4] A Reality Check for Data Snooping | Econometric Society / Halbert White | 2000-09 | https://doi.org/10.1111/1468-0262.00152 ↩
- [S5] Backtest Overfitting in Financial Markets | eScholarship, University of California / David H. Bailey, Jonathan M. Borwein, Marcos López de Prado, Amir Salehipour, Qiji Jim Zhu | 2016 | https://escholarship.org/uc/item/4hn4t174 ↩
- [S6] Leakage in data mining: Formulation, detection, and avoidance | Association for Computing Machinery / Shachar Kaufman, Saharon Rosset, Claudia Perlich, Ori Stitelman | 2012-12 | https://dl.acm.org/doi/10.1145/2382577.2382579 ↩
- [S7] Request for Comments on Measures To Improve Disclosure of Mutual Fund Transaction Costs | U.S. Securities and Exchange Commission | 2003-12-19 | https://www.sec.gov/rules-regulations/2003/12/request-comments-measures-improve-disclosure-mutual-fund-transaction-costs ↩
- [S8] … and the Cross-Section of Expected Returns | The Review of Financial Studies / Campbell R. Harvey, Yan Liu, Heqing Zhu | 2015-10-09 | https://academic.oup.com/rfs/article-abstract/29/1/5/1843824 ↩
- [S9] Replicating Anomalies | The Review of Financial Studies / Kewei Hou, Chen Xue, Lu Zhang | 2018-12-10 | https://doi.org/10.1093/rfs/hhy131 ↩
- [S10] Publication Bias and the Cross-Section of Stock Returns | Board of Governors of the Federal Reserve System / Andrew Y. Chen, Tom Zimmermann | 2020-01-09 | https://www.federalreserve.gov/econres/feds/publication-bias-and-the-cross-section-of-stock-returns.htm ↩
- [S11] FACTOR CROWDING AND LIQUIDITY EXHAUSTION | Journal of Financial Research / Joseph M. Marks, Chenguang Shang | 2018-12-04 | https://doi.org/10.1111/jfir.12165 ↩
- [S12] Does Academic Research Destroy Stock Return Predictability? | The Journal of Finance / R. David McLean, Jeffrey Pontiff | 2015-10-13 | https://onlinelibrary.wiley.com/doi/abs/10.1111/jofi.12365 ↩
Report an error or share feedback
Open a draft with this article’s title and URL. Review the message and recipient before sending.
Open email draftIf no email app opens, copy these details into your usual email service.
Related posts
Quant & Data Research Can Logistic Regression Beat the Baseline? What Makes a Model Comparison Fair
Define timing, data splits, retraining, and cost rules for a fair comparison of logistic regression and baseline strategies, before testing actual performance.
Quant & Data Research Why shift(1) Is Not Enough: A Time Contract for Signals, Execution, and Labels
Separate signal, execution, return, and label timing. Use synthetic unit tests to check information leakage and preprocessing boundaries.