What Makes a Good Backtest Different? Designing the Validation Process Before Looking at Returns
Evidence and scope — Artificial-series validation example
Literature and a fixed-seed artificial time-series example illustrate time splits and evaluation procedures. This is not an experiment establishing real-asset performance; costs and risk evaluation are outside the example’s scope.
A practical introduction to preregistering hypotheses, choosing baseline strategies, splitting data chronologically, using walk-forward validation, and protecting the test period
At first, I thought the goal of backtesting was to run many strategies and select the chart with the highest cumulative return. If the curve looked smooth and ended with a large gain, it felt as though I had discovered a good strategy.
But backtest returns are hypothetical results calculated from historical data, not actual investment performance.[S2] Their meaning can change simply because the comparison benchmark was changed or a favorable period was selected.[S2] Worse, if enough strategies, features, periods, and parameters are tested, it becomes possible to select a candidate that stands out because it was lucky rather than genuinely superior.[S9][S10][S11]
That led me to change the question.
Before asking how high the result was, how effectively did the process prevent future information and after-the-fact human choices from influencing it?
A good backtest is not a procedure for selecting the highest return. It is a procedure for making a fair attempt to falsify a predefined question without using future information.
The tables and code in this article are educational and research examples designed to explain backtest validation. They are not recommendations of specific securities or instructions to trade, and they do not guarantee returns. Differences observed in the example scores must not be generalized as evidence that one real asset, market, period, or investment strategy is superior.[S2][S8]
After Aligning the Data Clock, Align the Experiment Clock
Part 3 centered on one question:
Was that value actually observable when the trading decision had to be made?
In this installment, we will assume that the data’s creation and observation times have been aligned correctly. The next step is to decide how to divide that data, which period will be used to make choices, and which period will remain unseen until the end.
Even when individual features contain no future values, information from the test period can enter the training process if you calculate scaling parameters or imputation rules over the full period, select features, and only then split the data.[S4] Conversely, even a chronological split will not eliminate the leakage discussed in Part 3 if a feature’s observation time is incorrect.
Chronological splitting is an important line of defense, but it is not a button that automatically eliminates every form of leakage.[S4][S5]
A Good Backtest Records the Question Before the Result
Defining the research question and analysis plan before seeing the results makes it easier to distinguish an original prediction from an explanation created afterward.[S1] The same principle can be applied to a research log for financial backtesting.
This does not mean that every independent researcher must publicly preregister every experiment. Plans and revision histories can be preserved in timestamped notes, Git commits, or experiment-tracking tools. Based on the current evidence, however, I cannot say which recording method is the most reliable.
| Item to record in advance | Question to answer |
|---|---|
| Hypothesis | What information do I expect to be related to the next period’s target, and why? |
| Falsification condition | What result would lead me to conclude that the hypothesis is not supported? |
| Data timing | Was each feature available before the decision had to be made? |
| Baseline | What must the method beat before I can claim an incremental improvement? |
| Split rules | What are the dates and chronological order of the training, validation, and final test periods? |
| Evaluation criteria | Which metric and aggregation rule will be used to select candidates? |
| Search budget | What is the maximum number of strategy, feature, period, and parameter combinations I will test? |
| Final test | When will I open it, and what will I refuse to change after seeing the result? |
The purpose of recording a plan in advance is not to prohibit exploration. Exploratory analysis inspired by the results can also be useful. The important point is to distinguish it from the original plan rather than presenting it as though it had been a confirmatory hypothesis from the beginning.[S1]
Preregistration does not make a flawed hypothesis correct.[S1] It does, however, make it harder to change the question to fit the result and later remember the revised question as the one you tested all along.
Choose the Baseline Strategy Before Building a Complex Model
Before building a complex model, define a baseline that is comparable to the subject of the research. Comparing a backtested strategy with a benchmark representing a different market or type of investment can undermine the meaning of the comparison.[S2]
Buy and hold is not the only valid baseline for every study. The appropriate comparison depends on the research question.
| Research question | Possible baseline strategy |
|---|---|
| Is this better than simply holding the asset? | Buy and hold for the same asset and period |
| Is a predictive model necessary? | Carry forward the previous value, use a long-term average, or always make the same prediction |
| Is the complex model better? | A simple rule or reduced model addressing the same target |
For example, when evaluating a model that classifies upward and downward moves, a simple prediction rule using the same sample and target can serve as the comparison. When comparing a complex signal with holding the asset, the benchmark should cover the same asset and period.
If the baseline is replaced after seeing the results with one that makes the strategy look better, the choice of baseline becomes part of the selection process. The comparison should therefore be recorded in advance alongside the research question.[S1][S2][S9]
Training, Validation, and Testing: Studying, Taking a Practice Exam, and Opening the Sealed Final
The three datasets should be distinguished by their roles, not merely their names.
Training periodFit model coefficients and preprocessing statistics.Validation periodSelect features, rules, parameters, window lengths, and candidates.Final test periodInspect the final results only after all selections are complete.Training data is used to fit the model. Validation data is used to choose among candidates. Test data is used only after all selection is complete, to evaluate final generalization performance.[S3]
A useful analogy is studying, taking a practice exam, and sitting a sealed final exam. It is perfectly normal to revise your study method after reviewing a practice exam. But if you inspect the final exam questions and score, revise your answers, and then take the same exam again, it is no longer an independent final evaluation.
In financial time series, chronological order matters in addition to those roles.
Past Future|-------- Training --------|---- Validation ----|-- Final test --|The scaler’s mean and standard deviation, imputation values, feature-selection criteria, and model coefficients must all be estimated using only the relevant training period. Fitting preprocessing on the complete dataset allows test information to enter the training pipeline and can make the evaluation overly optimistic.[S4]
Example ratios for training, validation, and testing should not be treated as universally optimal.[S3] The specific length of each period must be designed and evaluated separately for each experiment.
The Moment You See the Test Result, You Become Part of the Training Process
The following sequence is common.
Inspect the final test→ Find the result unsatisfactory→ Modify parameters, features, or rules→ Recheck the same testEven if the code never directly calls fit on the test rows, the researcher learns something from the first test result. If that information is then used to change features or parameters, the test data has indirectly influenced model selection.[S3][S4]
From that point onward, the period is effectively validation data rather than an independent final test. One of the following responses is then necessary:
- Designate a later, previously unused period as the new test period.
- Conduct forward validation on observations that become available as real time passes.
- If no new data is available, label the result as “exploration adapted to the existing test” rather than “final validation completed.”
The current evidence does not establish a universally appropriate length for a new, unused test period. Its length must therefore remain an explicitly documented, unresolved design choice for each study rather than being treated as a validated general standard.
Looking at the test result more than once does not invalidate the entire research project. It does mean that the period’s status must be changed honestly. Whether data is truly a test set is determined not by its filename or date, but by whether it was actually used during the selection process.
Why Random Splitting Changes the Question for Financial Time Series
Many introductory machine-learning examples shuffle observations and use some for training and the rest for evaluation. By default, scikit-learn’s train_test_split also shuffles the data.[S6]
Applying that default directly to time-indexed data can create a structure like this:
Random splitTraining: 1, 3, 4, 7, 9Evaluation: 2, 5, 6, 8, 10Chronological splitTraining: 1, 2, 3, 4, 5, 6, 7Evaluation: 8, 9, 10Under a random split, the model could be evaluated at time 2 after being trained on later observations from times 3, 4, 7, and 9. Standard cross-validation can create folds that train on future observations and evaluate on past ones if it is applied to time series without modification.[S5]
The real-world forecasting question is usually closer to this:
Knowing only what was available up to that point, how well could I have predicted the next period?
To reproduce that question, observations later than the evaluation point must be excluded from training.[S5][S7] Even if no individual feature directly contains a future value, a randomly assembled training set can contain distributions and relationships observed during later market regimes. In other words, the model may be exposed to later-period evidence that would not have been available at the actual prediction time.
This does not mean random splitting is wrong for every analysis. It may be appropriate for some research questions, including cross-sectional problems that are reasonably close to being independent and identically distributed.[S6] The problem is using shuffling uncritically when the experiment is supposed to reproduce forecasting from the past into the future.
Nor does random splitting invariably produce a higher score. An official lagged-feature example found that random splitting produced a more optimistic error estimate than time-based evaluation, but that observation came from bicycle-demand data.[S8] The direction and magnitude of the difference cannot be generalized to financial data.
How Is a Single Split Different from Walk-Forward Validation?
The simplest chronological design trains once on a past period and evaluates once on the period that follows.
Fixed split[──────── Training ────────][──── Evaluation ────]Walk-forward validation repeatedly moves the forecast origin forward, trains on observations preceding each evaluation window, and evaluates on the next period.[S5][S7][S13]
Fold 1 [Training────][Evaluation]Fold 2 [Training────────][Evaluation]Fold 3 [Training────────────][Evaluation]The key feature of this procedure is that it allows errors from multiple forecast points to be aggregated.[S7] Two common ways of constructing the windows are the expanding window and the rolling window.
| Method | Chronological structure | Conditional interpretation and design questions |
|---|---|---|
| Fixed split | One past training period followed by one future evaluation period | Only one evaluation period is observed. Whether that period is representative must be assessed separately. |
| Walk-forward | Repeatedly move the forecast origin, train on the past, and evaluate on the next period | Results from multiple points in time are observed. Computational cost and relationships between folds can vary by implementation and dataset. |
| Expanding window | Keep the training start fixed while moving the endpoint forward and accumulating past observations | Earlier observations remain included. When structural change is present, the effect of retaining older data must be evaluated separately. |
| Rolling window | Move a fixed-length recent window forward and evaluate on the following period | Older observations are excluded. Sensitivity to recent changes and the appropriate window length must be established empirically. |
The third column does not imply that the cited sources prove one method to be superior. It presents conditional design interpretations based on the chronological structure of each split and the possibility of structural change in financial data.[S5][S7][S13]
With the default structure of TimeSeriesSplit, each later training set is a superset of the one before it, so observations accumulate in a way similar to an expanding window.[S5] A rolling window can instead fit the model on a moving, fixed-length period and evaluate it on the period that follows.[S13]
Financial data may undergo structural change, but this does not mean that either expanding or rolling windows are always superior.[S13] If several window lengths are tested, each length must be counted as another candidate in the search.[S9][S10]
One possible design is:
Compare candidates through walk-forward validation inside the development period, while reserving the latest unused period as a separate final test.
This is a research-design recommendation synthesized from the sources, not the only correct design for every backtest.[S3][S5][S7][S13]
The More You Search, the More Likely You Are to Find a Strategy That Looks Good by Chance
Even candidates with no genuine skill can occasionally look successful if enough of them are tested. In a backtest, each change to any of the following effectively creates another candidate:
- Strategy rules and input features
- Data start and end dates
- Holding periods and entry or exit criteria
- Models and parameters
- Window methods and lengths
- Candidate-selection metrics
When the same data is used repeatedly for inference and model selection, an apparently satisfying result may arise from chance rather than genuine methodological superiority. This is known as the data-snooping problem.[S9]
As more strategy configurations are tested, the chance of discovering impressive simulated performance by accident—and overfitting the backtest—increases.[S10][S11] If the number of attempted configurations is not disclosed, readers cannot know whether the final result was the first idea tested or the best of hundreds.[S10]
Multiple testing presents the same issue from a statistical perspective. When many financial factors are tested, decision thresholds designed for a single hypothesis may not adequately account for the risk of false discoveries. Correlations among the tests must also be considered.[S12]
The selection problem does not disappear merely because an individual backtest does not calculate formal p-values. At the same time, not every exploratory result is false, and the probability of a false discovery cannot be calculated automatically from the number of attempts alone. Relationships among candidates, the sample, the distribution of performance, and the selection rule all matter.[S10][S12]
Record the Entire Search Path, Not Just the Winning Strategy
A research log should include more than the final winner’s settings. It should also record:
- The number of candidates planned in advance and the number actually run
- The search ranges for features, periods, parameters, and window lengths
- Candidates abandoned along the way and failed experiments
- The number of charts and results inspected manually
- The candidate-selection criteria
- When the search stopped and why
- When the test period was first examined
- Any changes to code, features, or rules after examining the test
- The distinction between prior hypotheses and subsequent exploration
Failed experiments show the selection path through which the final result survived.[S9][S10][S11] If the search budget needs to be expanded, do not erase the original plan. Add the time and reason for the change.[S1]
Recordkeeping does not eliminate overfitting. It does, however, make it possible to evaluate the selection process behind the result.[S1][S10][S11]
A Small Reproducible Experiment: Same Model, Different Split
The purpose of this experiment is not to find a model with a high score. It is to keep the time-series feature and model unchanged, vary only the split method, and examine the different question each evaluation answers.
Instead of using the price of a real security, the example uses a synthetic time series with a fixed random seed. It is constructed so that the relationship between lag1 and the target changes between the earlier and later periods. Preprocessing is placed inside a Pipeline so that it is fitted only on each training set.[S4]
import numpy as npimport pandas as pdfrom sklearn.linear_model import LogisticRegressionfrom sklearn.metrics import accuracy_scorefrom sklearn.pipeline import make_pipelinefrom sklearn.preprocessing import StandardScalerrng = np.random.default_rng(42)n = 600# Synthetic time series for educational usex = rng.normal(size=n)noise = rng.normal(scale=0.8, size=n)# Change the relationship between lag1 and the target after the midpoint.score = np.empty(n)score[:400] = 0.9 * np.roll(x, 1)[:400] + noise[:400]score[400:] = -0.4 * np.roll(x, 1)[400:] + noise[400:]df = pd.DataFrame({"x": x})df["lag1"] = df["x"].shift(1)df["target"] = (score > 0).astype(int)df["regime"] = np.where(df.index < 400, "early", "late")df = df.dropna().reset_index(drop=True)X = df[["lag1"]]y = df["target"]def make_model(): return make_pipeline( StandardScaler(), LogisticRegression(), )# 1) Random splitidx = rng.permutation(len(df))cut = int(len(df) * 0.8)random_train = idx[:cut]random_test = idx[cut:]random_model = make_model()random_model.fit(X.iloc[random_train], y.iloc[random_train])random_accuracy = accuracy_score( y.iloc[random_test], random_model.predict(X.iloc[random_test]),)# 2) Chronological splittime_cut = int(len(df) * 0.8)time_train = np.arange(time_cut)time_test = np.arange(time_cut, len(df))time_model = make_model()time_model.fit(X.iloc[time_train], y.iloc[time_train])time_accuracy = accuracy_score( y.iloc[time_test], time_model.predict(X.iloc[time_test]),)print({ "random_split_accuracy": random_accuracy, "time_split_accuracy": time_accuracy, "random_test_index_range": ( int(random_test.min()), int(random_test.max()), ), "time_test_index_range": ( int(time_test.min()), int(time_test.max()), ),})for name, rows in { "random_train": random_train, "random_test": random_test, "time_train": time_train, "time_test": time_test,}.items(): print(f"\n{name}") print(df.iloc[rows]["regime"].value_counts())scikit-learn’s default random split can mix observations from different times across the two sets, while a chronological split can be designed to train on an earlier period and evaluate on the one that follows.[S5][S6] An official lagged-feature example likewise demonstrates that the same features and model can produce different generalization errors under the two splitting methods.[S8]
After running the code, fill in this table before focusing on the numerical scores.
| Item to inspect | Random split | Chronological split |
|---|---|---|
| Training indices | Scattered throughout the full period | Concentrated in the earlier period |
| Evaluation indices | Scattered throughout the full period | Concentrated in the final period |
| Later-regime observations in training | May be included | Includes only observations before the split point |
| Evaluation question | Generalization across a mixed sample | Forecasting a future period after training on the past |
| Accuracy | Record the observed value | Record the observed value |
Inspect the Timing and Regime Composition of the Data
The random evaluation indices are scattered throughout the complete period, and the random training set may contain rows from the later regime. The model does not learn the future’s exact sequence, but it is exposed to observations from a later regime that would not yet have been seen at the actual prediction time.
The chronological evaluation set occupies the final 20%, while training uses only earlier observations. This more closely matches the question of learning from the past and then evaluating on a previously unseen future period.[S5]
A difference between the two accuracy scores does not establish the superiority of the model or feature. What changed was the timing and regime composition of the training and evaluation data. If the random score is higher, examine both the inclusion of later-regime observations in its training set and the difference in regime composition between the two sets. This experiment alone cannot isolate how much each factor contributed.
The experiment has not failed if the random score is not higher. The result of the official example cannot be generalized to financial markets. The essential question is which split reproduces the condition of “knowing only the past and predicting the future.”[S8]
Leakage can still remain after chronological splitting if preprocessing or feature selection is performed over the full period.[S4] The current evidence also does not establish a universal gap size between splits when labels span overlapping periods.[S5]
How My View Changed: From the Highest Return to a Falsifiable Process
| Stage | What I would write today |
|---|---|
| Previous view | I thought the goal of backtesting was to find the chart with the highest cumulative return. But a backtest reports hypothetical performance, not actual investment results.[S2] |
| What I learned this time | The more candidates I test, the easier it becomes to select a lucky result. Repeatedly examining the same test set also brings that set into the selection process.[S3][S4][S9][S10][S11][S12] |
| Current judgment | The hypothesis, baseline, split, evaluation criteria, search budget, and final-test rules must be fixed first if I want to explain what the result means.[S1][S2][S3][S9][S10] |
| What I still do not know | I do not know which window length or market regime will remain relevant, or how well a single historical sample represents the future.[S8][S13] |
My thinking shifted from “find a better model” to “protect the model’s opportunity to be proven wrong.” I have not found a universal answer for which window method and length will work in the future, or how long a new test period should be.[S13]
Acknowledging that uncertainty is itself part of honest validation.
A One-Page Research Log to Complete Before Your Next Backtest
The following template is not a report to complete after the results arrive. It is a design document to write before running the first backtest.
My hypothesis:Result I would accept as evidence against it:Availability times of the data and each feature:Baseline strategy:Training period:Validation period and walk-forward rules:Choice of an expanding or rolling window, and rationale:Final test period:Evaluation metrics and candidate selection criteria:Strategies, features, periods, and parameters to explore:Maximum number of search trials:Where failed experiments will be recorded:Conditions for opening the final test for the first time:Items that will not be modified after inspecting the test:New unused period to obtain if modifications are needed:How prior hypotheses and post hoc exploration will be distinguished:When using it, preserve the original record and add the time and reason for every change. Count not only candidates executed in code, but also manual changes made after inspecting results. If the strategy changes after the test has been examined, reclassify that period as validation data. If no new unused period is available, state that limitation explicitly.[S3][S4]
Fees, spreads, market impact, and risk metrics are essential to interpreting real-world performance, but they are the subject of Part 5. This installment is limited to deciding how to protect the future period used for evaluation.
Conclusion: A Good Backtest Protects the Opportunity to Be Wrong
Before starting my next backtest, I plan to record five things:
- The hypothesis and its falsification conditions.[S1]
- A baseline strategy comparable to the subject of the research.[S2]
- The roles and chronological order of the training, validation, and final test periods.[S3][S4][S5]
- The total number of searches, including failed experiments.[S9][S10][S11][S12]
- The conditions for opening the final test and the rules for handling any changes made afterward.[S3][S4]
At first, I thought the goal was to find the highest curve. I now believe that the curve’s meaning can be explained only when the hypothesis, baseline, chronological split, search count, and final-test rules have been fixed before seeing the result.
A good backtest is not a procedure for selecting the highest return. It is a fair attempt to falsify a predefined question without using future information. [S1][S3][S4][S5]
Sources
- [S1] The preregistration revolution | Proceedings of the National Academy of Sciences / Brian A. Nosek et al. | 2018 | https://doi.org/10.1073/pnas.1708274114 ↩
- [S2] Investor Bulletin: Performance Claims | U.S. Securities and Exchange Commission, Office of Investor Education and Advocacy | 2022-02-17 | https://www.investor.gov/introduction-investing/general-resources/news-alerts/alerts-bulletins/investor-bulletins-47 ↩
- [S3] Datasets: Dividing the original dataset | Google for Developers | 2025 | https://developers.google.com/machine-learning/crash-course/overfitting/dividing-datasets ↩
- [S4] Common pitfalls and recommended practices | scikit-learn developers | Publication/update date not specified; stable version 1.9.0 documentation at the time of research | https://scikit-learn.org/stable/common_pitfalls.html ↩
- [S5] TimeSeriesSplit | scikit-learn developers | Publication/update date not specified; stable version 1.9.0 documentation at the time of research | https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.TimeSeriesSplit.html ↩
- [S6] train_test_split | scikit-learn developers | Publication/update date not specified; stable version 1.9.0 documentation at the time of research | https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.train_test_split.html ↩
- [S7] Time series cross-validation | Rob J. Hyndman, George Athanasopoulos / OTexts | 3rd edition, 2021; online edition continuously updated | https://otexts.com/fpp3/tscv.html ↩
- [S8] Lagged features for time series forecasting | scikit-learn developers | Publication/update date not specified; stable version 1.9.0 example at the time of research | https://scikit-learn.org/stable/auto_examples/applications/plot_time_series_lagged_features.html ↩
- [S9] A Reality Check for Data Snooping | Econometrica / Halbert White | 2000-09 | https://doi.org/10.1111/1468-0262.00152 ↩
- [S10] Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance | Notices of the American Mathematical Society / David H. Bailey, Jonathan M. Borwein, Marcos López de Prado, Qiji Jim Zhu | 2014-05; SSRN last revised 2014-04-14 | https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2308659 ↩
- [S11] The probability of backtest overfitting | The Journal of Computational Finance / David H. Bailey, Jonathan M. Borwein, Marcos López de Prado, Qiji Jim Zhu | 2017 | https://escholarship.org/uc/item/4w1110bb ↩
- [S12] … and the Cross-Section of Expected Returns | Campbell R. Harvey, Yan Liu, Heqing Zhu / National Bureau of Economic Research | NBER Working Paper 20592, 2014-10; Peer-reviewed version, 2016 | https://www.nber.org/papers/w20592 ↩
- [S13] Backtesting & Simulation | CFA Institute | 2026 | https://www.cfainstitute.org/insights/professional-learning/refresher-readings/2026/backtesting-and-simulation ↩
Report an error or share feedback
Open a draft with this article’s title and URL. Review the message and recipient before sending.
Open email draftIf no email app opens, copy these details into your usual email service.
Related posts
Quant & Data Research Can Logistic Regression Beat the Baseline? What Makes a Model Comparison Fair
Define timing, data splits, retraining, and cost rules for a fair comparison of logistic regression and baseline strategies, before testing actual performance.
Quant & Data Research Would Adding Transaction Costs Change the Conclusion? Checking Accounting and Risk Paths in a Synthetic Portfolio
Use a synthetic portfolio to check transaction-cost accounting and distinguish what ending returns, turnover, and maximum drawdown reveal.