Forge Fate
Quant & Data Research

What Makes a Good Backtest Different? Designing the Validation Process Before Looking at Returns

16 min read

Evidence and scope — Artificial-series validation example

Literature and a fixed-seed artificial time-series example illustrate time splits and evaluation procedures. This is not an experiment establishing real-asset performance; costs and risk evaluation are outside the example’s scope.

A practical introduction to preregistering hypotheses, choosing baseline strategies, splitting data chronologically, using walk-forward validation, and protecting the test period

At first, I thought the goal of backtesting was to run many strategies and select the chart with the highest cumulative return. If the curve looked smooth and ended with a large gain, it felt as though I had discovered a good strategy.

But backtest returns are hypothetical results calculated from historical data, not actual investment performance.[S2] Their meaning can change simply because the comparison benchmark was changed or a favorable period was selected.[S2] Worse, if enough strategies, features, periods, and parameters are tested, it becomes possible to select a candidate that stands out because it was lucky rather than genuinely superior.[S9][S10][S11]

That led me to change the question.

Before asking how high the result was, how effectively did the process prevent future information and after-the-fact human choices from influencing it?

A good backtest is not a procedure for selecting the highest return. It is a procedure for making a fair attempt to falsify a predefined question without using future information.

The tables and code in this article are educational and research examples designed to explain backtest validation. They are not recommendations of specific securities or instructions to trade, and they do not guarantee returns. Differences observed in the example scores must not be generalized as evidence that one real asset, market, period, or investment strategy is superior.[S2][S8]

After Aligning the Data Clock, Align the Experiment Clock

Part 3 centered on one question:

Was that value actually observable when the trading decision had to be made?

In this installment, we will assume that the data’s creation and observation times have been aligned correctly. The next step is to decide how to divide that data, which period will be used to make choices, and which period will remain unseen until the end.

Even when individual features contain no future values, information from the test period can enter the training process if you calculate scaling parameters or imputation rules over the full period, select features, and only then split the data.[S4] Conversely, even a chronological split will not eliminate the leakage discussed in Part 3 if a feature’s observation time is incorrect.

Chronological splitting is an important line of defense, but it is not a button that automatically eliminates every form of leakage.[S4][S5]

A Good Backtest Records the Question Before the Result

Defining the research question and analysis plan before seeing the results makes it easier to distinguish an original prediction from an explanation created afterward.[S1] The same principle can be applied to a research log for financial backtesting.

This does not mean that every independent researcher must publicly preregister every experiment. Plans and revision histories can be preserved in timestamped notes, Git commits, or experiment-tracking tools. Based on the current evidence, however, I cannot say which recording method is the most reliable.

Item to record in advanceQuestion to answer
HypothesisWhat information do I expect to be related to the next period’s target, and why?
Falsification conditionWhat result would lead me to conclude that the hypothesis is not supported?
Data timingWas each feature available before the decision had to be made?
BaselineWhat must the method beat before I can claim an incremental improvement?
Split rulesWhat are the dates and chronological order of the training, validation, and final test periods?
Evaluation criteriaWhich metric and aggregation rule will be used to select candidates?
Search budgetWhat is the maximum number of strategy, feature, period, and parameter combinations I will test?
Final testWhen will I open it, and what will I refuse to change after seeing the result?

The purpose of recording a plan in advance is not to prohibit exploration. Exploratory analysis inspired by the results can also be useful. The important point is to distinguish it from the original plan rather than presenting it as though it had been a confirmatory hypothesis from the beginning.[S1]

Preregistration does not make a flawed hypothesis correct.[S1] It does, however, make it harder to change the question to fit the result and later remember the revised question as the one you tested all along.

Choose the Baseline Strategy Before Building a Complex Model

Before building a complex model, define a baseline that is comparable to the subject of the research. Comparing a backtested strategy with a benchmark representing a different market or type of investment can undermine the meaning of the comparison.[S2]

Buy and hold is not the only valid baseline for every study. The appropriate comparison depends on the research question.

Research questionPossible baseline strategy
Is this better than simply holding the asset?Buy and hold for the same asset and period
Is a predictive model necessary?Carry forward the previous value, use a long-term average, or always make the same prediction
Is the complex model better?A simple rule or reduced model addressing the same target

For example, when evaluating a model that classifies upward and downward moves, a simple prediction rule using the same sample and target can serve as the comparison. When comparing a complex signal with holding the asset, the benchmark should cover the same asset and period.

If the baseline is replaced after seeing the results with one that makes the strategy look better, the choice of baseline becomes part of the selection process. The comparison should therefore be recorded in advance alongside the research question.[S1][S2][S9]

Training, Validation, and Testing: Studying, Taking a Practice Exam, and Opening the Sealed Final

The three datasets should be distinguished by their roles, not merely their names.

text
Training periodFit model coefficients and preprocessing statistics.Validation periodSelect features, rules, parameters, window lengths, and candidates.Final test periodInspect the final results only after all selections are complete.

Training data is used to fit the model. Validation data is used to choose among candidates. Test data is used only after all selection is complete, to evaluate final generalization performance.[S3]

A useful analogy is studying, taking a practice exam, and sitting a sealed final exam. It is perfectly normal to revise your study method after reviewing a practice exam. But if you inspect the final exam questions and score, revise your answers, and then take the same exam again, it is no longer an independent final evaluation.

In financial time series, chronological order matters in addition to those roles.

text
Past                                                      Future|-------- Training --------|---- Validation ----|-- Final test --|

The scaler’s mean and standard deviation, imputation values, feature-selection criteria, and model coefficients must all be estimated using only the relevant training period. Fitting preprocessing on the complete dataset allows test information to enter the training pipeline and can make the evaluation overly optimistic.[S4]

Example ratios for training, validation, and testing should not be treated as universally optimal.[S3] The specific length of each period must be designed and evaluated separately for each experiment.

The Moment You See the Test Result, You Become Part of the Training Process

The following sequence is common.

text
Inspect the final test→ Find the result unsatisfactory→ Modify parameters, features, or rules→ Recheck the same test

Even if the code never directly calls fit on the test rows, the researcher learns something from the first test result. If that information is then used to change features or parameters, the test data has indirectly influenced model selection.[S3][S4]

From that point onward, the period is effectively validation data rather than an independent final test. One of the following responses is then necessary:

  • Designate a later, previously unused period as the new test period.
  • Conduct forward validation on observations that become available as real time passes.
  • If no new data is available, label the result as “exploration adapted to the existing test” rather than “final validation completed.”

The current evidence does not establish a universally appropriate length for a new, unused test period. Its length must therefore remain an explicitly documented, unresolved design choice for each study rather than being treated as a validated general standard.

Looking at the test result more than once does not invalidate the entire research project. It does mean that the period’s status must be changed honestly. Whether data is truly a test set is determined not by its filename or date, but by whether it was actually used during the selection process.

Why Random Splitting Changes the Question for Financial Time Series

Many introductory machine-learning examples shuffle observations and use some for training and the rest for evaluation. By default, scikit-learn’s train_test_split also shuffles the data.[S6]

Applying that default directly to time-indexed data can create a structure like this:

text
Random splitTraining: 1, 3, 4, 7, 9Evaluation: 2, 5, 6, 8, 10Chronological splitTraining: 1, 2, 3, 4, 5, 6, 7Evaluation: 8, 9, 10

Under a random split, the model could be evaluated at time 2 after being trained on later observations from times 3, 4, 7, and 9. Standard cross-validation can create folds that train on future observations and evaluate on past ones if it is applied to time series without modification.[S5]

The real-world forecasting question is usually closer to this:

Knowing only what was available up to that point, how well could I have predicted the next period?

To reproduce that question, observations later than the evaluation point must be excluded from training.[S5][S7] Even if no individual feature directly contains a future value, a randomly assembled training set can contain distributions and relationships observed during later market regimes. In other words, the model may be exposed to later-period evidence that would not have been available at the actual prediction time.

This does not mean random splitting is wrong for every analysis. It may be appropriate for some research questions, including cross-sectional problems that are reasonably close to being independent and identically distributed.[S6] The problem is using shuffling uncritically when the experiment is supposed to reproduce forecasting from the past into the future.

Nor does random splitting invariably produce a higher score. An official lagged-feature example found that random splitting produced a more optimistic error estimate than time-based evaluation, but that observation came from bicycle-demand data.[S8] The direction and magnitude of the difference cannot be generalized to financial data.

How Is a Single Split Different from Walk-Forward Validation?

The simplest chronological design trains once on a past period and evaluates once on the period that follows.

text
Fixed split[──────── Training ────────][──── Evaluation ────]

Walk-forward validation repeatedly moves the forecast origin forward, trains on observations preceding each evaluation window, and evaluates on the next period.[S5][S7][S13]

text
Fold 1  [Training────][Evaluation]Fold 2  [Training────────][Evaluation]Fold 3  [Training────────────][Evaluation]

The key feature of this procedure is that it allows errors from multiple forecast points to be aggregated.[S7] Two common ways of constructing the windows are the expanding window and the rolling window.

MethodChronological structureConditional interpretation and design questions
Fixed splitOne past training period followed by one future evaluation periodOnly one evaluation period is observed. Whether that period is representative must be assessed separately.
Walk-forwardRepeatedly move the forecast origin, train on the past, and evaluate on the next periodResults from multiple points in time are observed. Computational cost and relationships between folds can vary by implementation and dataset.
Expanding windowKeep the training start fixed while moving the endpoint forward and accumulating past observationsEarlier observations remain included. When structural change is present, the effect of retaining older data must be evaluated separately.
Rolling windowMove a fixed-length recent window forward and evaluate on the following periodOlder observations are excluded. Sensitivity to recent changes and the appropriate window length must be established empirically.

The third column does not imply that the cited sources prove one method to be superior. It presents conditional design interpretations based on the chronological structure of each split and the possibility of structural change in financial data.[S5][S7][S13]

With the default structure of TimeSeriesSplit, each later training set is a superset of the one before it, so observations accumulate in a way similar to an expanding window.[S5] A rolling window can instead fit the model on a moving, fixed-length period and evaluate it on the period that follows.[S13]

Financial data may undergo structural change, but this does not mean that either expanding or rolling windows are always superior.[S13] If several window lengths are tested, each length must be counted as another candidate in the search.[S9][S10]

One possible design is:

Compare candidates through walk-forward validation inside the development period, while reserving the latest unused period as a separate final test.

This is a research-design recommendation synthesized from the sources, not the only correct design for every backtest.[S3][S5][S7][S13]

The More You Search, the More Likely You Are to Find a Strategy That Looks Good by Chance

Even candidates with no genuine skill can occasionally look successful if enough of them are tested. In a backtest, each change to any of the following effectively creates another candidate:

  • Strategy rules and input features
  • Data start and end dates
  • Holding periods and entry or exit criteria
  • Models and parameters
  • Window methods and lengths
  • Candidate-selection metrics

When the same data is used repeatedly for inference and model selection, an apparently satisfying result may arise from chance rather than genuine methodological superiority. This is known as the data-snooping problem.[S9]

As more strategy configurations are tested, the chance of discovering impressive simulated performance by accident—and overfitting the backtest—increases.[S10][S11] If the number of attempted configurations is not disclosed, readers cannot know whether the final result was the first idea tested or the best of hundreds.[S10]

Multiple testing presents the same issue from a statistical perspective. When many financial factors are tested, decision thresholds designed for a single hypothesis may not adequately account for the risk of false discoveries. Correlations among the tests must also be considered.[S12]

The selection problem does not disappear merely because an individual backtest does not calculate formal p-values. At the same time, not every exploratory result is false, and the probability of a false discovery cannot be calculated automatically from the number of attempts alone. Relationships among candidates, the sample, the distribution of performance, and the selection rule all matter.[S10][S12]

Record the Entire Search Path, Not Just the Winning Strategy

A research log should include more than the final winner’s settings. It should also record:

  • The number of candidates planned in advance and the number actually run
  • The search ranges for features, periods, parameters, and window lengths
  • Candidates abandoned along the way and failed experiments
  • The number of charts and results inspected manually
  • The candidate-selection criteria
  • When the search stopped and why
  • When the test period was first examined
  • Any changes to code, features, or rules after examining the test
  • The distinction between prior hypotheses and subsequent exploration

Failed experiments show the selection path through which the final result survived.[S9][S10][S11] If the search budget needs to be expanded, do not erase the original plan. Add the time and reason for the change.[S1]

Recordkeeping does not eliminate overfitting. It does, however, make it possible to evaluate the selection process behind the result.[S1][S10][S11]

A Small Reproducible Experiment: Same Model, Different Split

The purpose of this experiment is not to find a model with a high score. It is to keep the time-series feature and model unchanged, vary only the split method, and examine the different question each evaluation answers.

Instead of using the price of a real security, the example uses a synthetic time series with a fixed random seed. It is constructed so that the relationship between lag1 and the target changes between the earlier and later periods. Preprocessing is placed inside a Pipeline so that it is fitted only on each training set.[S4]

python
import numpy as npimport pandas as pdfrom sklearn.linear_model import LogisticRegressionfrom sklearn.metrics import accuracy_scorefrom sklearn.pipeline import make_pipelinefrom sklearn.preprocessing import StandardScalerrng = np.random.default_rng(42)n = 600# Synthetic time series for educational usex = rng.normal(size=n)noise = rng.normal(scale=0.8, size=n)# Change the relationship between lag1 and the target after the midpoint.score = np.empty(n)score[:400] = 0.9 * np.roll(x, 1)[:400] + noise[:400]score[400:] = -0.4 * np.roll(x, 1)[400:] + noise[400:]df = pd.DataFrame({"x": x})df["lag1"] = df["x"].shift(1)df["target"] = (score > 0).astype(int)df["regime"] = np.where(df.index < 400, "early", "late")df = df.dropna().reset_index(drop=True)X = df[["lag1"]]y = df["target"]def make_model():    return make_pipeline(        StandardScaler(),        LogisticRegression(),    )# 1) Random splitidx = rng.permutation(len(df))cut = int(len(df) * 0.8)random_train = idx[:cut]random_test = idx[cut:]random_model = make_model()random_model.fit(X.iloc[random_train], y.iloc[random_train])random_accuracy = accuracy_score(    y.iloc[random_test],    random_model.predict(X.iloc[random_test]),)# 2) Chronological splittime_cut = int(len(df) * 0.8)time_train = np.arange(time_cut)time_test = np.arange(time_cut, len(df))time_model = make_model()time_model.fit(X.iloc[time_train], y.iloc[time_train])time_accuracy = accuracy_score(    y.iloc[time_test],    time_model.predict(X.iloc[time_test]),)print({    "random_split_accuracy": random_accuracy,    "time_split_accuracy": time_accuracy,    "random_test_index_range": (        int(random_test.min()),        int(random_test.max()),    ),    "time_test_index_range": (        int(time_test.min()),        int(time_test.max()),    ),})for name, rows in {    "random_train": random_train,    "random_test": random_test,    "time_train": time_train,    "time_test": time_test,}.items():    print(f"\n{name}")    print(df.iloc[rows]["regime"].value_counts())

scikit-learn’s default random split can mix observations from different times across the two sets, while a chronological split can be designed to train on an earlier period and evaluate on the one that follows.[S5][S6] An official lagged-feature example likewise demonstrates that the same features and model can produce different generalization errors under the two splitting methods.[S8]

After running the code, fill in this table before focusing on the numerical scores.

Item to inspectRandom splitChronological split
Training indicesScattered throughout the full periodConcentrated in the earlier period
Evaluation indicesScattered throughout the full periodConcentrated in the final period
Later-regime observations in trainingMay be includedIncludes only observations before the split point
Evaluation questionGeneralization across a mixed sampleForecasting a future period after training on the past
AccuracyRecord the observed valueRecord the observed value

Inspect the Timing and Regime Composition of the Data

The random evaluation indices are scattered throughout the complete period, and the random training set may contain rows from the later regime. The model does not learn the future’s exact sequence, but it is exposed to observations from a later regime that would not yet have been seen at the actual prediction time.

The chronological evaluation set occupies the final 20%, while training uses only earlier observations. This more closely matches the question of learning from the past and then evaluating on a previously unseen future period.[S5]

A difference between the two accuracy scores does not establish the superiority of the model or feature. What changed was the timing and regime composition of the training and evaluation data. If the random score is higher, examine both the inclusion of later-regime observations in its training set and the difference in regime composition between the two sets. This experiment alone cannot isolate how much each factor contributed.

The experiment has not failed if the random score is not higher. The result of the official example cannot be generalized to financial markets. The essential question is which split reproduces the condition of “knowing only the past and predicting the future.”[S8]

Leakage can still remain after chronological splitting if preprocessing or feature selection is performed over the full period.[S4] The current evidence also does not establish a universal gap size between splits when labels span overlapping periods.[S5]

How My View Changed: From the Highest Return to a Falsifiable Process

StageWhat I would write today
Previous viewI thought the goal of backtesting was to find the chart with the highest cumulative return. But a backtest reports hypothetical performance, not actual investment results.[S2]
What I learned this timeThe more candidates I test, the easier it becomes to select a lucky result. Repeatedly examining the same test set also brings that set into the selection process.[S3][S4][S9][S10][S11][S12]
Current judgmentThe hypothesis, baseline, split, evaluation criteria, search budget, and final-test rules must be fixed first if I want to explain what the result means.[S1][S2][S3][S9][S10]
What I still do not knowI do not know which window length or market regime will remain relevant, or how well a single historical sample represents the future.[S8][S13]

My thinking shifted from “find a better model” to “protect the model’s opportunity to be proven wrong.” I have not found a universal answer for which window method and length will work in the future, or how long a new test period should be.[S13]

Acknowledging that uncertainty is itself part of honest validation.

A One-Page Research Log to Complete Before Your Next Backtest

The following template is not a report to complete after the results arrive. It is a design document to write before running the first backtest.

text
My hypothesis:Result I would accept as evidence against it:Availability times of the data and each feature:Baseline strategy:Training period:Validation period and walk-forward rules:Choice of an expanding or rolling window, and rationale:Final test period:Evaluation metrics and candidate selection criteria:Strategies, features, periods, and parameters to explore:Maximum number of search trials:Where failed experiments will be recorded:Conditions for opening the final test for the first time:Items that will not be modified after inspecting the test:New unused period to obtain if modifications are needed:How prior hypotheses and post hoc exploration will be distinguished:

When using it, preserve the original record and add the time and reason for every change. Count not only candidates executed in code, but also manual changes made after inspecting results. If the strategy changes after the test has been examined, reclassify that period as validation data. If no new unused period is available, state that limitation explicitly.[S3][S4]

Fees, spreads, market impact, and risk metrics are essential to interpreting real-world performance, but they are the subject of Part 5. This installment is limited to deciding how to protect the future period used for evaluation.

Conclusion: A Good Backtest Protects the Opportunity to Be Wrong

Before starting my next backtest, I plan to record five things:

  1. The hypothesis and its falsification conditions.[S1]
  2. A baseline strategy comparable to the subject of the research.[S2]
  3. The roles and chronological order of the training, validation, and final test periods.[S3][S4][S5]
  4. The total number of searches, including failed experiments.[S9][S10][S11][S12]
  5. The conditions for opening the final test and the rules for handling any changes made afterward.[S3][S4]

At first, I thought the goal was to find the highest curve. I now believe that the curve’s meaning can be explained only when the hypothesis, baseline, chronological split, search count, and final-test rules have been fixed before seeing the result.

A good backtest is not a procedure for selecting the highest return. It is a fair attempt to falsify a predefined question without using future information. [S1][S3][S4][S5]

Sources

Report an error or share feedback

Open a draft with this article’s title and URL. Review the message and recipient before sending.

To: [email protected]

Open email draft

If no email app opens, copy these details into your usual email service.

Contact information