Can Logistic Regression Beat the Baseline? What Makes a Model Comparison Fair
Evidence and scope — Experiment design and unperformed work
Real-price acquisition, model training, monthly backtests, and net-of-cost performance validation have not been performed. The probabilities are illustrative synthetic assumptions; the article defines timing and evaluation rules for a fair comparison.
This is Part 6 of “Building Your Own Quant Research System.” Can logistic regression outperform the baseline strategies? We cannot tell yet. SPY, IEF, and GLD are research candidates only. For this article, we have not acquired actual price data, trained a model, run monthly backtests, calculated performance after costs, or tested whether one strategy outperforms another.
Part 5 used synthetic examples to check whether our cost, turnover, and drawdown accounting followed consistent rules. Those checks did not establish logistic regression’s predictive power or profitability after costs. This article’s goal is to define the timeline and evaluation rules before we obtain real time series, so we can compare our first ML model fairly against the baselines.
A Probability Estimates an Event, Not a Return
Logistic regression is a classification model that estimates the conditional probability of a binary event. [S1] If we define the label as “Was the return positive over the next period in which we could actually hold the asset?”, the model’s output estimates the probability of that event.
A probability of 0.60 therefore does not mean an expected return of 60%. It does not mean a profit is guaranteed, and it does not represent the strategy’s overall return. The probability is an output from a classification problem. Strategy performance is the economic result of translating that signal into trades under specific execution-price, allocation, cost, and cash rules. Treating these as interchangeable obscures what we are testing. [S1][S3][S4]
Put Observations, Signals, Execution, and Labels on One Timeline
In financial time series, “When could we have known this?” matters as much as the feature value itself. Information leakage can occur when information finalized or published later is incorporated into an earlier signal. Evaluation must therefore respect chronological order. [S2][S3]
For this design, we assume that features use historical returns and volatility through the month-end close, and that signals are generated at that close. Orders are placed and executed at the next trading day’s close.
| Stage | Example rule to fix in advance | Information allowed at that point |
|---|---|---|
| Feature observation cutoff | Calculate historical returns and volatility through the month-end close | Inputs finalized by the month-end close |
| Signal generation | Estimate the probability of a positive return over the next holding period using month-end information | Information available by the feature observation cutoff |
| Order placement and execution | Assume execution at the next trading day’s close | Use only the month-end signal |
| Label period | Define the outcome using the return from one execution to the next | Exclude the return between signal generation and execution |
| Label finalization and retraining | Make the label eligible for training after the next execution time | Use only finalized historical labels |
On this timeline, the label is not a loosely defined “one-month return after the month-end signal.” It indicates whether the return is positive from execution at the next trading day’s close through the following execution—the period during which the position could actually be held. Price movements between the signal and execution are excluded from the label return because the position was not yet held. This is a conservative design choice inferred from the need to keep feature observations, target periods, and information availability properly aligned. [S3]
A label also cannot be finalized before its holding period ends. If the holding period for this month’s signal is still open when next month’s retraining occurs, that label cannot enter the training set. Calling something “historical data” is not enough; we must record when its label became final. [S3]
Separate Development, Validation, and Final Testing
One possible design assigns 2007~2017 to development, 2018~2020 to validation, and 2021~2025 to final testing. These dates illustrate a plan to fix in advance. They do not imply that the data has been acquired or that any experiment has been run.
At both the development–validation boundary and the validation–final-test boundary, this design removes samples from the earlier segment if their label holding periods extend into the later segment. This prevents price movements from the later segment from entering training or model selection in the earlier one. [S3][S4][S5] The chronological split defines the evaluation boundaries; the monthly retraining rule, fixed in advance, defines when the training window may expand during testing. The next section spells out those conditions. [S3][S4][S5] This is also why we preserve chronological order instead of splitting randomly. Standard cross-validation assumptions may break down for time series, making evaluation on a later period necessary to reflect actual use. [S2][S5]
Select the feature list, scaling method, and probability threshold using only the development and validation periods. If we reselect features or thresholds because they look good on the final test, that test is no longer an independent final evaluation. Keeping test data out of model and preprocessing selection helps reduce overly optimistic results. [S4][S5]
Respect the Boundaries in Preprocessing and Monthly Retraining
Standardization is useful, but even a mean or standard deviation can contain future information. At each retraining date, fit the scaler only on the training window. For validation and test observations, apply only transform using those fitted parameters. [S4][S5]
Monthly expanding-window training enlarges the training set as more historical data becomes available. This is different from redesigning the model in response to test performance. During the test period, retraining follows a monthly schedule fixed in advance, adding only historical labels whose holding periods have ended and whose outcomes are final by the current signal-generation time. The feature list, scaling method, and threshold remain as selected during development and validation; test results must not drive reselection. [S3][S4][S5]
The code below is an unexecuted educational example showing the structure of a reproducible workflow. It does not indicate that installation, data preparation, model fitting, or implementation of the full system has been completed.
from sklearn.pipeline import Pipelinefrom sklearn.preprocessing import StandardScalerfrom sklearn.linear_model import LogisticRegressionmodel = Pipeline([ ("scaler", StandardScaler()), ("classifier", LogisticRegression(max_iter=1000)),])# At each retraining date:# Include only past labels that are finalized by that date in X_train and y_train.# model.fit(X_train, y_train)# p_up = model.predict_proba(X_current)[:, 1]Actual runs must also record library versions and configuration settings. A scikit-learn 1.9.1 release has been announced, so a record that merely says “used scikit-learn” may not capture enough information to reproduce a run. [S6]
Set the Insufficient-Data Policy in Advance
Deciding how to handle small samples, a single class, or missing features after seeing the results can introduce selection bias. [S3][S4][S5] The following rules are examples of policies to document before an experiment, not claims about optimal settings.
- The minimum training sample size is a design parameter whose numerical value has not yet been set. Set it before running the experiment. If the number of finalized labels falls below that minimum, skip model training and signal generation for that month.
- If the training window contains only one label class—up or down—skip model training and signal generation for that month.
- If any feature required for training or current signal generation is missing, skip model training and signal generation for that month.
- Whenever signals are skipped for one of these reasons, set the target allocation to 100% cash. Sell any existing holdings at the scheduled next-trading-day closing execution and apply the common cost rules.
- This move-to-cash policy illustrates consistent handling; it has not been validated as the best approach. It follows the principle that exception rules should not change in response to observed results. [S3][S4][S5]
The point is to decide what happens when the model cannot run before it happens, rather than switching afterward to whichever fallback would have performed best.
An Illustrative Example: From Probabilities to Weights
The probabilities below are neither trained-model outputs nor ETF experiment results. They are synthetic assumptions used solely to illustrate how a probability threshold translates into portfolio weights.
Assume a fixed threshold of 0.55 and synthetic probabilities of 0.60, 0.45, and 0.70 for SYN_A, SYN_B, and SYN_C, respectively.
| Asset | Illustrative synthetic probability—not a model output | Meets threshold? | Illustrative weight |
|---|---|---|---|
| SYN_A | 0.60 | Yes | 0.5 |
| SYN_B | 0.45 | No | 0 |
| SYN_C | 0.70 | Yes | 0.5 |
Under these assumptions, SYN_A and SYN_C each receive a weight of 0.5. If every asset falls below the threshold, the portfolio holds cash. This educational example assumes a cash return of 0. We have not established whether 0.55 is an effective threshold, whether equal weighting is appropriate, or whether this approach produces useful performance after costs on actual SPY, IEF, and GLD data. [S1][S3][S5]
Fair Comparisons Start with Common Accounting
A general validation principle is that evaluation should reflect how a model will actually be used. [S5] Applying that principle, this article adopts a comparison design in which buy-and-hold, momentum, and ML strategies share the same start and end dates, initial capital, execution prices, cost rates, cash treatment, and end-of-period valuation rules. These specific ETF trading assumptions are design choices for this article, not rules prescribed directly by the cited literature.
The cost formula based on actual traded value and the turnover definition established in Part 5 must also apply equally to every strategy. A comparison is unfair if initial entry costs apply only to the ML strategy, or if liquidation costs and cash rules differ across strategies. If a longer warm-up period delays a strategy’s first trade, disclose the warm-up period separately from the common evaluation period.
In particular, we must not place Part 5’s synthetic, single-period accounting results beside real ML performance from a different future evaluation period and declare a winner. The performance comparisons defined here require the same data, period, execution rules, and cost accounting across strategies.
Accuracy and probability evaluation ask, “How well did the model predict the labels?” Returns after costs, maximum drawdown, and turnover ask, “What happened when those signals were translated into trading rules?” Improvement in one does not automatically guarantee improvement in the other. [S1][S3][S4][S5]
What to Record in the Next Run
This installment has laid out the comparison timeline, evaluation principles, and example exception policies. The numerical minimum for training sample size remains unset, and empirical validation is still incomplete. If Part 7 takes us into an actual run, we will need to preserve the following alongside the results:
- Run ID and data identifiers
- Data period and feature list
- Label start, end, and finalization times
- Development, validation, and test boundaries, including labels removed because they crossed those boundaries
- Retraining schedule, cost formula, execution assumptions, and cash rules
- Package versions and random seeds
- Records of failed runs, excluded data, and exception handling
These records make it possible to revisit the same experiment whether its results are good or bad. Until then, declaring that logistic regression either beats or fails to beat the baselines would be premature.
This article is a design document for education and research. It is not investment advice or a guarantee of returns.
Sources
- [S1] Statistics review 14: Logistic regression | Viv Bewick, Liz Cheek, Jonathan Ball; Critical Care | 2005-01-13 | https://pmc.ncbi.nlm.nih.gov/articles/PMC1065119/ ↩
- [S2] On the use of cross-validation for time series predictor evaluation | Christoph Bergmeir and José M. Benítez; Information Sciences | 2012-05-15 | https://doi.org/10.1016/j.ins.2011.12.028 ↩
- [S3] Information leakage in financial machine learning research | Zachary David; Algorithmic Finance | 2019-12-19 | https://journals.sagepub.com/doi/10.3233/AF-190900 ↩
- [S4] Being Aware of Data Leakage and Cross-Validation Scaling in Chemometric Model Validation | Péter Király and Gergely Tóth; Journal of Chemometrics | 2025-04-01 | https://analyticalsciencejournals.onlinelibrary.wiley.com/doi/10.1002/cem.70026 ↩
- [S5] A Set of Rules for Model Validation | José Camacho; Journal of Chemometrics | 2026-02-17 | https://analyticalsciencejournals.onlinelibrary.wiley.com/doi/10.1002/cem.70110 ↩
- [S6] Scikit-learn 1.9.1 | scikit-learn project | 2026-09-11 | https://github.com/scikit-learn/scikit-learn/releases/tag/1.9.1 ↩
Report an error or share feedback
Open a draft with this article’s title and URL. Review the message and recipient before sending.
Open email draftIf no email app opens, copy these details into your usual email service.
Related posts
Quant & Data Research Part 7. First Machine Learning Model: Testing the Probability of Gains in the Next Rebalancing Period with Logistic Regression
Design features, labels, time splits, and monthly retraining for a logistic-regression baseline, keeping predictive accuracy separate from investment returns.
Quant & Data Research The Risk of Mistaking Chance Results for Skill in Financial Machine Learning—and How to Challenge Them
Examine noise, changing markets, repeated searches, leakage, costs, and selection bias, and define ways to challenge promising financial ML results.