Part 7. First Machine Learning Model: Testing the Probability of Gains in the Next Rebalancing Period with Logistic Regression
Evidence and scope — Machine-learning baseline design
This article defines features, labels, time splits, retraining, and trading rules. It is not a performance report based on a fixed ETF data snapshot and presents no actual return or accuracy figures.
Accuracy Is Not Return: Designing a First ML Baseline with Logistic Regression
In Part 6, we built a simple baseline: at each month-end, hold only the ETFs whose 12-month momentum is positive, and otherwise stay in cash. This time, we keep the comparison conditions unchanged and replace only the signal-generation method. Instead of using the sign of past returns, we use the probability that logistic regression assigns to a gain over the next rebalancing period.
At first, I expected a machine-learning model to make better decisions than simple momentum simply because it could consider several features at once. Building the first model changed the question that mattered most.
The first question is not whether the model is more accurate. It is whether the model was compared with the baseline under the same conditions without using future information.
This article is not a performance report based on a particular ETF dataset. Reporting accuracy or return figures before fixing a verified snapshot—including the data provider, price-field definition, ETF universe, and observation period—would produce a result that readers could not reproduce. Instead, this article fixes the features, label, chronological split, retraining process, trading rule, and decision criteria needed for a real experiment.
The code and explanations are for educational and research purposes. They do not recommend buying or selling any ETF or security, and they do not guarantee returns or future performance. A backtest is hypothetical performance calculated under fixed data and assumptions. [S7]
What Changes—and What Does Not
For a fair comparison, the conditions established in Part 6 remain fixed.
| Item | Part 6: simple momentum | Part 7: logistic regression |
|---|---|---|
| ETF universe and price data | Fixed in advance | Use the same snapshot |
| Signal time | Month-end | Same |
| When new weights take effect | From the next trading day | Same |
| Rebalancing | Monthly | Same |
| Position constraint | Long or cash | Same |
| Allocation among selected assets | Equal weight | Same |
| Cost sensitivity | 0, 5, 10, and 20 bp one-way | Same |
| Signal | Sign of the trailing 12-month return | Estimated probability and a fixed threshold |
Only the last row changes. If we change the ETF universe, period, costs, or execution timing at the same time as the model, we cannot tell what caused the performance difference.
1. Define the Prediction Problem in One Sentence
The model in this article answers one question.
Using only price, volatility, and volume features observable by month-end, can we classify whether the return over the next monthly holding period will be positive?
The positive class means that the next holding-period return is greater than zero.
label_t = 1 if next_holding_return_t > 0
label_t = 0 otherwisescikit-learn's LogisticRegression is a classifier, and predict_proba returns an estimated probability for each class. [S1] A model output of 0.62 does not establish a 62% real-world chance that the asset will rise. It is better understood as a score calculated under the training data and model assumptions.
This distinction matters throughout the experiment.
| Term | Meaning here |
|---|---|
| Probability of a gain | The model's estimated probability for the positive class |
| Classification accuracy | The share of predicted classes that match the labels |
| Strategy return | Hypothetical performance after converting predictions into trading rules |
| Investability | A separate judgment that also considers costs, liquidity, execution, and risk |
A classifier can be slightly more accurate and still produce worse post-cost performance if its losing trades are larger or it trades too frequently. Accuracy and investment performance therefore belong in separate evaluations.
2. Separate Feature Time from Label Time
Every feature candidate is restricted to information already observable at month-end.
| Feature | Calculation window | Observable at month-end? |
|---|---|---|
| Most recent 1-day return | One trading day through month-end | Yes |
| Most recent 5-day return | Five trading days through month-end | Yes |
| Most recent 20-day return | Twenty trading days through month-end | Yes |
| Moving-average spread | 20-day and 120-day moving averages | Yes |
| Recent volatility | Standard deviation of daily returns over 20 days | Yes |
| Volume change | Difference between recent volume and its prior average | Yes |
| Market return | 20-day return of a market proxy | Yes |
| Next holding-period return | Through the next rebalance | No; label only |
Values appearing in the same data row were not necessarily knowable at the same time. Attaching next month's return to the current row creates a training label; it does not make that return a valid input at prediction time.
monthly["forward_return"] = (
monthly.groupby("symbol")["execution_close"]
.transform(lambda s: s.shift(-1).div(s).sub(1.0))
)
monthly["label"] = monthly["forward_return"].gt(0).astype("int8")The negative shift aligns the next row's outcome with the current signal row. [S4] The final observation for each asset has no completed next holding period, so it must be excluded from training and evaluation.
There is another timing condition. When the model is retrained at month-end t, even an earlier signal row is ineligible if its holding period has not ended and its label was not yet known.
eligible_train = samples[
(samples["signal_date"] < current_signal_date)
& (samples["label_end_date"] < current_signal_date)
]Without this condition, code can select rows that look historical while still learning from outcomes that were unavailable on the prediction date.
3. Preserve Calendar Order Instead of Randomly Splitting
Randomly shuffling time-series samples into training and test sets can allow future information to influence an evaluation of the past. TimeSeriesSplit supports chronological folds and a gap between the training and validation segments. [S2]
This experiment follows a simple calendar structure.
Past Future
|──────── Training ────────|── Validation ──|──── Test ────|
Fit Finalize choices Evaluate once- Fit the model and preprocessing rules in the training segment.
- Compare only a small, predeclared set of thresholds and model settings in validation.
- Evaluate the test segment once after all choices are fixed.
- If features or thresholds change after inspecting test performance, that segment is no longer a test set.
The date boundaries should be recorded after examining the real dataset's available history and market regimes. The important point is not a particular calendar year, but fixing the boundaries before evaluating performance and never reversing the direction of time.
4. Learn Preprocessing Only from the Training Window
Scaling can leak information even though it does not look predictive on its own. If the scaler is fitted on the full dataset, the mean and variance of the test period influence training. scikit-learn recommends fitting preprocessing only on training data and applying the learned transformation elsewhere. [S3]
The simplest implementation is to combine preprocessing and the model in a Pipeline.
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
model = Pipeline(
steps=[
("scale", StandardScaler()),
(
"classifier",
LogisticRegression(
C=1.0,
max_iter=2_000,
random_state=7,
),
),
]
)
model.fit(train[feature_columns], train["label"])
probability = model.predict_proba(current[feature_columns])
classes = model.named_steps["classifier"].classes_
positive_class_index = list(classes).index(1)
current["up_probability"] = probability[:, positive_class_index]The code locates class 1 in classes_ instead of assuming that the positive-class probability is always the second column. Before fitting, it should check not only for missing values but also for positive and negative infinity. A volume-change calculation can produce infinity when its prior average is zero; stopping with a data error is safer than silently replacing it.
5. Walk Forward Month by Month
If the strategy were used at a real month-end, it could access only the information finalized by that date. The test period should therefore be processed as a monthly walk-forward exercise rather than predicted all at once.
predictions = []
for signal_date in test_signal_dates:
train = samples[
(samples["signal_date"] < signal_date)
& (samples["label_end_date"] < signal_date)
].copy()
current = samples[samples["signal_date"] == signal_date].copy()
model.fit(train[feature_columns], train["label"])
probability = model.predict_proba(current[feature_columns])
classes = model.named_steps["classifier"].classes_
positive_index = list(classes).index(1)
current["up_probability"] = probability[:, positive_index]
predictions.append(current)An expanding window, in which past observations accumulate over time, is the default design here. A fixed-length window is also possible, but its length should be chosen before performance is inspected. Monthly retraining is not intended to maximize returns; it is intended to reproduce the information set that would actually have been available each month.
6. Convert Probabilities into a Trading Rule
Model scores alone do not define a backtest. The experiment must specify the probability threshold, the allocation when several ETFs qualify, and the position when none qualify.
The first rule is deliberately simple.
At month-end, estimate the probability of a gain for each ETF.
If one or more ETFs have an estimated probability of at least 0.50,
hold those ETFs at equal weight.
If no ETF meets the threshold,
hold 100% cash.
Apply the new weights from the next trading day.The 0.50 threshold is not claimed to be optimal for financial markets. It is the first threshold fixed before seeing performance. Repeatedly changing thresholds and keeping only the best historical outcome increases the chance that the selection is fitted to noise. As the number of backtests grows, apparently strong performance is less likely to reproduce out of sample. [S6]
If additional exploration is justified, compare only a limited, predeclared set—such as 0.45, 0.50, and 0.55—inside the validation segment. Once the test period begins, the selected threshold remains unchanged.
7. Evaluate Accuracy and Investment Performance Separately
The evaluation has two layers.
The first layer measures classification behavior.
- Number of test observations
- Accuracy
- Share of positive predictions
- Actual positive-class frequency
- Distribution of predicted probabilities
The second layer measures the economic behavior of the trading rule.
- Cumulative return before and after costs
- Annualized return
- Volatility
- Maximum drawdown
- Sharpe ratio and its calculation assumptions
- Total turnover
- Share of time held in cash
Empirical financial-classification research also reports classification accuracy separately from cumulative return. Results from a particular market and sampling frequency, however, cannot be generalized directly to a monthly ETF strategy. [S5]
No single metric determines the verdict.
| Observation | Question to investigate |
|---|---|
| Higher accuracy but lower post-cost return | Did loss size or turnover increase? |
| Higher return with similar accuracy | Did a small number of large rallies dominate the outcome? |
| Lower drawdown with more time in cash | Did risk fall mainly because market exposure fell? |
| Better at 0 bp but worse from 10 bp | Is the signal's economic edge smaller than its trading cost? |
| Better only in one subperiod | Does the model depend excessively on one market regime? |
The purpose of this table is not to construct a story in which the model wins. It is to trace why its behavior differs from the baseline.
8. Why This Article Does Not Report Performance Figures
This article does not fix a specific ETF universe, provider, adjusted-price methodology, raw-file hash, or observation period as one verified data snapshot. Filling in numerical performance without those inputs would create a claim that readers could not independently reproduce.
The completed deliverable of this part is therefore the following experiment contract, not a return figure.
- Use only features observable by month-end.
- Define the label as the sign of the next holding-period return.
- Train only on labels whose holding periods ended before the current signal date.
- Fit preprocessing separately within each training window.
- Finalize the threshold during validation and never change it during testing.
- Use the same assets, period, execution timing, and costs as the Part 6 baseline.
- Evaluate classification quality and economic performance separately.
When numerical findings are published, they should be accompanied by the ETF list, data provider, price-field definition, download timestamp, SHA-256 hash, date boundaries, code version, and cost assumptions. Trading costs depend on order size, market conditions, execution, and more than the bid-ask spread, so fixed basis-point values are sensitivity assumptions rather than universal estimates. [S8]
The absence of numbers is neither a successful nor a failed experiment. The honest conclusion at this stage is narrower: the comparison rules are now fixed.
9. What the Model Still Cannot Answer
Reducing leakage risk does not resolve every limitation.
- Whether the price data correctly accounts for distributions and splits
- Whether the ETF universe contains survivorship bias
- Whether the model's probabilities are calibrated to observed frequencies
- How closely a close-price assumption resembles actual execution
- Whether fixed costs reflect an individual's commissions and taxes
- Whether the same relationship persists across market regimes
- Whether a favorable configuration was selected from too many experiments
Historical backtests are hypothetical and do not predict or guarantee future returns. [S7] If these limitations are omitted, a more sophisticated model can create false confidence rather than better evidence.
10. How My View Changed After Building the First ML Model
I chose logistic regression not because it is the newest or most powerful model, but because the relationship between its inputs and probability output is relatively easy to inspect. It also provides a useful reference point for evaluating more complex models later.
Before this exercise, I assumed that a model using several features would naturally exploit more information than a simple momentum rule. Now I put the questions in a different order.
Before adding more information, verify that every input was actually observable at the time of the decision.
Higher accuracy also does not establish that a strategy is superior. To say that the model beat the baseline, it must improve the balance of post-cost return, maximum drawdown, and turnover under the same data and execution rules. If the advantage appears in only one test period, the conclusion should be correspondingly limited.
The outcome of this part is not a formula for generating returns. It is a standard for questioning the next model.
Next: Are More Complex Models Actually Better?
Part 8 will compare a decision tree and either a random forest or a gradient-boosting model with logistic regression under the same features, chronological split, and portfolio rules.
The investigation will go beyond asking whether a tree model produced a higher return.
- Did the benefit of learning nonlinear relationships remain in the test period?
- How many additional hyperparameter choices were explored?
- Did the advantage survive trading costs?
- Did one subperiod or a small number of ETFs dominate the difference?
- Did greater complexity reduce explainability and reproducibility?
A complex model earns its place only when it can surpass a simple baseline consistently under the same conditions. That judgment should come from the full experiment record, not from accuracy alone.
Sources
- [S1] LogisticRegression — scikit-learn documentation | scikit-learn developers | https://scikit-learn.org/stable/modules/generated/sklearn.linear_model.LogisticRegression.html ↩
- [S2] TimeSeriesSplit — scikit-learn documentation | scikit-learn developers | https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.TimeSeriesSplit.html ↩
- [S3] Common pitfalls and recommended practices — scikit-learn documentation | scikit-learn developers | https://scikit-learn.org/stable/common_pitfalls.html ↩
- [S4] pandas.DataFrame.shift — pandas documentation | pandas development team | https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.shift.html ↩
- [S5] Evaluating machine learning classification for financial trading: an empirical approach | Eduardo A. Gerlein, Martin McGinnity, Ammar Belatreche, Sarah Coleman | 2016 | https://irep.ntu.ac.uk/id/eprint/28062/ ↩
- [S6] Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance | David H. Bailey, Jonathan M. Borwein, Marcos López de Prado, Qiji Jim Zhu | 2014 | https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2308659 ↩
- [S7] Investor Bulletin: Performance Claims | U.S. Securities and Exchange Commission, Office of Investor Education and Advocacy | 2022 | https://www.investor.gov/introduction-investing/general-resources/news-alerts/alerts-bulletins/investor-bulletins-47 ↩
- [S8] Exchange-Traded Funds | U.S. Securities and Exchange Commission | 2019 | https://www.sec.gov/files/rules/final/2019/33-10695.pdf ↩
Report an error or share feedback
Open a draft with this article’s title and URL. Review the message and recipient before sending.
Open email draftIf no email app opens, copy these details into your usual email service.
Related posts
Quant & Data Research Can Logistic Regression Beat the Baseline? What Makes a Model Comparison Fair
Define timing, data splits, retraining, and cost rules for a fair comparison of logistic regression and baseline strategies, before testing actual performance.
Quant & Data Research Why shift(1) Is Not Enough: A Time Contract for Signals, Execution, and Labels
Separate signal, execution, return, and label timing. Use synthetic unit tests to check information leakage and preprocessing boundaries.