Forge Fate
Quant & Data Research

Part 7. First Machine Learning Model: Testing the Probability of Gains in the Next Rebalancing Period with Logistic Regression

10 min read

Evidence and scope — Machine-learning baseline design

This article defines features, labels, time splits, retraining, and trading rules. It is not a performance report based on a fixed ETF data snapshot and presents no actual return or accuracy figures.

Accuracy Is Not Return: Designing a First ML Baseline with Logistic Regression

In Part 6, we built a simple baseline: at each month-end, hold only the ETFs whose 12-month momentum is positive, and otherwise stay in cash. This time, we keep the comparison conditions unchanged and replace only the signal-generation method. Instead of using the sign of past returns, we use the probability that logistic regression assigns to a gain over the next rebalancing period.

At first, I expected a machine-learning model to make better decisions than simple momentum simply because it could consider several features at once. Building the first model changed the question that mattered most.

The first question is not whether the model is more accurate. It is whether the model was compared with the baseline under the same conditions without using future information.

This article is not a performance report based on a particular ETF dataset. Reporting accuracy or return figures before fixing a verified snapshot—including the data provider, price-field definition, ETF universe, and observation period—would produce a result that readers could not reproduce. Instead, this article fixes the features, label, chronological split, retraining process, trading rule, and decision criteria needed for a real experiment.

The code and explanations are for educational and research purposes. They do not recommend buying or selling any ETF or security, and they do not guarantee returns or future performance. A backtest is hypothetical performance calculated under fixed data and assumptions. [S7]

What Changes—and What Does Not

For a fair comparison, the conditions established in Part 6 remain fixed.

ItemPart 6: simple momentumPart 7: logistic regression
ETF universe and price dataFixed in advanceUse the same snapshot
Signal timeMonth-endSame
When new weights take effectFrom the next trading daySame
RebalancingMonthlySame
Position constraintLong or cashSame
Allocation among selected assetsEqual weightSame
Cost sensitivity0, 5, 10, and 20 bp one-waySame
SignalSign of the trailing 12-month returnEstimated probability and a fixed threshold

Only the last row changes. If we change the ETF universe, period, costs, or execution timing at the same time as the model, we cannot tell what caused the performance difference.

1. Define the Prediction Problem in One Sentence

The model in this article answers one question.

Using only price, volatility, and volume features observable by month-end, can we classify whether the return over the next monthly holding period will be positive?

The positive class means that the next holding-period return is greater than zero.

text
label_t = 1  if next_holding_return_t > 0
label_t = 0  otherwise

scikit-learn's LogisticRegression is a classifier, and predict_proba returns an estimated probability for each class. [S1] A model output of 0.62 does not establish a 62% real-world chance that the asset will rise. It is better understood as a score calculated under the training data and model assumptions.

This distinction matters throughout the experiment.

TermMeaning here
Probability of a gainThe model's estimated probability for the positive class
Classification accuracyThe share of predicted classes that match the labels
Strategy returnHypothetical performance after converting predictions into trading rules
InvestabilityA separate judgment that also considers costs, liquidity, execution, and risk

A classifier can be slightly more accurate and still produce worse post-cost performance if its losing trades are larger or it trades too frequently. Accuracy and investment performance therefore belong in separate evaluations.

2. Separate Feature Time from Label Time

Every feature candidate is restricted to information already observable at month-end.

FeatureCalculation windowObservable at month-end?
Most recent 1-day returnOne trading day through month-endYes
Most recent 5-day returnFive trading days through month-endYes
Most recent 20-day returnTwenty trading days through month-endYes
Moving-average spread20-day and 120-day moving averagesYes
Recent volatilityStandard deviation of daily returns over 20 daysYes
Volume changeDifference between recent volume and its prior averageYes
Market return20-day return of a market proxyYes
Next holding-period returnThrough the next rebalanceNo; label only

Values appearing in the same data row were not necessarily knowable at the same time. Attaching next month's return to the current row creates a training label; it does not make that return a valid input at prediction time.

python
monthly["forward_return"] = (
    monthly.groupby("symbol")["execution_close"]
    .transform(lambda s: s.shift(-1).div(s).sub(1.0))
)
monthly["label"] = monthly["forward_return"].gt(0).astype("int8")

The negative shift aligns the next row's outcome with the current signal row. [S4] The final observation for each asset has no completed next holding period, so it must be excluded from training and evaluation.

There is another timing condition. When the model is retrained at month-end t, even an earlier signal row is ineligible if its holding period has not ended and its label was not yet known.

python
eligible_train = samples[
    (samples["signal_date"] < current_signal_date)
    & (samples["label_end_date"] < current_signal_date)
]

Without this condition, code can select rows that look historical while still learning from outcomes that were unavailable on the prediction date.

3. Preserve Calendar Order Instead of Randomly Splitting

Randomly shuffling time-series samples into training and test sets can allow future information to influence an evaluation of the past. TimeSeriesSplit supports chronological folds and a gap between the training and validation segments. [S2]

This experiment follows a simple calendar structure.

text
Past                                               Future
|──────── Training ────────|── Validation ──|──── Test ────|
          Fit                   Finalize choices    Evaluate once
  • Fit the model and preprocessing rules in the training segment.
  • Compare only a small, predeclared set of thresholds and model settings in validation.
  • Evaluate the test segment once after all choices are fixed.
  • If features or thresholds change after inspecting test performance, that segment is no longer a test set.

The date boundaries should be recorded after examining the real dataset's available history and market regimes. The important point is not a particular calendar year, but fixing the boundaries before evaluating performance and never reversing the direction of time.

4. Learn Preprocessing Only from the Training Window

Scaling can leak information even though it does not look predictive on its own. If the scaler is fitted on the full dataset, the mean and variance of the test period influence training. scikit-learn recommends fitting preprocessing only on training data and applying the learned transformation elsewhere. [S3]

The simplest implementation is to combine preprocessing and the model in a Pipeline.

python
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler

model = Pipeline(
    steps=[
        ("scale", StandardScaler()),
        (
            "classifier",
            LogisticRegression(
                C=1.0,
                max_iter=2_000,
                random_state=7,
            ),
        ),
    ]
)

model.fit(train[feature_columns], train["label"])
probability = model.predict_proba(current[feature_columns])

classes = model.named_steps["classifier"].classes_
positive_class_index = list(classes).index(1)
current["up_probability"] = probability[:, positive_class_index]

The code locates class 1 in classes_ instead of assuming that the positive-class probability is always the second column. Before fitting, it should check not only for missing values but also for positive and negative infinity. A volume-change calculation can produce infinity when its prior average is zero; stopping with a data error is safer than silently replacing it.

5. Walk Forward Month by Month

If the strategy were used at a real month-end, it could access only the information finalized by that date. The test period should therefore be processed as a monthly walk-forward exercise rather than predicted all at once.

python
predictions = []

for signal_date in test_signal_dates:
    train = samples[
        (samples["signal_date"] < signal_date)
        & (samples["label_end_date"] < signal_date)
    ].copy()
    current = samples[samples["signal_date"] == signal_date].copy()

    model.fit(train[feature_columns], train["label"])
    probability = model.predict_proba(current[feature_columns])
    classes = model.named_steps["classifier"].classes_
    positive_index = list(classes).index(1)

    current["up_probability"] = probability[:, positive_index]
    predictions.append(current)

An expanding window, in which past observations accumulate over time, is the default design here. A fixed-length window is also possible, but its length should be chosen before performance is inspected. Monthly retraining is not intended to maximize returns; it is intended to reproduce the information set that would actually have been available each month.

6. Convert Probabilities into a Trading Rule

Model scores alone do not define a backtest. The experiment must specify the probability threshold, the allocation when several ETFs qualify, and the position when none qualify.

The first rule is deliberately simple.

text
At month-end, estimate the probability of a gain for each ETF.

If one or more ETFs have an estimated probability of at least 0.50,
hold those ETFs at equal weight.

If no ETF meets the threshold,
hold 100% cash.

Apply the new weights from the next trading day.

The 0.50 threshold is not claimed to be optimal for financial markets. It is the first threshold fixed before seeing performance. Repeatedly changing thresholds and keeping only the best historical outcome increases the chance that the selection is fitted to noise. As the number of backtests grows, apparently strong performance is less likely to reproduce out of sample. [S6]

If additional exploration is justified, compare only a limited, predeclared set—such as 0.45, 0.50, and 0.55—inside the validation segment. Once the test period begins, the selected threshold remains unchanged.

7. Evaluate Accuracy and Investment Performance Separately

The evaluation has two layers.

The first layer measures classification behavior.

  • Number of test observations
  • Accuracy
  • Share of positive predictions
  • Actual positive-class frequency
  • Distribution of predicted probabilities

The second layer measures the economic behavior of the trading rule.

  • Cumulative return before and after costs
  • Annualized return
  • Volatility
  • Maximum drawdown
  • Sharpe ratio and its calculation assumptions
  • Total turnover
  • Share of time held in cash

Empirical financial-classification research also reports classification accuracy separately from cumulative return. Results from a particular market and sampling frequency, however, cannot be generalized directly to a monthly ETF strategy. [S5]

No single metric determines the verdict.

ObservationQuestion to investigate
Higher accuracy but lower post-cost returnDid loss size or turnover increase?
Higher return with similar accuracyDid a small number of large rallies dominate the outcome?
Lower drawdown with more time in cashDid risk fall mainly because market exposure fell?
Better at 0 bp but worse from 10 bpIs the signal's economic edge smaller than its trading cost?
Better only in one subperiodDoes the model depend excessively on one market regime?

The purpose of this table is not to construct a story in which the model wins. It is to trace why its behavior differs from the baseline.

8. Why This Article Does Not Report Performance Figures

This article does not fix a specific ETF universe, provider, adjusted-price methodology, raw-file hash, or observation period as one verified data snapshot. Filling in numerical performance without those inputs would create a claim that readers could not independently reproduce.

The completed deliverable of this part is therefore the following experiment contract, not a return figure.

  1. Use only features observable by month-end.
  2. Define the label as the sign of the next holding-period return.
  3. Train only on labels whose holding periods ended before the current signal date.
  4. Fit preprocessing separately within each training window.
  5. Finalize the threshold during validation and never change it during testing.
  6. Use the same assets, period, execution timing, and costs as the Part 6 baseline.
  7. Evaluate classification quality and economic performance separately.

When numerical findings are published, they should be accompanied by the ETF list, data provider, price-field definition, download timestamp, SHA-256 hash, date boundaries, code version, and cost assumptions. Trading costs depend on order size, market conditions, execution, and more than the bid-ask spread, so fixed basis-point values are sensitivity assumptions rather than universal estimates. [S8]

The absence of numbers is neither a successful nor a failed experiment. The honest conclusion at this stage is narrower: the comparison rules are now fixed.

9. What the Model Still Cannot Answer

Reducing leakage risk does not resolve every limitation.

  • Whether the price data correctly accounts for distributions and splits
  • Whether the ETF universe contains survivorship bias
  • Whether the model's probabilities are calibrated to observed frequencies
  • How closely a close-price assumption resembles actual execution
  • Whether fixed costs reflect an individual's commissions and taxes
  • Whether the same relationship persists across market regimes
  • Whether a favorable configuration was selected from too many experiments

Historical backtests are hypothetical and do not predict or guarantee future returns. [S7] If these limitations are omitted, a more sophisticated model can create false confidence rather than better evidence.

10. How My View Changed After Building the First ML Model

I chose logistic regression not because it is the newest or most powerful model, but because the relationship between its inputs and probability output is relatively easy to inspect. It also provides a useful reference point for evaluating more complex models later.

Before this exercise, I assumed that a model using several features would naturally exploit more information than a simple momentum rule. Now I put the questions in a different order.

Before adding more information, verify that every input was actually observable at the time of the decision.

Higher accuracy also does not establish that a strategy is superior. To say that the model beat the baseline, it must improve the balance of post-cost return, maximum drawdown, and turnover under the same data and execution rules. If the advantage appears in only one test period, the conclusion should be correspondingly limited.

The outcome of this part is not a formula for generating returns. It is a standard for questioning the next model.

Next: Are More Complex Models Actually Better?

Part 8 will compare a decision tree and either a random forest or a gradient-boosting model with logistic regression under the same features, chronological split, and portfolio rules.

The investigation will go beyond asking whether a tree model produced a higher return.

  • Did the benefit of learning nonlinear relationships remain in the test period?
  • How many additional hyperparameter choices were explored?
  • Did the advantage survive trading costs?
  • Did one subperiod or a small number of ETFs dominate the difference?
  • Did greater complexity reduce explainability and reproducibility?

A complex model earns its place only when it can surpass a simple baseline consistently under the same conditions. That judgment should come from the full experiment record, not from accuracy alone.

Sources

Report an error or share feedback

Open a draft with this article’s title and URL. Review the message and recipient before sending.

To: [email protected]

Open email draft

If no email app opens, copy these details into your usual email service.

Contact information