Part 8. Are More Complex Models Better? Comparing Logistic Regression and Tree Models on Equal Terms
Evidence and scope — Model-comparison design
The article designs shared comparison conditions for momentum, logistic regression, and tree models. Without ETF price snapshots, fixed period boundaries, execution data, and run records, it does not rank models by accuracy, returns, or drawdown.
This time, we add decision trees and random forests to the logistic-regression baseline introduced in Part 7. The point is not to use a more impressive-sounding model. A complex model is meaningful only if it is compared with simple momentum and logistic regression using the same data, time order, trading rules, and cost assumptions.
This article does not report actual ETF performance. Without an ETF universe, price snapshot, period boundaries, execution data, and trading records, we cannot conclude that any model has better accuracy, returns, or drawdowns. Instead, we will define a fair experimental contract for increasing model complexity. This is an educational and research-oriented design, not a recommendation to buy or sell any particular ETF.
Putting the Four Comparisons on the Same Starting Line
The 12-month momentum strategy keeps its existing rule. Only the other three models use the same ML features and the same up-versus-not-up label.
| Comparison | How it generates signals | Complexity | Role |
|---|---|---|---|
| 12-month momentum | Existing 12-month momentum rule | Predefined rule | Simple baseline |
| Logistic regression | Estimates a positive-class score from ML features | Regularized linear classification baseline [S9] | First ML baseline |
DecisionTreeClassifier | Sequentially splits features at thresholds [S1] | Nonlinear splits, one tree | First increase in complexity |
RandomForestClassifier | Combines predictions from many trees [S2] | Subsampling and averaging [S2] | Second increase in complexity |
The predict_proba outputs from logistic regression and tree-based models are model scores before they enter the trading rule. They should not be interpreted as calibrated probabilities of investment success or guarantees of future returns. [S1] [S2] [S9]
There is one question the experiment should answer:
Under identical conditions, does a more complex model improve both classification performance and economic outcomes after costs?
You cannot know the answer before running the experiment.
What Nonlinear Splits Add
Logistic regression treats relationships among features as a relatively simple baseline. A decision tree, by contrast, splits nodes using features and thresholds under criteria such as Gini impurity, entropy, or log loss. Repeating those splits lets it represent more complex decision regions. [S1] [S9]
For example, a tree can express combinations of conditions like this:
if return_20d > threshold_a:
if volatility_20d < threshold_b:
predict higher up score
else:
predict lower up score
else:
predict lower up scoreThis means you can encode a hypothesis such as, “Even when recent momentum is positive, treat unusually high volatility differently.” It is not evidence that the structure will predict ETF performance more accurately. The ability to represent nonlinear relationships and superior results in a test period are separate claims.
Overfitting in a Single Tree and Averaging in a Forest
With its default settings, DecisionTreeClassifier can produce fully grown, unpruned trees. As trees get deeper and leaves contain fewer samples, they are more likely to memorize details of the training window. That is why candidate values for max_depth and min_samples_leaf should be limited before looking at results. [S1]
RandomForestClassifier fits many decision trees on different subsamples and averages their predictions. scikit-learn describes this as a way to improve predictive accuracy and control overfitting, while Breiman’s original paper explains that generalization error depends on the strength of individual trees and the correlation among them. [S2] [S3]
Here, too, separate the method from the result.
- Confirmed methodology: A random forest combines and averages predictions from multiple trees. [S2] [S3]
- Unconfirmed result: A random forest produces stronger test performance than momentum, logistic regression, or a single tree in this ETF experiment.
Even if you enable an OOB score, do not treat it as a replacement for chronological training, validation, and test splits. [S2] The evidence used here does not establish that OOB evaluation preserves time order.
The Fair-Comparison Contract
To isolate the effect of complexity, hold the following conditions constant across all four comparisons.
| Item | Condition to fix once for all four comparisons |
|---|---|
| Data and inputs | Use one shared ETF universe and price snapshot. Momentum retains its existing rule; the three ML models receive the same features and labels. |
| Time and training | Define training, validation, and test boundaries by date before viewing results. Assign all ETF rows with the same signal_date to the same split, and train only on rows where label_end_date < signal_date. This is a design choice that applies chronological principles to an ETF panel. [S4] [S5] |
| Execution | Rebalance at the next trading day’s close after a month-end signal, then apply new weights only to returns after that point. This is an educational approximation that omits execution frictions. Selected ETFs are equally weighted, and unselected capital remains in cash. |
| Selection and evaluation | Fit preprocessing only within each training window. Do not reselect choices in the test period after fixing them in validation. Evaluate accuracy separately from net returns, maximum drawdown, and turnover under a default threshold of 0.50 and one-way cost scenarios of 0/5/10/20 bp. [S5] [S6] [S7] |
Treating Selection Bias Like a Cost
Tree-based models create more choices: tree depth and leaf size, for example, or the number of trees and feature-selection settings in a forest. More candidates also increase the chance of selecting a configuration that happened to fit the validation period. Limit and record the candidates and search runs before viewing results. [S6] [S7]
| Comparison | Choices to constrain in advance | What to record |
|---|---|---|
| Logistic regression | Limited candidates from the existing baseline | Number of candidates and runs |
| Decision tree | A small set of max_depth and min_samples_leaf combinations | Number of combinations and validation selection criterion |
| Random forest | A small set of n_estimators, max_depth, min_samples_leaf, and max_features combinations | Number of combinations, seed, and actual run count |
| Trading threshold | Default value of 0.50 | Whether additional tuning occurred and how many runs it used |
The default threshold of 0.50 and one-way costs of 0/5/10/20 bp are prior assumptions, not market truths. If you add candidates for depth, feature count, or thresholds after seeing test performance, record that work separately as a new search. [S6] [S7]
Minimal Code for Adding Trees and Forests to the Same Workflow
The input table contains one row per ETF and signal date. signal_date is the observation date, label_end_date is the date on which the future label becomes known, symbol identifies the ETF, feature columns must be available on the signal date, and label is binary. Monthly training windows are skipped when they cannot satisfy the shared class-1 probability output contract. The output includes each eligible model's up_probability, selected portfolio weights, and predictions for evaluation.
from sklearn.ensemble import RandomForestClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.tree import DecisionTreeClassifier
FEATURES = [
"return_1d",
"return_5d",
"return_20d",
"ma_gap_20_120",
"volatility_20d",
"volume_change",
"market_return_20d",
]
THRESHOLD = 0.50
# Candidate lists must be declared before validation is inspected.
TREE_CANDIDATES = [
{"max_depth": 3, "min_samples_leaf": 20},
{"max_depth": 5, "min_samples_leaf": 40},
]
FOREST_CANDIDATES = [
{
"n_estimators": 300,
"max_depth": 5,
"min_samples_leaf": 20,
"max_features": "sqrt",
},
]
def make_models(selected_tree, selected_forest):
return {
"logistic": Pipeline(
[
("scale", StandardScaler()),
("classifier", LogisticRegression(max_iter=2_000, random_state=7)),
]
),
"tree": DecisionTreeClassifier(random_state=7, **selected_tree),
"forest": RandomForestClassifier(
random_state=7,
n_jobs=-1,
**selected_forest,
),
}
def positive_probability(model, X):
positive_index = list(model.classes_).index(1)
return model.predict_proba(X)[:, positive_index]
def select_equal_weights(frame, threshold=THRESHOLD):
selected = frame["up_probability"].ge(threshold)
frame["weight"] = 0.0
if selected.any():
frame.loc[selected, "weight"] = 1.0 / selected.sum()
return frame
def walk_forward_predictions(
samples,
test_signal_dates,
selected_tree,
selected_forest,
):
predictions = []
for signal_date in test_signal_dates:
# Only labels known before this signal date may enter training.
train = samples[
(samples["signal_date"] < signal_date)
& (samples["label_end_date"] < signal_date)
].copy()
current = samples.loc[
samples["signal_date"].eq(signal_date)
].copy()
# Skip windows that cannot satisfy the shared class-1 probability output contract.
if train["label"].nunique() != 2:
continue
# Each pipeline fits preprocessing only on the historical window.
for model_name, model in make_models(
selected_tree, selected_forest
).items():
model.fit(train[FEATURES], train["label"])
result = current[["signal_date", "symbol", "label"]].copy()
result["model"] = model_name
result["up_probability"] = positive_probability(
model, current[FEATURES]
)
predictions.append(select_equal_weights(result))
return predictionsDuring validation, select exactly one predeclared candidate for each model. Apply only those fixed selections during the test period, and do not reselect tree depth, forest settings, or thresholds. [S6] [S7]
This is a minimal contract example, not a complete backtest. It loops through date-defined test periods, trains only on labels that ended before each signal date, and fits logistic-regression scaling inside the pipeline for every historical window. The three ML models receive the same feature set and label. [S1] [S2] [S5] [S9]
Read Results in Separate Classification and Economic Layers
The classification layer describes the distribution of predictions. The economic layer examines after-cost outcomes under the same trading rule.
| Observation | Question to investigate |
|---|---|
| Accuracy is high, but net returns are low | Did turnover or losses during weak periods increase? |
| The model leads only at 0 bp | Is the result highly sensitive to trading-cost assumptions? |
| Maximum drawdown is lower and cash weight is higher | Did risk decline simply because market exposure declined? |
| Validation is strong but testing is weak | Could selection bias or regime dependence be involved? |
| A single tree and forest differ materially | Does the difference reflect averaging, or chance fit to a particular sample that needs further scrutiny? |
Combining accuracy and after-cost performance into one number makes causes harder to identify. Because there are no actual data or execution results here, no claim can be made about model superiority or persistence after costs. [S6] [S7]
What to Fix Before Increasing Complexity
A tree’s nonlinear splits and a random forest’s averaging can produce predictions that differ from logistic regression, but they do not imply superiority. [S1] [S2] [S3]
Before changing complexity, fix the data, timing, costs, and search budget. Do not use test results as the basis for the next model-selection decision.
Part 9 examines failure modes in financial ML—such as leakage, selection bias, and regime dependence—that can make luck look like skill.
Sources
-
[S1] DecisionTreeClassifier — scikit-learn 1.9.0 documentation | scikit-learn developers | 2007–2026 | https://scikit-learn.org/stable/modules/generated/sklearn.tree.DecisionTreeClassifier.html
↩ -
[S2] RandomForestClassifier — scikit-learn 1.9.0 documentation | scikit-learn developers | 2007–2026 | https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.RandomForestClassifier.html
↩ -
[S3] Random Forests | Leo Breiman | October 2001 | https://doi.org/10.1023/A:1010933404324
↩ -
[S4] TimeSeriesSplit — scikit-learn 1.9.0 documentation | scikit-learn developers | 2007–2026 | https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.TimeSeriesSplit.html
↩ -
[S5] 12. Common pitfalls and recommended practices — scikit-learn 1.9.0 documentation | scikit-learn developers | 2007–2026 | https://scikit-learn.org/stable/common_pitfalls.html
↩ -
[S6] On Over-fitting in Model Selection and Subsequent Selection Bias in Performance Evaluation | Gavin C. Cawley and Nicola L. C. Talbot | July 2010 | https://www.jmlr.org/beta/papers/v11/cawley10a.html
↩ -
[S7] Nested versus non-nested cross-validation — scikit-learn 1.9.0 documentation | scikit-learn developers | 2007–2026 | https://scikit-learn.org/stable/auto_examples/model_selection/plot_nested_cross_validation_iris.html
↩ -
[S9] LogisticRegression — scikit-learn 1.9.0 documentation | scikit-learn developers | 2007–2026 | https://scikit-learn.org/stable/modules/generated/sklearn.linear_model.LogisticRegression.html
↩
Report an error or share feedback
Open a draft with this article’s title and URL. Review the message and recipient before sending.
Open email draftIf no email app opens, copy these details into your usual email service.
Related posts
Quant & Data Research Can Logistic Regression Beat the Baseline? What Makes a Model Comparison Fair
Define timing, data splits, retraining, and cost rules for a fair comparison of logistic regression and baseline strategies, before testing actual performance.
Quant & Data Research Why shift(1) Is Not Enough: A Time Contract for Signals, Execution, and Labels
Separate signal, execution, return, and label timing. Use synthetic unit tests to check information leakage and preprocessing boundaries.