Forge Fate
Quant & Data Research

Part 8. Are More Complex Models Better? Comparing Logistic Regression and Tree Models on Equal Terms

6 min read

Evidence and scope — Model-comparison design

The article designs shared comparison conditions for momentum, logistic regression, and tree models. Without ETF price snapshots, fixed period boundaries, execution data, and run records, it does not rank models by accuracy, returns, or drawdown.

This time, we add decision trees and random forests to the logistic-regression baseline introduced in Part 7. The point is not to use a more impressive-sounding model. A complex model is meaningful only if it is compared with simple momentum and logistic regression using the same data, time order, trading rules, and cost assumptions.

This article does not report actual ETF performance. Without an ETF universe, price snapshot, period boundaries, execution data, and trading records, we cannot conclude that any model has better accuracy, returns, or drawdowns. Instead, we will define a fair experimental contract for increasing model complexity. This is an educational and research-oriented design, not a recommendation to buy or sell any particular ETF.

Putting the Four Comparisons on the Same Starting Line

The 12-month momentum strategy keeps its existing rule. Only the other three models use the same ML features and the same up-versus-not-up label.

ComparisonHow it generates signalsComplexityRole
12-month momentumExisting 12-month momentum rulePredefined ruleSimple baseline
Logistic regressionEstimates a positive-class score from ML featuresRegularized linear classification baseline [S9]First ML baseline
DecisionTreeClassifierSequentially splits features at thresholds [S1]Nonlinear splits, one treeFirst increase in complexity
RandomForestClassifierCombines predictions from many trees [S2]Subsampling and averaging [S2]Second increase in complexity

The predict_proba outputs from logistic regression and tree-based models are model scores before they enter the trading rule. They should not be interpreted as calibrated probabilities of investment success or guarantees of future returns. [S1] [S2] [S9]

There is one question the experiment should answer:

Under identical conditions, does a more complex model improve both classification performance and economic outcomes after costs?

You cannot know the answer before running the experiment.

What Nonlinear Splits Add

Logistic regression treats relationships among features as a relatively simple baseline. A decision tree, by contrast, splits nodes using features and thresholds under criteria such as Gini impurity, entropy, or log loss. Repeating those splits lets it represent more complex decision regions. [S1] [S9]

For example, a tree can express combinations of conditions like this:

text
if return_20d > threshold_a:
    if volatility_20d < threshold_b:
        predict higher up score
    else:
        predict lower up score
else:
    predict lower up score

This means you can encode a hypothesis such as, “Even when recent momentum is positive, treat unusually high volatility differently.” It is not evidence that the structure will predict ETF performance more accurately. The ability to represent nonlinear relationships and superior results in a test period are separate claims.

Overfitting in a Single Tree and Averaging in a Forest

With its default settings, DecisionTreeClassifier can produce fully grown, unpruned trees. As trees get deeper and leaves contain fewer samples, they are more likely to memorize details of the training window. That is why candidate values for max_depth and min_samples_leaf should be limited before looking at results. [S1]

RandomForestClassifier fits many decision trees on different subsamples and averages their predictions. scikit-learn describes this as a way to improve predictive accuracy and control overfitting, while Breiman’s original paper explains that generalization error depends on the strength of individual trees and the correlation among them. [S2] [S3]

Here, too, separate the method from the result.

  • Confirmed methodology: A random forest combines and averages predictions from multiple trees. [S2] [S3]
  • Unconfirmed result: A random forest produces stronger test performance than momentum, logistic regression, or a single tree in this ETF experiment.

Even if you enable an OOB score, do not treat it as a replacement for chronological training, validation, and test splits. [S2] The evidence used here does not establish that OOB evaluation preserves time order.

The Fair-Comparison Contract

To isolate the effect of complexity, hold the following conditions constant across all four comparisons.

ItemCondition to fix once for all four comparisons
Data and inputsUse one shared ETF universe and price snapshot. Momentum retains its existing rule; the three ML models receive the same features and labels.
Time and trainingDefine training, validation, and test boundaries by date before viewing results. Assign all ETF rows with the same signal_date to the same split, and train only on rows where label_end_date < signal_date. This is a design choice that applies chronological principles to an ETF panel. [S4] [S5]
ExecutionRebalance at the next trading day’s close after a month-end signal, then apply new weights only to returns after that point. This is an educational approximation that omits execution frictions. Selected ETFs are equally weighted, and unselected capital remains in cash.
Selection and evaluationFit preprocessing only within each training window. Do not reselect choices in the test period after fixing them in validation. Evaluate accuracy separately from net returns, maximum drawdown, and turnover under a default threshold of 0.50 and one-way cost scenarios of 0/5/10/20 bp. [S5] [S6] [S7]

Treating Selection Bias Like a Cost

Tree-based models create more choices: tree depth and leaf size, for example, or the number of trees and feature-selection settings in a forest. More candidates also increase the chance of selecting a configuration that happened to fit the validation period. Limit and record the candidates and search runs before viewing results. [S6] [S7]

ComparisonChoices to constrain in advanceWhat to record
Logistic regressionLimited candidates from the existing baselineNumber of candidates and runs
Decision treeA small set of max_depth and min_samples_leaf combinationsNumber of combinations and validation selection criterion
Random forestA small set of n_estimators, max_depth, min_samples_leaf, and max_features combinationsNumber of combinations, seed, and actual run count
Trading thresholdDefault value of 0.50Whether additional tuning occurred and how many runs it used

The default threshold of 0.50 and one-way costs of 0/5/10/20 bp are prior assumptions, not market truths. If you add candidates for depth, feature count, or thresholds after seeing test performance, record that work separately as a new search. [S6] [S7]

Minimal Code for Adding Trees and Forests to the Same Workflow

The input table contains one row per ETF and signal date. signal_date is the observation date, label_end_date is the date on which the future label becomes known, symbol identifies the ETF, feature columns must be available on the signal date, and label is binary. Monthly training windows are skipped when they cannot satisfy the shared class-1 probability output contract. The output includes each eligible model's up_probability, selected portfolio weights, and predictions for evaluation.

python
from sklearn.ensemble import RandomForestClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.tree import DecisionTreeClassifier

FEATURES = [
    "return_1d",
    "return_5d",
    "return_20d",
    "ma_gap_20_120",
    "volatility_20d",
    "volume_change",
    "market_return_20d",
]
THRESHOLD = 0.50

# Candidate lists must be declared before validation is inspected.
TREE_CANDIDATES = [
    {"max_depth": 3, "min_samples_leaf": 20},
    {"max_depth": 5, "min_samples_leaf": 40},
]
FOREST_CANDIDATES = [
    {
        "n_estimators": 300,
        "max_depth": 5,
        "min_samples_leaf": 20,
        "max_features": "sqrt",
    },
]

def make_models(selected_tree, selected_forest):
    return {
        "logistic": Pipeline(
            [
                ("scale", StandardScaler()),
                ("classifier", LogisticRegression(max_iter=2_000, random_state=7)),
            ]
        ),
        "tree": DecisionTreeClassifier(random_state=7, **selected_tree),
        "forest": RandomForestClassifier(
            random_state=7,
            n_jobs=-1,
            **selected_forest,
        ),
    }

def positive_probability(model, X):
    positive_index = list(model.classes_).index(1)
    return model.predict_proba(X)[:, positive_index]

def select_equal_weights(frame, threshold=THRESHOLD):
    selected = frame["up_probability"].ge(threshold)
    frame["weight"] = 0.0
    if selected.any():
        frame.loc[selected, "weight"] = 1.0 / selected.sum()
    return frame

def walk_forward_predictions(
    samples,
    test_signal_dates,
    selected_tree,
    selected_forest,
):
    predictions = []

    for signal_date in test_signal_dates:
        # Only labels known before this signal date may enter training.
        train = samples[
            (samples["signal_date"] < signal_date)
            & (samples["label_end_date"] < signal_date)
        ].copy()
        current = samples.loc[
            samples["signal_date"].eq(signal_date)
        ].copy()

        # Skip windows that cannot satisfy the shared class-1 probability output contract.
        if train["label"].nunique() != 2:
            continue

        # Each pipeline fits preprocessing only on the historical window.
        for model_name, model in make_models(
            selected_tree, selected_forest
        ).items():
            model.fit(train[FEATURES], train["label"])

            result = current[["signal_date", "symbol", "label"]].copy()
            result["model"] = model_name
            result["up_probability"] = positive_probability(
                model, current[FEATURES]
            )
            predictions.append(select_equal_weights(result))

    return predictions

During validation, select exactly one predeclared candidate for each model. Apply only those fixed selections during the test period, and do not reselect tree depth, forest settings, or thresholds. [S6] [S7]

This is a minimal contract example, not a complete backtest. It loops through date-defined test periods, trains only on labels that ended before each signal date, and fits logistic-regression scaling inside the pipeline for every historical window. The three ML models receive the same feature set and label. [S1] [S2] [S5] [S9]

Read Results in Separate Classification and Economic Layers

The classification layer describes the distribution of predictions. The economic layer examines after-cost outcomes under the same trading rule.

ObservationQuestion to investigate
Accuracy is high, but net returns are lowDid turnover or losses during weak periods increase?
The model leads only at 0 bpIs the result highly sensitive to trading-cost assumptions?
Maximum drawdown is lower and cash weight is higherDid risk decline simply because market exposure declined?
Validation is strong but testing is weakCould selection bias or regime dependence be involved?
A single tree and forest differ materiallyDoes the difference reflect averaging, or chance fit to a particular sample that needs further scrutiny?

Combining accuracy and after-cost performance into one number makes causes harder to identify. Because there are no actual data or execution results here, no claim can be made about model superiority or persistence after costs. [S6] [S7]

What to Fix Before Increasing Complexity

A tree’s nonlinear splits and a random forest’s averaging can produce predictions that differ from logistic regression, but they do not imply superiority. [S1] [S2] [S3]

Before changing complexity, fix the data, timing, costs, and search budget. Do not use test results as the basis for the next model-selection decision.

Part 9 examines failure modes in financial ML—such as leakage, selection bias, and regime dependence—that can make luck look like skill.

Sources

Report an error or share feedback

Open a draft with this article’s title and URL. Review the message and recipient before sending.

To: [email protected]

Open email draft

If no email app opens, copy these details into your usual email service.

Contact information