Forge Fate
Quant & Data Research

The First Baseline: Building a Simple Momentum Strategy

14 min read

Evidence and scope — Educational baseline code

The code specifies signal, execution-lag, weight, and cost rules for a simple momentum baseline. No validated actual-performance figures are presented. It is not a production system reproducing real orders, quotes, or partial fills.

To evaluate an AI model's performance, you must first fix a simple rule that was not changed after seeing the results, along with identical comparison conditions. The momentum strategy built in this article is not an answer that guarantees returns; it is the starting line that later experiments will share.

In Part 5, rather than judging a strategy by returns alone, we defined an evaluation table that also looks at maximum drawdown, turnover, and cost sensitivity. This article builds the first shared baseline that framework will actually be applied to: a simple momentum strategy that later AI and ML models will be measured against.

At first I assumed that an AI model using more information would naturally become a better strategy. But complex results are hard to interpret without a simple point of comparison. Even if AI shows a higher cumulative return, you cannot tell what that number means unless you can verify it beats a simple rule on the same ETF universe, the same period, the same trading costs, and the same execution timing.

So the goal of this article is not to find "the strategy that looks best." It is to first lock in a baseline that anyone can rerun with the same data snapshot and settings. The intended reader is an individual investor with basic Python at the pandas level who is building a reproducible backtest for the first time.

"Momentum" covers several different strategies. Cross-sectional momentum, which buys relatively strong assets and sells relatively weak ones, must be distinguished from time-series momentum, which uses each asset's own past return as a signal. [S1] [S2] This article uses a simplified version of the latter, adapted to a rule that holds ETFs or waits in cash without any short positions.

All examples and code in this article are for educational and research purposes and are not a recommendation to buy or sell any particular ETF, nor a guarantee of future returns.

1. Baseline Rule: "Hold If the 12-Month Signal Is Positive; Otherwise Hold Cash"

This strategy is neither an academically validated optimal solution nor a claim that it will outperform in the future. It is an intentionally simplified project rule, designed to keep later experiments comparable under stable conditions.

ItemFixed rule in this article
SignalTrailing 12-month return at month-end
Buy conditionETFs with a positive trailing 12-month return
AllocationEqual weight across ETFs with positive signals
Cash rule100% cash if no ETF has a positive signal
Initial lookback window0% ETFs, 100% cash until the first valid 12-month signal; still included in the evaluation
RebalancingOnce per month
Execution timingApplied from the next trading day after the month-end signal is confirmed
Position constraintLong or cash only; no shorting or leverage
CostsOne-way sensitivity scenarios of 0, 5, 10, and 20 bp

The asset classes and implementation used in time-series momentum research differ from the ETF long/cash rule here. [S2] So a 12-month lookback, monthly rebalancing, positive signals, and equal weighting are not the only standard, nor the optimal settings. The rule can be summarized as follows.

text
At month-end, compute each ETF's trailing 12-month return.  If at least one ETF is positive  -> hold those ETFs at equal weight  If no ETF is positive            -> 100% cashApply the month-end signal starting from the next trading day.

For example, if the candidate universe has four ETFs and the month-end signals are as below, you hold the positive A and C at half each, while B, D, and cash are 0%. If all four signals are negative, cash is 100%.

ETFTrailing 12-month returnTarget weight next period
ETF_APositive50%
ETF_BNegative0%
ETF_CPositive50%
ETF_DNegative0%

What matters here is less the "hold only positive ETFs" rule itself and more that the rule is fixed before you see the results. If you keep changing 12 months to 11, monthly to weekly, or equal weight to arbitrary weights because a backtest looked bad, the baseline stops functioning as a point of comparison.

2. Data and Liquidity: Verify Through Records, Not Claims

A single sentence like "we used sufficiently liquid ETFs" is not enough. ETF trades can incur brokerage commissions and additional transaction costs, and the bid-ask spread is itself a cost that lowers potential returns. The SEC notes that ETFs with higher liquidity and trading volume tend to have narrower spreads. [S3] Reproducibility, likewise, is not just publishing code. You also have to fix which file you used as input, what that file is and when it was downloaded, and what the price column reflects. For instance, one ETF total-return disclosure assumes dividends are reinvested but does not include brokerage commissions. [S5] In other words, how the price data handles dividends and splits must be recorded separately from how the strategy deducts trading costs.

For each real backtest, fill in the table below.

Record itemContent
ETF listTicker, selection date, inception date, exclusion reason — fixed before checking performance
Observation periodStart and end dates, plus each ETF's first usable date
Provider and sourceData provider name, source URL or file path, download time (UTC)
Price column definitionWhether it is adjusted close or a total-return index
Dividend and split treatmentProvider definition or a verifiable document
Missing-data ruleWhether rows are dropped, held, or reindexed, and why
Raw archiveCSV filename, SHA-256 hash, code version
Liquidity evidenceTrading value by sub-period, median bid-ask spread, any trading halts
Survivorship checkWhether the ETF was investable at the time, and whether only post-inception data was used

This article does not rule that any particular ETF is always sufficiently liquid. Without directly checking trading value by sub-period, spreads, inception dates, and trading halts, that judgment cannot be generalized. Whether this article's candidate set and dataset have survivorship bias or point-in-time investability issues has also not yet been verified.

3. Timing: Information Learned at Month-End Is Used From the Next Trading Day

The thing to be especially careful about is not mixing the moment you learned information with the moment you traded on it. If you computed a trailing 12-month return from the month-end close, you must not assume you already knew that signal and traded on it before that close was final. This article uses daily adjusted-price returns from the previous close to the current close, so it fixes the order of events as follows.

text
Month-end price is final  -> compute the trailing 12-month return  -> identify ETFs with a positive signal, derive target weights  -> keep current weights through the close of the EXECUTION_LAG_TRADING_DAYS-th trading day after the signal  -> after that close, deduct costs and rebalance to the target weights  -> from the following trading day, apply daily returns at the new weights

With EXECUTION_LAG_TRADING_DAYS = 1, you rebalance after the close of the first trading day following the signal. That execution day's previous-close-to-current-close return still applies to the old weights, and the new target weights apply to returns from the following trading day onward. This convention is a close-only educational approximation, not a reproduction of actual fill prices, order types, or partial fills.

4. Put the Benchmark Under the Same Conditions: Equal-Weight Buy and Hold

To see what the momentum strategy changed, you need a benchmark that shares the same asset set, period, data, and cost definition. This article's benchmark is equal-weight buy and hold. It uses the same ETF universe, the same observation start and end dates, and the same adjusted-price data; it allocates equal weight at the start and never rebalances. Both the strategy and the benchmark compute pre- and post-cost results, but the benchmark reflects only the initial entry cost.

This comparison is not meant to predetermine whether momentum or buy and hold is better. It is a reference for reading, under the same conditions, what changed when the signal rule changed — later including AI models. When looking at the results table, ask:

  • Did the strategy change the result, or did the ETF composition and period choice change it?
  • How does the pre- vs. post-cost difference relate to turnover?
  • What was the cost and the benefit of spending more months in cash?
  • Did a lower maximum drawdown come with a lower return or a larger cash allocation?

5. Trading Costs: Lock the Part 5 Evaluation Table Into This Baseline

The cost sensitivity and turnover evaluation table defined in Part 5 is not re-explained here. This article fixes that definition into the code's execution order so it applies unchanged to later ML experiments.

text
turnover_t = Sum_i |ETF weight_i just before execution - target weight_i|cost_t     = one-way cost rate x turnover_t x net asset value just before execution

Entry from an all-cash state is also counted in turnover, and 0, 5, 10, 20 bp are not common real-world values but research sensitivity scenarios. On the execution day, the cost-deducted net assets are reallocated to the target weights, so the cost keeps compounding into later weights. This weight-change measure is a project-internal definition, different from the portfolio turnover used in disclosures. [S8] Because ETF trades can incur commissions and transaction costs, the assumptions and their scope must be disclosed. [S3] [S4] The cash return is fixed at 0% in the code; this is a simplifying assumption, not a proxy for market rates.

6. Reproducible Backtest Code

The input is a single wide-format daily price CSV. It has a date column and one price column per ETF, and every ETF column must be the same kind of price whose dividend and split treatment you have verified — all adjusted close, or all total-return index.

text
date,ETF_A,ETF_B,ETF_C,ETF_D2020-01-02,100.12,50.43,...2020-01-03,99.80,50.71,...

6.1 Settings to Fix Before Running

These are values fixed before seeing any performance. The source, download time, price column definition, and missing-data rule are recorded alongside as separate metadata.

momentum_baseline.pypython
from pathlib import Pathimport numpy as npimport pandas as pdCODE_VERSION = "baseline-momentum-v1.0"RAW_PRICE_FILE = Path("data/prices_adjusted_close.csv")# ETF list fixed before checking performanceUNIVERSE = ("ETF_A", "ETF_B", "ETF_C", "ETF_D")LOOKBACK_MONTHS = 12REBALANCE_FREQUENCY = "ME"            # month-endEXECUTION_LAG_TRADING_DAYS = 1        # execute after the next trading day's closeCOST_SCENARIOS_BP = [0, 5, 10, 20]    # one-way, research sensitivityALLOW_SHORT = FalseCASH_RETURN_ASSUMPTION = 0.0          # daily return on the cash balance (simplifying assumption)RISK_FREE_RETURN_DAILY = 0.0          # risk-free rate for the Sharpe ratioTRADING_DAYS_PER_YEAR = 252           # annualization assumption# No randomness is used. The same input and settings must produce the same result.

6.2 Signal to Monthly Target Weights

Compute the 12-month return from month-end prices and build target weights that equal-weight only the ETFs with a positive signal. The month-ends before the first valid signal are left at 0% ETFs and 100% cash.

momentum_baseline.pypython
def make_monthly_targets(prices: pd.DataFrame) -> pd.DataFrame:    month_end = prices.resample(REBALANCE_FREQUENCY).last()    momentum_12m = month_end.pct_change(periods=LOOKBACK_MONTHS, fill_method=None)    targets = pd.DataFrame(0.0, index=momentum_12m.index, columns=prices.columns)    for signal_date, row in momentum_12m.iterrows():        winners = row.index[row.gt(0) & row.notna()]        if len(winners) > 0:            targets.loc[signal_date, winners] = 1.0 / len(winners)        # if there are no winners, ETF weights stay 0 and cash is 100%    return targets

6.3 Signal Date to Execution Date (Next-Trading-Day Lag)

Shift the month-end signal onto the date where it is actually executed — after the close of the EXECUTION_LAG_TRADING_DAYS-th trading day. This is the point that keeps future information out of the signal.

momentum_baseline.pypython
def move_targets_to_execution_dates(    prices: pd.DataFrame, monthly_targets: pd.DataFrame) -> pd.DataFrame:    rows = []    for signal_date, target in monthly_targets.iterrows():        first_after = prices.index.searchsorted(signal_date, side="right")        exec_pos = first_after + EXECUTION_LAG_TRADING_DAYS - 1        if exec_pos < len(prices.index):            rows.append((prices.index[exec_pos], target))    if not rows:        raise ValueError("No executable rebalancing date.")    out = pd.DataFrame(        [t for _, t in rows], index=[d for d, _ in rows], columns=prices.columns    )    if out.index.has_duplicates:        raise ValueError("Multiple target weights generated for the same execution date.")    return out

6.4 Daily Portfolio: Returns, Then Turnover, Then Cost

Apply each trading day's close return to the pre-rebalance holdings first; on an execution day, deduct a turnover-proportional cost from net assets and reallocate to the target weights. The cost keeps compounding into later weights.

momentum_baseline.pypython
def run_portfolio(    daily_returns: pd.DataFrame,    target_on_execution: pd.DataFrame,    one_way_cost_rate: float,) -> pd.DataFrame:    columns = daily_returns.columns    target_lookup = {d: target_on_execution.loc[d] for d in target_on_execution.index}    asset_values = pd.Series(0.0, index=columns)    cash_value = 1.0    rows = []    for date, asset_return in daily_returns.iterrows():        value_at_open = asset_values.sum() + cash_value        # previous-to-current close return on existing holdings; assumed daily return on cash        asset_values = asset_values * (1.0 + asset_return)        cash_value = cash_value * (1.0 + CASH_RETURN_ASSUMPTION)        value_before_trade = asset_values.sum() + cash_value        if value_before_trade <= 0:            raise ValueError("Portfolio value fell to zero or below after applying returns.")        pre_trade_weights = asset_values / value_before_trade        turnover = 0.0        trading_cost = 0.0        if date in target_lookup:            target = target_lookup[date]            if not np.isclose(target.sum(), 0.0) and not np.isclose(target.sum(), 1.0):                raise ValueError("The sum of ETF target weights must be 0 or 1.")            # project-internal definition that also counts entry from an all-cash state            turnover = float((target - pre_trade_weights).abs().sum())            trading_cost = value_before_trade * one_way_cost_rate * turnover            value_after_trade = value_before_trade - trading_cost            if value_after_trade <= 0:                raise ValueError("Portfolio value fell to zero or below after deducting trading cost.")            asset_values = target * value_after_trade            cash_value = value_after_trade - asset_values.sum()        else:            value_after_trade = value_before_trade        rows.append({            "date": date,            "net_return": value_after_trade / value_at_open - 1.0,            "turnover": turnover,            "trading_cost": trading_cost,            "cash_weight": cash_value / value_after_trade,        })    return pd.DataFrame(rows).set_index("date")

6.5 The Rest of the Wiring

The full script wires the functions above in this order. For length, only the core is shown here; the items below translate directly into code.

  • Price-load validation — check that dates are strictly increasing, have no duplicates, have no missing values, and that prices are positive; stop the run if any of these fail. Do not drop missing rows or reindex/interpolate dates.
  • Benchmark target weights — a DataFrame with a single 1/N row at the start date.
  • Cost-scenario loop — for each value in COST_SCENARIOS_BP, run run_portfolio from scratch to produce one performance-table row.
  • Performance metrics — from post-cost returns, compute cumulative return, CAGR, annualized volatility (√252), maximum drawdown, Sharpe (risk-free rate RISK_FREE_RETURN_DAILY), total turnover, months in cash, and average cash weight.
  • Experiment record — alongside the results, save the code version, the raw CSV's SHA-256 hash, the Python / pandas / numpy versions, and the full settings as JSON. This is for reproduction checks.

This code is an educational baseline. It is not an order system or an execution engine for live use, and it does not directly model opening prices, quotes, partial fills, taxes, or market impact.

7. How to Read the Performance Table: Not "Who Won?" but "What Was Computed Under the Same Conditions?"

This article does not include unverified real performance figures or a "beats AI" conclusion. The reader generates a table of the form below with a fixed data snapshot and the code.

StrategyCostCumulative returnCAGRAnn. volatilityMax drawdownSharpeTotal turnoverMonths in cashAvg. cash weight
Simple momentum0, 5, 10, 20 bp……………………
Equal-weight buy and hold0, 5, 10, 20 bp……………………

Below the table, fix the calculation conditions too. Even identical numbers do not compare if the following differ.

  • Shared evaluation start and end dates for the strategy and the benchmark (first and last rows of the common trading-day file)
  • Initial lookback handling: 100% cash until the first valid signal, not excluded from the evaluation
  • Months in cash: the number of months whose last trading day has a cash weight above 0, including the initial cash window
  • The ETF list used / price column definition / raw file SHA-256 / code version / return frequency
  • The formulas and annualization for CAGR, volatility, maximum drawdown, and Sharpe, plus TRADING_DAYS_PER_YEAR
  • The turnover formula, whether initial entry is included, how the one-way cost is interpreted, and when the cost is deducted
  • The cash return assumption CASH_RETURN_ASSUMPTION and the missing-data rule

The Sharpe ratio's calculation assumptions and the limits of judging by it alone were covered in Part 5. [S7] This table is only a historical summary computed from a fixed data snapshot and assumptions.

8. Validation Checklist: The Simpler the Strategy, the More You Should Suspect Implementation Errors

The difficulty of simple momentum is not that it has many rules. It is that small timing errors, differences in data handling, and differences in the cost definition change the result. Before attaching the code to real data, check the following.

  • Were the ETF list and selection criteria decided before checking performance?
  • Was each ETF actually listed at the backtest start date?
  • Did you record whether the price column is adjusted price or total return, and the dividend/split treatment?
  • Did you save the raw file, its SHA-256 hash, the code version, and the settings alongside the results?
  • Is the month-end signal applied from the next trading day, with no future prices mixed into the signal calculation?
  • Does the cash weight become 100% when there is no positive signal?
  • Is initial entry included in turnover and cost?
  • Do the strategy and benchmark use the same data, period, and cost definition, computing pre- and post-cost with the same metrics?
  • Does running twice with the same input and settings produce the same result?

With small synthetic inputs you can also check: with a single ETF whose signal stays positive, does the strategy keep a 100% weight; in a month where all signals are negative, does cash become 100%; when moving from cash to an ETF, do turnover and cost occur; does buy and hold have zero turnover after the initial entry.

9. How My View Changed: From a Vague Expectation to Fixed Comparison Conditions

StageNote
Earlier beliefI vaguely thought an AI model using more information was likely to beat a simple rule.
What I confirmed this timeEven a simple rule cannot serve as a point of comparison unless the data column definitions, the month-end signal, the execution lag, missing values, turnover, and the cost-deduction timing all line up exactly.
Current viewFirst build a simple baseline anyone can rerun, then compare AI models under the same asset set, data, costs, and evaluation table.
What I still do not knowWhether this baseline or a later AI model will outperform in the future, or survive real fill costs, cannot be known from a backtest alone.

10. Questions This Baseline Cannot Answer

Fixing a baseline makes the comparison more honest, but it does not remove the uncertainty of a backtest.

  • Cross-sectional and time-series momentum have different signals and portfolio construction, so results should not be mixed just because the name is the same. [S1] [S2]
  • The fill limits of a fixed-bp cost were laid out in Part 5. This article does not claim to solve them; it only applies the same cost scenarios to the baseline and later ML experiments. [S3] [S4]
  • The adjusted-price formula, revision history, and delisted-ETF inclusion of the price provider to be used were not verified as a single dataset. Whether this article's candidate set was investable at each past point in time was also not confirmed.
  • Rebalancing after the next trading day's close is an approximation to fix the order of events, not a model of actual fill prices, order types, or partial fills.
  • Taxes, account conditions, exchanges, order size, and market impact were not verified case by case in this study.
  • A historical performance table does not guarantee future returns and is not a basis for buying or selling any particular ETF.

In particular, if you repeatedly test many rule, period, and ETF combinations and then adopt only the combination that looks best, the risk that you have found a result fitted to the past data grows. [S6] So in the next ML experiment, "it looks better" alone is not a reason to change the following.

Kept fixedChanged (prediction model only)
ETF set, data snapshot (file, hash, adjustment method), observation periodModel input variables
Monthly rebalancing, next-trading-day execution, long/cash constraintTraining and validation windows
Turnover definition and the 0, 5, 10, 20 bp cost scenariosRetraining frequency
Performance-table metrics, benchmark (buy and hold + this baseline)How the signal is turned into a score or probability

Conclusion: A Good Baseline Is Not a Flashy Strategy but a Fixed Set of Comparison Conditions

This article's baseline reduces to one sentence.

Hold, at equal weight, only the ETFs whose trailing 12-month return at month-end is positive; when no signal is positive, stay in cash.

This rule is neither an optimized answer nor a strategy that promises future performance. Its value is that it lets you see what a later AI model actually changes under the same data, costs, execution timing, and risk evaluation table. The starting point for a complex model is not a complex explanation; it is a simple baseline that anyone can rerun.

All examples and code in this article are for educational and research purposes and are not investment advice. They do not recommend buying or selling any particular ETF or guarantee returns.

Sources

Report an error or share feedback

Open a draft with this article’s title and URL. Review the message and recipient before sending.

To: [email protected]

Open email draft

If no email app opens, copy these details into your usual email service.

Contact information