The First Baseline: Building a Simple Momentum Strategy
Evidence and scope — Educational baseline code
The code specifies signal, execution-lag, weight, and cost rules for a simple momentum baseline. No validated actual-performance figures are presented. It is not a production system reproducing real orders, quotes, or partial fills.
To evaluate an AI model's performance, you must first fix a simple rule that was not changed after seeing the results, along with identical comparison conditions. The momentum strategy built in this article is not an answer that guarantees returns; it is the starting line that later experiments will share.
In Part 5, rather than judging a strategy by returns alone, we defined an evaluation table that also looks at maximum drawdown, turnover, and cost sensitivity. This article builds the first shared baseline that framework will actually be applied to: a simple momentum strategy that later AI and ML models will be measured against.
At first I assumed that an AI model using more information would naturally become a better strategy. But complex results are hard to interpret without a simple point of comparison. Even if AI shows a higher cumulative return, you cannot tell what that number means unless you can verify it beats a simple rule on the same ETF universe, the same period, the same trading costs, and the same execution timing.
So the goal of this article is not to find "the strategy that looks best." It is to first lock in a baseline that anyone can rerun with the same data snapshot and settings. The intended reader is an individual investor with basic Python at the pandas level who is building a reproducible backtest for the first time.
"Momentum" covers several different strategies. Cross-sectional momentum, which buys relatively strong assets and sells relatively weak ones, must be distinguished from time-series momentum, which uses each asset's own past return as a signal. [S1] [S2] This article uses a simplified version of the latter, adapted to a rule that holds ETFs or waits in cash without any short positions.
All examples and code in this article are for educational and research purposes and are not a recommendation to buy or sell any particular ETF, nor a guarantee of future returns.
1. Baseline Rule: "Hold If the 12-Month Signal Is Positive; Otherwise Hold Cash"
This strategy is neither an academically validated optimal solution nor a claim that it will outperform in the future. It is an intentionally simplified project rule, designed to keep later experiments comparable under stable conditions.
| Item | Fixed rule in this article |
|---|---|
| Signal | Trailing 12-month return at month-end |
| Buy condition | ETFs with a positive trailing 12-month return |
| Allocation | Equal weight across ETFs with positive signals |
| Cash rule | 100% cash if no ETF has a positive signal |
| Initial lookback window | 0% ETFs, 100% cash until the first valid 12-month signal; still included in the evaluation |
| Rebalancing | Once per month |
| Execution timing | Applied from the next trading day after the month-end signal is confirmed |
| Position constraint | Long or cash only; no shorting or leverage |
| Costs | One-way sensitivity scenarios of 0, 5, 10, and 20 bp |
The asset classes and implementation used in time-series momentum research differ from the ETF long/cash rule here. [S2] So a 12-month lookback, monthly rebalancing, positive signals, and equal weighting are not the only standard, nor the optimal settings. The rule can be summarized as follows.
At month-end, compute each ETF's trailing 12-month return. If at least one ETF is positive -> hold those ETFs at equal weight If no ETF is positive -> 100% cashApply the month-end signal starting from the next trading day.For example, if the candidate universe has four ETFs and the month-end signals are as below, you hold the positive A and C at half each, while B, D, and cash are 0%. If all four signals are negative, cash is 100%.
| ETF | Trailing 12-month return | Target weight next period |
|---|---|---|
| ETF_A | Positive | 50% |
| ETF_B | Negative | 0% |
| ETF_C | Positive | 50% |
| ETF_D | Negative | 0% |
What matters here is less the "hold only positive ETFs" rule itself and more that the rule is fixed before you see the results. If you keep changing 12 months to 11, monthly to weekly, or equal weight to arbitrary weights because a backtest looked bad, the baseline stops functioning as a point of comparison.
2. Data and Liquidity: Verify Through Records, Not Claims
A single sentence like "we used sufficiently liquid ETFs" is not enough. ETF trades can incur brokerage commissions and additional transaction costs, and the bid-ask spread is itself a cost that lowers potential returns. The SEC notes that ETFs with higher liquidity and trading volume tend to have narrower spreads. [S3] Reproducibility, likewise, is not just publishing code. You also have to fix which file you used as input, what that file is and when it was downloaded, and what the price column reflects. For instance, one ETF total-return disclosure assumes dividends are reinvested but does not include brokerage commissions. [S5] In other words, how the price data handles dividends and splits must be recorded separately from how the strategy deducts trading costs.
For each real backtest, fill in the table below.
| Record item | Content |
|---|---|
| ETF list | Ticker, selection date, inception date, exclusion reason — fixed before checking performance |
| Observation period | Start and end dates, plus each ETF's first usable date |
| Provider and source | Data provider name, source URL or file path, download time (UTC) |
| Price column definition | Whether it is adjusted close or a total-return index |
| Dividend and split treatment | Provider definition or a verifiable document |
| Missing-data rule | Whether rows are dropped, held, or reindexed, and why |
| Raw archive | CSV filename, SHA-256 hash, code version |
| Liquidity evidence | Trading value by sub-period, median bid-ask spread, any trading halts |
| Survivorship check | Whether the ETF was investable at the time, and whether only post-inception data was used |
This article does not rule that any particular ETF is always sufficiently liquid. Without directly checking trading value by sub-period, spreads, inception dates, and trading halts, that judgment cannot be generalized. Whether this article's candidate set and dataset have survivorship bias or point-in-time investability issues has also not yet been verified.
3. Timing: Information Learned at Month-End Is Used From the Next Trading Day
The thing to be especially careful about is not mixing the moment you learned information with the moment you traded on it. If you computed a trailing 12-month return from the month-end close, you must not assume you already knew that signal and traded on it before that close was final. This article uses daily adjusted-price returns from the previous close to the current close, so it fixes the order of events as follows.
Month-end price is final -> compute the trailing 12-month return -> identify ETFs with a positive signal, derive target weights -> keep current weights through the close of the EXECUTION_LAG_TRADING_DAYS-th trading day after the signal -> after that close, deduct costs and rebalance to the target weights -> from the following trading day, apply daily returns at the new weightsWith EXECUTION_LAG_TRADING_DAYS = 1, you rebalance after the close of the first trading day following the signal. That execution day's previous-close-to-current-close return still applies to the old weights, and the new target weights apply to returns from the following trading day onward. This convention is a close-only educational approximation, not a reproduction of actual fill prices, order types, or partial fills.
4. Put the Benchmark Under the Same Conditions: Equal-Weight Buy and Hold
To see what the momentum strategy changed, you need a benchmark that shares the same asset set, period, data, and cost definition. This article's benchmark is equal-weight buy and hold. It uses the same ETF universe, the same observation start and end dates, and the same adjusted-price data; it allocates equal weight at the start and never rebalances. Both the strategy and the benchmark compute pre- and post-cost results, but the benchmark reflects only the initial entry cost.
This comparison is not meant to predetermine whether momentum or buy and hold is better. It is a reference for reading, under the same conditions, what changed when the signal rule changed — later including AI models. When looking at the results table, ask:
- Did the strategy change the result, or did the ETF composition and period choice change it?
- How does the pre- vs. post-cost difference relate to turnover?
- What was the cost and the benefit of spending more months in cash?
- Did a lower maximum drawdown come with a lower return or a larger cash allocation?
5. Trading Costs: Lock the Part 5 Evaluation Table Into This Baseline
The cost sensitivity and turnover evaluation table defined in Part 5 is not re-explained here. This article fixes that definition into the code's execution order so it applies unchanged to later ML experiments.
turnover_t = Sum_i |ETF weight_i just before execution - target weight_i|cost_t = one-way cost rate x turnover_t x net asset value just before executionEntry from an all-cash state is also counted in turnover, and 0, 5, 10, 20 bp are not common real-world values but research sensitivity scenarios. On the execution day, the cost-deducted net assets are reallocated to the target weights, so the cost keeps compounding into later weights. This weight-change measure is a project-internal definition, different from the portfolio turnover used in disclosures. [S8] Because ETF trades can incur commissions and transaction costs, the assumptions and their scope must be disclosed. [S3] [S4] The cash return is fixed at 0% in the code; this is a simplifying assumption, not a proxy for market rates.
6. Reproducible Backtest Code
The input is a single wide-format daily price CSV. It has a date column and one price column per ETF, and every ETF column must be the same kind of price whose dividend and split treatment you have verified — all adjusted close, or all total-return index.
date,ETF_A,ETF_B,ETF_C,ETF_D2020-01-02,100.12,50.43,...2020-01-03,99.80,50.71,...6.1 Settings to Fix Before Running
These are values fixed before seeing any performance. The source, download time, price column definition, and missing-data rule are recorded alongside as separate metadata.
from pathlib import Pathimport numpy as npimport pandas as pdCODE_VERSION = "baseline-momentum-v1.0"RAW_PRICE_FILE = Path("data/prices_adjusted_close.csv")# ETF list fixed before checking performanceUNIVERSE = ("ETF_A", "ETF_B", "ETF_C", "ETF_D")LOOKBACK_MONTHS = 12REBALANCE_FREQUENCY = "ME" # month-endEXECUTION_LAG_TRADING_DAYS = 1 # execute after the next trading day's closeCOST_SCENARIOS_BP = [0, 5, 10, 20] # one-way, research sensitivityALLOW_SHORT = FalseCASH_RETURN_ASSUMPTION = 0.0 # daily return on the cash balance (simplifying assumption)RISK_FREE_RETURN_DAILY = 0.0 # risk-free rate for the Sharpe ratioTRADING_DAYS_PER_YEAR = 252 # annualization assumption# No randomness is used. The same input and settings must produce the same result.6.2 Signal to Monthly Target Weights
Compute the 12-month return from month-end prices and build target weights that equal-weight only the ETFs with a positive signal. The month-ends before the first valid signal are left at 0% ETFs and 100% cash.
def make_monthly_targets(prices: pd.DataFrame) -> pd.DataFrame: month_end = prices.resample(REBALANCE_FREQUENCY).last() momentum_12m = month_end.pct_change(periods=LOOKBACK_MONTHS, fill_method=None) targets = pd.DataFrame(0.0, index=momentum_12m.index, columns=prices.columns) for signal_date, row in momentum_12m.iterrows(): winners = row.index[row.gt(0) & row.notna()] if len(winners) > 0: targets.loc[signal_date, winners] = 1.0 / len(winners) # if there are no winners, ETF weights stay 0 and cash is 100% return targets6.3 Signal Date to Execution Date (Next-Trading-Day Lag)
Shift the month-end signal onto the date where it is actually executed — after the close of the EXECUTION_LAG_TRADING_DAYS-th trading day. This is the point that keeps future information out of the signal.
def move_targets_to_execution_dates( prices: pd.DataFrame, monthly_targets: pd.DataFrame) -> pd.DataFrame: rows = [] for signal_date, target in monthly_targets.iterrows(): first_after = prices.index.searchsorted(signal_date, side="right") exec_pos = first_after + EXECUTION_LAG_TRADING_DAYS - 1 if exec_pos < len(prices.index): rows.append((prices.index[exec_pos], target)) if not rows: raise ValueError("No executable rebalancing date.") out = pd.DataFrame( [t for _, t in rows], index=[d for d, _ in rows], columns=prices.columns ) if out.index.has_duplicates: raise ValueError("Multiple target weights generated for the same execution date.") return out6.4 Daily Portfolio: Returns, Then Turnover, Then Cost
Apply each trading day's close return to the pre-rebalance holdings first; on an execution day, deduct a turnover-proportional cost from net assets and reallocate to the target weights. The cost keeps compounding into later weights.
def run_portfolio( daily_returns: pd.DataFrame, target_on_execution: pd.DataFrame, one_way_cost_rate: float,) -> pd.DataFrame: columns = daily_returns.columns target_lookup = {d: target_on_execution.loc[d] for d in target_on_execution.index} asset_values = pd.Series(0.0, index=columns) cash_value = 1.0 rows = [] for date, asset_return in daily_returns.iterrows(): value_at_open = asset_values.sum() + cash_value # previous-to-current close return on existing holdings; assumed daily return on cash asset_values = asset_values * (1.0 + asset_return) cash_value = cash_value * (1.0 + CASH_RETURN_ASSUMPTION) value_before_trade = asset_values.sum() + cash_value if value_before_trade <= 0: raise ValueError("Portfolio value fell to zero or below after applying returns.") pre_trade_weights = asset_values / value_before_trade turnover = 0.0 trading_cost = 0.0 if date in target_lookup: target = target_lookup[date] if not np.isclose(target.sum(), 0.0) and not np.isclose(target.sum(), 1.0): raise ValueError("The sum of ETF target weights must be 0 or 1.") # project-internal definition that also counts entry from an all-cash state turnover = float((target - pre_trade_weights).abs().sum()) trading_cost = value_before_trade * one_way_cost_rate * turnover value_after_trade = value_before_trade - trading_cost if value_after_trade <= 0: raise ValueError("Portfolio value fell to zero or below after deducting trading cost.") asset_values = target * value_after_trade cash_value = value_after_trade - asset_values.sum() else: value_after_trade = value_before_trade rows.append({ "date": date, "net_return": value_after_trade / value_at_open - 1.0, "turnover": turnover, "trading_cost": trading_cost, "cash_weight": cash_value / value_after_trade, }) return pd.DataFrame(rows).set_index("date")6.5 The Rest of the Wiring
The full script wires the functions above in this order. For length, only the core is shown here; the items below translate directly into code.
- Price-load validation — check that dates are strictly increasing, have no duplicates, have no missing values, and that prices are positive; stop the run if any of these fail. Do not drop missing rows or reindex/interpolate dates.
- Benchmark target weights — a DataFrame with a single
1/Nrow at the start date. - Cost-scenario loop — for each value in
COST_SCENARIOS_BP, runrun_portfoliofrom scratch to produce one performance-table row. - Performance metrics — from post-cost returns, compute cumulative return, CAGR, annualized volatility (√252), maximum drawdown, Sharpe (risk-free rate
RISK_FREE_RETURN_DAILY), total turnover, months in cash, and average cash weight. - Experiment record — alongside the results, save the code version, the raw CSV's SHA-256 hash, the Python / pandas / numpy versions, and the full settings as JSON. This is for reproduction checks.
This code is an educational baseline. It is not an order system or an execution engine for live use, and it does not directly model opening prices, quotes, partial fills, taxes, or market impact.
7. How to Read the Performance Table: Not "Who Won?" but "What Was Computed Under the Same Conditions?"
This article does not include unverified real performance figures or a "beats AI" conclusion. The reader generates a table of the form below with a fixed data snapshot and the code.
| Strategy | Cost | Cumulative return | CAGR | Ann. volatility | Max drawdown | Sharpe | Total turnover | Months in cash | Avg. cash weight |
|---|---|---|---|---|---|---|---|---|---|
| Simple momentum | 0, 5, 10, 20 bp | … | … | … | … | … | … | … | … |
| Equal-weight buy and hold | 0, 5, 10, 20 bp | … | … | … | … | … | … | … | … |
Below the table, fix the calculation conditions too. Even identical numbers do not compare if the following differ.
- Shared evaluation start and end dates for the strategy and the benchmark (first and last rows of the common trading-day file)
- Initial lookback handling: 100% cash until the first valid signal, not excluded from the evaluation
- Months in cash: the number of months whose last trading day has a cash weight above 0, including the initial cash window
- The ETF list used / price column definition / raw file SHA-256 / code version / return frequency
- The formulas and annualization for CAGR, volatility, maximum drawdown, and Sharpe, plus
TRADING_DAYS_PER_YEAR - The turnover formula, whether initial entry is included, how the one-way cost is interpreted, and when the cost is deducted
- The cash return assumption
CASH_RETURN_ASSUMPTIONand the missing-data rule
The Sharpe ratio's calculation assumptions and the limits of judging by it alone were covered in Part 5. [S7] This table is only a historical summary computed from a fixed data snapshot and assumptions.
8. Validation Checklist: The Simpler the Strategy, the More You Should Suspect Implementation Errors
The difficulty of simple momentum is not that it has many rules. It is that small timing errors, differences in data handling, and differences in the cost definition change the result. Before attaching the code to real data, check the following.
- Were the ETF list and selection criteria decided before checking performance?
- Was each ETF actually listed at the backtest start date?
- Did you record whether the price column is adjusted price or total return, and the dividend/split treatment?
- Did you save the raw file, its SHA-256 hash, the code version, and the settings alongside the results?
- Is the month-end signal applied from the next trading day, with no future prices mixed into the signal calculation?
- Does the cash weight become 100% when there is no positive signal?
- Is initial entry included in turnover and cost?
- Do the strategy and benchmark use the same data, period, and cost definition, computing pre- and post-cost with the same metrics?
- Does running twice with the same input and settings produce the same result?
With small synthetic inputs you can also check: with a single ETF whose signal stays positive, does the strategy keep a 100% weight; in a month where all signals are negative, does cash become 100%; when moving from cash to an ETF, do turnover and cost occur; does buy and hold have zero turnover after the initial entry.
9. How My View Changed: From a Vague Expectation to Fixed Comparison Conditions
| Stage | Note |
|---|---|
| Earlier belief | I vaguely thought an AI model using more information was likely to beat a simple rule. |
| What I confirmed this time | Even a simple rule cannot serve as a point of comparison unless the data column definitions, the month-end signal, the execution lag, missing values, turnover, and the cost-deduction timing all line up exactly. |
| Current view | First build a simple baseline anyone can rerun, then compare AI models under the same asset set, data, costs, and evaluation table. |
| What I still do not know | Whether this baseline or a later AI model will outperform in the future, or survive real fill costs, cannot be known from a backtest alone. |
10. Questions This Baseline Cannot Answer
Fixing a baseline makes the comparison more honest, but it does not remove the uncertainty of a backtest.
- Cross-sectional and time-series momentum have different signals and portfolio construction, so results should not be mixed just because the name is the same. [S1] [S2]
- The fill limits of a fixed-bp cost were laid out in Part 5. This article does not claim to solve them; it only applies the same cost scenarios to the baseline and later ML experiments. [S3] [S4]
- The adjusted-price formula, revision history, and delisted-ETF inclusion of the price provider to be used were not verified as a single dataset. Whether this article's candidate set was investable at each past point in time was also not confirmed.
- Rebalancing after the next trading day's close is an approximation to fix the order of events, not a model of actual fill prices, order types, or partial fills.
- Taxes, account conditions, exchanges, order size, and market impact were not verified case by case in this study.
- A historical performance table does not guarantee future returns and is not a basis for buying or selling any particular ETF.
In particular, if you repeatedly test many rule, period, and ETF combinations and then adopt only the combination that looks best, the risk that you have found a result fitted to the past data grows. [S6] So in the next ML experiment, "it looks better" alone is not a reason to change the following.
| Kept fixed | Changed (prediction model only) |
|---|---|
| ETF set, data snapshot (file, hash, adjustment method), observation period | Model input variables |
| Monthly rebalancing, next-trading-day execution, long/cash constraint | Training and validation windows |
| Turnover definition and the 0, 5, 10, 20 bp cost scenarios | Retraining frequency |
| Performance-table metrics, benchmark (buy and hold + this baseline) | How the signal is turned into a score or probability |
Conclusion: A Good Baseline Is Not a Flashy Strategy but a Fixed Set of Comparison Conditions
This article's baseline reduces to one sentence.
Hold, at equal weight, only the ETFs whose trailing 12-month return at month-end is positive; when no signal is positive, stay in cash.
This rule is neither an optimized answer nor a strategy that promises future performance. Its value is that it lets you see what a later AI model actually changes under the same data, costs, execution timing, and risk evaluation table. The starting point for a complex model is not a complex explanation; it is a simple baseline that anyone can rerun.
All examples and code in this article are for educational and research purposes and are not investment advice. They do not recommend buying or selling any particular ETF or guarantee returns.
Sources
- [S1] Returns to Buying Winners and Selling Losers: Implications for Stock Market Efficiency | Narasimhan Jegadeesh and Sheridan Titman; The Journal of Finance | March 1993 | https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1540-6261.1993.tb04702.x ↩
- [S2] Time series momentum | Tobias J. Moskowitz, Yao Hua Ooi, and Lasse Heje Pedersen; Journal of Financial Economics | May 2012 | https://doi.org/10.1016/j.jfineco.2011.11.003 ↩
- [S3] Updated Investor Bulletin: Exchange-Traded Funds (ETFs) | U.S. Securities and Exchange Commission, Office of Investor Education and Advocacy | 2023-02-23 | https://www.investor.gov/introduction-investing/general-resources/news-alerts/alerts-bulletins/investor-bulletins-24 ↩
- [S4] GIPS® Standards Handbook for Firms | CFA Institute and Global Investment Performance Standards | November 2020 | https://www.gipsstandards.org/standards/gips-standards-for-firms/gips-standards-handbook-for-firms/ ↩
- [S5] State Street® SPDR® S&P 500® ETF Trust Prospectus | State Street Global Advisors Trust Company and PDR Services LLC, filed with the U.S. Securities and Exchange Commission | 2026-01-26 | https://www.sec.gov/Archives/edgar/data/884394/000119312526023648/d935960d497.htm ↩
- [S6] The Probability of Backtest Overfitting | David H. Bailey, Jonathan M. Borwein, Marcos López de Prado, and Qiji Jim Zhu | 2015-02-27 | https://www.davidhbailey.com/dhbpapers/backtest-prob.pdf ↩
- [S7] The Sharpe Ratio | William F. Sharpe; Stanford University reprint of The Journal of Portfolio Management article | Fall 1994 | https://web.stanford.edu/~wfsharpe/art/sr/SR.htm ↩
- [S8] Form N-1A Registration Statement — Voya VACS Index Series S Portfolio | Voya Investors Trust, filed with the U.S. Securities and Exchange Commission | 2026-05-01 | https://www.sec.gov/Archives/edgar/data/837276/000083727626000042/f44756d1.htm ↩
Report an error or share feedback
Open a draft with this article’s title and URL. Review the message and recipient before sending.
Open email draftIf no email app opens, copy these details into your usual email service.
Related posts
Quant & Data Research Would Adding Transaction Costs Change the Conclusion? Checking Accounting and Risk Paths in a Synthetic Portfolio
Use a synthetic portfolio to check transaction-cost accounting and distinguish what ending returns, turnover, and maximum drawdown reveal.
Quant & Data Research Start a Backtest with Positions, Not Returns: One-Period Accounting for Buy and Hold vs. Momentum
Check holdings, cash, and ending values for buy-and-hold and momentum using three synthetic assets, establishing accounting rules before market testing.