Forge Fate
Quant & Data Research

What to Decide Before Backtesting: Questions and Rules for U.S. ETF Quant Research

8 min read

Evidence and scope — Research questions and advance planning

This article sets out questions, candidate assets, and comparison rules for a US ETF study. Neither the earlier methodology series nor this plan is presented as evidence of investment performance or productivity gains.

If you have learned the basics of Python, you may be eager to download price data and start building strategies right away. But in your first quant research project, the more important deliverable is not a return chart—it is a written question and set of comparison rules defined before you see the results. That helps reduce the temptation to select only favorable outcomes and makes failed experiments useful for future decisions.

This article is the starting point for a new eight-part series focused exclusively on U.S.-listed ETFs. The earlier 12-part series, Quant Investing and AI, covered validation methodology. It did not demonstrate actual investment performance or productivity gains. This series focuses on how to design and document a small research project in advance.

What This Series Covers

PartTopicCore question
Part 1Questions and experiment scopeWhat are we testing?
Part 2DataWhere can we obtain data, under what permissions, and with what price definition?
Part 3Point-in-time testingCould the signal and label truly have been known at that time?
Part 4Baseline backtestingDid we compare it with a simple baseline under the same conditions?
Part 5Costs and riskDid we consider trading costs, drawdowns, and turnover together?
Part 6Logistic comparisonDid we compare a simple rule and a classification model using the same time splits?
Part 7Record automationHow should code, data, and run records be preserved?
Part 8System and AI evaluationWhat should be evaluated, including the cost of using and reviewing AI?

The goal of this article is not to complete a strategy. It is to record the following four items in a research notebook first:

  • The baseline and hypothesis to test
  • Predefined rules for signals, trading, and costs
  • Time splits and falsification conditions to lock in before viewing performance
  • Data providers, price definitions, and preservation requirements to verify in the next article

Define the Question in a Sentence Before Choosing ETFs

“What will go up?” may be interesting, but it is too broad to serve as a research question. The question used here needs to be narrower and falsifiable.

Among SPY, IEF, and GLD, how does a rule that holds only ETFs with positive trailing 12-month returns at equal monthly weights compare with simple buy-and-hold over the same evaluation period in after-cost returns, drawdowns, and turnover?

This sentence does not promise a profit. It also defines a falsification condition: the rule will not be adopted if its after-cost returns are consistently worse than the baseline, or if its drawdowns or turnover impose an unacceptable trade-off.

Positive 12-month momentum is only a starting hypothesis. Prior research reported return persistence from one to 12 months across 58 liquid equity-index, currency, commodity, and bond-futures markets. But that sample was not made up of ETFs, and it does not guarantee results for the SPY–IEF–GLD combination, an equal-weight rule, or after-cost performance. [S4]

The Three ETFs Are an Educational Scope, Not a Recommendation List

This study limits its scope to three U.S.-listed ETFs. They are selected for educational purposes—to compare exposure to equities, U.S. Treasuries, and gold prices within one experiment—not as an optimal candidate universe or a recommendation to buy.

ETFAsset exposure in official materialsRole in this study
SPYExposure to the S&P 500 Index, which measures the U.S. large-cap equity segment [S1]One equity-risk sleeve
IEFExposure to U.S. Treasury bonds with 7–10 years remaining maturity [S2]One intermediate-term U.S. Treasury sleeve
GLDSeeks to reflect the performance of the price of gold bullion, less expenses [S3]One gold-price-exposure sleeve

SPY generally tracks the price and return performance of the S&P 500, an index measuring the U.S. large-cap equity segment. SPY was launched on January 22, 1993. [S1] IEF provides exposure to an index of U.S. Treasury bonds with 7–10 years remaining maturity and was launched on July 22, 2002. [S2] GLD seeks to reflect the performance of the price of gold bullion, less expenses, and was launched on November 18, 2004. [S3]

It is possible to verify that all three products launched before 2006. That does not mean reproducible daily price data from 2006 through 2025 has already been obtained. Price-adjustment methods, dividend treatment, missing observations, trading-day calendars, usage rights, and redistribution terms must be verified separately. [S1] [S2] [S3]

There is also a survivorship-bias limitation in applying products that still exist today to the past. Research suggests that a survivor-truncated sample can appear more predictable than it really was, so this article does not claim that these three ETFs represent the complete investable universe at the time or were historically optimal. [S5]

Write Down the Trading Rules Before Seeing the Results

These rules are not market facts. They are educational approximations intended to make the research reproducible. They are written down now so they are not changed later to fit the results once data becomes available.

The signal date is the end of each month. An ETF is included as a holding candidate for the following month when its trailing 12-month return is positive. Signals calculated at month-end are assumed to be executed at the next trading day’s close, and the new weights apply only to returns after that execution.

If one or more ETFs are selected, they receive equal weights. If none are selected, the portfolio holds 100% cash. Cash is assumed to earn 0%. One-way trading costs are tested in four scenarios: 0, 5, 10, and 20 basis points.

These rules do not claim to reproduce real order execution. Reported ETF prices, net asset value, and actual market execution prices can differ, so the next step must assess whether the price definition and execution approximation fit the research purpose. [S2] [S3]

Lock In the Baseline and Evaluation Criteria Together

The momentum rule should not be judged “good” in isolation. It needs to be compared with a simple baseline over the same period and within the same cost framework.

The baseline is a simple buy-and-hold portfolio: buy SPY, IEF, and GLD at equal weights at the start of the period and hold them without further rebalancing. The comparison is between this baseline and the 12-month positive-momentum rule. Both strategies must use the same price definition and evaluation period. Under the cost scenarios, one-way costs should apply only to each strategy’s actual trades.

Evaluation is limited to three measures:

  • After-cost return: Does a difference remain after costs are deducted?
  • Maximum drawdown: Is the loss path worth accepting, not merely the possibility of higher returns?
  • Turnover: Does the amount of portfolio change justify its costs and execution complexity?

This comparison does not reproduce prior research exactly or forecast future performance. It tests whether the idea of return persistence observed in futures markets is worth examining within this limited ETF universe. [S4]

Decide the Time Splits Before Looking at Performance

The target data period is 2006 through 2025. However, this is a design proposal made before data availability has been verified; it does not treat unavailable periods as though the data already exists.

Under the design, 2006 is a preparation period for calculating signals. The initial training and development period runs from 2007 through 2017, the validation period from 2018 through 2020, and the final test period from 2021 through 2025.

This time order will also be maintained when logistic regression is introduced in Part 6. The availability time of signals and labels must be checked separately, and preprocessing such as missing-data handling and standardization, along with model fitting, must occur only within the applicable training window. Performance from the final test period will not be repeatedly examined during development or used as evidence for decisions until at least the comparison design in Part 6 is in place.

Having No Data Yet Is Part of the Research

There is currently no price CSV on hand and no confirmed data service. Only the presence of local Python 3.14.5 has been verified; the virtual environment, package compatibility, and runnable state have not yet been tested.

In an earlier small-scale access attempt, a Yahoo Chart response with status 429 was observed. Stooq displayed HTML in a browser, but a CSV was not obtained. These were observations from that project at that time; they do not establish a permanent outage for either service or general availability.

Accordingly, the data provider, usage rights, redistribution conditions, definitions of daily closing and adjusted prices, dividend treatment, missing-data handling, and preservation method for source files remain undecided. Part 2 will verify these conditions using actual providers and files, then determine the price definition, preservation requirements, and quality-check criteria. Product inception dates do not substitute for those data requirements. [S1] [S2] [S3]

Rules for Preserving a Research Trail, Not Just Code

A reproducible system is not created by strategy code alone. Going forward, preserve the code, raw data and metadata, execution environment, dependency packages, configuration values, and result summaries together.

When using AI, record not only its inputs and outputs but also the human review and revisions, including the time they required. Failed data-access attempts, assumption changes, failed quality checks, and reasons for reruns should also be retained as failure records. Actual performance and any time savings from AI use have not yet been measured, so no improvement claim is being made.

The research project is intended to be managed separately from Content OS, but it has not been set up yet. Part 7 will cover how to automate this record-keeping workflow.

The First Research Contract

What this article establishes is not whether the strategy wins or loses, but the boundaries for interpreting its results. The research scope is limited to the U.S.-listed ETFs SPY, IEF, and GLD. These products form an educational scope for comparing equity, intermediate-term U.S. Treasury, and gold-price exposure; they are not investment recommendations. [S1] [S2] [S3]

The hypothesis being tested is positive 12-month momentum. It will be compared with buy-and-hold on after-cost return, maximum drawdown, and turnover, while treating the limitations of applying currently surviving products to the past and the still-unconfirmed data situation as prerequisites for interpreting the results. [S4] [S5]

The next article will not make the strategy more complicated. It will first verify, using actual files, where data can be obtained, under what rights and price definitions, and under what conditions it should be preserved and checked.

This series is a design record for educational and research purposes. It is not an instruction to place actual orders or a solicitation to buy or sell any specific product.

Sources

Report an error or share feedback

Open a draft with this article’s title and URL. Review the message and recipient before sending.

To: [email protected]

Open email draft

If no email app opens, copy these details into your usual email service.

Contact information