What to Decide Before Backtesting: Questions and Rules for U.S. ETF Quant Research
Evidence and scope — Research questions and advance planning
This article sets out questions, candidate assets, and comparison rules for a US ETF study. Neither the earlier methodology series nor this plan is presented as evidence of investment performance or productivity gains.
If you have learned the basics of Python, you may be eager to download price data and start building strategies right away. But in your first quant research project, the more important deliverable is not a return chart—it is a written question and set of comparison rules defined before you see the results. That helps reduce the temptation to select only favorable outcomes and makes failed experiments useful for future decisions.
This article is the starting point for a new eight-part series focused exclusively on U.S.-listed ETFs. The earlier 12-part series, Quant Investing and AI, covered validation methodology. It did not demonstrate actual investment performance or productivity gains. This series focuses on how to design and document a small research project in advance.
What This Series Covers
| Part | Topic | Core question |
|---|---|---|
| Part 1 | Questions and experiment scope | What are we testing? |
| Part 2 | Data | Where can we obtain data, under what permissions, and with what price definition? |
| Part 3 | Point-in-time testing | Could the signal and label truly have been known at that time? |
| Part 4 | Baseline backtesting | Did we compare it with a simple baseline under the same conditions? |
| Part 5 | Costs and risk | Did we consider trading costs, drawdowns, and turnover together? |
| Part 6 | Logistic comparison | Did we compare a simple rule and a classification model using the same time splits? |
| Part 7 | Record automation | How should code, data, and run records be preserved? |
| Part 8 | System and AI evaluation | What should be evaluated, including the cost of using and reviewing AI? |
The goal of this article is not to complete a strategy. It is to record the following four items in a research notebook first:
- The baseline and hypothesis to test
- Predefined rules for signals, trading, and costs
- Time splits and falsification conditions to lock in before viewing performance
- Data providers, price definitions, and preservation requirements to verify in the next article
Define the Question in a Sentence Before Choosing ETFs
“What will go up?” may be interesting, but it is too broad to serve as a research question. The question used here needs to be narrower and falsifiable.
Among SPY, IEF, and GLD, how does a rule that holds only ETFs with positive trailing 12-month returns at equal monthly weights compare with simple buy-and-hold over the same evaluation period in after-cost returns, drawdowns, and turnover?
This sentence does not promise a profit. It also defines a falsification condition: the rule will not be adopted if its after-cost returns are consistently worse than the baseline, or if its drawdowns or turnover impose an unacceptable trade-off.
Positive 12-month momentum is only a starting hypothesis. Prior research reported return persistence from one to 12 months across 58 liquid equity-index, currency, commodity, and bond-futures markets. But that sample was not made up of ETFs, and it does not guarantee results for the SPY–IEF–GLD combination, an equal-weight rule, or after-cost performance. [S4]
The Three ETFs Are an Educational Scope, Not a Recommendation List
This study limits its scope to three U.S.-listed ETFs. They are selected for educational purposes—to compare exposure to equities, U.S. Treasuries, and gold prices within one experiment—not as an optimal candidate universe or a recommendation to buy.
| ETF | Asset exposure in official materials | Role in this study |
|---|---|---|
| SPY | Exposure to the S&P 500 Index, which measures the U.S. large-cap equity segment [S1] | One equity-risk sleeve |
| IEF | Exposure to U.S. Treasury bonds with 7–10 years remaining maturity [S2] | One intermediate-term U.S. Treasury sleeve |
| GLD | Seeks to reflect the performance of the price of gold bullion, less expenses [S3] | One gold-price-exposure sleeve |
SPY generally tracks the price and return performance of the S&P 500, an index measuring the U.S. large-cap equity segment. SPY was launched on January 22, 1993. [S1] IEF provides exposure to an index of U.S. Treasury bonds with 7–10 years remaining maturity and was launched on July 22, 2002. [S2] GLD seeks to reflect the performance of the price of gold bullion, less expenses, and was launched on November 18, 2004. [S3]
It is possible to verify that all three products launched before 2006. That does not mean reproducible daily price data from 2006 through 2025 has already been obtained. Price-adjustment methods, dividend treatment, missing observations, trading-day calendars, usage rights, and redistribution terms must be verified separately. [S1] [S2] [S3]
There is also a survivorship-bias limitation in applying products that still exist today to the past. Research suggests that a survivor-truncated sample can appear more predictable than it really was, so this article does not claim that these three ETFs represent the complete investable universe at the time or were historically optimal. [S5]
Write Down the Trading Rules Before Seeing the Results
These rules are not market facts. They are educational approximations intended to make the research reproducible. They are written down now so they are not changed later to fit the results once data becomes available.
The signal date is the end of each month. An ETF is included as a holding candidate for the following month when its trailing 12-month return is positive. Signals calculated at month-end are assumed to be executed at the next trading day’s close, and the new weights apply only to returns after that execution.
If one or more ETFs are selected, they receive equal weights. If none are selected, the portfolio holds 100% cash. Cash is assumed to earn 0%. One-way trading costs are tested in four scenarios: 0, 5, 10, and 20 basis points.
These rules do not claim to reproduce real order execution. Reported ETF prices, net asset value, and actual market execution prices can differ, so the next step must assess whether the price definition and execution approximation fit the research purpose. [S2] [S3]
Lock In the Baseline and Evaluation Criteria Together
The momentum rule should not be judged “good” in isolation. It needs to be compared with a simple baseline over the same period and within the same cost framework.
The baseline is a simple buy-and-hold portfolio: buy SPY, IEF, and GLD at equal weights at the start of the period and hold them without further rebalancing. The comparison is between this baseline and the 12-month positive-momentum rule. Both strategies must use the same price definition and evaluation period. Under the cost scenarios, one-way costs should apply only to each strategy’s actual trades.
Evaluation is limited to three measures:
- After-cost return: Does a difference remain after costs are deducted?
- Maximum drawdown: Is the loss path worth accepting, not merely the possibility of higher returns?
- Turnover: Does the amount of portfolio change justify its costs and execution complexity?
This comparison does not reproduce prior research exactly or forecast future performance. It tests whether the idea of return persistence observed in futures markets is worth examining within this limited ETF universe. [S4]
Decide the Time Splits Before Looking at Performance
The target data period is 2006 through 2025. However, this is a design proposal made before data availability has been verified; it does not treat unavailable periods as though the data already exists.
Under the design, 2006 is a preparation period for calculating signals. The initial training and development period runs from 2007 through 2017, the validation period from 2018 through 2020, and the final test period from 2021 through 2025.
This time order will also be maintained when logistic regression is introduced in Part 6. The availability time of signals and labels must be checked separately, and preprocessing such as missing-data handling and standardization, along with model fitting, must occur only within the applicable training window. Performance from the final test period will not be repeatedly examined during development or used as evidence for decisions until at least the comparison design in Part 6 is in place.
Having No Data Yet Is Part of the Research
There is currently no price CSV on hand and no confirmed data service. Only the presence of local Python 3.14.5 has been verified; the virtual environment, package compatibility, and runnable state have not yet been tested.
In an earlier small-scale access attempt, a Yahoo Chart response with status 429 was observed. Stooq displayed HTML in a browser, but a CSV was not obtained. These were observations from that project at that time; they do not establish a permanent outage for either service or general availability.
Accordingly, the data provider, usage rights, redistribution conditions, definitions of daily closing and adjusted prices, dividend treatment, missing-data handling, and preservation method for source files remain undecided. Part 2 will verify these conditions using actual providers and files, then determine the price definition, preservation requirements, and quality-check criteria. Product inception dates do not substitute for those data requirements. [S1] [S2] [S3]
Rules for Preserving a Research Trail, Not Just Code
A reproducible system is not created by strategy code alone. Going forward, preserve the code, raw data and metadata, execution environment, dependency packages, configuration values, and result summaries together.
When using AI, record not only its inputs and outputs but also the human review and revisions, including the time they required. Failed data-access attempts, assumption changes, failed quality checks, and reasons for reruns should also be retained as failure records. Actual performance and any time savings from AI use have not yet been measured, so no improvement claim is being made.
The research project is intended to be managed separately from Content OS, but it has not been set up yet. Part 7 will cover how to automate this record-keeping workflow.
The First Research Contract
What this article establishes is not whether the strategy wins or loses, but the boundaries for interpreting its results. The research scope is limited to the U.S.-listed ETFs SPY, IEF, and GLD. These products form an educational scope for comparing equity, intermediate-term U.S. Treasury, and gold-price exposure; they are not investment recommendations. [S1] [S2] [S3]
The hypothesis being tested is positive 12-month momentum. It will be compared with buy-and-hold on after-cost return, maximum drawdown, and turnover, while treating the limitations of applying currently surviving products to the past and the still-unconfirmed data situation as prerequisites for interpreting the results. [S4] [S5]
The next article will not make the strategy more complicated. It will first verify, using actual files, where data can be obtained, under what rights and price definitions, and under what conditions it should be preserved and checked.
This series is a design record for educational and research purposes. It is not an instruction to place actual orders or a solicitation to buy or sell any specific product.
Sources
- [S1] Fact Sheet:State Street® SPDR® S&P 500® ETF Trust, Jun2026 | State Street Investment Management | 2026-06-30 | https://www.ssga.com/library-content/products/factsheets/etfs/us/factsheet-us-en-spy.pdf ↩
- [S2] iShares 7-10 Year Treasury Bond ETF | BlackRock Fund Advisors / iShares | 2026-06-30 | https://www.ishares.com/us/literature/fact-sheet/ief-ishares-7-10-year-treasury-bond-etf-fund-fact-sheet-en-us.pdf ↩
- [S3] Fact Sheet:SPDR® Gold Shares, Jun2026 | State Street Investment Management | 2026-06-30 | https://www.ssga.com/library-content/products/factsheets/etfs/us/factsheet-us-en-gld.pdf ↩
- [S4] Time series momentum | Tobias J. Moskowitz, Yao Hua Ooi, Lasse Heje Pedersen; Elsevier | 2012-05-01 | https://doi.org/10.1016/j.jfineco.2011.11.003 ↩
- [S5] Survivorship Bias in Performance Studies | Stephen J. Brown, William Goetzmann, Roger G. Ibbotson, Stephen A. Ross; Oxford University Press | 2015-05-19 | https://academic.oup.com/rfs/article-abstract/5/4/553/1590264 ↩
Report an error or share feedback
Open a draft with this article’s title and URL. Review the message and recipient before sending.
Open email draftIf no email app opens, copy these details into your usual email service.
Related posts
Quant & Data Research Does AI Actually Help with Quantitative Investing? A Conditional Conclusion and Standards for Validation
Separate AI's research value from investment profitability and define evidence for usefulness through chronological testing, baselines, and reproducible records.
Quant & Data Research Verify the Research, Not the Price: Where LLMs Belong in Quant Research
Distinguish document transformation, validation support, and decision boundaries in LLM-assisted quant research. Check drafts against primary sources and raw data.