Forge Fate
Quant & Data Research

Part 2. What to Check Before Trusting Investment Data: Acquiring and Validating U.S. ETF Data

7 min read

Evidence and scope — Data-contract design

Official sources inform the design of data access, preservation, and quality criteria. No real ETF price rows, file hashes, cleaned dataset, or inspection logs are supplied; this is not a completed data-acquisition or validation report.

The SPY, IEF, and GLD universe and the 2006–2025 period established in Part 1 are research conditions. They do not mean that daily price files for that period have already been obtained and validated. At this stage, the goal is not to calculate returns; it is to create a data contract that determines which data will be accepted into the research.

SPY’s official name changed to State Street SPDR S&P 500 ETF Trust on January 26, 2026, while its SPY ticker and NYSE Arca listing remained unchanged. IEF and GLD can likewise be identified through disclosures from their respective issuers. This is why your records should include more than just the ticker in a filename: keep the product name, provider symbol, and download timestamp as well. [S1] [S2] [S3]

This article does not present actual ETF price rows, file hashes, cleaned datasets, or validation logs. It is therefore not a “validation complete” report. Instead, it lays out a complete framework for the access requirements, preservation practices, and validation criteria to apply before obtaining real files.

Check the Data Contract Before Looking for Free Data

When choosing a market-data source, the first question should not be “Is it free?” It should be: “What am I receiving, under what rights, and how is it defined?” The items below are contract terms to verify for each provider before collecting data; they are not a list confirmed in full by this research.

  • Whether an account or API key is required
  • What is free versus paid, and which historical periods are available
  • Whether personal research, commercial use, and redistribution are each permitted
  • How Close, Adjusted Close, dividends, splits, and volume are defined
  • Which time zone is used and whether regular and extended sessions are included
  • How long raw files may be retained

The only provider terms confirmed here through dated official evidence are Massive’s. Massive’s market-data terms assume an account and personal, non-commercial use by non-professional individuals. Without separate written consent or a third-party agreement, they restrict third-party redistribution and commercial use of the data and data-derived works. [S4] This article does not establish whether free or paid tiers provide long historical data for SPY, IEF, and GLD, how price fields and adjustments are defined, or the terms for retaining raw files. Confirm the provider’s official documentation and the applicable agreement again at the time of collection.

Close, Adjusted Close, NAV, and Execution Prices Are Different

Identical-looking column names do not guarantee identical economic meaning. [S3] [S5] [S6] The table below is a working glossary to use before a provider specification has been selected. Once research begins, replace it with definitions from the received file and the provider’s documentation.

ItemWorking definitionInterpretation boundary
CloseThe price reported by the provider for that dateIt is not automatically the official close, a consolidated-tape value, or any other calculation
Adjusted CloseA provider-calculated adjusted field reflecting dividends, splits, or similar eventsIt is not an actual order execution price or realized execution performance
DividendA cash-distribution recordReinvestment timing, price, and tax assumptions require separate rules
SplitA record of the split ratio and effective dateConfirm the provider’s method for adjusting historical prices
VolumeTrading volume defined by the provider and its session scopeIt does not represent executable size or complete market microstructure

ETF market prices and NAV must also be kept separate. GLD’s NAV is calculated by subtracting estimated expenses and liabilities from the value of gold and other assets, then dividing by shares outstanding. That is not the same concept as the market price observed on an exchange. [S3]

The consolidated tape is not a replica of every order either. The SEC explains that it generally includes trades in listed equities and listed products, but odd-lot trades of fewer than 100 shares are generally not reported and information about orders beyond the best quote is not provided. [S5] The S&P 500 methodology uses primary-exchange closing prices for official end-of-day calculations and consolidated-tape values for the real-time index. This does not define an ETF provider’s closing-price methodology; it illustrates that a price calculation can have different purposes and contexts. [S6]

For that reason, results calculated from adjusted prices should be described as “research results based on adjusted prices as defined by the provider.” They should not be relabeled as the execution prices or realized execution performance of actual orders. SPY’s disclosed total-return calculation also notes that NAV reinvestment may differ from the actual reinvestment price. [S1]

Preserve the Raw Data Unchanged, and Record Decisions Separately

A clean research dataframe is not automatically reproducible. Start by separating the following flow:

immutable raw preservation → normalization → quality checks → report

Store raw files exactly as downloaded. At a minimum, metadata alongside each file should record the provider name, URL, UTC access time, request parameters, symbol, requested period, original filename, SHA-256, license-check date, and the location of the documentation defining the columns. In the normalized version, retain the raw-row identifier, column-name changes, date-parsing method, and the reason for any removal, replacement, or hold. The key rule is not to overwrite raw files with normalized versions.

A hash can be used as an integrity tool to confirm that a file matches a known version. It does not, however, prove the economic accuracy of prices, compliance with an agreement, or whether the values were actually available at that point in time. [S7] Preserving digital objects requires considerations related to both storage media and the object itself, so retain the context in which a file was obtained—not only the file. [S8]

Historical values can change later. The S&P 500 methodology lists corrected closing prices, missing or incorrectly applied corporate actions, and late corporate-action announcements as reasons for recalculation. This is not evidence that a specific SPY, IEF, or GLD data provider follows the same revision policy. It is, however, a practical reason to record the download date and raw-file hash. [S6]

Quality Checks Are Discovery Tools, Not Certificates

Validation is a process for finding structural issues in data. It does not prove the absence of future-information leakage, nor does it prove that real orders could have been executed. At a minimum, check every collected dataset for the following:

  • Date format and ascending dates within each symbol
  • Duplicate symbol-date composite keys
  • Missing required values
  • Finite, positive prices
  • Valid volume ranges based on the volume definition, session scope, and permitted format
  • Date gaps compared with the trading calendar
  • Rows that appear to precede an instrument’s listing
  • Candidates for unusually large upward or downward moves

For date gaps, first distinguish market holidays from missing records on trading days. If you have not yet obtained the relevant trading calendar and each fund’s tradable start date, you cannot conclude that there are “no missing values.” Large moves are not confirmed errors either. Keep them as investigation candidates, then revisit the raw rows, corporate actions, adjustment method, and session definition.

The code below is not actual ETF data. It is an example using a complete, deliberately flawed synthetic sample. The results shown here were not executed for this article.

python
import numpy as npimport pandas as pd# Synthetic input only; not market data.raw = pd.DataFrame(    {        "ticker": ["SPY", "SPY", "SPY", "IEF"],        "date": ["2025-01-03", "2025-01-02", "2025-01-02", "2025-01-02"],        "close": [600.0, np.nan, 599.0, -101.0],        "volume": [100, 100, -1, 50.5],    })data = raw.assign(date=pd.to_datetime(raw["date"], errors="coerce"))checks = {    "bad_date": data["date"].isna(),    "not_ascending": data.groupby("ticker")["date"].diff().lt(pd.Timedelta(0)),    "duplicate_key": data.duplicated(["ticker", "date"], keep=False),    "missing_required": data[["ticker", "date", "close", "volume"]].isna().any(axis=1),    "bad_price": data["close"].notna()    & (~np.isfinite(data["close"]) | data["close"].le(0)),    # Synthetic input assumes non-negative integer volume.    "bad_volume": data["volume"].lt(0) | data["volume"].mod(1).ne(0),}

For this input, the expected findings are one SPY date in reverse order, two duplicate SPY-2025-01-02 records, one missing close value, one negative price, one negative volume, and one fractional volume value. For real collected data, do not record only the count from each check. Also retain the raw identifier for each problematic row, the handling decision, and the reason for that decision.

Do Not Fill Missing Data for Convenience

A missing row should not automatically be filled with the prior day’s value. The appropriate treatment depends on whether the date was a market holiday, a trading day for which the provider omitted a row, or a row with only certain fields missing.

Whether you remove, replace, or hold a value, record the raw value, replacement value, rule applied, time of application, and raw-row identifier separately. In particular, using adjusted and unadjusted prices interchangeably to fill gaps can blur the meaning of later return calculations. When a decision remains unresolved, it is better to keep the row under investigation than force it into an artificially clean dataset.

The candidate universe also has limits. A teaching design fixed to the three currently identifiable instruments—SPY, IEF, and GLD—does not reconstruct the full set of assets that could have been invested in at each historical date. Research using a mutual-fund sample that substantially addressed survivorship bias found a relationship between poor performance and the probability of disappearance. That study does not estimate the magnitude of survivorship bias or performance for these three ETFs. [S9]

Completion Criteria for This Stage

The output of this stage is not a return chart. After obtaining actual files, retain the price definition, access terms, usage rights, raw-data metadata, normalization rules, and validation results as one package. If permission to publish raw data has not been confirmed, do not share the CSV itself; share only the process, metadata structure, hashes, validation rules, and publishable synthetic examples. [S4] [S7]

Passing quality checks does not establish that there is no future-information leakage. It also does not guarantee real execution, including order availability, quotes, spreads, or market impact. Part 3 will separate the signal date, label-calculation date, and order time to test whether only information actually knowable at that time was used.

This article documents a data-design process for educational and research purposes. It is not a recommendation to buy or sell any ETF, nor a guarantee of investment performance.

Sources

Report an error or share feedback

Open a draft with this article’s title and URL. Review the message and recipient before sending.

To: [email protected]

Open email draft

If no email app opens, copy these details into your usual email service.

Contact information