Forge Fate
Quant & Data Research

Why shift(1) Is Not Enough: A Time Contract for Signals, Execution, and Labels

7 min read

Evidence and scope — Synthetic unit tests and supplied run records

The PASS results are project-supplied records of four synthetic tests using Python 3.14.5 and unittest. They are neither independent reruns nor proof that the entire system is leakage-free. Real-price, order, and execution-data validation remains outstanding.

The data-quality checks in Part 2 asked whether an incoming price file was structurally suitable for research. But even after resolving duplicate dates and missing values, we still need to verify separately that tomorrow’s information has not slipped into today’s decision.

This article builds synthetic unit tests for an individual research project considering SPY, IEF, and GLD before any actual CSV files are available. There are not yet real price files, order or execution records, a trading-day calendar, or a completed backtest. The results here therefore do not represent ETF performance or prove that a strategy works. They are educational and research-oriented checks that turn timing rules into code.

Time Is Not One Date but Several Contracts

In financial time series, information leakage occurs when a simulated decision at a particular time includes information that became available only afterward. A one-row mismatch between features and labels, preprocessing statistics calculated over the full sample, and assigning information to a date before it was actually published are different routes for leakage. [S1]

This project keeps the next-trading-day close execution rule established in Part 1. We finalize a signal after observing data through the close on day t, then assume the order is executed at the close on t+1. The position’s first close-to-close return is attributed only to the interval ending from t+1 to t+2.

TimeWhat happensInformation not yet available
After the close on tFinalize the signal using observable featuresPrices, returns, and labels after t+1
Close on t+1Assume the order is executedReturns after t+2
Close on t+2The first close-to-close return interval after execution endsAny value that feeds that calculated return back into the signal at t
Label end dateThe outcome used for training becomes finalLabels ending after the signal date assigned to earlier training rows

shift(1) can correct one kind of one-row alignment error. It does not, by itself, define when a signal is finalized, when next-day execution occurs, when return attribution starts, how market holidays are handled, or what window a label observes. This is an interpretation of leakage types applied to this project’s execution rules. [S1]

The key is not to store only a “signal date.” Each row should distinguish, at minimum, the time the signal is finalized, the order execution time, the return-attribution interval, the label end time, and the feature availability time.

Four Checks to Build Before Using Real Data

According to the project’s own verification record, the following four synthetic checks passed on 2026-09-15 using local Python 3.14.5 and the standard unittest library. This is neither an independently reproduced result nor proof that the full system is free of leakage. Python 3.14.5 is an officially released maintenance version, [S2] and standard unittest provides assertRaises() for checking expected exceptions. [S3]

python
import unittestfrom datetime import datedef make_signal(prices):    # Use only the current and prior price.    return [0] + [int(prices[i] > prices[i - 1]) for i in range(1, len(prices))]def held_weight(signal):    # Signal at t is executed at t+1 and earns from t+1 to t+2.    return [0 if i < 2 else signal[i - 2] for i in range(len(signal))]def allowed_training_rows(label_end, signal_date):    # Exclude missing labels and require a strictly earlier end date.    return [        i for i, end in enumerate(label_end)        if end is not None and date.fromisoformat(end) < signal_date    ]class TimingTests(unittest.TestCase):    def test_future_price_does_not_change_prior_signals(self):        prices = [100, 110, 105, 120]        changed = [100, 110, 105, 999]        self.assertEqual(make_signal(prices), [0, 1, 0, 1])        self.assertEqual(make_signal(prices)[:3], make_signal(changed)[:3])    def test_execution_delay_and_wrong_weights(self):        signal = [0, 1, 0, 1]        held = held_weight(signal)        self.assertEqual(held, [0, 0, 0, 1])        with self.assertRaises(AssertionError):            self.assertEqual(held, signal)        with self.assertRaises(AssertionError):            self.assertEqual(held, [0, 0, 1, 0])    def test_label_end_is_strictly_before_signal_date(self):        label_end = ["2025-01-02", "2025-01-03", "2025-01-06", None]        signal_date = date.fromisoformat("2025-01-06")        self.assertEqual(allowed_training_rows(label_end, signal_date), [0, 1])    def test_fit_uses_training_window_only(self):        train = [100, 110]        mixed_with_future = [100, 110, 120, 999]        self.assertEqual(sum(train) / len(train), 105)        self.assertNotEqual(            sum(train) / len(train),            sum(mixed_with_future) / len(mixed_with_future),        )

The arrays are short, but the purpose is clear: the correct implementation should pass, while an implementation that deliberately violates the time contract should fail. These four checks are not a reduced substitute for a monthly 12-month strategy or a real backtest.

Check 1: Does Changing the Final Price Leave Earlier Signals Unchanged?

The first check changes only the final value in prices=[100,110,105,120] to 999. The original signal is [0,1,0,1], and the first three signals must remain unchanged after the edit.

This check has a narrow scope. It establishes only that this small make_signal() function does not read the final price when producing signals for the first three points in time. It says nothing about whether signals are valid for actual SPY, IEF, or GLD prices, economically meaningful, or tradable in practice. It directly targets the basic rule that future observations must not enter past decisions. [S1]

These invariance checks are especially useful after refactoring. Even after vectorizing indicators or changing the order of dataframe merges, they preserve the minimum rule that editing a future row must not alter earlier signals.

Check 2: Are Signal, Execution, and Return Attribution Kept Off the Same Row?

For signal=[0,1,0,1], the expected held exposure under the time contract is held=[0,0,0,1]. The signal is finalized at t, executed at the close on t+1, and begins receiving returns only for the interval ending from t+1 to t+2.

ComparisonValueTest expectation
Held exposure following the time contract[0,0,0,1]Matches the correct implementation
Using the same-row signal as held exposure[0,1,0,1]Detect an AssertionError
Using held exposure shifted by only one row[0,0,1,0]Detect an AssertionError

The two mistakes look similar but are different. Using the same-row signal directly as held exposure can make a signal learned at the close on t appear to have already earned that day’s return. A value shifted by only one row, by contrast, collapses the execution time and the start of the first return-attribution interval into a single shift.

assertRaises() explicitly verifies that an incorrect implementation actually fails. [S3] A successful test here means that a comparison designed to catch an error produced the intended exception; it does not mean that actual ETF execution prices, spreads, or market impact have been modeled.

Check 3: Are Unfinished Labels Excluded from Historical Training?

The third check uses the following input.

python
label_end = ["2025-01-02", "2025-01-03", "2025-01-06", None]signal_date = "2025-01-06"

This article adopts the conservative rule end < signal_date. Therefore, only indices [0,1], whose end dates are strictly earlier than the signal date and are not missing, are eligible for training. The third label, which ends on 2025-01-06, is excluded because it ends on the same date, and the fourth label is excluded because it is None.

This boundary is not the only correct answer. The actual rule should be chosen after documenting exactly which return interval or event a label observes and when features became genuinely available. In particular, a value released after the market close and a value available intraday can have different availability times even when they share a date. Assigning information earlier than its real publication time can itself create a leakage path. [S1]

For actual data, it is therefore safer to record separate times for separate roles rather than relying on a single date column: feature_available_at, signal_at, execution_at, label_start, and label_end.

Check 4: Is the Preprocessing Mean Fitted Only Within the Training Window?

The final check confirms that the mean of training values [100,110] is 105. It then checks that the mean changes when future values 120 and 999 are mixed in, producing [100,110,120,999].

This example does not validate a complete scaler implementation. It also does not guarantee the behavior of a particular library or every preprocessing step. Instead, it illustrates the minimum rule that statistics estimated through fit—such as means, standard deviations, and missing-value imputation benchmarks—must be calculated only within the relevant training window. Normalization statistics calculated over the full sample can carry future information into past training. [S1]

When extending this into actual walk-forward research, fit a new preprocessor for each training period and apply only transformations to validation and test periods. As the training window moves, the mean must be recalculated with that window.

What Passed, and What Still Needs Checking

Only four synthetic checks have an execution record in this article.

  • Invariance of past signals after changing a future price
  • The delay from signal to execution and from execution to return attribution
  • A strict eligibility condition based on label end dates
  • An example confirming that values outside the training window are not included in the mean

The following items still need additional checks once actual CSV files and a trading-day calendar are available.

  • How signals, held exposure, and returns are handled in the first and last rows
  • Whether a missing date is a market holiday or missing trading-day data
  • How SPY, IEF, and GLD are aligned by actual trading date rather than row number
  • How to split a date-based panel after calculating features separately for each ETF
  • Whether label end times and feature availability times truly do not overlap
  • Whether preprocessing fit uses only values inside every training window

In particular, the three ETFs should not simply be forced into arrays of equal length. Their actual trading dates, observation availability, and reasons for missing values must be reviewed before aligning them by date. Because real CSV files, calendars, and execution records are not yet available, we cannot verify whether next-trading-day-close execution was feasible in the market or whether the return calculation reflects market reality.

Backtesting must address limited data, repeated exploration, overfitting, and data-mining risk, and backtest results should not be treated as identical to live performance. [S4] Four passing checks do not replace that research protocol.

Next Step: Write the Time Contract for Every Row

The next action is not to rush into plotting returns. First, write these five items in the research notes or data contract.

  • Signal finalization time
  • Order execution time
  • Return-attribution interval
  • Label end time
  • Feature availability time

Then, do not test only expected correct outputs. Add tests in which deliberately incorrect implementations fail: same-row execution, overlapping labels, or preprocessing that mixes in future values. The moment code that should fail passes silently is the moment a timing error has been missed.

The baseline backtest in Part 4 can begin only after real SPY, IEF, and GLD CSV files, price-column definitions, trading-day alignment, cost and execution assumptions, and the execution environment and logs are ready. The synthetic tests in this article do not replace that preparation. They are a time-validation design for educational and research purposes, not a recommendation to buy or sell any ETF or instructions for placing live orders.

Sources

Report an error or share feedback

Open a draft with this article’s title and URL. Review the message and recipient before sending.

To: [email protected]

Open email draft

If no email app opens, copy these details into your usual email service.

Contact information