What the Reruns Confirmed: Where the Quant Research System Stands and AI’s Role
The previous article introduced a tool for recording successful and failed runs with synthetic inputs. In this eighth and final installment of “Building Your Own Quant Research System,” we compare those records with runs repeated the following day.
This series follows a Python research project that connects research questions, data, checks, and execution records for U.S. ETFs. According to the project’s status summary, actual source data for SPY, IEF, and GLD has not yet been obtained. Model training, monthly backtesting, and empirical comparisons of strategy performance also remain unfinished. Completing the series is different from completing an investment system.
The status overview below comes from the summary supplied by the project. The rerun details come from records that a verifier working on the host machine checked directly and passed along. The original code and verification files were neither inspected directly nor rerun during the writing of this article.
What the Eight Parts Produced
The table separates completed work from unfinished or unverified work, based on the supplied project records. Detailed code and execution logs for Parts 1–6 were not provided for this review, so “completed” refers only to the scope described in those records.
| Part | Completed according to the supplied records | Unfinished or unverified |
|---|---|---|
| Parts 1–2 | Defined the research questions, U.S. market scope, and limits on obtaining data | Obtaining actual source data for SPY, IEF, and GLD |
| Part 3 | Tested for look-ahead leakage using synthetic data | Verifying that all actual data respects information availability at each point in time |
| Parts 4–5 | Tested holdings and cost accounting using synthetic data | Monthly backtesting with actual execution and cost assumptions |
| Part 6 | Designed a comparison between the model and benchmark strategies | Model training and comparison of the three strategies’ performance |
| Part 7 | Built a tool to record successful and failed runs with synthetic inputs | Preserving records across all failure conditions |
| Part 8 | Reran the cases the next day on the same host, in new folders, and compared specified items | Independent reproduction in a new environment, on another OS, or by another person |
Computational reproducibility is described as obtaining consistent computational results using the same inputs, computational steps, methods, code, and conditions of analysis. [S10] In this review, the synthetic checks, comparison design, and record comparisons each count as distinct deliverables. Testing the research questions against actual market data remains unfinished.
What Matched in the Next Day’s Reruns
According to the host verifier’s report, the reruns took place on 2026-09-22, using Python 3.14.5 on the same host as the previous day. The existing Part 7 code was loaded unchanged with runpy. The previous inputs.json files were read, and the successful and failed cases were each rerun in a new UUID folder. The reported check result was checks PASS.
runpy executes code in the current process. [S3] Running the code in new folders therefore does not establish reproduction in a new environment, on another OS, or independently by another person.
| Comparison item | Successful case | Failed case |
|---|---|---|
| Previous run ID | d711f515-8380-4bca-9112-23ef651f6bf6 | 330fd103-4ddc-41ce-9620-7ae15bae1003 |
| New run ID | 70cb18b9-370d-4aa6-ba77-f9d2e49c3f17 | 1cddaea1-88a4-4340-9fae-bd846e05a188 |
| Synthetic input | SYN_A=0.60, SYN_B=0.45, SYN_C=0.70 | SYN_A=1.20 |
| Configuration or input validation | Threshold of 0.55 | Violates the allowed probability range of [0, 1] |
| Resulting weights | SYN_A: 0.5, SYN_B: 0, SYN_C: 0.5, cash_weight=0 | No calculated result |
| Status and error | Successful; matched the previous status | FAILED, ValueError |
| Failure reason | Not applicable | probabilities must be nonempty and within [0, 1] |
result.json | Present and byte-for-byte identical to the previous file | Absent from both the previous and new runs |
These probabilities are synthetic inputs, not outputs from a trained model or predictions for ETFs. cash_weight=0 represents the cash allocation. The table describes the supplied inputs and results; it does not reconstruct the full structure of the original JSON.
For the successful case, the comparison checked whether the same weights were recorded. For the failed case, it checked whether the same status and error reason were recorded and whether the result file was absent. Both cases were compared with their respective previous records.
Comparing Results Separately from Execution Metadata
The supplied verification record distinguishes the items that matched from those treated separately:
- The SHA-256 hashes of the inputs, configuration, and code matched the previous run, as did the Python version and the
status,error_type, andreasonvalues. - The presence or absence of
inputs.json,config.json,result.json, andreport.mdmatched. Every file that existed was byte-for-byte identical to its previous counterpart. - All files in the existing run folders remained unchanged before and after the reruns.
- Each new
run_iddiffered from the previous value.started_atandfinished_atwere excluded from the comparison.
The criterion here was exact byte equality for the specified files, not numerical agreement within a tolerance. This does not mean that every new record file, including record.json, was identical in full to its predecessor. To interpret the matches correctly, computational outputs must be distinguished from run identifiers and timestamps.
The following is a newly written, unexecuted example illustrating the file comparison rules. It is neither an excerpt from the preserved verifier nor the code that received checks PASS. Path.read_bytes() returns a file’s contents as bytes. [S4]
from pathlib import Pathdef saved_bytes(path): try: return path.read_bytes() except FileNotFoundError: return Nonedef compare_files(previous_dir, replay_dir): names = ("inputs.json", "config.json", "result.json", "report.md") for name in names: before = saved_bytes(Path(previous_dir) / name) after = saved_bytes(Path(replay_dir) / name) if before != after: raise AssertionError(f"{name}: presence or bytes differ")This example only compares file absence on either side and differences in file bytes. Other read errors propagate. It does not check run IDs, metadata expected to remain constant, or whether the existing folders stayed unchanged during the reruns.
The example uses conditional statements to raise explicit exceptions. Python removes assert statements when run with the -O optimization option. [S7] The full command line and optimization settings for the actual verification were not available, so the reported passing result cannot be treated as confirmation of all execution conditions.
SHA-256 computes a hash from byte input. [S5] Taken together, that function and the definition of reproducibility provide no basis for treating matching hashes alone as proof of source-data accuracy, economic validity, or complete reproducibility. [S5][S10] Whether incorrect data was preserved unchanged and whether the data is correct are separate questions.
The project supplied the following local record locations:
- Execution code:
data/manual-revisions/quant-research-07-records.py - Verification tool name:
quant-research-08-replay.py - Preserved verification record:
data/manual-revisions/quant-research-08-runs/c9e4c439-40ab-4d17-9f49-f6c851f60728/verification.json
The parent path for the verification tool was not provided. These references identify local files; they are not public download links. They also do not indicate that the complete code has been publicly released.
What AI Did—and What Was Not Measured
According to the supplied conversation summary, AI structured the article brief, wrote synthetic verification code and assert statements, and incorporated execution records into the writing requirements. The user determined the U.S. market scope and how the series would progress. The summary also reported that Parts 1–7 had been marked approved in the system. That approval status, however, is not evidence of human review of every sentence or numerical claim.
The project has no measurements of time spent, no control experiment without AI, and no complete log of accepted and discarded work. It therefore cannot quantify how much time AI saved or how much it improved accuracy. The supplied task history documents the work AI performed; the size of its effect was not measured.
Measurement is also a central issue in external research. METR stated that it was changing the design of a follow-up developer productivity experiment because of problems involving participant and task selection and the measurement of working time. [S11] A separate self-report survey discussed differences between perceived speed and the value of output, along with response consistency and the possibility of overstatement. [S12] These studies provide context for the difficulty of measurement. They do not measure AI’s effect on this project.
Likewise, the report that AI-written code passed the checks applies only to the synthetic cases and comparison items tested. It does not establish a model’s predictive power, a strategy’s performance, or the validity of a research conclusion.
Priorities for Moving into Actual ETF Research
The remaining work can proceed in the following order. This is a sequence for testing the research questions once actual data is available, not a list of completed tasks.
- Obtain source data and processing specifications: Secure usable U.S. ETF source data, the necessary licenses and usage rights, and specifications for handling dividends and splits.
- Connect timing checks with accounting: Integrate checks on when actual data became available with a common framework for costs and trade execution accounting.
- Fix the comparison rules, then evaluate: Set the data splits and comparison rules before evaluating the three strategies, then use the resulting evidence to assess their relative performance.
- Strengthen environment management and failure records: Specify how to recreate the environment and verify what records survive interruptions and write failures.
According to the supplied implementation status, the only exception currently handled is the ValueError raised during input validation. Handling disk errors and forced termination, atomic writes, and environment locking have not been implemented.
For environment management, venv can isolate packages, but its environments are generally not portable and need to be recreated at the destination. [S8] For file writes, os.replace provides an atomic rename when successful, but it may fail across different filesystems. [S9] Neither capability should be taken to mean that the entire environment has been fixed or that recovery of multiple files as a unit has been implemented.
What Can Now Be Verified
The supplied records establish a specific milestone: two synthetic cases were rerun on the same host, specified files, statuses, and errors were compared, and the existing folders were checked for changes. Actual ETF performance and reproduction in an independent environment remain unverified.
In your own experiments, rerun both successful and failed cases in new folders. Compare results produced from the same inputs and settings, while treating identifiers and timestamps that change with each run separately. The final takeaway from this series is the habit of recording what matched, what was not compared, and what remains unfinished.
The tools and exercises in this article are for education and research. They are not investment recommendations or guarantees of returns.
Sources
- [S3] runpy — Locating and executing Python modules — Python 3.14.7 documentation | Python Software Foundation | 2026-09-21 | https://docs.python.org/3.14/library/runpy.html ↩
- [S4] pathlib — Object-oriented filesystem paths — Python 3.14.7 documentation | Python Software Foundation | 2026-09-22 | https://docs.python.org/3.14/library/pathlib.html ↩
- [S5] hashlib — Secure hashes and message digests — Python 3.14.7 documentation | Python Software Foundation | 2026-09-21 | https://docs.python.org/3.14/library/hashlib.html ↩
- [S7] 7. Simple statements — Python 3.14.7 documentation | Python Software Foundation | 2026-09-21 | https://docs.python.org/3.14/reference/simple_stmts.html ↩
- [S8] venv — Creation of virtual environments — Python 3.14.7 documentation | Python Software Foundation | 2026-09-21 | https://docs.python.org/3.14/library/venv.html ↩
- [S9] os — Miscellaneous operating system interfaces — Python 3.14.7 documentation | Python Software Foundation | 2026-09-22 | https://docs.python.org/3.14/library/os.html ↩
- [S10] Reproducibility and Replicability in Science | National Academies of Sciences, Engineering, and Medicine; The National Academies Press | 2019 | https://www.nationalacademies.org/read/25303/chapter/2 ↩
- [S11] We are Changing our Developer Productivity Experiment Design | METR; Joel Becker, Nate Rush, Tom Cunningham, David Rein, Khalid Mahamud | 2026-02-24 | https://metr.org/blog/2026-02-24-uplift-update/ ↩
- [S12] Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity | METR; Joel Becker | 2026-05-11 | https://metr.org/blog/2026-05-11-ai-usage-survey/ ↩
Report an error or share feedback
Open a draft with this article’s title and URL. Review the message and recipient before sending.
Open email draftIf no email app opens, copy these details into your usual email service.
Related posts
Quant & Data Research Does AI Actually Help with Quantitative Investing? A Conditional Conclusion and Standards for Validation
Separate AI's research value from investment profitability and define evidence for usefulness through chronological testing, baselines, and reproducible records.
Quant & Data Research 11. A Safe Way to Do Quant Research with LLMs: From Questions to Validation and Research Records
Organize LLM-assisted quant research around questions, evidence, code checks, timing, and experiment records, with human review and documented failures.