Automatically Saving Experiment Records and Reports: A Minimal Tool for Preserving Successes and Failures
This is Part 7 of “Building Your Own Quant Research System.” In the previous installment, we defined the conditions for comparing a model with baseline strategies. This time, we examine a small tool that processes synthetic inputs and preserves what went in, what came out, and what went wrong.
First, let’s be clear about what was run. According to the execution records supplied by the project, actual price data for SPY, IEF, and GLD has still not been obtained, and the model training and backtests described in Part 6 have not been run. What was actually executed this time was a tool that converts synthetic probabilities into portfolio weights and saves execution records and reports.
The supplied records report checks PASS on 2026-09-21, using the standard library on a host running Python 3.14.5. The execution details below come from the project records included in this request. They are not the results of an independent inspection or rerun of the preserved code and folders. The official Python documentation serves a separate purpose: it provides external references for the behavior of the APIs used.
Reading a Successful Run and a Failed Run Side by Side
Comparing the two cases in the supplied project records shows what evidence this tool preserves.
| Item | Successful run | Failed run |
|---|---|---|
| Run ID | d711f515-8380-4bca-9112-23ef651f6bf6 | 330fd103-4ddc-41ce-9620-7ae15bae1003 |
| Input | SYN_A=0.60, SYN_B=0.45, SYN_C=0.70 | Separate input: A=1.20 |
| Threshold and validation | Threshold of 0.55 applied | Value outside [0, 1] |
| Resulting weights | A 0.5, B 0, C 0.5, cash 0 | No result |
| Status and exception | Weight conversion succeeded | FAILED, ValueError |
| Failure reason | Not applicable | probabilities must be nonempty and within [0, 1] |
result.json | Created | Not created |
These probabilities are synthetic inputs, not model outputs or ETF forecasts. A, B, and C are shorthand for SYN_A, SYN_B, and SYN_C, respectively. The exact JSON key names have not been verified because the original files were not inspected. Cash 0 also means a cash weight, not a cash return or a strategy return.
Read the successful result and the failure reason together. The failed run produced no calculated result. Its report should therefore show the failure status and reason rather than fill the gap with a return of 0. The point of this example is to trace inputs through to either results or errors, then preserve that outcome in files. It does not test whether the threshold has investment value.
Create One Folder per Run
The supplied records describe the following relative folder structure. The exact parent path of runs has not been confirmed.
runs/├── d711f515-8380-4bca-9112-23ef651f6bf6/│ ├── inputs.json│ ├── config.json│ ├── record.json│ ├── result.json│ └── report.md└── 330fd103-4ddc-41ce-9620-7ae15bae1003/ ├── inputs.json ├── config.json ├── record.json └── report.mdinputs.json preserves the inputs, while config.json preserves settings such as the threshold. record.json connects the run identifier with its environment, status, and errors. result.json is created only when the run succeeds. report.md is generated by reading the saved record and result back from disk. These are the storage rules supplied by the project.
Each run folder uses an ID generated by uuid4() and is created with mkdir(exist_ok=False). uuid4() generates a random UUID, and this mkdir setting raises FileExistsError if the target folder already exists.[S4][S1]
Separating runs into folders helps prevent a rerun from accidentally overwriting existing artifacts. This folder-creation rule does not prevent files from being modified later.[S1] Using a UUID also does not justify claiming that collisions are impossible or that the files are protected against tampering.[S4]
Core Code: Recording Identifiers, Bytes, and Timestamps
The examples below were written to illustrate the supplied storage rules. They are not excerpts from the preserved script or the code that received checks PASS. They demonstrate a few small operations needed for recordkeeping rather than provide a complete runner.
First, here is the run-folder creation step. This function assumes that the parent runs_dir already exists.
import uuiddef create_run_dir(runs_dir): run_id = str(uuid.uuid4()) run_dir = runs_dir / run_id run_dir.mkdir(exist_ok=False) return run_id, run_dirGenerating an identifier and detecting a folder collision serve separate purposes. The UUID supplies a name; mkdir creates the folder while checking whether that name already exists.[S4][S1]
Next, we prepare the JSON bytes to save and calculate their hash.
import hashlibimport jsondef json_bytes(value): text = json.dumps( value, sort_keys=True, ensure_ascii=False, indent=2 ) return (text + "\n").encode("utf-8")def save_inputs_and_config(run_dir, inputs, config, code_path): input_bytes = json_bytes(inputs) config_bytes = json_bytes(config) (run_dir / "inputs.json").write_bytes(input_bytes) (run_dir / "config.json").write_bytes(config_bytes) return { "input_hash": hashlib.sha256(input_bytes).hexdigest(), "config_hash": hashlib.sha256(config_bytes).hexdigest(), "code_hash": hashlib.sha256(code_path.read_bytes()).hexdigest(), }sort_keys=True sorts the keys. ensure_ascii=False outputs non-ASCII characters directly, except where escaping is required. indent=2 uses two-space indentation.[S2] Adding a final newline and encoding the text as UTF-8 are additional project rules. The example uses the same bytes for both storage and hashing.
sha256 calculates a hash from bytes, and hexdigest() returns it as a hexadecimal string.[S3] Under the supplied rules, inputs and settings are hashed using the specified JSON representation, while the code is hashed directly from its file bytes.
Hashes should be used to identify and compare content. This calculation alone does not establish the accuracy or provenance of source data, preserve the execution environment, or demonstrate complete research reproducibility.[S3] A separate security mechanism to address situations where both the content and its hash are changed is also outside this implementation’s scope.
Timestamps can be represented as follows.
from datetime import datetime, timezonedef utc_now(): return datetime.now(timezone.utc).isoformat()This code produces a timezone-aware UTC timestamp as an ISO 8601 string.[S5] It should be called separately at the start and end of a run. Using UTC does not, by itself, guarantee that the host’s clock is accurate.[S5]
According to the supplied records, record.json contains the following fields.
| Group | Recorded information |
|---|---|
| Identity and timing | run_id, started_at, finished_at |
| Environment and content identification | Python version; three hashes covering inputs, settings, and code |
| Status and errors | status, error_type, reason |
The actual hash strings and start and finish timestamps were not supplied, so no values are invented here. The supplied records identify the execution environment as 3.14.5, while the official documentation consulted was for 3.14.7. The official release pages confirm that both versions exist, but they do not establish which version ran on the project’s host.[S6][S7]
Generate Reports from Saved Files
According to the supplied project records, the report is generated by rereading record.json and, for a successful run, result.json. Calculated values are not copied separately from memory into the report.
Saved record.json ─────────────┐ ├── report.mdSaved result.json (on success) ┘This is what the one-way flow of record → result file → report means. It does not require writing each of the three files exactly once in a fixed sequence. The design principle is that the report’s status and numbers come from saved artifacts.
The key is to avoid a workflow in which someone retypes the successful run’s weight of 0.5 into the report. Saved JSON can be read with json.load or json.loads.[S2] The report-generation step should focus on presenting the results it reads.
A failure report uses the status, exception, and reason stored in the record, without a result file. For this failed run, its body would show FAILED, ValueError, and the reason describing the range violation. There is no need for an empty performance table or an arbitrary number.
What checks PASS Covered
The supplied project execution records state that the following assert checks passed.
- The successful run’s weights matched the expected values.
- The failed run had status
FAILEDand exception typeValueError. - The failed run’s folder contained no
result.json. - The successful and failed runs used different folders.
- Every file in the first folder remained byte-for-byte unchanged after the second run.
- The success report included the same content as the saved result.
- The failure report included the failure reason.
The fifth check is evidence that this particular rerun did not alter the first run’s artifacts. It does not mean that later manual edits or changes by someone else were prevented.
The supplied preservation locations are data/manual-revisions/quant-research-07-records.py and the runs folder. These are locations reported by the project, not a list of files inspected directly for this article. The passing checks should also be understood as covering the two supplied cases.
What the Tool Does Not Yet Handle
According to the supplied records, the current implementation handles only input-validation ValueError exceptions. It does not handle disk errors or forced termination, implement atomic writes or environment locking, preserve snapshots of actual market data, or provide complete research reproducibility. An interrupted run may also leave its status as RUNNING.
The statement that the tool “preserves failures” therefore applies to the input-validation failure handled here. It cannot be extended into a claim that complete records survive every kind of failure.
Successful JSON serialization must also be distinguished from successful probability validation. Python’s default json settings can allow nonstandard numeric representations such as NaN and infinity.[S2] The supplied records contain no test results for NaN, infinity, strings, or the range of the threshold itself.
The project’s policy is to include only explicitly allowed fields in public records and exclude API keys, tokens, and private source data. The inputs in this example are synthetic and contain no secrets. This policy does not mean that automatic secret detection has been implemented.
What to Compare in Your Next Run
Start the exercise with two small synthetic cases: one successful run and one failed run using an out-of-range value such as A=1.20. Compare the folders’ IDs, statuses, error reasons, and presence or absence of a result file. Then generate reports from the saved records and results. After rerunning, also compare the earlier folder’s bytes and check that the report content matches the saved result.
For future experiments using actual market data, additional fields should capture feature and label timestamps, data periods and identifiers, split boundaries, retraining schedules, and execution and cost settings. These are proposed extensions, not fields already implemented in the current tool.
Part 8, the final installment of the core series, will review the artifacts actually obtained, what can be rerun, observations about AI’s contribution, and the unfinished work involving real market data. It will close the series based on the evidence available, without promising completed empirical market validation or measured productivity gains.
This tool and exercise are for education and research. They are not investment recommendations or guarantees of returns.
Sources
- [S1] pathlib — Object-oriented filesystem paths | Python Software Foundation | 2026-09-21 | https://docs.python.org/3.14/library/pathlib.html ↩
- [S2] json — JSON encoder and decoder | Python Software Foundation | 2026-09-21 | https://docs.python.org/3.14/library/json.html ↩
- [S3] hashlib — Secure hashes and message digests | Python Software Foundation | 2026-09-21 | https://docs.python.org/3.14/library/hashlib.html ↩
- [S4] uuid — UUID objects according to RFC 9562 | Python Software Foundation | 2026-09-21 | https://docs.python.org/3.14/library/uuid.html ↩
- [S5] datetime — Basic date and time types | Python Software Foundation | 2026-09-21 | https://docs.python.org/3.14/library/datetime.html ↩
- [S6] Python Release Python 3.14.5 | Python.org | Python Software Foundation | 2026-05-10 | https://www.python.org/downloads/release/python-3145/ ↩
- [S7] Python Release Python 3.14.7 | Python.org | Python Software Foundation | 2026-08-05 | https://www.python.org/downloads/release/python-3147/ ↩
Report an error or share feedback
Open a draft with this article’s title and URL. Review the message and recipient before sending.
Open email draftIf no email app opens, copy these details into your usual email service.
Related posts
Quant & Data Research Does AI Actually Help with Quantitative Investing? A Conditional Conclusion and Standards for Validation
Separate AI's research value from investment profitability and define evidence for usefulness through chronological testing, baselines, and reproducible records.
Quant & Data Research 11. A Safe Way to Do Quant Research with LLMs: From Questions to Validation and Research Records
Organize LLM-assisted quant research around questions, evidence, code checks, timing, and experiment records, with human review and documented failures.