Forge Fate
Quant & Data Research

11. A Safe Way to Do Quant Research with LLMs: From Questions to Validation and Research Records

9 min read

Evidence and scope — Research procedure and record keeping

The article describes question setting, checks against source data, code validation, and research records. LLM proposals are distinguished from validated results. Live-order implementation and measured productivity gains from this procedure are outside its scope.

In Part 10, we set boundaries for using LLMs not as substitutes for investment judgment, but as tools for drafting and assisting validation against primary sources and raw data. This article turns those boundaries into a practical research workflow.

This is an educational and research-oriented guide. It is not investment advice or a recommendation to buy or sell any asset. It also does not cover live-order implementation.

In this workflow, you ask an LLM to draft research items, falsification conditions, and candidate code tests, then review and finalize them yourself. A suggestion is not evidence or a validation result. Documenting the AI’s role, human oversight, validation outcomes, and remaining risks aligns with core principles of AI risk management. [S1] [S2]

The process rests on three pillars:

Question and evidence → code and time validation → execution and records

The LLM produces drafts; the researcher confirms facts, figures, code, and sources.

Fix the Full Research Workflow First

Safe research is less about finding a promising result quickly than about defining the inputs and passing conditions at every stage.

StagePrimary ownerInputOutputCondition to proceedIf validation fails
Define the questionHumanObject of observation, period, constraintsComplete question and falsification conditionsTiming, universe, costs, and decision criteria are stated in writingReturn to question definition
Propose research itemsLLM, reviewed by humanQuestion and constraintsList of primary sources, variables, and falsification candidatesHuman finalizes the verification listRevise the question or research items
Obtain materialsHumanVerification listSource locations, raw data, access date, data descriptionSource, period, and price definition have been cross-checkedReturn to material collection
Draft codeLLM, reviewed by humanConfirmed rules and data dictionaryCalculation functions and test draftsHuman can read and confirm the mapping between rules and codeRevise rules or code
Validate timing and small samplesHumanCode and small examplesTest results and timing checksSignal, label, and trade timing are separatedReturn to code or data
Run the backtestHumanFixed rules, period, and costsResults and intermediate outputsRun followed the pre-specified rulesRevisit execution settings or an earlier stage
Review anomaliesLLM proposes candidates; human decidesResult tables, logs, and selected codeCandidates for recheckingHuman cross-checks against raw data, code, and sourcesReturn to the causal stage
Record and defer conclusionsHumanRuns including successes, failures, and stopsResearch logLimitations and change history are documentedRegister the change as a new experiment

This structure prevents fluent LLM explanations from skipping validation steps. In particular, temporal leakage, preprocessing leakage, and dependence between training and test data can produce overly optimistic conclusions and failed reproduction attempts, so they require separate checks. [S3]

A Hypothetical Example: Ask “Is the Timing Rule Correct?” Before “Will It Make Money?”

The following is a hypothetical educational example, not a real ETF study, market-data analysis, or execution result. Assume you have daily closing and opening prices for six fictional ETFs. A five-trading-day return signal is calculated using information available through the close on trading day (t), then hypothetically applied at the next trading day’s open. Set the cost assumptions before running the analysis, and do not change them after seeing results.

Write the research question with no blanks left to fill:

“For six fictional ETFs, when a hypothetical position is set at the next trading day’s open using a five-trading-day return signal calculated through each trading day’s close, can the signal confirmation time, holding-return interval, and cost assumptions be calculated without mixing them together?”

Put falsification conditions alongside the question:

  1. Reject the rule if the signal calculation includes the next trading day’s open or any later price.
  2. Reject the rule if the calculation treats a trade as though it occurred before the signal was confirmed.
  3. Stop interpreting results if costs, price definitions, or missing-data rules cannot be fixed before execution.
  4. If you change rules, preprocessing, or parameters after viewing the test period, do not retain the previous conclusion; record it as a new experiment.

A good question does not assert an expected return. It states both what information was used when and what would show that the rule is wrong. [S3]

Questions and Evidence: The LLM Builds a Research List, the Human Obtains the Primary Sources

First, the researcher fixes the universe, observation unit, price definition, data period, access date, cost assumptions, baseline, and rejection conditions. In the hypothetical example, this must explicitly include: “Calculate the signal using the closing price and use the next trading day’s open as the hypothetical execution time.”

Rather than asking an LLM to recommend a strategy, use a prompt such as:

“To test the question below, propose the required data columns, definitions to verify in primary sources, possible leakage risks, and falsification conditions. Label anything that cannot be verified as ‘unverifiable.’”

Then verify external facts directly at their primary sources and exact locations, and reconcile every figure against raw data and calculation code. Releases and macroeconomic data may be revised or updated, so preserve the data vintage that was available at the time along with the access date. [S5] Recording input snapshots, access dates, periods, and price definitions is a conservative documentation practice for reproducible computational research. [S5] [S6]

If a data definition is ambiguous or the primary-source location cannot be found, do not proceed to the backtest. Do not merely supplement the LLM response; return to the material-collection stage.

Code and Time Validation: Check the One-Row Shift Before Performance

Give the LLM the confirmed data dictionary and timing rule, then ask for a code draft. The first goal is not to demonstrate performance; it is to confirm that the signal and position are separated by one row in time.

python
import pandas as pdsignal = pd.Series(    [False, True],    index=pd.to_datetime(["2026-01-05", "2026-01-06"]))position_at_open = signal.shift(1)assert position_at_open.loc["2026-01-06"] == False

In this example, the position at the January 6 open uses only a signal confirmed through the January 5 close. This test checks only one rule: the one-row shift. Passing it does not prove that the entire codebase is free of future information.

Item to checkTiming in the hypothetical exampleValidation question
Signal confirmationAfter the close on trading day (t)Does the calculation include any price after (t)?
Label (subsequent outcome) and holding returnDefined interval after the (t+1) openDo the signal and outcome share the same future information?
Hypothetical execution(t+1) openIs it being treated as an execution before the signal was confirmed?
Preprocessing fitWithin the training periodDid standardization, missing-value treatment, or feature selection avoid seeing the test period?
SplitBy dateAre ETF rows from the same date kept out of both training and test sets?

The first-row missing value, final-row unrealized return, ascending date order, market holidays, and missing-data treatment should also be tested separately. Here, a label means an outcome observed after the signal is generated, so it must not be constructed from information available at the signal’s own time. Fit preprocessing only within the training window, and do not reuse the test period as justification for changing rules after the initial evaluation. Breaking these information boundaries can create leakage. [S3]

Splitting an ETF panel by date—keeping all ETF rows from a given date together—is a conservative practical inference that applies temporal-leakage and non-independence principles to an ETF panel. This research did not identify a primary paper that directly validates this exact rule for ETF panels.

Execution and Records: A Backtest Is Not a Stage for Selecting Results

Immediately before a backtest, fix the data snapshot and hash, access date and period, price definitions such as close, open, and adjusted price, and cost assumptions. Also record the code version, package versions, parameters, intermediate outputs, and random seed. In computational research, inputs, parameters, code, external software versions, intermediate results, and seeds should be preserved so the process that generated a result can be reviewed later. [S6]

If you use the LLM again, it is better suited to narrowing down anomaly candidates than to interpreting performance:

“Using the result table and execution logs below, identify only candidate dates or areas requiring rechecking: differences in missing-value treatment, possible omissions in cost application, and unexpectedly sharp moves. Do not determine causes or interpret performance.”

For each candidate, recalculate the reported figures from raw data and code, then inspect the price definition and missing-value treatment for that date. Confirm that the code version and runtime environment match the research log, and revisit the original context of any source-backed fact used in the analysis.

If you explored many strategies, parameter combinations, or prompt variants, do not keep only the result that looks best. Repeatedly searching across strategy and parameter combinations and then selecting only the best-looking result can create selection bias and backtest-overfitting risk. [S4] Preserving prompt variants, the number of candidates and searches, failed attempts, and reasons for stopping is a conservative recordkeeping design supported by [S4] [S6] [S7].

Keep Failures, and Register Changes as New Experiments

Your research log should retain the question and falsification conditions; data snapshots, hashes, access dates, periods, and price definitions; as well as code and package versions, parameters, intermediate outputs, and seeds. If an LLM was used, preserve the model identifier, prompt, request settings, actual inputs and outputs, and execution time.

Fixing the model and seed does not guarantee fully reproducible LLM responses. Nondeterminism can arise from factors such as the computational environment and operation order, so preserving actual inputs, outputs, and execution context is safer. [S7]

If you change even one of the question, data, price definition, split, preprocessing, costs, parameters, or prompt, do not overwrite the prior version as a revision. Record it as a new experiment. This research did not identify a single standard that mandates hashes, full prompt text, candidate and search counts, and failure reasons together. It is a conservative documentation design intended to address selection bias, weak reproducibility, and LLM nondeterminism. [S4] [S6] [S7]

Separate the Research Environment from the Order Environment

The hypothetical example and backtest in this article remain entirely within a research environment. Broker connections, order implementation, and automated order code are outside its scope.

If live orders exist separately, a useful educational safety boundary is an explicit human approval step after a person has reviewed the evidence, figures, code, and sources. In the regulatory context, automation errors in market-access orders can accumulate and propagate quickly, making risk controls and review important. [S8] However, this source is staff guidance on U.S. broker-dealer rules; it does not establish a legal requirement for individual investors to obtain human approval. [S8]

Choose one small research question and keep its question, falsification conditions, data definition, one-row-shift test, and reason for stopping in the same record. This article has not measured whether LLMs reduce real research time, find more errors, or improve returns. Part 12 will examine how to evaluate AI’s usefulness using evidence about productivity, error detection, chronological validation, post-cost economic value, and reproducibility.

Sources

Report an error or share feedback

Open a draft with this article’s title and URL. Review the message and recipient before sending.

To: [email protected]

Open email draft

If no email app opens, copy these details into your usual email service.

Contact information