Forge Fate
Quant & Data Research

Verify the Research, Not the Price: Where LLMs Belong in Quant Research

10 min read

Evidence and scope — LLM usage scope and unmeasured effects

The scope is LLM-generated drafts and questions that can be checked against original material. No measurements of task accuracy, time savings, or return improvements are presented, nor is a price-prediction model evaluated.

In Part 9, we examined why strong results from financial machine learning do not necessarily translate into an investing edge. Here, we apply the same caution to generative AI. Asking an LLM, “Which stock will rise?” is fundamentally different from asking, “What should I verify again in this material?”

LLMs can help summarize large documents, organize research reports, and locate issuer information in SEC filings or earnings materials. At the same time, accuracy, bias, and privacy concerns still matter. [S5] Because generative AI can produce plausible, confident errors, contradictory explanations, and citations that appear to be evidence, it is better used to create drafts that people can check against source material—not to settle an answer. [S1]

This is an educational and research-focused article. It does not recommend trades in particular securities or explain how to place orders. It does not present measured results for task accuracy, time savings, or improved returns.

Three Useful Categories

An LLM’s role in quant research can be organized around three ideas: information transformation, verification support, and decision boundaries. AI risk management emphasizes separating human roles and oversight responsibilities, with processes for testing, evaluation, and validation. [S2]

ConceptWork an LLM can supportWork that should remain with people
Information transformationSummarizing provided documents, organizing tables and lists, explaining terms, drafting code commentsChecking source locations and wording; identifying omissions or distortions
Verification supportSuggesting comparison points, flagging potential anomalies, drafting code-review questions and testsRecalculating from source data, running tests, reviewing timing rules and edge cases
Decision boundariesOrganizing relevant materials and verification questionsAccepting investment facts or figures; choosing securities, positions, or trades; approving live orders

This article does not cover how to design or evaluate price-prediction models. Its scope is limited to LLM drafts and verification questions that can be checked against provided source text and raw data. Because LLM outputs can be wrong, people should set the strength of verification according to the potential impact. [S1] [S2]

Risk Depends on Verifiability and Error Impact, Not Fluency

The low-, medium-, and high-risk categories below are not universal or official ratings. In this article, they are conditional categories based on two questions: Can the AI output be directly checked against source text or raw data? And how much harm could an error cause if it flows into an investment judgment, position size, or order? Even lower-risk work should never be accepted as fact without verification. [S1] [S2] [S7]

Putting a long document into a model does not mean every part of it has been used reliably. One study found that the language models it evaluated became weaker at using relevant information located in the middle of long contexts. That result reflects a particular study design and the models evaluated at that time, so it should not be generalized to every model or every financial document to the same degree. [S3]

Automated orders require a stricter boundary. In the context of U.S. broker-dealer market-access rules, an SEC staff FAQ explains that risk-management controls also apply to automatically generated orders and that automation errors can accumulate and propagate rapidly. This does not establish a legal procedure for individual investors, but it highlights the high impact of mistakes at the order stage. [S7]

Relatively Low Risk: Drafting Work with a Clear Comparison Target

When the original text, raw data, or completed result tables are available for comparison, LLMs can be used at relatively lower risk to draft organization and comparison work. Financial-document summarization and filing-based information retrieval are cited as possible use contexts, alongside concerns about accuracy. [S5]

TaskInput → AI outputHuman verification
Long-material summary draftA provided report or meeting record → summary of key claims, issues, and source locationsCompare every sentence with the source; check missing conditions, exceptions, and figures
Data descriptions and research logsVariable definitions, collection times, and change records → readable documentation and log draftCheck consistency with the data dictionary and execution records
Code comments and test draftsA provided function and its intended purpose → comments, input/output description, and test casesRun the code; check edge cases, date rules, and expected values
Comparing completed strategy result tablesTwo or more provided result tables → list of differences and follow-up checksRecalculate from raw data and code; confirm comparison criteria

Use Long-Document Summaries as a Checklist with Source Locations

When requesting a summary, you can ask for the supporting location of each item and for uncertain points to be marked so that the original text is easy to check. Location labels do not replace verification, however. You still need to compare the wording and confirm that the cited location actually supports the claim. Generative AI can produce errors or fabricated citations, and weaknesses tied to information position in long inputs have also been observed. [S1] [S3]

Treat Research Logs, Code, and Result Tables as Drafts for Review

LLMs can organize variable definitions and execution records into a research log, draft code comments and test cases, and propose questions for comparing result tables. But a written description does not prove when data was available, how it was processed, or how it changed. Likewise, code suggestions are not trustworthy without execution and review. For result tables, confirm that the period, costs, benchmark, and formulas are the same, then recalculate figures from raw data and executable code. [S1] [S2] [S4]

Hypothetical Educational Example: A Summary Is a Verification List, Not an Answer Key

The following is hypothetical educational material, not an actual filing or investment experiment.

Hypothetical Company A delayed its new-product launch until the next quarter. The company cited delays in component procurement as the reason. This material does not include a revenue forecast.

You could make this request: “Summarize the facts, the reason stated, and information not contained in the material in three sentences, and attach the supporting sentence for each.”

A suitable draft would look like this:

  1. Fact: The new-product launch was postponed until the next quarter.
  2. Stated reason: The company cited delays in component procurement.
  3. Cannot be determined: This material alone does not show whether the revenue outlook changed.

A person should check the surrounding context to determine exactly what “the next quarter” refers to, and compare the language to see whether the reason is an established fact or the company’s explanation. They should also confirm that the draft did not add information absent from the material, such as a revenue forecast, competitor reactions, or stock-price effects. This comparison addresses the risk that AI may fill in facts or present plausible inferences as facts. [S1]

Medium Risk: Candidate Generation Is Useful, but Judgment Returns to the Source Material

Extracting candidate risk factors from filings, classifying news topics, structuring companies, products, and events, identifying potential data anomalies, and reviewing backtest code can all help broaden the set of items a person should check. Information retrieval from filings or earnings materials is recognized as a possible use context, but this article does not establish common empirical evidence for how accurate these tasks are in financial practice. [S5]

TaskInput → AI outputHuman verification
Candidate extraction for filing risk factorsA provided filing → candidate risk-factor paragraphs and source locationsFull paragraph context, duplication, conditions and exceptions, actual significance
News-topic classificationA provided group of news items → categories for topics, events, and unverified itemsHeadline, body text, date, source, and consistency of classification criteria
Structuring companies, products, and eventsProvided materials → relationship list and chronological organizationSupporting sentences, dates, and whether names refer to different people or products
Potential data-anomaly explorationProvided data descriptions and samples → values to check and reasonsRaw data, collection errors, units, splits, and missing-value handling
Backtest code reviewProvided code → suspicious areas and test questionsUnit tests, edge cases, timing rules, and actual execution results

For candidates, you can request source locations, quoted sentences, and the reason for classification. For anomalies, request the value and the comparison criterion. A candidate list is an index, not a judgment of importance, and a news event’s factual content should be kept separate from conclusions about investment impact. Convert code review into test questions about future information, date alignment, and the first and last rows, then verify them in the results. [S1] [S4]

Hypothetical Educational Example: Start Backtest Code Review with Timing Questions

This code is also a hypothetical educational review target. It does not show the performance of an actual strategy or document a real error.

python
signal = price.pct_change()position = (signal > 0).astype(int)strategy_return = position * price.pct_change()

You could ask the model: “Explain whether the signal date and return date could be mixed on the same row, and propose two tests to check it.” The questions to verify are straightforward:

  1. Was the position applied on the same day after observing that day’s return?
  2. Do the results change when the position is applied one row later?

A person should confirm that dates are in ascending order, state the timing of the position and return explicitly, and run the rule with a one-row shift. They should also inspect the missing value in the first row and the handling of the final row. The point of this example is not whether the explanation sounds persuasive; the verification standard is the timing rule and the actual execution result. [S1] [S4]

High Risk: Keep Fact Acceptance and Investment Decisions Outside the Boundary

Under this article’s criteria, the following uses are high risk.

Out-of-bounds useWhy it is high riskPrinciple in this article
Using investment facts or figures without verificationIncorrect facts, citations, or calculations can become premises for later decisionsDo not adopt them before checking the source text, raw data, and code
Selecting securities based only on an LLM responsePlausible errors and claims of guaranteed selection require cautionDo not choose without independent evidence and verification
Automating position sizing or trading decisionsErrors flow directly into capital allocation and trading outcomesKeep decision criteria and final approval with people
Live orders without human approvalAutomation errors can accumulate and propagate quicklyExcluded from this article’s scope

A joint investor alert from the SEC, NASAA, and FINRA warns investors to be cautious about claims that AI can provide certain returns or guaranteed security selection. This does not mean that every use of AI is fraud; it means that guaranteed language should not become the basis for an investment decision. [S6]

Human approval is not merely the act of pressing the final button. It should involve responsibly reviewing the supporting text, calculations, assumptions, possible omissions, and impact of errors. AI risk-management principles concerning human roles and validation responsibility support this distinction. [S2]

Evaluate Value Through Revision and Verifiability, Not Speed Alone

This article does not present verified results on how much LLMs reduce research time or improve performance. Rather than describing their value as a specific percentage of time saved, it is more honest to observe four things:

  • What revisions were needed, and how much work did they require?
  • Does each sentence match the source text or raw data?
  • Did the process surface omissions a person might otherwise have missed?
  • Did it reveal plausible but incorrect content, figures, or code?

These items are not claims that effectiveness has already been proven. They are a recordkeeping framework for evaluating your own work. Keeping the input, AI output, and human verification criteria together makes the result easier to revisit or challenge. [S1] [S2]

Start with a task that is easy to check against the source material: a summary draft, data description, result-table comparison, or small code test. Define the input → AI output → human verification flow and the verification criteria. Use LLMs for information transformation and verification questions; leave judgment and orders to people. Part 11 will connect this work into a repeatable process spanning question definition, verification, and research records.

Sources

Report an error or share feedback

Open a draft with this article’s title and URL. Review the message and recipient before sending.

To: [email protected]

Open email draft

If no email app opens, copy these details into your usual email service.

Contact information