Verify the Research, Not the Price: Where LLMs Belong in Quant Research
Evidence and scope — LLM usage scope and unmeasured effects
The scope is LLM-generated drafts and questions that can be checked against original material. No measurements of task accuracy, time savings, or return improvements are presented, nor is a price-prediction model evaluated.
In Part 9, we examined why strong results from financial machine learning do not necessarily translate into an investing edge. Here, we apply the same caution to generative AI. Asking an LLM, “Which stock will rise?” is fundamentally different from asking, “What should I verify again in this material?”
LLMs can help summarize large documents, organize research reports, and locate issuer information in SEC filings or earnings materials. At the same time, accuracy, bias, and privacy concerns still matter. [S5] Because generative AI can produce plausible, confident errors, contradictory explanations, and citations that appear to be evidence, it is better used to create drafts that people can check against source material—not to settle an answer. [S1]
This is an educational and research-focused article. It does not recommend trades in particular securities or explain how to place orders. It does not present measured results for task accuracy, time savings, or improved returns.
Three Useful Categories
An LLM’s role in quant research can be organized around three ideas: information transformation, verification support, and decision boundaries. AI risk management emphasizes separating human roles and oversight responsibilities, with processes for testing, evaluation, and validation. [S2]
| Concept | Work an LLM can support | Work that should remain with people |
|---|---|---|
| Information transformation | Summarizing provided documents, organizing tables and lists, explaining terms, drafting code comments | Checking source locations and wording; identifying omissions or distortions |
| Verification support | Suggesting comparison points, flagging potential anomalies, drafting code-review questions and tests | Recalculating from source data, running tests, reviewing timing rules and edge cases |
| Decision boundaries | Organizing relevant materials and verification questions | Accepting investment facts or figures; choosing securities, positions, or trades; approving live orders |
This article does not cover how to design or evaluate price-prediction models. Its scope is limited to LLM drafts and verification questions that can be checked against provided source text and raw data. Because LLM outputs can be wrong, people should set the strength of verification according to the potential impact. [S1] [S2]
Risk Depends on Verifiability and Error Impact, Not Fluency
The low-, medium-, and high-risk categories below are not universal or official ratings. In this article, they are conditional categories based on two questions: Can the AI output be directly checked against source text or raw data? And how much harm could an error cause if it flows into an investment judgment, position size, or order? Even lower-risk work should never be accepted as fact without verification. [S1] [S2] [S7]
Putting a long document into a model does not mean every part of it has been used reliably. One study found that the language models it evaluated became weaker at using relevant information located in the middle of long contexts. That result reflects a particular study design and the models evaluated at that time, so it should not be generalized to every model or every financial document to the same degree. [S3]
Automated orders require a stricter boundary. In the context of U.S. broker-dealer market-access rules, an SEC staff FAQ explains that risk-management controls also apply to automatically generated orders and that automation errors can accumulate and propagate rapidly. This does not establish a legal procedure for individual investors, but it highlights the high impact of mistakes at the order stage. [S7]
Relatively Low Risk: Drafting Work with a Clear Comparison Target
When the original text, raw data, or completed result tables are available for comparison, LLMs can be used at relatively lower risk to draft organization and comparison work. Financial-document summarization and filing-based information retrieval are cited as possible use contexts, alongside concerns about accuracy. [S5]
| Task | Input → AI output | Human verification |
|---|---|---|
| Long-material summary draft | A provided report or meeting record → summary of key claims, issues, and source locations | Compare every sentence with the source; check missing conditions, exceptions, and figures |
| Data descriptions and research logs | Variable definitions, collection times, and change records → readable documentation and log draft | Check consistency with the data dictionary and execution records |
| Code comments and test drafts | A provided function and its intended purpose → comments, input/output description, and test cases | Run the code; check edge cases, date rules, and expected values |
| Comparing completed strategy result tables | Two or more provided result tables → list of differences and follow-up checks | Recalculate from raw data and code; confirm comparison criteria |
Use Long-Document Summaries as a Checklist with Source Locations
When requesting a summary, you can ask for the supporting location of each item and for uncertain points to be marked so that the original text is easy to check. Location labels do not replace verification, however. You still need to compare the wording and confirm that the cited location actually supports the claim. Generative AI can produce errors or fabricated citations, and weaknesses tied to information position in long inputs have also been observed. [S1] [S3]
Treat Research Logs, Code, and Result Tables as Drafts for Review
LLMs can organize variable definitions and execution records into a research log, draft code comments and test cases, and propose questions for comparing result tables. But a written description does not prove when data was available, how it was processed, or how it changed. Likewise, code suggestions are not trustworthy without execution and review. For result tables, confirm that the period, costs, benchmark, and formulas are the same, then recalculate figures from raw data and executable code. [S1] [S2] [S4]
Hypothetical Educational Example: A Summary Is a Verification List, Not an Answer Key
The following is hypothetical educational material, not an actual filing or investment experiment.
Hypothetical Company A delayed its new-product launch until the next quarter. The company cited delays in component procurement as the reason. This material does not include a revenue forecast.
You could make this request: “Summarize the facts, the reason stated, and information not contained in the material in three sentences, and attach the supporting sentence for each.”
A suitable draft would look like this:
- Fact: The new-product launch was postponed until the next quarter.
- Stated reason: The company cited delays in component procurement.
- Cannot be determined: This material alone does not show whether the revenue outlook changed.
A person should check the surrounding context to determine exactly what “the next quarter” refers to, and compare the language to see whether the reason is an established fact or the company’s explanation. They should also confirm that the draft did not add information absent from the material, such as a revenue forecast, competitor reactions, or stock-price effects. This comparison addresses the risk that AI may fill in facts or present plausible inferences as facts. [S1]
Medium Risk: Candidate Generation Is Useful, but Judgment Returns to the Source Material
Extracting candidate risk factors from filings, classifying news topics, structuring companies, products, and events, identifying potential data anomalies, and reviewing backtest code can all help broaden the set of items a person should check. Information retrieval from filings or earnings materials is recognized as a possible use context, but this article does not establish common empirical evidence for how accurate these tasks are in financial practice. [S5]
| Task | Input → AI output | Human verification |
|---|---|---|
| Candidate extraction for filing risk factors | A provided filing → candidate risk-factor paragraphs and source locations | Full paragraph context, duplication, conditions and exceptions, actual significance |
| News-topic classification | A provided group of news items → categories for topics, events, and unverified items | Headline, body text, date, source, and consistency of classification criteria |
| Structuring companies, products, and events | Provided materials → relationship list and chronological organization | Supporting sentences, dates, and whether names refer to different people or products |
| Potential data-anomaly exploration | Provided data descriptions and samples → values to check and reasons | Raw data, collection errors, units, splits, and missing-value handling |
| Backtest code review | Provided code → suspicious areas and test questions | Unit tests, edge cases, timing rules, and actual execution results |
For candidates, you can request source locations, quoted sentences, and the reason for classification. For anomalies, request the value and the comparison criterion. A candidate list is an index, not a judgment of importance, and a news event’s factual content should be kept separate from conclusions about investment impact. Convert code review into test questions about future information, date alignment, and the first and last rows, then verify them in the results. [S1] [S4]
Hypothetical Educational Example: Start Backtest Code Review with Timing Questions
This code is also a hypothetical educational review target. It does not show the performance of an actual strategy or document a real error.
signal = price.pct_change()position = (signal > 0).astype(int)strategy_return = position * price.pct_change()You could ask the model: “Explain whether the signal date and return date could be mixed on the same row, and propose two tests to check it.” The questions to verify are straightforward:
- Was the position applied on the same day after observing that day’s return?
- Do the results change when the position is applied one row later?
A person should confirm that dates are in ascending order, state the timing of the position and return explicitly, and run the rule with a one-row shift. They should also inspect the missing value in the first row and the handling of the final row. The point of this example is not whether the explanation sounds persuasive; the verification standard is the timing rule and the actual execution result. [S1] [S4]
High Risk: Keep Fact Acceptance and Investment Decisions Outside the Boundary
Under this article’s criteria, the following uses are high risk.
| Out-of-bounds use | Why it is high risk | Principle in this article |
|---|---|---|
| Using investment facts or figures without verification | Incorrect facts, citations, or calculations can become premises for later decisions | Do not adopt them before checking the source text, raw data, and code |
| Selecting securities based only on an LLM response | Plausible errors and claims of guaranteed selection require caution | Do not choose without independent evidence and verification |
| Automating position sizing or trading decisions | Errors flow directly into capital allocation and trading outcomes | Keep decision criteria and final approval with people |
| Live orders without human approval | Automation errors can accumulate and propagate quickly | Excluded from this article’s scope |
A joint investor alert from the SEC, NASAA, and FINRA warns investors to be cautious about claims that AI can provide certain returns or guaranteed security selection. This does not mean that every use of AI is fraud; it means that guaranteed language should not become the basis for an investment decision. [S6]
Human approval is not merely the act of pressing the final button. It should involve responsibly reviewing the supporting text, calculations, assumptions, possible omissions, and impact of errors. AI risk-management principles concerning human roles and validation responsibility support this distinction. [S2]
Evaluate Value Through Revision and Verifiability, Not Speed Alone
This article does not present verified results on how much LLMs reduce research time or improve performance. Rather than describing their value as a specific percentage of time saved, it is more honest to observe four things:
- What revisions were needed, and how much work did they require?
- Does each sentence match the source text or raw data?
- Did the process surface omissions a person might otherwise have missed?
- Did it reveal plausible but incorrect content, figures, or code?
These items are not claims that effectiveness has already been proven. They are a recordkeeping framework for evaluating your own work. Keeping the input, AI output, and human verification criteria together makes the result easier to revisit or challenge. [S1] [S2]
Start with a task that is easy to check against the source material: a summary draft, data description, result-table comparison, or small code test. Define the input → AI output → human verification flow and the verification criteria. Use LLMs for information transformation and verification questions; leave judgment and orders to people. Part 11 will connect this work into a repeatable process spanning question definition, verification, and research records.
Sources
- [S1] Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile | National Institute of Standards and Technology; Chloe Autio, Reva Schwartz, Jesse Dunietz, Shomik Jain, Martin Stanley, Elham Tabassi, Patrick Hall, Kamie Roberts | 2024-07-26 | https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence ↩
- [S2] Artificial Intelligence Risk Management Framework (AI RMF 1.0) | National Institute of Standards and Technology; Elham Tabassi | 2023-01-26 | https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-ai-rmf-10 ↩
- [S3] Lost in the Middle: How Language Models Use Long Contexts | Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, Percy Liang | 2023-11-20 | https://arxiv.org/abs/2307.03172 ↩
- [S4] Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions | Hammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt, Ramesh Karri | 2021-12-16 | https://arxiv.org/abs/2108.09293 ↩
- [S5] FINRA Reminds Members of Regulatory Obligations When Using Generative Artificial Intelligence and Large Language Models | Financial Industry Regulatory Authority | 2024-06-27 | https://www.finra.org/rules-guidance/notices/24-09 ↩
- [S6] Artificial Intelligence (AI) and Investment Fraud | Financial Industry Regulatory Authority, U.S. Securities and Exchange Commission Office of Investor Education and Advocacy, North American Securities Administrators Association | 2024-01-25 | https://www.finra.org/investors/insights/artificial-intelligence-and-investment-fraud ↩
- [S7] Responses to Frequently Asked Questions Concerning Risk Management Controls for Brokers or Dealers with Market Access | U.S. Securities and Exchange Commission, Division of Trading and Markets | 2017-10-11 | https://www.sec.gov/rules-regulations/staff-guidance/trading-markets-frequently-asked-questions/divisionsmarketregfaq-0 ↩
Report an error or share feedback
Open a draft with this article’s title and URL. Review the message and recipient before sending.
Open email draftIf no email app opens, copy these details into your usual email service.
Related posts
Quant & Data Research What to Decide Before Backtesting: Questions and Rules for U.S. ETF Quant Research
Set the research question, trading rules, baselines, and time splits before a first U.S. ETF study. Record the experiment contract before seeing results.
AI Getting the Most from AI Coding Agents: Design the Delegation, Not Just the Tool Comparison
Compare AI coding agents and their controls, then plan task boundaries, verification, approvals, and recovery for practical delegation.