Does AI Actually Help with Quantitative Investing? A Conditional Conclusion and Standards for Validation
Evidence and scope — Literature-based methodology review
Cited research and documentation inform evaluation criteria and validation steps for AI use. External findings are distinguished from proposed future comparisons; this article is not an empirical investment-performance report.
AI can help with quantitative investing—but that help should not be judged by the impression that it “predicted returns.” Its research value can only be established when we can verify that it helped formulate better questions, set up experiments more efficiently, or catch errors earlier. Investment performance, meanwhile, requires separate evidence: time-ordered and out-of-sample testing, comparison against simple benchmarks, and trading costs. [S3] [S7] [S8]
Because generative AI can produce false information, plausible but inaccurate reasoning, and unreliable citations, its outputs cannot themselves be accepted as evidence or conclusions. [S1] In fact, the study in [S5] found that ChatGPT code-generation results could vary even for the same prompt, and setting the temperature to zero did not guarantee determinism.
This article is a methodological discussion for educational and research purposes. It is not investment advice or a recommendation for any specific security.
Validation Comes Before Performance Claims
This series covered expectations, tools, and data in Parts 1–3; validation, risk, and simple benchmarks in Parts 4–6; comparative designs and failure modes for numerical machine learning in Parts 7–9; and the boundaries and validation procedures for LLM use in Parts 10–11. This is only a summary of the series context provided; it does not mean that actual investment performance or implementation results were established at each stage.
The central lesson of Season 1 is not the claim that “AI beat the benchmark.” It is the order of operations: before examining results, fix the question, data definitions, timing rules, benchmark, costs, and recordkeeping. The best-looking in-sample result cannot be assumed to outperform out of sample, and selecting only the strongest result after repeated exploration raises the risk of backtest overfitting. [S3]
This article is therefore not a success story. It sets out how AI’s usefulness can be tested or challenged without turning a research design into empirical evidence.
LLMs and Numerical ML Have Different Jobs
LLMs are closer to research assistants whose work must be verified. They can help organize items to check against source material, draft code, and flag possible anomalies in results. But facts, figures, citations, and code proposed by an LLM must be confirmed against original sources, raw data, and a rerun of the code under the stated analysis conditions. [S1] [S2] [S5]
Numerical machine learning, by contrast, is something to test in a predictive role. Performance should be compared using fixed data, a defined signal-availability time, training and evaluation periods, cost assumptions, and a simple benchmark. A model’s complexity is not evidence that it is better than the benchmark. Good results obtained without time-ordered and out-of-sample comparison are not evidence of economic value. [S3]
When applying LLMs to historical financial information, one must also ask separately: “Was this information genuinely available at that time?” If a model’s pretrained knowledge overlaps with the historical evaluation period, timing bias can arise because the evaluation may include information that was not knowable at the time. [S9] More generally, using future information in a backtest can overstate performance and understate risk. [S8]
The question should no longer be, “Does AI give us the answer?” It should be, “What rules let us challenge the drafts and model results produced with AI?”
Three Dimensions for Assessing Usefulness
AI’s usefulness should not be judged by a single return figure or a user’s intuition. The three dimensions below are not measured results from this season; they are an evaluation design to be tested going forward.
| Assessment dimension | Question to test | Items to include in the assessment |
|---|---|---|
| Research productivity | Did AI help complete the research faster? | Total time, including research and coding as well as review and revision |
| Validation quality | Did AI reduce errors or help find them earlier? | Agreement with source material, omissions, detection of factual and code errors, reproducibility |
| Economic value | Is an AI-assisted strategy actually better? | Post-cost returns, drawdowns, turnover, and stability across periods relative to a benchmark under identical conditions |
For research productivity, saving time on a first draft is not enough. If review and revision take longer, total time may increase instead. Randomized studies have reported both settings in which AI tools increased task time and settings in which they reduced it. [S4] [S6] Therefore, “AI saves time” should not be treated as a universal proposition, but as a hypothesis to test by comparing total time on the same task.
For validation quality, the key is not how quickly a plausible output is produced. It is how early incorrect facts, citations, and code are found, and whether the corrections can be reproduced. Reproducibility in computational research is tied to whether consistent results can be obtained again using the same inputs, methods, code, and analysis conditions. [S2]
Economic value should be assessed last. Strong in-sample results do not imply out-of-sample superiority, and transaction costs can change a strategy’s economics. [S3] In one walk-forward cryptocurrency study, unlike some configurations with positive pre-cost performance, a simple directional strategy failed after a 10bp transaction cost. [S7] This example says nothing about ETFs or the results of this series. It does, however, clearly show why economic value cannot be claimed without accounting for costs and turnover.
What Can Be Said Now—and What Cannot
| Category | What this article can say | What it still cannot say |
|---|---|---|
| Methodology established | LLM outputs cannot be adopted as judgments without checking them against original sources, data, and code; reproducible research requires records of data, code, methods, and environment. [S1] [S2] [S5] | These principles do not demonstrate improved investment performance or reduced working time in this series. |
| Results not measured | This research scope contains no actual ETF backtest, no post-cost return, drawdown, turnover, or multi-period stability relative to a benchmark, and no data on working time, error counts, or revision volume with and without AI. | We cannot claim that “AI beat the benchmark this season,” “AI reduced the time required,” or “AI found errors earlier.” |
| Evidence needed next | Small experiments should preserve fixed data and definitions, code and environment, timing tests, exploration history, post-cost results, failure records, and the models, prompts, and actual inputs and outputs used. This is a proposal addressing reproducibility, overfitting, and LLM-validation risks. [S1] [S2] [S3] [S5] [S7] | This record bundle is not a completed result; it is a minimum requirement for future evaluation. |
This distinction is not an admission of failure. It is the boundary of the research. If returns, time savings, or earlier error detection were not measured, presenting them as already established would undermine the validation process itself.
Conditions for Calling AI “Useful”
To say that AI was useful, at least the following conditions should be confirmed:
- On the same task, did total time—including research, coding, review, and revision—fall relative to the comparison condition? [S4] [S6]
- Were factual, code, and omission errors found earlier, and can the review process and corrections be rerun? [S1] [S2] [S5]
- Did results remain more stable than a simple benchmark while preserving time-ordered testing? [S3]
- Did meaningful results remain after transaction costs, with turnover and drawdowns explained alongside them? [S7]
- Were the findings robust across multiple periods rather than dependent on a single period or asset? [S3]
- Can a person explain and challenge the assumptions, falsification conditions, and reasons for failure? [S1] [S2]
Failing to meet these conditions does not mean AI is always useless. It means that research value and economic value have not yet been established separately.
Conversely, practical use should be deferred when warning signs appear. If rules, parameters, or prompts were repeatedly changed after seeing results and no exploration history exists, the risk of retaining only the best-looking outcome should be examined. [S3] If evidence comes only from a short period or a single asset, generalization should be withheld. If the timing of information availability is unclear, first check whether future information or pretrained knowledge may have leaked into the evaluation. [S8] [S9] If an LLM’s facts, citations, or code have not been verified against source material and execution, practical use should be deferred. [S1] [S5] If post-cost outperformance disappears or results are not more stable than a simple benchmark, it is not yet time to claim economic value. [S3] [S7]
A Small Comparative Experiment for the Next Decision
The value of using AI should be compared on the same research task. The following is a proposed evaluation design, not a result.
First, predefine the data, question, completion criteria, timing rules, and benchmark. Then assign AI-use and no-AI conditions randomly or in a balanced manner. For each condition, record total time spent on research, coding, review, and revision; factual and code errors; revision volume; and whether the work can be rerun. If investment performance is also evaluated, compare post-cost returns, drawdowns, turnover, and stability across periods against the same benchmark. Preserve not only adopted results, but also failures, abandoned paths, change histories, and the LLM’s actual inputs and outputs.
The fact that separate randomized studies found opposite effects of AI on task time means that an average result from external research should not be treated as the result of one’s own research. [S4] [S6] Those studies are precedents for comparative measurement, not predictions of what a quantitative research project will find.
Conclusion: Research Value Must Be Verifiable, and Profitability Must Be Proven Separately
Season 1’s answer to “Does AI actually help with quantitative investing?” is conditional. AI can help research when it is possible to verify that it improved questions, experiments, or error checking. LLM outputs are objects of verification, while numerical ML predictions must be tested through time-ordered comparisons under a benchmark. [S1] [S3] [S5]
But research usefulness and investment profitability are not the same claim. To discuss profitability, information availability must be controlled, out-of-sample comparison and transaction costs must be incorporated, and stability across multiple periods must be demonstrated separately. [S3] [S7] [S8] [S9]
The next step is not to declare greater confidence. It is to challenge or refine this conditional conclusion through small experiments with fixed data, validation procedures, and failure records. Whether AI is useful is determined not by the fluency of its answers, but by whether people can measure, reproduce, and challenge the help it provides.
Sources
- [S1] Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile | National Institute of Standards and Technology | 2024-07-25 | https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf ↩
- [S2] Reproducibility and Replicability in Science | National Academies of Sciences, Engineering, and Medicine | 2019 | https://www.nationalacademies.org/read/25303/chapter/6 ↩
- [S3] THE PROBABILITY OF BACKTEST OVERFITTING | David H. Bailey, Jonathan M. Borwein, Marcos López de Prado, Qiji Jim Zhu | 2015-02-27 | https://www.davidhbailey.com/dhbpapers/backtest-prob.pdf ↩
- [S4] Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity | Joel Becker, Nate Rush, Elizabeth Barnes, David Rein | 2025-07-12 | https://arxiv.org/abs/2507.09089 ↩
- [S5] An Empirical Study of the Non-determinism of ChatGPT in Code Generation | Shuyin Ouyang, Jie M. Zhang, Mark Harman, Meng Wang | 2023-08-05 | https://arxiv.org/abs/2308.02828 ↩
- [S6] How much does AI impact development speed? An enterprise-based randomized controlled trial | Elise Paradis, Kate Grey, Quinn Madison, Daye Nam, Andrew Macvean, Vahid Meimand, Nan Zhang, Ben Ferrari-Church, Satish Chandra | 2024-10-16 | https://arxiv.org/abs/2410.12944 ↩
- [S7] Machine Learning-Based Bitcoin Trading Under Transaction Costs: Evidence From Walk-Forward Forecasting | Andrei Bysik, Robert Ślepaczuk | 2026-05-19 | https://arxiv.org/abs/2606.00060 ↩
- [S8] Look-Ahead Benchmark Bias in Portfolio Performance Evaluation | Gilles Daniel, Didier Sornette, Peter Wohrmann | 2008-10-10 | https://arxiv.org/abs/0810.1922 ↩
- [S9] Assessing Look-Ahead Bias in Stock Return Predictions Generated By GPT Sentiment Analysis | Paul Glasserman, Caden Lin | 2023-09-29 | https://arxiv.org/abs/2309.17322 ↩
Report an error or share feedback
Open a draft with this article’s title and URL. Review the message and recipient before sending.
Open email draftIf no email app opens, copy these details into your usual email service.
Related posts
Quant & Data Research Can Logistic Regression Beat the Baseline? What Makes a Model Comparison Fair
Define timing, data splits, retraining, and cost rules for a fair comparison of logistic regression and baseline strategies, before testing actual performance.
Quant & Data Research Why shift(1) Is Not Enough: A Time Contract for Signals, Execution, and Labels
Separate signal, execution, return, and label timing. Use synthetic unit tests to check information leakage and preprocessing boundaries.