Forge Fate
Quant & Data Research

Rules, Machine Learning, Deep Learning, and LLMs: Four Ways to Answer the Same Investment Question

17 min read

Evidence and scope — Concept comparison and hypothetical examples

Sources support a comparison of inputs, outputs, and evaluation for rules, machine learning, deep learning, and LLMs. Investment scenarios are hypothetical, not experiments validating relative profitability or general LLM productivity effects.

Why you should define the outcome, data, and evaluation criteria before choosing an AI tool

Part 1 framed AI not as an automatic stock-picking machine, but as a tool for testing investment ideas. That raised a new question:

What should we ask machine learning, deep learning, generative AI, and LLMs to do—and how do their roles differ?

Does a moving-average crossover strategy become machine learning simply because a computer executes it automatically? Is ChatGPT’s ability to describe market conditions in natural language the same as predicting next month’s return? Because deep learning is more advanced than other forms of machine learning, will it always produce better results?

At first, I thought of all these technologies as variations on a “smarter stock-selection tool.” The more I studied them, however, the more I realized that the important distinction was not their name or novelty. It was what input each tool receives, what output it produces, and how that output should be validated.

AI is a broad concept encompassing systems that produce outputs such as predictions, content, recommendations, and decisions. Within that broad category, however, we need to distinguish between systems that execute rules already completed by a person and machine-learning systems that learn relationships from training data.[S1]

This article will not treat these terms as isolated dictionary entries to memorize. Instead, we will give the same investment question to a rule-based strategy, machine learning, deep learning, and an LLM, then examine how each one handles it differently.

All investment examples in this article are hypothetical and intended only to explain the concepts. They do not recommend any security, provide trading instructions, or guarantee returns from any method.

Let’s Start With One Question

We will use the following question throughout the article:

How should we use recent price, volume, and volatility data to approach the next month?

It sounds reasonable, but it is still far too vague to serve as a research question.

What exactly does “approach the next month” mean?

  • Predict next month’s return as a number.
  • Classify next month as an up or down period.
  • Find past market environments similar to the current one.
  • Summarize relevant documents in natural language.
  • Create rules that turn a prediction into a trading signal.

Even with the same price and volume data, the task changes depending on whether the desired output is a number, a category, or a grouping of observations. Regression, classification, and clustering each require different evaluation methods, and the score of a complex model is meaningful only when compared with a simple baseline.[S2][S3][S4]

We will therefore compare each method using the following framework.

DimensionQuestion to ask
ProblemWhat exactly are we trying to learn?
InputWhat data will we provide?
OutputDo we want a rule-based signal, number, category, cluster, or language draft?
Human roleWhich conditions, labels, models, or interpretations must a person define?
EvaluationWhat comparison will determine success or failure?
LimitationDoes this output alone constitute a complete investment strategy?

First, Let’s Draw the Conceptual Map

For beginners, the following mental model is useful.

text
AI: A broad category of systems that generate predictions, content, recommendations, or decisions├─ Rule-based approaches: May form part of automation or an AI system│  └─ Execute human-defined conditions; automatic execution alone is not machine learning└─ Machine learning: Learns relationships or patterns from data   └─ Deep learning: Learns representations through multiple processing layersGenerative AI: A category of AI that generates new content such as text, images, and audio└─ LLM: A large-scale model primarily trained to learn and generate language

A rule-based system may fall under automation or even AI in a broad definition, but there is no universally accepted boundary under which every rule-based automation system is automatically considered AI. The key distinction is clear: a system does not become machine learning merely because it automatically executes conditions that a person has fully specified.[S1]

Machine learning is a collection of methods that learn patterns and relationships between inputs and outcomes from data instead of requiring a person to explicitly write every relationship.[S1] Deep learning is a machine-learning approach that uses multiple processing layers to learn representations at progressively different levels.[S5] This relationship can be understood as AI ⊃ machine learning ⊃ deep learning.

Putting generative AI and LLMs on that same straight line, however, can be misleading. Generative AI is a broad category of systems that create new content—such as text, images, video, and audio—by modeling the structure and characteristics of their input data.[S8] LLMs are large models within this space that primarily work with language, and many modern LLMs use deep learning and Transformer-based architectures.[S9]

Generative AI and LLMs are therefore not synonyms. An image-generation model can be a form of generative AI without being an LLM.[S8] Conversely, not every AI system is generative AI or an LLM.[S1][S8]

Once we understand these relationships, we can stop asking, “Which technology is more advanced?” and start asking, “Which technology fits the output I need?”

The First Approach: In a Rule-Based Strategy, a Person Defines the Relationship

Let’s first turn our question into a moving-average crossover strategy.

Suppose a person writes the following conditions:

  • Generate a signal when the short-term moving average crosses above the long-term moving average.
  • Add an asset to the candidate set when its recent return exceeds a predetermined threshold.
  • Ignore the signal when volatility is above a certain level.

Price and volume data are used to calculate these conditions. But the computer has not examined the data and independently learned the meaning of the moving-average periods or thresholds. A person completed the relationship; the computer merely executes it consistently.

DimensionRule-based strategy
Who defines the relationship?A person
Role of the dataCalculate predetermined conditions
Typical outputSignals such as enter, exit, or remain on the sidelines
What a person must definePeriods, thresholds, position rules, and execution conditions
Evaluation questionDoes the rule remain effective on unused periods after costs are included?

Automatic execution alone does not make something machine learning. A system that follows human-written instructions is structurally different from one that learns patterns from training data.[S1]

For this article, I did not find evidence from a comparison of rule-based strategies and machine learning under identical markets, periods, costs, and evaluation conditions that would justify declaring either approach universally superior. The point we can establish here is not which one performs better, but the structural difference in who defines the relationship.

The Second Approach: Machine Learning Defines the Question Through Data

Now let’s give the same question to machine learning.

Starting a machine-learning project does not mean choosing an algorithm first. We must begin by deciding what will be used as input and what will count as the correct outcome.

Suppose each row of our dataset represents the end of a month:

TimeInputsOutcome observed later
End of a given monthRecent return, volatility, change in volumeNext month’s return or up/down direction

This example naturally introduces several core terms.

  • A feature is an input used to make a prediction. Here, recent return, volatility, and the change in trading volume are features.
  • A label is the outcome the model is asked to predict. The label might be next month’s return or whether the next month is up or down.
  • Training is the process of adjusting a model’s internal parameters so that they reflect relationships between historical features and labels.
  • Inference is the process of giving new features to a trained model and obtaining a prediction.
  • Evaluation means measuring performance with predetermined metrics on data that was not used to generate the predictions, then comparing that performance with a baseline.

Supervised learning uses labeled examples to learn a mathematical relationship between features and labels. Applying the trained model to new inputs is called inference.[S2] Evaluation should use metrics appropriate to the task and compare the model with a simple predictor or other baseline.[S4]

The next article will examine specific ways to preserve chronological order and prevent data leakage.

Regression: Predicting Next Month’s Return as a Number

Let’s make the question more precise:

Can recent return, volatility, and changes in trading volume predict next month’s return?

This is a regression problem. Regression is a supervised-learning task that predicts a numerical label.[S2]

  • Features: recent return, volatility, and change in volume
  • Label: next month’s return
  • Output: expected return or another continuous prediction score
  • Evaluation: how much the regression error improves over a simple baseline prediction

A simple baseline might predict that next month’s return will always equal the historical average, or it might always predict zero. The appropriate baseline depends on the research objective. Without one, however, it is difficult to tell whether the error reported by a complex model is actually good.[S4]

There is another important distinction. Building a model that outputs an expected return does not mean that we have completed an investment strategy.

We still need separate decision rules covering the expected-return threshold for taking a position, position size, exit timing, risk management, and trading costs. A regression model predicts a number; it is not the complete strategy that turns that number into action.

Classification: Distinguishing an Up Month From a Down Month

Now let’s change the question:

Can recent return, volatility, and changes in trading volume tell us whether next month will be up or down?

The features can be the same as in the regression example. What has changed is the label.

  • Regression label: a number such as next month’s return of 2.1% or -1.4%
  • Classification label: a category such as next month being up or down

Classification predicts a categorical label. Its output may be a single category or the probability of belonging to each category.[S2]

Even if a model assigns a 60% probability to an up month, a buy decision does not automatically follow. A person must separately decide the probability threshold for calling the month “up,” how differences in probability should affect position size, and how to constrain downside risk and costs.

Nor can we conclude that better classification accuracy automatically produces better investment performance after costs. Classification metrics and a strategy’s return and risk metrics measure different things.[S4]

Ultimately, the key distinction between regression and classification is not the impressive name of the algorithm. It is what we defined as the correct answer.

What If We Do Not Provide an Answer? Unsupervised Learning and Clustering

This time, suppose we provide neither next month’s return nor an up/down label.

Instead, we use features such as the following to group similar periods:

  • Volatility
  • Recent trend
  • Changes in trading volume
  • Correlations among assets

An approach that explores structure within data without labels is called unsupervised learning. Clustering groups similar observations according to the features and similarity criteria chosen by the user.[S3]

DimensionClassificationClustering
Predetermined labelYesNo
Central questionWhich category does this belong to?Which observations resemble one another?
Example outputUp or downCluster 1, 2, or 3
How meaning is assignedDetermined when labels are createdInterpreted by a person after seeing the results
Major riskPoor labels or evaluationResults changing with the choice of features, scales, or similarity measure

Suppose the clustering results place several periods in “Cluster 1.” The model has not automatically determined that this is a bull market or a crisis regime. The cluster number has no inherent economic meaning. A person must examine each cluster’s average volatility, returns, correlations, and other properties before deciding what interpretation—if any—is appropriate.[S3]

Changing feature scales or the method used to calculate similarity can also change the clustering results.[S3] The mere existence of clusters therefore does not prove that we have discovered genuine market regimes or that the results will improve future investment performance. We must separately test whether similar structures persist in other periods, whether their interpretation is stable, and whether they are useful for the actual research objective.

Clustering is less a machine that reveals the market’s hidden correct answer than a tool for exploring data according to the criteria we selected.

The Third Approach: Deep Learning Learns More Complex Representations

Deep learning is not a separate technology competing with machine learning. It is a machine-learning approach that uses multiple processing layers to learn representations at different levels of abstraction.[S5]

Using deep learning therefore does not automatically change the underlying question.

  • Predicting next month’s return is still regression.
  • Predicting the probability of an up month is still classification.
  • What changes is the structure and complexity of the model used to represent the relationship between inputs and outcomes.

Deep-learning models can learn progressively more complex representations across multiple layers and have delivered major advances in processing high-dimensional, unstructured data such as images, video, speech, and text.[S5] In investment research, they can be designed to use tabular price and volume data or to learn representations from less structured inputs such as news text and chart images.[S5]

Let’s make the phrase “complex nonlinear relationship” more concrete with a hypothetical example.

A simple model might assume that, as recent return rises, the expected return for the next month increases at a constant rate. A more flexible model could represent a relationship in which the same recent return leads to different outcomes depending on whether volatility is low or high.

This is only an illustration of the shapes of relationships that different models can represent. It does not establish a proven investment rule that momentum necessarily works better when volatility is low.

In one large study of US equity returns, tree-based models and neural networks used nonlinear interactions among predictors to produce better out-of-sample results than some linear methods.[S6] Those findings, however, were specific to the US market from 1957 through 2016 and to the researchers’ chosen variables, models, and portfolio designs. They cannot simply be transferred to another market or to an individual investor’s circumstances.[S6]

If Deep Learning Is More Sophisticated, Why Not Start There?

Even in that US equity study, making a neural network deeper did not lead to continuously improving performance. Among the configurations compared, performance peaked at around three hidden layers and declined in deeper models.[S6] The study describes individual stock returns as having a low signal-to-noise ratio because they are strongly affected by unpredictable news.[S6]

This does not mean that every financial dataset is small or that every deep-learning model performs worse than a simpler model. The volume and structure of data vary with the market, frequency, and data source, and I did not find evidence supporting such a universal claim.

The conclusion is therefore not that we should declare deep learning superior or inferior in advance. It is that we should define the objective and evaluation metrics first, then compare it with a simple baseline under the same conditions. General machine-learning practice also recommends establishing baseline performance and behavior with a simple model before adding complexity as needed.[S7] Applying that principle to investment research is a practical inference, not a leaderboard that guarantees the performance of any particular strategy.

The reason to choose deep learning should not be that its name sounds more like “real AI.” It should be evidence that, on the same problem, it provides meaningful additional value over a simpler baseline.

The Fourth Approach: Generative AI and LLMs Draft Language-Based Work

Generative AI is a category of AI models that create new synthetic content such as text, images, video, and audio.[S8] LLMs are large models within that category that primarily learn and generate language, and many modern implementations use deep learning and Transformer-based architectures.[S9]

When we apply an LLM to the same investment question, it is more appropriate to treat it as a potential tool for drafting language-based work that a person will review—not to assume it is a model that can automatically predict next month’s return. For example, we might ask it to draft:

  • A summary of the key points in news and regulatory filings
  • A structured set of research and falsification questions
  • A framework for data-processing or experimental code
  • Research notes documenting experimental conditions, changes, and results

This list is not evidence that LLMs have been proven to improve the accuracy or productivity of investment research. I did not find sufficient evidence in this review to determine how much these tasks generally improve. Each output should therefore be treated as a candidate whose usefulness must be measured and a draft that requires review.

An LLM’s fluency does not automatically imply factual accuracy or predictive power. Generative AI can confidently produce content that sounds plausible but is factually wrong or internally inconsistent, and NIST identifies this as a major risk.[S10] Researchers have also documented the tendency of language models to generate fluent but incorrect answers.[S11] Investor-protection authorities similarly warn investors not to rely solely on AI-generated information and to verify its underlying sources.[S12]

Facts, figures, and quotations in generated summaries or explanations must therefore be checked against primary materials, and generated code must be verified by actually running it.[S10][S11][S12] A natural-sounding sentence is not a regression or classification prediction evaluated on separate data. Nor is it an investment strategy with defined signals, positions, entries, and exits.

Parts 10 and 11 will cover specific verification procedures and prompt design. For now, it is enough to remember this boundary: an LLM can draft language-based work, but fluent output and a validated prediction are different kinds of results.

The Same Question, Different Roles

The following table brings the distinctions together.

ApproachWhat a person defines firstWhat the data doesTypical outputCentral evaluation question
Rule-based strategyConditions such as moving-average periods and momentum thresholdsCalculates predetermined conditionsTrading signalDoes it remain effective on unused periods after accounting for costs?
Supervised regressionFeatures and a numerical labelLearns the relationship between inputs and a numerical outcomeEstimated return for the next periodDoes regression error improve over a simple baseline?
Supervised classificationFeatures and a categorical labelLearns the relationship between inputs and categoriesUp category or probabilityDo classification metrics improve over the baseline?
Unsupervised clusteringFeatures and similarity criteriaFinds similar observations without labelsClustersAre the results interpretable and stable across periods?
Deep learningObjective, inputs, architecture, and evaluation criteriaLearns complex representations across multiple layersNumbers, categories, or learned representationsDoes it add value over a simpler model?
LLMLanguage task and scope of reviewGenerates context-appropriate languageLanguage draft for reviewDoes it fulfill the requested language task in a form that can be checked?

The key point is that these outputs cannot all be reduced to a single measure of “AI investment performance.” Regression, classification, and clustering require different evaluation criteria. Better prediction metrics must also be distinguished from better performance by an actual investment strategy.[S4] People must interpret the economic meaning of clusters,[S3] while an LLM’s natural-sounding response is not the same kind of output as a validated prediction score.[S10][S11]

What I Changed Was Not the Tool, but the Order of the Questions

At first, I saw AI, machine learning, deep learning, and ChatGPT as different versions of a smarter stock-selection tool. I vaguely assumed that the newest and most complex technology would uncover more opportunities in the market.

The first shift came when I understood the difference between moving-average or momentum strategies and machine learning. In a rule-based strategy, a person defines the relationship and the computer executes it. In machine learning, a person designs the features, labels, and evaluation method, after which the model estimates relationships from the data.[S1][S2]

The second shift was realizing that regression, classification, and clustering are not merely entries in a list of algorithms. They answer different kinds of questions—about numbers, categories, and unlabeled structures—and must be evaluated differently.[S2][S3][S4]

The third shift was learning not to see deep learning as an automatic upgrade. The ability to learn complex representations and evidence that a model outperforms simpler alternatives on a particular financial problem are two separate claims.[S5][S6] I now define the objective and metrics and establish a simple baseline before trying a more complex model.[S7]

The final shift was repositioning the LLM as a language tool rather than a numerical predictor or a finished strategy. The ability to explain something smoothly is not the same as the ability measured by regression error, classification performance, or investment returns.[S10][S11]

I no longer begin by opening the most fashionable tool. I first state the problem in one sentence, separate the available inputs from the required answer, and define the desired output and success criteria. Only then do I select the simplest suitable candidate from rules, regression, classification, clustering, deep learning, or an LLM. I choose a more complex tool only when it demonstrates additional value over the baseline under the same conditions.[S4][S7]

The biggest change I gained from studying AI was not a desire to find more complex tools. It was learning to write down the problem and validation criteria before choosing a tool.

Scope of This Research and Remaining Gaps

This article is a conceptual map of how these tools relate to one another, how their outputs differ, and how they should be evaluated. The available evidence does not establish universal superiority between rule-based strategies and machine learning, the relative performance of deep learning across every financial problem, the economic reality of clusters, or the general effect of LLMs on the accuracy or productivity of investment research. Findings from a particular US market study also cannot be generalized directly to other markets.[S3][S6] The investment examples in this article are therefore hypothetical illustrations of how each tool addresses a different question—not verified investment laws or recommendations for any particular market.

Conclusion: Choose the Problem Before Choosing AI

Part 1 concluded that “AI is not an automatic stock-picking machine, but a tool that helps with validation.” This article adds that there is no single tool called AI.

A rule-based approach executes conditions defined by a person. Machine learning learns relationships between inputs and outcomes from data. Deep learning is a machine-learning approach that uses multiple layers to learn more complex representations. Clustering explores structure according to selected criteria without being given a correct answer, while LLMs primarily generate language-based outputs.

These are not performance tiers that can be ranked on a single scale. They answer different questions, and their failures must be identified in different ways.

The important question is not, “Should I use AI?” First decide, “What problem am I trying to solve, with what data, and according to what criteria?” Only then can you decide what role to assign to rules, machine learning, deep learning, or an LLM.

The next article will examine why chronological order matters when fairly testing the problems and models defined here, and how to distinguish data leakage from repeated experimentation. Specific workflows for using LLMs safely, along with prompt design, will be covered separately in Parts 10 and 11.

Sources

Report an error or share feedback

Open a draft with this article’s title and URL. Review the message and recipient before sending.

To: [email protected]

Open email draft

If no email app opens, copy these details into your usual email service.

Contact information