Artificial intelligence has acquired an impressive list of supposed abilities.
It can write software, summarize legal documents, generate images, translate languages and, depending on the advertisement you are reading, identify tomorrow’s winning stocks before the opening bell.
Apparently, coffee is optional.
That last claim deserves a little caution.
Stock markets are not chessboards. The pieces do not follow fixed rules, the board occasionally changes shape, and central banks have an inconvenient habit of announcing interest-rate decisions halfway through the game.
AI can still be useful. Very useful, in fact.
It can process more information than a person could reasonably review, apply the same standards to thousands of securities and identify combinations of factors that are difficult to spot manually. It can help turn a market full of noise into a smaller list of companies worth researching.
It can also overfit old data, mistake coincidence for insight and produce a probability score with the confidence of someone who has never had to explain a bad trade.
We have seen both sides while developing StockScreen.art.
Some experiments strengthened our belief that AI belongs in modern investment research. Others reminded us that a model can be sophisticated, mathematically impressive and completely unaware that half its evidence is missing.
The useful question is not:
Can AI pick stocks?
That is too broad. It is a bit like asking whether a calculator can run a business.
A better question is:
Which parts of stock selection can AI improve, and which parts still require structure, scepticism and human judgement?
That is the question we have been trying to answer.
AI is very good at doing the boring work
One of AI’s biggest strengths is not especially glamorous.
It does not get tired.
A person can review a limited number of companies carefully. Even a dedicated investor will usually spend most of their time looking at familiar names: stocks already on a watchlist, companies appearing in the news or businesses they have followed for years.
That is understandable.
There are only so many hours in a day, and examining several thousand charts is not most people’s idea of a relaxing evening.
A machine has no such objection.
It can evaluate a broad universe of securities using the same criteria every time. It can review price trends, momentum, volume, volatility, liquidity, relative strength and risk/reward without gradually lowering its standards because lunch is approaching.
It does not become distracted.
It does not decide, after ticker number 437, that “close enough” is now a legitimate technical indicator.
Suppose a research process requires that a stock have:
- sufficient trading liquidity;
- usable price history;
- a credible long-term trend;
- positive relative strength;
- acceptable volatility;
- enough expected upside;
- a sensible relationship between potential reward and downside risk.
A system can apply those requirements to every candidate in exactly the same way.
It does not make an exception because the company has a famous chief executive, an exciting product launch or an unusually energetic fan club online.
That does not mean the rules are correct.
It means they are being applied consistently.
In investment research, honest consistency is already a meaningful advantage.
Unfortunately, AI can do the wrong thing very efficiently
Consistency has a less attractive cousin.
A model will apply a bad rule just as faithfully as a good one.
If the data is wrong, the model can process the wrong data at remarkable speed.
If the research target is poorly designed, the model can become extremely effective at answering the wrong question.
If an important group of securities is excluded from the universe, the model will never complain. It will simply act as though those companies do not exist.
Automation does not magically remove human decisions.
Every model still reflects choices made by people:
- which securities were included;
- which dates were used;
- which variables were selected;
- how success was defined;
- how missing values were handled;
- how risk was measured;
- how the final candidates were ranked.
AI can make a process more consistent.
It cannot make the process correct by itself.
A machine can be wrong with extraordinary discipline.
More data creates more opportunities—and more ways to fool yourself
The ability to analyze thousands of stocks and dozens of variables sounds impressive.
It can also become dangerous.
The more combinations a model tests, the more likely it is to find something that worked historically by accident.
Try enough indicators, thresholds, holding periods, sectors and market conditions, and eventually one combination will produce a beautiful backtest.
The equity curve may rise smoothly.
The statistics may look reassuring.
The chart may use a tasteful shade of green.
None of that guarantees that the relationship will survive contact with the future.
This is one of the central problems in quantitative research.
A model can discover patterns that exist only because the past contains noise. The model does not know the difference. It sees a relationship and uses it.
The research process must ask whether the relationship is real, stable, economically plausible, available at the historical decision time and still useful after realistic costs.
AI can search more broadly than a person.
It can also search more broadly for ways to be wrong.
The experiment that looked suspiciously wonderful
We encountered this directly in our market-regime research.
The idea was reasonable.
Stocks, industries and sectors do not behave identically in every market environment. A company may perform differently during a calm uptrend than during a volatile decline.
If we could measure those historical relationships, perhaps a regime-aware overlay could improve MLAlpha’s rankings.
The first result looked excellent.
The overlay appeared to improve MLAlpha by almost four percentage points.
That is not a tiny improvement hiding in the fourth decimal place.
That is the kind of number that makes everyone lean closer to the screen and briefly consider whether lunch can wait.
It is also the kind of number that should make a research team nervous.
We had two options.
Option one was to accept the result, give the feature a powerful name and add several upward-pointing arrows to the website.
Option two was to audit it.
We chose option two, which was less exciting but considerably cheaper than discovering the problem in production.
The audit found that the saved result represented only two of the four intended walk-forward test periods.
One of those periods had performed extremely well and was doing much of the heavy lifting.
The conclusion looked broad.
The evidence underneath it was not.
Once we reconstructed all four periods correctly, the apparent advantage disappeared. The complete result was essentially flat relative to the existing MLAlpha model.
The breakthrough had politely returned itself to the shelf.
This was a useful lesson in what AI can and cannot do.
The system was capable of studying a complicated relationship involving market states, sectors, industries and individual securities.
That is a genuine strength.
But the model did not pause and say:
I seem to be missing half the experiment. Perhaps we should delay the victory parade.
It calculated the answer it was asked to calculate.
The research process had to notice that the answer was not yet trustworthy.
That is why strong results often deserve more scrutiny than ordinary ones.
False results rarely arrive wearing a warning label.
They usually arrive looking terrific.
AI can find relationships people would miss
Despite that warning, pattern recognition remains one of AI’s strongest advantages.
Markets are influenced by many interacting variables.
Momentum may matter differently depending on volatility. Relative strength may be more useful during a healthy long-term trend than during a declining market.
A moderate RSI reading may be constructive in one context and irrelevant in another.
These interactions can be difficult to capture with a short list of fixed rules.
Machine-learning models can compare many variables at the same time and identify combinations associated with different historical outcomes.
A model might discover that two stocks with similar trend characteristics behave differently depending on liquidity, recent volatility, sector strength, momentum persistence, market breadth or distance from major moving averages.
A person might notice some of these relationships after years of experience.
A model can test them systematically across a much larger dataset—and it does not need to remember where it left its notes.
That is genuinely useful.
AI has patterns. Humans still have “why?”
A model may discover that a particular combination of variables was followed by stronger returns.
It does not necessarily understand why.
The relationship may reflect a lasting economic effect.
It may also reflect one unusual interest-rate cycle, temporary investor enthusiasm, a single dominant sector, a commodity shock, a handful of extreme winners, or pure coincidence wearing a respectable suit.
The model does not sit back and wonder whether the relationship makes financial sense.
It has no quiet moment of reflection.
It has coefficients.
Researchers still need to ask whether the relationship has a plausible explanation, survives different market environments, remains stable across time and would still matter after realistic trading costs.
AI is good at saying:
These things appeared together.
It is less capable of saying:
Here is why this relationship should still matter three years from now.
Sometimes it will attempt the second sentence anyway.
That is why we verify.
More complicated models did not automatically help us
There is a natural assumption in AI research that a more sophisticated model should perform better.
Neural networks sound advanced.
Deep learning sounds even more advanced.
Add enough layers and acronyms, and the model begins to sound as though it should manage a hedge fund, write the annual report and reserve a table for dinner.
Markets remain stubbornly unimpressed.
We tested an LSTM neural network against the existing MLAlpha approach.
LSTMs are designed to learn from sequences, which seems attractive for market data. Prices unfold through time, and the order of observations can matter.
The LSTM did not produce a meaningful improvement.
We also tested a vision-oriented approach that represented market information differently.
It was interesting.
It gave us another way to view the data.
It still did not clearly outperform the existing system.
These were useful experiments precisely because they were not spectacular.
They answered a practical question:
Does the added complexity improve the decision enough to justify itself?
In these cases, the answer was no.
A more complicated model brings more computing power, more tuning, more opportunities to overfit, more components to monitor, more obscure ways to fail, and a higher probability that something breaks at 2:00 a.m. for a reason documented only in a forum post from 2019.
Unless the added complexity creates a clear, repeatable advantage, it is not sophistication.
It is overhead with excellent branding.
The existing MLAlpha model stayed in place.
Sometimes the smartest technical decision is leaving the working model alone.
Engineers do not always enjoy this conclusion, but hard drives tend to appreciate it.
AI is often better at ranking than predicting
Stock prediction is frequently framed in dramatic terms.
Will the stock rise? Will it beat the market? Will it reach a particular target?
Those questions suggest a level of certainty markets rarely provide.
AI may be more useful when the task is comparative.
Instead of asking whether one stock is guaranteed to perform well, the model can ask:
Among the securities that already satisfy our standards, which appear more favourable relative to the others?
That is a ranking problem.
The model does not need to know the exact future return.
It only needs to provide useful separation among credible candidates.
This is how MLAlpha fits into the StockScreen.art architecture.
AlphaEngine first applies transparent qualification standards. It looks at trend, momentum, relative strength, liquidity, volatility, expected upside and trade structure.
MLAlpha then compares the candidates that have already passed those tests.
It is not being asked to wander freely through the entire market and announce which ticker will make everyone wealthy by Thursday.
That would be easier to advertise.
It would be much harder to defend.
A narrower task is less dramatic, but often more useful.
A probability score is not a promise wearing a percentage sign
Machine-learning systems frequently produce probabilities.
A candidate may receive a 68%, 72% or 81% estimated chance of meeting a defined outcome.
The number looks precise.
The future is not.
A probability depends on how success was defined, how many comparable examples existed, which period was used for training, whether the model was calibrated and whether current conditions resemble the historical sample.
It is an estimate.
It is not a contract signed by the market.
Seventy-two percent looks scientific.
“Somewhat more favourable than the other qualified candidates under the model’s historical assumptions” is probably more accurate, but admittedly difficult to fit inside a badge.
That is why probabilities should appear beside confidence, volatility, expected upside, downside risk, risk/reward, sector context, data quality and model limitations.
A probability should contribute to the decision.
It should not put on a crown and become the decision.
Sometimes the model is ready and the data is still looking for its shoes
Our sector-regime experiment produced one of the clearest examples of AI’s dependence on data quality.
The idea was to evaluate sectors independently.
Energy, financials, technology, healthcare and real estate do not move through the same economic cycle at the same time.
It seemed reasonable to determine whether each sector was operating in a bullish, neutral or bearish environment.
Strong securities in confirmed bullish sectors might then receive additional consideration.
The first run produced no eligible promotions.
Across 56,200 candidate rows:
- none occurred in a confirmed sector-bull state;
- none passed the complete promotion process;
- none changed the final Top-5 selections.
Zero is a wonderfully clear number.
It is also sometimes a sign that the pipes are not connected.
A separate audit found that the experiment had stopped before the hypothesis was meaningfully tested.
The local sector ETF histories required for market confirmation were missing. Point-in-time economic data had not yet been built. Several sector labels were not mapped correctly. One confidence threshold differed from the research design, and a sparse-history fallback had not been implemented.
The experiment had not shown that sector regimes were useless.
It had shown that we had not yet supplied enough valid information to calculate them.
AI cannot invent missing benchmark history.
It cannot reconstruct an economic release that was never stored.
It cannot safely guess which historical information investors knew on a particular day several years ago.
It cannot fix a disconnected pipe by staring at it more intelligently.
The model may receive most of the attention.
The data infrastructure usually does most of the work.
Garbage in, beautifully formatted garbage out
The old computing phrase is “garbage in, garbage out.”
AI has introduced a modern variation:
Garbage in, confident answer out.
Sometimes the answer also includes a chart.
A trustworthy historical experiment needs to know when information was actually released, whether later revisions have been accidentally included, which securities were truly eligible at the time, whether delisted companies remain represented, and whether sector classifications are historically accurate.
Using today’s corrected information to simulate yesterday’s decision is one of the easiest ways to create an impressive result that nobody could actually have achieved.
The AI model cannot object if the future has quietly been added to its training data.
It may simply become remarkably “insightful” about events that already happened.
AI can reduce emotion—while humans find new places to put it
Human investors are not perfectly consistent.
We become attached to companies. We chase stocks after large gains. We hold losing positions because selling would make the loss feel official.
A systematic process can help.
It can require the same qualification standards for every candidate. It can calculate trade levels before emotion takes over. It can identify when the original reasons for owning a security have weakened.
The machine does not feel regret.
It does not become excited because a company used the phrase “artificial intelligence” 47 times during an earnings call.
That objectivity can be useful.
But automation does not eliminate emotional mistakes.
It can simply move them upstairs.
Researchers may keep adjusting a model until the backtest looks attractive. Investors may trust a system too much after a strong period. A company may market an AI score more confidently than the evidence deserves.
A flawed manual decision affects one trade.
A flawed automated system can repeat the decision hundreds of times with admirable punctuality.
AI can adapt—and occasionally overreact like the rest of us
Markets change.
Volatility rises and falls. Correlations shift. Sector leadership rotates. A signal that worked well during one environment may become less effective during another.
AI systems can be designed to monitor those changes and update their estimates.
That adaptability is promising.
It is also difficult.
If a model updates too slowly, it may remain anchored to conditions that no longer exist.
If it updates too quickly, every short-term fluctuation begins to look like the start of a new economic era.
Sometimes the right response to new information is adaptation.
Sometimes it is patience.
Patience, unfortunately, is difficult to encode as a feature without the model interpreting it as missing data.
AI can support risk management, but the market still sends invoices
AI is often promoted as a return-prediction technology.
Its role in risk management may ultimately be just as useful.
A system can monitor rising volatility, weakening momentum, increasing correlation, shrinking liquidity, sector concentration, deteriorating breadth and unusual downside behaviour.
It can help identify when a portfolio is becoming more fragile.
What it cannot do is prevent every loss.
The events investors worry about most are often the events with the least reliable training data.
Crashes, liquidity shocks, geopolitical events and sudden regulatory changes are uncommon. Each also unfolds differently.
That is why position limits, diversification and risk controls remain necessary.
A sophisticated model is not a substitute for sensible exposure.
It is still possible to lose money with excellent software.
The market does not offer a technology discount.
Explainability matters—and can become performance art
AI does not have to be a black box.
A well-designed system can show which factors supported a candidate, which factors weakened it, how it compared with other candidates, whether its sector was supportive and why it qualified or was excluded.
That is more useful than a mysterious Buy or Sell label.
At StockScreen.art, we separate concepts that are often compressed into one score:
- technical recommendation;
- quality;
- confidence;
- expected upside;
- risk/reward;
- opportunity type.
A strong company can still be a Hold because the current entry is unattractive or the expected upside is limited.
That is not inconsistency.
It is nuance, which is less exciting than a flashing Buy signal but tends to age better.
Still, explanations need to be treated carefully.
AI systems can produce narratives that sound convincing even when the narrative is not tied to the real decision mechanism.
That is explanation theatre.
The prose is elegant.
The evidence is backstage looking confused.
True explainability should connect directly to the actual model inputs, rule thresholds, score contribution and historical evidence.
Why we use layers instead of one giant genius box
Our research has pushed us toward a layered architecture.
No single model is expected to do everything.
Universe construction
Determines which securities are eligible for serious analysis.
AlphaEngine qualification
Applies transparent standards for trend, momentum, relative strength, liquidity, volatility and trade quality.
MLAlpha ranking
Compares already qualified candidates using learned historical relationships.
Market and sector context
Examines whether the surrounding environment appears supportive, neutral or weak.
Portfolio construction
Considers overlap, correlation, concentration, risk contribution and position size.
Human oversight
Reviews unusual results, data defects, research changes and practical constraints.
Each layer has a defined responsibility.
This makes the system easier to test and explain. It also prevents one model from quietly becoming responsible for everything from deciding which securities exist to determining how much money should be invested in them.
Systems that claim to do everything often explain very little.
When something goes wrong, they tend to point at a confidence score and leave the room.
What AI is genuinely good at
Based on our research, AI is genuinely useful for:
- reviewing large candidate sets;
- applying analytical rules consistently;
- comparing complicated combinations of variables;
- ranking already qualified opportunities;
- monitoring changing relationships;
- identifying concentration and overlap;
- helping researchers test ideas that would be impractical to study manually.
It does not need to be an oracle to be valuable.
A research assistant that never sleeps, follows instructions consistently and can examine thousands of cases is already useful.
Even if it occasionally needs to be reminded not to celebrate an incomplete experiment.
Where AI still struggles
AI is much weaker at:
- knowing when an experiment was designed badly;
- recognizing that historical information was unavailable at the time;
- understanding why a statistical relationship exists;
- deciding whether added complexity is worthwhile;
- questioning a result because it looks too convenient;
- accounting for events that have never occurred before;
- taking responsibility for the final decision.
Those responsibilities still belong to people.
The biggest danger is not that AI will suddenly become too intelligent.
It is that people will trust it before it has earned that trust.
The real opportunity
The future of AI in investing is probably not an all-knowing machine that delivers tomorrow’s winning ticker every morning.
That is a compelling fantasy.
It would also make financial television considerably shorter.
The more realistic opportunity is to improve the decision process.
AI can help investors search more broadly, apply standards more consistently, compare candidates more intelligently and recognize risks that are easy to miss.
It can help turn a market full of possibilities into a smaller set of questions worth asking.
But it works best when surrounded by reliable data, transparent rules, independent testing, realistic trading assumptions, portfolio constraints, human scepticism and, on particularly difficult days, coffee.
At StockScreen.art, that is how we are approaching the work.
AI is not being treated as a replacement for investment theory, risk management or judgement.
It is being developed as one layer inside a larger research system.
The goal is not to build an oracle.
Oracles have a poor record of publishing reproducible test results.
The goal is to build a better process.
Because AI’s real strength is not that it knows the future.
It is that, when used carefully, it can help us examine the present more thoroughly—and occasionally remind us just how much work remains before a confident answer deserves to be believed.