AlphaStocks

By Maksym Lytvynov, Founder of AlphaStocks | Last updated: September 2026

Backtest Results

Does the AlphaStocks scoring formula pick stocks that beat the market? We tested it properly, found and fixed two look-ahead biases in our own backtest, and publish what is left. A portfolio of the top-10 highest-scoring S&P 500 stocks returned +136.3%versus +100.3% for the S&P 500 from January 2021 through February 2026. That excess is not statistically significant, and most of it is explained by market exposure rather than stock selection.

Earlier versions of this page reported a larger out-of-sample alpha. That figure did not survive the fixes described below, and we have withdrawn it. AlphaStocks is a research tool that turns five classic investor models into one score for shortlisting — not a proven return predictor. Below you will find the methodology, the full results, limitations, and the disclaimer.

What Is Walk-Forward Validation?

Walk-forward validation is the gold standard for testing quantitative investment strategies. The concept is straightforward: you split your data into two non-overlapping periods. The first period (in-sample) is used to design, calibrate, and optimize your model. The second period (out-of-sample) is used to test it on data the model has never seen.

This approach directly addresses the most dangerous problem in quantitative finance: overfitting. With enough parameters, any model can be made to look spectacular on historical data. A model that “predicts the past” with 99% accuracy might be completely useless going forward because it memorized noise rather than discovering genuine patterns.

Walk-forward validation prevents this by separating discovery from confirmation. If a strategy performs well on data it was specifically designed for, that tells you very little. If it performs well on data it has never seen, that is meaningful evidence that the underlying signals are real.

Our own test falls short of this standard in one important way: the current weights were chosen on a window that overlaps the out-of-sample period. We explain what that means for the results in the in-sample vs. out-of-sample section below.

How We Structured Our Test

AlphaStocks split the January 2021 through February 2026 period into two distinct windows:

PhasePeriodPurpose
In-sampleJan 2021 – Dec 2023Reference window spanning varied regimes (post-COVID rally, 2022 bear, 2023 recovery) used to study the scoring formula.
Out-of-sampleJan 2024 – Feb 2026Later window used to check whether the result holds. The current v22 weights were chosen on a window that overlaps it, so it is not truly unseen data for them.

Each month, every S&P 500 stock is scored using only information that was public on that date: financial statements as of their SEC filing date (not the period end they describe), prices up to that day, and the index membership at the time. The 10 highest-scoring stocks are bought in equal weight and held until the next monthly rebalance.

Two Look-Ahead Biases We Found and Fixed

A look-ahead bias lets a backtest use information that did not exist yet on the trading date. It is the most common way a backtest looks better than reality. Auditing our own code, we found two:

  1. Today's price in historical scores. The Graham, Buffett, and Lynch models compare a company's value with its price. The backtest read the live quote instead of the price on the scoring date, so a 2021 score was partly a function of the 2026 price.
  2. Today's index membership. The backtest selected from the current S&P 500 list for the whole period, so companies that joined the index later — often after strong runs — were eligible before they joined.

The two biases pushed the results in opposite directions, which is why the earlier figures looked plausible. Every backtest figure on this page has both fixed. One upward bias remains unfixed — companies that left the index — and is listed under limitations.

Scoring Formula Used (v22)

The backtest uses the v22 composite scoring formula, which blends four axes into a single 0–10 rating for every S&P 500 stock:

Composite = Quality × 0.40 + Value × 0.10 + Momentum × 0.35 + Timing × 0.15
AxisWeightWhat It Measures
Quality40%Fundamental strength via Piotroski F-Score (35%) and Buffett-style quality assessment (65%)
Value10%Valuation attractiveness via Graham Fair Value (45%), Lynch PEG (25%), Greenblatt Magic Formula (30%)
Momentum35%6-month price performance, percentile-ranked within the S&P 500
Timing15%min(Value, Momentum) — the value-trap killer, requires both cheapness and rising prices

These weights are intentional. Quality gets the largest share because long-term business strength is central to the investor models it draws on. Momentum is second because price trends capture information not yet reflected in quarterly filings. Value is deliberately modest to avoid chasing value traps. Timing ties them together. For a full explanation of each axis and the five investment models behind them, see the methodology page.

Top-10 Portfolio: Results

The test: a portfolio of the 10 highest-scoring S&P 500 stocks, rebalanced monthly and equally weighted, using only point-in-time data available on each scoring date. Over the full January 2021 through February 2026 window it returned +136.3% versus +100.3% for the S&P 500.

MeasureTop-10*S&P 500 (SPY)
Total return, full period (2021–2026)+136.3%+100.3%
Annual alpha, full period+3.9% (t = 0.71)
Annual alpha, in-sample (2021–2023)+4.8%
Annual alpha, out-of-sample (2024–Feb 2026)+2.4% (t = 0.36)
CAPM alpha, out-of-sample+0.4% (beta 1.11)
Sharpe ratio, full period0.810.73
Sharpe ratio, out-of-sample1.281.65
Maximum drawdown, full period-24.8%-23.9%

* Hypothetical backtested results. Top-10 portfolio = 10 highest-scoring S&P 500 stocks, equally weighted, rebalanced monthly on point-in-time data (no look-ahead). Simulated, not actual trading. Does not include transaction costs, taxes, or slippage.

The excess return is not statistically significant. A t-statistic measures how far a result sits from zero relative to its noise; below about 2, it cannot be told apart from luck. The full-period alpha has t = 0.71. The out-of-sample alpha of +2.4% has t = 0.36, with a 95% confidence interval of roughly ±13 percentage points — wide enough to include both a clearly negative and a clearly positive alpha.

Most of it is market exposure. The portfolio had a beta of 1.11 out-of-sample: it moved more than the market, and in a rising market that alone produces higher returns. Adjusting for that (CAPM alpha) leaves +0.4% a year — almost nothing attributable to stock selection.

Risk-Adjusted Returns: The Sharpe Ratio

Raw returns tell only half the story. The Sharpe ratio measures return per unit of risk. A higher Sharpe means the portfolio delivered more return for each unit of volatility endured.

Over the full 2021–2026 period, the top-10 portfolio's Sharpe ratio was 0.81 versus 0.73 for the S&P 500 — broadly in line with the market. In the out-of-sample window the market was the stronger risk-adjusted performer: a Sharpe of 1.65 versus the portfolio's 1.28. We do not claim better risk-adjusted returns.

For context, a Sharpe ratio above 1.0 is generally considered good for a long-only equity strategy. The S&P 500's historical Sharpe ratio typically falls between 0.3 and 0.7 depending on the measurement period.

Drawdown Analysis

Maximum drawdown measures the largest peak-to-trough decline during the test period. It answers the question every investor actually cares about: how bad did it get at the worst point?

PortfolioMaximum Drawdown
Top-10 Portfolio (full period)-24.8%
S&P 500 (SPY, full period)-23.9%

Over the full period the top-10 portfolio's worst peak-to-trough decline was -24.8% versus -23.9%for the S&P 500 — slightly deeper than the market. Despite the 40% weight on the Quality axis, a concentrated 10-stock portfolio did not protect capital better than the index in this test.

In-Sample vs. Out-of-Sample: Why It Matters

The distinction between in-sample and out-of-sample performance is the single most important concept in quantitative finance testing. Many backtests you encounter online are exclusively in-sample: the model was designed and tested on the same data. This guarantees nothing about future performance.

Here is what the split shows for our formula:

The edge shrank out-of-sample. Annual alpha fell from +4.8% in-sample to +2.4% out-of-sample. That is the pattern overfitting produces, although with this much noise the two numbers cannot be reliably told apart.
A genuine walk-forward earned nothing. The v22 weights were chosen on a window that overlaps the out-of-sample period, so their out-of-sample result is not truly out-of-sample. When we chose weights strictly on in-sample data and judged them on the later window, they earned -0.3% annual alpha out-of-sample.
The sign depends on basket size. A robust signal should help whether you hold the top 10, 20, or 30 stocks. Ours does not:
Basket (out-of-sample)Annual Alpha
Top-10+2.6%
Top-20-5.2%
Top-30-3.7%

From a separate basket-size comparison, which puts the top-10 basket at +2.6% rather than +2.4%. Every one of these results has a confidence interval that spans zero.

Known Weaknesses and Limitations

No investment strategy works in all market conditions. Transparency requires being explicit about where this formula and this test fall short:

  1. No statistically significant alpha. None of the alpha figures above can be distinguished from zero, and the out-of-sample excess is mostly market exposure.
  2. Can lag in sideways, range-bound markets. The formula relies on momentum as its second-largest axis (35%). When prices move sideways, momentum signals become noisy and unreliable, which can lead to weaker stock selection and stretches where the strategy trails a flat market. It is strongest in trending markets.
  3. Hypothetical backtested results only. The backtest is a simulation, not a track record of actual trades. No real money was invested during the test period. Simulated results are inherently more optimistic than real-world results.
  4. No transaction costs, taxes, or slippage modeled. Monthly rebalancing of 10 stocks is high turnover, and its costs plausibly exceed the small remaining excess return. Commission-free brokers still impose bid-ask spreads. Taxes on short-term capital gains (from monthly turnover) would reduce net returns. Market impact when entering and exiting positions is not reflected.
  5. Monthly rebalancing assumed. The backtest rebalances on the first trading day of each month. Real investors may rebalance at different frequencies or in response to different triggers, producing different results.
  6. Survivorship bias remains. The backtest respects when companies joined the S&P 500, but companies that were later removed (due to bankruptcy, acquisition, or relegation) are missing from our data, so their returns are not included. This biases every number on this page upward.
  7. The backtested score is not identical to the live score. The backtest ranks momentum and the Greenblatt model within the S&P 500; the site ranks them across all 1,500+ stocks it covers. The two composites differ.
  8. Past performance does not guarantee future results. Market regimes change. Correlation structures shift, monetary policy evolves, and new market dynamics emerge that historical models cannot anticipate.

What These Results Mean

The backtest does not show that the composite score predicts returns. A top-10 basket ended ahead of the S&P 500, but the gap is within the noise, most of it came from taking more market risk, and it did not hold for larger baskets or under a genuine walk-forward.

What the score is for: a consistent, transparent shortlist. The composite applies five published investor frameworks — Piotroski, Buffett, Graham, Lynch, and Greenblatt — the same way to every stock, with sector-aware calibration, and shows you why each stock scores the way it does. That saves research time. It does not replace research.

Why we publish this anyway. A backtest that has been audited for look-ahead bias and reported with its t-statistics, including what did not work, is more useful to you than a headline number. If a future test shows something different, we will update this page.

How to Use Backtest Results in Your Research

Backtest results are a starting point, not a conclusion. Here is how to incorporate them into a sound investment process:

  1. Use scores as a screening tool, not a decision tool. The stock screener and rankings help you narrow 500+ stocks to a manageable shortlist. The final investment decision should incorporate your own research, risk tolerance, and financial situation.
  2. Understand why a stock scores well, not just that it does. A composite score of 8.5 is meaningless without understanding which axes are driving it. A stock scoring 8.5 because of strong quality and momentum is a very different investment than one scoring 8.5 because of extreme cheapness.
  3. Do not ignore the limitations. If you are investing during a sideways market, the formula's momentum-heavy weighting may produce weaker signals. Be aware of when the strategy's known weaknesses are most likely to manifest.
  4. Diversify beyond the formula. Even a strategy with a strong backtest can underperform for extended periods, and this one has not shown a significant edge. Do not allocate 100% of your portfolio to any single approach.
  5. Consult a qualified financial adviser. AlphaStocks provides research tools, not personalized investment advice. Your individual circumstances — tax situation, time horizon, risk tolerance, income needs — should drive your investment decisions.

Frequently Asked Questions

What is walk-forward validation and why is it important?

Walk-forward validation is a backtesting method where a model is designed using one time period (in-sample) and then tested on a separate, future time period (out-of-sample) that was never seen during development. It matters because it reveals whether a model captures real patterns or merely fits historical noise. When we ran a genuine walk-forward on our own weight choice — picking weights on in-sample data and judging them out-of-sample — it earned -0.3% annual alpha out-of-sample.

Is the AlphaStocks backtest overfit to historical data?

We cannot rule it out, and the evidence leans that way. Annual alpha fell from +4.8% in-sample to +2.4% out-of-sample, the current v22 weights were chosen on a window that overlaps the out-of-sample period, and a genuine walk-forward of the weight choice earned -0.3% out-of-sample. None of these numbers is statistically distinguishable from zero.

Does the scoring formula always outperform the S&P 500?

No. The formula can lag in sideways, range-bound markets, where momentum signals turn noisy. The strategy works best in trending markets — both up and down — where quality and momentum factors can express themselves. No strategy outperforms in all market conditions, and anyone claiming otherwise should be treated with skepticism.

Can past backtest results predict future performance?

No. Past performance, backtested or real, does not guarantee future returns. Our own backtest shows no statistically significant alpha, so it is not evidence that the scores predict returns. Use the scores to shortlist stocks for research, not as a forecast.

Does the backtest include transaction costs and taxes?

No. The backtest is hypothetical and does not include transaction costs, taxes, slippage, or market impact. Real-world returns would be lower. Monthly rebalancing of 10 stocks generates meaningful trading costs that are not reflected in the reported numbers. For tax-advantaged accounts (like IRAs), the tax impact is reduced. For taxable accounts, the monthly turnover would generate short-term capital gains taxed at ordinary income rates.

Related Pages

Disclaimer

Important Disclaimer: All backtest results presented on this page are hypothetical and do not represent actual trading or real portfolio returns. The AlphaStocks scoring formula was backtested using historical data and simulated portfolio construction. Results do not include transaction costs, commissions, taxes, slippage, or market impact. Actual results would differ, likely unfavorably. Monthly rebalancing was assumed; actual rebalancing frequency and timing would affect results.

Past performance, whether actual or backtested, does not guarantee future results. Investment in securities involves risk, including the possible loss of principal. AlphaStocks provides algorithm-generated research tools, not personalized investment advice. Always conduct your own due diligence and consult a qualified financial adviser before making investment decisions.

Model names reference the published investment methodologies of their respective authors. AlphaStocks is not affiliated with, endorsed by, or sponsored by Warren Buffett, Berkshire Hathaway, Peter Lynch, Fidelity Investments, Joel Greenblatt, Gotham Asset Management, or the estates of Benjamin Graham or Joseph Piotroski. Data sourced from SEC EDGAR filings and Alpaca Markets.

← Back to Methodology