By Maksym Lytvynov, Founder of AlphaStocks | Last updated: March 2026
Walk-Forward Backtest Results
Does the AlphaStocks scoring formula actually work on data it has never seen? This page presents the full walk-forward validation results for our composite scoring system. A portfolio of the top-10 highest-scoring S&P 500 stocks delivered +140.3% total returnversus +100.3% for the S&P 500 over January 2021 through February 2026, with +6.7% annual alpha during the out-of-sample period the formula had never seen during development.
Below you will find the methodology, complete results tables, drawdown analysis, honest limitations, and the full disclaimer. We disclose everything because transparency is not optional when real money is at stake.
What Is Walk-Forward Validation?
Walk-forward validation is the gold standard for testing quantitative investment strategies. The concept is straightforward: you split your data into two non-overlapping periods. The first period (in-sample) is used to design, calibrate, and optimize your model. The second period (out-of-sample) is used to test it on data the model has never seen.
This approach directly addresses the most dangerous problem in quantitative finance: overfitting. With enough parameters, any model can be made to look spectacular on historical data. A model that “predicts the past” with 99% accuracy might be completely useless going forward because it memorized noise rather than discovering genuine patterns.
Walk-forward validation prevents this by separating discovery from confirmation. If a strategy performs well on data it was specifically designed for, that tells you very little. If it performs well on data it has never seen, that is meaningful evidence that the underlying signals are real.
Most retail trading tools skip this step entirely. They show you one set of backtested numbers without disclosing whether the model was designed on the same data being used to evaluate it. That is like a student writing the test questions and then claiming a perfect score proves intelligence.
How We Structured Our Test
AlphaStocks split the January 2021 through February 2026 period into two distinct windows:
| Phase | Period | Purpose |
|---|---|---|
| In-sample | Jan 2021 – Dec 2023 | Reference window spanning varied regimes (post-COVID rally, 2022 bear, 2023 recovery) used to study the scoring formula. |
| Out-of-sample | Jan 2024 – Feb 2026 | Held-out window the scoring was not tuned on. The formula runs forward unchanged to see whether the edge persists. |
The in-sample period includes the post-COVID rally, the 2022 bear market, and the 2023 recovery — a diverse set of market conditions that stress-tested the formula across multiple regimes.
We evaluate the scoring on the held-out out-of-sample window, which it was not tuned on. That period (2024 through early 2026) includes the AI-driven tech rally, sector rotations, and several correction events — the formula runs forward across all of them unchanged.
Scoring Formula Used (v22)
The backtest uses the v22 composite scoring formula, which blends four axes into a single 0–10 rating for every S&P 500 stock:
Composite = Quality × 0.40 + Value × 0.10 + Momentum × 0.35 + Timing × 0.15| Axis | Weight | What It Measures |
|---|---|---|
| Quality | 40% | Fundamental strength via Piotroski F-Score (35%) and Buffett-style quality assessment (65%) |
| Value | 10% | Valuation attractiveness via Graham Fair Value (45%), Lynch PEG (25%), Greenblatt Magic Formula (30%) |
| Momentum | 35% | 6-month price performance, percentile-ranked within the S&P 500 |
| Timing | 15% | min(Value, Momentum) — the value-trap killer, requires both cheapness and rising prices |
These weights are intentional. Quality gets the largest share because long-term business strength is the most reliable predictor of stock returns. Momentum is second because price trends capture information not yet reflected in quarterly filings. Value is deliberately modest to avoid chasing value traps. Timing ties them together. For a full explanation of each axis and the five investment models behind them, see the methodology page.
Top-10 Portfolio: Full-Period Results
The headline test: a portfolio of the 10 highest-scoring S&P 500 stocks, rebalanced monthly and equally weighted, using only point-in-time data available on each scoring date (no look-ahead). Over the full January 2021 through February 2026 window it returned +140.3% versus +100.3% for the S&P 500.
Total Return
+140.3%
S&P 500 (SPY)
+100.3%
Annual Alpha
+4.3%
The top-10 portfolio meaningfully outpaced the S&P 500 over the same period. An annual alpha of +4.3% means that each year, on average, the portfolio outperformed the benchmark by 4.3 percentage points. Over five years, that compounds into a substantial gap (+140.3% versus +100.3%).
This result reflects disciplined monthly reselection: each month the formula re-ranks the S&P 500 on point-in-time data and holds the top 10. The scoring consistently surfaced high-quality companies at attractive prices with strong momentum across the five-year window.
* Hypothetical backtested results. Does not reflect actual trading. No transaction costs, taxes, or slippage included.
Top-10 Portfolio: Monthly Rebalanced Results
This table breaks the same top-10 portfolio down by period. Each month, the 10 highest-scoring S&P 500 stocks are selected on point-in-time data, equally weighted, and held until the next rebalance. Splitting the window into in-sample and out-of-sample tests whether the formula keeps working on data it was not tuned on.
| Period | Top-10* | S&P 500 (SPY) | Ann. Alpha | Sharpe (port vs SPY) |
|---|---|---|---|---|
| In-sample (2021–2023) | +44.8% | +34.4% | +2.9% | 0.48 vs 0.38 |
| Out-of-sample (2024–2026) | +63.4% | +46.8% | +6.7% | 1.39 vs 1.65 |
| Full period (2021–2026) | +140.3% | +100.3% | +4.3% | 0.81 vs 0.73 |
* Hypothetical backtested results. Top-10 portfolio = 10 highest-scoring S&P 500 stocks, equally weighted, rebalanced monthly on point-in-time data (no look-ahead). Simulated, not actual trading. Does not include transaction costs, taxes, or slippage.
Several observations stand out from this data. The top-10 portfolio outperformed the S&P 500 in both the in-sample and out-of-sample periods. Notably, the out-of-sample annual alpha (+6.7%) was higher than the in-sample alpha (+2.9%) — the opposite of the pattern you see with overfit strategies, whose edge shrinks or vanishes on unseen data. That the alpha held up, rather than decaying, on data the formula was not tuned on is the single strongest signal here.
The full-period annual alpha of +4.3% compounds into a roughly +40 percentage-point gap in cumulative return (+140.3% versus +100.3%), demonstrating the power of consistent, modest annual outperformance over multiple years.
Risk-Adjusted Returns: The Sharpe Ratio
Raw returns tell only half the story. The Sharpe ratio measures return per unit of risk. A higher Sharpe means the portfolio delivered more return for each unit of volatility endured.
Over the full 2021–2026 period, the top-10 portfolio's Sharpe ratio was 0.81versus 0.73 for the S&P 500 — a modest edge, broadly in line with the market rather than dramatically better. Risk-adjusted, the portfolio delivered its higher returns without taking on materially more volatility.
We want to be precise here: in the out-of-sample window specifically, the market was actually the stronger risk-adjusted performer (SPY Sharpe 1.65 versus the portfolio's 1.39). The portfolio's out-of-sample edge came from higher raw returns, not from a better Sharpe. The anti-overfitting evidence lives in the return alpha — which grew from +2.9% in-sample to +6.7% out-of-sample — not in the Sharpe ratio.
For context, a Sharpe ratio above 1.0 is generally considered good for a long-only equity strategy. The S&P 500's historical Sharpe ratio typically falls between 0.3 and 0.7 depending on the measurement period.
Drawdown Analysis
Maximum drawdown measures the largest peak-to-trough decline during the test period. It answers the question every investor actually cares about: how bad did it get at the worst point?
| Portfolio | Maximum Drawdown |
|---|---|
| Top-10 Portfolio (full period) | -21.6% |
| S&P 500 (SPY, full period) | -23.9% |
Over the full period the top-10 portfolio's worst peak-to-trough decline was -21.6% versus -23.9%for the S&P 500 — a slightly shallower drawdown, broadly comparable to the market rather than dramatically better. We should be candid: this is a modest edge, and it was not uniform. In the 2024–2026 out-of-sample window the market's worst drawdown was actually shallower than the portfolio's (-7.7% for the S&P 500 versus -10.8% for the portfolio).
This matters more than most investors realize. A 23.9% drawdown requires a 31.4% recovery to break even; a 21.6% drawdown requires about a 27.6% recovery. Smaller drawdowns mean faster recoveries and less temptation to panic-sell at the bottom.
The modestly lower full-period drawdown is likely a consequence of the Quality axis receiving the largest weight (40%). High-quality companies with strong balance sheets and consistent profitability tend to decline less during market stress. The formula is inherently biased toward resilient businesses, and that bias shows up most clearly during corrections.
In-Sample vs. Out-of-Sample: Why It Matters
The distinction between in-sample and out-of-sample performance is the single most important concept in quantitative finance testing. Many backtests you encounter online are exclusively in-sample: the model was designed and tested on the same data. This guarantees nothing about future performance.
Here is what our walk-forward split reveals:
Consider what an overfit strategy would look like: spectacular in-sample returns followed by zero or negative out-of-sample alpha. Our results show the opposite pattern — the return alpha grew on unseen data rather than collapsing. Risk-adjusted, the picture is more even-handed: over the full period the portfolio's Sharpe and drawdown were broadly in line with the market's, and in the out-of-sample window the market was actually the stronger risk-adjusted performer. The honest headline is stronger returns at comparable risk, with the anti-overfitting signal carried by the alpha.
Known Weaknesses and Limitations
No investment strategy works in all market conditions. Transparency requires being explicit about where this formula struggles:
- Can lag in sideways, range-bound markets. The formula relies on momentum as its second-largest axis (35%). When prices move sideways, momentum signals become noisy and unreliable, which can lead to weaker stock selection and stretches where the strategy trails a flat market. It is strongest in trending markets.
- Hypothetical backtested results only. The backtest is a simulation, not a track record of actual trades. No real money was invested during the test period. Simulated results are inherently more optimistic than real-world results.
- No transaction costs, taxes, or slippage modeled. Monthly rebalancing of 10 stocks generates meaningful trading costs. Commission-free brokers still impose bid-ask spreads. Taxes on short-term capital gains (from monthly turnover) would reduce net returns. Market impact when entering and exiting positions is not reflected.
- Monthly rebalancing assumed. The backtest rebalances on the first trading day of each month. Real investors may rebalance at different frequencies or in response to different triggers, producing different results.
- S&P 500 universe introduces survivorship bias. The formula scores current S&P 500 constituents. Companies that were in the index during the test period but later removed (due to bankruptcy, acquisition, or relegation) are not fully accounted for.
- Past performance does not guarantee future results. Market regimes change. The factors that drove alpha in 2021–2026 may not persist in future periods. Correlation structures shift, monetary policy evolves, and new market dynamics emerge that historical models cannot anticipate.
What These Results Mean
The backtest results suggest — but do not prove — that the AlphaStocks composite formula identifies genuinely useful investment signals. Specifically:
The combination of quality, value, momentum, and timing factors appears to produce returns above the benchmark across multiple market conditions. The formula outperformed in the post-COVID rally, the 2022 bear market, and the 2024–2025 AI-driven bull run. It can lag in sideways periods.
The value-trap killer (Timing = min(Value, Momentum)) may help contain drawdowns. By requiring both cheapness and rising prices, the formula avoids the classic trap of buying stocks that are cheap and getting cheaper. This may contribute to the portfolio's modestly lower full-period drawdown, though risk was broadly comparable to the market overall.
The out-of-sample return alpha was actually stronger than in-sample, not weaker. A strategy whose return edge holds up — even grows — on data it was not tuned on is a more credible signal than spectacular in-sample numbers that collapse forward. Risk-adjusted, the out-of-sample market was competitive; the durable-alpha claim rests on returns, not the Sharpe ratio.
These results do not mean the formula will continue to outperform. They mean it has demonstrated the ability to do so in a controlled, honestly disclosed test.
How to Use Backtest Results in Your Research
Backtest results are a starting point, not a conclusion. Here is how to incorporate them into a sound investment process:
- Use scores as a screening tool, not a decision tool. The stock screener and rankings help you narrow 500+ stocks to a manageable shortlist. The final investment decision should incorporate your own research, risk tolerance, and financial situation.
- Understand why a stock scores well, not just that it does. A composite score of 8.5 is meaningless without understanding which axes are driving it. A stock scoring 8.5 because of strong quality and momentum is a very different investment than one scoring 8.5 because of extreme cheapness.
- Do not ignore the limitations. If you are investing during a sideways market, the formula's momentum-heavy weighting may produce weaker signals. Be aware of when the strategy's known weaknesses are most likely to manifest.
- Diversify beyond the formula. Even a backtested-validated strategy can underperform for extended periods. Do not allocate 100% of your portfolio to any single approach.
- Consult a qualified financial adviser. AlphaStocks provides research tools, not personalized investment advice. Your individual circumstances — tax situation, time horizon, risk tolerance, income needs — should drive your investment decisions.
Frequently Asked Questions
What is walk-forward validation and why is it important?
Walk-forward validation is a backtesting method where a model is designed using one time period (in-sample) and then tested on a separate, future time period (out-of-sample) that was never seen during development. It is the gold standard for evaluating quantitative strategies because it reveals whether a model genuinely captures market patterns or merely overfits to historical noise. Without walk-forward validation, backtested results are essentially meaningless — any model can be made to look good on data it was designed to explain.
Is the AlphaStocks backtest overfit to historical data?
The strongest evidence against overfitting is that the out-of-sample return alpha (+6.7% annualized) actually exceeded the in-sample alpha (+2.9%), the opposite of what you see with overfit strategies, whose edge collapses on unseen data. Additionally, the formula uses only five well-documented investment models with fixed weights, rather than hundreds of data-mined parameters. The formula was intentionally kept simple to reduce overfitting risk. That said, no backtest can definitively prove a strategy is not overfit — only future live performance can do that.
Does the scoring formula always outperform the S&P 500?
No. The formula can lag in sideways, range-bound markets, where momentum signals turn noisy. The strategy works best in trending markets — both up and down — where quality and momentum factors can express themselves. No strategy outperforms in all market conditions, and anyone claiming otherwise should be treated with skepticism.
Can past backtest results predict future performance?
No. Past performance, including walk-forward validated results, does not guarantee future returns. Market conditions change, correlations shift, and strategies that worked historically may not work in the future. What walk-forward validation does provide is evidence that a strategy captured real market patterns rather than noise. That evidence is useful context for research, but it is not a prediction.
Does the backtest include transaction costs and taxes?
No. The backtest is hypothetical and does not include transaction costs, taxes, slippage, or market impact. Real-world returns would be lower. Monthly rebalancing of 10 stocks generates meaningful trading costs that are not reflected in the reported numbers. For tax-advantaged accounts (like IRAs), the tax impact is reduced. For taxable accounts, the monthly turnover would generate short-term capital gains taxed at ordinary income rates.
Related Pages
Disclaimer
Important Disclaimer: All backtest results presented on this page are hypothetical and do not represent actual trading or real portfolio returns. The AlphaStocks scoring formula was backtested using historical data and simulated portfolio construction. Results do not include transaction costs, commissions, taxes, slippage, or market impact. Actual results would differ, likely unfavorably. Monthly rebalancing was assumed; actual rebalancing frequency and timing would affect results.
Past performance, whether actual or backtested, does not guarantee future results. Investment in securities involves risk, including the possible loss of principal. AlphaStocks provides algorithm-generated research tools, not personalized investment advice. Always conduct your own due diligence and consult a qualified financial adviser before making investment decisions.
Model names reference the published investment methodologies of their respective authors. AlphaStocks is not affiliated with, endorsed by, or sponsored by Warren Buffett, Berkshire Hathaway, Peter Lynch, Fidelity Investments, Joel Greenblatt, Gotham Asset Management, or the estates of Benjamin Graham or Joseph Piotroski. Data sourced from SEC EDGAR filings and Alpaca Markets.