Research Validation: the retired programme's walk-forward record
Every number on this page comes from the 2026-04-29 run, the last one made before the programme was retired. Nothing is carried forward from earlier runs and nothing is synthesised. Every figure is hypothetical, pre-cost, and attached to the series, window and sample size that produced it. The unedited output file, including the statistical tests recorded at that run, is published at /data/signals.json.
- Status
- Retired after the last run
- Last run
- 2026-04-29 · V12
- Evaluation window
- 2026-04-29 → 2026-05-29 (closed)
- Realised outcome
- Not computed; no evaluation published
Sample: 30 U.S. mega-caps; training lookback about three years of daily data; walk-forward out-of-sample panel 2024-09-30 → 2026-03-31 (19 months); prediction horizon one month. The page date is not a model date. Definitions and provenance of every metric are on the Research Validation page, and the unedited output file is published at /data/signals.json.
HYPOTHETICAL PERFORMANCE RESULTS have many inherent limitations. No representation is being made that any account will or is likely to achieve profits or losses similar to those shown. Past performance, whether actual or hypothetical, is not indicative of future results.
Read full disclaimerEvery CAGR, Sharpe, drawdown and Calmar figure below is from a hypothetical long-short equities portfolio: long the top quintile of estimates, short the bottom quintile, equally weighted, monthly rebalance. No leverage, no stock borrow cost, no slippage, no spread. They are pre-cost upper bounds.
The options overlay on the Lab page is a forward Monte Carlo from the implied-vol surface at run time, not a historical test. A historical options backtest was attempted on 2026-04-24 against a broker's historical options API and was blocked at contract resolution for expired contracts, so no historical options figure is published anywhere on this site.
Metric provenance
Two Sharpe-like numbers exist in the output file and they are not the same statistic: the long-short Sharpe is computed on monthly portfolio returns; the validation battery uses the Sharpe of the monthly IC series. The factor panel is longer than the OOS panel because factor ICs are measured over every month a factor can be computed, while the walk-forward needs a training window before its first out-of-sample month.
| Metric | Value | Series | Window | n | Method |
|---|---|---|---|---|---|
| L/S Sharpe | 0.74 | monthly top-quintile minus bottom-quintile return, equal-weight, pre-cost | 2024-09-30 → 2026-03-31 | 19 | mean / sd × √12 |
| CAGR · Max drawdown · Calmar · Total return | 19.93% · -17.58% · 1.13 · 33.34% | same monthly L/S series, compounded | 2024-09-30 → 2026-03-31 | 19 | geometric, 12/n annualisation |
| Mean IC · t-stat | 0.0271 · 0.41 | monthly rank correlation of estimate vs realised 1-month return | 2024-09-30 → 2026-03-31 | 19 | t = mean / (sd/√n) |
| Sharpe of the IC series (validation-battery input) | 0.322 (recomputed 0.322) | the monthly IC series above, not the L/S return series | 2024-09-30 → 2026-03-31 | 19 | mean / sd × √12; a different statistic from the L/S Sharpe |
| Hit rate · RMSE | 53.7% · 0.1123 | stock-month sign accuracy and prediction error | 2024-09-30 → 2026-03-31 | 570 stock-months | pooled |
| Factor mean IC · p-value | per factor, table below | monthly IC of each raw factor over the factor panel | factor panel, longer than the OOS panel | 25 | t-test on the monthly series; "significant" = p < 0.10 |
Walk-forward OOS panel: equities long-short
Monthly OOS information coefficient
19 months · mean = 0.027Information coefficient: rank correlation between the one-month estimate and the realised cross-sectional return. M1 is 2024-09-30; M19 is 2026-03-31. With 19 data points, individual months carry little statistical weight.
Hypothetical long-short equities portfolio
Long top quintile, short bottom quintile, equally weighted within each leg, rebalanced monthly. No leverage, no execution costs, no borrow cost. Stocks only. Point estimates; with 19 months every interval is wide.
Cumulative path
start = 1.00Walk-forward path, not a smoothed backtest. A single month (2026-03-31) contributes most of the total return; remove it and the picture changes materially.
Regime-conditional performance
Point-in-time VIX| Regime | Months | % of panel | Mean monthly L/S | Annualised Sharpe | Hit rate |
|---|---|---|---|---|---|
| Moderate | 19 | 100% | +1.86% | +0.74 | 53% |
VIX regime at the end of each prediction month. Every month in this panel fell in the Moderate regime, so the decomposition is a single row and says nothing about behaviour in other regimes.
Factor information coefficients
4 kept · 7 dropped · 25-month factor panel| Factor | Mean IC | t-stat | p-value | n months | p < 0.10 | Kept |
|---|---|---|---|---|---|---|
| low_vol_60d | -0.112 | -1.70 | 0.101 | 25 | no | ✓ |
| beta_residual | +0.102 | 1.84 | 0.079 | 25 | yes | ✓ |
| sector_neutral_momentum | +0.092 | 1.66 | 0.110 | 25 | no | ✓ |
| momentum_12_1 | +0.088 | 1.38 | 0.179 | 25 | no | ✓ |
| price_acceleration | -0.033 | -0.58 | 0.569 | 25 | no | ✗ |
| volume_trend | +0.031 | 0.77 | 0.447 | 25 | no | ✗ |
| ret_5d | -0.030 | -0.55 | 0.588 | 25 | no | ✗ |
| momentum_3_1 | +0.028 | 0.47 | 0.641 | 25 | no | ✗ |
| volume_shock | -0.022 | -0.53 | 0.603 | 25 | no | ✗ |
| short_term_reversal | -0.022 | -0.37 | 0.711 | 25 | no | ✗ |
| volatility_ratio | +0.021 | 0.74 | 0.468 | 25 | no | ✗ |
Each factor is tested cross-sectionally per month over a 25-month factor panel. “p < 0.10” is the significance rule the screen actually applies (a t-test on the monthly IC series at the 10% level), and “Kept” is a separate criterion: p < 0.10 or |mean IC| ≥ 0.05. At this sample size most factors fall short of individual significance even when an ensemble has value.
Note on low_vol_60d: defined as the negative of 60-day realised vol so that, under the classical low-vol anomaly, high factor values would map to low-vol stocks and positive forward returns. The panel IC came out negative: in this mega-cap sample high-vol names outperformed over the panel. The screen kept the factor on the magnitude criterion (|IC| ≥ 0.05), not the direction. Treat it as regime-specific.
Ensemble and universe at the run
Model weights
Source: inverse_oos_mae