Validation & Digitisation

← back to chart · machine-readable evidence: target · predictions · candidates · summary

Validation terminology (corrected)

The model was fit and compared across candidates on 2010-01 → 2025-06. The window 2025-07 → 2026-06 was used to select the final 4-component model over alternatives (including a 5-component variant with risk-parity). It is therefore a holdout model-selection period, not an untouched final test. Because it was used for selection, no untouched final test remains, and the holdout metrics below are selection-period estimates, not pristine out-of-sample figures.

Selected model (4 components, equal weight, no offset)

PeriodRMSECorrelationNRole
2010–2025-060.4140.892756fit / candidate comparison
2025-07–2026-060.4190.78352holdout model-selection period
Full (2010–2026-06)0.4140.890809descriptive

Date-level predictions, residuals, split labels and digitisation uncertainty are in equity_positioning_predictions.csv; all candidate metrics are in equity_positioning_candidate_results.json.

Candidate comparison (full-period correlation)

ModelTrain corrHoldout corrFull corr
4-comp (selected): CTA+VolControl+CFTC+NAAIM0.8920.7830.890
3-comp: CTA+VolControl+CFTC0.8970.7150.894
5-comp (with risk-parity)0.8350.6910.831
All 6 (incl. hedge-fund beta)~0.85
Baseline: CFTC only0.65
Baseline: NAAIM only0.70

The 4-component model was selected because it removed the risk-parity solver artifact and improved holdout correlation (0.78 vs 0.69), while keeping the components economically interpretable.

Correction 1 — composite offset removed

An earlier build re-centred the composite to full-sample zero mean, introducing a constant ~−0.10 offset. That centring used future data (lookahead bias) and was undocumented. It has been removed: the published composite is now the literal equal-weight mean of the component Z-scores, and a pipeline assertion enforces this.

Correction 2 — risk-parity component dropped

The risk-parity ERC solver collapsed to degenerate 0.0/1.0 corner solutions (1682 days at 0.0, 436 at 1.0) during volatile correlation regimes (e.g. Jul-2011, Mar-2020). Week-to-week jumps of ±1.0 are solver instability, not signal. The component was removed, which improved every split's correlation.

Fitting target: digitised chart

The fitting target is a digitised published chart of Deutsche Bank's "Consolidated Equity Positioning" line (ISABELNET capture). Pixel calibration: +1/0/−1 reference lines at y=259/366.5/473 px (107 px/unit); 17 year ticks (Jan-10…Jan-26) → 59.66 px/year. 862 weekly points (2009-12-29 → 2026-06-30). Per-point value σ ≈ 0.07, date σ ≈ ±6 days (see the target file, which carries these uncertainties per row). Median extraction residual 0 px.

Known limitation — April 2025

The reconstruction captures the April-2025 drawdown (model trough ≈ −0.82 late-April vs digitised target ≈ −1.04 mid-April) but roughly 0.2 shallower and ~2 weeks later. This is expected: the CTA and vol-control components react to price with a lag, and no fast sentiment/flow inputs (fund flows, AAII, short interest) are available from free point-in-time sources. Capturing the speed of that specific capitulation would require inputs this public-data reconstruction deliberately does not fabricate.

Honest limitations

This is an approximation, not a recovery of Deutsche Bank's proprietary series. Correlation ~0.89 captures broad turning points and crisis behaviour, but individual weekly values differ. DB's exact component set, lookbacks and weights are proprietary. The fitting target is itself a digitised image with quantified uncertainty. Not investment advice.