BlackArbs LLC · Alpha Lab Research
Corrfilter forensic reconstruction · July 2026

Deconstructing a Failed Strategy

Was Anything Worth Saving?

2025Strategy stopped after performance discrepancies undermined confidence.
55Distinct market stress periods examined. One statistic kept hiding two opposing signals.
18Portfolio designs locked before the 2022 test. All 18 lost to SPY.
Thesis

Corrfilter was not worth saving, but its failure revealed a safer allocation rule and proved the research process could be trusted.

The safer allocation rule

  • Hedge quality may reduce the proposed TMF position.
  • It may never increase leveraged-equity exposure in UPRO or TQQQ.

The research process we could trust

  • Every allocation, return, and benchmark became traceable.
  • Deconstruction traced the failure from signal design to capital allocation.
  • A strict train-validation-test split let us reject the strategy in validation before burning the pure, unseen test set reserved for one final use.

The strategy failed. The allocation lesson and the evidence standard survived.

What was worth saving
Corrfilter in one minute

A rule for deciding how much leveraged market risk to own

1 · What it watched

  • UPRO and TQQQ: leveraged U.S. equity funds.
  • TMF: leveraged long-term Treasuries, intended to diversify equity losses.
  • It measured how these assets moved together.

2 · What it decided

  • A single “correlation pressure” score decided how much of the portfolio could remain in leveraged assets.
  • More perceived danger meant less money in leveraged assets and more in SHY, a short-term Treasury fund.

3 · Why we reopened it

  • The reported results could not be trusted consistently.
  • But the underlying question remained useful: when do several positions become one concentrated equity bet?

Corrfilter was not a market forecast. It was a portfolio risk-control rule. This research asks whether any part of that rule deserves to survive.

Working glossary
BlackArbs LLC · Research transition
I
Section I · Rebuild trust

Before testing the idea, prove the evidence.

Can every return, allocation, and benchmark be traced to the same run?

Evidence before interpretation

First, prove the results describe the same run

Missing Benchmark
If the leveraged-ETF strategy (LETF) using Corrfilter could not beat SPY, the signal was not worth the added complexity. The verification backtest never included SPY. Without the benchmark, we could not know.
Missing Actions
The strategy generated daily target weights for UPRO, TQQQ, TMF, and SHY. The reporting layer dropped part of that allocation record. We could see the performance, but not the portfolio decisions that produced it.
Mismatched Benchmark
The strategy ran a benchmark, but the report ID could not prove which benchmark belonged to that run. It could attach SPY data from another calculation, including a different time period. The comparison looked valid even when it was not.

Before interpreting performance, we had to prove one chain: market data → target weights → executed portfolio → returns → same-period SPY → final report.

Why trust broke
The wider repair

The benchmark was only one broken link

Could Not Launch
The platform was upgraded. Corrfilter was not. The current system rejected some of the files required to run it.
Could Not Read
The backtest finished, but its results used an old file format that the current application could not open.
Wrong Interface
The strategy and reporting system disagreed about what a portfolio decision looked like. Valid allocations disappeared, and the application displayed analyses the strategy never created.

The same strategy identity, inputs, portfolio decisions, and returns had to survive intact from launch through the final report.

System reconstruction
BlackArbs LLC · Research transition
II
Section II · The tempting redesign

The evidence held. The signal still mixed two opposing facts.

Splitting them looked smarter. Did it improve the signal, or merely add equity risk?

Signal autopsy

One number was answering two different questions

Equity co-movement

Are the risky holdings becoming one crowded trade?

TMF hedge quality

Is Treasury exposure behaving like insurance?

The legacy statistic averaged these facts together. A stronger hedge could make elevated equity concentration look safer.

Cancellation
The 2020 contradiction

Risk concentration was high while the legacy warning stayed quiet

Median daily reading, Feb to Apr 2020Value
Equity pairwise correlation0.925
Equity stress0.861
TMF/equity correlation-0.581
TMF hedge quality0.791
Old mixed-signal pressure-0.409
Legacy binary signal0

What capital actually did

  • Defensive on March 12.
  • Fully risk-on again March 17.
  • Defensive again March 19.
  • Strategy trough: March 19.
  • SPY trough: March 23.

The portfolio followed the signal exactly. The signal cut risk late, restored it while equity stress stayed high, then cut again near the bottom.

Mechanism
2020 signal timing

The gate stayed risk-on until most of the damage was done

Equity stress, legacy binary gate, and equity sleeve drawdown

Equity stress stayed high through the selloff. The binary gate remained risk-on through most of the decline, then flipped twice near the trough.

Late defense
Descriptive edge

The old signal mixed equity stress with hedge quality. It erased both return signals.

Spearman time-series Rank IC by forward horizon

Higher equity-stress ranks preceded weaker 5- and 20-session returns. Higher TMF hedge-quality ranks preceded stronger 20- and 60-session returns. Mixing them compressed both associations toward zero.

2018–2021 in sample · descriptive, not causal
The conflicted state, in plain language

The old score let good insurance hide rising equity danger.

Equity danger rose

  • UPRO and TQQQ were moving together.
  • Several leveraged positions were becoming one concentrated equity bet.

The hedge still worked

  • TMF was still moving against equities.
  • Treasury exposure looked capable of absorbing part of an equity loss.

The old score averaged both

  • The useful hedge offset the equity warning inside one number.
  • The combined signal could stay quiet while equity concentration remained high.

A useful hedge should influence where defensive capital goes. It should not cancel the warning that determines how much equity risk to own.

Separate responsibilities
BlackArbs LLC · Research side quest
Open every historical conflict

We traced every 2018–2021 period when both conditions were high.

What happened to the leveraged equity basket from the day before equity stress turned high while hedge quality remained strong?

Historical event paths

The conflict usually began with a gain. Longer conflicts could become expensive.

Twenty-four observed equal-weight UPRO and TQQQ return paths during high-equity-stress and high-hedge-quality periods

Seventeen of 24 paths rose on the first conflicted day, averaging +0.96%. But that return had already happened when the state became observable. The paths then diverged sharply.

2018–2021 · historical paths, not simulations
BlackArbs LLC · Research side quest
Could timing rescue the warning?

The first conflicted day usually gained. That return was already gone.

Would waiting for the conflict to persist filter one-day noise and isolate the longer deterioration?

Timing side quest · verdict

Waiting did not isolate a cleaner warning.

Five-session and twenty-session returns after immediate, two-of-three, and three-consecutive confirmation rules

Waiting three sessions discarded 8 of 24 warnings without improving what followed. The timing shortcut failed. We returned to the structural repair.

Side quest closed
BlackArbs LLC · Main research resumes
Side quest closed

Timing could not rescue a score with two jobs.

So we returned to the structural test: let equity stress size leveraged exposure, while hedge quality handles the hedge.

Back to the main test

Six rules changed one decision: how much leveraged equity risk should stress permit?

Same input

  • Each rule received the same equity-stress reading.
  • Hedge quality no longer canceled the equity warning.

One variable

  • Only the mapping from stress to permitted leveraged exposure changed.
  • Six rules expressed six different sizing responses.

Everything else fixed

  • Same one-day delay and inverse-volatility inputs.
  • Same costs, accounting, and portfolio construction.

This was a controlled sizing experiment: six exposure rules, not six signals or six standalone strategies.

Main experiment resumes
Six exposure rules

How the six rules translate stress into portfolio exposure

Controls

  • Old mixed signal: combines equity stress with hedge quality to reproduce the old throttle.
  • Full risk/no action: ignores stress; exposure remains 100%.

Equity-stress mappings

  • Linear: continuous inverse response.
  • Sparse piecewise-linear: flat regions with linear transitions.
  • Shifted logistic: smooth S-shaped throttle.

Distribution-aware

  • Historical-percentile rule: reduces exposure according to how extreme the stress reading was in the development data.
  • All thresholds and calibration settings were locked before the 2022 test.

Every rule used the same data, one-day delay, inverse-volatility-based weights, trading costs, and accounting. Only the exposure rule changed.

Common economics
2018–2021 · In sample

The redesign looked smarter. It was mostly longer equities.

Matched equity-linear architecture comparison of return, drawdown, and average equity weight

Return and the Geometric Martin ratio, which rewards return and penalizes deep drawdowns, both improved. But average equity weight rose from 42% to 71%, while maximum drawdown nearly doubled.

Architecture comparison · descriptive, not causal
BlackArbs LLC · Research transition
III
Section III · Unseen data exposes hidden leverage

One rule looked spectacular in development. All six were locked before validation.

Would the ranking survive the untouched 2022 test, or was the redesign a backtest-only winner?

2022 · Out-of-sample test · rules locked beforehand

2022 broke the redesign. All six rules got rekt.

Losses from the original and redesigned portfolios across six rules locked before the 2022 test

Every redesigned rule lost more and drew down further than its matching original rule. Six rules tested. Six matched failures.

2022 · Out-of-sample test · descriptive, not causal
The failure mechanism

About 22 percentage points moved from SHY into leveraged equities

Average 2022 portfolio weight · legacy

UPRO + TQQQ share of $100$34.2
TMF share of $100$22.2
SHY share of $100$43.6
2022 cumulative net return-61.1%
2022 maximum drawdown-62.2%

Average 2022 portfolio weight · separated

UPRO + TQQQ share of $100$56.4
TMF share of $100$22.8
SHY share of $100$20.9
2022 cumulative net return-73.6%
2022 maximum drawdown-75.4%

The allocation logic was correct. The risk budget changed meaning: the same 56.4% that funded leveraged equities and Treasuries in the legacy design funded leveraged equities alone in the redesign.

Same number. More equity risk.
BlackArbs LLC · Research transition
IV
Section IV · Preserve the capital invariant

Validation found the real mistake. The redesign moved equity capital.

Could we leave equity exposure unchanged and reduce only the proposed Treasury hedge?

The capital rule we had to preserve

Keep equities unchanged. Adjust only the proposed Treasury hedge.

proposal = allocator(...) weights = proposal.copy() # Equities stay exactly where the allocator put them. weights["UPRO"] = proposal["UPRO"] weights["TQQQ"] = proposal["TQQQ"] # Hedge quality can only reduce proposed TMF. weights["TMF"] = proposal["TMF"] * hedge_quality rejected_tmf = proposal["TMF"] - weights["TMF"] # Rejected hedge capital moves to safety, not equities. weights["SHY"] = proposal["SHY"] + rejected_tmf assert sum(weights.values()) == 1.0

What can never happen

  • Rejected TMF cannot increase UPRO.
  • Rejected TMF cannot increase TQQQ.
  • TMF cannot exceed its original proposal.
  • Rejected weight can only move to the safe destination.

The correction separated equity risk from hedge quality without accidentally increasing equity exposure.

Capital invariant
2022 · Out-of-sample test · rules locked beforehand

The correction recovered 5 to 7 percentage points of 2022 return. The best rule still trailed SPY by 21.

Corrected TMF routing versus matched legacy policies and the SPY production benchmark in 2022

Every correction helped. None came close to the benchmark. SPY is an external comparison, not another strategy design.

2022 · Out-of-sample test · no promotion
Risk-adjusted reality

The bootstrap changed the ranking. It did not change the verdict.

Bootstrap Geometric Martin uncertainty for six corrected policies and the external SPY benchmark

After repeated resampling, all six medians were negative. Five of six 90% uncertainty ranges remained entirely negative. The sixth included zero.

2022 · Out-of-sample test · resampling uncertainty
BlackArbs LLC · Research transition
V
Section V · Reject and protect

Repairing the boundary improved the answer. It did not make the strategy investable.

Reject the candidate strategy, preserve the process, and protect the unopened holdout.

Current conclusion

The research process survived the 2022 test. Corrfilter did not.

1 · Audit

  • One evidence chain tied source data, allocations, returns, and SPY to the same run.
  • The repaired pipeline launched, rendered, and reproduced a real signal candidate.
  • Bad evidence was fixed before performance was interpreted.

2 · Isolate

  • Equity stress and TMF hedge quality were decomposed before capital was moved.
  • The in-sample redesign looked smarter by owning 29 percentage points more equity.
  • The architecture changed the answer. The chart exposed it.

3 · Gate

  • The untouched 2022 test rejected all six redesigned rules.
  • The corrected allocation logic helped, but the best rule still trailed SPY by 21 percentage points.
  • No candidate earned access to the true holdout.

The missing gate was investor usability: an active strategy must justify its costs and drawdown against buy-and-hold SPY.

Validation gate closed
2023–mid-2025 · Final test remains unopened

2022 killed the strategy. Do not spend the final test trying to revive it.

1 · Define what investors need

  • Keep maximum drawdown near the period-matched SPY benchmark.
  • Set a hard drawdown ceiling for the intended investor.
  • Make costs and recovery time part of the pass or fail decision.

2 · Prove the strategy earns its complexity

  • Measure daily correlation, beta, and downside beta to SPY.
  • Test drawdown overlap, active return, tracking error, and recovery.
  • Compare against unlevered and partially leveraged controls.

3 · Protect the final answer

  • Lock the signal, allocation rules, benchmark, and pass or fail thresholds before looking.
  • Write down exactly what would earn promotion before opening 2023 through mid-2025.
  • Open the final test once. Do not tune the strategy after seeing the result.

The final test gets one job: answer a capital question written in advance. It is not another chance to repair the strategy.

Final test remains unopened
BlackArbs LLC · Validation verdict
What the experiment actually proved

We did not rescue Corrfilter. We proved the pipeline could repair the evidence, expose leverage disguised as alpha, and reject the candidate strategy before opening the true holdout.

AuditMake every return, allocation, and benchmark traceable to one run.
IsolateSeparate what the signal measures, how the portfolio uses it, and how much market risk it creates.
GateDefine investor-usable risk first. Open the true holdout once.
← → · Space · Home · End · Click charts to zoom