BlackLeafwatch the watchmen
Backtest forensics

The edge that wasn't: how a trading signal 'worth' 485 basis points a month died of one median

Summary

A short-selling signal built from free FINRA data appeared to beat the market by nearly 5% a month, with the statistical significance quant funds advertise. One robustness check — a median where the mean had been — showed it was volatility wearing a costume.

By Marcus Aurelius · July 9, 2026

Every day, FINRA publishes a file listing, for every US stock, how many of the day's shares were sold short. It is free, public, and machine-readable — exactly the kind of data a thousand retail quants and a hundred newsletter backtests are built on. We built the obvious signal from it: each stock's short-sale share of volume, scored against its own recent history, across 121,148 ticker-days — 260 US stocks over 501 trading sessions from July 2024 through July 2026. Then we asked the only question that matters: did the stocks the signal ranked highest go on to beat the stocks it ranked lowest?

Measured the way backtests are usually reported, the answer was spectacular.

Apparent edge (mean spread)
+485 bps
Its statistical significance
t = 3.07
The typical outcome (median)
+9 bps

The finding that asked to be believed

Sort the 260 stocks each day by the signal. Buy the top fifth, short the bottom fifth, hold 21 sessions, measure against the S&P 500 (SPY). The mean gap between the top and bottom buckets came to +484.8 basis points per 21 sessions — and not as statistical noise: with standard errors corrected for the overlapping holding windows (Newey & West, 1987), the t-statistic was 3.07. Finance has spent a decade arguing that the traditional significance bar of t = 2 is too low precisely because so many strategies are tried; Harvey, Liu & Zhu (2016) proposed demanding t > 3. This cleared even that.

A signal from free public data, beating the market by five percent a month, at t > 3. If a fund's marketing deck showed you this, it would not be lying about a single number.

One line of arithmetic

The same portfolio, on the same days, summarized with a median instead of a mean, earned +8.9 basis points — a rounding error. And the daily rank correlation between the signal and what stocks actually did next — the statistic that measures whether a ranking ranks — was +0.0017 (t = 0.28): zero. A ranking with no ranking power was somehow producing five percent a month. Both numbers cannot be telling the truth about skill.

The same portfolio, measured three ways
Top-minus-bottom quintile spread, bps per 21 sessions, Jul 2024 – Jun 2026
Mean spread (as a backtest would report it)
484.8
Median spread (same portfolio, same days)
8.9
Mean spread, 68 most-liquid names only
-12.7
Source: FINRA Reg SHO daily short sale volume files; daily closes via Databento EQUS.MINI; SPY-excess returns; 462 evaluation days
View data as table
Spread summaries and the rank IC
Mean Q5−Q1 spread, 260 names+484.8 bps / 21 sessionsNW t = 3.07
Median Q5−Q1 spread, 260 names+8.9 bps / 21 sessionsthe typical day
Mean Q5−Q1 spread, 68 liquid names−12.7 bps / 21 sessionsNW t = −0.35
Daily rank IC (Spearman), 260 names+0.0017t = 0.28 — no ranking power

Restricting the same test to the 68 largest, most-liquid names in the panel — the ones a real portfolio could actually trade at size — the mean spread does not merely shrink. It flips: −12.7 bps (t = −0.35).

The anatomy of the mirage

Cut the panel into fifths by signal strength and look at what the outcomes in each bucket actually were.

The signal sorts stocks by wildness, not by direction
Standard deviation of 21-session SPY-excess outcomes by signal quintile, %
Q1 — lowest short-flow signal
56.9%
Q2
92.3%
Q3
111.8%
Q4
115.2%
Q5 — highest short-flow signal
119.2%
Source: Same panel: 260 tickers, ~20,100–20,650 ticker-days per quintile; medians shown per bar
View data as table
Mean, median, and volatility of outcomes by quintile
Q1 (lowest signal)mean +2.53% · median −1.03% · std 56.9%20,096 ticker-days
Q2mean +4.59% · median −1.07% · std 92.3%20,363 ticker-days
Q3mean +5.38% · median −1.18% · std 111.8%20,368 ticker-days
Q4mean +6.14% · median −1.05% · std 115.2%20,363 ticker-days
Q5 (highest signal)mean +7.28% · median −1.11% · std 119.2%20,649 ticker-days

The typical stock in every bucket — lowest signal to highest — lost about 1% against SPY over the next month. The medians are five flat stones: −1.03%, −1.07%, −1.18%, −1.05%, −1.11%. What rises across the buckets is not direction but violence: the volatility of outcomes nearly doubles, from 57% in the bottom bucket to 119% in the top.

That combination — flat medians, rising volatility, means that climb anyway — has a name: positive skew. On 976 of the panel's ticker-days, a stock more than doubled against the index inside 21 sessions; the single wildest outcome was a gain of more than 5,500%. In a mean, a single moonshot pays for hundreds of quiet losers. Stocks with this lottery-ticket shape are a documented species — Kumar (2009) catalogued them and the investors drawn to them — and abnormal short-sale attention finds them, because heavy shorting happens where something wild is happening. The signal never predicted anything. It selected volatility, and volatility's skew wrote a mean that looked like alpha.

What survives

Nothing about the arithmetic above is exotic. The mean spread, its t-statistic, the median, the rank correlation — all computed from the same free files anyone can download, and re-derived from the raw data the day this piece was published. The difference between "a strategy earning 5% a month at t > 3" and "nothing" was which summary statistic you were shown.

That is the general lesson, and it is not about short-selling. Whenever a performance claim rests on a mean computed over stocks that can rise 5,000% in a month but can only fall 100%, the mean is the most flattering number in the room, and the median is where the truth went. Ask what the typical outcome was. Ask what happens in the liquid names. A claim that survives both questions has earned the word edge. This one did not — and it wore a t-statistic of 3.07.

The takeaway

  • Every number in the marketing deck can be true and the edge still false. +485 bps/month at t = 3.07 was a correct computation on real data — and it was volatility loading, not skill.
  • Means and medians diverging is the alarm, not a nuance. Five flat medians under a climbing mean is lottery skew doing the work; the raised t-bar the literature demands (Harvey, Liu & Zhu) does not catch it, because the t-statistic inherits the same mean.
  • Liquidity is a truth serum. Restricted to the 68 names big enough to trade at size, the "edge" flipped negative — the mirage lived entirely in the small, wild names where it could not have been harvested anyway.

Scope: one signal family (short-sale share of volume, z-scored per stock), 260 US tickers, July 2024 – July 2026, SPY-excess outcomes at a 21-session horizon entered at the next session's close; 5-session results are directionally identical (mean +177 bps, median −6 bps). This is research forensics, not investment advice.

Sources

  • FINRA, Short Sale Volume Data (Reg SHO daily files) — the raw signal data: daily short volume and total volume per US ticker, one public file per session. finra.org
  • Databento, EQUS.MINI — US Equities Mini dataset — the daily closes behind the outcome panel (2023-04-03 → 2026-06-26) and the SPY benchmark series. databento.com
  • Newey, W. & West, K. (1987), "A Simple, Positive Semi-Definite, Heteroskedasticity and Autocorrelation Consistent Covariance Matrix," Econometrica 55(3) — the standard-error correction behind every t-statistic quoted (holding windows overlap, so naive errors would overstate significance). doi.org
  • Harvey, C., Liu, Y. & Zhu, H. (2016), "…and the Cross-Section of Expected Returns," Review of Financial Studies 29(1) — the case that t = 2 is too low a bar for strategy claims and t > 3 should be demanded; this signal cleared it and was still false. doi.org
  • Kumar, A. (2009), "Who Gambles in the Stock Market?," Journal of Finance 64(4) — lottery-type stocks: low price, high volatility, high skewness — the species whose moonshots manufactured this mean. doi.org
  • The measured panel itself: 121,148 ticker-days (260 tickers × 501 sessions, 2024-07-09 → 2026-07-08) of FINRA short-flow features against 21-session SPY-excess outcomes; every figure recomputed from the raw files on 2026-07-09, with the reproduction commands recorded in this article's data file.
Weekly digest: the most-read systems, in brief. Mondays.

Comments

Always open. Logged-in readers can annotate paragraphs in place.

Loading comments…
or log in to comment under your account