The government tells you who it's paying, months before the earnings report. We tested whether that's worth money.
Summary
Eight registered tests across three free public data streams — short-selling files, Wikipedia attention, federal contract awards — produced seven failures and, this week, one pass: contract flow ranked contractors' next-month returns even after the multiple-testing penalty. Here is the number, the attempts to kill it, and why it still earns a stopwatch instead of a dollar.
So we built the map. Three public data streams, each earlier in the information chain than a securities filing: daily short-selling volume from FINRA, public attention from Wikipedia's pageview logs, and federal contract flow from USAspending. Every test registered before looking — signal, horizon, universe written down first — because eight tests means eight chances to fool yourself, and the significance bar must rise to match: |t| ≥ 2.73, the Bonferroni correction at eight trials on top of the raised standard the literature already demands (Harvey, Liu & Zhu, 2016).
Seven of the eight failed.
Three streams, eight tests, seven bodies
The failures are the context that gives the pass its meaning. Short-sale flow produced a mean spread of +485 bps a month at t = 3.07 that dissolved under one median — volatility masquerading as alpha, dissected in an earlier piece. Wikipedia attention showed nothing on large caps at any horizon. Contract flow measured against the S&P 500 came close twice and cleared nothing.
View data as table
| usaspending_award_z_xs, 21d | IC t = +2.91 | PASS (bar 2.73); net +155 bps/21 sessions |
|---|---|---|
| usaspending_award_z_wide, 21d | IC t = +2.34 | fails (S&P-excess label) |
| usaspending_award_z, 21d / 63d | t = +1.44 / +0.64 | fails |
| wiki_attention_z, 5d / 21d | t = +0.52 / +0.11 | fails |
| finra_short_z, 5d / 21d | t = −0.64 / +0.28 | fails (mean-spread mirage) |
The one that crossed: rank the 32 publicly traded federal contractors in the panel each day by their trailing 21-day civilian contract flow, z-scored against each company's own year-long baseline. Measure each stock not against the market but against the median contractor in the same panel — so a defense-sector rally cannot masquerade as signal; the ranking has to work within the industry. Top quintile minus bottom quintile: +195 bps per 21 sessions on the mean, +146 on the median — the two summaries agree, which is precisely what the short-sale mirage could not manage — and +155 bps net of a double round-trip cost haircut. Rank IC +0.064, t = 2.91, on 859 evaluation days.
Two hygiene details carry the honesty. Every transaction enters the panel only seven days after its action date — later than the three-day rule requires, deliberately. And every Department of Defense row is thrown away: DoD withholds its awards from public view for 90 days (FPDS data-availability notice; USAspending, About the Data), so a backtest that could see them would be reading data the public did not have. The signal is built only from what a stranger with a laptop could have downloaded that morning.
Trying to kill it
A desk that has buried seven signals does not celebrate the eighth. It runs the diagnostics that would expose it.
View data as table
| Full window (n=859d) | IC +0.064, t = +2.91 | the pass |
|---|---|---|
| Common window (n=790d) | t = +2.52 | below the 2.73 bar on its own |
| Early-2023 extension (n=71d) | IC +0.154, t = +1.83 | 3× the full-sample IC |
| Every cut (horizons, universes, windows) | positive | artifacts here haven't looked like this |
| Mean vs median spread | +195 vs +146 bps | signs agree — not lottery skew |
The decomposition found something worth saying out loud: the passing variant's margin over the bar comes from 69 extra sessions of early 2023 that its sector-neutral label unlocked. On the 790 days both label constructions share, the t-statistic is +2.52 — below the bar. In those 69 early sessions the signal was three times as strong as its full-sample average. Nothing about that is disqualifying — every cut of the data is positive: both horizons, all three universes, early window and late, mean and median. The artifacts this desk has caught did not look like that; they inverted somewhere. But a pass whose margin lives in its oldest 8% of days has not earned confidence. It has earned measurement.
A stopwatch, not a dollar
What an in-sample pass proves is limited: one history, one panel, with the multiple-testing penalty as the only guard. It says nothing about the next 859 days. So the desk's response is not a position — it is a forward clock. Every night, the current ranking of all contractors is written to an append-only, timestamped log before the outcomes exist; every week, matured snapshots are scored with exactly the in-sample metric. The benchmark to beat is written down: IC +0.064, +195 bps per 21 sessions. The first matured readings arrive in mid-August 2026. Until the forward numbers hold up, the standing rule applies: no capital — paper trading included.
The takeaway
- The chain position is the thesis. A securities filing is the end of the information sequence; a contract award is the beginning of revenue. The one signal that passed came from the earliest data in the map — and the two latest streams (short flow, attention) showed nothing.
- The bar did its job twice. It killed a spectacular fake at t = 3.07 (the short-sale mirage) and it withheld judgment on two near-misses (t = 2.34, 2.52) that a lower bar would have anointed. One pass in eight registered attempts is what honesty looks like on free data.
- In-sample is an audition, not a verdict. The pass earns a forward clock with the benchmark written down in advance — the only test that cannot be fooled by anything done to the past.
Scope: 814,796 federal contract transactions for 35 recipients (2022-07-10 → 2026-07-07); the passing test evaluates 32 publicly traded contractors over 859 days at a 21-session horizon, sector-neutral label (return minus same-date contractor-panel median), entry at the next session's close, 7-day publication lag, DoD rows excluded. In-sample; forward validation began 2026-07-10. This is research forensics, not investment advice.
Sources
- USAspending.gov and its public API — the raw contract-transaction data: recipient, amount, action date for every reported federal award. usaspending.gov · api.usaspending.gov
- Federal Acquisition Regulation 4.604 — the rule requiring civilian contract actions to be reported to FPDS within three business days of award. acquisition.gov
- FPDS, DoD Data Availability notice — Department of Defense awards are withheld from public access for 90 days (the reason DoD rows are excluded from the signal). fpds.gov
- USAspending.gov, About the Data — the official documentation of agency reporting schedules and the DoD delay as it appears downstream. usaspending.gov
- FINRA, Short Sale Volume Data (Reg SHO daily files) — the failed tier-3 stream. finra.org
- Wikimedia Pageviews API — the failed tier-2 attention stream. wikimedia.org
- Databento, EQUS.MINI dataset — daily closes behind every outcome panel. databento.com
- Newey, W. & West, K. (1987), Econometrica 55(3) — the HAC standard errors behind every t-statistic (overlapping holding windows). doi.org
- Harvey, C., Liu, Y. & Zhu, H. (2016), "…and the Cross-Section of Expected Returns," Review of Financial Studies 29(1) — why a t bar near 2 is too low once many strategies are tried; the basis for the raised, trials-adjusted bar used here. doi.org
- The measured panels themselves: 814,796 USAspending transactions (35 recipients, 2022–2026), 121k FINRA ticker-days, 49k Wikipedia pageview-days — every figure recomputed from the raw files on 2026-07-10, with reproduction commands recorded in this article's data file.
Comments
Always open. Logged-in readers can annotate paragraphs in place.
When a federal agency signs a contract, it does not wait for the vendor's quarterly earnings call to tell you. Civilian agencies must report every contract action to the government's procurement database within three business days (FAR 4.604), and USAspending.gov publishes it — recipient, amount, date — through a free public API. Revenue that will not appear in a 10-Q for one or two quarters is sitting in a government database now, named and dated. The question a research desk should ask is not whether that is interesting. It is whether it is priced.