Six forensic checks for the failure modes that turn a paper edge into a live loss. One independent verdict, in writing.
Start a Check → Read the findings ↓One person runs every check. One person signs the verdict. Named, and liable for it.
Illustrative summary, not a client report. Every check gets a verdict, a measured haircut where possible, and a fix. Read a full sample verdict (PDF).
Everything here is checkable:
pip install backtest-bias ·
The notebook that reproduces +3.0pp on public data ·
Published audit findings ·
A complete sample verdict (PDF)
All real, all from the last few weeks: two from public-data audits, two from my own pipeline. I publish what I catch in my own work for the same reason a good lab publishes its failed experiments. The newest receipts come from reviewing one client's 35 works in a single week: 44 findings, 13 passed clean, verified exact, and two of my own claims retracted on the record along the way. The census is public; the client is not.
A delivery-percentage signal showed t = 7.61, monotone across buckets. The field is only knowable after the close. Lagged one day: t = 0.66, ordering reversed.
Illiquid-stock reversal in India: +28.8%/yr gross at t = 11.65, beautifully monotone. Then you pay to trade it.
From a review engagement: a bar chart whose shares summed to the mid-80s against its own footer's 100, a bias formula quoted from a regression with an intercept beside a regression without one, and a "25%" that recomputes to 27 from the post's own printed values. 44 findings across one client's 35 works; 13 passed clean. The census is published in full; the client is not named.
Popular free sources drop delisted names. Rebuilt with every dead company kept in the universe, one realistic strategy lost +4.0pp/yr of its claimed return. A public 99-name version reproduces +3.0pp, runnable on Kaggle, so you can check the method before you trust it.
Two fields in a widely-used panel are exactly 0.0 before 2022 rather than NaN, so notna() passes and any ranking built on them silently sorts "has the field" from "doesn't". Produced a fake t = 7.6 result before it was caught.
Every audit runs all six. Each gets a pass, flag or fail, the estimated performance haircut where measurable, and a fix.
Is your universe quietly missing the companies that died?
Does any signal use information that didn't exist on the trade date? Adjusted prices, restated fundamentals, index membership as of today.
Are commissions, slippage, impact and taxes at levels a real fill would actually pay?
How many configurations did you try, over how many years? Deflated metrics and a parameter-sensitivity map.
Is the edge distinguishable from luck at your sample size? Streaks and drawdowns checked against the distribution.
Could the market have absorbed your orders at those prices? Volume participation, limit fills, gap handling.
Trade list and equity curve, a platform export, or written rules precise enough to act on. Code welcome, never required.
Starts when payment and a complete folder are both in. Inside it, I reconstruct your test, rerun it independently, and cross-check it against my own panel. If something is missing I ask once and the clock pauses. A 24-hour rush lane exists when you need it, at plus fifty percent.
Pass, flag or fail on each check, the measured haircut where possible, and a prioritised fix list.
Not a grade for your ego. An engineering document for your next iteration. If a check fails in a way that needs deeper work than the fixed scope covers, I tell you what it would take before doing anything further. No surprise invoices.
None of these people are clients. They are researchers and founders who engaged with the method in the open, corrected it where it was wrong, and adopted it where it held. That is a more useful thing than a testimonial, and a harder one to manufacture.
"Effectively audits PhD-level quantitative work from very little information. Given a chart and a few hundred words, without the underlying code or data, he independently derives the figures and identifies what does not reconcile. His review has materially improved the published version of several analyses in my series."
Dr. Heather Dempsey, quantitative researcher, amended a published piece after an exchange about a discrepancy, and noted the change in the body of the post rather than editing quietly. Correcting a published piece in public is rarer than the catch that prompts it.
A simulation showing that ranking on the same window that selected a candidate can manufacture apparent mean reversion was published as a follow-up piece with named credit, and the disjoint-window recompute became stated practice. The useful test is whether a method travels beyond the person who proposed it.
On multiple testing under correlated search, we each corrected the other in the same thread: an equicorrelation figure that did not survive checking, and a threshold approximation of mine that ran high at practical sample sizes. Both corrections sit in the public exchange.
Two trading-platform founders have taken methodology suggestions from public threads into their own products, and one invited me in as an independent voice on data honesty for their users. The checks are not specific to my own book.
Everything above is public and linkable, and quoted with permission where a person is named.
Audit before your capital finds out, review before your audience does. The quick list, for the skimming and the undecided:
Errors caught here never reach your audience, your investors, or your reviewers. The embarrassing version of every finding dies in a confidential report.
Every statistic recomputed from your material alone, every chart checked against its own text. Findings you can defend line by line.
Written verdict in 48 hours standard, 24-hour rush when it matters. No calls required.
Fixed menu, scope agreed in writing before work starts, no surprise invoices.
A named human runs every check and signs the verdict. My own retractions sit on the public record; the review gets reviewed too.
Never shared, never reused, deleted on request. Only you decide whether anyone ever sees a verdict.
The full six-check audit. Written verdict in 48 hours. For most backtests this is the only step needed. See a full sample verdict (PDF).
Guaranteed: the intake asks what you already believe is wrong, in numbers. If the verdict doesn't go beyond what you wrote, the fee comes back (how "beyond" is judged).
Pay $250 by card, or ask for an invoice (PayPal or bank transfer).
Start a Check →A Check found real problems and you want them repaired: the leak closed, the cost model rebuilt, the test rerun clean.
Scoped in writing after a Check. 7-day delivery window.
Full pipeline engagement: your data rebuilt point-in-time correct, the audit run at source level, the report your next strategies inherit.
Scoped in writing. 7-day delivery window.
48h turnaround Fixed price No code required NDA on request
"He has reproduced my statistics from the posted material alone, caught a mechanism error I had already published, and retracted one of his own corrections unprompted when his model of my simulation turned out wrong. ... He's fast and very good."
I use AI heavily myself; it's why the price is a few hundred dollars and not a consultant's week. But three things don't come with the free prompt. The reference data: AI cannot conjure a survivorship-correct, point-in-time panel to check yours against, and mine took months to build. The adversarial stance: a model in the strategy owner's hands tends to validate the owner. And liability: a named human signs this verdict. In one recent week my own tooling produced four significant-looking results, up to t = 7.6, and every one was a bug. The catching is the product.
No. A trade list and an equity curve are enough for all six checks. The audit tests whether your test is trustworthy; it doesn't need to know why you trade what you trade. Code only enters at the Fix and Deep tiers, under an engagement letter with custody and deletion terms. NDA on request.
The six checks are method checks: they test how your backtest was built, so they work on equities, futures and options from any market. The data-level cross-check against my own reference panel is a different matter: it is deepest for Indian equities, partial for US indices, and elsewhere I verify your data's internal consistency rather than compare it against an independent panel. I say which kind of check produced each finding.
I'll tell you whether your test can be trusted. That's the part you cannot see from inside, and the part everything else depends on. The market grades the strategy.
Always. Reviews and materials are never shared, never reused, and deleted on request. Only you decide whether anyone ever sees a verdict. Any public tally of findings is counts only: no names, no content.
Check tier only. Your intake asks what you already believe about each check: which fields might leak, what your costs really are, how many configurations you tried. That document is the baseline. If the verdict doesn't go beyond it, tell me and the fee comes back. Beyond means a named mechanism, a measured magnitude, or a fix that isn't in your answers. You hold both documents, so the comparison is checkable from your side, not just mine. No guarantee on Fix and Deep, which are scoped in writing before any money moves.
Then you get a verdict saying so, per check, and that document is worth something: an independent audit you can show a funder, a prop firm, or yourself at 2am before going live. A clean pass from a hostile reviewer is not the same as your own backtest saying it's fine.
48 hours means 48 clock hours, Indian time, from the moment payment and a complete folder are both in. Current capacity is three Checks a week; if you land in a queue you are told your position when you order, and your clock starts when your slot opens. A missing-information request pauses the clock once until you reply.
Whatever you send stays on my machine, is never reused or redistributed, and is deleted on request, or in any case within 90 days of the engagement closing. If you'd rather not send code at all, don't. A trade list and an equity curve are enough for all six checks.
For transparency: I run my own systematic book in Indian markets. Nothing a client sends is ever reused in it or in my research, and no verdict is influenced by my own positioning. The audit is QA on your test, not an opinion on your trade.
Quants about to size real money on a backtest. Prop-challenge takers, where an audit costs less than a second attempt and tells you which one to make. Developers selling strategies who want independent verification to point at. Algo providers preparing for SEBI's PaRRVA regime, where a research report per algo is a standing requirement. And anyone whose equity curve looks a little too clean to be true.