How We Test

Anyone can sell a screenshot. We sell what survived.

Most indicators are sold on one pretty backtest — one symbol, one window, no costs, nothing trying to break it. We put every strategy through raw tick data, a 7-stage gauntlet, thousands of Monte Carlo runs, and a clean-room rebuild. Can't survive all of it? We don't ship it.

3B+
Raw ticks in the research set
7
Adversarial trials
5–20k
Monte Carlo paths per test
0
Unverified results shipped
01 — The Retail Gap

A backtest built to confirm will always find a reason to.

The TradingView strategy tester is a fine tool — but the way most indicators are sold, it's used to flatter a strategy, not stress it. Here's what a screenshot leaves out.

One symbol, one window

Tuned on the exact chart and date range where it looks best. Change either and the edge often disappears.

Repainting signals

Arrows that move after the bar closes. The backtest "knew" the outcome — in real time it never fired that way.

Zero-cost fills

No slippage, no commission, no spread. Frictionless math that no live account has ever experienced.

In-sample only

Optimized and "verified" on the same data. That's a memory of the past, not a prediction of the future.

One lucky regime

Great in a single bull run or volatility spike, then quietly dead the moment the market changes character.

Nothing adversarial

The strategy is never once attacked. It's shown at its best angle and sold before anyone asks how it breaks.

02 — The Testing Stack

Four layers. An edge has to survive every one.

Each layer is designed to expose a different way a "profitable" strategy can be fake. They run in sequence, and any single failure ends it.

01 · Raw tick foundation

No vendor-aggregated bars. We test on raw exchange tick data — exchange-native, not re-aggregated — across multiple instruments and 10+ years of history, including long stretches of untouched out-of-sample. The clock is DST-corrected and the data is verified bit-for-bit, so a result is never an artifact of a broken bar or a mislabeled hour.

02 · The Gauntlet 7 trials

A fixed battery of attacks, each engineered to kill the strategy a different way. The kill bar is set before any test runs and never moves. Every edge is worthless until it survives all of them. Detailed below.

03 · Monte Carlo 5,000–20,000 paths

One equity curve is one roll of the dice. We resample the strategy's trades thousands of times to build the full distribution of outcomes it could produce — median return, worst-case drawdown, share of losing years, and whether the account survives the bad paths. We judge by the range of futures, not the single luckiest line.

04 · The Clean-Room Rebuild built twice

Built twice. If the two builds disagree, it dies. Before anything ships, a second engineer rebuilds the entire strategy from scratch in a clean room — no access to the original code, nothing carried over. It's the only reliable way to catch look-ahead: an edge secretly built on information it wouldn't have had live. If the two builds don't produce the same number, the edge was a mirage, and it's discarded.

03 — Inside the Gauntlet

Seven trials, run in order.

Each stage isolates one failure mode and asks a question the equity curve alone cannot answer. A single kill ends the run — worthless until it survives every one.

01

Random-limit placebo

Fabricate fake levels at the same distances, directions, and times of day as the real signals, then run identical fill mechanics. If real levels don't beat the fakes by a wide margin, the edge was the fill — not the level.

KILL IF the measured edge sits inside the placebo cloud.
02

Cost & slippage stress

Re-run across a sweep of worsening execution — slippage from half a point to three, cost from two points to four. A real edge degrades gracefully and stays above water.

KILL IF net expectancy ≤ 0 at any realistic friction.
03

Lag test

Enter zero, one, three, and six bars after the trigger. Graceful decay means a real property of the level. Instant collapse means it only captured an unrepeatable exact-fill tick.

KILL IF it collapses to negative one bar off the touch.
04

Market-fill (the chase)

Force a market entry at the next bar's open instead of waiting for the limit. Reveals how much of the edge lives in patient, passive execution versus chasing price.

READ negative ⇒ "requires the limit," not a kill.
05

Walk-forward

Optimize every parameter on an in-sample window, freeze the winning settings, and test them on untouched out-of-sample history the optimizer never saw. This is where most beautiful backtests quietly die.

KILL IF out-of-sample expectancy ≤ 0.
06

Sign-flip permutation

Randomly flip the sign of each day's realized P&L thousands of times to build the distribution pure chance could produce. Flipping whole days respects same-day trade correlation.

KILL IF p > 0.05 — indistinguishable from luck.
07 · THE ACID TEST hardest to fake

Cross-instrument replication

Re-run the whole strategy, ATR-scaled, on a different but related market — the same structural logic on independent price data. A true market structure shows up in both; a pattern mined from one instrument's noise flips sign or vanishes. You cannot overfit a market you didn't touch.

KILL IF the edge flips sign across instruments — the tell of an artifact.

Most don't make it

Most candidates enter the Gauntlet; almost none walk out. A protocol that passes everything is measuring nothing — the value is in how early and how honestly the rest are killed.

Read the full Gauntlet breakdown → Visit the Graveyard →

Every stage in detail — plus a real case that passed all seven and still got killed, and the full log of what didn't survive.

04 — Monte Carlo

We don't ask "did it work?" We ask "how often would it fail?"

A single backtest is one path through history — the one that actually happened. But the order of trades is partly luck. Resample them thousands of times and you get the full cone of outcomes the same edge could have delivered.

That cone is what matters for real money. The median tells you the typical result; the lower tail tells you the drawdown you have to survive to ever see it. A strategy with a great average and a fifth percentile that blows the account is not tradeable — and only Monte Carlo shows you that before it happens to you.

Median
Typical annual outcome
P5 / P95
Drawdown envelope
% down years
Across all paths
Survives?
Account-ruin check
Monte Carlo · 400 resampled pathsequity →

Illustrative visualization of the method — not the performance of any specific product.

05 — The Difference, Side by Side

What a screenshot skips, and what we require.

DimensionTypical retail indicatorMXT ALGOS standard
DataVendor bars, one symbolRaw exchange ticks, multiple instruments, 10+ yrs
Test windowThe range that looks bestFrozen out-of-sample the strategy never saw
CostsDefault or zeroReal slippage, commission, spread, worst-case fills
RobustnessA single equity curveThousands of Monte Carlo paths
Adversarial testingNone7-stage gauntlet built to kill it
Look-ahead checkUndetectedClean-Room Rebuild — built twice, must match
Bar to ship"It looked good"Survived every test — or discarded
06 — What This Does & Doesn't Promise

Rigor is a filter, not a guarantee.

We won't tell you a strategy will print money — no honest tester can, and anyone who does is selling the screenshot again. Markets change, edges decay, and every position carries risk. What our process guarantees is narrower and more valuable: that the edge was real at the time of testing, that it survived being actively attacked, that its worst-case behavior was measured before you risked a dollar, and that no result was a data-mining fluke or a look-ahead mirage. That's the difference between "here's a great backtest" and "here's what we did to try to prove it wrong, and it held."

See exactly how a strategy earns its place.

Every tool we ship carries the testing record behind it. Ask to see the gauntlet results for any strategy.