KenKem Journal

How I Tell a Robust Edge From a Lucky One

Β· #robustness #backtesting #research-method #trading-psychology #systematic-trading

There is a specific moment in research that I have learned to distrust. A backtest finishes, the equity curve slopes the right way, and for about thirty seconds it feels like the work is done.

It almost never is. Most of what I have built over the past year is not a strategy, it is a set of obstacles designed to take that feeling away from me before I act on it. This is an honest account of what those obstacles are, why I needed them, and what they keep telling me.

Why does a good backtest feel so convincing?

Because wishful thinking feels productive. That is the trap.

When a result looks good, the natural next move is to explain why it makes sense. You can spend an entire evening building a story around a weak setup, and the whole time it feels like research. It has the shape of work. It produces no new information at all.

I have caught myself doing this, and the cost is not the wasted evening. The cost is that a story delays the real fix. Every hour spent defending a number is an hour not spent finding out whether the number is real. So I try to keep belief in the right place, which is after the evidence rather than before it. Measure first, then decide what to believe.

That sounds obvious written down. It is much harder when the number is one you wanted.

What does robustness actually mean?

A result that only works in one narrow pocket is not an edge, it is a coincidence with good timing.

So the engine I developed, Dquants, pushes an idea along four axes before I take it seriously:

Time. Does the behaviour hold across different windows, or is it carried by one unusually kind stretch of history?

Regime. Trending and balanced markets are different problems. A rule that only survives one of them has a much narrower job than its headline number suggests.

Parameters. I want a plateau, not a peak. If a setting has to be exactly right to work, then what I found was the shape of my own search, not the shape of the market.

Cost. Spread and commission are the difference between a marginal result and a losing one. Stress them upward and see what is left.

An equity curve that is less spectacular but survives all four is worth more to me than one that is beautiful exactly once. Robustness is not a bonus round after the research. It is most of the research.

Why does the sample matter more than the headline number?

A small sample can mislead badly while looking exciting, and the headline metric will not warn you.

So alongside the usual numbers I track how many trades a result came from and how much of the map it covers. If one quarter or one regime is doing most of the lifting, I treat the whole thing as a narrow finding rather than a broad one, regardless of how the summary reads.

This is also where the statistics earn their place. A raw Sharpe ratio says nothing about how many variants you tried before you found it, and if you test enough ideas, one of them will look good by accident. Deflated Sharpe adjusts for exactly that. Minimum track record length asks a related question from the other direction: given this much variance, how many trades would I actually need before the result stops being consistent with luck? If the sample is smaller than that number, the honest answer is that I do not know yet.

Those two checks have killed ideas I liked. That is the point of them.

What do the failures actually look like?

There is a graveyard behind this work, and it is large. Somewhere around thirty five parameter sweeps have been tested to conclusion and rejected. In one recent stretch I built and tested five separate ideas in a row and rejected every single one.

They tend to fail in three recognisable ways.

Curve-fit. The idea can be tuned to look good on the training window and falls apart outside it. One variant I was fond of fit its training data nicely and then produced an out-of-sample drawdown north of eighty percent. That is not a strategy with a rough patch. That is a strategy that never existed.

Regime-dependent. The behaviour is real but unconditioned. It works while one kind of market persists and grinds down when that market changes. This one is genuinely hard, because the result is not fake, it is just incomplete. It needs a condition attached, not a parameter tweak.

Feed-fictional. The edge exists in my data and not in the world. This is the most humbling category, because it means the tooling lied to me.

Naming the three modes changed how I work more than any single test did. When something fails now, the first question is which of the three it was, and that usually points at the fix.

Where do my own numbers still flatter me?

In the costs, and I would rather say so plainly than have someone find it later.

A tick engine that models spread and commission but leaves slippage, latency and rollover at zero is not simulating trading, it is simulating a frictionless version of trading. I know the spread in one of my historical feeds is roughly ten times tighter than what I actually pay live. A thin result under those conditions can be a losing one in reality.

The exit side has its own debt. Three separate times, a change that looked flat or positive in my engine came out clearly worse in MetaTrader. The engine cannot know the path price took inside a tick, so it systematically over-credits trades that are allowed to run. Entry-side numbers I trust. Exit-side numbers I hold loosely.

Building a realistic friction layer and re-ranking everything underneath it is the largest open item on my list. Until that is done, every number I have is provisional, and Master Volume Profile has no live track record to point at either way.

What happens when a result does survive?

It is genuinely satisfying, and I try to keep the satisfaction in proportion.

When something clears a tougher test than I expected, what I have earned is not proof. It is a reason to run a harder test. A validation gate is a checkpoint, not a finish line, and treating a pass as the end of the story is how a robust process quietly turns back into a hopeful one.

The results I want are clean, repeatable and explainable. If I cannot reproduce a result on demand, I probably do not understand it well enough to rely on it.

Why is "not yet" the most useful answer I have?

Because most of the time it is the true one.

Not yet validated. Not yet robust across regimes. Not yet tested under realistic costs. Not yet ready. None of that is failure, and none of it is a soft no. It is accurate progress reporting, and I have come to respect it far more than confident certainty that has not been earned.

The frustration is real. Long stretches of this work are dead ends, build errors and corrections, and progress is often invisible until one day the process is noticeably more stable than it was. But that is where the weak assumptions get exposed, and I would rather find them in research than in a live account.

What is the honest takeaway?

A strategy is not strong because it made money once. It is strong because it keeps its shape under pressure.

If you are coming to markets from engineering, my honest expectation is that most of your ideas will not survive clean data and real costs, and that the ones which do will be smaller than you hoped. That is not the process letting you down. That is the only version of the process worth having.

Waiting for more evidence is almost always cheaper than deploying too early.


This article was composed from the KenKem build-log series. Educational purpose only. Not financial advice. Past performance does not guarantee future results.

← All journal articles

Chat