KenKem Journal

Why Is Early Progress in Trading Not Proof?

· #trading-psychology #validation #systematic-trading #risk-management #backtesting

A thin ribbon of vapour curling against a near black background
A good stretch and a real edge look identical from the inside. Photo: Marek Piwnicki / Pexels.

Early progress in trading is not proof, because a good stretch and a real edge look identical from the inside. Proof is what survives conditions the idea was never tuned for: different costs, different regimes, different years. My own research has a clean example of that gap, and in the worst year it cost 67 percent.

What follows is how I tell momentum apart from evidence, including the results that failed and the one number I refuse to quote without its ugly twin.

Why does a good month say so little?

Because a month is a sample size of roughly one market condition.

Progress is worth noticing. It is not worth locking anything in by itself. A few encouraging signals mean an idea has earned more testing, not that the testing is over. The moment I promote an early signal to evidence, it stops being a research lead and becomes a liability, because I will start defending it instead of attacking it.

That distinction has a name in my notes: momentum versus proof. Momentum is pleasant and cheap. Proof is what is left after the idea has met conditions I did not choose for it.

What is the difference between hope and a plan?

Hope says "I think this will work." A plan says "if A happens I do B, and if it breaks at C I stop."

I am not against hope. Hope is what keeps a hard project moving. But hope with no structure has no way of being wrong, and anything that cannot be wrong cannot be improved. A plan gives hope boundaries: written rules, a test that can fail it, and a risk limit that exists in code rather than in my intentions on the day.

The practical version is that every rule in Master Volume Profile is explicit enough to debug. When a trade goes badly I want to trace it back to a named input, not to a mood.

Why did a strategy with a 3.78 Sharpe still lose 67 percent?

Because it met a year it was never built for, and I had not yet proven the edge outside its habitat.

In the 2025 to 2026 volatility regime, my gold configuration returned a 3.78 annualized Sharpe with a 19.7 percent maximum drawdown in the research engine. That is the flattering number. Run the same fixed configuration through 2024 on real MetaTrader 5 fills and it returns minus 67.1 percent, with a 74.2 percent maximum drawdown and a Sharpe of 0.10. In 2024 the spread ran around 8.6 percent of ATR, roughly twice normal, and the strategy simply bled.

Across the full 2024 to 2026 cycle the honest Sharpe is 1.77, not 3.78. My research notes carry a blunter line: no shipped sizing survives a 2024-type year. All of these are backtest and confirmation-run figures from my own gold scorecard, dated 2026-07. None of them is a live track record.

That result is the reason this article exists. Had I stopped measuring at the end of the good window, I would have had a beautiful curve, a confident story, and no idea that the edge had a habitat.

What does a rejection actually buy me?

A boundary I know about in advance, instead of one that finds me later with money on the line.

The clearest example in my work is crypto. Testing Master Volume Profile on BTC three-minute bars, the best configuration I could find reached a profit factor of 1.132 with a 21.2 percent drawdown on the training window. That is a respectable-looking result, and it turned out to be worth nothing. Every train-positive configuration in that family collapsed out of sample, landing at profit factors between 0.72 and 0.83 with drawdowns of 57 to 75 percent.

The two windows were not merely weaker out of sample, they were anti-correlated. The regimes disagreed with each other. So I rejected the whole three-minute family rather than tune it further, because a result that reverses on fresh data is not a weak edge, it is a fitted one. The note I kept was more useful than the configuration: cost and structural assumptions that hold for gold do not automatically transfer to another asset. That is expensive feedback and I would rather pay for it in research than in an account.

To be precise about what survived, since a rejection is easy to overstate: a five-minute BTC configuration with a much longer profile window and a stronger trend filter did stay positive on both windows, and I treat it as a lower-conviction forward-test candidate rather than a lock. The three-minute idea is the one that died, and it died in exactly the way this article is about.

Why does consistency matter more than cleverness?

Because a good process repeated well beats a clever process repeated badly, and because inconsistency makes results unreadable.

If the process changes every week, the data gets noisy fast and I lose the ability to tell whether a change helped. Consistency is not perfection. It is keeping behavior inside the same envelope long enough that the numbers mean something.

That is why the constraint lives in code rather than in discipline. Master Volume Profile has a coded cap on trades per session, so the system's behavior stays inside the same envelope day to day whether or not I am having a good week.

Why do I trust process more than emotion?

Emotion is a useful sensor and a poor execution engine.

Fear and excitement are reasonable signals that something matters. They are terrible at deciding what to do next, because they change on every candle while the plan does not. So the exits are written down and coded, including an ATR-based trailing stop, specifically so that a trade in progress is carried by a rule instead of by a hand reaching for the mouse.

This is the part that is genuinely hard to explain to someone still trading manually. Handing execution to a rule set feels like losing control. In practice it is the opposite: I keep control at the point where I set the rules and the risk, and I give up control at the point where I am least reliable, which is in the middle of an open position.

What does durable improvement look like?

Boring, mostly. One better decision, one cleaner test, one more honest rejection.

I am not looking for a lucky month. I am looking for a process that improves without depending on one kind regime. That is why I judge walk-forward results by the worst fold rather than the average, and why I plan around the 95th percentile of the Monte Carlo distribution rather than the median. For the gold configuration that means sizing for a 30 to 40 percent peak drawdown, because the distribution puts the median near 24 percent and the 95th percentile near 38 percent.

One detail I find quietly reassuring, and I mention it because it is the rare test that came back better than expected: in the out-of-sample fold the profit factor was 1.327 against 1.203 in training. Out-of-sample performing above in-sample is the opposite of a curve-fitting signature. It is not proof either. It is one gate passed out of several.

Why does serious progress look boring before it looks impressive?

Because the work that produces durable results is repetitive, and repetition photographs badly.

Nothing in the honest version of this journey makes a good screenshot. Cleaner logs. A test that failed for a reason I can name. A configuration that looked good in training and reversed on fresh data. A drawdown number revised upward because the first estimate was too kind. That is what the process looks like from the inside, and the loud version of trading almost never shows it.

I am still hopeful. I am also careful not to call anything proof before it has earned the word. Master Volume Profile has not cleared the live gate yet, and until it does, everything I publish is a hypothesis with evidence attached rather than a result.

Frequently asked questions

What is the difference between progress and proof in trading? Progress is a result you have seen; proof is a result that survived conditions you did not choose. The practical test is whether the outcome holds across different cost regimes, time periods, and parameter neighborhoods. My gold configuration looked strong in its 2025 to 2026 habitat and lost 67.1 percent on real fills in 2024, which is the difference in one line.

How long does it take to know if a trading strategy works? Longer than most people are willing to wait, and the answer is a trade count rather than a calendar. The minimum track record length depends on the strategy's own return distribution, so a noisier strategy needs a bigger sample to say the same thing. For my gold configuration that floor was 192 trades and the validated sample was 1,423. I used to call that seven times the minimum. Once I adjusted for a skew of 3.49 and excess kurtosis near 19.5, the Gaussian equivalent effective sample came out around 326, so the honest way to say it is comfortably above the floor rather than a multiple of it.

Why reject a strategy that was profitable in testing? Because profitable in testing and robust are different claims. My BTC three-minute configuration reached a 1.132 profit factor on the training window and then collapsed to between 0.72 and 0.83 out of sample, with drawdowns of 57 to 75 percent. When two windows disagree that sharply, the training result was describing that window rather than an edge, and tuning it further would only have fitted it harder.

What drawdown should I plan for? The unflattering percentile, not the average one. In my own research the Monte Carlo distribution puts maximum drawdown near 24 percent at the median and 38 percent at the 95th percentile, so my notes say to size for a 30 to 40 percent peak. Anyone planning for the median is planning to be surprised at the worst possible moment.

Does a high Sharpe ratio mean a strategy is safe? No. A Sharpe ratio is a description of one sample, and it inherits every limitation of the window it was measured on. The same fixed configuration in my research reads 3.78 in its favorable regime and 1.77 across the full cycle, which is the same strategy and the same code measured over a longer, less kind stretch of history.

Why is out-of-sample performance above in-sample a good sign? Because curve-fitting normally produces the reverse. A configuration tuned to the training data usually degrades when it meets fresh data, so an out-of-sample profit factor of 1.327 against a training figure of 1.203 argues against the parameters having been fitted to noise. It is evidence in one direction, not a clean bill of health.

Does Master Volume Profile have a live track record? No, and I will not imply otherwise. Everything published so far is research: backtests on roughly 160 million real broker ticks plus MetaTrader 5 confirmation runs. Treat every figure here as historical evidence about a hypothesis, not as a prediction. One further caveat: spread is charged on every trade in the model, but slippage, latency, and swap are not yet modeled, so live should be expected to be thinner than backtest.

Does any of this remove risk? No. It removes some specific ways of being wrong, which is a far smaller claim. Trading involves risk including the risk of loss, every figure above is historical, and no amount of validation converts a hypothesis into a certainty.


Written by KenKem, a software engineer and founder of twenty years, learning quantitative trading in the open and publishing the process, rejections included.

This article was composed from posts 49 to 60 of the KenKem build-log series. Educational purpose only. Not financial advice. All figures cited are backtest or confirmation-run results, not a live track record. Past performance does not guarantee future results.

← All journal articles

Chat