The question comes up again and again and almost nobody answers it. The short answer is that it matters less how many years you have than how many pieces you split them into — and here is a case of ours where splitting them in half changed the decision, and then changed it again, against us.
Two improvements to our system raised the total result by 12 % and 17 %. When the same data was split into two halves, neither of them improved in the first. And the one that ended up inside the system is one of them.
Our system trades stocks. We added a piece that looks at the general state of the market and classifies each day into three situations: normal, choppy and bad. The idea was to use that classification to risk less when the market is choppy and not trade when it's bad.
There are two ways to use it, and we measured both separately, along with not using it at all:
The result is measured in R, which is the natural unit here: one R is what you were willing to lose on that trade. Making 3 R is making three times what you risked. And since a version that risks less also falls less, all figures are adjusted to the same maximum drawdown, so the comparison doesn't simply reward the one that takes the fewest chances.
| Version | Whole period | First half | Second half |
|---|---|---|---|
| Without the piece (reference) | 251,66 | 139,24 | 113,17 |
| Full version | 282,60 | 106,19 | 142,45 |
| Short version | 293,40 | 135,30 | 134,60 |
Looking at the other two, neither of them wins in both halves. The full version clearly loses against using nothing in the first (106.19 versus 139.24) and all its advantage is in the second. And the short one, which is the one that went into the system, doesn't win in the first either: 135.30 versus 139.24.
That difference of 3.94 is small — the same system, run twenty-five times changing only the order in which it handles the same day's marks, swings between 16 and 19 in that stretch — so the honest thing isn't to say that the short one is worse: it's to say that in the first half it can't be told apart from using nothing. Which is exactly what this rule asks you to check, and exactly what stops it being presented as an improvement.
The one that went into the system was the short one. The full one, which gave the best total number, was left out. With these numbers in front of us, the short one got in with less support than we thought: its advantage also lives in the second half.
Correction — 17 August 2026
The two half columns were wrong, and the conclusion with them. The first column — the whole period — reproduces exactly (251.66 · 282.60 · 293.40). The halves don't: the report didn't even say where it split, and the run that produced them wasn't saved. When we measured them again — same engine, same data, same drawdown adjustment, splitting at the series' midpoint date (26-Apr-2018) and saying so — the numbers above come out.
And they change the reading. We had published that the short version won in both halves, by 12 % in the first. It doesn't win: it stays at 135.30 versus 139.24, within the noise. So the version we chose has the same problem we saw in the other one — its advantage lives in just one half. The report's lesson doesn't change, it gets stronger: splitting the data changed the decision, and this time it changed it against us. We're putting it in writing instead of fixing it quietly.
A total is an average, and an average can be put together in many ways. A rule that works very well during a specific stretch — a few years with a type of market that's gone — drags the total up even if it gets in the way the rest of the time. When that happens, what you've measured isn't that the rule works: it's that there was a season when it worked.
Splitting the data in two separates those two things. If the advantage shows up in both halves, it comes from something that repeats. If it only shows up in one, it comes from that season, and the future has no reason to look like it.
And there's a reason why this mistake is so common: nobody makes it on purpose. You test several versions, you pick the one with the highest number, and that choice has already brought in the lucky season without anyone deciding it.
It gives you a rule you can apply in ten minutes, with no statistics needed: split the history in half and look at both parts. If your strategy only works in one, you don't have a validated strategy: you have a strategy and an era.
It also answers the original question better than a number of years. Having twenty years of data and looking at them all at once tells you less than having eight and looking at them in two stretches of four. What validates isn't the amount: it's that the advantage shows up again.
And it gives a third thing, the one that saves the most: when someone shows you a curve, the useful question isn't how much it rises. It's in which stretches it rises. If the answer is “mostly between such and such a year”, you know what you're looking at.
It doesn't say the full version is bad, or that it won't work in the future. It says that with this data you can't claim it adds anything, which is different. It doesn't say two halves are enough: they're the minimum check, not the final one, and splitting into three or four is better when there's data for it. It says nothing about what anyone should do with their money. And the figures are measured without the cost of holding positions open overnight, which in a real account is paid and isn't small.
With a check you can do yourself: split your history in two at the midpoint and calculate the result of each part separately. If the sign changes, or if one half ends up below doing nothing, you know more than before you started.
And with an idea that holds even if this particular piece falls apart: a total result hides where it comes from. The cheapest way to find out is to split it, and almost nobody does because the total number is always prettier.
If you want to keep pulling the thread
The other two reports in this series measure the same thing from another angle: why a high win rate isn't good news and why the same test repeated twenty-five times gives twenty-five different results. They are all at sophronepsis.com/informes.html.
And if you want to learn to look at this on your own, the six-day walkthrough is at app.sophronepsis.com/empieza.
A report is a snapshot of one day. Inside the platform the measurements are redone when new data arrives, every strategy comes with its test alongside it, and the mentor answers whatever you ask, at any hour. sophronepsis.com/mentoria.html
This is research, not a sermon. If you find a flaw in the method or in the numbers, write to us at hola@sophronepsis.com and we'll correct it in public. Sophronepsis · Wait. Observe. Execute.
148 stocks, 16.6 years of daily data (4-Jan-2010 to 12-Aug-2026), exit at 3 R, 25 shuffles per configuration and median published. R adjusted to the same maximum drawdown across versions, and each half adjusted to its own reference's drawdown. The halves are split at 26-Apr-2018, the series' midpoint date: 11,988 trades in total, 5,944 in the first half and 6,046 in the second. The halves were re-measured on 17-Aug-2026 with herramientas/mitades_hmm.py, which leaves the result in MITADES-HMM-3R.json. Measured over 148 stocks and 16.6 years of daily data, with costs of 0.05 % commission and 0.05 % slippage, and without broker costs for holding the position open overnight. Each configuration is run 25 times shuffling the tie-break between same-day marks, and the median is published. Educational content, not financial advice. We don't sell signals or manage third-party capital. Familia FVR · Sophronepsis.