Overfitting is the error of tuning a strategy so tightly to a particular stretch of past data that it captures the noise in that data rather than a durable effect. It looks excellent on the history it was fitted to and performs poorly on data it has never seen; the gap between the two is its measurable signature.
Also seen as: Curve-fitting, curve fitting
How does Fincanva handle it?
You define the rules. Fincanva has no automatic optimizer: it does not search parameter values or tune your thresholds against history for you, so a backtest reports what your exact stated rules would have done and nothing else — you stay in control of how closely they are shaped to history. The start-date sensitivity view re-runs one strategy across many entry dates and holding windows and reports the spread, a direct robustness check on a single flattering run. A backtest always runs to the latest available market close, so recent history is never held back automatically: a chronological in-sample / out-of-sample split is one you construct yourself from the periods you compare, primarily via the simulation start year.
What counts as a robust result?
A result is robust when it does not depend on one exact set of parameter values, one start date, or one sample of instruments. In practice that is a broad plateau instead of a sharp peak — small changes to a threshold move the outcome slightly rather than destroying it — plus a narrow spread of outcomes across many start dates. A result that exists only at one precise parameter setting is the textbook profile of an overfitted one, and a result that only survives with costs left out carries cost-ignoring bias as well, so repeat the check on the net figures. Robustness describes how a result behaves under perturbation; no level of it makes a strategy safe or a future return likely.
Why does an overfitted strategy fail out of sample?
An overfitted strategy is not a strategy that was wrong about the past — it described the past extremely well, which is exactly the problem: it encoded accidents that will not recur. Overfitting is an error of the fitting step: it lives in the strategy's own complexity, not in the sample you chose to test or the number of variants you tried before this one.
An overfitted strategy fails out of sample because every historical price series contains two components — a repeatable one and an unrepeatable one — and fitting cannot tell them apart. As you add rules, thresholds, and exceptions, the strategy gains the flexibility to describe finer and finer detail in the sample. Some of that detail is a real effect that will show up again; the rest is coincidence specific to those particular dates. Past a certain point, each extra degree of freedom buys mostly coincidence, so in-sample performance keeps improving while out-of-sample performance flattens and then deteriorates.
This is the standard bias–variance trade-off applied to trading rules: a very simple rule may systematically miss part of the real effect, while a very flexible one reproduces the sample almost exactly and generalises badly. The complexity that performs best on the fitted sample is therefore almost never the complexity that performs best afterwards.
How many parameter combinations does a search really cover?
The number of distinct strategies a parameter search covers is the product of the values tried for each parameter, so it grows multiplicatively, not additively.
- the number of distinct strategy variants the search covers
- the number of values tried for each of the k tunable parameters
Four parameters with ten candidate values each is not forty tests — it is distinct strategies, of which the best-looking one will look very good on that sample whether or not any real effect exists. This is the mechanical link between overfitting and data-snooping bias: the more variants a search covers, the more the winner's apparent quality is explained by the size of the search.
The two are the same statistical problem seen from two ends: overfitting describes the strategy, which has too many degrees of freedom for the data and has absorbed noise; data-snooping describes the search, in which many candidates were tried and only the winner is reported. Most real cases are both — how data-snooping differs from overfitting separates them case by case.
How can a 12-parameter strategy be perfect in sample and dead out of sample?
Consider two versions of the same idea, both fitted on 2000–2012 and then run unchanged on 2013–2025.
| Rules / parameters | 2000–2012 (fitted) | 2013–2025 (unseen) | |
|---|---|---|---|
| Simple version | 2 | +11% CAGR, −28% max drawdown | +9% CAGR, −31% max drawdown |
| Tuned version | 12 | +19% CAGR, −14% max drawdown | +2% CAGR, −39% max drawdown |
On the fitted window the tuned version is clearly the better strategy on every figure, and that is the result a single backtest would have shown. On the unseen window it collapses, while the simple version behaves roughly as it did before. The 12 parameters did not discover a better strategy; they described 2000–2012 more precisely — including the parts of it that were accidents. Note also which figure moved most: the tuned version's drawdown nearly tripled, because tuning had quietly removed the specific historical declines it was fitted to avoid rather than teaching the strategy to avoid declines in general.
What is an in-sample / out-of-sample split?
An in-sample / out-of-sample split is the standard defence against overfitting: the history is divided in two, the strategy is designed and tuned on the first part only (in sample), and the second part (out of sample) is then run once, unchanged, as a test of whether the result survives on data that played no part in shaping it. The out-of-sample result is the one that carries information, because it is the only one the rules were not fitted to.
Two conditions make the split meaningful. The out-of-sample period must be genuinely untouched — each time you look at it, adjust the rules, and look again, it becomes part of the fitting sample and stops being a test. And the split must be chronological rather than random, so the earlier data trains and the later data tests, which also keeps the exercise free of look-ahead bias. Repeating the split as a series of rolling train-then-test windows is walk-forward validation, the sequential form of the same idea.