Tools · Track record significance
How long before a record means something?
A Sharpe ratio measured over a short period carries a wide margin of error. This works out how many years a record of a given quality needs before the result can be separated from chance, and how strong the evidence is at the length you already have.
Calculator
This page is showing a worked example. The figures below are for an annualised Sharpe ratio of 0.80 observed over 3 years, tested at 95% confidence. Enable JavaScript to enter your own figures; the table, curve and method on this page do not need it.
- Years needed
- 4.2
- at the chosen confidence
- t‑statistic
- 1.39
- over 3 years
- Probability from luck
- 0.083
- one‑sided, if the true ratio were zero
The curve is an inverse square: halving the Sharpe ratio multiplies the record length needed by four. It is capped at 50 years for drawing.
Reference
Years of record needed.
Read down the Sharpe column and across to the confidence you want. A Sharpe of 0.5 needs about 10.8 years at 95%; a Sharpe of 1.0 needs about 2.7.
| Sharpe ratio | 90% confidence | 95% confidence | 99% confidence |
|---|---|---|---|
| 0.25 | 26.3 | 43.3 | 86.6 |
| 0.50 | 6.6 | 10.8 | 21.6 |
| 0.75 | 2.9 | 4.8 | 9.6 |
| 1.00 | 1.6 | 2.7 | 5.4 |
| 1.50 | 0.7 | 1.2 | 2.4 |
| 2.00 | 0.4 | 0.7 | 1.4 |
This is why long records are asked for, and why a strong short record is not the same as a strong manager. Most track records that exist are shorter than the length their own ratio would require.
Method
How it is calculated.
The standard test for whether an observed Sharpe ratio differs from zero, under the assumption that returns are independent and identically distributed.
The t-statistic
t = Sharpe × √years
The standard error of a Sharpe ratio falls with the square root of the length of the record, so the evidence grows with √years rather than with years. A Sharpe of 0.80 over 3 years gives t = 1.39.
The record length needed
years = (z ÷ Sharpe)²
Rearranged for the length at which the t-statistic reaches the critical value z: 1.2816 at 90%, 1.6449 at 95%, 2.3263 at 99%, one-sided. The probability shown is 1 − Φ(t), computed from a normal approximation accurate to about 1.5 × 10⁻⁷.
Limits
What this test assumes.
- Independent returns. The test assumes each month is independent of the last. Real returns are autocorrelated, particularly in less liquid holdings, and autocorrelation makes a record look steadier than it is. Where it is present, the true record length needed is longer than this.
- Normality. The critical values come from a normal distribution. Return series with fat tails or strong skew need more evidence, not less.
- One hypothesis, chosen in advance. The test asks whether one ratio differs from zero. Selecting the best record out of many and then testing it is a different question with a much higher bar, and this calculator does not answer it.
- Statistical significance is not merit. Clearing the threshold says the result is unlikely to be chance alone. It says nothing about whether the process is repeatable, understood, or suited to any particular purpose.
This calculator is provided for general information only and is directed to wholesale and professional investors. It is not personal advice: it does not take into account the objectives, financial situation or needs of any person, and it is not an offer, invitation or recommendation to acquire any financial product. Investing involves risk, including the possible loss of capital. Past performance is not a reliable indicator of future performance.
Definitions
Terms on this page.
Each links to the glossary entry, which states the convention the term assumes as well as what it means.
- t-statistic — how many standard errors from zero.
- Statistical significance — evidence against chance, not a measure of merit.
- Standard error — how much an estimate varies from sample to sample.
- Sharpe ratio — the ratio being tested.
Questions
Common questions.
How many years of track record are needed to prove skill?
It depends entirely on the Sharpe ratio. At 95% confidence a Sharpe of 0.5 needs about 10.8 years, a Sharpe of 1.0 about 2.7 years, and a Sharpe of 2.0 under a year. The relationship is an inverse square, so weaker results need disproportionately longer records. And statistical significance is evidence against pure chance, not proof of skill.
Why does the evidence grow with the square root of time?
Because the standard error of the mean falls with √n. Doubling the length of a record does not double the confidence in it; it multiplies the t-statistic by about 1.41. To double the t-statistic you need four times the record.
What is a t-statistic of 2?
Roughly the 97.7th percentile of a normal distribution one-sided, so about a 2.3% probability of seeing a result at least that strong if the true ratio were zero. It is a common informal threshold, and it sits between the 95% and 99% critical values in this calculator.
Does a short record mean a manager has no skill?
No. It means the record alone cannot settle the question. That is an argument for looking at process, risk architecture and the reasoning behind positions, rather than for treating a short record as either proof or disproof.