Skip to content

fix(returns): use non-excess kurtosis and omit NaNs in deflated Sharpe ratio - #872

Merged
polakowo merged 1 commit into
polakowo:masterfrom
Cashubski:fix/deflated-sharpe-kurtosis
Sep 25, 2026
Merged

polakowo merged 1 commit into
polakowo:masterfrom
Cashubski:fix/deflated-sharpe-kurtosis

Conversation

@Cashubski

Copy link
Copy Markdown
Contributor

Summary

ReturnsAccessor.deflated_sharpe_ratio deviates from the formula it implements in two ways.

1. Excess kurtosis where the formula needs non-excess kurtosis. The DSR of Bailey & López de Prado uses the Sharpe ratio standard error

sqrt((1 - γ3·SR + (γ4 - 1)/4 · SR²) / (T - 1))

where γ4 is the non-excess kurtosis: for Gaussian returns γ4 = 3 and the expression reduces to the classic Lo (2002) standard error sqrt((1 + SR²/2) / (T - 1)). The accessor passes scipy.stats.kurtosis(returns), which returns the excess kurtosis by default (0 for Gaussian returns), so the term is off by 3/4 · SR². For Gaussian returns the variance becomes 1 - SR²/4 instead of 1 + SR²/2, and for a large per-period SR the argument of the square root can go negative, which is exactly why the existing test expected NaN for two of the three columns of the fixture.

2. Missing returns counted as zero-return periods. Before computing skewness and kurtosis, NaNs were replaced by 0.0, and backtest_horizon used the full number of rows. sharpe_ratio ignores NaNs, so the DSR mixed a NaN-aware Sharpe ratio with moments and a horizon computed on a padded series. On a series with 50 missing days out of 300, the result differs from the same series with the missing days removed, which it should not.

Changes

  • kurtosis(..., fisher=False) so the non-excess kurtosis enters the formula.
  • nan_policy="omit" for skewness and kurtosis, and a per-column backtest_horizon = number of observed returns, consistent with sharpe_ratio. metrics.deflated_sharpe_ratio already broadcasts, so it now accepts a per-column horizon (type hint and docstring updated, with the reference).
  • Tests: the stored fixture values are updated (they were the previous output, including the NaNs) and two tests are added. One checks the accessor against the reference formula computed independently, and that Gaussian returns recover Lo's standard error; the other checks that a missing return is treated as absent rather than as a zero return.

Reproduction on the current master, Gaussian returns with skew ≈ 0 and kurtosis ≈ 3 (per-period SR ≈ 0.05):

import numpy as np, pandas as pd, vectorbt as vbt
from scipy.stats import norm, skew, kurtosis
from vectorbt.returns.metrics import approx_exp_max_sharpe

rng = np.random.default_rng(1); T = 2520
r = pd.DataFrame(rng.normal(0.0005, 0.01, (T, 3)), index=pd.bdate_range("2015-01-01", periods=T), columns=list("abc"))
acc = r.vbt.returns(freq="D", year_freq="252 days")
sr = acc.sharpe_ratio().values / np.sqrt(252)
sr0 = approx_exp_max_sharpe(0, np.var(acc.sharpe_ratio().values, ddof=1) / 252, 3)
g3, g4 = skew(r.values, axis=0), kurtosis(r.values, axis=0, fisher=False)
print(acc.deflated_sharpe_ratio().values)                                              # [0.992681 0.792372 0.410755]
print(norm.cdf((sr - sr0) * np.sqrt(T - 1) / np.sqrt(1 - g3 * sr + (g4 - 1) / 4 * sr**2)))  # [0.992586 0.792238 0.410766]

The gap is small at daily frequency because SR² is tiny per day; it grows with the per-period Sharpe ratio (monthly data) and with kurtosis, and the NaN handling can move the result by more than the kurtosis term (see the new test).

References

  • Bailey, D. H., & López de Prado, M. (2014). The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality. The Journal of Portfolio Management, 40(5).
  • Bailey, D. H., & López de Prado, M. (2012). The Sharpe Ratio Efficient Frontier. Journal of Risk, 15(2).
  • Lo, A. W. (2002). The Statistics of Sharpe Ratios. Financial Analysts Journal, 58(4).

…e ratio

The Deflated Sharpe Ratio standard error of Bailey and López de Prado,
sqrt((1 - g3 * SR + (g4 - 1) / 4 * SR^2) / (T - 1)), takes g4 as the
non-excess kurtosis (3 for Gaussian returns, which recovers the Lo (2002)
standard error sqrt((1 + SR^2 / 2) / (T - 1))). scipy.stats.kurtosis
returns the excess kurtosis by default, so the term was off by 3 / 4 * SR^2
and could turn the variance negative, yielding NaN.

Missing returns were also replaced by zero-return periods before computing
the moments and the horizon, diluting skewness and kurtosis and inflating
the horizon, while sharpe_ratio ignores them. They are now omitted from the
moments and the per-column horizon.

Tests use the reference formula directly rather than only stored values.
@polakowo
polakowo merged commit bdb7f8b into polakowo:master Sep 25, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants