Skip to content

v1.2.0: temporal stability, quote sensitivities, variance events, market context, economic arbitrage - #37

Open
marwinsteiner wants to merge 78 commits into
mainfrom
feat/trading-context
Open

marwinsteiner wants to merge 78 commits into
mainfrom
feat/trading-context

Conversation

@marwinsteiner

Copy link
Copy Markdown
Owner

The trading layer: the workflows a desk runs ON a fitted surface --
stability across snapshots, quote-level risk, known events, coherent
conventions, and the economic meaning of a diagnostic flag.

  • Temporal stability: fit(quotes, prior=yesterdays_fit, anchor=...)
    warm-starts from the prior and adds a Tikhonov pull toward the prior
    SURFACE SHAPE (w on the penalty grid, deliberately not raw
    parameters -- near-degenerate SVI coordinates must not be pinned
    arbitrarily). anchor=0 reproduces the unanchored fit; identical
    quotes + identical prior give identical parameters, bitwise. On the
    real snapshot, a ~5bp iv perturbation moves the unanchored fit by
    0.24 in parameter space (a basin hop) and the anchored fit by ~1e-6.
  • Quote-to-surface Jacobians: quote_sensitivity /
    surface_sensitivity / iv_surface_sensitivity -- implicit-function
    Gauss-Newton sensitivities at the optimum, propagated through
    dw/dtheta: bump a quote, read the whole smile's response without
    recalibrating. FD-verified against bump-and-recalibrate.
  • Variance events: VarianceEvent(time, variance) decomposes
    w(k,T) = w_cont(k,T) + sum of events by T. Fitting subtracts,
    interpolation acts on the continuous component, evaluation adds the
    jump back on the correct side -- the two-expiry fixture reproduces
    the jump exactly where the naive blend smears it, calendar
    diagnostics raise no false violation across an event, events
    inconsistent with quoted variance are rejected, and events
    serialize with the surface. implied_event_variances reports the
    quoted ATM jump straddling each event as an upper bound.
  • MarketContext: one coherent source of valuation time, spot,
    rate/dividend views (any rate form) and day count (ACT/365F,
    ACT/360, ACT/365.25, BUS/252 + holidays).
    OptionChain.from_dataframe(context=...) takes real expiry DATES;
    put-call parity of surface prices holds against the context's
    discount factors by construction; mixing a context with separate
    numeraire inputs raises; day-count choice demonstrably changes T
    and the inverted vols.
  • Economic vs mathematical arbitrage + domain-of-validity status:
    iv(..., return_status=True) labels every point observed /
    interpolated / extrapolated from the quoted ranges on the fit
    report. classify_arbitrage(surface, panel=...) classifies each
    finding: extrapolation_risk (outside the quoted range),
    executable (a static butterfly at the quoted bid/ask prices has
    negative worst-case cost), quote_consistent (inside bid/ask
    uncertainty), or mathematical (no band supplied). On the real SPY
    snapshot every wing violation classifies as extrapolation risk --
    exactly the distinction the raw numbers could not make.
  • 21 new tests; 347 total, green on both backends. Docs and examples
    updated in the same PR per the maintenance policy: new example
    06_trading_workflows.py runs the whole layer on the real snapshot;
    new market-context docs page; sections on the calibration,
    identifiability, surface, and arbitrage pages.

Stacked on #33 (-> #32 -> #20 -> #19); the diff narrows as those merge.

Closes #27. Closes #28. Closes #29. Closes #30. Closes #31.

VolSurface.fit calibrates every maturity of an option panel through one
call for all seven parametrizations (per-slice theta/T/F derived
automatically, calibration controls pass through), or builds directly
from calibrated slices. Exposes vectorized evaluation (iv, total
variance, forward, ATM vol/skew/curvature, params), arbitrage
verification via the diagnostics module, and a Black-76 layer on the
slice forwards: price, forward delta, gamma, vega, and annualized theta
under sticky-strike conventions with a flat discount rate.
Fits all seven models from one panel; prices verified against py_vollib
and all four Greeks against finite differences of the Black price;
put-call parity and put theta via parity; construction validation.
Shared SURFACE_RATE constant keeps the fixture forwards and the pricing
assertions on one rate.
…d maturity interpolation

calibrate_surface orders expiries, threads each fitted slice's total
variance into the next slice's NO_CALENDAR penalty on the exact penalty
grid, clips SSVI/eSSVI theta(T) monotone, fits the eSSVI term structure
jointly across slices (identical shape parameters on every slice), and
warns once per surface against the Gatheral-Jacquier sufficient
no-butterfly bounds. VolSurface now interpolates between fitted
maturities: linear total-variance blend at fixed log-moneyness by
default (calendar-free between ordered slices, exact at fitted
maturities), or parametric theta interpolation for SSVI/eSSVI with
slice_at returning synthetic params; forwards interpolate log-linearly;
extrapolation raises.
Calendar-free verification without manual w_prev for four models, joint
eSSVI shape-sharing, repair of calendar-arbitrageable input data,
DirectSVI rejection, exact reproduction at quoted maturities, blend
monotonicity, log-linear forwards, pricing parity at interpolated
maturities, and the theta method round-trip.
VolSurface.fit now records per-slice fit evidence (status, quote
accounting, iv residual stats, quoted k-range) plus settings and
provenance in surface.fit_report; failed or rejected slices are a
visible property of the result, not a buried warning. diagnose()
combines the fit report with the arbitrage diagnostics, evaluated on
the quoted range by default, into one formatted result block.
# Conflicts:
#	src/pysvi/__init__.py
#	src/pysvi/surface.py
Sequential and joint-eSSVI calibration record per-slice evidence
(including insufficient and failed slices) with calendar_enforced set,
so diagnose() covers calendar-aware fits too.
calculate_implied_forward accepts a flat float or a callable T -> r(T);
flat behaviour unchanged.
from_dataframe cleans quotes, derives per-expiry implied forwards from
put-call parity (median over both-leg strikes, spot*exp((r-q)T)
fallback), selects OTM legs, and inverts Black-76 implied vols for mid,
bid, and ask on the implied forward - no spot or dividend assumption
needed when both legs quote. chain.fit goes straight to a VolSurface
(calendar-aware on request), and rejected quotes surface in the fit
report.
Captures model and arbitrage condition, per-slice parameters and
forwards, rate, interpolation method, and the fit report; load
validates schema_version and model name and reproduces evaluation
bitwise.
…rity

Hand-derived closed-form anchors per parametrization (provenance in
docstrings), pinned regression tables, boundary-parameter battery on
both backends, degenerate-calibration behaviour, and numba/NumPy parity
at the decision boundaries the diagnostics depend on, including verdict
stability for a near-boundary fit.
Runnable walkthrough (ingest, calendar-aware calibrate, audit with the
literal diagnose() output block, evaluate, price, persist); linked from
quickstart; chain module in the API reference.
When a strike quotes only the ITM leg, choose_leg falls back to it, but
the Black-76 inversion still used the OTM flag. An ITM price inverted
under the wrong flag yields an absurd-but-finite implied vol (observed
775% on real weekend SPY data) instead of a clean NaN, silently
corrupting the panel. The flag now follows the leg that was priced.
Found by the real-data examples; regression test included.
Top-level examples/ directory: five runnable scripts taking a real SPY
option chain (yfinance) from raw quotes to a priced, verified,
serialized surface, together covering every public endpoint and
parameter. 01 fetches a contemporaneous snapshot with the no-lookahead
discipline made mechanical (single recorded timestamp, snapshot-time
filters, persisted raw data; closed-book fallback to last prints when
Yahoo zeroes the weekend book); 02 walks forwards/IV/slice helpers and
surveys inversion methods (Black-76, Jaeckel's Let's Be Rational --
py_vollib's engine and thus what svi-py uses -- and Schadner's explicit
Volfi, with sources); 03 fits all seven models to one real slice and
tours the analytics API; 04 measures every calibration control on real
data; 05 runs the OptionChain -> calibrate_surface -> VolSurface
pipeline end to end. A committed sample snapshot makes 02-05 runnable
offline; regenerated snapshots and artifacts are gitignored.
New docs/examples.md documenting the examples/ scripts, the
no-lookahead snapshot discipline, the implied-vol inversion method
survey (Black-76 / Let's Be Rational / Volfi, with sources), and the
maintenance policy: examples and docs change in the same PR as any
public API change. Cross-linked from index, quickstart, the synthetic
walkthrough, the calibration pipeline page, and the README.
Milestone-driven weekly release train, PR conventions (title = tag
subject, body = release notes, release-ready gate, stacked branches),
tag-driven hatch-vcs versioning, and the tag-triggered CI/CD pipeline
(test, GitHub Release from the tag annotation, PyPI trusted
publishing) -- the entire SDLC, so contributors know how a change
reaches PyPI. Also states the examples+docs same-PR policy.
The jw exception is obsolete: multi_start's best-of now includes the
default start under default-path settings, so it matches or beats the
default initialization on every model.
…trols

Three code-review findings for v0.9.0:

- theta was pinned to the smile MINIMUM (nanmin of iv^2 T over raw
  rows), but SSVI fixes w(0) = theta exactly: on any skewed smile the
  minimum sits away from k = 0 and systematically understates the ATM
  level (~8% low fitted ATM vol on the repo's own fixture), which the
  (rho, eta) fit can never correct. theta is now w interpolated at
  k = 0 on the CLEANED slice inputs (prepare_slice), so a junk quote
  can no longer drive theta toward zero either; the raw minimum
  remains only as the fallback for slices too thin to clean.
  Regression test pins fitted ATM vol to the market ATM vol per slice.
- DirectSVI silently swallowed the calibration-control kwargs
  (objective/loss/f_scale/initialization/w_prev) while the fit report
  stamped them as applied; the closed-form fit now logs which controls
  it ignored.
- The six copy-pasted rank_objective closures collapse into a shared
  _rank_on_plain helper.
Three code-review findings for v0.10.0, all in _calibrate_essvi_global:

- f_scale for robust losses now follows the documented policy
  (explicit kwarg, else 1.4826 * MAD of the joint mode-space residuals
  at a pilot l2 fit) instead of silently defaulting to 1.0, under
  which huber/soft_l1/cauchy were numerically identical to l2.
- The multi-start call gains the validity filter (eta > 0) and the
  platform-stable plain-kernel ranking every per-slice path already
  has -- near-tied basins no longer resolve by fastmath runner noise.
- The butterfly/calendar grid comes from _penalty_grid instead of a
  duplicated inline linspace, so the grid policy has one owner.
The module docstring promised 'bid/ask IV bands ... chain.fit(...)
goes straight to a VolSurface', but nothing ever converted the
panel's iv_bid/iv_ask columns into the per-slice w_bid/w_ask arrays
the objective demands -- objective='bid_ask' through chain.fit or the
surface fitters raised deep in _prepare_loss_inputs, and passing one
band array through model_kwargs could never match every slice's
cleaned quote count.

prepare_slice gains return_index=True (positional indices of the
surviving rows), and both VolSurface.fit and calibrate_surface derive
per-slice bands from iv_bid/iv_ask filtered identically to the quotes
(_band_kwargs); rows with a missing or crossed band degenerate to a
zero-width band at the mid. Explicit w_bid/w_ask kwargs still win;
the joint eSSVI path still rejects bid_ask as documented. Example 05
now fits the real chain inside its quoted band; docs updated.
- calculate_implied_forward masks invalid rows BEFORE evaluating the
  rate view: curve and model objects raise on non-positive times where
  the float path just NaN'd the row -- a same-day/expired row no
  longer crashes the whole call under a DiscountCurve or model rate.
- compute_ivs_vectorized restores its NaN-per-row contract, broken by
  the narrowed exception tuple: an unrecognized flag now yields NaN
  (KeyError caught) instead of aborting the batch, and 'call'/'put'
  spellings normalize to 'c'/'p'.
- calibrate_surface honors mode for the cross-slice warnings: the
  monotone-theta clip raises under strict (a calendar-arbitrageable
  ATM term structure is bad input, naming the maturities), stays a
  warning under warn, and is silent under lenient; the SSVI
  admissibility warning is silent under lenient.
- monotone_cubic interpolators are cached per k-grid (bounded LRU) on
  the immutable surface: pricing a book at repeated strikes no longer
  re-evaluates every slice and rebuilds the PCHIP setup per call.
- identifiability_report reuses the Jacobian and residuals
  parameter_uncertainty computes instead of recomputing both (halves
  the FD-Jacobian cost for models without analytic overrides).
- warm_up honors its documented contract: returns 0.0 immediately
  when numba is not installed.
- RateLike is a real alias (float | Callable | zero_rate protocol)
  instead of Union[float, 'callable', object], which collapsed to
  object and typechecked nothing.
Five features spanning issues #27 #28 #29 #30 #31:

Temporal stability (#27): prior= warm-starts any iterative calibration
from a previous fit (with multi_start it becomes the base start) and
anchor= adds a Tikhonov pull toward the prior SURFACE SHAPE (w on the
penalty grid, not raw parameters -- near-degenerate coordinates must
not be pinned arbitrarily). anchor=0 reproduces the unanchored
objective; identical quotes + identical prior give identical
parameters. Surface-level prior=/anchor= on VolSurface.fit and
calibrate_surface match prior slices by maturity (joint eSSVI fits
unanchored with a warning).

Quote-to-surface Jacobians (#28): quote_sensitivity (dtheta/dquote via
the implicit-function Gauss-Newton system), surface_sensitivity
(propagated through dw/dtheta to evaluation points) and
iv_surface_sensitivity (vol-in, vol-out units) -- bump a quote, read
the smile's response without recalibrating.

Variance events (#29): VarianceEvent(time, variance, label); fitting
subtracts cumulative event variance (slices store the continuous
component; events inconsistent with quoted variance are rejected,
strict raises), interpolation acts on the continuous component, and
evaluation adds events back on the correct side of the event time --
the jump is reproduced instead of smeared and calendar diagnostics
raise no false violation across an event. Events serialize with the
surface; implied_event_variances reports the quoted ATM jump as an
upper bound per event.

MarketContext (#30): one coherent source of valuation time, spot,
rate/dividend views, and day count (ACT/365F, ACT/360, ACT/365.25,
BUS/252 with holidays). OptionChain.from_dataframe(context=...)
resolves real expiry DATES through the context's day count; mixing a
context with separate numeraire inputs raises.

Economic arbitrage + evaluation status (#31): iv(...,
return_status=True) labels every point observed / interpolated /
extrapolated from the fit report's quoted ranges;
classify_arbitrage(surface, panel=...) classifies findings as
extrapolation_risk / executable (static butterfly at quoted bid/ask
prices has negative worst-case cost) / quote_consistent /
mathematical.
21 new tests covering the acceptance criteria: anchored drift bounded
under a tiny perturbation with anchor=0 reproducing the unanchored
objective and bitwise determinism; sensitivity Jacobians verified
against bump-and-recalibrate; event-aware interpolation reproducing
the jump the naive blend smears, calendar clean across events,
oversized events rejected, serialization round-trip; MarketContext
date resolution, put-call parity against the context discount
factors, loud mixed-convention rejection, day-count effects on T and
iv; evaluation status labels including the no-report error path; and
the classification fixtures (quote_consistent vs executable on
crossed quotes, extrapolation risk outside the quoted range).
Temporal-stability section on the calibration page; quote-to-surface
sensitivities on the identifiability page; variance-events and
evaluation-status sections on the surface page; economic-vs-
mathematical classification on the arbitrage page; new market-context
page in the toctree and API reference. New example
06_trading_workflows.py runs the whole layer on the real SPY snapshot
-- anchored recalibration under a perturbation, ATM quote-bump
response, a synthetic event between real expiries, a dated
MarketContext chain, status-labeled evaluation, and the
classification labeling this data's wing violations as extrapolation
risk -- per the examples maintenance policy.
@marwinsteiner marwinsteiner added this to the v1.2.0 milestone Sep 24, 2026
@read-the-docs-community

read-the-docs-community Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

The FD reference recalibrates, so it carries optimizer-termination
noise that differs per platform under fastmath; a 5e-5 bump left no
headroom on the Linux runner. Larger bump, vector-norm agreement.
@codecov-commenter

codecov-commenter commented Sep 24, 2026 •

Copy link
Copy Markdown

⚠️ Please install the 'codecov app svg image' to ensure uploads and comments are reliably processed by Codecov.

Codecov Report

❌ Patch coverage is 92.41595% with 97 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
src/pysvi/surface.py 91.72% 51 Missing ⚠️
src/pysvi/chain.py 90.35% 11 Missing ⚠️
src/pysvi/diagnostics.py 82.81% 11 Missing ⚠️
src/pysvi/context.py 81.48% 10 Missing ⚠️
src/pysvi/models.py 95.96% 5 Missing ⚠️
src/pysvi/calibration.py 87.09% 4 Missing ⚠️
src/pysvi/report.py 97.24% 3 Missing ⚠️
src/pysvi/_kernels.py 98.14% 1 Missing ⚠️
src/pysvi/identifiability.py 99.05% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

Three release blockers plus accuracy fixes for v0.9.0:

- One junk expiry (all-NaN ivs) no longer poisons the eSSVI
  theta_ref and takes down every slice: the reference is the median
  of the finite, positive per-slice thetas only; the junk slice fails
  individually and stays visible on the report.
- The 'slices that fail to calibrate are skipped with a warning'
  contract now holds for EXCEPTIONS raised inside per-slice
  calibration (e.g. a singular matrix in DirectSVI's closed form on a
  flat junk slice), not only for a None return.
- diagnose() no longer reports false ARBITRAGE DETECTED on clean
  SABR/DirectSVI fits: the diagnostics default tolerance is
  model-aware (new Parametrization.diagnostics_tol, 1e-8 analytic,
  1e-4 for the finite-difference models, matching the documented FD
  noise) and an explicit tol= still overrides.
- quoted_range() only unions slices that made it into the surface: a
  failed slice's strike range must not widen the domain the
  diagnostics certify.
- np.interp seeds in JW and SABR calibration argsort k first
  (np.interp silently returns garbage on descending input).
- Docs and docstring no longer claim theta_ref applies to SSVI (it is
  eSSVI-only); fit_report annotated Optional.
…ception wrap, theta_ref guard in _data_thetas)
- Caller bugs surface instead of dissolving: a ValueError raised
  inside per-slice calibration (unknown loss/objective, missing band
  or prior kwargs) propagates, instead of every slice being skipped
  and the user getting 'no slice calibrated successfully' for a
  one-character typo; non-ValueError exceptions (numerical failures)
  keep the skip-and-record contract.
- An explicitly requested NO_CALENDAR arbitrage condition implies
  calendar enforcement: previously enforce_calendar=False silently
  disabled the chaining and (for eSSVI) the check the flag asked for,
  while the butterfly flag was honored -- an invisible asymmetry.
- Theta interpolation of eSSVI slices carrying only rho_theta (a
  params shape the model documents as sufficient) keeps the blended
  rho_theta instead of raising KeyError on rho0/rho1/alpha.
- The joint eSSVI plain-kernel rank objective is now fully plain:
  calendar w_prev chaining inside the rank path uses the plain
  essvi_w, closing the last fastmath leak into candidate ranking.
- Exact duplicate maturities hit the explicit ValueError instead of a
  TypeError from sorted() comparing params dicts.
- interp_method is validated at calibrate_surface entry, not after
  the entire multi-slice calibration.
- Dead scaffolding removed from test_interpolated_slice_calendar_free.
numpy's LinAlgError subclasses ValueError, so the caller-bug re-raise
would have propagated exactly the singular-matrix numerical failures
the skip-and-record contract exists for; they are now caught first.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment