Repository navigation
Add TimesFM 3.0 adapters + forecast-quality diagnostics - #1
Conversation
TimesFM 3.0 (Aug 2026) shipped after the M5 Gen5 work, and the existing `timesfm` adapter targeted the 1.x API (`TimesFm`/`TimesFmHparams`) which no longer exists in the current pip package — it would have failed on any installable version. Models (CLAUDE.md §3 Gen5 forecast-residual adapters): - `timesfm3` native multivariate joint forecast (variate attention) - `timesfm3_ci` channel-independent control, same weights/settings - `timesfm3_prob` probabilistic score (CRPS over the 9 quantiles) - `timesfm` repointed at the 2.5 API, 1.x kept as fallback The three 3.0 variants are separate registered names on purpose: result aggregation dedups on (model, dataset, channel, seed), so same-name variants differing only in model_params would overwrite each other in the leaderboard. Shared forecast-residual protocol (foundation.py), now fixed across adapters so they stay comparable: zero-shot fit keeps the train tail as initial context (no more discarded test head), score slides with stride=horizon predicting each test step exactly once, and context/horizon/score_mode/batch_size are configurable and recorded in the run config. Two backend flags matter for TSAD and are pinned + tested: make_positive=False (3.0) and infer_is_positive=False (2.5) — z-scored input is signed, and the benchmark defaults would clamp forecasts to non-negative. Forecast-quality metrics (evaluation/forecast_metrics.py): MASE, WQL, quantile-CRPS — the metrics GIFT-Eval/fev-bench/TIME rank on — recorded as `fc_*` next to `vus_pr` on every forecast-based run, so "does a rank-1 forecaster detect better?" is a join rather than a new experiment. Benchmarks: GIFT-Eval, fev-bench and TIME are forecasting benchmarks with no anomaly labels, so they are NOT added as datasets — manufacturing labels for them is exactly the flaw Wu & Keogh (TKDE 2021) warn about. TSB-AD stays the TSAD counterpart. docs/benchmarks/forecasting-benchmarks.md (en+ko) records that reasoning and why a better forecaster can be a worse residual detector. Licensing: timesfm code is Apache-2.0 (verified from the packaged LICENSE) and used as a pip dependency only. TimesFM 3.0 *weights* are under timesfm-non-commercial-license-v1.0 — the adapters warn at fit() and THIRD_PARTY_NOTICES.md gains a pretrained-weight license section. Not run here: HuggingFace egress is blocked in this environment, so there are no TimesFM rows in benchmarks/results/ and no leaderboard entries. Adapters are covered by tests against a stubbed backend pinning array contracts, scoring modes and step coverage; configs/foundation.yaml runs the matrix where weights are reachable. 218 passed, 5 skipped; ruff/black/mypy/i18n-parity clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016zQhp1DpFGAypbAhhxkTiE
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 9102a9afe2
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| tail = getattr(self, "train_tail_", np.empty((0, D))) | ||
| if tail.shape[1] != D: # 채널 수가 다르면 꼬리를 쓰지 않는다 | ||
| tail = np.empty((0, D)) | ||
| ext = np.vstack([tail, X]) if len(tail) else X |
There was a problem hiding this comment.
Avoid wrapping calibration context from the train suffix
When spot, dspot, or conformal is selected, runner/experiment.py calls model.score(train) to obtain calibration scores. This code prepends train_tail_—which is the end of that same array—to the beginning of train, so forecasts for roughly the first context training points condition on future observations from the end of the series. The resulting leaked scores can materially distort the calibration distribution and threshold; scoring the fitted training array needs a chronological path that does not reuse its suffix as prior context.
Useful? React with 👍 / 👎.
| for k, v in (getattr(model, "diagnostics_", None) or {}).items(): | ||
| metrics[str(k)] = float(v) |
There was a problem hiding this comment.
Preserve test diagnostics before calibration scoring
For runs using spot, dspot, or conformal, the earlier model.score(train) calibration call replaces diagnostics_ with metrics computed on the training array, because every _ForecastResidualBase.score() assignment overwrites that attribute. Consequently this loop silently records training forecast quality as fc_*, despite the new metric contract and documentation saying these values describe the anomaly-containing test split. Snapshot the diagnostics immediately after score(test) or prevent calibration scoring from replacing them.
Useful? React with 👍 / 👎.
The torch-free CI `test` job failed at collection: tests/test_timesfm.py imports the gen5 package, whose __init__ pulls in mamba_tsad -> torch. The adapters themselves need no torch (each backend carries its own dependency), but the registry path is coupled to it, so the test module now guards with importorskip like tests/test_gen5_models.py. The CI `test-dl` job installs torch and runs it for real — it passed on the previous commit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016zQhp1DpFGAypbAhhxkTiE
What
TimesFM 3.0 (Aug 2026) shipped after the M5 Gen5 work, so it was missing. While adding it, the existing
timesfmadapter turned out to target the 1.x API (TimesFm/TimesFmHparams), which no longer exists in the current pip package — it would have failed on any installable version.Models (CLAUDE.md §3 Gen5 forecast-residual adapters)
timesfm3google/timesfm-3.0-pytorchtimesfm3_citimesfm3_probtimesfmgoogle/timesfm-2.5-200m-pytorchSeparate registered names on purpose: result aggregation dedups on
(model, dataset, channel, seed), so same-name variants differing only inmodel_paramswould silently overwrite each other in the leaderboard.Shared forecast-residual protocol
Fixed across all forecast adapters so they stay comparable: zero-shot
fit()keeps the train tail as initial context (no more discarded test head),score()slides with stride =horizonpredicting each test step exactly once, andcontext/horizon/score_mode/batch_sizeare configurable and recorded in the run config.Two backend flags matter for TSAD and are pinned and tested:
make_positive=False(3.0) andinfer_is_positive=False(2.5) — z-scored input is signed, and the benchmark defaults would clamp forecasts to non-negative.Forecast-quality diagnostics
evaluation/forecast_metrics.pyadds MASE, WQL and quantile-CRPS — the metrics GIFT-Eval / fev-bench / TIME rank on — recorded asfc_*next tovus_pron every forecast-based run. "Does a rank-1 forecaster detect better?" becomes a join rather than a new experiment. The interesting cell is lowfc_*with lowvus_pr: the model predicted the anomalies well, which is what a residual detector cannot afford.Benchmarks: what was deliberately not added
GIFT-Eval, fev-bench and TIME are forecasting benchmarks with no anomaly labels, so they are not added as datasets — manufacturing labels for them is exactly the flaw Wu & Keogh (TKDE 2021) warn about. TSB-AD stays the TSAD counterpart.
docs/benchmarks/forecasting-benchmarks.md(en + ko) records that reasoning and why a better forecaster can be a worse residual detector.Licensing
timesfmcode is Apache-2.0 (verified from the packaged LICENSE) and used as a pip dependency only — no vendoring. TimesFM 3.0 weights are undertimesfm-non-commercial-license-v1.0(non-commercial, non-production); the adapters emit aUserWarningnaming it atfit(), and THIRD_PARTY_NOTICES.md gains a pretrained-weight license section. 2.5 weights stay Apache-2.0 for commercial/production comparison.Not run here
HuggingFace egress is blocked in the environment this was written in, so there are no TimesFM rows in
benchmarks/results/and no leaderboard entries — claiming numbers without running them would violate CLAUDE.md §9. The adapters are covered by tests against a stubbed backend pinning array contracts (variate-axis transposition, quantile-axis alignment), scoring modes and step coverage.configs/foundation.yamlruns the matrix on the same entities as theliteprofile where weights are reachable:Verification
218 passed, 5 skipped; ruff, black, mypy and the docs i18n parity check all clean.
tests/test_timesfm.pyadded to the CI torch job.🤖 Generated with Claude Code
https://claude.ai/code/session_016zQhp1DpFGAypbAhhxkTiE
Generated by Claude Code