Skip to content

Add TimesFM 3.0 adapters + forecast-quality diagnostics - #1

Merged
Denny-Hwang merged 2 commits into
mainfrom
claude/timesfm-model-benchmark-2l3gd0
Sep 2, 2026
Merged

Denny-Hwang merged 2 commits into
mainfrom
claude/timesfm-model-benchmark-2l3gd0

Conversation

@Denny-Hwang

Copy link
Copy Markdown
Owner

What

TimesFM 3.0 (Aug 2026) shipped after the M5 Gen5 work, so it was missing. While adding it, the existing timesfm adapter turned out to target the 1.x API (TimesFm / TimesFmHparams), which no longer exists in the current pip package — it would have failed on any installable version.

Models (CLAUDE.md §3 Gen5 forecast-residual adapters)

Registered name Checkpoint Isolates
timesfm3 google/timesfm-3.0-pytorch native multivariate joint forecast (variate attention), point residual
timesfm3_ci same weights channel-independent control — variate attention off
timesfm3_prob same weights probabilistic score (CRPS over the 9 quantiles)
timesfm google/timesfm-2.5-200m-pytorch repointed at the 2.5 API, 1.x kept as fallback

Separate registered names on purpose: result aggregation dedups on (model, dataset, channel, seed), so same-name variants differing only in model_params would silently overwrite each other in the leaderboard.

Shared forecast-residual protocol

Fixed across all forecast adapters so they stay comparable: zero-shot fit() keeps the train tail as initial context (no more discarded test head), score() slides with stride = horizon predicting each test step exactly once, and context / horizon / score_mode / batch_size are configurable and recorded in the run config.

Two backend flags matter for TSAD and are pinned and tested: make_positive=False (3.0) and infer_is_positive=False (2.5) — z-scored input is signed, and the benchmark defaults would clamp forecasts to non-negative.

Forecast-quality diagnostics

evaluation/forecast_metrics.py adds MASE, WQL and quantile-CRPS — the metrics GIFT-Eval / fev-bench / TIME rank on — recorded as fc_* next to vus_pr on every forecast-based run. "Does a rank-1 forecaster detect better?" becomes a join rather than a new experiment. The interesting cell is low fc_* with low vus_pr: the model predicted the anomalies well, which is what a residual detector cannot afford.

Benchmarks: what was deliberately not added

GIFT-Eval, fev-bench and TIME are forecasting benchmarks with no anomaly labels, so they are not added as datasets — manufacturing labels for them is exactly the flaw Wu & Keogh (TKDE 2021) warn about. TSB-AD stays the TSAD counterpart. docs/benchmarks/forecasting-benchmarks.md (en + ko) records that reasoning and why a better forecaster can be a worse residual detector.

Licensing

timesfm code is Apache-2.0 (verified from the packaged LICENSE) and used as a pip dependency only — no vendoring. TimesFM 3.0 weights are under timesfm-non-commercial-license-v1.0 (non-commercial, non-production); the adapters emit a UserWarning naming it at fit(), and THIRD_PARTY_NOTICES.md gains a pretrained-weight license section. 2.5 weights stay Apache-2.0 for commercial/production comparison.

Not run here

HuggingFace egress is blocked in the environment this was written in, so there are no TimesFM rows in benchmarks/results/ and no leaderboard entries — claiming numbers without running them would violate CLAUDE.md §9. The adapters are covered by tests against a stubbed backend pinning array contracts (variate-axis transposition, quantile-axis alignment), scoring modes and step coverage. configs/foundation.yaml runs the matrix on the same entities as the lite profile where weights are reachable:

pip install "tsad-forge[foundation]" "timesfm[torch]" chronos-forecasting momentfm
python benchmarks/run_all.py --profile configs/foundation.yaml

Verification

218 passed, 5 skipped; ruff, black, mypy and the docs i18n parity check all clean. tests/test_timesfm.py added to the CI torch job.

🤖 Generated with Claude Code

https://claude.ai/code/session_016zQhp1DpFGAypbAhhxkTiE


Generated by Claude Code

TimesFM 3.0 (Aug 2026) shipped after the M5 Gen5 work, and the existing
`timesfm` adapter targeted the 1.x API (`TimesFm`/`TimesFmHparams`) which no
longer exists in the current pip package — it would have failed on any
installable version.

Models (CLAUDE.md §3 Gen5 forecast-residual adapters):
- `timesfm3`      native multivariate joint forecast (variate attention)
- `timesfm3_ci`   channel-independent control, same weights/settings
- `timesfm3_prob` probabilistic score (CRPS over the 9 quantiles)
- `timesfm`       repointed at the 2.5 API, 1.x kept as fallback

The three 3.0 variants are separate registered names on purpose: result
aggregation dedups on (model, dataset, channel, seed), so same-name variants
differing only in model_params would overwrite each other in the leaderboard.

Shared forecast-residual protocol (foundation.py), now fixed across adapters
so they stay comparable: zero-shot fit keeps the train tail as initial context
(no more discarded test head), score slides with stride=horizon predicting each
test step exactly once, and context/horizon/score_mode/batch_size are
configurable and recorded in the run config.

Two backend flags matter for TSAD and are pinned + tested: make_positive=False
(3.0) and infer_is_positive=False (2.5) — z-scored input is signed, and the
benchmark defaults would clamp forecasts to non-negative.

Forecast-quality metrics (evaluation/forecast_metrics.py): MASE, WQL,
quantile-CRPS — the metrics GIFT-Eval/fev-bench/TIME rank on — recorded as
`fc_*` next to `vus_pr` on every forecast-based run, so "does a rank-1
forecaster detect better?" is a join rather than a new experiment.

Benchmarks: GIFT-Eval, fev-bench and TIME are forecasting benchmarks with no
anomaly labels, so they are NOT added as datasets — manufacturing labels for
them is exactly the flaw Wu & Keogh (TKDE 2021) warn about. TSB-AD stays the
TSAD counterpart. docs/benchmarks/forecasting-benchmarks.md (en+ko) records
that reasoning and why a better forecaster can be a worse residual detector.

Licensing: timesfm code is Apache-2.0 (verified from the packaged LICENSE) and
used as a pip dependency only. TimesFM 3.0 *weights* are under
timesfm-non-commercial-license-v1.0 — the adapters warn at fit() and
THIRD_PARTY_NOTICES.md gains a pretrained-weight license section.

Not run here: HuggingFace egress is blocked in this environment, so there are
no TimesFM rows in benchmarks/results/ and no leaderboard entries. Adapters are
covered by tests against a stubbed backend pinning array contracts, scoring
modes and step coverage; configs/foundation.yaml runs the matrix where weights
are reachable.

218 passed, 5 skipped; ruff/black/mypy/i18n-parity clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016zQhp1DpFGAypbAhhxkTiE

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 9102a9afe2

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +174 to +177
tail = getattr(self, "train_tail_", np.empty((0, D)))
if tail.shape[1] != D: # 채널 수가 다르면 꼬리를 쓰지 않는다
tail = np.empty((0, D))
ext = np.vstack([tail, X]) if len(tail) else X

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Avoid wrapping calibration context from the train suffix

When spot, dspot, or conformal is selected, runner/experiment.py calls model.score(train) to obtain calibration scores. This code prepends train_tail_—which is the end of that same array—to the beginning of train, so forecasts for roughly the first context training points condition on future observations from the end of the series. The resulting leaked scores can materially distort the calibration distribution and threshold; scoring the fitted training array needs a chronological path that does not reuse its suffix as prior context.

Useful? React with 👍 / 👎.

Comment on lines +140 to +141
for k, v in (getattr(model, "diagnostics_", None) or {}).items():
metrics[str(k)] = float(v)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve test diagnostics before calibration scoring

For runs using spot, dspot, or conformal, the earlier model.score(train) calibration call replaces diagnostics_ with metrics computed on the training array, because every _ForecastResidualBase.score() assignment overwrites that attribute. Consequently this loop silently records training forecast quality as fc_*, despite the new metric contract and documentation saying these values describe the anomaly-containing test split. Snapshot the diagnostics immediately after score(test) or prevent calibration scoring from replacing them.

Useful? React with 👍 / 👎.

The torch-free CI `test` job failed at collection: tests/test_timesfm.py
imports the gen5 package, whose __init__ pulls in mamba_tsad -> torch.

The adapters themselves need no torch (each backend carries its own
dependency), but the registry path is coupled to it, so the test module now
guards with importorskip like tests/test_gen5_models.py. The CI `test-dl` job
installs torch and runs it for real — it passed on the previous commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016zQhp1DpFGAypbAhhxkTiE
@Denny-Hwang
Denny-Hwang merged commit de2df00 into main Sep 2, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants