Skip to content

feat(licensing): add default-deny engine and model licensing audit - #57

Open
Pushkraj-Space wants to merge 1 commit into
october-dev:mainfrom
Pushkraj-Space:issue-33-licensing-audit-engine-dependencies-and
Open

Pushkraj-Space wants to merge 1 commit into
october-dev:mainfrom
Pushkraj-Space:issue-33-licensing-audit-engine-dependencies-and

Conversation

@Pushkraj-Space

Copy link
Copy Markdown
Contributor

Summary

Refs #33. This PR does not close the issue: qualified legal review (criterion 10) is an external gate, and the final
code audit still has four open findings
(see "Audit result").

Murmur is about to integrate FluidAudio (#42) and sherpa-onnx (#26). An SDK's license is not the license of the
weights, tokenizers, phonemizers, voices, and prebuilt binaries it fetches or links. This PR adds a machine-readable,
default-deny licensing record and a CI gate, so that no engine or model artifact can be packaged or downloaded without
an exact, pinned, reviewed entry.

Changed files: 4 new and 5 modified.

File Change
licensing/manifest.json New: the inventory (71 entries)
licensing/README.md New: contract, policy, download rule, findings, update procedure, counsel questions
tool/check_licensing.py New: stdlib-only checker
tool/test_check_licensing.py New: 100 unittest cases
Makefile New check-licensing target, added to check
.github/workflows/ci.yml The protocol job runs check-licensing
THIRD_PARTY_NOTICES.md Pointer section only
docs/voice-runtime.md Model packs resolve only approved entries
CONTRIBUTING.md Licensing update procedure for engine, model and fixture changes

Implementation

Inventory (licensing/manifest.json)

The manifest has 71 entries: 52 bundle candidates, 5 user-download candidates, 12 blocked, and 2 fixture sets.

  • Engines. There is one named build configuration per engine, and requires is complete for exactly that
    configuration. Both configurations are provisional until [FluidAudio] Build the Apple Silicon Murmur engine #42 and [Offline] Add a cross-platform sherpa-onnx engine #26 confirm them.
    • FluidAudio v0.17.4 (21493f8) is the SwiftPM default-traits build on macOS 14+ arm64.
    • sherpa-onnx v1.13.8 (11afbd0) is a recognition-only static build.
      • SHERPA_ONNX_USE_PRE_INSTALLED_ONNXRUNTIME_IF_AVAILABLE=OFF, so no unhashed ONNX Runtime from the build machine
        is linked.
      • SHERPA_ONNX_LINK_LIBSTDCPP_STATICALLY=OFF and SHERPA_ONNX_USE_STATIC_CRT=OFF, so C/C++ runtimes come from the
        OS and are not packaged.
  • Native closure.
    • NemoTextProcessing (xcframework, SHA-256 5fa8c10d…) statically links 37 Rust crates, the Rust 1.98.1 std,
      and compiler-builtins. The closure comes from cargo tree --locked -e normal,no-proc-macro over the seven Apple
      targets of its build script. It was cross-checked against panic locations and the rustc version string in every
      release-binary slice. Each crate is pinned by its crates.io archive hash (equal to its Cargo.lock checksum) and its
      VCS commit.
    • FluidAudio vendored code: fastcluster, VBx, and the Japanese G2P code. Upstream hops cover the projects its
      library ports: FunASR, NeMo, misaki, ZipVoice, Chatterbox, StyleTTS2 and mobius.
    • sherpa-onnx sources: kaldi-native-fbank, kaldi-decoder, kaldifst, OpenFst, Eigen, simple-sentencepiece and
      nlohmann/json. Every CMake source archive was downloaded and hash-verified against sherpa-onnx's pins.
    • ONNX Runtime rebuild: four per-platform archives, each hash-matched to its GitHub digest and CMake pin.
  • Models, tokenizers, phonemizers and voices for [FluidAudio] Build the Apple Silicon Murmur engine #42: Parakeet TDT v3 with its vocabulary, Parakeet EOU with its
    vocabulary, Silero VAD, Kokoro ANE English with its vocabulary, the af_heart voice, the lexicon and the G2P, and the
    two legacy diarization models. Each is pinned to a Hugging Face commit and split per component wherever provenance
    differs.
  • Recorded for completeness: espeak-ng and piper-phonemize, which are GPL-3.0 and unreachable from the chosen
    configuration.
  • Fixtures: conformance-audio-fixtures and conformance-json-fixtures cover every tracked file under
    conformance/fixtures/, each with a synthetic method.

Blocked, each with a recorded reason:

  • luxtts-en-us-g2p-lexicon: a lexicon harvested from espeak-ng, bundled into every FluidAudio build, with no stated
    license. This gates fluidaudio.
  • kokoro-english-lexicon and kokoro-english-g2p: undocumented source data and weights. These gate
    kokoro-82m-coreml-ane.
  • parakeet-realtime-eou-120m-coreml and its vocabulary: the NVIDIA Open Model License text has not been captured.
  • pyannote-segmentation-legacy-coreml and wespeaker-v2-legacy-coreml: explicitly outside their repository's
    CC-BY-4.0 scope, with no recorded source.
  • kokoro-spanish-french-g2p: espeak-generated pronunciations with no data license.
  • sherpa-onnx: it compiles a Buckwalter table from an unlicensed repository and a StackOverflow snippet, alongside
    Apache-2.0 Kaldi/k2/icefall/CATT code and zlib-style cpp-base64 code.
  • onnxruntime: a third-party rebuild with no declared license or reproducible provenance.
  • espeak-ng and piper-phonemize: GPL-3.0, TTS only.

Operational finding. FluidAudio's model downloader fetches Hugging Face main for every repository except the
diarizer, so #42 must fetch from the manifest's pinned files.

Checker (tool/check_licensing.py)

The checker is stdlib-only, in the style of check_conformance.py. It exposes a pure
validate(manifest, notices, tracked_files), and --fingerprint <id> prints the value to record on approval.

  • Shape.
    • Unknown fields and duplicate JSON keys are rejected.
    • Ids and enums are exact, and URLs are https://.
    • requires must resolve, be acyclic, and never point to an engine or fixture.
  • Pins.
    • Every entry needs a 40-hex revision, byte-hashed downloads, or both.
    • A Hugging Face download must use the canonical host and the path /<source repo>/resolve/<entry revision>/.
    • A commit embedded in source must equal revision.
  • Evidence.
    • Evidence must be a single file (blob/raw, or resolve on Hugging Face) on github.com, gitlab.com or
      huggingface.co.
    • It must be in the entry's vcs-or-source repository at its revision; for a hop, in the hop's repository at
      its revision.
  • Ambiguity forces blocked:
    • an unknown permission or download term
    • an unidentified kind-specific license
    • missing evidence
    • a custom license without an exact name, at top level or on a hop
    • a weak upstream hop
    • a rejected review.
  • Approval requires:
    • a legal reference and date
    • a matching SHA-256 fingerprint of every reviewed field, so any bump invalidates the approval
    • commercialUse: allowed and a recorded attribution
    • every requirement approved
    • for a bundle: a notice heading, plus byte-hashed downloads for content kinds
    • for a user-download: a terms URL and byte-hashed downloads.
  • Fixtures and tracked files.
    • Tracked fixture files must equal the union of fixture paths.
    • Any tracked .onnx, .mlmodelc/, .so, .wav or similar file must be registered. The match ignores case.

Tests

  • make check-licensing: pass, with 100 of 100 unittest cases. The checker reports "71 entries: 52 bundle,
    5 user-download, 12 blocked, 2 fixture; 0 legally approved".
  • make check-conformance check-python: pass.
  • git diff --check and python3 -m py_compile on both tool files: pass.
  • The cases cover:
    • pins and shape
    • download and evidence binding, including repository/revision mismatches, host variants, traversal, and
      directory/pathless evidence
    • ambiguity rules
    • custom upstream licenses
    • fingerprint invalidation for each of 17 reviewed fields
    • legal-status rules
    • notices
    • content bundles without hashes
    • fixture provenance
    • case-insensitive tracked-file scanning
    • the sherpa-onnx runtime and ONNX Runtime selection.
  • Fuzzing: malformed-type fuzz passes over test and real entries produced no crashes.
  • Manual verification:
    • Every source, evidence, and download URL was checked (HTTP 200). The 37 crates.io pages reject automated clients,
      so those crate versions were checked through the crates.io API instead.
    • The xcframework, ONNX Runtime and CMake source archives were hash-checked.
  • make check-protocol was not run locally, because protoc is not installed. Its inputs are unchanged, and CI
    runs it.

Audit result

Three audit-and-fix rounds resolved 14 findings:

  • Hugging Face download binding, the EOU vocabulary, and Kokoro lexicon provenance.
  • An exact NemoTextProcessing closure.
  • The diarization split and case-insensitive scanning.
  • ONNX Runtime pre-install pinning.
  • Approval fingerprints, evidence binding, and custom upstream license names.
  • Content-bundle hashes, sherpa-onnx vendored code, runtime linkage, and file-only evidence.

The final audit verdict is FAIL, with four findings still open in this branch:

  1. HIGH: tracked-artifact kind compatibility. A tracked .onnx or .mlmodelc can be registered under a code entry
    (for example an approved engine), so it never gets a license.content review. The fix is scan-category and kind
    compatibility checks, with negative tests.
  2. HIGH: prebuilt code without byte hashes. An approved native-library pinned only by revision validates even
    when it stands for a fetched release binary. The fix is to represent source-built versus prebuilt code explicitly
    and require downloads for approved prebuilt code.
  3. MEDIUM: license strings are not validated as SPDX. Values such as "TBD" count as identified. The fix is to
    validate SPDX identifiers and expressions, and require the custom-license form for anything else.
  4. MEDIUM: vcs commit consistency. A commit embedded in vcs is not compared with revision. The fix is to
    reject a mismatch, with a test.

None of these affects what can ship today, because nothing is approved and the gate still blocks every artifact. Each
one weakens the gate for a future approval, so they should be fixed before merge or tracked as immediate follow-ups.

Known open issues and follow-ups

🤖 Generated with Claude Code

Murmur is about to integrate FluidAudio (october-dev#42) and sherpa-onnx (october-dev#26), and
their SDK licenses do not cover the model weights, tokenizers,
phonemizers, voices, and prebuilt binaries they fetch or link. october-dev#33 needs
a machine-readable, fail-closed audit record before any engine lands.

Add licensing/manifest.json, a single default-deny inventory of 71
entries:
- the FluidAudio v0.17.4 and sherpa-onnx v1.13.8 engines, each with one
  named build configuration
- their native closures: NemoTextProcessing plus the 37 Rust crates,
  std, and compiler-builtins it links statically, verified against the
  release binary; fastcluster; VBx; Japanese G2P code;
  kaldi-native-fbank, kaldi-decoder, kaldifst, OpenFst, Eigen,
  simple-sentencepiece, nlohmann/json; and the ONNX Runtime rebuild
- the model, tokenizer, phonemizer, and voice components october-dev#42 needs
- the two synthetic conformance-fixture sets.

Every non-fixture entry is pinned by a 40-hex revision and/or
byte-hashed downloads, cites license evidence in its own repository at
its own commit, records upstream hops, terms, and download-presentation
obligations, and is classified bundle, user-download, or blocked.
Nothing is legally approved.

Add tool/check_licensing.py, a stdlib-only checker, with 100 unittest
cases, wired into `make check` and the CI protocol job. It enforces:
- exact field shapes and an acyclic `requires` graph
- ambiguity forcing `blocked`, and custom licenses needing exact names
- Hugging Face downloads bound to the entry's repository and revision
- evidence that is a file in the reviewed repository at the reviewed
  commit
- approvals bound to a SHA-256 fingerprint of the reviewed fields
- notices for approved bundles, and byte hashes for approved downloads
  and content bundles
- synthetic-fixture provenance, and a case-insensitive scan that fails
  on any unregistered tracked model, native binary, or audio file.

Document the contract, policy, download rule, runtime boundary,
findings, update procedure, and questions for counsel in
licensing/README.md. Point THIRD_PARTY_NOTICES.md,
docs/voice-runtime.md, and CONTRIBUTING.md at it.

Key findings, all recorded as blocked:
- FluidAudio bundles an espeak-ng-harvested LuxTTS lexicon with no
  license.
- The Kokoro English G2P and lexicon sources are undocumented.
- The Parakeet EOU terms and the legacy diarization models are
  unresolved.
- sherpa-onnx compiles an unlicensed table and a StackOverflow snippet.
- The ONNX Runtime rebuild declares no license.

Refs october-dev#33

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@harshsaver harshsaver left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks. The inventory work is careful, and default-deny is the right policy while Murmur ships nothing. check-protocol, check-conformance and check-licensing (100/100) pass locally, and I spot-checked FluidAudio, text-processing-rs, Silero VAD Core ML, Kokoro ANE, Parakeet EOU and the diarization legacy models against their sources — all correct. A few things need to change before this can merge.

Blocking

  1. Adding fixtures breaks CI. check_licensing.py:568 requires every tracked file under conformance/fixtures/ to be listed in a fixture entry's paths (manifest.json:3746). Merging #56 into this branch makes check_licensing.py fail with "fixture file has no provenance entry" for all of #56's new fixtures, and #51 adds more. Hand-written synthetic JSONL carries no licensed content. Please scope the fixture rule to the scanned binary and audio suffixes (or register a directory prefix), so protocol PRs don't have to edit the licensing manifest.
  2. The description's own audit is FAIL, with four open findings (kind compatibility for tracked models, prebuilt code without byte hashes, non-SPDX license strings, vcs commit consistency). Please don't merge a gate that is known to be weak. Preferred: cut the approval path down to what's needed today. Nothing can be approved until legal review, so check_approval, the 17-field fingerprint (check_licensing.py:150, :467) and the approval rules can land with the first real approval (#42 or #16), designed against a real case. Today's checker then only needs to verify the schema, pins, blocked/pending everywhere, ambiguity forcing blocked, and no unregistered model or binary files in the tree. (Alternatively, fix all four findings with tests.)
  3. Data accuracy:
    • manifest.json:517 (parakeet-tdt-0.6b-v3-coreml): the evidence README at 7dd20fe says license: cc-by-4.0 in its front-matter but "License: Apache 2.0" in the body. Under this PR's own "ambiguous means blocked" rule, record the contradiction in review.notes, or block the entry until upstream fixes it.
    • licensing/README.md:311 calls piper-phonemize GPL-3.0, but its LICENSE.md at f3ff95a is MIT (the manifest correctly says MIT). It's effectively GPL only through linking espeak-ng. Please reword.

Non-blocking

  1. Size: about 3.8k of the 5.5k lines are manifest data, including 37 pinned Rust crates for a FluidAudio configuration marked "provisional until #42 confirms". Consider inventorying the transitive crate closure in #42, when the configuration is real, and keeping this PR to the direct engine, model, tokenizer, voice and phonemizer entries plus the known blockers.
  2. Once item 1 is fixed, please state explicitly in CONTRIBUTING.md that JSON conformance fixtures don't need a manifest change.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants