Skip to content

feat(protocol): add engine, TTS, wake, speaker, and batch contracts - #56

Open
Pushkraj-Space wants to merge 1 commit into
october-dev:mainfrom
Pushkraj-Space:issue-40-protocol-add-desktop-engine-tts-wake-spe
Open

Pushkraj-Space wants to merge 1 commit into
october-dev:mainfrom
Pushkraj-Space:issue-40-protocol-add-desktop-engine-tts-wake-spe

Conversation

@Pushkraj-Space

Copy link
Copy Markdown
Contributor

Summary

Extends murmur.v1 to protocol 1.1 with the engine, voice-pack, speech-output (TTS), microphone, provider-lifecycle, input-gate (half duplex), wake-phrase, speaker-verification, and batch-transcription contracts that a production desktop voice runtime needs.

The change is wire behavior only, and every existing v1 field, enum value, and meaning is unchanged. Engine implementations, SDK runtime APIs, default models, and product settings are out of scope.

Closes #40.

Design

Scopes. Engine readiness, voice-pack preparation, microphone permission, and speech output can exist outside any capture session, so they cannot honestly live in RuntimeEvent/SessionControl, which require a non-empty session_id. One new envelope pair carries them. It mirrors the session pair field for field: protocol, routing id, ordering key, and time.

Envelope Routing key Ordering key Carries
EngineEvent engine_id sequence engine_status, voice_pack_status, microphone_status, speech_output, error
EngineControl engine_id request_sequence voice_pack, speak, cancel_speech
RuntimeEvent (extended) session_id sequence + provider_status, input_gate_status, wake_phrase, batch_progress
SessionControl (extended) session_id request_sequence + speaker_verification, start_batch

Capability gating. Capabilities are gated by open string tokens rather than simulated:

  • EngineDescriptor.capabilities uses documented tokens: streaming_transcription, graceful_finalization, batch_transcription, word_timings, speech_output, echo_cancellation, wake_phrase, speaker_verification, and voice_packs.
  • Unknown tokens are ignored for negotiation and round-trip unchanged.
  • A session that names an engine_id is gated by that engine's capabilities.
  • A session without one keeps v1 behavior unchanged.

Rejection vs outcome.

  • A refused command yields only a MurmurError on the stream's generic error arm. It carries metadata["request_sequence"] and, for capability_unsupported, metadata["capability"].
  • Every accepted operation reaches exactly one terminal state:
    • provider sessions (streaming or batch): FINALIZED, CANCELLED, or FAILED
    • speech output: COMPLETED, CANCELLED, or FAILED
  • FAILED requires an error, and other states forbid one. The one exception is an engine that is NOT_READY, which may carry an error as its reason.
  • A new error-code registry lists the stable MurmurError.code values.

Half duplex and echo clearance.

  • InputGateStatus.closed_by is the set of active causes: HOST, FINALIZATION, SPEECH_OUTPUT, ECHO_CLEARANCE, and INTERRUPTION. An empty set means open.
  • Speech output only adds and removes its own causes, so a TTS tail never reopens input that another cause holds.
  • Every terminal speech state that follows playback passes through ECHO_CLEARING.

Speaker rejection cannot be mistaken for acceptance.

  • A speaker-rejected final uses the existing TRANSCRIPT_KIND_REJECTED, with speaker_verification set to REJECTED or UNCERTAIN (fail-closed).
  • A per-message kind/result table forbids a rejection result on FINAL.
  • Bypass covers one utterance and is legal only while verification is ENFORCED.

Batch.

  • A batch job is a session: start_batch → AudioFrames → finalize → batch_progress and FINAL transcripts with words → provider_status FINALIZED.
  • stop cancels.
  • Offsets are uint32 milliseconds on the submitted-media timeline. Sample 0 is the first submitted sample, and frames are contiguous by sample count.
  • Each word covers [start, end).
  • frame_duration_ms is required for Opus.

Excluded from the protocol: product commands, UI state, entitlement or trial fields, analytics, model or provider choice, product settings (including wake-phrase configuration), credentials, device names, and paths or URLs.

Implementation

  • Protocol:
    • New engine.proto, voice_pack.proto, and speech_output.proto.
    • session.proto adds StartSession.engine_id = 4, SetSpeakerVerification (arm 14), and StartBatchTranscription (arm 15).
    • events.proto adds arms 18–21, Transcript.words = 5, Transcript.speaker_verification = 6, and a doc comment on TRANSCRIPT_KIND_REJECTED (not deprecated).
    • Every new message and field is documented.
  • Docs: spec/README.md covers the version, scope tables, capability registry, rejection and outcome rules, error codes, voice packs, the input gate, microphones, speaker verification, wake phrases, batch, evolution rules, and exclusions. conformance/README.md gains the new messages, reasons, and the full validation table, including the verbatim kind/result table.
  • Conformance:
    • The manifest is now {1, 1}. Existing accepted fixtures changed only their minor version, and the forward-compatible sets moved to minor 2.
    • There are 16 new accepted and 22 new invalid sets.
    • Six new reasons: missing-engine-command, ambiguous-engine-command, invalid-identifier, invalid-progress, invalid-word-timing, and invalid-outcome.
  • Checker (tool/check_conformance.py):
    • Learns the engine envelopes, with an empty or missing engineId diagnosed as invalid-protocol-version just like sessionId.
    • Adds the per-payload rules and a one-routing-id-per-set check.
    • Extends the synthetic-text check to speak.text and words[].text.
    • Enum checks no longer crash on list or object values.
  • SDKs: each implements the same validation profile, learns EngineEvent/EngineControl, and bumps its current protocol to 1.1.
    • Dart:
      • Typed SetSpeakerVerification, StartBatchTranscription, StartSession.engineId, and EngineControl with VoicePackRequest, SpeakRequest, and CancelSpeech.
      • An EngineEvent model.
      • An internal validation.dart.
      • A CHANGELOG entry.
    • TypeScript: parseEngineEvent/engineEventToJson and parseEngineControl/engineControlToJson.
    • Python: EngineEvent, EngineControl, and their kind enums.
    • Rust: EngineEvent/EnginePayload and EngineControl/EngineCommand.
  • No generated code and no new dependencies.

Notes for reviewers

  • Dart consumers with exhaustive switches over SessionCommand or RuntimePayloadKind must handle the new cases, as the CHANGELOG notes.
  • engineStatus forbids an error on PREPARING/READY, allows one on NOT_READY, and requires one on FAILED.
  • Validation is per message only. Lifecycle order, cross-stream timing, batch-only words, and the request_sequence key inside the opaque MurmurError body are documented and covered by accepted fixtures, not validated.
  • The engine-side capability rejection lives in its own engine-command-rejection set with a limited engine, because engine-events advertises every capability.

Tests

Check Result
Proto compilation of all files with imports (make check-protocol invocation) pass
make check-conformance pass: 53 sets, 246 lines, and every invalid line diagnosed with exactly its reason
make check-python pass
make check-typescript (tsc --noEmit + tests) pass
make check-dart (dart format --set-exit-if-changed, dart analyze, dart test, including runtime_test.dart after the 1.1 bump) pass
make check-rust pass
check-flutter and check-flutter-plugin (analyze + app tests) pass

Differential fuzzing. The SDK runners assert only that a line is rejected, not why, so the change was also fuzzed outside the repo:

  • 13,472 deletion, type, null, and boundary mutations of every accepted fixture were run against the checker and all four SDK parsers.
  • This found and fixed two gaps:
    • TypeScript and Dart accepted an explicit null for new optional fields.
    • The checker crashed on non-string enum values.
  • The remaining divergences come from existing code or documented SDK behaviour:
    • Dart accepts non-negative JSON-number uint64 values.
    • TypeScript and Dart accept null for existing VoiceSource lists.

Dart, Rust, and Flutter checks ran in containers locally. CI pins protoc 29.x, Python 3.13, and Flutter 3.35.x, and the change uses nothing newer.

Audit result

Verdict: PASS. The independent code audit found no blocking or non-blocking defects. It compared the full diff against the issue and plan:

  • scope split, capability gating, and terminal outcomes
  • batch time base, speaker-rejection representation, and field numbering
  • the 1.1 bump, evolution rules, and exclusions

It inspected every new fixture set, the manifest dispatch, exact checker diagnoses, ordering and routing, unknown-field handling, and four-SDK round-trips. It also ran 17,440 mutations against the checker and the Python SDK with zero verdict divergences, and git diff --check passed.

Acceptance criteria

  • New messages live under murmur.v1, with documented field semantics and evolution rules.
  • Optional behavior is gated by capability tokens rather than simulated.
  • Streaming and batch operations have explicit cancellation, completion, and typed failure outcomes.
  • TTS and microphone state can express half-duplex and echo-clearance behavior.
  • Speaker rejection cannot be confused with an accepted transcript.
  • Word timings have an unambiguous time base (submitted media) and unit (ms).
  • Product commands, entitlement fields, analytics, and UI state stay out of the protocol.
  • Valid and invalid ProtoJSON fixtures cover every new payload.
  • Every SDK passes the conformance checks.

🤖 Generated with Claude Code

Extend murmur.v1 to protocol 1.1 with the wire contracts a production
desktop voice runtime needs. This is an additive v1 change. SDK runtime
APIs and engine implementations remain out of scope.

Engine-scoped concepts can exist before, after, or across capture
sessions, so they cannot live in the session envelopes. A new
EngineEvent/EngineControl pair carries them. It mirrors
RuntimeEvent/SessionControl, but engine_id is the routing key and the
ordering scope. It covers:
- engine identity, locality, platform, readiness, and open string
  capability tokens (engine.proto)
- voice-pack snapshots with components, byte progress, and typed
  errors, plus prepare/retry/repair requests (voice_pack.proto)
- speech output requests, cancellation, and lifecycle
  (speech_output.proto)
- microphone permission, route, interruption, and device readiness
  per source

Session-scoped concepts extend the existing envelopes:
- StartSession.engine_id binds a session to an engine. Unbound
  sessions keep unchanged v1 behavior.
- SetSpeakerVerification sets the mode (enforce, disable, or
  per-utterance bypass). Transcript.speaker_verification reports the
  outcome. A speaker-rejected final reuses TRANSCRIPT_KIND_REJECTED, and
  a per-message kind/result table keeps a rejection off accepted
  finals.
- StartBatchTranscription runs a batch through the same session
  mechanics: AudioFrame input, finalize to end input, stop to cancel,
  BatchProgress, and Transcript.words. Word offsets are uint32
  milliseconds on the submitted-media timeline.
- ProviderStatus gives every provider session, streaming or batch,
  exactly one terminal state (FINALIZED, CANCELLED, or FAILED).
- InputGateStatus reports the effective input gate as a set of
  closing causes, so half-duplex speech output and its echo tail can
  never reopen input held by another cause.
- WakePhraseEvent reports wake detection, activation, and dismissal.

Command rejection has one representation. A MurmurError on the stream's
generic error arm carries metadata["request_sequence"]. It is distinct
from operation outcomes, where FAILED requires an error and other
states forbid one. spec/README.md documents the capability gating
table, error-code registry, half-duplex rules, batch time base,
evolution rules, and exclusions (product commands, entitlements,
analytics, UI state, settings, credentials, device names, paths).

Conformance moves to protocol {1, 1} with 38 new fixture sets (16
accepted, 22 invalid) and six new rejection reasons. The checker and
the Dart, TypeScript, Python, and Rust SDKs enforce one identical
single-message validation profile. Required discriminators reject
*_UNSPECIFIED. Optional enums treat an explicit *_UNSPECIFIED as
absent. Unknown capability tokens round-trip. Each SDK learns the
engine envelopes and bumps its current protocol constant. Dart adds
typed SetSpeakerVerification, StartBatchTranscription, and
EngineControl commands, and a new EngineEvent model.

The checker also stops crashing on non-string enum values.

Closes october-dev#40

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@harshsaver harshsaver left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Solid, well-specified change against #40: numbering only appends (new oneof arms 18–21 and 14–15), codes are snake_case, no new dependencies, and protoc, conformance (53 sets, 246 lines), Python, Dart and TypeScript all pass locally. I spot-checked the speaker kind/result table across the checker and all four SDKs and they match.

Blocking (you need to rebase for the #55 CHANGELOG conflict anyway, so please fold these in so the two layers land consistent):

  1. Graceful finalization: baseline or capability? spec/README.md:81 and engine.proto:43 add a graceful_finalization token that makes finalize optional per engine, but #55's provider contract (now merged) makes finalize baseline, and docs/voice-runtime.md invariant 3 ("Keep the tail") requires it. Suggest dropping the token and making finalize part of streaming_transcription.
  2. One error-code registry. spec/README.md:129 says "new contracts use these", but omits codes already in use: cancelled (runtime.dart), plus unsupported_format, unauthorized, rate_limited and provider_disconnected (#55's provider.dart:233). Please make this the single registry covering all stable codes.

Non-blocking
3. spec/README.md:206-207: "fail-open"/"fail-closed" decide where UNCERTAIN lands, but nothing exposes which policy an engine uses. Either say it's engine-defined and hosts must handle both, or drop the policy wording.
4. session.proto:53 / :103: engine_id lives on both StartSession and StartBatchTranscription, and format duplicates requested_format. Fine as is; one sentence saying a batch session never sends start would help readers.
5. spec/README.md:28: the 1.0 -> 1.1 bump rewrote every accepted fixture's minor, so no accepted fixture now proves a 1.1 reader accepts a minor-0 message. Keep one minor-0 accepted line, or stay on 1.0 until release.
6. Growable sets now use string tokens while VoiceSource.capabilities stays a closed enum. Pre-release, consider converging on one approach now.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Protocol] Add desktop engine, TTS, wake, speaker, and batch contracts

2 participants