|
| 1 | +# CTranslate2 migration: implementation status |
| 2 | + |
| 3 | +**Date:** 2026-07-05 |
| 4 | +**Branch:** `feat/native-stt-whispercpp` |
| 5 | +**Worktree:** `stt-migration` |
| 6 | + |
| 7 | +Supersedes the earlier [whisper.cpp-based plan](./transcription-engine-migration.md) for the |
| 8 | +word-timestamp problem. The decision doc lives at |
| 9 | +[stt-ctranslate2-migration.md](./stt-ctranslate2-migration.md). |
| 10 | + |
| 11 | +--- |
| 12 | + |
| 13 | +## What works |
| 14 | + |
| 15 | +### TypeScript / Electron side (committed) |
| 16 | + |
| 17 | +| Module | Status | |
| 18 | +| --- | --- | |
| 19 | +| `electron/stt/ctranslate2Server.ts` | Replaces `whisperServer.ts`. Same wire contract (`POST /inference` → verbose JSON), same lifecycle (port allocation + `/` poll + multipart upload). VAD flags gone; word timestamps expected absolute. | |
| 20 | +| `electron/stt/wav.ts` | Extracted `writeSamplesAsWav` + `cleanupWav` to a shared file, engine-agnostic. | |
| 21 | +| `electron/stt/gpuDetector.ts` | Simplified to CUDA-or-CPU only. Metal/Vulkan probes deleted. New env var `OPENSCREEN_CT2_SERVER_EXE`. | |
| 22 | +| `electron/stt/modelManager.ts` | Downloads + unpacks a CTranslate2-format model directory (`.tar.gz` archive → SHA-256 verify → `node:tar` extract). Model URL points at `SYSTRAN/faster-whisper-small` on HuggingFace. | |
| 23 | +| `electron/stt/transcriptionContract.ts` | `SttBackend = "ctranslate2-cuda" | "ctranslate2-cpu"`. | |
| 24 | +| `electron/stt/index.ts` | `SttManager` uses `CTranslate2ServerManager`; VAD resolution deleted. | |
| 25 | +| `electron/stt/index.test.ts` | Updated mocks for the new server. | |
| 26 | +| `electron/stt/ctranslate2Server.test.ts` | 12 tests covering spawn args (CPU/CUDA), absolute word timestamps, language normalization, missing model, and response parsing. | |
| 27 | +| `electron/stt/{gpuDetector,modelManager}.test.ts` | Updated to new backend types. | |
| 28 | +| `electron/stt/vadModel.ts` | DELETED (behaviour inlined in the engine). | |
| 29 | +| `electron/stt/whisperServer.ts` + `.test.ts` | DELETED. | |
| 30 | + |
| 31 | +### Build + CI (committed) |
| 32 | + |
| 33 | +| File | Status | |
| 34 | +| --- | --- | |
| 35 | +| `scripts/build-ctranslate2-server.sh` | Build script for the native C++ helper (CUDA + CPU). | |
| 36 | +| `.github/workflows/build-ctranslate2-server.yml` | CI workflow for building + uploading the server binary. | |
| 37 | +| `scripts/build-whisper-binaries.sh` | DELETED. | |
| 38 | +| `.github/workflows/build-whisper-binaries.yml` | DELETED. | |
| 39 | +| `scripts/fetch-vad-model.{sh,ps1}` | DELETED (VAD was tied to whisper.cpp). | |
| 40 | +| `.github/workflows/build.yml` | VAD fetch steps removed across all platform jobs. | |
| 41 | +| `electron-builder.json5` | `extraResources` no longer ships `electron/native/models/`. | |
| 42 | +| `package.json` | `tar@^7.4.3` added; old `setup:vad:*` scripts removed. | |
| 43 | + |
| 44 | +### Native C++ server (committed, partially tested) |
| 45 | + |
| 46 | +The C++ source files are **real, compilable** code tested on an x64 Windows host: |
| 47 | + |
| 48 | +| File | Function | |
| 49 | +| --- | --- | |
| 50 | +| `electron/native/ctranslate2-server/CMakeLists.txt` | Pulls CTranslate2 (FetchContent, v4.4.0), cpp-httplib (v0.18.1), nlohmann/json (v3.11.3), and links them with ruy + OpenBLAS for CPU SGEMM. | |
| 51 | +| `electron/native/ctranslate2-server/include/wav.h` | WAV parser — handles RIFF header, ffmpeg's LIST/INFO metadata chunks, validates 16 kHz mono 16-bit PCM. | |
| 52 | +| `electron/native/ctranslate2-server/include/mel.h` | Log-mel spectrogram featurizer — STFT (KissFFT) + 80-band Slaney mel filterbank + log10/normalize. | |
| 53 | +| `electron/native/ctranslate2-server/src/mel.cpp` | Implementation of the above. | |
| 54 | +| `electron/native/ctranslate2-server/include/tokenizer.h` | Whisper tokenizer decoder — parses `tokenizer.json` (vocab + added_tokens), GPT-2 byte decoder for text rendering, special-token lookup for prompt construction, timestamp-aware `split_segments()` + `decode_tokens()`. | |
| 55 | +| `electron/native/ctranslate2-server/src/main.cpp` | HTTP server — `GET /` (200), `POST /inference` (multipart, WAV decode → mel features → CT2 encode → `generate()` → split into phrase segments → JSON). | |
| 56 | +| `third_party/kissfft/` | Vendored KissFFT (BSD-3-Clause) for the STFT. | |
| 57 | + |
| 58 | +### What was verified end-to-end |
| 59 | + |
| 60 | +On a Windows x64 host with Visual Studio 2022 + CMake + Ninja: |
| 61 | + |
| 62 | +1. **Configure** → FetchContent downloads CTranslate2 v4.4.0, including all submodules, plus cpp-httplib + nlohmann/json. |
| 63 | +2. **Build** → `ctranslate2-server.exe` (435 KB) + `ctranslate2.dll` (6.9 MB) + `libopenblas.dll` (OpenBLAS for CPU SGEMM). |
| 64 | +3. **Boot** → Server starts, loads the 462 MB model, logs `listening on 127.0.0.1:20199`. |
| 65 | +4. **Readiness** → `GET /` returns `200 ok`. |
| 66 | +5. **Inference** → `POST /inference` with a 7-second WAV returns a valid JSON transcription with segments and word-level timestamps. **Language auto-detection works correctly.** |
| 67 | + |
| 68 | +Example response from a 7s (5s silence + 2s tone) WAV: |
| 69 | + |
| 70 | +```json |
| 71 | +{ |
| 72 | + "detected_language": "<|en|>", |
| 73 | + "language": "<|en|>", |
| 74 | + "segments": [ |
| 75 | + {"end": 0.05, "id": 0, "start": 0.0, "text": " you"} |
| 76 | + ] |
| 77 | +} |
| 78 | +``` |
| 79 | + |
| 80 | +(Text is " you" because the audio is a beep tone — `faster-whisper` ran on a frequency sweep that Whisper interpreted as a short English utterance — expected.) |
| 81 | + |
| 82 | +**All TypeScript unit tests pass** (72 files, 601 tests). |
| 83 | + |
| 84 | +--- |
| 85 | + |
| 86 | +## What remains |
| 87 | + |
| 88 | +### 1. Ship the pre-built OpenBLAS binary for Windows (blocker for Mac/Linux CI) |
| 89 | + |
| 90 | +OpenBLAS is required by CTranslate2 for CPU SGEMM — without it the server throws |
| 91 | +`No SGEMM backend on CPU`. The manual `choco` install path works on a dev machine but |
| 92 | +needs to be part of the CI build workflow: |
| 93 | + |
| 94 | +- [ ] Add OpenBLAS to `scripts/build-ctranslate2-server.sh` (download + rename `libopenblas.lib` to `openblas.lib` + ship alongside the exe). |
| 95 | +- [ ] On macOS: replace OpenBLAS with `Apple Accelerate` (`-DWITH_ACCELERATE=ON`). |
| 96 | +- [ ] On Linux: replace OpenBLAS deps with `apt install libopenblas-dev`. |
| 97 | + |
| 98 | +### 2. .align() integration for word-level timestamps |
| 99 | + |
| 100 | +The current server produces only phrase-level segments (via Whisper's timestamps in the |
| 101 | +generated token sequence). For precise **word-level timestamps** CTranslate2 provides |
| 102 | +`WhisperReplica::align()` (DTW over cross-attention weights — the main reason for the |
| 103 | +migration). Not wired yet: |
| 104 | + |
| 105 | +- [ ] After `generate()`, call `encode()` to get the encoder output, then call `align()` |
| 106 | + with the start sequence + emitted text token IDs. |
| 107 | +- [ ] Use the returned `WhisperAlignmentResult` to split each phrase into words with |
| 108 | + correct absolute start/end times (the 5-second-leading-silence regression test). |
| 109 | +- [ ] Wire a `config.json` field for `alignment_heads` (needed by CT2's align; already |
| 110 | + present in the SYSTRAN model's `config.json`). |
| 111 | + |
| 112 | +### 3. Model download + conversion/hosting story |
| 113 | + |
| 114 | +- [ ] `modelManager.ts` uses a placeholder URL (`example.invalid`) — replace with the |
| 115 | + real SYSTRAN HuggingFace URL (`https://huggingface.co/SYSTRAN/faster-whisper-small/resolve/main/{model.bin,config.json,tokenizer.json,vocabulary.txt}`) |
| 116 | + or a single-archive tarball from a CI-managed release. Also pin the SHA-256. |
| 117 | +- [ ] Decide on INT8 quantization to reduce the 462 MB download (whisper-small |
| 118 | + `Systran/faster-whisper-small` is fp16 — system `ct2-transformers-converter |
| 119 | + --quantization int8` if desired). Currently out of scope; functional with fp16. |
| 120 | + |
| 121 | +### 4. .align() word timestamps ≠ wire JSON shape (renderer side) |
| 122 | + |
| 123 | +Once .align() lands, the JSON emitted by the server will need `words[]` arrays |
| 124 | +per segment (matching the `SttWordSegment` shape the renderer expects). Currently |
| 125 | +only phrase-level `segments[]` with `id`, `start`, `end`, `text` are output. |
| 126 | + |
| 127 | +### 5. Chunking for recordings > 30 seconds |
| 128 | + |
| 129 | +The current code pads/trims to exactly 30 seconds of features. For longer recordings |
| 130 | +we need the explicit chunk → parallel decode → merge pipeline described in |
| 131 | +`stt-ctranslate2-migration.md` § "Long-recording handling". |
| 132 | + |
| 133 | +### 6. Tests for the C++ server |
| 134 | + |
| 135 | +- [ ] Unit tests for the WAV reader. |
| 136 | +- [ ] Unit tests for the mel filterbank (compare against Python FasterWhisper output). |
| 137 | +- [ ] Integration test: HTTP POST a known WAV, assert segment count + word boundaries. |
| 138 | +- [ ] CI: build the server and run the integration test on a matrix (macOS ARM, Ubuntu, |
| 139 | + Windows x64). |
| 140 | + |
| 141 | +### 7. Clean up debug logging |
| 142 | + |
| 143 | +`main.cpp` has several `std::cerr << "[ct2] ..."` diagnostic lines that should be |
| 144 | +removed or gated behind a verbose flag before merging. |
| 145 | + |
| 146 | +--- |
| 147 | + |
| 148 | +## Build instructions (dev) |
| 149 | + |
| 150 | +```bash |
| 151 | +# Pre-requisites (Windows) |
| 152 | +choco install cmake ninja |
| 153 | +# OR download OpenBLAS from |
| 154 | +# https://github.com/OpenMathLib/OpenBLAS/releases/tag/v0.3.33 |
| 155 | +# extracts under .cache/openblas/ |
| 156 | + |
| 157 | +# Build |
| 158 | +powershell -ExecutionPolicy Bypass -File scripts/configure-ct2-build.ps1 |
| 159 | + |
| 160 | +# Run |
| 161 | +set OPENSCREEN_CT2_MODEL_DIR=<path-to-whisper-small-ct2> |
| 162 | +.cache/ctranslate2-build/ctranslate2-server.exe --port 20199 --threads 4 |
| 163 | + |
| 164 | +# Test |
| 165 | +curl -X POST -F "file=@test.wav" -F "language=auto" http://127.0.0.1:20199/inference |
| 166 | +``` |
0 commit comments