Skip to content

feat(web): OpenAI-compatible STT endpoint for self-hosted transcription (whisper.cpp, faster-whisper) - #2397

Open
kovashikawa wants to merge 2 commits into
CapSoftware:mainfrom
kovashikawa:feat/openai-compatible-stt
Open

kovashikawa wants to merge 2 commits into
CapSoftware:mainfrom
kovashikawa:feat/openai-compatible-stt

Conversation

@kovashikawa

@kovashikawa kovashikawa commented Sep 30, 2026 •

Copy link
Copy Markdown

Summary

Adds an OpenAI-compatible transcription option so self-hosters can transcribe recorded videos without AssemblyAI, using whisper.cpp whisper-server, faster-whisper-server, speaches, or any endpoint that serves /audio/transcriptions with word timestamps.

STT_BASE_URL=http://whisper:9000/v1
STT_MODEL=large-v3-turbo
STT_API_KEY=            # optional

This is the STT half of #1877, rebuilt on top of the provider layer from #2113 (which already covers the LLM side via AI_PROVIDER=openai-compatible).

Behavior

  • STT_BASE_URL set: recorded-video transcription and editable-transcript backfills use it. Otherwise AssemblyAI is used exactly as today.
  • Live transcription during recording stays AssemblyAI-only.
  • The request is response_format=verbose_json with timestamp_granularities[]=word. The org's AI language is passed as language unless it is auto-detect.
  • Word timings go through the existing createEditTranscript pipeline, so a single pass produces both the caption VTT and the editable word transcript, the same as the AssemblyAI path.
  • whisper.cpp reports BPE tokens as "words" (" timest" + "amps", " Hello" + ",") and nests them under segments. These are folded back into words. For scripts written without spaces (CJK), only punctuation is folded. Its full language name ("english") is mapped to the code via language_probabilities.
  • An empty transcript is treated as no spoken audio (NO_AUDIO), matching the AssemblyAI no-speech handling.
  • The transcription gates in get-status, the share page, lib/transcribe, and finalize-desktop-recording now check for either provider instead of ASSEMBLY_API_KEY alone.

Testing

  • __tests__/unit/stt.test.ts: provider selection, the parser against a response captured from a real whisper-server (large-v3-turbo), OpenAI-shaped responses, CJK, missing word timestamps, and request shape and errors.
  • __tests__/integration/transcribe-workflow.test.ts: full workflow through the STT endpoint (VTT plus encrypted edit transcript, AssemblyAI not called), and silent audio leading to NO_AUDIO.
  • Existing transcription suites still pass. bun run typecheck is clean.
  • Ran it end to end on a self-hosted stack (Docker web plus media server, native whisper-server on Apple Silicon with Metal). A 6m44s recording produced the VTT, transcription.edit.v3.json, and an AI summary. STT took 13.7s, versus about 140s with faster-whisper on CPU in Docker, with 99.3% word agreement between the two transcripts.

RetriggerConfidence Score: 3/5

The PR is not ready to merge because supported word responses can lose boundaries, and missing duration can yield empty artifacts marked complete.

Findings

  1. P1 Distinct words merge together ▶
  2. P1 Missing duration discards speech ▶
  3. P2 Zero-duration speech disappears ▶
Fix with agent prompt
### Issue 1
apps/web/lib/stt.ts:71-78
If an OpenAI-compatible provider returns segment words without leading spaces, such as `"Hello"` and `"world"` for the segment text `"Hello world"`, this code treats `"world"` as part of the first word. The saved captions and editable transcript then contain `"Helloworld"` with one combined timing instead of two timed words.

### Issue 2
apps/web/workflows/transcribe.ts:854-857
If the STT response omits `duration` and the video has no positive stored duration, this passes zero to `createEditTranscript`. That function clamps every timed word to zero and drops it. The workflow can then mark transcription complete while saving an empty caption VTT and editable transcript despite receiving spoken words.

### Issue 3
apps/web/lib/stt.ts:76-78
The captured whisper.cpp response gives `"Um"` and its comma identical start and end times. Merging them still leaves a zero-duration word, which `createEditTranscript` discards. As a result, that spoken word is absent from both saved captions and the editable transcript.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

This PR adds an OpenAI-compatible provider for recorded-video transcription and editable-transcript backfills, with configuration gates, self-hosting instructions, and tests.

  • The parser also folds whisper.cpp tokens into timed words for the existing caption and edit-transcript pipeline.
  • The review found word-boundary corruption for another supported response shape, an empty-artifact path when duration is unavailable, and a dropped word in the captured whisper.cpp response.

Reviews (1) · Last reviewed commit: "feat(web): OpenAI-compatible STT endpoin..."

…ption

Adds STT_BASE_URL / STT_MODEL / STT_API_KEY so self-hosters can transcribe
with whisper.cpp, faster-whisper-server, speaches, or any endpoint that
serves /audio/transcriptions with word timestamps. When set it takes
precedence over AssemblyAI for recorded videos and editable-transcript
backfills; live transcription stays AssemblyAI-only.

Word timings feed the same createEditTranscript pipeline, so captions and
the editable transcript are produced from one pass like the AssemblyAI path.
whisper.cpp BPE tokens are folded back into words.
Comment thread apps/web/lib/stt.ts Outdated
Comment on lines +71 to +78
const continuesPrevious =
previous !== undefined &&
!/^\s/.test(word.text) &&
(spaceDelimited || PUNCTUATION_ONLY.test(word.text));

if (previous && continuesPrevious) {
previous.text += word.text;
previous.end = Math.max(previous.end, word.end);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Distinct words merge together

If an OpenAI-compatible provider returns segment words without leading spaces, such as "Hello" and "world" for the segment text "Hello world", this code treats "world" as part of the first word. The saved captions and editable transcript then contain "Helloworld" with one combined timing instead of two timed words.

Prompt To Fix With AI
This is a comment left during a code review.
Path: apps/web/lib/stt.ts
Line: 71-78

Comment:
**Distinct words merge together**

If an OpenAI-compatible provider returns segment words without leading spaces, such as `"Hello"` and `"world"` for the segment text `"Hello world"`, this code treats `"world"` as part of the first word. The saved captions and editable transcript then contain `"Helloworld"` with one combined timing instead of two timed words.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Comment on lines +854 to +857
// original media timeline) and the caption VTT derived from those words.
const durationMs =
videoDurationMs > 0
? videoDurationMs
: (transcript.audio_duration ?? 0) * 1000;
const editTranscript = createEditTranscript(transcript, durationMs);
videoDurationMs > 0 ? videoDurationMs : audioDurationSeconds * 1000;
const editTranscript = createEditTranscript(result, durationMs);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Missing duration discards speech

If the STT response omits duration and the video has no positive stored duration, this passes zero to createEditTranscript. That function clamps every timed word to zero and drops it. The workflow can then mark transcription complete while saving an empty caption VTT and editable transcript despite receiving spoken words.

Prompt To Fix With AI
This is a comment left during a code review.
Path: apps/web/workflows/transcribe.ts
Line: 854-857

Comment:
**Missing duration discards speech**

If the STT response omits `duration` and the video has no positive stored duration, this passes zero to `createEditTranscript`. That function clamps every timed word to zero and drops it. The workflow can then mark transcription complete while saving an empty caption VTT and editable transcript despite receiving spoken words.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Comment thread apps/web/lib/stt.ts Outdated
Comment on lines +76 to +78
if (previous && continuesPrevious) {
previous.text += word.text;
previous.end = Math.max(previous.end, word.end);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Zero-duration speech disappears

The captured whisper.cpp response gives "Um" and its comma identical start and end times. Merging them still leaves a zero-duration word, which createEditTranscript discards. As a result, that spoken word is absent from both saved captions and the editable transcript.

Prompt To Fix With AI
This is a comment left during a code review.
Path: apps/web/lib/stt.ts
Line: 76-78

Comment:
**Zero-duration speech disappears**

The captured whisper.cpp response gives `"Um"` and its comma identical start and end times. Merging them still leaves a zero-duration word, which `createEditTranscript` discards. As a result, that spoken word is absent from both saved captions and the editable transcript.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

- Fold whisper.cpp BPE tokens per token instead of per segment, so one-word
  segments (" timest" + "amps") stay whole and providers that don't mark word
  starts keep their words separate.
- Stretch zero-length words (whisper.cpp fillers like "Um") toward the next
  word so the edit transcript keeps them for filler removal.
- Fall back to the last word end when the response has no duration, so words
  are not clamped away on videos without a stored duration.
- Throw FatalError for 4xx responses, non-JSON bodies, and missing word
  timestamps so deterministic misconfiguration is not retried; name the
  endpoint on network errors and refuse redirects.
- Accept a blank STT_BASE_URL instead of failing env validation.
- Map full language names ("english") to codes.
- Docs: --vad for whisper-server, host.docker.internal on Apple Silicon,
  STT_MODEL default, no AssemblyAI fallback.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant