feat(web): OpenAI-compatible STT endpoint for self-hosted transcription (whisper.cpp, faster-whisper) - #2397
kovashikawa wants to merge 2 commits into
Conversation
…ption Adds STT_BASE_URL / STT_MODEL / STT_API_KEY so self-hosters can transcribe with whisper.cpp, faster-whisper-server, speaches, or any endpoint that serves /audio/transcriptions with word timestamps. When set it takes precedence over AssemblyAI for recorded videos and editable-transcript backfills; live transcription stays AssemblyAI-only. Word timings feed the same createEditTranscript pipeline, so captions and the editable transcript are produced from one pass like the AssemblyAI path. whisper.cpp BPE tokens are folded back into words.
| const continuesPrevious = | ||
| previous !== undefined && | ||
| !/^\s/.test(word.text) && | ||
| (spaceDelimited || PUNCTUATION_ONLY.test(word.text)); | ||
|
|
||
| if (previous && continuesPrevious) { | ||
| previous.text += word.text; | ||
| previous.end = Math.max(previous.end, word.end); |
There was a problem hiding this comment.
If an OpenAI-compatible provider returns segment words without leading spaces, such as "Hello" and "world" for the segment text "Hello world", this code treats "world" as part of the first word. The saved captions and editable transcript then contain "Helloworld" with one combined timing instead of two timed words.
Prompt To Fix With AI
This is a comment left during a code review.
Path: apps/web/lib/stt.ts
Line: 71-78
Comment:
**Distinct words merge together**
If an OpenAI-compatible provider returns segment words without leading spaces, such as `"Hello"` and `"world"` for the segment text `"Hello world"`, this code treats `"world"` as part of the first word. The saved captions and editable transcript then contain `"Helloworld"` with one combined timing instead of two timed words.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.| // original media timeline) and the caption VTT derived from those words. | ||
| const durationMs = | ||
| videoDurationMs > 0 | ||
| ? videoDurationMs | ||
| : (transcript.audio_duration ?? 0) * 1000; | ||
| const editTranscript = createEditTranscript(transcript, durationMs); | ||
| videoDurationMs > 0 ? videoDurationMs : audioDurationSeconds * 1000; | ||
| const editTranscript = createEditTranscript(result, durationMs); |
There was a problem hiding this comment.
Missing duration discards speech
If the STT response omits duration and the video has no positive stored duration, this passes zero to createEditTranscript. That function clamps every timed word to zero and drops it. The workflow can then mark transcription complete while saving an empty caption VTT and editable transcript despite receiving spoken words.
Prompt To Fix With AI
This is a comment left during a code review.
Path: apps/web/workflows/transcribe.ts
Line: 854-857
Comment:
**Missing duration discards speech**
If the STT response omits `duration` and the video has no positive stored duration, this passes zero to `createEditTranscript`. That function clamps every timed word to zero and drops it. The workflow can then mark transcription complete while saving an empty caption VTT and editable transcript despite receiving spoken words.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.| if (previous && continuesPrevious) { | ||
| previous.text += word.text; | ||
| previous.end = Math.max(previous.end, word.end); |
There was a problem hiding this comment.
Zero-duration speech disappears
The captured whisper.cpp response gives "Um" and its comma identical start and end times. Merging them still leaves a zero-duration word, which createEditTranscript discards. As a result, that spoken word is absent from both saved captions and the editable transcript.
Prompt To Fix With AI
This is a comment left during a code review.
Path: apps/web/lib/stt.ts
Line: 76-78
Comment:
**Zero-duration speech disappears**
The captured whisper.cpp response gives `"Um"` and its comma identical start and end times. Merging them still leaves a zero-duration word, which `createEditTranscript` discards. As a result, that spoken word is absent from both saved captions and the editable transcript.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.- Fold whisper.cpp BPE tokens per token instead of per segment, so one-word
segments (" timest" + "amps") stay whole and providers that don't mark word
starts keep their words separate.
- Stretch zero-length words (whisper.cpp fillers like "Um") toward the next
word so the edit transcript keeps them for filler removal.
- Fall back to the last word end when the response has no duration, so words
are not clamped away on videos without a stored duration.
- Throw FatalError for 4xx responses, non-JSON bodies, and missing word
timestamps so deterministic misconfiguration is not retried; name the
endpoint on network errors and refuse redirects.
- Accept a blank STT_BASE_URL instead of failing env validation.
- Map full language names ("english") to codes.
- Docs: --vad for whisper-server, host.docker.internal on Apple Silicon,
STT_MODEL default, no AssemblyAI fallback.
Summary
Adds an OpenAI-compatible transcription option so self-hosters can transcribe recorded videos without AssemblyAI, using whisper.cpp
whisper-server, faster-whisper-server, speaches, or any endpoint that serves/audio/transcriptionswith word timestamps.STT_BASE_URL=http://whisper:9000/v1 STT_MODEL=large-v3-turbo STT_API_KEY= # optionalThis is the STT half of #1877, rebuilt on top of the provider layer from #2113 (which already covers the LLM side via
AI_PROVIDER=openai-compatible).Behavior
STT_BASE_URLset: recorded-video transcription and editable-transcript backfills use it. Otherwise AssemblyAI is used exactly as today.response_format=verbose_jsonwithtimestamp_granularities[]=word. The org's AI language is passed aslanguageunless it is auto-detect.createEditTranscriptpipeline, so a single pass produces both the caption VTT and the editable word transcript, the same as the AssemblyAI path." timest"+"amps"," Hello"+",") and nests them under segments. These are folded back into words. For scripts written without spaces (CJK), only punctuation is folded. Its full language name ("english") is mapped to the code vialanguage_probabilities.NO_AUDIO), matching the AssemblyAI no-speech handling.get-status, the share page,lib/transcribe, andfinalize-desktop-recordingnow check for either provider instead ofASSEMBLY_API_KEYalone.Testing
__tests__/unit/stt.test.ts: provider selection, the parser against a response captured from a realwhisper-server(large-v3-turbo), OpenAI-shaped responses, CJK, missing word timestamps, and request shape and errors.__tests__/integration/transcribe-workflow.test.ts: full workflow through the STT endpoint (VTT plus encrypted edit transcript, AssemblyAI not called), and silent audio leading toNO_AUDIO.bun run typecheckis clean.whisper-serveron Apple Silicon with Metal). A 6m44s recording produced the VTT,transcription.edit.v3.json, and an AI summary. STT took 13.7s, versus about 140s with faster-whisper on CPU in Docker, with 99.3% word agreement between the two transcripts.The PR is not ready to merge because supported word responses can lose boundaries, and missing duration can yield empty artifacts marked complete.
Findings
Fix with agent prompt
Summary
This PR adds an OpenAI-compatible provider for recorded-video transcription and editable-transcript backfills, with configuration gates, self-hosting instructions, and tests.
Reviews (1) · Last reviewed commit: "feat(web): OpenAI-compatible STT endpoin..."