(tab icon tooltip: "SRT caption editor"): a segmented control at the top switches between two independent tools that share one caption format:
- AUTO CAPTIONS (beta): transcribes speech into caption layers, offline.
- SRT IMPORTER: imports/edits/creates captions from an existing
.srt, or reads back captions built by Auto Captions.
10.1 Auto Captions
Runs OpenAI's Whisper models locally, through whisper.cpp. Your audio never leaves your computer: no account, no per-minute fee, works offline once the model is downloaded. (Everything here runs on your machine. No Python, no account, and nothing is uploaded.)
One-time setup
The first time you open Auto Captions (or whenever the engine, decoder, or every model is missing), you land on a setup screen instead of the transcribe screen:
- Speech engine: a DOWNLOAD button installs both the whisper.cpp engine and the shared audio decoder (FFmpeg) together, about 110 MB total. A CHECK FOR ENGINE UPDATE button tells you whether this version of Cepter EFX comes with a newer engine build than the one installed. It checks locally with no network call and never touches your models; when there is one, the setup button reads UPDATE and installs it.
- Model folder: shows where model files are kept, with BROWSE (to put them on a different drive: they're large) and DEFAULT (back to the folder inside the extension).
- Models: a scrollable list of 10 Whisper builds you can download individually (see the full table below), each with a live download progress bar (percentage, MB, elapsed time) and a CANCEL button. Deleting an installed model needs two clicks (the button becomes SURE? for a few seconds, then deletes).
- A BACK button appears once the engine, decoder and at least one model are all installed, so you can jump straight back to transcribing.
Whisper models are never silently "updated" once installed: a model is a fixed, permanent snapshot; improvements ship as a new named model (e.g. Large v3 → Large v3 Turbo), not a background update to one you already have.
Choosing what to transcribe
- DETECT SELECTED: reads the source file straight off whatever layer(s) you have selected in the timeline. Every qualifying selected layer becomes its own segment (sorted into timeline order and joined), so a chopped-up scene across several layers still transcribes as one continuous pass, with each piece's timing tracked back to its own position. Layers without a real source file (solids, precomps, text) are skipped automatically.
- BROWSE: pick any video/audio file from disk instead (whole file, starts at the beginning of the comp).
A hint line under these two buttons always states exactly what's currently loaded (which layers, how much audio, what portion of the timeline it covers).
Transcribe options
- Spoken language: 34 languages plus Auto detect (the default).
- Output: Original language (default, i.e. plain transcription) or Translate to English.
- Model: only shows models you've actually downloaded. Picking an English-only (
.en) model automatically locks language to English and output to "Original language" (both dropdowns grey out): those builds can't translate and can't transcribe anything else. - A rough time estimate appears once a model is picked, e.g. "1:24 of audio · roughly 0:12–0:37 on 7 threads": an estimate only, not a guarantee.
TRANSCRIBE
Renders the selected audio, then runs it through Whisper locally. After Effects is busy extracting the audio at the start of this, but nothing is polled afterward: the transcription itself runs fully in the background. A PROGRESS LOG panel shows live phase/progress, with COPY LOG (for sending a full diagnostic dump if something goes wrong) and CANCEL buttons. Large timeline selections are capped at 40 pieces in one pass: precompose first if you're over that.
The transcript editor
Once transcription finishes, you get an editable transcript with a live word count. Fix any misheard words here before building the layers: edits keep each word's original timing. (Editing works by re-splitting the text on whitespace and re-pairing it word-for-word against the original timings, so it's best for correcting individual words, not for casually adding/removing several words at once if precise timing matters.)
- Words per caption slider (1–20, default 4): the maximum words a single caption will ever show. It's a hard ceiling, not a target: a caption still breaks earlier than the limit at the end of a sentence or a natural pause, so one line never bleeds into the next.
- Max tail (seconds) slider (0.5–10.0, default 1.5): if there's silence right after a caption, how long it's allowed to keep showing before disappearing (never past where the next caption starts).
- GENERATE CAPTIONS builds bottom-centred text layers on the timeline, sized to the comp, coloured with the Captions tab's base colour (set in SRT Importer, see below). START OVER discards the transcript.
Generating captions leaves you on the Auto Captions pane: it does not jump you to SRT Importer automatically. To reword or recolour what was just built, switch to SRT Importer and press Sync From Timeline.
10.2 SRT Importer
An editable caption list you can build three ways: importing a .srt file, syncing captions Auto Captions already built, or just typing your own.
IMPORT .SRT FILE
Loads a standard .srt file's timings and text into the list below. Nothing touches your timeline until you press Create Captions.
Colouring captions
- Base colour swatch: sets the colour for every caption. Opens the shared colour picker.
- Highlight row: select (drag-highlight) text inside any caption below, then click a colour to recolour just that selection. Nine fixed one-click swatches (white, red, orange, yellow, green, cyan, blue, purple, pink) plus a custom-colour swatch (opens the full picker), a strip of recently-used custom highlight colours, and a reset swatch to clear a highlight back to the base colour.
The caption list
One row per caption: its time range, an editable text field you can type directly into, and a × to delete that row from the list (this doesn't touch anything already on the timeline: it only edits the in-panel list until you press Create Captions).
CREATE CAPTIONS
Builds bottom-centred AE text layers from the list: same placement/sizing rule as Auto Captions. Highlighted portions become their own colour via a text animator, so a single caption can show multiple colours at once.
SYNC FROM TIMELINE
Reads caption layers already on the timeline back into this editor, including anything you've since moved, retimed, retyped or deleted directly in After Effects. This reliably reads back layers Auto Captions built; captions created by SRT Importer's own Create Captions are not tagged for this same round-trip, so treat "sync" as primarily an Auto-Captions companion feature.
EXPORT
Saves the current caption list to a folder you choose, in any combination of five formats (tick as many as you want):
| Format | What it is |
|---|
.srt (default) | SubRip: the usual subtitle format |
|---|
.vtt | WebVTT: for web video |
|---|
.json | Timings and text as structured data |
|---|
.txt | Plain text, no timings |
|---|
.tsv | Tab-separated, for spreadsheets |
|---|
Files are written as captions.<ext> in the chosen folder.