Live captions on your Intel NPU
vinoWhisper transcribes whatever your Linux laptop is playing, or your microphone, with whisper-small.en running on the NPU through OpenVINO GenAI. Captions go to your terminal’s scrollback or to a Wayland overlay, and holding Meta+H types what you say into the focused window. No cloud, no API key, no account. The same idea as Live Captions on Copilot+ PCs, for Linux.
$ curl -fsSL https://raw.githubusercontent.com/karanshukla/vinoWhisper/main/scripts/install.sh | bash
Read install.sh (opens in a new tab) first, or use pip install vinowhisper && vinowhisper-setup. Python 3.11–3.13, PipeWire or PulseAudio, systemd. MIT licensed.
The octopus has no bones at all, so any gap wider than its beak is a door. It tests the gap with one arm first, then pours the rest of its body through in a few seconds. Keepers who house them in tanks learn to weigh the lids down, and to count their animals every morning.
Your laptop has an NPU. It is probably idle.
Recent Intel laptop chips ship with an NPU, and on Linux almost nothing uses it. Captioning suits one: continuous, sensitive to lag, and small enough to fit. whisper-small.en runs on it through OpenVINO GenAI, so transcription costs no CPU, no GPU, no fan and no network.
Measured on one Wildcat Lake laptop, 2026-09-12. The same window takes 0.95s on the Intel iGPU and 2.30s on the CPU.
- 0.70s
- per 12s window on the NPU, 0.85s p90
- ~1.5s
- behind the audio, at best
- 0.204s
- to the first streamed token
- 0
- idle cost: the server exits itself
The transcript lives in your scrollback
A word is printed only once two overlapping windows agree on it (LocalAgreement-2), and it is never rewritten after that. So the transcript lives in your terminal’s own scrollback: it survives quitting, and search and selection still work. Words waiting on a second opinion sit on the hearing… line under it. --plain and --json are there for pipes.
[3:02]The octopus has no bones at all, so any gap wider than its beak is a door. It tests the gap with one arm first, then pours the rest of its body through in a few seconds.
[3:44]Most of the energy in a thunderstorm goes into lifting water, not into lightning. One storm cloud can hold as much water as a small lake, and all of it has to come down somewhere, usually within the hour.
Overlay and voice typing
vinowhisper-gui floats the captions above every window, fullscreen video included, with a tray icon and a global shortcut. It renders the same event stream as the terminal, from vinowhisper-caption --json.
- Hold Meta+H, talk, and let go to type into the focused window. One NPU decode per utterance.
- If typing fails, the text is left on the clipboard.
- A 6.3MB Rust binary that links only libc, checked against a pinned sha256 on install.
Standup moved to 10:00, same room.
Are you still coming?
Running ten minutes late, start without me and I’ll catch up on the notes.
The parts that aren’t Whisper
Words are never rewritten
A word is committed once two overlapping windows agree on it (LocalAgreement-2). After that it stays put.
Scale to zero
A systemd socket unit owns the port with nothing running. The server holds the 10–30s NPU load and exits after 30 idle minutes.
Verified downloads
The model export is hashed against pinned digests before anything loads it, and the overlay binary against a sha256 in the wheel.
Loud fallbacks
NPU, then GPU, then CPU. A fallback shows in the journal, in /health, in the doctor and as a red border on the status bar.
Errors that say what to run
Nearly every error prints the fix in your distro’s own package names, read from /etc/os-release.
Replayable sessions
--record a session, then re-run the stitcher offline or sweep window sizes with vinowhisper-replay.
Install
The installer sets up uv (opens in a new tab), clones the repo and hands over to vinowhisper-setup, which picks your capture tool, NPU driver and model export and generates systemd units against paths that exist. It prints each command before running it. --dry-run prints the plan and changes nothing.
- Accelerator
- An Intel NPU. An Intel GPU or the CPU also works, with more lag. AMD NPUs are detected and not usable: OpenVINO has no plugin for them.
- Python
- 3.11 to 3.13. 3.14 cannot export the model yet.
- Audio
- PipeWire (
pw-record) or PulseAudio (parec), picked automatically. - Overlay
- wlr-layer-shell: KDE Plasma 6, Sway, Hyprland, niri, COSMIC. Not GNOME.
- Disk
- About 1.5GB for the model export.
$ vinowhisper-caption # caption system audio $ vinowhisper-caption --source mic # caption yourself $ vinowhisper-caption --list-targets # one app instead of the whole sink $ vinowhisper-setup --gui # the overlay and voice typing $ vinowhisper-doctor # devices, model, digests, audio $ vinowhisper-replay ~/sess --sweep 8,12,20
| Family | Includes | Package names |
|---|---|---|
| Fedora | RHEL, Alma, Rocky, Nobara, Bazzite, Silverblue | Built and run here |
| Debian | Ubuntu, Pop!_OS, Mint, elementary | From the package index |
| Arch | CachyOS, EndeavourOS, Manjaro, Garuda | From the package index |
| openSUSE | Tumbleweed, Leap, SLES | From the package index |
| Gentoo, Void, Alpine | Derivatives matched through ID_LIKE | From the package index |
| NixOS | hardware.intel-npu on unstable | Config snippets |
Known limits
- Every benchmark is n=1: one laptop, early-silicon NPU drivers. The GPU and CPU fallbacks have run nowhere else, and 2 of 15 GPU model loads segfaulted inside OpenVINO’s GPU plugin.
- Captions trail the audio by about twice the cycle time. That is the cost of a two-cycle commit, not a bug.
- The overlay is tested on KDE Plasma 6.7 only. On GNOME, use the terminal.
- Voice typing goes to whatever has focus, because Wayland does not say what that is. Outside KDE it types through the compositor’s virtual keyboard, which has not been tried on a live compositor yet.
- Only Fedora’s package names have been used for real, and the PulseAudio backend has never run against a real PulseAudio server.
- The NPU export needs
transformers<5.4. That version has two open CVEs, both of which need you to export a malicious model repo.
Run it on hardware I haven’t
If your laptop is not a Wildcat Lake, I want the report, working or not. An issue with vinowhisper-doctor --json pasted in is worth more than any benchmark I can run here.