ByteScope

Live Captions

Captions for what you say — or for what your computer is playing: subtitle your own talk as you speak, or float live captions and a translation over a stream in a language you don't speak.

Recognition audio does leave your browser — it is transcribed by the browser vendor's servers.

Where your audio actually goes

While captions are running the audio being captioned — your microphone, or the tab or screen you shared — is streamed to Google's (or, in Edge, Microsoft's) speech service, and the text comes back over the network. Nothing this page does can pull that audio back once it has left. Treat a confidential meeting accordingly.

Translation is on-device either way: it runs on a model downloaded to your own machine, so caption text is never sent anywhere to be translated. ByteScope itself stores nothing — captions live in this tab's memory and are gone when you close it.

Caption window

Sample loop

Waiting for the first words…

This is a canned sample, not your microphone. Set the size and display mode until it reads well from the back of the room, then start.

Your browser will ask for the microphone. Stopping releases it again — the recording dot goes out.

The floating window needs its own click: browsers only open one straight from a button press. It floats over full-screen slides, and you can drag it or pull its edges to resize.

Caption setup

What to caption

Caption a person talking into a microphone, or caption the sound this computer is playing — a video, a livestream, a call in another tab. The second one opens your browser's share picker when you press start.

The captions listen to this input. Picking the wrong one raises no error at all — the captions still appear, just of the wrong microphone — so check it before you go on. You can also switch part-way through a talk.

Tick this if the input above is a virtual device carrying your system sound — BlackHole, Loopback, VB-Cable. It turns off echo cancellation, noise suppression and automatic gain control, which are tuned for a person in a room and quietly eat music, applause and anything that is not a voice.

Pick the language that will actually be spoken. Chrome cannot work it out while you talk, and the wrong pick does not raise an error — it invents confident nonsense in the language you chose. You can also switch part-way through a talk.

The translated line is the one your audience reads, so it is drawn larger and in the accent colour.

Where recognition runs

Chrome 139 and later can transcribe with a model on your own machine. Left off when the model still has to be downloaded, so a talk about to start is never held up by a fetch nobody asked for.

Checking whether this machine can recognise speech on its own…

Window size

Starting size of the floating window in pixels. You can still drag its edges afterwards, and the type scales with the window.

Show

Translation

Checking what this device can translate…

Browser support

Four separate browser features can be involved. Captions need the first one; the rest only make them better or more private.

FeatureChromeEdgeFirefoxSafari
Live captions (Web Speech API)Safari ships the prefixed API but ends a session after every utterance and ignores continuous mode, so a long talk stutters. Where the audio goes is the next row's question, not this one's.SupportedSupportedNot supportedPartial
On-device recognition (Chrome 139+)Transcribes with a model on your own machine, so no audio is uploaded. Which languages have a model is per machine and per platform, so the page asks at runtime rather than listing them. Edge is Chromium-based and very likely has it; that has not been verified here.SupportedPartialNot supportedNot supported
Captured playback audio (getDisplayMedia)Captions the sound of a tab, a window or the whole screen. Firefox and Safari show a share picker and then hand the page video only, which is the entire point of this row. Tab audio works wherever the row does; a native app's sound needs Chrome 141+ on macOS 14.2+, and the page measures that below rather than guessing from the browser.SupportedPartialNot supportedNot supported
On-device translation (Translator API)Chrome 138 and later. Detected at runtime — without it the page runs captions only.SupportedPartialNot supportedNot supported
Floating window (Document Picture-in-Picture)Chrome and Edge 116 and later. Without it, captions stay in the page.SupportedSupportedNot supportedNot supported

Captured audio on this machine

Checking what this machine can share…

This session

Recognition audio does leave your browser — it is transcribed by the browser vendor's servers.

Detected in this browser

  • SpeechRecognition checking…
  • SpeechRecognition.available checking…
  • getDisplayMedia checking…
  • Translator checking…
  • documentPictureInPicture checking…

About this tool

One pipeline, two jobs. Point it at a microphone and your talk is captioned as you speak; point it at the audio your computer is playing and the livestream, video or call you're watching gets live captions — the obvious case being a Japanese stream you'd like Chinese subtitles for. Recognition happens in the browser's built-in Web Speech API (SpeechRecognition) — no install, no account, no API key, no per-minute fee — and each finished sentence is translated by Chrome's built-in Translator API into one of 21 target languages, on your machine. For a talk, the mic picker matters on stage: a lapel or USB microphone is often not the system default, and choosing it explicitly beats captioning the laptop's built-in mic from across the room. Either way, one more click pops the caption bar out into a Document Picture-in-Picture window — a real always-on-top window, on both macOS and Windows, that floats over full-screen Google Slides, Keynote and PowerPoint when you're presenting, and just as happily over a full-screen video player when you're watching.

Captioning what your computer is playing

The played-audio route goes through the browser's own share picker: press start, choose a tab — or a window, or the whole screen — and tick the audio checkbox. Sharing a tab with “share tab audio” ticked is the good path: the page gets a clean digital copy of the tab's sound, with no room noise and no microphone in the loop, and the audio keeps playing out loud the whole time — captioning never mutes what you're watching. The single most common mistake is forgetting that checkbox; the page notices it has been handed a silent share and says so, but it can't tick the box for you. Tab audio is Chrome and Edge territory; capturing a native app's sound by sharing a window or the whole screen has worked on Windows and ChromeOS for years, but on a Mac it needs Chrome 141 and macOS 14.2 or later — macOS itself didn't let apps capture system audio before that — and on an older Mac, sharing a tab still works. Firefox and Safari can't hand a page audio from screen sharing at all. There is also a route that skips the picker entirely: a virtual audio device such as BlackHole or Loopback shows up as an ordinary entry in the microphone list, so system sound routed through it can be picked like any mic. Flip the “played-back audio” switch when you do — it turns off the browser's echo cancellation and noise suppression, processing tuned for a live voice that would otherwise fight the very audio you're trying to caption.

Where the audio actually goes

This site normally promises that nothing leaves your browser, and this page used to be the one exception. Since Chrome 139, it's a switch. The default route is still cloud recognition: the browser uploads your microphone audio to Google's servers (Chrome) or Microsoft's (Edge) — the same path the browsers' own dictation features use — and while that route is active, the upload is as real as it ever was. Captioning a stream raises the stakes on that route, because the audio being uploaded is then someone else's — the video you're watching, the far side of a call — not just your own voice. The new route is on-device recognition: the page asks SpeechRecognition whether a local model exists (processLocally), installs a language pack once, and from then on your speech is transcribed on your machine and the audio goes nowhere. Not every language has a local model, and availability varies by machine and platform, so the page checks at runtime and shows which route you're actually on rather than leaving you to guess. Translation was already local — Chrome 138+ downloads a model once, then every sentence is translated on-device — so with local recognition switched on, the whole pipeline runs on your laptop and the site's usual promise genuinely holds again. Cloud stays the default on purpose: minutes before a talk is no time to be stuck behind a model download. If the content is confidential — an internal roadmap, a client's numbers — or simply isn't yours to upload, flip the switch and confirm the route label before you start.

Why you pick the language yourself

There is no language auto-detection, and that is deliberate. A browser speech recogniser has to be told what language it is listening to before it starts; there is no reliable way to detect it in-browser, and letting the recogniser guess produces confident gibberish — the speech force-fitted into the wrong language's words rather than an admission of uncertainty. So the spoken language is a manual choice. For a talk, it is the one setting worth double-checking before you walk on stage; for a stream, remember it means the language being spoken in the video, not the one you want to read — a Japanese stream with the recogniser set to English yields fluent nonsense, not an error message. The target language is a separate pick from the 21 on offer, and the two can be swapped mid-session.

A window that outranks full screen

The floating captions use documentPictureInPicture — the same always-on-top machinery as a floating video, carrying a whole document instead. The operating system's compositor keeps that window above everything, which is exactly what a full-screen app cannot claim: Keynote, PowerPoint and Google Slides all take the screen over, and so does a video player — the caption bar stays on top of every one of them. That matters twice here: presenting, it keeps the translation over your slides; watching, it is what turns floating captions into actual subtitles over a full-screen stream. The requirements are stated where they bite: Chrome or Edge 116+ gets you captions and the floating window; the on-device translator needs Chrome 138+, on-device speech recognition Chrome 139+, and system-audio capture on a Mac needs Chrome 141+ on macOS 14.2+ (Windows and ChromeOS have no such gate) — and without them the tool degrades honestly, to cloud recognition, tab sharing or captions-only, rather than pretending. Safari and Firefox ship none of it — the page shows a support matrix and says so, instead of failing quietly.

Frequently asked questions

Why do I have to pick the speaking language myself — can't it auto-detect?

Because reliable in-browser auto-detection doesn't exist. SpeechRecognition must be told the language before it starts, and forcing it to guess produces gibberish captions — the recogniser will happily fit your speech to the wrong language rather than admit it doesn't know. A talk has one known speaking language anyway, so the honest design is a manual choice made once and checked before you start, not an automatic one that fails on stage.

How do I caption a video, livestream or call I'm watching?

Switch the audio source from microphone to the audio your computer is playing, and the browser opens its share picker. The good path is choosing the tab the video plays in and ticking “share tab audio”: the page gets a clean digital copy of the sound, with no room noise, and the audio keeps playing out loud — nothing gets muted. Set the spoken language to the language of the stream, pick your translation target, and pop out the floating window so the subtitles sit over the full-screen player. Two support lines to know: tab audio works in Chrome and Edge only, and capturing a native app's sound (system audio) works out of the box on Windows and ChromeOS but on a Mac needs Chrome 141 and macOS 14.2 or later — on an older Mac, playing the stream in a tab and sharing that tab still works. Firefox and Safari don't hand pages audio from screen sharing at all.

I shared a tab but no captions appear — what went wrong?

Almost always: the audio checkbox. The browser's share picker treats sound as opt-in, and with “share tab audio” unticked it hands the page a perfectly valid, perfectly silent stream — the single most common mistake with this source. The page detects that it is receiving silence and says so, rather than leaving you staring at an empty caption bar, but it cannot tick the box for you: stop the share, share again, and tick it this time. If there is no audio option to tick at all, you're on a browser or OS combination that can't deliver that audio — window and whole-screen shares carry sound in fewer configurations than tab shares, and on a Mac system audio needs Chrome 141 on macOS 14.2 or later.

Is my speech processed on my machine, or uploaded somewhere?

That is now a choice, and the page shows you which route is active. The default is cloud recognition, and there the audio really is uploaded: Chrome sends your microphone audio to Google's servers and Edge sends it to Microsoft's — while you're on that route, treat this page like any other cloud transcription service, because that is what it is. Captioning a stream makes that route a bigger deal, not a smaller one: the audio being uploaded is then someone else's — the video you're watching, the far side of a call — not just your own voice, and on-device recognition avoids the upload entirely. Since Chrome 139 there is a way out: switch to on-device recognition and the browser installs a language pack once, then transcribes entirely on your machine (processLocally), with nothing sent anywhere. Not every language has a local model and availability varies by machine, so the page checks at runtime and labels the active route instead of letting you assume. Translation runs locally either way. Cloud remains the default deliberately — a presenter shouldn't be stuck behind a model download minutes before walking on stage — so if confidentiality matters, flip the switch and check the label before you speak.

How big is the translation model download, and does it repeat?

On the order of a few tens of megabytes per language pair, downloaded by Chrome the first time that pair is requested and cached by the browser itself — later sessions, and any other site using the same pair, reuse it without downloading again. The page shows the download progress and runs captions-only until the model is ready. The on-device speech language pack works the same way — downloaded once, cached by Chrome — with one wrinkle: its download reports no progress, so the page can only say it's preparing. Conference Wi-Fi being what it is, open the page, pick your languages, and if you want the local route, let both installs finish before the talk starts.

How can a browser window stay above a full-screen slideshow?

Because it isn't an ordinary window. A Document Picture-in-Picture window is kept on top by the browser and the OS compositor, the same way a floating video stays visible over everything — full-screen Keynote, PowerPoint and Google Slides included, on both macOS and Windows. Drag it to where subtitles belong and resize it like any window. The microphone and the recognition live in the original tab, so the captions keep flowing while the presentation has focus.

Why doesn't this work in Safari or Firefox?

The fatal gap is speech recognition itself. Firefox does not implement the Web Speech API's SpeechRecognition at all, so there is nothing to make captions from — full stop. Safari ships a prefixed webkitSpeechRecognition, but it ends the session after each utterance instead of streaming continuously, so a talk becomes a stutter of restarts rather than a caption feed. On top of that, neither browser has the built-in Translator API or Document Picture-in-Picture — so even working captions could not be translated or float over a full-screen presentation. The played-audio source is just as closed: neither Firefox nor Safari will hand a page audio from screen sharing at all. Rather than half-run, this page shows a support matrix and names exactly what is missing. For an actual talk, use Chrome or Edge 116 or newer — Chrome 138+ if you want the translation, and 139+ if you want recognition to run on-device.

Can I show only the translation, without the original captions?

Yes — the caption bar has a display toggle: original only, translation only, or both stacked. Translation-only is the usual pick when the audience doesn't share your language, and it frees the whole bar for larger text. One caveat: on a browser without the Translator API (anything below Chrome 138), translation-only has nothing to show, so the tool falls back to the original captions rather than presenting a blank bar.

Can I caption a different microphone than the system default?

Yes — the page lists every microphone the browser can see and lets you pick which one feeds the captions. On stage that is not a detail: the mic that is actually on you is often a lapel mic or a USB interface that is not the system default, and captioning the laptop's built-in mic from three metres away produces exactly the mush you would expect. Pick the device explicitly, say a test sentence, and check the captions read correctly before the audience does. The same list is where a virtual audio device such as BlackHole or Loopback shows up if you route system sound through one — pick it like any mic, and flip the played-back audio switch so the browser stops running voice-tuned echo cancellation and noise suppression against the very audio you want captioned.