๐ŸŽ™๏ธ Live Transcript โ€” Help

โ† Back to the app

Quick start

  1. Pick ๆ—ฅๆœฌ่ชž / English / ไธญๆ–‡ on the Live tab.
  2. Press โ–ถ Start and allow the microphone.
  3. Talk โ€” recognized text appears in the box as you go.
  4. โธ Pause to hold, โ–ถ Resume to carry on. โน Stop when done โ€” nothing is saved automatically.
  5. Press ๐Ÿ’พ Save to keep the text (and audio, if ๐Ÿ”Š is ticked).

It transcribes; it does not translate the transcript. For a running translation turn on ๐Ÿ’ฌ Live translation under the transcript.

Why phone captions are slower than a computer

Speech recognition runs in the browser, not on our server โ€” and the browser hands the audio to whichever speech service its maker provides:

  • Desktop Chrome โ†’ Google's speech service. Continuous listening works well.
  • Desktop Edge โ†’ Microsoft Azure speech. Usually fine, but it has had spells lately where it returns no captions at all across whole regions. If Edge shows nothing, switch to Chrome (Google's engine) โ€” this is not a usage limit on you, and no setting needs changing, just the browser.
  • iPhone / iPad โ€” any browser (Safari, Chrome, Edgeโ€ฆ) โ†’ Apple's speech engine. iOS does not allow third-party speech engines, so every browser is the same here. It is weaker for long dictation, often stops after about a minute, and sometimes stops sending text without any error.
  • Android Chrome / Edge โ†’ Google's speech service (both are Chromium). Use Chrome โ€” its toolbar is at the top and doesn't cover the bottom navigation.

So on a phone, especially an iPhone, expect shorter useful sessions. This is a platform limit, not something a web page can fix.

Does speaking faster hurt accuracy?

Yes. Fast, run-on, or overlapping speech is harder for any engine. Speak in phrases with small pauses.

Same audio, recognized twice โ€” same result?

For a fixed audio file, roughly yes, and running it again does not improve accuracy โ€” the engine doesn't learn the clip. For live speech, two takes usually differ a little because the sound, noise, and levels change each time.

Captions freeze, lag, or stop

Phone speech engines get slower the longer one session runs, so the app quietly swaps in a fresh recognizer on a timer. You can tune this:

  1. Open Settings โ†’ ๐ŸŽ™๏ธ Caption refresh.
  2. Pick a shorter value (try 45โ€“60 seconds). It recovers sooner, at the cost of a sub-second gap on each swap.
  3. Or tick Auto-adjust this when captions stall: the app shortens the interval by itself after a 20-second silence, and lengthens it again after a few minutes of steady captions.
  4. โ†บ Auto resets it to the default (180s on phones, 50s on desktop).
If a session goes silent for 20s+ on a phone, a "Captions look stuck" hint appears. Press โน then โ–ถ to restart the recognizer.

On an iPhone the most reliable habit is to keep sessions short and press Stop then Start every few minutes.

Duplicated lines

On some Android phones the same sentence could appear twice because the engine restarts often. The app now drops a line that exactly repeats the one before it within 3 seconds.

Edge shows no captions at all

Edge sends audio to Microsoft's servers, which have had wide outages recently. Switch to Chrome โ€” it uses Google's engine, is usually faster, and needs no settings or code changes. This is not a usage limit on your account.

The in-progress line flickers or strobes while you speak fast

The last, not-yet-finished line is redrawn every time the speech engine revises its guess โ€” many times a second when you talk quickly. Open Settings โ†’ ๐ŸŽ™๏ธ Recognition โ†’ Caption smoothing and raise it: Extra (280ms) or Maximum (400ms) redraw that line less often, so it looks calm. Finished sentences still appear instantly. Takes effect the next time you press โ–ถ Start.

Desktop layout vs mobile layout

There are two layouts sharing the same engine:

  • Mobile layout (web.html) โ€” bottom tab bar, full-width transcript, tuned for touch. Opening the site on a phone lands here automatically.
  • Desktop layout (index.html) โ€” top tabs, side-by-side transcript/translation, drag-to-resize. Best on a computer.

Switch anytime from Settings โ†’ ๐Ÿ“ Layout. On a phone the mobile layout is recommended; the desktop layout is not designed for touch.

Microphones

Recognition accuracy and recording quality both depend on the sound reaching the microphone. The built-in phone mic is far from your mouth and picks up the room.

iPhone, best value first

  • Wired earbud mic โ€” USB-C EarPods (iPhone 15+) or Lightning EarPods (older). The capsule sits near your mouth. Cheapest and very effective.
  • Wireless lavalier with a plug-in receiver โ€” DJI Mic Mini (best value), DJI Mic 2 / Mic 3 (32-bit float + noise cancelling, good outdoors), RODE Wireless Micro. The receiver plugs into the USB-C / Lightning port and acts as a USB audio device.
  • Do not pair a Bluetooth headset or AirPods as the microphone โ€” iOS switches it to call quality (8โ€“16 kHz), which hurts recognition.

Search by model name on your local Amazon; exact links change, model names don't.

Transcribing a video call or computer audio

The app can only hear your device's current recording input (the microphone). It cannot hear the audio your computer or phone is playing โ€” the other person in a Meet / Teams call โ€” unless you route that audio into a recording device.

On a phone there is essentially no way to do this. Use the meeting app's own live captions instead โ€” Google Meet and Microsoft Teams both have them, with translation.

On a Windows PC โ€” Stereo Mix (no install)

  1. Right-click the volume icon in the taskbar โ†’ Sound settings.
  2. Scroll down โ†’ More sound settings (the classic panel).
  3. Recording tab โ†’ right-click empty space โ†’ Show Disabled Devices.
  4. Find Stereo Mix โ†’ right-click โ†’ Enable โ†’ right-click โ†’ Set as Default Device.
  5. Back on this page press โ–ถ Start; when the browser asks for a microphone choose Stereo Mix.
  6. When finished, set your normal microphone back as the default.

Japanese-language Windows: the menus are ใ‚ตใ‚ฆใƒณใƒ‰ใฎ่จญๅฎš โ†’ ใ‚ตใ‚ฆใƒณใƒ‰ใฎ่ฉณ็ดฐ่จญๅฎš โ†’ ้Œฒ้Ÿณ โ†’ ็„กๅŠนใชใƒ‡ใƒใ‚คใ‚นใฎ่กจ็คบ โ†’ ใ‚นใƒ†ใƒฌใ‚ช ใƒŸใ‚ญใ‚ตใƒผ โ†’ ๆœ‰ๅŠน โ†’ ๆ—ขๅฎšใฎใƒ‡ใƒใ‚คใ‚นใจใ—ใฆ่จญๅฎš.

No Stereo Mix โ€” VB-CABLE (one free tool)

  1. Install VB-CABLE from vb-audio.com (free) and restart.
  2. Sound settings โ†’ set Output to CABLE Input.
  3. Sound settings โ†’ set Input to CABLE Output.
  4. Press โ–ถ Start here and choose CABLE Output as the microphone.
  5. To also hear it yourself: CABLE Output properties โ†’ Listen to this device.
  6. Restore your normal Output / Input when finished.

This is a Windows sound-device setting; a web page has no permission to do it for you with a button.

Japanese comes out as Chinese (some Android phones)

Seen on Redmi / Xiaomi (MIUI) phones bought in China: choosing ๆ—ฅๆœฌ่ชž still produces Chinese text. The phone's speech engine has no Japanese pack, so the request for Japanese is ignored. A web page can't fix this โ€” add the Japanese voice pack in the phone's system settings.

Try this first

  1. Install Google and Speech Services by Google from the Play Store.
  2. Phone Settings โ†’ Additional settings โ†’ Languages & input โ†’ Voice input.
  3. Choose Speech Services by Google as the voice input service.
  4. Open its Languages โ†’ download ๆ—ฅๆœฌ่ชž (Japanese).
  5. Connect to Wi-Fi (the first Japanese recognition may need it).
  6. Open the site in Chrome, pick ๆ—ฅๆœฌ่ชž, press โ–ถ Start.

If that isn't enough

  1. Settings โ†’ Additional settings โ†’ Languages โ†’ add ๆ—ฅๆœฌ่ชž (it doesn't have to be first).
  2. Gboard โ†’ Voice typing โ†’ Languages โ†’ add Japanese.
  3. Restart the phone and try again.

Still Chinese? Try another phone or a computer to confirm it's this device's limit.

Recording, saving, History

  • โบ๏ธ Record audio (in the โ‹ฏ More panel on mobile) only decides whether the session's sound is also saved to a file. It has no effect on recognition speed or accuracy.
  • โน Stop never saves. Press ๐Ÿ’พ Save and tick ๐Ÿ“ Text and/or ๐Ÿ”Š Audio.
  • Saved recordings appear on the History tab โ€” play back, export subtitles (.srt / .vtt), copy, download, or delete.
  • Choosing a real folder to save into needs desktop Chrome / Edge. Elsewhere files download to your Downloads folder and History keeps a copy in the browser.

Translation & furigana

  • ๐Ÿ’ฌ Live translation (under the transcript) translates each finished line into one language. It runs in the background and never slows the transcript; a line can't be translated shows in red.
  • ใ‚ Furigana (โ‹ฏ More panel) shows hiragana readings above kanji when the recognition language is Japanese. First use downloads a ~17 MB dictionary once, then works offline.
  • Live translation needs the network. Recognition and captions do not.

Offline & privacy

  • Once the app has loaded once, it opens and transcribes even if our server is down โ€” recognition goes straight from your browser to Google / Microsoft / Apple, never through us.
  • What needs our server: the Live translation panel and the "check for updates" button.
  • Recordings and transcripts stay on your device. Nothing is uploaded to us.
  • The audio for recognition is sent by your browser to its speech provider (Google / Microsoft / Apple) โ€” that's how the Web Speech API works in every browser.