Voice Input and Read Aloud | LanMate User Guide
Using the Qianxin edition of LanMate?View Qianxin Edition User Guide

Voice Input and Read Aloud

LanMate supports three voice workflows: hold-to-talk (speech to text), voice conversation (continuous dialogue with playback), and read-aloud (text to speech). Local engines first — transcribed and synthesized text never leaves your machine.

Hold to Talk

The microphone button on the right of the input: hold to speak, release to transcribe:

Great for when your hands are busy, or when speaking is faster than typing.

Voice Conversation

The “Voice Conversation” button on the input enters a full-screen voice mode:

  1. Speak; a detected pause sends automatically
  2. Agent status shows while thinking
  3. Replies are spoken automatically
  4. Speak during playback to interrupt — no need to wait

The whole thing flows like a phone call. Live captions show everything you said and every reply on screen.

Read Aloud

The read-aloud button on a message speaks the text. Settings offer:

Engine Setup

Settings → Voice Transcription & Synthesis:

Engine Used for Notes
Transcription model Hold-to-talk, voice attachments Local Whisper one-click install (deps + model), or a FunASR-compatible endpoint
Real-time transcription Voice conversation Continuous recognition, low-latency; off disables the voice-conversation entry
Synthesis model Read aloud Local engine (works offline once set up), or a third-party service

Local Whisper install: when unconfigured, the settings page shows an “Install” button that installs dependencies into the shared environment and streams the model download (auto-falls-back to a mirror on failure), with live progress.

Lansenger Voice Messages

Voice messages received via Lansenger channels are auto-transcribed and enter context with the attachment. Send voice from the channel; the desktop handles it.