EN/Contact
Lydia / Works
AI Products2026

SayType Voice Input

Reliable voice-to-text for any macOS input field, shaped for the current context

Local-first Desktop Product Case / 02

SayType turns voice input into a reliable macOS system loop: start speaking in the original app, transcribe locally, shape the text for context, and paste it back where the user started.

This is neither meeting transcription nor a thin interface around Whisper. The real product problem is completing input across chat, email, documents, and AI tools without losing work when permissions, connectivity, models, or automatic paste fail.

01

Designed for macOS users who write primarily in Chinese or switch between Chinese and English; everyday input should not require repeatedly opening the main window.

02

The current build implements local recording, Whisper transcription, a global shortcut, a floating recorder, separate transcript and final output, automatic paste, and local history.

03

Speech recognition and text shaping are separate layers; if AI rewriting fails, the product still returns a locally corrected transcript.

04

The formal v1.0 remains in planning: provider routing, per-app modes, a full dictionary, onboarding, diagnostics export, signing, and notarization are not yet complete.

The Problem Is Not Accuracy Alone

Everyday input happens in different contexts: chat should be concise, email complete, documents structured, and AI instructions explicit. Users do not want to open a transcription tool and carry text back; they want to finish the thought at the current point of focus.

SayType therefore measures its job not as producing a transcript, but as completing an input. Shortcut, recording state, task preservation, text shaping, and paste-back must operate as one product chain.

Product Loop: Speak to Paste

A global shortcut starts recording from any input field and a floating bar shows state. Audio is preserved as a traceable job; local Whisper produces the raw transcript, deterministic corrections and mode-based shaping produce the final text, and the result is pasted back into the original app.

The main window is a configuration and recovery center, not a required daily stop. It surfaces microphone, Whisper, and shortcut readiness, local history, and failed-task recovery.

Separate Transcription from Rewriting

SayType always preserves the raw transcript and generates final text separately. Shaping may remove filler words, restore punctuation, and improve structure, but must not add facts the user did not say. Users can compare outputs, and the raw version remains a safe fallback.

The current build includes intelligent, daily, structured, spoken, verbatim, and custom modes. Work email, AI-prompt modes, and automatic per-app selection remain future capabilities rather than delivered claims.

Failure Must Still Complete the Job

Reliability is not one happy path but a degradation map. If LLM shaping fails, return locally corrected text; if paste fails, preserve the clipboard and guide manual paste; if permissions are missing, show a precise repair action; if local recognition fails, preserve the audio and task for recovery.

History is local by default, with no account, cloud history, or team workspace in the current scope. As more cloud providers are added, the product must keep disclosing what leaves the device and how users can return to a local path.

PROCESS

Key Decisions

01

Local transcription is the foundation

Users should still receive usable text when offline or when AI shaping fails. Cloud capability is an enhancement layer, not the only path to completing input.

02

Keep raw and final text side by side

This makes AI edits visible, comparable, and reversible, while separating recognition errors from over-aggressive rewriting during diagnosis.

03

Keep the main window out of the daily path

The main window owns setup, history, and recovery. Daily use stays in the original app through a shortcut, floating state, and paste-back, reducing context switching.

04

Delay voice commands until they are explicit and safe

Ordinary dictation must not accidentally trigger deletion, sending, or clipboard actions. Commands need an explicit mode, confirmation, and undo, so they remain outside the current foundation release.

Role & Collaboration

I led product and experience, reframing the objective around completing input and designing the journey, information architecture, transcription/shaping layers, mode system, degradation paths, privacy boundary, and formal release gates.

AI supported implementation, testing, documentation, and debugging. Product positioning, priorities, mode boundaries, safety tradeoffs, release judgment, and final experience acceptance remained my responsibility.

Validation, Outcome & Reflection

The private beta is live with the main loop working across recording, local transcription, shortcut activation, floating state, text shaping, failure fallback, paste-back, and local history.

This case does not present the private beta as a formal v1.0 release. Until clean-machine installation, signing and notarization, onboarding, diagnostics export, per-app rules, and a complete personal dictionary are accepted, the product remains in beta.

A voice-input product succeeds not when a transcript appears, but when correct final text reaches the place where the user intended to type. Permissions, focus, clipboard behavior, and recovery are therefore as important as model accuracy.

Next, the foundation should become installable, diagnosable, and recoverable before personalization expands through dictionaries, per-app modes, and multi-provider routing.

DEMO / 02

Product demo video

A 13-second public excerpt showing local readiness, the voice-input interface, and output modes. Private conversations and provider-configuration segments were removed from the supplied 720p recording.