Standalone macOS research build

Your voice.
Still yours.

Echo is a local-first system for private dictation and evidence-grounded meeting intelligence. Speech stays on your Mac.

Echo · local sessionPrivate
See how it works
4.74%generated WER11 generated speech fixtures
≈85 msaverage model timeproduction Parakeet route
0 / 2non-speech false positivessilence + low background noise
6.19%*AA public-subset WER634 released public samples
The product goal

One private voice layer.
Two focused workflows.

Echo combines fast, cross-app dictation with evidence-grounded meeting intelligence in a standalone Mac app—without turning private speech into a cloud dependency.

01
Private dictation

Speak naturally.
Text appears quickly.

Multilingual Parakeet runs through Core ML, then Echo applies local vocabulary and spoken-command rules before safe delivery.

  • Production route: Parakeet TDT 0.6B v3
  • Deterministic vocabulary—no weight fine-tuning claim
  • Hold-to-talk or press-to-toggle capture
02
Meeting intelligence

Grounded notes.
Traceable decisions.

Local transcripts are mapped into candidate facts, checked against evidence, and reduced into structured notes by Qwen3.5-4B MLX 4-bit.

  • Map → evidence validation → reduce
  • Five-domain 5/5 and corrected financial-call 9/9
  • Whisper large-v3-turbo benchmark and fallback
Production dictation architecture

Designed as a system,
not just a model.

Recognition quality matters. So do capture reliability, private correction, active-app safety, latency, and the evidence written after every test.

01

Capture

Function-key audio, permissions, VAD, silence handling

02

Recognize

Parakeet TDT 0.6B v3, converted for Core ML

03

Resolve

Private vocabulary and spoken commands, deterministically

04

Deliver

Insert or copy with focus and approval safety gates

Production

Parakeet TDT 0.6B v3

Multilingual, fast, local, and deployed through Core ML for the current dictation path.

Rejected challenger

Qwen3-ASR 0.6B INT8

13.19% generated WER, ≈2.94 s average, and severe technical-vocabulary weakness. It did not replace Parakeet.

Meeting fallback

Whisper large-v3-turbo

Retained as a meeting transcription benchmark and fallback—not presented as Echo’s production dictation route.

Measured, published, bounded

Benchmarks with
the footnotes attached.

Direct Echo comparisons share audio and scoring. Published model and cloud results remain context. A number never crosses protocols without a label.

Download the evidence package
Echo benchmark highlights showing WER, speed, and API fee contextOpen full resolution ↗
Common WER: Identical 100-clip public-audio track
Common WERIdentical 100-clip public-audio track
Difficulty: Clean and harder public-audio slices
DifficultyClean and harder public-audio slices
Processing time: Model-only and whole-run timing
Processing timeModel-only and whole-run timing
Short-form context: Published upstream model results
Short-form contextPublished upstream model results
Long-form context: Published meeting-transcription results
Long-form contextPublished meeting-transcription results
AA-WER context: Echo’s measured public-subset reconstruction
AA-WER contextEcho’s measured public-subset reconstruction
Meeting intelligence: Grounded-output contract results
Meeting intelligenceGrounded-output contract results
Evidence status: What is proven—and what remains open
Evidence statusWhat is proven—and what remains open
* About the 6.19% result

Echo ran all 634 released AA-WER v2 public samples: 3.25% on VoxPopuli-Cleaned-AA and 9.13% on Earnings22-Cleaned-AA. The private AgentTalk half and Artificial Analysis’ complete current normalizer are unavailable, so this is a public-subset reconstruction—not an official AA-WER score or rank.

Documentation

Using Echo,
from capture to evidence.

01

Getting started

Echo is currently a standalone macOS research build. Open the app and grant Microphone, Accessibility, and Input Monitoring permissions only when the corresponding workflow requires them.

Current release boundaryThe automatic-delivery acceptance evidence is incomplete. Treat this build as measured research software, not a finished public release.
02

Using private dictation

  1. Choose a capture mode. Hold the shortcut while speaking, or use press-to-toggle.
  2. Speak in the target app. Echo records locally and captures the destination safely.
  3. Release or stop. Parakeet transcribes through the local Core ML route.
  4. Review delivery. Echo inserts when the target is safe; otherwise it preserves the text on the clipboard.

Echo should never send an external message automatically. Recognition, rewriting, and delivery are separate gates.

03

Using meeting intelligence

  1. Start or import a recording. Audio is handled locally and source retention is minimized.
  2. Transcribe. Parakeet supplies the local transcript; Whisper remains a benchmark/fallback route.
  3. Map facts. Qwen3.5-4B proposes decisions, actions, owners, and open questions.
  4. Validate evidence. Unsupported claims are rejected before the final reduction.
≈3.06 GB MLX 4-bit model≈3–4 s warm processing5/5 · 9/9 measured contracts
04

Vocabulary and spoken commands

Private vocabulary is a deterministic spoken → written mapping applied after recognition. It can preserve product names, people, tools, file paths, and preferred casing without retraining or modifying model weights.

spokencore M LwrittenCore ML

Spoken commands such as “new paragraph,” “comma,” and correction cues are processed locally and are recorded as pipeline stages in benchmark evidence.

05

Privacy model

Local inferenceDictation and meeting intelligence execute on the Mac.
No API dependencyThe production path does not require sending speech to a transcription provider.
Private correctionVocabulary and command processing remain deterministic and device-local.
Evidence-awareClaims retain protocol, model, timing, and limitation boundaries.
06

What is proven—and what is not

Evidence gateStatusBoundary
Generated ASR + non-speechPassed4.74% · 0/2 false positives
Core ML dictation pathImplementedParakeet TDT 0.6B v3
Model-only latencyMeasured≈85 ms average
Deterministic vocabulary + commandsImplementedLocal rule-based processing
100-clip common trackMeasuredSampled public audio
AA-WER public subsetMeasured6.19% · not official AA-WER
Meeting intelligencePassed5/5 and corrected 9/9
Local-only production pathImplementedNo transcription API dependency
Automatic deliveryMissing0 successful current-build traces

The honest current claim: a fast, private, evidence-backed local voice pipeline—with automatic delivery still awaiting current-build release evidence.

Local by design

The future of voice can feel instant
without feeling invasive.

Echo is an evidence-first exploration of what private, useful, local voice software can become.