EdgeSpeak
A local-first desktop app that transcribes audio with word-level timestamps and semantic segments – built for creators and developers who value privacy and workflow integration.
Tool overview
EdgeSpeak is a macOS desktop app (optimized for Apple Silicon) focused on turning audio and video into accurate transcripts. Simply drag a file or record with your microphone, and you get results with semantic segmentation and precise word-level timestamps. Export options include JSON, SRT, and Markdown, making it easy to produce subtitles, summaries, or structured data.
In terms of performance, on an M4 chip it achieves around 40× real-time processing (1 minute of audio transcribed in ~1.5 seconds) with about 1 GB of memory usage. Since everything runs locally, no media files are uploaded to the cloud, addressing privacy concerns. The app also exposes an OpenAI Audio API-compatible interface, meaning it can serve as a local audio understanding service that replaces cloud APIs. For developers, it offers CLI and MCP Server support, allowing integration into automated workflows (e.g., with Claude or Codex).
The feature set is still growing: speaker diarization and text-to-speech (voice generation) are planned but not yet available. Pricing is a one-time purchase with a limited-time early bird price (exact amount not published), and a single license can activate up to 4 devices.