Cache key
Every transcription is keyed by file hash + model size:- Same file, same model = instant cache hit
- Same file, different model = new transcription
- Modified file = different hash = new transcription
Storage location
Everything lives in~/.augent/memory/:
What gets cached
Each type of data is cached independently. You can diarize with different speaker counts without re-transcribing. You can run semantic search without re-computing embeddings on subsequent queries.
Model caching
Whisper models stay loaded in memory between tool calls. The MCP server is a long-lived process — once a model is loaded for the first transcription, subsequent transcriptions with the same model size are faster because there’s no model loading overhead. The sentence-transformer model (all-MiniLM-L6-v2, ~80MB) is also loaded once and kept in memory.
Translations
When you translate a non-English transcription, the English version is stored as a sibling(eng) markdown file alongside the original. Both appear in the Memory Explorer and Web UI.
Managing memory
Or use the Web UI to browse, search, and delete individual transcriptions.

