Echo
A completely-local-compute, voice-to-voice AI assistant — no cloud, the whole STT → LLM → TTS loop running on your own GPU.
I think it'd be the most helpful to look at my old blog entries around Echo, to really get a sense of what inspired it, and what I was driving towards.
I have an old Discord I was using to dump all my model learning. I had a few of my friends brought in at different points.
Essentially, I knew I wanted a second me to talk to. I had been getting into a few things contemporaneously. Reading nonfiction led to me carrying loose leaf paper for notes. Eventually, recording the notes into Obsidian by hand got a bit, well, it's a nice process, but it's slow. I bought a Supernote so that my notes would be backed up easily, and I could carry all my books with me at once. Then I realized I basically wanted a pipeline: book snippets and my own commentary, and marginalia, automatically funneled into an Obsidian repo, and then funnel that into some kind of LLM compat DB, which I would then funnel search results into the LLM context to have conversations with my own thoughts.
At this point, I was also tracking Llamafile/cosmopolitan, Justine Tunney's truly universal compiler and the llama.cpp compilation using it. Llamafile got me thinking: I could run models local to my home. I spent a 36-hour weekend doing research on the pipeline:
- Wake word invocation (it would not be until my Echo Prismatic project that I'd get around to VAD and AEC) via openWakeWord. Trained an ONNX file for "echo" as the word using a Google GPU notebook.
- STT model (whisper.cpp, v3 large)
- LLM for local inference via llama.cpp (some q_4 Mistral 8B or so)
- TTS via Piper (best at the moment speech synthesis)
Each one of those, except Piper, sitting in VRAM, orchestrated by a Python program that did HTTP calls between them. Small optimizations like passing iostreams/bytes in memory instead of writing files helped. Tinkering with audio bitrates. There were lots of little things.
This worked fine; I had to kludge things to make it run on my AMD Radeon VII, which was blessedly one of the rare consumer-grade but totally compatible architectures to ROCm, AMD's native ML suite. Thankfully, someone made an AUR repo for ROCm (Arch btw), and so I was able to get it running. I spent the next few months making the system easier to bring up. The real kicker started being: Windows installs were a nightmare. Basically, it could work, but it was overall untenable.
When I intuited that you just feed the previous context back into the model, plus the new user input, that felt silly. Turns out that was basically how chat models work! It still feels silly.
At the end, I knew what semantic search was, about ModernBERT, vector DB, graphDBs, reranker models. I knew about quantization, and even offshoots like ik_llama.cpp (which btw had MoE compat before vanilla llama.cpp, not that I had the VRAM for DeepSeek). I had even toyed with getting Mistral to talk about no-go topics in my conversations with it.
I even have some delisted YouTube videos.
https://youtu.be/Qtin1-jJcfA
What kinda killed it: stopped being unemployed, and I kinda fell off it. It was an immense amount of legwork, and every inch was unknown territory to me. I would go fairly in-depth researching off-the-shelf FOSS components, because I wanted to focus full-local compute.
I know what all of these words mean, and can traverse Hugging Face a lil. I learned a lot. What I'm doing now builds on this in so many ways.
Echo is something I won't return to, except to glance. It is finished, in that way that most projects are finished: I got what I needed from it. It's a finished prototype for a local-only voice pipeline. It could respond in as few as 5 seconds, as many as 20 (never got around to streaming inference into voice synth).