GPT-Live Meets Hermes: My Tools, My Choice of LLM
OpenAI already lets you talk to AI agents. What I wanted was a voice interface for my Hermes agents, with my tools and whichever LLM I choose. That could be an OpenAI model, another provider’s model, or a model running on my own hardware.
So I built live-hermes-bridge, a small Python application that connects OpenAI’s GPT-Live to Hermes. GPT-Live handles listening and speaking. Hermes keeps responsibility for reasoning, tools, memory, and skills. The model powering Hermes is an independent choice.
Why Live changed the architecture
OpenAI made GPT-Live 1 (gpt-live-1) generally available on 10 September 2026, according to its API changelog.
The earlier Realtime API combined speech, reasoning, and tool selection in one voice-model session. It supported custom tools, and an external agent could be connected through custom orchestration. Live adds a native delegation interface that separates the voice frontend from the backend doing the work. OpenAI’s migration guide explains that change.
GPT-Live can listen and speak at the same time, and keep the conversation going while a backend works. With client delegation, my application chooses that backend and its provider.
I used that interface to connect Hermes, keeping my existing tools, workflows, and choice of reasoning model.
How I connected them

From the browser, I select a Hermes backend and start a conversation. Audio travels directly between the browser and OpenAI over WebRTC. A FastAPI server creates the session and attaches a separate WebSocket for transcripts, delegation events, and replies.
When GPT-Live delegates, the bridge sends recent conversation context to the selected Hermes agent through its /v1/responses endpoint. Hermes runs its normal workflow and returns an answer for GPT-Live to say aloud.
One implementation detail mattered: the delegation event contains an identifier, but no task text. The bridge has to collect transcript updates itself and build the backend’s context. It then attaches the answer to the original delegation identifier. Agent requests run asynchronously, so receiving conversation updates can continue while Hermes works.
Backend configuration lives in YAML, credentials stay on the server, and one bridge can serve several agents. The browser lists the configured backends; each can point to a different Hermes instance with its own tools and underlying LLM.
What this makes possible
With the relevant tools configured in Hermes, I can ask things like:
- “Check why that service is failing and give me the short version.”
- “Find the notes from our last discussion about this project.”
- “Look through this repository and explain where authentication happens.”
Those requests go to the agent that already has access to my environment. Its permissions and approval policy still govern the work. I can change Hermes’s reasoning model without rebuilding the voice interface or moving my tools into another agent framework.
The voice layer remains OpenAI-hosted, even when Hermes uses a local LLM: audio and the returned agent summaries reach OpenAI’s cloud. The backend can use any LLM supported by Hermes, including OpenAI’s.
It is a prototype for a trusted network, released under AGPL-3.0-or-later. The repository includes the browser client, example configuration, and a startup script to try it locally.