The ChatGPT desktop app makes a surprisingly good frontend for your local LLMs

The weak link in local LLMs has never been the models — it’s the frontends. Ollama and vLLM run beautifully on homelab hardware, but the chat UIs range from serviceable to science project. Which is why an XDA piece this month caught my attention: the author has been running local models — Qwen, GLM-Flash — inside the ChatGPT desktop app, with OpenAI’s client none the wiser, using a free open-source proxy called opencodex. That’s a trick worth understanding even if you never run it.

A proxy, not a fork — and that’s the clever part

opencodex works because the ChatGPT desktop app’s coding side reads its API endpoint from a local config file. The tool edits ~/.codex/config.toml and points it at itself:

# Auto-injected by opencodex
openai_base_url = "http://127.0.0.1:10100/v1"

From then on, every request — listing models, sending a chat — hits the local proxy first. It aggregates your configured providers into one model list, so your local models appear in the official model picker alongside the real OpenAI ones, prefixed by backend (a vLLM-served model shows up as vllm/glm-5.3-flash, say). When you pick one and send a message, the proxy routes the request wherever it actually needs to go.

The reason this is implemented as a proxy rather than a patched app matters: app updates keep landing and the integration keeps working. It survives upgrades because it never touched the app binary, only the config it reads. Anyone who’s maintained a fork of anything will recognize how much grief that design decision saves.

Setup is genuinely short

It’s an npm install. Run ocx init for a guided setup that writes the config and offers to inject the base URL into the config file. Out of the box it auto-detects the three runtimes most homelab people already run on their default ports — Ollama, vLLM, and LM Studio — and anything speaking an OpenAI-compatible API or Anthropic’s Messages API can be added manually. There’s a web dashboard for wiring up providers and, usefully for anyone who bills or budgets by tokens, per-model usage and cost statistics.

The homelab version of this

The interesting move for us isn’t laptop-local models — it’s that the proxy doesn’t care where the provider lives. XDA routes to models on multiple machines over a private network, which maps exactly onto how a homelab wants to run: vLLM serving something big on the GPU box in the rack, a smaller model on the workstation, all surfaced as one model list in a polished desktop client on whatever machine you’re actually sitting at. The frontend problem solves itself, and your model-serving architecture stays exactly as distributed as it already was.

Two honest caveats

Before you point a paying client at this, know where the seams are.

Tool-calling is the real compatibility test. The app’s agent tooling uses its own calling conventions — freeform patch edits, shell access, namespaced tools — and models trained on different conventions can flounder badly. In the XDA testing, one model got completely lost while others handled it fine. The fix is empirical: test your model on the tasks you actually do before trusting it, and don’t conclude “local models are broken” when it’s really “this model wasn’t trained for this client’s dialect.”

Not every request stays local. This is the one that matters for the privacy-motivated. The built-in web search tool still executes through OpenAI’s servers under your ChatGPT login, and if your local model is text-only, image requests get described by an OpenAI model first, with the text handed back to your model. Plain chat and code stay on your wire; search and vision do not. If your rule is “nothing leaves the house,” either avoid those features or accept exactly which calls bypass the proxy — because they will.

Verdict

opencodex is the missing piece a lot of local-LLM setups have been waiting for: a first-party-quality client pointed at your own hardware, with a design (config-level proxying) that won’t fall over on every update. Test your model’s tool-calling, keep an eye on which features phone home, and the best chat frontend for your homelab models might just be one you already have installed.

Leave a Reply

Your email address will not be published. Required fields are marked *

WordPress Appliance - Powered by TurnKey Linux