Skip to content
← Projects

project

Local AI Lab

Local models, search, image workflows, and Home Assistant on hardware I own.

status
active
period
2026 to present
core stack
Qwen3.8 · llama.cpp · CUDA · SearXNG · ComfyUI · Home Assistant

This is where most of my tinkering time goes. I run local language, vision, speech, image, and video models, then connect them to tools I actually use. It all stays on my own hardware.

Local inference

My main model is Qwen3.8-27B UD-Q4_K_XL, served through llama.cpp and llama-swap. It runs with a 114,688-token context on an RTX 5090. llama-swap lets several models share one OpenAI-compatible endpoint and loads the requested one when needed.

I size each model against the VRAM that is actually free while the desktop and speech services are running. MTP speculative decoding took Qwen from 62 to 111 tokens per second without cutting the 114K context. Other slots handle Home Assistant, embeddings, and vision.

Tools around the model

Open WebUI keeps the chat interface and history on my homelab. The agent can search through SearXNG, open pages in a headless Playwright browser, and run code in a sandboxed terminal. I also expose search over MCP for coding agents.

The browser and terminal stay isolated, and the model endpoints stay on the LAN. I check changes against the running service instead of trusting a command just because it exited cleanly.

ComfyUI

ComfyUI runs on the same GPU workstation. I keep six image and video workflows that I have actually tested and documented, including their models, custom nodes, VRAM needs, and recovery steps. That makes them much easier to return to after an update than a folder full of one-off graphs.

Voice in Home Assistant

Home Assistant uses a local Qwen3.5-9B model for device control and SearXNG search. Parakeet handles speech recognition and Kokoro handles speech synthesis through a Wyoming bridge. The warm round trip is about 1.32 seconds.

A small proxy wakes the GPU machine when it is needed. GPU-backed services stay on the workstation, while Open WebUI, search, browser automation, and chat data live on the homelab.