Vlad Babii

Feature SmithFull stack
idea → spec → ship → review

click ↑ to go back
LinkedIn · opens in a new tab

Private Local AI Platform

AI & LLMsArchitectureDevelopmentInfrastructure

A self-hosted AI platform combining local GPU inference, autonomous agents, MCP services, browser and device control, asynchronous jobs, scheduling and persistent knowledge into one private execution environment.

From AI models to an AI system

The goal is not simply to run an LLM locally. The platform combines local inference with agents, tools, real devices, background jobs and persistent knowledge so an AI system can reason, act, remember and continue work without requiring every task to remain inside an interactive chat.

Two servers, separate responsibilities

GPU-heavy inference is separated from agent orchestration and control.

  • AI Compute ServerGPU-focused local inference on two GPUs: GPU1, and GPU2 with lower performance.
  • Agent & Control ServerRuns Herd, Roamgate, MCPProxy, gateways, queues, scheduling and supporting services.
  • Edge / Device NodesRaspberry Pi services, Android devices, Linux desktop control and browser execution.

Asymmetric GPU workloads

Different models and GPUs are assigned according to the type of work rather than treating the machine as one homogeneous inference resource.

GPU1 is used for heavy local reasoning, including long-context coding, analysis and larger local models.

GPU2, the lower-performance card, handles specialized workloads such as vision, speech transcription and lightweight supervisory inference.

The small supervisor models support tool calling, so they run in the same harness as the big model and can take over lighter steps without a separate integration.

This allows perception and lightweight control workloads to run independently while the larger GPU concentrates on expensive reasoning.

  • GPU1
  • GPU2 — lower performance
  • Local llama.cpp inference
  • Vision inference
  • Whisper speech recognition
  • Small supervisor models with tool calling
  • Model specialization

Agent runtime

Agents are treated as workers inside a larger execution platform rather than as the platform itself.

Herd provides the headless environment for running multiple AI coding agents in persistent sessions.

Roamgate is the web interface to the agent environment: from any browser, a user can start and follow agent sessions, review their output and step in when a decision needs a human.

Agents can use Claude, OpenAI, Gemini, local models or other CLI-capable systems depending on the task.

Long-running work can be moved into durable jobs so an interactive conversation does not have to remain occupied for hours.

  • Herd
  • Roamgate web interface
  • Pi / CLI agents
  • Local models
  • Remote AI providers
  • Persistent agent sessions

MCP capability fabric

MCP provides one stable capability surface between agents and the services around them.

Agents communicate through MCPProxy and an MCP gateway instead of needing direct knowledge of every underlying service.

The fabric can expose productivity systems, source control, search, knowledge, desktop control, Android devices, browser automation, infrastructure operations and other services.

This keeps the agent-facing interface stable even when individual service implementations change.

  • MCPProxy
  • MCP gateway
  • HTTP and stdio service integration
  • Productivity integrations
  • Source control
  • Search
  • Knowledge services
  • Device control
  • Infrastructure control
  • Code sandbox
MCP capability fabricMCP capability fabric
MCP capability fabric

Physical devices as AI interfaces

The platform treats real hardware as an extension of the agent environment.

A physical Android phone can act as both an event source and an output channel. Notifications and incoming messages can create structured jobs, while agents can send replies or perform controlled device actions.

Linux desktop sessions can be exposed through a dedicated automation account rather than giving autonomous agents the human user's credentials.

Raspberry Pi and other edge nodes can bridge physical hardware and peripherals into the same capability model.

  • Android notifications → jobs
  • Agent → Android replies
  • USB / ADB device access
  • Linux desktop automation
  • Raspberry Pi edge nodes
  • Controlled device APIs

Browser automation

Real Chromium is used when the real web interface is the required interface.

ChrAPI exposes a real Chromium environment through an HTTP-facing control layer.

It supports JavaScript-heavy websites, selectors, wait conditions, multi-step sessions and persistent browser state.

Successful operations can be cached server-side to avoid unnecessary repeated browser work.

VNC provides a human-visible fallback for inspection or manual intervention.

  • Real Chromium
  • JavaScript execution
  • Selector and condition waits
  • Persistent sessions
  • Server-side caching
  • VNC human fallback
  • LXC isolation

Jobs, queues and scheduling

Background work uses the same execution infrastructure as interactive requests, but with explicit priority and scheduling rules.

Events and requests become durable job definitions which enter priority queues before being assigned to workers.

Jobs can have calendar rules, allowed execution windows, restrictions and resource requirements.

Interactive work can outrank background workloads, while low-priority tasks can run during otherwise unused compute capacity.

Recurring jobs can evolve into named skills with their own schedule, restrictions and expected output.

  • Priority queues
  • Calendar-aware scheduling
  • Execution windows
  • Resource restrictions
  • Recurring jobs
  • Off-hours execution
  • Persistent artifacts
  • Job logs and status
Jobs and scheduling flowJobs and scheduling flow
Jobs and scheduling flow

Durable knowledge

Knowledge is stored independently from the agent conversation and retrieved progressively when needed.

Documents, videos, meeting recordings, notes and conversations can enter an ingestion pipeline.

Skills extract facts, decisions, concepts and useful context and write them into a cross-linked Markdown knowledge base.

The Markdown/wiki layer remains the durable source of truth while a search index provides fast retrieval.

Agents first discover compact relevant results and expand into full documents only when additional context is required.

  • Markdown knowledge store
  • Source provenance
  • Canonical pages
  • Incremental indexing
  • Keyword search
  • Semantic search
  • Hybrid retrieval
  • Progressive context loading
Knowledge ingestion and retrievalKnowledge ingestion and retrieval
Knowledge ingestion and retrieval

Multimodal pipeline

Perception workloads are separated from the main reasoning workload.

Audio can pass through Whisper for transcription while visual inputs are processed by the dedicated vision model.

The resulting observations become normalized context for the main reasoning model.

The same architecture can process screenshots, documents, video frames and other multimodal inputs.

  • Audio / video
  • Whisper
  • Vision model
  • Normalized observations
  • Local reasoning model

Linux identity and isolation

Human interaction, desktop automation and AI execution use separate identities.

The human uses the normal Linux user account for interactive work.

A dedicated agent account owns desktop and device automation sessions.

A separate AI account runs model-facing agent processes and AI workloads.

Long-running infrastructure services use dedicated service accounts where appropriate.

The objective is to expose capabilities through explicit interfaces rather than handing autonomous processes unrestricted host credentials.

  • user — human interactive session
  • agent — desktop and device automation
  • ai — AI agent runtime
  • svc-* — infrastructure services

The Debian 13 and user/agent/ai separation is an implementation design direction; the technical documentation explicitly distinguishes proposed operating-system structure from components already deployed.

End-to-end workflow

The components form one execution path from human intent or external events to autonomous action.

A user or external device generates an event.

The control layer creates or routes a job.

Herd assigns the work to an agent.

The agent uses MCP to access models, search, knowledge, source control, devices or infrastructure.

Results become responses, artifacts, durable knowledge or further jobs.

The system can communicate the outcome back through the web interface, terminal, Android phone or another configured channel.

Example: Android message to autonomous response

A notification arrives on the physical phone and is converted into a structured event.

The event creates a job which enters the priority queue.

An agent retrieves relevant context through Google, Slack, Wiki, GitHub or other MCP services.

The agent reasons over the context and produces a response.

Depending on the workflow, the response can be returned automatically or routed to a human for approval before being sent through the phone.

Example: research assistant

An agent searches the web through the self-hosted search service, identifies useful sources and reads only the relevant portions.

The local reasoning model synthesizes the evidence into an answer, comparison or recommendation.

Useful findings can be written into the persistent knowledge base so later agents do not need to repeat the research.

Example: software development

A development task enters the agent environment.

The agent searches Git repositories, issues, documentation and project knowledge through MCP.

Code changes are made inside a bounded workspace and executed through a sandbox.

Tests and validation run before the result is reported.

Long-running work can remain as a durable job with logs, artifacts and status.

Example: self-learning knowledge pipeline

Videos, meetings, notes, conversations and documents are ingested by the appropriate skill.

Content is transcribed or extracted and reduced into reusable facts, concepts, decisions and references.

A knowledge-writer maintains canonical Markdown pages with provenance, aliases, links and history.

Search indexes are updated incrementally, providing agents with small, targeted context instead of entire source documents.

Architecture principles

  • Separate compute from control
  • Specialize models by workload
  • Expose capabilities instead of credentials
  • Separate human, agent and AI identities
  • Keep interactive and background work separate
  • Schedule resources rather than conversations
  • Treat real interfaces as real interfaces
  • Search before loading large context
  • Keep durable knowledge separate from retrieval indexes
  • Make every component replaceable
Platform architecturePlatform architecture
Platform architecture

The result

The result is a small private AI operating environment rather than a standalone model server.

Local models provide reasoning, perception and speech capabilities. Agents provide autonomous work. MCP provides the capability fabric. Queues and scheduling provide time and resource management. Persistent knowledge provides memory. Real devices and browsers provide a bridge into the physical and human world.

Stack
Debian · C++ / llama.cpp · Local LLMs · GPU1 · GPU2 · Herd · Roamgate · MCP · MCPProxy · Node.js · Python · SearXNG · Meilisearch · Markdown · Git · Docker · LXC · Chromium · VNC · Android / ADB · Raspberry Pi
Hardware
GPU1 · GPU2 (lower performance) · 48 GB system RAM · 1000 W PSU · Android phone · Raspberry Pi edge nodes