Private Local AI Platform
AI & LLMsArchitectureDevelopmentInfrastructure
A self-hosted AI platform combining local GPU inference, autonomous agents, MCP services, browser and device control, asynchronous jobs, scheduling and persistent knowledge into one private execution environment.
From AI models to an AI system
The goal is not simply to run an LLM locally. The platform combines local inference with agents, tools, real devices, background jobs and persistent knowledge so an AI system can reason, act, remember and continue work without requiring every task to remain inside an interactive chat.
Two servers, separate responsibilities
GPU-heavy inference is separated from agent orchestration and control.
Human / Client→Agent & Control→MCP Fabric→AI Compute / Devices / Services
- AI Compute ServerGPU-focused local inference on two GPUs: GPU1, and GPU2 with lower performance.
- Agent & Control ServerRuns Herd, Roamgate, MCPProxy, gateways, queues, scheduling and supporting services.
- Edge / Device NodesRaspberry Pi services, Android devices, Linux desktop control and browser execution.
Asymmetric GPU workloads
Different models and GPUs are assigned according to the type of work rather than treating the machine as one homogeneous inference resource.
GPU1 is used for heavy local reasoning, including long-context coding, analysis and larger local models.
GPU2, the lower-performance card, handles specialized workloads such as vision, speech transcription and lightweight supervisory inference.
The small supervisor models support tool calling, so they run in the same harness as the big model and can take over lighter steps without a separate integration.
This allows perception and lightweight control workloads to run independently while the larger GPU concentrates on expensive reasoning.
- GPU1
- GPU2 — lower performance
- Local llama.cpp inference
- Vision inference
- Whisper speech recognition
- Small supervisor models with tool calling
- Model specialization
Agent runtime
Agents are treated as workers inside a larger execution platform rather than as the platform itself.
Herd provides the headless environment for running multiple AI coding agents in persistent sessions.
Roamgate is the web interface to the agent environment: from any browser, a user can start and follow agent sessions, review their output and step in when a decision needs a human.
Agents can use Claude, OpenAI, Gemini, local models or other CLI-capable systems depending on the task.
Long-running work can be moved into durable jobs so an interactive conversation does not have to remain occupied for hours.
- Herd
- Roamgate web interface
- Pi / CLI agents
- Local models
- Remote AI providers
- Persistent agent sessions
MCP capability fabric
MCP provides one stable capability surface between agents and the services around them.
Agents communicate through MCPProxy and an MCP gateway instead of needing direct knowledge of every underlying service.
The fabric can expose productivity systems, source control, search, knowledge, desktop control, Android devices, browser automation, infrastructure operations and other services.
This keeps the agent-facing interface stable even when individual service implementations change.
- MCPProxy
- MCP gateway
- HTTP and stdio service integration
- Productivity integrations
- Source control
- Search
- Knowledge services
- Device control
- Infrastructure control
- Code sandbox


Physical devices as AI interfaces
The platform treats real hardware as an extension of the agent environment.
A physical Android phone can act as both an event source and an output channel. Notifications and incoming messages can create structured jobs, while agents can send replies or perform controlled device actions.
Linux desktop sessions can be exposed through a dedicated automation account rather than giving autonomous agents the human user's credentials.
Raspberry Pi and other edge nodes can bridge physical hardware and peripherals into the same capability model.
- Android notifications → jobs
- Agent → Android replies
- USB / ADB device access
- Linux desktop automation
- Raspberry Pi edge nodes
- Controlled device APIs
Browser automation
Real Chromium is used when the real web interface is the required interface.
ChrAPI exposes a real Chromium environment through an HTTP-facing control layer.
It supports JavaScript-heavy websites, selectors, wait conditions, multi-step sessions and persistent browser state.
Successful operations can be cached server-side to avoid unnecessary repeated browser work.
VNC provides a human-visible fallback for inspection or manual intervention.
- Real Chromium
- JavaScript execution
- Selector and condition waits
- Persistent sessions
- Server-side caching
- VNC human fallback
- LXC isolation
Jobs, queues and scheduling
Background work uses the same execution infrastructure as interactive requests, but with explicit priority and scheduling rules.
Events and requests become durable job definitions which enter priority queues before being assigned to workers.
Jobs can have calendar rules, allowed execution windows, restrictions and resource requirements.
Interactive work can outrank background workloads, while low-priority tasks can run during otherwise unused compute capacity.
Recurring jobs can evolve into named skills with their own schedule, restrictions and expected output.
- Priority queues
- Calendar-aware scheduling
- Execution windows
- Resource restrictions
- Recurring jobs
- Off-hours execution
- Persistent artifacts
- Job logs and status


Durable knowledge
Knowledge is stored independently from the agent conversation and retrieved progressively when needed.
Documents, videos, meeting recordings, notes and conversations can enter an ingestion pipeline.
Skills extract facts, decisions, concepts and useful context and write them into a cross-linked Markdown knowledge base.
The Markdown/wiki layer remains the durable source of truth while a search index provides fast retrieval.
Agents first discover compact relevant results and expand into full documents only when additional context is required.
- Markdown knowledge store
- Source provenance
- Canonical pages
- Incremental indexing
- Keyword search
- Semantic search
- Hybrid retrieval
- Progressive context loading


Multimodal pipeline
Perception workloads are separated from the main reasoning workload.
Audio can pass through Whisper for transcription while visual inputs are processed by the dedicated vision model.
The resulting observations become normalized context for the main reasoning model.
The same architecture can process screenshots, documents, video frames and other multimodal inputs.
- Audio / video
- Whisper
- Vision model
- Normalized observations
- Local reasoning model
Linux identity and isolation
Human interaction, desktop automation and AI execution use separate identities.
The human uses the normal Linux user account for interactive work.
A dedicated agent account owns desktop and device automation sessions.
A separate AI account runs model-facing agent processes and AI workloads.
Long-running infrastructure services use dedicated service accounts where appropriate.
The objective is to expose capabilities through explicit interfaces rather than handing autonomous processes unrestricted host credentials.
- user — human interactive session
- agent — desktop and device automation
- ai — AI agent runtime
- svc-* — infrastructure services
The Debian 13 and user/agent/ai separation is an implementation design direction; the technical documentation explicitly distinguishes proposed operating-system structure from components already deployed.
End-to-end workflow
The components form one execution path from human intent or external events to autonomous action.
A user or external device generates an event.
The control layer creates or routes a job.
Herd assigns the work to an agent.
The agent uses MCP to access models, search, knowledge, source control, devices or infrastructure.
Results become responses, artifacts, durable knowledge or further jobs.
The system can communicate the outcome back through the web interface, terminal, Android phone or another configured channel.
User / Event→Job→Queue→Agent→MCP→Tools / Models / Devices→Result→Artifact / Knowledge / Reply
Example: Android message to autonomous response
A notification arrives on the physical phone and is converted into a structured event.
The event creates a job which enters the priority queue.
An agent retrieves relevant context through Google, Slack, Wiki, GitHub or other MCP services.
The agent reasons over the context and produces a response.
Depending on the workflow, the response can be returned automatically or routed to a human for approval before being sent through the phone.
Example: research assistant
An agent searches the web through the self-hosted search service, identifies useful sources and reads only the relevant portions.
The local reasoning model synthesizes the evidence into an answer, comparison or recommendation.
Useful findings can be written into the persistent knowledge base so later agents do not need to repeat the research.
Example: software development
A development task enters the agent environment.
The agent searches Git repositories, issues, documentation and project knowledge through MCP.
Code changes are made inside a bounded workspace and executed through a sandbox.
Tests and validation run before the result is reported.
Long-running work can remain as a durable job with logs, artifacts and status.
Example: self-learning knowledge pipeline
Videos, meetings, notes, conversations and documents are ingested by the appropriate skill.
Content is transcribed or extracted and reduced into reusable facts, concepts, decisions and references.
A knowledge-writer maintains canonical Markdown pages with provenance, aliases, links and history.
Search indexes are updated incrementally, providing agents with small, targeted context instead of entire source documents.
Architecture principles
- Separate compute from control
- Specialize models by workload
- Expose capabilities instead of credentials
- Separate human, agent and AI identities
- Keep interactive and background work separate
- Schedule resources rather than conversations
- Treat real interfaces as real interfaces
- Search before loading large context
- Keep durable knowledge separate from retrieval indexes
- Make every component replaceable


The result
The result is a small private AI operating environment rather than a standalone model server.
Local models provide reasoning, perception and speech capabilities. Agents provide autonomous work. MCP provides the capability fabric. Queues and scheduling provide time and resource management. Persistent knowledge provides memory. Real devices and browsers provide a bridge into the physical and human world.
- Stack
- Debian · C++ / llama.cpp · Local LLMs · GPU1 · GPU2 · Herd · Roamgate · MCP · MCPProxy · Node.js · Python · SearXNG · Meilisearch · Markdown · Git · Docker · LXC · Chromium · VNC · Android / ADB · Raspberry Pi
- Hardware
- GPU1 · GPU2 (lower performance) · 48 GB system RAM · 1000 W PSU · Android phone · Raspberry Pi edge nodes

