A calm map of the machinery

Our AI Stack

The pieces are intentionally visible: delivery flows into a local runtime, the runtime leans on dependable services, and AI connectivity stays replaceable.

flow and delivery explanation available
01 / Delivery

A repeatable path from change to running system.

Infrastructure is described as code, delivered through a predictable pipeline, and hosted close to the people and devices it serves.

GitLab CI

Tests, packages, and deploys changes with a visible history.

Terraform

Defines the infrastructure and the boundary between managed services.

Cloud-native local hosting environment

Runs the stack as observable containers on local infrastructure, keeping ownership and latency clear.

02 / Application stack

Five layers work together as one stack.

Follow the path from user interfaces through orchestration and services to persistent data and operational visibility.

User interfaces
Telegram

Conversation entry point for Hermes.

Open WebUI

Browser interface for model conversations through LiteLLM.

Nextcloud UI

Human interface for files and collaboration.

Pebble Smartwatch

Glanceable interface for Gotchi watchface state.

People interact through chat, a browser, collaboration tools, and a glanceable watch face. Each interface connects to the service suited to that task.

Agentic orchestration
Hermes

Coordinates tools, services, and model requests.

n8n

Supports repeatable workflow automation.

Hermes routes assistant requests to tools, services, and model providers; n8n runs repeatable multi-step workflows. This layer coordinates work without owning the data those services store.

Services
Hindsight

Provides long-term memory and retrieval, using PostgreSQL with pgvector for storage.

Nextcloud

Provides files, calendars, and collaboration services.

Pushover

Delivers notifications from Hermes.

Vikunja

Provides task and work management.

These services provide capabilities the orchestration layer can call. Hindsight is a memory service; its stored memory lives in PostgreSQL with pgvector.

Storage
PostgreSQL + pgvector

Stores structured service data and Hindsight's vector-backed memory.

Nextcloud Redis

Supports fast coordination for Nextcloud; durable files remain in separate storage.

Prometheus

Stores scraped time-series metrics for later queries and dashboards.

Persistent databases, caches, and metrics history give services and operators the state they need across requests and restarts.

Monitoring
Grafana

Turns collected metrics into readable dashboards.

cAdvisor

Exposes container metrics for Prometheus to scrape.

Node Exporter

Exposes host metrics for Prometheus to scrape.

Exporters expose host and container health, while Grafana presents the metrics Prometheus stores as dashboards.

03 / Agent orchestration

Hermes brings the right tools and models into one answer.

One request can draw on the right model and the right services. Hermes gathers context, calls what the task needs, and returns a useful response.

Follow a request from conversation to a complete answer
request User device
orchestrator Hermes Understand and plan Compose response
services as needed Hindsight · recalled context Nextcloud · files and calendar Vikunja · tasks Pushover · notifications
model options LiteLLM Local AI server OpenRouter OpenAI Gemini Anthropic

Hermes can recall relevant context, choose a suitable model through LiteLLM, and use services such as Nextcloud, Vikunja, or Pushover when the request calls for them. The results come back together as one response. A request uses only the services and model routes it needs.

04 / AI connectivity

One calm gateway, many model choices.

LiteLLM gives the orchestration layer a consistent interface to local models and named hosted providers.

LiteLLM

A consistent gateway that routes each request to the best available local or hosted model.

LiteLLM → Local AI server
Local AI server

Qwen, llama.cpp, and Firecrawl support private, local-first workflows where they fit best.

LiteLLM → OpenRouter
OpenRouter

A hosted model gateway for broad model selection behind one interface.

LiteLLM → OpenAI
OpenAI

Hosted reasoning and language models for workloads that benefit from them.

LiteLLM → Gemini
Gemini

Hosted multimodal and language capabilities available through the same routing boundary.

LiteLLM → Anthropic
Anthropic

Hosted language models for careful, capable agent interactions.

Want to build something dependable?

Tell us what is taking too much of your time. We will help you find the first useful step.

contact@autom8ai.ca  ↗