Files
server-infra/docs/OBSERVABILITY.md
2026-07-10 06:50:52 -05:00

963 lines
28 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Observability — Alloy → Loki / Prometheus → Grafana
Centralized **logs + metrics** for the homelab. Every managed host runs
**Grafana Alloy** (collector). **Loki** (logs), **Prometheus** (metrics), and
**Grafana** (UI / dashboards / alerts) run on **ai-server-4080**. Grafana is
exposed through Nginx Proxy Manager at `10.0.0.230` on an external domain.
```mermaid
flowchart LR
subgraph hosts ["All webservers"]
A["adama\nAlloy"]
R["roslin\nAlloy"]
C["ai-server-4080\nAlloy"]
end
subgraph stack ["ai-server-4080"]
L["Loki :3100"]
P["Prometheus :9090"]
G["Grafana :3000"]
end
NPM["NPM\n10.0.0.230"]
DNS["grafana.YOURDOMAIN"]
A -->|logs| L
R -->|logs| L
C -->|logs| L
A -->|metrics| P
R -->|metrics| P
C -->|metrics| P
L --> G
P --> G
G --> NPM
DNS --> NPM
```
| Piece | Where | Role |
|-------|--------|------|
| **Alloy** | every host (`adama`, `roslin`, `ai-server-4080`) | Ship journald + Docker **logs** to Loki; scrape host + container **metrics** → Prometheus |
| **Loki** | `ai-server-4080` only | Store and index logs |
| **Prometheus** | `ai-server-4080` only | Store metrics (CPU, RAM, disk, container health) |
| **Grafana** | `ai-server-4080` only | Explore logs/metrics, dashboards, alerts |
| **NPM** | `10.0.0.230` (external) | TLS + public hostname → Grafana |
**What you get in Grafana**
| View | Datasource | Examples |
|------|------------|----------|
| App / host logs | Loki | `{host="adama", env="prod", app="dta_service"}` |
| Container health | Prometheus | Up/down, restarts, CPU %, memory, network I/O per container |
| System status | Prometheus | Host CPU, RAM, disk, load, network |
This guide is written for a first-time setup. Do the **manual bootstrap** on
ai-server-4080 first so you can see logs **and** metrics flowing. Then wire
Alloy into Ansible provision (`site.yml`) so new hosts get it automatically.
---
## Prerequisites
- Docker + Compose already on all three hosts (`roles/docker` via `./scripts/provision.sh`)
- LAN reachability: hosts can talk to `10.0.0.128:3100` (Loki) and
`10.0.0.128:9090` (Prometheus remote-write)
- Access to Nginx Proxy Manager at `10.0.0.230`
- A DNS name you control (example used below: `grafana.aimloperations.com` — replace with yours)
- SMTP2GO account (for alerts; covered in [GRAFANA_USAGE.md](GRAFANA_USAGE.md))
Suggested ports (keep free on ai-server-4080):
| Service | Host port | Who can reach it |
|---------|-----------|------------------|
| Loki | `3100` | LAN only (`10.0.0.0/24`) |
| Prometheus | `9090` | LAN only (`10.0.0.0/24`) — Alloy remote-write + Grafana |
| Grafana | `3000` | LAN (NPM proxies to it); do **not** open to the public internet directly |
---
## Part 1 — Central stack on ai-server-4080 (Loki + Prometheus + Grafana)
Run these steps **on ai-server-4080** as `westfarn`.
### 1.1 Create directories
```bash
sudo mkdir -p /opt/apps/observability/{loki,prometheus,grafana,alloy}
sudo chown -R westfarn:westfarn /opt/apps/observability
mkdir -p /opt/apps/observability/loki/{data,rules}
mkdir -p /opt/apps/observability/prometheus/data
mkdir -p /opt/apps/observability/grafana/data
```
### 1.2 Loki config
Create `/opt/apps/observability/loki/loki-config.yml`:
```yaml
# Managed manually for first bootstrap; later: roles/observability template.
auth_enabled: false
server:
http_listen_port: 3100
grpc_listen_port: 9096
log_level: info
common:
instance_addr: 127.0.0.1
path_prefix: /loki
storage:
filesystem:
chunks_directory: /loki/chunks
rules_directory: /loki/rules
replication_factor: 1
ring:
kvstore:
store: inmemory
schema_config:
configs:
- from: "2024-01-01"
store: tsdb
object_store: filesystem
schema: v13
index:
prefix: index_
period: 24h
# Keep labels low-cardinality. host / env / app / job are fine.
# Do NOT put request IDs, user IDs, or full paths in labels.
limits_config:
reject_old_samples: true
reject_old_samples_max_age: 168h
ingestion_rate_mb: 16
ingestion_burst_size_mb: 32
max_query_series: 500
retention_period: 744h # 31 days
compactor:
working_directory: /loki/compactor
compaction_interval: 10m
retention_enabled: true
retention_delete_delay: 2h
delete_request_store: filesystem
ruler:
alertmanager_url: http://localhost:9093 # unused for now; Grafana handles alerts
```
### 1.3 Prometheus config
Create `/opt/apps/observability/prometheus/prometheus.yml`:
```yaml
# Managed manually for first bootstrap; later: roles/observability template.
global:
scrape_interval: 15s
evaluation_interval: 15s
# Alloy on each host *pushes* metrics here (remote_write).
# We still keep a tiny local scrape so Prometheus has a self-health target.
scrape_configs:
- job_name: prometheus
static_configs:
- targets: ["localhost:9090"]
labels:
host: ai-server-4080
```
Prometheus must accept remote-write from Alloy — that is enabled with the
`--web.enable-remote-write-receiver` flag in compose (below).
### 1.4 Docker Compose for Loki + Prometheus + Grafana
Create `/opt/apps/observability/docker-compose.yml`:
```yaml
# Managed manually for first bootstrap; later: roles/observability template.
services:
loki:
image: grafana/loki:3.4.2
container_name: loki
restart: unless-stopped
user: "0:0"
command: -config.file=/etc/loki/loki-config.yml
ports:
# Bind to all interfaces so Alloy on adama/roslin can push.
# Firewall (UFW) should restrict who can connect — see 1.6.
- "3100:3100"
volumes:
- ./loki/loki-config.yml:/etc/loki/loki-config.yml:ro
- ./loki/data:/loki
healthcheck:
test: ["CMD-SHELL", "wget -qO- http://127.0.0.1:3100/ready || exit 1"]
interval: 15s
timeout: 5s
retries: 10
prometheus:
image: prom/prometheus:v3.2.1
container_name: prometheus
restart: unless-stopped
command:
- --config.file=/etc/prometheus/prometheus.yml
- --storage.tsdb.path=/prometheus
- --storage.tsdb.retention.time=31d
- --web.enable-remote-write-receiver
- --web.enable-lifecycle
ports:
- "9090:9090"
volumes:
- ./prometheus/prometheus.yml:/etc/prometheus/prometheus.yml:ro
- ./prometheus/data:/prometheus
healthcheck:
test: ["CMD-SHELL", "wget -qO- http://127.0.0.1:9090/-/ready || exit 1"]
interval: 15s
timeout: 5s
retries: 10
grafana:
image: grafana/grafana:11.5.2
container_name: grafana
restart: unless-stopped
depends_on:
loki:
condition: service_healthy
prometheus:
condition: service_healthy
ports:
# LAN only is fine; NPM on 10.0.0.230 will proxy to this.
- "3000:3000"
environment:
GF_SECURITY_ADMIN_USER: admin
# Change immediately after first login. Prefer a secrets file later.
GF_SECURITY_ADMIN_PASSWORD: "CHANGE_ME_ON_FIRST_LOGIN"
GF_USERS_ALLOW_SIGN_UP: "false"
GF_SERVER_ROOT_URL: "https://grafana.aimloperations.com"
GF_SERVER_DOMAIN: "grafana.aimloperations.com"
# SMTP is configured in the UI or via GF_SMTP_* — see GRAFANA_USAGE.md
volumes:
- ./grafana/data:/var/lib/grafana
- ./grafana/provisioning:/etc/grafana/provisioning:ro
```
Replace `grafana.aimloperations.com` with your real hostname in
`GF_SERVER_ROOT_URL` and `GF_SERVER_DOMAIN`.
### 1.5 Provision Grafana datasources (Loki + Prometheus)
Create provisioning so Grafana auto-connects on every restart:
```bash
mkdir -p /opt/apps/observability/grafana/provisioning/datasources
```
Create `/opt/apps/observability/grafana/provisioning/datasources/datasources.yml`:
```yaml
apiVersion: 1
datasources:
- name: Loki
type: loki
access: proxy
url: http://loki:3100
isDefault: false
editable: false
jsonData:
maxLines: 1000
- name: Prometheus
type: prometheus
access: proxy
url: http://prometheus:9090
isDefault: true
editable: false
jsonData:
timeInterval: 15s
```
Prometheus is the default datasource so new panels default to metrics;
switch to Loki in Explore when searching logs.
### 1.6 Firewall — allow Loki + Prometheus from LAN only
On **ai-server-4080**, open ingest ports to the LAN (not the world):
```bash
sudo ufw allow from 10.0.0.0/24 to any port 3100 proto tcp comment 'Loki ingest from Alloy'
sudo ufw allow from 10.0.0.0/24 to any port 9090 proto tcp comment 'Prometheus remote-write from Alloy'
sudo ufw allow from 10.0.0.0/24 to any port 3000 proto tcp comment 'Grafana for NPM'
sudo ufw status numbered
```
Later, when this is Ansible-managed, add these ports to a host-specific UFW
allow list in `inventory/host_vars/ai-server-4080.yml` rather than opening
them on every webserver via `ufw_allowed_tcp_ports`.
### 1.7 Start the stack
```bash
cd /opt/apps/observability
docker compose up -d
docker compose ps
curl -s http://127.0.0.1:3100/ready # expect: ready
curl -s http://127.0.0.1:9090/-/ready # expect: Prometheus Server is Ready.
curl -s -o /dev/null -w "%{http_code}\n" http://127.0.0.1:3000/login # expect: 200
```
Open `http://10.0.0.128:3000` from your LAN, log in with `admin` /
`CHANGE_ME_ON_FIRST_LOGIN`, and set a strong password when prompted.
Confirm both datasources: **Connections → Data sources** → Loki and
Prometheus both **Save & test** green.
---
## Part 2 — Expose Grafana on an external domain (NPM)
Nginx Proxy Manager lives at **10.0.0.230** and is **not** managed by this
repo. Do this in the NPM UI.
### 2.1 DNS
Create an A (or CNAME) record for your chosen hostname pointing at the public
IP that reaches NPM (same pattern as your other `*.aimloperations.com` hosts).
Example: `grafana.aimloperations.com` → your public IP / NPM.
### 2.2 Proxy Host in NPM
1. Log into Nginx Proxy Manager.
2. **Hosts → Proxy Hosts → Add Proxy Host**.
3. **Details**
- Domain Names: `grafana.aimloperations.com` (your name)
- Scheme: `http`
- Forward Hostname / IP: `10.0.0.128` (ai-server-4080)
- Forward Port: `3000`
- Cache Assets: off
- Block Common Exploits: on
- Websockets Support: **on** (Grafana live features need this)
4. **SSL**
- Request a new Let's Encrypt certificate
- Force SSL: on
- HTTP/2: on
5. Save.
### 2.3 Confirm
```bash
curl -sI https://grafana.aimloperations.com/login
```
You should get HTTP 200 (or a redirect to login). Update
`GF_SERVER_ROOT_URL` if the hostname differs, then:
```bash
cd /opt/apps/observability && docker compose up -d grafana
```
---
## Part 3 — Alloy on each machine (logs + metrics)
Alloy runs on **every** host in `webservers`. It:
1. Reads **systemd journal** (host / service logs) → Loki
2. Reads **Docker container logs** → Loki
3. Scrapes **host metrics** (`prometheus.exporter.unix` ≈ node_exporter) → Prometheus
4. Scrapes **container metrics** (`prometheus.exporter.cadvisor`) → Prometheus
5. Attaches labels: `host`, `env`, `app`, `job`, `unit` / `container` / `name`
6. Pushes logs to `http://10.0.0.128:3100/loki/api/v1/push`
7. Remote-writes metrics to `http://10.0.0.128:9090/api/v1/write`
### Labels we care about
**Logs (Loki)**
| Label | Source | Example values |
|-------|--------|----------------|
| `host` | Ansible inventory hostname | `adama`, `roslin`, `ai-server-4080` |
| `job` | collector name | `systemd`, `docker` |
| `env` | Docker Compose project suffix | `beta`, `prod`, `host`, `infra` |
| `app` | Compose project prefix | `company_site`, `dta_service`, `dta_webapp`, … |
| `unit` | systemd unit (journal only) | `docker.service`, `ssh.service` |
| `container` | Docker container name | `company_site_prod-web-1` |
**Metrics (Prometheus)** — same `host` / `env` / `app` idea, plus cAdvisor fields:
| Label | Source | Example values |
|-------|--------|----------------|
| `host` | Added by Alloy relabel | `adama`, `roslin`, `ai-server-4080` |
| `job` | scrape job name | `node`, `cadvisor` |
| `name` | container name (cAdvisor) | `company_site_prod-web-1` |
| `container_label_com_docker_compose_project` | Compose project | `dta_service_prod` |
| `env` / `app` | Parsed from Compose project by Alloy | `prod` / `dta_service` |
| `instance` | scrape target | usually the Alloy exporter address |
Your deploy convention already names Compose projects `<app>_<env>`
(e.g. `company_site_prod`). Alloy parses that for both logs and metrics:
```logql
{host="adama", env="prod", app="dta_service"}
```
```promql
container_memory_usage_bytes{host="adama", env="prod", app="dta_service"}
```
### 3.1 Install Alloy (manual — one host)
Repeat on `adama`, `roslin`, and `ai-server-4080`. Example for **adama**:
```bash
sudo mkdir -p /opt/apps/observability/alloy
sudo chown -R westfarn:westfarn /opt/apps/observability
```
Create `/opt/apps/observability/alloy/config.alloy`. **Change every
`host = "adama"`** on each machine (`adama` / `roslin` / `ai-server-4080`):
```river
// Grafana Alloy — logs → Loki, metrics → Prometheus.
// host label MUST match inventory hostname.
// =====================================================================
// LOGS
// =====================================================================
// ---- systemd journal ----
loki.source.journal "system" {
forward_to = [loki.process.journal_labels.receiver]
relabel_rules = discovery.relabel.journal.rules
labels = {
job = "systemd",
host = "adama", // <<< CHANGE PER HOST
}
}
discovery.relabel "journal" {
targets = []
rule {
source_labels = ["__journal__systemd_unit"]
target_label = "unit"
}
}
loki.process "journal_labels" {
forward_to = [loki.write.default.receiver]
stage.static_labels {
values = {
env = "host",
app = "system",
}
}
}
// ---- Docker containers (logs) ----
discovery.docker "containers" {
host = "unix:///var/run/docker.sock"
}
discovery.relabel "docker" {
targets = discovery.docker.containers.targets
rule {
source_labels = ["__meta_docker_container_name"]
regex = "/(.*)"
target_label = "container"
}
rule {
source_labels = ["__meta_docker_container_label_com_docker_compose_project"]
target_label = "compose_project"
}
rule {
source_labels = ["__meta_docker_container_label_com_docker_compose_service"]
target_label = "compose_service"
}
}
loki.source.docker "containers" {
host = "unix:///var/run/docker.sock"
targets = discovery.relabel.docker.output
forward_to = [loki.process.docker_labels.receiver]
labels = {
job = "docker",
host = "adama", // <<< CHANGE PER HOST
}
}
loki.process "docker_labels" {
forward_to = [loki.write.default.receiver]
// compose_project is like "company_site_prod" or "dta_service_beta"
stage.regex {
source = "compose_project"
expression = "^(?P<app>[a-z0-9_]+)_(?P<env>beta|prod)$"
}
stage.labels {
values = {
app = "",
env = "",
}
}
// Containers not named <app>_<env> (loki, grafana, web-static, …)
stage.template {
source = "env"
template = `{{ if .env }}{{ .env }}{{ else }}infra{{ end }}`
}
stage.template {
source = "app"
template = `{{ if .app }}{{ .app }}{{ else }}{{ .compose_project }}{{ end }}`
}
stage.labels {
values = {
app = "",
env = "",
}
}
}
loki.write "default" {
endpoint {
url = "http://10.0.0.128:3100/loki/api/v1/push"
}
}
// =====================================================================
// METRICS — host (CPU / RAM / disk / load / network)
// =====================================================================
prometheus.exporter.unix "node" {
// When Alloy runs in Docker, point at the mounted host filesystem.
// Matches the volume mounts in alloy/docker-compose.yml below.
procfs_path = "/host/proc"
sysfs_path = "/host/sys"
rootfs_path = "/host/root"
}
prometheus.scrape "node" {
targets = prometheus.exporter.unix.node.targets
forward_to = [prometheus.relabel.add_host.receiver]
scrape_interval = "15s"
job_name = "node"
}
// =====================================================================
// METRICS — Docker containers (cAdvisor)
// =====================================================================
prometheus.exporter.cadvisor "docker" {
docker_host = "unix:///var/run/docker.sock"
docker_only = true
storage_duration = "5m"
// Keep cardinality down: only promote Compose labels we care about.
store_container_labels = false
allowlisted_container_labels = [
"com.docker.compose.project",
"com.docker.compose.service",
]
}
prometheus.scrape "cadvisor" {
targets = prometheus.exporter.cadvisor.docker.targets
forward_to = [prometheus.relabel.cadvisor_labels.receiver]
scrape_interval = "15s"
job_name = "cadvisor"
}
// Parse compose project → app + env; stamp inventory hostname.
prometheus.relabel "cadvisor_labels" {
forward_to = [prometheus.remote_write.default.receiver]
rule {
target_label = "host"
replacement = "adama" // <<< CHANGE PER HOST
}
// container_label_com_docker_compose_project → compose_project helper
rule {
source_labels = ["container_label_com_docker_compose_project"]
target_label = "compose_project"
}
// company_site_prod → app=company_site
rule {
source_labels = ["compose_project"]
regex = "^([a-z0-9_]+)_(beta|prod)$"
target_label = "app"
replacement = "${1}"
}
// company_site_prod → env=prod
rule {
source_labels = ["compose_project"]
regex = "^([a-z0-9_]+)_(beta|prod)$"
target_label = "env"
replacement = "${2}"
}
// Non-app containers (loki, grafana, alloy, web-static, …)
rule {
source_labels = ["env"]
regex = "^$"
target_label = "env"
replacement = "infra"
}
rule {
source_labels = ["app"]
regex = "^$"
target_label = "app"
replacement = "infra"
}
}
// Stamp host on node metrics too.
prometheus.relabel "add_host" {
forward_to = [prometheus.remote_write.default.receiver]
rule {
target_label = "host"
replacement = "adama" // <<< CHANGE PER HOST
}
}
prometheus.remote_write "default" {
endpoint {
url = "http://10.0.0.128:9090/api/v1/write"
}
}
```
### 3.2 Run Alloy with Docker Compose (per host)
cAdvisor + node metrics need host filesystem mounts and privileged mode.
Create `/opt/apps/observability/alloy/docker-compose.yml` on **each** host:
```yaml
services:
alloy:
image: grafana/alloy:v1.7.1
container_name: alloy
restart: unless-stopped
privileged: true
pid: host
command:
- run
- /etc/alloy/config.alloy
- --storage.path=/var/lib/alloy/data
- --server.http.listen-addr=0.0.0.0:12345
volumes:
- ./config.alloy:/etc/alloy/config.alloy:ro
- alloy-data:/var/lib/alloy/data
# Logs
- /var/run/docker.sock:/var/run/docker.sock:ro
- /var/log/journal:/var/log/journal:ro
- /etc/machine-id:/etc/machine-id:ro
- /run/systemd/journal:/run/systemd/journal:ro
# Metrics (node + cAdvisor) — host views
- /proc:/host/proc:ro
- /sys:/host/sys:ro
- /:/host/root:ro
- /var/run:/var/run:ro
- /var/lib/docker:/var/lib/docker:ro
- /dev/disk:/dev/disk:ro
# No published ports required — Alloy pushes outbound to Loki + Prometheus.
# Optional: expose Alloy's own UI/metrics on LAN for debugging:
# ports:
# - "12345:12345"
volumes:
alloy-data:
```
Start:
```bash
cd /opt/apps/observability/alloy
docker compose up -d
docker compose logs -f --tail=50
```
Healthy Alloy logs mention connecting / sending without repeated
`connection refused` to either `:3100` or `:9090`.
### 3.3 Verify logs in Grafana (Loki)
1. Open Grafana → **Explore** → datasource **Loki**.
2. Query:
```logql
{host=~".+"}
```
3. Confirm label browser shows `host`, `env`, `app`, `job`.
4. Spot-check each machine:
```logql
{host="adama"}
{host="roslin"}
{host="ai-server-4080"}
```
5. Spot-check env / app:
```logql
{env="prod", app="company_site"}
{env="beta", app="dta_service"}
{job="systemd", host="adama"}
```
If Docker logs show `env="infra"`, that container is not a `<app>_<env>` Compose
project (expected for `loki`, `grafana`, `alloy`, `web-static`).
### 3.4 Verify metrics in Grafana (Prometheus)
1. **Explore** → datasource **Prometheus**.
2. Host metrics present:
```promql
up{job="node"}
node_load1{host=~".+"}
100 - (avg by (host) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)
```
3. Container metrics present:
```promql
up{job="cadvisor"}
container_memory_usage_bytes{host=~".+", name!=""}
rate(container_cpu_usage_seconds_total{host=~".+", name!=""}[5m])
```
4. Labels `host`, `env`, `app` on containers:
```promql
container_memory_usage_bytes{env="prod", app="dta_service"}
```
5. Quick check from the shell on ai-server-4080:
```bash
curl -sG 'http://127.0.0.1:9090/api/v1/label/host/values' | head
curl -sG 'http://127.0.0.1:9090/api/v1/query' \
--data-urlencode 'query=count(up{job="cadvisor"})'
```
You should see all three hostnames under `host`, and `up` series for
`job="node"` and `job="cadvisor"` from each Alloy.
Dashboard build steps (container health, system status, imports) are in
[GRAFANA_USAGE.md](GRAFANA_USAGE.md) §4–§5.
---
## Part 4 — Ansible provision (implemented)
Roles are in the repo. `./scripts/provision.sh` runs `site.yml`, which:
1. Base roles on every host (`common``tianji`)
2. **`observability`** on `ai-server-4080` when `observability_stack: true`
(Loki + Prometheus + Grafana + UFW for `3100`/`9090`/`3000`)
3. **`alloy`** on every host (logs → Loki, metrics → Prometheus)
Alloy is Grafanas Promtail successor; one agent covers logs **and** metrics.
### 4.1 Layout
```
roles/
├── alloy/ # every webserver
│ ├── defaults/main.yml
│ ├── tasks/main.yml
│ └── templates/
│ ├── config.alloy.j2
│ └── docker-compose.yml.j2
└── observability/ # ai-server-4080 only
├── defaults/main.yml
├── tasks/main.yml
└── templates/
├── docker-compose.yml.j2
├── loki-config.yml.j2
├── prometheus.yml.j2
└── grafana-datasources.yml.j2
```
### 4.2 Inventory
Vars in `inventory/group_vars/all.yml` (`loki_url`, `prometheus_url`,
`grafana_public_url`, …). Flag in `inventory/host_vars/ai-server-4080.yml`:
```yaml
observability_stack: true
```
`playbooks/site.yml` order: base → observability (4080) → alloy (all).
### 4.3 Secrets
- Default Grafana password is `CHANGE_ME_ON_FIRST_LOGIN` (role default). Change
on first login, or set `observability_grafana_admin_password` in host_vars /
vault (never commit real passwords).
- SMTP2GO: see [GRAFANA_USAGE.md](GRAFANA_USAGE.md).
### 4.4 Apply
```bash
./scripts/provision.sh ai-server-4080 --check
./scripts/provision.sh ai-server-4080
./scripts/provision.sh adama
./scripts/provision.sh roslin
# or all (site.yml orders stack before Alloy):
./scripts/provision.sh
```
---
## Part 5 — Day-2 operations
### Useful LogQL starters
```logql
# Everything from one host
{host="adama"}
# One app in prod across all hosts
{app="dta_service", env="prod"}
# Errors (text match — tune per app)
{app="company_site", env="prod"} |= "ERROR"
# Rate of log lines per host (Explore → Build a dashboard)
sum by (host) (rate({job="docker"}[5m]))
```
### Useful PromQL starters
```promql
# Host CPU % busy
100 - (avg by (host) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)
# Host memory % used
100 * (1 - (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes))
# Disk % used (root)
100 - ((node_filesystem_avail_bytes{mountpoint="/",fstype!="rootfs"}
* 100) / node_filesystem_size_bytes{mountpoint="/",fstype!="rootfs"})
# Container CPU cores used
sum by (host, name, app, env) (
rate(container_cpu_usage_seconds_total{name!=""}[5m])
)
# Container memory working set
sum by (host, name, app, env) (
container_memory_working_set_bytes{name!=""}
)
# Containers that disappeared / stopped reporting (up == 0)
up{job="cadvisor"} == 0
```
### Check Alloy → Loki / Prometheus path
On a shipper host:
```bash
cd /opt/apps/observability/alloy && docker compose logs --tail=100 alloy
```
On ai-server-4080:
```bash
curl -s "http://127.0.0.1:3100/loki/api/v1/label/host/values"
curl -s "http://127.0.0.1:9090/api/v1/label/host/values"
curl -sG "http://127.0.0.1:9090/api/v1/query" \
--data-urlencode 'query=count by (host, job) (up)'
```
### Restart / update images
```bash
cd /opt/apps/observability && docker compose pull && docker compose up -d
cd /opt/apps/observability/alloy && docker compose pull && docker compose up -d
```
### Disk
Loki + Prometheus retention are both ~31 days in the sample configs. Watch:
```bash
du -sh /opt/apps/observability/loki/data \
/opt/apps/observability/prometheus/data
```
---
## Part 6 — Checklist
### Central stack (ai-server-4080)
- [ ] `/opt/apps/observability` created
- [ ] Loki + Prometheus configs + Grafana datasource provisioning in place
- [ ] `docker compose up -d` — Loki ready, Prometheus ready, Grafana login on `:3000`
- [ ] Prometheus started with `--web.enable-remote-write-receiver`
- [ ] UFW allows `3100`, `9090`, and `3000` from `10.0.0.0/24` only
- [ ] Admin password changed from default
- [ ] DNS + NPM proxy host with SSL + websockets
- [ ] `GF_SERVER_ROOT_URL` matches public URL
- [ ] Both Loki and Prometheus datasources green in Grafana
### Alloy (each of adama, roslin, ai-server-4080)
- [ ] `config.alloy` has correct `host = "..."` (or Ansible `inventory_hostname`)
- [ ] Alloy container running privileged with host `/proc` `/sys` mounts
- [ ] Can reach `10.0.0.128:3100` and `10.0.0.128:9090`
- [ ] Grafana Explore (Loki) shows that hosts log labels
- [ ] Grafana Explore (Prometheus) shows `up{job="node"}` and `up{job="cadvisor"}` for that host
- [ ] Docker apps show `env` of `beta` or `prod` when Compose project is `<app>_<env>`
### Dashboards / alerts
- [ ] Container health + system status dashboards — [GRAFANA_USAGE.md](GRAFANA_USAGE.md)
- [ ] SMTP2GO + alert rules for down containers / disk / CPU
### Follow-ups
- [x] Encode `roles/alloy` + `roles/observability` and add to `site.yml`
- [ ] Optional: drop or keep Tianji side-by-side (they do not conflict)
- [ ] NPM proxy + change Grafana admin password
- [ ] SMTP2GO + dashboards — [GRAFANA_USAGE.md](GRAFANA_USAGE.md)
---
## Troubleshooting
| Symptom | Likely cause | Fix |
|---------|--------------|-----|
| Alloy: connection refused to Loki | Loki down or UFW blocking | `docker compose ps` on 4080; `ufw status`; `curl 10.0.0.128:3100/ready` from shipper |
| Alloy: connection refused to Prometheus | Prometheus down, no remote-write flag, or UFW | Confirm `--web.enable-remote-write-receiver`; `curl 10.0.0.128:9090/-/ready`; open `9090` from LAN |
| No container metrics / empty cAdvisor | Missing privileged / mounts | Alloy compose needs `privileged: true`, `/sys`, `/var/lib/docker`, docker.sock |
| Node metrics look like containers own tiny FS | Wrong `rootfs_path` / mounts | Use `/host/proc`, `/host/sys`, `/host/root` mounts as in §3.2 |
| No `env`/`app` on Docker logs/metrics | Container not from Compose, or missing compose labels | `docker inspect <id>` for `com.docker.compose.project` |
| Wrong `host` label | Hardcoded wrong in `config.alloy` | Fix host string or use Ansible template |
| Grafana behind NPM shows blank / websocket errors | Websockets off in NPM | Enable Websockets on the Proxy Host |
| Grafana redirects to wrong host | `GF_SERVER_ROOT_URL` mismatch | Set to `https://your.domain` and recreate container |
| Disk filling on 4080 | Retention / volume growth | Lower Loki `retention_period` / Prometheus `--storage.tsdb.retention.time` |
---
## Related docs
- [GRAFANA_USAGE.md](GRAFANA_USAGE.md) — Explore, container/system dashboards, alerts, SMTP2GO
- [IMPLEMENTATION.md](../IMPLEMENTATION.md) — overall Ansible architecture
- [README.md](../README.md) — provision quick start