doc plan update
This commit is contained in:
@@ -32,6 +32,8 @@ ansible adama -m ping
|
||||
|
||||
See [IMPLEMENTATION.md](IMPLEMENTATION.md) for full architecture, CI/CD plan, and phase breakdown.
|
||||
|
||||
Observability (Alloy → Loki / Prometheus → Grafana): [docs/OBSERVABILITY.md](docs/OBSERVABILITY.md) · [docs/GRAFANA_USAGE.md](docs/GRAFANA_USAGE.md)
|
||||
|
||||
## Servers
|
||||
|
||||
| Host | IP | Role |
|
||||
|
||||
@@ -0,0 +1,774 @@
|
||||
# Grafana — Using Logs, Metrics, Dashboards, and Alerts
|
||||
|
||||
Companion to [OBSERVABILITY.md](OBSERVABILITY.md). Assumes Alloy is shipping
|
||||
**logs → Loki** and **metrics → Prometheus**, and Grafana is reachable (LAN
|
||||
or your public URL).
|
||||
|
||||
This guide is for day-to-day use: Explore logs and metrics, build **container
|
||||
health** and **system status** dashboards, and send alerts via **SMTP2GO**
|
||||
(email and SMS-via-email gateways).
|
||||
|
||||
---
|
||||
|
||||
## 1. First login and basics
|
||||
|
||||
1. Open Grafana (`https://grafana.YOURDOMAIN` or `http://10.0.0.128:3000`).
|
||||
2. Log in as `admin` (password you set during bootstrap).
|
||||
3. Confirm **Connections → Data sources**:
|
||||
- **Prometheus** — green “Data source is working”
|
||||
- **Loki** — green “Data source is working”
|
||||
|
||||
### UI map (what you will use most)
|
||||
|
||||
| Area | Path | Use for |
|
||||
|------|------|---------|
|
||||
| Explore | left menu → **Explore** | Ad-hoc log search (Loki) or metric queries (Prometheus) |
|
||||
| Dashboards | **Dashboards** | Container health, system status, log overviews |
|
||||
| Alerting | **Alerting** | Rules, contact points, notification policies |
|
||||
| Admin | **Administration** | Users, SMTP, org settings |
|
||||
|
||||
Pick the datasource at the top of Explore / each panel:
|
||||
|
||||
| Want… | Datasource |
|
||||
|-------|------------|
|
||||
| Log lines, error text | **Loki** |
|
||||
| CPU / RAM / disk / container up | **Prometheus** |
|
||||
|
||||
---
|
||||
|
||||
## 2. Labels you can filter on
|
||||
|
||||
Alloy attaches these labels (see OBSERVABILITY.md). Prefer **label filters**
|
||||
first; use text search (`|=`, `|~`) on logs second.
|
||||
|
||||
### Logs (Loki)
|
||||
|
||||
| Label | Meaning | Examples |
|
||||
|-------|---------|----------|
|
||||
| `host` | Inventory hostname | `adama`, `roslin`, `ai-server-4080` |
|
||||
| `env` | Deploy environment | `beta`, `prod`, `host` (journal), `infra` (stack containers) |
|
||||
| `app` | App / service name | `company_site`, `dta_service`, `dta_webapp`, `system` |
|
||||
| `job` | Collector | `docker`, `systemd` |
|
||||
| `unit` | systemd unit (journal) | `docker.service`, `ssh.service` |
|
||||
| `container` | Docker container name | `company_site_prod-web-1` |
|
||||
|
||||
### Metrics (Prometheus)
|
||||
|
||||
| Label | Meaning | Examples |
|
||||
|-------|---------|----------|
|
||||
| `host` | Inventory hostname (stamped by Alloy) | `adama`, `roslin`, `ai-server-4080` |
|
||||
| `job` | Scrape job | `node` (host), `cadvisor` (containers) |
|
||||
| `env` / `app` | Parsed from Compose project `<app>_<env>` | `prod` / `dta_service` |
|
||||
| `name` | Container name (cAdvisor) | `company_site_prod-web-1` |
|
||||
| `device` / `mountpoint` | Disk (node metrics) | `sda`, `/` |
|
||||
|
||||
### Mental model
|
||||
|
||||
```text
|
||||
LOGS: {host, env, app, job} ← filters
|
||||
|= "ERROR" ← text search after labels
|
||||
|
||||
METRICS: metric_name{host, env, app, name, job} ← PromQL selectors
|
||||
rate(...[5m]) / sum by (...) ← aggregations
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Explore — logs (LogQL)
|
||||
|
||||
Open **Explore** → pick **Loki** → switch to **Code** mode (easier while learning).
|
||||
|
||||
### 3.1 Starter queries
|
||||
|
||||
```logql
|
||||
# All recent logs from one host
|
||||
{host="adama"}
|
||||
|
||||
# One app in prod on every host (active/active)
|
||||
{app="dta_service", env="prod"}
|
||||
|
||||
# Beta only on roslin
|
||||
{host="roslin", env="beta"}
|
||||
|
||||
# Host OS / systemd (not containers)
|
||||
{job="systemd", host="ai-server-4080"}
|
||||
|
||||
# Docker only
|
||||
{job="docker", host="adama"}
|
||||
```
|
||||
|
||||
### 3.2 Text filters (after labels)
|
||||
|
||||
```logql
|
||||
# Lines containing ERROR (case-sensitive)
|
||||
{app="company_site", env="prod"} |= "ERROR"
|
||||
|
||||
# Case-insensitive
|
||||
{app="dta_service", env="prod"} |~ "(?i)error|exception|traceback"
|
||||
|
||||
# Exclude noise
|
||||
{job="docker", app="dta_webapp"} != "healthcheck" != "favicon"
|
||||
```
|
||||
|
||||
### 3.3 Parse JSON logs (if an app logs JSON)
|
||||
|
||||
```logql
|
||||
{app="dta_service", env="prod"}
|
||||
| json
|
||||
| level="ERROR"
|
||||
```
|
||||
|
||||
Only works when the **line** is JSON. If not, stick to `|=` / `|~`.
|
||||
|
||||
### 3.4 Metrics from logs (for graphs / alerts)
|
||||
|
||||
```logql
|
||||
# Log lines per second, by host
|
||||
sum by (host) (rate({job="docker"}[5m]))
|
||||
|
||||
# Error-ish lines per minute for one app
|
||||
sum by (host) (
|
||||
count_over_time({app="dta_service", env="prod"} |~ "(?i)error|exception"[1m])
|
||||
)
|
||||
```
|
||||
|
||||
Tips:
|
||||
|
||||
- Time range picker (top right): start with **Last 15 minutes**.
|
||||
- Click a log line → expand → see all labels.
|
||||
- **Add to dashboard** (panel menu) once a query looks useful.
|
||||
- Live tail: Explore → enable live (good while reproducing a bug).
|
||||
|
||||
---
|
||||
|
||||
## 4. Explore — metrics (PromQL)
|
||||
|
||||
Open **Explore** → pick **Prometheus** → **Code** mode.
|
||||
|
||||
### 4.1 Are agents reporting?
|
||||
|
||||
```promql
|
||||
# 1 = healthy scrape path from Alloy → Prometheus
|
||||
up{job="node"}
|
||||
up{job="cadvisor"}
|
||||
|
||||
# Count series by host
|
||||
count by (host, job) (up)
|
||||
```
|
||||
|
||||
Every host should show `up == 1` for both `node` and `cadvisor`.
|
||||
|
||||
### 4.2 System status (per host)
|
||||
|
||||
```promql
|
||||
# CPU % busy
|
||||
100 - (avg by (host) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)
|
||||
|
||||
# Load (1 / 5 / 15 minute)
|
||||
node_load1
|
||||
node_load5
|
||||
node_load15
|
||||
|
||||
# Memory % used
|
||||
100 * (1 - (
|
||||
node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes
|
||||
))
|
||||
|
||||
# Root disk % used (path may be /host/root when Alloy uses rootfs_path)
|
||||
100 - (
|
||||
(node_filesystem_avail_bytes{fstype!~"tmpfs|overlay|squashfs"}
|
||||
/ node_filesystem_size_bytes{fstype!~"tmpfs|overlay|squashfs"})
|
||||
* 100
|
||||
)
|
||||
|
||||
# Network receive/transmit bytes/sec (exclude lo)
|
||||
sum by (host, device) (
|
||||
rate(node_network_receive_bytes_total{device!="lo"}[5m])
|
||||
)
|
||||
sum by (host, device) (
|
||||
rate(node_network_transmit_bytes_total{device!="lo"}[5m])
|
||||
)
|
||||
```
|
||||
|
||||
If disk queries return nothing, open Explore → Metrics browser → search
|
||||
`node_filesystem` and check which `mountpoint` label your hosts use
|
||||
(sometimes `/`, sometimes `/host/root`).
|
||||
|
||||
### 4.3 Container health and resource usage
|
||||
|
||||
cAdvisor exposes many series. Filter with `name!=""` to skip blank/cgroup roots.
|
||||
|
||||
```promql
|
||||
# Memory working set (best "RSS-like" view) per container
|
||||
sum by (host, name, app, env) (
|
||||
container_memory_working_set_bytes{name!=""}
|
||||
)
|
||||
|
||||
# CPU usage in cores (1.0 = one full core)
|
||||
sum by (host, name, app, env) (
|
||||
rate(container_cpu_usage_seconds_total{name!=""}[5m])
|
||||
)
|
||||
|
||||
# Network I/O
|
||||
sum by (host, name) (
|
||||
rate(container_network_receive_bytes_total{name!=""}[5m])
|
||||
)
|
||||
sum by (host, name) (
|
||||
rate(container_network_transmit_bytes_total{name!=""}[5m])
|
||||
)
|
||||
|
||||
# One app across the fleet
|
||||
sum by (host, name) (
|
||||
container_memory_working_set_bytes{app="dta_service", env="prod"}
|
||||
)
|
||||
|
||||
# Container last-seen / presence (1 while metrics flow)
|
||||
up{job="cadvisor"}
|
||||
```
|
||||
|
||||
**“Is the container healthy?”** in this stack means:
|
||||
|
||||
| Signal | How to read it |
|
||||
|--------|----------------|
|
||||
| Metrics still arriving for `name=…` | Container (or cgroup) is running and Alloy can see it |
|
||||
| CPU / memory suddenly zero then missing | Container stopped or host/Alloy down |
|
||||
| Memory climbing without bound | Leak / need limit |
|
||||
| Logs show crash loop + memory resets | Restart loop — correlate with Loki |
|
||||
|
||||
Docker’s own `healthcheck` status is **not** a first-class Prometheus metric
|
||||
from cAdvisor. For true HEALTHCHECK state later you can add
|
||||
`docker events` scraping or an exporter; for now combine **metrics presence +
|
||||
resource graphs + Loki errors**.
|
||||
|
||||
### 4.4 Split-pane: metrics + logs for one incident
|
||||
|
||||
1. Explore → Prometheus → CPU/memory query for `{app="dta_service", env="prod"}`.
|
||||
2. Click **Split** (top) → second pane → Loki →
|
||||
`{app="dta_service", env="prod"} |~ "(?i)error|traceback"`.
|
||||
3. Align time ranges — spike on left should line up with errors on right.
|
||||
|
||||
---
|
||||
|
||||
## 5. Dashboards
|
||||
|
||||
Create a folder **Homelab** (Dashboards → New → New folder). Build these three
|
||||
first.
|
||||
|
||||
### 5.1 Dashboard: System status (fleet)
|
||||
|
||||
**Name:** `Fleet — System status`
|
||||
**Datasource:** Prometheus
|
||||
|
||||
1. **Dashboards → New → New dashboard**.
|
||||
2. Gear → **Variables**:
|
||||
|
||||
| Name | Type | Definition |
|
||||
|------|------|------------|
|
||||
| `host` | Query | Prometheus → Label values → `host` (from metric `up`) · Multi · Include All |
|
||||
|
||||
3. Add panels (all Prometheus, filter `host=~"$host"`):
|
||||
|
||||
| Panel | Type | Query |
|
||||
|-------|------|-------|
|
||||
| CPU % busy | Time series | `100 - (avg by (host) (rate(node_cpu_seconds_total{mode="idle", host=~"$host"}[5m])) * 100)` |
|
||||
| Load 1m | Time series | `node_load1{host=~"$host"}` |
|
||||
| Memory % used | Time series | `100 * (1 - (node_memory_MemAvailable_bytes{host=~"$host"} / node_memory_MemTotal_bytes{host=~"$host"}))` |
|
||||
| Disk % used | Time series | See §4.2 disk query; add `host=~"$host"` |
|
||||
| Network RX/TX | Time series | `sum by (host, device) (rate(node_network_receive_bytes_total{host=~"$host", device!="lo"}[5m]))` (and transmit) |
|
||||
| Agent up | Stat | `up{job="node", host=~"$host"}` — thresholds: 1 green, 0 red |
|
||||
|
||||
4. Save.
|
||||
|
||||
Optional unit settings: CPU → `percent (0-100)`; memory/disk → percent;
|
||||
network → `bytes/sec (SI)`.
|
||||
|
||||
### 5.2 Dashboard: Docker container health
|
||||
|
||||
**Name:** `Fleet — Docker containers`
|
||||
**Datasource:** Prometheus
|
||||
|
||||
Variables:
|
||||
|
||||
| Name | Type | Definition |
|
||||
|------|------|------------|
|
||||
| `host` | Query | label_values(`host`) · Multi · All |
|
||||
| `env` | Query | label_values(`env`) · Multi · All |
|
||||
| `app` | Query | label_values(`app`) · Multi · All |
|
||||
|
||||
Common selector used below:
|
||||
|
||||
```text
|
||||
host=~"$host", env=~"$env", app=~"$app", name!=""
|
||||
```
|
||||
|
||||
| Panel | Type | Query / notes |
|
||||
|-------|------|----------------|
|
||||
| Containers reporting | Stat | `count(container_memory_working_set_bytes{host=~"$host", env=~"$env", app=~"$app", name!=""})` |
|
||||
| CPU by container | Time series | `sum by (host, name, app, env) (rate(container_cpu_usage_seconds_total{host=~"$host", env=~"$env", app=~"$app", name!=""}[5m]))` · Legend: `{{host}} / {{name}}` · Unit: short / cores |
|
||||
| Memory working set | Time series | `sum by (host, name, app, env) (container_memory_working_set_bytes{host=~"$host", env=~"$env", app=~"$app", name!=""})` · Unit: bytes (IEC) |
|
||||
| Memory top table | Table | Same memory query → Transform **Sort by** → Instant query · useful “who’s fattest” |
|
||||
| Network RX | Time series | `sum by (host, name) (rate(container_network_receive_bytes_total{host=~"$host", env=~"$env", app=~"$app", name!=""}[5m]))` |
|
||||
| Network TX | Time series | transmit counterpart |
|
||||
| Prod apps only (row) | — | Set variable `env=prod` or duplicate panels with `env="prod"` hard-coded |
|
||||
|
||||
**Layout tip:** top row = Stat “how many containers” + Stat “hosts with cadvisor up”;
|
||||
middle = CPU + Memory; bottom = Network.
|
||||
|
||||
### 5.3 Dashboard: Per-app (logs + metrics mixed)
|
||||
|
||||
**Name:** `App — dta_service` (clone per app)
|
||||
|
||||
Variables: `host`, `env` (default `prod`).
|
||||
|
||||
| Panel | DS | Query |
|
||||
|-------|-----|-------|
|
||||
| CPU | Prometheus | `sum by (host, name) (rate(container_cpu_usage_seconds_total{app="dta_service", env=~"$env", host=~"$host", name!=""}[5m]))` |
|
||||
| Memory | Prometheus | `sum by (host, name) (container_memory_working_set_bytes{app="dta_service", env=~"$env", host=~"$host", name!=""})` |
|
||||
| Error log rate | Loki | `sum by (host) (count_over_time({app="dta_service", env=~"$env", host=~"$host"} \|~ "(?i)error\|exception\|traceback"[1m]))` |
|
||||
| Recent errors | Loki Logs | `{app="dta_service", env=~"$env", host=~"$host"} \|~ "(?i)error\|exception\|traceback"` |
|
||||
|
||||
Clone dashboard → change `app="…"` in every panel for `company_site` /
|
||||
`dta_webapp`.
|
||||
|
||||
### 5.4 Dashboard: Fleet logs (optional)
|
||||
|
||||
Same as before — Loki-only overview:
|
||||
|
||||
1. Logs panel: `{env="prod"}`
|
||||
2. Time series: `sum by (host) (rate({job="docker", env="prod"}[5m]))`
|
||||
3. Errors: `{env="prod"} |~ "(?i)error|exception"`
|
||||
|
||||
### 5.5 Import community dashboards (fast start)
|
||||
|
||||
**Dashboards → New → Import** → enter ID → choose **Prometheus** datasource.
|
||||
|
||||
| ID | Name | Notes |
|
||||
|----|------|-------|
|
||||
| [1860](https://grafana.com/grafana/dashboards/1860) | Node Exporter Full | Excellent host CPU/RAM/disk/net. Our Alloy stamps `host`; you may need to edit panel queries from `instance` → `host`, or set the dashboard’s instance variable to match. |
|
||||
| [14282](https://grafana.com/grafana/dashboards/14282) | Cadvisor exporter | Container CPU/mem. Filter/adjust labels to `name`, `host`, `app`, `env`. |
|
||||
| [193](https://grafana.com/grafana/dashboards/193) | Docker monitoring (older) | Often needs label tweaks; prefer 14282 or the hand-built §5.2 board. |
|
||||
|
||||
After import:
|
||||
|
||||
1. Open a panel → Edit → fix label matchers (`instance` → `host=~"$host"`).
|
||||
2. Delete panels you do not care about (keeps the board readable).
|
||||
3. Save under folder **Homelab**.
|
||||
|
||||
Hand-built §5.1 / §5.2 boards already use your `host` / `env` / `app` labels —
|
||||
prefer those if imports fight you.
|
||||
|
||||
### 5.6 What “healthy” looks like on the Docker dashboard
|
||||
|
||||
| Green | Investigate |
|
||||
|-------|-------------|
|
||||
| Container count stable for that env | Count dropped → container exited / host down |
|
||||
| CPU flat or gently varying | CPU pegged at ~cores available → thrash / loop |
|
||||
| Memory flat or sawtooth (GC) | Steady climb over hours → leak |
|
||||
| Network quiet except traffic spikes | Continuous TX flood → scrape/abuse/misconfig |
|
||||
| Loki error rate ~0 | Sustained errors → open split Explore |
|
||||
|
||||
You do **not** need a separate “monitor” product — a **dashboard panel**
|
||||
visualizes; an **alert rule** watches a query and notifies you.
|
||||
|
||||
---
|
||||
|
||||
## 6. Alerting overview
|
||||
|
||||
Grafana Alerting has three pieces:
|
||||
|
||||
```text
|
||||
Alert rule → what to watch (LogQL / metric query + threshold)
|
||||
Contact point → how to notify (email, webhook, …)
|
||||
Notification policy → which rules go to which contact points
|
||||
```
|
||||
|
||||
Flow:
|
||||
|
||||
1. Configure **SMTP** (SMTP2GO) so Grafana can send email.
|
||||
2. Create a **Contact point** (email, and optionally SMS-via-email).
|
||||
3. Create **Alert rules** that fire on PromQL (preferred for health) or LogQL rates.
|
||||
4. Route them with a **Notification policy** (default route is enough at first).
|
||||
|
||||
---
|
||||
|
||||
## 7. SMTP2GO — email (and SMS) from Grafana
|
||||
|
||||
SMTP2GO is an SMTP relay. Grafana sends ordinary email through it. SMS is
|
||||
usually “email to carrier gateway” or SMTP2GO’s own SMS product — both end up
|
||||
as an email destination Grafana can use.
|
||||
|
||||
### 7.1 Gather SMTP2GO settings
|
||||
|
||||
From the SMTP2GO dashboard you need:
|
||||
|
||||
| Setting | Typical SMTP2GO value |
|
||||
|---------|------------------------|
|
||||
| Host | `mail.smtp2go.com` |
|
||||
| Port | `587` (STARTTLS) or `465` (TLS) |
|
||||
| Username | SMTP2GO SMTP user |
|
||||
| Password | SMTP2GO SMTP password |
|
||||
| From address | A verified sender in SMTP2GO (e.g. `alerts@yourdomain.com`) |
|
||||
|
||||
Confirm the **From** domain/sender is verified in SMTP2GO or mail will be rejected.
|
||||
|
||||
### 7.2 Configure SMTP in Grafana (UI)
|
||||
|
||||
1. **Administration → Default preferences** is not enough — use config or env.
|
||||
2. Easiest on this host: set env vars in
|
||||
`/opt/apps/observability/docker-compose.yml` under `grafana.environment`,
|
||||
then recreate the container.
|
||||
|
||||
Example additions:
|
||||
|
||||
```yaml
|
||||
environment:
|
||||
GF_SECURITY_ADMIN_USER: admin
|
||||
GF_SECURITY_ADMIN_PASSWORD: "…already set…"
|
||||
GF_USERS_ALLOW_SIGN_UP: "false"
|
||||
GF_SERVER_ROOT_URL: "https://grafana.aimloperations.com"
|
||||
GF_SERVER_DOMAIN: "grafana.aimloperations.com"
|
||||
|
||||
# --- SMTP2GO ---
|
||||
GF_SMTP_ENABLED: "true"
|
||||
GF_SMTP_HOST: "mail.smtp2go.com:587"
|
||||
GF_SMTP_USER: "your-smtp2go-username"
|
||||
GF_SMTP_PASSWORD: "your-smtp2go-password"
|
||||
GF_SMTP_FROM_ADDRESS: "alerts@yourdomain.com"
|
||||
GF_SMTP_FROM_NAME: "Grafana Alerts"
|
||||
GF_SMTP_STARTTLS_POLICY: "OpportunisticStartTLS"
|
||||
```
|
||||
|
||||
**Do not commit** the password. Prefer an env file:
|
||||
|
||||
```bash
|
||||
# /opt/apps/observability/grafana/smtp.env (mode 600, not in git)
|
||||
GF_SMTP_ENABLED=true
|
||||
GF_SMTP_HOST=mail.smtp2go.com:587
|
||||
GF_SMTP_USER=...
|
||||
GF_SMTP_PASSWORD=...
|
||||
GF_SMTP_FROM_ADDRESS=alerts@yourdomain.com
|
||||
GF_SMTP_FROM_NAME=Grafana Alerts
|
||||
GF_SMTP_STARTTLS_POLICY=OpportunisticStartTLS
|
||||
```
|
||||
|
||||
Wire it in compose:
|
||||
|
||||
```yaml
|
||||
grafana:
|
||||
env_file:
|
||||
- ./grafana/smtp.env
|
||||
environment:
|
||||
GF_SECURITY_ADMIN_USER: admin
|
||||
# …
|
||||
```
|
||||
|
||||
Apply:
|
||||
|
||||
```bash
|
||||
chmod 600 /opt/apps/observability/grafana/smtp.env
|
||||
cd /opt/apps/observability && docker compose up -d grafana
|
||||
```
|
||||
|
||||
### 7.3 Contact point — email
|
||||
|
||||
1. **Alerting → Contact points → Add contact point**.
|
||||
2. Name: `email-ops`.
|
||||
3. Integration: **Email**.
|
||||
4. Addresses: your inbox (comma-separated for several people).
|
||||
5. **Test** → you should receive a message via SMTP2GO.
|
||||
6. Save.
|
||||
|
||||
### 7.4 Contact point — SMS via email gateway
|
||||
|
||||
Carriers expose addresses like `number@txt.att.net`. Grafana still sends
|
||||
**email**; the carrier turns it into SMS.
|
||||
|
||||
Examples (US — confirm with your carrier):
|
||||
|
||||
| Carrier | Gateway pattern |
|
||||
|---------|-----------------|
|
||||
| AT&T | `10digit@txt.att.net` |
|
||||
| T-Mobile | `10digit@tmomail.net` |
|
||||
| Verizon | `10digit@vtext.com` |
|
||||
|
||||
1. Add another contact point, e.g. `sms-ryan`.
|
||||
2. Integration: **Email**.
|
||||
3. Addresses: `5551234567@tmomail.net` (your number + gateway).
|
||||
4. Keep the message short — SMS truncates.
|
||||
5. **Test**.
|
||||
|
||||
Alternatively, if you use **SMTP2GO SMS** (separate product/API), use a
|
||||
Grafana **Webhook** contact point to their SMS API instead of Email. Email-to-SMS
|
||||
is simpler and enough for most homelabs.
|
||||
|
||||
### 7.5 Notification policy
|
||||
|
||||
1. **Alerting → Notification policies**.
|
||||
2. Default policy → set **Contact point** to `email-ops`.
|
||||
3. Optional child policy:
|
||||
- Matcher: `severity = critical`
|
||||
- Contact point: `sms-ryan` (and/or email)
|
||||
4. Save.
|
||||
|
||||
At small scale, routing everything to email and only critical rules to SMS is
|
||||
enough.
|
||||
|
||||
---
|
||||
|
||||
## 8. Create alert rules (practical examples)
|
||||
|
||||
Path: **Alerting → Alert rules → New alert rule**.
|
||||
|
||||
Prefer **Prometheus** rules for health/resources; use **Loki** for error-text
|
||||
spikes.
|
||||
|
||||
### 8.1 Anatomy of an alert
|
||||
|
||||
1. **Rule name**: e.g. `host high CPU` or `dta_service prod errors`
|
||||
2. **Query**: PromQL or LogQL that returns a **number** (not a raw log stream).
|
||||
3. **Reduce**: usually **Last** or **Mean**.
|
||||
4. **Threshold**: e.g. `IS ABOVE 90` or `IS BELOW 1`.
|
||||
5. **Evaluation**: folder + group; interval e.g. `1m`; pending period e.g. `5m`.
|
||||
6. **Labels**: `severity=warning|critical` for routing.
|
||||
7. **Summary**: include `{{ $labels.host }}` / `{{ $labels.name }}`.
|
||||
|
||||
### 8.2 Example A — Alloy / host metrics missing (critical)
|
||||
|
||||
**Name:** `host metrics down`
|
||||
**Datasource:** Prometheus
|
||||
|
||||
```promql
|
||||
up{job="node"}
|
||||
```
|
||||
|
||||
**Condition:** WHEN last of A IS BELOW `1`
|
||||
**For:** `5m`
|
||||
**Labels:** `severity=critical`
|
||||
**Summary:** `No node metrics from {{ $labels.host }} — host or Alloy down`
|
||||
|
||||
Repeat pattern for cAdvisor:
|
||||
|
||||
```promql
|
||||
up{job="cadvisor"}
|
||||
```
|
||||
|
||||
### 8.3 Example B — expected prod container missing
|
||||
|
||||
After you know the stable container `name` (from the Docker dashboard legend),
|
||||
alert when its memory series disappears.
|
||||
|
||||
**Name:** `dta_service prod container missing`
|
||||
**Datasource:** Prometheus
|
||||
|
||||
```promql
|
||||
sum by (host) (
|
||||
container_memory_working_set_bytes{
|
||||
app="dta_service",
|
||||
env="prod",
|
||||
name!=""
|
||||
}
|
||||
)
|
||||
```
|
||||
|
||||
**Condition:** WHEN last of A IS BELOW `1` (bytes — effectively “no series” /
|
||||
near-zero; better: use **Alerting → No data** handling = **Alerting** on this
|
||||
query, or per-host:
|
||||
|
||||
```promql
|
||||
sum(
|
||||
container_memory_working_set_bytes{
|
||||
host="adama",
|
||||
app="dta_service",
|
||||
env="prod",
|
||||
name!=""
|
||||
}
|
||||
)
|
||||
```
|
||||
|
||||
**For:** `5m`
|
||||
**Labels:** `severity=critical`, `app=dta_service`
|
||||
**Summary:** `dta_service prod not reporting memory on {{ $labels.host }}`
|
||||
|
||||
Create one rule per host you care about until you are comfortable with
|
||||
`sum by (host)` + no-data behaviour.
|
||||
|
||||
### 8.4 Example C — host CPU high
|
||||
|
||||
```promql
|
||||
100 - (avg by (host) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)
|
||||
```
|
||||
|
||||
**Condition:** IS ABOVE `90`
|
||||
**For:** `10m`
|
||||
**severity:** `warning`
|
||||
**Summary:** `CPU {{ $values.A }}% on {{ $labels.host }}`
|
||||
|
||||
### 8.5 Example D — host disk almost full
|
||||
|
||||
```promql
|
||||
100 - (
|
||||
(
|
||||
node_filesystem_avail_bytes{fstype!~"tmpfs|overlay|squashfs"}
|
||||
/ node_filesystem_size_bytes{fstype!~"tmpfs|overlay|squashfs"}
|
||||
) * 100
|
||||
)
|
||||
```
|
||||
|
||||
**Condition:** IS ABOVE `90`
|
||||
**For:** `15m`
|
||||
**severity:** `critical`
|
||||
**Summary:** `Disk {{ $values.A }}% full on {{ $labels.host }} ({{ $labels.mountpoint }})`
|
||||
|
||||
Confirm `mountpoint` labels in Explore first; restrict with
|
||||
`mountpoint="/"` or `mountpoint="/host/root"` if needed.
|
||||
|
||||
### 8.6 Example E — container memory high
|
||||
|
||||
```promql
|
||||
sum by (host, name, app, env) (
|
||||
container_memory_working_set_bytes{env="prod", name!=""}
|
||||
)
|
||||
```
|
||||
|
||||
**Condition:** IS ABOVE `1073741824` (1 GiB) — tune per app
|
||||
**For:** `15m`
|
||||
**severity:** `warning`
|
||||
**Summary:** `{{ $labels.app }} {{ $labels.name }} using high memory on {{ $labels.host }}`
|
||||
|
||||
### 8.7 Example F — app error spike (Loki)
|
||||
|
||||
**Name:** `dta_service prod error spike`
|
||||
**Datasource:** Loki
|
||||
|
||||
```logql
|
||||
sum by (host) (
|
||||
count_over_time(
|
||||
{app="dta_service", env="prod"} |~ "(?i)error|exception|traceback"[5m]
|
||||
)
|
||||
)
|
||||
```
|
||||
|
||||
**Condition:** WHEN last of A IS ABOVE `5`
|
||||
**For:** `5m`
|
||||
**Labels:** `severity=warning`, `app=dta_service`
|
||||
**Summary:** `dta_service prod errors on {{ $labels.host }}`
|
||||
|
||||
### 8.8 Example G — SSH auth failures (Loki / journal)
|
||||
|
||||
```logql
|
||||
sum by (host) (
|
||||
count_over_time(
|
||||
{job="systemd", unit="ssh.service"} |= "Failed password"[5m]
|
||||
)
|
||||
)
|
||||
```
|
||||
|
||||
Threshold: `IS ABOVE 20` over `5m`, severity `warning`.
|
||||
|
||||
### 8.9 After saving a rule
|
||||
|
||||
1. Wait for evaluation (**Normal** / **Pending** / **Firing**).
|
||||
2. Contact point **Test** already proved SMTP.
|
||||
3. Optionally lower threshold briefly to force a fire, confirm email/SMS, restore.
|
||||
|
||||
---
|
||||
|
||||
## 9. Alert hygiene (avoid pager fatigue)
|
||||
|
||||
- Prefer PromQL health rules (`up`, disk, CPU) over noisy log greps.
|
||||
- Always **label-filter** before text-matching on Loki.
|
||||
- Use a **For** duration (`5m`–`15m`) so blips do not page you.
|
||||
- Start with **email only**; add SMS for `severity=critical` after a week.
|
||||
- Separate **beta** and **prod** rules — beta is noisier.
|
||||
- Mute windows: **Alerting → Silence** during planned deploys.
|
||||
- Revisit thresholds after a few days of real volume.
|
||||
|
||||
---
|
||||
|
||||
## 10. Users and access (optional)
|
||||
|
||||
**Administration → Users**:
|
||||
|
||||
- Invite a read-only viewer for others who only need Explore/Dashboards.
|
||||
- Keep `admin` for you; do not reuse the admin password elsewhere.
|
||||
- If Grafana is on the public internet via NPM, consider also:
|
||||
- Strong admin password (required)
|
||||
- Disabling sign-up (`GF_USERS_ALLOW_SIGN_UP=false` — already in compose)
|
||||
- Optional: NPM Access List / Authelia later if you harden further
|
||||
|
||||
---
|
||||
|
||||
## 11. Workflow cheat sheet
|
||||
|
||||
### Debug a production issue
|
||||
|
||||
1. Open **Fleet — Docker containers** → set `app` / `env=prod` — CPU/mem weird?
|
||||
2. Explore → Prometheus → same selectors for detail.
|
||||
3. Split Explore → Loki → `{app="dta_service", env="prod"} |~ "(?i)error"`
|
||||
4. Narrow host if active/active: add `host="adama"`
|
||||
5. Optional: pin panels on an incident dashboard for the postmortem
|
||||
|
||||
### After deploying a new app env
|
||||
|
||||
1. Confirm Compose project is `<app>_<env>` so Alloy labels `app` / `env`
|
||||
2. Docker dashboard → container appears with CPU/mem
|
||||
3. Loki Explore → `{app="new_app", env="beta"}` shows lines
|
||||
4. Clone alert rules; change `app` matcher
|
||||
|
||||
### Weekly health pass
|
||||
|
||||
1. **Fleet — System status** — CPU / RAM / disk sane on all hosts?
|
||||
2. **Fleet — Docker containers** — expected prod count; no memory climbers
|
||||
3. **Alerting → Alert rules** — anything lingering in Firing?
|
||||
4. Disk on 4080:
|
||||
|
||||
```bash
|
||||
du -sh /opt/apps/observability/loki/data \
|
||||
/opt/apps/observability/prometheus/data
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 12. Troubleshooting alerts / mail
|
||||
|
||||
| Symptom | Check |
|
||||
|---------|--------|
|
||||
| Contact point Test fails | SMTP user/pass, From address verified in SMTP2GO, port 587 not blocked outbound |
|
||||
| PromQL rule never fires | Query in Explore returns data? Threshold direction correct (`ABOVE` vs `BELOW`)? |
|
||||
| Loki rule never fires | Needs a **number** — wrap with `count_over_time` / `rate` |
|
||||
| Rule always firing | Threshold too low; noisy pattern; raise threshold or lengthen **For** |
|
||||
| Container “missing” false alarm | Deploy renamed container; update `name`/`app` matchers |
|
||||
| SMS never arrives | Gateway address wrong; carrier blocks; try email first |
|
||||
| Duplicate notifications | Multiple contact points on default policy; simplify policy tree |
|
||||
| No labels in alert text | Use `{{ $labels.host }}`; ensure query uses `sum by (host)` / `by (name)` |
|
||||
|
||||
Grafana container logs:
|
||||
|
||||
```bash
|
||||
cd /opt/apps/observability && docker compose logs -f --tail=100 grafana
|
||||
```
|
||||
|
||||
SMTP2GO activity log (in their web UI) shows whether the message was accepted
|
||||
and delivered.
|
||||
|
||||
---
|
||||
|
||||
## 13. What this stack does *not* cover (yet)
|
||||
|
||||
| Need | Later addition |
|
||||
|------|----------------|
|
||||
| Docker HEALTHCHECK status as a metric | Small exporter / script, or Docker events → Alloy |
|
||||
| Blackbox HTTP uptime (public URL checks) | Blackbox exporter, Grafana Synthetic Monitoring, or keep Tianji |
|
||||
| On-call scheduling | Grafana OnCall / external PagerDuty |
|
||||
| Distributed traces | Tempo + app instrumentation |
|
||||
| Long-term cheap metrics at huge scale | Mimir / remote storage (not needed at 3 hosts) |
|
||||
|
||||
You already have host + container **resource metrics**, **presence**, and
|
||||
**logs** — enough for solid health dashboards and alerts.
|
||||
|
||||
---
|
||||
|
||||
## Related docs
|
||||
|
||||
- [OBSERVABILITY.md](OBSERVABILITY.md) — install Alloy / Loki / Prometheus / Grafana, NPM, Ansible plan
|
||||
- [IMPLEMENTATION.md](../IMPLEMENTATION.md) — server roles and deploy conventions (`<app>_<env>`)
|
||||
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user