Files
server-infra/docs/GRAFANA_USAGE.md
2026-07-09 17:11:14 -05:00

775 lines
25 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Grafana — Using Logs, Metrics, Dashboards, and Alerts
Companion to [OBSERVABILITY.md](OBSERVABILITY.md). Assumes Alloy is shipping
**logs → Loki** and **metrics → Prometheus**, and Grafana is reachable (LAN
or your public URL).
This guide is for day-to-day use: Explore logs and metrics, build **container
health** and **system status** dashboards, and send alerts via **SMTP2GO**
(email and SMS-via-email gateways).
---
## 1. First login and basics
1. Open Grafana (`https://grafana.YOURDOMAIN` or `http://10.0.0.128:3000`).
2. Log in as `admin` (password you set during bootstrap).
3. Confirm **Connections → Data sources**:
- **Prometheus** — green “Data source is working”
- **Loki** — green “Data source is working”
### UI map (what you will use most)
| Area | Path | Use for |
|------|------|---------|
| Explore | left menu → **Explore** | Ad-hoc log search (Loki) or metric queries (Prometheus) |
| Dashboards | **Dashboards** | Container health, system status, log overviews |
| Alerting | **Alerting** | Rules, contact points, notification policies |
| Admin | **Administration** | Users, SMTP, org settings |
Pick the datasource at the top of Explore / each panel:
| Want… | Datasource |
|-------|------------|
| Log lines, error text | **Loki** |
| CPU / RAM / disk / container up | **Prometheus** |
---
## 2. Labels you can filter on
Alloy attaches these labels (see OBSERVABILITY.md). Prefer **label filters**
first; use text search (`|=`, `|~`) on logs second.
### Logs (Loki)
| Label | Meaning | Examples |
|-------|---------|----------|
| `host` | Inventory hostname | `adama`, `roslin`, `ai-server-4080` |
| `env` | Deploy environment | `beta`, `prod`, `host` (journal), `infra` (stack containers) |
| `app` | App / service name | `company_site`, `dta_service`, `dta_webapp`, `system` |
| `job` | Collector | `docker`, `systemd` |
| `unit` | systemd unit (journal) | `docker.service`, `ssh.service` |
| `container` | Docker container name | `company_site_prod-web-1` |
### Metrics (Prometheus)
| Label | Meaning | Examples |
|-------|---------|----------|
| `host` | Inventory hostname (stamped by Alloy) | `adama`, `roslin`, `ai-server-4080` |
| `job` | Scrape job | `node` (host), `cadvisor` (containers) |
| `env` / `app` | Parsed from Compose project `<app>_<env>` | `prod` / `dta_service` |
| `name` | Container name (cAdvisor) | `company_site_prod-web-1` |
| `device` / `mountpoint` | Disk (node metrics) | `sda`, `/` |
### Mental model
```text
LOGS: {host, env, app, job} ← filters
|= "ERROR" ← text search after labels
METRICS: metric_name{host, env, app, name, job} ← PromQL selectors
rate(...[5m]) / sum by (...) ← aggregations
```
---
## 3. Explore — logs (LogQL)
Open **Explore** → pick **Loki** → switch to **Code** mode (easier while learning).
### 3.1 Starter queries
```logql
# All recent logs from one host
{host="adama"}
# One app in prod on every host (active/active)
{app="dta_service", env="prod"}
# Beta only on roslin
{host="roslin", env="beta"}
# Host OS / systemd (not containers)
{job="systemd", host="ai-server-4080"}
# Docker only
{job="docker", host="adama"}
```
### 3.2 Text filters (after labels)
```logql
# Lines containing ERROR (case-sensitive)
{app="company_site", env="prod"} |= "ERROR"
# Case-insensitive
{app="dta_service", env="prod"} |~ "(?i)error|exception|traceback"
# Exclude noise
{job="docker", app="dta_webapp"} != "healthcheck" != "favicon"
```
### 3.3 Parse JSON logs (if an app logs JSON)
```logql
{app="dta_service", env="prod"}
| json
| level="ERROR"
```
Only works when the **line** is JSON. If not, stick to `|=` / `|~`.
### 3.4 Metrics from logs (for graphs / alerts)
```logql
# Log lines per second, by host
sum by (host) (rate({job="docker"}[5m]))
# Error-ish lines per minute for one app
sum by (host) (
count_over_time({app="dta_service", env="prod"} |~ "(?i)error|exception"[1m])
)
```
Tips:
- Time range picker (top right): start with **Last 15 minutes**.
- Click a log line → expand → see all labels.
- **Add to dashboard** (panel menu) once a query looks useful.
- Live tail: Explore → enable live (good while reproducing a bug).
---
## 4. Explore — metrics (PromQL)
Open **Explore** → pick **Prometheus****Code** mode.
### 4.1 Are agents reporting?
```promql
# 1 = healthy scrape path from Alloy → Prometheus
up{job="node"}
up{job="cadvisor"}
# Count series by host
count by (host, job) (up)
```
Every host should show `up == 1` for both `node` and `cadvisor`.
### 4.2 System status (per host)
```promql
# CPU % busy
100 - (avg by (host) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)
# Load (1 / 5 / 15 minute)
node_load1
node_load5
node_load15
# Memory % used
100 * (1 - (
node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes
))
# Root disk % used (path may be /host/root when Alloy uses rootfs_path)
100 - (
(node_filesystem_avail_bytes{fstype!~"tmpfs|overlay|squashfs"}
/ node_filesystem_size_bytes{fstype!~"tmpfs|overlay|squashfs"})
* 100
)
# Network receive/transmit bytes/sec (exclude lo)
sum by (host, device) (
rate(node_network_receive_bytes_total{device!="lo"}[5m])
)
sum by (host, device) (
rate(node_network_transmit_bytes_total{device!="lo"}[5m])
)
```
If disk queries return nothing, open Explore → Metrics browser → search
`node_filesystem` and check which `mountpoint` label your hosts use
(sometimes `/`, sometimes `/host/root`).
### 4.3 Container health and resource usage
cAdvisor exposes many series. Filter with `name!=""` to skip blank/cgroup roots.
```promql
# Memory working set (best "RSS-like" view) per container
sum by (host, name, app, env) (
container_memory_working_set_bytes{name!=""}
)
# CPU usage in cores (1.0 = one full core)
sum by (host, name, app, env) (
rate(container_cpu_usage_seconds_total{name!=""}[5m])
)
# Network I/O
sum by (host, name) (
rate(container_network_receive_bytes_total{name!=""}[5m])
)
sum by (host, name) (
rate(container_network_transmit_bytes_total{name!=""}[5m])
)
# One app across the fleet
sum by (host, name) (
container_memory_working_set_bytes{app="dta_service", env="prod"}
)
# Container last-seen / presence (1 while metrics flow)
up{job="cadvisor"}
```
**“Is the container healthy?”** in this stack means:
| Signal | How to read it |
|--------|----------------|
| Metrics still arriving for `name=…` | Container (or cgroup) is running and Alloy can see it |
| CPU / memory suddenly zero then missing | Container stopped or host/Alloy down |
| Memory climbing without bound | Leak / need limit |
| Logs show crash loop + memory resets | Restart loop — correlate with Loki |
Dockers own `healthcheck` status is **not** a first-class Prometheus metric
from cAdvisor. For true HEALTHCHECK state later you can add
`docker events` scraping or an exporter; for now combine **metrics presence +
resource graphs + Loki errors**.
### 4.4 Split-pane: metrics + logs for one incident
1. Explore → Prometheus → CPU/memory query for `{app="dta_service", env="prod"}`.
2. Click **Split** (top) → second pane → Loki →
`{app="dta_service", env="prod"} |~ "(?i)error|traceback"`.
3. Align time ranges — spike on left should line up with errors on right.
---
## 5. Dashboards
Create a folder **Homelab** (Dashboards → New → New folder). Build these three
first.
### 5.1 Dashboard: System status (fleet)
**Name:** `Fleet — System status`
**Datasource:** Prometheus
1. **Dashboards → New → New dashboard**.
2. Gear → **Variables**:
| Name | Type | Definition |
|------|------|------------|
| `host` | Query | Prometheus → Label values → `host` (from metric `up`) · Multi · Include All |
3. Add panels (all Prometheus, filter `host=~"$host"`):
| Panel | Type | Query |
|-------|------|-------|
| CPU % busy | Time series | `100 - (avg by (host) (rate(node_cpu_seconds_total{mode="idle", host=~"$host"}[5m])) * 100)` |
| Load 1m | Time series | `node_load1{host=~"$host"}` |
| Memory % used | Time series | `100 * (1 - (node_memory_MemAvailable_bytes{host=~"$host"} / node_memory_MemTotal_bytes{host=~"$host"}))` |
| Disk % used | Time series | See §4.2 disk query; add `host=~"$host"` |
| Network RX/TX | Time series | `sum by (host, device) (rate(node_network_receive_bytes_total{host=~"$host", device!="lo"}[5m]))` (and transmit) |
| Agent up | Stat | `up{job="node", host=~"$host"}` — thresholds: 1 green, 0 red |
4. Save.
Optional unit settings: CPU → `percent (0-100)`; memory/disk → percent;
network → `bytes/sec (SI)`.
### 5.2 Dashboard: Docker container health
**Name:** `Fleet — Docker containers`
**Datasource:** Prometheus
Variables:
| Name | Type | Definition |
|------|------|------------|
| `host` | Query | label_values(`host`) · Multi · All |
| `env` | Query | label_values(`env`) · Multi · All |
| `app` | Query | label_values(`app`) · Multi · All |
Common selector used below:
```text
host=~"$host", env=~"$env", app=~"$app", name!=""
```
| Panel | Type | Query / notes |
|-------|------|----------------|
| Containers reporting | Stat | `count(container_memory_working_set_bytes{host=~"$host", env=~"$env", app=~"$app", name!=""})` |
| CPU by container | Time series | `sum by (host, name, app, env) (rate(container_cpu_usage_seconds_total{host=~"$host", env=~"$env", app=~"$app", name!=""}[5m]))` · Legend: `{{host}} / {{name}}` · Unit: short / cores |
| Memory working set | Time series | `sum by (host, name, app, env) (container_memory_working_set_bytes{host=~"$host", env=~"$env", app=~"$app", name!=""})` · Unit: bytes (IEC) |
| Memory top table | Table | Same memory query → Transform **Sort by** → Instant query · useful “whos fattest” |
| Network RX | Time series | `sum by (host, name) (rate(container_network_receive_bytes_total{host=~"$host", env=~"$env", app=~"$app", name!=""}[5m]))` |
| Network TX | Time series | transmit counterpart |
| Prod apps only (row) | — | Set variable `env=prod` or duplicate panels with `env="prod"` hard-coded |
**Layout tip:** top row = Stat “how many containers” + Stat “hosts with cadvisor up”;
middle = CPU + Memory; bottom = Network.
### 5.3 Dashboard: Per-app (logs + metrics mixed)
**Name:** `App — dta_service` (clone per app)
Variables: `host`, `env` (default `prod`).
| Panel | DS | Query |
|-------|-----|-------|
| CPU | Prometheus | `sum by (host, name) (rate(container_cpu_usage_seconds_total{app="dta_service", env=~"$env", host=~"$host", name!=""}[5m]))` |
| Memory | Prometheus | `sum by (host, name) (container_memory_working_set_bytes{app="dta_service", env=~"$env", host=~"$host", name!=""})` |
| Error log rate | Loki | `sum by (host) (count_over_time({app="dta_service", env=~"$env", host=~"$host"} \|~ "(?i)error\|exception\|traceback"[1m]))` |
| Recent errors | Loki Logs | `{app="dta_service", env=~"$env", host=~"$host"} \|~ "(?i)error\|exception\|traceback"` |
Clone dashboard → change `app="…"` in every panel for `company_site` /
`dta_webapp`.
### 5.4 Dashboard: Fleet logs (optional)
Same as before — Loki-only overview:
1. Logs panel: `{env="prod"}`
2. Time series: `sum by (host) (rate({job="docker", env="prod"}[5m]))`
3. Errors: `{env="prod"} |~ "(?i)error|exception"`
### 5.5 Import community dashboards (fast start)
**Dashboards → New → Import** → enter ID → choose **Prometheus** datasource.
| ID | Name | Notes |
|----|------|-------|
| [1860](https://grafana.com/grafana/dashboards/1860) | Node Exporter Full | Excellent host CPU/RAM/disk/net. Our Alloy stamps `host`; you may need to edit panel queries from `instance``host`, or set the dashboards instance variable to match. |
| [14282](https://grafana.com/grafana/dashboards/14282) | Cadvisor exporter | Container CPU/mem. Filter/adjust labels to `name`, `host`, `app`, `env`. |
| [193](https://grafana.com/grafana/dashboards/193) | Docker monitoring (older) | Often needs label tweaks; prefer 14282 or the hand-built §5.2 board. |
After import:
1. Open a panel → Edit → fix label matchers (`instance``host=~"$host"`).
2. Delete panels you do not care about (keeps the board readable).
3. Save under folder **Homelab**.
Hand-built §5.1 / §5.2 boards already use your `host` / `env` / `app` labels —
prefer those if imports fight you.
### 5.6 What “healthy” looks like on the Docker dashboard
| Green | Investigate |
|-------|-------------|
| Container count stable for that env | Count dropped → container exited / host down |
| CPU flat or gently varying | CPU pegged at ~cores available → thrash / loop |
| Memory flat or sawtooth (GC) | Steady climb over hours → leak |
| Network quiet except traffic spikes | Continuous TX flood → scrape/abuse/misconfig |
| Loki error rate ~0 | Sustained errors → open split Explore |
You do **not** need a separate “monitor” product — a **dashboard panel**
visualizes; an **alert rule** watches a query and notifies you.
---
## 6. Alerting overview
Grafana Alerting has three pieces:
```text
Alert rule → what to watch (LogQL / metric query + threshold)
Contact point → how to notify (email, webhook, …)
Notification policy → which rules go to which contact points
```
Flow:
1. Configure **SMTP** (SMTP2GO) so Grafana can send email.
2. Create a **Contact point** (email, and optionally SMS-via-email).
3. Create **Alert rules** that fire on PromQL (preferred for health) or LogQL rates.
4. Route them with a **Notification policy** (default route is enough at first).
---
## 7. SMTP2GO — email (and SMS) from Grafana
SMTP2GO is an SMTP relay. Grafana sends ordinary email through it. SMS is
usually “email to carrier gateway” or SMTP2GOs own SMS product — both end up
as an email destination Grafana can use.
### 7.1 Gather SMTP2GO settings
From the SMTP2GO dashboard you need:
| Setting | Typical SMTP2GO value |
|---------|------------------------|
| Host | `mail.smtp2go.com` |
| Port | `587` (STARTTLS) or `465` (TLS) |
| Username | SMTP2GO SMTP user |
| Password | SMTP2GO SMTP password |
| From address | A verified sender in SMTP2GO (e.g. `alerts@yourdomain.com`) |
Confirm the **From** domain/sender is verified in SMTP2GO or mail will be rejected.
### 7.2 Configure SMTP in Grafana (UI)
1. **Administration → Default preferences** is not enough — use config or env.
2. Easiest on this host: set env vars in
`/opt/apps/observability/docker-compose.yml` under `grafana.environment`,
then recreate the container.
Example additions:
```yaml
environment:
GF_SECURITY_ADMIN_USER: admin
GF_SECURITY_ADMIN_PASSWORD: "…already set…"
GF_USERS_ALLOW_SIGN_UP: "false"
GF_SERVER_ROOT_URL: "https://grafana.aimloperations.com"
GF_SERVER_DOMAIN: "grafana.aimloperations.com"
# --- SMTP2GO ---
GF_SMTP_ENABLED: "true"
GF_SMTP_HOST: "mail.smtp2go.com:587"
GF_SMTP_USER: "your-smtp2go-username"
GF_SMTP_PASSWORD: "your-smtp2go-password"
GF_SMTP_FROM_ADDRESS: "alerts@yourdomain.com"
GF_SMTP_FROM_NAME: "Grafana Alerts"
GF_SMTP_STARTTLS_POLICY: "OpportunisticStartTLS"
```
**Do not commit** the password. Prefer an env file:
```bash
# /opt/apps/observability/grafana/smtp.env (mode 600, not in git)
GF_SMTP_ENABLED=true
GF_SMTP_HOST=mail.smtp2go.com:587
GF_SMTP_USER=...
GF_SMTP_PASSWORD=...
GF_SMTP_FROM_ADDRESS=alerts@yourdomain.com
GF_SMTP_FROM_NAME=Grafana Alerts
GF_SMTP_STARTTLS_POLICY=OpportunisticStartTLS
```
Wire it in compose:
```yaml
grafana:
env_file:
- ./grafana/smtp.env
environment:
GF_SECURITY_ADMIN_USER: admin
# …
```
Apply:
```bash
chmod 600 /opt/apps/observability/grafana/smtp.env
cd /opt/apps/observability && docker compose up -d grafana
```
### 7.3 Contact point — email
1. **Alerting → Contact points → Add contact point**.
2. Name: `email-ops`.
3. Integration: **Email**.
4. Addresses: your inbox (comma-separated for several people).
5. **Test** → you should receive a message via SMTP2GO.
6. Save.
### 7.4 Contact point — SMS via email gateway
Carriers expose addresses like `number@txt.att.net`. Grafana still sends
**email**; the carrier turns it into SMS.
Examples (US — confirm with your carrier):
| Carrier | Gateway pattern |
|---------|-----------------|
| AT&T | `10digit@txt.att.net` |
| T-Mobile | `10digit@tmomail.net` |
| Verizon | `10digit@vtext.com` |
1. Add another contact point, e.g. `sms-ryan`.
2. Integration: **Email**.
3. Addresses: `5551234567@tmomail.net` (your number + gateway).
4. Keep the message short — SMS truncates.
5. **Test**.
Alternatively, if you use **SMTP2GO SMS** (separate product/API), use a
Grafana **Webhook** contact point to their SMS API instead of Email. Email-to-SMS
is simpler and enough for most homelabs.
### 7.5 Notification policy
1. **Alerting → Notification policies**.
2. Default policy → set **Contact point** to `email-ops`.
3. Optional child policy:
- Matcher: `severity = critical`
- Contact point: `sms-ryan` (and/or email)
4. Save.
At small scale, routing everything to email and only critical rules to SMS is
enough.
---
## 8. Create alert rules (practical examples)
Path: **Alerting → Alert rules → New alert rule**.
Prefer **Prometheus** rules for health/resources; use **Loki** for error-text
spikes.
### 8.1 Anatomy of an alert
1. **Rule name**: e.g. `host high CPU` or `dta_service prod errors`
2. **Query**: PromQL or LogQL that returns a **number** (not a raw log stream).
3. **Reduce**: usually **Last** or **Mean**.
4. **Threshold**: e.g. `IS ABOVE 90` or `IS BELOW 1`.
5. **Evaluation**: folder + group; interval e.g. `1m`; pending period e.g. `5m`.
6. **Labels**: `severity=warning|critical` for routing.
7. **Summary**: include `{{ $labels.host }}` / `{{ $labels.name }}`.
### 8.2 Example A — Alloy / host metrics missing (critical)
**Name:** `host metrics down`
**Datasource:** Prometheus
```promql
up{job="node"}
```
**Condition:** WHEN last of A IS BELOW `1`
**For:** `5m`
**Labels:** `severity=critical`
**Summary:** `No node metrics from {{ $labels.host }} — host or Alloy down`
Repeat pattern for cAdvisor:
```promql
up{job="cadvisor"}
```
### 8.3 Example B — expected prod container missing
After you know the stable container `name` (from the Docker dashboard legend),
alert when its memory series disappears.
**Name:** `dta_service prod container missing`
**Datasource:** Prometheus
```promql
sum by (host) (
container_memory_working_set_bytes{
app="dta_service",
env="prod",
name!=""
}
)
```
**Condition:** WHEN last of A IS BELOW `1` (bytes — effectively “no series” /
near-zero; better: use **Alerting → No data** handling = **Alerting** on this
query, or per-host:
```promql
sum(
container_memory_working_set_bytes{
host="adama",
app="dta_service",
env="prod",
name!=""
}
)
```
**For:** `5m`
**Labels:** `severity=critical`, `app=dta_service`
**Summary:** `dta_service prod not reporting memory on {{ $labels.host }}`
Create one rule per host you care about until you are comfortable with
`sum by (host)` + no-data behaviour.
### 8.4 Example C — host CPU high
```promql
100 - (avg by (host) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)
```
**Condition:** IS ABOVE `90`
**For:** `10m`
**severity:** `warning`
**Summary:** `CPU {{ $values.A }}% on {{ $labels.host }}`
### 8.5 Example D — host disk almost full
```promql
100 - (
(
node_filesystem_avail_bytes{fstype!~"tmpfs|overlay|squashfs"}
/ node_filesystem_size_bytes{fstype!~"tmpfs|overlay|squashfs"}
) * 100
)
```
**Condition:** IS ABOVE `90`
**For:** `15m`
**severity:** `critical`
**Summary:** `Disk {{ $values.A }}% full on {{ $labels.host }} ({{ $labels.mountpoint }})`
Confirm `mountpoint` labels in Explore first; restrict with
`mountpoint="/"` or `mountpoint="/host/root"` if needed.
### 8.6 Example E — container memory high
```promql
sum by (host, name, app, env) (
container_memory_working_set_bytes{env="prod", name!=""}
)
```
**Condition:** IS ABOVE `1073741824` (1 GiB) — tune per app
**For:** `15m`
**severity:** `warning`
**Summary:** `{{ $labels.app }} {{ $labels.name }} using high memory on {{ $labels.host }}`
### 8.7 Example F — app error spike (Loki)
**Name:** `dta_service prod error spike`
**Datasource:** Loki
```logql
sum by (host) (
count_over_time(
{app="dta_service", env="prod"} |~ "(?i)error|exception|traceback"[5m]
)
)
```
**Condition:** WHEN last of A IS ABOVE `5`
**For:** `5m`
**Labels:** `severity=warning`, `app=dta_service`
**Summary:** `dta_service prod errors on {{ $labels.host }}`
### 8.8 Example G — SSH auth failures (Loki / journal)
```logql
sum by (host) (
count_over_time(
{job="systemd", unit="ssh.service"} |= "Failed password"[5m]
)
)
```
Threshold: `IS ABOVE 20` over `5m`, severity `warning`.
### 8.9 After saving a rule
1. Wait for evaluation (**Normal** / **Pending** / **Firing**).
2. Contact point **Test** already proved SMTP.
3. Optionally lower threshold briefly to force a fire, confirm email/SMS, restore.
---
## 9. Alert hygiene (avoid pager fatigue)
- Prefer PromQL health rules (`up`, disk, CPU) over noisy log greps.
- Always **label-filter** before text-matching on Loki.
- Use a **For** duration (`5m``15m`) so blips do not page you.
- Start with **email only**; add SMS for `severity=critical` after a week.
- Separate **beta** and **prod** rules — beta is noisier.
- Mute windows: **Alerting → Silence** during planned deploys.
- Revisit thresholds after a few days of real volume.
---
## 10. Users and access (optional)
**Administration → Users**:
- Invite a read-only viewer for others who only need Explore/Dashboards.
- Keep `admin` for you; do not reuse the admin password elsewhere.
- If Grafana is on the public internet via NPM, consider also:
- Strong admin password (required)
- Disabling sign-up (`GF_USERS_ALLOW_SIGN_UP=false` — already in compose)
- Optional: NPM Access List / Authelia later if you harden further
---
## 11. Workflow cheat sheet
### Debug a production issue
1. Open **Fleet — Docker containers** → set `app` / `env=prod` — CPU/mem weird?
2. Explore → Prometheus → same selectors for detail.
3. Split Explore → Loki → `{app="dta_service", env="prod"} |~ "(?i)error"`
4. Narrow host if active/active: add `host="adama"`
5. Optional: pin panels on an incident dashboard for the postmortem
### After deploying a new app env
1. Confirm Compose project is `<app>_<env>` so Alloy labels `app` / `env`
2. Docker dashboard → container appears with CPU/mem
3. Loki Explore → `{app="new_app", env="beta"}` shows lines
4. Clone alert rules; change `app` matcher
### Weekly health pass
1. **Fleet — System status** — CPU / RAM / disk sane on all hosts?
2. **Fleet — Docker containers** — expected prod count; no memory climbers
3. **Alerting → Alert rules** — anything lingering in Firing?
4. Disk on 4080:
```bash
du -sh /opt/apps/observability/loki/data \
/opt/apps/observability/prometheus/data
```
---
## 12. Troubleshooting alerts / mail
| Symptom | Check |
|---------|--------|
| Contact point Test fails | SMTP user/pass, From address verified in SMTP2GO, port 587 not blocked outbound |
| PromQL rule never fires | Query in Explore returns data? Threshold direction correct (`ABOVE` vs `BELOW`)? |
| Loki rule never fires | Needs a **number** — wrap with `count_over_time` / `rate` |
| Rule always firing | Threshold too low; noisy pattern; raise threshold or lengthen **For** |
| Container “missing” false alarm | Deploy renamed container; update `name`/`app` matchers |
| SMS never arrives | Gateway address wrong; carrier blocks; try email first |
| Duplicate notifications | Multiple contact points on default policy; simplify policy tree |
| No labels in alert text | Use `{{ $labels.host }}`; ensure query uses `sum by (host)` / `by (name)` |
Grafana container logs:
```bash
cd /opt/apps/observability && docker compose logs -f --tail=100 grafana
```
SMTP2GO activity log (in their web UI) shows whether the message was accepted
and delivered.
---
## 13. What this stack does *not* cover (yet)
| Need | Later addition |
|------|----------------|
| Docker HEALTHCHECK status as a metric | Small exporter / script, or Docker events → Alloy |
| Blackbox HTTP uptime (public URL checks) | Blackbox exporter, Grafana Synthetic Monitoring, or keep Tianji |
| On-call scheduling | Grafana OnCall / external PagerDuty |
| Distributed traces | Tempo + app instrumentation |
| Long-term cheap metrics at huge scale | Mimir / remote storage (not needed at 3 hosts) |
You already have host + container **resource metrics**, **presence**, and
**logs** — enough for solid health dashboards and alerts.
---
## Related docs
- [OBSERVABILITY.md](OBSERVABILITY.md) — install Alloy / Loki / Prometheus / Grafana, NPM, Ansible plan
- [IMPLEMENTATION.md](../IMPLEMENTATION.md) — server roles and deploy conventions (`<app>_<env>`)