2e3fe13b391560f7db8e910320fc2d368fda52d7
Sync runner checkout / sync (push) Successful in 7s
## Summary The django deploy sets `COMPOSE_PROFILES` on the **start** step but not on the **build** step, so `docker compose build` skips profile-gated services. `up -d` then reuses whatever image already exists and the container silently keeps running old code. This adds `COMPOSE_PROFILES` to the build step so it matches the start step. Only affects hosts that set `host_apps.compose_profiles` (today: the `monica_site` dj-queue worker on adama). ## Symptom this fixes On adama the beta worker image was 20 hours stale while web was current: | Image | Built | Postgres driver | |---|---|---| | `monica_site_beta-web` | today | psycopg 3.3.4 | | `monica_site_beta-worker` | Aug 8 | psycopg2 2.9.12 | So the worker crash-looped on LISTEN/NOTIFY (`TypeError: 'list' object is not callable` in `dj_queue/runtime/notify.py`) long after `monica_site` had moved to psycopg3, because its image was never rebuilt. Branch is merged up with `master`, which already carries the worker auto-start from [#18](#18); the diff here is just the build step. ## Test plan - [x] Manual `COMPOSE_PROFILES=worker docker compose build worker` on adama produced an image with psycopg 3.3.4 and the notify errors stopped. - [ ] Beta deploy from this branch recreates the worker with a fresh image, no manual rebuild. - [ ] Hosts without `compose_profiles` (roslin, ai-server-4080) still build/start web only.Reviewed-on: #19
server-infra
Ansible provisioning and deployment for homelab web servers.
Quick Start
# Install collections (once)
ansible-galaxy collection install -r requirements.yml
# Bootstrap SSH key to each host (one-time, before Ansible)
ssh-copy-id westfarn@10.0.0.77
ssh-copy-id westfarn@10.0.0.176
# First-time only: passwordless sudo on each new host (before first provision)
ssh -t westfarn@10.0.0.176 # repeat for each host IP
# on the host:
echo 'westfarn ALL=(ALL) NOPASSWD:ALL' | sudo tee /etc/sudoers.d/westfarn
sudo chmod 440 /etc/sudoers.d/westfarn
exit
# Test connectivity to one host
ansible adama -m ping
# Provision one host (dry run first) — includes Alloy log/metrics agent
./scripts/provision.sh adama --check
./scripts/provision.sh adama
# Control node: Alloy + Loki + Prometheus + Grafana
./scripts/provision.sh ai-server-4080
# Provision all hosts
./scripts/provision.sh
See IMPLEMENTATION.md for full architecture, CI/CD plan, and phase breakdown.
Observability (Alloy → Loki / Prometheus → Grafana): docs/OBSERVABILITY.md · docs/GRAFANA_USAGE.md
Provision installs Alloy on every host. On ai-server-4080 it also starts
the central Loki / Prometheus / Grafana stack (observability_stack: true).
Servers
| Host | IP | Role |
|---|---|---|
| adama | 10.0.0.77 | app host |
| roslin | 10.0.0.176 | app host |
| ai-server-4080 | 10.0.0.128 | control node + act runner |
Languages
Jinja
63.3%
Shell
36.7%