Files
server-infra/IMPLEMENTATION.md
westfarn 7e48c22e3e
Sync runner checkout / sync (push) Successful in 7s
Deploy SearxNG on ai-server-4080 for chat_backend grounded search (#10) (#11)
## Summary
- Closes [#10](#10).
- Deploys **SearxNG** on **ai-server-4080** (`10.0.0.128`) for [chat_backend#62](ai_ml_operations/chat_backend#62) / [PR #65](ai_ml_operations/chat_backend#65) grounded search.
- New `roles/searxng/` (compose + JSON-enabled `settings.yml`), gated by `searxng_stack: true`, wired into `site.yml`.
- **Host port 8088** (not 8080 — that is `dta_webapp` on this host). UFW allows `10.0.0.0/24` → `8088/tcp` only.

## Ops after merge

```bash
./scripts/provision.sh ai-server-4080
# or targeted:
ansible-playbook playbooks/site.yml --limit ai-server-4080 --tags never  # full site play includes searxng when searxng_stack
```

Then set in `chat_backend_prod.env` / `chat_backend_beta.env`:

```text
SEARCH_PROVIDER=searxng
SEARCH_FAILOVER_PROVIDER=ddgs
SEARXNG_BASE_URL=http://10.0.0.128:8088
```

Smoke test from any app host:

```bash
curl -sG 'http://10.0.0.128:8088/search' --data-urlencode 'q=test' -d 'format=json' | head
```

## Test plan
- [ ] Provision ai-server-4080; confirm `docker ps` shows `searxng`
- [ ] Confirm `:8088` responds with JSON; `:8080` still serves dta_webapp
- [ ] Confirm UFW rule is LAN-only
- [ ] From adama/roslin container network, curl SearxNG succeeds
- [ ] Update chat_backend secrets to `:8088` and redeploy betaReviewed-on: #11
2026-08-02 11:33:38 -07:00

18 KiB

Server Infrastructure — Implementation Guide

Ansible-based provisioning and deployment for homelab web servers.

Architecture

flowchart TB
    subgraph provision ["Provisioning (manual / rare)"]
        Control1["ai-server-4080\n(control node)"]
        Control1 -->|ansible-playbook site.yml| Adama
        Control1 -->|ansible-playbook site.yml| Roslin
    end

    subgraph cicd ["CI/CD (every merge to master)"]
        Gitea["Gitea push"]
        Gitea --> Tests["Act: unit tests"]
        Tests --> Deploy["Act: deploy job"]
        Deploy --> AnsibleDeploy["ansible-playbook deploy-apps.yml"]
        AnsibleDeploy --> Adama2["adama"]
        AnsibleDeploy --> Roslin2["roslin"]
    end
Pipeline When Playbook Where it runs
Provision New VM, OS change, firewall, Docker install site.yml ai-server-4080 — run manually
Deploy Green unit tests on master deploy-apps.yml Gitea Act runner on ai-server-4080

Both pipelines share the same inventory (inventory/hosts.yml).

Servers

Name IP Role
adama 10.0.0.77 Ubuntu Server VM (Proxmox) — app host
roslin 10.0.0.176 Ubuntu Server VM (Proxmox) — app host
ai-server-4080 10.0.0.128 Control node + Gitea act runner + Ollama + SearxNG + observability; also runs app replicas

Hostname on this machine: ryan-development-1

Repo Layout

server-infra/
├── IMPLEMENTATION.md       # This file
├── README.md               # Quick start
├── ansible.cfg
├── requirements.yml        # Ansible Galaxy collections
├── inventory/
│   ├── hosts.yml
│   ├── group_vars/
│   │   └── all.yml         # vars + app_catalog
│   └── host_vars/
│       ├── adama.yml       # host_apps (django + dta_webapp)
│       ├── roslin.yml      # host_apps (mirrors adama)
│       └── ai-server-4080.yml  # control node / act runner / SearxNG / observability
├── playbooks/
│   ├── site.yml            # Phase 1: provision
│   └── deploy-apps.yml     # Phase 2: CI deploy
├── roles/
│   ├── common/             # Base packages
│   ├── ufw/                # Firewall
│   ├── docker/             # Docker CE + compose plugin
│   ├── nodejs/             # Node.js + npm + npx (NodeSource)
│   ├── gitea-key/          # per-server SSH key + Gitea access probe
│   ├── tianji/             # Monitoring reporter
│   ├── observability/      # Loki + Prometheus + Grafana (ai-server-4080)
│   ├── searxng/            # SearxNG JSON API for chat_backend (#10)
│   ├── alloy/              # log/metrics shipper
│   ├── app-deploy/         # django (docker) + node-static deploy
│   └── web-static/         # nginx container serving /var/www builds
└── scripts/
    ├── provision.sh        # Wrapper with --limit support
    └── deploy.sh           # Wrapper for deploy playbook

Prerequisites (One-Time Bootstrap)

Ansible needs SSH + sudo on each target before playbooks work.

  1. Create westfarn on each VM with sudo membership.
  2. Copy your SSH public key from the control node (ai-server-4080):
    ssh-copy-id westfarn@10.0.0.77
    ssh-copy-id westfarn@10.0.0.176
    
  3. Confirm passwordless SSH:
    ssh westfarn@10.0.0.77
    ssh westfarn@10.0.0.176
    
  4. First-time only — grant passwordless sudo on each new host before the first provision.sh run. Ubuntu 26.04 ships sudo-rs by default; Ansible's --ask-become-pass does not recognize its password prompt, so bootstrap sudo manually over SSH instead:
    ssh -t westfarn@10.0.0.176   # repeat for each host IP
    
    On the host:
    echo 'westfarn ALL=(ALL) NOPASSWD:ALL' | sudo tee /etc/sudoers.d/westfarn
    sudo chmod 440 /etc/sudoers.d/westfarn
    exit
    
    The common role writes the same file on later runs; this one-time step is only needed before Ansible can escalate privileges the first time.
  5. On ai-server-4080 (control node), install Ansible:
    sudo apt update && sudo apt install -y ansible
    # or: pip install ansible
    
  6. Install Galaxy collections:
    cd ~/Documents/repos/server-infra
    ansible-galaxy collection install -r requirements.yml
    
  7. Update inventory/host_vars/ai-server-4080.yml with this machine's LAN IP (ansible_host).

Testing on a Single Server

Use --limit to target one host without touching the others. Helper scripts wrap this.

Ping one host

./scripts/provision.sh adama --check   # dry run
ansible adama -m ping

Provision one host

New hosts need the one-time passwordless sudo bootstrap in Prerequisites before the first run.

# Dry run (no changes)
./scripts/provision.sh adama --check

# Apply for real
./scripts/provision.sh adama

# Same for other hosts
./scripts/provision.sh roslin
./scripts/provision.sh ai-server-4080

Provision all hosts

./scripts/provision.sh

Deploy to one host (Phase 2)

./scripts/deploy.sh adama
./scripts/deploy.sh --check roslin

Under the hood, scripts pass --limit <hostname> to ansible-playbook.

Phase 1: Provision (site.yml)

Applies roles in order to the webservers group:

Role Purpose
common apt update, git, python3, pip, curl, ca-certificates
ufw Firewall: SSH from LAN only, HTTP/HTTPS public
docker Docker CE, compose plugin, add westfarn to docker group
nodejs Node.js + npm + npx (NodeSource) for dta_webapp builds
gitea-key Per-server SSH key + Gitea access probe
tianji Monitoring reporter

UFW rules

Port Source Purpose
22 10.0.0.0/24 SSH (LAN only)
80 anywhere HTTP
443 anywhere HTTPS
default deny incoming Block everything else

Warning: Test UFW on one host first (./scripts/provision.sh adama). Keep a Proxmox console session open in case SSH rules lock you out.

After Docker install, re-SSH so the docker group membership takes effect.

Phase 2: CI Deploy (deploy-apps.yml)

Apps

App Type Hosts Envs Notes
company_site django (docker) adama + roslin (+ ai-server-4080) prod active/active behind NPM; beta port reserved
dta_service django (docker) adama + roslin + ai-server-4080 beta + prod active/active behind NPM
dta_webapp node/vite static adama + roslin (+ ai-server-4080) beta + prod active/active; built to /var/www/<env>.app.ditchtheagent/html, served by web-static nginx
scha django (docker) adama + roslin + ai-server-4080 prod active/active behind NPM; beta port reserved
chat_web_app node-static (CRA) adama + roslin + ai-server-4080 beta + prod active/active; built to /var/www/<env>.chat.aimloperations/html, served by web-static nginx
chat_backend django (docker) adama + roslin + ai-server-4080 beta + prod active/active behind NPM; Ollama http://10.0.0.128:11434; SearxNG http://10.0.0.128:8088 (SEARXNG_BASE_URL)

Django apps use a shared external Postgres (via DATABASE_URL in each host's env file) so active/active replicas share one database. Beta and prod never share a DB.

Data model

  • app_catalog (group_vars/all.yml) — how each app is built (repo, type, compose file, migrate cmd).
  • host_apps (host_vars/<host>.yml) — which app+env+port runs on that host.
  • Django app = one compose project per env: project name <app>_<env>, host port from host_apps. Ports match across adama/roslin so NPM can balance adama:PORT + roslin:PORT.

Ports

Reserved host ports for NPM upstreams. Ports must match across every host that serves the same app+env. Rows marked not deployed keep the port free for a future beta replica.

App beta prod Deployed on
company_site 8010 (not deployed) 8000 adama, roslin, ai-server-4080
dta_service 8011 8001 adama, roslin, ai-server-4080
scha 8012 (not deployed) 8002 adama, roslin, ai-server-4080
chat_backend 8013 8003 adama, roslin, ai-server-4080
dta_webapp (nginx) 8081 8080 adama, roslin, ai-server-4080
chat_web_app (nginx) 8083 8082 adama, roslin, ai-server-4080
SearxNG (LAN only) 8088 ai-server-4080 only (searxng_stack); not an NPM upstream

Host-local services on ai-server-4080 (not balanced by NPM):

Service Port Notes
Ollama 11434 Not Ansible-managed today; GPU host
SearxNG 8088 roles/searxng (#10); JSON API for chat_backend grounded search
Loki 3100 roles/observability
Prometheus 9090 roles/observability
Grafana 3000 roles/observability

Port clash warning: do not bind SearxNG to 8080 — that is dta_webapp prod. chat_backend secrets must use SEARXNG_BASE_URL=http://10.0.0.128:8088.

Flow

  1. Gitea push to master → repo's .gitea/workflows runs tests.
  2. On green, deploy job on the self-hosted runner calls:
    ~/Documents/repos/server-infra/scripts/deploy.sh \
      --app company_site --env prod --ref "${{ gitea.sha }}"
    
  3. deploy-apps.yml runs against webservers; each host deploys only the matching app+env from its host_apps.

app-deploy role behavior

  • django: push per-app secret from control node {{ secrets_dir }}/<app>/<app>_<env>.env to host {{ apps_env_dir }} → git checkout at ref → copy .env into checkout → docker compose buildup -d → migrate (run once, shared DB).
  • node-static: git checkout at ref → npm cinpm run build:<env> (writes to the app's webroot_pattern, e.g. /var/www/{env}.app.ditchtheagent/html or /var/www/{env}.chat.aimloperations/html).
  • web-static role: one nginx container per app host serving the static roots on their ports (from host_apps); NPM balances across hosts.

Reverse proxy / load balancing (NPM at 10.0.0.230)

Ansible does not manage NPM. It only guarantees stable host ports. In NPM you point each domain at the backend(s):

  • Single host: standard Proxy Host → adama:PORT.
  • Active/active: jc21 NPM's UI Proxy Host is single-target. To balance adama+roslin you need the Advanced tab with a custom upstream {} block (or a real LB). Confirm this before relying on active/active.
App Domains Backends
company_site aimloperations.com (+ www) adama:8000 + roslin:8000
dta_service (see DTA NPM hosts) adama:8001 / 8011 + same on roslin / ai-server-4080
dta_webapp (see DTA NPM hosts) adama:8080 / 8081 + same on roslin
scha schawheaton.aimloperations.com, schawheaton.com (+ www) adama:8002 + roslin:8002 (+ ai-server-4080:8002)
chat_web_app chat.aimloperations.com (+ www); beta.chat.aimloperations.com adama:8082 / 8083 + same on roslin / ai-server-4080
chat_backend chatbackend.aimloperations.com; beta.chatbackend.aimloperations.com adama:8003 / 8013 + same on roslin / ai-server-4080

Required changes IN each app repo (owned separately)

  • docker-compose.prod.yml: drop the bundled db service; web reads DATABASE_URL / DB_HOST pointing at the shared external Postgres.
  • Each app has its own database + user on the shared Postgres.
  • .gitea/workflows/deploy.yml: replace the local scripts/deploy.sh step with a call to server-infra/scripts/deploy.sh --app <name> --env <env> --ref <sha> (keep the test/docker jobs).
  • dta_webapp: npm run build:beta / build:prod output to /var/www/beta.app.ditchtheagent/html / /var/www/prod.app.ditchtheagent/html.
  • chat_web_app: npm run build:beta / build:prod output to /var/www/beta.chat.aimloperations/html / /var/www/prod.chat.aimloperations/html.

Companion chat_web_app frontend is already registered in this infrastructure repo.

Shared Postgres (10.0.0.230, same box as NPM)

One shared instance; each app+env gets its own database (beta and prod MUST NOT share a DB — active/active replicas of the same env share one DB, different envs do not).

app env database DATABASE_URL
company_site prod company_site postgres://westfarn:<pw>@10.0.0.230:5432/company_site
company_site beta company_site_beta postgres://westfarn:<pw>@10.0.0.230:5432/company_site_beta
dta_service prod dta_service postgres://westfarn:<pw>@10.0.0.230:5432/dta_service
dta_service beta dta_service_beta postgres://westfarn:<pw>@10.0.0.230:5432/dta_service_beta
scha prod scha postgres://westfarn:<pw>@10.0.0.230:5432/scha
scha beta scha_beta postgres://westfarn:<pw>@10.0.0.230:5432/scha_beta
chat_backend prod chat_backend postgres://westfarn:<pw>@10.0.0.230:5432/chat_backend
chat_backend beta chat_backend_beta postgres://westfarn:<pw>@10.0.0.230:5432/chat_backend_beta

Server prereqs on 10.0.0.230: create each DB + grant westfarn; listen_addresses covers LAN; pg_hba.conf allows 10.0.0.0/24; firewall opens 5432 to 10.0.0.0/24 only.

One-time host bootstrap (per target)

  • Gitea SSH key: the gitea-key role (in site.yml) generates a key per server, configures SSH for port 30009, probes access, and — if the server can't reach Gitea yet — prints the public key to add and stops. Add the key (Gitea user SSH keys, or repo Deploy Keys) and re-run provisioning.
  • Create control-node secrets {{ secrets_dir }}/<app>/<app>_<env>.env (default ~/Documents/secrets/<app>/<app>_<env>.env) with DATABASE_URL (see table), DJANGO_ENV, DJANGO_SECRET_KEY, WEB_PORT (matching the port table). Deploy pushes these to /opt/apps/env/<app>_<env>.env (mode 600) on adama + roslin. Never committed to git.
  • Node.js/npm/npx for the dta_webapp build — installed by the nodejs role in site.yml (NodeSource, node_major default 20).

Gitea Act Runner

Recommended: Single self-hosted runner on ai-server-4080.

  • One orchestration point.
  • App hosts (adama/roslin) run the workloads; no runner needed on them for deploy fan-out.
  • Runner needs: Ansible, this repo checked out, SSH key to all hosts, vault password (later).

Runner requirements on ai-server-4080

Requirement Why
Ansible Run deploy-apps.yml
server-infra checkout Playbooks + inventory
SSH key to adama + roslin Deploy fan-out

On every push or merged PR to master, .gitea/workflows/sync-checkout.yml fast-forward pulls this repo at ~/Documents/repos/server-infra on the Act runner so playbooks and inventory stay current without a manual git pull.

SSH Keys for CI Deploy

Key Used by Purpose
Personal key You Manual provisioning
Deploy key (runner) Act → Ansible → hosts Automated deploy

Consider a dedicated deploy user with limited sudo (docker only) — future hardening step.

Secrets (Phase 2)

Use Ansible Vault for production secrets. Do not commit plaintext.

ansible-vault create inventory/group_vars/webservers/vault.yml
ansible-playbook playbooks/site.yml --ask-vault-pass

Store vault password for CI in a file readable only by the Act runner (e.g. ~/.ansible-vault-pass, mode 600).

Implementation Order

# Task Status
1 Create server-infra repo Done
2 Inventory with all 3 hosts Done
3 Bootstrap SSH to adama + roslin Manual
4 site.yml → common, ufw, docker Done
5 Verify ansible webservers -m ping Manual
6 Test on single server: ./scripts/provision.sh adama Manual
7 Provision all: ./scripts/provision.sh Manual
8 Deploy SSH key for Act runner Future
8a Auto-sync runner checkout on master (.gitea/workflows/sync-checkout.yml) Done
9 Stub deploy-apps.yml + update company_site workflow Future
10 Dockerize company_site Future (separate ticket)
10a Register + deploy scha (all webservers, port 8002) In progress (scha#19)
10b Register + deploy chat_web_app (node-static, ports 8082/8083) Done (prod); beta (#7, chat_web_app#35)
10c Register + deploy chat_backend (django, ports 8003/8013) Done (prod); beta (#7, chat_backend#26)
11 Gitea container registry (optional) Future

Open Decisions

  1. Deploy userwestfarn vs dedicated deploy for CI.
  2. NPM load balancing — confirm jc21 NPM can express adama+roslin upstreams (Advanced tab), else active/active is just two independent instances.
  3. Secrets — Ansible Vault vs per-host env files (currently per-host /opt/apps/env/*.env).

Adding a New VM

  1. Add host to inventory/hosts.yml under webservers.
  2. Bootstrap SSH: ssh-copy-id westfarn@<new-ip>.
  3. Test: ./scripts/provision.sh <hostname> --check.
  4. Provision: ./scripts/provision.sh <hostname>.
  5. Deploys automatically include new host once in webservers group.