# Server Infrastructure — Implementation Guide Ansible-based provisioning and deployment for homelab web servers. ## Architecture ```mermaid flowchart TB subgraph provision ["Provisioning (manual / rare)"] Control1["ai-server-4080\n(control node)"] Control1 -->|ansible-playbook site.yml| Adama Control1 -->|ansible-playbook site.yml| Roslin end subgraph cicd ["CI/CD (every merge to master)"] Gitea["Gitea push"] Gitea --> Tests["Act: unit tests"] Tests --> Deploy["Act: deploy job"] Deploy --> AnsibleDeploy["ansible-playbook deploy-apps.yml"] AnsibleDeploy --> Adama2["adama"] AnsibleDeploy --> Roslin2["roslin"] end ``` | Pipeline | When | Playbook | Where it runs | |----------|------|----------|---------------| | **Provision** | New VM, OS change, firewall, Docker install | `site.yml` | ai-server-4080 — run manually | | **Deploy** | Green unit tests on `master` | `deploy-apps.yml` | Gitea Act runner on ai-server-4080 | Both pipelines share the same inventory (`inventory/hosts.yml`). ## Servers | Name | IP | Role | |------|-----|------| | adama | 10.0.0.77 | Ubuntu Server VM (Proxmox) — app host | | roslin | 10.0.0.176 | Ubuntu Server VM (Proxmox) — app host | | ai-server-4080 | 10.0.0.128 | Control node + Gitea act runner (no app workloads) | Hostname on this machine: `ryan-development-1` ## Repo Layout ``` server-infra/ ├── IMPLEMENTATION.md # This file ├── README.md # Quick start ├── ansible.cfg ├── requirements.yml # Ansible Galaxy collections ├── inventory/ │ ├── hosts.yml │ ├── group_vars/ │ │ └── all.yml # vars + app_catalog │ └── host_vars/ │ ├── adama.yml # host_apps (django + dta_webapp) │ ├── roslin.yml # host_apps (mirrors adama) │ └── ai-server-4080.yml # control node / act runner, no workloads ├── playbooks/ │ ├── site.yml # Phase 1: provision │ └── deploy-apps.yml # Phase 2: CI deploy ├── roles/ │ ├── common/ # Base packages │ ├── ufw/ # Firewall │ ├── docker/ # Docker CE + compose plugin │ ├── nodejs/ # Node.js + npm + npx (NodeSource) │ ├── gitea-key/ # per-server SSH key + Gitea access probe │ ├── tianji/ # Monitoring reporter │ ├── app-deploy/ # django (docker) + node-static deploy │ └── web-static/ # nginx container serving /var/www builds └── scripts/ ├── provision.sh # Wrapper with --limit support └── deploy.sh # Wrapper for deploy playbook ``` ## Prerequisites (One-Time Bootstrap) Ansible needs SSH + sudo on each target before playbooks work. 1. Create `westfarn` on each VM with sudo membership. 2. Copy your SSH public key from the control node (ai-server-4080): ```bash ssh-copy-id westfarn@10.0.0.77 ssh-copy-id westfarn@10.0.0.176 ``` 3. Confirm passwordless SSH: ```bash ssh westfarn@10.0.0.77 ssh westfarn@10.0.0.176 ``` 4. **First-time only** — grant passwordless sudo on each new host before the first `provision.sh` run. Ubuntu 26.04 ships `sudo-rs` by default; Ansible's `--ask-become-pass` does not recognize its password prompt, so bootstrap sudo manually over SSH instead: ```bash ssh -t westfarn@10.0.0.176 # repeat for each host IP ``` On the host: ```bash echo 'westfarn ALL=(ALL) NOPASSWD:ALL' | sudo tee /etc/sudoers.d/westfarn sudo chmod 440 /etc/sudoers.d/westfarn exit ``` The `common` role writes the same file on later runs; this one-time step is only needed before Ansible can escalate privileges the first time. 5. On ai-server-4080 (control node), install Ansible: ```bash sudo apt update && sudo apt install -y ansible # or: pip install ansible ``` 6. Install Galaxy collections: ```bash cd ~/Documents/repos/server-infra ansible-galaxy collection install -r requirements.yml ``` 7. Update `inventory/host_vars/ai-server-4080.yml` with this machine's LAN IP (`ansible_host`). ## Testing on a Single Server Use `--limit` to target one host without touching the others. Helper scripts wrap this. ### Ping one host ```bash ./scripts/provision.sh adama --check # dry run ansible adama -m ping ``` ### Provision one host New hosts need the one-time passwordless sudo bootstrap in [Prerequisites](#prerequisites-one-time-bootstrap) before the first run. ```bash # Dry run (no changes) ./scripts/provision.sh adama --check # Apply for real ./scripts/provision.sh adama # Same for other hosts ./scripts/provision.sh roslin ./scripts/provision.sh ai-server-4080 ``` ### Provision all hosts ```bash ./scripts/provision.sh ``` ### Deploy to one host (Phase 2) ```bash ./scripts/deploy.sh adama ./scripts/deploy.sh --check roslin ``` Under the hood, scripts pass `--limit ` to `ansible-playbook`. ## Phase 1: Provision (`site.yml`) Applies roles in order to the `webservers` group: | Role | Purpose | |------|---------| | `common` | apt update, git, python3, pip, curl, ca-certificates | | `ufw` | Firewall: SSH from LAN only, HTTP/HTTPS public | | `docker` | Docker CE, compose plugin, add `westfarn` to `docker` group | | `nodejs` | Node.js + npm + npx (NodeSource) for `dta_webapp` builds | | `gitea-key` | Per-server SSH key + Gitea access probe | | `tianji` | Monitoring reporter | ### UFW rules | Port | Source | Purpose | |------|--------|---------| | 22 | `10.0.0.0/24` | SSH (LAN only) | | 80 | anywhere | HTTP | | 443 | anywhere | HTTPS | | default | deny incoming | Block everything else | **Warning:** Test UFW on one host first (`./scripts/provision.sh adama`). Keep a Proxmox console session open in case SSH rules lock you out. After Docker install, re-SSH so the `docker` group membership takes effect. ## Phase 2: CI Deploy (`deploy-apps.yml`) ### Apps | App | Type | Hosts | Envs | Notes | |-----|------|-------|------|-------| | `company_site` | django (docker) | adama + roslin (+ ai-server-4080) | prod | active/active behind NPM; beta port reserved | | `dta_service` | django (docker) | adama + roslin + ai-server-4080 | beta + prod | active/active behind NPM | | `dta_webapp` | node/vite static | adama + roslin (+ ai-server-4080) | beta + prod | active/active; built to `/var/www/.app.ditchtheagent/html`, served by web-static nginx | | `scha` | django (docker) | adama + roslin + ai-server-4080 | prod | active/active behind NPM; beta port reserved | | `chat_web_app` | node-static (CRA) | adama + roslin + ai-server-4080 | prod | active/active; built to `/var/www/.chat.aimloperations/html`, served by web-static nginx; beta port reserved | | `chat_backend` | django (docker) | adama + roslin + ai-server-4080 | prod | active/active behind NPM; Ollama via `OLLAMA_BASE_URL=http://10.0.0.128:11434`; beta port reserved | Django apps use a **shared external Postgres** (via `DATABASE_URL` in each host's env file) so active/active replicas share one database. Beta and prod never share a DB. ### Data model - `app_catalog` (`group_vars/all.yml`) — how each app is built (repo, type, compose file, migrate cmd). - `host_apps` (`host_vars/.yml`) — which app+env+port runs on that host. - Django app = one compose project per env: project name `_`, host port from `host_apps`. Ports match across adama/roslin so NPM can balance `adama:PORT` + `roslin:PORT`. ### Ports Reserved host ports for NPM upstreams. Ports must match across every host that serves the same app+env. Rows marked *not deployed* keep the port free for a future beta replica. | App | beta | prod | Deployed on | |-----|------|------|-------------| | company_site | 8010 (*not deployed*) | 8000 | adama, roslin, ai-server-4080 | | dta_service | 8011 | 8001 | adama, roslin, ai-server-4080 | | scha | 8012 (*not deployed*) | 8002 | adama, roslin, ai-server-4080 | | chat_backend | 8013 (*not deployed*) | 8003 | adama, roslin, ai-server-4080 | | dta_webapp (nginx) | 8081 | 8080 | adama, roslin, ai-server-4080 | | chat_web_app (nginx) | 8083 (*not deployed*) | 8082 | adama, roslin, ai-server-4080 | ### Flow 1. Gitea push to `master` → repo's `.gitea/workflows` runs tests. 2. On green, deploy job on the self-hosted runner calls: ```bash ~/Documents/repos/server-infra/scripts/deploy.sh \ --app company_site --env prod --ref "${{ gitea.sha }}" ``` 3. `deploy-apps.yml` runs against `webservers`; each host deploys only the matching app+env from its `host_apps`. ### `app-deploy` role behavior - **django**: push per-app secret from control node `{{ secrets_dir }}//_.env` to host `{{ apps_env_dir }}` → git checkout at ref → copy `.env` into checkout → `docker compose build` → `up -d` → migrate (run once, shared DB). - **node-static**: git checkout at ref → `npm ci` → `npm run build:` (writes to the app's `webroot_pattern`, e.g. `/var/www/{env}.app.ditchtheagent/html` or `/var/www/{env}.chat.aimloperations/html`). - **web-static** role: one nginx container per app host serving the static roots on their ports (from `host_apps`); NPM balances across hosts. ### Reverse proxy / load balancing (NPM at 10.0.0.230) Ansible does **not** manage NPM. It only guarantees stable host ports. In NPM you point each domain at the backend(s): - Single host: standard Proxy Host → `adama:PORT`. - Active/active: jc21 NPM's UI Proxy Host is single-target. To balance adama+roslin you need the **Advanced** tab with a custom `upstream {}` block (or a real LB). Confirm this before relying on active/active. | App | Domains | Backends | |-----|---------|----------| | company_site | aimloperations.com (+ www) | `adama:8000` + `roslin:8000` | | dta_service | (see DTA NPM hosts) | `adama:8001` / `8011` + same on roslin / ai-server-4080 | | dta_webapp | (see DTA NPM hosts) | `adama:8080` / `8081` + same on roslin | | scha | `schawheaton.aimloperations.com`, `schawheaton.com` (+ www) | `adama:8002` + `roslin:8002` (+ `ai-server-4080:8002`) | | chat_web_app | `chat.aimloperations.com` (+ www) | `adama:8082` + `roslin:8082` (+ `ai-server-4080:8082`) | | chat_backend | `chatbackend.aimloperations.com` | `adama:8003` + `roslin:8003` + `ai-server-4080:8003` | ### Required changes IN each app repo (owned separately) - [ ] `docker-compose.prod.yml`: drop the bundled `db` service; `web` reads `DATABASE_URL` / `DB_HOST` pointing at the shared external Postgres. - [ ] Each app has its own database + user on the shared Postgres. - [ ] `.gitea/workflows/deploy.yml`: replace the local `scripts/deploy.sh` step with a call to `server-infra/scripts/deploy.sh --app --env --ref ` (keep the test/docker jobs). - [ ] `dta_webapp`: `npm run build:beta` / `build:prod` output to `/var/www/beta.app.ditchtheagent/html` / `/var/www/prod.app.ditchtheagent/html`. - [ ] `chat_web_app`: `npm run build:beta` / `build:prod` output to `/var/www/beta.chat.aimloperations/html` / `/var/www/prod.chat.aimloperations/html`. Companion `chat_web_app` frontend is already registered in this infrastructure repo. ### Shared Postgres (10.0.0.230, same box as NPM) One shared instance; each app+env gets its own database (beta and prod MUST NOT share a DB — active/active replicas of the same env share one DB, different envs do not). | app | env | database | DATABASE_URL | |-----|-----|----------|--------------| | company_site | prod | `company_site` | `postgres://westfarn:@10.0.0.230:5432/company_site` | | company_site | beta | `company_site_beta` | `postgres://westfarn:@10.0.0.230:5432/company_site_beta` | | dta_service | prod | `dta_service` | `postgres://westfarn:@10.0.0.230:5432/dta_service` | | dta_service | beta | `dta_service_beta` | `postgres://westfarn:@10.0.0.230:5432/dta_service_beta` | | scha | prod | `scha` | `postgres://westfarn:@10.0.0.230:5432/scha` | | scha | beta | `scha_beta` | `postgres://westfarn:@10.0.0.230:5432/scha_beta` | | chat_backend | prod | `chat_backend` | `postgres://westfarn:@10.0.0.230:5432/chat_backend` | | chat_backend | beta | `chat_backend_beta` | `postgres://westfarn:@10.0.0.230:5432/chat_backend_beta` | Server prereqs on 10.0.0.230: create each DB + grant `westfarn`; `listen_addresses` covers LAN; `pg_hba.conf` allows `10.0.0.0/24`; firewall opens 5432 to `10.0.0.0/24` only. ### One-time host bootstrap (per target) - [x] Gitea SSH key: the `gitea-key` role (in `site.yml`) generates a key per server, configures SSH for port 30009, probes access, and — if the server can't reach Gitea yet — prints the public key to add and stops. Add the key (Gitea user SSH keys, or repo Deploy Keys) and re-run provisioning. - [ ] Create control-node secrets `{{ secrets_dir }}//_.env` (default `~/Documents/secrets//_.env`) with `DATABASE_URL` (see table), `DJANGO_ENV`, `DJANGO_SECRET_KEY`, `WEB_PORT` (matching the port table). Deploy pushes these to `/opt/apps/env/_.env` (mode 600) on adama + roslin. Never committed to git. - [x] Node.js/npm/npx for the `dta_webapp` build — installed by the `nodejs` role in `site.yml` (NodeSource, `node_major` default 20). ## Gitea Act Runner **Recommended:** Single self-hosted runner on ai-server-4080. - One orchestration point. - App hosts (adama/roslin) run the workloads; no runner needed on them for deploy fan-out. - Runner needs: Ansible, this repo checked out, SSH key to all hosts, vault password (later). ### Runner requirements on ai-server-4080 | Requirement | Why | |-------------|-----| | Ansible | Run `deploy-apps.yml` | | `server-infra` checkout | Playbooks + inventory | | SSH key to adama + roslin | Deploy fan-out | On every push or merged PR to `master`, `.gitea/workflows/sync-checkout.yml` fast-forward pulls this repo at `~/Documents/repos/server-infra` on the Act runner so playbooks and inventory stay current without a manual `git pull`. ## SSH Keys for CI Deploy | Key | Used by | Purpose | |-----|---------|---------| | Personal key | You | Manual provisioning | | Deploy key (runner) | Act → Ansible → hosts | Automated deploy | Consider a dedicated `deploy` user with limited sudo (docker only) — future hardening step. ## Secrets (Phase 2) Use Ansible Vault for production secrets. Do not commit plaintext. ```bash ansible-vault create inventory/group_vars/webservers/vault.yml ansible-playbook playbooks/site.yml --ask-vault-pass ``` Store vault password for CI in a file readable only by the Act runner (e.g. `~/.ansible-vault-pass`, mode 600). ## Implementation Order | # | Task | Status | |---|------|--------| | 1 | Create `server-infra` repo | Done | | 2 | Inventory with all 3 hosts | Done | | 3 | Bootstrap SSH to adama + roslin | Manual | | 4 | `site.yml` → common, ufw, docker | Done | | 5 | Verify `ansible webservers -m ping` | Manual | | 6 | Test on single server: `./scripts/provision.sh adama` | Manual | | 7 | Provision all: `./scripts/provision.sh` | Manual | | 8 | Deploy SSH key for Act runner | Future | | 8a | Auto-sync runner checkout on `master` (`.gitea/workflows/sync-checkout.yml`) | Done | | 9 | Stub `deploy-apps.yml` + update `company_site` workflow | Future | | 10 | Dockerize `company_site` | Future (separate ticket) | | 10a | Register + deploy `scha` (all webservers, port 8002) | In progress ([scha#19](https://git.aimloperations.com/ai_ml_operations/scha/issues/19)) | | 10b | Register + deploy `chat_web_app` (node-static, port 8082) | In progress ([chat_web_app#13](https://git.aimloperations.com/ai_ml_operations/chat_web_app/issues/13)) | | 10c | Register + deploy `chat_backend` (django, port 8003) | In progress ([chat_backend#6](https://git.aimloperations.com/ai_ml_operations/chat_backend/issues/6)) | | 11 | Gitea container registry (optional) | Future | ## Open Decisions 1. **Deploy user** — `westfarn` vs dedicated `deploy` for CI. 2. **NPM load balancing** — confirm jc21 NPM can express adama+roslin upstreams (Advanced tab), else active/active is just two independent instances. 3. **Secrets** — Ansible Vault vs per-host env files (currently per-host `/opt/apps/env/*.env`). ## Adding a New VM 1. Add host to `inventory/hosts.yml` under `webservers`. 2. Bootstrap SSH: `ssh-copy-id westfarn@`. 3. Test: `./scripts/provision.sh --check`. 4. Provision: `./scripts/provision.sh `. 5. Deploys automatically include new host once in `webservers` group.