westfarn a049e4f685
Unit Tests / test (push) Successful in 9s
Track token in/out per prompt on PromptMetric (#18)
Closes #15

## Summary
- Add nullable `tokens_in` / `tokens_out` `IntegerField`s to `PromptMetric` to record real prompt/completion token counts per turn.
- New `extract_token_usage()` helper parses provider usage payloads (LangChain `usage_metadata`, OpenAI-style `prompt_tokens`/`completion_tokens`, Ollama `prompt_eval_count`/`eval_count`). When a provider reports no usage, the fields stay **null** — counts are never estimated/fabricated.
- `create_prompt_metric` / `finish_prompt_metric` in both `consumers.py` and `consumers_graph.py` accept and persist optional `tokens_in` / `tokens_out` (added to `update_fields` only when present).
- Admin panel (this ticket's deliverable):
  - `PromptMetricAdmin` lists `tokens_in` / `tokens_out` and adds `event` / `model_name` / `has_file` filters.
  - `ConversationAdmin` shows summed `tokens_in` / `tokens_out` / `tokens_total` per conversation.
- Migration `0023_promptmetric_tokens_in_promptmetric_tokens_out` (existing rows remain valid — null).

## Note on live capture
The streaming chat path uses LangChain `StrOutputParser`, which yields plain string chunks with no usage metadata, so live turns currently persist `null` tokens (honest, per acceptance criteria — no fabricated counts). The plumbing + helper are in place so wiring real provider usage is a drop-in once the services expose it.

## Follow-ups
- #16 — Show token in/out in chat web app UI (FE + API exposure)
- #17 — Token-based billing, quotas, and enforcement

## Test plan
- [x] `uv run python manage.py test` — full suite green (266 tests, 6 skipped)
- [x] Model: token fields default null + persist when set
- [x] `extract_token_usage`: LangChain / OpenAI / Ollama key variants, attribute sources, bool/float handling, missing usage → (None, None)
- [x] Metric lifecycle: tokens persist when provided, stay null when absent (both consumers)
- [x] Admin: conversation token totals sum across metrics and ignore other conversationsReviewed-on: #18
2026-07-26 08:22:34 -07:00
2026-07-25 06:17:46 -07:00
2025-12-07 06:31:06 -06:00
2025-03-07 12:23:00 -06:00

Chat Backend

Django + Channels API for AIML Operations chat (chatbackend.aimloperations.com). Packaging via uv; production deploy via server-infra.

Companion frontend: chat_web_app (node-static, not Docker).

Ticket: chat_backend#6

Layout

chat_backend/                 ← repo root (Dockerfile, compose, pyproject, .gitea)
├── llm_be/                   ← Django project root (manage.py)
│   ├── manage.py
│   ├── llm_be/               ← settings, urls, asgi/wsgi
│   └── chat_backend/         ← app (models, consumers, services, storage)
├── scripts/
│   ├── docker-entrypoint.sh
│   └── validate-env.sh
├── docker-compose.yml        ← local/CI (bundled Postgres)
└── docker-compose.prod.yml   ← server-infra (external DATABASE_URL)

Local development

Prerequisites

  • Python 3.12+
  • uv
  • Docker + Docker Compose (optional, recommended)
  • Ollama reachable at OLLAMA_BASE_URL for LLM features

uv (host)

cp .env.example .env
uv sync
cd llm_be
uv run python manage.py migrate
uv run python manage.py runserver 0.0.0.0:8003

Without DATABASE_URL / DB_HOST, settings fall back to SQLite (llm_be/db.sqlite3).

Tests:

cd llm_be
uv run python manage.py test

The suite is offline by default — the custom test runner (llm_be/test_runner.py) sets SKIP_RAG_INIT=1 and a cheap password hasher, and LLM chains are faked, so no Ollama, Chroma, SMTP or network access is needed. Tests live in llm_be/chat_backend/tests/ (models, storage, serializers, views, services, signals, consumers).

Non-deterministic checks against a real model server are opt-in:

cd llm_be
RUN_LIVE_OLLAMA_TESTS=1 uv run python manage.py test chat_backend.tests.test_live_ollama

Docker (dev, bundled Postgres)

docker compose up --build

App: http://localhost:8003 — Postgres via bundled db (postgres://chat_backend:chat_backend@db:5432/chat_backend). Compose does not read host DATABASE_URL (avoids CI/prod leaks); override with COMPOSE_DATABASE_URL if needed.

Environment variables

Variable Dev default Prod required Notes
DJANGO_ENV dev prod / beta
DJANGO_SECRET_KEY insecure default yes Must be real in prod/beta
DJANGO_DEBUG true when dev false
DJANGO_ALLOWED_HOSTS localhost + chat hosts yes Comma-separated
DJANGO_CSRF_TRUSTED_ORIGINS derived from hosts optional Full origins
DATABASE_URL SQLite fallback yes Shared Postgres in prod
WEB_PORT n/a (compose maps 8003) 8003 Host port for prod compose
OLLAMA_BASE_URL http://127.0.0.1:11434 yes GPU host in prod: http://10.0.0.128:11434
OLLAMA_MODEL / OLLAMA_EMBED_MODEL from DEBUG optional Override model names
EMAIL_HOST_* empty yes (prod/beta) SMTP2GO
CAPTCHA_SECRET_KEY empty recommended
CORS_ALLOWED_ORIGINS local + chat FE set in prod Frontend origin
USE_TLS_PROXY false (dev) true behind NPM Sets SECURE_PROXY_SSL_HEADER
GUNICORN_WORKERS / GUNICORN_BIND 2 / 0.0.0.0:8000 optional Entrypoint
SKIP_RAG_INIT unset CI/migrate often 1 Skip Chroma/Ollama boot work

Templates: .env.example (local), .env.prod.example (control-node secret).

Control-node secret path (server-infra on ai-server-4080):

~/Documents/secrets/chat_backend/chat_backend_prod.env

Validate with:

./scripts/validate-env.sh ~/Documents/secrets/chat_backend/chat_backend_prod.env

If DATABASE_URL password contains $, escape each as $$ for Compose.

Ollama

All clients (ollama.Client, OllamaLLM, OllamaEmbeddings, ChatOllama) use OLLAMA_BASE_URL — never hardcoded localhost in deployed code.

Env Typical URL
Local (Ollama on same machine) http://127.0.0.1:11434
prod / beta (containers on adama/roslin/ai-server) http://10.0.0.128:11434

Firewall / Ollama listen on ai-server-4080 must allow 10.0.0.0/24:11434.

File storage

Prompt attachments and workspace documents use DatabaseStorage (chat_backend.StoredFile BinaryField in Postgres). Blobs are not written to the container filesystem under media/.

RAG loaders that need a path materialize a short-lived temp file, then delete it. Chromas vector index may still use a volume (chroma_db); that is embeddings metadata, not the original upload.

Production (docker-compose.prod.yml)

  • Single web service; no bundled DB — DATABASE_URL → shared Postgres (10.0.0.230).
  • Host port from WEB_PORT (catalog: 8003; beta reserved 8013).
  • Entrypoint: wait DB → migrate → collectstatic → gunicorn + UvicornWorker (ASGI for HTTP and WebSockets).
  • Active/active on adama + roslin + ai-server-4080; NPM balances upstreams.
  • Deployed by:
~/Documents/repos/server-infra/scripts/deploy.sh \
  --app chat_backend --env prod --ref <sha>

CI / CD (Gitea Actions)

Workflow Trigger Action
unittests.yml push + PR → master uv sync + manage.py test
ci.yml PR → master same unit tests
deploy.yml after Unit Tests succeeds on master push docker build + tests on ephemeral compose Postgresdeploy.sh

Deploy never runs on PRs.

Security note

Secrets previously hardcoded in settings.py (email password, captcha, Django secret) must live only in the control-node env file. Rotate anything that was ever committed; never commit .env or ~/Documents/secrets/.

S
Description
Backend django project for the chat site
Readme
1.2 MiB
Languages
Python 98.9%
HTML 0.7%
Shell 0.3%