Closes Phases 1–3 of #62 (Phase 4 eval harness left for a follow-up).
Accuracy: Retrieval is decided every turn (GroundingDecider, fails open). FAST no longer skips search — it only selects OLLAMA_MODEL_FAST. Search failures surface an explicit error instead of hallucinating from parametric memory.
Search: Pluggable services/search/ with SearxNG primary + DDGS failover, ranking/dedupe/rumour filtering, numbered dated source blocks, citations persisted on Prompt.citations and emitted as {"v":1,"type":"citations",...} after stream end.
Models: Role-scoped OLLAMA_MODEL_THINKING / _FAST / _UTILITY / OLLAMA_EMBED_MODEL=nomic-embed-text, configurable num_ctx, real model name on PromptMetric, reindex_embeddings management command + loud embedding-dimension mismatch.
SearxNG (ops)
See README SearxNG section. Short version: run searxng/searxng on the GPU host, enable json in settings.yml, set SEARXNG_BASE_URL=http://10.0.0.128:8080 in prod/beta secrets, open :8080 on the LAN firewall like Ollama.
Test plan
SKIP_RAG_INIT=1 python manage.py test chat_backend.tests — 442 OK (6 skipped)
Deploy beta with updated secrets (OLLAMA_MODEL_*, OLLAMA_EMBED_MODEL=nomic-embed-text, SEARXNG_BASE_URL)
After embed change: python manage.py reindex_embeddings
Verify did Taylor Swift get married in FAST and THINKING returns grounded answer with citations frame
Kill SearxNG and confirm factual turns return search_unavailable (not Joe Alwyn hallucination); non-factual chat still works
## Summary
- Closes Phases 1–3 of [#62](https://git.aimloperations.com/ai_ml_operations/chat_backend/issues/62) (Phase 4 eval harness left for a follow-up).
- **Accuracy:** Retrieval is decided every turn (`GroundingDecider`, fails open). `FAST` no longer skips search — it only selects `OLLAMA_MODEL_FAST`. Search failures surface an explicit error instead of hallucinating from parametric memory.
- **Search:** Pluggable `services/search/` with **SearxNG primary** + DDGS failover, ranking/dedupe/rumour filtering, numbered dated source blocks, citations persisted on `Prompt.citations` and emitted as `{"v":1,"type":"citations",...}` after stream end.
- **Models:** Role-scoped `OLLAMA_MODEL_THINKING` / `_FAST` / `_UTILITY` / `OLLAMA_EMBED_MODEL=nomic-embed-text`, configurable `num_ctx`, real model name on `PromptMetric`, `reindex_embeddings` management command + loud embedding-dimension mismatch.
## SearxNG (ops)
See README **SearxNG** section. Short version: run `searxng/searxng` on the GPU host, enable `json` in `settings.yml`, set `SEARXNG_BASE_URL=http://10.0.0.128:8080` in prod/beta secrets, open `:8080` on the LAN firewall like Ollama.
## Test plan
- [x] `SKIP_RAG_INIT=1 python manage.py test chat_backend.tests` — 442 OK (6 skipped)
- [ ] Deploy beta with updated secrets (`OLLAMA_MODEL_*`, `OLLAMA_EMBED_MODEL=nomic-embed-text`, `SEARXNG_BASE_URL`)
- [ ] After embed change: `python manage.py reindex_embeddings`
- [ ] Verify `did Taylor Swift get married` in FAST and THINKING returns grounded answer with citations frame
- [ ] Kill SearxNG and confirm factual turns return search_unavailable (not Joe Alwyn hallucination); non-factual chat still works
Phases 1–3: split THINKING/FAST/UTILITY/EMBED models, structured search with
SearxNG primary + DDGS failover, fail-open grounding, citations on Prompt + WS
frames, single history window with prompt budgeting, and reindex_embeddings.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Summary
GroundingDecider, fails open).FASTno longer skips search — it only selectsOLLAMA_MODEL_FAST. Search failures surface an explicit error instead of hallucinating from parametric memory.services/search/with SearxNG primary + DDGS failover, ranking/dedupe/rumour filtering, numbered dated source blocks, citations persisted onPrompt.citationsand emitted as{"v":1,"type":"citations",...}after stream end.OLLAMA_MODEL_THINKING/_FAST/_UTILITY/OLLAMA_EMBED_MODEL=nomic-embed-text, configurablenum_ctx, real model name onPromptMetric,reindex_embeddingsmanagement command + loud embedding-dimension mismatch.SearxNG (ops)
See README SearxNG section. Short version: run
searxng/searxngon the GPU host, enablejsoninsettings.yml, setSEARXNG_BASE_URL=http://10.0.0.128:8080in prod/beta secrets, open:8080on the LAN firewall like Ollama.Test plan
SKIP_RAG_INIT=1 python manage.py test chat_backend.tests— 442 OK (6 skipped)OLLAMA_MODEL_*,OLLAMA_EMBED_MODEL=nomic-embed-text,SEARXNG_BASE_URL)python manage.py reindex_embeddingsdid Taylor Swift get marriedin FAST and THINKING returns grounded answer with citations frame