Retrieval-Augmented Knowledge Base for Business Decisions
A vector knowledge base that lets an AI answer strategy, pricing and risk questions grounded in a specific library of thinkers rather than generic output. Built on Postgres with pgvector and OpenAI embeddings, with relational metadata linking authors, sources and concepts, and a query path using LLM query expansion, dual concept-and-document matching, and priority weighting. The relevance threshold is set from measured similarity distributions rather than intuition, so questions the corpus genuinely doesn't cover return an honest "not strongly covered" instead of a confident answer assembled from weak matches. Delivered after an evidenced architecture change away from a file-pipeline design that could not work within the hosting platform's hard limits.
The problem
A business owner wanted AI answers to strategy, pricing and risk questions grounded in a specific library — behavioural science, persuasion, decision-making — rather than generic model output.
What was built
A Postgres + pgvector store with relational metadata (authors → sources → concepts → retrieval rules → synthesis rules), an ingestion path through an n8n form for transcripts, Python scripts for bulk ingestion, and a query path with LLM query expansion, dual matching against curated concepts and document chunks, a relevance floor, and author-priority weighting. Surfaces include a Claude Code skill, an n8n workflow, and a browser console.
The hard part
Saying the architecture was wrong. The original design had a workflow watching a cloud drive folder, downloading files, parsing, embedding, and writing to the database. Three constraints killed it: the source material is largely audiobooks — 200 MB-plus MP3s — and the managed platform caps form data around 200 MB with no way to raise it; the cloud drive was the wrong store on quota and cost; and the document parser has its own ceilings and cannot process audio at all, since audio must be transcribed before it can be embedded.
Then the requirement simplified to transcript-only — at which point the size problem disappeared entirely and the rejected tool became the right answer again, because the form was the friendliest path for the non-technical person doing the uploads. That sequence is worth telling in full: reject on evidence, re-adopt when the constraint changes.
Setting the relevance floor from measurement, not intuition. Strong in-domain questions score 0.30–0.38; off-domain questions score 0.24–0.25. The distributions separate cleanly, so the floor sits at 0.28 with query expansion on. Below it, the system returns an honest "not strongly covered yet" rather than an answer assembled from weak matches — which, in a decision-support tool, is the entire point.
What can be verified
- Retrieval validated end to end with correct ranking for in-domain questions
- floor tuned against the measured distributions above
- first full ingest of 575 chunks for roughly one cent
- 12 curated concepts seeded
- the context packet assembles correctly in the UI path, with only the final model call failing on billing.
The workflow
(select to enlarge)
On numbers: every figure above is an artefact count or a measured technical value. No business-outcome metric, whether time saved, revenue or conversion, was captured on these engagements, so none is claimed.
On status: reflects repository evidence and platform backups, not a live systems check.