Overview
Overview¶
concierge.knowledge is a standalone bounded context that ingests
Markdown files into a pgvector-backed
vector store via langchain-postgres.
It follows the same clean-architecture layering used by the other services in
this repository.
The bounded context exposes a single CLI entry point (knowledge-cli)
with two subcommand groups:
ingest(run/stats/drop) — write side that creates the pgvector table, splits + embeds Markdown chunks, and manages the collection.search(run) — read side that runssimilarity_searchagainst an existing collection using the same embeddings factory, with both human-readable and--jsonoutput for piping.
All embedding / persistence / loader concerns are wired through
factories in concierge.knowledge.infrastructure, so the application
layer never imports a framework directly.
flowchart LR
CLI["Typer CLI<br/>knowledge-cli"]
Ingest["ingest run/stats/drop"]
Search["search run"]
CLI --> Ingest
CLI --> Search
Ingest --> IngestUC["IngestMarkdown /<br/>DeleteCollection (use cases)"]
Search --> SearchUC["SearchKnowledge (use case)"]
IngestUC --> Repo[KnowledgeRepository protocol]
SearchUC --> Repo
Repo --> PG[PgVectorKnowledgeRepository]
PG --> LC[langchain-postgres PGVectorStore]
LC --> Docker[("pgvector / PostgreSQL<br/>compose service")]
LC --> Azure[("Azure Database for PostgreSQL<br/>Flexible Server + pgvector")]
IngestUC --> Loader[load_markdown_documents]
IngestUC --> Splitter[split_documents<br/>RecursiveCharacterTextSplitter]
IngestUC --> Emb["create_embeddings()<br/>(Foundry / Fake)"]
SearchUC --> Emb
Directory Layout¶
concierge/knowledge/
domain/
entities.py # KnowledgeDocument, KnowledgeChunk, KnowledgeSearchResult
value_objects.py # CollectionName, ChunkId, ContentHash
exceptions.py # CollectionValidationError
application/
repositories.py # KnowledgeRepository protocol
use_cases.py # IngestMarkdown / DeleteCollection / SearchKnowledge
infrastructure/
cli/app.py # knowledge-cli (Typer)
embeddings/factory.py # create_embeddings() (foundry|fake)
loaders/markdown.py # load_markdown_documents() / split_documents()
persistence/
factory.py # get_knowledge_repository()
pgvector.py # PgVectorKnowledgeRepository (langchain-postgres)
The import direction infrastructure -> application -> domain is enforced
by import-linter (knowledge-layers, knowledge-domain-no-frameworks,
knowledge-application-no-infrastructure, knowledge-no-agents-coupling
in pyproject.toml).
Minimal Quickstart (Docker Compose, fake embeddings)¶
The fastest end-to-end smoke test. No Azure credentials required.
fake embeddings are for plumbing tests only
KNOWLEDGE_EMBEDDING_PROVIDER=fake produces deterministic but
semantically meaningless vectors, so search rankings are effectively
random. Use it only to verify the ingest/search wiring without Azure. For
real semantic retrieval (RAG agents, realtime voice) use foundry and
re-ingest — see Troubleshooting.
# 1. Boot local pgvector
docker compose up -d postgres
# 2. Use deterministic fake embeddings so no Foundry call happens
export KNOWLEDGE_EMBEDDING_PROVIDER=fake
# 3. Ingest this repository's docs/ folder into a fresh collection
uv run knowledge-cli ingest run --collection demo_md docs
# 4. Verify the row count
uv run knowledge-cli ingest stats --collection demo_md
# 5. Try a query against the same collection
uv run knowledge-cli search run --collection demo_md "vector store" --k 3
# 6. Tear down the collection when you are done
uv run knowledge-cli ingest drop --collection demo_md --yes
Expected output of step 3 looks like:
Minimal Quickstart (Azure Database for PostgreSQL + Foundry)¶
For the managed target, reuse the AZURE_* and AZURE_AI_PROJECT_ENDPOINT
variables described in
Step 3 – PostgreSQL (pgvector) CRUD
and Step 2 – Observability.
# 1. Sign in so DefaultAzureCredential can fetch tokens for both
# Azure PostgreSQL (Entra auth) and Foundry embeddings.
az login
# 2. Make sure your Flexible Server has the vector extension enabled
# and your Entra principal is mapped to a PostgreSQL role
# (see the Tutorial Step 3 page for the SQL snippet).
# 3. Ingest a Markdown directory into Azure pgvector using Foundry embeddings
uv run knowledge-cli ingest run \
--collection demo_md \
--target azure \
docs
# 4. Inspect / search / drop as needed
uv run knowledge-cli ingest stats --collection demo_md --target azure
uv run knowledge-cli search run --collection demo_md --target azure "vector store"
uv run knowledge-cli ingest drop --collection demo_md --target azure --yes
Default collection
Omitting --collection falls back to KNOWLEDGE_DEFAULT_COLLECTION
(default knowledge_default). Useful when you want to keep all
Markdown in a single table.
Configuration¶
All knowledge settings are read by concierge.settings.KnowledgeSettings
with the KNOWLEDGE_ prefix. PostgreSQL connection settings are reused
from PostgresSettings (POSTGRES_*) for --target docker and
AzurePostgresSettings (AZURE_*) for --target azure.
| Variable | Default | Description |
|---|---|---|
KNOWLEDGE_EMBEDDING_PROVIDER |
foundry |
foundry uses Azure AI Foundry with DefaultAzureCredential; fake uses DeterministicFakeEmbedding (no network calls). |
KNOWLEDGE_EMBEDDING_MODEL |
text-embedding-3-small |
Foundry deployment name passed to init_embeddings("azure_ai:<model>"). |
KNOWLEDGE_VECTOR_SIZE |
1536 |
Embedding dimension used when creating the pgvector table. Must match the embedding model. |
KNOWLEDGE_VECTOR_BACKEND |
pgvector |
Vector store backend. Only pgvector is implemented today. |
KNOWLEDGE_DEFAULT_COLLECTION |
knowledge_default |
Table name used when --collection is omitted. Must match ^[A-Za-z0-9_]+$. |
KNOWLEDGE_CHUNK_SIZE |
1000 |
RecursiveCharacterTextSplitter chunk_size. |
KNOWLEDGE_CHUNK_OVERLAP |
200 |
RecursiveCharacterTextSplitter chunk_overlap. |
AZURE_AI_PROJECT_ENDPOINT |
"" |
Required when KNOWLEDGE_EMBEDDING_PROVIDER=foundry. The CLI derives the OpenAI-compatible /openai/v1 endpoint automatically. |
--target azure additionally consumes AZURE_DBHOST, AZURE_DBNAME,
AZURE_DBUSER, AZURE_USE_ENTRA_AUTH, and (when Entra auth is disabled)
AZURE_DBPASSWORD. See
Step 3 – PostgreSQL (pgvector) CRUD
for the full description and provisioning steps.
Troubleshooting¶
Search returns irrelevant chunks (or an agent says "no results" despite hits)¶
KNOWLEDGE_EMBEDDING_PROVIDER=fake uses DeterministicFakeEmbedding, which
hashes text into vectors that carry no semantic meaning. It only exists to
smoke-test the ingest/search plumbing without calling Azure. With fake,
similarity_search returns essentially arbitrary chunks — every cosine distance
clusters around the same value (~0.9) — so callers that ground their answers on
the results (RAG agents, the realtime voice tool) report that nothing relevant
was found even though hits came back.
Switch to Foundry embeddings for any real retrieval:
KNOWLEDGE_EMBEDDING_PROVIDER=foundry
KNOWLEDGE_EMBEDDING_MODEL=text-embedding-3-small
AZURE_AI_PROJECT_ENDPOINT=https://<resource>.services.ai.azure.com/api/projects/<project>
Confirm the fix with a query whose answer you know is in the corpus — the obviously-relevant document should now rank first:
Switched the provider or model but results did not change¶
Embeddings are computed at ingest time and stored in the table. Changing
KNOWLEDGE_EMBEDDING_PROVIDER, KNOWLEDGE_EMBEDDING_MODEL, or
KNOWLEDGE_VECTOR_SIZE does not retro-fit existing rows, and ingest run
appends instead of replacing. Drop and re-ingest so the whole collection
shares one embedding space:
uv run knowledge-cli ingest drop --collection knowledge_default --yes
uv run knowledge-cli ingest run --collection knowledge_default docs
Restart long-running servers after editing .env
App surfaces such as chat-web load .env and cache settings at startup,
so a provider change in .env is not picked up until you stop and restart
the process.
OperationalError / connection refused on port 5432¶
pgvector is not running. Start it before ingesting or searching (the healthcheck must pass first):
Programmatic use¶
You can also drive the same use cases from Python without going through the
CLI. This is how agents / RAG callers wire a retriever on top of an
existing collection:
from concierge.knowledge.application.use_cases import SearchKnowledge
from concierge.knowledge.domain.value_objects import CollectionName
from concierge.knowledge.infrastructure.embeddings.factory import create_embeddings
from concierge.knowledge.infrastructure.persistence.factory import get_knowledge_repository
from concierge.settings import KnowledgeTarget
collection = CollectionName("demo_md")
repository = get_knowledge_repository(
collection=collection,
target=KnowledgeTarget.DOCKER,
embeddings=create_embeddings(),
)
results = SearchKnowledge(repository).execute(collection, query="vector store", k=3)
for result in results:
print(result.metadata.get("source"), result.content[:80])
See the Knowledge CLI Reference for every command and flag.
Related: for runtime retrieval from agents, see
Shared Agent Runtime
(AGENTS_KNOWLEDGE__*-driven tool registration).