Overview
RAG Chatbot
Angular component <wuic-rag-chatbot> that allows querying the WUIC codebase
in natural language, with two operating modes: pure retrieval
(top-K code snippets) and RAG + LLM (response generated by Claude using
the top-K as context). Automatically degrades to retrieval-only if the Claude
API key is not configured on the server.
Architecture
Three layers, no external process:
- Layer 1 -- .NET RAG engine
WuicRagEngine.dll(rag-engine/folder),
loaded in-process by RagEngineLoader when rag-use-dotnet-engine is
true. With ONNX Runtime (CUDA GPU when available, otherwise CPU) it runs
the hybrid BM25 + bge-m3 + cross-encoder reranker retrieval and the call to
the configured LLM provider. Models, tokenizer and index are downloaded on
first start into rag-engine/artifacts/.
- Layer 2 --
RagControllerin KonvergenceCore (/api/Rag/Query,
/api/Rag/Chat, /api/Rag/Health, /api/Rag/Reload), authenticated via
the k-user cookie; it resolves the metadata tools and applies the proposed
actions.
- Layer 3 -- Angular standalone component
<wuic-rag-chatbot>exported
from wuic-framework-lib.
The Python FastAPI server rag_server.py (127.0.0.1:8765) is the legacy
stack: it is still supported as an optional fallback with
rag-use-dotnet-engine = "false", but it is not shipped in release packages.
Page context and dynamic metadata retrieval
When the user chats from an app page, the component injects a **minimal page
context** into the prompt: only the current route/page (e.g. cities/list,
cities/edit, designer). It does NOT inline the column list, SQL identity or
lookup details — inlining them bloated every request regardless of the prompt and
saturated the context window.
The missing details are fetched on-demand by the model via the non-terminal
tool request_metadata_detail, resolved on the backend by
RagController.ResolveMetadataDetail (the engine has no access to the metadata DB).
The model calls the tool, receives the result as a tool_result, and ONLY THEN
emits the propose_* with the real names. The multi-turn loop is handled by the
engine (max 3 retrieval turns).
Supported detail:
detail | Returns | When |
|---|---|---|
columns | SQL identity (schema/table + full qualifier) and the real column list (name, type, lookup, required, pk, SQL name) | the real column names of a route are needed |
lookup_columns | join_alias, value_field, text_field, related_route, related_columns[] of a lookupByID column | composing a SQL snippet on a lookup |
Join-alias convention (lookup): the WUIC auto-generated query joins the related
table of a lookupByID column with alias <column>_<entity> (e.g. the StateProvinceID
column looking up stateprovinces has alias StateProvinceID_stateprovinces). In SQL
snippets the related table fields are referenced as [<join_alias>].[<col>], e.g.
[StateProvinceID_stateprovinces].[StateProvinceName]. The exact alias value must
always be obtained via request_metadata_detail{detail:'lookup_columns'}, never
deduced by hand.
The mechanism is extensible: new detail kinds (e.g. physical_columns, related_routes,
enum_values) are added in the resolver without touching the Angular component or
bloating the context.
Modes
mode input | Behavior |
|---|---|
auto (default) | Uses chat if the backend exposes a Claude API key, otherwise retrieval |
chat | Forces RAG + LLM. If the API key is missing, the backend returns mode=retrieval-only with a visible warning |
retrieval | Forces retrieval-only, skips the LLM call (useful to reduce costs) |
Inputs
| Input | Type | Default | Description |
|---|---|---|---|
title | string | 'Assistente codebase WUIC' | Header label |
mode | 'auto' | 'chat' | 'retrieval' | 'auto' | Operating mode |
topK | number | 5 | Number of chunks to retrieve from the RAG |
showSources | boolean | true | Shows/hides source chips in assistant messages |
maxHistory | number | 20 | Maximum number of turns kept in memory |
placeholder | string | 'Chiedi qualcosa...' | Input placeholder |
model | string | 'claude-haiku-4-5-20251001' | Claude model for chat mode |
chatHeight | string | '420px' | Fixed chat area height |
showClearButton | boolean | true | Shows the "Clear history" button |
Outputs
| Output | Payload | Description |
|---|---|---|
resultSelected | RagSource | Click on a source chip; the component also attempts a deep-link vscode://file/... |
errorOccurred | {message, details?} | Non-recoverable HTTP errors |
turnAdded | RagChatbotTurn | Emitted after each turn (user or assistant) added to the history |
Usage Example
Standalone import in the parent component:
- HTML selector:
<wuic-rag-chatbot mode="auto" [topK]="5" (resultSelected)="onSrc($event)"></wuic-rag-chatbot> - TypeScript import:
import { WuicRagChatbotComponent, RagSource } from 'wuic-framework-lib'; - Add
WuicRagChatbotComponentto the standalone parent component'simports.
For a complete demo page with debug aside and event handling, see
WuicTest/wwwroot/src/app/component/rag-chatbot-demo-page/.
Auth
All HTTP calls are sent with withCredentials: true and the C# bridge
applies the same authentication checks as other WUIC controllers
(session cookie k-user, rule 10 of AGENTS). The RAG engine runs in-process inside the backend: there is no extra
service to expose or protect.
Endpoint access since 1.7.13:
| Endpoint | Access |
|---|---|
/api/Rag/Query, /api/Rag/Health | no login |
/api/Rag/Chat | valid session (guest too, if enableGuest is on) |
/api/Rag/MetadataDetail | no login for metadata details; sample_records, lookup_value, db_connections, db_databases, db_tables, db_columns administrators only |
/api/Rag/Reload | administrators only |
The LLM key configured on the server is never sent to a provider or baseUrl given in the request: whoever overrides them passes their own apiKey, otherwise the answer is retrieval-only.
MCP server wuic-rag and user wuic_assistant
The MCP server scripts/mcp/wuic-rag-mcp.mjs (WUIC Assistant, Claude Code, Cursor) uses these endpoints. When the backend answers 401 it opens a session with MetaService.login, using WUIC_USER / WUIC_PASSWORD from the environment (.mcp.json) or, if missing, from scripts/mcp/wuic-assistant.credentials.json. The login is needed by wuic_ask and, with an administrator user, by wuic_sample_records / wuic_lookup_value; wuic_codebase_search, wuic_route_columns, wuic_route_metadata and wuic_lookup_columns stay without login.
- The first-run wizard creates the dedicated user
wuic_assistant(superadmin, separate from the interactive admin) with a password generated per installation, and writes it toscripts/mcp/wuic-assistant.credentials.jsonwith a.gitignorenext to it: the file must not be committed or copied to other installations. - The VS Code extension uses the same file when
wuicAssistant.metadataAdminUserandwuicAssistant.metadataAdminPasswordare empty (default); set them only for different credentials. - Upgrading from 1.7.12 or earlier: the fixed password
wuic_assistanthad up to 1.7.12 is refused at login. An administrator sets a new password forwuic_assistantand writes it to the credentials file ({"user":"wuic_assistant","password":"..."}) or to the extension settings.
WuicRagService Service
Typed API over RagController, exported from wuic-framework-lib.
Three main methods:
query(text, {topK, useLora}) -> Observable<RagQueryResponse>chat(text, history, {topK, model}) -> Observable<RagChatResponse>health() -> Observable<RagHealthResponse>reload() -> Observable<{status, ...}>(post RAG rebuild)
*Async variants return Promise via firstValueFrom().
All interfaces RagSource, RagQueryResponse, RagChatResponse,
RagHealthResponse, RagChatTurn are exported.
Claude Model
Default: claude-haiku-4-5-20251001 (fast, native Italian, ~$0.001 per
5-chunk query). Override possible via the component's [model] input
or by passing options.model to the service's chat() method.
System prompt used server-side:
> You are an expert assistant for the WUIC codebase. Answer the user's question
> using EXCLUSIVELY the provided context. If the answer is not in the context,
> reply 'I did not find enough information in the codebase to answer.'
> Always cite relevant files in square brackets in the format
> [file.ext::OptionalSymbol]. Reply in Italian unless another language is
> explicitly requested. Do not make up APIs or method names: if they are not
> in the context, say so explicitly.
Automatic Fallback
When the backend detects one of these conditions, the response includes
mode: 'retrieval-only' + warning + sources:
rag-llm-provider/rag-llm-api-keynot configured inAppSettings- Claude call failed (HTTP error, rate limit, unknown model, etc.)
- request with its own
providerorbaseUrlbut noapiKey(the server key is not used towards providers or hosts chosen by the caller)
The Angular component interprets response.mode === 'retrieval-only' and shows
in the assistant turn a textual summary of the top-K chunks + the warning banner,
so the user still sees useful results.
Runtime Prerequisites
- KonvergenceCore running with
rag-use-dotnet-engine = "true"(exposes/api/Rag/...) rag-engine/folder withWuicRagEngine.dllnext to the app (shipped in release packages; in dev override withrag-engine-dll-path)- Internet access on first start to download models + index (~4.5 GB) from
rag-engine-models-url - Login with a valid
k-usersession cookie - (Optional) NVIDIA GPU with CUDA 12.x + cuDNN 9 (
rag-engine-device = autouses it when present) - (Optional)
rag-llm-provider+rag-llm-api-keyinAppSettingsfor chat mode
Getting Started (first use)
The first time the rag-chatbot route is opened (or at the end of the first
run) the backend starts the engine bootstrap in the background: it downloads
the ONNX models, tokenizer and index (~4.5 GB, 1-5 min depending on
bandwidth) into rag-engine/artifacts/, loads the satellite and warms it up
(~30 s). The user gets an in-app notification at start and at completion;
until then the component shows the "RAG server unreachable" banner with
the RAG offline status. If the backend restarts mid-download, the download
resumes at the next boot.
There is no script to run. The only choices live in appsettings.json
(AppSettings section):
{
"rag-use-dotnet-engine": "true",
"rag-engine-device": "auto",
"rag-engine-profile": "release",
"rag-engine-models-url": "https://wuic-framework.com/rag-models",
"rag-llm-provider": "anthropic",
"rag-llm-api-key": "sk-ant-..."
rag-engine-device:autouses the NVIDIA GPU when found (CUDA 12.x +
cuDNN 9, ~1 s/query), otherwise CPU (~15-25 s/query). rag-engine-cuda-path
points to the CUDA DLL folder when they are not installed system-wide.
rag-llm-provider/rag-llm-api-key/rag-llm-base-url/
rag-llm-default-chat-model: chat provider (anthropic, openai,
openrouter, ollama). Without a provider the chatbot stays retrieval-only.
Keys are hot-reloaded: details on the AppSettings page.
Python fallback (legacy)
scripts/rag-setup.ps1 creates the venv and starts rag_server.py on
127.0.0.1:8765; it is used only with rag-use-dotnet-engine = "false", for
example to split the GPU load onto a dedicated host. It is not shipped in
release packages.
Hot-reload After RAG Rebuild
After regenerating the index/LoRA, just
call:
POST /api/Rag/Reload(via WuicRagService.reload(), administrators only)- or restart the backend (the engine reloads
index/at boot)
to reload the new index without application downtime.
References
- .NET engine:
rag-engine/WuicRagEngine.dll - C# bridge:
/api/Rag/Query,/api/Rag/Chat,/api/Rag/Health