Overview

RAG Chatbot

Angular component <wuic-rag-chatbot> that allows querying the WUIC codebase

in natural language, with two operating modes: pure retrieval

(top-K code snippets) and RAG + LLM (response generated by Claude using

the top-K as context). Automatically degrades to retrieval-only if the Claude

API key is not configured on the server.

Architecture

Three layers, no external process:

  • Layer 1 -- .NET RAG engine WuicRagEngine.dll (rag-engine/ folder),

loaded in-process by RagEngineLoader when rag-use-dotnet-engine is

true. With ONNX Runtime (CUDA GPU when available, otherwise CPU) it runs

the hybrid BM25 + bge-m3 + cross-encoder reranker retrieval and the call to

the configured LLM provider. Models, tokenizer and index are downloaded on

first start into rag-engine/artifacts/.

  • Layer 2 -- RagController in KonvergenceCore (/api/Rag/Query,

/api/Rag/Chat, /api/Rag/Health, /api/Rag/Reload), authenticated via

the k-user cookie; it resolves the metadata tools and applies the proposed

actions.

  • Layer 3 -- Angular standalone component <wuic-rag-chatbot> exported

from wuic-framework-lib.

The Python FastAPI server rag_server.py (127.0.0.1:8765) is the legacy

stack: it is still supported as an optional fallback with

rag-use-dotnet-engine = "false", but it is not shipped in release packages.

Page context and dynamic metadata retrieval

When the user chats from an app page, the component injects a **minimal page

context** into the prompt: only the current route/page (e.g. cities/list,

cities/edit, designer). It does NOT inline the column list, SQL identity or

lookup details — inlining them bloated every request regardless of the prompt and

saturated the context window.

The missing details are fetched on-demand by the model via the non-terminal

tool request_metadata_detail, resolved on the backend by

RagController.ResolveMetadataDetail (the engine has no access to the metadata DB).

The model calls the tool, receives the result as a tool_result, and ONLY THEN

emits the propose_* with the real names. The multi-turn loop is handled by the

engine (max 3 retrieval turns).

Supported detail:

detailReturnsWhen
columnsSQL identity (schema/table + full qualifier) and the real column list (name, type, lookup, required, pk, SQL name)the real column names of a route are needed
lookup_columnsjoin_alias, value_field, text_field, related_route, related_columns[] of a lookupByID columncomposing a SQL snippet on a lookup

Join-alias convention (lookup): the WUIC auto-generated query joins the related

table of a lookupByID column with alias <column>_<entity> (e.g. the StateProvinceID

column looking up stateprovinces has alias StateProvinceID_stateprovinces). In SQL

snippets the related table fields are referenced as [<join_alias>].[<col>], e.g.

[StateProvinceID_stateprovinces].[StateProvinceName]. The exact alias value must

always be obtained via request_metadata_detail{detail:'lookup_columns'}, never

deduced by hand.

The mechanism is extensible: new detail kinds (e.g. physical_columns, related_routes,

enum_values) are added in the resolver without touching the Angular component or

bloating the context.

Modes

mode inputBehavior
auto (default)Uses chat if the backend exposes a Claude API key, otherwise retrieval
chatForces RAG + LLM. If the API key is missing, the backend returns mode=retrieval-only with a visible warning
retrievalForces retrieval-only, skips the LLM call (useful to reduce costs)

Inputs

InputTypeDefaultDescription
titlestring'Assistente codebase WUIC'Header label
mode'auto' | 'chat' | 'retrieval''auto'Operating mode
topKnumber5Number of chunks to retrieve from the RAG
showSourcesbooleantrueShows/hides source chips in assistant messages
maxHistorynumber20Maximum number of turns kept in memory
placeholderstring'Chiedi qualcosa...'Input placeholder
modelstring'claude-haiku-4-5-20251001'Claude model for chat mode
chatHeightstring'420px'Fixed chat area height
showClearButtonbooleantrueShows the "Clear history" button

Outputs

OutputPayloadDescription
resultSelectedRagSourceClick on a source chip; the component also attempts a deep-link vscode://file/...
errorOccurred{message, details?}Non-recoverable HTTP errors
turnAddedRagChatbotTurnEmitted after each turn (user or assistant) added to the history

Usage Example

Standalone import in the parent component:

  • HTML selector: <wuic-rag-chatbot mode="auto" [topK]="5" (resultSelected)="onSrc($event)"></wuic-rag-chatbot>
  • TypeScript import: import { WuicRagChatbotComponent, RagSource } from 'wuic-framework-lib';
  • Add WuicRagChatbotComponent to the standalone parent component's imports.

For a complete demo page with debug aside and event handling, see

WuicTest/wwwroot/src/app/component/rag-chatbot-demo-page/.

Auth

All HTTP calls are sent with withCredentials: true and the C# bridge

applies the same authentication checks as other WUIC controllers

(session cookie k-user, rule 10 of AGENTS). The RAG engine runs in-process inside the backend: there is no extra

service to expose or protect.

Endpoint access since 1.7.13:

EndpointAccess
/api/Rag/Query, /api/Rag/Healthno login
/api/Rag/Chatvalid session (guest too, if enableGuest is on)
/api/Rag/MetadataDetailno login for metadata details; sample_records, lookup_value, db_connections, db_databases, db_tables, db_columns administrators only
/api/Rag/Reloadadministrators only

The LLM key configured on the server is never sent to a provider or baseUrl given in the request: whoever overrides them passes their own apiKey, otherwise the answer is retrieval-only.

MCP server wuic-rag and user wuic_assistant

The MCP server scripts/mcp/wuic-rag-mcp.mjs (WUIC Assistant, Claude Code, Cursor) uses these endpoints. When the backend answers 401 it opens a session with MetaService.login, using WUIC_USER / WUIC_PASSWORD from the environment (.mcp.json) or, if missing, from scripts/mcp/wuic-assistant.credentials.json. The login is needed by wuic_ask and, with an administrator user, by wuic_sample_records / wuic_lookup_value; wuic_codebase_search, wuic_route_columns, wuic_route_metadata and wuic_lookup_columns stay without login.

  • The first-run wizard creates the dedicated user wuic_assistant (superadmin, separate from the interactive admin) with a password generated per installation, and writes it to scripts/mcp/wuic-assistant.credentials.json with a .gitignore next to it: the file must not be committed or copied to other installations.
  • The VS Code extension uses the same file when wuicAssistant.metadataAdminUser and wuicAssistant.metadataAdminPassword are empty (default); set them only for different credentials.
  • Upgrading from 1.7.12 or earlier: the fixed password wuic_assistant had up to 1.7.12 is refused at login. An administrator sets a new password for wuic_assistant and writes it to the credentials file ({"user":"wuic_assistant","password":"..."}) or to the extension settings.

WuicRagService Service

Typed API over RagController, exported from wuic-framework-lib.

Three main methods:

  • query(text, {topK, useLora}) -> Observable<RagQueryResponse>
  • chat(text, history, {topK, model}) -> Observable<RagChatResponse>
  • health() -> Observable<RagHealthResponse>
  • reload() -> Observable<{status, ...}> (post RAG rebuild)

*Async variants return Promise via firstValueFrom().

All interfaces RagSource, RagQueryResponse, RagChatResponse,

RagHealthResponse, RagChatTurn are exported.

Claude Model

Default: claude-haiku-4-5-20251001 (fast, native Italian, ~$0.001 per

5-chunk query). Override possible via the component's [model] input

or by passing options.model to the service's chat() method.

System prompt used server-side:

> You are an expert assistant for the WUIC codebase. Answer the user's question

> using EXCLUSIVELY the provided context. If the answer is not in the context,

> reply 'I did not find enough information in the codebase to answer.'

> Always cite relevant files in square brackets in the format

> [file.ext::OptionalSymbol]. Reply in Italian unless another language is

> explicitly requested. Do not make up APIs or method names: if they are not

> in the context, say so explicitly.

Automatic Fallback

When the backend detects one of these conditions, the response includes

mode: 'retrieval-only' + warning + sources:

  • rag-llm-provider / rag-llm-api-key not configured in AppSettings
  • Claude call failed (HTTP error, rate limit, unknown model, etc.)
  • request with its own provider or baseUrl but no apiKey (the server key is not used towards providers or hosts chosen by the caller)

The Angular component interprets response.mode === 'retrieval-only' and shows

in the assistant turn a textual summary of the top-K chunks + the warning banner,

so the user still sees useful results.

Runtime Prerequisites

  • KonvergenceCore running with rag-use-dotnet-engine = "true" (exposes /api/Rag/...)
  • rag-engine/ folder with WuicRagEngine.dll next to the app (shipped in release packages; in dev override with rag-engine-dll-path)
  • Internet access on first start to download models + index (~4.5 GB) from rag-engine-models-url
  • Login with a valid k-user session cookie
  • (Optional) NVIDIA GPU with CUDA 12.x + cuDNN 9 (rag-engine-device = auto uses it when present)
  • (Optional) rag-llm-provider + rag-llm-api-key in AppSettings for chat mode

Getting Started (first use)

The first time the rag-chatbot route is opened (or at the end of the first

run) the backend starts the engine bootstrap in the background: it downloads

the ONNX models, tokenizer and index (~4.5 GB, 1-5 min depending on

bandwidth) into rag-engine/artifacts/, loads the satellite and warms it up

(~30 s). The user gets an in-app notification at start and at completion;

until then the component shows the "RAG server unreachable" banner with

the RAG offline status. If the backend restarts mid-download, the download

resumes at the next boot.

There is no script to run. The only choices live in appsettings.json

(AppSettings section):

Snippet 1JSON
{
  "rag-use-dotnet-engine": "true",
  "rag-engine-device": "auto",
  "rag-engine-profile": "release",
  "rag-engine-models-url": "https://wuic-framework.com/rag-models",
  "rag-llm-provider": "anthropic",
  "rag-llm-api-key": "sk-ant-..."
  • rag-engine-device: auto uses the NVIDIA GPU when found (CUDA 12.x +

cuDNN 9, ~1 s/query), otherwise CPU (~15-25 s/query). rag-engine-cuda-path

points to the CUDA DLL folder when they are not installed system-wide.

  • rag-llm-provider / rag-llm-api-key / rag-llm-base-url /

rag-llm-default-chat-model: chat provider (anthropic, openai,

openrouter, ollama). Without a provider the chatbot stays retrieval-only.

Keys are hot-reloaded: details on the AppSettings page.

Python fallback (legacy)

scripts/rag-setup.ps1 creates the venv and starts rag_server.py on

127.0.0.1:8765; it is used only with rag-use-dotnet-engine = "false", for

example to split the GPU load onto a dedicated host. It is not shipped in

release packages.

Hot-reload After RAG Rebuild

After regenerating the index/LoRA, just

call:

  • POST /api/Rag/Reload (via WuicRagService.reload(), administrators only)
  • or restart the backend (the engine reloads index/ at boot)

to reload the new index without application downtime.

References

  • .NET engine: rag-engine/WuicRagEngine.dll
  • C# bridge: /api/Rag/Query, /api/Rag/Chat, /api/Rag/Health