Documentation
Getting Started
Learn how to get started with llmkb — create your first space, ingest documents, and query your knowledge base from the web UI, CLI, or any MCP-compatible editor.
Getting Started with llmkb
llmkb turns scattered documents into a connected, queryable knowledge graph — without any manual organization. Drop in PDFs, code repos, meeting recordings, or web clippings, and the platform automatically extracts entities, maps relationships, and generates browsable wiki pages.
Think of it as a living wiki that stays current as your sources change, paired with a knowledge graph your team can search in natural language — from the browser, the terminal, or directly inside their editor via MCP.
What You'll Build
By the end of this guide, you'll have:
- A space with documents ingested and processed
- A knowledge graph with entities and relationships
- Auto-generated wiki pages you can browse
- MCP tools connected to your AI assistant (Claude Code, Cursor, Zed)
- CLI access for sync and search from the terminal
How llmkb Works
llmkb follows a four-layer model inspired by the LLM Wiki pattern:
┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌─────────────┐
│ Source Layer│ ──► │ Wiki Layer │ ──► │Knowledge │ ──► │ Query Layer │
│ Raw files │ │ Markdown │ │ Graph │ │ MCP, Web, │
│ (PDF, code, │ │ pages, cross│ │ (entities, │ │ CLI, API │
│ audio, web)│ │ references │ │ relations) │ │ │
└──────────────┘ └──────────────┘ └──────────────┘ └─────────────┘
- Source Layer — You upload or sync raw documents. They're stored immutably with SHA256 content-addressed deduplication, so unchanged files are never re-processed.
- Knowledge Graph — An LLM reads each source and extracts entities (people, concepts, functions, decisions) and their typed relationships (
calls,extends,depends_on,implements). Each relationship carries a confidence label:extracted(directly from source),inferred(by the LLM), orambiguous. - Wiki Layer — Structured markdown pages are auto-generated from the knowledge graph. Pages cross-reference each other through entity links, creating a navigable documentation set that updates as sources evolve.
- Query Layer — Access your knowledge through multiple channels: the web UI, 31 MCP tools (Claude Code, Cursor, Zed), the CLI, or the REST API.
Ingestion Pipeline
Your documents flow through an async three-stage pipeline:
| Stage | What Happens |
|---|---|
| Parsing | Text extraction — PyMuPDF for PDFs, LibreOffice for office files, Whisper for audio transcription, tree-sitter for code AST |
| Analyzing | Entity extraction, relationship mapping, and vector embedding (4096d via qwen3-embedding-8b) |
| Generating | Wiki page generation, knowledge graph wiring, and embedding storage |
Progress is visible in real-time through the web UI. You can follow along with SSE events from the CLI.
Real-World Use Cases
llmkb is designed for teams who need to make sense of large document collections without manual organization:
| Scenario | How llmkb Helps |
|---|---|
| Onboarding | New hire asks "how does our auth flow work?" — llmkb surfaces the exact wiki page with linked ADRs, not 47 Slack threads from 2024 |
| Debugging Legacy Code | Ingest a monolith's code and architecture docs. Query "who calls PaymentService.validate()?" and get a call graph, not grep output |
| Research & Analysis | Drop 200 pages of research PDFs. llmkb extracts entities, links concepts, and builds a browsable knowledge base |
| AI Assistant Context | Connect llmkb to Claude Code via MCP. Your AI assistant now knows your internal APIs, auth patterns, and deployment process |
| Living Design Docs | ADRs, meeting notes, architecture decisions — ingested, cross-linked, and searchable. SHA256 dedup keeps them current |
Step 1 — Create Your First Space
A space is an isolated knowledge base with its own sources, wiki pages, and knowledge graph. Create one from the web UI at llmkb.ai.
Each space supports multiple team members with four roles:
| Role | Permissions |
|---|---|
| owner | Full control, can delete space, manage members |
| admin | Manage members and settings |
| editor | Upload sources, modify wiki pages |
| viewer | Browse wiki, search, query graph |
After creating a space, you'll get a space UUID and a slug — either one works for the CLI and MCP configuration in later steps.
Step 2 — Get an Access Token
An access token authenticates your MCP client and CLI with the llmkb server.
- Sign in at llmkb.ai
- Navigate to Settings → Access Tokens
- Click Create Token and give it a name (e.g.,
claude-code) - Copy the token — it will look like
llmkb_ut_3qRw70ewabl4BoUm...
Tokens are stored in your OS keychain, not in config files. If you're on a headless server, use the
LLMKB_TOKENenvironment variable instead.
Step 3 — Upload Your First Sources
There are four ways to bring content into a space:
Option A — Upload Files
Drop files directly through the web UI. The ingestion pipeline will:
- PDFs — Extract text with PyMuPDF (handles scanned pages, multi-column layouts)
- Office files (DOCX, XLSX, PPTX) — Convert via LibreOffice, extract text and tables
- Audio (MP3, WAV, M4A) — Transcribe with Whisper, then analyze transcript
- Images — Extract any OCR-readable text
- Code (15 languages) — Parse AST with tree-sitter, extract functions, classes, and structural relationships
Option B — Web Clip
Paste a URL to scrape and ingest a web page. Useful for documentation sites, blog posts, and articles.
Option C — Git Sync
Connect a Git repository to automatically ingest code. Changes are detected via SHA256 — only new or modified files trigger re-processing.
Option D — CLI Sync
Upload from the terminal with content-addressed sync. Only new and changed files are uploaded:
llmkb sync src/ # Sync a directory
llmkb sync --watch # Watch mode — auto-sync on file changes
Progress
You can monitor ingestion progress in real-time through the web UI. Each document shows which stage it's in (parsing → analyzing → generating) and when it completed.
Step 4 — Explore Your Knowledge
Once ingestion completes, you have three ways to interact with your knowledge:
Web UI
The simplest way to explore. Navigate to your space and:
- Browse wiki pages — Generated markdown pages with cross-references between entities
- Explore the knowledge graph — Interactive force-directed layout powered by AntV G6
- Search — Hybrid semantic search combining pgvector vector search with full-text BM25 (fused via Reciprocal Rank Fusion)
- Upload more sources — Drag and drop to keep your knowledge base growing
MCP Tools (Editor Integration)
Connect any MCP-compatible editor to query your knowledge base in real-time. 31 tools are available across seven categories:
| Category | Tools | Example |
|---|---|---|
| Search | space_search, space_search_all, code_search | "search for auth patterns" |
| Read | space_read, space_answer, space_synthesize | "read the page about deployment" |
| Graph | space_graph, space_trace, space_entity_relations | "show me the graph for UserService" |
| Code | code_context, code_graph, code_impact | "what depends on this function?" |
| Browse | space_list_entities, space_list_wiki_pages, space_guide | "list all entities" |
| Write | space_entities, space_write | "create an entity for this concept" |
| Metadata | space_history, space_provenance, space_impact | "what does this affect?" |
Try it: In your editor, ask your AI assistant: "How does our authentication flow work?" — it will query llmkb via MCP and return a synthesized answer with source citations.
CLI
The @llmkb/claude-code CLI lets you sync files and search your knowledge base from the terminal. See Step 5 below for installation.
Step 5 — Set Up the Claude Code Plugin
The @llmkb/claude-code plugin scaffolds your project with MCP config, skills, and hooks so Claude Code can search, read, and write your knowledge base.
Install and Initialize
# 1. Install the CLI globally
npm install -g @llmkb/claude-code
# 2. Move to your project
cd my-project
# 3. Scaffold configuration files
llmkb init
# 4. Authenticate and set the default project space
llmkb login --project-space <slug-or-uuid>
# 5. Verify everything is healthy
llmkb doctor
The llmkb init command creates:
| File | Purpose |
|---|---|
.mcp.json | MCP server registration (HTTP transport for remote, Docker stdio for local) |
.llmkb/spaces.yml | Space configuration (space UUIDs and names, no tokens) |
.llmkb/config.yml | General plugin configuration |
.claude-plugin/plugin.json | Plugin metadata for Claude Code |
Register the MCP Server
For the hosted llmkb service at api.llmkb.ai:
{
"mcpServers": {
"llmkb-prod": {
"description": "LLMKB MCP Server",
"command": "npx",
"args": [
"mcp-remote",
"https://api.llmkb.ai/mcp/",
"--header",
"Authorization:Bearer YOUR_ACCESS_TOKEN"
]
}
}
}
Replace YOUR_ACCESS_TOKEN with the token from Step 2.
For a local development server:
{
"mcpServers": {
"llmkb-local": {
"description": "LLMKB MCP Server",
"command": "npx",
"args": [
"mcp-remote",
"http://localhost:3011/mcp/",
"--header",
"Authorization:Bearer YOUR_ACCESS_TOKEN"
]
}
}
}
Available Commands
| Command | Description |
|---|---|
llmkb init | Scaffold config in your project |
llmkb login --project-space <slug-or-uuid> | Authenticate, set the default project space |
llmkb login --guest | Authenticate without a project space — query public spaces only (read-only) |
llmkb logout | Remove access token from keychain |
llmkb use <slug-or-uuid> | Switch the active project space |
llmkb add --space <slug-or-uuid> | Add a related space |
llmkb remove --space <slug-or-uuid> | Remove a related space |
llmkb update --project-space <slug-or-uuid> | Change the project space |
llmkb spaces | List configured spaces |
llmkb spaces add --public <slug-or-uuid> | Add a public space to a guest session |
llmkb spaces add --list | List public spaces available for guest sessions |
llmkb spaces sync-related | Re-sync related spaces from the server |
llmkb sync [path] | Upload local files (content-addressed, only new/changed) |
llmkb sync --watch | Watch for file changes and auto-sync |
llmkb query "search text" | Search your knowledge base from the terminal |
llmkb doctor | Run diagnostic checks |
llmkb status | Show version stamps and config summary |
Using Skills (Optional)
If you installed with llmkb init --with-skills, guided workflows are available via slash commands:
| Skill | Command | What It Does |
|---|---|---|
| llmkb-query | /llmkb-query | Search with guided follow-up questions |
| llmkb-exploring | /llmkb-exploring | Explore code relationships and execution flows |
| llmkb-admin | /llmkb-admin | Manage spaces, tokens, and configuration |
| llmkb-guide | /llmkb-guide | Space overview — entities, clusters, recent changes |
| llmkb-sync | /llmkb-sync | Check sync status and trigger file synchronization |
How It Works Under the Hood
npx mcp-remote
This acts as a stdio-to-HTTP bridge — Claude Code speaks MCP over stdio, and mcp-remote forwards calls to your llmkb server over HTTP.
--header flags
mcp-remote reads each request's headers from repeated --header "Name:Value" flags in args[]. We emit the colon WITHOUT a surrounding space (e.g. Authorization:Bearer ...) for Windows args-escaping compat — both forms parse identically. Earlier versions of this guide used an MCP_REMOTE_HEADERS env var; mcp-remote does not read that env var, so any .mcp.json written with the env block will silently fail to authenticate. Use --header flags as shown above.
Content-Addressed Storage
Every document is fingerprinted with SHA256 before upload. If a file hasn't changed, it's skipped — no re-uploads, no re-processing. This makes incremental sync fast and efficient.
Headless / CI Environments
In environments without a keychain (CI runners, Linux servers), use environment variables:
export LLMKB_ENDPOINT=https://llmkb.example.com
export LLMKB_SPACE=my-project
export LLMKB_TOKEN=llmkb_ut_abc123...
Guest Mode (Public-Only Access)
For exploration, onboarding, or external projects that don't need a private space, you can log in without creating a project space and query public spaces only. Guest mode is read-only — you can search, read, and explore public knowledge bases, but you can't upload, sync, or write.
# 1. Install the CLI globally (same as Step 5)
npm install -g @llmkb/claude-code
# 2. Move to your project
cd my-project
# 3. Scaffold configuration files
llmkb init
# 4. Log in as a guest (no project space required)
llmkb login --guest
# → Prompts for your account token (or uses LLMKB_TOKEN env var)
# → Fetches the public spaces available to your account
# → Writes the guest .mcp.json + .llmkb/spaces.yml
# 5. Add public spaces you want to query
llmkb spaces add --public llmkb-docs
llmkb spaces add --public nuxt-public-docs
# or list everything available
llmkb spaces add --list
# 6. Verify the guest setup
llmkb doctor
# → Confirms is_guest: true + which public spaces are in scope
# 7. Query from your editor
# "How do I deploy a Nuxt 4 app on Coolify?"
# → Uses MCP tools to read the public nuxt-public-docs space
Guest Mode Invariants
| Property | Value |
|---|---|
| Account | Required (same Better-Auth account as standard login) |
| Project space | Empty — no X-Project-Space header sent |
| Browse public | true — public tier is the only tier available |
| Write access | None — sync, sync --watch, space_write, and space_entities are refused |
| Read access | Full read access to public spaces added via llmkb spaces add --public |
| Rate limit | Same per-bearer rate limit as standard login (no anonymous bucket) |
When to Use Guest Mode
- Documentation-only projects — repos with no internal docs but plenty of public references (a marketing site, an open-source library, a personal sandbox).
- External / open-source contributors — collaborators from outside your tenant who only need read access to public spaces.
- Evaluation & onboarding — kicking the tires before committing to a private space.
If you later need to switch to a full project-based setup, run llmkb login --project-space <slug-or-uuid> --force to overwrite the guest config. The guest's public spaces are lost on overwrite unless you re-add them via llmkb spaces add --public.
Browse public spaces visually: Public spaces have unauthenticated detail, graph, and wiki pages at
/spaces/p/<slug>on your llmkb instance — nollmkb loginrequired, just open the URL in a browser. See M3.6-T13/T14/T15 for the public-space pages.
Typical Workflow
# First-time setup
cd my-project
llmkb init
llmkb login --project-space abc123...
llmkb doctor
# Daily use
llmkb sync src/ # Sync code changes
llmkb query "auth patterns" # Search from the terminal
Once your space has content, the real power comes from the combination of the CLI and MCP:
- Sync your code or documents with
llmkb sync - Browse generated wiki pages in the web UI to verify quality
- Query from your editor — ask Claude Code about your codebase
- Iterate — as your sources change, sync again and the wiki updates automatically
Next Steps
- Upload documents through the web UI to populate your space
- Explore the knowledge graph — switch to the graph view in your space to see entities and relationships
- Connect your editor — follow Step 5 above to integrate with Claude Code, Cursor, or Zed
- Invite team members — go to your space settings and add collaborators with appropriate roles
- Run
llmkb doctorto verify your setup is healthy