Skip to content

Stop re-explaining your project to your AI.

llmkb reads your docs, code, and meeting recordings and builds a knowledge graph where every link is labeled by confidence. Claude Code, Cursor, or any MCP client queries it directly. Every answer cites the paragraph it came from.

space://acme/billing 218 sources · synced
validate()AuthServiceSessionGuardTokenRotatorRateLimitAuditLogPIIadr-014.mdsprint-22.mp3vendor-dpa.pdf
MCPspace_answer( "" )

3 × EXTRACTEDcited: adr-014.md §2

Demo: llmkb answers questions about a synced workspace by tracing its knowledge graph and citing sources. Example: "Who calls AuthService.validate()?" — answer: Three callers: AuthService.boot, SessionGuard.require, and TokenRotator.refresh.

MCP tools for AI clients
35
Source types
upload · clip · scrape · git sync · research
Embeddings
pgvector · 4096-d
Access control
RBAC per space

From pile of files to cited answers, in four steps.

The LLM does the reading and the writing. The pipeline does everything else.

  1. 1

    Sources

    Drop in PDFs, repos, recordings, or web clips. Content-hash dedup means only changes re-run.

    PyMuPDF · tree-sitter · faster-whisper · Tesseract

  2. 2

    Graph

    Entities and typed relationships, each labeled EXTRACTED, INFERRED, or AMBIGUOUS.

    pgvector · Leiden communities

  3. 3

    Wiki

    The LLM writes cross-linked markdown pages on top of the graph, versioned per change.

    two-pass generation · per-page history

  4. 4

    Query

    Web UI, JSON:API, or MCP from your AI client. Answers cite the paragraph they came from.

    hybrid retrieval · RBAC enforced

Who uses llmkb

Five jobs where re-finding context eats the week. Each one is a different way in. Pick the one that sounds like your team.

Engineering lead

Onboarding without the archaeology

Ingest the monolith, the ADRs, and the on-call wiki. The new hire asks Claude Code how auth works and gets the real call path, not a plausible guess, with a citation for every hop.

› how does our auth flow work?

space_trace · AuthService.validate 4 hops
  1. 0

    AuthService.boot()

    src/auth/service.py:41

    EXTRACTED
  2. 1

    SessionGuard.require()

    src/auth/middleware.py:88

    EXTRACTED
  3. 2

    TokenRotator.refresh()

    src/auth/token.py:130

    EXTRACTED
  4. 3

    Billing.retryAuth()

    src/billing/retry.py:57

    INFERRED 0.62

every hop cites the file and line it came from

Research analyst

Two hundred PDFs become one map

Drop a folder of papers, reports, or competitor docs. llmkb extracts the methods, datasets, and people, links them across documents, and writes a summary page per paper: searchable, traversable, and sourced.

› which papers use base editing in vivo?

214 papers · one entity map

base editing

method · 7 papers

CRISPR-Cas9in vivoHEK293off-target assaymouse model

cross-paper links no single PDF mentions

Solo founder

An AI pair that actually knows your repo

Sync the codebase, the deploy runbook, and the OpenAPI spec. The next suggestion from your AI follows your conventions and remembers the one production gotcha, because it read them first.

› add a webhook to billing.subscriptions

claude code · mcp: llmkb

› add a webhook to billing.subscriptions

Reading src/billing/webhooks.py. Follows your publish_event() convention.

Found openapi.yaml spec + 2 merged PRs on retry policy.

Draft ready. EXTRACTED · your code, not boilerplate

zero glue code. One .mcp.json entry.

Tech writer

Docs that update when the code does

Connect the docs repo and the SDK source. When a signature changes, the affected wiki pages regenerate incrementally. The old version stays for diff, and every claim still points at its source.

› what changed in the billing API this week?

wiki / billing / webhooks v14v15

Webhook retries

- max_retries: 3

+ max_retries: 5 · commit 8a3f

Regenerated automatically when the SDK changed. v14 kept for diff and audit.

the wiki tracks the source. Not the other way round.

Compliance & legal

A paper trail for every answer

Every claim cites the paragraph it came from, and every relationship is labeled by confidence. Reviewers sign off on what the AI used, and see exactly what it guessed.

› which vendors process PII?

confidence report · vendor overview exportable
EXTRACTEDINFERREDAMBIGUOUS
  • “Acme Analytics is a data processor”vendor-dpa.pdf · p.3 ¶2
  • “PII fields are encrypted at rest”security-whitepaper.md · §5
  • “Retention: 90 days”needs review · AMBIGUOUS

what the AI used vs what it guessed. Visible, per claim.

Why the answers hold up

provenance

Auditable by construction

Every relationship carries EXTRACTED, INFERRED, or AMBIGUOUS, plus a pointer to the source paragraph. Trust is a page feature, not a promise.

mcp-native

Built for AI clients, not retrofitted

35 tools over the Model Context Protocol. Claude Code, Cursor, Zed, Windsurf: connect once, and the agent searches, reads, and traverses the graph itself.

incremental

Never stale, never re-read twice

Content-hash change detection re-processes only what changed, and tree-sitter reads real code structure: calls, imports, extends, not LLM guesses.

If you can read it, llmkb can index it.

Drop a folder, paste a URL, or sync a repo. Every source is parsed by the right tool for the job, deduplicated by content hash, and re-processed only when it changes.

  • PDFs & Office

    PyMuPDF · LibreOffice

  • Code & repos

    tree-sitter AST

  • Audio & video

    faster-whisper

  • Scans & images

    Tesseract OCR

  • Web clips & URLs

    web clipper

  • Git sync

    repo + webhook

Your AI is smart. It just doesn't know your business yet.

Create a space, drop in your first sources, and watch the graph build itself. Free to start, nothing to self-host.