The Real Problem With Document Relationships
You've got 200 documents. A contract references a statement of work. That SOW depends on a discovery document. The discovery doc cites three research papers. A client email thread mentions all of them.
Traditional document systems treat each file as an island. They're stored, tagged, maybe full-text indexed. But the relationships between them? That's your problem. You have to remember them, or spend 20 minutes following breadcrumbs.
At AiFiler, we built the knowledge graph to solve this at the architecture level. It's not a feature bolted on top—it's the foundation. This is how we did it.
What We're Actually Building
A knowledge graph is a database of entities and the relationships between them. In AiFiler's case:
- Entities are documents, contacts, projects, and concepts
- Edges are the relationships: "references," "depends on," "authored by," "related to," etc.
- The graph is the queryable structure that connects them
The challenge isn't building a graph. It's building one that:
- Stays fast as you add thousands of documents
- Handles 8 different relationship types without collision
- Lets you query across relationships in real time
- Doesn't require you to manually define connections
We solved this by making the graph emergent—relationships surface automatically from document content, metadata, and user behavior.
The Architecture: Three Layers
┌─────────────────────────────────────────┐
│ Query Layer (Universal Command) │
│ - Intent routing (87 handlers) │
│ - Search parser + heuristics │
└──────────────┬──────────────────────────┘
│
┌──────────────▼──────────────────────────┐
│ Graph Processing Layer │
│ - Edge inference (8 types) │
│ - Relationship scoring │
│ - Real-time updates via webhooks │
└──────────────┬──────────────────────────┘
│
┌──────────────▼──────────────────────────┐
│ Storage Layer (Supabase PostgreSQL) │
│ - Document metadata │
│ - Edge tables (one per relationship) │
│ - Vector embeddings (Anthropic) │
└─────────────────────────────────────────┘
The data flows like this: When you upload a document, the ingest pipeline (lib/ingest/parseFile.ts) extracts text and metadata. That feeds into edge inference, which asks Claude: "What documents does this reference? What does it depend on? Who created it?" Those answers become edges. The query layer then uses those edges to answer questions like "Show me everything related to this contract."
The Edge Types: Eight Ways Documents Connect
We identified eight relationship types that cover 95% of real document workflows:
| Edge Type | Meaning | Example |
|---|---|---|
references | Document A cites Document B | Contract references SOW |
depends_on | Document A requires Document B | Proposal depends on discovery |
authored_by | Document created by person/team | Report authored by Sarah |
related_to | Thematic or topical connection | Two case studies on same client |
version_of | Document B is a version of A | Draft → Final contract |
contains | Document A includes Document B | Folder contains files |
mentions | Document A names/discusses Document B | Email mentions project |
supersedes | Document A replaces Document B | New policy supersedes old |
Each edge type lives in its own Supabase table with the same schema:
CREATE TABLE edges_references (
id UUID PRIMARY KEY,
source_doc_id UUID REFERENCES documents(id),
target_doc_id UUID REFERENCES documents(id),
confidence FLOAT (0-1 score from Claude),
extracted_at TIMESTAMP,
metadata JSONB
);
This design has three advantages:
- No collision: Each relationship type is isolated, so a document can be both "authored by" someone and "references" another doc without ambiguity
- Queryable: You can ask "show me all documents this contract references" with a single table scan
- Scalable: Indexes on (source_doc_id, target_doc_id) make lookups O(log n) even with millions of edges
How Edges Get Created: The Inference Pipeline
When you upload a document via the Knowledge View, here's what happens:
Step 1: Parse the file
The parseFile.ts module extracts text, metadata, and structure. For Word docs, we get paragraphs and formatting. For spreadsheets, we get cell values and relationships. For PDFs, we get text and layout information.
Step 2: Send to Claude with context
We use the Anthropic Files API (lib/ai/fileStore.ts) to upload the document, then ask Claude to identify relationships. The prompt is specific:
Given this document, identify:
1. Other documents it explicitly references (by name, date, or ID)
2. Documents it depends on to be valid
3. The author or team who created it
4. Related documents (same client, same topic, same project)
5. Documents this might supersede or be a version of
For each relationship, provide:
- The document identifier (if known)
- A confidence score (0-1)
- The evidence (quoted text)
Claude's response is structured JSON. We parse it and create edges with the confidence scores.
Step 3: Enrich with vector search
If Claude couldn't find explicit relationships, we fall back to semantic search. We compute embeddings for the document and query similar documents in the vector store. If the similarity is above a threshold (typically 0.75), we create a related_to edge.
Step 4: Store and index Edges are written to their respective tables. Supabase triggers fire webhooks that invalidate caches in the frontend (via SWR with localStorage prefixing for offline-first data fetching).
Why This Matters for Your Queries
When you open Universal Command (Ctrl+Shift+A) and ask "Show me everything related to the Acme contract," here's what happens:
- The intent router recognizes this as a "find related documents" intent
- The search parser extracts "Acme contract" as the query target
- The system finds the contract document
- It queries all eight edge tables where this document is the source OR target
- It returns a ranked list (by confidence score and recency)
This happens in ~200ms, even with 10,000 documents, because:
- Edges are indexed on both source and target
- We cache frequently-accessed subgraphs in memory
- Supabase uses connection pooling to avoid round-trip overhead
The Confidence Score: Why We Don't Trust Claude Blindly
Claude is good at identifying relationships, but it's not perfect. A document might mention another without actually referencing it. So every edge has a confidence score (0-1) that we compute three ways:
- Explicit mention: If Claude finds a direct reference ("See Appendix A: SOW_2024.pdf"), confidence = 0.95
- Semantic similarity: If vector embeddings match, confidence = similarity_score (0.6-0.9)
- Metadata alignment: If documents share project ID, client, or date range, confidence += 0.1
When you query, you can filter by confidence. Universal Command defaults to showing edges above 0.7, but power users can adjust this threshold in settings.
Real-World Example: The Contract Workflow
Let's trace a concrete scenario. You upload a new contract for Acme Corp.
- Ingest: parseFile extracts text, identifies it as a contract, extracts dates and parties
- Inference: Claude identifies that it references a statement of work (SOW) and depends on a discovery document
- Edges created:
referencesedge from contract → SOW (confidence 0.98)depends_onedge from contract → discovery (confidence 0.85)authored_byedge from contract → Legal team (confidence 0.99)
- Enrichment: Vector search finds three related case studies (confidence 0.72 each)
- Storage: All edges written to Supabase, webhooks fire
Now, when a team member opens the contract in the Knowledge View, they see a sidebar showing:
- The SOW it references (one click to open)
- The discovery it depends on
- Related case studies
- The legal team who authored it
No manual linking. No metadata forms. The relationships emerged from the document itself.
What We Learned Building This
Lesson 1: Confidence scores are non-negotiable Early versions treated all edges equally. Users got frustrated when a document that merely mentioned another appeared as "related." Confidence thresholds fixed this—now we show high-confidence edges by default, with an option to see lower-confidence suggestions.
Lesson 2: Edge types must be mutually exclusive We originally had 12 edge types. Some overlapped ("related_to" vs "mentions"). This created ambiguity in queries. We consolidated to 8 and made the definitions precise. Now the system is predictable.
Lesson 3: Real-time updates matter more than you'd think When you upload a document, you want to see its relationships immediately. We use Supabase webhooks to trigger cache invalidation and SWR refetches. The latency dropped from 3-5 seconds to <500ms.
Lesson 4: Don't over-infer We tried creating edges for every semantic connection. The graph became noisy. Now we're conservative: explicit references and high-confidence semantic matches only. Better to miss a connection than show false positives.
The Performance Reality
With 10,000 documents and ~50,000 edges:
- Edge creation: ~2-3 seconds per document (Claude inference + vector search)
- Query latency: <200ms for "show me all related documents"
- Storage footprint: ~50MB for edge tables (metadata + confidence scores)
- Cache hit rate: 85% (most queries are repeats)
We use Supabase's built-in connection pooling and query optimization. The real bottleneck is Claude inference time, not the database. That's why we batch edge creation for bulk uploads and offer an async option in the API.
What This Means for You
The knowledge graph isn't a visualization tool (though we have one). It's infrastructure. It means:
- You don't have to remember relationships. The system finds them and shows them to you
- Your documents are discoverable by context, not just keywords. "Show me everything related to this contract" works because the graph knows what's related
- Bulk operations work smarter. When you move 500 documents to a new project, the graph moves their relationships too
- Search is more intelligent. Universal Command uses the graph to expand queries: searching for a contract also shows related SOWs and discovery docs
This is why we built it into the core architecture instead of bolting it on. Every feature in AiFiler—Batch Operations, Knowledge View, Universal Command—relies on the graph being fast and accurate.
The architecture is built for scale. As you add more documents, the graph gets richer, not slower. That's the design.
Enjoyed this article?
Get more articles like this delivered to your inbox. No spam, unsubscribe anytime.