You've got 5,000 documents. A client asks, "Show me everything related to the Q3 budget review." Most tools would search for "Q3 budget review," return 200 results, and call it a day. AiFiler does something different: it traverses relationships.
A document isn't just a file. It's a node in a graph. And that node connects to other nodes through edges—relationships that encode meaning. A contract relates to a client. A meeting note relates to a decision. A budget spreadsheet relates to a fiscal year. When you ask for "everything related to Q3 budget," the knowledge graph doesn't search; it walks.
This is the architecture that makes AiFiler's retrieval work at scale. Here's how we built it.
The Problem With Traditional Document Organization
Before we designed the knowledge graph, we asked ourselves a hard question: why do document tools fail at retrieval?
The answer: they treat documents as isolated objects. You store them, tag them, maybe full-text search them. But documents don't exist in isolation. They're connected. A contract references a statement of work. A meeting note references a decision. A budget revision references a previous budget. Those relationships are where meaning lives.
Traditional tools force you to encode those relationships manually—through folder hierarchies, naming conventions, or metadata fields. It doesn't scale. By the time you've manually tagged 500 documents, you've already lost the thread.
The knowledge graph inverts that problem. Instead of asking "How do I organize this document?" we ask "What does this document relate to?" The relationships emerge from the document content itself, and the graph becomes a queryable index of those relationships.
The Edge Model: Eight Types of Connection
At the core of AiFiler's knowledge graph is the edge model. An edge is a directed relationship between two nodes (documents, people, projects, dates, whatever). We defined eight edge types that cover 95% of real-world document relationships:
- REFERENCES: Document A cites or quotes Document B
- RELATES_TO: General thematic connection (no stronger relationship applies)
- DEPENDS_ON: Document A requires Document B to be understood or executed
- SUPERSEDES: Document A replaces or updates Document B
- AUTHORED_BY: Document A was created by Person B
- ASSIGNED_TO: Document A is assigned to Person B or Project B
- CREATED_FOR: Document A was created for Client B or Purpose B
- DATED: Document A is associated with Time Period B
Why eight, not three or thirty? We tested. Fewer types and you lose semantic precision—everything becomes "related." More types and you hit the real problem: the AI model that infers edges gets confused. Eight types is the sweet spot where the inference model (Claude) stays accurate without hallucinating relationships.
How Edges Get Inferred
When you upload a document to AiFiler, it doesn't immediately become part of the graph. Here's what actually happens:
-
Parsing: The document is parsed (DOCX, XLSX, PDF, PPTX) using our file parsing system in
lib/ingest/parseFile.ts. We extract text, metadata, and structure. -
Chunking: The parsed content is split into semantic chunks—not arbitrary 512-token windows, but coherent sections that preserve meaning. A section header, its content, and relevant metadata stay together.
-
Inference: Each chunk is sent to Claude (via
lib/ai/client.ts) with a system prompt that asks: "What edges should this chunk create?" The model responds with a JSON object listing inferred edges, their types, and confidence scores. -
Deduplication & Storage: We deduplicate edges (if two chunks suggest the same relationship, we keep one), score them by confidence, and store them in Supabase. The edge table includes:
source_id(the document that makes the claim)target_id(the document/entity being related to)edge_type(one of the eight types)confidence(0-1 score from the inference model)metadata(the text snippet that justified the edge)
The key architectural decision: we store the evidence (the metadata field). This lets us show users why a relationship exists. Click on an edge in the UI, and you see the sentence that justified it.
The Query Path: From Intent to Results
When you use Universal Command (Ctrl+Shift+A) to ask "Show me all documents related to the Q3 budget," here's what happens:
-
Intent Parsing: The query is parsed by
lib/intentHeuristics.tsandlib/searchParser.ts. The system recognizes this as a graph traversal intent, not a text search. -
Entity Recognition: The system identifies "Q3 budget" as an entity. It searches the graph for nodes matching that entity (either by document name, content, or inferred metadata).
-
Traversal: Starting from the matched node, the system traverses outbound edges. It collects all documents connected via REFERENCES, RELATES_TO, DEPENDS_ON, and CREATED_FOR edges (configurable by intent type).
-
Ranking: Results are ranked by:
- Edge confidence (high-confidence edges bubble up)
- Traversal depth (direct connections rank higher than second-degree)
- Recency (newer documents rank higher, configurable)
-
Return: Results are returned to the UI with edge metadata displayed, so you see not just the documents, but how they're related.
This entire flow—from intent to ranked results—takes <200ms on a typical 5,000-document workspace. We achieve this through:
- Indexed queries: The edge table is indexed on
source_id,target_id, andedge_type. Traversal queries hit those indexes. - Lazy loading: We don't load all edges upfront. We load the first degree of relationships, then load deeper levels on demand.
- Caching: Frequently traversed paths are cached in Redis (via Supabase's connection pooling).
Why Confidence Matters
Not every inferred edge is correct. Claude is smart, but it hallucinates. A document about "Q3 strategy" might get incorrectly linked to "Q3 budget" if the inference prompt isn't precise.
We handle this with confidence scoring. Every edge has a 0-1 confidence score assigned by the inference model. In the UI, you can filter results by minimum confidence:
- 0.8+: High confidence. These relationships are almost certainly correct.
- 0.6-0.8: Medium confidence. Likely correct, but worth a quick glance.
- <0.6: Low confidence. Shown only if the user explicitly requests them.
This is why the metadata field matters. When you see an edge with 0.65 confidence, you can read the snippet that justified it and decide: "Yeah, that's related" or "Nope, false positive."
The Edge Type Distribution
In a typical AiFiler workspace, the edge distribution looks like this:
- RELATES_TO: 40% (the catch-all for thematic connections)
- REFERENCES: 25% (explicit citations and quotes)
- DEPENDS_ON: 15% (prerequisite relationships)
- CREATED_FOR: 10% (purpose-based relationships)
- SUPERSEDES: 5% (version and update relationships)
- AUTHORED_BY, ASSIGNED_TO, DATED: 5% combined
This distribution tells us something important: most relationships are thematic, not structural. That's why we didn't try to build a rigid schema. The graph needs flexibility.
The Scaling Challenge: Inference at 10K Documents
Here's where the architecture gets interesting. When you have 10,000 documents, you don't want to re-infer edges every time you upload a new document. That's prohibitively expensive.
Our approach:
-
Incremental inference: Only new documents get full edge inference. Existing documents' edges are static.
-
Bidirectional linking: When Document A is inferred to relate to Document B, we also create a reverse edge (B relates to A). This doubles storage but halves query time.
-
Batch processing: Edge inference runs in a background job queue (via Supabase's pg_cron extension). High-priority workspaces get inference within minutes; lower-priority ones within hours.
-
Pruning: We periodically remove low-confidence edges (confidence <0.4) that haven't been accessed in 30 days. This keeps the graph lean.
At 10,000 documents with 8 edges per document on average, you're looking at 80,000 edges. Our Supabase instance handles that comfortably. At 100,000 documents, we'd need to shard the edge table by workspace. We haven't hit that threshold yet, but the architecture is designed for it.
Why This Matters for You
The knowledge graph isn't an abstract data structure. It's what makes three concrete things possible:
1. Retrieval without search: You don't need to remember keywords. You ask "Show me everything related to X," and the graph walks the relationships. This is faster and more accurate than full-text search for relationship-heavy queries.
2. Discovery: When you open a document, the sidebar shows related documents (via the graph). You discover connections you didn't know existed. That's how you find that budget revision you forgot about, or the client email that clarifies a contract clause.
3. Audit trails: Because edges store metadata (the text that justified them), you can trace why two documents are connected. This matters for compliance and for sanity-checking the AI's work.
The architecture is built for scale, but it's also built for transparency. You can see the edges, understand them, and trust them.
What's Next
We're exploring two extensions to the edge model:
- Weighted edges: Instead of binary relationships, edges could carry weights (how strongly related are these documents?). This would improve ranking.
- Temporal edges: Edges could encode time relationships ("Document A was revised after Document B"). This would help with version tracking and audit trails.
Both would add complexity. We're testing them on a subset of workspaces before rolling them out broadly.
For now, the eight-edge model works. It's simple enough to be maintainable, expressive enough to capture real relationships, and fast enough to query at scale. That's the goal of any good architecture: solve the problem you have, not the problem you might have someday.
Enjoyed this article?
Get more articles like this delivered to your inbox. No spam, unsubscribe anytime.