Beyond Vector Search: How Nebula OS Integrates Graph Databases for Deeper Context

Beyond Vector Search: How Nebula OS Integrates Graph Databases for Deeper Context
Beyond Vector Search: How Nebula OS Integrates Graph Databases for Deeper Context

Ask a RAG system "what does this contract say about liability" and vector search works well — it's a similarity problem, and embeddings are good at similarity. Ask it "which of our vendor contracts would be affected if we terminated the Meridian agreement" and vector search quietly fails, not because the answer isn't in the data, but because the question isn't really about similarity at all.

It's about relationships — which documents reference which, which obligations depend on which, which entities connect to which. That's a graph problem wearing a search problem's clothes.

What Vector Search Is Actually Good At — and Where It Runs Out

Vector search finds chunks of text that are semantically close to a query. This is genuinely powerful for a specific class of question: "find me the passage that talks about X," where X can be phrased many different ways and the system needs to match meaning rather than exact keywords. It's the right tool for open-ended, fuzzy retrieval.

It runs out of road on questions that require traversal rather than similarity — multi-hop reasoning where the answer depends on following a chain of relationships the embedding space doesn't represent. "Which of our customers are affected by this vendor's outage" isn't a semantic similarity question; it's a graph traversal: customer → depends on → service → provided by → vendor.

No amount of better embeddings fixes this, because the information a graph traversal needs — explicit, typed relationships between entities — isn't what vector similarity encodes in the first place.

This is the same structural gap we described in the Graph RAG advantage: flat context retrieval treats every chunk as an independent island, when much of the value in enterprise data lives in how those islands connect.

Why Graph Databases Fill the Gap

A graph database stores entities (people, documents, contracts, systems) and typed relationships between them (reports to, supersedes, depends on, references) as first-class structure, rather than leaving those relationships implicit in unstructured text.

Once relationships are explicit, a query can traverse them directly — follow a chain from one entity to another through defined edges — instead of hoping semantic similarity happens to surface the connection.

This matters for a specific, common category of enterprise question: anything that starts with "which," "who else," or "what would be affected if." These questions are inherently relational, and they're exactly the questions flat vector retrieval handles worst and graph traversal handles best.

How Nebula OS Combines the Two

Rather than treating graph and vector as competing retrieval strategies, Nebula OS runs them as complementary layers over the same underlying data:

  • Ingestion builds both representations simultaneously. As documents and structured data enter the system, an extraction pass identifies entities and relationships (this contract references that vendor; this employee reports to that manager) and writes them into the graph layer, while the same content is chunked and embedded into the vector layer. Neither representation is an afterthought bolted onto the other.
  • A query planner decides which layer — or which combination — a given question needs. A fuzzy, open-ended question routes primarily to vector search. A relational question routes primarily to graph traversal. Many real questions need both: vector search to find the relevant starting entities, graph traversal to follow the relationships that matter from there, and vector search again to pull the relevant content attached to whatever the traversal surfaces.
  • Results carry provenance from both layers. When an answer depends on a graph traversal, the system can show the actual path — this contract, superseded by that amendment, which triggered this obligation — rather than presenting a synthesized answer with no visible chain of reasoning. This matters enormously for anything compliance-adjacent, where "how did the system arrive at this answer" is a question someone will eventually ask.

What This Looks Like in Practice

Consider an enterprise knowledge base spanning contracts, org charts, and support tickets. A flat vector index handles "summarize our termination clause" well. It handles "if we lose this key engineer, which active customer commitments have no backup owner" quite badly, because that question requires traversing from employee, to project assignments, to customer commitments, to backup ownership — four hops through explicit relationships that don't live in any single document's embedding.

With a combined vector-and-graph layer, the same question resolves cleanly: vector search identifies the employee record, graph traversal follows the assignment and ownership edges, and the system returns not just an answer but the specific chain of commitments actually at risk — each one traceable back to the relationship that connects it.

The Cost of Getting This Wrong

Teams that rely on vector search alone tend to discover this gap the hard way: the system performs well in early testing (where questions tend to be open-ended and exploratory) and then quietly underperforms in production, where a meaningful share of real user questions turn out to be relational.

The failure mode is subtle — the system doesn't error out, it just returns a plausible-sounding but incomplete answer, because it retrieved the semantically closest chunks rather than the relationally correct ones. This is a harder failure to catch than an outright error, which is exactly why it's worth designing against from the start rather than patching in after users notice.

Conclusion

Vector search and graph traversal aren't competing retrieval paradigms — they answer different categories of question, and most enterprise knowledge bases contain both categories in volume.

A context layer built around vector search alone will always have a blind spot around relational, multi-hop questions, no matter how good the underlying embeddings get.

Building the graph layer in from the start — rather than retrofitting it once the gap becomes visible in production — is what lets a system answer both "what does this say" and "what does this affect" with the same confidence.

Curious how Nebula OS's combined vector-and-graph context layer would handle the relational questions your current retrieval setup struggles with? Book a call with our team or explore the documentation to get started.

Learn more at