Files
semantica/docs/reference/kg.md
T
KaifAhmad1 5eefadaa7f docs: apply full Mintlify component overhaul to all 27 reference pages and concepts.md
Replace plain markdown in every docs/reference/ file and docs/concepts.md with
rich Mintlify JSX components — CardGroup, Steps, Tabs, AccordionGroup, Tip,
Warning, Note, and CodeGroup — for a consistent, navigable, production-grade
developer experience.
2026-05-23 23:02:03 +05:30

14 KiB
Raw Blame History

title, description, icon
title description icon
Knowledge Graph Module Graph construction, temporal models, analytics, and distance intelligence. diagram-project

semantica.kg transforms extracted entities and relationships into structured, queryable knowledge graphs. It includes temporal support, a full suite of graph analytics algorithms, node embeddings, and Distance Intelligence (v0.5.0).

What You Get

Construct graphs from entities and relationships with automatic entity merging. Time-aware edges (`valid_from`/`valid_until`) and point-in-time queries (v0.4.0). Semantic neighborhoods, N×N distance matrices, and distance band classification (v0.5.0). PageRank, degree, betweenness, closeness, and eigenvector centrality. Louvain, Leiden, Label Propagation, and K-Clique community detection. Dijkstra, A\*, BFS, and K-Shortest path algorithms. For conflict detection and advanced entity resolution, use `semantica.conflicts` and `semantica.deduplication` alongside this module.

<img src="/assets/img/diagrams/kg-structure.svg" alt="Knowledge graph entity and relation structure: Person, Organization, Location, Date nodes with typed labeled edges" style={{ width: '100%', borderRadius: '12px', margin: '0 0 24px' }} />

Quick Start

```python from semantica.kg import GraphBuilder
builder = GraphBuilder(merge_entities=True)
kg      = builder.build(entities=entities, relationships=relationships)

print(f"Nodes: {kg.node_count}, Edges: {kg.edge_count}")
```
```python from semantica.kg import CentralityCalculator
calc     = CentralityCalculator()
pagerank = calc.calculate_pagerank(kg, damping_factor=0.85)
top_10   = calc.get_top_nodes(pagerank, top_k=10)

for node_id, score in top_10:
    print(f"  {node_id}: {score:.4f}")
```
```python from semantica.kg import CommunityDetector
detector    = CommunityDetector()
communities = detector.detect_communities(kg, algorithm="louvain")
metrics     = detector.calculate_community_metrics(kg, communities)

print(f"Communities: {len(communities)}")
```
```python from semantica.graph_store import GraphStore
store = GraphStore(backend="neo4j", uri="bolt://localhost:7687",
                   user="neo4j", password="password")
store.add_nodes_bulk(kg.entities,      batch_size=1000)
store.add_edges_bulk(kg.relationships, batch_size=1000)
```

GraphBuilder

Constructs knowledge graphs from extracted entities and relationships:

from semantica.kg import GraphBuilder

builder = GraphBuilder(merge_entities=True)
kg      = builder.build(entities=entities, relationships=relationships)
Method Description
build(sources) Build graph from multiple data sources
build_single_source(data) Build graph from a single data source
merge_entities() Deduplicate and merge entities during construction
Always use `merge_entities=True` in production. Without it, "Steve Jobs" extracted from five different documents creates five separate person nodes. `GraphBuilder(merge_entities=True)` uses edit distance matching to consolidate them at build time.

Temporal Knowledge Graphs (v0.4.0)

Attach valid_from / valid_until time windows to nodes and edges for point-in-time queries and historical analysis:

from semantica.kg import TemporalKnowledgeGraph, TemporalGraphQuery
from datetime import datetime

tkg = TemporalKnowledgeGraph()

tkg.add_node("ceo_role",  valid_from=datetime(2020, 1, 1), valid_until=datetime(2023, 6, 1))
tkg.add_edge(
    "alice", "acme_corp", "ceo_of",
    valid_from=datetime(2020, 1, 1),
    valid_until=datetime(2023, 6, 1)
)

# Point-in-time snapshot
snapshot = tkg.at(datetime(2021, 6, 15))

# Query and diff via TemporalGraphQuery
query         = TemporalGraphQuery(tkg)
snap_2020     = query.query_at_time(datetime(2020, 1, 1))
snap_2023     = query.query_at_time(datetime(2023, 1, 1))
added         = [r for r in snap_2023.relationships if r not in snap_2020.relationships]
print(f"New edges since 2020: {len(added)}")
Edges added without `valid_from`/`valid_until` are treated as **always-valid**. For historical data, always attach timestamps — otherwise point-in-time queries return misleading results.

Distance Intelligence (v0.5.0)

Semantic neighborhood exploration for any entity in the graph:

from semantica.kg import DistanceCalculator

calc = DistanceCalculator(kg)

# Semantic neighborhood of a single node
neighborhood = calc.semantic_neighborhood("Apple Inc.", radius=0.4)

# N×N pairwise distance matrix
matrix = calc.distance_matrix(["Apple Inc.", "Google", "Microsoft"])

# Classify nodes into distance bands: "near" | "mid" | "far"
bands = calc.classify_bands(neighborhood)
`DistanceCalculator` is expensive at large scale — the N×N matrix requires embedding all entities and computing pairwise cosine similarities. Cache the result between runs and only recompute for changed entities.

Graph Analytics

Identify the most structurally important nodes in your graph:
```python
from semantica.kg import CentralityCalculator

calc = CentralityCalculator()

pagerank    = calc.calculate_pagerank(kg, damping_factor=0.85)
degree      = calc.calculate_degree_centrality(kg)
betweenness = calc.calculate_betweenness_centrality(kg)
closeness   = calc.calculate_closeness_centrality(kg)
eigenvector = calc.calculate_eigenvector_centrality(kg)
all_metrics = calc.calculate_all_centrality(kg)

top_nodes = calc.get_top_nodes(pagerank, top_k=10)
```

| Measure | Best For |
| ------- | -------- |
| PageRank | Overall importance (link-based) |
| Degree | Most connected nodes |
| Betweenness | Bridge / bottleneck nodes |
| Closeness | Fastest to reach all others |
| Eigenvector | Connected to other important nodes |
Partition the graph into thematically dense clusters:
```python
from semantica.kg import CommunityDetector

detector = CommunityDetector()

# Louvain — fast, high quality (default)
communities = detector.detect_communities(kg, algorithm="louvain")

# Leiden — higher quality, slower
leiden_communities = detector.detect_communities_leiden(kg, resolution=1.2)

metrics = detector.calculate_community_metrics(kg, communities)
print(f"Communities: {len(communities)}")
```

Algorithms available: **Louvain**, **Leiden**, **Label Propagation**, **K-Clique Communities**.

<Tip>
  Community detection finds thematic clusters — often corresponding to real-world subject groups. Use cluster membership as context boundaries for GraphRAG retrieval.
</Tip>
Find shortest and alternative paths between nodes:
```python
from semantica.kg import PathFinder

finder = PathFinder()

path    = finder.dijkstra_shortest_path(kg, "node_a", "node_b")
paths   = finder.all_shortest_paths(kg, "source", "target")
k_paths = finder.find_k_shortest_paths(kg, "source", "target", k=3)
```

Algorithms: **Dijkstra**, **A\***, **BFS**, **All Shortest Paths**, **K-Shortest Paths**.
Analyse graph structure — components, bridges, and density:
```python
from semantica.kg import ConnectivityAnalyzer

analyzer   = ConnectivityAnalyzer()
components = analyzer.find_connected_components(kg)
density    = analyzer.calculate_density(kg)
bridges    = analyzer.find_bridges(kg)

print(f"Components: {len(components)}, Largest: {len(components[0])} nodes")
print(f"Density:    {density:.4f}")
print(f"Bridges:    {bridges}")
```

| Method | Returns | Description |
| ------ | ------- | ----------- |
| `find_connected_components(kg)` | `List[List[str]]` | Groups of mutually reachable nodes |
| `calculate_density(kg)` | `float` | Edge density (actual / possible edges) |
| `find_bridges(kg)` | `List[str]` | Nodes whose removal disconnects the graph |
Predict which edges are likely missing from the graph:
```python
from semantica.kg import LinkPredictor

predictor = LinkPredictor(method="preferential_attachment")
links     = predictor.predict_links(kg, top_k=20)
score     = predictor.score_link(kg, "node_a", "node_b")
```

Algorithms: **Preferential Attachment**, **Common Neighbors**, **Jaccard**, **Adamic-Adar**, **Resource Allocation**.
Compute structural embeddings for similarity search and downstream ML:
```python
from semantica.kg import NodeEmbedder

embedder      = NodeEmbedder(method="node2vec", embedding_dimension=128)
embeddings    = embedder.compute_embeddings(graph_store, ["Entity"], ["RELATED_TO"])
similar_nodes = embedder.find_similar_nodes(graph_store, "entity_123", top_k=10)
```

Algorithms: **Node2Vec**, **DeepWalk**, **Word2Vec**.

Algorithm Summary

Category Algorithms Use Cases
Node Embeddings Node2Vec, DeepWalk, Word2Vec Structural similarity, node representation
Path Finding Dijkstra, A*, BFS, K-Shortest Route planning, network analysis
Link Prediction Preferential Attachment, Jaccard, Adamic-Adar Network completion
Centrality Degree, Betweenness, Closeness, PageRank Influence analysis
Community Detection Louvain, Leiden, Label Propagation Social clustering
Connectivity Components, Bridges, Density Network robustness

SeedManager

Load and inject curated seed data into a knowledge graph:

from semantica.kg import SeedManager

manager    = SeedManager()
seed_data  = manager.load_seed("seeds/domain_entities.json")
normalized = manager.normalize(seed_data, source="manual_curation_v1")

builder = GraphBuilder(merge_entities=True)
kg      = builder.build(normalized + extracted_sources)

MethodRegistry

Register custom KG construction methods and dispatch by name:

from semantica.kg import method_registry

def my_kg_builder(entities, relationships, **kwargs):
    filtered = [e for e in entities if e["confidence"] >= 0.9]
    return {"entities": filtered, "relationships": relationships}

method_registry.register("build", "high_confidence", my_kg_builder)

from semantica.kg import build_knowledge_graph
kg = build_knowledge_graph(sources, method="high_confidence")

ProvenanceTracker

Track entity and relationship lineage within a knowledge graph:

from semantica.kg import ProvenanceTracker

tracker = ProvenanceTracker()

tracker.track_entity(
    entity_id="apple_inc",
    source="sec_filing_2024q1.pdf",
    source_location="page 3, paragraph 2",
    source_quote="Apple Inc. reported revenue of...",
    confidence=0.98,
)

lineage = tracker.get_lineage("apple_inc")
for entry in lineage.entries:
    print(f"  Source: {entry.source}  ({entry.timestamp})")

For full W3C PROV-O compliance and provenance export, see the Provenance module.

Configuration

kg:
  resolution:
    threshold: 0.9
    strategy: semantic

  temporal:
    enabled: true
    default_validity: infinite

Tips and Common Pitfalls

**Deduplicate before `GraphBuilder`, not after.** It's far easier to merge entities before they become nodes than to update all relationship endpoints after the fact. Run `DuplicateDetector` on extracted entities before calling `builder.build()`. **PageRank identifies your most connected, important nodes.** If you're not sure which entities in your graph are the most structurally significant, `CentralityCalculator.calculate_pagerank()` gives you a ranked list — useful for GraphRAG context anchoring. **Community detection finds thematic clusters.** `CommunityDetector` with Louvain partitions your graph into clusters of densely-connected nodes — often corresponding to real-world thematic groups. Use these clusters for exploratory analysis and to scope GraphRAG retrieval. **`ProvenanceTracker` links entities back to their source documents.** Use it during graph construction so you can always answer "where did this fact come from?" — critical for compliance and for debugging incorrect graph data. Persist graphs in Neo4j, FalkorDB, or Apache AGE. Source of entities and relationships fed to GraphBuilder. Visualize knowledge graphs interactively. Conflict detection and resolution.

Cookbooks