Replace plain markdown in every docs/reference/ file and docs/concepts.md with rich Mintlify JSX components — CardGroup, Steps, Tabs, AccordionGroup, Tip, Warning, Note, and CodeGroup — for a consistent, navigable, production-grade developer experience.
17 KiB
title, description, icon
| title | description | icon |
|---|---|---|
| Deduplication Module | Entity deduplication v1/v2 — similarity scoring, blocking, merging, and cluster-based batch processing. | copy |
semantica.deduplication detects and merges duplicate entities across sources to produce a clean, single-source-of-truth knowledge graph. v2 strategies (blocking_v2, hybrid_v2, semantic_v2) are up to 7x faster than v1 with fine-grained result control.
Why Deduplicate?
Real-world data sources disagree on names. "Apple Inc.", "Apple Computer", and "Apple International Ltd" can all refer to the same company — but if they land in your knowledge graph as separate nodes, every query that should return one entity returns three. Relationships, analytics, and retrieval all degrade.
Deduplication solves this before it reaches the graph:
- Cross-source ingestion — Wikipedia calls it "OpenAI", SEC filings call it "OpenAI, Inc.", your CRM calls it "OpenAI LLC"
- Data entry variation — "Steve Jobs", "Steven P. Jobs", "S. Jobs" are the same person
- Transliteration differences — Cyrillic, Chinese, or Arabic names romanized inconsistently across sources
- Abbreviation drift — "US", "U.S.", "United States", "USA" all mean the same thing
What You Get
Pairwise and batch duplicate detection with configurable strategies and result filtering. Merge duplicate groups with configurable property-level merge policies. Multi-factor similarity: Levenshtein, Jaro-Winkler, cosine, Jaccard, and embedding. Union-Find and hierarchical clustering for large-scale batch deduplication. Reusable per-property merge rule configurations — define once, apply across operations. `blocking_v2`, `hybrid_v2`, `semantic_v2` — up to 7x faster than v1 equivalents.Quick Start
from semantica.deduplication import detect_duplicates, merge_entities
# Detect
duplicates = detect_duplicates(entities, method="hybrid_v2", similarity_threshold=0.85)
# Merge
merged = merge_entities(entities, duplicates, method="keep_most_complete")
print(f"Reduced {len(entities)} → {len(merged)} entities")
Choosing a Strategy
Combines blocking + string similarity + semantic embedding — best default for production:```python
from semantica.deduplication import DuplicateDetector
detector = DuplicateDetector(similarity_threshold=0.85)
duplicates = detector.detect_duplicates(entities, strategy="hybrid_v2")
```
Best for: general production use — handles 95% of real-world cases without GPU. Catches string variants ("Apple Inc." / "Apple Inc") and semantic aliases ("Machine Learning" / "ML").
```python
detector = DuplicateDetector(similarity_threshold=0.85)
duplicates = detector.detect_duplicates(entities, strategy="semantic_v2")
```
Best for: entities that use different words for the same concept ("ML" vs "Machine Learning"), cross-language entity matching, and abbreviation expansion. Requires embedding model.
```python
detector = DuplicateDetector(similarity_threshold=0.85)
duplicates = detector.detect_duplicates(entities, strategy="blocking_v2")
```
Best for: large datasets (500K+ entities) where speed is critical and entities are in the same language. No GPU required.
| Strategy | Algorithm | Speed | Accuracy | Best For |
| -------- | --------- | ----- | -------- | -------- |
| `jaro_winkler` | String similarity only (v1) | Fast | Medium | Small datasets, names, single-source |
| `blocking_v2` | Blocking + Jaro-Winkler | Very fast | Medium | Large datasets, speed-critical, CPU-only |
| `hybrid_v2` | Blocking + string + semantic | Fast | High | General production use — best default |
| `semantic_v2` | Embedding similarity | Medium | Highest | Semantic aliases, cross-language, abbreviations |
**Rules of thumb:**
- Start with `hybrid_v2` — it handles 95% of real-world cases without GPU
- Use `semantic_v2` when entities use different words for the same concept
- Use `blocking_v2` when processing >500k entities and speed matters most
- Never use v1 strategies on new projects — they exist for backwards compatibility only
Threshold Tuning
| Domain | Recommended Threshold | Notes |
|---|---|---|
| Person names | 0.85–0.90 | Names vary a lot; too high misses "Steve" / "Steven" |
| Organization names | 0.80–0.88 | Corporate suffixes create variation; lower threshold helps |
| Product names | 0.88–0.95 | Product names are more stable |
| Medical terms | 0.90–0.95 | High precision required; false merges are dangerous |
| General entities | 0.85 | Safe default starting point |
Start at 0.85, inspect false positives and false negatives, then adjust ±0.05.
DuplicateDetector
**v0.5.0 fix:** `DuplicateDetector` no longer produces duplicate definition errors when the same entity appears in multiple sources with identical definitions.from semantica.deduplication import DuplicateDetector
detector = DuplicateDetector(similarity_threshold=0.85)
duplicates = detector.detect_duplicates(entities)
for dup in duplicates:
print(f"{dup.entity_a} ≈ {dup.entity_b} ({dup.similarity:.2f})")
Fine-grained control over strategy, thresholds, and result size:
duplicates = detector.detect_duplicates(
entities,
strategy="hybrid_v2",
min_similarity=0.85,
top_k_per_entity=3, # max candidates per entity — avoids false-positive floods
max_results=100,
sort_by="similarity", # "similarity" | "entity_id" | "cluster_size"
)
Key behaviours:
- Defaults to
hybrid_v2when no strategy is specified top_k_per_entity=3prevents one entity from flooding results by being a near-match to everything- Pairs are returned once — never both
(A, B)and(B, A) sort_by="similarity"puts highest-confidence duplicates first for faster manual review
EntityMerger
Merge detected duplicate groups into canonical entities, preserving provenance:
from semantica.deduplication import EntityMerger
merger = EntityMerger()
merged_entities = merger.merge_duplicates(
entities,
strategy="keep_most_complete",
preserve_provenance=True,
)
print(f"Merged to: {len(merged_entities)} canonical entities")
Merge Strategies
| Strategy | Behavior | When to Use |
|---|---|---|
keep_first |
Keep the first entity in each group | Source order is meaningful (most authoritative first) |
keep_last |
Keep the most recently seen entity | Most recent source is most accurate |
keep_most_complete |
Keep the entity with the most non-null properties | Default — maximizes data richness |
union |
Merge all properties; combine non-conflicting fields | You want every known alias, tag, and label |
voting |
Most common property value wins | Multiple semi-reliable sources |
Per-Property Merge Rules
from semantica.deduplication import EntityMerger, PropertyMergeRule
merger = EntityMerger(
property_rules={
"name": PropertyMergeRule.KEEP_FIRST,
"aliases": PropertyMergeRule.UNION,
"description": PropertyMergeRule.KEEP_LONGEST,
"confidence": PropertyMergeRule.MAX,
"created_at": PropertyMergeRule.KEEP_FIRST,
"updated_at": PropertyMergeRule.KEEP_LAST,
}
)
merged_entities = merger.merge_duplicates(entities, preserve_provenance=True)
| Rule | Behaviour |
|---|---|
KEEP_FIRST |
Value from the first entity in the group |
KEEP_LAST |
Value from the last entity |
KEEP_LONGEST |
Longest non-null string value |
KEEP_MOST_COMPLETE |
Entity with the most non-null fields (default fallback) |
UNION |
Combine all unique values into a list |
MAX |
Numerically largest value |
MIN |
Numerically smallest value |
VOTING |
Most frequently occurring value |
SimilarityCalculator
Compute multi-factor similarity scores — useful for debugging why two entities were (or were not) detected as duplicates:
from semantica.deduplication import SimilarityCalculator
calc = SimilarityCalculator()
score = calc.calculate_similarity(entity_a, entity_b)
print(score.score) # overall score 0.0–1.0
print(score.components["label"]) # label similarity contribution
print(score.components["embedding"]) # semantic similarity contribution
print(score.components["property"]) # property overlap contribution
lev = calc.levenshtein("Apple Inc.", "Apple Inc")
jaro = calc.jaro_winkler("Steve Jobs", "Steven Jobs")
cos = calc.cosine_similarity(embedding_a, embedding_b)
jacc = calc.jaccard({"founded", "tech"}, {"founded", "technology"})
ClusterBuilder
Build entity clusters for large-scale batch deduplication:
from semantica.deduplication import ClusterBuilder
builder = ClusterBuilder(algorithm="union_find") # or "hierarchical"
result = builder.build_clusters(entities, similarity_threshold=0.85)
print(f"Original: {len(entities)} entities")
print(f"Clusters: {len(result.clusters)} groups")
print(f"Singletons: {result.singleton_count}")
for cluster in result.clusters:
print(f" [{cluster.id}] members={cluster.members} cohesion={cluster.cohesion:.2f}")
Union-Find vs Hierarchical
# Union-Find — O(n·α(n)), scales to millions; use for production
builder = ClusterBuilder(algorithm="union_find")
# Hierarchical — tighter clusters; use when quality matters more than speed
builder = ClusterBuilder(algorithm="hierarchical", linkage="average")
result = builder.build_clusters(entities, similarity_threshold=0.85)
Key behaviours:
- Union-Find groups entities transitively: if A≈B and B≈C, all three land in one cluster even if A and C are only 0.70 similar
- Hierarchical with
linkage="average"avoids chaining by requiring average similarity across all pairs to meet the threshold
Cluster Quality Metrics
print(f"Silhouette score: {result.quality.silhouette_score:.3f}")
# → -1.0 to 1.0; above 0.5 is good, above 0.7 is excellent
print(f"Avg cohesion: {result.quality.avg_cohesion:.3f}")
print(f"Avg separation: {result.quality.avg_separation:.3f}")
MergeStrategyManager
Define complex, reusable merge configurations once and apply them across multiple operations:
from semantica.deduplication import MergeStrategyManager, PropertyMergeRule
manager = MergeStrategyManager()
manager.add_rule("name", PropertyMergeRule.KEEP_FIRST)
manager.add_rule("aliases", PropertyMergeRule.UNION)
manager.add_rule("description", PropertyMergeRule.KEEP_LONGEST)
manager.add_rule("confidence", PropertyMergeRule.MAX)
manager.add_rule("sources", PropertyMergeRule.UNION)
manager.add_rule("created_at", PropertyMergeRule.KEEP_FIRST)
manager.add_rule("updated_at", PropertyMergeRule.KEEP_LAST)
manager.set_default_rule(PropertyMergeRule.KEEP_MOST_COMPLETE)
merged_entity = manager.merge(duplicate_group)
Blocking Strategies
detector = DuplicateDetector(
blocking_strategy="token", # "token" | "phonetic" | "ngram"
blocking_threshold=0.6,
similarity_threshold=0.85
)
| Blocking Strategy | How It Works | Best For |
|---|---|---|
token |
Shared token overlap (default) | General entity names |
phonetic |
Soundex/Metaphone phonetic codes | Names with spelling variations |
ngram |
Character n-gram overlap | Short strings, typos |
Custom Similarity Functions
from semantica.deduplication import method_registry
def drug_name_similarity(entity_a, entity_b) -> float:
compound_a = entity_a.properties.get("active_compound", "")
compound_b = entity_b.properties.get("active_compound", "")
return 1.0 if compound_a == compound_b else 0.0
method_registry.register("similarity", "drug_name", drug_name_similarity)
detector = DuplicateDetector(similarity_method="drug_name", similarity_threshold=0.90)
Schemas
@dataclass
class Cluster:
id: str
members: List[str] # entity IDs in this duplicate group
cohesion: float # mean pairwise similarity within the cluster (0–1)
centroid: str # member ID closest to the cluster centroid
@dataclass
class ClusterResult:
clusters: List[Cluster]
singleton_count: int # entities with no duplicates found
merge_candidates: int # clusters with > 1 member
quality: ClusterQuality
@dataclass
class ClusterQuality:
silhouette_score: float # -1 to 1; above 0.5 is good
avg_cohesion: float # mean within-cluster similarity
avg_separation: float # mean between-cluster distance
End-to-End Pipeline
```python entities = load_entities_from_sources(["crunchbase", "wikipedia", "internal_db"]) print(f"Loaded: {len(entities)} raw entities") ``` ```python from semantica.deduplication import DuplicateDetectordetector = DuplicateDetector(similarity_threshold=0.85)
duplicates = detector.detect_duplicates(entities, strategy="hybrid_v2")
print(f"Found: {len(duplicates)} duplicate pairs")
```
manager = MergeStrategyManager()
manager.add_rule("name", PropertyMergeRule.KEEP_FIRST)
manager.add_rule("aliases", PropertyMergeRule.UNION)
manager.add_rule("description", PropertyMergeRule.KEEP_LONGEST)
manager.set_default_rule(PropertyMergeRule.KEEP_MOST_COMPLETE)
```
merger = EntityMerger()
merged = merger.merge_duplicates(entities, preserve_provenance=True)
print(f"Result: {len(merged)} canonical entities")
for entity in merged:
if hasattr(entity, "source_entities"):
print(f"{entity.label} merged from: {entity.source_entities}")
```