- Rewrote index.md to match README (tagline, badges, Problem/Solution text) - Improved getting-started, concepts, quickstart, installation, faq, use-cases, contributing, glossary, learning-more, examples, modules, architecture, cookbook, deep-dive pages: tighter prose, fixed headings/bullets, removed inconsistencies and duplicate sections - Removed overuse of emojis from headings in integration pages (docling, snowflake) - Fixed change_management reference page: closed unclosed JSON code block that broke the right TOC, demoted noisy sub-headings to bold text - CSS layout: widened content area (max-width 1440px grid, left sidebar 11rem, right TOC narrowed to 11rem for broader content), tightened TOC spacing and font size, fixed word-wrap/overflow on TOC links - Added mkdocs_local.yml for local serving without mkdocs-jupyter plugin Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
6.1 KiB
Deep Dive
Internals, advanced concepts, and extension points for contributors and power users.
!!! tip "Just getting started?" Read Architecture for a higher-level overview first.
Pipeline Internals
The full data flow through a Semantica pipeline:
graph TB
A[Data Sources] --> B[Ingestion Layer]
B --> C[Parsing Layer]
C --> D[Extraction Layer]
D --> E[Normalization Layer]
E --> F[Conflict Resolution]
F --> G[Knowledge Graph Builder]
G --> H[Embedding Generator]
H --> I[Export Layer]
D --> D1[Entity Extractor]
D --> D2[Relationship Extractor]
D --> D3[Triplet Extractor]
G --> G1[Graph Validator]
G --> G2[Graph Analyzer]
H --> H1[Text Embeddings]
H --> H2[Graph Embeddings]
Sequence Diagram
sequenceDiagram
participant User
participant Semantica
participant Ingestor
participant Parser
participant Extractor
participant Resolver
participant GraphBuilder
participant Exporter
User->>Semantica: build_knowledge_base(sources)
Semantica->>Ingestor: ingest(sources)
Ingestor->>Parser: parse(documents)
Parser->>Extractor: extract(text)
Extractor->>Resolver: resolve_conflicts(entities)
Resolver->>GraphBuilder: build_graph(resolved_data)
GraphBuilder->>Exporter: export(graph)
Exporter->>User: return result
System Components
Ingestion Layer
Handles input from any source:
- FileIngestor — PDF, DOCX, HTML, JSON, CSV, archives
- WebIngestor — URL crawling and scraping
- DBIngestor / SnowflakeIngestor — SQL databases
- StreamIngestor — Kafka and real-time feeds
Parsing Layer
Converts raw data to structured text:
- Text and metadata extraction from documents
- OCR for scanned content
- Layout analysis (via Docling for tables and columns)
Extraction Layer
Core semantic processing pipeline:
text → Tokenization → NER → Entity Linking → Entity Validation
Components: Named Entity Recognition, Relationship Extraction, Triplet Extraction, Coreference Resolution.
Normalization Layer
Standardizes extracted data: entity names, date formats, numbers, encodings, and language normalization.
Conflict Resolution
Handles contradictory facts from multiple sources:
graph LR
A[Multiple Sources] --> B[Conflict Detection]
B --> C{Resolution Strategy}
C --> D[Voting]
C --> E[Credibility Weighted]
C --> F[Most Recent]
C --> G[Highest Confidence]
D --> H[Resolved Entity]
E --> H
F --> H
G --> H
Knowledge Graph Builder
- Entity resolution across sources
- Edge creation (typed relationships)
- Property assignment with confidence scores
- Graph validation and quality checks
Embedding Generator
- Text embeddings (Sentence-Transformers, FastEmbed, OpenAI, BGE)
- Graph embeddings (Node2Vec, GraphSAGE)
Advanced Concepts
Entity Resolution Algorithm
def resolve_entities(entities, threshold=0.85):
clusters = []
for entity in entities:
matched = False
for cluster in clusters:
if similarity(entity, cluster.representative) > threshold:
cluster.add(entity)
matched = True
break
if not matched:
clusters.append(EntityCluster(entity))
return clusters
Relationship Inference
Semantica's reasoning engines can derive implicit relationships:
- Transitive — if A→B and B→C, infer A→C
- Temporal — before, after, during from timestamped facts
- Causal — IF/THEN rules via
Reasoner - Hierarchical — subclass/instance inference via
OntologyReasoner
Batch Processing for Large Datasets
def process_large_dataset(sources, batch_size=100):
for i in range(0, len(sources), batch_size):
batch = sources[i : i + batch_size]
result = semantica.build_knowledge_base(batch)
save_result(result)
del result
gc.collect()
Extension Points
Custom Plugin
from semantica.core import Plugin
class CustomPlugin(Plugin):
def process(self, data):
# Your custom processing logic
return processed_data
Custom Extractor
from semantica.semantic_extract import BaseExtractor
class DomainSpecificExtractor(BaseExtractor):
def extract(self, text):
# Domain-specific entity extraction logic
return entities
Custom Ingestor
from semantica.ingest import BaseIngestor
class CustomIngestor(BaseIngestor):
def ingest(self, source):
# Load and return document dicts
return documents
Internal APIs
| API | Purpose |
|---|---|
Semantica.build_knowledge_base() |
Main orchestration entry point |
GraphBuilder.build() |
Graph construction |
ConflictResolver.resolve() |
Conflict resolution |
EmbeddingGenerator.generate() |
Embedding generation |
Extension hooks: plugin registration, custom extractor registration, custom exporter registration, event hooks.
Design Decisions
Why modular architecture? Each component is independently testable and swappable. You can use NERExtractor alone without pulling in graph storage or pipelines.
Why built-in conflict resolution? Multi-source data always has contradictions. Ignoring them produces garbage graphs. Explicit resolution strategies give you control over data quality.
Why W3C PROV-O for provenance? It's an industry standard with tooling support. Using a custom format would make lineage data non-portable.
Why multiple reasoning engines? Different problems need different reasoning: forward chaining for rule application, SPARQL for graph queries, abductive for hypothesis generation. No single engine fits all cases.
Further Reading
- Architecture — high-level three-layer overview
- Modules — every module with code examples
- API Reference — complete technical reference
- Contributing — how to extend the framework