* docs(learning-more): fix broken pipeline and dedup snippets * docs(learning-more): address review feedback on snippets and config defaults - Define sample tasks in concurrency snippet to avoid NameError - Define document_text in batch extraction snippet - Define sample entities in deduplication snippet - Correct GRAPH_STORE_DEFAULT_BACKEND default to neo4j - Replace ineffective SEMANTICA_PORT with SEMANTICA_API_KEY in config table
9.8 KiB
title, description, icon
| title | description | icon |
|---|---|---|
| Learning More | Structured learning paths, configuration reference, troubleshooting, and performance guidance. | graduation-cap |
Whether you're running your first pipeline or deploying Semantica in production, this page gives you a structured path forward: from beginner to enterprise-grade usage.
Learning Paths
- Beginner (1–2 hrs): new to Semantica and knowledge graphs. Start with Installation →
- Intermediate (4–6 hrs): comfortable with basics, building real applications. Start with Modules →
- Advanced (8+ hrs): enterprise deployments, customization, and extension. Start with Architecture →
<Steps>
<Step title="Set up your environment">
[Installation Guide](/installation): virtual environments, optional extras, platform-specific fixes.
</Step>
<Step title="Understand the core ideas">
[Core Concepts](/concepts): what knowledge graphs are, how embeddings work, what extraction does.
</Step>
<Step title="Run your first example">
[Getting Started](/getting-started): 5-minute code walkthrough with pattern-based extraction (no API key needed).
</Step>
<Step title="Build your first knowledge graph">
[Quickstart Tutorial](/quickstart): full 6-step pipeline from ingestion to visualization.
</Step>
<Step title="Explore interactively">
[Welcome to Semantica notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/01_Welcome_to_Semantica.ipynb): Jupyter walkthrough of every module.
</Step>
</Steps>
<Steps>
<Step title="Learn every module">
[Modules Guide](/modules): all 27 modules with code examples and common pipeline chains.
</Step>
<Step title="Build production knowledge graphs">
[Building Knowledge Graphs notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/07_Building_Knowledge_Graphs.ipynb): multi-source, deduplication, conflict resolution.
</Step>
<Step title="Add semantic search">
[Embedding Generation notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/12_Embedding_Generation.ipynb): generating embeddings, provider and model switching, dimensions. Then [Vector Store notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/13_Vector_Store.ipynb): storing and searching vectors for retrieval.
</Step>
<Step title="Multi-source integration">
[Multi-Source Data Integration notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/advanced/06_Multi_Source_Data_Integration.ipynb) for multi-source patterns.
</Step>
</Steps>
<Steps>
<Step title="Understand the architecture">
[Architecture Guide](/architecture): four-layer design, extension points, and design decisions.
</Step>
<Step title="Temporal intelligence">
[Temporal Graphs notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/advanced/10_Temporal_Knowledge_Graphs.ipynb): `valid_from`/`valid_until`, Allen interval algebra, point-in-time queries.
</Step>
<Step title="Ontology-driven knowledge bases">
[Ontology notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/14_Ontology.ipynb): auto-generation, SHACL validation, Ontology Hub.
</Step>
<Step title="Advanced visualization">
[Complete Visualization Suite notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/advanced/03_Complete_Visualization_Suite.ipynb): UMAP, t-SNE, community layouts, embedding projections.
</Step>
<Step title="Enterprise export">
[Multi-Format Export notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/advanced/05_Multi_Format_Export.ipynb): RDF with PROV-O, Parquet, Neo4j Cypher, Arrow, OWL.
</Step>
</Steps>
Configuration Reference
All settings can be overridden with environment variables: no code changes needed.
| Setting | Environment Variable | Default |
|---|---|---|
| OpenAI API Key | OPENAI_API_KEY |
None |
| Groq API Key | GROQ_API_KEY |
None |
| Anthropic API Key | ANTHROPIC_API_KEY |
None |
| Graph Store Backend | GRAPH_STORE_DEFAULT_BACKEND |
"neo4j" |
| Vector Store Backend | VECTOR_STORE_DEFAULT_BACKEND |
"faiss" |
| Server Host | SEMANTICA_HOST |
"127.0.0.1" |
| Server API Key | SEMANTICA_API_KEY |
None |
Troubleshooting
Verify installation and that the correct Python environment is active:
pip list | grep semantica
pip install --upgrade semantica
For optional features, install the relevant extra:
pip install "semantica[llm-openai]" # OpenAI provider
pip install "semantica[gpu]" # GPU acceleration
Set your API key as an environment variable (never hardcode keys in source files):
export OPENAI_API_KEY="sk-..."
export GROQ_API_KEY="gsk_..."
Switch from the default in-memory NetworkX backend to a persistent graph database:
from semantica.graph_store import FalkorDBStore
from semantica.kg import GraphBuilder
store = FalkorDBStore(host="localhost", port=6379)
builder = GraphBuilder(merge_entities=True, graph_store=store)
Also reduce batch sizes and enable streaming ingestion for large corpora.
Enable parallel execution and GPU acceleration:
from semantica.pipeline import ParallelismManager, Task
# Run pipeline tasks concurrently across worker threads
manager = ParallelismManager(max_workers=8)
tasks = [
Task("task_1", lambda: "process part 1"),
Task("task_2", lambda: "process part 2"),
]
results = manager.execute_parallel(tasks)
pip install "semantica[gpu]" # CUDA-backed embeddings
Upgrade to the latest release:
pip install --upgrade semantica
Or install extras individually: pip install semantica, then add [llm-openai], [gpu], etc. as needed.
Set the encoding environment variable:
set PYTHONIOENCODING=utf-8
Performance Optimization
| Operation | NetworkX (default) | Neo4j / FalkorDB |
|---|---|---|
| Graph construction | Fast | Moderate |
| Query performance | Moderate | Fast |
| Scalability | In-memory only | Persistent, production-scale |
| Recommended for | Development, small graphs | Production, large corpora |
Use NetworkX for local development and prototyping. Switch to a persistent backend before deploying to production.
Process documents in batches rather than one at a time. Split large texts into chunks and extract entities in batches:
from semantica.split import TextSplitter
from semantica.semantic_extract import NERExtractor
document_text = "Acme Corp announced record revenue in Seattle. CEO Jane Doe presented results."
splitter = TextSplitter(chunk_size=1000, chunk_overlap=100)
chunks = splitter.split(document_text)
extractor = NERExtractor()
batch_entities = extractor.extract_entities_batch([c.text for c in chunks])
If deduplication is a bottleneck, use candidate blocking to reduce O(n²) comparisons before similarity scoring:
from semantica.deduplication import DuplicateDetector, EntityMerger
entities = [
{"id": "1", "name": "Acme Corp", "type": "Company"},
{"id": "2", "name": "Acme Corporation", "type": "Company"},
{"id": "3", "name": "Globex", "type": "Company"},
]
# Fast candidate blocking for large entity sets
detector = DuplicateDetector(similarity_threshold=0.8)
duplicates = detector.detect_duplicates(entities, candidate_strategy="blocking_v2")
merger = EntityMerger()
merged = merger.merge_duplicates(entities, strategy="keep_most_complete")
The blocking_v2 and hybrid_v2 candidate strategies filter candidate pairs before calculating fine-grained similarity.
Security Best Practices
-
API keys: store in environment variables or a secrets manager; never commit them to version control; rotate on a schedule
-
Sensitive data: use local embedding models (Ollama, HuggingFace) for PII or classified content; avoid sending sensitive data to external APIs without data handling agreements
-
Graph exports: encrypt sensitive exports at rest; use SSRF-safe
base_urlvalidation when configuring custom LLM gateways -
XML ingestion: always use
XMLIngestor, which uses the XXE-safe lxml backend; never parse untrusted XML with the standard library parser -
Cookbook: interactive Jupyter notebooks from beginner to advanced.
-
FAQ: common questions answered.
-
API Reference: complete technical documentation.