- Fix What's new → link in Info banner (now a proper <a> tag, always clickable) - Replace 4-stat CardGroup on index with inline premium stats row - Convert every <CardGroup>/<Card> block site-wide to markdown bullet lists: content sections → bold-title bullets with sub-bullets, nav cards → [Title](href) — description - Add cursor-animated list item hover effects to custom.css: green inset left border, subtle background tint, marker color change on hover - Affects index, getting-started, quickstart, concepts, modules, faq, architecture, installation, cookbook, glossary, learning-more, explorer-setup, cli-setup, community, contributing-guide, governance, citation, project-license, all integrations pages, and all 20+ reference module pages
8.8 KiB
title, description, icon
| title | description | icon |
|---|---|---|
| Learning More | Structured learning paths, configuration reference, troubleshooting, and performance guidance. | graduation-cap |
Whether you're running your first pipeline or deploying Semantica in production, this page gives you a structured path forward: from beginner to enterprise-grade usage.
Learning Paths
- Beginner (1–2 hrs) — New to Semantica and knowledge graphs. Start with Installation →
- Intermediate (4–6 hrs) — Comfortable with basics, building real applications. Start with Modules →
- Advanced (8+ hrs) — Enterprise deployments, customization, and extension. Start with Architecture →
<Steps>
<Step title="Set up your environment">
[Installation Guide](installation): virtual environments, optional extras, platform-specific fixes.
</Step>
<Step title="Understand the core ideas">
[Core Concepts](concepts): what knowledge graphs are, how embeddings work, what extraction does.
</Step>
<Step title="Run your first example">
[Getting Started](getting-started): 5-minute code walkthrough with pattern-based extraction (no API key needed).
</Step>
<Step title="Build your first knowledge graph">
[Quickstart Tutorial](quickstart): full 6-step pipeline from ingestion to visualization.
</Step>
<Step title="Explore interactively">
[Welcome to Semantica notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/01_Welcome_to_Semantica.ipynb): Jupyter walkthrough of every module.
</Step>
</Steps>
<Steps>
<Step title="Learn every module">
[Modules Guide](modules): all 27 modules with code examples and common pipeline chains.
</Step>
<Step title="Build production knowledge graphs">
[Building Knowledge Graphs notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/07_Building_Knowledge_Graphs.ipynb): multi-source, deduplication, conflict resolution.
</Step>
<Step title="Add semantic search">
[Embeddings notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/09_Embeddings.ipynb): providers, pooling strategies, vector stores.
</Step>
<Step title="Multi-source integration">
[Multi-Source Data Integration notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/advanced/06_Multi_Source_Data_Integration.ipynb) for multi-source patterns.
</Step>
</Steps>
<Steps>
<Step title="Understand the architecture">
[Architecture Guide](architecture): four-layer design, extension points, and design decisions.
</Step>
<Step title="Temporal intelligence">
[Temporal Graphs notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/advanced/04_Temporal_Graphs.ipynb): `valid_from`/`valid_until`, Allen interval algebra, point-in-time queries.
</Step>
<Step title="Ontology-driven knowledge bases">
[Ontology notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/14_Ontology.ipynb): auto-generation, SHACL validation, Ontology Hub (v0.5.0).
</Step>
<Step title="Advanced visualization">
[Complete Visualization Suite notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/advanced/03_Complete_Visualization_Suite.ipynb): UMAP, t-SNE, community layouts, embedding projections.
</Step>
<Step title="Enterprise export">
[Multi-Format Export notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/advanced/05_Multi_Format_Export.ipynb): RDF with PROV-O, Parquet, Neo4j Cypher, Arrow, OWL.
</Step>
</Steps>
Configuration Reference
All settings can be overridden with environment variables: no code changes needed.
| Setting | Environment Variable | Default |
|---|---|---|
| OpenAI API Key | OPENAI_API_KEY |
None |
| Groq API Key | GROQ_API_KEY |
None |
| Anthropic API Key | ANTHROPIC_API_KEY |
None |
| Embedding Provider | SEMANTICA_EMBEDDING_PROVIDER |
"openai" |
| Graph Backend | SEMANTICA_GRAPH_BACKEND |
"networkx" |
| Log Level | SEMANTICA_LOG_LEVEL |
"INFO" |
| Log Format | SEMANTICA_LOG_FORMAT |
"text" |
Troubleshooting
Verify installation and that the correct Python environment is active:
pip list | grep semantica
pip install --upgrade semantica
For optional features, install the relevant extra:
pip install "semantica[llm-openai]" # OpenAI provider
pip install "semantica[gpu]" # GPU acceleration
Set your API key as an environment variable — never hardcode keys in source files:
export OPENAI_API_KEY="sk-..."
export GROQ_API_KEY="gsk_..."
Switch from the default in-memory NetworkX backend to a persistent graph database:
from semantica.graph_store import FalkorDBStore
from semantica.kg import GraphBuilder
store = FalkorDBStore(host="localhost", port=6379)
builder = GraphBuilder(merge_entities=True, graph_store=store)
Also reduce batch sizes and enable streaming ingestion for large corpora.
Enable parallel execution and GPU acceleration:
from semantica.pipeline import Pipeline
pipeline = Pipeline(workers=8, batch_size=32)
pipeline.run(sources)
pip install "semantica[gpu]" # CUDA-backed embeddings
Fixed in v0.5.0. Upgrade:
pip install --upgrade semantica
Or install extras individually: pip install "semantica[core]", then add [llm-openai], [gpu], etc. as needed.
Fixed in v0.5.0. For earlier versions, set the encoding environment variable:
set PYTHONIOENCODING=utf-8
Performance Optimization
| Operation | NetworkX (default) | Neo4j / FalkorDB |
|---|---|---|
| Graph construction | Fast | Moderate |
| Query performance | Moderate | Fast |
| Scalability | In-memory only | Persistent, production-scale |
| Recommended for | Development, small graphs | Production, large corpora |
Use NetworkX for local development and prototyping. Switch to a persistent backend before deploying to production.
Process documents in batches rather than one at a time. Configure chunk_size based on available RAM: a good starting point is 1,000 documents per batch on a 16 GB machine.
from semantica.pipeline import Pipeline
pipeline = Pipeline(workers=8, batch_size=32)
pipeline.run(sources)
If deduplication is a bottleneck, switch from v1 strategies to the v2 engine:
resolver = EntityResolver()
merged = resolver.resolve(entities, strategy="semantic_v2") # up to 7x faster
The blocking_v2, hybrid_v2, and semantic_v2 strategies reduce O(n²) comparisons via candidate blocking before similarity scoring.
Security Best Practices
-
API keys: store in environment variables or a secrets manager; never commit them to version control; rotate on a schedule
-
Sensitive data: use local embedding models (Ollama, HuggingFace) for PII or classified content; avoid sending sensitive data to external APIs without data handling agreements
-
Graph exports: encrypt sensitive exports at rest; use the v0.5.0 SSRF-safe
base_urlvalidation when configuring custom LLM gateways -
XML ingestion: always use
XMLIngestor(v0.5.0), which uses the XXE-safe lxml backend; never parse untrusted XML with the standard library parser -
Cookbook — Interactive Jupyter notebooks from beginner to advanced.
-
FAQ — Common questions answered.
-
API Reference — Complete technical documentation.