Files
semantica/docs/learning-more.md
T
KaifAhmad1 9113ef3428 docs: premium overhaul of all reference pages and core docs
- Rewrote all 26 reference module pages: removed blockquote taglines and
  horizontal rule separators, added "What You Get" bullet summaries,
  added constructor/method parameter tables, expanded thin files
  (graph_store, triplet_store, visualization, provenance) with full API
  coverage, added backend comparison tables and real-world usage patterns
- Renamed Modules tab from "API Reference" and group from "Context &
  Knowledge" to "Context & Intelligence" in docs.json
- Fixed logo: copied "Semantica Logo.png" to web-safe semantica-logo.png
  and updated all 4 references in docs.json
- Improved core docs (index, modules, concepts, quickstart, installation,
  getting-started) with better fonts, bullet points, and complete module
  listings (mcp_server, evals, core, utils previously missing)
- Rewrote community pages (community, community-projects, contributing-guide,
  use-cases, architecture, faq, learning-more, glossary) with heading
  hierarchy fixes, expanded definitions, and better structure
- Fixed markdown linter warnings: MD036 bold-as-heading, MD001 heading
  skips, MD040 missing code fence language, MD032 blank lines around lists
2026-05-23 13:10:09 +05:30

6.9 KiB

title, description, icon
title description icon
Learning More Structured learning paths, configuration reference, troubleshooting, and performance guidance. graduation-cap

Whether you're running your first pipeline or deploying Semantica in production, this page gives you a structured path forward — from beginner to enterprise-grade usage.

Learning Paths

New to Semantica and knowledge graphs. [Start with Installation →](installation) Comfortable with basics, building real applications. [Start with Modules →](modules) Enterprise deployments, customization, and extension. [Start with Architecture →](architecture)

Beginner Path

  1. Installation Guide — set up your environment
  2. Core Concepts — understand KGs, embeddings, and extraction
  3. Getting Started — first working example
  4. Quickstart Tutorial — build your first knowledge graph
  5. Welcome to Semantica notebook — interactive introduction to all modules

Intermediate Path

  1. Modules Guide — every module with code examples
  2. Building Knowledge Graphs notebook
  3. Embeddings notebook
  4. GraphRAG Complete notebook
  5. Multi-Source Data Integration notebook
  6. Use Cases — domain-specific examples with notebooks

Advanced Path

  1. Architecture Guide — three-layer system, extension points, design decisions
  2. Temporal Graphs notebook — v0.4.0 temporal intelligence
  3. Ontology notebook — v0.5.0 Ontology Hub
  4. Complete Visualization Suite notebook
  5. Multi-Format Export notebook

Configuration Reference

All settings can be overridden with environment variables — no code changes needed.

Setting Environment Variable Default
OpenAI API Key OPENAI_API_KEY None
Groq API Key GROQ_API_KEY None
Anthropic API Key ANTHROPIC_API_KEY None
Embedding Provider SEMANTICA_EMBEDDING_PROVIDER "openai"
Graph Backend SEMANTICA_GRAPH_BACKEND "networkx"
Log Level SEMANTICA_LOG_LEVEL "INFO"
Log Format SEMANTICA_LOG_FORMAT "text"

Troubleshooting

ModuleNotFoundError: No module named 'semantica'

Verify installation and that the correct Python environment is active:

pip list | grep semantica
pip install --upgrade semantica

For optional features, install the relevant extra:

pip install "semantica[llm-openai]"   # OpenAI provider
pip install "semantica[gpu]"          # GPU acceleration

AuthenticationError

Set your API key as an environment variable — never hardcode keys in source files:

export OPENAI_API_KEY="sk-..."
export GROQ_API_KEY="gsk_..."

MemoryError or OOM crashes

Switch from the default in-memory NetworkX backend to a persistent graph database:

from semantica.graph_store import FalkorDBStore
from semantica.kg import GraphBuilder

store   = FalkorDBStore(host="localhost", port=6379)
builder = GraphBuilder(merge_entities=True, graph_store=store)

Also reduce batch sizes and enable streaming ingestion for large corpora.

Slow processing on large datasets

Enable parallel execution and GPU acceleration:

from semantica.pipeline import Pipeline

pipeline = Pipeline(workers=8, batch_size=32)
pipeline.run(sources)
pip install "semantica[gpu]"  # CUDA-backed embeddings

Windows [all] installation fails

Fixed in v0.5.0. Upgrade:

pip install --upgrade semantica

Or install extras individually: pip install "semantica[core]", then add [llm-openai], [gpu], etc. as needed.

cp1252 encoding crash on Windows

Fixed in v0.5.0. For earlier versions, pass encoding explicitly or set the environment variable:

set PYTHONIOENCODING=utf-8

Performance Optimization

Backend Selection

Operation NetworkX (default) Neo4j / FalkorDB
Graph construction Fast Moderate
Query performance Moderate Fast
Scalability Low — in-memory only High — persistent
Recommended for Development, small graphs Production, large corpora

Use NetworkX for local development and prototyping. Switch to a persistent backend before deploying to production.

Batch Processing

Process documents in batches rather than one at a time. Configure chunk_size based on available RAM — a good starting point is 1,000 documents per batch on a 16 GB machine.

Deduplication v2

If deduplication is a bottleneck, switch from v1 strategies to v2:

resolver = EntityResolver()
merged   = resolver.resolve(entities, strategy="semantic_v2")  # up to 7x faster

Security Best Practices

  • API keys — store in environment variables or a secrets manager; never commit them to version control; rotate on a schedule
  • Sensitive data — use local embedding models (Ollama, HuggingFace) for PII or classified content; avoid sending sensitive data to external APIs without data handling agreements
  • Graph exports — encrypt sensitive exports at rest; use the v0.5.0 SSRF-safe base_url validation when configuring custom LLM gateways
  • XML ingestion — always use XMLIngestor (v0.5.0), which uses the XXE-safe lxml backend; never parse untrusted XML with the standard library parser
Interactive Jupyter notebooks from beginner to advanced. Common questions answered. Complete technical documentation. Domain-specific examples with notebooks.