- Rewrote index.md to match README (tagline, badges, Problem/Solution text) - Improved getting-started, concepts, quickstart, installation, faq, use-cases, contributing, glossary, learning-more, examples, modules, architecture, cookbook, deep-dive pages: tighter prose, fixed headings/bullets, removed inconsistencies and duplicate sections - Removed overuse of emojis from headings in integration pages (docling, snowflake) - Fixed change_management reference page: closed unclosed JSON code block that broke the right TOC, demoted noisy sub-headings to bold text - CSS layout: widened content area (max-width 1440px grid, left sidebar 11rem, right TOC narrowed to 11rem for broader content), tightened TOC spacing and font size, fixed word-wrap/overflow on TOC links - Added mkdocs_local.yml for local serving without mkdocs-jupyter plugin Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
16 KiB
Semantica Cookbook
Interactive Jupyter notebooks covering everything from your first knowledge graph to production GraphRAG systems.
!!! tip "Where to start" - New to Semantica — begin with Core Tutorials - Building an application — see Advanced Concepts or Industry Use Cases - Need installation help — see the Installation Guide
!!! note "Prerequisites" Python 3.8+, Jupyter, and an OpenAI API key (for most examples).
Featured Recipes
-
:material-graph: Your First Knowledge Graph
Go from raw text to a queryable knowledge graph in 20 minutes.
Topics: Extraction, Graph Construction, Visualization · Difficulty: Beginner
-
:material-robot: GraphRAG Complete
Build a production-ready Graph Retrieval Augmented Generation system with hybrid retrieval and logical inference.
Topics: RAG, LLMs, Vector Search, Graph Traversal · Difficulty: Advanced
-
:material-scale-balance: RAG vs. GraphRAG Comparison
Side-by-side benchmark of standard RAG vs. GraphRAG on real-world data.
Topics: RAG, GraphRAG, Benchmarking, Reasoning Gap · Difficulty: Intermediate
-
:material-shield-alert: Real-Time Anomaly Detection
Detect anomalies in streaming data using dynamic knowledge graphs.
Topics: Streaming, Security, Dynamic Graphs · Difficulty: Advanced
Core Tutorials
Essential guides to master the Semantica framework.
-
:material-hand-wave: Welcome to Semantica
An interactive introduction to the framework's core philosophy and all modules including ingestion, parsing, extraction, knowledge graphs, embeddings, and more.
Topics: Framework Overview, Architecture, All Modules
Difficulty: Beginner
-
:material-database-import: Data Ingestion
Techniques for loading data from multiple sources using FileIngestor, WebIngestor, FeedIngestor, StreamIngestor, RepoIngestor, EmailIngestor, DBIngestor, and MCPIngestor.
Topics: File Ingestion, Web Scraping, Database Integration, Streams, Feeds, Repositories, Email, MCP
Difficulty: Beginner
-
:material-file-document-outline: Document Parsing
Extracting clean text from complex formats like PDF, DOCX, and HTML.
Topics: OCR, PDF Parsing, Text Extraction
Difficulty: Beginner
-
:material-broom: Data Normalization
Pipelines for cleaning, normalizing, and preparing text.
Topics: Text Cleaning, Unicode, Formatting
Difficulty: Beginner
-
:material-account-search: Entity Extraction
Using NER to identify people, organizations, and custom entities.
Topics: NER, Spacy, LLM Extraction
Difficulty: Beginner
-
:material-relation-many-to-many: Relation Extraction
Discovering and classifying relationships between entities.
Topics: Relation Classification, Dependency Parsing
Difficulty: Beginner
-
:material-vector-square: Embedding Generation
Creating and managing vector embeddings for semantic search.
Topics: Embeddings, OpenAI, HuggingFace
Difficulty: Intermediate
-
:material-database-search: Vector Store
Setting up vector stores for similarity search and retrieval.
Difficulty: Intermediate
-
:material-database-settings: Graph Store
Persisting knowledge graphs in Neo4j or FalkorDB.
Topics: Neo4j, Cypher, Persistence
Difficulty: Intermediate
-
:material-sitemap: Ontology
Defining domain schemas and ontologies to structure your data.
Topics: OWL, RDF, Schema Design
Difficulty: Intermediate
Advanced Concepts
Deep dive into advanced features, customization, and complex workflows.
-
:material-flask: Advanced Extraction
Custom extractors, LLM-based extraction, and complex pattern matching.
Topics: Custom Models, Regex, LLMs
Difficulty: Advanced
-
:material-chart-network: Advanced Graph Analytics
Centrality, community detection, and pathfinding algorithms.
Topics: PageRank, Louvain, Shortest Path
Difficulty: Advanced
-
:material-brain: Advanced Context Engineering
Build a production-grade memory system for AI agents using persistent Vector (FAISS) and Graph (Neo4j) stores.
Topics: Agent Memory, GraphRAG, Entity Injection, Lifecycle Management
Difficulty: Advanced
-
:material-monitor-dashboard: Complete Visualization Suite
Creating interactive, publication-ready visualizations of your graphs.
Topics: PyVis, NetworkX, D3.js
Difficulty: Intermediate
-
:material-scale-balance: Conflict Resolution
Strategies for handling contradictory information from multiple sources.
Topics: Truth Discovery, Voting, Confidence
Difficulty: Advanced
-
:material-export: Multi-Format Export
Exporting to RDF, OWL, JSON-LD, and NetworkX formats.
Topics: Serialization, Interoperability
Difficulty: Intermediate
-
:material-source-merge: Multi-Source Integration
Merging data from disparate sources into a unified graph.
Topics: Entity Resolution, Merging, Fusion
Difficulty: Advanced
-
:material-pipe: Pipeline Orchestration
Building robust, automated data processing pipelines.
Topics: Workflows, Automation, Error Handling
Difficulty: Advanced
-
:material-brain: Reasoning and Inference
Using logical reasoning to infer new knowledge from existing facts.
Topics: Logic Rules, Inference Engines
Difficulty: Advanced
-
:material-layers: Semantic Layer Construction
Building a semantic layer over your data warehouse or lake.
Topics: Semantic Layer, Data Warehouse
Difficulty: Advanced
-
:material-clock-outline: Temporal Knowledge Graphs
Modeling and querying data that changes over time.
Topics: Time Series, Temporal Logic
Difficulty: Advanced
Industry Use Cases
Real-world examples and end-to-end applications across various industries.
Biomedical
-
:material-pill: Drug Discovery Pipeline
Accelerating drug discovery by connecting genes, proteins, and drugs using PubMed RSS feeds, entity-aware chunking, GraphRAG, and vector similarity search.
Topics: Bioinformatics, KG Construction, GraphRAG, Vector Search
Difficulty: Advanced
-
:material-dna: Genomic Variant Analysis
Analyzing genomic variants and their implications for disease using bioRxiv RSS feeds, temporal knowledge graphs, deduplication, and pathway analysis.
Topics: Genomics, Variant Calling, Temporal KGs, Graph Analytics
Difficulty: Advanced
Finance
-
:material-finance: Financial Data Integration MCP
Merging financial data from Alpha Vantage API, MCP servers, RSS feeds, and market feeds with seed data integration.
Topics: Finance, Data Fusion, MCP Integration, Real-Time Ingestion
Difficulty: Intermediate
-
:material-incognito: Fraud Detection
Identifying fraudulent activities and patterns in transaction networks using temporal knowledge graphs, conflict detection, and pattern recognition.
Topics: Anomaly Detection, Graph Mining, Temporal Analysis, Pattern Detection
Difficulty: Advanced
Blockchain
-
:material-bitcoin: DeFi Protocol Intelligence
Analyzing decentralized finance protocols and transaction flows using CoinDesk RSS feeds, ontology-aware chunking, conflict detection, and ontology generation.
Topics: Blockchain, DeFi, Smart Contracts, Ontology, Conflict Resolution
Difficulty: Advanced
-
:material-network: Transaction Network Analysis
Mapping and analyzing blockchain transaction networks using blockchain APIs, deduplication, and network pattern detection.
Topics: Blockchain Analytics, Network Analysis, Deduplication, Pattern Detection
Difficulty: Advanced
Cybersecurity
-
:material-shield-alert: Real-Time Anomaly Detection
Detecting anomalies in real-time network traffic streams using CVE RSS feeds, Kafka streams, temporal knowledge graphs, and sentence chunking.
Topics: Network Security, Streaming, Temporal KGs, Pattern Detection
Difficulty: Advanced
-
:material-robot-angry: Threat Intelligence Hybrid RAG
Combining enhanced GraphRAG with threat intelligence for security insights using security RSS feeds, entity-aware chunking, deduplication, and temporal knowledge graphs.
Topics: Threat Intelligence, GraphRAG, Security, Hybrid Retrieval
Difficulty: Advanced
Intelligence
-
:material-account-network: Criminal Network Analysis
Analyze criminal networks with graph analytics and key player detection using OSINT RSS feeds, deduplication, and network centrality analysis.
Topics: Forensics, Social Network Analysis, Deduplication, Graph Analytics
Difficulty: Advanced
-
:material-file-search: Intelligence Analysis Orchestrator Worker
Comprehensive intelligence analysis using pipeline orchestrator with multiple RSS feeds, conflict detection, and multi-source integration.
Topics: Intelligence Analysis, Pipeline Orchestration, Multi-Source Integration, Conflict Resolution
Difficulty: Advanced
Renewable Energy
-
:material-wind-turbine: Energy Market Analysis
Analyzing trends and pricing in the renewable energy market using energy RSS feeds, EIA API, temporal knowledge graphs, TemporalPatternDetector, and seed data integration.
Topics: Energy, Time Series, Temporal Analysis, Trend Prediction
Difficulty: Intermediate
Supply Chain
-
:material-truck-delivery: Supply Chain Data Integration
Integrating supply chain data to optimize logistics and reduce risk using logistics RSS feeds, deduplication, and multi-source relationship mapping.
Topics: Logistics, Risk Management, Data Integration, Deduplication
Difficulty: Advanced
How to Run
To run these notebooks locally:
-
Install Semantica from PyPI (recommended):
pip install semantica[all] pip install jupyter -
Or install from source (for development):
git clone https://github.com/Hawksight-AI/semantica.git cd semantica pip install -e .[all] pip install jupyter -
Launch Jupyter:
jupyter notebook
!!! tip "Using Docker"
You can also run the cookbook using Docker:
bash docker run -p 8888:8888 hawksight/semantica-cookbook