Files
semantica/docs/cookbook/introduction/Your_First_Knowledge_Graph.ipynb
T
KaifAhmad1 89a64a483b docs: enhance documentation structure and notebook formatting
- Update mkdocs.yml navigation to use new cookbook index
- Create comprehensive docs/cookbook.md index
- Refactor Welcome_to_Semantica.ipynb to use Markdown cells
- Enhance markdown formatting in Your_First_Knowledge_Graph.ipynb and Financial_Data_Integration.ipynb
- Update custom.css and version-selector.js for better styling
- Populate modules.md, concepts.md, and getting-started.md with detailed content
2025-11-23 14:17:36 +05:30

302 lines
13 KiB
Plaintext

{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# 🚀 Your First Knowledge Graph\n",
"\n",
"## Overview\n",
"\n",
"This notebook walks you through creating your first knowledge graph from a simple document. You'll learn the complete end-to-end workflow from ingesting a file to visualizing the resulting knowledge graph.\n",
"\n",
"> [!TIP]\n",
"> This is the perfect starting point if you are new to Semantica. No prior knowledge of knowledge graphs is required!\n",
"\n",
"### 🎯 Learning Objectives\n",
"\n",
"- **Understand the Workflow**: Learn the `File → Parse → Extract → Graph` pipeline\n",
"- **Ingest Data**: Load documents using `FileIngestor`\n",
"- **Parse Content**: Extract text using `DocumentParser`\n",
"- **Extract Knowledge**: Identify entities using `NERExtractor`\n",
"- **Build Graph**: Construct a graph using `GraphBuilder`\n",
"- **Visualize**: See your graph come to life with `KGVisualizer`\n",
"\n",
"---\n",
"\n",
"## 🔄 Simple End-to-End Workflow\n",
"\n",
"The complete workflow consists of four main steps:\n",
"\n",
"1. **📥 Ingest** - Load data from files or other sources\n",
"2. **📄 Parse** - Extract and structure content from documents\n",
"3. **⛏️ Extract** - Identify entities and relationships\n",
"4. **🕸️ Build Graph** - Construct the knowledge graph\n",
"\n",
"Each step is demonstrated in the code cells below.\n",
"\n",
"---\n",
"\n",
"## 📂 Step 1: Ingest a File\n",
"\n",
"In this step, we'll use `FileIngestor` to load a document. The ingestor supports various file formats including PDF, DOCX, TXT, and more.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from semantica.ingest import FileIngestor\n",
"from pathlib import Path\n",
"\n",
"# Initialize the ingestor\n",
"ingestor = FileIngestor()\n",
"\n",
"# Create a sample document for demonstration\n",
"sample_text = \"\"\"\n",
"Apple Inc. is a technology company founded by Steve Jobs, Steve Wozniak, and Ronald Wayne in 1976.\n",
"The company is headquartered in Cupertino, California.\n",
"Tim Cook is the current CEO of Apple Inc.\n",
"Apple designs and manufactures consumer electronics, software, and online services.\n",
"\"\"\"\n",
"\n",
"sample_file = Path(\"sample_document.txt\")\n",
"sample_file.write_text(sample_text)\n",
"\n",
"print(\"Sample document created:\")\n",
"print(f\"File: {sample_file}\")\n",
"print(f\"Content length: {len(sample_text)} characters\")\n",
"\n",
"# Ingest the file\n",
"try:\n",
" file_object = ingestor.ingest_file(sample_file, read_content=True)\n",
" print(f\"\\n✓ File ingested successfully!\")\n",
" print(f\" File name: {file_object.name}\")\n",
" print(f\" File type: {file_object.file_type}\")\n",
" print(f\" Content available: {file_object.content is not None}\")\n",
"except Exception as e:\n",
" print(f\"\\n✗ Error ingesting file: {e}\")\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## 📄 Step 2: Parse the Document\n",
"\n",
"After ingesting the file, we need to parse it to extract the text content. The `DocumentParser` handles various file formats and extracts structured content.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from semantica.parse import DocumentParser\n",
"\n",
"parser = DocumentParser()\n",
"\n",
"try:\n",
" # Parse the document to extract text\n",
" if 'file_object' in locals():\n",
" parsed_content = parser.parse_document(str(sample_file))\n",
" print(\"✓ Document parsed successfully!\")\n",
" print(f\" Parsed content length: {len(parsed_content) if parsed_content else 0} characters\")\n",
" print(f\" Preview: {parsed_content[:200] if parsed_content else 'N/A'}...\")\n",
" else:\n",
" # Fallback if ingestion failed\n",
" parsed_content = parser.parse_document(str(sample_file))\n",
" print(\"✓ Document parsed successfully!\")\n",
" print(f\" Parsed content length: {len(parsed_content) if parsed_content else 0} characters\")\n",
"except Exception as e:\n",
" print(f\"✗ Error parsing document: {e}\")\n",
" parsed_content = sample_text\n",
" print(\"Using raw text as fallback\")\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## ⛏️ Step 3: Extract Entities\n",
"\n",
"Now we'll extract entities from the parsed text using Named Entity Recognition (NER). This identifies people, organizations, locations, dates, and other entities in the text.\n",
"\n",
"> [!NOTE]\n",
"> In a real scenario, you would use `NERExtractor` with an LLM or model backend. Here we simulate the output for demonstration purposes.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from semantica.semantic_extract import NamedEntityRecognizer, NERExtractor\n",
"\n",
"try:\n",
" ner = NamedEntityRecognizer()\n",
" extractor = NERExtractor()\n",
" \n",
" print(\"Extracting entities from text...\")\n",
" print(f\"\\nText: {parsed_content[:100]}...\")\n",
" \n",
" # Simulated extraction results\n",
" expected_entities = [\n",
" {\"text\": \"Apple Inc.\", \"type\": \"Organization\", \"start\": 0, \"end\": 10},\n",
" {\"text\": \"Steve Jobs\", \"type\": \"Person\", \"start\": 50, \"end\": 60},\n",
" {\"text\": \"Steve Wozniak\", \"type\": \"Person\", \"start\": 62, \"end\": 75},\n",
" {\"text\": \"Ronald Wayne\", \"type\": \"Person\", \"start\": 81, \"end\": 93},\n",
" {\"text\": \"1976\", \"type\": \"Date\", \"start\": 97, \"end\": 101},\n",
" {\"text\": \"Cupertino, California\", \"type\": \"Location\", \"start\": 130, \"end\": 151},\n",
" {\"text\": \"Tim Cook\", \"type\": \"Person\", \"start\": 153, \"end\": 161},\n",
" ]\n",
" \n",
" print(f\"\\n✓ Found {len(expected_entities)} entities:\")\n",
" for entity in expected_entities:\n",
" print(f\" - {entity['text']} ({entity['type']})\")\n",
" \n",
"except Exception as e:\n",
" print(f\"✗ Error extracting entities: {e}\")\n",
" expected_entities = []\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## 🕸️ Step 4: Build the Knowledge Graph\n",
"\n",
"Using the extracted entities and relationships, we'll construct a knowledge graph. The graph represents entities as nodes and relationships as edges.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from semantica.kg import GraphBuilder\n",
"import networkx as nx\n",
"\n",
"builder = GraphBuilder()\n",
"\n",
"# Prepare data for graph construction\n",
"entities_data = [\n",
" {\"id\": f\"entity_{i}\", \"name\": entity[\"text\"], \"type\": entity[\"type\"]}\n",
" for i, entity in enumerate(expected_entities)\n",
"]\n",
"\n",
"relationships_data = [\n",
" {\"source\": \"entity_0\", \"target\": \"entity_1\", \"type\": \"founded_by\"},\n",
" {\"source\": \"entity_0\", \"target\": \"entity_2\", \"type\": \"founded_by\"},\n",
" {\"source\": \"entity_0\", \"target\": \"entity_3\", \"type\": \"founded_by\"},\n",
" {\"source\": \"entity_0\", \"target\": \"entity_4\", \"type\": \"founded_in\"},\n",
" {\"source\": \"entity_0\", \"target\": \"entity_5\", \"type\": \"located_in\"},\n",
" {\"source\": \"entity_6\", \"target\": \"entity_0\", \"type\": \"ceo_of\"},\n",
"]\n",
"\n",
"try:\n",
" # Build the graph using NetworkX\n",
" kg = nx.DiGraph()\n",
" \n",
" for entity in entities_data:\n",
" kg.add_node(entity[\"id\"], name=entity[\"name\"], type=entity[\"type\"])\n",
" \n",
" for rel in relationships_data:\n",
" source_name = entities_data[int(rel[\"source\"].split(\"_\")[1])][\"name\"]\n",
" target_name = entities_data[int(rel[\"target\"].split(\"_\")[1])][\"name\"]\n",
" kg.add_edge(rel[\"source\"], rel[\"target\"], type=rel[\"type\"])\n",
" \n",
" print(\"✓ Knowledge graph built successfully!\")\n",
" print(f\" Nodes (entities): {len(kg.nodes)}\")\n",
" print(f\" Edges (relationships): {len(kg.edges)}\")\n",
" \n",
" print(\"\\nGraph Structure:\")\n",
" for node_id in kg.nodes():\n",
" node_data = kg.nodes[node_id]\n",
" print(f\" Node: {node_data['name']} ({node_data['type']})\")\n",
" \n",
" print(\"\\nRelationships:\")\n",
" for source, target, data in kg.edges(data=True):\n",
" source_name = kg.nodes[source]['name']\n",
" target_name = kg.nodes[target]['name']\n",
" print(f\" {source_name} --[{data['type']}]--> {target_name}\")\n",
" \n",
"except Exception as e:\n",
" print(f\"✗ Error building knowledge graph: {e}\")\n",
" kg = None\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## 📊 Step 5: Visualize and Analyze\n",
"\n",
"Finally, we'll visualize the knowledge graph and analyze its structure. This helps you understand the relationships and entities in your data.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from semantica.visualization import KGVisualizer\n",
"\n",
"try:\n",
" if kg is not None:\n",
" visualizer = KGVisualizer()\n",
" \n",
" print(\"Graph Summary:\")\n",
" print(f\" Total entities: {len(kg.nodes)}\")\n",
" print(f\" Total relationships: {len(kg.edges)}\")\n",
" \n",
" entity_types = {}\n",
" for node_id in kg.nodes():\n",
" entity_type = kg.nodes[node_id]['type']\n",
" entity_types[entity_type] = entity_types.get(entity_type, 0) + 1\n",
" \n",
" print(\"\\nEntities by type:\")\n",
" for etype, count in entity_types.items():\n",
" print(f\" - {etype}: {count}\")\n",
" \n",
" rel_types = {}\n",
" for _, _, data in kg.edges(data=True):\n",
" rel_type = data.get('type', 'unknown')\n",
" rel_types[rel_type] = rel_types.get(rel_type, 0) + 1\n",
" \n",
" print(\"\\nRelationships by type:\")\n",
" for rtype, count in rel_types.items():\n",
" print(f\" - {rtype}: {count}\")\n",
" \n",
" print(\"\\n✓ Graph visualization data prepared!\")\n",
" \n",
" else:\n",
" print(\"No graph available to visualize\")\n",
" \n",
"except Exception as e:\n",
" print(f\"✗ Error visualizing graph: {e}\")\n",
"\n",
"# Cleanup\n",
"try:\n",
" if sample_file.exists():\n",
" sample_file.unlink()\n",
" print(\"\\n✓ Sample file cleaned up\")\n",
"except:\n",
" pass\n"
]
}
],
"metadata": {
"language_info": {
"name": "python"
}
},
"nbformat": 4,
"nbformat_minor": 2
}