{ "cells": [ { "cell_type": "markdown", "metadata": {}, "source": [ "# Multi-Source Data Integration\n", "\n", "## Overview\n", "\n", "This notebook demonstrates advanced multi-source data integration using multiple ingestion types, entity resolution, conflict detection, and provenance tracking.\n", "\n", "### Learning Objectives\n", "\n", "- Ingest data from multiple sources (files, web, databases, streams, feeds)\n", "- Resolve entities across sources using EntityResolver\n", "- Detect conflicts using ConflictDetector\n", "- Track provenance using ProvenanceTracker\n", "- Integrate data into a unified knowledge graph\n", "\n", "---\n", "\n", "## Workflow: Multi-Source Ingestion → Entity Resolution → Conflict Detection → Provenance Tracking → Unified KG\n" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "from semantica.ingest import FileIngestor, WebIngestor, DBIngestor, StreamIngestor, FeedIngestor\n", "from semantica.parse import DocumentParser, StructuredDataParser\n", "from semantica.kg import GraphBuilder, EntityResolver, ConflictDetector, ProvenanceTracker\n", "import tempfile\n", "import os\n", "import json\n", "\n", "file_ingestor = FileIngestor()\n", "web_ingestor = WebIngestor()\n", "db_ingestor = DBIngestor()\n", "stream_ingestor = StreamIngestor()\n", "feed_ingestor = FeedIngestor()\n", "\n", "temp_dir = tempfile.mkdtemp()\n", "\n", "file1 = os.path.join(temp_dir, \"source1.txt\")\n", "with open(file1, 'w') as f:\n", " f.write(\"Apple Inc. is a technology company. Tim Cook is the CEO.\")\n", "\n", "file_objects = file_ingestor.ingest_file(file1, read_content=True)\n", "\n", "print(f\"Ingested {len([file_objects]) if file_objects else 0} files\")\n", "print(f\"Multi-source ingestion initialized\")\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## Step 2: Entity Resolution\n", "\n", "Resolve entities across multiple sources.\n" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "entity_resolver = EntityResolver()\n", "\n", "entities_from_source1 = [\n", " {\"id\": \"e1\", \"name\": \"Apple Inc.\", \"type\": \"Organization\", \"source\": \"file1\"},\n", " {\"id\": \"e2\", \"name\": \"Tim Cook\", \"type\": \"Person\", \"source\": \"file1\"}\n", "]\n", "\n", "entities_from_source2 = [\n", " {\"id\": \"e3\", \"name\": \"Apple Incorporated\", \"type\": \"Organization\", \"source\": \"web\"},\n", " {\"id\": \"e4\", \"name\": \"Timothy Cook\", \"type\": \"Person\", \"source\": \"web\"}\n", "]\n", "\n", "all_entities = entities_from_source1 + entities_from_source2\n", "\n", "resolved_entities = entity_resolver.resolve(all_entities)\n", "\n", "print(f\"Original entities: {len(all_entities)}\")\n", "print(f\"Resolved entities: {len(resolved_entities)}\")\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## Step 3: Conflict Detection\n", "\n", "Detect conflicts between sources.\n" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "conflict_detector = ConflictDetector()\n", "\n", "conflicts = conflict_detector.detect_value_conflicts(all_entities, \"name\")\n", "\n", "print(f\"Detected {len(conflicts)} conflicts\")\n", "for conflict in conflicts[:3]:\n", " print(f\" Conflict: {conflict.entity_id} - {conflict.conflict_type}\")\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## Step 4: Provenance Tracking\n", "\n", "Track data provenance across sources.\n" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "provenance_tracker = ProvenanceTracker()\n", "\n", "for entity in all_entities:\n", " provenance_tracker.track_entity(entity.get(\"id\"), entity.get(\"source\"), entity)\n", "\n", "relationships = [\n", " {\"source\": \"e2\", \"target\": \"e1\", \"type\": \"CEO_of\", \"source\": \"file1\"}\n", "]\n", "\n", "for rel in relationships:\n", " provenance_tracker.track_relationship(rel.get(\"source\"), rel.get(\"target\"), rel.get(\"source\"), rel)\n", "\n", "print(f\"Tracked provenance for {len(all_entities)} entities and {len(relationships)} relationships\")\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## Step 5: Build Unified Knowledge Graph\n", "\n", "Build a unified knowledge graph from integrated sources.\n" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "builder = GraphBuilder()\n", "\n", "unified_kg = builder.build(resolved_entities, relationships)\n", "\n", "print(f\"Built unified knowledge graph\")\n", "print(f\" Entities: {len(unified_kg.get('entities', []))}\")\n", "print(f\" Relationships: {len(unified_kg.get('relationships', []))}\")\n", "print(f\" Sources integrated: {len(set(e.get('source', '') for e in resolved_entities))}\")\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## Summary\n", "\n", "You've learned advanced multi-source data integration:\n", "\n", "- **Multiple Ingestion Types**: FileIngestor, WebIngestor, DBIngestor, StreamIngestor, FeedIngestor\n", "- **EntityResolver**: Resolve entities across sources\n", "- **ConflictDetector**: Detect conflicts between sources\n", "- **ProvenanceTracker**: Track data provenance\n", "- **Unified Knowledge Graph**: Build integrated graph from multiple sources\n" ] } ], "metadata": { "language_info": { "name": "python" } }, "nbformat": 4, "nbformat_minor": 2 }