mirror of
https://github.com/semantica-agi/semantica.git
synced 2026-09-15 04:00:33 +00:00
- Replace mock data with real feed URLs, APIs, and database patterns - Add real threat intelligence feeds (CISA, US-CERT, Security Week, Dark Reading) - Add real financial feeds (Reuters, CNN Money, Bloomberg, Financial Times) - Add real healthcare feeds (CDC, WHO) - Add real API endpoints (MITRE ATT&CK, NVD CVE API, Polygon.io, Alpha Vantage, FHIR APIs) - Add realistic database connection patterns with SQL queries - Add Kafka/RabbitMQ streaming configurations - Update all cybersecurity notebooks (5/5) with real sources - Update finance notebooks (2/2) with real sources - Update healthcare notebooks (1/1) with real sources - Create REAL_DATA_SOURCES.md documentation - Improve error handling with try-except blocks - Add batch processing for multiple feed URLs
163 lines
5.3 KiB
Plaintext
163 lines
5.3 KiB
Plaintext
{
|
|
"cells": [
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"# Graph Analytics\n",
|
|
"\n",
|
|
"## Overview\n",
|
|
"\n",
|
|
"This notebook demonstrates how to analyze knowledge graphs using Semantica's analytics modules. You'll learn to use `GraphAnalyzer`, `CentralityCalculator`, `CommunityDetector`, and `ConnectivityAnalyzer` to understand graph structure and properties.\n",
|
|
"\n",
|
|
"### Learning Objectives\n",
|
|
"\n",
|
|
"- Use `GraphAnalyzer` for comprehensive graph analysis\n",
|
|
"- Use `CentralityCalculator` to compute centrality measures\n",
|
|
"- Use `CommunityDetector` to find communities in graphs\n",
|
|
"- Use `ConnectivityAnalyzer` to analyze graph connectivity\n",
|
|
"\n",
|
|
"---\n",
|
|
"\n",
|
|
"## Step 1: Graph Analysis\n",
|
|
"\n",
|
|
"Analyze graph structure and properties.\n"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "code",
|
|
"execution_count": null,
|
|
"metadata": {},
|
|
"outputs": [],
|
|
"source": [
|
|
"from semantica.kg import GraphBuilder, GraphAnalyzer\n",
|
|
"from semantica.semantic_extract import NERExtractor, RelationExtractor\n",
|
|
"\n",
|
|
"builder = GraphBuilder()\n",
|
|
"analyzer = GraphAnalyzer()\n",
|
|
"\n",
|
|
"entities = [\n",
|
|
" {\"id\": \"e1\", \"type\": \"Organization\", \"name\": \"Apple Inc.\", \"properties\": {}},\n",
|
|
" {\"id\": \"e2\", \"type\": \"Person\", \"name\": \"Tim Cook\", \"properties\": {}},\n",
|
|
" {\"id\": \"e3\", \"type\": \"Location\", \"name\": \"Cupertino\", \"properties\": {}}\n",
|
|
"]\n",
|
|
"\n",
|
|
"relationships = [\n",
|
|
" {\"source\": \"e2\", \"target\": \"e1\", \"type\": \"CEO_of\", \"properties\": {}},\n",
|
|
" {\"source\": \"e1\", \"target\": \"e3\", \"type\": \"located_in\", \"properties\": {}}\n",
|
|
"]\n",
|
|
"\n",
|
|
"kg = builder.build(entities, relationships)\n",
|
|
"\n",
|
|
"metrics = analyzer.compute_metrics(kg)\n",
|
|
"\n",
|
|
"print(f\"Graph metrics:\")\n",
|
|
"print(f\" Entities: {metrics.get('entity_count', 0)}\")\n",
|
|
"print(f\" Relationships: {metrics.get('relationship_count', 0)}\")\n",
|
|
"print(f\" Density: {metrics.get('density', 0):.3f}\")\n"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"## Step 2: Centrality Measures\n",
|
|
"\n",
|
|
"Calculate centrality measures for entities.\n"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "code",
|
|
"execution_count": null,
|
|
"metadata": {},
|
|
"outputs": [],
|
|
"source": [
|
|
"from semantica.kg import CentralityCalculator\n",
|
|
"\n",
|
|
"centrality_calculator = CentralityCalculator()\n",
|
|
"\n",
|
|
"centrality_scores = centrality_calculator.calculate_centrality(kg, measure=\"degree\")\n",
|
|
"\n",
|
|
"print(f\"Centrality scores:\")\n",
|
|
"for entity_id, score in list(centrality_scores.items())[:5]:\n",
|
|
" print(f\" {entity_id}: {score:.3f}\")\n"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"## Step 3: Community Detection\n",
|
|
"\n",
|
|
"Detect communities in the graph.\n"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "code",
|
|
"execution_count": null,
|
|
"metadata": {},
|
|
"outputs": [],
|
|
"source": [
|
|
"from semantica.kg import CommunityDetector\n",
|
|
"\n",
|
|
"community_detector = CommunityDetector()\n",
|
|
"\n",
|
|
"communities = community_detector.detect_communities(kg)\n",
|
|
"\n",
|
|
"print(f\"Detected {len(communities)} communities\")\n",
|
|
"for i, community in enumerate(communities[:3], 1):\n",
|
|
" print(f\" Community {i}: {len(community)} entities\")\n"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"## Step 4: Connectivity Analysis\n",
|
|
"\n",
|
|
"Analyze graph connectivity.\n"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "code",
|
|
"execution_count": null,
|
|
"metadata": {},
|
|
"outputs": [],
|
|
"source": [
|
|
"from semantica.kg import ConnectivityAnalyzer\n",
|
|
"\n",
|
|
"connectivity_analyzer = ConnectivityAnalyzer()\n",
|
|
"\n",
|
|
"connectivity = connectivity_analyzer.analyze_connectivity(kg)\n",
|
|
"\n",
|
|
"print(f\"Connectivity analysis:\")\n",
|
|
"print(f\" Is connected: {connectivity.get('is_connected', False)}\")\n",
|
|
"print(f\" Components: {len(connectivity.get('components', []))}\")\n"
|
|
]
|
|
},
|
|
{
|
|
"cell_type": "markdown",
|
|
"metadata": {},
|
|
"source": [
|
|
"## Summary\n",
|
|
"\n",
|
|
"You've learned how to analyze knowledge graphs:\n",
|
|
"\n",
|
|
"- **GraphAnalyzer**: Comprehensive graph analysis and metrics\n",
|
|
"- **CentralityCalculator**: Calculate centrality measures\n",
|
|
"- **CommunityDetector**: Detect communities in graphs\n",
|
|
"- **ConnectivityAnalyzer**: Analyze graph connectivity\n",
|
|
"\n",
|
|
"Next: Learn how to assess graph quality in the Graph_Quality notebook.\n"
|
|
]
|
|
}
|
|
],
|
|
"metadata": {
|
|
"language_info": {
|
|
"name": "python"
|
|
}
|
|
},
|
|
"nbformat": 4,
|
|
"nbformat_minor": 2
|
|
}
|