Files
semantica/docs/reference/explorer.md
T
KaifAhmad1 5eefadaa7f docs: apply full Mintlify component overhaul to all 27 reference pages and concepts.md
Replace plain markdown in every docs/reference/ file and docs/concepts.md with
rich Mintlify JSX components — CardGroup, Steps, Tabs, AccordionGroup, Tip,
Warning, Note, and CodeGroup — for a consistent, navigable, production-grade
developer experience.
2026-05-23 23:02:03 +05:30

14 KiB
Raw Blame History

title, description, icon
title description icon
Explorer Interactive FastAPI dashboard for knowledge graph exploration and the Ontology Hub. map

semantica.explorer is a browser-based dashboard for exploring knowledge graphs, managing ontologies, and running visual analyses — no code required after launch.

What You Get

Interactive node/edge search, filtering, path highlighting, and neighborhood expansion. Indexed search at 0.004ms on 118k-node graphs. Visual ontology editor, SHACL Studio, alignment authoring, health dashboard, and version control — all in the browser. Semantic similarity search, ego-mode neighborhood views, N×N distance heatmaps, and distance band classification. 15+ endpoints for graph data, path finding, embeddings, semantic search, analytics, and export — fully documented at `/docs`. Long-running exports and analyses stream progress events in real time — no polling required. `semantica-explorer --graph my_graph.json` for instant local startup without writing any Python.

Installation

pip install "semantica[explorer]"

Requires uvicorn and fastapi. Included automatically with pip install semantica[all].

Launch

```python from semantica.explorer import start_explorer from semantica.kg import GraphBuilder from semantica.ontology import OntologyManager
kg       = GraphBuilder().build(entities=entities, relationships=relationships)
ontology = OntologyManager()

start_explorer(
    graph=kg,
    ontology=ontology,     # optional — enables Ontology Hub tab
    port=8080,
    host="127.0.0.1",
    open_browser=True,
)
# → Serving at http://127.0.0.1:8080
```
```bash # Start on a saved graph semantica-explorer --graph my_graph.json
# Custom host and port
semantica-explorer --graph my_graph.json --host 0.0.0.0 --port 8080

# Skip auto-opening the browser
semantica-explorer --graph my_graph.json --no-browser
```
```python start_explorer( graph=kg, port=8080, enable_auth=True, api_key="my-secret-key", cors_origins=["https://app.example.com"], session_timeout=1800, # 30-minute inactivity timeout ) ``` ```bash curl -X POST http://localhost:8080/api/import \ -H "Content-Type: multipart/form-data" \ -F "file=@updated_graph.json" # Browser dashboard reloads automatically ```

start_explorer() Parameters

Parameter Type Default Description
graph KnowledgeGraph or ContextGraph (required) The graph to load into Explorer
ontology OntologyManager None Ontology to load into the Ontology Hub tab
port int 8000 Port to bind the server
host str "127.0.0.1" Host to bind. Use "0.0.0.0" for network access
open_browser bool True Auto-open the dashboard in the default browser
session_timeout int 3600 Session inactivity timeout in seconds; None disables
enable_auth bool False Require X-API-Key header on all API requests
api_key str None API key value when enable_auth=True
cors_origins list[str] ["*"] Allowed CORS origins. Restrict in production
log_level str "info" Uvicorn log level ("debug" / "info" / "warning")

CLI Reference

Flag Default Description
--graph, -g (required) Path to a saved graph JSON file
--port, -p 8000 Port to bind the server
--host 127.0.0.1 Host to bind the server
--no-browser false Skip auto-opening the browser
--enable-auth false Require X-API-Key header on all requests
--api-key None API key value when --enable-auth is set
--cors-origins "*" Comma-separated list of allowed CORS origins
--log-level "info" Uvicorn log level

Features

Core dashboard for navigating knowledge graphs:
- **Indexed search** — find any node by label or type; 0.004ms on 118k-node graphs (v0.5.0)
- **Bidirectional path finding** — trace paths between any two nodes
- **Neighbor expansion** — click any node to expand its connections
- **Filter by entity type** — focus on Person, Organization, Event, or any custom type
- **Edge label display** — relationship types shown on all edges
- **Graph declutter** — workspace layout controls for dense graphs
Full ontology lifecycle management in the browser:
- **Visual ontology editor** — drag-and-drop class and property authoring
- **SHACL Studio** — create, validate, and test SHACL shapes with live feedback
- **Alignment authoring** — author ontology alignments across schemas
- **Health dashboard** — graph quality metrics, validation status, coverage reports
- **Version control** — snapshot, diff, and restore ontology versions
Semantic neighborhood analysis centered on any node:
- **N×N distance matrices** — pairwise semantic distances across a set of nodes
- **Ego-mode visualization** — focus on a single node's semantic neighborhood
- **Distance band classification** — nodes grouped as `near` / `mid` / `far`
- **Embedding cache** — optimized embedding reuse for large graphs
Thread-safe sessions with rollback protection:
```python
start_explorer(
    graph=kg,
    session_timeout=1800,   # 30-minute inactivity timeout
    enable_auth=True,
    api_key="my-secret-key",
)
```

- Sessions are per connected browser tab
- Write operations (annotate, import) roll back automatically on failure
- All writes appended to audit trail at `/api/provenance/audit`
- Session state held in memory — use `/api/export/json` to persist between restarts

API Endpoints

Full interactive docs at http://localhost:8000/docs. All endpoints available via REST.

| Endpoint | Method | Description |
| -------- | ------ | ----------- |
| `/api/graph/summary` | `GET` | Node count, edge count, entity type distribution |
| `/api/graph/search` | `GET` | Indexed full-text and type-filtered node search |
| `/api/graph/node/{id}` | `GET` | Fetch a single node with all properties |
| `/api/graph/neighbors` | `GET` | Neighbors of a node — `?node_id=&depth=2` |
| `/api/graph/path` | `GET` | Bidirectional shortest path — `?source=&target=` |
| `/api/graph/subgraph` | `POST` | Extract a subgraph by node IDs or type filter |
| `/api/graph/annotate` | `POST` | Add a user annotation to a node or edge |
| `/api/graph/annotations` | `GET` | List all annotations on the graph |
| Endpoint | Method | Description |
| -------- | ------ | ----------- |
| `/api/ontology/classes` | `GET` | List all ontology classes and properties |
| `/api/ontology/validate` | `POST` | Run SHACL validation; returns violations |
| `/api/ontology/hierarchy` | `GET` | Class hierarchy as a tree structure |
| `/api/ontology/vocabulary` | `GET` | SKOS vocabulary terms and alt labels |
| `/api/ontology/align` | `POST` | Submit two ontologies for alignment |
| `/api/ontology/diff` | `POST` | Diff two ontology versions |
**Provenance:**

| Endpoint | Method | Description |
| -------- | ------ | ----------- |
| `/api/provenance/entity/{id}` | `GET` | Full provenance lineage for an entity |
| `/api/provenance/source/{id}` | `GET` | All entities sourced from a document |
| `/api/provenance/audit` | `GET` | Full audit trail of Explorer operations |

**Decisions:**

| Endpoint | Method | Description |
| -------- | ------ | ----------- |
| `/api/decisions/list` | `GET` | Paginated list of recorded decisions |
| `/api/decisions/{id}` | `GET` | Single decision with causal chain |
| `/api/decisions/search` | `GET` | Precedent search — `?query=&limit=5` |
| `/api/decisions/influence/{id}` | `GET` | Downstream influence of a decision |

**Analytics:**

| Endpoint | Method | Description |
| -------- | ------ | ----------- |
| `/api/analytics/centrality` | `GET` | Degree, betweenness, PageRank scores |
| `/api/analytics/communities` | `GET` | Community detection result |
| `/api/analytics/distance` | `POST` | N×N distance matrix for a node list |
| `/api/analytics/neighborhood` | `GET` | Semantic neighborhood — `?node=&radius=0.4` |
**SPARQL & Temporal:**

| Endpoint | Method | Description |
| -------- | ------ | ----------- |
| `/api/sparql` | `POST` | Execute a SPARQL SELECT query |
| `/api/temporal/snapshot` | `GET` | Graph snapshot at a point in time — `?at=ISO8601` |
| `/api/temporal/range` | `GET` | Nodes/edges active in a time range |
| `/api/temporal/diff` | `POST` | Diff two temporal snapshots |

**Export & Import:**

| Endpoint | Method | Description |
| -------- | ------ | ----------- |
| `/api/export/{format}` | `GET` | Export in: `turtle`, `json-ld`, `ntriples`, `rdf-xml`, `parquet`, `aql`, `csv`, `owl`, `arrow`, `lpg`, `yaml`, `distance-matrix` |
| `/api/import` | `POST` | Import a graph from file (replaces current graph in session) |

WebSocket Progress

Long-running operations stream progress events over WebSocket at ws://localhost:8000/ws/progress:

import asyncio, websockets, json

async def watch_progress():
    async with websockets.connect("ws://localhost:8000/ws/progress") as ws:
        async for message in ws:
            event = json.loads(message)
            print(f"[{event['operation']}] {event['step']}{event['progress_pct']:.0f}%")
            if event["status"] in ("completed", "failed"):
                break

asyncio.run(watch_progress())

WebSocket event schema:

{
  "operation":    "export",
  "step":         "serializing nodes",
  "current":      3500,
  "total":        10000,
  "progress_pct": 35.0,
  "status":       "running",
  "message":      "Serializing 10000 nodes to Turtle...",
  "error":        null
}

status values: "running" | "completed" | "failed" | "cancelled"

Performance

Scenario Latency
Node search (118k nodes, indexed) 0.004ms
Neighbor expansion (depth 2) < 5ms
Bidirectional path (118k nodes) < 50ms
SPARQL SELECT (simple pattern) < 20ms
N×N distance matrix (100 nodes) ~2s (with embedding cache)

The node search index is built on startup. For graphs > 500k nodes, pass index_build_timeout=120 to start_explorer() to allow more time.

Tips and Common Pitfalls

**Set `max_nodes` when loading large graphs.** `start_explorer(graph=kg, max_nodes=50000)` limits the rendered node count — Explorer's force-directed layout becomes unusable above ~10k nodes without limiting. Use `graph.filter(node_type="Organization")` first to focus on what matters. **Use authentication in shared environments.** Pass `enable_auth=True, api_key="..."` whenever Explorer is accessible to more than one person. Without auth, anyone who can reach the port can write to the graph via the annotate and import endpoints. **Export before the session ends.** Session state lives in memory and is lost on server restart. Call `/api/export/json` or `/api/export/turtle` to persist the current state before shutting down. Explorer does not auto-save. **Use WebSocket progress for long operations.** Export, analysis, and large SPARQL queries stream progress to `ws://localhost:8000/ws/progress`. Polling the REST endpoints instead gives no progress signal — use the WebSocket client so users see incremental updates. **Pass `session_timeout` for demos and shared notebooks.** The default session never expires. In Jupyter or shared environments, set `session_timeout=1800` (30 minutes) so stale sessions don't hold large graphs in memory. **Use the REST API for automation, Explorer UI for exploration.** Explorer's REST endpoints are a stable programmatic API — pipe them into scripts to automate batch annotation, SPARQL querying, or exports. The browser UI is for interactive exploration and sharing; they use the same server. Build and save the ContextGraph that Explorer loads. Programmatic ontology management and SHACL generation. Programmatic graph rendering without the Explorer server. Export to RDF, Parquet, AQL without launching a server.