mirror of
https://github.com/semantica-agi/semantica.git
synced 2026-09-04 04:01:07 +00:00
Compare commits
17
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
98900af751 | ||
|
|
8bceff105c | ||
|
|
f45499b5a7 | ||
|
|
b574e2e6b4 | ||
|
|
c04adcd1a9 | ||
|
|
a85cf913a5 | ||
|
|
6ba433fea0 | ||
|
|
f6a0e4a32e | ||
|
|
111bcf997e | ||
|
|
6a07ad29be | ||
|
|
064f0eccad | ||
|
|
afec253451 | ||
|
|
dd1e654047 | ||
|
|
6b8437781e | ||
|
|
ba85215aea | ||
|
|
5809418421 | ||
|
|
40efab6796 |
@@ -28,7 +28,13 @@
|
||||
|
||||
[](https://github.com/semantica-agi/semantica) [](https://github.com/semantica-agi/semantica/network/members) [](https://github.com/semantica-agi/semantica/graphs/contributors) [](https://pypi.org/project/semantica/) [](https://pepy.tech/project/semantica) [](https://www.python.org/) [](https://opensource.org/licenses/MIT) [](https://github.com/semantica-agi/semantica/actions) [](https://github.com/semantica-agi/semantica/actions/workflows/install-matrix.yml) [](https://scorecard.dev/viewer/?uri=github.com/semantica-agi/semantica) [](https://deepwiki.com/semantica-agi/semantica)
|
||||
|
||||
[](https://getsemantica.ai/) [](https://docs.getsemantica.ai/) [](https://discord.gg/sV34vps5hH) [](https://x.com/BuildSemantica) [](https://www.youtube.com/watch?v=QfnNZg4-dZA) [](CHANGELOG.md)
|
||||
[](https://getsemantica.ai/)
|
||||
[](https://docs.getsemantica.ai/)
|
||||
[](https://discord.gg/sV34vps5hH)
|
||||
[](https://x.com/BuildSemantica)
|
||||
|
||||
[](https://www.youtube.com/watch?v=QfnNZg4-dZA)
|
||||
|
||||
|
||||
```bash
|
||||
pip install semantica
|
||||
|
||||
+10
-10
@@ -8,16 +8,16 @@ icon: "book-open"
|
||||
New here? Start with [Getting Started](/getting-started) for hands-on examples, then return here for deeper understanding.
|
||||
</Info>
|
||||
|
||||
Semantica transforms unstructured data: documents, web pages, reports, databases: into **knowledge graphs**: structured representations that AI systems can query, reason about, and trace back to sources.
|
||||
Semantica transforms unstructured data (documents, web pages, reports, databases) into **knowledge graphs**: structured representations that AI systems can query, reason about, and trace back to sources.
|
||||
|
||||
At its core, Semantica adds a **context and accountability layer** on top of your existing AI stack. It doesn't replace LangChain, LlamaIndex, or your LLM provider: it makes their outputs **grounded**, **traceable**, and **auditable**.
|
||||
At its core, Semantica adds a context and semantic layer on top of your existing AI stack. It doesn't replace LangChain, LlamaIndex, or your LLM provider. It makes their outputs grounded, traceable, and auditable.
|
||||
|
||||
- **Context Layer** — Knowledge graphs, GraphRAG retrieval, semantic embeddings, and temporal intelligence ground every LLM response in structured, queryable facts.
|
||||
- **Accountability Layer** — Provenance tracking, decision intelligence, conflict detection, and W3C PROV-O compliance make every claim in your AI stack auditable and explainable.
|
||||
- **Extension Layer** — `PluginRegistry` and `MethodRegistry` let you replace or augment any component: ingestors, extractors, reasoning engines, backends: without changing framework code.
|
||||
- **Context Layer.** Knowledge graphs, GraphRAG retrieval, semantic embeddings, and temporal intelligence ground every LLM response in structured, queryable facts.
|
||||
- **Accountability Layer.** Provenance tracking, decision intelligence, conflict detection, and W3C PROV-O compliance make every claim in your AI stack auditable and explainable.
|
||||
- **Extension Layer.** `PluginRegistry` and `MethodRegistry` let you replace or augment any component (ingestors, extractors, reasoning engines, backends) without changing framework code.
|
||||
|
||||
<Warning>
|
||||
**This is system-level explainability, not foundation-model explainability.** Semantica does not expose, reconstruct, or explain what happens *inside* the LLM/foundation model — its internal reasoning or chain-of-thought stays opaque, as it does for any external system. What Semantica explains is *outside* the model: the context and data fed in, the decision produced, its provenance, the relevant relationships, the policies applied, and the full execution trail. In short, Semantica explains and audits *what the AI system did*, not the foundation model's private internal reasoning.
|
||||
**This is system-level explainability, not foundation-model explainability.** Semantica does not expose, reconstruct, or explain what happens *inside* the LLM/foundation model. Its internal reasoning or chain-of-thought stays opaque, as it does for any external system. What Semantica explains is *outside* the model: the context and data fed in, the decision produced, its provenance, the relevant relationships, the policies applied, and the full execution trail. In short, Semantica explains and audits *what the AI system did*, not the foundation model's private internal reasoning.
|
||||
</Warning>
|
||||
|
||||
## Knowledge Graphs
|
||||
@@ -30,7 +30,7 @@ The foundation of everything in Semantica. A knowledge graph stores information
|
||||
- **Edges (relationships)**: `works_for`, `located_in`, `founded_by`
|
||||
- **Properties**: name, date, confidence score, source URL
|
||||
|
||||
This structure makes knowledge **searchable**, **connectable**, **queryable**, and: critically: **explainable**: every answer can be traced back to the facts and relationships that produced it.
|
||||
This structure makes knowledge searchable, connectable, and queryable. Critically, it's explainable: every answer can be traced back to the facts and relationships that produced it.
|
||||
|
||||
|
||||
## Entity Extraction (NER)
|
||||
@@ -513,6 +513,6 @@ Semantica is designed for extension. Any component: ingestor, extractor, graph b
|
||||
</Accordion>
|
||||
</AccordionGroup>
|
||||
|
||||
- [Quickstart Tutorial](/quickstart) — Build a full pipeline with code.
|
||||
- [Modules Guide](/modules) — Every module explained with examples.
|
||||
- [API Reference](/reference/context) — Complete technical reference.
|
||||
- [Quickstart Tutorial](/quickstart): build a full pipeline with code.
|
||||
- [Modules Guide](/modules): every module explained with examples.
|
||||
- [API Reference](/reference/context): complete technical reference.
|
||||
|
||||
+2
-2
@@ -2,7 +2,7 @@
|
||||
"$schema": "https://mintlify.com/docs.json",
|
||||
"theme": "mint",
|
||||
"name": "Semantica",
|
||||
"description": "The Accountability and Context Layer for AI — Context Graphs · Decision Intelligence · Full Provenance",
|
||||
"description": "The Context and Semantic Layer for AI in High-Stakes Domains — Context Graphs · Decision Intelligence · Full Provenance",
|
||||
"colors": {
|
||||
"primary": "#10B981",
|
||||
"light": "#10B981",
|
||||
@@ -43,7 +43,7 @@
|
||||
"raiseIssue": true
|
||||
},
|
||||
"metadata": {
|
||||
"og:title": "Semantica — Accountability & Context Layer for AI",
|
||||
"og:title": "Semantica — Context & Semantic Layer for AI in High-Stakes Domains",
|
||||
"og:description": "Build explainable, auditable knowledge graphs with full provenance. Open source. MIT licensed.",
|
||||
"og:image": "/assets/img/semantica-logo.png",
|
||||
"twitter:card": "summary_large_image",
|
||||
|
||||
+55
-42
@@ -1,9 +1,9 @@
|
||||
---
|
||||
title: "GraphRAG — Graph-Augmented Retrieval"
|
||||
title: "GraphRAG: Graph-Augmented Retrieval"
|
||||
description: "Go beyond vector search: retrieve facts, trace reasoning paths, and ground LLM responses in your knowledge graph."
|
||||
---
|
||||
|
||||
GraphRAG combines vector similarity with knowledge graph traversal so retrieval finds structurally connected facts, not just text that sounds related. When a `ContextGraph` is attached to `AgentContext`, every retrieval call automatically blends semantic search with multi-hop graph expansion — and `query_with_reasoning()` returns an auditable reasoning path alongside the LLM answer.
|
||||
GraphRAG combines vector similarity with knowledge graph traversal so retrieval finds structurally connected facts, not just text that sounds related. When a `ContextGraph` is attached to `AgentContext`, every retrieval call automatically blends semantic search with multi-hop graph expansion, and `query_with_reasoning()` returns an auditable reasoning path alongside the LLM answer.
|
||||
|
||||
## What Is GraphRAG?
|
||||
|
||||
@@ -11,7 +11,7 @@ GraphRAG (Graph-Augmented Retrieval-Augmented Generation) enhances traditional R
|
||||
|
||||
**GraphRAG vs. traditional vector-only RAG:** Vector RAG finds documents similar to your query text. GraphRAG finds documents similar to your query AND documents connected to those through entity relationships, even if they don't mention your query terms directly.
|
||||
|
||||
**The role of graph traversal:** Starting from entities found in vector-similar documents, GraphRAG expands outward through relationship edges to discover related facts. This reveals connections that pure text similarity would miss — like finding that a threat actor targets healthcare by following the path: Actor → Tool → Victim Organization → Industry Sector.
|
||||
**The role of graph traversal:** Starting from entities found in vector-similar documents, GraphRAG expands outward through relationship edges to discover related facts. This reveals connections that pure text similarity would miss, like finding that a threat actor targets healthcare by following the path: Actor → Tool → Victim Organization → Industry Sector.
|
||||
|
||||
## Why Use GraphRAG?
|
||||
|
||||
@@ -96,7 +96,7 @@ context = AgentContext(
|
||||
)
|
||||
```
|
||||
|
||||
Now ingest your documents. `store()` with `extract_entities=True` runs the full extraction pipeline internally — Named Entity Recognition (NER), relation extraction, and entity linking — and populates both the vector index and the graph simultaneously:
|
||||
Now ingest your documents. `store()` with `extract_entities=True` runs the full extraction pipeline internally (Named Entity Recognition, relation extraction, and entity linking) and populates both the vector index and the graph simultaneously:
|
||||
|
||||
```python
|
||||
intel_documents = [
|
||||
@@ -132,16 +132,17 @@ stats = context.store(
|
||||
print("Graph built: {} nodes, {} edges".format(
|
||||
stats["graph_nodes"], stats["graph_edges"]
|
||||
))
|
||||
# Graph built: 18 nodes, 14 edges
|
||||
# Nodes: APT29, HAMMERTOSS, NATO, LifeCare, AS59796, CISA Sector 6, ...
|
||||
# Edges: deployed, observed_on, classified_as, targets, operates_in, ...
|
||||
```
|
||||
|
||||
The graph now contains a connected subgraph linking APT29 to healthcare infrastructure across four document boundaries — something that would be invisible to a pure vector search.
|
||||
`store()` returns a dict with `stored_count`, `memory_ids`, `graph_nodes`, and
|
||||
`graph_edges`. The extracted nodes (APT29, HAMMERTOSS, LifeCare, AS59796, …) and
|
||||
edges (`deployed`, `observed_on`, `classified_as`, …) now span all four documents.
|
||||
|
||||
The graph now contains a connected subgraph linking APT29 to healthcare infrastructure across four document boundaries, something that would be invisible to a pure vector search.
|
||||
|
||||
## Retrieving the relevant subgraph
|
||||
|
||||
With the graph populated, a plain `retrieve()` call already does more than vector search. When `use_graph=True`, the retriever seeds the graph traversal from the top-k vector matches and expands outward by following edges, collecting connected facts within `max_hops`:
|
||||
With the graph populated, a plain `retrieve()` call already does more than vector search. When `use_graph=True`, the retriever seeds the graph traversal from the top-k vector matches and expands outward by following edges. Expansion depth is set once, by `max_expansion_hops` on the `AgentContext` constructor:
|
||||
|
||||
```python
|
||||
results = context.retrieve(
|
||||
@@ -149,7 +150,6 @@ results = context.retrieve(
|
||||
use_graph=True,
|
||||
max_results=10,
|
||||
expand_graph=True,
|
||||
max_hops=3,
|
||||
)
|
||||
|
||||
for r in results:
|
||||
@@ -169,17 +169,25 @@ Notice the top results: while pure vector search might rank connected facts lowe
|
||||
When you know specifically which entity you want to anchor the traversal to, pass `anchor_node`:
|
||||
|
||||
```python
|
||||
# Anchor on APT29 explicitly — proximity scores are calculated from this node
|
||||
# Anchor on APT29 explicitly: proximity scores are calculated from this node
|
||||
apt29_intel = context.retrieve(
|
||||
"C2 infrastructure beaconing patterns",
|
||||
use_graph=True,
|
||||
anchor_node="APT29",
|
||||
proximity_weight=0.7, # strongly favour nodes close to APT29
|
||||
max_hops=3,
|
||||
max_hops=3, # with an anchor, this bounds the proximity radius
|
||||
max_results=8,
|
||||
)
|
||||
```
|
||||
|
||||
<Note>
|
||||
`max_hops` on `retrieve()` only takes effect when `anchor_node` is set: it
|
||||
bounds the proximity radius used for scoring and drops results farther than
|
||||
`max_hops` from the anchor. Without an `anchor_node` it is ignored. It does
|
||||
**not** change how far graph expansion reaches: that is fixed by
|
||||
`max_expansion_hops` on the constructor.
|
||||
</Note>
|
||||
|
||||
## Getting a grounded LLM answer with a reasoning path
|
||||
|
||||
`retrieve()` gives you the grounded context. `query_with_reasoning()` goes one step further: it passes that subgraph context to an LLM and returns the answer together with the multi-hop path the retrieval system traced through the graph. That path is your audit trail.
|
||||
@@ -187,7 +195,7 @@ apt29_intel = context.retrieve(
|
||||
```python
|
||||
from semantica.llms import LiteLLM
|
||||
|
||||
llm = LiteLLM(model="anthropic/claude-sonnet-4-20250514")
|
||||
llm = LiteLLM(model="anthropic/claude-sonnet-5")
|
||||
|
||||
result = context.query_with_reasoning(
|
||||
"What are APT29's known TTPs against healthcare infrastructure, "
|
||||
@@ -197,7 +205,7 @@ result = context.query_with_reasoning(
|
||||
max_hops=3,
|
||||
)
|
||||
|
||||
# The LLM answer — grounded in graph-retrieved context, not training memory
|
||||
# The LLM answer, grounded in graph-retrieved context, not training memory
|
||||
print(result["response"])
|
||||
|
||||
# The multi-hop trace: APT29 → deployed → HAMMERTOSS → observed_on → LifeCare → ...
|
||||
@@ -213,7 +221,7 @@ for src in result["sources"]:
|
||||
print(" [{:.3f}] {}".format(src["score"], src["content"][:80]))
|
||||
```
|
||||
|
||||
The `reasoning_path` field is what separates GraphRAG from a black-box LLM call. When an analyst asks "how do you know APT29 targeted healthcare?", you can show them the exact traversal the system made across your own documents — not a claim the model generated from training data.
|
||||
The `reasoning_path` field is what separates GraphRAG from a black-box LLM call. When an analyst asks "how do you know APT29 targeted healthcare?", you can show them the exact traversal the system made across your own documents, not a claim the model generated from training data.
|
||||
|
||||
The full return structure from `query_with_reasoning()`:
|
||||
|
||||
@@ -232,11 +240,11 @@ The full return structure from `query_with_reasoning()`:
|
||||
|
||||
<Tabs>
|
||||
|
||||
<Tab title="Defense — CTI/Threat">
|
||||
<Tab title="Defense: CTI/Threat">
|
||||
|
||||
Multi-INT intelligence fusion: OSINT threat feeds, NVD CVE data, and HUMINT summaries ingested into a single graph, then queried with multi-hop reasoning to trace C2 infrastructure chains and attribute campaigns to specific actors.
|
||||
|
||||
In classified environments the graph can be partitioned by data handling caveat — each `AgentContext` operates over the subset of documents cleared for the querying user. The `reasoning_path` output doubles as a sanitisable audit trail for downgraded reporting.
|
||||
In classified environments the graph can be partitioned by data handling caveat: each `AgentContext` operates over the subset of documents cleared for the querying user. The `reasoning_path` output doubles as a sanitisable audit trail for downgraded reporting.
|
||||
|
||||
```python
|
||||
from semantica.context import AgentContext, ContextGraph
|
||||
@@ -273,7 +281,7 @@ context.store(
|
||||
link_entities=True,
|
||||
)
|
||||
|
||||
llm = LiteLLM(model="anthropic/claude-sonnet-4-20250514")
|
||||
llm = LiteLLM(model="anthropic/claude-sonnet-5")
|
||||
result = context.query_with_reasoning(
|
||||
"Trace the C2 infrastructure chain for APT29 operations targeting "
|
||||
"ITAR-controlled contractors in 2025. Include IP ranges, ASNs, and TTPs.",
|
||||
@@ -300,11 +308,11 @@ proximate = context.retrieve(
|
||||
|
||||
</Tab>
|
||||
|
||||
<Tab title="Security — SOC/Incident">
|
||||
<Tab title="Security: SOC/Incident">
|
||||
|
||||
Security operations: real-time alert triage against a graph containing hosts, CVEs, user accounts, runbooks, and historical incidents. GraphRAG retrieves the relevant runbook and similar past incidents in a single call, reducing mean-time-to-respond.
|
||||
|
||||
The `decision_tracking=True` flag records every triage query as an auditable decision, with the full context that was provided to the LLM — essential for post-incident review and SOC metrics.
|
||||
The `decision_tracking=True` flag records every triage query as an auditable decision, with the full context that was provided to the LLM. That's essential for post-incident review and SOC metrics.
|
||||
|
||||
```python
|
||||
from semantica.context import AgentContext, ContextGraph
|
||||
@@ -343,7 +351,7 @@ Parent: wmiprvse.exe
|
||||
Sigma match: T1053.005 Scheduled Task/Job
|
||||
"""
|
||||
|
||||
llm = LiteLLM(model="anthropic/claude-sonnet-4-20250514")
|
||||
llm = LiteLLM(model="anthropic/claude-sonnet-5")
|
||||
triage = soc_context.query_with_reasoning(
|
||||
"Triage this SIEM alert and identify the correct response runbook:\n{}".format(alert_text),
|
||||
llm_provider=llm,
|
||||
@@ -369,7 +377,7 @@ for inc in similar:
|
||||
|
||||
</Tab>
|
||||
|
||||
<Tab title="Life Science — Clinical/Pharma">
|
||||
<Tab title="Life Science: Clinical/Pharma">
|
||||
|
||||
Clinical decision support: FDA drug labels, clinical guidelines, and trial summaries ingested into a graph where drug-enzyme-metabolite-interaction chains become traversable paths. A three-hop query (drug → enzyme → metabolite → contraindication) surfaces interaction risks that no single document would make explicit.
|
||||
|
||||
@@ -417,7 +425,7 @@ Patient: 68F, AF, CKD stage 3b (eGFR 32). On warfarin (INR target 2.0–3.0).
|
||||
Presenting for elective hip replacement. Concurrent: amiodarone 200mg, atorvastatin 40mg.
|
||||
"""
|
||||
|
||||
llm = LiteLLM(model="anthropic/claude-sonnet-4-20250514")
|
||||
llm = LiteLLM(model="anthropic/claude-sonnet-5")
|
||||
answer = clinical_context.query_with_reasoning(
|
||||
"What is the evidence-based warfarin bridging protocol for this patient "
|
||||
"given CKD and amiodarone interaction risk?\n\n{}".format(patient_context),
|
||||
@@ -443,7 +451,7 @@ contra_chain = clinical_context.retrieve(
|
||||
|
||||
</Tab>
|
||||
|
||||
<Tab title="Banking — Risk/Compliance">
|
||||
<Tab title="Banking: Risk/Compliance">
|
||||
|
||||
Regulatory compliance: Basel III (CRE20), BCBS 239, SR 11-7, and EBA IRRBB guidelines ingested as a graph where regulation articles cross-reference each other as edges. Multi-hop queries traverse those cross-references automatically, so a question about commercial real estate RWA pulls the relevant CRE20 paragraphs and the BCBS 239 data quality requirements that govern their calculation in a single call.
|
||||
|
||||
@@ -466,12 +474,17 @@ compliance_context = AgentContext(
|
||||
retention_days=2555, # 7-year regulatory retention
|
||||
)
|
||||
|
||||
# In production these come from ingest_file() — shown as strings here for brevity
|
||||
basel_cre20_text = "CRE20.32: For income-producing real estate where repayment depends on "
|
||||
"property cash flows, RWA = exposure × risk weight, where risk weight "
|
||||
"is determined by LTV bucket per Table CRE20.3..."
|
||||
bcbs239_text = "Principle 3: Risk data should be accurate and have a single authoritative source. "
|
||||
"Where data is aggregated across systems, reconciliation must be documented..."
|
||||
# In production the text comes from a parsed file, e.g. FileIngestor().ingest_file(path).text;
|
||||
# inline strings here for brevity
|
||||
basel_cre20_text = (
|
||||
"CRE20.32: For income-producing real estate where repayment depends on "
|
||||
"property cash flows, RWA = exposure × risk weight, where risk weight "
|
||||
"is determined by LTV bucket per Table CRE20.3..."
|
||||
)
|
||||
bcbs239_text = (
|
||||
"Principle 3: Risk data should be accurate and have a single authoritative source. "
|
||||
"Where data is aggregated across systems, reconciliation must be documented..."
|
||||
)
|
||||
|
||||
compliance_context.store(
|
||||
[
|
||||
@@ -482,7 +495,7 @@ compliance_context.store(
|
||||
extract_relationships=True,
|
||||
)
|
||||
|
||||
llm = LiteLLM(model="anthropic/claude-sonnet-4-20250514")
|
||||
llm = LiteLLM(model="anthropic/claude-sonnet-5")
|
||||
answer = compliance_context.query_with_reasoning(
|
||||
"Under Basel III CRE20, what are the RWA calculation requirements for "
|
||||
"commercial real estate exposures with LTV > 80%? "
|
||||
@@ -496,7 +509,7 @@ print(answer["response"])
|
||||
print("Regulatory sources cited: {}".format(answer["num_sources"]))
|
||||
print("Confidence: {:.1%}".format(answer["confidence"]))
|
||||
|
||||
# The reasoning path is the audit log — show it to the regulator
|
||||
# The reasoning path is the audit log: show it to the regulator
|
||||
print("\n--- Reasoning Path (audit log) ---")
|
||||
print(answer["reasoning_path"])
|
||||
```
|
||||
@@ -524,18 +537,18 @@ The `hybrid_alpha` parameter set in the `AgentContext` constructor establishes a
|
||||
When targeting a specific `anchor_node`, you can apply `proximity_weight` in `retrieve()` to dynamically blend structural distance from the anchor into the final score:
|
||||
|
||||
```python
|
||||
# Anchor node provided — let vector semantics lead, graph proximity only slightly boosts
|
||||
# Anchor node provided: let vector semantics lead, graph proximity only slightly boosts
|
||||
results = context.retrieve(
|
||||
query, use_graph=True, anchor_node="APT29", proximity_weight=0.2
|
||||
)
|
||||
|
||||
# Known-entity tracing — topology drives the retrieval
|
||||
# Known-entity tracing: topology drives the retrieval
|
||||
results = context.retrieve(
|
||||
query, use_graph=True, anchor_node="APT29", proximity_weight=0.8
|
||||
)
|
||||
```
|
||||
|
||||
Each additional hop in `max_hops` exponentially increases the subgraph size. Practical defaults by domain:
|
||||
Each additional expansion hop exponentially increases the subgraph size. Practical defaults by domain:
|
||||
|
||||
```text
|
||||
General Q&A max_expansion_hops=2 (95% of useful facts within 2 hops)
|
||||
@@ -544,7 +557,7 @@ Drug interactions max_expansion_hops=3 (drug → enzyme → metabolite
|
||||
Regulatory cross-ref max_expansion_hops=2 (rule → article → article)
|
||||
```
|
||||
|
||||
Set globally in the constructor; override per call with the `max_hops` argument to `retrieve()`.
|
||||
Expansion depth is a constructor setting only (`max_expansion_hops`); there is no per-call override on `retrieve()`. `query_with_reasoning()` does take a per-call `max_hops` argument.
|
||||
|
||||
## How GraphRAG works internally
|
||||
|
||||
@@ -576,9 +589,9 @@ The vector search and graph traversal run independently, then their scores are f
|
||||
|
||||
## Related Guides
|
||||
|
||||
- [Semantic Extraction](/guides/semantic-extraction) — build the graph from raw unstructured text
|
||||
- [Agent Memory](/guides/agent-memory) — store, retrieve, and persist agent memories
|
||||
- [Context Graphs](/guides/context-graphs) — build and traverse the knowledge graph directly
|
||||
- [Reasoning](reasoning) — derive new facts and run inference rules over the graph
|
||||
- [Decision Intelligence](/guides/decision-intelligence) — causal chains, policy enforcement, decision tracking
|
||||
- [LLM Integrations](/guides/llm-integrations) — connect Groq, OpenAI, Anthropic, HuggingFace, and 100+ more
|
||||
- [Semantic Extraction](/guides/semantic-extraction): build the graph from raw unstructured text
|
||||
- [Agent Memory](/guides/agent-memory): store, retrieve, and persist agent memories
|
||||
- [Context Graphs](/guides/context-graphs): build and traverse the knowledge graph directly
|
||||
- [Reasoning](/guides/reasoning): derive new facts and run inference rules over the graph
|
||||
- [Decision Intelligence](/guides/decision-intelligence): causal chains, policy enforcement, decision tracking
|
||||
- [LLM Integrations](/guides/llm-integrations): connect Groq, OpenAI, Anthropic, HuggingFace, and 100+ more
|
||||
|
||||
@@ -275,20 +275,20 @@ print(data)
|
||||
|
||||
**LiteLLM** is a universal adapter that provides a single interface to over 100 different LLM providers, including Anthropic Claude, Azure OpenAI, AWS Bedrock, Google Vertex AI, and local Ollama instances. It acts as a translation layer, converting your unified API calls into provider-specific requests, enabling easy switching between providers without code changes.
|
||||
|
||||
`LiteLLM` is the Swiss Army knife. It wraps the `litellm` library, which speaks to every major provider using a unified completion API. The model string encodes both provider and model name: `"anthropic/claude-sonnet-4-20250514"`, `"azure/gpt-4o"`, `"bedrock/anthropic.claude-3-5-sonnet-20241022-v2:0"`, `"ollama/llama3.2"`. Change the string, change the provider — no other code changes needed.
|
||||
`LiteLLM` is the Swiss Army knife. It wraps the `litellm` library, which speaks to every major provider using a unified completion API. The model string encodes both provider and model name: `"anthropic/claude-sonnet-5"`, `"azure/gpt-4o"`, `"bedrock/anthropic.claude-sonnet-4-5-20250929-v1:0"`, `"ollama/llama3.2"`. Change the string, change the provider — no other code changes needed.
|
||||
|
||||
```python
|
||||
from semantica.llms import LiteLLM
|
||||
|
||||
# Anthropic Claude — highest accuracy for complex reasoning
|
||||
llm = LiteLLM(model="anthropic/claude-sonnet-4-20250514")
|
||||
llm = LiteLLM(model="anthropic/claude-sonnet-5")
|
||||
# Reads ANTHROPIC_API_KEY from environment
|
||||
|
||||
# Azure OpenAI — compliance and data-residency requirements
|
||||
llm = LiteLLM(model="azure/gpt-4o", api_key="YOUR_AZURE_KEY")
|
||||
|
||||
# AWS Bedrock — existing cloud agreement, no new vendor
|
||||
llm = LiteLLM(model="bedrock/anthropic.claude-3-5-sonnet-20241022-v2:0")
|
||||
llm = LiteLLM(model="bedrock/anthropic.claude-sonnet-4-5-20250929-v1:0")
|
||||
|
||||
# Google Vertex AI
|
||||
llm = LiteLLM(model="vertex_ai/gemini-1.5-pro")
|
||||
@@ -306,7 +306,7 @@ The environment-variable convention for each provider: `ANTHROPIC_API_KEY`, `AZU
|
||||
import os
|
||||
|
||||
PROVIDER_MAP = {
|
||||
"prod": "anthropic/claude-sonnet-4-20250514",
|
||||
"prod": "anthropic/claude-sonnet-5",
|
||||
"staging": "openai/gpt-4o-mini",
|
||||
"local": "ollama/llama3.2",
|
||||
"azure": "azure/gpt-4o",
|
||||
@@ -378,7 +378,7 @@ print("FAST: {} (conf={:.0%})".format(fast_result["response"], fast_result["con
|
||||
|
||||
# Tier 2: deep answer with Claude if confidence is below threshold
|
||||
if fast_result["confidence"] < 0.85:
|
||||
deep_llm = LiteLLM(model="anthropic/claude-sonnet-4-20250514")
|
||||
deep_llm = LiteLLM(model="anthropic/claude-sonnet-5")
|
||||
deep_result = context.query_with_reasoning(
|
||||
query, llm_provider=deep_llm, max_results=15, max_hops=3
|
||||
)
|
||||
@@ -574,7 +574,7 @@ print("TRIAGE: {} (conf={:.0%})".format(triage["response"], triage["confidence"]
|
||||
|
||||
# Tier 2: escalate to Claude for deep analysis if Tier 1 is uncertain
|
||||
if triage["confidence"] < 0.88:
|
||||
deep_llm = LiteLLM(model="anthropic/claude-sonnet-4-20250514")
|
||||
deep_llm = LiteLLM(model="anthropic/claude-sonnet-5")
|
||||
deep = context.query_with_reasoning(
|
||||
"Full MITRE ATT&CK analysis of this alert: identify the attack chain, "
|
||||
"blast radius, affected systems, and recommended containment steps.",
|
||||
@@ -630,7 +630,7 @@ for d in drugs:
|
||||
# trastuzumab (conf=0.98), pertuzumab (conf=0.97), docetaxel (conf=0.96)
|
||||
|
||||
# Report synthesis with Claude — switch to azure/gpt-4o for HIPAA by changing one string
|
||||
report_llm = LiteLLM(model="anthropic/claude-sonnet-4-20250514")
|
||||
report_llm = LiteLLM(model="anthropic/claude-sonnet-5")
|
||||
# For HIPAA-constrained Azure deployment:
|
||||
# report_llm = LiteLLM(model="azure/gpt-4o", api_key="YOUR_AZURE_KEY")
|
||||
|
||||
@@ -682,7 +682,7 @@ question = (
|
||||
|
||||
# Two-provider consensus — same query, same graph, different LLMs
|
||||
gpt4o = OpenAI(model="gpt-4o", api_key="YOUR_OAI_KEY")
|
||||
claude = LiteLLM(model="anthropic/claude-sonnet-4-20250514")
|
||||
claude = LiteLLM(model="anthropic/claude-sonnet-5")
|
||||
|
||||
answer_a = context.query_with_reasoning(question, llm_provider=gpt4o, max_results=10)
|
||||
answer_b = context.query_with_reasoning(question, llm_provider=claude, max_results=10)
|
||||
|
||||
@@ -197,7 +197,7 @@ reasoning_agent.load("./pipeline/enriched_intel/")
|
||||
# All memories, graph nodes, and vector embeddings from both ingestion agents are now available.
|
||||
|
||||
# Use a high-capability model for the synthesis step
|
||||
llm = LiteLLM(model="anthropic/claude-sonnet-4-20250514")
|
||||
llm = LiteLLM(model="anthropic/claude-sonnet-5")
|
||||
|
||||
synthesis = reasoning_agent.query_with_reasoning(
|
||||
"Summarize the APT29 exploitation of CVE-2024-3400: affected products, "
|
||||
@@ -428,7 +428,7 @@ tier1.store(
|
||||
|
||||
# --- Tier 2: deep investigation when Tier 1 confidence is low ---
|
||||
if triage["confidence"] < 0.90:
|
||||
deep_llm = LiteLLM(model="anthropic/claude-sonnet-4-20250514")
|
||||
deep_llm = LiteLLM(model="anthropic/claude-sonnet-5")
|
||||
|
||||
investigation = tier2.query_with_reasoning(
|
||||
"Full MITRE ATT&CK analysis of incident {}. "
|
||||
@@ -533,7 +533,7 @@ t1.start(); t2.start()
|
||||
t1.join(); t2.join()
|
||||
|
||||
# Chief agent synthesizes across literature and experimental data
|
||||
llm = LiteLLM(model="anthropic/claude-sonnet-4-20250514")
|
||||
llm = LiteLLM(model="anthropic/claude-sonnet-5")
|
||||
|
||||
synthesis = chief.query_with_reasoning(
|
||||
"Identify the top two candidate compounds for KRAS G12C NSCLC that show "
|
||||
@@ -576,7 +576,7 @@ credit_officer = make_desk_agent()
|
||||
committee_chair = make_desk_agent()
|
||||
|
||||
app_id = "LOAN-2025-88421"
|
||||
llm = LiteLLM(model="anthropic/claude-sonnet-4-20250514")
|
||||
llm = LiteLLM(model="anthropic/claude-sonnet-5")
|
||||
|
||||
# --- Risk Desk: PD/LGD/EL analysis ---
|
||||
risk_desk.store(
|
||||
|
||||
@@ -477,7 +477,7 @@ regs = [
|
||||
]
|
||||
|
||||
# Use an LLM to extract the conceptual model from regulatory prose
|
||||
llm_gen = LLMOntologyGenerator(provider="anthropic", model="claude-sonnet-4-20250514")
|
||||
llm_gen = LLMOntologyGenerator(provider="anthropic", model="claude-sonnet-5")
|
||||
ontology = llm_gen.generate_ontology_from_text(
|
||||
"\n\n".join(r.text[:8000] for r in regs) # token-safe excerpt per document
|
||||
)
|
||||
|
||||
@@ -127,7 +127,7 @@ engine = ExecutionEngine(max_workers=4, retry_on_failure=True)
|
||||
result = engine.execute_pipeline(pipeline)
|
||||
|
||||
print(f"Success: {result.success}")
|
||||
print(f"Output: {result.output}") # {"node_count": 312, "edge_count": 847}
|
||||
print(f"Output: {result.output}") # the final step's return value, e.g. {"node_count": ..., "edge_count": ...}
|
||||
print(f"Duration: {result.metrics['execution_time']:.2f}s")
|
||||
print(f"Steps completed: {result.metrics['steps_executed']}")
|
||||
```
|
||||
@@ -197,7 +197,9 @@ engine = ExecutionEngine(
|
||||
max_workers = 4,
|
||||
retry_on_failure = True,
|
||||
)
|
||||
# The engine uses handler.get_retry_policy(step.step_type) when a step fails
|
||||
# ExecutionEngine builds its own FailureHandler; replace it with the configured one
|
||||
engine.failure_handler = handler
|
||||
# The engine now calls engine.failure_handler.get_retry_policy(step.step_type) on failure
|
||||
```
|
||||
|
||||
`handler.classify_error()` distinguishes `ValidationError` (low severity, usually don't retry), `ProcessingError` (high severity), and timeout/connection errors (medium severity, always retry). You can inspect the classification:
|
||||
|
||||
@@ -100,14 +100,15 @@ ner = NamedEntityRecognizer(
|
||||
methods=["llm", "ml", "pattern"],
|
||||
confidence_threshold=0.75,
|
||||
provider="anthropic",
|
||||
llm_model="claude-sonnet-4-6",
|
||||
llm_model="claude-sonnet-5",
|
||||
)
|
||||
entities = ner.extract_entities(report)
|
||||
|
||||
for e in entities:
|
||||
print("[{:>5.2f}] {:15s} {}".format(e.confidence, e.label, e.text))
|
||||
|
||||
# Expected output (abbreviated):
|
||||
# Illustrative output — exact labels and scores depend on the method and model.
|
||||
# Abbreviated:
|
||||
# [ 0.94] THREAT_ACTOR GAMMA-7
|
||||
# [ 0.91] THREAT_ACTOR DELTA-3
|
||||
# [ 0.97] MALWARE HAMMERTOSS
|
||||
@@ -262,16 +263,18 @@ from semantica.semantic_extract import TripletExtractor
|
||||
tri = TripletExtractor(
|
||||
method="llm",
|
||||
provider="anthropic",
|
||||
llm_model="claude-sonnet-4-6",
|
||||
llm_model="claude-sonnet-5",
|
||||
include_temporal=True, # attach time context to triplets when available
|
||||
include_provenance=True, # embed source document reference in each triplet
|
||||
validate=False, # return raw triplets; validate explicitly below
|
||||
)
|
||||
|
||||
# Feed in the entities and relations you already extracted — the extractor
|
||||
# uses them to constrain and validate what it produces
|
||||
# uses them to constrain what it produces
|
||||
triplets = tri.extract_triplets(report, entities, relations)
|
||||
|
||||
# Filter malformed triplets before serialisation
|
||||
# (extract_triplets validates automatically unless validate=False, as above)
|
||||
valid = tri.validate_triplets(triplets)
|
||||
print("Valid: {}/{}".format(len(valid), len(triplets)))
|
||||
|
||||
@@ -320,7 +323,7 @@ def ingest_intel_report(
|
||||
methods=[method, "pattern"],
|
||||
confidence_threshold=0.70,
|
||||
provider="anthropic",
|
||||
llm_model="claude-sonnet-4-6",
|
||||
llm_model="claude-sonnet-5",
|
||||
)
|
||||
entities = ner.extract_entities(text)
|
||||
classified = ner.classify_entities(entities)
|
||||
@@ -335,7 +338,7 @@ def ingest_intel_report(
|
||||
relation_types=["deployed", "targets", "exploits", "operates_from", "provided_to"],
|
||||
confidence_threshold=0.65,
|
||||
provider="anthropic",
|
||||
llm_model="claude-sonnet-4-6",
|
||||
llm_model="claude-sonnet-5",
|
||||
)
|
||||
relations = rel.extract_relations(text, entities)
|
||||
|
||||
@@ -347,9 +350,10 @@ def ingest_intel_report(
|
||||
tri = TripletExtractor(
|
||||
method=method,
|
||||
provider="anthropic",
|
||||
llm_model="claude-sonnet-4-6",
|
||||
llm_model="claude-sonnet-5",
|
||||
include_temporal=True,
|
||||
include_provenance=True,
|
||||
validate=False, # keep raw triplets so the summary can report rejections
|
||||
)
|
||||
triplets = tri.extract_triplets(text, entities, relations)
|
||||
valid = tri.validate_triplets(triplets)
|
||||
@@ -377,6 +381,7 @@ def ingest_intel_report(
|
||||
"coref_chains": len(chains),
|
||||
"relations": len(relations),
|
||||
"events": len(events),
|
||||
"triplets_total": len(triplets),
|
||||
"triplets_valid": len(valid),
|
||||
"graph_nodes": graph_stats.get("graph_nodes", 0),
|
||||
"graph_edges": graph_stats.get("graph_edges", 0),
|
||||
@@ -402,7 +407,7 @@ for text, doc_id in reports:
|
||||
summary["relations"],
|
||||
summary["events"],
|
||||
summary["triplets_valid"],
|
||||
len(summary["rdf_turtle"]),
|
||||
summary["triplets_total"],
|
||||
))
|
||||
```
|
||||
|
||||
@@ -421,7 +426,7 @@ ner = NamedEntityRecognizer(
|
||||
methods=["llm", "pattern"],
|
||||
confidence_threshold=0.75,
|
||||
provider="anthropic",
|
||||
llm_model="claude-sonnet-4-6",
|
||||
llm_model="claude-sonnet-5",
|
||||
)
|
||||
entities = ner.extract_entities(fintel_text)
|
||||
grouped = ner.classify_entities(entities)
|
||||
@@ -438,14 +443,14 @@ rel = RelationExtractor(
|
||||
relation_types=["operates_from", "deployed", "targets", "exploits"],
|
||||
confidence_threshold=0.70,
|
||||
provider="anthropic",
|
||||
llm_model="claude-sonnet-4-6",
|
||||
llm_model="claude-sonnet-5",
|
||||
)
|
||||
relations = rel.extract_relations(fintel_text, entities)
|
||||
|
||||
tri = TripletExtractor(
|
||||
method="llm",
|
||||
provider="anthropic",
|
||||
llm_model="claude-sonnet-4-6",
|
||||
llm_model="claude-sonnet-5",
|
||||
include_temporal=True,
|
||||
include_provenance=True,
|
||||
)
|
||||
@@ -544,14 +549,14 @@ rel = RelationExtractor(
|
||||
relation_types=["treats", "causes_adverse_event", "has_efficacy", "evaluated_in"],
|
||||
confidence_threshold=0.65,
|
||||
provider="anthropic",
|
||||
llm_model="claude-sonnet-4-6",
|
||||
llm_model="claude-sonnet-5",
|
||||
)
|
||||
relations = rel.extract_relations(paper, entities)
|
||||
|
||||
tri = TripletExtractor(
|
||||
method="llm",
|
||||
provider="anthropic",
|
||||
llm_model="claude-sonnet-4-6",
|
||||
llm_model="claude-sonnet-5",
|
||||
triplet_types=["treats", "has_efficacy", "causes_adverse_event"],
|
||||
include_temporal=True,
|
||||
include_provenance=True,
|
||||
@@ -595,7 +600,7 @@ ner = NamedEntityRecognizer(
|
||||
methods=["llm", "ml", "pattern"],
|
||||
confidence_threshold=0.70,
|
||||
provider="anthropic",
|
||||
llm_model="claude-sonnet-4-6",
|
||||
llm_model="claude-sonnet-5",
|
||||
)
|
||||
entities = ner.extract_entities(credit_memo)
|
||||
grouped = ner.classify_entities(entities)
|
||||
@@ -612,14 +617,14 @@ rel = RelationExtractor(
|
||||
relation_types=["guaranteed_by", "secured_by", "classified_as", "exposed_to"],
|
||||
confidence_threshold=0.65,
|
||||
provider="anthropic",
|
||||
llm_model="claude-sonnet-4-6",
|
||||
llm_model="claude-sonnet-5",
|
||||
)
|
||||
relations = rel.extract_relations(credit_memo, entities)
|
||||
|
||||
tri = TripletExtractor(
|
||||
method="llm",
|
||||
provider="anthropic",
|
||||
llm_model="claude-sonnet-4-6",
|
||||
llm_model="claude-sonnet-5",
|
||||
include_temporal=True,
|
||||
include_provenance=True,
|
||||
)
|
||||
|
||||
+25
-304
@@ -1,109 +1,31 @@
|
||||
---
|
||||
title: "Semantica"
|
||||
description: "The Accountability and Context Layer for AI: Context Graphs · Decision Intelligence · Full Provenance"
|
||||
title: "Welcome to Semantica"
|
||||
description: "The Context and Semantic Layer for AI in High-Stakes Domains: Context Graphs · Decision Intelligence · Full Provenance"
|
||||
---
|
||||
|
||||
```bash
|
||||
pip install semantica
|
||||
```
|
||||
|
||||
Your AI agent just made a decision. Now someone needs to explain it.
|
||||
Most AI agents run on embeddings, not meaning. A similarity score has no structure, no relationships, and no way to explain why a result came back.
|
||||
|
||||
*What did it know at the time? Which facts shaped the outcome? Where did those facts come from? Has it made the same call before: and did that go well?*
|
||||
Semantica is the semantic and context layer underneath your LLM, vector store, and agent framework: deterministic infrastructure, not a model. Graph construction, reasoning, and provenance all run without an LLM in the loop. It turns fragmented enterprise data into a structured, queryable context graph and knowledge graph, governed by ontologies, taxonomies, and controlled vocabularies (OWL, SHACL, SKOS), so your data's meaning is explicit rather than approximated by an embedding.
|
||||
|
||||
If your stack can't answer those questions with a traceable record, you have a gap. Not a capability gap: an **accountability gap**. It's the reason AI hasn't landed at scale in healthcare, finance, legal, and government. And it's why teams building for those markets keep rebuilding the same guardrails from scratch.
|
||||
Provenance and audit trails aren't a bolt-on. They fall out naturally once your data has that structure, so the same graph that powers retrieval and reasoning also gives you a straight answer when a regulator asks why.
|
||||
|
||||
**Semantica closes that gap.** It's the context and accountability layer that sits beneath your existing agent framework: not a replacement for LangChain or LlamaIndex, but the infrastructure that makes their outputs trustworthy.
|
||||
## What you get
|
||||
|
||||
|
||||
## The Problem Every Production AI Team Hits
|
||||
|
||||
Powerful agents aren't automatically trustworthy ones. Five structural blind spots make modern AI systems impossible to deploy in regulated environments:
|
||||
|
||||
**No memory structure** — agents store embeddings, not meaning
|
||||
- No way to ask *why* a fact was recalled
|
||||
- No link from a recalled fact back to its source document
|
||||
- Context is a black box that resets on every run
|
||||
|
||||
**No decision trail** — agents act continuously but record nothing
|
||||
- No history to hand to a regulator or auditor
|
||||
- No way to replay or reproduce a past decision
|
||||
- Debugging means re-running, not reviewing
|
||||
|
||||
**No provenance** — outputs can't be traced to source facts
|
||||
- In healthcare, finance, and legal: this is a hard compliance blocker
|
||||
- No lineage from inference back to the original document
|
||||
- Impossible to demonstrate what the agent actually relied on
|
||||
|
||||
**No reasoning transparency** — black-box answers with no explanation
|
||||
- Impossible to validate the reasoning path
|
||||
- Impossible to contest a specific conclusion
|
||||
- No basis for improving or correcting future behavior
|
||||
|
||||
**No conflict detection** — contradictory facts silently coexist in vector stores
|
||||
- No detection when two sources disagree
|
||||
- Outputs become inconsistent and unpredictable over time
|
||||
- Silent failures compound as the knowledge base grows
|
||||
|
||||
<Note>
|
||||
These aren't edge cases. They're why enterprise AI pilots stall: and why your compliance team keeps saying *not yet*.
|
||||
</Note>
|
||||
|
||||
|
||||
## What Semantica Adds to Your Stack
|
||||
|
||||
Semantica gives every agent the infrastructure it needs to be accountable. Drop it into your existing setup in minutes:
|
||||
|
||||
**Context Graphs** — a structured, queryable graph of everything your agent knows, decides, and reasons about
|
||||
- Persistent across agent runs: no context loss between sessions
|
||||
- Queryable with SPARQL and full graph algorithms
|
||||
- Temporal model with `valid_from` / `valid_until` on nodes and edges
|
||||
- Point-in-time snapshots of the full knowledge state
|
||||
|
||||
**Decision Intelligence** — every decision is a first-class object in your system
|
||||
- `record_decision()` captures full lifecycle and causal chain
|
||||
- Hybrid precedent search over past decisions for consistency
|
||||
- `analyze_decision_impact()` shows downstream consequences
|
||||
- Causal chain visualization from trigger to outcome
|
||||
|
||||
**Full Provenance** — every fact links to its source document and ingestion event
|
||||
- W3C PROV-O compliant lineage across all modules
|
||||
- Full traceability from raw input to final inference
|
||||
- `recorded_at` stamping with OWL-Time export
|
||||
- Audit-ready for HIPAA, SOX, GDPR, FDA 21 CFR Part 11
|
||||
|
||||
**Reasoning Engines** — explainable reasoning paths, not black boxes
|
||||
- Forward chaining, Rete, deductive, abductive
|
||||
- SPARQL query-based inference over RDF graphs
|
||||
- Datalog with recursive Horn clause rules
|
||||
- Every conclusion backed by a traceable derivation path
|
||||
|
||||
**Temporal Intelligence** — your graph knows not just *what*, but *when*
|
||||
- Allen interval algebra: all 13 temporal relations
|
||||
- Point-in-time queries over historical graph states
|
||||
- Temporal provenance stamping on every fact
|
||||
- OWL-Time export for standards-compliant archiving
|
||||
|
||||
**Ontology Hub** — full ontology lifecycle in the browser
|
||||
- Visual editor for schema design and editing
|
||||
- SHACL Studio for constraint authoring and validation
|
||||
- Alignment authoring across multiple ontologies
|
||||
- Health dashboard and version control built in
|
||||
- **[Context graphs](/guides/context-graphs)**: a persistent, queryable graph of everything your agent knows, decides, and reasons about
|
||||
- **Decision intelligence**: `record_decision()` captures the full lifecycle and causal chain of every decision
|
||||
- **[Full provenance](/guides/provenance)**: every fact links back to its source, W3C PROV-O compliant and audit-ready for HIPAA, SOX, and GDPR
|
||||
- **[Explainable reasoning](/guides/reasoning)**: forward chaining, Datalog, and SPARQL, each with a derivation path you can inspect
|
||||
- **Temporal intelligence**: Allen interval algebra and point-in-time snapshots, so the graph knows not just *what* but *when*
|
||||
|
||||
<Tip>
|
||||
Works alongside any LLM provider and any agent framework: add it to an existing stack without changing your architecture.
|
||||
Works alongside any LLM provider and any agent framework, and ingests directly from enterprise data platforms like Databricks, SAP, Salesforce, and Snowflake. Add it to an existing stack without changing your architecture.
|
||||
</Tip>
|
||||
|
||||
<img src="/assets/img/diagrams/architecture-overview.svg" alt="Semantica four-layer architecture: Ingestion → Processing → Intelligence → Application" style={{ width: '100%', borderRadius: '12px', margin: '24px 0' }} />
|
||||
|
||||
|
||||
## See It In Action
|
||||
|
||||
One pip install. A few lines to connect your agent. Everything else becomes traceable.
|
||||
|
||||
```bash
|
||||
pip install semantica
|
||||
```
|
||||
## Try it
|
||||
|
||||
<CodeGroup>
|
||||
|
||||
@@ -185,229 +107,28 @@ decision_id = context.record_decision(
|
||||
|
||||
</CodeGroup>
|
||||
|
||||
- [Full Quickstart](/quickstart) — Step-by-step pipeline walkthrough
|
||||
- [Cookbook](/cookbook) — 40+ real-world Jupyter notebooks
|
||||
- [Join Discord](https://discord.gg/sV34vps5hH) — Community chat and support
|
||||
|
||||
|
||||
## Built for Where Mistakes Have Consequences
|
||||
|
||||
Semantica was designed for domains where every decision must be explainable and every fact must be traceable.
|
||||
|
||||
<Warning>
|
||||
**This is system-level explainability, not foundation-model explainability.** Semantica does not expose, reconstruct, or explain what happens *inside* the LLM/foundation model — its internal reasoning or chain-of-thought stays opaque, as it does for any external system. What Semantica explains is *outside* the model: the context and data fed in, the decision produced, its provenance, the relevant relationships, the policies applied, and the full execution trail. See [Core Concepts](/concepts) for the full scope note.
|
||||
</Warning>
|
||||
|
||||
**Healthcare & Life Sciences**
|
||||
- Clinical decision support with full audit trails
|
||||
- Drug interaction and contraindication graphs
|
||||
- Patient safety event tracking and root-cause analysis
|
||||
- HIPAA-compliant provenance chains out of the box
|
||||
|
||||
**Finance & Risk**
|
||||
- Fraud detection knowledge graphs
|
||||
- Risk assessment trails built to survive an audit
|
||||
- SOX, GDPR, and MiFID II compliance infrastructure
|
||||
- Model decision lineage for regulatory reporting
|
||||
|
||||
**Legal & Compliance**
|
||||
- Evidence-backed research with every cited fact provenance-linked
|
||||
- Contract analysis with traceable clause extraction
|
||||
- Regulatory change tracking across jurisdictions
|
||||
- Full reasoning paths ready for court-admissible documentation
|
||||
|
||||
**Cybersecurity**
|
||||
- Threat attribution graphs linking actors, TTPs, and indicators
|
||||
- Incident response timelines with full event provenance
|
||||
- Security audit trails across the complete kill chain
|
||||
- MITRE ATT&CK-aligned knowledge graph integration
|
||||
|
||||
**Government & Defense**
|
||||
- Policy decision trails from brief to outcome
|
||||
- Classified information handling with provenance chains
|
||||
- Chain-of-custody scrutiny for intelligence reporting
|
||||
- Air-gapped deployment with local LLM support
|
||||
|
||||
**Critical Infrastructure**
|
||||
- Power grid state tracking with temporal intelligence
|
||||
- Transportation safety event graphs
|
||||
- Emergency response coordination with decision audit trails
|
||||
- Consequence modeling for high-stakes operational decisions
|
||||
|
||||
|
||||
## Start Here
|
||||
## Start here
|
||||
|
||||
<Steps>
|
||||
<Step title="Install Semantica">
|
||||
<Step title="Install">
|
||||
```bash
|
||||
pip install semantica
|
||||
```
|
||||
See [Installation](/installation) for optional extras (`[all]`, `[neo4j]`, `[pinecone]`) and environment setup.
|
||||
Optional extras: `[all]`, `[neo4j]`, `[pinecone]`. See [Installation](/installation).
|
||||
</Step>
|
||||
<Step title="Run the Quickstart">
|
||||
Build a complete knowledge graph pipeline in [5 minutes](/quickstart):
|
||||
- Ingest documents from any source
|
||||
- Extract entities and relationships
|
||||
- Build and query the graph
|
||||
- Record and trace a decision
|
||||
<Step title="Build a pipeline">
|
||||
Follow the [Quickstart](/quickstart) to ingest documents, extract entities, build a graph, and record a decision in 5 minutes.
|
||||
</Step>
|
||||
<Step title="Learn the mental model">
|
||||
[Core Concepts](/concepts) covers:
|
||||
- Knowledge graphs vs. vector stores: when to use each
|
||||
- What GraphRAG is and how Semantica implements it
|
||||
- How provenance and decision tracking work together
|
||||
- The accountability layer architecture
|
||||
<Step title="Learn the model">
|
||||
[Core Concepts](/concepts) covers knowledge graphs vs. vector stores, GraphRAG, and how provenance and decisions fit together.
|
||||
</Step>
|
||||
<Step title="Go deep on any module">
|
||||
Every module has a dedicated [reference page](/reference/context) with:
|
||||
- Full class and method documentation
|
||||
- Parameter tables with types and defaults
|
||||
- Runnable code examples for each feature
|
||||
<Step title="Go deep">
|
||||
Every module has a [reference page](/reference/context) with full API docs and runnable examples.
|
||||
</Step>
|
||||
</Steps>
|
||||
|
||||
- [Installation](/installation) — Get Semantica installed in under a minute
|
||||
- [Quickstart](/quickstart) — Build a complete knowledge graph pipeline in 5 minutes
|
||||
- [Core Concepts](/concepts) — The mental model behind the API
|
||||
- [API Reference](/reference/context) — Exact module, class, and method details
|
||||
- [Cookbook](/cookbook) — Domain notebooks for real-world use cases
|
||||
- [Changelog](https://github.com/semantica-agi/semantica/releases) — Release history
|
||||
|
||||
|
||||
## Full Capabilities
|
||||
|
||||
<AccordionGroup>
|
||||
|
||||
<Accordion title="Context & Decision Intelligence" icon="brain">
|
||||
|
||||
### Context Graphs
|
||||
|
||||
- Structured, persistent graph of entities, relationships, and decisions
|
||||
- Temporal model with `valid_from` / `valid_until` on every node and edge
|
||||
- Point-in-time queries across historical graph states
|
||||
- Distance Intelligence: semantic neighborhoods and N×N distance matrices
|
||||
|
||||
### Decision Tracking
|
||||
|
||||
- `record_decision()` with full lifecycle management and causal chains
|
||||
- Hybrid similarity search over past decisions for consistency enforcement
|
||||
- `analyze_decision_impact()` and `analyze_decision_influence()` for consequence modeling
|
||||
- Ego-mode exploration for targeted neighborhood investigation
|
||||
More: the [Cookbook](/cookbook) for real-world notebooks, [Discord](https://discord.gg/sV34vps5hH) for help.
|
||||
|
||||
<Accordion title="Full module list">
|
||||
`semantica.ingest`, `semantica.parse`, `semantica.split`, `semantica.normalize`, `semantica.semantic_extract`, `semantica.kg`, `semantica.ontology`, `semantica.reasoning`, `semantica.embeddings`, `semantica.vector_store`, `semantica.graph_store`, `semantica.triplet_store`, `semantica.context`, `semantica.provenance`, `semantica.change_management`, `semantica.deduplication`, `semantica.conflicts`, `semantica.export`, `semantica.visualization`, `semantica.pipeline`, `semantica.seed`, `semantica.llms`, `semantica.mcp_server`, `semantica.explorer`, `semantica.evals`, `semantica.utils`, `semantica.core`. See the [API Reference](/reference/context) for full docs on each.
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Knowledge Engineering" icon="diagram-project">
|
||||
|
||||
### Entity & Relation Extraction
|
||||
|
||||
- Named entity recognition: pattern, ML, or LLM methods
|
||||
- Typed triplet extraction via LLM or rule-based pipelines
|
||||
- Event extraction with temporal and causal linking
|
||||
|
||||
### Ontology & Schema
|
||||
|
||||
- Ontology Hub: visual editor, SHACL Studio, alignments, health dashboard
|
||||
- Deduplication v2: `blocking_v2`, `hybrid_v2`, `semantic_v2`: up to 7x faster
|
||||
- Datalog reasoning: recursive Horn clause rules with fixpoint semantics
|
||||
- SPARQL reasoning: query-based inference over RDF graphs
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Provenance & Auditability" icon="shield-check">
|
||||
|
||||
### Lineage Tracking
|
||||
|
||||
- W3C PROV-O lineage across all modules: every fact has a source
|
||||
- `recorded_at` stamping with full OWL-Time export
|
||||
- Change management with SHA-256 checksums and version control
|
||||
- Full audit trails from ingestion event to final inference
|
||||
|
||||
### Compliance Infrastructure
|
||||
|
||||
- HIPAA: patient data handling with audit-ready provenance chains
|
||||
- SOX / MiFID II: financial decision records with full traceability
|
||||
- GDPR: data lineage for subject access and right-to-erasure workflows
|
||||
- FDA 21 CFR Part 11: electronic records and signature compliance
|
||||
|
||||
</Accordion>
|
||||
|
||||
<Accordion title="Data Ingestion & Export" icon="database">
|
||||
|
||||
### Ingestion Formats
|
||||
|
||||
- Documents: PDF, DOCX, HTML, PPTX, Docling layout analysis
|
||||
- Structured data: JSON, CSV, Excel, Parquet, XML
|
||||
- Sources: web crawl, SQL, Snowflake, feeds, email, code repositories, MCP
|
||||
|
||||
### Vector Stores
|
||||
|
||||
- FAISS, Pinecone, Weaviate, Qdrant, Milvus, PgVector, in-memory
|
||||
|
||||
### Graph Stores
|
||||
|
||||
- Neo4j, FalkorDB, Apache AGE, Amazon Neptune
|
||||
|
||||
### Export Formats
|
||||
|
||||
- RDF: Turtle, JSON-LD, N-Triples, RDF/XML
|
||||
- Tabular: Parquet, CSV, Arrow
|
||||
- Graph: GraphML, GEXF, DOT, ArangoDB AQL
|
||||
- Ontology: OWL, SKOS, SHACL
|
||||
|
||||
</Accordion>
|
||||
|
||||
</AccordionGroup>
|
||||
|
||||
|
||||
## Module Reference
|
||||
|
||||
| Module | What it provides |
|
||||
| :-------- | :----------------- |
|
||||
| `semantica.context` | Context graphs, agent memory, decision tracking, causal analysis, precedent search |
|
||||
| `semantica.kg` | KG construction, graph algorithms, temporal model, Allen interval algebra |
|
||||
| `semantica.semantic_extract` | NER, relation extraction, event extraction, triplet generation |
|
||||
| `semantica.reasoning` | Forward chaining, Rete, deductive, abductive, SPARQL, Datalog |
|
||||
| `semantica.ontology` | SHACL, SKOS, alignments, diff/migration, auto-generation, OWL/RDF |
|
||||
| `semantica.explorer` | FastAPI Knowledge Explorer, Ontology Hub, Distance Intelligence, SHACL Studio |
|
||||
| `semantica.mcp_server` | MCP stdio server: 15 tools for Claude Desktop, VS Code, Cursor, Windsurf, Cline |
|
||||
| `semantica.vector_store` | FAISS, Pinecone, Weaviate, Qdrant, Milvus, PgVector |
|
||||
| `semantica.graph_store` | Neo4j, FalkorDB, Apache AGE, Amazon Neptune |
|
||||
| `semantica.triplet_store` | In-memory and persistent RDF triple store with SPARQL |
|
||||
| `semantica.ingest` | Files, web, feeds, databases, Snowflake, Parquet, XML, MCP |
|
||||
| `semantica.parse` | Document parsing: PDF, DOCX, HTML, PPTX, Docling layout analysis |
|
||||
| `semantica.split` | Text chunking: sentence, paragraph, token, semantic boundary strategies |
|
||||
| `semantica.normalize` | Text normalization, entity canonicalization, whitespace and encoding cleanup |
|
||||
| `semantica.embeddings` | Sentence-Transformers, FastEmbed, OpenAI, BGE, Ollama local embeddings |
|
||||
| `semantica.pipeline` | Pipeline DSL, parallel workers, retry policies, failure handling |
|
||||
| `semantica.export` | RDF, Parquet, ArangoDB AQL, CSV, OWL, Arrow, GraphML, GEXF, DOT |
|
||||
| `semantica.visualization` | Programmatic graph rendering: force, hierarchical, circular, spring layouts |
|
||||
| `semantica.deduplication` | Entity deduplication v1/v2, similarity scoring, blocking, merging |
|
||||
| `semantica.conflicts` | Conflict detection and resolution across overlapping knowledge sources |
|
||||
| `semantica.provenance` | W3C PROV-O lineage tracking, source attribution, audit trails |
|
||||
| `semantica.change_management` | Version control with SHA-256 checksums, diff, rollback |
|
||||
| `semantica.llms` | Groq, OpenAI, Anthropic, Gemini, Ollama, DeepSeek, Novita AI, LiteLLM, HuggingFace |
|
||||
| `semantica.seed` | Foundation graph seeding from CSV, JSON, SQL, API, and RDF sources |
|
||||
| `semantica.evals` | Evaluation harness: KG quality, extraction F1, pipeline benchmarking, regression tracking |
|
||||
| `semantica.core` | Orchestration, ConfigManager, LifecycleManager, PluginRegistry, MethodRegistry |
|
||||
| `semantica.utils` | Logging, validation, progress tracking, hash utilities, nested dict helpers |
|
||||
|
||||
|
||||
## Why Semantica?
|
||||
|
||||
**Open Source, MIT** — No vendor lock-in. No paywalled features.
|
||||
- Full source available on GitHub
|
||||
- Every line auditable by your security team
|
||||
- Fork, extend, and self-host with no restrictions
|
||||
- No telemetry, no usage reporting
|
||||
|
||||
**Production Ready** — Built for teams that can't afford surprises.
|
||||
- 1,000+ passing tests with full regression coverage
|
||||
- `PipelineValidator` catches configuration errors at startup
|
||||
- `FailureHandler` with exponential backoff and dead-letter queues
|
||||
- Ongoing security hardening: fixes shipped in every release ([CHANGELOG](https://github.com/semantica-agi/semantica/blob/main/CHANGELOG.md))
|
||||
|
||||
**Modular by Design** — Import only what you need.
|
||||
- Use `NERExtractor` without a graph store
|
||||
- Use `ContextGraph` without vector storage
|
||||
- Every component independently swappable and testable
|
||||
- No framework lock-in: works with any agent stack
|
||||
|
||||
@@ -12,13 +12,13 @@ icon: "link"
|
||||
pip install "semantica[langchain]"
|
||||
```
|
||||
|
||||
Requires `langchain-core >= 0.3`. If langchain-core is not installed, the integration still imports — every class carries the full Semantica API and degrades gracefully (`build()` returns `None`; branch on `LANGCHAIN_AVAILABLE`).
|
||||
Requires `langchain-core >= 0.3`. If langchain-core is not installed, the integration still imports. Every class carries the full Semantica API and degrades gracefully (`build()` returns `None`; branch on `LANGCHAIN_AVAILABLE`).
|
||||
|
||||
## Components at a Glance
|
||||
|
||||
- **SemanticaRetriever** — `BaseRetriever`: hybrid-search seeds retrieval, then graph edges are walked `hops` steps (default 2) for GraphRAG-style results.
|
||||
- **SemanticaVectorStore** — `VectorStore`: `add_texts` / `similarity_search` / `similarity_search_with_score` / `from_texts` over `HybridSearch`.
|
||||
- **SemanticaKGTool** / **SemanticaDecisionTool** — `BaseTool` subclasses: `semantica_query_graph` and `semantica_query_decisions` for LangGraph / tool-calling agents.
|
||||
- **SemanticaRetriever** (`BaseRetriever`): hybrid-search seeds retrieval, then graph edges are walked `hops` steps (default 2) for GraphRAG-style results.
|
||||
- **SemanticaVectorStore** (`VectorStore`): `add_texts` / `similarity_search` / `similarity_search_with_score` / `from_texts` over `HybridSearch`.
|
||||
- **SemanticaKGTool** / **SemanticaDecisionTool** (`BaseTool` subclasses): `semantica_query_graph` and `semantica_query_decisions` for LangGraph / tool-calling agents.
|
||||
|
||||
## Component Details
|
||||
|
||||
|
||||
+163
-122
@@ -28,7 +28,9 @@ Semantica is organized into **27 modules** across six logical layers. Each modul
|
||||
|
||||
### Ingest
|
||||
|
||||
Loads data from files, web, databases, and streams into a unified `SourceDocument` format.
|
||||
Loads data from files, web, databases, and streams. Each ingestor returns its own
|
||||
result type (`FileIngestor` → `FileObject`, `WebIngestor` → `WebContent`, …);
|
||||
document-oriented ones expose a `.text` payload and `.metadata`.
|
||||
|
||||
```python
|
||||
from semantica.ingest import FileIngestor, WebIngestor, ParquetIngestor, XMLIngestor, DatabricksIngestor
|
||||
@@ -37,7 +39,7 @@ from semantica.ingest import FileIngestor, WebIngestor, ParquetIngestor, XMLInge
|
||||
ingestor = FileIngestor()
|
||||
documents = ingestor.ingest_directory("data/")
|
||||
|
||||
# Web crawl
|
||||
# Web page: returns a WebContent with .text, .title, .links, .metadata
|
||||
web_ingestor = WebIngestor()
|
||||
page = web_ingestor.ingest_url("https://example.com")
|
||||
|
||||
@@ -67,13 +69,13 @@ Extracts structured text and layout metadata from raw documents.
|
||||
```python
|
||||
from semantica.parse import DocumentParser, DoclingParser
|
||||
|
||||
# Standard parser: all common formats
|
||||
# Standard parser: all common formats. parse() takes a path, returns a dict
|
||||
parser = DocumentParser()
|
||||
parsed = parser.parse_document("document.pdf")
|
||||
parsed = parser.parse("document.pdf") # {"full_text": ..., "metadata": ..., ...}
|
||||
|
||||
# Advanced parser: multi-column PDFs, merged-cell tables, OCR
|
||||
parser = DoclingParser(extract_tables=True, extract_images=True, output_format="markdown")
|
||||
parsed = parser.parse("data/annual_report.pdf")
|
||||
# Advanced parser (pip install semantica[parse-docling]): tables, OCR, layout
|
||||
parser = DoclingParser(export_format="markdown", enable_ocr=True)
|
||||
parsed = parser.parse("data/annual_report.pdf") # dict with full_text, tables, pages
|
||||
```
|
||||
|
||||
**Available parsers:** `DocumentParser`, `DoclingParser`, `CodeParser`, `CSVParser`, `DocxParser`, `EmailParser`, `ExcelParser`, `HTMLParser`, `ImageParser`, `JSONParser`, `MCPParser`, `MediaParser`, `PDFParser`, `PPTXParser`, `StructuredDataParser`, `WebParser`, `XMLParser`
|
||||
@@ -85,11 +87,12 @@ Chunks text for embedding and RAG pipelines with awareness of semantic boundarie
|
||||
```python
|
||||
from semantica.split import TextSplitter
|
||||
|
||||
splitter = TextSplitter(method="semantic_transformer")
|
||||
chunks = splitter.split(text, chunk_size=1000, chunk_overlap=200)
|
||||
# chunk_size / chunk_overlap are constructor arguments
|
||||
splitter = TextSplitter(method="semantic_transformer", chunk_size=1000, chunk_overlap=200)
|
||||
chunks = splitter.split(text)
|
||||
```
|
||||
|
||||
**Chunking strategies:** `recursive`, `semantic_transformer`, `entity_aware`, `relation_aware`, `sliding_window`, `structural`
|
||||
**Chunking methods:** `recursive`, `token`, `sentence`, `paragraph`, `semantic_transformer`, `entity_aware`, `relation_aware`, `graph_based`, `ontology_aware`, `hierarchical`, `community_detection`, `centrality_based`, `llm`
|
||||
|
||||
### Normalize
|
||||
|
||||
@@ -115,17 +118,18 @@ Named entity recognition, relation extraction, and triplet generation.
|
||||
```python
|
||||
from semantica.semantic_extract import NERExtractor, RelationExtractor, TripletExtractor
|
||||
|
||||
ner = NERExtractor(method="llm", llm_provider=llm)
|
||||
entities = ner.extract("Apple Inc. was founded by Steve Jobs.")
|
||||
# LLM method: provider + llm_model select the backend; the API key comes from the env
|
||||
ner = NERExtractor(method="llm", provider="groq", llm_model="llama-3.3-70b-versatile")
|
||||
entities = ner.extract("Apple Inc. was founded by Steve Jobs.") # list[Entity]
|
||||
|
||||
rel = RelationExtractor(method="llm", llm_provider=llm)
|
||||
relationships = rel.extract(text, entities=entities)
|
||||
rel = RelationExtractor(method="llm", provider="groq", llm_model="llama-3.3-70b-versatile")
|
||||
relationships = rel.extract(text, entities=entities) # list[Relation]
|
||||
|
||||
trip = TripletExtractor(method="llm", llm_provider=llm)
|
||||
triplets = trip.extract(text)
|
||||
trip = TripletExtractor(method="pattern")
|
||||
triplets = trip.extract(text) # list[Triplet]
|
||||
```
|
||||
|
||||
**Extraction methods:** `"pattern"` (no API key), `"ml"` (local model), `"llm"` (any of the 8 supported providers)
|
||||
**Extraction methods:** `"pattern"` (no API key), `"ml"` (local spaCy model), `"llm"` (any of the 9 supported providers)
|
||||
|
||||
**Additional extractors:** `CoreferenceResolver`, `EventDetector`, `SemanticAnalyzer`, `SemanticNetworkExtractor`
|
||||
|
||||
@@ -137,17 +141,17 @@ Graph construction, graph algorithms, temporal model, and distance intelligence.
|
||||
from semantica.kg import GraphBuilder, GraphAnalyzer, TemporalGraphQuery, SimilarityCalculator
|
||||
from datetime import datetime
|
||||
|
||||
# Build
|
||||
# Build: build() takes a {"entities": ..., "relationships": ...} dict
|
||||
builder = GraphBuilder(merge_entities=True)
|
||||
kg = builder.build(entities=entities, relationships=relationships)
|
||||
kg = builder.build({"entities": entities, "relationships": relationships})
|
||||
|
||||
# Temporal graphs (v0.4.0)
|
||||
query_engine = TemporalGraphQuery(enable_temporal_reasoning=True)
|
||||
snapshot = query_engine.query_at_time(kg, query="", at_time=datetime(2021, 6, 15))
|
||||
|
||||
# Semantic similarity (v0.5.0)
|
||||
calc = SimilarityCalculator()
|
||||
scores = calc.calculate_similarity(entity_a, entity_b)
|
||||
# Semantic similarity (v0.5.0): operates on embedding vectors
|
||||
calc = SimilarityCalculator(method="cosine")
|
||||
score = calc.cosine_similarity(vec_a, vec_b)
|
||||
```
|
||||
|
||||
**Graph algorithms available:** centrality calculation, community detection, connectivity analysis, entity resolution, link prediction, path finding, similarity calculation
|
||||
@@ -175,19 +179,23 @@ Derives new facts from existing knowledge using multiple inference strategies.
|
||||
```python
|
||||
from semantica.reasoning import Reasoner, DatalogReasoner
|
||||
|
||||
# Rule-based reasoning
|
||||
# Forward chaining: facts and rules as predicate(args) / IF-THEN strings
|
||||
engine = Reasoner()
|
||||
engine.apply_transitivity("located_in")
|
||||
engine.apply_symmetry("knows")
|
||||
result = engine.infer()
|
||||
engine.add_fact("Manager(Alice)")
|
||||
engine.add_rule("IF Manager(?x) THEN HasAuthority(?x)")
|
||||
results = engine.forward_chain() # list[InferenceResult] with .conclusion, .rule_used
|
||||
|
||||
# Datalog: recursive Horn clause rules (v0.4.0)
|
||||
datalog = DatalogEngine()
|
||||
datalog = DatalogReasoner()
|
||||
datalog.add_fact("parent(tom, bob)")
|
||||
datalog.add_fact("parent(bob, ann)")
|
||||
datalog.add_rule("ancestor(X, Y) :- parent(X, Y).")
|
||||
datalog.add_rule("ancestor(X, Z) :- parent(X, Y), ancestor(Y, Z).")
|
||||
results = datalog.query("ancestor(alice, ?)")
|
||||
datalog.derive_all()
|
||||
results = datalog.query("ancestor(tom, ?Z)") # [{"Z": "bob"}, {"Z": "ann"}], order not guaranteed
|
||||
```
|
||||
|
||||
**Engines:** forward chaining, Rete network, deductive, abductive, SPARQL, Datalog: all produce explainable inference paths
|
||||
**Engines:** `Reasoner` (forward/backward chaining), `ReteEngine`, `SPARQLReasoner`, `DatalogReasoner`, `TemporalReasoningEngine`, `GraphReasoner` (LLM)
|
||||
|
||||
|
||||
## Storage
|
||||
@@ -199,9 +207,9 @@ Generates and manages vector embeddings for semantic similarity.
|
||||
```python
|
||||
from semantica.embeddings import EmbeddingGenerator
|
||||
|
||||
generator = EmbeddingGenerator(model="sentence-transformers")
|
||||
embeddings = generator.generate(["text1", "text2"])
|
||||
similarity = generator.similarity(embeddings[0], embeddings[1])
|
||||
generator = EmbeddingGenerator()
|
||||
embeddings = generator.generate_embeddings(["text1", "text2"]) # np.ndarray
|
||||
similarity = generator.compare_embeddings(embeddings[0], embeddings[1])
|
||||
```
|
||||
|
||||
**Supported models:** Sentence-Transformers, FastEmbed, OpenAI, BGE
|
||||
@@ -215,12 +223,18 @@ Multi-backend vector database with hybrid search support.
|
||||
```python
|
||||
from semantica.vector_store import VectorStore
|
||||
|
||||
store = VectorStore(backend="faiss", dimension=768)
|
||||
store.add_vectors(embeddings, ids)
|
||||
results = store.search(query_vector, top_k=10)
|
||||
store = VectorStore(backend="faiss", dimension=768)
|
||||
|
||||
# Raw vectors
|
||||
ids = store.store_vectors(embeddings) # returns generated ids
|
||||
hits = store.search_vectors(query_vector, k=10)
|
||||
|
||||
# Or store text and let the store embed it
|
||||
store.add_documents(["Apple was founded in 1976.", "Google was founded in 1998."])
|
||||
results = store.search("tech company founding dates", limit=10)
|
||||
```
|
||||
|
||||
**Backends:** FAISS, Pinecone, Weaviate, Qdrant, Milvus, PgVector, in-memory
|
||||
**Backends:** FAISS, Pinecone, Weaviate, Qdrant, Milvus, PgVector, SQLite, in-memory
|
||||
|
||||
**Search modes:** semantic top-k, hybrid (vector + keyword), metadata-filtered
|
||||
|
||||
@@ -232,8 +246,8 @@ Connects to graph databases for persistent, query-able storage.
|
||||
from semantica.graph_store import GraphStore
|
||||
|
||||
store = GraphStore(backend="neo4j")
|
||||
store.add_nodes(entities)
|
||||
store.add_edges(relationships)
|
||||
store.add_nodes([{"id": "acme", "type": "Organization", "properties": {"name": "Acme"}}])
|
||||
store.add_edges([{"source": "alice", "target": "acme", "type": "works_for"}])
|
||||
results = store.query("MATCH (n)-[r]->(m) RETURN n, r, m")
|
||||
```
|
||||
|
||||
@@ -246,9 +260,9 @@ RDF triple-based storage with SPARQL query support.
|
||||
```python
|
||||
from semantica.triplet_store import TripletStore
|
||||
|
||||
store = TripletStore(backend="blazegraph")
|
||||
store.add_triplets(subject, predicate, obj)
|
||||
results = store.sparql("SELECT ?s ?p ?o WHERE { ?s ?p ?o }")
|
||||
store = TripletStore(backend="oxigraph")
|
||||
store.add_triplets(triplets) # list of Triplet objects (or add_triplet for one)
|
||||
results = store.execute_query("SELECT ?s ?p ?o WHERE { ?s ?p ?o }")
|
||||
```
|
||||
|
||||
**Backends:** Oxigraph (embedded), Blazegraph, Apache Jena, RDF4J
|
||||
@@ -261,15 +275,18 @@ results = store.sparql("SELECT ?s ?p ?o WHERE { ?s ?p ?o }")
|
||||
Detects, scores, and merges duplicate entities across sources.
|
||||
|
||||
```python
|
||||
from semantica.deduplication import EntityResolver
|
||||
from semantica.deduplication import DuplicateDetector, EntityMerger
|
||||
|
||||
resolver = EntityResolver()
|
||||
merged = resolver.resolve(entities, strategy="semantic_v2")
|
||||
detector = DuplicateDetector(similarity_threshold=0.85)
|
||||
candidates = detector.detect_duplicates(entities)
|
||||
|
||||
merger = EntityMerger()
|
||||
operations = merger.merge_duplicates(entities, strategy="keep_most_complete")
|
||||
```
|
||||
|
||||
**v2 strategies** (`blocking_v2`, `hybrid_v2`, `semantic_v2`) are up to 7x faster than v1.
|
||||
**v2 candidate-generation modes** (`blocking_v2`, `hybrid_v2`, `semantic_v2`) are up to 7x faster than v1.
|
||||
|
||||
**Components:** `EntityResolver`, `DuplicateDetector`, `EntityMerger`, `SimilarityCalculator`, `ClusterBuilder`
|
||||
**Components:** `DuplicateDetector`, `EntityMerger`, `ClusterBuilder`, `MergeStrategyManager`
|
||||
|
||||
**`DuplicateDetector` options:** `max_results`, `top_k_per_entity`, `min_similarity`, `sort_by`
|
||||
|
||||
@@ -278,14 +295,13 @@ merged = resolver.resolve(entities, strategy="semantic_v2")
|
||||
Detects and resolves fact conflicts across overlapping knowledge sources.
|
||||
|
||||
```python
|
||||
from semantica.conflicts import ConflictDetector
|
||||
from semantica.conflicts import ConflictDetector, ConflictResolver
|
||||
|
||||
detector = ConflictDetector()
|
||||
conflicts = detector.detect_conflicts(kg)
|
||||
resolved = detector.resolve(conflicts, strategy="most_recent")
|
||||
conflicts = ConflictDetector().detect_conflicts(entities) # list of entity dicts
|
||||
resolved = ConflictResolver().resolve_conflicts(conflicts, strategy="most_recent")
|
||||
```
|
||||
|
||||
**Detection types:** value conflicts, type conflicts, temporal conflicts, logical conflicts
|
||||
**Detection types:** value conflicts, type conflicts, relationship conflicts, temporal conflicts, logical conflicts
|
||||
|
||||
**Resolution strategies:** prefer most recent, prefer most reliable source, majority vote, flag for manual review
|
||||
|
||||
@@ -298,6 +314,7 @@ Agent context graphs, decision tracking, causal chains, and precedent search.
|
||||
|
||||
```python
|
||||
from semantica.context import AgentContext, ContextGraph
|
||||
from semantica.vector_store import VectorStore
|
||||
|
||||
context = AgentContext(
|
||||
vector_store=VectorStore(backend="faiss", dimension=768),
|
||||
@@ -328,7 +345,7 @@ W3C PROV-O compliant lineage tracking across all modules.
|
||||
from semantica.provenance import ProvenanceManager
|
||||
|
||||
manager = ProvenanceManager()
|
||||
manager.track_entity("entity_1", "document.pdf", "person")
|
||||
manager.track_entity("entity_1", source="document.pdf", metadata={"type": "person"})
|
||||
lineage = manager.get_lineage("entity_1")
|
||||
```
|
||||
|
||||
@@ -364,8 +381,8 @@ RDFExporter().export(graph, file_path="graph.ttl", format="turtle")
|
||||
# Analytics
|
||||
ParquetExporter().export(graph, file_path="output/graph.parquet")
|
||||
|
||||
# ArangoDB
|
||||
aql = ArangoAQLExporter().export(graph)
|
||||
# ArangoDB: writes AQL INSERT statements to the given path
|
||||
ArangoAQLExporter().export(graph, file_path="graph.aql")
|
||||
```
|
||||
|
||||
**Export formats:** RDF (Turtle, JSON-LD, N-Triples, XML), Parquet, ArangoDB AQL, CSV, OWL, Arrow, LPG, YAML, distance matrices
|
||||
@@ -390,16 +407,24 @@ viz.visualize_network(graph, output="html", file_path="graph.html")
|
||||
Pipeline DSL with parallel workers, retry policies, and failure handling.
|
||||
|
||||
```python
|
||||
from semantica.pipeline import Pipeline
|
||||
from semantica.pipeline import PipelineBuilder, ExecutionEngine
|
||||
from semantica.ingest import FileIngestor
|
||||
from semantica.semantic_extract import NERExtractor
|
||||
|
||||
pipeline = Pipeline()
|
||||
pipeline.add_step("ingest", FileIngestor())
|
||||
pipeline.add_step("extract", NERExtractor())
|
||||
pipeline.add_step("build", GraphBuilder())
|
||||
result = pipeline.run("data/")
|
||||
builder = PipelineBuilder()
|
||||
|
||||
# Each step type dispatches to a handler you register (or supply explicitly)
|
||||
builder.register_step_handler("ingest", lambda data, **c: FileIngestor().ingest(c["source"]))
|
||||
builder.register_step_handler("extract", lambda docs, **c: NERExtractor(method="pattern").extract(docs[0].text))
|
||||
|
||||
builder.add_step("ingest", step_type="ingest", source="data/")
|
||||
builder.add_step("extract", step_type="extract")
|
||||
|
||||
pipeline = builder.connect_steps("ingest", "extract").build(name="docs_to_entities")
|
||||
result = ExecutionEngine().execute_pipeline(pipeline)
|
||||
```
|
||||
|
||||
**Components:** `Pipeline`, `PipelineBuilder`, `ExecutionEngine`, `FailureHandler`, `PipelineValidator`, `ParallelismManager`, `ResourceScheduler`
|
||||
**Components:** `PipelineBuilder`, `Pipeline`, `ExecutionEngine`, `FailureHandler`, `PipelineValidator`, `ParallelismManager`, `ResourceScheduler`
|
||||
|
||||
### Explorer
|
||||
|
||||
@@ -428,7 +453,7 @@ llm = OpenAI(model="gpt-4o", api_key=os.getenv("OPENAI_API_KEY"))
|
||||
llm = LiteLLM(model="anthropic/claude-opus-4-7", api_key=os.getenv("ANTHROPIC_API_KEY"))
|
||||
```
|
||||
|
||||
**Supported providers:** OpenAI, Anthropic, Google Gemini, Groq, Ollama, DeepSeek, Novita AI, LiteLLM (20+ models via one interface)
|
||||
**Supported providers:** OpenAI, Anthropic, Google Gemini, Groq, Ollama, DeepSeek, Novita AI, HuggingFace, plus LiteLLM (100+ models via one interface)
|
||||
|
||||
### MCP Server
|
||||
|
||||
@@ -445,44 +470,43 @@ python -m semantica.mcp_server
|
||||
Bootstrap knowledge graphs from verified structured sources: fixed-point reference data, controlled vocabularies, and domain anchors.
|
||||
|
||||
```python
|
||||
from semantica.seed import SeedManager
|
||||
from semantica.seed import SeedDataManager
|
||||
|
||||
seed = SeedManager()
|
||||
seed.populate(kg, dataset="companies", count=100)
|
||||
seed = SeedDataManager()
|
||||
|
||||
# Load domain seeds from file or built-in datasets
|
||||
seed.load_from_file("seed_data/industries.json")
|
||||
seed.inject(kg) # merges seed nodes without duplicating existing entities
|
||||
# Load trusted reference data from CSV / JSON / a database / an API
|
||||
seed_data = seed.load_from_csv("seed_data/industries.csv", entity_type="Industry")
|
||||
|
||||
# Merge seed data with extraction output (seed values win on conflict by default)
|
||||
combined = seed.integrate_with_extracted(
|
||||
{"entities": seed_data, "relationships": []},
|
||||
{"entities": extracted_entities, "relationships": extracted_relationships},
|
||||
merge_strategy="seed_first",
|
||||
)
|
||||
```
|
||||
|
||||
**Use cases:** anchoring extraction with known entities, pre-populating ontology classes, deterministic test graph generation.
|
||||
|
||||
### Evals
|
||||
|
||||
Evaluation framework for measuring KG quality, extraction accuracy, and pipeline performance.
|
||||
Scores decision-intelligence outputs (decision records, audit trails, reasoning
|
||||
text) with a registry of deterministic and model-backed evaluators plus a small
|
||||
run harness.
|
||||
|
||||
```python
|
||||
from semantica.evals import KGEvaluator, ExtractionEvaluator, PipelineEvaluator, RegressionTracker
|
||||
from semantica.evals import evaluate, list_evaluators
|
||||
|
||||
# KG quality
|
||||
report = KGEvaluator().evaluate(kg, ontology=ontology)
|
||||
print(f"Completeness: {report.completeness:.2%} Consistency: {report.consistency:.2%}")
|
||||
list_evaluators()
|
||||
# ['decision_scores', 'exact_match', 'keyword_check', 'length_range',
|
||||
# 'levenshtein', 'llm_as_judge', 'numeric_range', 'regex_match', 'rouge',
|
||||
# 'temporal_range']
|
||||
|
||||
# Extraction accuracy
|
||||
report = ExtractionEvaluator().evaluate_ner(predictions=extracted, gold_standard=annotated)
|
||||
print(f"Precision: {report.precision:.3f} Recall: {report.recall:.3f} F1: {report.f1:.3f}")
|
||||
|
||||
# Pipeline throughput and latency
|
||||
metrics = PipelineEvaluator().benchmark(pipeline, data="data/", bench_runs=5)
|
||||
print(f"Throughput: {metrics.docs_per_second:.1f} docs/sec")
|
||||
|
||||
# Regression tracking across runs
|
||||
tracker = RegressionTracker(db_path="eval_history.db")
|
||||
run_id = tracker.record_run(pipeline_version="v1.2.0", metrics=metrics)
|
||||
diff = tracker.compare(run_id, baseline_run_id="run_abc123")
|
||||
cases = [("apple", "aple"), ("night", "nacht")]
|
||||
summary = evaluate(cases, evaluators=["levenshtein"])
|
||||
print(summary.total, summary.passed, summary.pass_rate)
|
||||
```
|
||||
|
||||
**Components:** `KGEvaluator`, `ExtractionEvaluator`, `PipelineEvaluator`, `RegressionTracker`
|
||||
**Public API:** `evaluate(cases, evaluators, config=None)`, `list_evaluators()`, `get_evaluator(name)`, and the `EvalMetric` / `CaseResult` / `EvalSummary` result types. See the [Evals reference](/reference/evals).
|
||||
|
||||
### Core
|
||||
|
||||
@@ -491,20 +515,20 @@ Base classes, shared data models, and the plugin registry used across all module
|
||||
```python
|
||||
from semantica.core import Semantica, PluginRegistry, ConfigManager
|
||||
|
||||
# Top-level orchestrator
|
||||
sem = Semantica(config_path="config.yaml")
|
||||
# ConfigManager loads a Config; Config.get() does dotted lookups
|
||||
config = ConfigManager().load_from_file("config.yaml")
|
||||
batch = config.get("processing.batch_size", default=32)
|
||||
|
||||
# Top-level orchestrator: pass the Config object (or a dict), not a path
|
||||
sem = Semantica(config=config)
|
||||
sem.initialize()
|
||||
|
||||
# Plugin registry: register custom components
|
||||
# Plugin registry: register custom components under a name
|
||||
registry = PluginRegistry()
|
||||
registry.register("my_ingestor", MyCustomIngestor)
|
||||
|
||||
# Config management
|
||||
config = ConfigManager(config_path="config.yaml")
|
||||
batch = config.get("processing.batch_size", default=32)
|
||||
registry.register_plugin("my_ingestor", MyCustomIngestor, version="1.0.0")
|
||||
```
|
||||
|
||||
**Components:** `Semantica`, `PluginRegistry`, `ConfigManager`, `LifecycleManager`, `HealthMonitor`, `Config`
|
||||
**Components:** `Semantica`, `PluginRegistry`, `ConfigManager`, `Config`, `LifecycleManager`, `HealthStatus`, `MethodRegistry`
|
||||
|
||||
### Utils
|
||||
|
||||
@@ -532,11 +556,13 @@ from semantica.semantic_extract import NERExtractor, RelationExtractor
|
||||
from semantica.kg import GraphBuilder
|
||||
|
||||
sources = FileIngestor().ingest("data/")
|
||||
parsed = DocumentParser().parse(sources[0])
|
||||
entities = NERExtractor(method="llm", llm_provider=llm).extract(parsed)
|
||||
relationships = RelationExtractor(method="llm", llm_provider=llm).extract(parsed, entities=entities)
|
||||
text = DocumentParser().parse(sources[0].path)["full_text"]
|
||||
ner = NERExtractor(method="llm", provider="groq", llm_model="llama-3.3-70b-versatile")
|
||||
rel = RelationExtractor(method="llm", provider="groq", llm_model="llama-3.3-70b-versatile")
|
||||
entities = ner.extract(text)
|
||||
relationships = rel.extract(text, entities=entities)
|
||||
graph = GraphBuilder(merge_entities=True).build(
|
||||
entities=entities, relationships=relationships
|
||||
{"entities": entities, "relationships": relationships}
|
||||
)
|
||||
```
|
||||
|
||||
@@ -555,16 +581,20 @@ from semantica.vector_store import VectorStore
|
||||
context = AgentContext(
|
||||
vector_store=VectorStore(backend="faiss", dimension=768),
|
||||
knowledge_graph=ContextGraph(advanced_analytics=True),
|
||||
graph_expansion=True,
|
||||
)
|
||||
context.load_graph("company_kg.json")
|
||||
|
||||
result = context.query(
|
||||
# store() extracts entities and populates the graph + vector index
|
||||
context.store([{"content": "Steve Wozniak co-founded Apple with Steve Jobs."}])
|
||||
|
||||
# retrieve() blends vector similarity with multi-hop graph traversal
|
||||
results = context.retrieve(
|
||||
"What companies did Apple alumni found?",
|
||||
mode="graphrag",
|
||||
reasoning=True,
|
||||
use_graph=True,
|
||||
expand_graph=True,
|
||||
)
|
||||
for claim in result.claims:
|
||||
print(f"{claim.text} → {claim.source_node}")
|
||||
for r in results:
|
||||
print(f"[{r['score']:.3f}] {r['content']} (source: {r['source']})")
|
||||
```
|
||||
|
||||
**Best for:** question-answering systems, RAG with source attribution, research assistants
|
||||
@@ -606,18 +636,22 @@ precedents = context.find_precedents("model selection", limit=5)
|
||||
|
||||
```python
|
||||
from semantica.ingest import FileIngestor
|
||||
from semantica.parse import DocumentParser
|
||||
from semantica.semantic_extract import NERExtractor
|
||||
from semantica.kg import GraphBuilder
|
||||
from semantica.provenance import ProvenanceManager
|
||||
from semantica.export import RDFExporter
|
||||
|
||||
sources = FileIngestor().ingest("records/")
|
||||
entities = NERExtractor(method="llm", llm_provider=llm).extract(sources)
|
||||
graph = GraphBuilder(merge_entities=True).build(entities=entities, relationships=[])
|
||||
prov = ProvenanceManager()
|
||||
lineage = prov.get_entity_lineage("entity_id")
|
||||
ner = NERExtractor(method="llm", provider="groq", llm_model="llama-3.3-70b-versatile")
|
||||
entities = ner.extract(DocumentParser().parse(sources[0].path)["full_text"])
|
||||
graph = GraphBuilder(merge_entities=True).build({"entities": entities, "relationships": []})
|
||||
|
||||
RDFExporter(include_provenance=True).export(graph, file_path="audit.ttl", format="turtle")
|
||||
prov = ProvenanceManager()
|
||||
prov.track_entity("entity_id", source="records/filing.pdf", metadata={"extractor": "llm"})
|
||||
lineage = prov.get_lineage("entity_id")
|
||||
|
||||
RDFExporter().export(graph, file_path="audit.ttl", format="turtle")
|
||||
```
|
||||
|
||||
**Best for:** HIPAA, SOX, GDPR, FDA 21 CFR Part 11 deployments
|
||||
@@ -632,18 +666,25 @@ RDFExporter(include_provenance=True).export(graph, file_path="audit.ttl", format
|
||||
from semantica.ingest import WebIngestor
|
||||
from semantica.normalize import TextNormalizer
|
||||
from semantica.semantic_extract import NERExtractor, RelationExtractor
|
||||
from semantica.graph_store import Neo4jStore
|
||||
from semantica.graph_store import GraphStore
|
||||
from semantica.kg import GraphBuilder
|
||||
|
||||
pages = WebIngestor(max_depth=2).ingest("https://example.com")
|
||||
ingestor = WebIngestor()
|
||||
normalizer = TextNormalizer()
|
||||
store = Neo4jStore(uri="bolt://localhost:7687", user="neo4j", password="password")
|
||||
ner = NERExtractor(method="pattern")
|
||||
rel = RelationExtractor(method="pattern")
|
||||
|
||||
for page in pages:
|
||||
# The generic GraphStore wrapper exposes the add_nodes/add_edges interface
|
||||
# GraphBuilder persists through; a raw Neo4jStore does not
|
||||
store = GraphStore(backend="neo4j", uri="bolt://localhost:7687", user="neo4j", password="password")
|
||||
builder = GraphBuilder(merge_entities=True, graph_store=store)
|
||||
|
||||
for url in ["https://example.com/a", "https://example.com/b"]:
|
||||
page = ingestor.ingest_url(url) # WebContent, has .text
|
||||
text = normalizer.normalize_text(page.text)
|
||||
entities = NERExtractor().extract(text)
|
||||
relationships = RelationExtractor().extract(text, entities=entities)
|
||||
store.add_nodes(entities)
|
||||
store.add_edges(relationships)
|
||||
entities = ner.extract(text)
|
||||
relationships = rel.extract(text, entities=entities)
|
||||
builder.build({"entities": entities, "relationships": relationships})
|
||||
```
|
||||
|
||||
**Best for:** competitive intelligence, news monitoring, research aggregation
|
||||
@@ -692,8 +733,8 @@ versioner.create_snapshot(kg, "2024-Q1", author="user@example.com", description=
|
||||
| [vector_store](/reference/vector_store) | Vector database | `VectorStore` |
|
||||
| [graph_store](/reference/graph_store) | Graph database | `GraphStore` |
|
||||
| [triplet_store](/reference/triplet_store) | RDF triple store | `TripletStore` |
|
||||
| [deduplication](/reference/deduplication) | Entity resolution | `EntityResolver`, `DuplicateDetector`, `ClusterBuilder`, `MergeStrategyManager` |
|
||||
| [conflicts](/reference/conflicts) | Conflict resolution | `ConflictDetector` |
|
||||
| [deduplication](/reference/deduplication) | Entity resolution | `DuplicateDetector`, `EntityMerger`, `ClusterBuilder`, `MergeStrategyManager` |
|
||||
| [conflicts](/reference/conflicts) | Conflict resolution | `ConflictDetector`, `ConflictResolver`, `SourceTracker` |
|
||||
| [context](/reference/context) | Agent context & decisions | `AgentContext`, `ContextGraph` |
|
||||
| [provenance](/reference/provenance) | W3C PROV-O lineage | `ProvenanceManager` |
|
||||
| [change_management](/reference/change_management) | Version control | `TemporalVersionManager` |
|
||||
@@ -703,8 +744,8 @@ versioner.create_snapshot(kg, "2024-Q1", author="user@example.com", description=
|
||||
| [explorer](/reference/explorer) | Knowledge Explorer UI | `semantica-explorer --graph <file>` |
|
||||
| [llms](/reference/llms) | LLM providers | `Groq`, `OpenAI`, `create_provider` |
|
||||
| [mcp_server](/reference/mcp_server) | MCP stdio server | `python -m semantica.mcp_server` |
|
||||
| [seed](/reference/seed) | KG bootstrapping from structured sources | `SeedManager` |
|
||||
| [evals](/reference/evals) | Quality evaluation | `KGEvaluator`, `ExtractionEvaluator`, `PipelineEvaluator`, `RegressionTracker` |
|
||||
| [seed](/reference/seed) | KG bootstrapping from structured sources | `SeedDataManager` |
|
||||
| [evals](/reference/evals) | Decision-intelligence evaluation | `evaluate`, `list_evaluators`, `EvalSummary` |
|
||||
| [core](/reference/core) | Base classes & registry | `Semantica`, `ConfigManager`, `PluginRegistry`, `LifecycleManager` |
|
||||
| [utils](/reference/utils) | Shared utilities | `helpers`, `validators` |
|
||||
|
||||
|
||||
+21
-21
@@ -30,28 +30,28 @@ icon: "brain"
|
||||
|
||||
## What You Get
|
||||
|
||||
- **AgentContext** — Memory, decision tracking, and graph-backed retrieval behind one API
|
||||
- **AgentContext**: memory, decision tracking, and graph-backed retrieval behind one API
|
||||
- Conversation history and checkpoint diffing
|
||||
- Persist and restore full context state to disk
|
||||
- **ContextGraph** — Thread-safe in-memory knowledge graph
|
||||
- **ContextGraph**: thread-safe in-memory knowledge graph
|
||||
- PageRank, centrality, community detection, temporal validity
|
||||
- Cross-graph navigation and link traversal
|
||||
- **AgentMemory** — Embedding-backed memory with retention policy
|
||||
- **AgentMemory**: embedding-backed memory with retention policy
|
||||
- LRU eviction at configurable `max_memory_size`
|
||||
- Per-conversation history isolation
|
||||
- **DecisionRecorder** — Records decisions with causal chains and confidence scores
|
||||
- **DecisionRecorder**: records decisions with causal chains and confidence scores
|
||||
- Temporal validity windows (`valid_from` / `valid_until`)
|
||||
- Cross-system context capture on every decision
|
||||
- **PolicyEngine** — Versioned policy storage in the knowledge graph
|
||||
- **PolicyEngine**: versioned policy storage in the knowledge graph
|
||||
- Compliance checking against recorded decisions
|
||||
- Policy exception tracking with approver audit trail
|
||||
- **EntityLinker** — Maps entity text to stable URIs
|
||||
- **EntityLinker**: maps entity text to stable URIs
|
||||
- Creates typed links between entity IDs
|
||||
- Prevents "Apple", "Apple Inc.", "AAPL" becoming separate nodes
|
||||
- **ContextRetriever** — Fuses vector similarity, graph traversal, and agent memory
|
||||
- **ContextRetriever**: fuses vector similarity, graph traversal, and agent memory
|
||||
- Richer context than pure vector search
|
||||
- Configurable `hybrid_alpha` and expansion hops
|
||||
- **CausalChainAnalyzer** — Traces upstream causes and downstream effects of any decision
|
||||
- **CausalChainAnalyzer**: traces upstream causes and downstream effects of any decision
|
||||
- Explainability paths with relationship types
|
||||
- Configurable depth and direction
|
||||
|
||||
@@ -273,7 +273,7 @@ icon: "brain"
|
||||
</Tip>
|
||||
|
||||
<Tip>
|
||||
**Persist your context between runs.** `VectorStore` does not auto-persist — passing `index_path=` to its constructor is a no-op. Call `context.save("agent_state/")` to write memory, the vector index, and the graph to disk, and `context.load("agent_state/")` on the next process to restore them. See the "Persist & Restore" tab under [Real-World Patterns](#real-world-patterns) below.
|
||||
**Persist your context between runs.** `VectorStore` does not auto-persist; passing `index_path=` to its constructor is a no-op. Call `context.save("agent_state/")` to write memory, the vector index, and the graph to disk, and `context.load("agent_state/")` on the next process to restore them. See the "Persist & Restore" tab under [Real-World Patterns](#real-world-patterns) below.
|
||||
</Tip>
|
||||
|
||||
### Memory Methods
|
||||
@@ -449,7 +449,7 @@ print("Nodes: {}, Edges: {}".format(stats["node_count"], stats["edge_count"]))
|
||||
`ContextGraph` exposes a full Distance Intelligence API for exploring semantic neighborhoods and blending proximity into retrieval.
|
||||
|
||||
<Info>
|
||||
Full Distance Intelligence reference — distance matrices, API endpoints, embedding cache, Explorer UI — is covered in the dedicated [Distance Intelligence](/reference/distance) page. This section documents the context-layer API.
|
||||
Full Distance Intelligence reference (distance matrices, API endpoints, embedding cache, Explorer UI) is covered in the dedicated [Distance Intelligence](/reference/distance) page. This section documents the context-layer API.
|
||||
</Info>
|
||||
|
||||
### Neighbors with Distance Metadata
|
||||
@@ -480,7 +480,7 @@ for n in neighbors:
|
||||
| Added field | Type | Description |
|
||||
| :---------- | :---- | :----------- |
|
||||
| `distance_band` | `str` | `"direct"` (1 hop) / `"near"` (2) / `"mid-range"` (3–4) / `"distant"` (5+) |
|
||||
| `confidence_decay` | `float` | `edge_weight ^ hop_count` — decays with each hop |
|
||||
| `confidence_decay` | `float` | `edge_weight ^ hop_count`; decays with each hop |
|
||||
| `path_to_anchor` | `List[str]` | Shortest path from anchor node to this neighbor |
|
||||
| `hop_count` | `int` | BFS depth from anchor |
|
||||
|
||||
@@ -659,7 +659,7 @@ if not receipt.complete:
|
||||
```
|
||||
|
||||
<Warning>
|
||||
Check the receipt — the call returning is not proof the data is gone. FAISS,
|
||||
Check the receipt. The call returning is not proof the data is gone. FAISS,
|
||||
Milvus, and Weaviate expose no delete method, so erasure cannot be completed on
|
||||
those backends today; the receipt reports `unsupported` rather than a success it
|
||||
did not achieve.
|
||||
@@ -687,9 +687,9 @@ At least one store is required; a store that is not supplied reports
|
||||
|
||||
| Status | Meaning |
|
||||
| :--- | :--- |
|
||||
| `erased` | Reached, data removed. On the vectors leg this means the store accepted the delete for the ids given — backends offer no portable existence check, so it is not a count of embeddings that were really there |
|
||||
| `erased` | Reached, data removed. On the vectors leg this means the store accepted the delete for the ids given; backends offer no portable existence check, so it is not a count of embeddings that were really there |
|
||||
| `not_found` | Reached, held nothing for this entity |
|
||||
| `not_configured` | No such store was bound — normal, not a failure |
|
||||
| `not_configured` | No such store was bound: normal, not a failure |
|
||||
| `unsupported` | The store cannot delete at all; retrying will not help |
|
||||
| `failed` | The store was reached and the deletion did not succeed |
|
||||
|
||||
@@ -721,7 +721,7 @@ receipt.to_dict()
|
||||
# }
|
||||
```
|
||||
|
||||
Erasure runs outward-in — vectors, then memory, then the graph. The tombstone is
|
||||
Erasure runs outward-in: vectors, then memory, then the graph. The tombstone is
|
||||
the durable attestation that an erasure happened, so it is written last: a crash
|
||||
mid-cascade leaves the node present and the receipt incomplete, rather than a
|
||||
tombstone claiming more than actually happened. A store that raises is recorded
|
||||
@@ -1087,10 +1087,10 @@ class EntityLink:
|
||||
</Tab>
|
||||
</Tabs>
|
||||
|
||||
- [Vector Store](/reference/vector_store) — Embedding storage backend for memory retrieval.
|
||||
- [Knowledge Graph](/reference/kg) — Graph algorithms and analytics used inside ContextGraph.
|
||||
- [Reasoning](reasoning) — Logical inference layered on top of context.
|
||||
- [Provenance](provenance) — W3C PROV-O lineage for every stored fact.
|
||||
- [Vector Store](/reference/vector_store): embedding storage backend for memory retrieval.
|
||||
- [Knowledge Graph](/reference/kg): graph algorithms and analytics used inside ContextGraph.
|
||||
- [Reasoning](/guides/reasoning): logical inference layered on top of context.
|
||||
- [Provenance](/guides/provenance): W3C PROV-O lineage for every stored fact.
|
||||
|
||||
- [Context Module](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/19_Context_Module.ipynb) — Memory and decision tracking · Intermediate
|
||||
- [Advanced Context Engineering](https://github.com/semantica-agi/semantica/blob/main/cookbook/advanced/11_Advanced_Context_Engineering.ipynb) — Production FAISS + Neo4j setup · Advanced
|
||||
- [Context Module](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/19_Context_Module.ipynb): memory and decision tracking · Intermediate
|
||||
- [Advanced Context Engineering](https://github.com/semantica-agi/semantica/blob/main/cookbook/advanced/11_Advanced_Context_Engineering.ipynb): production FAISS + Neo4j setup · Advanced
|
||||
|
||||
@@ -129,7 +129,7 @@ from semantica.llms import Groq, OpenAI, LiteLLM, HuggingFaceLLM
|
||||
from semantica.llms import LiteLLM
|
||||
|
||||
llm = LiteLLM(
|
||||
model="anthropic/claude-sonnet-4-20250514",
|
||||
model="anthropic/claude-sonnet-5",
|
||||
api_key=os.getenv("ANTHROPIC_API_KEY"),
|
||||
temperature=0.0,
|
||||
)
|
||||
@@ -198,7 +198,7 @@ llm = Groq(api_key=os.getenv("GROQ_API_KEY"), model="llama-3.1-8b-instant")
|
||||
# Method 3: Multiple providers via LiteLLM
|
||||
providers = {
|
||||
"fast": LiteLLM(model="groq/llama-3.1-8b-instant", api_key=os.getenv("GROQ_API_KEY")),
|
||||
"smart": LiteLLM(model="anthropic/claude-sonnet-4-20250514", api_key=os.getenv("ANTHROPIC_API_KEY"))
|
||||
"smart": LiteLLM(model="anthropic/claude-sonnet-5", api_key=os.getenv("ANTHROPIC_API_KEY"))
|
||||
}
|
||||
```
|
||||
|
||||
@@ -252,7 +252,7 @@ from semantica.llms import LiteLLM
|
||||
# pip install "semantica[llm-litellm]"
|
||||
|
||||
# Anthropic Claude
|
||||
llm = LiteLLM(model="anthropic/claude-opus-4-5", api_key=os.getenv("ANTHROPIC_API_KEY"))
|
||||
llm = LiteLLM(model="anthropic/claude-opus-4-7", api_key=os.getenv("ANTHROPIC_API_KEY"))
|
||||
|
||||
# Google Gemini
|
||||
llm = LiteLLM(model="gemini/gemini-1.5-pro", api_key=os.getenv("GOOGLE_API_KEY"))
|
||||
@@ -267,7 +267,7 @@ llm = LiteLLM(model="deepseek/deepseek-chat", api_key=os.getenv("DEEP
|
||||
llm = LiteLLM(model="azure/gpt-4o", api_key=os.getenv("AZURE_API_KEY"))
|
||||
|
||||
# AWS Bedrock
|
||||
llm = LiteLLM(model="bedrock/anthropic.claude-3-5-sonnet-20241022-v2:0")
|
||||
llm = LiteLLM(model="bedrock/anthropic.claude-sonnet-4-5-20250929-v1:0")
|
||||
|
||||
# Novita AI
|
||||
llm = LiteLLM(model="novita/deepseek/deepseek-v3.2", api_key=os.getenv("NOVITA_API_KEY"))
|
||||
@@ -297,12 +297,12 @@ from semantica.llms import LiteLLM
|
||||
|
||||
# Pattern: LiteLLM(model="<provider>/<model-name>")
|
||||
providers = {
|
||||
"Anthropic": LiteLLM(model="anthropic/claude-opus-4-5", api_key=os.getenv("ANTHROPIC_API_KEY")),
|
||||
"Anthropic": LiteLLM(model="anthropic/claude-opus-4-7", api_key=os.getenv("ANTHROPIC_API_KEY")),
|
||||
"Gemini": LiteLLM(model="gemini/gemini-1.5-pro", api_key=os.getenv("GOOGLE_API_KEY")),
|
||||
"Ollama": LiteLLM(model="ollama/llama3.2:3b", api_base="http://localhost:11434"),
|
||||
"DeepSeek": LiteLLM(model="deepseek/deepseek-chat", api_key=os.getenv("DEEPSEEK_API_KEY")),
|
||||
"Azure": LiteLLM(model="azure/gpt-4o", api_key=os.getenv("AZURE_API_KEY")),
|
||||
"Bedrock": LiteLLM(model="bedrock/anthropic.claude-3-5-sonnet-20241022-v2:0"),
|
||||
"Bedrock": LiteLLM(model="bedrock/anthropic.claude-sonnet-4-5-20250929-v1:0"),
|
||||
"Cohere": LiteLLM(model="cohere/command-r-plus", api_key=os.getenv("COHERE_API_KEY")),
|
||||
"Novita AI": LiteLLM(model="novita/deepseek/deepseek-v3.2", api_key=os.getenv("NOVITA_API_KEY")),
|
||||
}
|
||||
@@ -416,7 +416,7 @@ for text in texts:
|
||||
| :---------- | :--------------------------- | :----------- |
|
||||
| **Entity Extraction** | `Groq("llama-3.3-70b-versatile")` | Fast, good accuracy for structured tasks |
|
||||
| **Relation Extraction** | `OpenAI("gpt-4o")` | Best at complex relationship reasoning |
|
||||
| **Complex Analysis** | `LiteLLM("anthropic/claude-sonnet-4-20250514")` | Highest reasoning capability |
|
||||
| **Complex Analysis** | `LiteLLM("anthropic/claude-sonnet-5")` | Highest reasoning capability |
|
||||
| **High Volume/Cost** | `LiteLLM("deepseek/deepseek-chat")` | Lowest cost per token |
|
||||
|
||||
### Error Handling
|
||||
|
||||
@@ -323,7 +323,7 @@ all_facts = datalog.derive_all()
|
||||
|
||||
# Query with variable pattern: variables start with uppercase or ?
|
||||
results = datalog.query("ancestor(alice, ?Z)")
|
||||
# → [{"Z": "bob"}, {"Z": "charlie"}, {"Z": "dave"}]
|
||||
# → a list of binding dicts: [{"Z": "bob"}, {"Z": "charlie"}, {"Z": "dave"}] (order not guaranteed)
|
||||
|
||||
# Clear and start over
|
||||
datalog.clear()
|
||||
|
||||
@@ -9,7 +9,7 @@
|
||||
"lint": "eslint .",
|
||||
"preview": "vite preview",
|
||||
"test:graph-store": "node --test tests/graphStore.multi-edge.test.mjs",
|
||||
"test:graph-workspace": "node --import tsx --test tests/markdownContentViewer.test.ts tests/graphSceneState.display.test.ts tests/temporalLifecycle.test.ts tests/deterministicExplorerRendering.test.ts tests/smallGraphLayout.test.ts tests/realtimeGraphAttributes.test.ts",
|
||||
"test:graph-workspace": "node --import tsx --test tests/markdownContentViewer.test.ts tests/graphSceneState.display.test.ts tests/temporalLifecycle.test.ts tests/deterministicExplorerRendering.test.ts tests/smallGraphLayout.test.ts tests/realtimeGraphAttributes.test.ts tests/ontologyEditorModel.test.ts",
|
||||
"test:deterministic-e2e": "node --import tsx --test tests/deterministicExplorerRendering.e2e.ts",
|
||||
"test:plugin-registry": "node --import tsx --test tests/pluginRegistry.temporal.test.mjs"
|
||||
},
|
||||
|
||||
+13
-1
@@ -93,6 +93,18 @@ const navItems: NavItem[] = [
|
||||
{ id: 'ontology-hub', label: 'Ontology Hub', hint: 'Schema governance, registry, and vocabulary management', icon: GitMerge },
|
||||
];
|
||||
|
||||
function readInitialWorkspace(): WorkspaceId {
|
||||
try {
|
||||
const params = new URLSearchParams(window.location.search);
|
||||
if (params.has("ontologyTab") || params.has("ontologyEntity")) {
|
||||
return "ontology-hub";
|
||||
}
|
||||
} catch {
|
||||
// Default to the welcome screen when URL state is unavailable.
|
||||
}
|
||||
return "welcome";
|
||||
}
|
||||
|
||||
const shellStyles = `
|
||||
:root {
|
||||
--app-bg: #07111f;
|
||||
@@ -1773,7 +1785,7 @@ function WelcomeScreen({
|
||||
}
|
||||
|
||||
export default function App() {
|
||||
const [activeWorkspace, setActiveWorkspace] = useState<WorkspaceId>('welcome');
|
||||
const [activeWorkspace, setActiveWorkspace] = useState<WorkspaceId>(readInitialWorkspace);
|
||||
const [exploreView, setExploreView] = useState<ExploreView>('graph');
|
||||
const [analyzeView, setAnalyzeView] = useState<AnalyzeView>('reasoning');
|
||||
const [enrichView, setEnrichView] = useState<EnrichView>('import');
|
||||
|
||||
@@ -8,8 +8,10 @@ import {
|
||||
useNodesState,
|
||||
useEdgesState,
|
||||
MarkerType,
|
||||
Handle,
|
||||
Position,
|
||||
} from "@xyflow/react";
|
||||
import type { Connection, Edge, Node } from "@xyflow/react";
|
||||
import type { Connection, Edge, Node, ReactFlowInstance } from "@xyflow/react";
|
||||
import "@xyflow/react/dist/style.css";
|
||||
import {
|
||||
Plus,
|
||||
@@ -22,10 +24,20 @@ import {
|
||||
Pencil,
|
||||
Trash2,
|
||||
} from "lucide-react";
|
||||
import { loadOntologyEntityOwner, loadOntologyGraph } from "./api";
|
||||
import type { OntologyGraphEdge, OntologyGraphNode } from "./api";
|
||||
import {
|
||||
classifyNodeType,
|
||||
inferOntologyUri,
|
||||
isEditableEntityType,
|
||||
ONTOLOGY_MINIMAP_THEME,
|
||||
} from "./ontologyEditorModel";
|
||||
import type { EditorEntityType, RegistryEntry } from "./ontologyEditorModel";
|
||||
|
||||
type OntologyNodeData = {
|
||||
label?: string;
|
||||
type?: string;
|
||||
entityType?: EditorEntityType;
|
||||
};
|
||||
|
||||
type OntologyNode = Node<OntologyNodeData>;
|
||||
@@ -34,12 +46,57 @@ type OntologyEdge = Edge<Record<string, unknown>>;
|
||||
const nodeTypes = {
|
||||
classNode: ({ data }: { data: OntologyNodeData }) => (
|
||||
<div style={classNodeStyle}>
|
||||
<Handle type="target" position={Position.Left} style={handleStyle} />
|
||||
<div style={classNodeHeader}>{data.label}</div>
|
||||
<div style={classNodeSub}>{data.type}</div>
|
||||
<Handle type="source" position={Position.Right} style={handleStyle} />
|
||||
</div>
|
||||
),
|
||||
};
|
||||
|
||||
const handleStyle: React.CSSProperties = {
|
||||
width: 8,
|
||||
height: 8,
|
||||
border: "1px solid rgba(235, 243, 255, 0.8)",
|
||||
background: "#4aa3ff",
|
||||
};
|
||||
|
||||
const ontologyFlowThemeCss = `
|
||||
.ontology-editor-flow .react-flow__controls {
|
||||
overflow: hidden;
|
||||
border: 1px solid rgba(127, 208, 255, 0.2);
|
||||
border-radius: 9px;
|
||||
background: rgba(6, 13, 26, 0.96);
|
||||
box-shadow: 0 8px 24px rgba(0, 0, 0, 0.38);
|
||||
}
|
||||
|
||||
.ontology-editor-flow .react-flow__controls-button {
|
||||
width: 30px;
|
||||
height: 30px;
|
||||
background: transparent;
|
||||
border-bottom-color: rgba(127, 208, 255, 0.14);
|
||||
color: #8fa8c6;
|
||||
transition: color 140ms ease, background 140ms ease;
|
||||
}
|
||||
|
||||
.ontology-editor-flow .react-flow__controls-button:hover {
|
||||
background: rgba(74, 163, 255, 0.14);
|
||||
color: #ebf3ff;
|
||||
}
|
||||
|
||||
.ontology-editor-flow .react-flow__controls-button:focus-visible {
|
||||
position: relative;
|
||||
z-index: 1;
|
||||
outline: 2px solid #7fd0ff;
|
||||
outline-offset: -2px;
|
||||
}
|
||||
|
||||
.ontology-editor-flow .react-flow__controls-button:disabled {
|
||||
background: rgba(3, 9, 18, 0.32);
|
||||
color: #40566f;
|
||||
}
|
||||
`;
|
||||
|
||||
const classNodeStyle: React.CSSProperties = {
|
||||
padding: "12px 16px",
|
||||
borderRadius: "8px",
|
||||
@@ -79,17 +136,95 @@ interface DraftDiff {
|
||||
annotation_changes: Record<string, Record<string, any>>;
|
||||
}
|
||||
|
||||
interface RegistryEntry {
|
||||
uri: string;
|
||||
name: string;
|
||||
function requestedEntityUri(): string {
|
||||
try {
|
||||
return new URLSearchParams(window.location.search).get("ontologyEntity") || "";
|
||||
} catch {
|
||||
return "";
|
||||
}
|
||||
}
|
||||
|
||||
function nodeLabel(node: OntologyGraphNode): string {
|
||||
const explicit = String(node.content || node.properties?.["rdfs:label"] || "").trim();
|
||||
if (explicit && explicit !== node.id) {
|
||||
return explicit;
|
||||
}
|
||||
const trimmed = node.id.replace(/[/#]+$/, "");
|
||||
return trimmed.split("#").pop() || trimmed.split("/").pop() || node.id;
|
||||
}
|
||||
|
||||
function classifyEditorNode(node: OntologyGraphNode): OntologyNodeData["entityType"] {
|
||||
return classifyNodeType(node.type);
|
||||
}
|
||||
|
||||
function layoutEditorNodes(inputNodes: OntologyNode[]): OntologyNode[] {
|
||||
const properties = inputNodes.filter((node) => node.data.entityType === "property");
|
||||
const targets = inputNodes.filter((node) => (
|
||||
node.data.entityType === "class" || node.data.entityType === "external"
|
||||
));
|
||||
const context = inputNodes.filter((node) => (
|
||||
node.data.entityType !== "property"
|
||||
&& node.data.entityType !== "class"
|
||||
&& node.data.entityType !== "external"
|
||||
));
|
||||
const height = Math.max(360, Math.max(properties.length, targets.length) * 180);
|
||||
const positions = new Map<string, { x: number; y: number }>();
|
||||
|
||||
properties.forEach((node, index) => {
|
||||
positions.set(node.id, { x: 0, y: ((index + 1) * height) / (properties.length + 1) });
|
||||
});
|
||||
targets.forEach((node, index) => {
|
||||
positions.set(node.id, { x: 600, y: ((index + 1) * height) / (targets.length + 1) });
|
||||
});
|
||||
context.forEach((node, index) => {
|
||||
positions.set(node.id, { x: 300 + index * 220, y: height + 120 });
|
||||
});
|
||||
|
||||
return inputNodes.map((node) => ({
|
||||
...node,
|
||||
position: positions.get(node.id) || node.position,
|
||||
}));
|
||||
}
|
||||
|
||||
function buildEditorElements(apiNodes: OntologyGraphNode[], apiEdges: OntologyGraphEdge[]) {
|
||||
const sortedNodes = [...apiNodes].sort((left, right) => {
|
||||
const typeDelta = left.type.localeCompare(right.type);
|
||||
return typeDelta || left.id.localeCompare(right.id);
|
||||
});
|
||||
const nodes = layoutEditorNodes(sortedNodes.map((node) => ({
|
||||
id: node.id,
|
||||
type: "classNode",
|
||||
position: { x: 0, y: 0 },
|
||||
data: {
|
||||
label: nodeLabel(node),
|
||||
type: node.type,
|
||||
entityType: classifyEditorNode(node),
|
||||
},
|
||||
})));
|
||||
const edges: OntologyEdge[] = apiEdges.map((edge, index) => ({
|
||||
id: edge.id || `${edge.source}:${edge.type}:${edge.target}:${index}`,
|
||||
source: edge.source,
|
||||
target: edge.target,
|
||||
label: edge.type,
|
||||
type: "default",
|
||||
markerEnd: { type: MarkerType.ArrowClosed },
|
||||
style: { stroke: "rgba(127, 208, 255, 0.72)", strokeWidth: 1.5 },
|
||||
labelStyle: { fill: "#c8dcf5", fontSize: 11, fontWeight: 600 },
|
||||
labelBgStyle: { fill: "#07111f", fillOpacity: 0.9 },
|
||||
}));
|
||||
return { nodes, edges };
|
||||
}
|
||||
|
||||
export function OntologyEditor() {
|
||||
const [nodes, setNodes, onNodesChange] = useNodesState<OntologyNode>([]);
|
||||
const [edges, setEdges, onEdgesChange] = useEdgesState<OntologyEdge>([]);
|
||||
const [selectedElement, setSelectedElement] = useState<OntologyNode | OntologyEdge | null>(null);
|
||||
const hasDetailPanel = selectedElement !== null;
|
||||
const [registry, setRegistry] = useState<RegistryEntry[]>([]);
|
||||
const [ontologyUri, setOntologyUri] = useState<string>("");
|
||||
const [flowInstance, setFlowInstance] = useState<ReactFlowInstance<OntologyNode, OntologyEdge> | null>(null);
|
||||
const [isLoadingGraph, setIsLoadingGraph] = useState(false);
|
||||
const [graphError, setGraphError] = useState("");
|
||||
const [draftDiff, setDraftDiff] = useState<DraftDiff>({
|
||||
added_classes: [],
|
||||
removed_classes: [],
|
||||
@@ -108,12 +243,18 @@ export function OntologyEditor() {
|
||||
|
||||
useEffect(() => {
|
||||
let cancelled = false;
|
||||
fetch("/api/ontology/registry")
|
||||
.then((response) => (response.ok ? response.json() : []))
|
||||
.then((entries: RegistryEntry[]) => {
|
||||
const requested = requestedEntityUri();
|
||||
Promise.all([
|
||||
fetch("/api/ontology/registry").then((response) => (response.ok ? response.json() : [])),
|
||||
requested
|
||||
? loadOntologyEntityOwner(requested).catch(() => undefined)
|
||||
: Promise.resolve(undefined),
|
||||
])
|
||||
.then(([entries, explicitOwner]: [RegistryEntry[], string | undefined]) => {
|
||||
if (cancelled) return;
|
||||
setRegistry(entries);
|
||||
setOntologyUri((current) => current || entries[0]?.uri || "");
|
||||
const inferredOntology = inferOntologyUri(entries, requested, explicitOwner);
|
||||
setOntologyUri((current) => current || inferredOntology || entries[0]?.uri || "");
|
||||
})
|
||||
.catch((error) => {
|
||||
console.error("Failed to load ontology registry:", error);
|
||||
@@ -123,6 +264,47 @@ export function OntologyEditor() {
|
||||
};
|
||||
}, []);
|
||||
|
||||
useEffect(() => {
|
||||
if (!ontologyUri) {
|
||||
setNodes([]);
|
||||
setEdges([]);
|
||||
setSelectedElement(null);
|
||||
return;
|
||||
}
|
||||
|
||||
const controller = new AbortController();
|
||||
setIsLoadingGraph(true);
|
||||
setGraphError("");
|
||||
loadOntologyGraph(ontologyUri, controller.signal)
|
||||
.then((payload) => {
|
||||
const elements = buildEditorElements(payload.nodes, payload.edges);
|
||||
setNodes(elements.nodes);
|
||||
setEdges(elements.edges);
|
||||
const requested = requestedEntityUri();
|
||||
setSelectedElement(elements.nodes.find((node) => node.id === requested) || null);
|
||||
})
|
||||
.catch((error) => {
|
||||
if (controller.signal.aborted) return;
|
||||
setNodes([]);
|
||||
setEdges([]);
|
||||
setSelectedElement(null);
|
||||
setGraphError(error instanceof Error ? error.message : "Failed to load ontology graph");
|
||||
})
|
||||
.finally(() => {
|
||||
if (!controller.signal.aborted) setIsLoadingGraph(false);
|
||||
});
|
||||
|
||||
return () => controller.abort();
|
||||
}, [ontologyUri, setEdges, setNodes]);
|
||||
|
||||
useEffect(() => {
|
||||
if (!flowInstance || nodes.length === 0) return;
|
||||
const frame = window.requestAnimationFrame(() => {
|
||||
void flowInstance.fitView({ padding: 0.22, duration: 320, maxZoom: 1.25 });
|
||||
});
|
||||
return () => window.cancelAnimationFrame(frame);
|
||||
}, [flowInstance, hasDetailPanel, nodes.length, ontologyUri]);
|
||||
|
||||
const onConnect = useCallback(
|
||||
(params: Connection) => setEdges((eds) => addEdge({ ...params, markerEnd: { type: MarkerType.ArrowClosed } }, eds)),
|
||||
[setEdges]
|
||||
@@ -134,7 +316,7 @@ export function OntologyEditor() {
|
||||
id: newId,
|
||||
type: "classNode",
|
||||
position: { x: Math.random() * 400, y: Math.random() * 300 },
|
||||
data: { label: "NewClass", type: "owl:Class" },
|
||||
data: { label: "NewClass", type: "owl:Class", entityType: "class" },
|
||||
};
|
||||
setNodes((nds) => [...nds, newNode]);
|
||||
setDraftDiff((prev) => ({
|
||||
@@ -170,7 +352,7 @@ export function OntologyEditor() {
|
||||
id: newId,
|
||||
type: "classNode",
|
||||
position: { x: Math.random() * 400, y: Math.random() * 300 },
|
||||
data: { label: "NewIndividual", type: "owl:NamedIndividual" },
|
||||
data: { label: "NewIndividual", type: "owl:NamedIndividual", entityType: "external" },
|
||||
};
|
||||
setNodes((nds) => [...nds, newNode]);
|
||||
}, [setNodes]);
|
||||
@@ -190,13 +372,21 @@ export function OntologyEditor() {
|
||||
}, []);
|
||||
|
||||
const autoLayout = useCallback(() => {
|
||||
const layoutNodes = nodes.map((node, index) => ({
|
||||
...node,
|
||||
position: { x: (index % 4) * 200, y: Math.floor(index / 4) * 150 },
|
||||
}));
|
||||
setNodes(layoutNodes);
|
||||
setNodes(layoutEditorNodes(nodes));
|
||||
}, [nodes, setNodes]);
|
||||
|
||||
const selectNode = useCallback((node: OntologyNode) => {
|
||||
setSelectedElement(node);
|
||||
try {
|
||||
const params = new URLSearchParams(window.location.search);
|
||||
params.set("ontologyTab", "editor");
|
||||
params.set("ontologyEntity", node.id);
|
||||
window.history.replaceState(null, "", `?${params.toString()}`);
|
||||
} catch {
|
||||
// URL state is optional; the editor selection still works without it.
|
||||
}
|
||||
}, []);
|
||||
|
||||
const saveDraft = useCallback(async () => {
|
||||
if (!ontologyUri) {
|
||||
alert("Please select an ontology first");
|
||||
@@ -247,12 +437,11 @@ export function OntologyEditor() {
|
||||
...prev,
|
||||
removed_properties: [...prev.removed_properties, target.id],
|
||||
}));
|
||||
} else {
|
||||
} else if (isEditableEntityType(target.data.entityType)) {
|
||||
setNodes((nds) => nds.filter((n) => n.id !== target.id));
|
||||
setDraftDiff((prev) => ({
|
||||
...prev,
|
||||
removed_classes: [...prev.removed_classes, target.id],
|
||||
}));
|
||||
setDraftDiff((prev) => target.data.entityType === "property"
|
||||
? { ...prev, removed_properties: [...prev.removed_properties, target.id] }
|
||||
: { ...prev, removed_classes: [...prev.removed_classes, target.id] });
|
||||
}
|
||||
setSelectedElement(null);
|
||||
}
|
||||
@@ -261,16 +450,21 @@ export function OntologyEditor() {
|
||||
|
||||
const renameSelected = useCallback(() => {
|
||||
const target = showContext?.element ?? selectedElement;
|
||||
if (target && !("source" in target)) {
|
||||
if (target && !("source" in target) && isEditableEntityType(target.data.entityType)) {
|
||||
const newLabel = prompt("Enter new name:", String(target.data.label ?? ""));
|
||||
if (newLabel) {
|
||||
setNodes((nds) =>
|
||||
nds.map((n) => (n.id === target.id ? { ...n, data: { ...n.data, label: newLabel } } : n))
|
||||
);
|
||||
setDraftDiff((prev) => ({
|
||||
...prev,
|
||||
modified_classes: { ...prev.modified_classes, [target.id]: { label: newLabel } },
|
||||
}));
|
||||
setDraftDiff((prev) => target.data.entityType === "property"
|
||||
? {
|
||||
...prev,
|
||||
modified_properties: { ...prev.modified_properties, [target.id]: { label: newLabel } },
|
||||
}
|
||||
: {
|
||||
...prev,
|
||||
modified_classes: { ...prev.modified_classes, [target.id]: { label: newLabel } },
|
||||
});
|
||||
}
|
||||
}
|
||||
setShowContext(null);
|
||||
@@ -339,11 +533,10 @@ export function OntologyEditor() {
|
||||
};
|
||||
|
||||
const detailPanelStyle: React.CSSProperties = {
|
||||
position: "absolute",
|
||||
right: 0,
|
||||
top: 0,
|
||||
bottom: 0,
|
||||
flex: "0 0 320px",
|
||||
width: "320px",
|
||||
minWidth: "320px",
|
||||
boxSizing: "border-box",
|
||||
background: "rgba(9, 19, 34, 0.95)",
|
||||
borderLeft: "1px solid rgba(140, 192, 255, 0.12)",
|
||||
padding: "20px",
|
||||
@@ -353,11 +546,24 @@ export function OntologyEditor() {
|
||||
|
||||
return (
|
||||
<div style={{ display: "flex", flexDirection: "column", height: "100%", background: "#07111f" }}>
|
||||
<style>{ontologyFlowThemeCss}</style>
|
||||
<div style={toolbarStyle}>
|
||||
<select
|
||||
aria-label="Active ontology"
|
||||
value={ontologyUri}
|
||||
onChange={(event) => setOntologyUri(event.target.value)}
|
||||
onChange={(event) => {
|
||||
setOntologyUri(event.target.value);
|
||||
setSelectedElement(null);
|
||||
try {
|
||||
// Drop the previous ontology's entity from the URL, or a reload
|
||||
// would resolve the stale ID and jump back to that ontology.
|
||||
const params = new URLSearchParams(window.location.search);
|
||||
params.delete("ontologyEntity");
|
||||
window.history.replaceState(null, "", `?${params.toString()}`);
|
||||
} catch {
|
||||
// URL state is optional; switching ontologies still works.
|
||||
}
|
||||
}}
|
||||
style={selectStyle}
|
||||
>
|
||||
<option value="">Select ontology...</option>
|
||||
@@ -398,43 +604,75 @@ export function OntologyEditor() {
|
||||
</button>
|
||||
</div>
|
||||
|
||||
<div style={{ flex: 1, position: "relative" }}>
|
||||
<ReactFlow
|
||||
nodes={nodes}
|
||||
edges={edges}
|
||||
onNodesChange={onNodesChange}
|
||||
onEdgesChange={onEdgesChange}
|
||||
onConnect={onConnect}
|
||||
onNodeClick={(_, node) => setSelectedElement(node)}
|
||||
onEdgeClick={(_, edge) => setSelectedElement(edge)}
|
||||
onNodeContextMenu={handleNodeContextMenu}
|
||||
onEdgeContextMenu={handleEdgeContextMenu}
|
||||
nodeTypes={nodeTypes}
|
||||
fitView
|
||||
style={{ background: "#07111f" }}
|
||||
>
|
||||
<Background color="#1a2d3d" gap={20} />
|
||||
<Controls />
|
||||
<MiniMap nodeColor="#4aa3ff" maskColor="rgba(0,0,0,0.6)" />
|
||||
</ReactFlow>
|
||||
<div style={{ display: "flex", flex: 1, minHeight: 0, minWidth: 0 }}>
|
||||
<div style={{ flex: 1, minHeight: 0, minWidth: 0, position: "relative" }}>
|
||||
<ReactFlow
|
||||
className="ontology-editor-flow"
|
||||
nodes={nodes}
|
||||
edges={edges}
|
||||
onNodesChange={onNodesChange}
|
||||
onEdgesChange={onEdgesChange}
|
||||
onConnect={onConnect}
|
||||
onInit={setFlowInstance}
|
||||
onNodeClick={(_, node) => selectNode(node)}
|
||||
onEdgeClick={(_, edge) => setSelectedElement(edge)}
|
||||
onNodeContextMenu={handleNodeContextMenu}
|
||||
onEdgeContextMenu={handleEdgeContextMenu}
|
||||
nodeTypes={nodeTypes}
|
||||
fitView
|
||||
style={{ background: "#07111f" }}
|
||||
>
|
||||
<Background color="#1a2d3d" gap={20} />
|
||||
<Controls />
|
||||
<MiniMap {...ONTOLOGY_MINIMAP_THEME} />
|
||||
</ReactFlow>
|
||||
|
||||
{showContext && (
|
||||
<div style={{ ...contextMenuStyle, left: showContext.x, top: showContext.y }}>
|
||||
<div style={contextItemStyle} onClick={renameSelected}>
|
||||
<Pencil size={14} />
|
||||
Rename
|
||||
{isLoadingGraph && (
|
||||
<div style={canvasMessageStyle}>Loading ontology structure…</div>
|
||||
)}
|
||||
{!isLoadingGraph && graphError && (
|
||||
<div style={{ ...canvasMessageStyle, color: "#ff9a8d" }}>{graphError}</div>
|
||||
)}
|
||||
{!isLoadingGraph && !graphError && ontologyUri && nodes.length === 0 && (
|
||||
<div style={canvasMessageStyle}>This ontology has no editable classes or properties.</div>
|
||||
)}
|
||||
|
||||
{showContext && (
|
||||
<div style={{ ...contextMenuStyle, left: showContext.x, top: showContext.y }}>
|
||||
{"source" in showContext.element || isEditableEntityType(showContext.element.data.entityType) ? (
|
||||
<>
|
||||
{!("source" in showContext.element) && (
|
||||
<div style={contextItemStyle} onClick={renameSelected}>
|
||||
<Pencil size={14} />
|
||||
Rename
|
||||
</div>
|
||||
)}
|
||||
<div style={contextItemStyle} onClick={deleteSelected}>
|
||||
<Trash2 size={14} />
|
||||
Delete
|
||||
</div>
|
||||
</>
|
||||
) : (
|
||||
<div style={{ ...contextItemStyle, cursor: "default", color: "#8fa8c6" }}>
|
||||
This term is read-only
|
||||
</div>
|
||||
)}
|
||||
</div>
|
||||
<div style={contextItemStyle} onClick={deleteSelected}>
|
||||
<Trash2 size={14} />
|
||||
Delete
|
||||
</div>
|
||||
</div>
|
||||
)}
|
||||
)}
|
||||
</div>
|
||||
|
||||
{selectedElement && (
|
||||
<div style={detailPanelStyle}>
|
||||
<h3 style={{ margin: "0 0 16px", color: "#ebf3ff", fontSize: "16px" }}>
|
||||
{"source" in selectedElement ? "Property Details" : "Class Details"}
|
||||
{"source" in selectedElement
|
||||
? "Relationship Details"
|
||||
: selectedElement.data.entityType === "property"
|
||||
? "Property Details"
|
||||
: selectedElement.data.entityType === "ontology"
|
||||
? "Ontology Details"
|
||||
: selectedElement.data.entityType === "external"
|
||||
? "External Term Details"
|
||||
: "Class Details"}
|
||||
</h3>
|
||||
<div style={{ marginBottom: "12px" }}>
|
||||
<label style={{ display: "block", color: "#8fa8c6", fontSize: "12px", marginBottom: "4px" }}>
|
||||
@@ -453,7 +691,9 @@ export function OntologyEditor() {
|
||||
<input
|
||||
type="text"
|
||||
value={String(selectedElement.data.label ?? "")}
|
||||
readOnly={!isEditableEntityType(selectedElement.data.entityType)}
|
||||
onChange={(e) => {
|
||||
if (!isEditableEntityType(selectedElement.data.entityType)) return;
|
||||
setNodes((nds) =>
|
||||
nds.map((n) =>
|
||||
n.id === selectedElement.id
|
||||
@@ -463,10 +703,19 @@ export function OntologyEditor() {
|
||||
);
|
||||
setDraftDiff((prev) => ({
|
||||
...prev,
|
||||
modified_classes: {
|
||||
...prev.modified_classes,
|
||||
[selectedElement.id]: { label: e.target.value },
|
||||
},
|
||||
...(selectedElement.data.entityType === "property"
|
||||
? {
|
||||
modified_properties: {
|
||||
...prev.modified_properties,
|
||||
[selectedElement.id]: { label: e.target.value },
|
||||
},
|
||||
}
|
||||
: {
|
||||
modified_classes: {
|
||||
...prev.modified_classes,
|
||||
[selectedElement.id]: { label: e.target.value },
|
||||
},
|
||||
}),
|
||||
}));
|
||||
}}
|
||||
style={{
|
||||
@@ -496,3 +745,17 @@ export function OntologyEditor() {
|
||||
</div>
|
||||
);
|
||||
}
|
||||
|
||||
const canvasMessageStyle: React.CSSProperties = {
|
||||
position: "absolute",
|
||||
left: "50%",
|
||||
top: "50%",
|
||||
transform: "translate(-50%, -50%)",
|
||||
padding: "10px 14px",
|
||||
borderRadius: "8px",
|
||||
border: "1px solid rgba(127, 208, 255, 0.18)",
|
||||
background: "rgba(3, 9, 18, 0.9)",
|
||||
color: "#8fa8c6",
|
||||
fontSize: "13px",
|
||||
pointerEvents: "none",
|
||||
};
|
||||
|
||||
@@ -9,6 +9,32 @@ import type {
|
||||
ShaclValidationResponse,
|
||||
} from "./types";
|
||||
|
||||
export type OntologyGraphNode = {
|
||||
id: string;
|
||||
type: string;
|
||||
content?: string;
|
||||
properties?: Record<string, unknown>;
|
||||
};
|
||||
|
||||
export type OntologyGraphEdge = {
|
||||
id?: string;
|
||||
source: string;
|
||||
target: string;
|
||||
type: string;
|
||||
weight?: number;
|
||||
properties?: Record<string, unknown>;
|
||||
};
|
||||
|
||||
export type OntologyGraphResponse = {
|
||||
uri: string;
|
||||
nodes: OntologyGraphNode[];
|
||||
edges: OntologyGraphEdge[];
|
||||
};
|
||||
|
||||
export type OntologyEntityOwner = {
|
||||
source_ontology?: string;
|
||||
};
|
||||
|
||||
async function parseResponse<T>(response: Response): Promise<T> {
|
||||
if (!response.ok) {
|
||||
let detail = `Request failed with status ${response.status}`;
|
||||
@@ -31,6 +57,18 @@ export async function loadOntologyRegistry(): Promise<OntologyEntry[]> {
|
||||
return parseResponse<OntologyEntry[]>(await fetch("/api/ontology/registry"));
|
||||
}
|
||||
|
||||
export async function loadOntologyGraph(uri: string, signal?: AbortSignal): Promise<OntologyGraphResponse> {
|
||||
return parseResponse<OntologyGraphResponse>(
|
||||
await fetch(`/api/ontology/graph?uri=${encodeURIComponent(uri)}`, { signal }),
|
||||
);
|
||||
}
|
||||
|
||||
export async function loadOntologyEntityOwner(uri: string): Promise<string | undefined> {
|
||||
const response = await fetch(`/api/ontology/entity/${encodeURIComponent(uri)}`);
|
||||
if (!response.ok) return undefined;
|
||||
return (await response.json() as OntologyEntityOwner).source_ontology;
|
||||
}
|
||||
|
||||
export async function loadAlignments(uri?: string): Promise<OntologyAlignment[]> {
|
||||
const query = uri ? `?uri=${encodeURIComponent(uri)}` : "";
|
||||
return parseResponse<OntologyAlignment[]>(await fetch(`/api/ontology/alignments${query}`));
|
||||
|
||||
@@ -38,6 +38,7 @@ function readTabParam(): OntologyHubTab {
|
||||
const params = new URLSearchParams(window.location.search);
|
||||
const raw = params.get(TAB_PARAM);
|
||||
if (raw && TABS.some((t) => t.id === raw)) return raw as OntologyHubTab;
|
||||
if (params.get("ontologyEntity")) return "editor";
|
||||
} catch {
|
||||
// ignore
|
||||
}
|
||||
@@ -116,4 +117,3 @@ export function OntologyWorkspace({ onJumpToGraphNode }: OntologyWorkspaceProps)
|
||||
</div>
|
||||
);
|
||||
}
|
||||
|
||||
|
||||
@@ -0,0 +1,70 @@
|
||||
export type EditorEntityType = "ontology" | "class" | "property" | "external";
|
||||
|
||||
export type RegistryEntry = {
|
||||
uri: string;
|
||||
name: string;
|
||||
};
|
||||
|
||||
export const ONTOLOGY_MINIMAP_THEME = {
|
||||
bgColor: "#0b1625",
|
||||
maskColor: "rgba(7, 17, 31, 0.72)",
|
||||
maskStrokeColor: "#5faeff",
|
||||
maskStrokeWidth: 2,
|
||||
nodeColor: "#2d7fd3",
|
||||
nodeStrokeColor: "#9acbff",
|
||||
nodeStrokeWidth: 1,
|
||||
style: {
|
||||
border: "1px solid #29435c",
|
||||
borderRadius: 6,
|
||||
boxShadow: "0 4px 16px rgba(0, 0, 0, 0.32)",
|
||||
},
|
||||
} as const;
|
||||
|
||||
// The backend emits node types in compact (owl:Class) or full IRI
|
||||
// (http://www.w3.org/2002/07/owl#Class) form; classification must accept both.
|
||||
const FULL_IRI_PREFIXES: Array<[string, string]> = [
|
||||
["http://www.w3.org/2002/07/owl#", "owl:"],
|
||||
["http://www.w3.org/2000/01/rdf-schema#", "rdfs:"],
|
||||
["http://www.w3.org/2004/02/skos/core#", "skos:"],
|
||||
];
|
||||
|
||||
export function compactNodeType(type: string): string {
|
||||
for (const [iri, prefix] of FULL_IRI_PREFIXES) {
|
||||
if (type.startsWith(iri)) {
|
||||
return `${prefix}${type.slice(iri.length)}`;
|
||||
}
|
||||
}
|
||||
return type;
|
||||
}
|
||||
|
||||
export function classifyNodeType(rawType: string): EditorEntityType {
|
||||
const type = compactNodeType(rawType);
|
||||
if (type === "owl:Ontology") return "ontology";
|
||||
if (type === "owl:Class" || type === "rdfs:Class") return "class";
|
||||
if (type.includes("Property")) return "property";
|
||||
return "external";
|
||||
}
|
||||
|
||||
function ownsByNamespace(entityUri: string, ontologyUri: string): boolean {
|
||||
const stem = ontologyUri.replace(/[/#]+$/, "");
|
||||
return entityUri === ontologyUri
|
||||
|| entityUri.startsWith(`${stem}#`)
|
||||
|| entityUri.startsWith(`${stem}/`);
|
||||
}
|
||||
|
||||
export function inferOntologyUri(
|
||||
entries: RegistryEntry[],
|
||||
entityUri: string,
|
||||
explicitOwner?: string,
|
||||
): string | undefined {
|
||||
if (explicitOwner && entries.some((entry) => entry.uri === explicitOwner)) {
|
||||
return explicitOwner;
|
||||
}
|
||||
return [...entries]
|
||||
.filter((entry) => ownsByNamespace(entityUri, entry.uri))
|
||||
.sort((left, right) => right.uri.length - left.uri.length)[0]?.uri;
|
||||
}
|
||||
|
||||
export function isEditableEntityType(entityType?: EditorEntityType): boolean {
|
||||
return entityType === "class" || entityType === "property";
|
||||
}
|
||||
@@ -0,0 +1,66 @@
|
||||
import assert from "node:assert/strict";
|
||||
import test from "node:test";
|
||||
|
||||
import {
|
||||
classifyNodeType,
|
||||
compactNodeType,
|
||||
inferOntologyUri,
|
||||
isEditableEntityType,
|
||||
ONTOLOGY_MINIMAP_THEME,
|
||||
} from "../src/workspaces/OntologyWorkspace/ontologyEditorModel";
|
||||
|
||||
const registry = [
|
||||
{ uri: "https://example.test/foo", name: "Foo" },
|
||||
{ uri: "https://example.test/foo/nested", name: "Nested" },
|
||||
];
|
||||
|
||||
test("ontology inference requires a URI delimiter and prefers the closest namespace", () => {
|
||||
assert.equal(inferOntologyUri(registry, "https://example.test/foobar/Class"), undefined);
|
||||
assert.equal(
|
||||
inferOntologyUri(registry, "https://example.test/foo/nested#Class"),
|
||||
"https://example.test/foo/nested",
|
||||
);
|
||||
});
|
||||
|
||||
test("explicit scheme ownership wins when an entity uses another namespace", () => {
|
||||
assert.equal(
|
||||
inferOntologyUri(registry, "https://vocabulary.test/Class", "https://example.test/foo"),
|
||||
"https://example.test/foo",
|
||||
);
|
||||
});
|
||||
|
||||
test("only draft-supported class and property nodes are editable", () => {
|
||||
assert.equal(isEditableEntityType("class"), true);
|
||||
assert.equal(isEditableEntityType("property"), true);
|
||||
assert.equal(isEditableEntityType("ontology"), false);
|
||||
assert.equal(isEditableEntityType("external"), false);
|
||||
});
|
||||
|
||||
test("the ontology minimap has an explicit dark, high-contrast theme", () => {
|
||||
assert.equal(ONTOLOGY_MINIMAP_THEME.bgColor, "#0b1625");
|
||||
assert.equal(ONTOLOGY_MINIMAP_THEME.maskStrokeColor, "#5faeff");
|
||||
assert.equal(ONTOLOGY_MINIMAP_THEME.nodeStrokeColor, "#9acbff");
|
||||
assert.match(ONTOLOGY_MINIMAP_THEME.style.border, /#29435c/);
|
||||
});
|
||||
|
||||
test("node types classify identically in compact and full IRI form", () => {
|
||||
const cases: Array<[string, string, string]> = [
|
||||
["owl:Ontology", "http://www.w3.org/2002/07/owl#Ontology", "ontology"],
|
||||
["owl:Class", "http://www.w3.org/2002/07/owl#Class", "class"],
|
||||
["rdfs:Class", "http://www.w3.org/2000/01/rdf-schema#Class", "class"],
|
||||
["owl:ObjectProperty", "http://www.w3.org/2002/07/owl#ObjectProperty", "property"],
|
||||
["owl:DatatypeProperty", "http://www.w3.org/2002/07/owl#DatatypeProperty", "property"],
|
||||
["owl:AnnotationProperty", "http://www.w3.org/2002/07/owl#AnnotationProperty", "property"],
|
||||
];
|
||||
for (const [compact, fullIri, expected] of cases) {
|
||||
assert.equal(classifyNodeType(compact), expected, compact);
|
||||
assert.equal(classifyNodeType(fullIri), expected, fullIri);
|
||||
}
|
||||
assert.equal(classifyNodeType("owl:NamedIndividual"), "external");
|
||||
assert.equal(classifyNodeType("http://www.w3.org/2004/02/skos/core#Concept"), "external");
|
||||
});
|
||||
|
||||
test("compactNodeType leaves unknown namespaces untouched", () => {
|
||||
assert.equal(compactNodeType("https://example.org/custom#Thing"), "https://example.org/custom#Thing");
|
||||
assert.equal(compactNodeType("owl:Class"), "owl:Class");
|
||||
});
|
||||
+36
-7
@@ -47,7 +47,11 @@ dependencies = [
|
||||
"numpy>=2.0.2",
|
||||
"pandas>=1.3.0",
|
||||
"scipy>=1.13.1",
|
||||
"scikit-learn>=1.7.2",
|
||||
# scikit-learn dropped Python 3.9 support at 1.7.0 (requires_python >=3.10),
|
||||
# so an unqualified >=1.7.2 floor is unsatisfiable on 3.9. Cap 3.9 to the
|
||||
# last 3.9-compatible release line; 3.10+ is left unconstrained.
|
||||
"scikit-learn>=1.6.1,<1.7.0; python_version < '3.10'",
|
||||
"scikit-learn>=1.7.2; python_version >= '3.10'",
|
||||
"umap-learn>=0.5.12",
|
||||
# thinc (spacy's core dep) dropped Python 3.9 wheels at 8.3.10, and later
|
||||
# spacy patch releases (3.8.8+) require thinc>=8.3.9-only-on-3.10+ ranges,
|
||||
@@ -66,24 +70,49 @@ dependencies = [
|
||||
"seaborn>=0.13.2",
|
||||
"plotly>=6.8.0",
|
||||
"ipywidgets>=8.0.0",
|
||||
"requests>=2.34.2",
|
||||
# requests dropped Python 3.9 support at 2.33.0 (requires_python >=3.10),
|
||||
# so an unqualified >=2.34.2 floor is unsatisfiable on 3.9. Cap 3.9 to the
|
||||
# last 3.9-compatible release; 3.10+ is left unconstrained.
|
||||
"requests>=2.32.5,<2.33.0; python_version < '3.10'",
|
||||
"requests>=2.34.2; python_version >= '3.10'",
|
||||
"GitPython>=3.1.58",
|
||||
"chardet>=7.4.3",
|
||||
# chardet dropped Python 3.9 support at 6.0.0 (requires_python >=3.10), so
|
||||
# an unqualified >=7.4.3 floor is unsatisfiable on 3.9. Cap 3.9 to the last
|
||||
# 3.9-compatible release; 3.10+ is left unconstrained.
|
||||
"chardet>=5.2.0,<6.0.0; python_version < '3.10'",
|
||||
"chardet>=7.4.3; python_version >= '3.10'",
|
||||
"protobuf>=5.29.1,<8.0",
|
||||
"grpcio>=1.81.1",
|
||||
# grpcio dropped Python 3.9 support at 1.81.0 (requires_python >=3.10), so
|
||||
# an unqualified >=1.81.1 floor is unsatisfiable on 3.9. Cap 3.9 to the last
|
||||
# 3.9-compatible release; 3.10+ is left unconstrained.
|
||||
"grpcio>=1.80.0,<1.81.0; python_version < '3.10'",
|
||||
"grpcio>=1.81.1; python_version >= '3.10'",
|
||||
"beautifulsoup4>=4.15.0",
|
||||
"lxml>=6.1.1",
|
||||
"python-docx>=1.2.0",
|
||||
"openpyxl>=3.1.5",
|
||||
"pillow>=12.2.0",
|
||||
# pillow dropped Python 3.9 support at 12.0.0 (requires_python >=3.10), so
|
||||
# an unqualified >=12.2.0 floor is unsatisfiable on 3.9. Cap 3.9 to the last
|
||||
# 3.9-compatible release; 3.10+ is left unconstrained.
|
||||
"pillow>=11.3.0,<12.0.0; python_version < '3.10'",
|
||||
"pillow>=12.2.0; python_version >= '3.10'",
|
||||
"librosa>=0.9.0",
|
||||
"opencv-python>=4.13.0.92",
|
||||
"faiss-cpu>=1.7.0",
|
||||
"fastembed>=0.2.0",
|
||||
"onnxruntime>=1.20.1",
|
||||
# onnxruntime stopped shipping cp39 wheels at 1.20.0 (its PyPI metadata
|
||||
# still claims requires_python >=3.9, but no matching wheel exists), so an
|
||||
# unqualified >=1.20.1 floor is unsatisfiable on 3.9. Cap 3.9 to the last
|
||||
# release with a cp39 wheel; 3.10+ is left unconstrained.
|
||||
"onnxruntime>=1.19.2,<1.20.0; python_version < '3.10'",
|
||||
"onnxruntime>=1.20.1; python_version >= '3.10'",
|
||||
"tokenizers>=0.15.0",
|
||||
"pydantic>=2.13.4",
|
||||
"click>=8.4.2",
|
||||
# click dropped Python 3.9 support at 8.2.0 (requires_python >=3.10), so an
|
||||
# unqualified >=8.4.2 floor is unsatisfiable on 3.9. Cap 3.9 to the last
|
||||
# 3.9-compatible release; 3.10+ is left unconstrained.
|
||||
"click>=8.1.8,<8.2.0; python_version < '3.10'",
|
||||
"click>=8.4.2; python_version >= '3.10'",
|
||||
"rich>=12.5.0",
|
||||
"tqdm>=4.68.3",
|
||||
"pyyaml>=6.0",
|
||||
|
||||
@@ -234,6 +234,12 @@ class EntityDetailResponse(BaseModel):
|
||||
properties: Dict[str, Any] = Field(default_factory=dict)
|
||||
|
||||
|
||||
class OntologyGraphResponse(BaseModel):
|
||||
uri: str
|
||||
nodes: List[Dict[str, Any]] = Field(default_factory=list)
|
||||
edges: List[Dict[str, Any]] = Field(default_factory=list)
|
||||
|
||||
|
||||
class SKOSScheme(BaseModel):
|
||||
uri: str
|
||||
title: str
|
||||
@@ -664,6 +670,7 @@ def _convert_ontology_to_graph(ontology_dict: Dict[str, Any]) -> Tuple[List[Dict
|
||||
"rdfs:label": cls.get("label", cls.get("name", "")),
|
||||
"rdfs:comment": cls.get("description", ""),
|
||||
"uri": cls_uri,
|
||||
"scheme_uri": ontology_uri,
|
||||
},
|
||||
}
|
||||
nodes.append(node)
|
||||
@@ -680,14 +687,21 @@ def _convert_ontology_to_graph(ontology_dict: Dict[str, Any]) -> Tuple[List[Dict
|
||||
# Add property nodes and edges
|
||||
for prop in ontology_dict.get("properties", []):
|
||||
prop_uri = prop.get("uri", f"temp:prop:{uuid.uuid4().hex[:12]}")
|
||||
property_type = {
|
||||
"object": "owl:ObjectProperty",
|
||||
"data": "owl:DatatypeProperty",
|
||||
"datatype": "owl:DatatypeProperty",
|
||||
"annotation": "owl:AnnotationProperty",
|
||||
}.get(str(prop.get("type", "object")).lower(), "owl:ObjectProperty")
|
||||
node = {
|
||||
"id": prop_uri,
|
||||
"type": f"owl:{prop.get('type', 'Object').title()}Property",
|
||||
"type": property_type,
|
||||
"content": prop.get("name", prop.get("label", "")),
|
||||
"properties": {
|
||||
"rdfs:label": prop.get("label", prop.get("name", "")),
|
||||
"rdfs:comment": prop.get("description", ""),
|
||||
"uri": prop_uri,
|
||||
"scheme_uri": ontology_uri,
|
||||
},
|
||||
}
|
||||
nodes.append(node)
|
||||
@@ -748,14 +762,38 @@ def _node_source_ontology(node: Dict[str, Any]) -> Optional[str]:
|
||||
)
|
||||
|
||||
|
||||
def _node_belongs_to_ontology(node: Dict[str, Any], ontology_uri: str) -> bool:
|
||||
def _node_belongs_to_ontology(
|
||||
node: Dict[str, Any],
|
||||
ontology_uri: str,
|
||||
known_ontology_uris: Optional[set[str]] = None,
|
||||
) -> bool:
|
||||
nid = node.get("id", "")
|
||||
if nid == ontology_uri:
|
||||
return True
|
||||
if _node_source_ontology(node) == ontology_uri:
|
||||
return True
|
||||
owner = _node_source_ontology(node)
|
||||
if owner:
|
||||
return owner == ontology_uri
|
||||
if known_ontology_uris:
|
||||
namespace_owners = [
|
||||
candidate
|
||||
for candidate in known_ontology_uris
|
||||
if nid == candidate
|
||||
or nid.startswith(
|
||||
(candidate.rstrip("#/") + "#", candidate.rstrip("#/") + "/")
|
||||
)
|
||||
]
|
||||
if namespace_owners and max(namespace_owners, key=len) != ontology_uri:
|
||||
return False
|
||||
stem = ontology_uri.rstrip("#/")
|
||||
return nid.startswith((stem + "#", stem + "/"))
|
||||
if not nid.startswith((stem + "#", stem + "/")):
|
||||
return False
|
||||
# Prefix ownership only extends to names minted directly in the
|
||||
# ontology's namespace (<stem>#Term or <stem>/Term). Any further
|
||||
# delimiter marks a nested vocabulary (<stem>/child#Term,
|
||||
# <stem>/child/Term), which must not be absorbed into the parent
|
||||
# until it is registered or carries an explicit owner.
|
||||
local_name = nid[len(stem) + 1 :]
|
||||
return "#" not in local_name and "/" not in local_name
|
||||
|
||||
|
||||
def _is_ontology_entity(node: Dict[str, Any]) -> bool:
|
||||
@@ -1221,7 +1259,8 @@ def _parse_rdf_sync(content: bytes, fmt: str) -> tuple:
|
||||
metadata.setdefault("description", str(obj))
|
||||
break
|
||||
|
||||
if "uri" not in metadata:
|
||||
synthetic_uri = "uri" not in metadata
|
||||
if synthetic_uri:
|
||||
metadata["uri"] = f"urn:semantica:onto:{uuid.uuid4().hex[:8]}"
|
||||
metadata.setdefault("name", metadata["uri"].rsplit("/", 1)[-1].rsplit("#", 1)[-1] or "Unnamed")
|
||||
metadata["triple_count"] = len(g)
|
||||
@@ -1268,6 +1307,20 @@ def _parse_rdf_sync(content: bytes, fmt: str) -> tuple:
|
||||
"weight": 1.0,
|
||||
})
|
||||
|
||||
if synthetic_uri:
|
||||
# No owl:Ontology / skos:ConceptScheme declaration exists, so the
|
||||
# synthetic registry URI shares no namespace with any node. Ownership
|
||||
# must be recorded explicitly, and the editor needs a matching graph
|
||||
# node, or the registered ontology resolves to an empty core and 404s.
|
||||
for node in nodes:
|
||||
node["properties"].setdefault("scheme_uri", metadata["uri"])
|
||||
nodes.append({
|
||||
"id": metadata["uri"],
|
||||
"type": "owl:Ontology",
|
||||
"content": metadata["name"],
|
||||
"properties": {"rdfs:label": metadata["name"], "uri": metadata["uri"]},
|
||||
})
|
||||
|
||||
return nodes, edges, metadata
|
||||
|
||||
|
||||
@@ -1768,6 +1821,107 @@ async def search_entities(
|
||||
return results
|
||||
|
||||
|
||||
@router.get("/graph", response_model=OntologyGraphResponse)
|
||||
async def get_ontology_graph(
|
||||
request: Request,
|
||||
uri: str = Query(..., min_length=1),
|
||||
session: GraphSession = Depends(get_session),
|
||||
):
|
||||
"""Return the editable schema subgraph for one registered ontology."""
|
||||
registry = _get_registry(request)
|
||||
ontology_nodes: List[Dict[str, Any]] = []
|
||||
for node_type in _ONTOLOGY_TYPES:
|
||||
nodes, _ = await asyncio.to_thread(
|
||||
session.get_nodes, node_type=node_type, skip=0, limit=2**63 - 1
|
||||
)
|
||||
ontology_nodes.extend(nodes)
|
||||
known_ontology_uris = set(registry) | {
|
||||
str(node.get("id", "")) for node in ontology_nodes if node.get("id")
|
||||
}
|
||||
if uri not in known_ontology_uris:
|
||||
raise HTTPException(status_code=404, detail="Ontology not found in registry.")
|
||||
|
||||
schema_types = _CLASS_TYPES | _PROPERTY_TYPES | _CONCEPT_TYPES | _ONTOLOGY_TYPES
|
||||
candidates_by_id: Dict[str, Dict[str, Any]] = {}
|
||||
for node_type in schema_types:
|
||||
nodes, _ = await asyncio.to_thread(
|
||||
session.get_nodes, node_type=node_type, skip=0, limit=2**63 - 1
|
||||
)
|
||||
candidates_by_id.update(
|
||||
(str(node.get("id", "")), node) for node in nodes if node.get("id")
|
||||
)
|
||||
|
||||
core_node_ids = {
|
||||
str(node.get("id", ""))
|
||||
for node in candidates_by_id.values()
|
||||
if _node_belongs_to_ontology(node, uri, known_ontology_uris)
|
||||
}
|
||||
if not core_node_ids:
|
||||
raise HTTPException(status_code=404, detail="Ontology graph not found.")
|
||||
|
||||
structure_edge_types = {
|
||||
"rdf:type",
|
||||
"rdfs:subClassOf",
|
||||
"rdfs:domain",
|
||||
"rdfs:range",
|
||||
"owl:disjointWith",
|
||||
"owl:equivalentClass",
|
||||
"owl:equivalentProperty",
|
||||
"owl:inverseOf",
|
||||
"skos:broader",
|
||||
"skos:narrower",
|
||||
"skos:related",
|
||||
}
|
||||
selected_edges: List[Dict[str, Any]] = []
|
||||
for edge_type in structure_edge_types:
|
||||
edges, _ = await asyncio.to_thread(
|
||||
session.get_edges,
|
||||
edge_type=edge_type,
|
||||
skip=0,
|
||||
limit=2**63 - 1,
|
||||
)
|
||||
# Keep only edges whose source is a core node: the requested ontology
|
||||
# may reference outward (e.g. rdfs:range to an external vocabulary),
|
||||
# but an unrelated ontology's property pointing at a core class must
|
||||
# not leak inward.
|
||||
selected_edges.extend(
|
||||
edge for edge in edges
|
||||
if str(edge.get("source", "")) in core_node_ids
|
||||
)
|
||||
if (
|
||||
len(core_node_ids) > _MAX_ANALYSIS_NODES
|
||||
or len(selected_edges) > _MAX_ANALYSIS_NODES
|
||||
):
|
||||
raise HTTPException(
|
||||
status_code=413,
|
||||
detail=(
|
||||
"Ontology editor graph exceeds the maximum size "
|
||||
f"({_MAX_ANALYSIS_NODES} nodes or edges)."
|
||||
),
|
||||
)
|
||||
|
||||
selected_node_ids = set(core_node_ids)
|
||||
for edge in selected_edges:
|
||||
selected_node_ids.add(str(edge.get("source", "")))
|
||||
selected_node_ids.add(str(edge.get("target", "")))
|
||||
|
||||
selected_nodes = [candidates_by_id[node_id] for node_id in core_node_ids]
|
||||
for node_id in selected_node_ids - core_node_ids:
|
||||
external = await asyncio.to_thread(session.get_node, node_id)
|
||||
if external is not None:
|
||||
selected_nodes.append(external)
|
||||
selected_nodes.sort(key=lambda node: str(node.get("id", "")))
|
||||
selected_edges.sort(
|
||||
key=lambda edge: (
|
||||
str(edge.get("source", "")),
|
||||
str(edge.get("type", "")),
|
||||
str(edge.get("target", "")),
|
||||
str(edge.get("id", "")),
|
||||
)
|
||||
)
|
||||
return OntologyGraphResponse(uri=uri, nodes=selected_nodes, edges=selected_edges)
|
||||
|
||||
|
||||
@router.get("/entity/{entity_uri:path}", response_model=EntityDetailResponse)
|
||||
async def get_entity_detail(
|
||||
entity_uri: str,
|
||||
|
||||
@@ -35,7 +35,7 @@ Example Usage:
|
||||
>>> llm = LiteLLM(model="openai/gpt-4o", api_key="your-key")
|
||||
>>> response = llm.generate("Hello, world!")
|
||||
>>> # Or use other providers via LiteLLM
|
||||
>>> llm = LiteLLM(model="anthropic/claude-sonnet-4-20250514")
|
||||
>>> llm = LiteLLM(model="anthropic/claude-sonnet-5")
|
||||
>>> response = llm.generate("Hello, world!")
|
||||
>>>
|
||||
>>> # Anthropic provider
|
||||
|
||||
@@ -31,7 +31,7 @@ class LiteLLM:
|
||||
Provides unified interface to 100+ LLM providers through LiteLLM library.
|
||||
Supports providers like OpenAI, Anthropic, Groq, Azure, Bedrock, Vertex AI, etc.
|
||||
|
||||
Model format: "provider/model-name" (e.g., "openai/gpt-4o", "anthropic/claude-sonnet-4-20250514", "groq/llama-3.1-8b-instant")
|
||||
Model format: "provider/model-name" (e.g., "openai/gpt-4o", "anthropic/claude-sonnet-5", "groq/llama-3.1-8b-instant")
|
||||
|
||||
Example:
|
||||
>>> from semantica.llms import LiteLLM
|
||||
@@ -39,7 +39,7 @@ class LiteLLM:
|
||||
>>> response = llm.generate("What is AI?")
|
||||
>>>
|
||||
>>> # Use with different providers
|
||||
>>> llm = LiteLLM(model="anthropic/claude-sonnet-4-20250514")
|
||||
>>> llm = LiteLLM(model="anthropic/claude-sonnet-5")
|
||||
>>> response = llm.generate("Hello!")
|
||||
"""
|
||||
|
||||
@@ -54,7 +54,7 @@ class LiteLLM:
|
||||
|
||||
Args:
|
||||
model: Model identifier in format "provider/model-name"
|
||||
Examples: "openai/gpt-4o", "anthropic/claude-sonnet-4-20250514",
|
||||
Examples: "openai/gpt-4o", "anthropic/claude-sonnet-5",
|
||||
"groq/llama-3.1-8b-instant", "azure/gpt-4", etc.
|
||||
api_key: API key (optional, can use environment variables)
|
||||
**kwargs: Additional LiteLLM options (temperature, max_tokens, etc.)
|
||||
|
||||
@@ -37,6 +37,29 @@ from .naming_conventions import NamingConventions
|
||||
from .relationship_utils import build_entity_aliases, resolve_relationship_endpoint_type
|
||||
|
||||
|
||||
# Top-level entity keys that describe structure or provenance rather than
|
||||
# business attributes. GraphBuilder and EntityMerger attach these to entity
|
||||
# dicts (relationships list, nested properties/metadata maps, merge history),
|
||||
# so they must not be inferred as datatype properties. Each key mirrors what
|
||||
# the framework actually writes to a merged entity top level
|
||||
# (see MergeStrategyManager._merge_entities merged_entity dict and GraphBuilder).
|
||||
_CONTROL_FIELDS = frozenset(
|
||||
{
|
||||
"id",
|
||||
"type",
|
||||
"entity_type",
|
||||
"text",
|
||||
"label",
|
||||
"confidence",
|
||||
"properties",
|
||||
"relationships",
|
||||
"metadata",
|
||||
"merged_from",
|
||||
"merge_strategy",
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
class PropertyGenerator:
|
||||
"""
|
||||
Property generation engine for ontologies.
|
||||
@@ -347,7 +370,7 @@ class PropertyGenerator:
|
||||
|
||||
for entity in entities:
|
||||
for key, value in entity.items():
|
||||
if key in ["id", "type", "entity_type", "text", "label", "confidence"]:
|
||||
if key in _CONTROL_FIELDS:
|
||||
continue
|
||||
|
||||
# Infer type
|
||||
|
||||
@@ -12,7 +12,11 @@ from semantica.context.context_graph import ContextGraph
|
||||
pytest.importorskip("fastapi")
|
||||
|
||||
from semantica.explorer.app import create_app # noqa: E402
|
||||
from semantica.explorer.routes.ontology import OntologyEntry # noqa: E402
|
||||
from semantica.explorer.routes.ontology import ( # noqa: E402
|
||||
OntologyEntry,
|
||||
_convert_ontology_to_graph,
|
||||
_node_belongs_to_ontology,
|
||||
)
|
||||
from semantica.explorer.session import GraphSession # noqa: E402
|
||||
|
||||
from starlette.testclient import TestClient # noqa: E402
|
||||
@@ -131,6 +135,181 @@ def test_health_returns_dimensions_and_issues(client):
|
||||
assert isinstance(payload["issues"], list)
|
||||
|
||||
|
||||
def test_ontology_graph_returns_editable_schema_nodes_and_edges(client):
|
||||
response = client.get(
|
||||
"/api/ontology/graph",
|
||||
params={"uri": "http://example.org/onto-a"},
|
||||
)
|
||||
|
||||
assert response.status_code == 200
|
||||
payload = response.json()
|
||||
node_ids = {node["id"] for node in payload["nodes"]}
|
||||
assert "http://example.org/onto-a" in node_ids
|
||||
assert "http://example.org/onto-a#Person" in node_ids
|
||||
assert "http://example.org/onto-a#name" in node_ids
|
||||
assert any(
|
||||
edge["source"] == "http://example.org/onto-a#name"
|
||||
and edge["target"] == "http://example.org/onto-a#Person"
|
||||
and edge["type"] == "rdfs:domain"
|
||||
for edge in payload["edges"]
|
||||
)
|
||||
|
||||
|
||||
def test_ontology_graph_rejects_unregistered_namespace(client):
|
||||
response = client.get(
|
||||
"/api/ontology/graph",
|
||||
params={"uri": "http://example.org"},
|
||||
)
|
||||
|
||||
assert response.status_code == 404
|
||||
|
||||
|
||||
def test_ontology_graph_excludes_separately_registered_nested_ontology(client):
|
||||
graph = client.app.state.session.graph
|
||||
nested = "http://example.org/onto-a/nested"
|
||||
nested_class = f"{nested}#PrivateClass"
|
||||
graph.add_node(nested, node_type="owl:Ontology", content="Nested Ontology")
|
||||
graph.add_node(nested_class, node_type="owl:Class", content="Private Class")
|
||||
|
||||
response = client.get(
|
||||
"/api/ontology/graph",
|
||||
params={"uri": "http://example.org/onto-a"},
|
||||
)
|
||||
|
||||
assert response.status_code == 200
|
||||
node_ids = {node["id"] for node in response.json()["nodes"]}
|
||||
assert nested not in node_ids
|
||||
assert nested_class not in node_ids
|
||||
|
||||
|
||||
def test_ontology_graph_prefers_explicit_ownership_over_uri_namespace(client):
|
||||
graph = client.app.state.session.graph
|
||||
explicit_member = "http://unrelated.example/Person"
|
||||
graph.add_node(
|
||||
explicit_member,
|
||||
node_type="owl:Class",
|
||||
content="Explicit Member",
|
||||
scheme_uri="http://example.org/onto-a",
|
||||
)
|
||||
|
||||
response = client.get(
|
||||
"/api/ontology/graph",
|
||||
params={"uri": "http://example.org/onto-a"},
|
||||
)
|
||||
|
||||
assert response.status_code == 200
|
||||
assert explicit_member in {node["id"] for node in response.json()["nodes"]}
|
||||
|
||||
|
||||
def test_ontology_graph_excludes_inward_edges_from_other_ontologies(client):
|
||||
graph = client.app.state.session.graph
|
||||
foreign_prop = "http://example.org/onto-b#recordOf"
|
||||
graph.add_node(
|
||||
foreign_prop,
|
||||
node_type="owl:ObjectProperty",
|
||||
content="record of",
|
||||
scheme_uri="http://example.org/onto-b",
|
||||
)
|
||||
# onto-b's property points its domain at onto-a's class: an inward
|
||||
# reference that must not pull the foreign property into onto-a's graph.
|
||||
graph.add_edge(foreign_prop, "http://example.org/onto-a#Person", edge_type="rdfs:domain")
|
||||
|
||||
response = client.get(
|
||||
"/api/ontology/graph",
|
||||
params={"uri": "http://example.org/onto-a"},
|
||||
)
|
||||
|
||||
assert response.status_code == 200
|
||||
payload = response.json()
|
||||
assert foreign_prop not in {node["id"] for node in payload["nodes"]}
|
||||
assert all(edge["source"] != foreign_prop for edge in payload["edges"])
|
||||
|
||||
|
||||
def test_ontology_graph_excludes_unregistered_nested_namespace(client):
|
||||
graph = client.app.state.session.graph
|
||||
nested_class = "http://example.org/onto-a/vocab#Term"
|
||||
graph.add_node(nested_class, node_type="owl:Class", content="Nested Term")
|
||||
|
||||
response = client.get(
|
||||
"/api/ontology/graph",
|
||||
params={"uri": "http://example.org/onto-a"},
|
||||
)
|
||||
|
||||
assert response.status_code == 200
|
||||
assert nested_class not in {node["id"] for node in response.json()["nodes"]}
|
||||
|
||||
|
||||
def test_node_belongs_to_ontology_nested_namespace_matrix():
|
||||
parent = "http://example.org/onto-a"
|
||||
child = "http://example.org/onto-a/nested"
|
||||
|
||||
def node(node_id):
|
||||
return {"id": node_id, "properties": {}}
|
||||
|
||||
assert _node_belongs_to_ontology(node(f"{parent}#Person"), parent, {parent})
|
||||
assert _node_belongs_to_ontology(node(f"{parent}/Person"), parent, {parent})
|
||||
# An unregistered nested namespace is not absorbed into the parent,
|
||||
# whether fragment-based or path-based
|
||||
assert not _node_belongs_to_ontology(node(f"{child}#Term"), parent, {parent})
|
||||
assert not _node_belongs_to_ontology(node(f"{child}/Term"), parent, {parent})
|
||||
# Once registered, the nested namespace owns its nodes
|
||||
assert not _node_belongs_to_ontology(node(f"{child}#Term"), parent, {parent, child})
|
||||
assert _node_belongs_to_ontology(node(f"{child}#Term"), child, {parent, child})
|
||||
assert _node_belongs_to_ontology(node(f"{child}/Term"), child, {parent, child})
|
||||
|
||||
|
||||
def test_load_fallback_import_without_declaration_is_editable(client):
|
||||
turtle = """
|
||||
@prefix ex: <http://data.example.org/people#> .
|
||||
@prefix rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#> .
|
||||
@prefix rdfs: <http://www.w3.org/2000/01/rdf-schema#> .
|
||||
|
||||
ex:Employee a rdfs:Class ;
|
||||
rdfs:label "Employee" .
|
||||
ex:manager a rdf:Property ;
|
||||
rdfs:label "manager" .
|
||||
"""
|
||||
with patch(
|
||||
"semantica.ingest.ontology_ingestor.OntologyIngestor.ingest_ontology",
|
||||
side_effect=RuntimeError("force fallback parser"),
|
||||
):
|
||||
loaded = client.post(
|
||||
"/api/ontology/load",
|
||||
json={"content": turtle, "format": "turtle"},
|
||||
)
|
||||
assert loaded.status_code == 200
|
||||
uri = loaded.json()["uri"]
|
||||
assert uri.startswith("urn:semantica:onto:")
|
||||
|
||||
response = client.get("/api/ontology/graph", params={"uri": uri})
|
||||
assert response.status_code == 200
|
||||
payload = response.json()
|
||||
node_ids = {node["id"] for node in payload["nodes"]}
|
||||
assert uri in node_ids
|
||||
assert "http://data.example.org/people#Employee" in node_ids
|
||||
|
||||
|
||||
def test_ontology_graph_ignores_unrelated_data_when_enforcing_size_limit(client):
|
||||
graph = client.app.state.session.graph
|
||||
for index in range(5_001):
|
||||
graph.add_node(
|
||||
f"urn:unrelated:{index}",
|
||||
node_type="owl:Class",
|
||||
content="Unrelated",
|
||||
scheme_uri="http://example.org/onto-b",
|
||||
)
|
||||
|
||||
response = client.get(
|
||||
"/api/ontology/graph",
|
||||
params={"uri": "http://example.org/onto-a"},
|
||||
)
|
||||
|
||||
assert response.status_code == 200
|
||||
assert "http://example.org/onto-a#Person" in {
|
||||
node["id"] for node in response.json()["nodes"]
|
||||
}
|
||||
|
||||
|
||||
def test_shacl_generate_and_shapes(client):
|
||||
response = client.post(
|
||||
"/api/ontology/shacl/generate",
|
||||
@@ -696,6 +875,29 @@ def test_ontology_load_does_not_swallow_422_from_ingestor_success_path(client):
|
||||
fallback_parse.assert_not_called()
|
||||
|
||||
|
||||
def test_convert_ontology_uses_standard_property_types_and_scheme_uri():
|
||||
ontology_uri = "http://example.org/onto"
|
||||
nodes, _ = _convert_ontology_to_graph(
|
||||
{
|
||||
"uri": ontology_uri,
|
||||
"name": "Example Ontology",
|
||||
"classes": [
|
||||
{"uri": f"{ontology_uri}#Person", "name": "Person"},
|
||||
],
|
||||
"properties": [
|
||||
{"uri": f"{ontology_uri}#name", "name": "name", "type": "data"},
|
||||
{"uri": f"{ontology_uri}#knows", "name": "knows", "type": "object"},
|
||||
],
|
||||
}
|
||||
)
|
||||
|
||||
by_id = {node["id"]: node for node in nodes}
|
||||
assert by_id[f"{ontology_uri}#Person"]["properties"]["scheme_uri"] == ontology_uri
|
||||
assert by_id[f"{ontology_uri}#name"]["type"] == "owl:DatatypeProperty"
|
||||
assert by_id[f"{ontology_uri}#knows"]["type"] == "owl:ObjectProperty"
|
||||
assert by_id[f"{ontology_uri}#name"]["properties"]["scheme_uri"] == ontology_uri
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# refresh_ontology — single combined add_nodes_and_edges() coverage (#775)
|
||||
# ---------------------------------------------------------------------------
|
||||
@@ -789,5 +991,3 @@ def test_refresh_ontology_missing_source_url_returns_422(client):
|
||||
response = client.post(f"/api/ontology/{encoded_uri}/refresh")
|
||||
assert response.status_code == 422
|
||||
assert "source url" in response.json()["detail"].lower()
|
||||
|
||||
|
||||
|
||||
@@ -0,0 +1,94 @@
|
||||
"""Framework/control entity keys must not be inferred as datatype properties.
|
||||
|
||||
The _CONTROL_FIELDS skip set mirrors exactly what the framework writes to a
|
||||
merged entity's top level (see MergeStrategyManager._merge_entities): no extra
|
||||
guesses, so business attributes that merely share a common name (e.g. source)
|
||||
keep getting inferred.
|
||||
"""
|
||||
|
||||
from semantica.deduplication.merge_strategy import MergeStrategyManager
|
||||
from semantica.ontology.property_generator import PropertyGenerator
|
||||
|
||||
|
||||
def _merged_entity():
|
||||
"""Run a real merge so the entity carries the framework's actual top-level keys."""
|
||||
manager = MergeStrategyManager(default_strategy="keep_most_complete")
|
||||
result = manager.merge_entities(
|
||||
[
|
||||
{"id": "b1", "name": "Hangzhou Branch", "type": "ORG", "employee_count": 120},
|
||||
{"id": "b2", "name": "Hangzhou Branch", "type": "ORG"},
|
||||
]
|
||||
)
|
||||
return result.merged_entity
|
||||
|
||||
|
||||
def test_framework_fields_not_inferred_as_data_properties():
|
||||
entity = _merged_entity()
|
||||
classes = [{"name": "Organization", "metadata": {"inferred_from": "ORG"}}]
|
||||
|
||||
properties = PropertyGenerator().infer_properties([entity], [], classes)
|
||||
|
||||
names = {p["name"] for p in properties}
|
||||
for framed in (
|
||||
"properties",
|
||||
"relationships",
|
||||
"metadata",
|
||||
"merged_from",
|
||||
"merge_strategy",
|
||||
):
|
||||
assert framed not in names, f"framework field {framed} leaked as a property"
|
||||
|
||||
|
||||
def test_business_attributes_still_inferred():
|
||||
entity = _merged_entity()
|
||||
classes = [{"name": "Organization", "metadata": {"inferred_from": "ORG"}}]
|
||||
|
||||
properties = PropertyGenerator().infer_properties([entity], [], classes)
|
||||
|
||||
names = {p["name"] for p in properties}
|
||||
assert "name" in names
|
||||
assert "metadata" not in names
|
||||
|
||||
|
||||
def test_source_field_still_inferred_as_business_attribute():
|
||||
"""A top-level 'source' is a business attribute, not a framework field."""
|
||||
entity = {
|
||||
"id": "b1",
|
||||
"name": "Hangzhou Branch",
|
||||
"type": "ORG",
|
||||
"source": "doc-42",
|
||||
}
|
||||
classes = [{"name": "Organization", "metadata": {"inferred_from": "ORG"}}]
|
||||
|
||||
properties = PropertyGenerator().infer_properties([entity], [], classes)
|
||||
|
||||
names = {p["name"] for p in properties}
|
||||
assert "source" in names
|
||||
|
||||
|
||||
def test_unmerged_graphbuilder_entities_infer_business_attributes():
|
||||
"""Flat entities from GraphBuilder (merge_entities=False) must not lose business
|
||||
attributes through _CONTROL_FIELDS: name and domain-specific fields must be
|
||||
inferred, and none of the framework keys should appear in the output."""
|
||||
from semantica.kg.graph_builder import GraphBuilder
|
||||
|
||||
builder = GraphBuilder(merge_entities=False, resolve_conflicts=False)
|
||||
graph = builder.build(
|
||||
{
|
||||
"entities": [
|
||||
{"id": "c1", "name": "Chengdu Plant", "type": "ORG", "headcount": 300},
|
||||
{"id": "c2", "name": "Wuhan Plant", "type": "ORG", "headcount": 450},
|
||||
],
|
||||
"relationships": [],
|
||||
}
|
||||
)
|
||||
entities = graph["entities"]
|
||||
classes = [{"name": "Organization", "metadata": {"inferred_from": "ORG"}}]
|
||||
|
||||
properties = PropertyGenerator().infer_properties(entities, [], classes)
|
||||
|
||||
names = {p["name"] for p in properties}
|
||||
assert "name" in names, "name must be inferred from flat GraphBuilder entities"
|
||||
assert "headcount" in names, "domain business attribute must be inferred"
|
||||
for framed in ("properties", "relationships", "metadata", "merged_from", "merge_strategy"):
|
||||
assert framed not in names, f"framework field {framed!r} must not appear"
|
||||
Reference in New Issue
Block a user