mirror of
https://github.com/semantica-agi/semantica.git
synced 2026-08-29 04:26:20 +00:00
* docs: replace Exported Classes import blocks with summary tables across all 25 modules * docs: add method/parameter tables to parse, ingest, ontology, normalize, triplet_store, change_management, conflicts, export, graph_store, provenance, and semantic_extract modules
14 KiB
14 KiB
title, description, icon
| title | description | icon |
|---|---|---|
| Export Module | Export knowledge graphs to RDF, Parquet, LPG, ArangoDB AQL, CSV, GraphML, OWL, JSON-LD, and vector formats. | file-export |
semantica.export serializes knowledge graphs to every downstream format — semantic web standards, analytics pipelines, graph databases, and vector stores. All exporters share a consistent export(graph, path, format) interface.
Exported Classes
| Class | Output formats | Notes |
|---|---|---|
RDFExporter |
Turtle, JSON-LD, N-Triples, RDF/XML | Optional PROV-O provenance embedding |
ParquetExporter |
.parquet |
PyArrow-typed, Hive-partition support |
LPGExporter |
Cypher CREATE/MERGE |
Neo4j and Memgraph compatible |
ArangoAQLExporter |
AQL INSERT |
Vertex and edge collections |
GraphExporter |
GraphML, GEXF, Graphviz DOT | Standard graph interchange formats |
OWLExporter |
OWL 2.0 in Turtle/XML/JSON-LD | Full ontology serialization |
CSVExporter |
.csv |
Flat nodes + edges tables |
VectorExporter |
JSON, NumPy .npy, FAISS index |
Embedding vector export |
ArrowExporter |
Apache Arrow IPC | Zero-copy transfer to Pandas/Polars/Spark |
DistanceExporter |
JSON/CSV matrix | Semantic distance matrices and ego-graphs |
ReportGenerator |
HTML, Markdown, JSON | Human-readable analytics reports |
NamespaceManager |
— | Register and resolve RDF namespace prefixes |
What You Get
Turtle, JSON-LD, N-Triples, RDF/XML with namespace management and optional PROV-O provenance embedding. Columnar storage for Spark, BigQuery, Databricks, and Snowflake with explicit PyArrow typing. Cypher CREATE/MERGE for Neo4j and Memgraph; AQL INSERT for ArangoDB vertex and edge collections. GraphML, GEXF, DOT for Gephi and Graphviz. OWL 2.0 ontology export in Turtle, XML, and JSON-LD. JSON, NumPy `.npy`, and FAISS index export for embedding vectors. Apache Arrow IPC for zero-copy transfer. Distance matrix CSV/JSON from Distance Intelligence (v0.5.0). HTML, Markdown, and JSON analytics reports.Quick Start
```python from semantica.export import RDFExporterexporter = RDFExporter()
```
export_rdf(graph, "output.ttl", format="turtle")
export_parquet(graph, "output/", compression="snappy")
export_csv(graph, "nodes.csv", target="nodes")
```
exporter = ParquetExporter(compression="snappy")
exporter.export_stream(graph, output_dir="output/", batch_size=10_000)
```
RDFExporter Constructor Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
namespace_manager |
NamespaceManager |
None |
Custom namespace prefix manager |
include_provenance |
bool |
False |
Embed W3C PROV-O lineage triples |
provenance_manager |
ProvenanceManager |
None |
Provenance source when include_provenance=True |
Exporters
Export to W3C RDF formats — Turtle, JSON-LD, N-Triples, and RDF/XML:```python
from semantica.export import RDFExporter
exporter = RDFExporter()
# Export to RDF string — write to file manually
rdf_str = exporter.export_to_rdf(graph, format="turtle") # Turtle (most readable)
rdf_str = exporter.export_to_rdf(graph, format="json-ld") # JSON-LD (APIs, Linked Data)
rdf_str = exporter.export_to_rdf(graph, format="nt") # N-Triples (streaming-friendly)
rdf_str = exporter.export_to_rdf(graph, format="xml") # RDF/XML (W3C standard)
with open("output.ttl", "w") as f:
f.write(exporter.export_to_rdf(graph, format="turtle"))
```
**Custom namespace management:**
```python
from semantica.export import NamespaceManager, RDFExporter
ns_manager = NamespaceManager()
ns_manager.register("ex", "http://example.org/")
ns_manager.register("schema", "https://schema.org/")
exporter = RDFExporter(namespace_manager=ns_manager)
```
**Export with PROV-O provenance:**
```python
from semantica.export import RDFExporter
from semantica.provenance import ProvenanceManager
provenance = ProvenanceManager()
# ... track entities during extraction ...
exporter = RDFExporter(include_provenance=True, provenance_manager=provenance)
exporter.export_to_file(graph, "output_with_prov.ttl", format="turtle")
# → Each entity's prov:wasGeneratedBy, prov:wasDerivedFrom, prov:hadPrimarySource
# triples are included alongside the entity data triples
```
```python
from semantica.export import ParquetExporter
exporter = ParquetExporter(compression="snappy")
# compression options: snappy | gzip | brotli | zstd | lz4
# Export nodes and edges as separate Parquet files
exporter.export_nodes(graph, "nodes.parquet")
exporter.export_edges(graph, "edges.parquet")
# Export full graph partitioned by node type
exporter.export(graph, output_dir="graph_parquet/", partition_by="node_type")
```
Schema is explicitly typed with PyArrow for clean Spark/BigQuery ingestion.
```python
from semantica.export import CSVExporter
exporter = CSVExporter(delimiter=",")
exporter.export_nodes(graph, "nodes.csv")
exporter.export_edges(graph, "edges.csv")
```
```python
from semantica.export import SemanticNetworkYAMLExporter
exporter = SemanticNetworkYAMLExporter()
exporter.export(graph, "graph.yaml")
yaml_str = exporter.to_string(graph)
```
```python
from semantica.export import LPGExporter
exporter = LPGExporter()
# CREATE statements
cypher = exporter.to_cypher(graph)
exporter.export(graph, "import.cypher", format="cypher")
# MERGE statements (idempotent — safe to re-run)
cypher_merge = exporter.to_cypher(graph, use_merge=True)
```
```python
from semantica.export import ArangoAQLExporter
exporter = ArangoAQLExporter(
vertex_collection="entities",
edge_collection="relationships"
)
aql = exporter.export(graph) # returns AQL string
exporter.export_to_file(graph, "import.aql")
```
```python
from semantica.export import GraphExporter
exporter = GraphExporter()
exporter.export(graph, "graph.graphml", format="graphml") # Gephi, yEd
exporter.export(graph, "graph.gexf", format="gexf") # Gephi streaming
exporter.export(graph, "graph.dot", format="dot") # Graphviz
```
```python
from semantica.export import OWLExporter
exporter = OWLExporter()
exporter.export(ontology, path="ontology.ttl", format="turtle")
exporter.export(ontology, path="ontology.owl", format="xml")
exporter.export(ontology, path="ontology.json", format="json-ld")
```
```python
from semantica.export import VectorExporter
exporter = VectorExporter()
exporter.export(embeddings, metadata, "vectors.json", format="json")
exporter.export(embeddings, metadata, "vectors.npy", format="numpy")
exporter.export(embeddings, metadata, "vectors.faiss", format="faiss")
```
```python
from semantica.export import ArrowExporter
exporter = ArrowExporter()
exporter.export(graph, "graph.arrow") # requires pyarrow
```
```python
from semantica.export import DistanceExporter
exporter = DistanceExporter()
exporter.export_matrix(distance_matrix, node_labels, "distances.csv")
exporter.export_ego(ego_neighborhood, center_node="Apple Inc.", path="ego.json")
```
```python
from semantica.export import ReportGenerator
generator = ReportGenerator()
generator.generate(graph, analytics_result, "report.html", format="html")
generator.generate(graph, analytics_result, "report.md", format="markdown")
generator.generate(graph, analytics_result, "report.json", format="json")
```
Streaming Export
For graphs too large to hold in memory, use streaming export — writes incrementally without buffering the full graph:
from semantica.export import RDFExporter, ParquetExporter
# Stream RDF — yields triples one at a time, no full-graph buffer
exporter = RDFExporter()
with exporter.stream(graph, format="turtle") as stream:
for triple_line in stream:
output_file.write(triple_line)
# Stream Parquet — writes row groups incrementally
exporter = ParquetExporter(compression="snappy")
exporter.export_stream(graph, output_dir="output/", batch_size=10_000)
Streaming is recommended for graphs with > 500k nodes.
Selective Export
# Export a subgraph
subgraph = graph.subgraph(node_ids=["apple_inc", "steve_jobs"])
export_rdf(subgraph, "subgraph.ttl", format="turtle")
# Export nodes by type
org_nodes = graph.filter(node_type="Organization")
export_parquet(org_nodes, "organizations.parquet")
Convenience Functions
from semantica.export import (
export_rdf, export_parquet, export_csv, export_lpg,
export_arango, export_graph, export_owl, export_vector,
export_arrow, generate_report
)
export_rdf(graph, "output.ttl", format="turtle")
export_parquet(graph, "output/", compression="snappy")
export_csv(graph, "nodes.csv", target="nodes")
export_lpg(graph, "import.cypher", method="cypher")
export_arango(graph, "import.aql")
export_graph(graph, "graph.graphml", format="graphml")
Format Reference
| Format | Exporter | Output | Best For |
|---|---|---|---|
turtle |
RDFExporter |
.ttl |
Readable RDF, ontology sharing |
json-ld |
RDFExporter |
.jsonld |
APIs, Linked Data, JSON pipelines |
nt |
RDFExporter |
.nt |
Streaming RDF, line-by-line processing |
xml |
RDFExporter |
.xml |
W3C RDF/XML, broadest compatibility |
parquet |
ParquetExporter |
.parquet |
Spark, BigQuery, Databricks, Snowflake |
cypher |
LPGExporter |
.cypher |
Neo4j, Memgraph import |
aql |
ArangoAQLExporter |
.aql |
ArangoDB vertex + edge collections |
graphml |
GraphExporter |
.graphml |
Gephi, yEd visualization |
gexf |
GraphExporter |
.gexf |
Gephi streaming format |
dot |
GraphExporter |
.dot |
Graphviz rendering |
owl |
OWLExporter |
.owl / .ttl |
OWL 2.0 ontology distribution |
csv |
CSVExporter |
.csv |
Spreadsheets, simple pipelines |
yaml |
SemanticNetworkYAMLExporter |
.yaml |
Human-readable, config-driven use |
arrow |
ArrowExporter |
.arrow |
Zero-copy inter-process transfer |
numpy |
VectorExporter |
.npy |
NumPy arrays from embeddings |
faiss |
VectorExporter |
.faiss |
Direct FAISS index files |
distance-matrix |
DistanceExporter |
.csv / .json |
Distance Intelligence matrices |
html |
ReportGenerator |
.html |
Human-readable analytics reports |
markdown |
ReportGenerator |
.md |
Documentation, GitHub |