mirror of
https://github.com/semantica-agi/semantica.git
synced 2026-08-29 04:26:20 +00:00
- Rewrote all 26 reference module pages: removed blockquote taglines and horizontal rule separators, added "What You Get" bullet summaries, added constructor/method parameter tables, expanded thin files (graph_store, triplet_store, visualization, provenance) with full API coverage, added backend comparison tables and real-world usage patterns - Renamed Modules tab from "API Reference" and group from "Context & Knowledge" to "Context & Intelligence" in docs.json - Fixed logo: copied "Semantica Logo.png" to web-safe semantica-logo.png and updated all 4 references in docs.json - Improved core docs (index, modules, concepts, quickstart, installation, getting-started) with better fonts, bullet points, and complete module listings (mcp_server, evals, core, utils previously missing) - Rewrote community pages (community, community-projects, contributing-guide, use-cases, architecture, faq, learning-more, glossary) with heading hierarchy fixes, expanded definitions, and better structure - Fixed markdown linter warnings: MD036 bold-as-heading, MD001 heading skips, MD040 missing code fence language, MD032 blank lines around lists
4.7 KiB
4.7 KiB
title, description, icon
| title | description | icon |
|---|---|---|
| Seed Module | Seed data management for initializing Knowledge Graphs from trusted, verified sources. | database |
semantica.seed provides a system for bootstrapping Knowledge Graphs with verified, structured data from trusted sources — taxonomies, reference tables, user lists, product catalogs — so you start with a reliable foundation rather than an empty graph.
What You Get
SeedDataManager— register sources, build a foundation graph, and integrate with extracted dataSeedDataSource— define individual data sources with format, path, and config- Merge strategies —
seed_first,extracted_first,smart_mergefor combining seed and extracted data - Validation — check data quality and schema compliance before loading
- Versioning — track and manage versions of seed data sources
- Bootstrapping — you have existing structured data (taxonomies, user lists, product catalogs) to build on
- Reference data — load immutable reference information (countries, ISO codes, ontology terms)
- Testing — load consistent, reproducible datasets for development and CI pipelines
SeedDataManager
The primary interface for seed data:
from semantica.seed import SeedDataManager
manager = SeedDataManager()
manager.register_source("countries", "csv", "data/countries.csv")
manager.register_source("taxonomy", "json", "data/taxonomy.json")
# Build a foundation KG from all registered sources
foundation_kg = manager.create_foundation_graph()
Core Methods
| Method | Description |
|---|---|
register_source(name, format, location) |
Add a data source to the registry |
create_foundation_graph() |
Build a KG from all registered sources |
validate_quality(seed_data) |
Check data quality and completeness |
integrate_with_extracted(seed, extracted) |
Merge seed and extracted graphs |
export_seed_data(path, format) |
Export seed data to RDF, JSON, or CSV |
SeedDataSource
Define a source with format-specific configuration:
from semantica.seed import SeedDataSource
source = SeedDataSource(
name="taxonomy",
type="json", # "csv" | "json" | "api" | "sql"
path="taxonomy.json",
config={"encoding": "utf-8"}
)
Merge Strategies
Control how seed data and extracted data are combined:
final_kg = manager.integrate_seed_extracted(
seed_graph=foundation_kg,
extracted_data=new_data,
strategy="seed_first" # see options below
)
| Strategy | Behavior |
|---|---|
seed_first |
Seed data wins on conflicts — use for authoritative reference data |
extracted_first |
Extracted data overrides seed — use when new data is more current |
smart_merge |
Property-level merging with conflict detection and resolution |
Bootstrapping a KG
Full example — load foundation data then merge with freshly ingested content:
from semantica.seed import SeedDataManager
from semantica.ingest import FileIngestor
from semantica.semantic_extract import NERExtractor
from semantica.llms import Groq
import os
# Build foundation from verified reference data
manager = SeedDataManager()
manager.register_source("taxonomy", "json", "taxonomy.json")
manager.register_source("employees", "csv", "employees.csv")
foundation_kg = manager.create_foundation_graph()
# Ingest and extract from new sources
llm = Groq(model="llama-3.3-70b-versatile", api_key=os.getenv("GROQ_API_KEY"))
ingestor = FileIngestor()
ner = NERExtractor(method="llm", llm_provider=llm)
sources = ingestor.ingest("news_articles/")
new_data = [ner.extract(s.content) for s in sources]
# Merge — seed data takes precedence for reference facts
final_kg = manager.integrate_seed_extracted(
seed_graph=foundation_kg,
extracted_data=new_data,
strategy="seed_first"
)
Configuration
seed:
sources:
- name: "employees"
type: "csv"
path: "./data/employees.csv"
- name: "taxonomy"
type: "json"
path: "./data/taxonomy.json"
merge:
strategy: "seed_first"
validation:
strict: true
Environment variable overrides:
export SEED_DATA_DIR=./data/seed
export SEED_MERGE_STRATEGY=seed_first