Files
semantica/docs/reference/seed.md
T
KaifAhmad1 9113ef3428 docs: premium overhaul of all reference pages and core docs
- Rewrote all 26 reference module pages: removed blockquote taglines and
  horizontal rule separators, added "What You Get" bullet summaries,
  added constructor/method parameter tables, expanded thin files
  (graph_store, triplet_store, visualization, provenance) with full API
  coverage, added backend comparison tables and real-world usage patterns
- Renamed Modules tab from "API Reference" and group from "Context &
  Knowledge" to "Context & Intelligence" in docs.json
- Fixed logo: copied "Semantica Logo.png" to web-safe semantica-logo.png
  and updated all 4 references in docs.json
- Improved core docs (index, modules, concepts, quickstart, installation,
  getting-started) with better fonts, bullet points, and complete module
  listings (mcp_server, evals, core, utils previously missing)
- Rewrote community pages (community, community-projects, contributing-guide,
  use-cases, architecture, faq, learning-more, glossary) with heading
  hierarchy fixes, expanded definitions, and better structure
- Fixed markdown linter warnings: MD036 bold-as-heading, MD001 heading
  skips, MD040 missing code fence language, MD032 blank lines around lists
2026-05-23 13:10:09 +05:30

4.7 KiB

title, description, icon
title description icon
Seed Module Seed data management for initializing Knowledge Graphs from trusted, verified sources. database

semantica.seed provides a system for bootstrapping Knowledge Graphs with verified, structured data from trusted sources — taxonomies, reference tables, user lists, product catalogs — so you start with a reliable foundation rather than an empty graph.

What You Get

  • SeedDataManager — register sources, build a foundation graph, and integrate with extracted data
  • SeedDataSource — define individual data sources with format, path, and config
  • Merge strategiesseed_first, extracted_first, smart_merge for combining seed and extracted data
  • Validation — check data quality and schema compliance before loading
  • Versioning — track and manage versions of seed data sources
**When to use the Seed Module:**
  • Bootstrapping — you have existing structured data (taxonomies, user lists, product catalogs) to build on
  • Reference data — load immutable reference information (countries, ISO codes, ontology terms)
  • Testing — load consistent, reproducible datasets for development and CI pipelines

SeedDataManager

The primary interface for seed data:

from semantica.seed import SeedDataManager

manager = SeedDataManager()
manager.register_source("countries", "csv",  "data/countries.csv")
manager.register_source("taxonomy",  "json", "data/taxonomy.json")

# Build a foundation KG from all registered sources
foundation_kg = manager.create_foundation_graph()

Core Methods

Method Description
register_source(name, format, location) Add a data source to the registry
create_foundation_graph() Build a KG from all registered sources
validate_quality(seed_data) Check data quality and completeness
integrate_with_extracted(seed, extracted) Merge seed and extracted graphs
export_seed_data(path, format) Export seed data to RDF, JSON, or CSV

SeedDataSource

Define a source with format-specific configuration:

from semantica.seed import SeedDataSource

source = SeedDataSource(
    name="taxonomy",
    type="json",          # "csv" | "json" | "api" | "sql"
    path="taxonomy.json",
    config={"encoding": "utf-8"}
)

Merge Strategies

Control how seed data and extracted data are combined:

final_kg = manager.integrate_seed_extracted(
    seed_graph=foundation_kg,
    extracted_data=new_data,
    strategy="seed_first"   # see options below
)
Strategy Behavior
seed_first Seed data wins on conflicts — use for authoritative reference data
extracted_first Extracted data overrides seed — use when new data is more current
smart_merge Property-level merging with conflict detection and resolution

Bootstrapping a KG

Full example — load foundation data then merge with freshly ingested content:

from semantica.seed import SeedDataManager
from semantica.ingest import FileIngestor
from semantica.semantic_extract import NERExtractor
from semantica.llms import Groq
import os

# Build foundation from verified reference data
manager = SeedDataManager()
manager.register_source("taxonomy",  "json", "taxonomy.json")
manager.register_source("employees", "csv",  "employees.csv")
foundation_kg = manager.create_foundation_graph()

# Ingest and extract from new sources
llm = Groq(model="llama-3.3-70b-versatile", api_key=os.getenv("GROQ_API_KEY"))
ingestor = FileIngestor()
ner = NERExtractor(method="llm", llm_provider=llm)

sources = ingestor.ingest("news_articles/")
new_data = [ner.extract(s.content) for s in sources]

# Merge — seed data takes precedence for reference facts
final_kg = manager.integrate_seed_extracted(
    seed_graph=foundation_kg,
    extracted_data=new_data,
    strategy="seed_first"
)

Configuration

seed:
  sources:
    - name: "employees"
      type: "csv"
      path: "./data/employees.csv"
    - name: "taxonomy"
      type: "json"
      path: "./data/taxonomy.json"
  merge:
    strategy: "seed_first"
  validation:
    strict: true

Environment variable overrides:

export SEED_DATA_DIR=./data/seed
export SEED_MERGE_STRATEGY=seed_first
Load unstructured data alongside seed data. The target graph that seed data populates. Handle duplicates during seed-extracted merge. Incorporate seed loading as a pipeline step.