Files
semantica/docs/reference/core.md
T
KaifAhmad1 9113ef3428 docs: premium overhaul of all reference pages and core docs
- Rewrote all 26 reference module pages: removed blockquote taglines and
  horizontal rule separators, added "What You Get" bullet summaries,
  added constructor/method parameter tables, expanded thin files
  (graph_store, triplet_store, visualization, provenance) with full API
  coverage, added backend comparison tables and real-world usage patterns
- Renamed Modules tab from "API Reference" and group from "Context &
  Knowledge" to "Context & Intelligence" in docs.json
- Fixed logo: copied "Semantica Logo.png" to web-safe semantica-logo.png
  and updated all 4 references in docs.json
- Improved core docs (index, modules, concepts, quickstart, installation,
  getting-started) with better fonts, bullet points, and complete module
  listings (mcp_server, evals, core, utils previously missing)
- Rewrote community pages (community, community-projects, contributing-guide,
  use-cases, architecture, faq, learning-more, glossary) with heading
  hierarchy fixes, expanded definitions, and better structure
- Fixed markdown linter warnings: MD036 bold-as-heading, MD001 heading
  skips, MD040 missing code fence language, MD032 blank lines around lists
2026-05-23 13:10:09 +05:30

5.5 KiB

title, description, icon
title description icon
Core Module Framework orchestration, lifecycle management, configuration, and plugin system. gear

semantica.core is the coordination layer for the framework. For most tasks you should use individual modules directly (semantica.ingest, semantica.kg, etc.). Reach for Core when you need application-level lifecycle management, centralized configuration, or a plugin registry.

What You Get

  • Semantica — orchestration class for coordinating complex multi-module workflows
  • ConfigManager — unified config loading, merging, and validation with environment variable overrides
  • LifecycleManager — startup/shutdown hooks and component health monitoring
  • PluginRegistry — dynamic plugin discovery, registration, and loading
  • MethodRegistry — register and dispatch custom orchestration methods
**Use individual modules directly** for the vast majority of use cases. Use the `Semantica` orchestration class only when you need application-level lifecycle management or a plugin system.

Semantica (Orchestration)

High-level entry point that coordinates the full KG construction pipeline:

from semantica.core import Semantica, ConfigManager

config_manager = ConfigManager()
config = config_manager.load_from_file("config.yaml")

framework = Semantica(config=config)
framework.initialize()

try:
    result = framework.build_knowledge_base(
        sources=["doc1.pdf", "doc2.docx"],
        embeddings=True,
        graph=True,
    )
    status = framework.get_status()
    print(f"State: {status['state']}")
finally:
    framework.shutdown(graceful=True)

Core Methods

Method Description
initialize() Initialize all framework components
build_knowledge_base(sources, **kwargs) Orchestrate full KG construction pipeline
run_pipeline(pipeline, data) Execute an existing Pipeline instance
get_status() Return system health and current state
shutdown(graceful=True) Graceful shutdown — waits for in-flight operations

ConfigManager

Centralized config loading with deep-merge and environment variable overrides:

from semantica.core import ConfigManager

manager = ConfigManager()
config = manager.load_from_file("config.yaml")

# Merge base config with environment-specific overrides
merged = manager.merge_configs(
    manager.load_from_file("base.yaml"),
    manager.load_from_file("prod.yaml"),
)

# Nested key access with dot notation
batch_size = config.get("processing.batch_size", default=16)
config.set("processing.batch_size", 64)
config.validate()

YAML Configuration

llm_provider:
  name: openai
  model: gpt-4o
  api_key: ${OPENAI_API_KEY}

processing:
  batch_size: 32
  max_workers: 4

quality:
  min_confidence: 0.7

logging:
  level: INFO

Environment variable overrides (prefix SEMANTICA_):

export SEMANTICA_PROCESSING_BATCH_SIZE=64
export SEMANTICA_LOG_LEVEL=DEBUG

LifecycleManager

Manages framework state with a defined state machine and ordered startup/shutdown hooks:

State machine: UNINITIALIZEDINITIALIZINGREADYRUNNINGSTOPPINGSTOPPED

from semantica.core import LifecycleManager

manager = LifecycleManager()

def init_db():
    print("Initializing database...")

def cleanup_db():
    print("Closing database connections...")

# Lower priority values run first during startup
# Higher priority values run first during shutdown
manager.register_startup_hook(init_db,     priority=10)
manager.register_shutdown_hook(cleanup_db, priority=10)

manager.startup()

# Component health monitoring
class DatabaseComponent:
    def health_check(self):
        return {"healthy": True, "message": "Connected"}

manager.register_component("database", DatabaseComponent())
summary = manager.get_health_summary()
# → {"database": {"healthy": True, "message": "Connected"}, ...}

manager.shutdown(graceful=True)

PluginRegistry

Register custom components that participate in the full pipeline — provenance tracking, retry policies, and parallel execution included:

from semantica.core import PluginRegistry

class MyPlugin:
    def initialize(self):
        print("Plugin initialized")

    def execute(self, data):
        return {"processed": True}

registry = PluginRegistry(plugin_paths=["./plugins"])
registry.register_plugin("my_plugin", MyPlugin, version="1.0.0")

plugin = registry.load_plugin("my_plugin", api_key="xxx")
result = plugin.execute("sample data")

for info in registry.list_plugins():
    print(f"{info['name']}: {info['version']}")

MethodRegistry

Register custom orchestration methods and dispatch them by name:

from semantica.core import method_registry

def fast_kb_builder(sources, **kwargs):
    # Custom logic — skip embeddings for speed
    ...

method_registry.register("knowledge_base", "fast", fast_kb_builder)

from semantica.core.methods import build_knowledge_base
result = build_knowledge_base(sources=["doc.pdf"], method="fast")
Pipeline execution and step orchestration. Shared utilities used by Core internally. Learn the basics before using Core. Configure LLM providers via ConfigManager.