Replace plain markdown in every docs/reference/ file and docs/concepts.md with rich Mintlify JSX components — CardGroup, Steps, Tabs, AccordionGroup, Tip, Warning, Note, and CodeGroup — for a consistent, navigable, production-grade developer experience.
13 KiB
title, description, icon
| title | description | icon |
|---|---|---|
| Core Module | Framework orchestration, lifecycle management, configuration, and plugin system. | gear |
semantica.core is the coordination layer for the framework. For most tasks you should use individual modules directly (semantica.ingest, semantica.kg, etc.). Reach for Core when you need application-level lifecycle management, centralized configuration, or a plugin registry.
What You Get
Orchestration class for coordinating complex multi-module workflows and full KG construction pipelines. Unified config loading, merging, and validation with environment variable overrides. Startup/shutdown hooks with priority ordering and component health monitoring. Dynamic plugin discovery, registration, loading, and unloading. Register and dispatch custom orchestration methods by name. Live configuration state — dot-notation access, update, validate, and serialize. **Use individual modules directly** for the vast majority of use cases. Use the `Semantica` orchestration class only when you need application-level lifecycle management or a plugin system.Quick Start
```python from semantica.core import ConfigManagermanager = ConfigManager()
config = manager.load_from_file("config.yaml")
# Override one key at runtime
config.set("processing.batch_size", 64)
```
framework = Semantica(config=config)
framework.initialize()
status = framework.get_status()
print(f"State: {status['state']}") # → "READY"
```
Semantica (Orchestration)
High-level entry point that coordinates the full KG construction pipeline:
from semantica.core import Semantica, ConfigManager
config_manager = ConfigManager()
config = config_manager.load_from_file("config.yaml")
framework = Semantica(config=config)
framework.initialize()
try:
result = framework.build_knowledge_base(
sources=["doc1.pdf", "doc2.docx"],
embeddings=True,
graph=True,
)
status = framework.get_status()
print(f"State: {status['state']}")
finally:
framework.shutdown(graceful=True)
| Method | Description |
|---|---|
initialize() |
Initialize all framework components |
build_knowledge_base(sources, **kwargs) |
Orchestrate full KG construction pipeline |
run_pipeline(pipeline, data) |
Execute an existing Pipeline instance |
get_status() |
Return system health and current state |
shutdown(graceful=True) |
Graceful shutdown — waits for in-flight operations |
ConfigManager
Centralized config loading with deep-merge and environment variable overrides:
from semantica.core import ConfigManager
manager = ConfigManager()
config = manager.load_from_file("config.yaml")
# Merge base config with environment-specific overrides
merged = manager.merge_configs(
manager.load_from_file("base.yaml"),
manager.load_from_file("prod.yaml"),
)
# Nested dot-notation access
batch_size = config.get("processing.batch_size", default=16)
config.set("processing.batch_size", 64)
config.update({"quality": {"min_confidence": 0.75}}, merge=True)
config.validate()
config_dict = config.to_dict()
Config Section Reference
| Section | Key Fields | Description |
|---|---|---|
llm_provider |
name, model, api_key, base_url |
LLM used for extraction and reasoning |
embedding_model |
provider, model, dimension, device |
Embedding provider and model |
vector_store |
backend, dimension, index_type |
Vector storage backend |
graph_db |
backend, uri, user, password |
Graph database connection |
processing |
batch_size, max_workers, chunk_size |
Parallelism and batching |
pipeline |
retry_max, backoff, failure_strategy |
Pipeline retry and failure policy |
logging |
level, format, file |
Logging configuration |
quality |
min_confidence, dedup_threshold |
Quality thresholds |
security |
redact_pii, allowed_domains |
Security and compliance settings |
custom |
any key | User-defined extension settings |
Environment Variable Overrides
Any config key can be overridden with a SEMANTICA_ prefix using double underscores for nesting:
export SEMANTICA_PROCESSING__BATCH_SIZE=64
export SEMANTICA_LLM_PROVIDER__MODEL=gpt-4o
export SEMANTICA_LOGGING__LEVEL=DEBUG
export SEMANTICA_QUALITY__MIN_CONFIDENCE=0.8
LifecycleManager
Manages framework state with a defined state machine and ordered startup/shutdown hooks:
from semantica.core import LifecycleManager
manager = LifecycleManager()
def init_db():
print("Initializing database...")
def cleanup_db():
print("Closing database connections...")
# Lower priority values run first during startup
# Higher priority values run first during shutdown
manager.register_startup_hook(init_db, priority=10)
manager.register_shutdown_hook(cleanup_db, priority=10)
manager.startup()
# Component health monitoring
class DatabaseComponent:
def health_check(self):
return {"healthy": True, "message": "Connected"}
manager.register_component("database", DatabaseComponent())
summary = manager.get_health_summary()
# → {"database": {"healthy": True, "message": "Connected"}, ...}
manager.shutdown(graceful=True)
PluginRegistry
Register custom components that participate in the full pipeline:
from semantica.core import PluginRegistry
class MyPlugin:
def initialize(self):
print("Plugin initialized")
def execute(self, data):
return {"processed": True}
registry = PluginRegistry(plugin_paths=["./plugins"])
registry.register_plugin(
"my_plugin", MyPlugin,
version="1.0.0",
description="Custom domain extractor",
author="team@example.com",
capabilities=["extract"],
)
plugin = registry.load_plugin("my_plugin", api_key="xxx")
result = plugin.execute("sample data")
# Inspect registered plugins
for info in registry.list_plugins():
print(f"{info['name']} v{info['version']} — {info['description']}")
# Unload when done
registry.unload_plugin("my_plugin")
MethodRegistry
Register custom orchestration methods and dispatch them by name:
from semantica.core import method_registry
def fast_kb_builder(sources, **kwargs):
# Custom logic — skip embeddings for speed
...
method_registry.register("knowledge_base", "fast", fast_kb_builder)
from semantica.core.methods import build_knowledge_base
result = build_knowledge_base(sources=["doc.pdf"], method="fast")
Schemas
from semantica.core import SystemState
SystemState.UNINITIALIZED # → startup() →
SystemState.INITIALIZING # → hooks complete →
SystemState.READY # → first operation →
SystemState.RUNNING # → shutdown() →
SystemState.STOPPING # → hooks complete →
SystemState.STOPPED
# Any unhandled exception during startup/shutdown →
SystemState.ERROR
Check current state at any time:
state = manager.get_state()
if manager.is_ready():
result = framework.build_knowledge_base(sources)
@dataclass
class HealthStatus:
component: str # component name
healthy: bool # True = operational
message: str # human-readable status
timestamp: datetime # time of last check
details: Dict[str, Any] # component-specific diagnostics
health = manager.health_check()
for name, status in health.items():
icon = "✓" if status.healthy else "✗"
print(f"{icon} {name}: {status.message}")
@dataclass
class PluginInfo:
name: str
version: str
plugin_class: Type
description: str
author: str
dependencies: List[str] # pip package names required
capabilities: List[str] # e.g. ["ingest", "extract"]
metadata: Dict[str, Any]
@dataclass
class LoadedPlugin:
info: PluginInfo
instance: Any # the live plugin object
config: Dict[str, Any] # config passed at load time
loaded_at: datetime
registry = PluginRegistry()
registry.register_plugin("my_plugin", MyPlugin, version="1.0.0")
if registry.is_plugin_loaded("my_plugin"):
plugin = registry.get_loaded_plugin("my_plugin")
else:
plugin = registry.load_plugin("my_plugin")
details = registry.get_plugin_info("my_plugin")
Complete Configuration Example
# config.yaml
llm_provider:
name: groq
model: llama-3.3-70b-versatile
api_key: ${GROQ_API_KEY}
embedding_model:
provider: sentence-transformers
model: all-mpnet-base-v2
dimension: 768
device: cpu # "cpu" | "cuda" | "mps"
vector_store:
backend: faiss
dimension: 768
index_type: hnsw # "flat" | "ivf" | "hnsw" | "pq"
graph_db:
backend: neo4j
uri: bolt://localhost:7687
user: neo4j
password: ${NEO4J_PASSWORD}
processing:
batch_size: 32
max_workers: 4
chunk_size: 512
pipeline:
retry_max: 3
backoff: exponential # "fixed" | "linear" | "exponential"
failure_strategy: skip # "skip" | "stop" | "retry"
quality:
min_confidence: 0.7
dedup_threshold: 0.85
logging:
level: INFO # DEBUG | INFO | WARNING | ERROR
format: "%(asctime)s %(name)s %(levelname)s %(message)s"
Load and use:
from semantica.core import ConfigManager, Semantica
manager = ConfigManager()
config = manager.load_from_file("config.yaml")
config.set("processing.batch_size", 64) # runtime override
framework = Semantica(config=config)
framework.initialize()
result = framework.build_knowledge_base(["doc1.pdf", "doc2.docx"])
framework.shutdown()