mirror of
https://github.com/semantica-agi/semantica.git
synced 2026-08-29 04:26:20 +00:00
* docs: replace Exported Classes import blocks with summary tables across all 25 modules * docs: add method/parameter tables to parse, ingest, ontology, normalize, triplet_store, change_management, conflicts, export, graph_store, provenance, and semantic_extract modules
8.4 KiB
8.4 KiB
title, description, icon
| title | description | icon |
|---|---|---|
| Utils Module | Shared utilities for logging, validation, error handling, progress tracking, and common operations. | wrench |
semantica.utils provides shared infrastructure used throughout Semantica. Most users won't call it directly, but its APIs are available when you need fine-grained control over logging, validation, progress tracking, or error handling.
Exported Classes
| Name | Type | Role |
|---|---|---|
setup_logging |
function | Configure root logger — level, format ("json" or "text") |
get_logger |
function | Get a named logger instance |
log_performance |
decorator | Logs function name, duration, and any exception |
validate_entity |
function | Validate entity dict structure — raises ValidationError on failure |
validate_config |
function | Validate config dict against schema — raises ValidationError on failure |
ProgressTracker |
class | Class-based progress tracker with ETA and step callbacks |
track_progress |
function | Wrap any iterable with a live progress bar |
clean_text |
function | Normalize whitespace and strip control characters |
hash_data |
function | Deterministic SHA-256 hash of any serializable object |
SemanticaError |
exception | Base exception for all Semantica errors |
ValidationError |
exception | Raised when input fails validation |
ProcessingError |
exception | Raised during extraction, graph build, or pipeline step |
What You Get
Structured logging with `@log_performance` decorator and quality metrics via environment variables. `validate_entity` and `validate_config` with a typed `ValidationError` carrying field and value context. `track_progress` wraps any iterable — auto-detects console vs Jupyter for the right renderer. `clean_text`, `hash_data`, `safe_filename`, and nested dict utilities used throughout the framework. `SemanticaError` → `ValidationError`, `ProcessingError` — typed exceptions for targeted recovery. `read_json_file` with `ProcessingError` on failure — no boilerplate try/except around JSON I/O.Logging
```python from semantica.utils import setup_logging, get_loggersetup_logging(level="INFO") # "DEBUG" | "INFO" | "WARNING" | "ERROR"
logger = get_logger(__name__)
```
@log_performance
def process_data(data):
logger.info(f"Processing {len(data)} items")
# Logs function name, duration, and any exception automatically
@log_execution_time
def expensive_step(data):
...
# Logs: "expensive_step completed in 2.34s"
```
Validation
from semantica.utils import validate_entity, validate_config, ValidationError
# Validate an entity dict
try:
validate_entity({"id": "1", "type": "PERSON", "text": "Alice"})
except ValidationError as e:
print(f"Invalid entity: {e.message}")
print(f" Field: {e.field}")
print(f" Value: {e.value}")
# Validate a configuration dict
try:
validate_config(config)
except ValidationError as e:
print(f"Invalid config: {e}")
| Function | Description |
|---|---|
validate_entity(data) |
Check entity dict has required fields and correct types |
validate_config(cfg) |
Check configuration dict against schema |
Progress Tracking
from semantica.utils import track_progress
# Wraps any iterable — auto-detects console vs Jupyter
for item in track_progress(items, desc="Processing documents"):
process(item)
Supports:
- Console — tqdm progress bar with ETA
- Jupyter — notebook-compatible widget (auto-detected)
- File — write progress to a log file
Helper Functions
from semantica.utils import clean_text, hash_data, safe_filename
# Normalize whitespace and strip control characters
clean = clean_text(" Hello World ") # → "Hello World"
# Deterministic SHA-256 hash of any JSON-serializable object
uid = hash_data({"key": "value"}) # → hex digest string
# Sanitize a string for use as a filename
fname = safe_filename("My File?.txt") # → "My_File_.txt"
Nested Dict Utilities
Helper functions for deep configuration access — used extensively inside Config and ConfigManager:
from semantica.utils import get_nested_value, set_nested_value, merge_dicts
config = {
"processing": {"batch_size": 32, "max_workers": 4},
"llm": {"provider": "groq", "model": "llama-3.3-70b-versatile"},
}
# Dot-notation read — returns default if key path is absent
batch = get_nested_value(config, "processing.batch_size", default=16)
# → 32
# Dot-notation write
set_nested_value(config, "processing.batch_size", 64)
# Deep merge — nested keys are merged recursively
base = {"a": {"x": 1, "y": 2}, "b": 3}
overrides = {"a": {"y": 99, "z": 4}, "c": 5}
merged = merge_dicts(base, overrides, deep=True)
# → {"a": {"x": 1, "y": 99, "z": 4}, "b": 3, "c": 5}
Exception Hierarchy
from semantica.utils import SemanticaError, ValidationError, ProcessingError
try:
run_pipeline(data)
except ValidationError as e:
# Input data did not pass schema validation
logger.error(f"Validation failed at field '{e.field}': {e.message}")
except ProcessingError as e:
# Failure during extraction or graph construction
logger.error(f"Processing failed at step {e.step}: {e}")
except SemanticaError as e:
# Catch-all for all Semantica framework errors
logger.error(f"Framework error: {e}")
| Exception | When Raised |
|---|---|
SemanticaError |
Base class — all framework errors inherit from this |
ValidationError |
Input data failed schema or type validation |
ProcessingError |
Failure during extraction, graph build, or pipeline step |
File Utilities
from semantica.utils import read_json_file
# Read and parse a JSON file — raises ProcessingError on failure
config = read_json_file("config.json")