mirror of
https://github.com/semantica-agi/semantica.git
synced 2026-08-29 04:26:20 +00:00
* docs: replace Exported Classes import blocks with summary tables across all 25 modules * docs: add method/parameter tables to parse, ingest, ontology, normalize, triplet_store, change_management, conflicts, export, graph_store, provenance, and semantic_extract modules
17 KiB
17 KiB
title, description, icon
| title | description | icon |
|---|---|---|
| Normalize Module | Text cleaning, entity canonicalization, date normalization, number conversion, language detection, and encoding repair — before extraction runs. | broom |
semantica.normalize standardizes raw data before extraction and graph construction. All normalizers expose both convenience functions (one-liners) and stateful class instances (full control over configuration and reuse).
Why Normalize Before Extraction
Unstructured data is inconsistent by nature. Without normalization, the same real-world entity appears as dozens of variants in your graph:
"Apple Inc.","Apple Computer Inc.","APPLE INC.","Apple, Inc."— four nodes, one company"Jan 1st, 2020","01/01/2020","2020-01-01"— three formats, one date"$1.2B","1,200,000,000","1.2 billion USD"— three strings, one number"Hello World"vs"Hello World"— a non-breaking space that breaks string matching
Normalization collapses these variants before any extractor, deduplicator, or graph builder sees the data — producing cleaner entities, fewer false duplicates, and more reliable downstream results.
Exported Classes
| Class | Role |
|---|---|
TextNormalizer |
Unicode forms (NFC/NFKC), whitespace collapse, HTML stripping, smart-quote and dash normalization |
EntityNormalizer |
Corporate suffixes, honorifics, alias resolution, and entity disambiguation |
DateNormalizer |
Parses any date string format → ISO 8601; handles relative dates and fiscal quarters |
NumberNormalizer |
"$1.2B" → 1200000000.0; unit conversion (km/h → m/s); currency parsing |
DataCleaner |
Remove duplicates, fill missing values, validate records against a schema |
LanguageDetector |
detect(text) → {language, confidence} using statistical n-gram models |
What You Get
Unicode forms, whitespace collapse, HTML stripping, smart-quote and dash replacement. Corporate suffix normalization, honorific removal, alias resolution, and disambiguation. Any date format → ISO 8601; relative dates, timezones, and date ranges. Currency, scientific notation, unit abbreviations, and percentages → float. 50+ languages with confidence scoring and batch detection. Encoding detection, UTF-8 conversion, BOM removal, and cp1252 repair. **v0.5.0 fix:** Encoding repair now handles cp1252 and latin-1 characters that previously caused crashes on Windows when processing documents with non-ASCII content.Recommended Processing Order
Broken bytes corrupt everything downstream. Always run this before anything else.```python
from semantica.normalize import EncodingHandler
handler = EncodingHandler()
utf8_text = handler.to_utf8(raw_bytes)
```
normalizer = TextNormalizer(strip_html=True, normalize_unicode=True)
clean_text = normalizer.normalize_text(utf8_text)
```
normalizer = EntityNormalizer()
canonical = normalizer.normalize_entity("Apple Computer Inc.", entity_type="Organization")
# → "Apple Inc."
```
date_norm = DateNormalizer(target_timezone="UTC")
num_norm = NumberNormalizer()
date = date_norm.normalize_date("Jan 1st, 2020") # → "2020-01-01"
num = num_norm.normalize_number("$1.2B") # → 1200000000.0
```
detector = LanguageDetector()
lang = detector.detect("Bonjour le monde")
# → {"language": "fr", "confidence": 0.98}
```
Convenience Functions
The fastest path — one import, one call:
from semantica.normalize import (
normalize_text, normalize_entity, normalize_date,
normalize_number, clean_data, detect_language, handle_encoding,
)
clean = normalize_text(" Hello, World!! \n\n") # → "Hello, World!!"
entity = normalize_entity("Apple Computer Inc.", entity_type="Organization") # → "Apple Inc."
date = normalize_date("Jan 1st, 2020") # → "2020-01-01"
num = normalize_number("$1.2B") # → 1200000000.0
lang = detect_language("Bonjour le monde") # → {"language": "fr", "confidence": 0.98}
TextNormalizer Constructor Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
lowercase |
bool |
False |
Convert to lowercase |
remove_punctuation |
bool |
False |
Strip all punctuation |
remove_extra_whitespace |
bool |
True |
Collapse tabs, newlines, non-breaking spaces |
strip_html |
bool |
False |
Remove HTML tags and decode entities |
normalize_unicode |
bool |
True |
Apply Unicode normal form |
fix_encoding |
bool |
True |
Repair cp1252/latin-1 mojibake |
form |
str |
"NFC" |
Unicode form: "NFC" / "NFD" / "NFKC" / "NFKD" |
Normalizers
Cleans raw text at the character and token level:```python
from semantica.normalize import TextNormalizer
normalizer = TextNormalizer(
lowercase=False,
remove_punctuation=False,
remove_extra_whitespace=True,
strip_html=True,
normalize_unicode=True,
fix_encoding=True,
form="NFC", # "NFC" | "NFD" | "NFKC" | "NFKD"
)
normalized = normalizer.normalize_text(raw_text)
```
| Parameter | Type | Default | Description |
| --------- | ---- | ------- | ----------- |
| `lowercase` | `bool` | `False` | Convert to lowercase — use for bag-of-words matching, not NER |
| `remove_punctuation` | `bool` | `False` | Strip all punctuation — use for keyword extraction only |
| `remove_extra_whitespace` | `bool` | `True` | Collapse tabs, newlines, non-breaking spaces into single spaces |
| `strip_html` | `bool` | `False` | Remove HTML tags and decode `&`, `<`, etc. |
| `normalize_unicode` | `bool` | `True` | Apply Unicode normal form |
| `fix_encoding` | `bool` | `True` | Repair common encoding mojibake (cp1252 / latin-1 → UTF-8) |
| `form` | `str` | `"NFC"` | Unicode normalization form: `"NFC"` / `"NFD"` / `"NFKC"` / `"NFKD"` |
**Unicode form guide:**
| Form | Use When |
| ---- | -------- |
| `NFC` | Default — best for storage and display |
| `NFKC` | Search indexing — normalises ligatures, fullwidth chars, and fractions |
| `NFD` | Stripping diacritics — split é → e + combining accent, then strip accents |
| `NFKD` | Same as NFD but also decomposes compatibility characters |
**Sub-normalizers for fine-grained control:**
```python
from semantica.normalize import UnicodeNormalizer, WhitespaceNormalizer, SpecialCharacterProcessor
unicode_norm = UnicodeNormalizer(form="NFC")
text = unicode_norm.normalize("café")
ws_norm = WhitespaceNormalizer()
text = ws_norm.normalize("Hello\t\t World\n\n") # → "Hello World"
processor = SpecialCharacterProcessor()
text = processor.process("'Hello'") # '' → '', -- → -
```
```python
from semantica.normalize import EntityNormalizer
normalizer = EntityNormalizer()
# Corporate name normalization
normalizer.normalize_entity("Apple Computer, Inc.", entity_type="Organization") # → "Apple Inc."
normalizer.normalize_entity("APPLE INC.", entity_type="Organization") # → "Apple Inc."
# Person name normalization
normalizer.normalize_entity("JOBS, STEVE", entity_type="Person") # → "Steve Jobs"
```
**Key behaviours:**
- Corporate suffix normalization handles: `Inc`, `Inc.`, `Incorporated`, `Ltd`, `Limited`, `Corp`, `Corporation`, `LLC`, `GmbH`, `PLC`, and 30+ more
- `entity_type="Person"` activates last-name-first reversal, honorific removal, and suffix stripping
**Sub-normalizers:**
```python
from semantica.normalize import AliasResolver, EntityDisambiguator, NameVariantHandler
resolver = AliasResolver(aliases={
"ML": "Machine Learning",
"NLP": "Natural Language Processing",
})
resolved = resolver.resolve("ML and NLP are subfields of AI")
disambiguator = EntityDisambiguator()
result = disambiguator.disambiguate(
"Apple",
context="Steve Jobs founded Apple in Cupertino in 1976",
)
# → {"entity": "Apple Inc.", "type": "Organization", "confidence": 0.96}
handler = NameVariantHandler()
canonical = handler.normalize("Dr. JOHN P. SMITH Jr.") # → "John P. Smith"
```
```python
from semantica.normalize import DateNormalizer
normalizer = DateNormalizer(
target_timezone="UTC",
output_format="%Y-%m-%d",
)
dates = [
"January 1st, 2020",
"01/01/2020",
"2020-01-01T00:00:00Z",
"yesterday",
"3 weeks ago",
"Q1 2024",
]
normalized = [normalizer.normalize_date(d) for d in dates]
```
**Sub-normalizers:**
```python
from semantica.normalize import TimeZoneNormalizer, RelativeDateProcessor, TemporalExpressionParser
from datetime import datetime
tz_norm = TimeZoneNormalizer(target_tz="UTC")
utc_dt = tz_norm.normalize("2024-01-01 09:00", source_tz="America/New_York")
# → datetime(2024, 1, 1, 14, 0, tzinfo=UTC)
processor = RelativeDateProcessor(reference_date=datetime(2025, 1, 15))
result = processor.process("3 days ago") # → datetime(2025, 1, 12)
result = processor.process("next quarter") # → {"start": "2025-04-01", "end": "2025-06-30"}
parser = TemporalExpressionParser()
result = parser.parse("from January 2020 to March 2021")
# → {"start": "2020-01-01", "end": "2021-03-31", "type": "range"}
result = parser.parse("Q2 2023")
# → {"start": "2023-04-01", "end": "2023-06-30", "type": "quarter"}
```
```python
from semantica.normalize import NumberNormalizer
normalizer = NumberNormalizer()
normalizer.normalize_number("$1,234.56") # → 1234.56
normalizer.normalize_number("€42K") # → 42000.0
normalizer.normalize_number("$1.2B") # → 1200000000.0
normalizer.normalize_number("3.14e-2") # → 0.0314
normalizer.normalize_number("42%") # → 0.42
normalizer.normalize_number("−7") # → -7.0 (minus sign, not hyphen)
```
**Unit and currency conversion:**
```python
from semantica.normalize import UnitConverter, CurrencyNormalizer
converter = UnitConverter()
result = converter.convert(100, from_unit="km/h", to_unit="m/s")
# → 27.78
categories = converter.list_categories()
# → ["length", "weight", "volume", "temperature", "speed", "area", "pressure", "energy"]
currency_norm = CurrencyNormalizer()
result = currency_norm.normalize("$42.50")
# → {"amount": 42.50, "currency": "USD", "raw": "$42.50"}
```
Identify the language of a text string. Used internally by the sentence splitter and chunker:
```python
from semantica.normalize import LanguageDetector
detector = LanguageDetector()
lang = detector.detect("Bonjour le monde")
# → {"language": "fr", "confidence": 0.98}
langs = detector.detect_top_n("This might be mixed", n=3)
# → [{"language": "en", "probability": 0.85}, ...]
results = detector.detect_batch(["Hello", "Hola", "Bonjour", "Ciao"])
supported = detector.list_supported_languages()
```
Supports 50+ languages: `en`, `de`, `fr`, `es`, `it`, `pt`, `nl`, `ru`, `zh`, `ja`, `ko`, `ar`, `hi`, `tr`, `pl`, `sv`, `da`, `no`, `fi`, and more.
### EncodingHandler
Detect and repair character encoding issues:
```python
from semantica.normalize import EncodingHandler
handler = EncodingHandler()
encoding = handler.detect_encoding(raw_bytes)
# → {"encoding": "windows-1252", "confidence": 0.73}
utf8_text = handler.to_utf8(raw_bytes)
clean = handler.remove_bom(text_with_bom)
repaired = handler.repair_encoding(garbled_text, source_encoding="cp1252")
```
**Key behaviours:**
- Encoding detection uses `chardet` internally — accuracy improves with longer input
- `to_utf8()` attempts cp1252 repair automatically when it detects mojibake patterns
- Always run `EncodingHandler` first — broken bytes cause cascading failures in every downstream normalizer
DataCleaner
Cleans structured record sets — useful before loading into a vector store or graph:
from semantica.normalize import DataCleaner, DataValidator
cleaner = DataCleaner()
deduped = cleaner.remove_duplicates(records, similarity_threshold=0.9)
filled = cleaner.fill_missing(
records,
strategy="mean", # "mean" | "median" | "mode" | "remove" | "constant"
constant_value=None,
)
validator = DataValidator()
result = validator.validate(records, schema={"name": str, "age": int, "active": bool})
print(f"Valid: {result.valid_count}")
print(f"Invalid: {result.error_count}")
Pipeline Integration
from semantica.pipeline import PipelineBuilder, ExecutionEngine
from semantica.ingest import FileIngestor
from semantica.normalize import TextNormalizer
from semantica.semantic_extract import NERExtractor
from semantica.llms import Groq
import os
llm = Groq(model="llama-3.3-70b-versatile", api_key=os.getenv("GROQ_API_KEY"))
ingestor = FileIngestor()
normalizer = TextNormalizer(strip_html=True, normalize_unicode=True)
extractor = NERExtractor(method="llm", llm_provider=llm)
builder = PipelineBuilder()
builder.add_step("ingest", "file_ingest", handler=ingestor.ingest)
builder.add_step("normalize", "text_normalize", handler=normalizer.normalize)
builder.add_step("extract", "ner_extract", handler=extractor.extract)
builder.connect_steps("ingest", "normalize")
builder.connect_steps("normalize", "extract")
pipeline = builder.build("normalize_pipeline")
result = ExecutionEngine().execute_pipeline(pipeline, data="data/documents/")