Files
semantica/docs/reference/normalize.md
T
Mohd Kaif 5d70d0c10d docs: replace Exported Classes import blocks with summary tables (all 25 modules) (#567)
* docs: replace Exported Classes import blocks with summary tables across all 25 modules

* docs: add method/parameter tables to parse, ingest, ontology, normalize, triplet_store, change_management, conflicts, export, graph_store, provenance, and semantic_extract modules
2026-05-24 15:49:58 +05:30

17 KiB
Raw Blame History

title, description, icon
title description icon
Normalize Module Text cleaning, entity canonicalization, date normalization, number conversion, language detection, and encoding repair — before extraction runs. broom

semantica.normalize standardizes raw data before extraction and graph construction. All normalizers expose both convenience functions (one-liners) and stateful class instances (full control over configuration and reuse).

Why Normalize Before Extraction

Unstructured data is inconsistent by nature. Without normalization, the same real-world entity appears as dozens of variants in your graph:

  • "Apple Inc.", "Apple Computer Inc.", "APPLE INC.", "Apple, Inc." — four nodes, one company
  • "Jan 1st, 2020", "01/01/2020", "2020-01-01" — three formats, one date
  • "$1.2B", "1,200,000,000", "1.2 billion USD" — three strings, one number
  • "Hello World" vs "Hello World" — a non-breaking space that breaks string matching

Normalization collapses these variants before any extractor, deduplicator, or graph builder sees the data — producing cleaner entities, fewer false duplicates, and more reliable downstream results.

Exported Classes

Class Role
TextNormalizer Unicode forms (NFC/NFKC), whitespace collapse, HTML stripping, smart-quote and dash normalization
EntityNormalizer Corporate suffixes, honorifics, alias resolution, and entity disambiguation
DateNormalizer Parses any date string format → ISO 8601; handles relative dates and fiscal quarters
NumberNormalizer "$1.2B"1200000000.0; unit conversion (km/hm/s); currency parsing
DataCleaner Remove duplicates, fill missing values, validate records against a schema
LanguageDetector detect(text){language, confidence} using statistical n-gram models

What You Get

Unicode forms, whitespace collapse, HTML stripping, smart-quote and dash replacement. Corporate suffix normalization, honorific removal, alias resolution, and disambiguation. Any date format → ISO 8601; relative dates, timezones, and date ranges. Currency, scientific notation, unit abbreviations, and percentages → float. 50+ languages with confidence scoring and batch detection. Encoding detection, UTF-8 conversion, BOM removal, and cp1252 repair. **v0.5.0 fix:** Encoding repair now handles cp1252 and latin-1 characters that previously caused crashes on Windows when processing documents with non-ASCII content. Broken bytes corrupt everything downstream. Always run this before anything else.
```python
from semantica.normalize import EncodingHandler

handler   = EncodingHandler()
utf8_text = handler.to_utf8(raw_bytes)
```
```python from semantica.normalize import TextNormalizer
normalizer  = TextNormalizer(strip_html=True, normalize_unicode=True)
clean_text  = normalizer.normalize_text(utf8_text)
```
```python from semantica.normalize import EntityNormalizer
normalizer = EntityNormalizer()
canonical  = normalizer.normalize_entity("Apple Computer Inc.", entity_type="Organization")
# → "Apple Inc."
```
```python from semantica.normalize import DateNormalizer, NumberNormalizer
date_norm = DateNormalizer(target_timezone="UTC")
num_norm  = NumberNormalizer()

date = date_norm.normalize_date("Jan 1st, 2020")   # → "2020-01-01"
num  = num_norm.normalize_number("$1.2B")          # → 1200000000.0
```
```python from semantica.normalize import LanguageDetector
detector = LanguageDetector()
lang     = detector.detect("Bonjour le monde")
# → {"language": "fr", "confidence": 0.98}
```

Convenience Functions

The fastest path — one import, one call:

from semantica.normalize import (
    normalize_text, normalize_entity, normalize_date,
    normalize_number, clean_data, detect_language, handle_encoding,
)

clean  = normalize_text("  Hello,   World!!  \n\n")         # → "Hello, World!!"
entity = normalize_entity("Apple Computer Inc.", entity_type="Organization")  # → "Apple Inc."
date   = normalize_date("Jan 1st, 2020")                    # → "2020-01-01"
num    = normalize_number("$1.2B")                          # → 1200000000.0
lang   = detect_language("Bonjour le monde")                # → {"language": "fr", "confidence": 0.98}

TextNormalizer Constructor Parameters

Parameter Type Default Description
lowercase bool False Convert to lowercase
remove_punctuation bool False Strip all punctuation
remove_extra_whitespace bool True Collapse tabs, newlines, non-breaking spaces
strip_html bool False Remove HTML tags and decode entities
normalize_unicode bool True Apply Unicode normal form
fix_encoding bool True Repair cp1252/latin-1 mojibake
form str "NFC" Unicode form: "NFC" / "NFD" / "NFKC" / "NFKD"

Normalizers

Cleans raw text at the character and token level:
```python
from semantica.normalize import TextNormalizer

normalizer = TextNormalizer(
    lowercase=False,
    remove_punctuation=False,
    remove_extra_whitespace=True,
    strip_html=True,
    normalize_unicode=True,
    fix_encoding=True,
    form="NFC",              # "NFC" | "NFD" | "NFKC" | "NFKD"
)

normalized = normalizer.normalize_text(raw_text)
```

| Parameter | Type | Default | Description |
| --------- | ---- | ------- | ----------- |
| `lowercase` | `bool` | `False` | Convert to lowercase — use for bag-of-words matching, not NER |
| `remove_punctuation` | `bool` | `False` | Strip all punctuation — use for keyword extraction only |
| `remove_extra_whitespace` | `bool` | `True` | Collapse tabs, newlines, non-breaking spaces into single spaces |
| `strip_html` | `bool` | `False` | Remove HTML tags and decode `&`, `<`, etc. |
| `normalize_unicode` | `bool` | `True` | Apply Unicode normal form |
| `fix_encoding` | `bool` | `True` | Repair common encoding mojibake (cp1252 / latin-1 → UTF-8) |
| `form` | `str` | `"NFC"` | Unicode normalization form: `"NFC"` / `"NFD"` / `"NFKC"` / `"NFKD"` |

**Unicode form guide:**

| Form | Use When |
| ---- | -------- |
| `NFC` | Default — best for storage and display |
| `NFKC` | Search indexing — normalises ligatures, fullwidth chars, and fractions |
| `NFD` | Stripping diacritics — split é → e + combining accent, then strip accents |
| `NFKD` | Same as NFD but also decomposes compatibility characters |

**Sub-normalizers for fine-grained control:**

```python
from semantica.normalize import UnicodeNormalizer, WhitespaceNormalizer, SpecialCharacterProcessor

unicode_norm = UnicodeNormalizer(form="NFC")
text = unicode_norm.normalize("café")

ws_norm = WhitespaceNormalizer()
text    = ws_norm.normalize("Hello\t\t World\n\n")  # → "Hello World"

processor = SpecialCharacterProcessor()
text      = processor.process("'Hello'")  # '' → '', -- → -
```
Canonicalises entity name variants — corporate suffixes, honorifics, case, and punctuation:
```python
from semantica.normalize import EntityNormalizer

normalizer = EntityNormalizer()

# Corporate name normalization
normalizer.normalize_entity("Apple Computer, Inc.", entity_type="Organization")  # → "Apple Inc."
normalizer.normalize_entity("APPLE INC.",           entity_type="Organization")  # → "Apple Inc."

# Person name normalization
normalizer.normalize_entity("JOBS, STEVE", entity_type="Person")  # → "Steve Jobs"
```

**Key behaviours:**
- Corporate suffix normalization handles: `Inc`, `Inc.`, `Incorporated`, `Ltd`, `Limited`, `Corp`, `Corporation`, `LLC`, `GmbH`, `PLC`, and 30+ more
- `entity_type="Person"` activates last-name-first reversal, honorific removal, and suffix stripping

**Sub-normalizers:**

```python
from semantica.normalize import AliasResolver, EntityDisambiguator, NameVariantHandler

resolver = AliasResolver(aliases={
    "ML":  "Machine Learning",
    "NLP": "Natural Language Processing",
})
resolved = resolver.resolve("ML and NLP are subfields of AI")

disambiguator = EntityDisambiguator()
result = disambiguator.disambiguate(
    "Apple",
    context="Steve Jobs founded Apple in Cupertino in 1976",
)
# → {"entity": "Apple Inc.", "type": "Organization", "confidence": 0.96}

handler   = NameVariantHandler()
canonical = handler.normalize("Dr. JOHN P. SMITH Jr.")  # → "John P. Smith"
```
Parses any date format and outputs ISO 8601 strings. Handles relative expressions, timezones, and date ranges:
```python
from semantica.normalize import DateNormalizer

normalizer = DateNormalizer(
    target_timezone="UTC",
    output_format="%Y-%m-%d",
)

dates = [
    "January 1st, 2020",
    "01/01/2020",
    "2020-01-01T00:00:00Z",
    "yesterday",
    "3 weeks ago",
    "Q1 2024",
]
normalized = [normalizer.normalize_date(d) for d in dates]
```

**Sub-normalizers:**

```python
from semantica.normalize import TimeZoneNormalizer, RelativeDateProcessor, TemporalExpressionParser
from datetime import datetime

tz_norm = TimeZoneNormalizer(target_tz="UTC")
utc_dt  = tz_norm.normalize("2024-01-01 09:00", source_tz="America/New_York")
# → datetime(2024, 1, 1, 14, 0, tzinfo=UTC)

processor = RelativeDateProcessor(reference_date=datetime(2025, 1, 15))
result    = processor.process("3 days ago")    # → datetime(2025, 1, 12)
result    = processor.process("next quarter")  # → {"start": "2025-04-01", "end": "2025-06-30"}

parser = TemporalExpressionParser()
result = parser.parse("from January 2020 to March 2021")
# → {"start": "2020-01-01", "end": "2021-03-31", "type": "range"}

result = parser.parse("Q2 2023")
# → {"start": "2023-04-01", "end": "2023-06-30", "type": "quarter"}
```
Converts number strings with units, currencies, and abbreviations to `float`:
```python
from semantica.normalize import NumberNormalizer

normalizer = NumberNormalizer()

normalizer.normalize_number("$1,234.56")   # → 1234.56
normalizer.normalize_number("€42K")         # → 42000.0
normalizer.normalize_number("$1.2B")        # → 1200000000.0
normalizer.normalize_number("3.14e-2")      # → 0.0314
normalizer.normalize_number("42%")          # → 0.42
normalizer.normalize_number("7")           # → -7.0  (minus sign, not hyphen)
```

**Unit and currency conversion:**

```python
from semantica.normalize import UnitConverter, CurrencyNormalizer

converter = UnitConverter()
result    = converter.convert(100, from_unit="km/h", to_unit="m/s")
# → 27.78

categories = converter.list_categories()
# → ["length", "weight", "volume", "temperature", "speed", "area", "pressure", "energy"]

currency_norm = CurrencyNormalizer()
result = currency_norm.normalize("$42.50")
# → {"amount": 42.50, "currency": "USD", "raw": "$42.50"}
```
### LanguageDetector
Identify the language of a text string. Used internally by the sentence splitter and chunker:

```python
from semantica.normalize import LanguageDetector

detector = LanguageDetector()

lang  = detector.detect("Bonjour le monde")
# → {"language": "fr", "confidence": 0.98}

langs = detector.detect_top_n("This might be mixed", n=3)
# → [{"language": "en", "probability": 0.85}, ...]

results   = detector.detect_batch(["Hello", "Hola", "Bonjour", "Ciao"])
supported = detector.list_supported_languages()
```

Supports 50+ languages: `en`, `de`, `fr`, `es`, `it`, `pt`, `nl`, `ru`, `zh`, `ja`, `ko`, `ar`, `hi`, `tr`, `pl`, `sv`, `da`, `no`, `fi`, and more.

### EncodingHandler

Detect and repair character encoding issues:

```python
from semantica.normalize import EncodingHandler

handler = EncodingHandler()

encoding  = handler.detect_encoding(raw_bytes)
# → {"encoding": "windows-1252", "confidence": 0.73}

utf8_text = handler.to_utf8(raw_bytes)
clean     = handler.remove_bom(text_with_bom)
repaired  = handler.repair_encoding(garbled_text, source_encoding="cp1252")
```

**Key behaviours:**
- Encoding detection uses `chardet` internally — accuracy improves with longer input
- `to_utf8()` attempts cp1252 repair automatically when it detects mojibake patterns
- Always run `EncodingHandler` first — broken bytes cause cascading failures in every downstream normalizer

DataCleaner

Cleans structured record sets — useful before loading into a vector store or graph:

from semantica.normalize import DataCleaner, DataValidator

cleaner = DataCleaner()

deduped = cleaner.remove_duplicates(records, similarity_threshold=0.9)

filled = cleaner.fill_missing(
    records,
    strategy="mean",      # "mean" | "median" | "mode" | "remove" | "constant"
    constant_value=None,
)

validator = DataValidator()
result = validator.validate(records, schema={"name": str, "age": int, "active": bool})
print(f"Valid:   {result.valid_count}")
print(f"Invalid: {result.error_count}")

Pipeline Integration

from semantica.pipeline import PipelineBuilder, ExecutionEngine
from semantica.ingest import FileIngestor
from semantica.normalize import TextNormalizer
from semantica.semantic_extract import NERExtractor
from semantica.llms import Groq
import os

llm       = Groq(model="llama-3.3-70b-versatile", api_key=os.getenv("GROQ_API_KEY"))
ingestor  = FileIngestor()
normalizer = TextNormalizer(strip_html=True, normalize_unicode=True)
extractor = NERExtractor(method="llm", llm_provider=llm)

builder = PipelineBuilder()
builder.add_step("ingest",    "file_ingest",    handler=ingestor.ingest)
builder.add_step("normalize", "text_normalize", handler=normalizer.normalize)
builder.add_step("extract",   "ner_extract",    handler=extractor.extract)
builder.connect_steps("ingest",    "normalize")
builder.connect_steps("normalize", "extract")

pipeline = builder.build("normalize_pipeline")
result   = ExecutionEngine().execute_pipeline(pipeline, data="data/documents/")

Tips and Common Pitfalls

**Run encoding repair before anything else.** A single cp1252 character in a UTF-8 stream silently corrupts the surrounding text. Run `EncodingHandler` or set `fix_encoding=True` on `TextNormalizer` first. **Don't lowercase before NER.** `normalize_text(lowercase=True)` before entity extraction destroys capitalization signals that NER relies on. Apply case normalization only after extraction if needed. **AliasResolver is order-sensitive.** If you register overlapping aliases (`"ML"` and `"ML model"`), the longer match wins. Sort aliases by length descending for predictable behaviour. **DateNormalizer and timezone.** Without `target_timezone`, dates without timezone information are returned as naive datetime strings. In regulated pipelines (HIPAA, SOX), always set `target_timezone="UTC"` for unambiguous timestamps. **`DataCleaner.remove_duplicates()` is not the same as `DuplicateDetector`.** `DataCleaner` operates on flat records (dicts/rows) by field-level Jaccard similarity. `DuplicateDetector` in the Deduplication module operates on graph entities with embedding-based matching. Use the latter for entity resolution. Parse documents before normalization. Chunk normalized text for embedding. Resolve duplicate entities after normalization. Include normalization as a named pipeline step.