Files
semantica/docs/reference/semantic_extract.md
T
KaifAhmad1 9113ef3428 docs: premium overhaul of all reference pages and core docs
- Rewrote all 26 reference module pages: removed blockquote taglines and
  horizontal rule separators, added "What You Get" bullet summaries,
  added constructor/method parameter tables, expanded thin files
  (graph_store, triplet_store, visualization, provenance) with full API
  coverage, added backend comparison tables and real-world usage patterns
- Renamed Modules tab from "API Reference" and group from "Context &
  Knowledge" to "Context & Intelligence" in docs.json
- Fixed logo: copied "Semantica Logo.png" to web-safe semantica-logo.png
  and updated all 4 references in docs.json
- Improved core docs (index, modules, concepts, quickstart, installation,
  getting-started) with better fonts, bullet points, and complete module
  listings (mcp_server, evals, core, utils previously missing)
- Rewrote community pages (community, community-projects, contributing-guide,
  use-cases, architecture, faq, learning-more, glossary) with heading
  hierarchy fixes, expanded definitions, and better structure
- Fixed markdown linter warnings: MD036 bold-as-heading, MD001 heading
  skips, MD040 missing code fence language, MD032 blank lines around lists
2026-05-23 13:10:09 +05:30

5.8 KiB

title, description, icon
title description icon
Semantic Extract Module Named entity recognition, relation extraction, event detection, and triplet generation. magnifying-glass-chart

semantica.semantic_extract extracts structured information from unstructured text — the foundation of every knowledge graph in Semantica. All extractors support three modes: pattern-based (no API key), ML-based, and LLM-based.

What You Get

  • NERExtractor — named entity recognition: Person, Organization, Location, Date, and custom types
  • RelationExtractor — typed semantic relationships between entities (founded_by, located_in, etc.)
  • TripletExtractor — direct (subject, predicate, object) triplet generation for RDF-ready output
  • EventExtractor — event detection with participants, temporal context, and confidence scores
  • CoreferenceResolver — resolve "Apple" and "the company" to the same entity across a document

NERExtractor

from semantica.semantic_extract import NERExtractor
from semantica.llms import Groq
import os

# Pattern-based — fast, no API key, good for standard entity types
ner = NERExtractor(method="pattern")
entities = ner.extract("Apple Inc. was founded by Steve Jobs in Cupertino.")

# ML-based — higher accuracy, no API cost
ner = NERExtractor(method="ml", model="dslim/bert-large-NER")
entities = ner.extract(text)

# LLM-based — best accuracy, handles complex schemas and custom types
llm = Groq(model="llama-3.3-70b-versatile", api_key=os.getenv("GROQ_API_KEY"))
ner = NERExtractor(method="llm", llm_provider=llm, max_retries=3)
entities = ner.extract(text)

Output format:

[
    {"text": "Apple Inc.",  "type": "ORGANIZATION", "confidence": 0.98, "start": 0,  "end": 10},
    {"text": "Steve Jobs",  "type": "PERSON",       "confidence": 0.99, "start": 27, "end": 37},
    {"text": "Cupertino",   "type": "LOCATION",     "confidence": 0.97, "start": 41, "end": 50}
]

Custom Entity Types

ner = NERExtractor(
    method="pattern",
    custom_entities={
        "DRUG": ["aspirin", "ibuprofen", "metformin"],
        "GENE": ["BRCA1", "TP53", "EGFR"]
    }
)
**v0.5.0 fix:** `NERExtractor(method="llm")` no longer silently falls back to pattern extraction on custom gateways. The `response_format=json_object` parameter is now conditionally omitted for incompatible gateways, with a plain `generate()` + JSON parsing fallback applied automatically.

RelationExtractor

from semantica.semantic_extract import RelationExtractor

rel = RelationExtractor(method="llm", llm_provider=llm, max_retries=3)
relationships = rel.extract(text, entities=entities)

Output format:

[
    {"subject": "Steve Jobs", "predicate": "founded",    "object": "Apple Inc.", "confidence": 0.92},
    {"subject": "Apple Inc.", "predicate": "located_in", "object": "Cupertino",  "confidence": 0.89}
]

Available methods: "rule" (pattern-based), "ml" (REBEL model), "llm".

TripletExtractor

Generate RDF-ready (subject, predicate, object) triplets directly from text:

from semantica.semantic_extract import TripletExtractor

trip = TripletExtractor(method="llm", llm_provider=llm)
triplets = trip.extract(text)
# → [{"subject": "Steve Jobs", "predicate": "founded", "object": "Apple Inc.", ...}]

Triplets are suitable for loading directly into a triplet store or knowledge graph.

EventExtractor

Detect events with participants and temporal context:

from semantica.semantic_extract import EventExtractor

extractor = EventExtractor(method="llm", llm_provider=llm)
events = extractor.extract(text)

Output includes: event type, participants (with roles), temporal information, location, and confidence score.

CoreferenceResolver

Resolve pronoun and alias references to canonical entities before extraction:

from semantica.semantic_extract import CoreferenceResolver

resolver = CoreferenceResolver()
resolved_text = resolver.resolve(
    "Apple Inc. was founded in 1976. The company is headquartered in Cupertino."
)
# "Apple Inc." replaces "The company" for consistent downstream extraction

Batch Processing

All extractors support batch input for efficient large-scale processing:

texts = ["Text 1...", "Text 2...", "Text 3..."]

ner = NERExtractor(method="llm", llm_provider=llm)
batch_results = ner.extract_batch(texts, batch_size=10)

Using All Extractors Together

The standard extraction pipeline — entities → relationships → triplets:

from semantica.semantic_extract import NERExtractor, RelationExtractor, TripletExtractor
from semantica.llms import Groq
import os

llm = Groq(model="llama-3.3-70b-versatile", api_key=os.getenv("GROQ_API_KEY"))

ner  = NERExtractor(method="llm",      llm_provider=llm, max_retries=3)
rel  = RelationExtractor(method="llm", llm_provider=llm, max_retries=3)
trip = TripletExtractor(method="llm",  llm_provider=llm, max_retries=3)

entities      = ner.extract(text)
relationships = rel.extract(text, entities=entities)
triplets      = trip.extract(text)

Extraction Method Comparison

Method Speed Cost Accuracy Custom Types
pattern Very fast Free Medium Yes (dictionary)
ml Fast Free High Limited
llm Medium API cost Highest Yes (schema)
Configure which LLM is used for extraction. Build graphs from extracted entities and relationships. Parse documents before extraction. Resolve duplicate entities after extraction.