mirror of
https://github.com/semantica-agi/semantica.git
synced 2026-08-29 04:26:20 +00:00
- Migrate from mint.json to docs.json (Mintlify v4) - Theme: maple, emerald green + near-black dark / cream light palette (#059669 primary, #0A0A0A dark bg, #FAF7F0 light bg) - Typography: Lexend headings, Inter body - 5-tab navigation: Documentation, Quick Start, API Reference, Cookbook, FAQ - Homepage: removed badge stickers, redundant h2, added blockquote tagline, full 27-module reference table with semantica.mcp_server added - quickstart.md: CodeGroup per pipeline step, pattern vs LLM options, AccordionGroup for patterns and troubleshooting - faq.md: full AccordionGroup structure across 5 sections - reference/explorer.md: NEW — FastAPI explorer, Ontology Hub, Distance Intelligence, CLI reference, REST API endpoints - reference/mcp_server.md: NEW — MCP stdio server, 12 tools with I/O examples, 3 resources, Claude Desktop/VS Code/Windsurf/Cline config - docs.json: explorer added to Output group, mcp_server to Utilities group - Chat, feedback (thumbs/suggest/raise), OG/Twitter metadata, search topbar - All reference pages reformatted with Mintlify JSX components Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
3.1 KiB
3.1 KiB
title, description, icon
| title | description | icon |
|---|---|---|
| Split Module | 15+ text chunking methods including recursive, semantic, entity-aware, and relation-aware splitting. | scissors |
Comprehensive document chunking for optimal RAG, embedding, and extraction pipelines.
Overview
The Split Module breaks documents into chunks while preserving context and semantic meaning — critical for embedding quality in RAG systems.
TextSplitter
from semantica.split import TextSplitter
splitter = TextSplitter(
method="semantic", # see methods below
chunk_size=1000,
overlap=200
)
chunks = splitter.split(text)
for chunk in chunks:
print(f"Chunk: {chunk.text[:80]}... ({chunk.token_count} tokens)")
Splitting Methods
| Method | Description | Best for |
|---|---|---|
recursive |
Split by paragraph → sentence → word | General purpose |
semantic |
Split at semantic boundaries (topic shifts) | RAG systems |
entity-aware |
Keep entity mentions intact across boundaries | NER pipelines |
relation-aware |
Keep relation triplets intact | KG construction |
sentence |
Split by sentence | Short content |
token |
Split by token count (tiktoken) | LLM context windows |
fixed |
Fixed character count with overlap | Batch processing |
markdown |
Split by Markdown headers | Documentation |
code |
Split by function/class boundaries | Code analysis |
Entity-Aware Chunking
from semantica.split import TextSplitter
from semantica.semantic_extract import NERExtractor
ner = NERExtractor()
entities = ner.extract(text)
splitter = TextSplitter(method="entity-aware")
chunks = splitter.split(text, entities=entities)
Entity mentions are never split across chunk boundaries, preserving context for downstream NER.
Relation-Aware Chunking
splitter = TextSplitter(method="relation-aware")
chunks = splitter.split(text, relationships=relationships)
Keeps subject–predicate–object triplets within the same chunk.
Semantic Chunking
from semantica.split import TextSplitter
from semantica.embeddings import EmbeddingGenerator
embedder = EmbeddingGenerator(model="sentence-transformers")
splitter = TextSplitter(
method="semantic",
embedder=embedder,
similarity_threshold=0.7 # split when topic similarity drops below this
)
chunks = splitter.split(text)
Chunk Object
@dataclass
class Chunk:
text: str
start_char: int
end_char: int
token_count: int
metadata: Dict # source_id, chunk_index, section_title, etc.
entities: List[Dict] # if entity-aware splitting was used