mirror of
https://github.com/semantica-agi/semantica.git
synced 2026-08-29 04:26:20 +00:00
Replace plain markdown in every docs/reference/ file and docs/concepts.md with rich Mintlify JSX components — CardGroup, Steps, Tabs, AccordionGroup, Tip, Warning, Note, and CodeGroup — for a consistent, navigable, production-grade developer experience.
17 KiB
17 KiB
title, description, icon
| title | description | icon |
|---|---|---|
| LLMs Module | Unified interface for Groq, OpenAI, Anthropic, Gemini, Ollama, DeepSeek, Novita AI, LiteLLM, and HuggingFace — swap providers with a one-line change. | microchip |
semantica.llms gives every Semantica module a single, consistent interface to 9+ LLM providers. Every extractor, reasoning engine, and context graph accepts any provider through the same llm_provider= parameter — swap Groq for Anthropic, or a cloud API for a local Ollama model, by changing one line.
What You Get
Groq, OpenAI, Anthropic, Gemini, Ollama, DeepSeek, Novita AI, LiteLLM, and HuggingFace — all behind one interface. `complete()`, `chat()`, and `stream()` work identically across all providers — swap with a one-line change. Instantiate any provider from a name string — drive provider selection entirely from YAML config or environment variables. Ollama and HuggingFace run fully on-premise — no API key, no data leaves your machine, air-gap compatible. Token-by-token output via `stream()` for responsive agent pipelines and live UI updates. Configurable `max_retries` with exponential backoff. Typed exceptions: `LLMAuthenticationError`, `LLMRateLimitError`, `LLMContextLengthError`.Installation
The base install includes Groq and DeepSeek. Other providers require optional extras:
| Provider | Install Command | API Key Required |
|---|---|---|
| Groq | pip install semantica |
Yes — GROQ_API_KEY |
| DeepSeek | pip install semantica |
Yes — DEEPSEEK_API_KEY |
| Novita AI | pip install semantica |
Yes — NOVITA_API_KEY |
| OpenAI | pip install "semantica[llm-openai]" |
Yes — OPENAI_API_KEY |
| Anthropic | pip install "semantica[llm-anthropic]" |
Yes — ANTHROPIC_API_KEY |
| Gemini | pip install "semantica[llm-gemini]" |
Yes — GOOGLE_API_KEY |
| Ollama | pip install "semantica[llm-ollama]" |
No — local server |
| LiteLLM | pip install "semantica[llm-litellm]" |
Varies by target |
| HuggingFace | pip install "semantica[llm-huggingface]" |
No — local weights |
| All providers | pip install "semantica[all]" |
Varies |
Quick Start
```python from semantica.llms import Groq import osllm = Groq(model="llama-3.3-70b-versatile", api_key=os.getenv("GROQ_API_KEY"))
```
ner = NERExtractor(method="llm", llm_provider=llm)
```
config = ConfigManager("config.yaml")
llm = create_provider(
config.get("llm_provider.name"),
model=config.get("llm_provider.model"),
api_key=config.get("llm_provider.api_key"),
)
```
Providers
from semantica.llms import Groq
import os
llm = Groq(
model="llama-3.3-70b-versatile", # default and recommended
api_key=os.getenv("GROQ_API_KEY"),
temperature=0.0,
max_tokens=64000,
max_retries=3,
timeout=60,
)
from semantica.llms import OpenAI
import os
llm = OpenAI(
model="gpt-4o",
api_key=os.getenv("OPENAI_API_KEY"),
temperature=0.0,
max_tokens=4096,
max_retries=3,
timeout=120,
organization=None, # optional org ID
)
from semantica.llms import Anthropic
import os
llm = Anthropic(
model="claude-opus-4-7",
api_key=os.getenv("ANTHROPIC_API_KEY"),
max_tokens=8192,
temperature=0.0,
max_retries=3,
timeout=120,
)
from semantica.llms import Gemini
import os
llm = Gemini(
model="gemini-1.5-pro",
api_key=os.getenv("GOOGLE_API_KEY"),
temperature=0.0,
max_tokens=8192,
timeout=120,
)
from semantica.llms import Ollama
llm = Ollama(
model="llama3.2",
base_url="http://localhost:11434", # default Ollama address
temperature=0.0,
timeout=180, # local models can be slower; increase for large models
)
# No API key — model runs entirely on your machine
from semantica.llms import DeepSeek
import os
llm = DeepSeek(
model="deepseek-chat",
api_key=os.getenv("DEEPSEEK_API_KEY"),
temperature=0.0,
max_tokens=4096,
max_retries=3,
)
from semantica.llms import NovitaAI
import os
llm = NovitaAI(
model="deepseek/deepseek-v3.2",
api_key=os.getenv("NOVITA_API_KEY"),
temperature=0.0,
max_tokens=4096,
)
# OpenAI-compatible endpoint; also accepts deepseek/deepseek-r1, meta-llama models
from semantica.llms import LiteLLM
import os
llm = LiteLLM(
model="gpt-4o", # any LiteLLM model string
api_key=os.getenv("OPENAI_API_KEY"),
temperature=0.0,
max_tokens=4096,
)
# Supports: OpenAI, Anthropic, Gemini, Cohere, Azure, Bedrock, Together AI, and 90+ more
# Use the LiteLLM model string format: "anthropic/claude-3-5-sonnet", "bedrock/anthropic.claude-v2"
from semantica.llms import HuggingFace
llm = HuggingFace(
model="mistralai/Mistral-7B-Instruct-v0.3",
device="cuda", # "cpu" | "cuda" | "mps" (Apple Silicon)
max_new_tokens=512,
temperature=0.1,
load_in_4bit=True, # enable 4-bit quantisation to reduce VRAM
)
# No API key — weights downloaded from Hugging Face Hub (or loaded from local path)
Constructor Parameters
Common Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
model |
str |
Provider default | Model identifier string |
api_key |
str |
None |
API key — reads from environment if omitted |
temperature |
float |
0.0 |
Sampling temperature: 0 = deterministic, 1 = creative |
max_tokens |
int |
Provider default | Maximum tokens in the response |
max_retries |
int |
3 |
Number of retry attempts on transient failures |
timeout |
int |
60 |
Request timeout in seconds |
base_url |
str |
Provider default | Override the API endpoint — useful for proxies and gateways |
Provider-Specific Parameters
| Provider | Parameter | Description |
|---|---|---|
OpenAI |
organization |
OpenAI organisation ID |
OpenAI |
project |
OpenAI project ID |
Ollama |
base_url |
Ollama server address (default: http://localhost:11434) |
HuggingFace |
device |
Compute device: "cpu" / "cuda" / "mps" |
HuggingFace |
load_in_4bit |
Enable 4-bit quantisation (requires bitsandbytes) |
HuggingFace |
max_new_tokens |
Maximum new tokens to generate (replaces max_tokens) |
LiteLLM |
model |
Full LiteLLM model string, e.g. "anthropic/claude-3-5-sonnet" |
Direct API Usage
Providers can be used directly — not just through Semantica modules:
from semantica.llms import Groq
import os
llm = Groq(model="llama-3.3-70b-versatile", api_key=os.getenv("GROQ_API_KEY"))
# Single completion
response = llm.complete("What is a knowledge graph?")
print(response.text) # answer string
print(response.input_tokens) # tokens consumed by the prompt
print(response.output_tokens) # tokens in the response
print(response.model) # model that served the request
# Multi-turn chat
messages = [
{"role": "system", "content": "You are a knowledge graph expert."},
{"role": "user", "content": "What is the difference between RDF and property graphs?"},
]
response = llm.chat(messages)
print(response.text)
# Streaming — token-by-token output
for token in llm.stream("Explain knowledge graph reasoning in 3 sentences."):
print(token, end="", flush=True)
LLMResponse Object
All three methods (complete, chat, stream) return a LLMResponse dataclass:
@dataclass
class LLMResponse:
text: str # the generated text
model: str # model identifier that served the request
input_tokens: int # tokens consumed by the prompt
output_tokens: int # tokens in the generated response
total_tokens: int # input_tokens + output_tokens
latency_ms: float # wall-clock time for the API call in milliseconds
finish_reason: str # "stop" | "length" | "content_filter" | "tool_calls"
Error Handling
from semantica.llms import Groq
from semantica.utils import SemanticaError
import os
try:
llm = Groq(model="llama-3.3-70b-versatile", api_key=os.getenv("GROQ_API_KEY"))
response = llm.complete("Summarise this document.")
print(response.text)
except LLMAuthenticationError as e:
# Invalid or expired API key
print(f"Authentication failed: {e}")
except LLMRateLimitError as e:
# Rate limit hit — Semantica retries automatically up to max_retries
print(f"Rate limited after retries: {e}")
except LLMContextLengthError as e:
# Prompt exceeds the model's context window
print(f"Prompt too long ({e.token_count} tokens, limit {e.context_limit}): {e}")
except LLMProviderError as e:
# General provider-side error (5xx, model unavailable, etc.)
print(f"Provider error: {e}")
except SemanticaError as e:
# Catch-all for all Semantica framework errors
print(f"Framework error: {e}")
| Exception | When Raised |
|---|---|
LLMAuthenticationError |
Invalid or missing API key |
LLMRateLimitError |
Rate limit exceeded after all retries |
LLMContextLengthError |
Prompt exceeds the model's context window |
LLMProviderError |
Provider-side error (5xx, unavailability) |
LLMTimeoutError |
Request exceeded the timeout parameter |
Custom and Enterprise Gateways
Any provider that exposes an OpenAI-compatible REST API can be used by passing base_url:
from semantica.llms import OpenAI
import os
# Internal LLM routing gateway
llm = OpenAI(
model="qwen2.5-72b",
api_key=os.getenv("GATEWAY_API_KEY"),
base_url="https://llm-gateway.internal.company.com/v1",
)
# Azure OpenAI Service
llm = OpenAI(
model="gpt-4o",
api_key=os.getenv("AZURE_OPENAI_API_KEY"),
base_url="https://my-resource.openai.azure.com/openai/deployments/gpt-4o",
)
# Self-hosted vLLM server
llm = OpenAI(
model="meta-llama/Llama-3.1-8B-Instruct",
api_key="not-needed",
base_url="http://localhost:8000/v1",
)
Using in Semantica Modules
Every module that uses an LLM accepts any provider through llm_provider=:
from semantica.llms import Groq, Anthropic
from semantica.semantic_extract import NERExtractor, RelationExtractor, TripletExtractor
from semantica.ontology import LLMOntologyGenerator
from semantica.reasoning import ReasoningEngine
from semantica.context import AgentContext, ContextGraph
from semantica.vector_store import VectorStore
import os
groq_llm = Groq(model="llama-3.3-70b-versatile", api_key=os.getenv("GROQ_API_KEY"))
claude_llm = Anthropic(model="claude-opus-4-7", api_key=os.getenv("ANTHROPIC_API_KEY"))
# Extraction — use fast Groq for high-throughput NER
ner = NERExtractor(method="llm", llm_provider=groq_llm)
rel = RelationExtractor(method="llm", llm_provider=groq_llm)
trip = TripletExtractor(method="llm", llm_provider=groq_llm)
# Complex reasoning — use Claude for accuracy
engine = ReasoningEngine(llm_provider=claude_llm)
# Ontology generation from natural language
gen = LLMOntologyGenerator(llm_provider=claude_llm)
Provider Comparison
| Provider | Speed | Cost | Local | Max Context | Best For |
|---|---|---|---|---|---|
| Groq | ⚡ Very fast | 💲 Low | No | 128k | High-throughput extraction, fast pipelines |
| OpenAI | Fast | 💲💲 Medium | No | 128k | General purpose, function calling, JSON mode |
| Anthropic | Fast | 💲💲 Medium | No | 200k | Complex reasoning, long documents, safety |
| Gemini | Fast | 💲 Low | No | 1M | Very long context, multimodal (text + image) |
| Ollama | Medium | Free | ✅ Yes | Varies | Privacy, air-gapped, no API key |
| DeepSeek | Fast | 💲 Very low | No | 64k | Coding tasks, structured analysis |
| Novita AI | Fast | 💲 Low | No | Varies | DeepSeek and LLaMA models, cost-effective |
| LiteLLM | Varies | Varies | Varies | Varies | Multi-provider routing, vendor abstraction |
| HuggingFace | Slow | Free | ✅ Yes | Varies | Custom and fine-tuned models, full local control |
Environment Variables
| Variable | Provider | Notes |
|---|---|---|
GROQ_API_KEY |
Groq | Required for Groq cloud |
OPENAI_API_KEY |
OpenAI, LiteLLM | Also used for OpenAI-compatible gateways |
ANTHROPIC_API_KEY |
Anthropic | Required for Claude |
GOOGLE_API_KEY |
Gemini | Required for Gemini |
DEEPSEEK_API_KEY |
DeepSeek | Required for DeepSeek cloud |
NOVITA_API_KEY |
Novita AI | Required for Novita AI |
HUGGINGFACE_HUB_TOKEN |
HuggingFace | Required for gated models (optional for public models) |
YAML Configuration
# config.yaml
llm_provider:
name: "groq"
model: "llama-3.3-70b-versatile"
api_key: "${GROQ_API_KEY}" # reads from environment
temperature: 0.0
max_tokens: 64000
max_retries: 3
timeout: 60
Load it with ConfigManager:
from semantica.core import ConfigManager
from semantica.llms import create_provider
config = ConfigManager("config.yaml")
llm = create_provider(
config.get("llm_provider.name"),
model=config.get("llm_provider.model"),
api_key=config.get("llm_provider.api_key"),
temperature=config.get("llm_provider.temperature", default=0.0),
)