docs: apply full Mintlify component overhaul to all 27 reference pages and concepts.md

Replace plain markdown in every docs/reference/ file and docs/concepts.md with
rich Mintlify JSX components — CardGroup, Steps, Tabs, AccordionGroup, Tip,
Warning, Note, and CodeGroup — for a consistent, navigable, production-grade
developer experience.
This commit is contained in:
KaifAhmad1
2026-05-23 23:02:03 +05:30
parent 11e8a2fc0d
commit 5eefadaa7f
29 changed files with 7838 additions and 2274 deletions
+205 -24
View File
@@ -8,33 +8,70 @@ icon: "database"
## What You Get
- **`VectorStore`** — unified interface across all backends
- **Backends** — FAISS, Pinecone, Weaviate, Qdrant, Milvus, PgVector, in-memory
- **Hybrid search** — combine dense vector similarity with sparse keyword/metadata filtering
- **Metadata filtering** — rich filter expressions: `eq`, `ne`, `gt`, `lt`, `in`, `contains`, `$and`, `$or`
- **Namespace isolation** — multi-tenant support via isolated namespaces
- **Batch operations** — bulk add, delete, and metadata updates
<CardGroup cols={2}>
<Card title="VectorStore" icon="database">
Unified interface across FAISS, Pinecone, Weaviate, Qdrant, Milvus, and PgVector.
</Card>
<Card title="HybridSearch" icon="magnifying-glass">
Combine dense vector similarity with sparse keyword/BM25 filtering and configurable fusion strategies.
</Card>
<Card title="MetadataStore" icon="table">
Rich metadata indexing and schema management — query by field values without a vector.
</Card>
<Card title="NamespaceManager" icon="folder-tree">
Multi-tenant namespace isolation — structural separation, not just metadata filters.
</Card>
<Card title="Batch Operations" icon="layer-group">
Bulk add, delete, and metadata updates — automatically chunked for memory efficiency.
</Card>
<Card title="FAISS Index Types" icon="chart-scatter">
Flat, IVF, HNSW, and PQ index types with full configuration control.
</Card>
</CardGroup>
## Basic Usage
## Quick Start
```python
from semantica.vector_store import VectorStore
<Steps>
<Step title="Create a vector store">
```python
from semantica.vector_store import VectorStore
# In-memory (development)
store = VectorStore(backend="inmemory", dimension=768)
# In-memory (development)
store = VectorStore(backend="inmemory", dimension=768)
# FAISS (local, production)
store = VectorStore(backend="faiss", dimension=768, index_path="store.faiss")
# Add vectors
store.add_vectors(embeddings=embeddings, ids=["doc1", "doc2"], metadata=[{}, {}])
# Semantic search
results = store.search(query_vector, top_k=10)
for r in results:
print(f"{r['id']} — score: {r['score']:.3f}")
print(f" metadata: {r['metadata']}")
```
# FAISS (local production — persists to disk)
store = VectorStore(backend="faiss", dimension=768, index_path="store.faiss")
```
</Step>
<Step title="Add vectors">
```python
store.add_vectors(
embeddings=embeddings,
ids=["doc1", "doc2"],
metadata=[{"title": "Document 1"}, {"title": "Document 2"}]
)
```
</Step>
<Step title="Search by semantic similarity">
```python
results = store.search(query_vector, top_k=10)
for r in results:
print(f"{r['id']} — score: {r['score']:.3f}")
print(f" metadata: {r['metadata']}")
```
</Step>
<Step title="Filter results by metadata">
```python
# Equality, range, and set filters
results = store.search(query_vector, filters={
"$and": [
{"category": "research"},
{"year": {"$gte": 2022}}
]
})
```
</Step>
</Steps>
## Backends
@@ -45,7 +82,7 @@ for r in results:
store = VectorStore(
backend="faiss",
dimension=768,
index_type="IVF", # "Flat" | "IVF" | "HNSW"
index_type="IVF", # "Flat" | "IVF" | "HNSW" | "PQ"
index_path="store.faiss"
)
```
@@ -197,6 +234,150 @@ store.update_metadata("doc1", {"status": "archived", "reviewed": True})
| PgVector | PostgreSQL | No | Limited | Postgres-native integration |
| In-memory | Process | No | No | Development, testing |
## HybridSearch
`HybridSearch` is the low-level class behind `store.hybrid_search()` — use it directly when you need custom result fusion logic:
```python
from semantica.vector_store import HybridSearch, VectorStore
store = VectorStore(backend="faiss", dimension=768)
hybrid = HybridSearch(vector_store=store)
results = hybrid.search(
query_vector=query_embedding,
query_text="machine learning frameworks",
top_k=20,
vector_weight=0.7, # weight for vector similarity leg
keyword_weight=0.3, # weight for BM25/keyword leg
fusion="rrf", # "rrf" (Reciprocal Rank Fusion) | "weighted_avg"
filters={"category": "research", "year": {"$gte": 2022}},
deduplicate=True,
)
for r in results:
print(f"{r['id']} vector_score={r['vector_score']:.3f} final_score={r['score']:.3f}")
```
| Fusion strategy | Description |
| --------------- | ----------- |
| `rrf` | Reciprocal Rank Fusion — rank-based combination, robust to score scale differences |
| `weighted_avg` | Weighted average of normalised scores — requires `vector_weight` + `keyword_weight` = 1.0 |
## MetadataStore
`MetadataStore` manages structured metadata attached to vectors — query by field values without a vector:
```python
from semantica.vector_store import MetadataStore
meta_store = MetadataStore()
meta_store.register_schema({
"author": "str",
"year": "int",
"category": "str",
"score": "float",
})
meta_store.add("doc1", {"author": "Alice", "year": 2024, "category": "research"})
meta_store.add("doc2", {"author": "Bob", "year": 2023, "category": "review"})
results = meta_store.filter({"category": "research", "year": {"$gte": 2023}})
meta = meta_store.get("doc1")
meta_store.update("doc1", {"score": 0.92})
```
## NamespaceManager
Isolates vector collections per tenant, project, or model version:
```python
from semantica.vector_store import NamespaceManager, VectorStore
base_store = VectorStore(backend="faiss", dimension=768)
ns_manager = NamespaceManager(vector_store=base_store)
ns_manager.create_namespace("tenant_a", description="Customer A data")
ns_manager.create_namespace("tenant_b", description="Customer B data")
ns_manager.add_vectors("tenant_a", embeddings_a, ids_a, metadata_a)
ns_manager.add_vectors("tenant_b", embeddings_b, ids_b, metadata_b)
# Search is scoped — tenant_a never sees tenant_b's data
results = ns_manager.search("tenant_a", query_vector, top_k=10)
for ns in ns_manager.list_namespaces():
print(f"{ns['name']}: {ns['vector_count']} vectors")
ns_manager.delete_namespace("tenant_a")
```
## FAISS Index Type Reference
| Index | Memory | Speed | Accuracy | When to Use |
| ----- | ------ | ----- | -------- | ----------- |
| `Flat` | High | Slow | Exact (100%) | < 100K vectors, correctness critical |
| `IVF` | Medium | Fast | ~9598% | 100K10M vectors, good balance |
| `HNSW` | Medium-High | Very fast | ~9799% | Low latency, production retrieval |
| `PQ` | Low | Fast | ~9095% | Millions of vectors, memory-constrained |
```python
# Flat — brute-force exact search
store = VectorStore(backend="faiss", dimension=768, index_type="Flat")
# IVF — inverted file index with nlist clusters
store = VectorStore(backend="faiss", dimension=768, index_type="IVF", nlist=100)
# HNSW — hierarchical navigable small world graph
store = VectorStore(backend="faiss", dimension=768, index_type="HNSW", M=16, ef_construction=200)
# PQ — product quantization for memory efficiency
store = VectorStore(backend="faiss", dimension=768, index_type="PQ", m=8)
```
## Similarity Metrics
| Metric | Constructor arg | Distance → Similarity | Best For |
| ------ | --------------- | --------------------- | -------- |
| Cosine | `metric="cosine"` | `1 - cosine_distance` | Text, embeddings |
| L2 (Euclidean) | `metric="l2"` | `1 / (1 + distance)` | Image features |
| Inner Product | `metric="ip"` | raw dot product | Recommendation systems |
```python
store = VectorStore(backend="faiss", dimension=768, metric="cosine")
```
## Tips and Common Pitfalls
<Warning>
**Match vector dimension to your embedding model.** The `dimension` parameter must exactly match your embedding model's output size — `all-MiniLM-L6-v2` = 384, `all-mpnet-base-v2` = 768, `bge-large-en-v1.5` = 1024. A mismatch raises an error at insert time, not at store creation.
</Warning>
<Tip>
**Use `Flat` index only for small datasets.** Flat (brute-force) search has perfect recall but O(n) query time. At 500K+ vectors, switch to `IVF` or `HNSW` — they sacrifice less than 5% recall for 1001000x speedup.
</Tip>
<Warning>
**Don't search without normalizing first.** If you disabled `normalize=True` in `EmbeddingGenerator`, compute cosine similarity with `metric="cosine"` (which normalizes internally). Raw dot product on un-normalized vectors produces incorrect similarity rankings.
</Warning>
<Tip>
**Use `hybrid_search` for precision-sensitive workloads.** Pure vector search finds semantically similar results but may miss keyword matches important to the user. Hybrid search (vector + BM25) combines both signals — especially valuable for domain-specific terminology.
</Tip>
<Tip>
**Use `NamespaceManager` for multi-tenant applications.** Storing all tenants' vectors in the same collection and filtering by metadata at query time is slow and leaks data if a filter is accidentally omitted. Namespace isolation is both faster (smaller search space) and safer (structural isolation).
</Tip>
<Warning>
**Persist FAISS indexes to disk.** `VectorStore(backend="faiss", index_path="store.faiss")` saves the index to disk on each write. Without a path, the index is in-memory only and is lost on process exit.
</Warning>
<Tip>
**Update metadata without re-embedding.** `store.update_metadata(id, {...})` changes attached fields (status, tags, review date) without re-running the embedding model. Use this for state changes that don't affect semantic content.
</Tip>
<CardGroup cols={2}>
<Card title="Embeddings" icon="vector-square" href="embeddings">
Generate the vectors stored here.