* fix(vector-store): prevent in-memory vector id reuse
* fix(vector-store): synchronize in-memory mutations
* fix(vector-store): clarify no-silent-overwrite guarantee for in-memory ID generation
The while-loop in store_vectors() already prevents any generated candidate
from landing on a live key: it breaks only when the candidate is absent
from both pre_existing (the live store at lock-entry) and within_batch
(IDs chosen earlier in the same call).
This commit:
- Renames 'existing' to 'pre_existing' and introduces 'within_batch' to
make the two-level de-dupe explicit and self-documenting.
- Tightens the break condition to 'not in pre_existing and not in
within_batch' so within-batch duplicates are also guarded.
- Adds test_auto_generated_id_never_silently_overwrites_live_vector:
stores 5, deletes 3, stores 2 more and asserts none of the new IDs
land on a surviving vec_N.
- Adds test_collision_detection_no_silent_overwrite_even_with_corrupted_counter:
rewinds _next_id to 0 with live vectors present and proves the loop
finds a free slot without overwriting either existing vector or its
metadata — the no-silent-overwrite invariant holds even under counter
corruption.
The skip-over-occupied-key behaviour for caller-inserted vec_N IDs is
intentional and preserved (test_counter_skips_explicit_vec_n_ids), matching
FAISSStore's identical pattern.
Closes reviewer finding: 'Vector collisions do not fail writes'.
* fix(vector-store): protect concurrent in-memory access
---------
Co-authored-by: Sameer Kadam <sameerkadam@Mac.lan>
Co-authored-by: Zohaib Hassnain <109234410+ZohaibHassan16@users.noreply.github.com>
* fix(parse): stop infinite recursion in default method dispatch
parse_document and its five sibling dispatchers (parse_web_content,
parse_structured_data, parse_email, parse_code, parse_media) were
registered in the method registry under their own task's "default"
name. Every dispatcher begins with method_registry.get(<task>, method),
so calling e.g. parse_document(file, method="default") found itself in
the registry and re-entered infinitely until RecursionError --
`semantica parse <any file>` crashed before any parsing ran.
Drop the six self-registrations. "default" remains the built-in code
path; users can still register their own "default" (or any other name)
to override it, and the existing "docling" registration is unaffected.
Verified: `semantica parse demo.pdf` now parses successfully (was
RecursionError). No repo code or tests consume
get_parse_method("document", "default"), so removing the entries
changes no behavior besides fixing the crash.
Co-Authored-By: Claude Code <noreply@anthropic.com>
* fix(deps): make resolution satisfiable on Python 3.9 and add parse-pdf extra
Dependency fixes so `uv lock` / `pip install semantica[all]` resolves
across the supported matrix (3.9-3.12):
- requires-python >=3.8 was unsatisfiable (numpy>=2.0.2 needs >=3.9) and
3.9.0/3.9.1 can never resolve (every cryptography release excludes
them) -> bump to >=3.9.2 and drop the 3.8 classifier.
- Split recently-raised floors that dropped 3.9 into marker pairs
(3.9-capped / 3.10-unconstrained), following the pattern already used
for scikit-learn/requests/etc.: pyarrow extras (>=24 needs 3.10),
pre-commit 4.6, snowflake-connector 4.6, fastapi 0.129 + starlette 0.53
(older fastapi caps starlette<0.53).
- Gate docling, litellm (its only 3.9 release pins
python-dotenv==1.0.1, conflicting with the >=1.2.1 core floor), and
crewai (no un-yanked 3.9 release) to >=3.10.
Also add a parse-pdf extra: the default PDFParser requires pdfplumber,
but no extra installed it, so `semantica parse file.pdf` failed on a
default install. Follows the parse-docling convention.
Co-Authored-By: Claude Code <noreply@anthropic.com>
* fix(vector-store): make the qdrant backend usable through VectorStore
The qdrant backend could not be used at all through the VectorStore
facade in 0.6.8 - every path raised:
1. store_vectors() only dispatched to backend add/add_vectors methods;
QdrantStore exposes insert_vectors, so vs.store() raised
NotImplementedError. Add an insert_vectors branch (uuid-generated
ids, metadata -> payloads).
2. QdrantStore required an explicit create_collection() before any read
or write, unlike FAISSStore's automatic index creation. Lazily attach
the configured collection (config key "collection", default
"semantica_default") on first insert/search, reusing an existing one.
3. search_points() called client.search(), removed from qdrant-client in
favor of query_points() - use it when available, fall back otherwise.
4. store() silently dropped plain-string documents (it only extracted
doc.metadata), so payloads lost the source text; keep them under
payload "document".
Verified end-to-end against a Qdrant 5 server (docker): embed ->
store -> semantic search returns correctly ranked results whose payloads
carry the original documents.
Co-Authored-By: Claude Code <noreply@anthropic.com>
* fix(vector-store): address Qodo review findings on the qdrant write path
Four findings from the Qodo review of #1508:
1. store_vectors() returned QdrantStore.insert_vectors()' upsert status
dict although the facade promises the stored vector IDs (decision
storage indexes the result at position 0). Return the generated or
caller-supplied IDs after a successful insert instead.
2. insert_vectors() pairs points with zip(vectors, ids), so a shorter
non-empty id list silently dropped the unpaired vectors while the
completion message still reported the full batch as inserted. Reject
the mismatch with ValidationError before any write.
3. _ensure_default_collection() looked up the legacy "collection" config
key, so the documented collection_name=... option was ignored and
lazy init always fell back to semantica_default. Prefer
collection_name, keep "collection" as an alias.
4. The new parse-pdf extra was missing from both aggregate "all"
bundles, so semantica[all] still shipped without pdfplumber and the
default PDFParser raised ProcessingError on first use.
Also removes the now-stale strict xfail for qdrant's write dispatch in
test_backend_facade_contract.py — that marker exists precisely to fail
once the wiring lands.
Co-Authored-By: Claude Code <noreply@anthropic.com>
* test(vector-store): migrate qdrant mocks to the query_points API
The qdrant search path now prefers client.query_points() (qdrant-client
removed client.search), but these tests still mocked the legacy call.
A MagicMock exposes query_points too, so the code took the modern path,
read .points off an unconfigured mock, and all three tests failed on
the branch. Return the hits in response.points as the real client does.
Co-Authored-By: Claude Code <noreply@anthropic.com>
* fix(ci): regenerate requirements-ci.txt for the parse-pdf extra
The CI lockfile check re-resolves pyproject.toml --extra all and diffs
the pinned versions against requirements-ci.txt. The parse-pdf extra
added pdfplumber (+pdfminer-six) to the 'all' bundle without refreshing
the lockfile, so the diff failed and the build job exited 1.
Regenerated with the command from the file header. Only the two new
pins and their 'via' comments changed - every other version is
identical. Verified locally with the CI's own check command (diff of
pkg==ver lines exits clean).
Co-Authored-By: Claude Code <noreply@anthropic.com>
* fix(tests): update stale default-method-registration assertion
test_parse_methods_dynamic_default_resolution asserted that "default"
resolves through the registry to parse_document, which was true only
because of the self-registration this PR removes (it's what caused the
infinite recursion in the first place). Update the assertion to match
the intended post-fix state: "default" is not registered at all, and
falls through to the built-in dispatch path unconditionally.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
---------
Co-authored-by: yanyu <yanyu@polixir.com>
Co-authored-by: Claude Code <noreply@anthropic.com>
Co-authored-by: Mohd Kaif <98801504+KaifAhmad1@users.noreply.github.com>
Co-authored-by: KaifAhmad1 <kaifahmad087@gmail.com>
Adds native vector deletion support for FAISS Flat indices via remove_ids
and IDSelectorBatch (Closes#1374):
- Adds FAISSIndex.delete_vectors() and FAISSStore.delete_vectors() to
translate external string IDs to sequential internal positions, remove
vectors via remove_ids, and keep vector_ids and metadata parallel with
the compacted index.
- Auto-saves index and metadata sidecar after deletion when the store was
loaded from disk via load_index(), skipping redundant disk writes for
no-op deletions.
- Rejects IVF and HNSW index deletions with NotImplementedError: IVF does
not compact internal labels upon remove_ids (causing search/offset
desynchronization), and HNSW lacks remove_ids entirely. Both cleanly
map to STATUS_UNSUPPORTED in ErasureCoordinator.
Robust ID Management & Invariant Hardening:
- Generates default IDs with a monotonic candidate-check loop against
existing vector_ids, eliminating collisions and metadata overwrites
when explicit vec_N IDs coexist.
- Persists next_id in the .meta.json sidecar and clamps on load() to
max(persisted, max(vec_N) + 1), preventing stale sidecars from
re-introducing collisions across restarts or legacy migrations.
- Guards against FAISS -1 sentinels (0 <= idx < len(vector_ids)) in
search results when k > ntotal, preventing negative index wrapping
to vector_ids[-1].
- Replaces bare post-delete assert with an explicit ProcessingError check
that survives python -O.
Adds 43 comprehensive tests in test_faiss_delete_vectors.py covering
deletion, search exclusion, metadata cleanup, persistence roundtrips,
IVF/HNSW rejection, ErasureCoordinator receipts, and mock-spied
deterministic no-op saves.
* feat(milvus): add delete_vectors to MilvusStore
Adds delete_vectors(ids) so the ErasureCoordinator can erase embeddings
on a Milvus backend. Ids are escaped with _format_milvus_value before
building the delete expression, and the backend delete count is returned
so a delete that removed nothing is distinguishable from a failure.
* test(milvus): cover delete_vectors incl. erasure integration
Unit tests assert the single-id equality expression, multi-id in
expression, id escaping, empty-id noop, missing-collection and backend
error paths. Integration tests bind MilvusStore as a VectorStore backend
and assert ErasureCoordinator reports the vector leg erased.
* test(milvus): cover vector store delete facade
---------
Co-authored-by: Mohd Kaif <98801504+KaifAhmad1@users.noreply.github.com>
Co-authored-by: Sameer Kadam <sskadam6305@gmail.com>
Co-authored-by: Sameer Kadam <sameerkadam@Mac.lan>
if not vector_ids: return fired before the continuation token was ever
checked. Pinecone's actual pagination contract is that a scan is only
exhausted when the response carries no pagination token -- a page can
legitimately list zero ids while pagination.next is still set (sparse
or filtered namespaces, eventual-consistency windows on serverless
indexes). This was flagged in review but the fix commit that followed
only addressed the separate repeated-token stall case, not this one.
Reproduced concretely against the unfixed code: a page with data,
followed by an empty page with a live token, followed by a page with
more data -- the last page was silently dropped with no error raised,
exactly the #1083 failure mode (store migrate reporting success after
copying only part of a collection).
Now the empty-page case skips the pointless fetch() call but still
falls through to the same next_token check every other path already
goes through, so a live token continues the scan and only a genuinely
absent token (or one that's stopped advancing) ends it.
Added test_continues_past_an_empty_page_with_a_live_token, the
"empty page + non-None next token" case the original review asked for
and that wasn't otherwise covered.
iter_all() unconditionally returned on any empty fetch_objects() page,
regardless of pagination mode. That's safe for offset/single_page (an
empty page there is a direct, unambiguous statement about live rows),
but not for cursor mode: `after` has no server-issued continuation
value of its own, it's derived client-side from the last object's uuid,
so an empty page gives nothing to advance it with. If Weaviate's cursor
walks internal storage position rather than strict uuid order, a batch
can in principle land entirely on a gap (e.g. tombstoned objects) with
live data past it -- the same risk already confirmed and fixed for
Qdrant's scroll cursor in #1316. Reproduced concretely against the
pre-fix code: a full page followed by an empty page followed by a page
with real data silently dropped that last page with no error raised.
iter_all() now falls back to offset pagination once when a cursor-mode
page comes back empty, rather than assuming that's the end. Offset
addresses live rows directly by position and has no equivalent gap, so
an empty page there (or in single_page mode) is trustworthy and still
ends the scan immediately.
Also updates test_iter_all_empty_collection_yields_nothing and
test_iter_all_requests_vectors, which needed a second empty page now
that a genuinely empty collection takes two calls (cursor, then the
confirming offset check) to report as such.
This branch was forked from an earlier commit of feat/vector-store-iter-all
(#1316), before that PR fixed a false-positive/silent-truncation bug in
QdrantStore.iter_all(): an empty scroll page with a still-advancing cursor
(e.g. a window landing entirely on tombstoned points) was treated as the
end of the collection instead of continuing. Syncing qdrant_store.py,
vector_store.py, and their tests to #1316's current tip (fa967983) so this
branch doesn't reintroduce the already-fixed bug once merged. Content-only
sync of the 4 shared files (verified via diff against origin/feat/vector-store-iter-all)
rather than a full branch merge, to avoid pulling in unrelated main drift
that has landed on that branch since this one diverged.
`FAISSIndex.save()` previously wrote only the raw FAISS index. `load()` then rebuilt the wrapper with empty `vector_ids` and `metadata`, so that state was never restored.
That made a save/load round trip effectively unusable through the wrapper API: `scan_vectors()` returned no vectors, `count()` returned `0`, and `get_vector(id)` returned `None` for IDs that were present in the underlying FAISS index.
This also affected migration. `semantica store migrate --from faiss` could load a valid FAISS index, see zero vectors through `scan_vectors()`, migrate nothing, and still exit successfully. Since FAISS is a supported migration source, this was a silent data-loss path rather than just a persistence bug.
The fix adds a `.meta.json` sidecar next to the FAISS binary. It stores:
* `vector_ids`
* `metadata`
* `dimension`
* `index_type`
The sidecar is written atomically using a temporary file and rename. `save()` also serializes the metadata before writing the FAISS binary, so a serialization error fails before either persistence artifact is created. That avoids leaving a valid-looking index file behind without the metadata needed to use it correctly.
On load, `dimension` and `index_type` come from the sidecar rather than the caller's arguments. This makes the reconstructed wrapper reflect the index that was actually saved instead of relying on the caller to provide matching values.
There are also explicit checks for incomplete or inconsistent persisted state. If the sidecar is missing, which can happen with indexes written by older versions or when only the FAISS binary was copied, `load()` emits a `RuntimeWarning` instead of silently returning an apparently usable wrapper with no IDs or metadata. If the number of saved vector IDs doesn't match the FAISS index's `ntotal`, `load()` raises `ProcessingError` rather than returning a state where vectors exist in FAISS but can't be reached through `scan_vectors()`.
Metadata serialization changed during review as well. The first version used `json.dumps(..., default=str)`. That avoided failures for values such as `datetime`, `UUID`, and `set`, but it was lossy: those values came back as strings instead of their original Python types.
That was replaced with a tagged encoder/decoder that preserves the supported types across a round trip. It currently handles sets, datetimes, dates, UUIDs, NumPy scalars and arrays, and bytes, with bytes stored as base64.
The decoder also uses an exact-schema check for tagged values. A normal dictionary that happens to contain a reserved tag key alongside other fields is left alone instead of being interpreted as an encoded type.
The final implementation was spread across fourteen commits, mostly following review feedback. Those changes included cleaning up conflict markers from an unfinished stash pop, expanding round-trip and retry coverage, adding the missing-sidecar warning, adding an end-to-end `scan_vectors()` persistence test, replacing lossy metadata serialization with the tagged format, checking FAISS/sidecar count mismatches, adding `bytes` support, and reordering `save()` so metadata serialization happens before the FAISS index is written.
`get_collection()` attached any collection right after the existence check, with no look at its schema. A collection with an INT64 primary key, or one missing the `metadata` field entirely, would attach without complaint and only fail later, inside `get_vector()` or `get_metadata()`, with an error that gave no hint the real problem was upstream at attach time.
This adds a schema check between the attach and the assignment to `self.collection`, so a mismatch is caught at the point of failure instead of surfacing three calls later as an unrelated-looking error. The check validates against exactly the shape `create_collection()` builds: a `VARCHAR` primary key named `id` with `auto_id=False`, a `FLOAT_VECTOR` field named `vector`, and a `JSON` field named `metadata`. Anything else, wrong dtype, wrong name, a missing field, or an auto-generated id, is rejected before the store ever holds a reference to it.
The auto_id and metadata-dtype checks were added in a second pass after review. A collection with `auto_id=True` still attached cleanly and only broke once the store tried to insert with the explicit ids it always sends, and a `metadata` field that existed but wasn't `JSON`-typed only broke during a later write or metadata filter, for the same reason: schema drift that looked fine at attach time and failed downstream instead of at the source.
Nine tests cover this: the one matching-schema case that should succeed, and each rejection path independently, wrong pk dtype, missing pk, wrong pk name, auto_id pk, missing vector field, wrong vector dtype, missing metadata field, and non-JSON metadata.
Closes#1331.
vector_store_config.get_all() always includes a "dimension" key, so
forwarding it via **config into VectorIndexer(dimension=dimension, **config)
raised "got multiple values for keyword argument 'dimension'" any time the
default index-creation path ran with the default config — including
`semantica embed index`, which is exactly the second half of the #994
quick-start pipeline this PR fixes.
* fix(vector_store): make VectorManager methods work on persistent backends (#855)
maintain_store() and collect_statistics() reached into VectorStore
internals (.vectors/.metadata), which only exist for the inmemory
backend — any persistent backend (FAISS, Qdrant, Pinecone, Milvus,
...) crashed with AttributeError.
Add a public backend-agnostic VectorStore.count() accessor following
the get_vector()/get_metadata() precedent (#843) and the
NotImplementedError-on-unsupported-capability precedent of
_filter_by_metadata() (#848): inmemory counts its dict, persistent
backends delegate to count() when available, and raise
NotImplementedError otherwise. VectorManager methods now go through
count(); maintain_store() keeps the exact inmemory semantics (separate
vector/metadata dict counts) and reports a 1:1 count for persistent
backends, where metadata is stored alongside each vector.
Tests: 10 hermetic unit tests covering inmemory, delegation and the
NotImplementedError path. Core vector_store suite: 40 passed.
* fix(vector_store): raise NotImplementedError when count() unavailable
Address Qodo review findings on #914:
- Persistent backend with no wrapped store no longer silently returns 0
(which masked a missing initialization as an empty, healthy store);
it now raises NotImplementedError like get_vector()/get_metadata().
- A mis-shaped adapter exposing a non-callable 'count' attribute now
surfaces a clean NotImplementedError instead of a TypeError, via a
getattr + callable() capability check.
Adds regression tests for both cases.
* fix(vector_store): implement count() on FAISS/SQLite/PgVector backends (#914)
- FAISSStore.count(): returns len(index.vector_ids); 0 when no index exists yet
- SQLiteVecStore.count(): delegates to get_stats()[vector_count] (SELECT COUNT(*))
- PgVectorStore.count(): delegates to get_stats()[vector_count] (SELECT COUNT(*))
- VectorStore.count(): fix misleading NotImplementedError message; now describes
how to add count() support to a backend adapter rather than claiming only the
inmemory backend can ever support counting
- VectorManager.maintain_store(): split inmemory and persistent paths:
* inmemory: independently reads len(vectors) and len(metadata) and compares
them as an integrity check (original semantics preserved)
* persistent: calls store.count(); returns metadata_count=None because
metadata is co-located with vectors in the backend and cannot be counted
independently; never manufactures metadata_count=vector_count as a vacuous
tautology (#914 Qodo review)
- Tests: rewrite test_vector_manager_persistent.py with 31 tests covering
dispatch logic, inmemory divergence detection, persistent metadata_count=None
invariant, FAISSStore/PgVectorStore via mocks, and SQLiteVecStore via real
in-memory SQLite (skipped when sqlite-vec absent)
* docs(changelog): document VectorManager persistent-backend count fix (#914, closes#855)
Records the VectorStore.count() accessor, the FAISS/SQLite/PgVector
implementations added during review, and the maintain_store()
metadata_count fix (no longer fabricates equality for persistent backends).
---------
Co-authored-by: Sameer6305 <sskadam6305@gmail.com>
Co-authored-by: KaifAhmad1 <kaifahmad087@gmail.com>
Co-authored-by: Mohd Kaif <98801504+KaifAhmad1@users.noreply.github.com>
- pinecone_store: call self.index.describe_index_stats() instead of the
nonexistent self.describe_index_stats(), and use a unit query vector
instead of an all-zero vector so filter_by_metadata() works on
cosine-metric indexes (the library's own default)
- pgvector_store: apply the existing lowercase true/false bool handling
to the list-filter branch too, and use the jsonb ?| operator so
list-valued metadata fields match on intersection instead of being
compared as a single JSON-text blob
- sqlite_vec_store: use json_each() with a json_type guard so list-valued
metadata fields match on intersection, mirroring the in-memory
backend's set-intersection semantics
- faiss_store: filter_by_metadata(limit=0) now returns [] instead of one
result
- milvus_store: reject NaN/Infinity filter values up front with a clear
ValidationError instead of building an invalid expression that gets
silently swallowed
- update the #848 FAISS NotImplementedError test to reflect that FAISS
now implements real filter_by_metadata() (this PR's whole point)
- add regression tests for each fix; sqlite tests run against the real
sqlite-vec extension
- vector_store.save(): use v.tolist() instead of list(v) so numpy float32
vectors round-trip through JSON instead of raising TypeError.
- ontology._fetch_url_sync(): resolve relative Location headers via urljoin
before re-validating (previously any relative redirect was rejected
outright), and close every response instead of leaking the connection
across redirect hops.
- sparql.execute_sparql(): move _build_rdflib_graph inside the handler's
error handling so the graph-size cap returns a clean SparqlResponse
error instead of an unhandled 500.
- add regression tests for all three.
The 1.0 / (1.0 + max(0.0, 1.0 - score)) normalization added in the last
commit clamped every raw score >= 1.0 to an identical 1.0, collapsing
result ranking for dot-product-metric indexes (unbounded), which cosine
(bounded to [-1, 1]) never exercised. Replaced with x/(1+|x|) rescaled
to (0, 1), which is strictly monotonic for any real score.
Also adds regression tests for scores >= 1 and a CHANGELOG entry.
Adds similarity_unavailable marker and warning logs to build_decision_context and explain_decision when a persistent backend (like FAISS) fails to reconstruct a vector. Updates docstrings to explicitly state this degraded-path behavior and guarantees schema stability. Adds regression tests to test vector retrieval failure behavior via caplog and context assertions.
- build_decision_context() and explain_decision(include_paths=True) both
accessed self.vectors directly, which is only initialized for the
inmemory backend, crashing with AttributeError on any persistent
backend (FAISS, Qdrant, Pinecone, etc.). Replaced with self.get_vector()
(#843's backend-agnostic accessor) + an is-not-None check — a verified
1:1 behavioral equivalent for the old 'decision_id in self.vectors'
guard on the inmemory path.
- Found a third, undocumented instance of the same bug during
verification: _filter_by_metadata() also accessed self.metadata/
self.vectors directly. Initial fix silently returned [] for persistent
backends, which was itself a new silent-failure bug (indistinguishable
from a genuine zero-match result). Reconciled to raise
NotImplementedError instead, matching the established precedent from
get_vector()/get_metadata() (#843) for 'backend exists but doesn't
support this operation' — confirmed via full grep of all 7 backend
wrapper classes that none currently implement filter_by_metadata,
so this path was previously dead-code-masked-as-working.
Tests: 14 new tests across two rounds — inmemory behavioral equivalence,
real (non-mocked) FAISS backend regression tests for all three methods,
and explicit coverage proving the NotImplementedError fires with a clear
message rather than the old silent-[] behavior. Full suite: 53 passed,
0 failed, 0 regressions across the 39 pre-existing tests.