Commit Graph
2764 Commits
Author SHA1 Message Date
Harsh Arora 45915e50a3 fix(context): vector_store=False must suppress AgentMemory's internal vector cascade in ErasureCoordinator (#1395)
* fix(erasure): ensure vector_store=False disables internal vector cascade in AgentMemory

* fix(erasure): ensure skip_vector=True does not orphan local vector ID tracking
2026-09-03 04:05:53 +05:00
Zohaib Hassnain b7b60d4a17 docs(quickstart): fix broken code against real APIs (#1401) 2026-09-03 04:05:19 +05:00
Zohaib Hassnain 25d2ea5fe9 docs: update stale latest version claims 2026-09-03 03:50:50 +05:00
Zohaib Hassnain 3c68cd12ad docs(mcp): correct tool count (#1399) 2026-09-03 03:40:52 +05:00
Sameer Kadam 798a7455e4 fix(mcp): complete persistence and setup fixes (#1394)
MCP's stdio transport uses stdout for JSON-RPC framing, so anything else written there corrupts every response after it. The original #1134 bug was progress-tracker output landing on stdout during tool calls that construct a `ContextGraph`, which is exactly what happens on any request that triggers reasoning or extraction. This PR closes out the remaining pieces of that fix: loading now goes through `load_from_file()` instead of the older `load()` path on the root graph, and mutations, `record_decision`, `add_entity`, `add_relationship`, now persist back to `SEMANTICA_KG_PATH` when it's configured, in both MCP server implementations (the root `mcp/` package and the packaged `semantica.mcp_server`), not just one.

Four things came out of review on top of that.

The stdio regression test originally exercised `get_graph_summary`, which doesn't touch the progress tracker at all, so it couldn't have caught the original bug. Swapped it for `run_reasoning`: `Reasoner.infer_with_results()` calls `progress_tracker.start_tracking()` directly, the exact call site that corrupted stdout before, so this is the minimal path that actually proves the fix. The test now spawns a real `python -m mcp` subprocess, sends it a `tools/call` for `run_reasoning`, and asserts every single line on stdout parses as JSON.

Loading a corrupt or unreadable `SEMANTICA_KG_PATH` used to fail silently and fall through to an empty graph, which meant the next mutation would happily save that empty graph over the original file. Both implementations now track whether the initial load actually succeeded. If it didn't, every mutation handler refuses to save and returns an error instead, so a broken file on disk stays broken rather than getting silently replaced with nothing. An empty file is treated differently: that's a fresh destination, not a corrupt one, and starts a normal empty graph without tripping the guard.

`save_to_file` used to `open(path, 'w')` and `json.dump` directly into the destination, so a crash or disk-full error mid-write could leave a truncated file as the only copy of the graph. It now writes to a temp file in the same directory, flushes, fsyncs, and only then `os.replace`s the destination, so the destination is always either the old contents or the new contents, never a partial write. The temp file gets cleaned up if anything fails before the replace.

And since a mutation is applied to the in-memory graph before the save happens, a save failure used to leave the in-memory graph ahead of what's on disk, an entity or decision the client thinks succeeded but that never made it to the file. `record_decision`, `add_entity`, and `add_relationship` all roll back the in-memory mutation now if `save_to_file` raises, so the client-visible state and the persisted state never diverge: either both hold the change or neither does.

104 tests passing across the MCP, persistence, and progress-tracking suites.
2026-09-03 03:35:20 +05:00
Mohd Kaif bd584b7402 Update features list in README
Removed 'Self-Hostable' and 'Auditable' from the features list.
2026-09-02 22:21:41 +05:30
Mohd Kaif a4500f5b20 Merge pull request #1396 from semantica-agi/readme-enterprise-connectors-update
docs: tighten README audience list, add SAP connector mentions, log Salesforce ingestor
2026-09-02 21:50:51 +05:30
KaifAhmad1 bdd12e8ac6 fix: correct JWT auth requirements in changelog, add missing SAP install extra
- CHANGELOG: JWT Bearer requires username too, not just consumer_key + private key
- README: add pip install semantica[ingest-sap] to the install-extras list, which was missing despite SAP appearing in the supported-sources lists

Addresses Qodo review feedback on #1396.
2026-09-02 21:45:24 +05:30
KaifAhmad1 48204d4e02 docs: tighten README audience list, add SAP to connector mentions, log unreleased Salesforce ingestor
- Trim "Who it's for" bullets in README for concision
- Propagate SAP OData connector mentions across README's integration lists (was only in the What's New section)
- Add missing CHANGELOG entry for the unreleased Salesforce ingestor (#1240)
- Remove sample `semantica doctor` output lines from the quickstart snippet
2026-09-02 21:23:29 +05:30
Zohaib Hassnain 1bc873cbbd Merge pull request #1328 from semantica-agi/feat/pinecone-iter-all
feat(vector_store): add pinecone iter_all
2026-09-02 20:07:54 +05:30
KaifAhmad1 4dd88375e1 fix(vector_store): don't short-circuit pinecone iter_all() on an empty page
if not vector_ids: return fired before the continuation token was ever
checked. Pinecone's actual pagination contract is that a scan is only
exhausted when the response carries no pagination token -- a page can
legitimately list zero ids while pagination.next is still set (sparse
or filtered namespaces, eventual-consistency windows on serverless
indexes). This was flagged in review but the fix commit that followed
only addressed the separate repeated-token stall case, not this one.

Reproduced concretely against the unfixed code: a page with data,
followed by an empty page with a live token, followed by a page with
more data -- the last page was silently dropped with no error raised,
exactly the #1083 failure mode (store migrate reporting success after
copying only part of a collection).

Now the empty-page case skips the pointless fetch() call but still
falls through to the same next_token check every other path already
goes through, so a live token continues the scan and only a genuinely
absent token (or one that's stopped advancing) ends it.

Added test_continues_past_an_empty_page_with_a_live_token, the
"empty page + non-None next token" case the original review asked for
and that wasn't otherwise covered.
2026-09-02 19:54:24 +05:30
Mohd Kaif fc899c6966 Merge pull request #1326 from semantica-agi/feat/milvus-iter-all
feat(vector_store): add milvus iter_all
2026-09-02 19:37:54 +05:30
KaifAhmad1 98bd632585 Merge remote-tracking branch 'origin/main' into feat/milvus-iter-all 2026-09-02 19:20:55 +05:30
KevinandSameer Kadam 30a91a3a78 feat(evals): add per-metric objective support to runner (closes #1091) (#1092)
* chore: ignore .worktrees directory

* feat(evals): add eval metric and result models

* feat(evals): add evaluator registry

* feat(evals): add exact/regex/range/length evaluators

* feat(evals): add keyword/levenshtein/rouge/llm-as-judge evaluators

* feat(evals): add decision_scores composite evaluator

* feat(evals): add evaluation runner

* feat(evals): expose public API and module proxy

* fix(evals): resolve __all__ names and repair usage example

* docs(evals): add usage docs and changelog entry

* style(evals): tidy evaluator metadata and wiring comments

* fix(evals): honor expected arg and classify error metrics

* fix(evals): export get_evaluator and fix shared meta default

* fix(evals): guard provenance check against non-dict metadata

* docs: add objective layer design spec for semantica.evals

* docs: refine objective spec for consistency with AIP Evals semantics

* docs: add implementation plan for evals objective layer

* docs: fix plan tests to use module-level pytest import

* feat(evals): add per-metric objective support to runner

* docs(evals): document per-metric objectives

* docs(evals): fix minimize example threshold to demonstrate pass

* fix(evals): validate objective config shape strictly

* docs(evals): clarify objective examples and Boolean semantics

* fix(evals): honor direction-only minimize, fail fast on objectives, deep-merge case config

- minimize without threshold is now a no-op, matching maximize (issue #1091
  requires thresholds to be optional for both directions)
- objective config is parsed for every case before any target_fn/evaluator
  runs, so an invalid per-case objective rejects the run up front
- per-case evaluator config deep-merges over the global config so a case
  that overrides one setting keeps the run-level objective
- regression tests for all three, plus updated docs/CHANGELOG

Addresses 3 of 4 Qodo findings on #1092 (the 4th, 'result models defined
twice', is a false positive: types live in types.py)

* fix: finalize eval objectives review

---------

Co-authored-by: Sameer Kadam <sskadam6305@gmail.com>
2026-09-02 18:47:28 +05:30
Mohd Kaif 23126106a3 Merge pull request #1317 from semantica-agi/feat/weaviate-iter-all
feat(vector_store): add weaviate iter_all
2026-09-02 18:39:57 +05:30
Mohd Kaif 4b001b4c9d Merge branch 'main' into feat/weaviate-iter-all 2026-09-02 18:29:14 +05:30
KaifAhmad1 6c9eb2296d fix(vector_store): don't treat an empty weaviate page as end of scan
iter_all() unconditionally returned on any empty fetch_objects() page,
regardless of pagination mode. That's safe for offset/single_page (an
empty page there is a direct, unambiguous statement about live rows),
but not for cursor mode: `after` has no server-issued continuation
value of its own, it's derived client-side from the last object's uuid,
so an empty page gives nothing to advance it with. If Weaviate's cursor
walks internal storage position rather than strict uuid order, a batch
can in principle land entirely on a gap (e.g. tombstoned objects) with
live data past it -- the same risk already confirmed and fixed for
Qdrant's scroll cursor in #1316. Reproduced concretely against the
pre-fix code: a full page followed by an empty page followed by a page
with real data silently dropped that last page with no error raised.

iter_all() now falls back to offset pagination once when a cursor-mode
page comes back empty, rather than assuming that's the end. Offset
addresses live rows directly by position and has no equivalent gap, so
an empty page there (or in single_page mode) is trustworthy and still
ends the scan immediately.

Also updates test_iter_all_empty_collection_yields_nothing and
test_iter_all_requests_vectors, which needed a second empty page now
that a genuinely empty collection takes two calls (cursor, then the
confirming offset check) to report as such.
2026-09-02 17:43:52 +05:30
KaifAhmad1 1ad17beaf6 fix(vector_store): sync qdrant iter_all() with #1316's stall-guard fix
This branch was forked from an earlier commit of feat/vector-store-iter-all
(#1316), before that PR fixed a false-positive/silent-truncation bug in
QdrantStore.iter_all(): an empty scroll page with a still-advancing cursor
(e.g. a window landing entirely on tombstoned points) was treated as the
end of the collection instead of continuing. Syncing qdrant_store.py,
vector_store.py, and their tests to #1316's current tip (fa967983) so this
branch doesn't reintroduce the already-fixed bug once merged. Content-only
sync of the 4 shared files (verified via diff against origin/feat/vector-store-iter-all)
rather than a full branch merge, to avoid pulling in unrelated main drift
that has landed on that branch since this one diverged.
2026-09-02 17:39:42 +05:30
Zohaib Hassnain 110f6deb1e Merge pull request #1316 from semantica-agi/feat/vector-store-iter-all
feat(vector_store): add iter_all enumeration for cursor-based backends
2026-09-02 17:37:11 +05:30
Zohaib Hassnain b8299b1427 chore: clean it 2026-09-02 17:19:19 +05:30
Zohaib Hassnain fa967983e6 fix(vector_store): dedupe qdrant record conversion, don't abort iter_all on a live cursor with an empty page 2026-09-02 17:19:19 +05:30
Zohaib Hassnain bbd423c50a fix(vector_store): raise on stalled pinecone pagination instead of truncating 2026-09-02 17:19:19 +05:30
Zohaib Hassnain fcdad56893 making it clean 2026-09-02 17:19:19 +05:30
Zohaib Hassnain 6b36379f15 feat(vector_store): add pinecone iter_all 2026-09-02 17:19:19 +05:30
Zohaib Hassnain f2e7b9ed75 fix(vector_store): raise instead of truncating when a qdrant scan cannot advance 2026-09-02 17:19:19 +05:30
Zohaib Hassnain 5e80ebd837 fix(vector_store): drop qdrant migrate wiring, keep iter_all only
VectorStore cannot actually migrate to or from qdrant yet. _init_backend_store constructs QdrantStore without connecting or selecting a collection, so reads raise a Collection not initialized error, and the facade store_vectors dispatches only to add/add_vectors while QdrantStore exposes insert_vectors, so writes raise NotImplementedError.

Both are pre-existing facade gaps that nothing had exposed, since migrate previously only allowed faiss/sqlite/pgvector. Adding qdrant to the allowlist claimed support that does not work end to end, so it is removed along with the dimension inference that only fires for backends missing a .dimension attribute. Tracked separately; this PR keeps just the iter_all primitive.
2026-09-02 17:19:19 +05:30
Zohaib Hassnain d175f894a4 feat(vector_store): add iter_all enumeration and wire up qdrant migration 2026-09-02 17:19:19 +05:30
Mohd Kaif 3d32254b07 Merge pull request #1390 from semantica-agi/fix/security-scan-pip-audit-migration
fix(ci): migrate security-scan from Safety to pip-audit
2026-09-02 16:39:01 +05:30
KaifAhmad1 b5199ae6e3 fix(ci): handle pip-audit skipped dependencies, restore manual trigger, fix stale docs
Addresses review feedback on this PR:

- Guard 2 and the PR-comment JS parser both required every dependency
  in pip-audit's report to carry an array-valued `vulns` field. A
  dependency pip-audit can't resolve/audit is reported instead as
  {"name": ..., "skip_reason": ...} with no `vulns` key at all (see
  pip_audit._format.json.JsonFormat._format_dep) - a normal, documented
  shape, not a malformed one. That made a single unauditable package
  hard-fail the whole job and show "Invalid report structure" in the PR
  comment, reintroducing the same class of scan-unrelated CI break this
  migration was meant to fix for Safety. Both now accept skipped
  entries, treat them as zero vulns, and surface them explicitly (job
  log + PR comment) instead of silently dropping or crashing on them.
  Verified the fixed jq queries and JS parse logic against synthetic
  pip-audit report fixtures covering the normal, skipped, and malformed
  shapes.

- Restored a `workflow_dispatch` trigger on security-scan.yml. Deleting
  security.yml (which had it) left no way to manually run a dependency
  audit on demand.

- Updated SECURITY.md, which still described security.yml as a live
  scanning workflow and Safety as an active scanner after this PR
  deletes both.
2026-09-02 16:22:16 +05:30
Mohd Kaif db48f73755 Merge branch 'main' into fix/security-scan-pip-audit-migration 2026-09-02 16:01:17 +05:30
Zohaib Hassnain 07113d2d2d fix(ci): migrate security scan from Safety to pip audit 2026-09-02 15:21:36 +05:00
Shubham SrivastavaandSameer Kadam 909ccf0ded test: install extractor dispatch mocks per test, not at module scope (#1337)
The module assigned MagicMocks into sys.modules at import time and never
removed them. pytest imports every test module during collection before
running anything, so those mocks were live while later modules were
imported and each bound them into its own globals.

132 tests passed alone and failed in a full-suite run as a result. Full
suite goes from 199 failed / 5506 passed to 67 failed / 5638 passed.

A tearDownModule cannot fix this: collection has already finished by the
time it runs. The extractors resolve 'from .methods import
get_entity_method' lazily inside their methods, so the stand-in only has
to be in sys.modules while a test executes - it is now installed per test
via patch.dict in setUp and removed by addCleanup.

Co-authored-by: Sameer Kadam <sskadam6305@gmail.com>
2026-09-02 15:31:38 +05:30
Guofang.Tang fb69b033be fix(ontology): resolve endpoints in direct property inference (#1229)
The relationship-endpoint fix merged in #1170 covers the main ontology generation pipeline, but the public property-inference path still had the same gap.

`OntologyGenerator.infer_properties()`, and the `PropertyGenerator` it delegates to, fell back to `owl:Thing` for both domain and range when a relationship used entity IDs or aliases instead of explicit `source_type` / `target_type` values. The pipeline resolved those endpoints correctly, but the public API path did not.

This moves the existing alias-building and endpoint-resolution logic out of `OntologyGenerator` and into a shared `relationship_utils` module:

* `build_entity_aliases`
* `get_relationship_endpoint`
* `resolve_relationship_endpoint_type`

`OntologyGenerator` now uses those shared helpers instead of keeping its own copies.

`PropertyGenerator._infer_object_properties()` now also receives the entity list, builds the same alias index, and uses the shared endpoint resolver. This replaces the old fallback:

```python
rel.get("source_type") or self._infer_class_from_entity(...)
```

which could only fall back to `owl:Thing` because `_infer_class_from_entity()` never actually resolved an entity.

There are two small behavior changes from centralizing the logic. `build_entity_aliases()` now converts `entity_type` to `str` before adding it to the alias set, avoiding mixed-type alias values. `resolve_relationship_endpoint_type()` returns `None` rather than `""` when there is no usable explicit type, since an empty string isn't a meaningful endpoint type.

The new `test_public_infer_properties_resolves_id_endpoints` covers the broken public API path directly. It creates entities and ID-based relationships through `infer_classes()` / `infer_properties()` and verifies that the inferred `worksFor` property resolves to `Person` for the domain and `Organization` for the range instead of falling back to `owl:Thing`.

That exercises the same endpoint-resolution behavior already covered by the pipeline tests, but through the public entry point that was still missing it.
2026-09-02 14:26:13 +05:00
Guofang.Tang c10090dc9b fix(ci): fail closed on malformed Safety reports (#1366)
The Security Scan workflow already scans `requirements-ci.txt` directly, but malformed Safety output could still be treated as a clean scan. If the report existed on disk but `vulnerabilities` was missing, `null`, or the wrong type, the workflow could end up counting it as zero findings.

This adds a structural check immediately after the report is written. `vulnerabilities` must be an array; otherwise the step fails closed with a clear error instead of treating a broken report as a successful scan.

There was a related problem in the PR reporting path. The comment step already knew how to render an `Invalid report structure` warning, but that branch was effectively unreachable. In GitHub Actions, a custom `if:` is implicitly gated by `success()` unless it includes a status function such as `always()` or `failure()`. Once the Safety step exited non-zero, the Upload and Comment steps were skipped, so the warning could never be posted.

Fixing that required changing how Safety failures flow through the job rather than just adding another guard. The Safety step now uses `continue-on-error: true`, which lets Bandit and Semgrep continue running and allows the Upload and Comment steps to process the failed or malformed Safety result.

Because `continue-on-error` means the Safety step no longer carries the job's final failure signal itself, the workflow now tracks that state explicitly with `SAFETY_SCAN_STATUS`. It is set to `failed` at the start of the Safety step, before any validation runs, and changes to `passed` only when the report is valid and contains zero vulnerabilities.

That default-failed behavior covers every other exit path: a missing report, malformed `vulnerabilities` field, invalid vulnerability count, Safety failure, or an actual vulnerability finding all leave the status as `failed`.

A final `Enforce Safety Gate` step checks `SAFETY_SCAN_STATUS` and fails the job unless it is exactly `passed`. This keeps the same merge-blocking behavior while still allowing the rest of the security checks and reporting steps to run after a Safety failure.

This is a follow-up to #1356. The overlapping Safety behavior changes and duplicate pip-audit path from that PR were dropped after `main` picked up the canonical fix for the underlying `cuda-toolkit` crash. This change keeps only the report-validation hardening that remains independent of that fix.
2026-09-02 14:12:33 +05:00
Zohaib Hassnain 28c96c2539 fix(ci): update actions/deploy-pages pin to current v5 (v5.0.1) (#1387) 2026-09-02 13:59:42 +05:00
Ahmad Bilal 170b4215a6 fix(vector_store): persist vector_ids and metadata across FAISS index save/load (#1272) (#1314)
`FAISSIndex.save()` previously wrote only the raw FAISS index. `load()` then rebuilt the wrapper with empty `vector_ids` and `metadata`, so that state was never restored.

That made a save/load round trip effectively unusable through the wrapper API: `scan_vectors()` returned no vectors, `count()` returned `0`, and `get_vector(id)` returned `None` for IDs that were present in the underlying FAISS index.

This also affected migration. `semantica store migrate --from faiss` could load a valid FAISS index, see zero vectors through `scan_vectors()`, migrate nothing, and still exit successfully. Since FAISS is a supported migration source, this was a silent data-loss path rather than just a persistence bug.

The fix adds a `.meta.json` sidecar next to the FAISS binary. It stores:

* `vector_ids`
* `metadata`
* `dimension`
* `index_type`

The sidecar is written atomically using a temporary file and rename. `save()` also serializes the metadata before writing the FAISS binary, so a serialization error fails before either persistence artifact is created. That avoids leaving a valid-looking index file behind without the metadata needed to use it correctly.

On load, `dimension` and `index_type` come from the sidecar rather than the caller's arguments. This makes the reconstructed wrapper reflect the index that was actually saved instead of relying on the caller to provide matching values.

There are also explicit checks for incomplete or inconsistent persisted state. If the sidecar is missing, which can happen with indexes written by older versions or when only the FAISS binary was copied, `load()` emits a `RuntimeWarning` instead of silently returning an apparently usable wrapper with no IDs or metadata. If the number of saved vector IDs doesn't match the FAISS index's `ntotal`, `load()` raises `ProcessingError` rather than returning a state where vectors exist in FAISS but can't be reached through `scan_vectors()`.

Metadata serialization changed during review as well. The first version used `json.dumps(..., default=str)`. That avoided failures for values such as `datetime`, `UUID`, and `set`, but it was lossy: those values came back as strings instead of their original Python types.

That was replaced with a tagged encoder/decoder that preserves the supported types across a round trip. It currently handles sets, datetimes, dates, UUIDs, NumPy scalars and arrays, and bytes, with bytes stored as base64.

The decoder also uses an exact-schema check for tagged values. A normal dictionary that happens to contain a reserved tag key alongside other fields is left alone instead of being interpreted as an encoded type.

The final implementation was spread across fourteen commits, mostly following review feedback. Those changes included cleaning up conflict markers from an unfinished stash pop, expanding round-trip and retry coverage, adding the missing-sidecar warning, adding an end-to-end `scan_vectors()` persistence test, replacing lossy metadata serialization with the tagged format, checking FAISS/sidecar count mismatches, adding `bytes` support, and reordering `save()` so metadata serialization happens before the FAISS index is written.
2026-09-02 13:44:42 +05:00
Zohaib Hassnain 5f600a3f36 fix(ci): ignore SFTY-20260723-60537 (CVE-2026-65918) in torchvision, unreachable transitive dep (#1385) 2026-09-02 13:37:03 +05:00
Mohd Kaif 1d18755a4e Delete cookbook/advanced/13_Manual_Ontology_Snowflake_Mapping.ipynb 2026-09-02 13:54:13 +05:30
Mohd Kaif 3acf801273 Merge pull request #1361 from taoche/fix/semantic-layer-basics-intro
docs(cookbook): rewrite Semantic Layer Basics as an introductory workflow
2026-09-02 13:22:42 +05:30
KaifAhmad1 796f181c75 docs(cookbook): fix stale RDFExporter claim, pin oxigraph install version
Step 5 explicitly avoids RDFExporter's compact projection and builds
the Turtle export from the TripletStore's own triples instead, but the
Summary cell still credited RDFExporter -- a leftover from before the
rdflib-based export replaced it. Correct the claim to match the code.

Also pin the install to >=0.6.7: earlier releases could return
ontology classes with an empty uri (#1103), which made entity_type
mappings silently resolve to None instead of raising, so the notebook
would appear to pass while never actually typing its instances.
2026-09-02 13:15:14 +05:30
Mohd Kaif 8e73ed8d4c Merge branch 'main' into fix/semantic-layer-basics-intro 2026-09-02 12:44:19 +05:30
Mohd Kaif 114641b39d Merge pull request #1359 from taoche/fix/cookbook-08-end-to-end
docs(cookbook): make notebook 08 a real rerunnable KG workflow
2026-09-02 12:26:00 +05:30
Mohd Kaif f71b711205 Merge branch 'main' into fix/cookbook-08-end-to-end 2026-09-02 12:11:43 +05:30
Mohd Kaif f9a661a4ed security(deps-dev): bump browserslist from 4.28.2 to 4.28.8 in /explorer (#1382)
Fixes GHSA-73wf-gq98-2v4g (prototype pollution / DoS via unguarded
browserslist-stats.json parsing) and GHSA-c83g-rgw3-j3cx (unbounded
cache growth leading to OOM), both patched upstream in 4.28.7.
2026-09-02 11:49:55 +05:30
Kevin 8e7aaee4f5 fix(vector-store): validate collection schema in MilvusStore.get_collection (#1344)
`get_collection()` attached any collection right after the existence check, with no look at its schema. A collection with an INT64 primary key, or one missing the `metadata` field entirely, would attach without complaint and only fail later, inside `get_vector()` or `get_metadata()`, with an error that gave no hint the real problem was upstream at attach time.

This adds a schema check between the attach and the assignment to `self.collection`, so a mismatch is caught at the point of failure instead of surfacing three calls later as an unrelated-looking error. The check validates against exactly the shape `create_collection()` builds: a `VARCHAR` primary key named `id` with `auto_id=False`, a `FLOAT_VECTOR` field named `vector`, and a `JSON` field named `metadata`. Anything else, wrong dtype, wrong name, a missing field, or an auto-generated id, is rejected before the store ever holds a reference to it.

The auto_id and metadata-dtype checks were added in a second pass after review. A collection with `auto_id=True` still attached cleanly and only broke once the store tried to insert with the explicit ids it always sends, and a `metadata` field that existed but wasn't `JSON`-typed only broke during a later write or metadata filter, for the same reason: schema drift that looked fine at attach time and failed downstream instead of at the source.

Nine tests cover this: the one matching-schema case that should succeed, and each rejection path independently, wrong pk dtype, missing pk, wrong pk name, auto_id pk, missing vector field, wrong vector dtype, missing metadata field, and non-JSON metadata.

Closes #1331.
2026-09-02 00:58:26 +05:00
Zohaib Hassnain af829f5f20 fix(vector_store): don't let iterator close() mask the real scan error, dedupe milvus result shaping, split unavailable/uninitialized messages 2026-09-01 23:09:53 +05:00
Zohaib Hassnain 930e7f9b71 docs(vector_store): note the milvus schema assumption 2026-09-01 23:09:53 +05:00
Zohaib Hassnain e335971dcd feat(vector_store): add milvus iter_all 2026-09-01 23:09:53 +05:00
Zohaib Hassnain 78682076d5 fix(vector_store): raise before yielding on a stalled weaviate cursor, extract v4 dict vectors, dedupe fallback ladder 2026-09-01 23:03:07 +05:00
Zohaib Hassnain 1227947be5 fix(vector_store): dedupe qdrant record conversion, don't abort iter_all on a live cursor with an empty page 2026-09-01 22:51:32 +05:00