Compare commits

...
Author SHA1 Message Date
KaifAhmad1 bf292ccbbc fix(docker): address terrascan findings, drop apt-get upgrade
Two terrascan/GHAS findings on the previous commit:
- AC_DOCKER_0052 (no apt-get upgrade in Dockerfiles): dropped it. It also
  wasn't fixing anything - Debian's openssl fix for CVE-2026-14456 is still
  in trixie-proposed-updates, not reachable via a normal upgrade. Pin both
  base images by digest instead (matches #1329's approach) so the docker
  Dependabot ecosystem bumps them once Debian ships a rebuilt image with
  the fix, and document why the QUIC DoS isn't reachable here regardless
  (HTTP-only via uvicorn).
- AC_DOCKER_0010 (pin pip package versions): setuptools was `>=78.1.1`;
  pinned to the exact 84.0.0 already used by pyproject.toml/requirements-ci.txt.

Also fixes two build breaks this introduces on its own: requirements-ci.txt
wasn't in .dockerignore's allowlist or container-scan.yml's path trigger,
so the COPY in the prior commit would have failed the image build outright.
2026-08-31 14:06:28 +05:30
KaifAhmad1 64d942503b fix(docker): extract requirements-ci.txt pins with Python instead of sed
The sed expression to strip requirements-ci.txt's line-continuation
backslash (`[\]$`) is valid POSIX/GNU sed - verified it exits 0 and
strips correctly - but it's easy to misread as broken (a bot reviewer
flagged it as an unterminated bracket expression), and the seemingly
more obvious `\$`/` \$` forms silently fail to match at all rather
than erroring. Swap to a small `re.findall` extraction so there's no
backslash-escaping judgment call left for a reader (bot or human) to
second-guess.
2026-08-31 14:06:02 +05:30
KaifAhmad1 471a420b10 fix(docker): resolve Trivy-flagged CVEs in the built image
Container Security Scan flagged five HIGH-severity findings against
semantica:scan:
- setuptools 70.3.0 (CVE-2025-47273, path traversal) - the base image's
  bundled copy, never touched by our own build. Upgraded explicitly.
- msgpack 1.1.2 (GHSA-6v7p-g79w-8964, OOB read/crash) - `pip install
  ".[explorer]"` re-resolved deps from scratch instead of reusing the
  audited, hash-pinned requirements-ci.txt (which already pins
  msgpack==1.2.1), so it landed on an unpatched transitive version. Now
  installs against a constraints file derived from requirements-ci.txt.
- openssl / libssl3t64 / openssl-provider-legacy (CVE-2026-14456, QUIC
  server DoS) - the Debian fix is still in trixie-proposed-updates, not
  yet promoted to trixie-security, so it can't be pulled via apt today.
  Added an apt upgrade step so the next image rebuild picks it up
  automatically once Debian ships it; documented why this image isn't
  actually exposed to it in the meantime (HTTP-only via uvicorn, no QUIC
  listener).
2026-08-31 14:06:02 +05:30
Mohd KaifandSameer6305 a4aa71ad87 fix(ci): unblock py3.9 install matrix and raise Scorecard pinning/signing (#1329)
* fix(ci): unblock py3.9 install matrix and raise Scorecard pinning/signing

pip install semantica failed on Python 3.9 across all three OSes because
spacy had no upper bound, so pip resolved spacy 3.8.16 whose thinc>=8.3.12
requirement has no cp39 wheels and no working sdist build path. Cap
spacy/thinc for python_version < '3.10' to the last wheel-compatible pair.

Also addresses the two OpenSSF Scorecard findings that were actually
fixable in code:
- Pinned-Dependencies: Dockerfile base images (node:26-alpine,
  python:3.13-slim) were unpinned by digest; pin both, and pin five
  previously-unversioned pip install calls in CI (build, safety, bandit,
  semgrep, jq, pip-audit).
- Signed-Releases: attest-build-provenance only publishes to the GH
  attestations API, which Scorecard doesn't inspect. Sign dist/* with
  Sigstore and attach the .sigstore.json bundles as release assets.

* fix(ci): correct Sigstore artifact inputs

---------

Co-authored-by: Sameer6305 <sskadam6305@gmail.com>
2026-08-31 14:04:28 +05:30
Mohd Kaif fa87a1a9be ci: add npm Dependabot ecosystem and container image scanning (#1286)
* ci: add npm Dependabot ecosystem and container image scanning

- dependabot.yml had no npm ecosystem entry for explorer/, so its
  lockfile was never watched - exactly why the brace-expansion/nanoid
  CVEs fixed in #1280 went undetected. Add it, mirroring the existing
  pip entry's schedule/labels/reviewer conventions.
- New container-scan.yml builds the root Dockerfile's image and scans
  it with Trivy (CRITICAL/HIGH OS+lib CVEs, SARIF to the Security tab)
  and Syft (SPDX SBOM artifact), on push to main, weekly, and manual
  dispatch. Neither the base-image scan nor an SBOM existed before -
  Dependabot's docker entry only bumps the base image tag, it doesn't
  scan built layers.
- Trivy runs report-only for now (no exit-code gate): this is its
  first run against the image, so the CRITICAL/HIGH baseline hasn't
  been triaged yet. Once reviewed, add exit-code: '1' to make it a
  hard gate, same as Safety/Bandit-HIGH in security-scan.yml.

* fix: run Trivy via digest-pinned image, not the aquasecurity/trivy-action wrapper

verify-action-pins.sh failed in CI: the aquasecurity GitHub org has an IP
allow list on its API that 403s the live tag->SHA resolution from
Actions-runner IPs (confirmed reproducible, not transient - resolves fine
from a non-blocked host). Rather than carve a skip exception into the pin
verifier for an org this script already flags as a past tag-repointing
target (see its "LiteLLM/Trivy 2026 incident" comment), pull Trivy as a
sha256-digest-pinned Docker Hub image instead. A digest is immutable and
verifiable independently of GitHub's API entirely, so it sidesteps the
IP block without weakening verification of the one action this repo
already treats as higher-risk. Confirmed the pinned digest
(aquasec/trivy@sha256:62b1e65e...) resolves live against Docker Hub's
registry API.

* fix: match container-scan.yml's push paths to what actually reaches the image

The path filter only watched explorer/package.json and package-lock.json,
but Dockerfile COPYs the whole explorer/ tree plus README.md, LICENSE, and
MANIFEST.in, and .dockerignore controls all of it. A frontend source change
or a README/LICENSE edit would change the built image without triggering a
scan, silently drifting until the next weekly run. Replace the filter with
exactly .dockerignore's opt-in list.
2026-08-30 21:39:30 +05:30
Mohd Kaif 08c78bfb40 fix(security): resolve Scorecard vulnerability and token-permission alerts (#1280)
- Bump explorer's brace-expansion (minimatch dep) 5.0.8 -> 5.0.9 and
  nanoid (postcss dep) 3.3.16 -> 3.3.18, fixing GHSA-rgw5-rvv9-x895 and
  GHSA-2v37-7h3g-55p8 (both DoS via unbounded input, both within the
  existing caret ranges declared by their parents).
- Move codeql.yml and defender-for-devops.yml's security-events: write
  (and codeql.yml's actions: read) from workflow-level down to their
  single job, matching Scorecard's Token-Permissions ideal of a
  read-only top-level default with sensitive scopes granted only where
  used.
2026-08-30 20:58:39 +05:30
Mohd Kaif 56b174781f ci: reusable install action, install-matrix, and release hardening (#1266)
Distribution and trust-signal infrastructure to make pip install semantica
frictionless in downstream CI, and to bring the release pipeline in line
with mature OSS practice.

- .github/actions/setup-semantica: reusable composite action other repos
  can call to install + verify semantica in one step
- install-matrix.yml: verifies the published package installs and imports
  cleanly across Ubuntu/macOS/Windows x Python 3.9-3.12, weekly and on
  release; backs a new README badge
- scorecard.yml: OpenSSF Scorecard analysis, weekly and on push to main,
  backing a new README badge
- release.yml: twine check gate before publish, catching a broken PyPI
  long-description render before it ships
- CITATION.cff: enables GitHub's native "Cite this repository" button
- examples/ci/: copy-paste GitHub Actions, GitLab CI, and CircleCI
  templates for projects adopting semantica
- GROWTH.md: tracked checklist of distribution channels, what's done vs
  outstanding, with guardrails against inflating metrics artificially

Fixes folded in along the way:

- Re-pinned softprops/action-gh-release to the immutable v3.0.3 tag
  instead of the floating v3, after verify-action-pins.sh caught the
  mutable tag had drifted to a newer commit
- setup-semantica now passes extras/version through env vars instead of
  interpolating ${{ inputs.* }} directly into the bash script, closing
  a script-injection vector for callers deriving these from event data
- install-matrix now triggers on the Release workflow's completion
  (workflow_run) instead of release: published, since the GitHub release
  is created before the PyPI upload runs and the old trigger could race
  the publish
- The workflow_run path derives the expected version from the triggering
  tag and passes it into setup-semantica's version input, so pip
  installs and verifies the exact release instead of whatever's latest
  on PyPI at the time
- setup-semantica's pip caching is now opt-in (default disabled), since
  actions/setup-python errors out with cache: 'pip' enabled when the
  caller repo has no requirements.txt/pyproject.toml to key on
- examples/ci/github-actions.yml pins actions/checkout and
  actions/setup-python to verified commit SHAs instead of mutable tags
- examples/ci templates guard the requirements.txt install step with
  -f requirements.txt and call out pyproject.toml/Poetry/Pipenv as
  alternatives, since not every project has a requirements.txt
2026-08-30 17:30:21 +05:00
Zohaib HassnainandSameer Kadam dfda4c561a feat(vector_store): add scan_vectors enumeration and wire up store mi… (#1264)
* feat(vector_store): add scan_vectors enumeration and wire up store migrate

* fix(vector_store): address Qodo finds

* fix(vector_store): make FAISS add_vectors idempotent for retried migrations

---------

Co-authored-by: Sameer Kadam <sskadam6305@gmail.com>
2026-08-30 17:31:57 +05:30
32 changed files with 1295 additions and 36 deletions
+1
View File
@@ -6,6 +6,7 @@
!README.md
!LICENSE
!MANIFEST.in
!requirements-ci.txt
!semantica/
!semantica/**
!integrations/
@@ -0,0 +1,56 @@
name: 'Setup Semantica'
description: 'Install Python, cache pip, and install the semantica package into a workflow'
author: 'Semantica'
inputs:
python-version:
description: 'Python version to set up'
required: false
default: '3.11'
version:
description: 'Version constraint to append to the pip spec, e.g. "==0.6.7" or ">=0.6,<0.7". Leave empty for the latest release.'
required: false
default: ''
extras:
description: 'Comma-separated extras to install, e.g. "explorer,all"'
required: false
default: ''
cache:
description: 'Pip cache mode passed straight to actions/setup-python ("pip" to enable). Left empty (disabled) by default because this action is meant to run standalone in any caller repo, and actions/setup-python errors out if it cannot find a requirements.txt/pyproject.toml/setup.py/poetry.lock to key the cache on. Opt in only when the caller repo has one of those files.'
required: false
default: ''
outputs:
version:
description: 'The installed semantica version'
value: ${{ steps.verify.outputs.version }}
runs:
using: 'composite'
steps:
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7
with:
python-version: ${{ inputs.python-version }}
cache: ${{ inputs.cache }}
- name: Install semantica
shell: bash
env:
SEMANTICA_EXTRAS: ${{ inputs.extras }}
SEMANTICA_VERSION: ${{ inputs.version }}
run: |
python -m pip install --upgrade pip
if [ -n "$SEMANTICA_EXTRAS" ]; then
spec="semantica[$SEMANTICA_EXTRAS]$SEMANTICA_VERSION"
else
spec="semantica$SEMANTICA_VERSION"
fi
python -m pip install -- "$spec"
- name: Verify install
id: verify
shell: bash
run: |
VERSION=$(python -c "import semantica; print(semantica.__version__)")
echo "Installed semantica $VERSION"
echo "version=$VERSION" >> "$GITHUB_OUTPUT"
+23
View File
@@ -101,6 +101,29 @@ updates:
allow:
- dependency-type: "production"
# Explorer frontend (npm)
- package-ecosystem: "npm"
directory: "/explorer"
schedule:
interval: "weekly"
day: "monday"
time: "03:30" # 3:30 AM UTC (9:00 AM IST)
open-pull-requests-limit: 10
reviewers:
- "KaifAhmad1"
assignees:
- "KaifAhmad1"
commit-message:
prefix: "security"
include: "scope"
labels:
- "dependencies"
- "javascript"
- "security"
allow:
- dependency-type: "production"
- dependency-type: "development"
# Docker dependencies (if you use Docker)
- package-ecosystem: "docker"
directory: "/"
+1 -1
View File
@@ -72,7 +72,7 @@ jobs:
diff \
<(grep -E '^[a-zA-Z0-9._-]+==' requirements-ci.txt | sed 's/ \\$//') \
<(grep -E '^[a-zA-Z0-9._-]+==' /tmp/requirements-ci-check.txt)
- run: pip install build
- run: pip install build==1.6.0
# wheel is build-time only (not in requirements-ci.txt) — install the
# same pinned version [build-system] declares so --no-isolation works.
- run: pip install wheel==0.48.0
+4 -2
View File
@@ -10,13 +10,15 @@ on:
permissions:
contents: read
security-events: write
actions: read
jobs:
analyze:
name: Analyze Python
runs-on: ubuntu-latest
permissions:
contents: read
security-events: write # for github/codeql-action/upload-sarif below
actions: read # for github/codeql-action/init's CodeQL bundle cache lookup
steps:
- name: Checkout repository
+74
View File
@@ -0,0 +1,74 @@
name: Container Security Scan
on:
push:
branches: [main]
# Mirrors .dockerignore's opt-in list exactly - anything not listed there
# can't reach the build context, so it can't change the built image.
paths:
- 'Dockerfile'
- '.dockerignore'
- 'pyproject.toml'
- 'README.md'
- 'LICENSE'
- 'MANIFEST.in'
- 'requirements-ci.txt'
- 'semantica/**'
- 'integrations/**'
- 'explorer/**'
- '.github/workflows/container-scan.yml'
schedule:
- cron: '30 2 * * 1' # weekly, catches new CVEs published against the base image between pushes
workflow_dispatch:
permissions:
contents: read
jobs:
scan:
runs-on: ubuntu-latest
permissions:
contents: read
security-events: write # for github/codeql-action/upload-sarif below
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- name: Build image
run: docker build -t semantica:scan .
# Run Trivy as a digest-pinned image rather than the aquasecurity/trivy-action
# marketplace wrapper: the aquasecurity GitHub org has an IP allow list on its
# API that 403s verify-action-pins.sh's live tag->SHA check from Actions-runner
# IPs, and this repo already treats Trivy's action pin as a known past target
# for tag-repointing (see the LiteLLM/Trivy 2026 incident note above). Pulling
# by sha256 digest from Docker Hub is immutable and verifiable independently of
# GitHub's API, so it sidesteps both problems at once instead of carving a skip
# exception into the pin verifier for an org already flagged as higher-risk.
#
# Report-only for now: this is Trivy's first run against this image, so we
# don't yet know the CRITICAL/HIGH baseline. Findings still land in the
# Security tab either way. Once triaged, add `--exit-code 1` (like
# Safety/Bandit-HIGH in security-scan.yml) to make it a hard gate.
- name: Scan image for vulnerabilities (Trivy)
run: |
docker run --rm \
-v /var/run/docker.sock:/var/run/docker.sock \
-v "$PWD:/output" \
aquasec/trivy@sha256:62b1e65e8869bc4b4c6aa4fa2b21595256c7c2f6018a9d9ad61caf87187c1969 \
image --format sarif --output /output/trivy-results.sarif \
--severity CRITICAL,HIGH --ignore-unfixed semantica:scan
- name: Upload Trivy SARIF
if: always()
uses: github/codeql-action/upload-sarif@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4
with:
sarif_file: trivy-results.sarif
category: trivy-container
- name: Generate SBOM (Syft)
if: always()
uses: anchore/sbom-action@3ad7283483fc7af8ff2b4ea19663c2d5ca935e26 # v0.24.2
with:
image: semantica:scan
format: spdx-json
output-file: semantica-sbom.spdx.json
+3 -1
View File
@@ -28,12 +28,14 @@ on:
permissions:
contents: read
security-events: write
jobs:
MSDO:
# currently only windows-latest is supported
runs-on: windows-latest
permissions:
contents: read
security-events: write # for github/codeql-action/upload-sarif below
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
+59
View File
@@ -0,0 +1,59 @@
name: Install Matrix
permissions:
contents: read
on:
schedule:
- cron: '0 6 * * 1' # weekly, catches upstream dependency breakage between releases
workflow_run:
# The Release workflow publishes the GitHub release *before* it uploads to
# PyPI (see release.yml), so triggering on `release: published` would race
# the PyPI upload and could pass by silently installing the prior version.
# workflow_run fires only after the whole Release workflow - including the
# PyPI publish step - has finished.
workflows: ['Release']
types: [completed]
workflow_dispatch:
jobs:
verify-install:
if: github.event_name != 'workflow_run' || github.event.workflow_run.conclusion == 'success'
name: pip install semantica (${{ matrix.os }}, py${{ matrix.python-version }})
runs-on: ${{ matrix.os }}
strategy:
fail-fast: false
matrix:
os: [ubuntu-latest, macos-latest, windows-latest]
python-version: ['3.9', '3.10', '3.11', '3.12']
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- name: Pin expected version for release-triggered runs
id: expected-version
if: github.event_name == 'workflow_run'
shell: bash
env:
EXPECTED_TAG: ${{ github.event.workflow_run.head_branch }}
run: |
expected="${EXPECTED_TAG#v}"
if [ -z "$expected" ]; then
echo "::error::Could not determine a release tag from the triggering workflow run (head_branch was empty)."
exit 1
fi
echo "constraint===$expected" >> "$GITHUB_OUTPUT"
- id: setup-semantica
uses: ./.github/actions/setup-semantica
with:
python-version: ${{ matrix.python-version }}
cache: 'pip'
version: ${{ steps.expected-version.outputs.constraint }}
- name: Smoke test import
shell: bash
run: |
python -c "
import semantica
print('semantica', semantica.__version__, 'installed and importable')
"
+22 -4
View File
@@ -16,7 +16,7 @@ jobs:
cancel-in-progress: false
permissions:
contents: write # for the GitHub Release
id-token: write # for PyPI Trusted Publishing (OIDC) and attestation signing
id-token: write # for PyPI Trusted Publishing (OIDC), attestation signing, and Sigstore
attestations: write # for SLSA build provenance
# If you add another job to this workflow, give it its own explicit
# `permissions:` block rather than relying on the workflow-level default
@@ -40,7 +40,7 @@ jobs:
# build runs against the same versions CI tests against.
- name: Install pinned build dependencies
run: pip install -r requirements-ci.txt
- run: pip install build
- run: pip install build==1.6.0
# wheel is build-time only (not in requirements-ci.txt) — install the
# same pinned version [build-system] declares so --no-isolation works.
- run: pip install wheel==0.48.0
@@ -63,11 +63,29 @@ jobs:
print("Explorer frontend is packaged")
PY
- name: Verify PyPI long-description will render
run: |
pip install twine==7.0.0
twine check dist/*
- name: Attest build provenance
uses: actions/attest-build-provenance@4d101475d8b20a2381f78447822ac1eab6504dd8 # v4
with:
subject-path: 'dist/*'
- uses: softprops/action-gh-release@3d0d9888cb7fd7b750713d6e236d1fcb99157228 # v3
# attest-build-provenance publishes to the GH attestations API only, which
# OpenSSF Scorecard's Signed-Releases check does not inspect - it looks for
# signature files attached as release assets. Sign here too so
# `dist/*.sigstore.json` bundles ship alongside the wheel/sdist on the
# GitHub Release itself.
- name: Sign artifacts with Sigstore
uses: sigstore/gh-action-sigstore-python@790bc6befb9d733738f18d8f895854b453640ec9 # v3.5.0
with:
files: dist/*
inputs: |
dist/*.whl
dist/*.tar.gz
- uses: softprops/action-gh-release@efb35369e0ad2afab669f228072c1b0d510eae64 # v3.0.3
with:
files: |
dist/*.whl
dist/*.tar.gz
dist/*.sigstore.json
- uses: pypa/gh-action-pypi-publish@dc37677b2e1c63e2034f94d8a5b11f265b73ba33 # release/v1
+45
View File
@@ -0,0 +1,45 @@
name: Scorecard supply-chain security
permissions: read-all
on:
branch_protection_rule:
schedule:
- cron: '30 1 * * 6' # weekly
push:
branches: [main]
jobs:
analysis:
name: Scorecard analysis
runs-on: ubuntu-latest
permissions:
security-events: write # to upload SARIF results
id-token: write # to publish results and get a badge
contents: read
actions: read # to detect GitHub Actions workflows
steps:
- name: Checkout code
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
with:
persist-credentials: false
- name: Run analysis
uses: ossf/scorecard-action@2d1146689b8cda280b9bc96326124645441f03bc # v2.4.4
with:
results_file: results.sarif
results_format: sarif
publish_results: true
- name: Upload artifact
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
with:
name: SARIF file
path: results.sarif
retention-days: 5
- name: Upload to code-scanning
uses: github/codeql-action/upload-sarif@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4
with:
sarif_file: results.sarif
+1 -1
View File
@@ -52,7 +52,7 @@ jobs:
# Tooling AFTER the pinned set: installing safety/bandit/semgrep/jq
# first lets the pinned requirements overwrite their transitive deps
# (e.g. rich), which breaks the safety CLI at runtime.
pip install safety bandit semgrep jq
pip install safety==3.8.1 bandit==1.9.4 semgrep==1.175.0 jq==1.12.0
- name: Run Safety Check (Package Vulnerabilities)
run: |
+1 -1
View File
@@ -37,6 +37,6 @@ jobs:
# pyproject.toml changes under review. The schedule/workflow_dispatch
# runs stay non-blocking until a full pass over pre-existing findings
# across the whole [all] tree has been done.
- run: pip install pip-audit
- run: pip install pip-audit==2.10.1
- run: pip-audit -r requirements-ci.txt
continue-on-error: ${{ github.event_name != 'pull_request' }}
+20
View File
@@ -0,0 +1,20 @@
cff-version: 1.2.0
message: "If you use this software, please cite it as below."
title: "Semantica: Graph-Native Infrastructure for Context and Accountable AI Systems"
type: software
authors:
- name: "Semantica"
repository-code: "https://github.com/semantica-agi/semantica"
url: "https://getsemantica.ai"
license: MIT
version: 0.6.7
date-released: 2026-08-28
keywords:
- knowledge-graph
- context-graph
- ai-agents
- llm
- decision-intelligence
- provenance
- explainability
- graph-rag
+30 -4
View File
@@ -1,5 +1,5 @@
# syntax=docker/dockerfile:1
FROM node:26-alpine AS frontend-builder
FROM node:26-alpine@sha256:2d984a15c9b54fd0aeb608b8e0d0d83529eb34d2966db27a1fb4f1edc3d298a3 AS frontend-builder
WORKDIR /app
COPY explorer/package*.json ./explorer/
@@ -9,7 +9,18 @@ RUN npm ci
COPY explorer/ ./
RUN mkdir -p /app/semantica && npm run build
FROM python:3.13-slim AS runtime
# CVE-2026-14456 (OpenSSL QUIC-server DoS, flagged against this base image's
# openssl/libssl3t64/openssl-provider-legacy): the Debian fix
# (3.5.7-1~deb13u2) is only in trixie-proposed-updates as of this writing,
# not yet promoted to trixie-security, so there's no package to pin here
# today. Deliberately NOT running `apt-get upgrade` to chase it - that
# breaks build reproducibility (terrascan AC_DOCKER_0052) and still
# wouldn't reach a proposed-updates-only package. Once Debian ships the fix
# and rebuilds this tag, the docker Dependabot ecosystem in
# .github/dependabot.yml opens a PR bumping the digest pin above. Also: this
# image only serves plain HTTP via uvicorn and never opens a QUIC listener,
# so the bug isn't reachable here regardless.
FROM python:3.13-slim@sha256:7ce4b6dfe35e55397b7cda544f8a13f191b7ae28dc5aad71fe664dbc9bc2623f AS runtime
ENV PYTHONDONTWRITEBYTECODE=1 \
PYTHONUNBUFFERED=1 \
@@ -22,12 +33,27 @@ WORKDIR /app
RUN groupadd --system semantica \
&& useradd --system --gid semantica --home-dir /app --shell /usr/sbin/nologin semantica
COPY pyproject.toml README.md LICENSE MANIFEST.in ./
COPY pyproject.toml README.md LICENSE MANIFEST.in requirements-ci.txt ./
COPY semantica/ ./semantica/
COPY integrations/ ./integrations/
COPY --from=frontend-builder /app/semantica/static ./semantica/static
RUN pip install --no-cache-dir ".[explorer]" \
# The base image ships an outdated setuptools (CVE-2025-47273); upgrade it
# explicitly since nothing in our own dependency tree otherwise pulls a
# newer copy. Pinned to the exact version requirements-ci.txt/pyproject.toml
# already build against, rather than a floor, per terrascan AC_DOCKER_0010.
# requirements-ci.txt itself carries the audited, CVE-checked pins for every
# transitive dependency (see security-scan.yml / security.yml) - feed them
# in as an unhashed constraints file (pip's hash-checking mode rejects the
# unhashable local source directory this installs) so the image lands on
# the same patched versions CI verified, e.g. msgpack>=1.2.1, rather than
# letting pip freely re-resolve and pick up an unpatched transitive version.
# (Extracted with Python's re module rather than sed/grep so there's no
# line-continuation-backslash stripping to get subtly wrong.)
RUN pip install --no-cache-dir "setuptools==84.0.0" \
&& python -c "import re, pathlib; pins = re.findall(r'^([A-Za-z0-9._-]+==\S+)', pathlib.Path('requirements-ci.txt').read_text(), re.M); pathlib.Path('/tmp/constraints.txt').write_text('\n'.join(pins))" \
&& pip install --no-cache-dir -c /tmp/constraints.txt ".[explorer]" \
&& rm -f /tmp/constraints.txt requirements-ci.txt \
&& chown -R semantica:semantica /app
USER semantica
+131
View File
@@ -0,0 +1,131 @@
# Growth & Distribution Playbook
North star: **10,000 developers who actually use Semantica in real projects**, not a raw PyPI download number. Downloads are a lagging indicator of distribution, not a target to optimize directly.
```
GitHub stars → Website visitors → PyPI installs → Weekly active users → Production deployments → Enterprise customers
```
The last two matter far more than the download count.
## Guardrails — do not do this
- No fake/looping CI jobs that repeatedly `pip install semantica` purely to inflate the graph. It's detectable, it produces zero real users, and it damages credibility with anyone doing diligence (investors, enterprise buyers, security reviewers).
- No package-splitting purely to multiply install counts — only split into `semantica-*` packages when there's a real architectural reason.
- No meaningless Docker pulls or notebook launches with no real content behind them.
- Every item below should get someone from "installed it" to "used it for something real." If a channel can't do that, it's not worth building.
## 30-day priority sprint
Ordered by leverage-to-effort ratio; do these first.
| # | Initiative | Target |
| - | ---------- | ------ |
| 1 | ✅ GitHub Actions example + reusable `setup-semantica` composite action + install-matrix badge | done |
| 2 | Google Colab notebooks | 10 |
| 3 | Docker images (RAG, Graph, Agent, API) | 4-5 |
| 4 | Hugging Face Spaces demos | 3-4 |
| 5 | LangChain integration + example | 1 |
| 6 | LlamaIndex integration + example | 1 |
| 7 | Vector/graph DB integrations (Qdrant, Weaviate, Neo4j) | 3 |
| 8 | MCP server + example | 1 (already have `mcp/` — package as a distributable example) |
| 9 | Production-quality starter repos (FastAPI, Streamlit, Gradio) | 3 |
| 10 | `awesome-rag` / `awesome-llm` / `awesome-knowledge-graph` list submissions | 3+ PRs |
Push everything through: GitHub → Discord (`sV34vps5hH`) → X (`@BuildSemantica`) → GitHub Discussions → Reddit → Hacker News → relevant newsletters.
## Full channel checklist
### CI/CD (highest-intent distribution — installs tied to real pipelines)
- [x] GitHub Actions example in `examples/ci/github-actions.yml`
- [x] Reusable composite GitHub Action — [`.github/actions/setup-semantica`](.github/actions/setup-semantica/action.yml), modeled on `actions/setup-python`; usable by any repo as `uses: semantica-agi/semantica/.github/actions/setup-semantica@main`
- [x] "pip install" status badge in the README, backed by [`.github/workflows/install-matrix.yml`](.github/workflows/install-matrix.yml) — verifies the *published* package installs cleanly on Ubuntu/macOS/Windows across Python 3.9-3.12, weekly + on every release
- [x] GitLab CI template — `examples/ci/gitlab-ci.yml`
- [x] CircleCI template — `examples/ci/circleci-config.yml`
- [ ] Jenkins, Azure DevOps, Bitbucket Pipelines, Buildkite, Travis CI equivalents
### Release pipeline hardening (already had Trusted Publishing/OIDC + SLSA attestation — this rounds it out to match top-tier OSS release practice)
- [x] `twine check` gate in `.github/workflows/release.yml` before publish — catches a broken PyPI long-description render before it goes live instead of after (a malformed README on the live PyPI page is a silent conversion killer)
- [x] `CITATION.cff` (see Academic & research below)
- [x] OpenSSF Scorecard (see Discoverability below)
- [ ] Considered and deliberately skipped: Release Drafter / auto-generated changelogs — this repo hand-curates `CHANGELOG.md` with far more detail (PR numbers, contributors, phase-1 limitations) than a bot would produce. Don't introduce this without checking with maintainers first.
- [ ] Renovate / Dependabot config templates that auto-bump the `semantica` version in downstream repos — real recurring CI runs on real adopters
- [ ] Nightly scheduled workflow template that tests a downstream project against `semantica@latest`
### Containers & dev environments
- [ ] Official Docker images: RAG, Graph, Agent, API, `+Postgres`, `+Neo4j`, `+Qdrant`
- [ ] `docker-compose` examples (repo already has `docker-compose.dev.yml` / `docker-compose.yml` as a base)
- [ ] `.devcontainer/devcontainer.json` for one-click "Reopen in Container"
- [ ] GitHub Codespaces-ready config
- [ ] Gitpod config
- [ ] "Use this template" GitHub repo button so new projects start with `semantica` in `requirements.txt`
### Notebooks & hosted demos
- [ ] 10-20 Google Colab notebooks (Graph RAG, agent memory, entity resolution, semantic search, document intelligence)
- [ ] Kaggle Notebooks/Kernels
- [ ] Binder / mybinder.org config for instant repo launch
- [ ] SageMaker Studio Lab / Databricks Community Edition / Paperspace Gradient examples
- [ ] Hugging Face Spaces (Streamlit/Gradio) demos with `semantica` in `requirements.txt`
- [ ] Public hosted playground (source on GitHub, install visible)
### Framework & data-store integrations
- [x] LangChain integration — `integrations/langchain/` (`SemanticaRetriever`, `SemanticaVectorStore`, `SemanticaKGTool`/`SemanticaDecisionTool`), `pip install semantica[langchain]`, shipped in 0.6.7
- [ ] LlamaIndex integration + example
- [ ] LangGraph example
- [ ] Neo4j integration/example (docs already list it as a supported graph store — turn into a runnable example repo)
- [ ] Vector DB examples: Qdrant, Weaviate, Milvus, Pinecone, Chroma, FAISS, pgvector, OpenSearch/Elasticsearch (FAISS/Pinecone/Weaviate/Qdrant/Milvus/PgVector already supported per `docs/community-projects.md` — package each as a standalone example)
- [ ] LLM provider quickstarts: OpenAI, Anthropic, Gemini, Groq, Ollama, HuggingFace, DeepSeek, LiteLLM (already-supported providers per docs — each gets its own copy-paste quickstart)
- [ ] CrewAI / Agno integration examples (already documented under `docs/integrations/`) — promote as standalone repos, not just docs pages
### Package managers & installers
- [ ] conda-forge feedstock
- [ ] Homebrew formula for the CLI
- [ ] Nix/nixpkgs packaging
- [ ] Chocolatey / Scoop (Windows)
- [ ] Document `uv add semantica` and `poetry add semantica` explicitly alongside `pip install`
### Downstream packages & CLI
- [ ] Genuinely useful `semantica-*` packages only where warranted (e.g. `semantica-rag`, `semantica-connectors`) — each pulls `semantica` as a real dependency
- [ ] Make sure `semantica init / ingest / index / query / serve` CLI flows are the default onboarding path in every tutorial
- [ ] VS Code extension wrapping the CLI (scaffold + run commands from the command palette)
- [ ] JetBrains plugin equivalent
### Templates & starters
- [ ] Cookiecutter templates: `cookiecutter-semantic-rag`, `cookiecutter-ai-agent`, `cookiecutter-enterprise-rag`
- [ ] Starter repos: FastAPI, Streamlit, Gradio, Next.js frontend + Semantica backend
- [ ] Cloud deploy templates: AWS, GCP, Azure, Modal, Railway, Render, Fly.io (repo already has `deploy/azure`, `deploy/gcp`, `deploy/fly`, `deploy/railway`, `deploy/render`, `deploy/kubernetes`, `deploy/helm` — link these prominently from the README/quickstart, they're already-built distribution surface)
- [ ] Terraform / Pulumi / Helm modules published to their respective registries
### Discoverability & curation
- [ ] Submit to `awesome-rag`, `awesome-llm`, `awesome-knowledge-graph`, `awesome-python`
- [ ] Pitch newsletters with engaged Python/AI audiences (Python Weekly, Import AI, TLDR AI, etc.)
- [x] PyPI trove classifiers/keywords and `project.urls` (Homepage/Docs/Repository/Changelog/Bug Tracker) — already complete in `pyproject.toml`
- [ ] Get listed on Papers With Code for any retrieval/graph-RAG benchmark work
- [x] [OpenSSF Scorecard](https://scorecard.dev/viewer/?uri=github.com/semantica-agi/semantica) badge + weekly workflow (`.github/workflows/scorecard.yml`) — a concrete trust signal security/procurement teams check before greenlighting adoption, which gates real (non-CI-bot) install growth at enterprises
### Academic & research
- [x] `CITATION.cff` at repo root — enables GitHub's native "Cite this repository" button, feeds Google Scholar/academic tooling; complements `docs/citation.md` (still needs a real Zenodo DOI to replace the `XXXXXXX` placeholder in both places once one is minted)
- [ ] arXiv paper if there's real architectural novelty to describe
- [ ] Zenodo DOI for citability (`docs/citation.md` already exists — make sure it points to a real DOI)
- [ ] Workshop/tutorial sessions at PyData/ODSC-style events with hands-on install steps
- [ ] University course material / bootcamp adoption outreach
### Content
- [ ] Reproducible benchmark repos (Graph RAG vs vector RAG, retrieval@k, enterprise-scale retrieval) with `pip install semantica && python benchmark.py`
- [ ] 20-30 real-world example applications (RAG, enterprise document intelligence, financial entity graphs, code knowledge graphs, research discovery, agent memory)
- [ ] Blog/tutorial posts on Dev.to, Medium, personal blogs — always with runnable code, not just prose
- [ ] Contribute integrations/PRs to other projects building RAG/agents/knowledge graphs — "I implemented Semantica support" beats "please use Semantica"
## Tracking
Don't just watch the raw PyPI number — use download analytics (e.g. PePy) to separate CI/bot traffic from real installs, and track the funnel above end-to-end where possible (stars → site visits → installs → weekly actives).
+15 -1
View File
@@ -26,7 +26,7 @@
#### Built for High-Stakes, Regulated Domains
[![GitHub Stars](https://img.shields.io/github/stars/semantica-agi/semantica?style=flat-square&color=FFD700&logo=github&logoColor=white&label=Stars)](https://github.com/semantica-agi/semantica) [![GitHub Forks](https://img.shields.io/github/forks/semantica-agi/semantica?style=flat-square&color=6E40C9&logo=github&logoColor=white&label=Forks)](https://github.com/semantica-agi/semantica/network/members) [![Contributors](https://img.shields.io/github/contributors/semantica-agi/semantica?style=flat-square&color=2EA043&logo=github&logoColor=white)](https://github.com/semantica-agi/semantica/graphs/contributors) [![PyPI](https://img.shields.io/pypi/v/semantica.svg?style=flat-square&color=0066CC&logo=pypi&logoColor=white)](https://pypi.org/project/semantica/) [![Total Downloads](https://static.pepy.tech/badge/semantica?style=flat-square)](https://pepy.tech/project/semantica) [![Python 3.8+](https://img.shields.io/badge/python-3.8+-3776AB?style=flat-square&logo=python&logoColor=white)](https://www.python.org/) [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg?style=flat-square)](https://opensource.org/licenses/MIT) [![CI](https://img.shields.io/github/actions/workflow/status/semantica-agi/semantica/ci.yml?style=flat-square&label=CI)](https://github.com/semantica-agi/semantica/actions) [![Ask DeepWiki](https://deepwiki.com/badge.svg)](https://deepwiki.com/semantica-agi/semantica)
[![GitHub Stars](https://img.shields.io/github/stars/semantica-agi/semantica?style=flat-square&color=FFD700&logo=github&logoColor=white&label=Stars)](https://github.com/semantica-agi/semantica) [![GitHub Forks](https://img.shields.io/github/forks/semantica-agi/semantica?style=flat-square&color=6E40C9&logo=github&logoColor=white&label=Forks)](https://github.com/semantica-agi/semantica/network/members) [![Contributors](https://img.shields.io/github/contributors/semantica-agi/semantica?style=flat-square&color=2EA043&logo=github&logoColor=white)](https://github.com/semantica-agi/semantica/graphs/contributors) [![PyPI](https://img.shields.io/pypi/v/semantica.svg?style=flat-square&color=0066CC&logo=pypi&logoColor=white)](https://pypi.org/project/semantica/) [![Total Downloads](https://static.pepy.tech/badge/semantica?style=flat-square)](https://pepy.tech/project/semantica) [![Python 3.8+](https://img.shields.io/badge/python-3.8+-3776AB?style=flat-square&logo=python&logoColor=white)](https://www.python.org/) [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg?style=flat-square)](https://opensource.org/licenses/MIT) [![CI](https://img.shields.io/github/actions/workflow/status/semantica-agi/semantica/ci.yml?style=flat-square&label=CI)](https://github.com/semantica-agi/semantica/actions) [![Install Matrix](https://img.shields.io/github/actions/workflow/status/semantica-agi/semantica/install-matrix.yml?style=flat-square&label=pip%20install)](https://github.com/semantica-agi/semantica/actions/workflows/install-matrix.yml) [![OpenSSF Scorecard](https://api.scorecard.dev/projects/github.com/semantica-agi/semantica/badge?style=flat-square)](https://scorecard.dev/viewer/?uri=github.com/semantica-agi/semantica) [![Ask DeepWiki](https://deepwiki.com/badge.svg)](https://deepwiki.com/semantica-agi/semantica)
[![Website](https://img.shields.io/badge/Website-getsemantica.ai-000000?style=flat-square&logo=googlechrome&logoColor=white)](https://getsemantica.ai/) [![Docs](https://img.shields.io/badge/Docs-docs.getsemantica.ai-0099FF?style=flat-square&logo=readthedocs&logoColor=white)](https://docs.getsemantica.ai/) [![Discord](https://img.shields.io/badge/Discord-Join%20Community-5865F2?style=flat-square&logo=discord&logoColor=white)](https://discord.gg/sV34vps5hH) [![Twitter/X](https://img.shields.io/badge/Follow-%40BuildSemantica-000000?style=flat-square&logo=x&logoColor=white)](https://x.com/BuildSemantica) [![YouTube](https://img.shields.io/badge/YouTube-Watch%20Demos-FF0000?style=flat-square&logo=youtube&logoColor=white)](https://www.youtube.com/watch?v=QfnNZg4-dZA) [![Changelog](https://img.shields.io/badge/Changelog-View-6E40C9?style=flat-square&logo=keepachangelog&logoColor=white)](CHANGELOG.md)
@@ -1534,6 +1534,20 @@ git clone https://github.com/semantica-agi/semantica.git
cd semantica && pip install -e ".[dev]" && pytest tests/
```
### CI & Deployment
Wiring `semantica` into your own CI is a two-minute job. On GitHub Actions, use the reusable composite action:
```yaml
- uses: semantica-agi/semantica/.github/actions/setup-semantica@main
with:
python-version: '3.11'
```
Copy-paste starting templates for GitHub Actions, GitLab CI, and CircleCI live in [examples/ci/](examples/ci/). The published package itself is verified installable across Ubuntu/macOS/Windows and Python 3.9-3.12 every week by the [Install Matrix workflow](.github/workflows/install-matrix.yml).
Ready-made deployment configs for AWS, GCP, Azure, Fly.io, Railway, Render, Kubernetes, and Helm are in [deploy/](deploy/).
---
## Enterprise
+36
View File
@@ -0,0 +1,36 @@
# CI templates
Copy-paste starting points for wiring `semantica` into your own project's CI. Each file is a
complete, working config — rename it into your project (see the comment at the top of each file
for the target path) and swap the smoke-test / test step for whatever your project does with
Semantica. Each template installs `semantica` unconditionally and your own project's dependencies
only if a `requirements.txt` is present; if your project uses `pyproject.toml`, Poetry, or Pipenv
instead, adjust the marked install line (each file calls it out inline).
| File | Target path in your repo |
| ---- | ------------------------- |
| [`github-actions.yml`](github-actions.yml) | `.github/workflows/semantica.yml` |
| [`gitlab-ci.yml`](gitlab-ci.yml) | `.gitlab-ci.yml` |
| [`circleci-config.yml`](circleci-config.yml) | `.circleci/config.yml` |
If your own project is hosted on GitHub, you can skip the setup boilerplate entirely and use
Semantica's reusable composite action instead:
```yaml
- uses: semantica-agi/semantica/.github/actions/setup-semantica@main
with:
python-version: '3.11'
# extras: 'explorer,all' # optional
# version: '==0.6.7' # optional, pin an exact release
# cache: 'pip' # optional, only if your repo has a requirements.txt/pyproject.toml/etc.
```
`@main` always tracks this repo's default branch, which is convenient but — like any mutable
ref — can change out from under you between runs. For production CI, pin it to a commit SHA
instead (find one via `git rev-parse` against a tagged release, or the commit history for
[`.github/actions/setup-semantica/`](../../.github/actions/setup-semantica/)) and update the pin
deliberately when you want to pick up changes, the same way this repo's own workflows are pinned
(see [`verify-action-pins.yml`](../../.github/workflows/verify-action-pins.yml)).
It installs Python, installs `semantica`, and verifies the import (pip caching is opt-in via `cache: 'pip'`, since not every caller repo has a requirements file to key the cache on) — see
[`.github/actions/setup-semantica/action.yml`](../../.github/actions/setup-semantica/action.yml).
+40
View File
@@ -0,0 +1,40 @@
# Drop this in as .circleci/config.yml in your own project.
version: 2.1
jobs:
test:
docker:
- image: cimg/python:3.11
steps:
- checkout
# A content-hashed cache key (e.g. `{{ checksum "requirements.txt" }}`)
# is more precise but breaks if that exact file doesn't exist in your
# project - swap in one matched to however you declare dependencies
# once you've adjusted the install step below.
- restore_cache:
keys:
- pip-cache-v1
- run:
name: Install dependencies
command: |
pip install --upgrade pip
pip install semantica
# Install your own project's dependencies however your project
# declares them - adjust this to match, e.g. `pip install -e .`
# for pyproject.toml / setup.cfg, or `poetry install`.
if [ -f requirements.txt ]; then pip install -r requirements.txt; fi
- save_cache:
key: pip-cache-v1
paths:
- ~/.cache/pip
- run:
name: Smoke test
command: python -c "import semantica; print('semantica', semantica.__version__)"
- run:
name: Run tests
command: pytest
workflows:
test:
jobs:
- test
+44
View File
@@ -0,0 +1,44 @@
# Drop this in as .github/workflows/semantica.yml in your own project.
#
# Installs Semantica and runs a smoke import + your test suite. Swap the
# smoke-test step for whatever your project actually does with Semantica
# (build a context graph, run an ingest pipeline, etc.).
#
# Third-party actions below are pinned to a commit SHA rather than a mutable
# tag - a moved tag can silently swap in different code. Update the pin (and
# the trailing "# vX" comment) deliberately when you want a newer version;
# see semantica-agi/semantica's own .github/workflows/verify-action-pins.yml
# for one way to keep pins honest automatically.
name: Semantica
on:
push:
branches: [main]
pull_request:
branches: [main]
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7
with:
python-version: '3.11'
cache: 'pip'
- name: Install dependencies
run: |
python -m pip install --upgrade pip
pip install semantica
# Install your own project's dependencies however your project
# declares them - adjust this to match. Examples:
# pip install -r requirements.txt
# pip install -e . # pyproject.toml / setup.cfg
# pip install -e ".[dev]"
# poetry install
if [ -f requirements.txt ]; then pip install -r requirements.txt; fi
- name: Run tests
run: pytest
+20
View File
@@ -0,0 +1,20 @@
# Drop this in as .gitlab-ci.yml in your own project.
semantica-test:
image: python:3.11-slim
cache:
paths:
- .cache/pip
variables:
PIP_CACHE_DIR: "$CI_PROJECT_DIR/.cache/pip"
script:
- pip install --upgrade pip
- pip install semantica
# Install your own project's dependencies however your project declares
# them - adjust this to match, e.g. `pip install -e .` for pyproject.toml
# / setup.cfg, or `poetry install`.
- if [ -f requirements.txt ]; then pip install -r requirements.txt; fi
- python -c "import semantica; print('semantica', semantica.__version__)"
- pytest
rules:
- if: '$CI_PIPELINE_SOURCE == "merge_request_event"'
- if: '$CI_COMMIT_BRANCH == "main"'
+6 -6
View File
@@ -2083,9 +2083,9 @@
}
},
"node_modules/brace-expansion": {
"version": "5.0.8",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.8.tgz",
"integrity": "sha512-JZyDyq3D4AUifKTPOB7DELf6XsB3WdPuNxCtob1vFXPsSXhdAiHBWJ/tJ8HAc9aH84BK+5JFZLNkJKx3G9kzQg==",
"version": "5.0.9",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.9.tgz",
"integrity": "sha512-ScQ4IuvIEF1TMlP7Zt+vjJ//9zlPb2SDcxWxM3bk8s6t6GGdJ7KO1dCcTidOPJKePW30LE/2cT7wCyPho9/Wxg==",
"dev": true,
"license": "MIT",
"dependencies": {
@@ -4250,9 +4250,9 @@
"license": "MIT"
},
"node_modules/nanoid": {
"version": "3.3.16",
"resolved": "https://registry.npmjs.org/nanoid/-/nanoid-3.3.16.tgz",
"integrity": "sha512-bzlKTyNJ7+LdGIIwy8ijFpIqEQIvafahV7eYykJ8Cvh42EdJeODoJ6gUJXpQJvej1BddH8OqTXZNE/KfbWAu8Q==",
"version": "3.3.18",
"resolved": "https://registry.npmjs.org/nanoid/-/nanoid-3.3.18.tgz",
"integrity": "sha512-DTg4MJbGMWkfi6VZFdNt2/caMbQy4Ou+Op/hJQvGEWcnVfoA1QA+xzRKAzw9jD6+GVOOeYr/mIcuDSdug6F6+w==",
"dev": true,
"funding": [
{
+8 -1
View File
@@ -49,7 +49,14 @@ dependencies = [
"scipy>=1.13.1",
"scikit-learn>=1.7.2",
"umap-learn>=0.5.12",
"spacy>=3.4.0",
# thinc (spacy's core dep) dropped Python 3.9 wheels at 8.3.10, and later
# spacy patch releases (3.8.8+) require thinc>=8.3.9-only-on-3.10+ ranges,
# which forces a source build that fails outright on 3.9 (see Install
# Matrix run history). Capping both keeps 3.9 on the last wheel-compatible
# pair; 3.10+ is left unconstrained to always get the latest spacy/thinc.
"spacy>=3.4.0,<3.8.8; python_version < '3.10'",
"spacy>=3.4.0; python_version >= '3.10'",
"thinc<8.3.5; python_version < '3.10'",
"transformers>=4.20.0",
"torch>=1.13.1",
"sentence-transformers>=2.2.0",
+114 -9
View File
@@ -3714,19 +3714,61 @@ def store_stats(cli_ctx: CLIContext, backend: str, fmt: str, local_json: bool) -
_run_with_error_handling(_action)
_MIGRATE_SUPPORTED_BACKENDS = {"faiss", "sqlite", "pgvector"}
_MIGRATE_BATCH_SIZE = 500
def _migrate_backend_config(vs_cfg: Dict[str, Any], backend: str) -> Dict[str, Any]:
"""Resolve per-backend config out of the vector_store config section.
Supports both a per-backend nested shape (``vector_store.faiss.dimension``)
and the common flat single-backend shape (``vector_store.backend`` +
sibling keys), since either can appear depending on how many backends a
user has configured.
"""
nested = vs_cfg.get(backend)
if isinstance(nested, dict):
return dict(nested)
if vs_cfg.get("backend") == backend:
return {k: v for k, v in vs_cfg.items() if k != "backend"}
return {}
def _require_faiss_index_path(cfg: Dict[str, Any], role: str) -> str:
"""FAISS has no server to hold state between commands: a fresh FAISSStore
starts empty and nothing outside the process persists it, so migration
needs an explicit on-disk index to read from or write to."""
index_path = cfg.get("index_path")
if not index_path:
raise click.ClickException(
f"faiss as migration {role} requires 'index_path' in the vector_store "
f"config (vector_store.faiss.index_path or vector_store.index_path "
f"when faiss is the configured backend)."
)
return index_path
@store.command("migrate")
@click.option("--from", "from_backend", required=True)
@click.option("--to", "to_backend", required=True)
@click.option("--namespace", default=None)
@click.option("--dry-run", "local_dry", is_flag=True, default=False)
@click.option("--json", "local_json", is_flag=True, default=False)
@click.pass_obj
def store_migrate(cli_ctx: CLIContext, from_backend: str, to_backend: str,
namespace: Optional[str], local_dry: bool) -> None:
namespace: Optional[str], local_dry: bool, local_json: bool) -> None:
"""Migrate data between backends.
Direct migration is only wired up between faiss, sqlite, and pgvector -
these are the backends whose storage contract supports paging through
every stored vector. Migrating to or from qdrant, pinecone, milvus, or
weaviate still needs the export/reindex workaround below, since each of
those needs its own enumeration design (Qdrant scroll, Pinecone list,
etc.) that hasn't been built yet.
\b
Example:
semantica store migrate --from faiss --to qdrant --namespace production --dry-run
semantica store migrate --from faiss --to sqlite --namespace production --dry-run
"""
cli_ctx = _require_ctx(cli_ctx)
@@ -3734,13 +3776,76 @@ def store_migrate(cli_ctx: CLIContext, from_backend: str, to_backend: str,
if _is_dry(cli_ctx, local_dry):
_dry(cli_ctx, "migrate", from_backend=from_backend, to_backend=to_backend)
return
raise click.ClickException(
f"Direct backend migration ({from_backend}{to_backend}) is not yet supported "
"by the vector store layer. To migrate, export your data first:\n"
" semantica export --format parquet --output dump.parquet\n"
f" semantica embed index dump.parquet --store {to_backend}"
+ (f" --namespace {namespace}" if namespace else "")
)
if from_backend not in _MIGRATE_SUPPORTED_BACKENDS or to_backend not in _MIGRATE_SUPPORTED_BACKENDS:
raise click.ClickException(
f"Direct backend migration ({from_backend}{to_backend}) is only supported "
f"between {', '.join(sorted(_MIGRATE_SUPPORTED_BACKENDS))}. To migrate involving "
"another backend, export your data first:\n"
" semantica export --format parquet --output dump.parquet\n"
f" semantica embed index dump.parquet --store {to_backend}"
+ (f" --namespace {namespace}" if namespace else "")
)
from .vector_store import VectorStore
vs_cfg = cli_ctx.config.to_dict().get("vector_store", {}) or {}
source_cfg = _migrate_backend_config(vs_cfg, from_backend)
dest_cfg = _migrate_backend_config(vs_cfg, to_backend)
source_index_path = None
if from_backend == "faiss":
source_index_path = _require_faiss_index_path(source_cfg, "source")
dest_index_path = None
if to_backend == "faiss":
dest_index_path = _require_faiss_index_path(dest_cfg, "destination")
source = VectorStore(backend=from_backend, config=source_cfg)
if source_index_path:
source._backend_store.load_index(source_index_path)
source_dimension = getattr(source._backend_store, "dimension", None)
if source_dimension and "dimension" not in dest_cfg:
dest_cfg["dimension"] = source_dimension
dest = VectorStore(backend=to_backend, config=dest_cfg)
if dest_index_path and Path(dest_index_path).exists():
dest._backend_store.load_index(dest_index_path)
migrated = 0
vectors_batch: List[Any] = []
metadata_batch: List[Dict[str, Any]] = []
ids_batch: List[str] = []
def _flush() -> None:
nonlocal migrated
if not vectors_batch:
return
dest.store_vectors(list(vectors_batch), list(metadata_batch), ids=list(ids_batch))
migrated += len(vectors_batch)
vectors_batch.clear()
metadata_batch.clear()
ids_batch.clear()
for item in source.iter_vectors(batch_size=_MIGRATE_BATCH_SIZE):
meta = dict(item.get("metadata") or {})
if namespace and "namespace" not in meta:
meta["namespace"] = namespace
vectors_batch.append(item["vector"])
metadata_batch.append(meta)
ids_batch.append(item["id"])
if len(vectors_batch) >= _MIGRATE_BATCH_SIZE:
_flush()
_flush()
if dest_index_path and migrated:
dest._backend_store.save_index(dest_index_path)
result = {"from": from_backend, "to": to_backend, "migrated": migrated}
if _is_json(cli_ctx, local_json):
_jecho(result)
else:
_ok(cli_ctx, f"Migrated {migrated} vectors from {from_backend} to {to_backend}")
_run_with_error_handling(_action)
+55 -4
View File
@@ -66,12 +66,32 @@ class FAISSIndex:
self.metadata: Dict[str, Dict[str, Any]] = {}
def add_vectors(self, vectors: np.ndarray, ids: Optional[List[str]] = None):
"""Add vectors to index."""
"""
Add vectors to index.
Skips any id already present in vector_ids rather than appending a
second physical vector under the same id. FAISS indices here don't
support removing or replacing a single vector in place, so an
"update" isn't possible; without this check, re-running an add for
ids that already exist (e.g. retrying an interrupted migration)
would silently duplicate vectors under the same id on every retry.
"""
if ids is None:
ids = [f"vec_{i}" for i in range(len(vectors))]
self.index.add(vectors.astype(np.float32))
self.vector_ids.extend(ids)
new_rows = []
new_ids = []
existing = set(self.vector_ids)
for row, vec_id in zip(vectors, ids):
if vec_id in existing:
continue
new_rows.append(row)
new_ids.append(vec_id)
existing.add(vec_id)
if new_rows:
self.index.add(np.array(new_rows, dtype=np.float32))
self.vector_ids.extend(new_ids)
def search(
self, query_vectors: np.ndarray, k: int = 10
@@ -305,6 +325,12 @@ class FAISSStore:
"""
Add vectors to index.
Any id that already exists in the index is skipped rather than
stored as a second physical vector under the same id (see
FAISSIndex.add_vectors), so calling this again with ids from a
previous call is safe and doesn't accumulate duplicates. Metadata
for those ids is still updated.
Args:
vectors: List of vectors or numpy array
ids: Vector IDs
@@ -312,7 +338,8 @@ class FAISSStore:
**options: Additional options
Returns:
List of vector IDs
List of vector IDs (including ids that were already present
and therefore not re-added as new vectors)
"""
num_vectors = len(vectors) if isinstance(vectors, (list, np.ndarray)) else 1
tracking_id = self.progress_tracker.start_tracking(
@@ -526,6 +553,30 @@ class FAISSStore:
return results
def scan_vectors(self, offset: int = 0, limit: int = 100) -> List[Dict[str, Any]]:
"""
Page through stored vectors in insertion order.
Args:
offset: Number of vectors to skip
limit: Maximum number of vectors to return
Returns:
List of result dicts with 'id', 'metadata', and 'vector'
"""
if self.index is None or limit <= 0:
return []
ids_page = self.index.vector_ids[offset:offset + limit]
return [
{
"id": vector_id,
"metadata": self.get_metadata(vector_id) or {},
"vector": self.get_vector(vector_id),
}
for vector_id in ids_page
]
def get_stats(self) -> Dict[str, Any]:
"""Get index statistics."""
if self.index is None:
+46
View File
@@ -656,6 +656,52 @@ class PgVectorStore:
self.logger.warning(f"Failed to get metadata for {vector_id}: {e}")
return None
def scan_vectors(self, offset: int = 0, limit: int = 100) -> List[Dict[str, Any]]:
"""
Page through stored vectors ordered by id.
Args:
offset: Number of rows to skip
limit: Maximum number of rows to return
Returns:
List of result dicts with 'id', 'metadata', and 'vector'
"""
if not PSYCOPG3_AVAILABLE and not PSYCOPG2_AVAILABLE:
raise ProcessingError(
"Neither psycopg3 nor psycopg2 is available. "
"Install with: pip install psycopg[binary] or psycopg2-binary"
)
if limit <= 0:
return []
scan_sql = psycopg_sql.SQL("""
SELECT id, vector, metadata
FROM {}
ORDER BY id
LIMIT %s OFFSET %s
""").format(psycopg_sql.Identifier(self.table_name))
with self._get_connection() as conn:
try:
cur = conn.cursor()
cur.execute(scan_sql, (limit, offset))
rows = cur.fetchall()
cur.close()
results = []
for row in rows:
vec_id, vec, meta = row
results.append({
"id": vec_id,
"metadata": meta if isinstance(meta, dict) else json.loads(meta) if meta else {},
"vector": np.array(vec) if vec is not None else None,
})
return results
except Exception as e:
raise ProcessingError(f"Failed to scan vectors: {str(e)}") from e
def filter_by_metadata(
self, filters: Dict[str, Any], limit: int = 10
) -> List[Dict[str, Any]]:
@@ -616,6 +616,49 @@ class SQLiteVecStore:
self.logger.warning(f"Failed to get metadata for {vector_id}: {e}")
return None
def scan_vectors(self, offset: int = 0, limit: int = 100) -> List[Dict[str, Any]]:
"""
Page through stored vectors ordered by id.
Args:
offset: Number of rows to skip
limit: Maximum number of rows to return
Returns:
List of result dicts with 'id', 'metadata', and 'vector'
"""
if limit <= 0:
return []
query_sql = f"""
SELECT id, embedding, metadata
FROM {self.table_name}
ORDER BY id
LIMIT ? OFFSET ?
"""
with self._lock, self._get_connection() as conn:
try:
cur = conn.cursor()
cur.execute(query_sql, (limit, offset))
rows = cur.fetchall()
cur.close()
results = []
for row in rows:
vec_id, embedding_blob, meta_json = row
vec = None
if embedding_blob:
vec = np.frombuffer(embedding_blob, dtype=np.float32).copy()
results.append({
"id": vec_id,
"metadata": json.loads(meta_json) if meta_json else {},
"vector": vec,
})
return results
except Exception as e:
raise ProcessingError(f"Failed to scan vectors: {str(e)}") from e
def filter_by_metadata(
self, filters: Dict[str, Any], limit: int = 10
) -> List[Dict[str, Any]]:
+58
View File
@@ -824,6 +824,64 @@ class VectorStore:
else:
raise NotImplementedError(f"Backend store {type(self._backend_store).__name__} does not implement get_metadata")
def scan_vectors(self, offset: int = 0, limit: int = 100) -> List[Dict[str, Any]]:
"""
Page through stored vectors, backend-agnostic.
Follows the get_vector()/get_metadata() precedent (#843): the inmemory
backend pages its local dict directly, a persistent backend delegates
to a scan_vectors() on the wrapped store when available, and one that
cannot enumerate its contents raises NotImplementedError rather than
silently returning an empty page.
Args:
offset: Number of vectors to skip
limit: Maximum number of vectors to return
Returns:
List of result dicts with 'id', 'metadata', and 'vector'
"""
if limit <= 0:
return []
if self.backend == "inmemory":
ids_page = list(self.vectors.keys())[offset:offset + limit]
return [
{
"id": vec_id,
"metadata": self.metadata.get(vec_id, {}),
"vector": self.vectors.get(vec_id),
}
for vec_id in ids_page
]
elif self._backend_store and hasattr(self._backend_store, "scan_vectors"):
return self._backend_store.scan_vectors(offset=offset, limit=limit)
else:
raise NotImplementedError(
f"Backend store {type(self._backend_store).__name__} does not "
"implement scan_vectors(). Add a scan_vectors() method to the "
"backend store adapter to enable enumeration for this backend."
)
def iter_vectors(self, batch_size: int = 500):
"""
Iterate over every stored vector, one page at a time.
Args:
batch_size: Number of vectors to fetch per underlying scan_vectors() call
Yields:
Result dicts with 'id', 'metadata', and 'vector', in scan order
"""
offset = 0
while True:
page = self.scan_vectors(offset=offset, limit=batch_size)
if not page:
return
for item in page:
yield item
offset += len(page)
def count(self) -> int:
"""Return the number of vectors in the store, backend-agnostic.
+95
View File
@@ -1359,6 +1359,101 @@ class TestStore:
result = runner.invoke(cli_module.main, ["store", "migrate", "--from", "faiss"])
assert result.exit_code != 0
def test_migrate_refuses_unsupported_backend_pair(self, runner):
result = runner.invoke(cli_module.main, ["store", "migrate",
"--from", "faiss", "--to", "qdrant"])
assert result.exit_code != 0
assert "faiss, pgvector, sqlite" in result.output
def _fake_migrate_store_module(self, source_items, stored, dest_configs=None):
class _FakeBackendStore:
def __init__(self, dimension=None):
self.dimension = dimension
class _FakeStore:
def __init__(self, backend, config=None, **kw):
self.backend = backend
self._config = config or {}
dim = self._config.get("dimension")
self._backend_store = _FakeBackendStore(dimension=dim)
if dest_configs is not None:
dest_configs[backend] = dict(self._config)
def iter_vectors(self, batch_size=500):
if self.backend == "sqlite":
yield from source_items
return
return
yield # pragma: no cover - makes this a generator for other backends
def store_vectors(self, vectors, metadata, ids=None):
for vec_id, meta in zip(ids, metadata):
stored[vec_id] = meta
return _fake_module(VectorStore=_FakeStore)
def test_migrate_runs_between_supported_backends(self, runner, monkeypatch):
source_items = [
{"id": "a", "vector": [0.1, 0.2], "metadata": {"tag": "x"}},
{"id": "b", "vector": [0.3, 0.4], "metadata": {}},
]
stored = {}
fake_vs = self._fake_migrate_store_module(source_items, stored)
monkeypatch.setitem(__import__("sys").modules, "semantica.vector_store", fake_vs)
result = runner.invoke(cli_module.main, ["store", "migrate",
"--from", "sqlite", "--to", "pgvector",
"--namespace", "prod", "--json"])
_ok(result)
data = _json_output(result)
assert data == {"from": "sqlite", "to": "pgvector", "migrated": 2}
assert stored == {"a": {"tag": "x", "namespace": "prod"}, "b": {"namespace": "prod"}}
def test_migrate_reports_zero_for_empty_source(self, runner, monkeypatch):
stored = {}
fake_vs = self._fake_migrate_store_module([], stored)
monkeypatch.setitem(__import__("sys").modules, "semantica.vector_store", fake_vs)
result = runner.invoke(cli_module.main, ["store", "migrate",
"--from", "sqlite", "--to", "pgvector", "--json"])
_ok(result)
assert _json_output(result)["migrated"] == 0
assert stored == {}
def test_migrate_inherits_source_dimension_into_dest(self, runner, monkeypatch):
source_items = [{"id": "a", "vector": [0.1, 0.2, 0.3], "metadata": {}}]
stored = {}
dest_configs: dict = {}
fake_vs = self._fake_migrate_store_module(source_items, stored, dest_configs)
monkeypatch.setitem(__import__("sys").modules, "semantica.vector_store", fake_vs)
monkeypatch.setattr(
cli_module.Config, "to_dict",
lambda self: {"vector_store": {"sqlite": {"dimension": 3}, "pgvector": {}}},
)
result = runner.invoke(cli_module.main, ["store", "migrate",
"--from", "sqlite", "--to", "pgvector", "--json"])
_ok(result)
assert dest_configs["pgvector"].get("dimension") == 3
def test_migrate_faiss_source_requires_index_path(self, runner, monkeypatch):
fake_vs = _fake_module(VectorStore=lambda **kw: MagicMock())
monkeypatch.setitem(__import__("sys").modules, "semantica.vector_store", fake_vs)
result = runner.invoke(cli_module.main, ["store", "migrate",
"--from", "faiss", "--to", "sqlite"])
assert result.exit_code != 0
assert "index_path" in result.output
def test_migrate_faiss_dest_requires_index_path(self, runner, monkeypatch):
fake_vs = _fake_module(VectorStore=lambda **kw: MagicMock())
monkeypatch.setitem(__import__("sys").modules, "semantica.vector_store", fake_vs)
result = runner.invoke(cli_module.main, ["store", "migrate",
"--from", "sqlite", "--to", "faiss"])
assert result.exit_code != 0
assert "index_path" in result.output
def test_flush_requires_confirm(self, runner):
result = runner.invoke(cli_module.main, ["store", "flush"])
assert result.exit_code != 0
+88 -1
View File
@@ -3,7 +3,7 @@ from unittest.mock import MagicMock
import numpy as np
import pytest
from semantica.vector_store.faiss_store import FAISSIndex
from semantica.vector_store.faiss_store import FAISSIndex, FAISSStore
def test_get_vector_reconstructs_from_flat_l2_index():
@@ -120,3 +120,90 @@ def test_get_vector_reconstructs_from_real_ivfflat_index_without_prior_direct_ma
result = index.get_vector("vec_target")
np.testing.assert_allclose(result, vectors[3], atol=1e-6)
def _store_with_fake_index(ids, metadata_by_id=None):
backend_index = MagicMock()
backend_index.reconstruct.side_effect = lambda idx: [float(idx)] * 3
index = FAISSIndex(backend_index, dimension=3)
index.vector_ids = list(ids)
index.metadata = dict(metadata_by_id or {})
store = FAISSStore(dimension=3)
store.index = index
return store
def test_scan_vectors_returns_all_across_pages():
store = _store_with_fake_index(["a", "b", "c", "d", "e"])
seen_ids = []
offset = 0
while True:
page = store.scan_vectors(offset=offset, limit=2)
if not page:
break
seen_ids.extend(p["id"] for p in page)
offset += len(page)
assert seen_ids == ["a", "b", "c", "d", "e"]
def test_scan_vectors_includes_vector_and_metadata():
store = _store_with_fake_index(["a"], {"a": {"tag": "only"}})
page = store.scan_vectors(offset=0, limit=10)
assert len(page) == 1
assert page[0]["id"] == "a"
assert page[0]["metadata"] == {"tag": "only"}
np.testing.assert_array_equal(page[0]["vector"], np.array([0.0, 0.0, 0.0], dtype=np.float32))
def test_scan_vectors_no_index_returns_empty_list():
store = FAISSStore(dimension=3)
assert store.scan_vectors(offset=0, limit=10) == []
def test_scan_vectors_zero_limit_returns_empty_list():
store = _store_with_fake_index(["a"])
assert store.scan_vectors(offset=0, limit=0) == []
def test_scan_vectors_offset_past_end_returns_empty_list():
store = _store_with_fake_index(["a"])
assert store.scan_vectors(offset=100, limit=10) == []
def test_add_vectors_retry_with_same_ids_does_not_duplicate():
"""Re-running add_vectors with ids already in the index (e.g. retrying
an interrupted migration) must not create a second physical vector
under the same id."""
backend_index = MagicMock()
store = FAISSStore(dimension=3)
store.index = FAISSIndex(backend_index, dimension=3)
vectors = np.array([[1, 2, 3], [4, 5, 6], [7, 8, 9], [10, 11, 12]], dtype=np.float32)
ids = ["a", "b", "c", "d"]
store.add_vectors(vectors, ids=ids, metadata=[{"i": i} for i in range(4)])
assert store.count() == 4
store.add_vectors(vectors, ids=ids, metadata=[{"i": i} for i in range(4)])
assert store.count() == 4
assert store.index.vector_ids == ids
def test_add_vectors_retry_with_partial_overlap_only_adds_new_ids():
backend_index = MagicMock()
store = FAISSStore(dimension=3)
store.index = FAISSIndex(backend_index, dimension=3)
store.add_vectors(np.array([[1, 2, 3], [4, 5, 6]], dtype=np.float32), ids=["a", "b"])
store.add_vectors(np.array([[1, 2, 3], [7, 8, 9]], dtype=np.float32), ids=["a", "c"])
assert store.index.vector_ids == ["a", "b", "c"]
second_call_vectors = backend_index.add.call_args[0][0]
assert second_call_vectors.shape[0] == 1
np.testing.assert_array_equal(second_call_vectors[0], np.array([7, 8, 9], dtype=np.float32))
+42
View File
@@ -467,6 +467,48 @@ class TestPgVectorStoreDelete:
assert success is True
class TestPgVectorStoreScan:
"""Test scan_vectors pagination."""
def test_scan_returns_all_vectors_across_pages(self, store):
vectors = [np.random.rand(128).astype(np.float32) for _ in range(5)]
ids = store.add(vectors, [{"index": i} for i in range(5)])
seen_ids = []
offset = 0
while True:
page = store.scan_vectors(offset=offset, limit=2)
if not page:
break
seen_ids.extend(p["id"] for p in page)
offset += len(page)
assert set(seen_ids) == set(ids)
assert len(seen_ids) == 5
def test_scan_page_includes_vector_and_metadata(self, store):
vectors = [np.random.rand(128).astype(np.float32)]
ids = store.add(vectors, [{"tag": "only"}])
page = store.scan_vectors(offset=0, limit=10)
assert len(page) == 1
assert page[0]["id"] == ids[0]
assert page[0]["metadata"] == {"tag": "only"}
assert page[0]["vector"] is not None
def test_scan_empty_store_returns_empty_list(self, store):
assert store.scan_vectors(offset=0, limit=10) == []
def test_scan_zero_limit_returns_empty_list(self, store):
store.add([np.random.rand(128).astype(np.float32)])
assert store.scan_vectors(offset=0, limit=0) == []
def test_scan_offset_past_end_returns_empty_list(self, store):
store.add([np.random.rand(128).astype(np.float32)])
assert store.scan_vectors(offset=100, limit=10) == []
class TestPgVectorStoreIndex:
"""Test index creation operations."""
@@ -415,6 +415,48 @@ class TestSQLiteVecStoreStats:
assert stats["vector_count"] == 4
class TestSQLiteVecStoreScan:
"""Test scan_vectors pagination."""
def test_scan_returns_all_vectors_across_pages(self, store):
vectors = [np.random.rand(128).astype(np.float32) for _ in range(5)]
ids = store.add(vectors, [{"index": i} for i in range(5)])
seen_ids = []
offset = 0
while True:
page = store.scan_vectors(offset=offset, limit=2)
if not page:
break
seen_ids.extend(p["id"] for p in page)
offset += len(page)
assert set(seen_ids) == set(ids)
assert len(seen_ids) == 5
def test_scan_page_includes_vector_and_metadata(self, store):
vectors = [np.random.rand(128).astype(np.float32)]
ids = store.add(vectors, [{"tag": "only"}])
page = store.scan_vectors(offset=0, limit=10)
assert len(page) == 1
assert page[0]["id"] == ids[0]
assert page[0]["metadata"] == {"tag": "only"}
assert page[0]["vector"] is not None
def test_scan_empty_store_returns_empty_list(self, store):
assert store.scan_vectors(offset=0, limit=10) == []
def test_scan_zero_limit_returns_empty_list(self, store):
store.add([np.random.rand(128).astype(np.float32)])
assert store.scan_vectors(offset=0, limit=0) == []
def test_scan_offset_past_end_returns_empty_list(self, store):
store.add([np.random.rand(128).astype(np.float32)])
assert store.scan_vectors(offset=100, limit=10) == []
class TestSQLiteVecStoreFilterByMetadata:
"""Test filter_by_metadata, including list-valued metadata handling."""
@@ -120,6 +120,78 @@ class VectorStoreCountTests(unittest.TestCase):
self.assertIn("count()", msg)
# ---------------------------------------------------------------------------
# VectorStore.scan_vectors() / iter_vectors() dispatch tests
# ---------------------------------------------------------------------------
class _ScanningBackendStore:
"""Fake persistent backend store that supports scan_vectors()."""
def __init__(self, items):
self._items = items
def scan_vectors(self, offset=0, limit=100):
return self._items[offset:offset + limit]
class _NonScanningBackendStore:
"""Fake persistent backend store without any scan capability."""
class VectorStoreScanVectorsTests(unittest.TestCase):
"""VectorStore.scan_vectors() / iter_vectors() backend-agnostic accessors."""
def setUp(self):
self.vectors = [np.array([1.0, 0.0]), np.array([0.0, 1.0]), np.array([1.0, 1.0])]
self.metadata = [{"type": "a"}, {"type": "b"}, {"type": "c"}]
def test_scan_inmemory_pages_through_all_vectors(self):
store = VectorStore(backend="inmemory", dimension=2)
ids = store.store_vectors(self.vectors, self.metadata)
page1 = store.scan_vectors(offset=0, limit=2)
page2 = store.scan_vectors(offset=2, limit=2)
self.assertEqual([p["id"] for p in page1], ids[:2])
self.assertEqual([p["id"] for p in page2], ids[2:])
self.assertEqual(page2[0]["metadata"], {"type": "c"})
def test_scan_inmemory_empty_store(self):
store = VectorStore(backend="inmemory", dimension=2)
self.assertEqual(store.scan_vectors(offset=0, limit=10), [])
def test_scan_zero_limit_returns_empty_list(self):
store = VectorStore(backend="inmemory", dimension=2)
store.store_vectors(self.vectors, self.metadata)
self.assertEqual(store.scan_vectors(offset=0, limit=0), [])
def test_scan_delegates_to_backend_store(self):
items = [{"id": "a", "metadata": {}, "vector": None}]
store = VectorStore(backend="inmemory", dimension=2)
store.backend = "faiss"
store._backend_store = _ScanningBackendStore(items)
self.assertEqual(store.scan_vectors(offset=0, limit=10), items)
def test_scan_raises_not_implemented_without_backend_support(self):
store = VectorStore(backend="inmemory", dimension=2)
store.backend = "faiss"
store._backend_store = _NonScanningBackendStore()
with self.assertRaises(NotImplementedError):
store.scan_vectors(offset=0, limit=10)
def test_iter_vectors_walks_every_page(self):
store = VectorStore(backend="inmemory", dimension=2)
ids = store.store_vectors(self.vectors, self.metadata)
collected = list(store.iter_vectors(batch_size=2))
self.assertEqual([item["id"] for item in collected], ids)
def test_iter_vectors_empty_store_yields_nothing(self):
store = VectorStore(backend="inmemory", dimension=2)
self.assertEqual(list(store.iter_vectors(batch_size=2)), [])
# ---------------------------------------------------------------------------
# VectorManager tests — inmemory backend
# ---------------------------------------------------------------------------