mirror of
https://github.com/semantica-agi/semantica.git
synced 2026-08-29 04:26:20 +00:00
e12eec40a1fd1e4f4b17dad90c608e418b0fe100
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e12eec40a1 |
refactor(ner): remove dead _extract_with_spacy method and unused self.nlp (#1220)
* test(ner): fix NER configuration tests for the typed LLM extraction API Two of the three failing tests tracked in #1059 were still red after #1070 was closed because the mocks targeted the pre-typed provider API: - test_ner_llm_config mocked generate_structured, but the LLM path now goes through generate_typed with a Pydantic schema. Mock the typed response (namespace items with .text/.label/.start/.end/.confidence) and expect extraction_method 'llm_typed'. - test_ner_pattern_config asserted 'Apple Inc' without the trailing dot, but the ORG pattern captures it via (?:\.|\b). Assert 'Apple Inc.' to match current production behavior. Verified locally: 8/8 pass in test_ner_configurations.py; the performance-test failures in tests/semantic_extract/ reproduce on a clean main checkout and are unrelated. Fixes #1059 Signed-off-by: Yunare Maia <yunare@gmail.com> * refactor(ner): remove dead _extract_with_spacy method and unused self.nlp _extract_with_spacy() had no callers: the ML dispatch path goes through get_entity_method('ml') -> extract_entities_ml(), which loads the spaCy model lazily via the process-level cache in methods.py. The instance attribute self.nlp was only read by that dead method, so __init__ now just validates the runtime (keeping the _ml_runtime_usable gate) instead of eagerly loading a model that was never used. Fixes #1058 Signed-off-by: Yunare Maia <yunare@gmail.com> * test(split): rewrite NERExtractor cache tests to not rely on removed .nlp attribute NERExtractor.nlp was removed in this PR as part of dead-code cleanup (the attribute was only used by the equally-dead _extract_with_spacy()). The three affected tests in TestNERExtractorSpacyModelCache previously verified cache behavior through .nlp identity comparisons; rewrite them to use load-call counts and direct se_methods.load_spacy_model() cache queries instead: - test_ner_extractor_reuses_cached_model_across_instances: drop the e1.nlp is e2.nlp is e3.nlp assertion; len(calls)==1 already proves reuse; add a cache query to confirm the cached object is non-None. - test_ner_extractor_distinct_model_names_load_separately: store each mock nlp in a dict keyed by name, then query the cache to assert sm_cached is loaded['en_core_web_sm'] and sm_cached is not lg_cached. - test_ner_extractor_failed_load_not_cached_and_retried: replace extractor.nlp is None/not None with is-not-None construction checks and a final cache query that verifies the recovered model is the exact object returned by working_load. All three tests still exercise the original behavioral contract (no crash on missing model, failures not cached / retried, successful load shared across instances); they just no longer rely on a private instance attribute that no longer exists. --------- Signed-off-by: Yunare Maia <yunare@gmail.com> Co-authored-by: Sameer Kadam <sskadam6305@gmail.com> |
||
|
|
4513b61e40 |
ci: pin Python dependencies in requirements-ci.txt for reproducible CI (#945)
* ci: pin Python dependencies in requirements-ci.txt for reproducible CI Adds a committed lockfile pinning all transitive dependencies at exact versions (uv pip compile, Python 3.11, all extras — 1581 lines), the Python equivalent of explorer/package-lock.json + npm ci. - CI installs from requirements-ci.txt before building the wheel - CI verifies the lockfile is byte-identical to a fresh compile (fails on staleness after pyproject.toml changes) - CONTRIBUTING documents the regeneration command Closes #938 Signed-off-by: Yunare Maia <yunare@gmail.com> * ci: address Qodo review — security scans use pinned deps, exclude gpu extras - security-scan.yml installs from requirements-ci.txt instead of "./[llm-litellm]" so Safety scans the exact CI/release dependency tree - security.yml runs pip-audit -r requirements-ci.txt for the same parity - lockfile regenerated with --extra all (the cross-platform set) instead of --all-extras, which pulled faiss-gpu/cupy from the Linux-only gpu extra and co-installed faiss-cpu + faiss-gpu in CI - uv pinned to 0.12.1 (the version that generated the lockfile) in CI and CONTRIBUTING so regeneration is deterministic Signed-off-by: Yunare Maia <yunare@gmail.com> * ci: make lockfile staleness check immune to upstream releases The previous check re-resolved pyproject.toml without constraints, so any upstream package release (e.g. boto3 1.43.69 -> 1.43.70) failed CI even when nothing in the repo changed — exactly the time-dependent drift Qodo flagged. The check now re-resolves with requirements-ci.txt as a constraint and compares only version lines, so it detects intentional pyproject.toml changes but ignores upstream releases. CONTRIBUTING updated to match. Signed-off-by: Yunare Maia <yunare@gmail.com> * ci: fix security workflows — install pip-audit; order tooling after pinned deps Security workflow: the pip-audit install step was lost in the rebase conflict merge — pip-audit was invoked but never installed (exit 127). Security-scan workflow: installing safety first let the pinned requirements-ci.txt overwrite its transitive deps (rich), breaking the safety CLI at runtime (RuntimeError: Type not yet supported). Tooling is now installed AFTER the pinned set. Signed-off-by: Yunare Maia <yunare@gmail.com> * fix(ci): address review — hashes, build isolation, release builds, docs (4/4) ZohaibHassan16's review flagged 4 supply-chain gaps; all addressed: 1. **Release builds now use the lockfile**: release.yml installs requirements-ci.txt and runs `python -m build --no-isolation` so the sdist/wheel is built against the exact tested dependency set. 2. **Build isolation pinned**: [build-system].requires is now setuptools==84.0.0 + wheel==0.48.0 (exact pins, no ranges). 3. **Hashes**: requirements-ci.txt regenerated with --generate-hashes (5,708 sha256 hashes, verified against PyPI). Staleness check updated to strip the `\` line continuations hashes introduce. 4. **CONTRIBUTING.md documents the separate environment**: hashes, never-install-into-dev note, build-system pins, --no-isolation release builds. Validated: stale-check diff clean, hash spot-check matches PyPI. Signed-off-by: Yunare Maia <yunare@gmail.com> * fix(ci): apply --no-isolation to CI build + align benchmark to Python 3.11 Follow-up to ZohaibHassan16's second review round: 1. ci.yml was still running `python -m build` with build isolation (unpinned setuptools/wheel from PyPI) — now `python -m build --no-isolation` against the pinned deps, matching release.yml. 2. benchmark.yml was on Python 3.12 while the lockfile is compiled for 3.11 — aligned to 3.11 so every workflow runs the same environment. Signed-off-by: Yunare Maia <yunare@gmail.com> * fix(ci): install pinned wheel before --no-isolation build python -m build --no-isolation failed with 'Missing dependencies: wheel==0.48.0' because wheel is build-time only — uv's lockfile excludes it, so installing requirements-ci.txt alone left the build env without it. Both ci.yml and release.yml now install wheel==0.48.0 (the same pin [build-system] declares) before building. Validated locally: wheel builds clean with --no-isolation. Signed-off-by: Yunare Maia <yunare@gmail.com> --------- Signed-off-by: Yunare Maia <yunare@gmail.com> Co-authored-by: Zohaib Hassnain <109234410+ZohaibHassan16@users.noreply.github.com> |
||
|
|
c5d13a45db |
feat(seed): allow_private_ips opt-in for trusted internal API sources (#959)
* feat(seed): add allow_private_ips opt-in for trusted internal API sources (Closes #943) SeedDataManager.load_from_api now delegates to the shared SSRF guard (semantica/ingest/ssrf.py, added in #906) instead of raw requests.get, gaining redirect validation and bounded DNS resolution for free. New config option allow_private_ips (parsed via the shared parse_bool helper) lets trusted internal deployments load from private APIs while the secure default (block private/loopback/link-local) is unchanged. Tests updated to mock request_with_ssrf_guard; new tests cover the block-by-default behavior and the opt-in flag reaching the guard. 19/19 green in test_seed_manager.py, 25/25 across both seed suites. Signed-off-by: Yunare Maia <yunare@gmail.com> * fix(ssrf): strip sensitive headers on cross-host redirects (Qodo finding) request_with_ssrf_guard reused the caller's headers on every redirect hop, so an Authorization bearer token from load_from_api could leak to a different redirect target host. Now strips Authorization and Proxy-Authorization when the redirect origin (netloc) changes, while keeping them for same-host hops (matching requests semantics). 2 new tests: cross-host redirect drops the credential; same-host keeps it. 37/37 green in test_ssrf_protection.py. load_from_api docstring now also documents cloud-metadata blocking and per-hop redirect validation. Signed-off-by: Yunare Maia <yunare@gmail.com> * fix(ssrf): strip credentials on https->http downgrade redirects (review feedback) _should_strip_auth now mirrors requests' should_strip_auth semantics: strip on hostname change, port change, or scheme downgrade; keep the credential only for the safe http->https upgrade on default ports. Previously only netloc was compared, so an https->http redirect on the same host replayed the Authorization header in cleartext. --------- Signed-off-by: Yunare Maia <yunare@gmail.com> |
||
|
|
43bac6170c |
fix(vector_store): make VectorManager methods work on persistent backends (#855) (#914)
* fix(vector_store): make VectorManager methods work on persistent backends (#855) maintain_store() and collect_statistics() reached into VectorStore internals (.vectors/.metadata), which only exist for the inmemory backend — any persistent backend (FAISS, Qdrant, Pinecone, Milvus, ...) crashed with AttributeError. Add a public backend-agnostic VectorStore.count() accessor following the get_vector()/get_metadata() precedent (#843) and the NotImplementedError-on-unsupported-capability precedent of _filter_by_metadata() (#848): inmemory counts its dict, persistent backends delegate to count() when available, and raise NotImplementedError otherwise. VectorManager methods now go through count(); maintain_store() keeps the exact inmemory semantics (separate vector/metadata dict counts) and reports a 1:1 count for persistent backends, where metadata is stored alongside each vector. Tests: 10 hermetic unit tests covering inmemory, delegation and the NotImplementedError path. Core vector_store suite: 40 passed. * fix(vector_store): raise NotImplementedError when count() unavailable Address Qodo review findings on #914: - Persistent backend with no wrapped store no longer silently returns 0 (which masked a missing initialization as an empty, healthy store); it now raises NotImplementedError like get_vector()/get_metadata(). - A mis-shaped adapter exposing a non-callable 'count' attribute now surfaces a clean NotImplementedError instead of a TypeError, via a getattr + callable() capability check. Adds regression tests for both cases. * fix(vector_store): implement count() on FAISS/SQLite/PgVector backends (#914) - FAISSStore.count(): returns len(index.vector_ids); 0 when no index exists yet - SQLiteVecStore.count(): delegates to get_stats()[vector_count] (SELECT COUNT(*)) - PgVectorStore.count(): delegates to get_stats()[vector_count] (SELECT COUNT(*)) - VectorStore.count(): fix misleading NotImplementedError message; now describes how to add count() support to a backend adapter rather than claiming only the inmemory backend can ever support counting - VectorManager.maintain_store(): split inmemory and persistent paths: * inmemory: independently reads len(vectors) and len(metadata) and compares them as an integrity check (original semantics preserved) * persistent: calls store.count(); returns metadata_count=None because metadata is co-located with vectors in the backend and cannot be counted independently; never manufactures metadata_count=vector_count as a vacuous tautology (#914 Qodo review) - Tests: rewrite test_vector_manager_persistent.py with 31 tests covering dispatch logic, inmemory divergence detection, persistent metadata_count=None invariant, FAISSStore/PgVectorStore via mocks, and SQLiteVecStore via real in-memory SQLite (skipped when sqlite-vec absent) * docs(changelog): document VectorManager persistent-backend count fix (#914, closes #855) Records the VectorStore.count() accessor, the FAISS/SQLite/PgVector implementations added during review, and the maintain_store() metadata_count fix (no longer fabricates equality for persistent backends). --------- Co-authored-by: Sameer6305 <sskadam6305@gmail.com> Co-authored-by: KaifAhmad1 <kaifahmad087@gmail.com> Co-authored-by: Mohd Kaif <98801504+KaifAhmad1@users.noreply.github.com> |