#1540 added GHSA-4j2p-28q2-5m79 to security-scan.yml's IGNORED_VULN_IDS to
suppress the open, unpatched accelerate<=1.14.0 path traversal advisory
(Semantica doesn't call load_checkpoint_in_model/load_checkpoint_and_dispatch).
That commit's own CI run still failed: pip-audit's OSV-backed report picked
CVE-2026-69112 as the vuln's canonical `id` and demoted the GHSA id to an
alias, but the shell/JS matching only ever compared against `.id`.
List both identifiers and match against `.id` plus `.aliases` (which
pip-audit includes by default for JSON output) in the audit gate, the
"Vulnerability details" printer, and the PR-comment script, so an ignored
advisory is excluded regardless of which alias the report surfaces as
canonical. Verified against the actual failing report from run 34296586683 -
the corrected filter now yields 0 actionable vulnerabilities.
Also add osv-scanner.toml (per GHSA-4j2p-28q2-5m79's Scorecard code-scanning
remediation) so the weekly Scorecard "Vulnerabilities" check stops flagging
the same accepted, unpatched advisory.
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Add GHSA-4j2p-28q2-5m79 to IGNORED_VULN_IDS in security-scan.yml.
accelerate<=1.14.0 has an open path traversal advisory (GHSA-4j2p-28q2-5m79) in load_checkpoint_in_model. 1.14.0 is currently the latest available release on PyPI, so no upstream patch exists yet. Semantica does not expose or call sharded checkpoint loading, making this non-actionable. Re-evaluate once an updated accelerate release is published.
- Remove thinc direct constraint from nlp-spacy in pyproject.toml and update README/CHANGELOG
- Raise ProcessingError with install hint when UMAP is requested but unavailable in EmbeddingVisualizer
- Guard PCA and TSNE dimensionality reducers against None with actionable error messages
- Prevent keyword collisions in EmbeddingVisualizer dimensionality reduction (_reduce_dimensions)
- Add core-only install & test step to .github/workflows/ci.yml using base-deps.txt
- Prevent AttributeError on module load in repo_ingestor.py and xml_ingestor.py when optional dependencies are absent
- Make TOML loading in tests/test_issue_1513_slim_core.py portable across Python versions via UTF-8 text decode and loads
- Expand slim-core test suite to 17 test cases covering 2D/3D projections, core imports, and options handling
- Reduce direct core dependencies in pyproject.toml from 44 to 22
- Move heavy and specialized packages into modular optional extras:
* models-huggingface: torch, transformers
* embeddings-local: sentence-transformers, fastembed, onnxruntime, tokenizers
* nlp-spacy: spacy, thinc
* viz: matplotlib, seaborn, plotly, ipywidgets, umap-learn (expanded)
* media: librosa, opencv-python
* vectorstore-faiss: faiss-cpu
* documents: python-docx, openpyxl, lxml, beautifulsoup4
* ingest-git: GitPython
* graph-embeddings: gensim
- Update semantica[all] and semantica[vectorstore-all] to encompass all extras
- Ensure lazy parser construction (DOCXParser, ExcelParser, HTMLParser, XMLParser) without error on __init__(), failing only inside .parse() with clear hints
- Add stdlib xml.etree fallback in XMLParser when lxml is missing
- Guard unguarded matplotlib imports in EmbeddingVisualizer and OntologyVisualizer
- Standardize user-facing error messages to point to pip install 'semantica[extra]'
- Bump version to 0.7.0 in pyproject.toml, semantica/__init__.py, and CITATION.cff
- Add migration notes to README.md and CHANGELOG.md
- Recompile CI and Docker requirements lockfiles
- Add dedicated test suite tests/test_issue_1513_slim_core.py
#1290 bumped the runtime stage to python:3.14-slim. gensim is a base
(non-extras-gated) dependency and publishes no cp314 wheel on PyPI, so
pip falls back to building it from source, which needs a C compiler
this slim image doesn't carry:
error: [Errno 2] No such file or directory: 'gcc'
This has broken the Docker build (and Container Security Scan) on
every push to main since #1290 merged. Confirmed via
`pip index versions gensim` / the PyPI files listing that gensim 4.4.0
ships cp313 wheels but no cp314 ones; the earlier lockfile fix in
#1290 only validated that `uv pip compile` could *resolve* dependencies
for 3.14, which reads sdist metadata and doesn't need a compiler --
it doesn't validate that `pip install` can actually *build* them,
which is what fails here.
Reverts to the exact base image digest that was building successfully
before #1290 and regenerates explorer-extra-py313.txt against today's
PyPI state. Once gensim (and anything else pulled in transitively)
ships cp314 wheels, the 3.14 bump can be retried.
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* docker(deps): bump python from 3.13-slim to 3.14-slim
Bumps python from 3.13-slim to 3.14-slim.
---
updated-dependencies:
- dependency-name: python
dependency-version: 3.14-slim
dependency-type: direct:production
...
Signed-off-by: dependabot[bot] <support@github.com>
* fix(docker): regenerate explorer-extra lockfile for Python 3.14
The base image bump to python:3.14-slim left explorer-extra-py313.txt
in place, a lockfile uv resolved specifically for Python 3.13 and
installed with --require-hashes. Re-resolving the explorer extra for
3.14 adds cloudpickle, which has no hash entry in the 3.13 file, so
the runtime stage's pip install would fail outright.
Regenerates the lockfile as explorer-extra-py314.txt and updates every
reference to it (Dockerfile, .dockerignore, container-scan.yml's path
trigger, ci.yml's comment, and the requirements README).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
---------
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Mohd Kaif <98801504+KaifAhmad1@users.noreply.github.com>
Co-authored-by: KaifAhmad1 <kaifahmad087@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Replace the static biomedical category key in Explorer with a live node color legend derived from the active display graph, sharing the canvas's baseColor -> color -> theme fallback resolver.
- Preserve original semantic colors on focused display clones via `semanticBaseColor` so interaction styling (selection, path, neighbor highlights) does not bleed into the legend.
- Suppress the semantic legend while distance visualization modes are active.
- Refactor plugin fallback legend panels to use the shared legend builder with composite (group, color) keys.
- Consolidate duplicate `"test:graph-workspace"` scripts in `explorer/package.json`, recovering 26 previously shadowed tests across markdown and explorer capability suites.
- Add `tests/graphColorLegend.test.ts` to workspace test runs and wire `test:graph-legend-e2e` into `.github/workflows/ci.yml`.
- Closes#1479
The v0.6.8 release run failed at the PyPI publish step:
Checking dist/semantica-0.6.8-py3-none-any.whl.sigstore.json: ERROR
InvalidDistribution: Unknown distribution format
pypa/gh-action-pypi-publish uploads everything under packages-dir
(default dist/) with no include/exclude filter, so once the Sigstore
signing step (added in #1329) started writing dist/*.sigstore.json
alongside the wheel/sdist, publish was broken for every release from
that point on - it just never ran, since v0.6.7 was tagged two days
before #1329 merged. Confirmed nothing was uploaded to PyPI before
failing (dist/*.whl checked and passed first; the sigstore.json file
failed validation before any upload began).
Fix: run pypi-publish immediately after the package build, before the
Sigstore/attest-build-provenance steps write anything else into dist/.
The GitHub Release upload (which needs the .sigstore.json files) still
runs after signing, unaffected by the reorder.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds Google ADK (Agent Development Kit) support to Semantica.
`semantica_kg_tools()` and `semantica_decision_tools()` expose entity/relation extraction, graph updates, and decision recording as ADK `FunctionTool`s. `SemanticaSessionService` implements ADK session storage on top of a Semantica `ContextGraph`, so session state, events, and knowledge graph data can live in the same graph instead of keeping sessions in memory.
There were also a number of dependency and CI fixes needed to get the integration working reliably. `google-adk` is pinned to a range that avoids the CI `websockets` conflict, the deprecated `pinecone-client` dependency was replaced with `pinecone`, Windows-only dependencies now have the appropriate platform markers, and `requirements-ci.txt` was regenerated to match. A `pip-audit` pass also required updates to `google-adk` and `starlette` for known CVEs.
Some unrelated `pyproject.toml` changes had slipped in during rebases, so the previous version, dependency bounds, `ingest-sap`/LangChain entries, and package-data settings were restored.
A few bugs in the initial ADK implementation were fixed during review:
* `extract_relations()` was calling `RelationExtractor.extract_entities()`, which doesn't exist on that extractor. The failure was being caught and returned in the tool's `error` field, leaving callers with an empty relation list. It now calls the correct extraction path.
* The repo's top-level `mcp/` package shadowed the third-party `mcp` package imported by `google.adk`, causing `google.adk` imports to fail from a normal repo checkout. The local package was moved to `semantica_mcp/mcp/`.
The MCP move needed a follow-up as well. `semantica/cli.py` and four existing tests were still importing from `mcp.*`, and the modules under `semantica_mcp/mcp/` still used the old absolute imports internally. `semantica_mcp` was also missing from the setuptools package include list and had no `__init__.py`, so it wouldn't have been included in an installed package. Those imports and packaging settings are fixed now.
The session service and ADK tools also had a few other problems:
* `list_sessions()` returned a plain list instead of ADK's `ListSessionsResponse`. The original import for that type doesn't work against the installed `google-adk` package, so it was silently falling back to a stub. `user_id` was also incorrectly required instead of being optional.
* Session node IDs were built by joining `app_name`, `user_id`, and `session_id` with unescaped colons, which allowed different identities to produce the same graph node ID. Each component is now encoded before joining.
* `kg_tools.py` and `decision_tools.py` each had their own lock registry and default graph instance. Sharing a graph between the two modules therefore didn't share the lock, and using both factories without an explicit graph produced two different defaults. The shared state now lives in one module used by both.
* `add_to_graph` had a `TypeError` compatibility fallback that couldn't succeed with the current `RelationExtractor` API and could hide the original extraction error. That fallback was removed.
* `append_event` persisted partial streaming events even though ADK's base session service skips them.
* `get_session()` ignored its `config` argument, so `num_recent_events` and `after_timestamp` had no effect.
* The async session-service methods performed synchronous graph scans while holding a `threading.RLock` on the event loop thread. That work now runs in worker threads with `asyncio.to_thread()` so a slow or contended graph operation doesn't block the loop.
---
Co-authored-by: Zohaib Hassnain [109234410+ZohaibHassan16@users.noreply.github.com](mailto:109234410+ZohaibHassan16@users.noreply.github.com)
Addresses review feedback on this PR:
- Guard 2 and the PR-comment JS parser both required every dependency
in pip-audit's report to carry an array-valued `vulns` field. A
dependency pip-audit can't resolve/audit is reported instead as
{"name": ..., "skip_reason": ...} with no `vulns` key at all (see
pip_audit._format.json.JsonFormat._format_dep) - a normal, documented
shape, not a malformed one. That made a single unauditable package
hard-fail the whole job and show "Invalid report structure" in the PR
comment, reintroducing the same class of scan-unrelated CI break this
migration was meant to fix for Safety. Both now accept skipped
entries, treat them as zero vulns, and surface them explicitly (job
log + PR comment) instead of silently dropping or crashing on them.
Verified the fixed jq queries and JS parse logic against synthetic
pip-audit report fixtures covering the normal, skipped, and malformed
shapes.
- Restored a `workflow_dispatch` trigger on security-scan.yml. Deleting
security.yml (which had it) left no way to manually run a dependency
audit on demand.
- Updated SECURITY.md, which still described security.yml as a live
scanning workflow and Safety as an active scanner after this PR
deletes both.
The Security Scan workflow already scans `requirements-ci.txt` directly, but malformed Safety output could still be treated as a clean scan. If the report existed on disk but `vulnerabilities` was missing, `null`, or the wrong type, the workflow could end up counting it as zero findings.
This adds a structural check immediately after the report is written. `vulnerabilities` must be an array; otherwise the step fails closed with a clear error instead of treating a broken report as a successful scan.
There was a related problem in the PR reporting path. The comment step already knew how to render an `Invalid report structure` warning, but that branch was effectively unreachable. In GitHub Actions, a custom `if:` is implicitly gated by `success()` unless it includes a status function such as `always()` or `failure()`. Once the Safety step exited non-zero, the Upload and Comment steps were skipped, so the warning could never be posted.
Fixing that required changing how Safety failures flow through the job rather than just adding another guard. The Safety step now uses `continue-on-error: true`, which lets Bandit and Semgrep continue running and allows the Upload and Comment steps to process the failed or malformed Safety result.
Because `continue-on-error` means the Safety step no longer carries the job's final failure signal itself, the workflow now tracks that state explicitly with `SAFETY_SCAN_STATUS`. It is set to `failed` at the start of the Safety step, before any validation runs, and changes to `passed` only when the report is valid and contains zero vulnerabilities.
That default-failed behavior covers every other exit path: a missing report, malformed `vulnerabilities` field, invalid vulnerability count, Safety failure, or an actual vulnerability finding all leave the status as `failed`.
A final `Enforce Safety Gate` step checks `SAFETY_SCAN_STATUS` and fails the job unless it is exactly `passed`. This keeps the same merge-blocking behavior while still allowing the rest of the security checks and reporting steps to run after a Safety failure.
This is a follow-up to #1356. The overlapping Safety behavior changes and duplicate pip-audit path from that PR were dropped after `main` picked up the canonical fix for the underlying `cuda-toolkit` crash. This change keeps only the report-validation hardening that remains independent of that fix.
* fix(ci): drop --ignore from Safety check, filter accepted CVEs in jq instead
The follow-up to #1370: adding `--ignore SFTY-20260120-40557` to the
`safety check` invocation reintroduced the exact crash #1131/#1157 had
just fixed - "Unhandled exception happened: 'cuda-toolkit'" - but only
once Safety actually has a live vulnerability match to apply the ignore
against (the plain, un-ignored scan against the same requirements-ci.txt
had already succeeded and correctly reported that same match on main,
per the run right before this one).
I couldn't reproduce this locally: my local Safety installation doesn't
surface the live cuda-toolkit CVE match at all (its open-source
vulnerability DB appears to lag CI's), so --ignore never had a real
match to crash on in my testing. That's on me - I should have caught
that my "0 vulnerabilities" local result meant the DB hadn't even seen
the finding yet, not that the fix worked.
Since I can't safely iterate against Safety's own --ignore path without
live-DB access, this moves the "should we still fail on ID X" decision
out of Safety entirely: run the plain scan (the one path an actual CI
run has now proven doesn't crash), then filter the accepted vulnerability
ID out of the report ourselves in jq before counting/printing. Verified
the jq expression directly against a synthetic report shaped like a real
one (id present + one other unrelated id): filters exactly the intended
entry, and - as a bonus - iterating over a null/missing "vulnerabilities"
key with jq now raises inside jq the way the existing guard comment always
assumed it did, rather than silently coming back as 0.
* fix(ci): apply the accepted-CVE exclusion list to the PR comment too
Qodo caught a real gap on this PR: the jq-based exclusion I added only
covers the CI gate (the VULNS count and the failure-path detail print).
The "Comment PR with Security Results" step reads safety-report.json
independently in its own JS, with no filtering at all, so a PR touching
only the accepted cuda-toolkit CVE would still get a comment saying
"Found 1" even though the gate itself correctly treats it as
non-actionable and passes.
Export IGNORED_VULN_IDS via $GITHUB_ENV from the shell step so the JS
step can read the same list, and filter data.vulnerabilities there
before rendering - with a footnote naming what was excluded and why,
so the comment stays transparent about the accepted finding rather than
just silently hiding it.
Verified the JS logic standalone against two synthetic reports: one with
the accepted CVE plus an unrelated real one (shows only the real one,
plus the footnote), and one with only the accepted CVE (shows "No
findings" plus the footnote, rather than misleadingly looking identical
to a clean scan with no explanation).
Merging #1357 surfaced a real (not crashed) Safety finding: cuda-toolkit
13.0.3.0 < 13.1.0 is affected by SFTY-20260120-40557 / CVE-2025-33228.
This can't be fixed with a version bump on our end: torch 2.13.0 (the
latest release on PyPI - there is no newer one) hard-pins
`cuda-toolkit[cublas,cudart,cufft,cufile,cupti,curand,cusolver,cusparse,
nvjitlink,nvrtc,nvtx]==13.0.3` on Linux via its own METADATA, not a loose
transitive requirement we control.
The underlying CVE is OS command injection in NVIDIA Nsight Systems'
gfx_hotspot recipe (process_nsys_rep_cli.py), which requires a human to
manually invoke that script with an attacker-supplied string. It isn't
reachable from any Semantica code path, and Nsight Systems isn't even
part of the extras torch requests here (cublas/cudart/cufft/cufile/
cupti/curand/cusolver/cusparse/nvjitlink/nvrtc/nvtx - no Nsight extra
among them).
Ignoring this one vulnerability ID only (not the whole package or a
blanket policy) so CI reflects actionable risk. Re-evaluate once torch
ships a release that pins a patched cuda-toolkit.
* feat(explorer): serve the RDF export formats the MCP tool already offers (#1131)
`POST /api/export` accepted only `json` and `csv` and answered 422 for everything
else, while the MCP `export_graph` tool resolved Turtle, N-Triples, RDF/XML,
JSON-LD, GraphML and Parquet through `semantica.export`. Two surfaces of one
product disagreeing about what the product can do — and for an RDF-native project,
a graph that loads as JSON-LD and cannot be exported as RDF is a one-way door.
The route now reaches the same exporters the MCP tool uses. Nothing is
reimplemented: `RDFExporter.export_to_rdf` and `GraphMLExporter.export` receive the
dict `session.build_graph_dict()` already builds.
The alias table is a copy of `mcp/tools/export.py::_FORMAT_ALIASES` plus the
spellings the issue mentioned (`ntriples`, `rdf-xml`), and a test asserts the two
tables agree — if either drifts, the formats a caller can use would depend on which
door they came through.
Media types and extensions per serializer, so a Turtle export is `text/turtle` and
not `application/json` with a `.json` name.
The 422 message now names what IS supported. The old one said only that the format
was unsupported, which reads as "this format does not exist" rather than "this door
does not open it" — that is what sent me looking through the library.
Parquet is left out on purpose: `ParquetExporter.export` writes a file and returns a
path, so serving it over HTTP is a different shape of change and deserves its own
review.
Tests, in `TestImportExport`: the seven RDF spellings, each **parsed with rdflib**
rather than asserted on strings — a response that merely looks like Turtle is what
lets this class of gap survive a suite. Plus the alias-agreement canary and the
error message. Three mutations (rejecting RDF again, breaking one alias, emptying
the message) each turn the matching tests red.
110 tests in `tests/explorer/test_explorer_api.py` pass.
* fix(explorer): complete RDF export support
* fix(explorer): secure GraphML temporary file handling
* fix(ci): scan declared dependencies with Safety
---------
Co-authored-by: 13g4d0 <13g4d0@users.noreply.github.com>
Co-authored-by: Sameer Kadam <sskadam6305@gmail.com>
* fix(ci): drop unpinnable benchmarks/requirements.txt install
Scorecard flagged this pip install as unpinned-by-hash (#6082). Can't
hash-pin it - benchmarks/requirements.txt doesn't exist in this repo, so
there's nothing to compile a lockfile from. Dropping it instead of
leaving it unpinned: the job already fails on the next real step
(benchmarks/benchmarks_runner.py, also missing), so this line wasn't
doing anything useful to begin with.
* fix(ci): hash-pin the spacy model download in benchmark.yml
Qodo review on this PR: dropping the benchmarks/requirements.txt install
(the previous failure point) let the job actually reach
`python -m spacy download en_core_web_sm`, which fetches an unpinned,
unhashed wheel from spacy-models' GitHub releases - undoing the point of
this PR by exposing a real unpinned-install path instead of a dead one.
Replaced with a hash-pinned direct-URL entry in benchmark-extra.in/.txt
for en_core_web_sm-3.8.0 (matches the spacy==3.8.15 already pinned in
base-deps.txt). uv independently computed the same sha256 I got via a
manual curl+sha256 of the release asset, and a --require-hashes dry-run
install verifies clean.
* fix(ci): stop checkov's suppressed checks from reopening as new alerts
Root cause found, not just worked around: checkov's SARIF exporter
includes every evaluated check as an ordinary result, including ones it
internally marked SKIPPED via the inline # checkov:skip= comments and
checkov.io/skipN annotations already on the Helm chart. It never uses
SARIF's own `suppressions` field and never drops them - so the exact same
already-suppressed finding reopens as a brand-new code scanning alert
number on every single run, forever (#6035/#6036, #6112-6115,
#6128-6131 are all the same 4 findings, manually dismissed 3 times now).
checkov's JSON output *does* correctly record which checks were skipped.
Added .github/scripts/filter_checkov_skipped.py, which cross-references
the JSON's skipped_checks against the SARIF's results (matched by check
ID + the last two path segments, since the two outputs use different path
roots) and drops anything checkov itself already decided to suppress,
before upload. Verified locally against a real checkov+helm run: removed
exactly the 4 known-suppressed helm chart results, left the 2 genuinely
real findings (deploy/gcp/cloudrun-service.yaml, deploy/kubernetes/
deployment.yaml) untouched.
* fix(ci): drop unpinnable benchmarks/requirements.txt install
Scorecard flagged this pip install as unpinned-by-hash (#6082). Can't
hash-pin it - benchmarks/requirements.txt doesn't exist in this repo, so
there's nothing to compile a lockfile from. Dropping it instead of
leaving it unpinned: the job already fails on the next real step
(benchmarks/benchmarks_runner.py, also missing), so this line wasn't
doing anything useful to begin with.
* fix(ci): hash-pin the spacy model download in benchmark.yml
Qodo review on this PR: dropping the benchmarks/requirements.txt install
(the previous failure point) let the job actually reach
`python -m spacy download en_core_web_sm`, which fetches an unpinned,
unhashed wheel from spacy-models' GitHub releases - undoing the point of
this PR by exposing a real unpinned-install path instead of a dead one.
Replaced with a hash-pinned direct-URL entry in benchmark-extra.in/.txt
for en_core_web_sm-3.8.0 (matches the spacy==3.8.15 already pinned in
base-deps.txt). uv independently computed the same sha256 I got via a
manual curl+sha256 of the release asset, and a --require-hashes dry-run
install verifies clean.
Dependabot flagged 12 aiohttp advisories (1 high, rest moderate/low - CVE
range covering request smuggling, websocket/parser bugs, cookie/redirect
issues) against aiohttp==3.13.5 pinned in checkov.txt. checkov==3.3.1
itself pinned `aiohttp<3.14.0`, which excludes every fixed release;
3.3.16 (latest) relaxes that to `<3.15.0`, so bumping checkov also lets
aiohttp resolve to 3.14.3 (fixes all of them).
Two alerts remain open, both genuinely blocked upstream rather than
something a version bump here can fix:
- asteval: checkov 3.3.16 (latest, still) hard-pins asteval==1.0.6 with
no range; the fix (1.0.9) is unresolvable without violating checkov's
own declared dependency - confirmed via `uv pip compile` refusing to
solve it. Needs checkov itself to bump the pin upstream.
- ecdsa: 0.19.2 is already the latest release; the Minerva timing-attack
advisory has no patched version, since python-ecdsa's maintainers have
stated side-channel attacks are out of scope for the project.
Both are checkov's own transitive deps, used only for local static IaC
analysis in defender-for-devops.yml (no network signing/cloud-auth calls
that would actually exercise ecdsa's signing path) - dismissing on
GitHub with that reasoning as a separate step.
* fix(docker): split explorer-extra.txt by Python version, fix broken build
main's container-scan.yml has been failing since PR #1338 merged:
ERROR: In --require-hashes mode, all requirements must have their
versions pinned with ==. These do not:
standard-aifc from .../standard_aifc-3.13.0-py3-none-any.whl
(from audioread==3.1.0->-r explorer-extra.txt (line 30))
Root cause: explorer-extra.txt was compiled with `--python-version 3.11`
but is installed on the Dockerfile's actual python:3.13-slim interpreter.
librosa's audioread dependency needs standard-aifc/standard-sunau only
under `python_version >= "3.13"` (Python 3.13 dropped aifc/sunau from
stdlib) - a file resolved for 3.11 has no hash for those packages at all,
so --require-hashes fails outright once pip resolves against the real
3.13 environment instead of silently under-pinning.
Splits the file in two: explorer-extra-py311.txt (ci.yml, unchanged
resolution) and explorer-extra-py313.txt (Dockerfile, newly compiled for
--python-version 3.13). They aren't interchangeable and shouldn't be
recombined - documented in .github/requirements/README.md, including how
to catch this class of bug before it ships again.
* fix(ci): correct stale -o path in explorer-extra-py311.txt header
Qodo review on this PR: the autogenerated header comment still said
-o .github/requirements/explorer-extra.txt (the pre-rename path), which
would silently regenerate the wrong file if someone copy-pasted it.
Qodo review on this PR: `pip install --no-deps -e .` / `pip install
--no-deps .` still leaves PEP 517 build isolation on by default, which
fetches [build-system] requires (setuptools==84.0.0, wheel==0.48.0)
completely outside any hash checking - the --require-hashes installs
right next to it didn't cover this at all.
Adds .github/requirements/pep517-build.txt, hash-locked to the exact
pyproject.toml [build-system] requires, and installs it before every
local-source install (Dockerfile, ci.yml, benchmark.yml) with
--no-build-isolation so pip reuses those hash-verified copies instead
of fetching its own.
Scorecard's Pinned-Dependencies check requires pip installs to be
hash-verified, not just version-pinned - our existing pkg==X.Y.Z pins
(and even pip install -r requirements-ci.txt, despite that file already
carrying hashes) still scored a 4 because the pin/hash isn't visible on
the command line itself.
Adds .github/requirements/*.txt: hash-locked files generated via
`uv pip compile --generate-hashes` for every pip target that isn't
already requirements-ci.txt, covering standalone CI tooling (build,
wheel, twine, uv, pip-audit, safety/bandit/semgrep/jq, pip/setuptools
bootstrap) and the project's own local-source installs. The latter
(`pip install -e ".[explorer]"`, `pip install -e .`) can't be hash-pinned
directly since there's nothing to hash for a local source tree; split
into `pip install --no-deps -e .` plus a separate hash-pinned install of
the actual fetched dependencies instead.
Also adds --require-hashes to every `-r requirements-ci.txt` install so
hash verification is enforced explicitly rather than only implied by the
file's own content.
Simplifies the Dockerfile in the process: it now installs from the same
pre-generated explorer-extra.txt (copied in at build time) instead of
extracting constraints from requirements-ci.txt at build time, which
also means setuptools gets its CVE-2025-47273 fix as a side effect of
the hash-pinned install rather than a separate upgrade step.
benchmark.yml: pip install -r benchmarks/requirements.txt is left
unpinned - that directory doesn't exist in this repo, so there's nothing
to generate hashes from. Pre-existing breakage, unrelated to this change.
Two terrascan/GHAS findings on the previous commit:
- AC_DOCKER_0052 (no apt-get upgrade in Dockerfiles): dropped it. It also
wasn't fixing anything - Debian's openssl fix for CVE-2026-14456 is still
in trixie-proposed-updates, not reachable via a normal upgrade. Pin both
base images by digest instead (matches #1329's approach) so the docker
Dependabot ecosystem bumps them once Debian ships a rebuilt image with
the fix, and document why the QUIC DoS isn't reachable here regardless
(HTTP-only via uvicorn).
- AC_DOCKER_0010 (pin pip package versions): setuptools was `>=78.1.1`;
pinned to the exact 84.0.0 already used by pyproject.toml/requirements-ci.txt.
Also fixes two build breaks this introduces on its own: requirements-ci.txt
wasn't in .dockerignore's allowlist or container-scan.yml's path trigger,
so the COPY in the prior commit would have failed the image build outright.
* fix(ci): unblock py3.9 install matrix and raise Scorecard pinning/signing
pip install semantica failed on Python 3.9 across all three OSes because
spacy had no upper bound, so pip resolved spacy 3.8.16 whose thinc>=8.3.12
requirement has no cp39 wheels and no working sdist build path. Cap
spacy/thinc for python_version < '3.10' to the last wheel-compatible pair.
Also addresses the two OpenSSF Scorecard findings that were actually
fixable in code:
- Pinned-Dependencies: Dockerfile base images (node:26-alpine,
python:3.13-slim) were unpinned by digest; pin both, and pin five
previously-unversioned pip install calls in CI (build, safety, bandit,
semgrep, jq, pip-audit).
- Signed-Releases: attest-build-provenance only publishes to the GH
attestations API, which Scorecard doesn't inspect. Sign dist/* with
Sigstore and attach the .sigstore.json bundles as release assets.
* fix(ci): correct Sigstore artifact inputs
---------
Co-authored-by: Sameer6305 <sskadam6305@gmail.com>
* ci: add npm Dependabot ecosystem and container image scanning
- dependabot.yml had no npm ecosystem entry for explorer/, so its
lockfile was never watched - exactly why the brace-expansion/nanoid
CVEs fixed in #1280 went undetected. Add it, mirroring the existing
pip entry's schedule/labels/reviewer conventions.
- New container-scan.yml builds the root Dockerfile's image and scans
it with Trivy (CRITICAL/HIGH OS+lib CVEs, SARIF to the Security tab)
and Syft (SPDX SBOM artifact), on push to main, weekly, and manual
dispatch. Neither the base-image scan nor an SBOM existed before -
Dependabot's docker entry only bumps the base image tag, it doesn't
scan built layers.
- Trivy runs report-only for now (no exit-code gate): this is its
first run against the image, so the CRITICAL/HIGH baseline hasn't
been triaged yet. Once reviewed, add exit-code: '1' to make it a
hard gate, same as Safety/Bandit-HIGH in security-scan.yml.
* fix: run Trivy via digest-pinned image, not the aquasecurity/trivy-action wrapper
verify-action-pins.sh failed in CI: the aquasecurity GitHub org has an IP
allow list on its API that 403s the live tag->SHA resolution from
Actions-runner IPs (confirmed reproducible, not transient - resolves fine
from a non-blocked host). Rather than carve a skip exception into the pin
verifier for an org this script already flags as a past tag-repointing
target (see its "LiteLLM/Trivy 2026 incident" comment), pull Trivy as a
sha256-digest-pinned Docker Hub image instead. A digest is immutable and
verifiable independently of GitHub's API entirely, so it sidesteps the
IP block without weakening verification of the one action this repo
already treats as higher-risk. Confirmed the pinned digest
(aquasec/trivy@sha256:62b1e65e...) resolves live against Docker Hub's
registry API.
* fix: match container-scan.yml's push paths to what actually reaches the image
The path filter only watched explorer/package.json and package-lock.json,
but Dockerfile COPYs the whole explorer/ tree plus README.md, LICENSE, and
MANIFEST.in, and .dockerignore controls all of it. A frontend source change
or a README/LICENSE edit would change the built image without triggering a
scan, silently drifting until the next weekly run. Replace the filter with
exactly .dockerignore's opt-in list.
- Bump explorer's brace-expansion (minimatch dep) 5.0.8 -> 5.0.9 and
nanoid (postcss dep) 3.3.16 -> 3.3.18, fixing GHSA-rgw5-rvv9-x895 and
GHSA-2v37-7h3g-55p8 (both DoS via unbounded input, both within the
existing caret ranges declared by their parents).
- Move codeql.yml and defender-for-devops.yml's security-events: write
(and codeql.yml's actions: read) from workflow-level down to their
single job, matching Scorecard's Token-Permissions ideal of a
read-only top-level default with sensitive scopes granted only where
used.
Distribution and trust-signal infrastructure to make pip install semantica
frictionless in downstream CI, and to bring the release pipeline in line
with mature OSS practice.
- .github/actions/setup-semantica: reusable composite action other repos
can call to install + verify semantica in one step
- install-matrix.yml: verifies the published package installs and imports
cleanly across Ubuntu/macOS/Windows x Python 3.9-3.12, weekly and on
release; backs a new README badge
- scorecard.yml: OpenSSF Scorecard analysis, weekly and on push to main,
backing a new README badge
- release.yml: twine check gate before publish, catching a broken PyPI
long-description render before it ships
- CITATION.cff: enables GitHub's native "Cite this repository" button
- examples/ci/: copy-paste GitHub Actions, GitLab CI, and CircleCI
templates for projects adopting semantica
- GROWTH.md: tracked checklist of distribution channels, what's done vs
outstanding, with guardrails against inflating metrics artificially
Fixes folded in along the way:
- Re-pinned softprops/action-gh-release to the immutable v3.0.3 tag
instead of the floating v3, after verify-action-pins.sh caught the
mutable tag had drifted to a newer commit
- setup-semantica now passes extras/version through env vars instead of
interpolating ${{ inputs.* }} directly into the bash script, closing
a script-injection vector for callers deriving these from event data
- install-matrix now triggers on the Release workflow's completion
(workflow_run) instead of release: published, since the GitHub release
is created before the PyPI upload runs and the old trigger could race
the publish
- The workflow_run path derives the expected version from the triggering
tag and passes it into setup-semantica's version input, so pip
installs and verifies the exact release instead of whatever's latest
on PyPI at the time
- setup-semantica's pip caching is now opt-in (default disabled), since
actions/setup-python errors out with cache: 'pip' enabled when the
caller repo has no requirements.txt/pyproject.toml to key on
- examples/ci/github-actions.yml pins actions/checkout and
actions/setup-python to verified commit SHAs instead of mutable tags
- examples/ci templates guard the requirements.txt install step with
-f requirements.txt and call out pyproject.toml/Poetry/Pipenv as
alternatives, since not every project has a requirements.txt
feat(explorer): add deterministic rendering E2E example and test (#1037)
Adds a deterministic Explorer graph baseline and coverage for the full
build -> persist -> API -> frontend hydration -> canvas rendering path,
so a regression anywhere along that chain shows up in CI instead manually.
examples/explorer_deterministic_rendering_example.py builds the
canonical 4-node, 3-edge graph (Alice -WORKS_AT-> Acme, Bob -KNOWS->
Alice, Acme -LOCATED_IN-> New York) with ContextGraph.add_node()/
add_edge(), persists it with save_to_file() and reloads it with
GraphSession.from_file(), printing the setup prerequisites and the
expected node/edge/label checklist for anyone running it by hand.
tests/explorer/test_explorer_deterministic_rendering_e2e.py covers
graph construction, the serialize/deserialize round trip, GraphSession
loading, and the Explorer API's /api/graph/* responses against the
exact expected nodes, edges, and labels, plus all three auth modes
(unconfigured, API-key required, anonymous opt-in).
fix(explorer): address Qodo review findings for deterministic rendering e2e (#1037)
- configure SEMANTICA_ALLOW_ANONYMOUS=true and document
SEMANTICA_API_KEY as the alternative in the reproduction
instructions, so the documented commands don't 503 on a clean
checkout
- add clean-checkout prerequisites and a visual verification
checklist to the example
- add edge-label (WORKS_AT, KNOWS, LOCATED_IN), zoom-tier, and
hover-interaction coverage to the frontend test
- add an explicit auth-enforcement integration test for the
deterministic graph endpoints
fix(explorer): connect deterministic rendering E2E path
The frontend test built its own node/edge objects directly with
batchMergeNodes()/batchMergeEdges(), bypassing the real loading path
entirely -- it never went through useLoadGraph, never mounted the
canvas, and its fixture didn't even carry the same fields the backend
actually returns (e.g. no color values), so a break in API hydration,
the edge.type -> edgeType mapping, or canvas label rendering could
still pass.
Adds deterministicExplorerRendering.e2e.ts, which mounts the real
Explorer app in Chromium, serves API-shaped /api/graph/nodes and
/api/graph/edges responses through route interception, drives the
app through its actual useLoadGraph hydration path into a real Sigma
canvas, and asserts on captured canvas fillText() calls that
WORKS_AT, KNOWS, and LOCATED_IN are genuinely drawn, both after load
and after Zoom In.
fix(explorer): preserve upstream markdown dependencies
ci(explorer): isolate deterministic backend test dependencies
Wires the new Python test into ci.yml as its own focused step (it
previously only ran manually), installs Playwright's Chromium
browser before the frontend suite, and keeps the deterministic
backend test's dependency install separate from the rest of the
pipeline so it doesn't pull in unrelated optional extras during
collection.
fix(explorer): remove redundant edge label hydration
An earlier commit in this PR added an explicit `label` field to
hydrated edge attributes on the theory that it was needed for edge
labels to render. Review traced through GraphCanvas.tsx's label
resolution (`attrs.edgeType || data.label || ""`, from the earlier
#1009 fix already on main) and found that `edgeType` is set
unconditionally on every edge during hydration, so it always wins the
`||` before `data.label` is ever consulted -- the added field and its
plumbing in useLoadGraph.ts and graphStore.ts never did anything.
Removed both; reran the real Chromium E2E test against the reverted
code and confirmed all three labels still render identically, closing
out the question of whether anything else was actually broken.
Add a Cite Us section to the README with BibTeX citation info, and
align it with docs/citation.md (author/organization: Semantica, 2026).
Update LICENSE and docs/project-license.md copyright holder to
Semantica, and replace the stale Hawksight-AI GitHub org slug with
semantica-agi across READMEs, plugin manifests, cookbook notebooks,
and GitHub templates.
* docs(contributing): formalize issue assignment and duplicate-PR triage workflow
Comments are no longer required before an issue can be assigned - maintainers
may assign directly based on recent activity. Also documents the duplicate-PR
priority order for triage (contributor PR, claimed issue, activity tiebreak,
late duplicates, overlapping scope).
* docs(contributing): clarify assignment precedence and define activity tiebreak
Addresses Qodo review feedback on PR #1030: the duplicate-PR priority list
now states these rules apply on top of the assignment workflow (opening a PR
pre-assignment doesn't grant priority), and the "most active" tiebreak now
specifies a concrete 60-day window and signals instead of being subjective.
The pin was 5595ccaf..., but upstream has since moved the v4 tag to
ff2f1c62.... The Verify Action Pins workflow flags this drift on every
PR that touches any workflow file, regardless of whether that PR
changed codeql.yml or defender-for-devops.yml.
Verified the new SHA against the GitHub API directly (not just the CI
error text) and confirmed .github/scripts/verify-action-pins.sh passes
clean locally (40/40 action references OK, exit 0).
* ci: pin Python dependencies in requirements-ci.txt for reproducible CI
Adds a committed lockfile pinning all transitive dependencies at exact
versions (uv pip compile, Python 3.11, all extras — 1581 lines), the
Python equivalent of explorer/package-lock.json + npm ci.
- CI installs from requirements-ci.txt before building the wheel
- CI verifies the lockfile is byte-identical to a fresh compile (fails
on staleness after pyproject.toml changes)
- CONTRIBUTING documents the regeneration command
Closes#938
Signed-off-by: Yunare Maia <yunare@gmail.com>
* ci: address Qodo review — security scans use pinned deps, exclude gpu extras
- security-scan.yml installs from requirements-ci.txt instead of
"./[llm-litellm]" so Safety scans the exact CI/release dependency tree
- security.yml runs pip-audit -r requirements-ci.txt for the same parity
- lockfile regenerated with --extra all (the cross-platform set) instead
of --all-extras, which pulled faiss-gpu/cupy from the Linux-only gpu
extra and co-installed faiss-cpu + faiss-gpu in CI
- uv pinned to 0.12.1 (the version that generated the lockfile) in CI and
CONTRIBUTING so regeneration is deterministic
Signed-off-by: Yunare Maia <yunare@gmail.com>
* ci: make lockfile staleness check immune to upstream releases
The previous check re-resolved pyproject.toml without constraints, so any
upstream package release (e.g. boto3 1.43.69 -> 1.43.70) failed CI even
when nothing in the repo changed — exactly the time-dependent drift Qodo
flagged. The check now re-resolves with requirements-ci.txt as a
constraint and compares only version lines, so it detects intentional
pyproject.toml changes but ignores upstream releases. CONTRIBUTING
updated to match.
Signed-off-by: Yunare Maia <yunare@gmail.com>
* ci: fix security workflows — install pip-audit; order tooling after pinned deps
Security workflow: the pip-audit install step was lost in the rebase
conflict merge — pip-audit was invoked but never installed (exit 127).
Security-scan workflow: installing safety first let the pinned
requirements-ci.txt overwrite its transitive deps (rich), breaking the
safety CLI at runtime (RuntimeError: Type not yet supported). Tooling is
now installed AFTER the pinned set.
Signed-off-by: Yunare Maia <yunare@gmail.com>
* fix(ci): address review — hashes, build isolation, release builds, docs (4/4)
ZohaibHassan16's review flagged 4 supply-chain gaps; all addressed:
1. **Release builds now use the lockfile**: release.yml installs
requirements-ci.txt and runs `python -m build --no-isolation` so the
sdist/wheel is built against the exact tested dependency set.
2. **Build isolation pinned**: [build-system].requires is now
setuptools==84.0.0 + wheel==0.48.0 (exact pins, no ranges).
3. **Hashes**: requirements-ci.txt regenerated with --generate-hashes
(5,708 sha256 hashes, verified against PyPI). Staleness check updated
to strip the `\` line continuations hashes introduce.
4. **CONTRIBUTING.md documents the separate environment**: hashes,
never-install-into-dev note, build-system pins, --no-isolation release
builds.
Validated: stale-check diff clean, hash spot-check matches PyPI.
Signed-off-by: Yunare Maia <yunare@gmail.com>
* fix(ci): apply --no-isolation to CI build + align benchmark to Python 3.11
Follow-up to ZohaibHassan16's second review round:
1. ci.yml was still running `python -m build` with build isolation
(unpinned setuptools/wheel from PyPI) — now `python -m build
--no-isolation` against the pinned deps, matching release.yml.
2. benchmark.yml was on Python 3.12 while the lockfile is compiled for
3.11 — aligned to 3.11 so every workflow runs the same environment.
Signed-off-by: Yunare Maia <yunare@gmail.com>
* fix(ci): install pinned wheel before --no-isolation build
python -m build --no-isolation failed with 'Missing dependencies:
wheel==0.48.0' because wheel is build-time only — uv's lockfile
excludes it, so installing requirements-ci.txt alone left the build
env without it. Both ci.yml and release.yml now install wheel==0.48.0
(the same pin [build-system] declares) before building. Validated
locally: wheel builds clean with --no-isolation.
Signed-off-by: Yunare Maia <yunare@gmail.com>
---------
Signed-off-by: Yunare Maia <yunare@gmail.com>
Co-authored-by: Zohaib Hassnain <109234410+ZohaibHassan16@users.noreply.github.com>
* security(deps): bump fastapi floor to >=0.109.1 (PYSEC-2024-38)
The [explorer] extra declared fastapi>=0.100.0, which allows the
vulnerable 0.109.0 (PYSEC-2024-38, HTTP response splitting). Raise the
floor to 0.109.1, the patched release. One-line change, no functional
impact -- the 0.109.x API is stable and backward-compatible.
Fixes#869
* fix(deps): bump fastapi to >=0.109.2 and python-multipart to >=0.0.7 for PYSEC-2024-38
PYSEC-2024-38 (CVE-2024-24762 / GHSA-2jv5-9r88-3w3p) is a ReDoS in
python-multipart < 0.0.7: an attacker sends a crafted Content-Type header
that causes catastrophic backtracking in the multipart regex, stalling the
event loop and causing a DoS on any endpoint that parses form data.
The original PR bumped fastapi to >=0.109.1, but that version pins
starlette<0.36.0,>=0.35.0 and cannot install starlette 0.36.2+ (which
contains the fix via python-multipart>=0.0.7). FastAPI 0.109.2 is the
first version that pins starlette>=0.36.3 (verified against PyPI metadata).
Two changes are necessary:
1. fastapi>=0.109.1 -> fastapi>=0.109.2: ensures starlette>=0.36.3 is
installed as a transitive dependency, which in turn pulls the fixed
python-multipart>=0.0.7.
2. python-multipart>=0.0.6 -> python-multipart>=0.0.7: closes the direct
dependency path. python-multipart is listed explicitly in the explorer
extra, so without this floor a resolver could still install 0.0.6 and
leave the vulnerability present even with the fastapi bump.
The fix targets only the 'explorer' optional dependency group, which is
the only code surface where FastAPI and form-data parsing are used.
No functional API changes between 0.109.1 and 0.109.2; 239 Explorer tests
pass without modification.
* ci(security): gate pip-audit on explorer-extra dependency PRs, add changelog entry for PYSEC-2024-38
The Security workflow's pip-audit job ran weekly against a bare Python
env with none of Semantica's optional extras installed, and always
continue-on-error'd -- it would never have flagged the vulnerable
fastapi/python-multipart floors this PR fixes, or the first attempt at
the fix that left python-multipart>=0.0.6 in place. security-scan.yml's
Safety check has the same blind spot (only installs [llm-litellm]).
pip-audit now also runs on pull_request when pyproject.toml changes,
installs semantica[all] so it can actually see extras like [explorer],
and fails the build on findings for that trigger. Scheduled/dispatch
runs stay non-blocking pending a full pass over the [all] tree.
Also documents the fix (#871, closes#869) in CHANGELOG.md, including
the correction made during review after the original fastapi-only bump
turned out not to close the vulnerability.
* fix(deps): raise setuptools floor to >=83.0.0 (CVE-2026-59890), harden audit env
The new pull_request pip-audit gate (previous commit) caught this on its
first run: pip install -e ".[all]" resolved setuptools==79.0.1, vulnerable
to CVE-2026-59890 / GHSA-h35f-9h28-mq5c / PYSEC-2026-3447 (Unicode
normalization lets a MANIFEST.in exclude/prune pattern be bypassed on
macOS APFS/HFS+, leaking excluded files into a built sdist). Fixed in
setuptools 83.0.0.
[build-system] requires had the same too-permissive floor this whole PR
is about (setuptools>=61.0). Raised to >=83.0.0. Also upgrade pip/
setuptools explicitly in the Security workflow before running pip-audit,
since [build-system] requires only governs isolated build environments,
not the ambient one actions/setup-python provisions and pip-audit scans.
---------
Co-authored-by: Sameer Kadam <sskadam6305@gmail.com>
Co-authored-by: KaifAhmad1 <kaifahmad087@gmail.com>
Upstream moved the v4 tag to 5595ccaf912efad79be6eef63a5619ff05969be3
(v4.37.6), which the repo's own verify-action-pins.sh now (correctly)
flags as a mismatch against the previously-pinned commit. Pre-existing
drift unrelated to #830/#836, but it was failing this PR's required
"verify" check, so fixing it here.
- Wire the Explorer frontend's node --test suites (test:graph-store,
test:graph-workspace, and the new test:plugin-registry regression
test) into CI. Previously only `npm run build` ran, so none of the
frontend tests -- including this fix's own regression coverage --
executed anywhere except a contributor's local machine.
- Broaden the diagnostics dedup's structureLayer comparison to also
cover disabledReason/curveCount/bridgeCurveCount/backboneCurveCount,
not just cacheKey/lastDrawAt/enabled, so a disabledReason-only
transition doesn't leave the dev diagnostics panel stale.
* security: SHA-pin all Actions, harden release pipeline, add pin verification
Hardens the CI/CD supply chain against the LiteLLM/Trivy-style attack (a
compromised third-party Action with a mutable tag stealing a long-lived
publishing token) and closes several related gaps found in an audit of the
actual repository state.
- Pin every third-party GitHub Action across all workflows to a full commit
SHA (tag kept as a trailing comment); add verify-action-pins.yml, a CI
check that confirms via the GitHub API that each pin still matches its
tag, on every workflow change, push to main, and weekly.
- Scope release.yml permissions to the job level (workflow defaults to
contents: read); add a concurrency group so simultaneous tag pushes can't
race the publish job.
- Add SLSA build provenance attestation (actions/attest-build-provenance)
for every released wheel.
- Fix a latent bug in security-scan.yml: the PR-comment step was missing
pull-requests: write and silently failing; add bounded artifact retention
for uploaded scan reports.
- Group github-actions Dependabot updates to cut review noise.
- Document the resulting posture in SECURITY.md for auditors/regulated
adopters, including what's enforced and what a fork needs to reconfigure
for itself (environment/branch protection, Trusted Publishing trust).
Also (via GitHub API, not in this diff): created a protected `pypi`
environment with a required reviewer restricted to v* tags, and enabled
branch protection on main (required review, required status checks, no
force-push/deletion).
* fix: harden verify-action-pins per PR #824 bot review
Addresses real findings from the automated review on #824:
- The script previously only matched uses: lines that already contained a
40-hex SHA, so a newly added mutable-tag action (e.g. some/action@v1)
would never be scanned at all and the check would pass silently. It now
matches every uses: line and hard-fails on any ref that isn't a full
commit SHA.
- A tag that fails to resolve via the GitHub API (rate limit, deleted tag)
previously only logged a warning and continued; that's now a hard
failure too, since an unverifiable pin is exactly the failure mode this
check exists to catch.
- verify-action-pins.yml only triggered on .github/workflows/** changes,
so an edit to the verifier script itself wouldn't run the check that
verifies it. Added the script path to both trigger filters.
The reviewer's claim that slash-containing tag comments (release/v1) break
the API lookup did not reproduce - tested directly against
pypa/gh-action-pypi-publish@release/v1 and GitHub's commits API resolves
multi-segment refs natively - so no change was needed there.
Verified with a synthetic test workflow containing a mutable-tag action,
a correctly-pinned SHA, and a deliberately mismatched SHA: the updated
script now catches the first and third cases and passes the second. Also
re-ran against the real workflow tree (40/40 pins still verify clean).
* fix: repair broken Safety scan and PR comment formatting
The "Comment PR with Security Results" step was producing garbled output
(literal \n characters instead of newlines, "undefined:" labels) because:
- Every line in the JS comment builder used \n (escaped backslash-n)
inside template literals, which JS renders as the literal two-character
string \n, not a newline.
- The Semgrep section read issue.rule_id, but Semgrep's JSON field is
check_id - hence "undefined: <path>" for every entry.
Rewrote the comment builder to construct each section as an array of
lines joined with a real '\n', with correct field names, and collapsed
long finding lists into a <details> block instead of a flat list.
Verified by extracting the exact script and running it under node against
synthetic fixtures matching each tool's real JSON schema (found/clean/
missing-report paths all render correctly).
While tracing the "undefined" and always-empty Safety section, found the
Safety step itself was silently broken:
- `safety check --json --output safety-report.json` is invalid in
Safety 3.x: --output now selects a console format (json/text/screen),
not a file path. The command errored on every run (swallowed by
`|| true`), so safety-report.json was never created and the PR comment
always fell back to a generic "scan completed" message. Switched to
`--save-json`, which is the correct flag for writing a JSON report to
disk, and confirmed against the real safety 3.8.1 CLI locally.
- Even with a report, the code read vuln.package - the real field is
package_name.
- The job never installed Semantica's own dependencies before scanning,
so `safety check` (which defaults to scanning the environment) was
auditing the scanner tools' own dependencies, not Semantica's. Added
`pip install -e ".[llm-litellm]"` so the project's actual dependency
tree - including the LiteLLM extra this whole hardening effort is
about - is what gets scanned.
Also updated the corresponding SECURITY.md bullet to describe what Safety
actually covers now.
* fix: remove unused pypdf2 dependency (CVE-2023-36464)
Now that the Safety scan step actually runs (see previous commit), it
correctly failed this PR's checks on CVE-2023-36464 in pypdf2==3.0.1 - a
real, pre-existing vulnerability that was invisible until the scan was
fixed.
PyPDF2 is not a patchable dependency here: the project is discontinued
(merged into `pypdf`), 3.0.1 is its final release, and there is no fixed
version to upgrade to. Grepping the repo for `import PyPDF2` / `from
PyPDF2` turns up nothing - it was never actually imported anywhere. Its
only presence outside pyproject.toml was in docstrings describing a
"PyPDF2.PdfReader() fallback" for PDF parsing that was never implemented
in code; pdfplumber is the library actually used. Removed the dependency
and corrected the stale docstrings in parse/__init__.py, parse/methods.py,
parse/pdf_parser.py, and ingest/email_ingestor.py accordingly.
* fix: suppress Bandit B324 false positives on non-cryptographic MD5 use
Same pattern as the previous pypdf2 commit: fixing the Safety scan
surfaced this PR's own Bandit HIGH-severity gate actually blocking on 10
pre-existing findings, all Bandit B324 ("Use of weak MD5 hash for
security").
Checked each of the 10 call sites: every one uses hashlib.md5() to build
a short deterministic cache key, entity ID, or IRI suffix from already-
non-secret input (query text, entity text/type, class/property names) -
none are used for passwords, tokens, or integrity verification of
untrusted data. This is exactly the case Bandit's own message points at
("Consider usedforsecurity=False").
Did not use usedforsecurity=False itself: that keyword argument was
added to hashlib in Python 3.9, and pyproject.toml declares
`requires-python = ">=3.8"` - adding it unconditionally risks a TypeError
on 3.8. Used a targeted `# nosec B324` comment with a one-line
justification instead, which suppresses only this specific check and
carries no runtime behavior change on any supported Python version.
Verified locally: bandit -r semantica/ -ll now reports 0 HIGH-severity
findings (was 10).
* docs: add CHANGELOG entry for #824 CI/CD supply-chain hardening
Covers the SHA-pinning + verify-action-pins.yml enforcement, release.yml
hardening (job-scoped permissions, concurrency, SLSA provenance), the
pypi environment/branch protection GitHub-side config, the
security-scan.yml Safety/comment-formatting fixes, and the two
vulnerabilities those fixes surfaced (pypdf2 CVE-2023-36464 removal,
Bandit B324 suppression).
* fix: close two remaining gaps missed by upstream bot-review fixes
verify-action-pins.sh:
- Quoted uses: lines (e.g. uses: owner/action@SHA) were not matched
by the existing regex, so a SHA-pinned action written with quotes would
silently skip verification. Updated the main ERE to accept an optional
leading/trailing single or double quote around the owner/action@ref
value, and excluded quote chars from the inner character classes so the
ref is still extracted cleanly.
- The grep input glob only covered *.yml. GitHub also treats *.yaml as a
valid workflow extension. Added *.yaml to the glob and a 2>/dev/null
guard so the command doesn't fail when no *.yaml files exist.
security-scan.yml (on top of Kaif's --save-json fix in 67c7ec2a):
- Kaif's fix kept the '|| echo 0' fallback on the VULNS= line, so all
five scanner-failure modes (file missing, empty file, malformed JSON,
valid JSON with no 'vulnerabilities' key, vulnerabilities: null) still
silently produce VULNS=0 or VULNS=null and pass the merge-blocker check.
- Added guard 1: '[ ! -s safety-report.json ]' fails loudly if Safety
crashed before writing a report (covers missing and empty-file cases).
- Dropped the '|| echo 0' fallback and added guard 2: '[[ ! VULNS =~
^[0-9]+$ ]]' fails loudly on non-integer VULNS (covers malformed JSON,
missing key, and null cases). Both guards emit ::error:: annotations.
- Verified with a 7-case simulation: all 5 failure modes now exit 1;
genuine zero-vuln and real-vuln cases still behave correctly.
* fix: correct bash [[ =~ ]] quoting that broke verify-action-pins.sh in CI
The regex for matching uses: lines was embedded directly inline in a
[[ =~ ]] test with literal \" and \' escape sequences. Bash's conditional-
expression parser interprets these as shell syntax rather than regex
literals, producing:
syntax error in conditional expression: unexpected token ')'
at line 27 on every CI run.
Fix: move the regex into a USES_PATTERN variable using safe single-quote
shell-string concatenation so the [[ =~ ]] parser receives an unquoted
variable reference ($USES_PATTERN) rather than a literal pattern containing
bash-special characters. The regex semantics are identical: optional
leading/trailing quote around owner/action@ref, quote chars excluded from
capture groups.
Verified in real bash 5.2.21 (Git for Windows):
No syntax error on the real 40-pin workflow tree (Checked 40)
Unquoted SHA pin: MATCH, correct repo+ref extracted
Double-quoted SHA pin: MATCH, correct repo+ref extracted
Single-quoted SHA pin: MATCH, correct repo+ref extracted
.yaml extension file: MATCH, correct repo+ref extracted
./local-action: NO MATCH (correct)
docker://: NO MATCH (correct)
* docs: add 3 missing items to fork-reconfiguration checklist in SECURITY.md
The checklist covered Trusted Publishing trust, protected environment,
branch protection, and Dependabot github-actions entry. Three non-forking
controls described elsewhere in SECURITY.md were omitted:
- GitHub secret scanning and push protection (repo settings, not copied
on fork)
- GitGuardian (GitHub App installation scoped to this specific repo,
requires separate install on any fork)
- CodeQL Default Setup vs Advanced Setup state (repo setting that affects
whether the upload-sarif step in codeql.yml does anything)
Added as items 5, 6, 7 matching the existing numbered bullet style.
* fix: update github/codeql-action pins to v4 tip (SHA drift caught by verify check)
verify-action-pins caught that github/codeql-action@v4 tag was re-pointed
upstream:
old: f205ea1c3313d32999d8d6a48b4f6530d4437b38
new: d1ba80a13dd99fba24a470575428917156a28b43
Updated all 8 occurrences across codeql.yml (init x3, autobuild, analyze,
upload-sarif) and defender-for-devops.yml (upload-sarif x2). Tag comment
# v4 unchanged — the tag itself hasn't changed, only what commit it points to.
---------
Co-authored-by: Sameer6305 <sskadam6305@gmail.com>
The CodeQL Analyze Python job failed on the #757 merge commit with
ECONNRESET while streaming the CodeQL bundle download in
codeql-action/init's "Setup CodeQL tools" step. This is unrelated to
the merged code — it's a known, currently-unaddressed gap in
codeql-action: the download error is retryable but the action doesn't
retry it internally (confirmed via codeql-action's issue tracker and
changelog).
Since a `uses:` step can't be wrapped by a shell-level retry action,
Initialize CodeQL now runs up to 3 times, cascading to the next
attempt only if the previous one failed, so the common case (success
on attempt 1) costs nothing extra.