Compare commits

..
137 changed files with 1350 additions and 26905 deletions
-3
View File
@@ -18,9 +18,6 @@
.git/**
.github
.github/**
!.github/requirements/
!.github/requirements/explorer-extra-py313.txt
!.github/requirements/pep517-build.txt
.claude
.claude/**
.codex
@@ -1,56 +0,0 @@
name: 'Setup Semantica'
description: 'Install Python, cache pip, and install the semantica package into a workflow'
author: 'Semantica'
inputs:
python-version:
description: 'Python version to set up'
required: false
default: '3.11'
version:
description: 'Version constraint to append to the pip spec, e.g. "==0.6.7" or ">=0.6,<0.7". Leave empty for the latest release.'
required: false
default: ''
extras:
description: 'Comma-separated extras to install, e.g. "explorer,all"'
required: false
default: ''
cache:
description: 'Pip cache mode passed straight to actions/setup-python ("pip" to enable). Left empty (disabled) by default because this action is meant to run standalone in any caller repo, and actions/setup-python errors out if it cannot find a requirements.txt/pyproject.toml/setup.py/poetry.lock to key the cache on. Opt in only when the caller repo has one of those files.'
required: false
default: ''
outputs:
version:
description: 'The installed semantica version'
value: ${{ steps.verify.outputs.version }}
runs:
using: 'composite'
steps:
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7
with:
python-version: ${{ inputs.python-version }}
cache: ${{ inputs.cache }}
- name: Install semantica
shell: bash
env:
SEMANTICA_EXTRAS: ${{ inputs.extras }}
SEMANTICA_VERSION: ${{ inputs.version }}
run: |
python -m pip install --upgrade pip
if [ -n "$SEMANTICA_EXTRAS" ]; then
spec="semantica[$SEMANTICA_EXTRAS]$SEMANTICA_VERSION"
else
spec="semantica$SEMANTICA_VERSION"
fi
python -m pip install -- "$spec"
- name: Verify install
id: verify
shell: bash
run: |
VERSION=$(python -c "import semantica; print(semantica.__version__)")
echo "Installed semantica $VERSION"
echo "version=$VERSION" >> "$GITHUB_OUTPUT"
-23
View File
@@ -101,29 +101,6 @@ updates:
allow:
- dependency-type: "production"
# Explorer frontend (npm)
- package-ecosystem: "npm"
directory: "/explorer"
schedule:
interval: "weekly"
day: "monday"
time: "03:30" # 3:30 AM UTC (9:00 AM IST)
open-pull-requests-limit: 10
reviewers:
- "KaifAhmad1"
assignees:
- "KaifAhmad1"
commit-message:
prefix: "security"
include: "scope"
labels:
- "dependencies"
- "javascript"
- "security"
allow:
- dependency-type: "production"
- dependency-type: "development"
# Docker dependencies (if you use Docker)
- package-ecosystem: "docker"
directory: "/"
-58
View File
@@ -1,58 +0,0 @@
# CI tool requirements
Hash-pinned `pip install` targets for CI/release/Dockerfile steps that install
something other than the project's own audited `requirements-ci.txt` set.
These exist because OpenSSF Scorecard's Pinned-Dependencies check flags any
`pip install` in a workflow or Dockerfile that isn't hash-verified, and
`requirements-ci.txt` alone doesn't cover build/release/security tooling or
the project's own local-source install.
Each `.txt` was generated from the adjacent `.in` (or, for `explorer-extra-py311.txt`,
`explorer-extra-py313.txt`, and `base-deps.txt`, from `pyproject.toml` directly) with:
```
uv pip compile <input> --python-version 3.11 --python-platform linux \
--constraint requirements-ci.txt --generate-hashes -o <output>.txt
```
(`--constraint requirements-ci.txt` is omitted for `bootstrap.txt`,
`build-tools.txt`, `uv-tool.txt`, `twine.txt`, `pip-audit.txt`, and
`security-scan-tools.txt`, since those install standalone tooling with no
version relationship to the project's own dependency tree.)
Regenerate a file the same way after bumping a pinned version, and re-run it
whenever `requirements-ci.txt` changes if the file used `--constraint` (see
each file's own autogenerated header comment for its exact command).
| File | Used by | Installs |
| --- | --- | --- |
| `bootstrap.txt` | security-scan.yml, benchmark.yml | pip, setuptools (upgrade before anything else) |
| `pep517-build.txt` | ci.yml, benchmark.yml, Dockerfile | exact `[build-system] requires` from `pyproject.toml` (setuptools, wheel) - installed with `--no-build-isolation` before any `pip install -e .` / `pip install .`, since `--no-deps` alone doesn't stop pip's PEP 517 build isolation from fetching those two *unhashed* |
| `explorer-extra-py311.txt` | ci.yml | semantica's base deps + the `explorer` extra, resolved for python 3.11 |
| `explorer-extra-py313.txt` | Dockerfile | the same, resolved for python 3.13 (the image's actual interpreter) |
| `pytest-tool.txt` | ci.yml | pytest, for the pre-all-extras deterministic test |
| `uv-tool.txt` | ci.yml | uv, to verify requirements-ci.txt is current |
| `build-tools.txt` | ci.yml, release.yml | build, wheel |
| `twine.txt` | release.yml | twine |
| `pip-audit.txt` | security-scan.yml | pip-audit |
| `security-scan-tools.txt` | security-scan.yml | bandit, semgrep, jq |
| `base-deps.txt` | benchmark.yml | semantica's base deps (no extras) |
| `benchmark-extra.txt` | benchmark.yml | the benchmark-only libs (neo4j, pdfplumber, etc.) |
`explorer-extra-py31{1,3}.txt` and `base-deps.txt` are large (they mirror
most of `requirements-ci.txt`) because semantica's `dependencies` list in
`pyproject.toml` isn't extras-gated - installing the package at all pulls
the full base set. That's expected, not a mistake.
`explorer-extra-py311.txt` and `explorer-extra-py313.txt` are **not**
interchangeable, and can't be collapsed into one file compiled for either
version: `librosa`'s `audioread` dependency needs `standard-aifc` /
`standard-sunau` only under `python_version >= "3.13"` (Python 3.13 dropped
`aifc`/`sunau` from stdlib). A file resolved for 3.11 simply omits those
packages' hashes, so installing it with `--require-hashes` on a real 3.13
interpreter (the Dockerfile's base image) fails outright rather than
silently under-pinning. Any other file shared across a 3.11 and 3.13
consumer would need the same split if it hits a similar stdlib-removal
edge case - check for `ERROR: In --require-hashes mode, all requirements
must have their versions pinned` on the *other* Python version before
assuming one `--python-version` covers every consumer.
File diff suppressed because it is too large Load Diff
-14
View File
@@ -1,14 +0,0 @@
rdflib
neo4j
faiss-cpu
torch
pyarrow
pdfplumber
python-pptx
openpyxl
lxml
python-docx
beautifulsoup4
chardet
langdetect
en-core-web-sm @ https://github.com/explosion/spacy-models/releases/download/en_core_web_sm-3.8.0/en_core_web_sm-3.8.0-py3-none-any.whl
File diff suppressed because it is too large Load Diff
-2
View File
@@ -1,2 +0,0 @@
pip
setuptools
-10
View File
@@ -1,10 +0,0 @@
# This file was autogenerated by uv via the following command:
# uv pip compile .github/requirements/bootstrap.in --generate-hashes --python-version 3.11 --python-platform linux -o .github/requirements/bootstrap.txt
pip==26.2.1 \
--hash=sha256:71138adf1f4ca900cdb7d289c21b7494329f2332b6d85f0e1c42108c0384ed3e \
--hash=sha256:f6ad667e89a1fe78046c8f13232b247200f5258d7828f3f7883d660878e0813f
# via -r .github/requirements/bootstrap.in
setuptools==84.0.0 \
--hash=sha256:51a52592b3b99e102b609654876bd65f19f999935166d1352678931132b0c670 \
--hash=sha256:f4695c21257f0d9b537ec2692c941d02ee143b7cc1276941349a546573b2ef73
# via -r .github/requirements/bootstrap.in
-2
View File
@@ -1,2 +0,0 @@
build==1.6.0
wheel==0.48.0
-20
View File
@@ -1,20 +0,0 @@
# This file was autogenerated by uv via the following command:
# uv pip compile .github/requirements/build-tools.in --generate-hashes --python-version 3.11 --python-platform linux -o .github/requirements/build-tools.txt
build==1.6.0 \
--hash=sha256:bd2c8afc603e7a2e0ce70e2ea85f0a6d02043bafbd307f5bada0f98669eca5af \
--hash=sha256:f7aaf1ebbb79178a02ba248bb524f2176b256017e17e8e4bd4289c7b38cc2bad
# via -r .github/requirements/build-tools.in
packaging==26.3 \
--hash=sha256:94edc256424af38762eb31306eed28beb9f0efc50a8837492c9d6fd6004aed79 \
--hash=sha256:d7193f7c8e4e93f444fde0262bf90af30e16fa0ad0ad44cb553c87339b23cd1c
# via
# build
# wheel
pyproject-hooks==1.2.0 \
--hash=sha256:1e859bd5c40fae9448642dd871adf459e5e2084186e8d2c2a79a824c970da1f8 \
--hash=sha256:9e5c6bfa8dcc30091c74b0cf803c81fdd29d94f01992a7707bc97babb1141913
# via build
wheel==0.48.0 \
--hash=sha256:3217dcc807155e45db462d7ef2431f5ddda0d7273b700d05a67b271ceb1287ab \
--hash=sha256:94800765601e9171bf5d58d066e640662842bcedcbab982b2c90787a2c987322
# via -r .github/requirements/build-tools.in
-1
View File
@@ -1 +0,0 @@
checkov==3.3.16
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
-2
View File
@@ -1,2 +0,0 @@
setuptools==84.0.0
wheel==0.48.0
-14
View File
@@ -1,14 +0,0 @@
# This file was autogenerated by uv via the following command:
# uv pip compile .github/requirements/pep517-build.in --generate-hashes --python-version 3.11 --python-platform linux -o .github/requirements/pep517-build.txt
packaging==26.3 \
--hash=sha256:94edc256424af38762eb31306eed28beb9f0efc50a8837492c9d6fd6004aed79 \
--hash=sha256:d7193f7c8e4e93f444fde0262bf90af30e16fa0ad0ad44cb553c87339b23cd1c
# via wheel
setuptools==84.0.0 \
--hash=sha256:51a52592b3b99e102b609654876bd65f19f999935166d1352678931132b0c670 \
--hash=sha256:f4695c21257f0d9b537ec2692c941d02ee143b7cc1276941349a546573b2ef73
# via -r .github/requirements/pep517-build.in
wheel==0.48.0 \
--hash=sha256:3217dcc807155e45db462d7ef2431f5ddda0d7273b700d05a67b271ceb1287ab \
--hash=sha256:94800765601e9171bf5d58d066e640662842bcedcbab982b2c90787a2c987322
# via -r .github/requirements/pep517-build.in
-1
View File
@@ -1 +0,0 @@
pip-audit==2.10.1
-423
View File
@@ -1,423 +0,0 @@
# This file was autogenerated by uv via the following command:
# uv pip compile .github/requirements/pip-audit.in --generate-hashes --python-version 3.11 --python-platform linux -o .github/requirements/pip-audit.txt
boolean-py==5.0 \
--hash=sha256:60cbc4bad079753721d32649545505362c754e121570ada4658b852a3a318d95 \
--hash=sha256:ef28a70bd43115208441b53a045d1549e2f0ec6e3d08a9d142cbc41c1938e8d9
# via license-expression
cachecontrol==0.14.4 \
--hash=sha256:b7ac014ff72ee199b5f8af1de29d60239954f223e948196fa3d84adaffc71d2b \
--hash=sha256:e6220afafa4c22a47dd0badb319f84475d79108100d04e26e8542ef7d3ab05a1
# via pip-audit
certifi==2026.7.22 \
--hash=sha256:62f22742b58a1a33014a2b6b706588a8d7e2a88ae7bd1a6ebe8c992928483775 \
--hash=sha256:741e2c3b351ddf169a738da9f2c048608ff7f2c5cc02f1ebc6b118bb090d5d55
# via requests
charset-normalizer==3.5.1 \
--hash=sha256:00668ebb0609751758682eb0b5857e7c35b9f00e84dfdef062e103244ec94d45 \
--hash=sha256:012a22b88a77ca2e59b98ac5889b0deb604147666032f45e6d6e217634d2550d \
--hash=sha256:01e93745f7f219b703b60ba7afead36cfc4242782be5af484673fc500df12da5 \
--hash=sha256:04368edf83514385ffc3e1cfd4546e595f4f1272dd23ba437a93a9cc3741d47b \
--hash=sha256:0722590aabf9dc6a6c0343d523c05458fa2b5047dbe6302fd526bb570600753f \
--hash=sha256:07ffd07412fc5d5e84cd8952acf9ff7e4ed7a708e69d1bada19d8ba91711353f \
--hash=sha256:09a7bba9f739468c8e78c36a75c33768e53cb1959fc638f510454c14683f00d5 \
--hash=sha256:0b2b1b3fa5670c127b246df1d0c059defd41f689a868a3b9d79df9b1cac42d22 \
--hash=sha256:0c6dfb5ca6723eeed15aa8e564a014d69fcb8812f94eef11fe3631e0508199f5 \
--hash=sha256:0d929fc574b4d6fd9e7c0f5c2ede8716a41911923aa7fa5fce38e0818aa4a1ac \
--hash=sha256:13e3afe97712e8887cd516e960c63f0b93122971e5b5e4b2622fe7701771e838 \
--hash=sha256:15f024313246a4ed976c60f440bb8d257815513a681d212ff74fd46f7d715a90 \
--hash=sha256:195ce897c6153c0700078142cf8efe3e6454ca4cf4357499e4078dfd83396626 \
--hash=sha256:19a3dd5aa73cef1c99687c4fc57db016a9c17104ae1185da88ba566a5d3bebe4 \
--hash=sha256:1d1c7a53a6c2103925cdd6d7229f8c567379f211c869793df679f2e9f738c369 \
--hash=sha256:1f5883d77fd409a261abb5dc8ccbe335720d798b1de4abb3b1d47ccbbc76b53b \
--hash=sha256:21b82d8082f6f5e7f456ef0bd16323d08de1266efbfeb476e64b2a91d1471a4e \
--hash=sha256:252d099029bcbea642f2a06c4ed5046bdf8b5a8150b64afa5e027e88b106e5ee \
--hash=sha256:256dd4d85d9e4dc595e2bc983c980e73f62ddeb3165c58b4c3dfe78c5c8548c1 \
--hash=sha256:26422d45fd13551cf564c58932f7d72b4f58b93b0fcf18c35ba6be12b46bb102 \
--hash=sha256:2679de311c7946dde5d3b6f44941844133ff5c7cb86099c0061ab1e8901c20a8 \
--hash=sha256:29880d17a8eb0b5cfdfd8944b468322928059aa35f1f5fa8ff22b149ec0b42f8 \
--hash=sha256:2bced4061f000f7187254a02ad3433ae17eaf991747ceea2f478422590a5bba9 \
--hash=sha256:2e9cf9253119d8e5d111f05d71626786fd3d6193817316eab1ca088cdb8593cf \
--hash=sha256:2f06b7eae9dbe77fe1d644ca244dad508de8d302870a43f3c559b521270938a0 \
--hash=sha256:2f293479cce755c75f1697e87c409b7ae4c555c7dfecb6e988ad13abba943031 \
--hash=sha256:329fc3ccb63ad22d867d84c2adea759a64079a37ba4a343433b02c7a2816871e \
--hash=sha256:343fb4f2821043bd87095f7b08a1a181febc8e36ac64212143bbfd0a0e1bc235 \
--hash=sha256:3588e376b3ea2eea84976f67273d679f229e24c66dce7b82ae45aef04ff6e072 \
--hash=sha256:35aea775dc2bd5f54cd84a1cd2696cc3207c479cb9cf0bd346f0d343e4300ddb \
--hash=sha256:35fe081843b35aad20ffeccec3eeffbe637b15d14f3fb22cc1b59cd8ec17e93c \
--hash=sha256:36047af20e17097c3bb9476c2b7655f2f7aa51322c0ba58c07695bedf755a950 \
--hash=sha256:3617ac3cfd8b9888f145ad89dd6e692285834b0201c6074a5eeaad3fd4d668c2 \
--hash=sha256:366ec70f5547c640d3ce1985722490f23faf4eb5216a7eeba78277490e78dacb \
--hash=sha256:394fea06235c8543390050ed5f529187074b029fb027213f6c46ac11ab5d950e \
--hash=sha256:3d27167433c0d5f18dc850f07d0b3816221984fecdc405d6c157a6f0b8f8e9e6 \
--hash=sha256:3e5e1224c0a6a90e05843e07adfec669edebec17801c67072f51e59561d63c0b \
--hash=sha256:41876ee62a3dddf48ff1121ad8f0798032aa03f2fd35f21f34a4cab14f18d8d2 \
--hash=sha256:433c5a81eade63b47e522303bad236f59dba55ea6951746f5558355eeed8c75d \
--hash=sha256:4582c27e8c889d64811987b5967fbd3ae0c823fe1fd933b543d55ac20bb475fa \
--hash=sha256:485a0d363cafefcd2538a73c7c838daa2035f09b2c9f9b5e3133f80c6aeb84c2 \
--hash=sha256:494b70049a4d69aec6e8137c13af4cf8db8c9f9820a1392ac293b0dd2987a818 \
--hash=sha256:496846868fea80e479324862fa877f02411f2fd0f83b79ccee2607aa68b2a032 \
--hash=sha256:4abdc5f9ad448c1ecbfae2974b820535d6bc6e7eef63babbab3d81cf46968c71 \
--hash=sha256:4b599739b93b2cbeded49645ae3c8d1405c29ddfbceac1545c87a3f9580a9e96 \
--hash=sha256:4bea7f8ebe90bbd7f0e4a2de42ca6924ba23e3e76418c408ff82f1d46fabd687 \
--hash=sha256:4c4fb141a727957c93edfe5c32a26ceb6b5f6461d67146e2d39f51e16170bea8 \
--hash=sha256:4c9548dc78002099910abaebc0a72ac58b7d30931869e0351c09b507dff4ece3 \
--hash=sha256:4d26f14f041e83dd8edfd61f4cd4fa7285d31798b5bf1f28e70c367ba6c41d61 \
--hash=sha256:4f298bdadb8f0b9e5672877f647d1be9373ef5320c9e2f049795e26cad28b6a9 \
--hash=sha256:52ec005752a56ae79547a05c0139ca2501a0c866390b6115008456b9f0e7cde1 \
--hash=sha256:55261ac0d2941c42f196dd576f543d87a8ee03cd6f5e30dfb4d807b2e3b9121a \
--hash=sha256:56490c595a28b1bb27dfc583e816152a9767721ef58b2c03b13f954d2f707420 \
--hash=sha256:58d3e12c88e0950bca850ae1f7c256055c097639c2edb9eb123af9807d8b15e4 \
--hash=sha256:58d4aa13a59c969dbfdf9e6a9560e242cbfd9e8a8f50c2747714df1a423adf65 \
--hash=sha256:59171c6e45bf07d0d5cab3b0bf81d945035530f6873398b3b531c31184d46663 \
--hash=sha256:5b6d1386bf0096d26d3a863dc0a487a5b4eb9aa93cf5ba69683d29dde6b9d60f \
--hash=sha256:5c0ea61a470e070686aa30892fed79e297d2c8d0ab46b8bcdf027d38c51da591 \
--hash=sha256:5c84bec0ab5ae0c64bfe73a7d2adcb5ce73b467523fc27fd6a28ab2aa6cbe35a \
--hash=sha256:5ca0555312ae2fe82715cada7fac375530c2f3349e1eaa1bcb33d0283ac79a18 \
--hash=sha256:5d8531a6569d025f68e2321e7638fb7978f23db58e5f69f56913837aae03816e \
--hash=sha256:5e2d0e146dcb57034f8b97dc58d2d512cb90aba253960ce449f695fec6a82c6f \
--hash=sha256:5fc45d653ea8c9a20479167e11d4a0f8cb2fa3470737ab6f9c827532313187b7 \
--hash=sha256:6117b84ea48435e5356dc737f5121485c30920ba43375fa7b434fd753df0eac3 \
--hash=sha256:6199d5606e2bbf2b096cf64d03f8b6790c91081d5ac866b8e7bb6422738cc60c \
--hash=sha256:62b55f6722735a6c472f88361cde6640608773d9443cebdbb51abf436a1fcdd3 \
--hash=sha256:687c9ca3035544b113bea2055e180af96fb63c0c476e22a9180f51925186e7b7 \
--hash=sha256:6b7430cf5728e68f6c462254009a6ef4086e1bea43cf2f57aa9c55fb4f50ff96 \
--hash=sha256:6ba32c4d2abf1d2fe7cf27d280f4cca5664233b0f885549c7761719eb977f486 \
--hash=sha256:6c9cdde8becb25a7fde49924511aa2644d6f8081cc8df8e9452724303348d8e3 \
--hash=sha256:6df0ec430f9a831772c23ca5a224cba36517a58a84bb32c32bb59a9fa67c47f6 \
--hash=sha256:6e2912d4babbc65196ac13c2f53468dc57fb8b9c25ef913e8c59ddf7c6dc0e1b \
--hash=sha256:6e5e4d73d588ca5ed09df1b7dcd1b203d1df3c542e3f50d126c947d432b10731 \
--hash=sha256:70055ff39b97c99e7ae40ea3e393fb62aa2e44dbd9b29f8d14f42fb0025c3959 \
--hash=sha256:706bfd38730a5ac7a365793269a00f4e988178cec121391f4248d84ad8c972e9 \
--hash=sha256:7235dc28fc6dd9d832ac7c7bce95367dedb85929f17368a0c2bee1e080b9acbf \
--hash=sha256:774d157f112367ff4abd29019f38f023c24e00e56edc7829c20e358a5a913ad8 \
--hash=sha256:77efcff2b23071c349402ac1066667a3d011f62398d81408c9b88ad991747c9e \
--hash=sha256:789b8982559ae28dad2356519f841655756cdcd96616410590ae0b17454ee64f \
--hash=sha256:7ac76cf9afd34929d76eb7fcb63be476a4853d8a96f0dcf2d0db68a0cbdf9885 \
--hash=sha256:7c0c10730342b0c9b35dd1d619beb8214e520bd96a1f870f452680b238aab3e0 \
--hash=sha256:823f82903d189af463d7df250ef1f7f696f3cee08cc8d91deb565e8d425f6506 \
--hash=sha256:838648accb3a7fd9803fd45c87bce8509648eb0c11bc34e216141300977244f2 \
--hash=sha256:854066be00447fa8de2ccbbe893e2ffc4b123ef16d897af794c1e18bd4a714b0 \
--hash=sha256:85d5855daafc240cc045c026d7a15fd198a09b0fc8ff6f5ecbb5297b509cb11e \
--hash=sha256:85de3134b5379856e323ba37c19c9256d39425f7b76a63af52b09fb4664c2e8f \
--hash=sha256:87e4f41d375c0b9be2fb5251aee4b8a689169e134535aed81bf085c3b647451e \
--hash=sha256:88ca277405c2d3b71c4e1c2ee0e7966e807bcba86a69d11e19ba199d18ae4491 \
--hash=sha256:88e85ab89cb822c1e635f51d6d32e488f94e002e70e2f492bdb8b945543f345a \
--hash=sha256:8ac8c94b6539074e0f40899301273ac8402b9b3e01c7b7ba269ff30340aaaf20 \
--hash=sha256:8fe532b3c966d1fb794e0698e4589d0444017ae77fc0b31edea13c0e35bcc449 \
--hash=sha256:9085f87b0e38a2b92b8923059b4e8789fe40d9279712d15dcc670048d77079af \
--hash=sha256:90b7481fb62fbe172c558bc6fd1c4c98d82004a54a7551f20e11ac9bf0b8708c \
--hash=sha256:92caef967d287a407085d61176fce4012b1dd62daed4eb6d5ceb26d3d2538712 \
--hash=sha256:9362dd90aa7dab48c0054a21187791ccf05473f7dba5d92b8033ae62164675e7 \
--hash=sha256:94d78ecec2605a8d0398b0f365d5f12a63248438516f5dac536a5eff7337df4a \
--hash=sha256:94fbf1c0c6cc0d3d5e50f9a9313a8cdca90dd696d34b381cd1704f8c9e939f20 \
--hash=sha256:950f23cb393f85543777b0433f082cddd25b51ab398eac7971146495679efe5f \
--hash=sha256:96eefc178f8636b9c760c5829345307fd81cfae9ab1e80997dbddeb0f54ee9a3 \
--hash=sha256:96fef3e886d6a9874b14f27fc193fbdc69d5d8035783d86aa4e1cea594e695f9 \
--hash=sha256:977cdbd483a9cff38179bea4fd754289a6f2195c7abd414aba85410b3e66cc5e \
--hash=sha256:978eab16f55b4ab2c2a745be9a0a840bf8f09a7f227d9c76eb30214d078865a5 \
--hash=sha256:994e883d17c559cdfd38c84003c8b27d25424a1077272a17e7cd27bfe0bf57b2 \
--hash=sha256:9ac4444d8d4fd4c4bd08bf451ed3167aa9e7ec6cdb41b648794f1d1103652e36 \
--hash=sha256:9b5db6052055d34d41230fb78d7c439c23dc536a9896f6cb039e8dd92cfc1263 \
--hash=sha256:9d9a0dc7cbe9bec24c3f767c9122c41fe5a1bc43f47cd099d00d393e09769de4 \
--hash=sha256:9dbdd9205662134957cf0c324f639bdc5031c0ca056e2369e238db75187c0f11 \
--hash=sha256:9eea3ab2597a5e65fe65296e2d6a84570845a6b55532d90333d740d48bbc850a \
--hash=sha256:a2028475ba855475b8b4d3cfeb4994269c967aea8b9892dfba907f4263a863a3 \
--hash=sha256:a3a370082ce34d0612f421e15fe011c53bb1feff21a26d06ad4fb244dab5a375 \
--hash=sha256:a545775cfe815855ea32d7c27731d79da358ef2055b4a25830231b1622dd18aa \
--hash=sha256:a5cbd90ecf0fc62e64726917ad083b73001f0563657a87ec3c0b504e277dc90d \
--hash=sha256:a6d095662e73e74f0a49988e0593373e243e3a52e27bfeea0a859e88acf4a0f5 \
--hash=sha256:a6dac12ff6b846103483683f60c5f8fee205121adc58ffd87e90a90a3af69e99 \
--hash=sha256:a951ad59cad9145664a730d3036b40b844e74d2d3683da40111463cd3a83845d \
--hash=sha256:aa1099b956fb795e686d073568f6dc002a0bb89765ea6d5b055dd7d9bf1b116c \
--hash=sha256:aa2bb0b37202dca27175591f761108b5d34096ade1191ffe4808bdf6b1571488 \
--hash=sha256:aae2ee51122d3ae968a3837d97dc24a0aeebb0dea23694422cd172bd30017cd6 \
--hash=sha256:ab743e9bc90c1f73552ec33e10e3331315acd2c397b36065b591b0181de533cc \
--hash=sha256:ac00177c4831ffa650f8609e4bdddd5fe09c03b1c0c47acece7e6ea20421598b \
--hash=sha256:ac13b004224fb341e1e25a1ed5e19d32f57cdb2a403e01f003b46f051a550f6f \
--hash=sha256:acaf604462bf330b0d07e7a07c1d6e4adac79e5fb13e9c5140590542cafacc00 \
--hash=sha256:ae31a1a1db2ee6cc2942fccaf695c934bc7f3db9f2133a3fef1f367cf1a4ab10 \
--hash=sha256:ae4a097991662cd4fff0ddc74e0fe7874f82e00042fa0ea00855645ed0c79598 \
--hash=sha256:aea996a6aba25260827c9ea511d1addfde2da9eb686ac961838509086188b7e6 \
--hash=sha256:b39b69b347e5e47a3b5b8cfc005c68c1ba347474e3960236c4944a8ecd174962 \
--hash=sha256:b54e7e13267d49ffbfe68e25b3cbd774dab38fa37238f71265e91b36146eb21c \
--hash=sha256:b9af956078716df40d985fb0dfeb2c2120c5ca92ba4ff4b388acfd01cdc14d08 \
--hash=sha256:ba2f37ee79e6338845261a3c5b1784e5d1acdff2c0785b284f1b633033d136ab \
--hash=sha256:ba501e667c17d8411f98e67a022d9604ef179aff0e459b7e292c796837c13573 \
--hash=sha256:baf3775a2635e5a11fbd5e4e64ee69c7e86875d224a5c72aca4c141064589a90 \
--hash=sha256:bb57753e36e4855b8ca375069482250a6246372331a3e4f3407eaebb007443f5 \
--hash=sha256:bd6c173f04743d483881bffa1478d5a4624475b8cd1d2194956a75548e191c18 \
--hash=sha256:be47f99644b208bff7766314013f9acf57b056b04191d570d68ad14022cf5b1d \
--hash=sha256:c010f5581d9c612804cc59fcf7b524b707fbcb72828551237ab545bb5c7034af \
--hash=sha256:c1dcc36dcb96abc02236e182d17e0f71430152a6c2c7447421da2d2dc144edea \
--hash=sha256:c428c6c31eb5f4277d7f8eccaf767fbd548ddd5ce3c8b4f4cbbfab3d96b5904c \
--hash=sha256:c658c50ac0c98cd755a2dd50b7977d3bca7df401dcc47fbdfa87db53ef7d4e8b \
--hash=sha256:c71fb0d56c920c269cd3e2e3fe7c610e3f1fdb21a6ce60efa6430ff63676cea6 \
--hash=sha256:c7b742bf31c88566b4bb6335a7f393bb322e580b6bb98df7bd0c25e6e3519ce8 \
--hash=sha256:cc0329df4caaceb950d2f580b5ac716a377f7059624a0bafaeaf8a218c6ed774 \
--hash=sha256:cc5d36d96478aa9c60654bd932525bf32964c62a7281eafdf16d85003a8d6004 \
--hash=sha256:ce854f5f478050ade5a238731c4ca985a7d3b3cb53ff600a9b5c3b689b5f0a7a \
--hash=sha256:ced3fdd71aaa83ce593746c2edb42b7a59cb4c19c8b5c407781c72e493aae55a \
--hash=sha256:cee5dd7c6fb5dd52a0fe2a740f9bc6e3593f5f8b1788bde49de02086f30182b2 \
--hash=sha256:cfa1c0cc3a8f9f53f1243a5a99ac36fd003880199383b37672e86ddda9cb07e2 \
--hash=sha256:d1ee1e296209fdce05b81b663250eefa02213a2da7b41bf26f7829b8ba3545aa \
--hash=sha256:d59b75732e9b6f27388e10c14b0259cc5f2e48c78627d185e6a177b58ad3cffe \
--hash=sha256:d63600d620ad0064c3a748b950ac5ea38a80190e5498532efefa4b7b3f1da1f3 \
--hash=sha256:dd732602a7009217f658d5863d12d79d373a4de0eebc111094bcdd3bb8e0a6cc \
--hash=sha256:e06efa066f7dbadbc84ebc126a97c452a6451dfcf589d89d788484949e1cf795 \
--hash=sha256:e199fb99720074809a7720f1c0b4d919eea8b87e88713e0f8f602f7bef543d9d \
--hash=sha256:e4b018dc5a0eee4676e38fe84a47a427816c590b93b55d9025274ec4d6ffc2dc \
--hash=sha256:e6621fb2a4988d6e53eedc455e5903e2679f3967b8acb3d639f1b63c14a2e893 \
--hash=sha256:e71c909f353863b2b89c83de2ebed71ea6d0df8a6ef65a128193c5e650766bef \
--hash=sha256:e90251c0c7bdd54a100a0dce3c07b7e637278c93af29dbf78ebb89a58c4bac7d \
--hash=sha256:e9fbdce1e47394b09bc9f26ab117dfc8d6491977a11d86f592bb42c779db2fda \
--hash=sha256:eb12fb2ba69ffa05f8695f61c69e591dc4b4a12ac3757ac8af8adb259bf56d17 \
--hash=sha256:eda059b6bc8bc0812d626fd91a7ce01bf583df0a61296eff390fd94141a34e30 \
--hash=sha256:f03ac127268b43ef4fe9e6ab6794a6794b49485a0cc0c1db79876d2f33f75bc7 \
--hash=sha256:f298e218441525d3794428b4c8b8fb8662c6d3ea79925d4807ee6b9a96a3bca5 \
--hash=sha256:f5542f9b941279d82d41eb0aa9f98eba36fe4df5c7086c651df7944935b37182 \
--hash=sha256:f6f7deae3feb4edfa2efaf7c574fe88cbf055038a6abdb40188e4fff66d5699f \
--hash=sha256:f9b1e28d0e8dbfa858abdba91d6b547beaf2df1a59bec6da6faae7b96a4991a9 \
--hash=sha256:f9f8405c2c758532c74fed975dbee57be1f31a6e865c031870c79a6ed3212ada \
--hash=sha256:fa48b1b63d639f9483e0633e092f5851e2348c352f1f9bb6c8182f87884ef876 \
--hash=sha256:fb78f6e7fcd8ad785d28cd577168bc1aaee827b25bb8755638f694794ea98f0a \
--hash=sha256:fbc597639158fd7c14d55e808718848319540f51b0e6746e3eefa59723a4a348 \
--hash=sha256:fce8cbd4997efeb450bd298b54f755dcdff18d496f7a5ddbb4867c6d7c88fdc3 \
--hash=sha256:fd0350afdc3aabd5576f60ea109228bd5538139713c7b094c5cd27c73a98bc6f \
--hash=sha256:fd0a274c0e5f9a21565cd9d3dd749b61f96b7aa1e20a93aa1ba4029518f2e5c0 \
--hash=sha256:fdb8a068947befafba9952162645dc2fecaeb400e64584829ed5e9b2fbe21a7f
# via requests
cyclonedx-python-lib==11.12.0 \
--hash=sha256:0e807521a921a5c3cb8ce1153f8a61d29eedfe76a46aac2796b7c6b573391a54 \
--hash=sha256:16767c4039de90c04e9f03348f8f0ed4b8ff842eaa7eefcad3a95685f970dacf
# via pip-audit
defusedxml==0.7.1 \
--hash=sha256:1bb3032db185915b62d7c6209c5a8792be6a32ab2fedacc84e01b52c51aa3e69 \
--hash=sha256:a352e7e428770286cc899e2542b6cdaedb2b4953ff269a210103ec58f6198a61
# via py-serializable
filelock==3.32.4 \
--hash=sha256:22e58ca3b1ae3b98993b762d7338367ae64fe50252bf78d59da3bfebcdf1cedd \
--hash=sha256:2bde2e4cf732e0153406d8a7bc80620ecf5e621fe0d25e41143c4e3b4733ff30
# via cachecontrol
idna==3.19 \
--hash=sha256:5e0811a4383b21dc5838069f801c4fb62113b7447663d2530d2bd6e77b49bf15 \
--hash=sha256:815e7be7a7806d54abb586dc943addc79e8b2ee16915059658cbeff4b1b43bf4
# via requests
license-expression==30.4.4 \
--hash=sha256:421788fdcadb41f049d2dc934ce666626265aeccefddd25e162a26f23bcbf8a4 \
--hash=sha256:73448f0aacd8d0808895bdc4b2c8e01a8d67646e4188f887375398c761f340fd
# via cyclonedx-python-lib
markdown-it-py==4.2.0 \
--hash=sha256:04a21681d6fbb623de53f6f364d352309d4094dd4194040a10fd51833e418d49 \
--hash=sha256:9f7ebbcd14fe59494226453aed97c1070d83f8d24b6fc3a3bcf9a38092641c4a
# via rich
mdurl==0.1.2 \
--hash=sha256:84008a41e51615a49fc9966191ff91509e3c40b939176e643fd50a5c2196b8f8 \
--hash=sha256:bb413d29f5eea38f31dd4754dd7377d4465116fb207585f97bf925588687c1ba
# via markdown-it-py
msgpack==1.2.2 \
--hash=sha256:06d95f61de7afe4f4ff908a6feebfcb070d0582ac87c9cf3cedf8551cf634516 \
--hash=sha256:0708afbf6a9587f0bfe479a9825c141d14d91e2f6a5c8103cf28bc96f4edb5d9 \
--hash=sha256:0883a1578168929fd1640fbbc4614773f1a130e419a8a817dc2918d9af1b651c \
--hash=sha256:0a652ceeededf71d3fa40c303a02a149d42338d310162367b91c539d4bd6e0a3 \
--hash=sha256:0dd9173c5ebaf5ecc5ca86e7ae1db92934e1d57b856f3dd90698941431f4fd77 \
--hash=sha256:0e3315de5a4b2920ccef48d96b4448025e064a10d0f5a250f6584477d839c8d4 \
--hash=sha256:0e91332144f69bc3018c91232fac26da580ef748fb8eaddd7914d4458001cc4f \
--hash=sha256:0fbc1bed8a535389b41882cfae66376e248cd1680eaa94fd83193c73e1d24986 \
--hash=sha256:11e8c421e117d1c36728b423d0402555cccbf0c6f53e288f0e75b6b12100d70f \
--hash=sha256:1510f24612d4b983dff6935d9273e02c320cfd525727fbcb58836a75f589fdbc \
--hash=sha256:1814f92306ae7862908e9ece7cfd90e0dc87ded3e89b6ae7ffdd1175d6376fdc \
--hash=sha256:1e8cdd1f3e7cc52c751092a9bf740e81e6919ab109cd376ae2d965dad0bbae34 \
--hash=sha256:1f3af0baafd184436501004828bb3df64eeb2fc49dfe9d89abcf604956094563 \
--hash=sha256:1f6b6f8deb07d49090e1808c6ef9cb7d23ca17bef3aa6ed3e5e03df16606e60c \
--hash=sha256:226a62ffe99fe54c5c61d910ec64c3449b7766c3280bd286bf6c94838dde239a \
--hash=sha256:29cc2d5291711a52956a79a51f41c732329df39ad727c886bd8f0b5b9237a808 \
--hash=sha256:336525cc2688e43ea77dfb1a4ce012c8cde561835913801dbfcfdcf4111d8abb \
--hash=sha256:34e83e345194a2a51d8bd447dea9de2104f91e75b247f4735f14f04529f0746b \
--hash=sha256:352ed831042549cca8be23780e1fe7c9177e65ff02bf183509c4b4d33f671782 \
--hash=sha256:3e915d390d7068b257ca8b62f3fc59fad135c8631d1017ab03b0b924b07c5367 \
--hash=sha256:419a45c67a5c04213172a14b1864657e014665b77d7081b107a51707923dd39e \
--hash=sha256:42fd9260416885b4815caca5bdd14dfd5dda6cdade732d6c09104ef8f6228761 \
--hash=sha256:46ec851571d8f1b6e29794ebb9dd36f785008da6d14f57c702e60781d6caf648 \
--hash=sha256:4710d881d8fb047deed2485707409116722af2b992d3fefd73c7667c4e350839 \
--hash=sha256:4955accbd87f27beebef5f3ecc27503aa74cb016fb4f640868e749fd93194a35 \
--hash=sha256:4a4348705be86e029d04e741cf9ed0dfe03e942d7d3b92e838fa80d3aa2c3ebc \
--hash=sha256:4b554d8164ebb526892194f71dcd96ef1fefe0c250087498785d3ffc04a80be3 \
--hash=sha256:4d9a562aec0a92fe536da2e533d313b3d2a6b929157b1dec7ff623446dc0a8ab \
--hash=sha256:51dd39d23cfdea0400ed3ff2d29d1e83bd951d3aea79dc89be5b701a09edfe23 \
--hash=sha256:53679573c75cce5f82359e0bd4e6a97809a6b9a9b7a48fd1ba592f4a82cddc84 \
--hash=sha256:55faa6f8395e23b848c535ad5dcb96b3462f37f5e7f4ac500d500434f7345da7 \
--hash=sha256:58ce37a4a54577115922385d37201d9a44d66d0167dfbbf4770a2e9bf8ea7ba3 \
--hash=sha256:59d5b93efa45fd09f620d0c9ba81cde339a2c9937af3eea42ee9653094ce6640 \
--hash=sha256:6195257a107bf25872ef84aab7295078271eea3ac6413f0506b631f6c9586ed5 \
--hash=sha256:652d1bf13d01bac8fd569def0fe76745e55bcda01e30aa6332d5947ea3788839 \
--hash=sha256:682804bf31e43d46e51a9a33bd575b51e839d715ce6bd5612c055f7b28ad637b \
--hash=sha256:68df2947921d449f6dcfeafd86cb2cdde13327a8b447534bbe4ee5aaf32a5695 \
--hash=sha256:6f53285f20d592ed309ee19e509cc4c77a3bda1db02ad67e8a0949bb227a5a6d \
--hash=sha256:73b0e05c32c3cfc3cd84994908e57430c0ebc6813abf905d3f18ff115d54df3f \
--hash=sha256:77c2e018417dc1d66f235e383877ee885b60ade9d29e494dd581e08af2cb1923 \
--hash=sha256:7826f16edc763e768404f55605ef85dfcf5857e729c1ed29e0d7c180be4fe6d8 \
--hash=sha256:7afa5431f6f3487c584187ca6c8e2a34e9b106529893b3e720eabb068f6ac970 \
--hash=sha256:7d095df2627e5dd59ac7b0c5ad627a671c76e6020171e03cbe4621a61f0562c3 \
--hash=sha256:7fe374ba76eb0ecca13a1703daa8fa85825a6ddddbb52d4c1a732fa524194683 \
--hash=sha256:82b1bdf293267afaadcc608b125e7fc6576bb0785a60c4fa7d07c7ab76ed76ec \
--hash=sha256:86f173a584f72f6164801f31866d22a581f60c991572cf922aed9ab8eb422b77 \
--hash=sha256:8b1415d02e9bf722672af8a90f90813265a0cd0b14163187261e54a5592bc949 \
--hash=sha256:8b2a281b556f120a43e591ea39915741b7ad54d4727b9c4350a0a11692252533 \
--hash=sha256:8c6321a414f8b4a8dc43976b2fa8349156434ca9adedd9a187b796f7e1d3d3fc \
--hash=sha256:8dc4487097571f7311188c3eca2a3e86cd1f1db4c37c7a017bcc3fd38486cbfe \
--hash=sha256:90986cc9aab9d7d1d8f38bcbf65d3f7ac83bdd90c35765db7d691b4829698cba \
--hash=sha256:9352e6cdb510a7b1a5d3ccaccec730e82e50cf3484a3af7bdaab19e23b9589ff \
--hash=sha256:935b1cfad9b908b0fa845010f4271df4c2f04e1cd26e3f18acd61a45f93c9e36 \
--hash=sha256:9b659d77f8726fa5e7038967dda6b68d53cf34472c094cfa5b845454713b90d5 \
--hash=sha256:9bd3d1557c3fe1a095068210708a03e3e4795973392af6f4047060e70abd9a6c \
--hash=sha256:9bf452ff4d4981f25a18e9476e002bcc9263e7928024aa4d7148e25f7be3f929 \
--hash=sha256:9d7fb25b4442fae0cb2590272d06ab4f6caa526ee36a994edb81e946b874813e \
--hash=sha256:9db1ba1c1e6a84245a9dd866265b56b8a1e9461549cc72ed296d8cbfbd32961b \
--hash=sha256:9eb0b0e602064527a045ea28c4f174ed69383587e29cebe28947e3b84106eb2a \
--hash=sha256:9fd7f32e2f0fb334e7ecc5adb5cf0458785bd3a9d9d86f950e1715f101cebce5 \
--hash=sha256:a378e12ccc06d76efde115caf4073b7e5ff3cc18291d1341f9e65fb882e3f754 \
--hash=sha256:a4161eee7799863aee237c35c90427861f7b994416dd81ae829f560b0a81bdcd \
--hash=sha256:a9b4cf3685a135666d27d0d7a73fece74e2fad01d9b508fded89e843512f0e90 \
--hash=sha256:aa1120c653b76d8eafa50423b5eba06b5c9737f8692c74fa3afe03e84b8978ea \
--hash=sha256:b07c03f0da7e5279170df7745ddc732d526c8a198208936ec1a95c11ed2b2d5f \
--hash=sha256:b13b59e66f107cca1ba708dd5307179870ca1b15b19fcee7ccf722e5308d9212 \
--hash=sha256:b542ffc0a5c531eedc40419f291f1bd659aa8d4223408a5b51c88a2796083fd3 \
--hash=sha256:b5c696ae7cd7166b3657261adb855b461ff31f07823fdbae9de8bf80adfccc21 \
--hash=sha256:b68614fba0570349833b7dd999ff0aed4e5cc8d9eb6e3a7d4527be33c65e33d3 \
--hash=sha256:b8dd6c71d20c28d2d0eb0c51e7cccf3584afde3b1364f6629596186c9025bd54 \
--hash=sha256:b9b0c1f2aa7b0026b4bd50718100e8b04175e4f36e160aa852502377b5e572e7 \
--hash=sha256:c522420d78db2431887d45b518e304d86e27b9ad0b30f24e3806a6ad5d8bdbfc \
--hash=sha256:ccfd880988f8438d1c91c77d7edc58e70f4d2012e999167bc154c64c6f06ea6b \
--hash=sha256:cdb6cc6e1127d15879c47a8b3270716243da82d3e7feab1f5946872c75b3d60f \
--hash=sha256:cf66fb38703e61a486b01b56d43bb1f50698fbe99b6bd90feba10f24fab60b3b \
--hash=sha256:d13d07efbf655f9ae7a2352b630c52727b359005b21ba08a507585c9ac8c0896 \
--hash=sha256:d242f3c4ccf55b056e6cf901720dccde58f1df117898f2bbf3bcd6e38ec7c248 \
--hash=sha256:d24b38a825bcca41bb956de50eb98451ef291304a8607fad99e619043d3e79b9 \
--hash=sha256:d3c247d457ae9079974c7ce3c665396754a6d2baff7eaa51332212a8a5a3f13b \
--hash=sha256:d886baa46b2532135e7320067e6a44edb09ba5883a6096b0f9c044533984b8a8 \
--hash=sha256:e05a94a0442de86818a30281c6cc2cb9cc7aa148386fd3541c4d4774b73cb3a9 \
--hash=sha256:e1b99ad34613d5f8477fa5cf99bc4eaeaf27965588007c102370cd9a78fe9de5 \
--hash=sha256:e2eb7ea0ac3911a7aac9d8aaa36d40f216d99455b3274cd3fac38181bcd910cf \
--hash=sha256:e497ee34e8a3342bbde51b27c22d8db05a651df3361dd3daef5b3ab0d66f3e04 \
--hash=sha256:f11e09f10210a91c169e39c7a5a1f9090eaa73ad75555fafad5023c3053c47ba \
--hash=sha256:f466049b8e1ec0854287bbe9a074316826fe0e08dcf707245f98b1ae49e92650 \
--hash=sha256:f80361592c13d7226b4379c8941529b63fe1a9d0e05d2de8f3306b70e522b53f \
--hash=sha256:ffdd2f4950daf7815490f23087963e3420175b9609520b7ff5df64d351159c22
# via cachecontrol
packageurl-python==0.17.6 \
--hash=sha256:1252ce3a102372ca6f86eb968e16f9014c4ba511c5c37d95a7f023e2ca6e5c25 \
--hash=sha256:31a85c2717bc41dd818f3c62908685ff9eebcb68588213745b14a6ee9e7df7c9
# via cyclonedx-python-lib
packaging==26.3 \
--hash=sha256:94edc256424af38762eb31306eed28beb9f0efc50a8837492c9d6fd6004aed79 \
--hash=sha256:d7193f7c8e4e93f444fde0262bf90af30e16fa0ad0ad44cb553c87339b23cd1c
# via
# pip-audit
# pip-requirements-parser
pip==26.2.1 \
--hash=sha256:71138adf1f4ca900cdb7d289c21b7494329f2332b6d85f0e1c42108c0384ed3e \
--hash=sha256:f6ad667e89a1fe78046c8f13232b247200f5258d7828f3f7883d660878e0813f
# via pip-api
pip-api==0.0.34 \
--hash=sha256:8b2d7d7c37f2447373aa2cf8b1f60a2f2b27a84e1e9e0294a3f6ef10eb3ba6bb \
--hash=sha256:9b75e958f14c5a2614bae415f2adf7eeb54d50a2cfbe7e24fd4826471bac3625
# via pip-audit
pip-audit==2.10.1 \
--hash=sha256:1eb4565d19ebe5d48996f4b770b4d2b32887e12cb12cfa637f1a064011b55ffc \
--hash=sha256:99ef3f600a317c1945f1e89e227ef26e1c2d618429b8bd3fa6f4f7c440c4611a
# via -r .github/requirements/pip-audit.in
pip-requirements-parser==32.0.1 \
--hash=sha256:4659bc2a667783e7a15d190f6fccf8b2486685b6dba4c19c3876314769c57526 \
--hash=sha256:b4fa3a7a0be38243123cf9d1f3518da10c51bdb165a2b2985566247f9155a7d3
# via pip-audit
platformdirs==4.11.5 \
--hash=sha256:89f8d42695853b89c7170bd49bc3dc593f98a71e695ede88e06a3b247bc4563b \
--hash=sha256:e8b31f4f8bcbbedef91a6b57a706255e4f148d2a4e01648382a0a47342539173
# via pip-audit
py-serializable==2.1.0 \
--hash=sha256:9d5db56154a867a9b897c0163b33a793c804c80cee984116d02d49e4578fc103 \
--hash=sha256:b56d5d686b5a03ba4f4db5e769dc32336e142fc3bd4d68a8c25579ebb0a67304
# via cyclonedx-python-lib
pygments==2.21.0 \
--hash=sha256:2363c69b61c4a97c838da3b130dcd6468f4848992b21a82f2a63ec34377137d9 \
--hash=sha256:610ca751c9bc2492b38eb9a38a7fbc93edbbb2d7182edaf34e66ae493dee5c8c
# via rich
pyparsing==3.3.2 \
--hash=sha256:850ba148bd908d7e2411587e247a1e4f0327839c40e2e5e6d05a007ecc69911d \
--hash=sha256:c777f4d763f140633dcb6d8a3eda953bf7a214dc4eff598413c070bcdc117cbc
# via pip-requirements-parser
requests==2.34.2 \
--hash=sha256:2a0d60c172f83ac6ab31e4554906c0f3b3588d37b5cb939b1c061f4907e278e0 \
--hash=sha256:f288924cae4e29463698d6d60bc6a4da69c89185ad1e0bcc4104f584e960b9ed
# via
# cachecontrol
# pip-audit
rich==15.0.0 \
--hash=sha256:33bd4ef74232fb73fe9279a257718407f169c09b78a87ad3d296f548e27de0bb \
--hash=sha256:edd07a4824c6b40189fb7ac9bc4c52536e9780fbbfbddf6f1e2502c31b068c36
# via pip-audit
sortedcontainers==2.4.0 \
--hash=sha256:25caa5a06cc30b6b83d11423433f65d1f9d76c4c6a0c90e3379eaa43b9bfdb88 \
--hash=sha256:a163dcaede0f1c021485e957a39245190e74249897e2ae4b2aa38595db237ee0
# via cyclonedx-python-lib
tomli==2.4.1 \
--hash=sha256:01f520d4f53ef97964a240a035ec2a869fe1a37dde002b57ebc4417a27ccd853 \
--hash=sha256:0d85819802132122da43cb86656f8d1f8c6587d54ae7dcaf30e90533028b49fe \
--hash=sha256:136443dbd7e1dee43c68ac2694fde36b2849865fa258d39bf822c10e8068eac5 \
--hash=sha256:1d8591993e228b0c930c4bb0db464bdad97b3289fb981255d6c9a41aedc84b2d \
--hash=sha256:2190f2e9dd7508d2a90ded5ed369255980a1bcdd58e52f7fe24b8162bf9fedbd \
--hash=sha256:2c1c351919aca02858f740c6d33adea0c5deea37f9ecca1cc1ef9e884a619d26 \
--hash=sha256:36d2bd2ad5fb9eaddba5226aa02c8ec3fa4f192631e347b3ed28186d43be6b54 \
--hash=sha256:3d48a93ee1c9b79c04bb38772ee1b64dcf18ff43085896ea460ca8dec96f35f6 \
--hash=sha256:47149d5bd38761ac8be13a84864bf0b7b70bc051806bc3669ab1cbc56216b23c \
--hash=sha256:4ab97e64ccda8756376892c53a72bd1f964e519c77236368527f758fbc36a53a \
--hash=sha256:4b605484e43cdc43f0954ddae319fb75f04cc10dd80d830540060ee7cd0243cd \
--hash=sha256:504aa796fe0569bb43171066009ead363de03675276d2d121ac1a4572397870f \
--hash=sha256:51529d40e3ca50046d7606fa99ce3956a617f9b36380da3b7f0dd3dd28e68cb5 \
--hash=sha256:52c8ef851d9a240f11a88c003eacb03c31fc1c9c4ec64a99a0f922b93874fda9 \
--hash=sha256:559db847dc486944896521f68d8190be1c9e719fced785720d2216fe7022b662 \
--hash=sha256:5a881ab208c0baf688221f8cecc5401bd291d67e38a1ac884d6736cbcd8247e9 \
--hash=sha256:5cb41aa38891e073ee49d55fbc7839cfdb2bc0e600add13874d048c94aadddd1 \
--hash=sha256:5e262d41726bc187e69af7825504c933b6794dc3fbd5945e41a79bb14c31f585 \
--hash=sha256:5ee18d9ebdb417e384b58fe414e8d6af9f4e7a0ae761519fb50f721de398dd4e \
--hash=sha256:7008df2e7655c495dd12d2a4ad038ff878d4ca4b81fccaf82b714e07eae4402c \
--hash=sha256:734e20b57ba95624ecf1841e72b53f6e186355e216e5412de414e3c51e5e3c41 \
--hash=sha256:7c7e1a961a0b2f2472c1ac5b69affa0ae1132c39adcb67aba98568702b9cc23f \
--hash=sha256:7f86fd587c4ed9dd76f318225e7d9b29cfc5a9d43de44e5754db8d1128487085 \
--hash=sha256:7f94b27a62cfad8496c8d2513e1a222dd446f095fca8987fceef261225538a15 \
--hash=sha256:88dceee75c2c63af144e456745e10101eb67361050196b0b6af5d717254dddf7 \
--hash=sha256:8a650c2dbafa08d42e51ba0b62740dae4ecb9338eefa093aa5c78ceb546fcd5c \
--hash=sha256:8d65a2fbf9d2f8352685bc1364177ee3923d6baf5e7f43ea4959d7d8bc326a36 \
--hash=sha256:96481a5786729fd470164b47cdb3e0e58062a496f455ee41b4403be77cb5a076 \
--hash=sha256:a120733b01c45e9a0c34aeef92bf0cf1d56cfe81ed9d47d562f9ed591a9828ac \
--hash=sha256:b1d22e6e9387bf4739fbe23bfa80e93f6b0373a7f1b96c6227c32bef95a4d7a8 \
--hash=sha256:b8c198f8c1805dc42708689ed6864951fd2494f924149d3e4bce7710f8eb5232 \
--hash=sha256:c2541745709bad0264b7d4705ad453b76ccd191e64aa6f0fc66b69a293a45ece \
--hash=sha256:c742f741d58a28940ce01d58f0ab2ea3ced8b12402f162f4d534dfe18ba1cd6a \
--hash=sha256:c7f2c7f2b9ca6bdeef8f0fa897f8e05085923eb091721675170254cbc5b02897 \
--hash=sha256:d312ef37c91508b0ab2cee7da26ec0b3ed2f03ce12bd87a588d771ae15dcf82d \
--hash=sha256:d4d8fe59808a54658fcc0160ecfb1b30f9089906c50b23bcb4c69eddc19ec2b4 \
--hash=sha256:da25dc3563bff5965356133435b757a795a17b17d01dbc0f42fb32447ddfd917 \
--hash=sha256:eab21f45c7f66c13f2a9e0e1535309cee140182a9cdae1e041d02e47291e8396 \
--hash=sha256:eb0dc4e38e6a1fd579e5d50369aa2e10acfc9cace504579b2faabb478e76941a \
--hash=sha256:ec9bfaf3ad2df51ace80688143a6a4ebc09a248f6ff781a9945e51937008fcbc \
--hash=sha256:ede3e6487c5ef5d28634ba3f31f989030ad6af71edfb0055cbbd14189ff240ba \
--hash=sha256:f3c6818a1a86dd6dca7ddcaaf76947d5ba31aecc28cb1b67009a5877c9a64f3f \
--hash=sha256:f758f1b9299d059cc3f6546ae2af89670cb1c4d48ea29c3cacc4fe7de3058257 \
--hash=sha256:f8f0fc26ec2cc2b965b7a3b87cd19c5c6b8c5e5f436b984e85f486d652285c30 \
--hash=sha256:fd0409a3653af6c147209d267a0e4243f0ae46b011aa978b1080359fddc9b6cf \
--hash=sha256:ff18e6a727ee0ab0388507b89d1bc6a22b138d1e2fa56d1ad494586d61d2eae9 \
--hash=sha256:ff2983983d34813c1aeb0fa89091e76c3a22889ee83ab27c5eeb45100560c049
# via pip-audit
tomli-w==1.2.0 \
--hash=sha256:188306098d013b691fcadc011abd66727d3c414c571bb01b1a174ba8c983cf90 \
--hash=sha256:2dd14fac5a47c27be9cd4c976af5a12d87fb1f0b4512f81d69cce3b35ae25021
# via pip-audit
typing-extensions==4.16.0 \
--hash=sha256:481caa481374e813c1b176ada14e97f1f67a4539ce9cfeb3f350d78d6370c2e8 \
--hash=sha256:dc983d19a509c94dba722ee6abd33940f7c05a89e243c47e907eb4db6f1a43e5
# via cyclonedx-python-lib
urllib3==2.7.0 \
--hash=sha256:231e0ec3b63ceb14667c67be60f2f2c40a518cb38b03af60abc813da26505f4c \
--hash=sha256:9fb4c81ebbb1ce9531cce37674bbc6f1360472bc18ca9a553ede278ef7276897
# via requests
-1
View File
@@ -1 +0,0 @@
pytest==9.1.1
-32
View File
@@ -1,32 +0,0 @@
# This file was autogenerated by uv via the following command:
# uv pip compile .github/requirements/pytest-tool.in --generate-hashes --python-version 3.11 --python-platform linux --constraint requirements-ci.txt -o .github/requirements/pytest-tool.txt
iniconfig==2.3.0 \
--hash=sha256:c76315c77db068650d49c5b56314774a7804df16fee4402c1f19d6d15d8c4730 \
--hash=sha256:f631c04d2c48c52b84d0d0549c99ff3859c98df65b3101406327ecc7d53fbf12
# via
# -c requirements-ci.txt
# pytest
packaging==26.3 \
--hash=sha256:94edc256424af38762eb31306eed28beb9f0efc50a8837492c9d6fd6004aed79 \
--hash=sha256:d7193f7c8e4e93f444fde0262bf90af30e16fa0ad0ad44cb553c87339b23cd1c
# via
# -c requirements-ci.txt
# pytest
pluggy==1.6.0 \
--hash=sha256:7dcc130b76258d33b90f61b658791dede3486c3e6bfb003ee5c9bfb396dd22f3 \
--hash=sha256:e920276dd6813095e9377c0bc5566d94c932c33b27a3e3945d8389c374dd4746
# via
# -c requirements-ci.txt
# pytest
pygments==2.20.0 \
--hash=sha256:6757cd03768053ff99f3039c1a36d6c0aa0b263438fcab17520b30a303a82b5f \
--hash=sha256:81a9e26dd42fd28a23a2d169d86d7ac03b46e2f8b59ed4698fb4785f946d0176
# via
# -c requirements-ci.txt
# pytest
pytest==9.1.1 \
--hash=sha256:1088fbde8f2b49d95a549a195707afa7a76a3ce9bcadc26b6d71f0ffda5fe313 \
--hash=sha256:37a86b45efb9a47a61a36449063e8e18d0cab3161329fc099eb21783169c4f0c
# via
# -c requirements-ci.txt
# -r .github/requirements/pytest-tool.in
@@ -1,3 +0,0 @@
bandit==1.9.4
semgrep==1.175.0
jq==1.12.0
File diff suppressed because it is too large Load Diff
-1
View File
@@ -1 +0,0 @@
twine==7.0.0
-470
View File
@@ -1,470 +0,0 @@
# This file was autogenerated by uv via the following command:
# uv pip compile .github/requirements/twine.in --generate-hashes --python-version 3.11 --python-platform linux -o .github/requirements/twine.txt
backports-tarfile==1.2.0 \
--hash=sha256:77e284d754527b01fb1e6fa8a1afe577858ebe4e9dad8919e34c862cb399bc34 \
--hash=sha256:d75e02c268746e1b8144c278978b6e98e85de6ad16f8e4b0844a154557eca991
# via jaraco-context
certifi==2026.7.22 \
--hash=sha256:62f22742b58a1a33014a2b6b706588a8d7e2a88ae7bd1a6ebe8c992928483775 \
--hash=sha256:741e2c3b351ddf169a738da9f2c048608ff7f2c5cc02f1ebc6b118bb090d5d55
# via requests
cffi==2.1.1 \
--hash=sha256:046bfc24911b37851ee1b51aab8bffe713d89c68c6a057b09484ce9fd5f69b4e \
--hash=sha256:06c72bb76605a4b0cd0aad6930b69d4baf7dd5d806cfc409b824191099700e66 \
--hash=sha256:0beceaabe56af686895136a2de78db54ecd8e4046b236b8fd6d6cb61389e9bf2 \
--hash=sha256:154852545011f779917b11c78db2358d095da62a9a172b78ad0a583ee5adc0d0 \
--hash=sha256:194cffa889098ced9976c3fc6340305e43f6303657d298da55366907c05c22d6 \
--hash=sha256:19ee6127ee34de7d83ce3d371ebc5ed91addbdcc39f9ab15ce4eb35a4e534971 \
--hash=sha256:1a18a57b58cfb21fc28d72e876acf10eaed67a1ed96226f92af4df681d571c4c \
--hash=sha256:1aa5645c30469b09530c4ebca77ebf8f17618293c58f8549cb1a543a50236e7d \
--hash=sha256:1dea0e4d7d4f11f619fe8c1d76caf49e24405b4b5743c0e3be16a500ecd930c9 \
--hash=sha256:208f941bb9d18e768138677f0a6d2ce01f590df56043dda1df1535ac57c88517 \
--hash=sha256:210019b6c7cf07f081b4c54635c8cf744377001350e29cc0f81c4377b4797735 \
--hash=sha256:246fa40ce8645a614ff682e0b70f37134e460eaf93a775e0cbe3cca585a67a80 \
--hash=sha256:25792eac27877609e7bb06d42ff88278a6624fff2ba9bbb523c09616b117e80f \
--hash=sha256:27350daa11d4f10c540e6e89dada4c54feb7256ad03e9a4dc075ebad7ba360d1 \
--hash=sha256:28907ab9bfb6aa13184cfc17c6b8e1023c5ab6fd7076d8c20a35e59fe04f8f29 \
--hash=sha256:2ae64be792b8966f2c69538199728b290e34726562896df1e5dc8ffd8d8188e8 \
--hash=sha256:31348097ff5bbe827ccc41795d4dd099d9f0625e7def00ee653c137a490c2a6c \
--hash=sha256:3143d81e29e1e20a9ce10901ec369012947876596f75a222235965f2b7ae832e \
--hash=sha256:3222ba5d678f80a030e6afbcc33dc1ae5cb45facabb61cee2c7016b8432fde48 \
--hash=sha256:3311ed60d36f83378794e1009ac6258bafbf81f7888b4caa7b35a521e3f95813 \
--hash=sha256:334644fbac4eff73d985a17a91226df55d0f394160c4cfb880e084c8f7161cac \
--hash=sha256:34e261f78cb6ceaaa36f42f2613f4380d94d9c759a9c73c769ee6e0247364632 \
--hash=sha256:363e05fa78e15116c3c32c210ee36884fd6b9afa6d440e47112c3bd511d64cb6 \
--hash=sha256:398aff33cee2767e3e781d2554c54bd0dff386bb437581e0d8011fde1a942ec1 \
--hash=sha256:3d22a20b1fb1632cc72c22f95f7b0d2961c3e1c235f245ba4c606c4771035659 \
--hash=sha256:42a494cee34437f05546455144f2b5d9ac09b1face62bcfce597d2e521066688 \
--hash=sha256:42e2f76b9455f5a9a844f770bf3e200ed3da0e15f5df3db9c31fe80b04b3d004 \
--hash=sha256:42f6930c31dc7f50732c9ae793c2786c7b6b044195967bbdde40bb9be81c4cc0 \
--hash=sha256:456a61fa52d579ebf9df2e9552ead5129855dbaff6c1e5a9b1bc408809bdc062 \
--hash=sha256:471cee653ae88de62096552e6d24ccb4a5adb8c8c9f10b5054d0122c15bf2779 \
--hash=sha256:49cbc70e6542d4ccccb936558d1064a8012541e78f821f955cff24e357776c94 \
--hash=sha256:4a7c934f7360e8cd64fe9efadcbd10c7c6364f531e432b9a4bf5ccbc9e0e8b50 \
--hash=sha256:4be96343e422f2dfcd12ab5c9f5aebe03f82f737c6bffeca6830b3875cb44aab \
--hash=sha256:4f42141fc14250de6dde5ee7ea4432be017252d91f19c5ad043c084cea629cac \
--hash=sha256:507a24c282e0f42f8ed737cf048572cbf580468da5555764a8331735e9c736b6 \
--hash=sha256:51b31d1c98274844cfd7838ce00bfc27c7423a4dc00fc0772fc3331c2cc90676 \
--hash=sha256:58acb8ab8e295e6c5ea12f888cbb13cf21511ef2a3303a23f4325c29d17fe5c1 \
--hash=sha256:5a59cc1c4442bc3d5c703bf720b51138d0bfc173618807c9ee2490a7541dd3d9 \
--hash=sha256:5bb4e7ea95dcd6a014a6fef62e62467d67d8e582326443f3d68e71d6320a9fcf \
--hash=sha256:5c58fe613dc5e5336357eff555824a314d8e43282600435c8d1cb6a7a2fedd13 \
--hash=sha256:5e7cecbaadb83884793e05828cee59b210b24583b9c7425d0ba6a754fe22eb4e \
--hash=sha256:616f097f2fe415bc92a247f02e11f634e1f9e9a83d327e3c915c15089c87869e \
--hash=sha256:63bbfd5ded17c4840ac07cd8f1c21ba9d9708141f840b324f422f41b207e3973 \
--hash=sha256:64faea20f4e2613363a1a9b9c7dd73058f3ecd00133a511e72ad7c511658f527 \
--hash=sha256:661c298b4821edebead0c91edd2b00374d67ad7c5a1f7a91d4442633b79d6a72 \
--hash=sha256:68e62fe11f30d5ca8289242866f0a5291402d8529ca2178ab8afc5c9694ae890 \
--hash=sha256:6a8dddef476fab96d066d578fc88526767b836ab5ab21754e1d5bf3879c31c7c \
--hash=sha256:6e192623c49c94421616a5778fba35cf0d5a8d000650c1967ef4448ee5cdd990 \
--hash=sha256:7225e4514edb64eb6740324353e0da0711954fd8d7da4576755b1c6e09b697cd \
--hash=sha256:75f80557d1389eddbd0de2681f6a390a0c5338c31ddaa821381c203fc3fd50d9 \
--hash=sha256:770de9db11e84213beec501cfcaa013b019820ca881e03344dea5844f7876d94 \
--hash=sha256:7750c6449dff7864bb9bb27ddfb0267756189201a3afc911d82b3caacd70dfc3 \
--hash=sha256:7bde5e4cc5c10140859842b9d383af292b22639a4dffb725314baf45968cef80 \
--hash=sha256:7ce713ace7c0e4520535b42b77eaa742c16dab813978064913e5a3cf82973b41 \
--hash=sha256:7da0c5eff80f0197f3b3d1232ec5a682a9325f4ae9016a78f5f5ca35f9ced1f5 \
--hash=sha256:7dbb61fe3a7699468030f71bbe5f8a0e326a151daa91beb11a6fc1f980c55e1c \
--hash=sha256:811bd1e21d32de12efca32393a0ab3f5133b54fce9bd44b8bd77ab07da14bf6a \
--hash=sha256:8ef53b2de9bcb9197d31854256575d59dbac0cba72ac627bb291ef5eceb74be4 \
--hash=sha256:937c0052c05a31ca1daf18de3158eed4dbfcb9cc107adbea227728d647be701e \
--hash=sha256:9d2055050ea716bd38b7f7f1579c275386646b4894c155a3e2f3cd62ed41b7c6 \
--hash=sha256:9f8d177621de5cb38ee3e731eda45d421db093ec0739f46a5594babda7987a98 \
--hash=sha256:a2d7755bef5a12ed488f4ef1f1b69ee9191d7396083b755a5d2295f6edb4768b \
--hash=sha256:a48d62ab9d6f4f98c983223a547af44be6ca3691074c31cecced6facd3ba2dc1 \
--hash=sha256:a4f00aa42f75d6e4595e8866e748cc1705adc0cddfeb2ca86d0d03993d63ba03 \
--hash=sha256:a6e721d4b0e45d5b65e87534470e67b18dcd092c83f68fba09f152b9cbc061af \
--hash=sha256:a730a083190634c65cca36ba5f489531576ebd79bcd5c8e172130f6453127231 \
--hash=sha256:a931079504ecc49efed7744c476a5c343a92fabf66dec2db95edb1b2fdc770e2 \
--hash=sha256:aa9511c62d14da7aacc9b4bf51f3f697a621e83b2d6919008243c3aad168eea3 \
--hash=sha256:ab36d55f9ed2d067327667c2fea18dda018eb628dd6347aa01dda6cf1f5d3836 \
--hash=sha256:ad2c86c495b899d862ea0f4b42891b8713a3bd45dd4105c7fd51c2a72f39f3a5 \
--hash=sha256:aeae0e330c9f6acd681f647d46cefd30c29f93e3392882e792e82080c9691399 \
--hash=sha256:b0431303acaea1089ad4b3e9ce4e6518193def1118d4073ca848635ee4ea2e96 \
--hash=sha256:b5bdfd1c873d4e093aabc0ca84c4ca6dbc4f752afb5c86f146d9742580c9da2e \
--hash=sha256:baed1e86cc735622097354b9d1281406caf42ff42a886d29faa8e8d1630333be \
--hash=sha256:c1453022f490d2459a11819d83ad1d586e9ff65a12ac3e705ffebd46d3685dcf \
--hash=sha256:c26608d2222fb1e94487e4a387d85f13eb55d5ed725cb25a0c589ac4ee60e7bc \
--hash=sha256:c7659f22557c5a0bc4855cd635f55edec690cc008a40768527762cb9fb263455 \
--hash=sha256:c8c69575568085ba0b1b10c0249d779a214aea6f6522e949a0fc9fb0fcb449d0 \
--hash=sha256:c8d2c9fd1f2d16f780d15127abb050d13d1a76c03a4bd87d7e4980e45e511e12 \
--hash=sha256:ca82be1a1d406ecfe1d25dc16cb33488e5a16bf4438c9fb590484ea29d92478b \
--hash=sha256:cc572dace3f60ef98d7b12ff411d20f5362feb31a0439eab0085bbfd349982d7 \
--hash=sha256:d18e5ac0f2f03f4f518d3e23db0f0cad7faa1da8620e9c09461d443bbf6e6692 \
--hash=sha256:d28630f5854ab07ab1fd4aba756de52326c82e6be15d414b12793f1975048b54 \
--hash=sha256:d9c275eaacd24aa73f94ffd6de08fc3f932424d8b6c376f4bed7cde376fe7bc3 \
--hash=sha256:da0e573f9f97159390c89d9f1a9e41908b66d408cc5b58d08cf3847d844c531b \
--hash=sha256:dd31f52ea1086513bb9df30f8fcee9b8918323ae067a3d5b78bc826a000712be \
--hash=sha256:dddad92b554513a31f272570678ba307fb9f618f05e3d4a5eacafff9eae03e1d \
--hash=sha256:df423d40ee8654634421812bc3b196da3f9bd7d32929da813f8394c4348a5358 \
--hash=sha256:df913725b79db7bcf03448f36b7bf8815363417d5b58deecf9305e3e30f0f21a \
--hash=sha256:e0bcb7e0f677f543555d2adff3bf19c05f66cdb4796e5ff602442ab2fe3c4ef7 \
--hash=sha256:e2d65b31f36619cda3999b78b2aa9632e76b78448e7a56fc4240824200e7c4fc \
--hash=sha256:e6e8cff14d6fb0be70a09c0bdc58096f501952d04624ebf867e0e56da2df8960 \
--hash=sha256:f16c709686a78c727bbbf059f92b0bf41c6fc60deec706d2dc19f529175a6125 \
--hash=sha256:f24fb43132a4c6b4cb4eb029492919b2db645be6808d738f244fd146c03c32cb \
--hash=sha256:f53e442b08449d42821fa4a4fba000095af9f62742a500f978a9f557ec44339a \
--hash=sha256:f5cfbc5fe74540d335175b656c725d74d90e3730c626d92575eea35029d9afaa \
--hash=sha256:f81b3b8f3d4e343550fa4baa0e479bba9f2d29ce9c2e9b51d1ce1718d7442fcf \
--hash=sha256:f8ec5e643a9a937f64e1999eb9f75d072263751912dc5cd06d3c85f8f44be7c3 \
--hash=sha256:fb92203a88b3d3053034db775110081c49d28be6551923805e039924093761e4 \
--hash=sha256:fcd22650c908d7b7da162bbfaab594a1227a15d1643a98c68b122ac642fa2264
# via cryptography
charset-normalizer==3.5.1 \
--hash=sha256:00668ebb0609751758682eb0b5857e7c35b9f00e84dfdef062e103244ec94d45 \
--hash=sha256:012a22b88a77ca2e59b98ac5889b0deb604147666032f45e6d6e217634d2550d \
--hash=sha256:01e93745f7f219b703b60ba7afead36cfc4242782be5af484673fc500df12da5 \
--hash=sha256:04368edf83514385ffc3e1cfd4546e595f4f1272dd23ba437a93a9cc3741d47b \
--hash=sha256:0722590aabf9dc6a6c0343d523c05458fa2b5047dbe6302fd526bb570600753f \
--hash=sha256:07ffd07412fc5d5e84cd8952acf9ff7e4ed7a708e69d1bada19d8ba91711353f \
--hash=sha256:09a7bba9f739468c8e78c36a75c33768e53cb1959fc638f510454c14683f00d5 \
--hash=sha256:0b2b1b3fa5670c127b246df1d0c059defd41f689a868a3b9d79df9b1cac42d22 \
--hash=sha256:0c6dfb5ca6723eeed15aa8e564a014d69fcb8812f94eef11fe3631e0508199f5 \
--hash=sha256:0d929fc574b4d6fd9e7c0f5c2ede8716a41911923aa7fa5fce38e0818aa4a1ac \
--hash=sha256:13e3afe97712e8887cd516e960c63f0b93122971e5b5e4b2622fe7701771e838 \
--hash=sha256:15f024313246a4ed976c60f440bb8d257815513a681d212ff74fd46f7d715a90 \
--hash=sha256:195ce897c6153c0700078142cf8efe3e6454ca4cf4357499e4078dfd83396626 \
--hash=sha256:19a3dd5aa73cef1c99687c4fc57db016a9c17104ae1185da88ba566a5d3bebe4 \
--hash=sha256:1d1c7a53a6c2103925cdd6d7229f8c567379f211c869793df679f2e9f738c369 \
--hash=sha256:1f5883d77fd409a261abb5dc8ccbe335720d798b1de4abb3b1d47ccbbc76b53b \
--hash=sha256:21b82d8082f6f5e7f456ef0bd16323d08de1266efbfeb476e64b2a91d1471a4e \
--hash=sha256:252d099029bcbea642f2a06c4ed5046bdf8b5a8150b64afa5e027e88b106e5ee \
--hash=sha256:256dd4d85d9e4dc595e2bc983c980e73f62ddeb3165c58b4c3dfe78c5c8548c1 \
--hash=sha256:26422d45fd13551cf564c58932f7d72b4f58b93b0fcf18c35ba6be12b46bb102 \
--hash=sha256:2679de311c7946dde5d3b6f44941844133ff5c7cb86099c0061ab1e8901c20a8 \
--hash=sha256:29880d17a8eb0b5cfdfd8944b468322928059aa35f1f5fa8ff22b149ec0b42f8 \
--hash=sha256:2bced4061f000f7187254a02ad3433ae17eaf991747ceea2f478422590a5bba9 \
--hash=sha256:2e9cf9253119d8e5d111f05d71626786fd3d6193817316eab1ca088cdb8593cf \
--hash=sha256:2f06b7eae9dbe77fe1d644ca244dad508de8d302870a43f3c559b521270938a0 \
--hash=sha256:2f293479cce755c75f1697e87c409b7ae4c555c7dfecb6e988ad13abba943031 \
--hash=sha256:329fc3ccb63ad22d867d84c2adea759a64079a37ba4a343433b02c7a2816871e \
--hash=sha256:343fb4f2821043bd87095f7b08a1a181febc8e36ac64212143bbfd0a0e1bc235 \
--hash=sha256:3588e376b3ea2eea84976f67273d679f229e24c66dce7b82ae45aef04ff6e072 \
--hash=sha256:35aea775dc2bd5f54cd84a1cd2696cc3207c479cb9cf0bd346f0d343e4300ddb \
--hash=sha256:35fe081843b35aad20ffeccec3eeffbe637b15d14f3fb22cc1b59cd8ec17e93c \
--hash=sha256:36047af20e17097c3bb9476c2b7655f2f7aa51322c0ba58c07695bedf755a950 \
--hash=sha256:3617ac3cfd8b9888f145ad89dd6e692285834b0201c6074a5eeaad3fd4d668c2 \
--hash=sha256:366ec70f5547c640d3ce1985722490f23faf4eb5216a7eeba78277490e78dacb \
--hash=sha256:394fea06235c8543390050ed5f529187074b029fb027213f6c46ac11ab5d950e \
--hash=sha256:3d27167433c0d5f18dc850f07d0b3816221984fecdc405d6c157a6f0b8f8e9e6 \
--hash=sha256:3e5e1224c0a6a90e05843e07adfec669edebec17801c67072f51e59561d63c0b \
--hash=sha256:41876ee62a3dddf48ff1121ad8f0798032aa03f2fd35f21f34a4cab14f18d8d2 \
--hash=sha256:433c5a81eade63b47e522303bad236f59dba55ea6951746f5558355eeed8c75d \
--hash=sha256:4582c27e8c889d64811987b5967fbd3ae0c823fe1fd933b543d55ac20bb475fa \
--hash=sha256:485a0d363cafefcd2538a73c7c838daa2035f09b2c9f9b5e3133f80c6aeb84c2 \
--hash=sha256:494b70049a4d69aec6e8137c13af4cf8db8c9f9820a1392ac293b0dd2987a818 \
--hash=sha256:496846868fea80e479324862fa877f02411f2fd0f83b79ccee2607aa68b2a032 \
--hash=sha256:4abdc5f9ad448c1ecbfae2974b820535d6bc6e7eef63babbab3d81cf46968c71 \
--hash=sha256:4b599739b93b2cbeded49645ae3c8d1405c29ddfbceac1545c87a3f9580a9e96 \
--hash=sha256:4bea7f8ebe90bbd7f0e4a2de42ca6924ba23e3e76418c408ff82f1d46fabd687 \
--hash=sha256:4c4fb141a727957c93edfe5c32a26ceb6b5f6461d67146e2d39f51e16170bea8 \
--hash=sha256:4c9548dc78002099910abaebc0a72ac58b7d30931869e0351c09b507dff4ece3 \
--hash=sha256:4d26f14f041e83dd8edfd61f4cd4fa7285d31798b5bf1f28e70c367ba6c41d61 \
--hash=sha256:4f298bdadb8f0b9e5672877f647d1be9373ef5320c9e2f049795e26cad28b6a9 \
--hash=sha256:52ec005752a56ae79547a05c0139ca2501a0c866390b6115008456b9f0e7cde1 \
--hash=sha256:55261ac0d2941c42f196dd576f543d87a8ee03cd6f5e30dfb4d807b2e3b9121a \
--hash=sha256:56490c595a28b1bb27dfc583e816152a9767721ef58b2c03b13f954d2f707420 \
--hash=sha256:58d3e12c88e0950bca850ae1f7c256055c097639c2edb9eb123af9807d8b15e4 \
--hash=sha256:58d4aa13a59c969dbfdf9e6a9560e242cbfd9e8a8f50c2747714df1a423adf65 \
--hash=sha256:59171c6e45bf07d0d5cab3b0bf81d945035530f6873398b3b531c31184d46663 \
--hash=sha256:5b6d1386bf0096d26d3a863dc0a487a5b4eb9aa93cf5ba69683d29dde6b9d60f \
--hash=sha256:5c0ea61a470e070686aa30892fed79e297d2c8d0ab46b8bcdf027d38c51da591 \
--hash=sha256:5c84bec0ab5ae0c64bfe73a7d2adcb5ce73b467523fc27fd6a28ab2aa6cbe35a \
--hash=sha256:5ca0555312ae2fe82715cada7fac375530c2f3349e1eaa1bcb33d0283ac79a18 \
--hash=sha256:5d8531a6569d025f68e2321e7638fb7978f23db58e5f69f56913837aae03816e \
--hash=sha256:5e2d0e146dcb57034f8b97dc58d2d512cb90aba253960ce449f695fec6a82c6f \
--hash=sha256:5fc45d653ea8c9a20479167e11d4a0f8cb2fa3470737ab6f9c827532313187b7 \
--hash=sha256:6117b84ea48435e5356dc737f5121485c30920ba43375fa7b434fd753df0eac3 \
--hash=sha256:6199d5606e2bbf2b096cf64d03f8b6790c91081d5ac866b8e7bb6422738cc60c \
--hash=sha256:62b55f6722735a6c472f88361cde6640608773d9443cebdbb51abf436a1fcdd3 \
--hash=sha256:687c9ca3035544b113bea2055e180af96fb63c0c476e22a9180f51925186e7b7 \
--hash=sha256:6b7430cf5728e68f6c462254009a6ef4086e1bea43cf2f57aa9c55fb4f50ff96 \
--hash=sha256:6ba32c4d2abf1d2fe7cf27d280f4cca5664233b0f885549c7761719eb977f486 \
--hash=sha256:6c9cdde8becb25a7fde49924511aa2644d6f8081cc8df8e9452724303348d8e3 \
--hash=sha256:6df0ec430f9a831772c23ca5a224cba36517a58a84bb32c32bb59a9fa67c47f6 \
--hash=sha256:6e2912d4babbc65196ac13c2f53468dc57fb8b9c25ef913e8c59ddf7c6dc0e1b \
--hash=sha256:6e5e4d73d588ca5ed09df1b7dcd1b203d1df3c542e3f50d126c947d432b10731 \
--hash=sha256:70055ff39b97c99e7ae40ea3e393fb62aa2e44dbd9b29f8d14f42fb0025c3959 \
--hash=sha256:706bfd38730a5ac7a365793269a00f4e988178cec121391f4248d84ad8c972e9 \
--hash=sha256:7235dc28fc6dd9d832ac7c7bce95367dedb85929f17368a0c2bee1e080b9acbf \
--hash=sha256:774d157f112367ff4abd29019f38f023c24e00e56edc7829c20e358a5a913ad8 \
--hash=sha256:77efcff2b23071c349402ac1066667a3d011f62398d81408c9b88ad991747c9e \
--hash=sha256:789b8982559ae28dad2356519f841655756cdcd96616410590ae0b17454ee64f \
--hash=sha256:7ac76cf9afd34929d76eb7fcb63be476a4853d8a96f0dcf2d0db68a0cbdf9885 \
--hash=sha256:7c0c10730342b0c9b35dd1d619beb8214e520bd96a1f870f452680b238aab3e0 \
--hash=sha256:823f82903d189af463d7df250ef1f7f696f3cee08cc8d91deb565e8d425f6506 \
--hash=sha256:838648accb3a7fd9803fd45c87bce8509648eb0c11bc34e216141300977244f2 \
--hash=sha256:854066be00447fa8de2ccbbe893e2ffc4b123ef16d897af794c1e18bd4a714b0 \
--hash=sha256:85d5855daafc240cc045c026d7a15fd198a09b0fc8ff6f5ecbb5297b509cb11e \
--hash=sha256:85de3134b5379856e323ba37c19c9256d39425f7b76a63af52b09fb4664c2e8f \
--hash=sha256:87e4f41d375c0b9be2fb5251aee4b8a689169e134535aed81bf085c3b647451e \
--hash=sha256:88ca277405c2d3b71c4e1c2ee0e7966e807bcba86a69d11e19ba199d18ae4491 \
--hash=sha256:88e85ab89cb822c1e635f51d6d32e488f94e002e70e2f492bdb8b945543f345a \
--hash=sha256:8ac8c94b6539074e0f40899301273ac8402b9b3e01c7b7ba269ff30340aaaf20 \
--hash=sha256:8fe532b3c966d1fb794e0698e4589d0444017ae77fc0b31edea13c0e35bcc449 \
--hash=sha256:9085f87b0e38a2b92b8923059b4e8789fe40d9279712d15dcc670048d77079af \
--hash=sha256:90b7481fb62fbe172c558bc6fd1c4c98d82004a54a7551f20e11ac9bf0b8708c \
--hash=sha256:92caef967d287a407085d61176fce4012b1dd62daed4eb6d5ceb26d3d2538712 \
--hash=sha256:9362dd90aa7dab48c0054a21187791ccf05473f7dba5d92b8033ae62164675e7 \
--hash=sha256:94d78ecec2605a8d0398b0f365d5f12a63248438516f5dac536a5eff7337df4a \
--hash=sha256:94fbf1c0c6cc0d3d5e50f9a9313a8cdca90dd696d34b381cd1704f8c9e939f20 \
--hash=sha256:950f23cb393f85543777b0433f082cddd25b51ab398eac7971146495679efe5f \
--hash=sha256:96eefc178f8636b9c760c5829345307fd81cfae9ab1e80997dbddeb0f54ee9a3 \
--hash=sha256:96fef3e886d6a9874b14f27fc193fbdc69d5d8035783d86aa4e1cea594e695f9 \
--hash=sha256:977cdbd483a9cff38179bea4fd754289a6f2195c7abd414aba85410b3e66cc5e \
--hash=sha256:978eab16f55b4ab2c2a745be9a0a840bf8f09a7f227d9c76eb30214d078865a5 \
--hash=sha256:994e883d17c559cdfd38c84003c8b27d25424a1077272a17e7cd27bfe0bf57b2 \
--hash=sha256:9ac4444d8d4fd4c4bd08bf451ed3167aa9e7ec6cdb41b648794f1d1103652e36 \
--hash=sha256:9b5db6052055d34d41230fb78d7c439c23dc536a9896f6cb039e8dd92cfc1263 \
--hash=sha256:9d9a0dc7cbe9bec24c3f767c9122c41fe5a1bc43f47cd099d00d393e09769de4 \
--hash=sha256:9dbdd9205662134957cf0c324f639bdc5031c0ca056e2369e238db75187c0f11 \
--hash=sha256:9eea3ab2597a5e65fe65296e2d6a84570845a6b55532d90333d740d48bbc850a \
--hash=sha256:a2028475ba855475b8b4d3cfeb4994269c967aea8b9892dfba907f4263a863a3 \
--hash=sha256:a3a370082ce34d0612f421e15fe011c53bb1feff21a26d06ad4fb244dab5a375 \
--hash=sha256:a545775cfe815855ea32d7c27731d79da358ef2055b4a25830231b1622dd18aa \
--hash=sha256:a5cbd90ecf0fc62e64726917ad083b73001f0563657a87ec3c0b504e277dc90d \
--hash=sha256:a6d095662e73e74f0a49988e0593373e243e3a52e27bfeea0a859e88acf4a0f5 \
--hash=sha256:a6dac12ff6b846103483683f60c5f8fee205121adc58ffd87e90a90a3af69e99 \
--hash=sha256:a951ad59cad9145664a730d3036b40b844e74d2d3683da40111463cd3a83845d \
--hash=sha256:aa1099b956fb795e686d073568f6dc002a0bb89765ea6d5b055dd7d9bf1b116c \
--hash=sha256:aa2bb0b37202dca27175591f761108b5d34096ade1191ffe4808bdf6b1571488 \
--hash=sha256:aae2ee51122d3ae968a3837d97dc24a0aeebb0dea23694422cd172bd30017cd6 \
--hash=sha256:ab743e9bc90c1f73552ec33e10e3331315acd2c397b36065b591b0181de533cc \
--hash=sha256:ac00177c4831ffa650f8609e4bdddd5fe09c03b1c0c47acece7e6ea20421598b \
--hash=sha256:ac13b004224fb341e1e25a1ed5e19d32f57cdb2a403e01f003b46f051a550f6f \
--hash=sha256:acaf604462bf330b0d07e7a07c1d6e4adac79e5fb13e9c5140590542cafacc00 \
--hash=sha256:ae31a1a1db2ee6cc2942fccaf695c934bc7f3db9f2133a3fef1f367cf1a4ab10 \
--hash=sha256:ae4a097991662cd4fff0ddc74e0fe7874f82e00042fa0ea00855645ed0c79598 \
--hash=sha256:aea996a6aba25260827c9ea511d1addfde2da9eb686ac961838509086188b7e6 \
--hash=sha256:b39b69b347e5e47a3b5b8cfc005c68c1ba347474e3960236c4944a8ecd174962 \
--hash=sha256:b54e7e13267d49ffbfe68e25b3cbd774dab38fa37238f71265e91b36146eb21c \
--hash=sha256:b9af956078716df40d985fb0dfeb2c2120c5ca92ba4ff4b388acfd01cdc14d08 \
--hash=sha256:ba2f37ee79e6338845261a3c5b1784e5d1acdff2c0785b284f1b633033d136ab \
--hash=sha256:ba501e667c17d8411f98e67a022d9604ef179aff0e459b7e292c796837c13573 \
--hash=sha256:baf3775a2635e5a11fbd5e4e64ee69c7e86875d224a5c72aca4c141064589a90 \
--hash=sha256:bb57753e36e4855b8ca375069482250a6246372331a3e4f3407eaebb007443f5 \
--hash=sha256:bd6c173f04743d483881bffa1478d5a4624475b8cd1d2194956a75548e191c18 \
--hash=sha256:be47f99644b208bff7766314013f9acf57b056b04191d570d68ad14022cf5b1d \
--hash=sha256:c010f5581d9c612804cc59fcf7b524b707fbcb72828551237ab545bb5c7034af \
--hash=sha256:c1dcc36dcb96abc02236e182d17e0f71430152a6c2c7447421da2d2dc144edea \
--hash=sha256:c428c6c31eb5f4277d7f8eccaf767fbd548ddd5ce3c8b4f4cbbfab3d96b5904c \
--hash=sha256:c658c50ac0c98cd755a2dd50b7977d3bca7df401dcc47fbdfa87db53ef7d4e8b \
--hash=sha256:c71fb0d56c920c269cd3e2e3fe7c610e3f1fdb21a6ce60efa6430ff63676cea6 \
--hash=sha256:c7b742bf31c88566b4bb6335a7f393bb322e580b6bb98df7bd0c25e6e3519ce8 \
--hash=sha256:cc0329df4caaceb950d2f580b5ac716a377f7059624a0bafaeaf8a218c6ed774 \
--hash=sha256:cc5d36d96478aa9c60654bd932525bf32964c62a7281eafdf16d85003a8d6004 \
--hash=sha256:ce854f5f478050ade5a238731c4ca985a7d3b3cb53ff600a9b5c3b689b5f0a7a \
--hash=sha256:ced3fdd71aaa83ce593746c2edb42b7a59cb4c19c8b5c407781c72e493aae55a \
--hash=sha256:cee5dd7c6fb5dd52a0fe2a740f9bc6e3593f5f8b1788bde49de02086f30182b2 \
--hash=sha256:cfa1c0cc3a8f9f53f1243a5a99ac36fd003880199383b37672e86ddda9cb07e2 \
--hash=sha256:d1ee1e296209fdce05b81b663250eefa02213a2da7b41bf26f7829b8ba3545aa \
--hash=sha256:d59b75732e9b6f27388e10c14b0259cc5f2e48c78627d185e6a177b58ad3cffe \
--hash=sha256:d63600d620ad0064c3a748b950ac5ea38a80190e5498532efefa4b7b3f1da1f3 \
--hash=sha256:dd732602a7009217f658d5863d12d79d373a4de0eebc111094bcdd3bb8e0a6cc \
--hash=sha256:e06efa066f7dbadbc84ebc126a97c452a6451dfcf589d89d788484949e1cf795 \
--hash=sha256:e199fb99720074809a7720f1c0b4d919eea8b87e88713e0f8f602f7bef543d9d \
--hash=sha256:e4b018dc5a0eee4676e38fe84a47a427816c590b93b55d9025274ec4d6ffc2dc \
--hash=sha256:e6621fb2a4988d6e53eedc455e5903e2679f3967b8acb3d639f1b63c14a2e893 \
--hash=sha256:e71c909f353863b2b89c83de2ebed71ea6d0df8a6ef65a128193c5e650766bef \
--hash=sha256:e90251c0c7bdd54a100a0dce3c07b7e637278c93af29dbf78ebb89a58c4bac7d \
--hash=sha256:e9fbdce1e47394b09bc9f26ab117dfc8d6491977a11d86f592bb42c779db2fda \
--hash=sha256:eb12fb2ba69ffa05f8695f61c69e591dc4b4a12ac3757ac8af8adb259bf56d17 \
--hash=sha256:eda059b6bc8bc0812d626fd91a7ce01bf583df0a61296eff390fd94141a34e30 \
--hash=sha256:f03ac127268b43ef4fe9e6ab6794a6794b49485a0cc0c1db79876d2f33f75bc7 \
--hash=sha256:f298e218441525d3794428b4c8b8fb8662c6d3ea79925d4807ee6b9a96a3bca5 \
--hash=sha256:f5542f9b941279d82d41eb0aa9f98eba36fe4df5c7086c651df7944935b37182 \
--hash=sha256:f6f7deae3feb4edfa2efaf7c574fe88cbf055038a6abdb40188e4fff66d5699f \
--hash=sha256:f9b1e28d0e8dbfa858abdba91d6b547beaf2df1a59bec6da6faae7b96a4991a9 \
--hash=sha256:f9f8405c2c758532c74fed975dbee57be1f31a6e865c031870c79a6ed3212ada \
--hash=sha256:fa48b1b63d639f9483e0633e092f5851e2348c352f1f9bb6c8182f87884ef876 \
--hash=sha256:fb78f6e7fcd8ad785d28cd577168bc1aaee827b25bb8755638f694794ea98f0a \
--hash=sha256:fbc597639158fd7c14d55e808718848319540f51b0e6746e3eefa59723a4a348 \
--hash=sha256:fce8cbd4997efeb450bd298b54f755dcdff18d496f7a5ddbb4867c6d7c88fdc3 \
--hash=sha256:fd0350afdc3aabd5576f60ea109228bd5538139713c7b094c5cd27c73a98bc6f \
--hash=sha256:fd0a274c0e5f9a21565cd9d3dd749b61f96b7aa1e20a93aa1ba4029518f2e5c0 \
--hash=sha256:fdb8a068947befafba9952162645dc2fecaeb400e64584829ed5e9b2fbe21a7f
# via requests
cryptography==50.0.1 \
--hash=sha256:01f41478cf33fc605a6a089cd56d28b45c6c0b45a1928b61797f2621a04bac71 \
--hash=sha256:05ba322c4da95b262a212c345af888ef2c37c88c0509756ea00a0e6d68850f23 \
--hash=sha256:16c5ecd954b3330ebfb6605eca4fd952da8bef376551d5cc264534e3770a9ee6 \
--hash=sha256:2a93d05e34d5f67fba6f891fe85d929999baa7195e853923ea6d7576c9e68c5e \
--hash=sha256:2b34d76a652ea2b6faf777c35df230c5637842cd904e04f16230c3f9f03e4361 \
--hash=sha256:2ebbfb0f1fed745e91796e3e1080a1440423fdae8ece1b995a1d80883a409054 \
--hash=sha256:30a125032e5642a21ff816e021152bd4e7e94f03eff3f4b7fca41cd22bc3110f \
--hash=sha256:330fbb252391c596f1ae42c5754449dc924e6ad012dca8efe0d703f9f2d12ec6 \
--hash=sha256:359e62deae718bce96170e223fdcb6357e4fbd3bb7a3a75f4430763532560e49 \
--hash=sha256:407fe2b6db00939c05c0e945e9914238f2f0a430974839429dafc82b1ee6bee5 \
--hash=sha256:42be3bb70596b3abe4ac097b75be223e8b3ab614a0e5de068e3dcc54d71d6149 \
--hash=sha256:4c4188f7c0cf655be5c06342b817ed0f9595b69ffa2b12026e5353eed29dea88 \
--hash=sha256:51593d180cf6d179bde5c5d065bed81386b1f381656ae7d042b7ffc87a9895ad \
--hash=sha256:51afcfceb15597cf2635068e4ac9a56b2abde622edde17f37d85fd7b5306497a \
--hash=sha256:53e279950892dc102c6b4e52af03ae5ea92fac572a1ddab78ca73a997f62b69f \
--hash=sha256:55d16b1ef3ee0958d893a977b19777887e546c9954ea81b200c3301a864013f2 \
--hash=sha256:5dd9bda1c12b4162f6ff568eeb5e0ff956c28d14406e875cfe8a63a2d414ff20 \
--hash=sha256:5fe002589592ed749ce77fe0695fcbd3500dd61d7d6db5858a7544c612fa8e45 \
--hash=sha256:5fe939deeb161024a6be98229c953b6591fef1f41214497a78fe793a244c017f \
--hash=sha256:693c99b49bd37d0d096e4334c10232c77248c415b98d35236094cdf96d57258b \
--hash=sha256:76de83fbd91ac49c0feaaa983d0748fd7a53176afac5fb3bf7478d244f0eb527 \
--hash=sha256:79bf008d1f9af6071c797ad133e39915dfee7614f18f18f4db9072eb715064a3 \
--hash=sha256:804728ce710890870f3aaa344b2e161172d258d768ac139d02cfd9092d0d94e6 \
--hash=sha256:8921d58f426793c5f1b47f0b59575780de9a095214958d0eb37d909593db8367 \
--hash=sha256:8df2de9102026855887e4587084f6eabd80ed0f345b8ad8a7ac27ab9bf4723e0 \
--hash=sha256:9cb3cb952cf5a8abd50c782a98a89d71699715e802fe349704b47f2425b42a94 \
--hash=sha256:9dde0a357190eb3b1da1bb9ab750e9c85cba82ca5977aa0836cbb94e92611239 \
--hash=sha256:9ebcdd5519be9b652a46f507817a74591774fc3d6923ac364e4dfa64e36b291b \
--hash=sha256:a0b1a59e3a089064a0ec309e9428c8e3ae4e161419d20ac33600767e83fc658a \
--hash=sha256:a255449073358275b64b67d3f595f268bbef70e72b6edb65e0c70c735bf739c9 \
--hash=sha256:a8f40ea47330e71b594a7e246898f93177c259490c63183dbaf9e571d71ed9a5 \
--hash=sha256:ac02b07824d4d1001bd4367599f839c19cb171924c796e52c23508ac14c2c0cc \
--hash=sha256:aed8db4f6d71c51efb89530e12d9464e7bf2923d46c3205dc794a2a93f8c0648 \
--hash=sha256:b8f852c65863251b9e3a1b8c150ce21e59b522dbb6a7d4bc80e680d38388e986 \
--hash=sha256:be224a65493ec5b74a158ff22a5522ce4a5ca1e543c647a3a4730d4a09e5f959 \
--hash=sha256:ca83d00d9e69cd5eb63f2e69c3a5a59e0cecae5ae14c6ae0b35830fe3b37bad0 \
--hash=sha256:cbf74a81765ee67413503ca6e26dcc4f6f5a519822436cc0a1b97aab6c1b8a17 \
--hash=sha256:d63ae8f6481fec907ac0f588eee8a90aefde112c633131fe540e5711ddbb5a4e \
--hash=sha256:e22dfed744bd4002e909464cb23d2f0b05c6f3113a79ef2e9864a53db737c733 \
--hash=sha256:e2ca8fd1b6b4b82a1c4cb02841d0837e3c12336c2e24b520ab8ab3b969733d8f \
--hash=sha256:e74591e283fe6eb956416c929eb58262a719fe0311fd9054c62c3350ed8760d8 \
--hash=sha256:f74455bb086a85d5e81246412602aaa97ed095e504cd40dd261ef50be42205bf \
--hash=sha256:fb4b9672d389c738b175c4166e78310f8a70358886aacd9173ee03a85ffdc671 \
--hash=sha256:fc3ed7ebd2a8c96f5b166de0ab9b624996bef3b07bbeb19364dfb78222c22c80 \
--hash=sha256:fd3718b960d0b5dd213cdf03f3bcb7000e69dda0de8b956061947ff6bcff5558 \
--hash=sha256:ff838d62ec1bfce4f9ba7fa16f4a7b554cd8d0c299e6be37502161a660c84eef
# via secretstorage
docutils==0.23 \
--hash=sha256:25d013af9bf23bc1c7b2b093dff4208166c53a94786c9e447808335ef1185fea \
--hash=sha256:746f5060322511280a1e50eb76846ed6bf2342984b2ac04dc42caa1a8d78799e
# via readme-renderer
id==1.6.1 \
--hash=sha256:d0732d624fb46fd4e7bc4e5152f00214450953b9e772c182c1c22964def1a069 \
--hash=sha256:f5ec41ed2629a508f5d0988eda142e190c9c6da971100612c4de9ad9f9b237ca
# via twine
idna==3.19 \
--hash=sha256:5e0811a4383b21dc5838069f801c4fb62113b7447663d2530d2bd6e77b49bf15 \
--hash=sha256:815e7be7a7806d54abb586dc943addc79e8b2ee16915059658cbeff4b1b43bf4
# via requests
importlib-metadata==9.0.1 \
--hash=sha256:ab830580bc0ef3db61ce8fae716389e5462b67e033018bab6d8f80ef17172f99 \
--hash=sha256:bba5600596a7e21f3eef53281cf28d6a5195634d2f2b78ff9501a3272c6eaab0
# via keyring
jaraco-classes==3.4.0 \
--hash=sha256:47a024b51d0239c0dd8c8540c6c7f484be3b8fcf0b2d85c13825780d3b3f3acd \
--hash=sha256:f662826b6bed8cace05e7ff873ce0f9283b5c924470fe664fff1c2f00f581790
# via keyring
jaraco-context==6.1.2 \
--hash=sha256:bf8150b79a2d5d91ae48629d8b427a8f7ba0e1097dd6202a9059f29a36379535 \
--hash=sha256:f1a6c9d391e661cc5b8d39861ff077a7dc24dc23833ccee564b234b81c82dfe3
# via keyring
jaraco-functools==4.6.0 \
--hash=sha256:880c577ec9720b3a052d5bc611fb9f2269b3d87902ef42440df443b88e443280 \
--hash=sha256:99e3dc0060c5cbe8fcd1cdb36258e2a65ca40f1566b2033b12abb1bb44dd3c30
# via keyring
jeepney==0.9.0 \
--hash=sha256:97e5714520c16fc0a45695e5365a2e11b81ea79bba796e26f9f1d178cb182683 \
--hash=sha256:cf0e9e845622b81e4a28df94c40345400256ec608d0e55bb8a3feaa9163f5732
# via
# keyring
# secretstorage
keyring==25.7.0 \
--hash=sha256:be4a0b195f149690c166e850609a477c532ddbfbaed96a404d4e43f8d5e2689f \
--hash=sha256:fe01bd85eb3f8fb3dd0405defdeac9a5b4f6f0439edbb3149577f244a2e8245b
# via twine
markdown-it-py==4.2.0 \
--hash=sha256:04a21681d6fbb623de53f6f364d352309d4094dd4194040a10fd51833e418d49 \
--hash=sha256:9f7ebbcd14fe59494226453aed97c1070d83f8d24b6fc3a3bcf9a38092641c4a
# via rich
mdurl==0.1.2 \
--hash=sha256:84008a41e51615a49fc9966191ff91509e3c40b939176e643fd50a5c2196b8f8 \
--hash=sha256:bb413d29f5eea38f31dd4754dd7377d4465116fb207585f97bf925588687c1ba
# via markdown-it-py
more-itertools==11.1.0 \
--hash=sha256:48e8f4d9e7e5878571ecf6f2b4e57634f93cd474cc8cfbd2376f2d11b396e30d \
--hash=sha256:4b65538ae22f6fed0ce4874efd317463a7489796a0939fa66824dd542125a192
# via
# jaraco-classes
# jaraco-functools
nh3==0.3.7 \
--hash=sha256:157ec1eb7a62f3d9a7badb8d82d89aa810e3e24e097eedfa481a25d0c8a99877 \
--hash=sha256:15f5fbf090f5c88d61c820e1fc1fceecb6520cca9fe85649c06b57ef9dc9ff62 \
--hash=sha256:18f4278ecd157d43cb35acd5aae9f35cfa79f546b4922bd86536adc0f6312102 \
--hash=sha256:19f288c938ec6eef1f5d2c6cab47838e71fef8097e1c1233802be5a6230ba086 \
--hash=sha256:4968fe8d2db97c6f047659bf46a449fd8ec377f44ebf3e0a1b96c0d3a333ae32 \
--hash=sha256:5ffdfcb9a686ffb12765376bcfb6b5b55728516d3c0ee317d29982381ded3df8 \
--hash=sha256:614dac4a4c36ad084e78447d16fe898dedd762e354a7ab9cda2984e82f67883d \
--hash=sha256:618e3059caf41ccdf5dcccb3fa9df4cf6e4efe23d1382a8bbfca272a8a4f8bfc \
--hash=sha256:6698a822132beedab80f131c08d8d0ac5a178ddeb488d02ca4b67716ecfac7af \
--hash=sha256:6c3aa50eb26e9228238271db9f983cbc3b006dfbfeca2d4dc34c33ddc6ac5ea5 \
--hash=sha256:6e4280115d44c3b278eef712a86748c1a723105cd79feec46952383117ab4e59 \
--hash=sha256:70f5ac8626e899a4bab0ef74ca2f5bd602f49c7b739e6e5026b4afc6d63dac42 \
--hash=sha256:71860d01c16f4d8c72e334e0674beb2b0899dbd0bf760de18932ef4390303848 \
--hash=sha256:808def0c8c07843e6e50dc84f532457bfa2cfd17417b219a5d9e7c773709331a \
--hash=sha256:874b7d67a067bd29a59223f6270fc30da4edd8e6d87fd219fc93bcbaa662c946 \
--hash=sha256:91a4dab4e94d9fc54b9f67b1adfb23e81fab7ab43f33c3b8c97be9aa38f789ba \
--hash=sha256:94fd6e59553fbb9ffd8ba71bbd5a54e3126ba01799a097ae30d5341d750bc6ac \
--hash=sha256:9b7279d43323a25225df23576af6594a16693f61431170848b8b2ac21ad4f174 \
--hash=sha256:bc42bb1193c1e28a1e74c2cabaca178e118a7103e8832699fef8a2b3e2496493 \
--hash=sha256:be53a4825585f701955cb9baf49f478f56eb81e20294329fe4bc689dd5dd81fa \
--hash=sha256:d56e76bd3cadb09b6b0cef364850811663734b348a25f5f587a2819c495367bd \
--hash=sha256:de2b2aab32ea303405debefdcfc58043d3e635fa3f67b9eb140d2b0e0c0d2563 \
--hash=sha256:e8fd1ab205258b29254f72db377d99e2c96aa7653ef3b015ccab0420b094b506 \
--hash=sha256:eae64328e46a25785535afcb6885b6f182ecaf5ee8c88f8c075422db8aacc65b \
--hash=sha256:f04b7d333b27f13ca439da3cf1c75c2fba34f104969f6ce4ac8e7079699c2f4a \
--hash=sha256:f266d3f1b3647449923a8e406524632220dd5d8b647078dfe45b885d33d10479 \
--hash=sha256:fd4a70efb45d5372174f718878eb7a35c12677626a63b2f103b23b833457dcac
# via readme-renderer
packaging==26.3 \
--hash=sha256:94edc256424af38762eb31306eed28beb9f0efc50a8837492c9d6fd6004aed79 \
--hash=sha256:d7193f7c8e4e93f444fde0262bf90af30e16fa0ad0ad44cb553c87339b23cd1c
# via twine
pycparser==3.0 \
--hash=sha256:600f49d217304a5902ac3c37e1281c9fe94e4d0489de643a9504c5cdfdfc6b29 \
--hash=sha256:b727414169a36b7d524c1c3e31839a521725078d7b2ff038656844266160a992
# via cffi
pygments==2.21.0 \
--hash=sha256:2363c69b61c4a97c838da3b130dcd6468f4848992b21a82f2a63ec34377137d9 \
--hash=sha256:610ca751c9bc2492b38eb9a38a7fbc93edbbb2d7182edaf34e66ae493dee5c8c
# via
# readme-renderer
# rich
readme-renderer==46.0 \
--hash=sha256:af3e964914f6310a33ff67b72a4bdd940bed8d7c3bdecd2d14f40edf284bfe90 \
--hash=sha256:d0dae1f74bb273b534770cb4cccb6bb78735540afdb03c2146f4e19dcd412560
# via twine
requests==2.34.2 \
--hash=sha256:2a0d60c172f83ac6ab31e4554906c0f3b3588d37b5cb939b1c061f4907e278e0 \
--hash=sha256:f288924cae4e29463698d6d60bc6a4da69c89185ad1e0bcc4104f584e960b9ed
# via
# requests-toolbelt
# twine
requests-toolbelt==1.0.0 \
--hash=sha256:7681a0a3d047012b5bdc0ee37d7f8f07ebe76ab08caeccfc3921ce23c88d5bc6 \
--hash=sha256:cccfdd665f0a24fcf4726e690f65639d272bb0637b9b92dfd91a5568ccf6bd06
# via twine
rfc3986==2.0.0 \
--hash=sha256:50b1502b60e289cb37883f3dfd34532b8873c7de9f49bb546641ce9cbd256ebd \
--hash=sha256:97aacf9dbd4bfd829baad6e6309fa6573aaf1be3f6fa735c8ab05e46cecb261c
# via twine
rich==15.0.0 \
--hash=sha256:33bd4ef74232fb73fe9279a257718407f169c09b78a87ad3d296f548e27de0bb \
--hash=sha256:edd07a4824c6b40189fb7ac9bc4c52536e9780fbbfbddf6f1e2502c31b068c36
# via twine
secretstorage==3.5.0 \
--hash=sha256:0ce65888c0725fcb2c5bc0fdb8e5438eece02c523557ea40ce0703c266248137 \
--hash=sha256:f04b8e4689cbce351744d5537bf6b1329c6fc68f91fa666f60a380edddcd11be
# via keyring
twine==7.0.0 \
--hash=sha256:85cdb29c518efef867360ae4acd4b0dfd61c8654a22fca08e6f8539f05022177 \
--hash=sha256:b854164df26db268af05f49aa5c0344b10e27a494343ff05b1e0bad3b135f5a7
# via -r .github/requirements/twine.in
urllib3==2.7.0 \
--hash=sha256:231e0ec3b63ceb14667c67be60f2f2c40a518cb38b03af60abc813da26505f4c \
--hash=sha256:9fb4c81ebbb1ce9531cce37674bbc6f1360472bc18ca9a553ede278ef7276897
# via
# id
# requests
# twine
zipp==4.1.0 \
--hash=sha256:25ad4e16390cd314347dd8f1de67a2ac538ae658ed4ab9db16029c07c188e97f \
--hash=sha256:4cb57381f544315db7688e976e922a2b18cdb513d21cc194eb42232ba2a3e602
# via importlib-metadata
-1
View File
@@ -1 +0,0 @@
uv==0.12.1
-23
View File
@@ -1,23 +0,0 @@
# This file was autogenerated by uv via the following command:
# uv pip compile .github/requirements/uv-tool.in --generate-hashes --python-version 3.11 --python-platform linux -o .github/requirements/uv-tool.txt
uv==0.12.1 \
--hash=sha256:04290ea4001dca31ac8a8324113a4930dccad69ce35dbf6eaae307d54880890d \
--hash=sha256:153ec0959a15397514438aefc1d7cd04235f335dd6bb53ea0f9e6e82c5a49f03 \
--hash=sha256:173ee216f17d89fc39f65339d311a53584fc7de4918d27c0f3c7edafabc6b54d \
--hash=sha256:1de49d9b04438f1ad2f41a1441dbbe19e230b94fca56d632818cfaed69e03bfc \
--hash=sha256:1e8fd95fe98768e29436ad57f9ef7b68dc294b7b9862ef63396af8b15ab85e6c \
--hash=sha256:27211df9b277f440dea438a4e525ba40250fb721ad39b8927eefc2d91f9aea15 \
--hash=sha256:29399e1e73b67ed24abe82bc971aa4eb8419c4de804784290f39cf681f0b51ce \
--hash=sha256:2e9b0b86e180abc5968b979c6e25203b32e85969abb5083ee1e8b88a5aa98a76 \
--hash=sha256:3bd5db002adc763aa8d277f5b44f8d6e3fd82d20f2e51225b0bbdae1badc7259 \
--hash=sha256:41b8fc2335f682312a1ca39a7b4abfd6af800992065c663582ca3e4d51cf9258 \
--hash=sha256:5bd04849dd5346517cc4e57b4b3aa0b01c67c423878260c04f5893a038fe25b6 \
--hash=sha256:6f7e72543264d2420ebb2ddc84696a751af2d6c5910046b7666589118f47292b \
--hash=sha256:71f86410264c69a3e8acd18171897dd8ab1a13350cf40f718e4def5db2b724be \
--hash=sha256:76d87de420213ca92fa403e87023c4c7c6956c6726c6b96d91c42cfe620173a3 \
--hash=sha256:9331dda0dc4990512c232f86e1d3a7b83c13f459777fcc2bd46030911b40eaaa \
--hash=sha256:b255ac23958e45f39f9c7a4cd65890df5ef46f539a3b14de03bd296bbba9cb60 \
--hash=sha256:bd02f2da212e6a983115dc64a6fc94e9256c2d60e056d6b669de0a6025aaec05 \
--hash=sha256:e35e0030480a8c3bf8ecd87ae4a6f6a224009e15e96a6fbb3634ac11ab75d582 \
--hash=sha256:ead7ad064f291a5df358c3ffa8ffab347a32bd5a75a6a068ca22254c2539a829
# via -r .github/requirements/uv-tool.in
-76
View File
@@ -1,76 +0,0 @@
"""Drop checkov-suppressed results from its SARIF output before upload.
checkov's SARIF exporter includes every evaluated check as an ordinary
result, including ones it internally marked SKIPPED via an inline
`# checkov:skip=` comment or a `checkov.io/skipN` resource annotation - it
never uses SARIF's `suppressions` field, and never drops them. checkov's
JSON output *does* correctly record which checks were skipped, so this
cross-references the two: any SARIF result whose (check_id, file) pair
appears in the JSON's skipped_checks is removed before GitHub ever sees it.
Without this, every already-suppressed finding reopens as a brand new code
scanning alert on every run, forever (see #6035/#6036, #6112-6115,
#6128-6131 for the pattern this was chasing before this script existed).
Usage: filter_checkov_skipped.py <json_path> <sarif_in_path> <sarif_out_path>
"""
import json
import sys
def path_suffix(path: str, segments: int = 2) -> str:
"""Last N path segments, normalized to forward slashes, lowercased.
checkov's JSON file_path and SARIF artifactLocation.uri are relative to
different roots (the scanned directory vs. a temp helm-render dir), so
they can't be compared directly - but the last couple of segments
(e.g. "templates/service.yaml") are stable across both and specific
enough in practice to avoid cross-file collisions.
"""
normalized = path.replace("\\", "/").strip("/")
return "/".join(normalized.split("/")[-segments:]).lower()
def main() -> None:
json_path, sarif_in_path, sarif_out_path = sys.argv[1:4]
with open(json_path, encoding="utf-8") as f:
checkov_json = json.load(f)
if isinstance(checkov_json, dict):
checkov_json = [checkov_json]
skipped = set()
for block in checkov_json:
for check in block.get("results", {}).get("skipped_checks", []):
skipped.add((check["check_id"], path_suffix(check["file_path"])))
with open(sarif_in_path, encoding="utf-8") as f:
sarif = json.load(f)
removed = 0
for run in sarif.get("runs", []):
kept = []
for result in run.get("results", []):
rule_id = result.get("ruleId")
locations = result.get("locations") or [{}]
uri = (
locations[0]
.get("physicalLocation", {})
.get("artifactLocation", {})
.get("uri", "")
)
if (rule_id, path_suffix(uri)) in skipped:
removed += 1
continue
kept.append(result)
run["results"] = kept
with open(sarif_out_path, "w", encoding="utf-8") as f:
json.dump(sarif, f)
print(f"Removed {removed} checkov-suppressed result(s) from the SARIF before upload.")
if __name__ == "__main__":
main()
+5 -31
View File
@@ -28,37 +28,11 @@ jobs:
BENCHMARK_REAL_LIBS: "1"
run: |
pip install -r .github/requirements/bootstrap.txt --require-hashes
# --no-deps + a hash-pinned install of the same base dependency set
# (rather than a bare `pip install -e .`) so every fetched package
# is hash-verified (Scorecard Pinned-Dependencies); the local
# editable install itself has nothing to hash.
#
# --no-deps only skips *runtime* dependency resolution - `-e .`
# still does a PEP 517 build, which by default creates an isolated
# build env and fetches [build-system] requires (setuptools,
# wheel) completely outside any hash checking. Install
# pep517-build.txt (pins that exact build-system.requires) first
# and pass --no-build-isolation so pip reuses those hash-verified
# copies instead of fetching its own.
pip install -r .github/requirements/pep517-build.txt --require-hashes
pip install --no-deps --no-build-isolation -e .
pip install -r .github/requirements/base-deps.txt --require-hashes
# NOTE: benchmarks/ does not currently exist in this repo (neither
# requirements.txt nor benchmarks_runner.py below), so this job
# already fails on any real invocation - pre-existing, unrelated to
# this pinning change. The `pip install -r benchmarks/requirements.txt`
# step that used to be here is dropped rather than fixed: there's
# nothing to hash-pin without knowing what that file should
# contain, and an unpinned install here would just re-trip
# Scorecard's Pinned-Dependencies check for no real benefit, since
# the job can't run to completion regardless.
#
# `python -m spacy download en_core_web_sm` fetches an unpinned,
# unhashed wheel from spacy-models' GitHub releases - replaced with
# a hash-pinned direct-URL install of the same 3.8.0 model (matches
# the spacy==3.8.15 pinned in base-deps.txt) via benchmark-extra.txt.
pip install -r .github/requirements/benchmark-extra.txt --require-hashes
python -m pip install --upgrade pip
pip install -e .
pip install -r benchmarks/requirements.txt
python -m spacy download en_core_web_sm
pip install rdflib neo4j faiss-cpu torch pyarrow pdfplumber python-pptx openpyxl lxml python-docx beautifulsoup4 chardet langdetect
- name: Execute Benchmarks (Real Mode)
env:
+7 -30
View File
@@ -52,39 +52,16 @@ jobs:
# environment is installed. The Explorer extra supplies the
# production API dependencies without importing optional vector
# providers such as Pinecone during test collection.
#
# --no-deps + a separate hash-pinned install (rather than the old
# `pip install -e ".[explorer]" pytest==9.1.1`) so every fetched
# package is hash-verified (Scorecard Pinned-Dependencies); the
# local editable install itself has nothing to hash.
# .github/requirements/explorer-extra-py311.txt is
# `uv pip compile pyproject.toml --extra explorer --python-version 3.11 --constraint requirements-ci.txt --generate-hashes`
# - regenerate it the same way if pyproject.toml's base/explorer
# deps change. Resolved specifically for this job's python 3.11
# (see the Dockerfile's explorer-extra-py313.txt for why this
# can't be shared with python 3.13: audioread needs extra
# standard-aifc/standard-sunau hashes only on 3.13+).
#
# --no-deps only skips *runtime* dependency resolution - `-e .`
# still does a PEP 517 build, which by default creates an isolated
# build env and fetches [build-system] requires (setuptools,
# wheel) completely outside any hash checking. Install
# pep517-build.txt (pins that exact build-system.requires) first
# and pass --no-build-isolation so pip reuses those hash-verified
# copies instead of fetching its own.
pip install -r .github/requirements/pep517-build.txt --require-hashes
pip install --no-deps --no-build-isolation -e .
pip install -r .github/requirements/explorer-extra-py311.txt --require-hashes
pip install -r .github/requirements/pytest-tool.txt --require-hashes
pip install -e ".[explorer]" pytest==9.1.1
- name: Test deterministic Explorer backend path
run: |
pytest -q tests/explorer/test_explorer_deterministic_rendering_e2e.py
- name: Install pinned Python dependencies
run: |
pip install -r requirements-ci.txt --require-hashes
pip install -r requirements-ci.txt
- name: Verify requirements-ci.txt is up to date
run: |
pip install -r .github/requirements/uv-tool.txt --require-hashes
pip install uv==0.12.1
# Re-resolve with the committed file as a constraint: upstream package
# releases must NOT fail CI (deps only change when pyproject.toml
# changes intentionally). Compare only version lines (pkg==ver),
@@ -95,10 +72,10 @@ jobs:
diff \
<(grep -E '^[a-zA-Z0-9._-]+==' requirements-ci.txt | sed 's/ \\$//') \
<(grep -E '^[a-zA-Z0-9._-]+==' /tmp/requirements-ci-check.txt)
# build is a dev-time dependency; wheel is build-time only (neither is
# in requirements-ci.txt) — install the same pinned versions
# [build-system] declares so --no-isolation works below.
- run: pip install -r .github/requirements/build-tools.txt --require-hashes
- run: pip install build
# wheel is build-time only (not in requirements-ci.txt) — install the
# same pinned version [build-system] declares so --no-isolation works.
- run: pip install wheel==0.48.0
- name: Build package (no isolation — pinned deps)
run: python -m build --no-isolation
- name: Verify Explorer frontend is packaged
+2 -4
View File
@@ -10,15 +10,13 @@ on:
permissions:
contents: read
security-events: write
actions: read
jobs:
analyze:
name: Analyze Python
runs-on: ubuntu-latest
permissions:
contents: read
security-events: write # for github/codeql-action/upload-sarif below
actions: read # for github/codeql-action/init's CodeQL bundle cache lookup
steps:
- name: Checkout repository
-75
View File
@@ -1,75 +0,0 @@
name: Container Security Scan
on:
push:
branches: [main]
# Mirrors .dockerignore's opt-in list exactly - anything not listed there
# can't reach the build context, so it can't change the built image.
paths:
- 'Dockerfile'
- '.dockerignore'
- 'pyproject.toml'
- 'README.md'
- 'LICENSE'
- 'MANIFEST.in'
- '.github/requirements/explorer-extra-py313.txt'
- '.github/requirements/pep517-build.txt'
- 'semantica/**'
- 'integrations/**'
- 'explorer/**'
- '.github/workflows/container-scan.yml'
schedule:
- cron: '30 2 * * 1' # weekly, catches new CVEs published against the base image between pushes
workflow_dispatch:
permissions:
contents: read
jobs:
scan:
runs-on: ubuntu-latest
permissions:
contents: read
security-events: write # for github/codeql-action/upload-sarif below
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- name: Build image
run: docker build -t semantica:scan .
# Run Trivy as a digest-pinned image rather than the aquasecurity/trivy-action
# marketplace wrapper: the aquasecurity GitHub org has an IP allow list on its
# API that 403s verify-action-pins.sh's live tag->SHA check from Actions-runner
# IPs, and this repo already treats Trivy's action pin as a known past target
# for tag-repointing (see the LiteLLM/Trivy 2026 incident note above). Pulling
# by sha256 digest from Docker Hub is immutable and verifiable independently of
# GitHub's API, so it sidesteps both problems at once instead of carving a skip
# exception into the pin verifier for an org already flagged as higher-risk.
#
# Report-only for now: this is Trivy's first run against this image, so we
# don't yet know the CRITICAL/HIGH baseline. Findings still land in the
# Security tab either way. Once triaged, add `--exit-code 1` (like
# Safety/Bandit-HIGH in security-scan.yml) to make it a hard gate.
- name: Scan image for vulnerabilities (Trivy)
run: |
docker run --rm \
-v /var/run/docker.sock:/var/run/docker.sock \
-v "$PWD:/output" \
aquasec/trivy@sha256:62b1e65e8869bc4b4c6aa4fa2b21595256c7c2f6018a9d9ad61caf87187c1969 \
image --format sarif --output /output/trivy-results.sarif \
--severity CRITICAL,HIGH --ignore-unfixed semantica:scan
- name: Upload Trivy SARIF
if: always()
uses: github/codeql-action/upload-sarif@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4
with:
sarif_file: trivy-results.sarif
category: trivy-container
- name: Generate SBOM (Syft)
if: always()
uses: anchore/sbom-action@3ad7283483fc7af8ff2b4ea19663c2d5ca935e26 # v0.24.2
with:
image: semantica:scan
format: spdx-json
output-file: semantica-sbom.spdx.json
+5 -23
View File
@@ -28,14 +28,12 @@ on:
permissions:
contents: read
security-events: write
jobs:
MSDO:
# currently only windows-latest is supported
runs-on: windows-latest
permissions:
contents: read
security-events: write # for github/codeql-action/upload-sarif below
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
@@ -68,7 +66,7 @@ jobs:
python-version: "3.12"
- name: Install Checkov
run: pip install -r .github/requirements/checkov.txt --require-hashes
run: python -m pip install checkov==3.3.1
- name: Run Checkov
shell: pwsh
@@ -76,28 +74,12 @@ jobs:
PYTHONUTF8: "1"
run: |
New-Item -ItemType Directory -Force reports | Out-Null
checkov --directory . --framework kubernetes helm dockerfile github_actions secrets bicep arm --soft-fail --output sarif --output json --output-file-path reports
if (-not (Test-Path reports/results_sarif.sarif)) {
checkov --directory . --framework kubernetes helm dockerfile github_actions secrets bicep arm --soft-fail --output sarif --output-file-path reports/checkov.sarif
if (-not (Test-Path reports/checkov.sarif)) {
$sarif = Get-ChildItem -Path reports -Recurse -Filter *.sarif | Select-Object -First 1
if ($null -eq $sarif) { throw "Checkov did not produce a SARIF file" }
Copy-Item $sarif.FullName reports/results_sarif.sarif
Copy-Item $sarif.FullName reports/checkov.sarif
}
if (-not (Test-Path reports/results_json.json)) {
$json = Get-ChildItem -Path reports -Recurse -Filter *.json | Select-Object -First 1
if ($null -eq $json) { throw "Checkov did not produce a JSON file" }
Copy-Item $json.FullName reports/results_json.json
}
# checkov's SARIF exporter includes checks it internally marked SKIPPED
# (via the inline `# checkov:skip=` comments / `checkov.io/skipN`
# annotations already on the Helm chart) as ordinary un-suppressed
# results - it never uses SARIF's own `suppressions` field, so GitHub
# opens a fresh alert for the same already-suppressed finding on every
# single run (see #6035/#6036, #6112-6115, #6128-6131). checkov's JSON
# output does correctly record the skip, so cross-reference it here
# instead of re-dismissing the same alerts by hand forever.
- name: Filter checkov's own suppressed checks out of the SARIF
run: python .github/scripts/filter_checkov_skipped.py reports/results_json.json reports/results_sarif.sarif reports/checkov.sarif
- name: Upload Checkov results to Security tab
uses: github/codeql-action/upload-sarif@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4
+1 -1
View File
@@ -65,4 +65,4 @@ jobs:
- name: Deploy to GitHub Pages
id: deployment
uses: actions/deploy-pages@368f82528645a54fb793d4d04e342629a3f51346 # v5
uses: actions/deploy-pages@cd2ce8fcbc39b97be8ca5fce6e763baed58fa128 # v5
-59
View File
@@ -1,59 +0,0 @@
name: Install Matrix
permissions:
contents: read
on:
schedule:
- cron: '0 6 * * 1' # weekly, catches upstream dependency breakage between releases
workflow_run:
# The Release workflow publishes the GitHub release *before* it uploads to
# PyPI (see release.yml), so triggering on `release: published` would race
# the PyPI upload and could pass by silently installing the prior version.
# workflow_run fires only after the whole Release workflow - including the
# PyPI publish step - has finished.
workflows: ['Release']
types: [completed]
workflow_dispatch:
jobs:
verify-install:
if: github.event_name != 'workflow_run' || github.event.workflow_run.conclusion == 'success'
name: pip install semantica (${{ matrix.os }}, py${{ matrix.python-version }})
runs-on: ${{ matrix.os }}
strategy:
fail-fast: false
matrix:
os: [ubuntu-latest, macos-latest, windows-latest]
python-version: ['3.9', '3.10', '3.11', '3.12']
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- name: Pin expected version for release-triggered runs
id: expected-version
if: github.event_name == 'workflow_run'
shell: bash
env:
EXPECTED_TAG: ${{ github.event.workflow_run.head_branch }}
run: |
expected="${EXPECTED_TAG#v}"
if [ -z "$expected" ]; then
echo "::error::Could not determine a release tag from the triggering workflow run (head_branch was empty)."
exit 1
fi
echo "constraint===$expected" >> "$GITHUB_OUTPUT"
- id: setup-semantica
uses: ./.github/actions/setup-semantica
with:
python-version: ${{ matrix.python-version }}
cache: 'pip'
version: ${{ steps.expected-version.outputs.constraint }}
- name: Smoke test import
shell: bash
run: |
python -c "
import semantica
print('semantica', semantica.__version__, 'installed and importable')
"
+8 -26
View File
@@ -16,7 +16,7 @@ jobs:
cancel-in-progress: false
permissions:
contents: write # for the GitHub Release
id-token: write # for PyPI Trusted Publishing (OIDC), attestation signing, and Sigstore
id-token: write # for PyPI Trusted Publishing (OIDC) and attestation signing
attestations: write # for SLSA build provenance
# If you add another job to this workflow, give it its own explicit
# `permissions:` block rather than relying on the workflow-level default
@@ -39,11 +39,11 @@ jobs:
# Install the pinned dependency set (with hashes) so the sdist/wheel
# build runs against the same versions CI tests against.
- name: Install pinned build dependencies
run: pip install -r requirements-ci.txt --require-hashes
# build is a dev-time dependency; wheel is build-time only (neither is
# in requirements-ci.txt) — install the same pinned versions
# [build-system] declares so --no-isolation works below.
- run: pip install -r .github/requirements/build-tools.txt --require-hashes
run: pip install -r requirements-ci.txt
- run: pip install build
# wheel is build-time only (not in requirements-ci.txt) — install the
# same pinned version [build-system] declares so --no-isolation works.
- run: pip install wheel==0.48.0
- name: Build package (no isolation — pinned deps)
run: python -m build --no-isolation
- name: Verify Explorer frontend is packaged
@@ -63,29 +63,11 @@ jobs:
print("Explorer frontend is packaged")
PY
- name: Verify PyPI long-description will render
run: |
pip install -r .github/requirements/twine.txt --require-hashes
twine check dist/*
- name: Attest build provenance
uses: actions/attest-build-provenance@4d101475d8b20a2381f78447822ac1eab6504dd8 # v4
with:
subject-path: 'dist/*'
# attest-build-provenance publishes to the GH attestations API only, which
# OpenSSF Scorecard's Signed-Releases check does not inspect - it looks for
# signature files attached as release assets. Sign here too so
# `dist/*.sigstore.json` bundles ship alongside the wheel/sdist on the
# GitHub Release itself.
- name: Sign artifacts with Sigstore
uses: sigstore/gh-action-sigstore-python@790bc6befb9d733738f18d8f895854b453640ec9 # v3.5.0
- uses: softprops/action-gh-release@3d0d9888cb7fd7b750713d6e236d1fcb99157228 # v3
with:
inputs: |
dist/*.whl
dist/*.tar.gz
- uses: softprops/action-gh-release@efb35369e0ad2afab669f228072c1b0d510eae64 # v3.0.3
with:
files: |
dist/*.whl
dist/*.tar.gz
dist/*.sigstore.json
files: dist/*
- uses: pypa/gh-action-pypi-publish@dc37677b2e1c63e2034f94d8a5b11f265b73ba33 # release/v1
-45
View File
@@ -1,45 +0,0 @@
name: Scorecard supply-chain security
permissions: read-all
on:
branch_protection_rule:
schedule:
- cron: '30 1 * * 6' # weekly
push:
branches: [main]
jobs:
analysis:
name: Scorecard analysis
runs-on: ubuntu-latest
permissions:
security-events: write # to upload SARIF results
id-token: write # to publish results and get a badge
contents: read
actions: read # to detect GitHub Actions workflows
steps:
- name: Checkout code
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
with:
persist-credentials: false
- name: Run analysis
uses: ossf/scorecard-action@2d1146689b8cda280b9bc96326124645441f03bc # v2.4.4
with:
results_file: results.sarif
results_format: sarif
publish_results: true
- name: Upload artifact
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
with:
name: SARIF file
path: results.sarif
retention-days: 5
- name: Upload to code-scanning
uses: github/codeql-action/upload-sarif@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4
with:
sarif_file: results.sarif
+41 -168
View File
@@ -3,7 +3,6 @@ name: Security Scan
on:
schedule:
- cron: '30 1 * * 1,4' # Mon/Thu 7 AM IST
workflow_dispatch:
push:
branches: [main]
paths-ignore:
@@ -45,101 +44,46 @@ jobs:
- name: Install dependencies
run: |
pip install -r .github/requirements/bootstrap.txt --require-hashes
# Install the pinned dependency set FIRST so pip-audit scans
# Semantica's exact CI/release dependency tree (requirements-ci.txt
# is generated from pyproject.toml extras, so this covers the
# project's real deps).
pip install -r requirements-ci.txt --require-hashes
# Tooling AFTER the pinned set: installing it first would let the
# pinned requirements overwrite the tooling's own transitive deps.
pip install -r .github/requirements/pip-audit.txt --require-hashes
pip install -r .github/requirements/security-scan-tools.txt --require-hashes
python -m pip install --upgrade pip
# Install the pinned dependency set FIRST so Safety scans Semantica's
# exact CI/release dependency tree (requirements-ci.txt is generated
# from pyproject.toml extras, so this covers the project's real deps).
pip install -r requirements-ci.txt
# Tooling AFTER the pinned set: installing safety/bandit/semgrep/jq
# first lets the pinned requirements overwrite their transitive deps
# (e.g. rich), which breaks the safety CLI at runtime.
pip install safety bandit semgrep jq
- name: Run pip-audit (Package Vulnerabilities)
continue-on-error: true
- name: Run Safety Check (Package Vulnerabilities)
run: |
# Keep publishing reports and the PR comment even when the audit
# gate fails. The final gate below preserves the failure status.
echo 'AUDIT_SCAN_STATUS=failed' >> "$GITHUB_ENV"
# NOTE: Safety 3.x repurposed --output to select a console format
# (json/text/screen/...), not a file path. Writing JSON to a file
# now requires --save-json; the previous `--output safety-report.json`
# usage was silently invalid and never produced a report.
safety check --save-json safety-report.json || true
# Same dependency tree Safety used to scan, and the same tool and
# invocation already proven reliable in security.yml.
pip-audit -r requirements-ci.txt --format=json --output=pip-audit-report.json || true
# Guard 1: fail loudly if pip-audit exited before writing a report
# at all (network error, tool crash). Without this check a missing
# or empty file causes jq to fall back to "0", making a broken
# Guard 1: fail loudly if Safety exited before writing a report at all
# (network error, API auth failure, tool crash). Without this check a
# missing or empty file causes jq to fall back to "0", making a broken
# scanner indistinguishable from a clean scan.
if [ ! -s pip-audit-report.json ]; then
echo "::error::pip-audit produced no report (pip-audit-report.json is missing or empty). Treating as failure — check for network errors or pip-audit crashes in the logs above."
exit 1
fi
# Guard 2: fail closed when the report doesn't have the shape the
# checks below assume: a non-empty dependencies array, each entry
# either carrying an array-valued vulns field or being a dependency
# pip-audit couldn't resolve/audit, which it reports as
# {"name": ..., "skip_reason": ...} with no vulns field at all
# (see pip_audit._format.json.JsonFormat._format_dep). That's a
# normal, documented report shape, not a malformed one — treating
# it as invalid would fail the whole job over a single unauditable
# package, the same kind of scan-unrelated CI break this migration
# away from Safety was meant to fix.
if ! jq -e '
(.dependencies | type == "array" and length > 0)
and all(.dependencies[]; type == "object" and ((.vulns | type == "array") or (.skip_reason | type == "string")))
' pip-audit-report.json >/dev/null 2>&1; then
echo "::error::pip-audit report has an invalid dependency structure. Expected a non-empty dependencies array where every entry has either a vulns array or a skip_reason. Treating as failure."
if [ ! -s safety-report.json ]; then
echo "::error::Safety scan produced no report (safety-report.json is missing or empty). Treating as failure — check for network errors, API auth failures, or Safety crashes in the logs above."
exit 1
fi
echo "Checking for package vulnerabilities..."
# Guard 2 above already confirmed pip-audit-report.json is valid
# JSON with a well-shaped dependencies array, so this count is
# always a plain non-negative integer.
SKIPPED=$(jq '[.dependencies[] | select(has("skip_reason"))] | length' pip-audit-report.json)
if [ "$SKIPPED" -gt 0 ]; then
echo "⚠️ pip-audit could not audit $SKIPPED dependencies (see pip-audit-report.json for skip_reason):"
jq -r '.dependencies[] | select(has("skip_reason")) | " - \(.name): \(.skip_reason)"' pip-audit-report.json
fi
# No || echo "0" fallback: if jq fails (malformed JSON, missing key,
# vulnerabilities:null) VULNS will be empty or "null" so guard 2 below
# catches it rather than silently treating the broken report as zero.
VULNS=$(jq '.vulnerabilities | length' safety-report.json 2>/dev/null)
# Vulnerability IDs reviewed and accepted as non-actionable for this
# project. Empty for now: pip-audit's OSV-backed database doesn't
# currently carry either of the findings Safety used to flag here
# (cuda-toolkit CVE-2025-33228, torchvision CVE-2026-65918), so
# there's nothing to exclude. Left in place so a future finding can
# be added the same way without restructuring this step - see git
# history on this file for the reasoning behind past entries.
IGNORED_VULN_IDS=""
# Exported so the "Comment PR with Security Results" step below can
# apply the same exclusion list to the raw report - it reads
# pip-audit-report.json independently in JS, so without this the PR
# comment would show an accepted finding as live even though this
# gate correctly treats it as non-actionable.
echo "IGNORED_VULN_IDS=$IGNORED_VULN_IDS" >> "$GITHUB_ENV"
# No []? / || echo "0" fallback: if jq fails (malformed JSON) VULNS
# will be empty or "null" so Guard 3 below catches it rather than
# silently treating the broken report as zero.
# `.vulns // []` guards against skipped dependencies, which carry
# no vulns field at all (see the skip_reason handling above) -
# without the fallback, iterating `null[]` raises inside jq and
# this whole computation silently evaluates to empty.
VULNS=$(jq --arg ignored "$IGNORED_VULN_IDS" '
($ignored | split(",") | map(select(length > 0))) as $ignore_list
| [.dependencies[] | (.vulns // [])[] | select(.id as $id | ($ignore_list | index($id)) | not)]
| length
' pip-audit-report.json 2>/dev/null)
# Guard 3: ensure VULNS is a non-negative integer before the -gt
# Guard 2: ensure VULNS is a non-negative integer before the -gt
# comparison. "null" (missing/null key) or "" (jq parse failure) would
# cause bash's -gt to throw an arithmetic error and fall through to the
# success branch — the same silent-pass bug as a missing file.
if ! [[ "$VULNS" =~ ^[0-9]+$ ]]; then
echo "::error::pip-audit report exists but dependency vulnerabilities are missing or non-numeric (got: '${VULNS}'). The report may be malformed or contain an error-only JSON response. Treating as failure."
echo "::error::Safety report exists but 'vulnerabilities' is missing or non-numeric (got: '${VULNS}'). The report may be malformed or Safety may have written an error-only JSON. Treating as failure."
exit 1
fi
@@ -148,18 +92,12 @@ jobs:
echo "CI will fail to prevent merging of vulnerable dependencies"
echo ""
echo "Vulnerability details:"
jq --arg ignored "$IGNORED_VULN_IDS" -r '
($ignored | split(",") | map(select(length > 0))) as $ignore_list
| .dependencies[] as $dependency
| ($dependency.vulns // [])[] | select(.id as $id | ($ignore_list | index($id)) | not)
| "- \($dependency.name)==\($dependency.version): \(.id)"
' pip-audit-report.json || true
jq -r '.vulnerabilities[] | "- \(.package_name)==\(.analyzed_version): \(.vulnerability_id) (\(.CVE // "no CVE assigned"))"' safety-report.json || true
exit 1
else
echo "✅ No actionable security vulnerabilities found${IGNORED_VULN_IDS:+ (ignored: $IGNORED_VULN_IDS)}"
echo 'AUDIT_SCAN_STATUS=passed' >> "$GITHUB_ENV"
echo "✅ No security vulnerabilities found"
fi
- name: Run Bandit (Code Security Linter)
run: |
bandit -r semantica/ -f json -o bandit-report.json || true
@@ -197,18 +135,17 @@ jobs:
fi
- name: Upload Security Reports
if: always()
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
with:
name: security-reports
retention-days: 14
path: |
pip-audit-report.json
safety-report.json
bandit-report.json
semgrep-report.json
- name: Comment PR with Security Results
if: always() && github.event_name == 'pull_request'
if: github.event_name == 'pull_request'
uses: actions/github-script@3a2844b7e9c422d3c10d287c895573f7108da1b3 # v9
with:
script: |
@@ -231,12 +168,6 @@ jobs:
}
const items = parse(data);
if (items === null) {
return [
'### ' + title,
'⚠️ Invalid report structure in ' + reportPath + ' — check the job logs.',
].join('\n');
}
if (items.length === 0) {
return [`### ${title}`, `✅ No findings.`].join('\n');
}
@@ -253,64 +184,14 @@ jobs:
return lines.join('\n');
}
// Mirrors the shell step's own IGNORED_VULN_IDS (passed through
// $GITHUB_ENV) so an accepted, non-actionable CVE that the CI
// gate already excluded doesn't reappear here as a live finding -
// this reads the same raw, unfiltered pip-audit-report.json.
const ignoredVulnIds = (process.env.IGNORED_VULN_IDS || '')
.split(',')
.map((id) => id.trim())
.filter(Boolean);
// A dependency pip-audit couldn't resolve/audit is reported as
// {"name": ..., "skip_reason": ...} with no vulns field at all
// (see pip_audit._format.json.JsonFormat._format_dep) - that's a
// normal report shape, not a malformed one, so it must not be
// treated as an invalid dependency below.
const isSkipped = (dependency) => typeof dependency.skip_reason === 'string';
let skippedDeps = [];
try {
const auditData = JSON.parse(fs.readFileSync('pip-audit-report.json', 'utf8'));
skippedDeps = (auditData.dependencies || []).filter(
(dependency) => dependency && typeof dependency === 'object' && isSkipped(dependency)
);
} catch (e) {
// Unreadable/unparseable report - renderSection's own
// report-missing branch below surfaces this.
}
const pipAuditSection = renderSection(
'pip-audit — dependency vulnerabilities',
'pip-audit-report.json',
(data) => {
if (
!Array.isArray(data.dependencies) ||
data.dependencies.length === 0 ||
data.dependencies.some(
(dependency) =>
!dependency ||
typeof dependency !== 'object' ||
(!Array.isArray(dependency.vulns) && !isSkipped(dependency))
)
) {
return null;
}
return data.dependencies.flatMap((dependency) =>
(dependency.vulns || [])
.filter((vulnerability) => !ignoredVulnIds.includes(vulnerability.id))
.map(
(vulnerability) => `- \`${dependency.name}==${dependency.version}\`: ${vulnerability.id}` +
(vulnerability.fix_versions?.length ? ` (fixed by ${vulnerability.fix_versions.join(', ')})` : '')
)
);
}
) + (ignoredVulnIds.length
? `\n\n_Excluded as accepted, non-actionable findings: ${ignoredVulnIds.join(', ')} — see the workflow file's inline comments for why._`
: '') + (skippedDeps.length
? `\n\n_Could not be audited: ${skippedDeps.map((d) => `\`${d.name}\` (${d.skip_reason})`).join(', ')}_`
: '');
const safetySection = renderSection(
'Safety — dependency vulnerabilities',
'safety-report.json',
(data) => (data.vulnerabilities || []).map(
(v) => `- \`${v.package_name}==${v.analyzed_version}\`: ${v.vulnerability_id}` +
(v.CVE ? ` (${v.CVE})` : '') + ` — ${v.advisory || 'no advisory text'}`
)
);
const banditSection = renderSection(
'Bandit — HIGH-severity code issues',
@@ -331,7 +212,7 @@ jobs:
const comment = [
'# 🔒 Security Scan Results',
'',
pipAuditSection,
safetySection,
'',
banditSection,
'',
@@ -341,7 +222,7 @@ jobs:
'',
'*This security scan runs automatically on source-code PRs and bi-weekly (skipped for doc/markdown-only changes).*',
'',
'📊 **Security Policy**: CI fails on pip-audit vulnerabilities and Bandit HIGH-severity findings. Semgrep findings above are informational and do not block merge.',
'📊 **Security Policy**: CI fails on Safety vulnerabilities and Bandit HIGH-severity findings. Semgrep findings above are informational and do not block merge.',
].join('\n');
try {
@@ -356,11 +237,3 @@ jobs:
console.log('⚠️ Could not post security comment:', error.message);
console.log('📋 Security scan results saved to artifacts');
}
- name: Enforce Audit Gate
if: always()
run: |
if [ "${AUDIT_SCAN_STATUS:-failed}" != "passed" ]; then
echo "::error::pip-audit scan failed. See the pip-audit output and uploaded reports above."
exit 1
fi
+42
View File
@@ -0,0 +1,42 @@
name: Security
on:
schedule:
- cron: '0 0 * * 1'
workflow_dispatch:
pull_request:
branches: [main]
paths:
- 'pyproject.toml'
- 'requirements-ci.txt'
- '.github/workflows/security.yml'
permissions:
contents: read
jobs:
audit:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7
with:
python-version: '3.11'
# Upgrade first: actions/setup-python's baked-in setuptools has been
# behind known-vulnerable floors before (e.g. PYSEC-2026-3447 /
# setuptools 75.1.0), so don't trust the preinstalled one.
- run: python -m pip install --upgrade pip setuptools
# Audit the pinned dependency set (requirements-ci.txt is compiled from
# pyproject.toml with --extra all — the same coverage as the [all]
# extra, minus the Linux-only gpu set — so this keeps scan parity with
# CI/release builds without a time-dependent resolution). This is the
# fix for PYSEC-2024-38 (#869): the bare-env job never had fastapi or
# python-multipart installed to look at.
- run: pip install -r requirements-ci.txt
# PR runs gate on findings, since they're scoped to actual
# pyproject.toml changes under review. The schedule/workflow_dispatch
# runs stay non-blocking until a full pass over pre-existing findings
# across the whole [all] tree has been done.
- run: pip install pip-audit
- run: pip-audit -r requirements-ci.txt
continue-on-error: ${{ github.event_name != 'pull_request' }}
BIN
View File
Binary file not shown.
-44
View File
@@ -9,37 +9,6 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## [Unreleased]
### Added
- **Salesforce ingestor** (#1240) by @Sameer6305
- New `SalesforceConnector` / `SalesforceData` / `SalesforceIngestor` (`semantica.ingest`, lazy export), following the same Connector + Data + Ingestor pattern already used for Snowflake/Databricks/SAP
- Auth covers both landscapes Salesforce actually uses: username + password + security token (SOAP login), session_id + instance_url (reusing an existing session), and username + consumer_key + private key (JWT Bearer); production and sandbox are selected via `domain`, and credentials can come from environment variables. Credential material is never intentionally written to logs, exceptions, or `repr()`
- `ingest_sobject()`, `ingest_query()`, `list_sobjects()`, `get_sobject_schema()`, `export_as_documents()` against standard sObjects, custom objects (`__c`), custom metadata (`__mdt`), platform events (`__e`), namespaced objects, and relationship-field traversal (e.g. `Owner.Name`); pagination follows `nextRecordsUrl`/`query_more()` and stops once a caller's `limit` is satisfied
- New `pip install semantica[db-salesforce]` extra (`simple-salesforce>=1.12.0`)
- New `tests/test_salesforce_ingestor.py`
- Docs: `docs/integrations/salesforce.md`
- **`ErasureCoordinator` completes the erasure workflow `purge_node()` only starts — the graph node was removed while the same content survived verbatim in `AgentMemory` and as an embedding** (closes #1018) by @pravit-amp
- New `semantica/context/erasure.py`, exporting `ErasureCoordinator` and `ErasureReceipt` from `semantica.context`. `purge_node()`/`purge_edge()` (#957) are graph-scope by design and their changelog entry documents this gap explicitly; the changelog also names GDPR Article 17 as the motivation, and an Article 17 erasure that removes the node while the content stays retrievable by similarity search is not an erasure — it is worse than not offering one, because `purge_node()` returns `True` and writes a tombstone attesting the content is gone
- The coordinator **composes** the existing public APIs — nothing in `context_graph.py` or `agent_memory.py` changes behaviorally, and `ContextGraph` keeps its documented graph-scope contract rather than acquiring references to `AgentMemory`/`vector_store` that would invert the dependency
- `erase_entity(entity_id, reason=..., at=..., vector_ids=...)` returns an `ErasureReceipt`; `erase_entities([...])` returns one receipt per entity, in order, so one entity's failure does not stop the rest
- **Honest partial reporting is the point.** Each store reports one of five statuses — `erased`, `not_found`, `not_configured` (store never bound; normal), `unsupported` (store cannot delete at all; retrying will not help), `failed` — and `receipt.complete` is `False` when any store reports `unsupported`/`failed`, with `receipt.incomplete_stores` naming them. A receipt reading `graph: erased, memory: 14 erased, vectors: unsupported on faiss` is actionable; a bare `True` is a compliance liability
- **Erasure runs outward-in: vectors → memory → graph.** The graph tombstone is the durable attestation that an erasure happened, so writing it first would let a crash mid-cascade leave a record claiming more than occurred. Erasing the graph last means a partial failure leaves the node present and the receipt incomplete — recoverable and honest; the reverse is neither
- **Partial failure is a result, not an exception**: a store that raises is recorded as `failed` (with the exception type) and the remaining legs still run, rather than aborting into a half-erased state with no record of which half
- **The memory sweep cannot be silently truncated.** `find_by_entity(entity_id, limit=10)` returned `results[:limit]`, so the obvious hand-rolled cascade erases the first ten items and reports success — an erasure check computed from a page already truncated by the very `limit` it was called with. The coordinator sweeps in pages until dry (deleting as it goes, so the next page is the remainder) rather than passing one large number that is only correct until someone exceeds it, then **re-queries once after the sweep** and reports `failed` with the residual count if anything survived. It also stops rather than spinning if `batch_delete` reports no progress on a non-empty page. Note `find_by_entity` returns items keyed `memory_id`, not `id`
- **`unsupported` vector backends are detected by probing, not by calling and catching.** `faiss_store.py`, `milvus_store.py` and `weaviate_store.py` expose no delete at all (FAISS cannot remove from a flat index without a rebuild), while the `VectorStore` facade declares `delete_vectors()` for *every* backend and only raises `NotImplementedError` once called — so probing the facade alone cannot tell a deletable backend from a delete-less one, and the coordinator looks at the backend it wraps. Probing also keeps a missing method distinguishable from an `AttributeError` raised *inside* a working one, which is exactly where guessing wrong produces a false clean bill of health. `NotImplementedError` at call time is still caught and reported as `unsupported`; a store returning `False` is reported as `failed`
- Backends are reached under either supported name — `delete_vectors(ids)` (pinecone/qdrant) or `delete(ids)` (pgvector/sqlite-vec) — and the receipt records which was used
- `vector_store` defaults to `memory.vector_store` when a memory is supplied, stays overridable for deployments binding a store the memory does not own, and accepts `False` to disable the vector leg. Vectors owned by memory items are removed by the memory leg's own `delete_memory()` cascade; the explicit vector leg covers entity-keyed embeddings written by something other than `AgentMemory`
- The receipt's `erased_at` is normalized through `ContextGraph`'s own temporal normalizer, so the receipt and the tombstone written by the same erasure cannot disagree about when it happened; an unparseable `at` is rejected before any store is touched rather than half way through the cascade
- `purge_node()`'s docstring now points at the coordinator, so callers reading the graph-scope caveat find the thing that completes the workflow
- New `tests/context/test_erasure_coordinator.py`: 48 tests against **real** `ContextGraph`/`AgentMemory` instances rather than mocks — the bug lives in the interaction between them, so mocking it away would test nothing. Covers the 25-items-on-one-entity regression that fails against a naive single `find_by_entity()` call, all three vector-backend shapes (`delete_vectors`/`delete`/neither) plus the facade-over-delete-less-backend shape, residual/no-progress/no-identifier memory failures, partial failure continuing the cascade, idempotency, receipt serialization, and `at` normalization
- Full `tests/context/` suite: 738 passed
- **Fixed during review** (Qodo): `erase_entity()` resolved `erased_at` up front but passed the caller's original `at` down to `purge_node()`, so on the default `at=None` path the coordinator and the graph each took their own `now()` and the receipt attested to a different instant than the tombstone it points at — breaking the one invariant this module states most loudly. The resolved timestamp is now passed to the graph. The existing test passed only because it supplied an explicit `at`, which hides the drift; a regression test now covers the `at=None` path that callers actually use
- **Fixed during review** (Qodo): the vectors leg treated any return value other than the literal `False` as success, but no in-repo backend returns a bool — Qdrant returns `{"status": <UpdateStatus>}` and Pinecone `{"deleted": True}`, so every dict was read as a success and the backend's own account of the delete was discarded. Delete results are now interpreted by shape (bool, dict with explicit failure markers, `None` for a void method, anything else at face value) and the backend payload is recorded in the receipt as `backend_result`, stringified so the receipt stays JSON-serializable as the audit record it is meant to be. Bool markers are matched by identity so a `0` count is not read as `False`, and string markers match as substrings so an enum rendering as `"UpdateStatus.FAILED"` is not read as a success
- **Fixed during review** (Qodo): the constructor's "at least one store" guard used `not vector_store`, rejecting a valid store whose `__bool__`/`__len__` makes an empty instance falsey, and reporting `vector_store=None` in the error when an object had been passed; it now distinguishes `None` (absent) from `False` (deliberately disabled) from any other value (provided), and echoes what it actually received
- **Fixed during review** (Qodo): `at` annotations accepted only `str`/`datetime` while the shared `ContextGraph` normalizer they delegate to also takes epoch seconds; widened to `int`/`float` with the docstrings updated, so the coordinator no longer advertises less than the graph API it wraps
- **Known limitation, unchanged by this PR**: erasure still cannot be *completed* on FAISS/Milvus/Weaviate — `delete_vectors()` is declared on the `VectorStore` facade (`vector_store.py:786`) but not implemented across the backend set, under at least three different names. That is worth its own issue; the coordinator ships reporting `unsupported` and starts reporting `erased` for those backends once it is fixed, with no API change here
## [0.6.7] - 2026-08-28
### Added
@@ -159,19 +128,6 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- **Also fixed, on the JSON-LD paths**: the first fix covered the Turtle, N-Triples and RDF/XML serializers, and left both JSON-LD writers interpolating the entity's own text into `f"semantica:entity/{text}"` and the endpoints into `f"semantica:rel/{source}_{target}"`. Three consequences, all live in 0.6.5: an entity whose text contained a space produced an invalid IRI, and a JSON-LD parser dropped that node in full rather than reporting it, so the entity disappeared from the export; every relationship carrying `source`/`target` rather than `source_id`/`target_id` minted the identical `semantica:rel/_`, collapsing all of them onto one node whose types and endpoints merged; and the JSON-LD `@id` disagreed with the Turtle IRI for the same entity, so the two serializations of one knowledge graph were two different graphs. Both JSON-LD writers now use `mint_entity_iri`/`mint_relationship_iri`, and `JSONExporter.export_entities`/`export_relationships` declare the `semantica` prefix their `@context` was already writing `semantica:entities` against — without it a processor reads that as an IRI in the scheme `semantica`, which is the original #1101 defect on a third path
- `tests/export/test_jsonld_iri_minting.py` parses each export with a real JSON-LD processor and asserts the entity survives, the relationships stay distinct, no term expands into the `semantica` scheme, and the JSON-LD `@id` equals the Turtle IRI
- 236 export and ontology tests pass
- **`semantica.evals` runner gains per-metric objectives** (#1091)
- `evaluate()` now accepts `config={"<evaluator>": {"objective": {"direction": "maximize"|"minimize", "threshold": X}}}` to override the evaluator's default pass verdict with a threshold; `{"objective": {"expect": bool}}` expresses a Boolean expectation
- `minimize` requires a `threshold` — omitting it or setting it to `None` raises `ValueError`; `maximize` without a threshold is a no-op (the evaluator's own verdict stands); `expect` cannot be combined with `direction`/`threshold`; invalid config raises `ValueError` before any evaluator runs
- Error metrics are never affected by objectives (error wins over fail)
- Backward compatible: no `objective` key → existing behavior unchanged
- New tests in `tests/evals/test_runner.py::TestObjective`
- **`semantica.evals` is now a fully implemented evaluation module** (was a "Coming Soon" stub in the package layout)
- `evaluate(cases, evaluators, config=None, target_fn=None)` runner with per-case `pass`/`fail`/`error` status and an aggregate `pass_rate`, using a registry of named evaluators (`list_evaluators()`)
- 10 built-in evaluators: `exact_match`, `regex_match`, `numeric_range`, `temporal_range`, `length_range`, `keyword_check`, `levenshtein` (edit-distance similarity), `rouge` (in-house token F1, no new dependencies), `llm_as_judge` (lazy: caller-supplied `judge_fn`), and `decision_scores` (composite over `semantica.context.Decision`)
- `decision_scores` validates field-level (expected outcome, confidence bounds, non-empty maker/reasoning/scenario) and governance-level (provenance record presence; opt-in `PolicyEngine.check_compliance`) checks, coercing dict inputs via `Decision(**actual)` and never crashing on malformed input; an interface slot for causal-chain/embedding checks is reserved and raises `NotImplementedError` (V2)
- `__version__` is `0.1.0`, and the module ships a usage guide at `semantica/evals/usage.md` with worked import/run/interpret examples
- `semantica.evals` is reachable through the root package lazy module proxy (`semantica.evals`)
- 99 unit tests in `tests/evals/` covering every evaluator, registry errors, runner aggregation, decision coercion, and per-metric objectives; `python -m pytest tests/evals -q` → 99 passed
- **First-class CrewAI integration** (#988, closes #962) by @Shindevrp
- New `pip install semantica[crewai]` extra (`crewai>=0.80.0`) — crewai core provides `BaseTool`/`BaseKnowledgeSource`, so `crewai-tools` is intentionally not included, and the extra is intentionally **not** part of the `all` bundle: crewai hard-requires `chromadb~=1.1.0`, which is affected by the unpatched pre-auth code-injection CVE-2026-45829 (see `integrations/crewai/README.md`)
-20
View File
@@ -1,20 +0,0 @@
cff-version: 1.2.0
message: "If you use this software, please cite it as below."
title: "Semantica: Graph-Native Infrastructure for Context and Accountable AI Systems"
type: software
authors:
- name: "Semantica"
repository-code: "https://github.com/semantica-agi/semantica"
url: "https://getsemantica.ai"
license: MIT
version: 0.6.7
date-released: 2026-08-28
keywords:
- knowledge-graph
- context-graph
- ai-agents
- llm
- decision-intelligence
- provenance
- explainability
- graph-rag
+4 -38
View File
@@ -1,5 +1,5 @@
# syntax=docker/dockerfile:1
FROM node:26-alpine@sha256:2d984a15c9b54fd0aeb608b8e0d0d83529eb34d2966db27a1fb4f1edc3d298a3 AS frontend-builder
FROM node:26-alpine AS frontend-builder
WORKDIR /app
COPY explorer/package*.json ./explorer/
@@ -9,18 +9,7 @@ RUN npm ci
COPY explorer/ ./
RUN mkdir -p /app/semantica && npm run build
# CVE-2026-14456 (OpenSSL QUIC-server DoS, flagged against this base image's
# openssl/libssl3t64/openssl-provider-legacy): the Debian fix
# (3.5.7-1~deb13u2) is only in trixie-proposed-updates as of this writing,
# not yet promoted to trixie-security, so there's no package to pin here
# today. Deliberately NOT running `apt-get upgrade` to chase it - that
# breaks build reproducibility (terrascan AC_DOCKER_0052) and still
# wouldn't reach a proposed-updates-only package. Once Debian ships the fix
# and rebuilds this tag, the docker Dependabot ecosystem in
# .github/dependabot.yml opens a PR bumping the digest pin above. Also: this
# image only serves plain HTTP via uvicorn and never opens a QUIC listener,
# so the bug isn't reachable here regardless.
FROM python:3.13-slim@sha256:7ce4b6dfe35e55397b7cda544f8a13f191b7ae28dc5aad71fe664dbc9bc2623f AS runtime
FROM python:3.13-slim AS runtime
ENV PYTHONDONTWRITEBYTECODE=1 \
PYTHONUNBUFFERED=1 \
@@ -33,35 +22,12 @@ WORKDIR /app
RUN groupadd --system semantica \
&& useradd --system --gid semantica --home-dir /app --shell /usr/sbin/nologin semantica
COPY pyproject.toml README.md LICENSE MANIFEST.in \
.github/requirements/explorer-extra-py313.txt .github/requirements/pep517-build.txt ./
COPY pyproject.toml README.md LICENSE MANIFEST.in ./
COPY semantica/ ./semantica/
COPY integrations/ ./integrations/
COPY --from=frontend-builder /app/semantica/static ./semantica/static
# explorer-extra-py313.txt is `uv pip compile pyproject.toml --extra explorer
# --python-version 3.13 --constraint requirements-ci.txt --generate-hashes`
# (see ci.yml's explorer-extra-py311.txt for the CI counterpart, resolved
# for CI's python 3.11 instead - the two aren't interchangeable: audioread
# (via librosa) needs standard-aifc/standard-sunau only on python>=3.13,
# since aifc/sunau left stdlib there, so a 3.11-resolved lockfile is
# missing hashes pip needs on this image's actual 3.13 interpreter and
# --require-hashes fails outright rather than silently under-pinning).
# Every fetched package is hash-verified (Scorecard Pinned-Dependencies)
# and pinned to the same versions CI audited, e.g. msgpack==1.2.1 and
# setuptools==84.0.0 (which also replaces the base image's vulnerable
# 70.3.0, CVE-2025-47273 - nothing else in the tree pulls a newer copy).
# --no-deps on the local package itself: it's our own source tree, not a
# fetch, so there's nothing to hash-pin there - but `pip install .` still
# does a PEP 517 build, which by default creates an *isolated* build env
# and fetches [build-system] requires (setuptools, wheel) completely
# outside any hash checking. pep517-build.txt pins that exact
# build-system.requires; installing it first and passing
# --no-build-isolation makes pip reuse those hash-verified copies instead
# of fetching its own.
RUN pip install --no-cache-dir -r explorer-extra-py313.txt -r pep517-build.txt --require-hashes \
&& pip install --no-cache-dir --no-deps --no-build-isolation . \
&& rm -f explorer-extra-py313.txt pep517-build.txt \
RUN pip install --no-cache-dir ".[explorer]" \
&& chown -R semantica:semantica /app
USER semantica
-131
View File
@@ -1,131 +0,0 @@
# Growth & Distribution Playbook
North star: **10,000 developers who actually use Semantica in real projects**, not a raw PyPI download number. Downloads are a lagging indicator of distribution, not a target to optimize directly.
```
GitHub stars → Website visitors → PyPI installs → Weekly active users → Production deployments → Enterprise customers
```
The last two matter far more than the download count.
## Guardrails — do not do this
- No fake/looping CI jobs that repeatedly `pip install semantica` purely to inflate the graph. It's detectable, it produces zero real users, and it damages credibility with anyone doing diligence (investors, enterprise buyers, security reviewers).
- No package-splitting purely to multiply install counts — only split into `semantica-*` packages when there's a real architectural reason.
- No meaningless Docker pulls or notebook launches with no real content behind them.
- Every item below should get someone from "installed it" to "used it for something real." If a channel can't do that, it's not worth building.
## 30-day priority sprint
Ordered by leverage-to-effort ratio; do these first.
| # | Initiative | Target |
| - | ---------- | ------ |
| 1 | ✅ GitHub Actions example + reusable `setup-semantica` composite action + install-matrix badge | done |
| 2 | Google Colab notebooks | 10 |
| 3 | Docker images (RAG, Graph, Agent, API) | 4-5 |
| 4 | Hugging Face Spaces demos | 3-4 |
| 5 | LangChain integration + example | 1 |
| 6 | LlamaIndex integration + example | 1 |
| 7 | Vector/graph DB integrations (Qdrant, Weaviate, Neo4j) | 3 |
| 8 | MCP server + example | 1 (already have `mcp/` — package as a distributable example) |
| 9 | Production-quality starter repos (FastAPI, Streamlit, Gradio) | 3 |
| 10 | `awesome-rag` / `awesome-llm` / `awesome-knowledge-graph` list submissions | 3+ PRs |
Push everything through: GitHub → Discord (`sV34vps5hH`) → X (`@BuildSemantica`) → GitHub Discussions → Reddit → Hacker News → relevant newsletters.
## Full channel checklist
### CI/CD (highest-intent distribution — installs tied to real pipelines)
- [x] GitHub Actions example in `examples/ci/github-actions.yml`
- [x] Reusable composite GitHub Action — [`.github/actions/setup-semantica`](.github/actions/setup-semantica/action.yml), modeled on `actions/setup-python`; usable by any repo as `uses: semantica-agi/semantica/.github/actions/setup-semantica@main`
- [x] "pip install" status badge in the README, backed by [`.github/workflows/install-matrix.yml`](.github/workflows/install-matrix.yml) — verifies the *published* package installs cleanly on Ubuntu/macOS/Windows across Python 3.9-3.12, weekly + on every release
- [x] GitLab CI template — `examples/ci/gitlab-ci.yml`
- [x] CircleCI template — `examples/ci/circleci-config.yml`
- [ ] Jenkins, Azure DevOps, Bitbucket Pipelines, Buildkite, Travis CI equivalents
### Release pipeline hardening (already had Trusted Publishing/OIDC + SLSA attestation — this rounds it out to match top-tier OSS release practice)
- [x] `twine check` gate in `.github/workflows/release.yml` before publish — catches a broken PyPI long-description render before it goes live instead of after (a malformed README on the live PyPI page is a silent conversion killer)
- [x] `CITATION.cff` (see Academic & research below)
- [x] OpenSSF Scorecard (see Discoverability below)
- [ ] Considered and deliberately skipped: Release Drafter / auto-generated changelogs — this repo hand-curates `CHANGELOG.md` with far more detail (PR numbers, contributors, phase-1 limitations) than a bot would produce. Don't introduce this without checking with maintainers first.
- [ ] Renovate / Dependabot config templates that auto-bump the `semantica` version in downstream repos — real recurring CI runs on real adopters
- [ ] Nightly scheduled workflow template that tests a downstream project against `semantica@latest`
### Containers & dev environments
- [ ] Official Docker images: RAG, Graph, Agent, API, `+Postgres`, `+Neo4j`, `+Qdrant`
- [ ] `docker-compose` examples (repo already has `docker-compose.dev.yml` / `docker-compose.yml` as a base)
- [ ] `.devcontainer/devcontainer.json` for one-click "Reopen in Container"
- [ ] GitHub Codespaces-ready config
- [ ] Gitpod config
- [ ] "Use this template" GitHub repo button so new projects start with `semantica` in `requirements.txt`
### Notebooks & hosted demos
- [ ] 10-20 Google Colab notebooks (Graph RAG, agent memory, entity resolution, semantic search, document intelligence)
- [ ] Kaggle Notebooks/Kernels
- [ ] Binder / mybinder.org config for instant repo launch
- [ ] SageMaker Studio Lab / Databricks Community Edition / Paperspace Gradient examples
- [ ] Hugging Face Spaces (Streamlit/Gradio) demos with `semantica` in `requirements.txt`
- [ ] Public hosted playground (source on GitHub, install visible)
### Framework & data-store integrations
- [x] LangChain integration — `integrations/langchain/` (`SemanticaRetriever`, `SemanticaVectorStore`, `SemanticaKGTool`/`SemanticaDecisionTool`), `pip install semantica[langchain]`, shipped in 0.6.7
- [ ] LlamaIndex integration + example
- [ ] LangGraph example
- [ ] Neo4j integration/example (docs already list it as a supported graph store — turn into a runnable example repo)
- [ ] Vector DB examples: Qdrant, Weaviate, Milvus, Pinecone, Chroma, FAISS, pgvector, OpenSearch/Elasticsearch (FAISS/Pinecone/Weaviate/Qdrant/Milvus/PgVector already supported per `docs/community-projects.md` — package each as a standalone example)
- [ ] LLM provider quickstarts: OpenAI, Anthropic, Gemini, Groq, Ollama, HuggingFace, DeepSeek, LiteLLM (already-supported providers per docs — each gets its own copy-paste quickstart)
- [ ] CrewAI / Agno integration examples (already documented under `docs/integrations/`) — promote as standalone repos, not just docs pages
### Package managers & installers
- [ ] conda-forge feedstock
- [ ] Homebrew formula for the CLI
- [ ] Nix/nixpkgs packaging
- [ ] Chocolatey / Scoop (Windows)
- [ ] Document `uv add semantica` and `poetry add semantica` explicitly alongside `pip install`
### Downstream packages & CLI
- [ ] Genuinely useful `semantica-*` packages only where warranted (e.g. `semantica-rag`, `semantica-connectors`) — each pulls `semantica` as a real dependency
- [ ] Make sure `semantica init / ingest / index / query / serve` CLI flows are the default onboarding path in every tutorial
- [ ] VS Code extension wrapping the CLI (scaffold + run commands from the command palette)
- [ ] JetBrains plugin equivalent
### Templates & starters
- [ ] Cookiecutter templates: `cookiecutter-semantic-rag`, `cookiecutter-ai-agent`, `cookiecutter-enterprise-rag`
- [ ] Starter repos: FastAPI, Streamlit, Gradio, Next.js frontend + Semantica backend
- [ ] Cloud deploy templates: AWS, GCP, Azure, Modal, Railway, Render, Fly.io (repo already has `deploy/azure`, `deploy/gcp`, `deploy/fly`, `deploy/railway`, `deploy/render`, `deploy/kubernetes`, `deploy/helm` — link these prominently from the README/quickstart, they're already-built distribution surface)
- [ ] Terraform / Pulumi / Helm modules published to their respective registries
### Discoverability & curation
- [ ] Submit to `awesome-rag`, `awesome-llm`, `awesome-knowledge-graph`, `awesome-python`
- [ ] Pitch newsletters with engaged Python/AI audiences (Python Weekly, Import AI, TLDR AI, etc.)
- [x] PyPI trove classifiers/keywords and `project.urls` (Homepage/Docs/Repository/Changelog/Bug Tracker) — already complete in `pyproject.toml`
- [ ] Get listed on Papers With Code for any retrieval/graph-RAG benchmark work
- [x] [OpenSSF Scorecard](https://scorecard.dev/viewer/?uri=github.com/semantica-agi/semantica) badge + weekly workflow (`.github/workflows/scorecard.yml`) — a concrete trust signal security/procurement teams check before greenlighting adoption, which gates real (non-CI-bot) install growth at enterprises
### Academic & research
- [x] `CITATION.cff` at repo root — enables GitHub's native "Cite this repository" button, feeds Google Scholar/academic tooling; complements `docs/citation.md` (still needs a real Zenodo DOI to replace the `XXXXXXX` placeholder in both places once one is minted)
- [ ] arXiv paper if there's real architectural novelty to describe
- [ ] Zenodo DOI for citability (`docs/citation.md` already exists — make sure it points to a real DOI)
- [ ] Workshop/tutorial sessions at PyData/ODSC-style events with hands-on install steps
- [ ] University course material / bootcamp adoption outreach
### Content
- [ ] Reproducible benchmark repos (Graph RAG vs vector RAG, retrieval@k, enterprise-scale retrieval) with `pip install semantica && python benchmark.py`
- [ ] 20-30 real-world example applications (RAG, enterprise document intelligence, financial entity graphs, code knowledge graphs, research discovery, agent memory)
- [ ] Blog/tutorial posts on Dev.to, Medium, personal blogs — always with runnable code, not just prose
- [ ] Contribute integrations/PRs to other projects building RAG/agents/knowledge graphs — "I implemented Semantica support" beats "please use Semantica"
## Tracking
Don't just watch the raw PyPI number — use download analytics (e.g. PePy) to separate CI/bot traffic from real installs, and track the funnel above end-to-end where possible (stars → site visits → installs → weekly actives).
+23 -32
View File
@@ -18,15 +18,15 @@
> Ingest your enterprise data, extract what matters, build a Context Graph and knowledge graph (KG), and run graph analytics and causal reasoning over all of it, with full decision provenance baked in. Explainable, traceable, and trustworthy by design.
**Context Management &nbsp;·&nbsp; Knowledge Modeling &nbsp;·&nbsp; Deterministic Reasoning &nbsp;·&nbsp; Ontology Management &nbsp;·&nbsp; Decision Intelligence &nbsp;·&nbsp; End-to-End Traceability**
**Decision Intelligence &nbsp;·&nbsp; Context Management &nbsp;·&nbsp; Deterministic Reasoning &nbsp;·&nbsp; Ontology Management &nbsp;·&nbsp; Knowledge Modeling &nbsp;·&nbsp; End-to-End Traceability**
**Open Source &nbsp;·&nbsp; Governed &nbsp;·&nbsp; Zero Vendor Lock-In**
**Open Source &nbsp;·&nbsp; Self-Hostable &nbsp;·&nbsp; Auditable &nbsp;·&nbsp; Governed &nbsp;·&nbsp; Zero Vendor Lock-In**
**Polyglot Graph Storage &nbsp;·&nbsp; RDF & LPG Support &nbsp;·&nbsp; W3C Standards &nbsp;·&nbsp; Interoperable**
#### Built for High-Stakes, Regulated Domains
[![GitHub Stars](https://img.shields.io/github/stars/semantica-agi/semantica?style=flat-square&color=FFD700&logo=github&logoColor=white&label=Stars)](https://github.com/semantica-agi/semantica) [![GitHub Forks](https://img.shields.io/github/forks/semantica-agi/semantica?style=flat-square&color=6E40C9&logo=github&logoColor=white&label=Forks)](https://github.com/semantica-agi/semantica/network/members) [![Contributors](https://img.shields.io/github/contributors/semantica-agi/semantica?style=flat-square&color=2EA043&logo=github&logoColor=white)](https://github.com/semantica-agi/semantica/graphs/contributors) [![PyPI](https://img.shields.io/pypi/v/semantica.svg?style=flat-square&color=0066CC&logo=pypi&logoColor=white)](https://pypi.org/project/semantica/) [![Total Downloads](https://static.pepy.tech/badge/semantica?style=flat-square)](https://pepy.tech/project/semantica) [![Python 3.8+](https://img.shields.io/badge/python-3.8+-3776AB?style=flat-square&logo=python&logoColor=white)](https://www.python.org/) [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg?style=flat-square)](https://opensource.org/licenses/MIT) [![CI](https://img.shields.io/github/actions/workflow/status/semantica-agi/semantica/ci.yml?style=flat-square&label=CI)](https://github.com/semantica-agi/semantica/actions) [![Install Matrix](https://img.shields.io/github/actions/workflow/status/semantica-agi/semantica/install-matrix.yml?style=flat-square&label=pip%20install)](https://github.com/semantica-agi/semantica/actions/workflows/install-matrix.yml) [![OpenSSF Scorecard](https://api.scorecard.dev/projects/github.com/semantica-agi/semantica/badge?style=flat-square)](https://scorecard.dev/viewer/?uri=github.com/semantica-agi/semantica) [![Ask DeepWiki](https://deepwiki.com/badge.svg)](https://deepwiki.com/semantica-agi/semantica)
[![GitHub Stars](https://img.shields.io/github/stars/semantica-agi/semantica?style=flat-square&color=FFD700&logo=github&logoColor=white&label=Stars)](https://github.com/semantica-agi/semantica) [![GitHub Forks](https://img.shields.io/github/forks/semantica-agi/semantica?style=flat-square&color=6E40C9&logo=github&logoColor=white&label=Forks)](https://github.com/semantica-agi/semantica/network/members) [![Contributors](https://img.shields.io/github/contributors/semantica-agi/semantica?style=flat-square&color=2EA043&logo=github&logoColor=white)](https://github.com/semantica-agi/semantica/graphs/contributors) [![PyPI](https://img.shields.io/pypi/v/semantica.svg?style=flat-square&color=0066CC&logo=pypi&logoColor=white)](https://pypi.org/project/semantica/) [![Total Downloads](https://static.pepy.tech/badge/semantica?style=flat-square)](https://pepy.tech/project/semantica) [![Python 3.8+](https://img.shields.io/badge/python-3.8+-3776AB?style=flat-square&logo=python&logoColor=white)](https://www.python.org/) [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg?style=flat-square)](https://opensource.org/licenses/MIT) [![CI](https://img.shields.io/github/actions/workflow/status/semantica-agi/semantica/ci.yml?style=flat-square&label=CI)](https://github.com/semantica-agi/semantica/actions) [![Ask DeepWiki](https://deepwiki.com/badge.svg)](https://deepwiki.com/semantica-agi/semantica)
[![Website](https://img.shields.io/badge/Website-getsemantica.ai-000000?style=flat-square&logo=googlechrome&logoColor=white)](https://getsemantica.ai/) [![Docs](https://img.shields.io/badge/Docs-docs.getsemantica.ai-0099FF?style=flat-square&logo=readthedocs&logoColor=white)](https://docs.getsemantica.ai/) [![Discord](https://img.shields.io/badge/Discord-Join%20Community-5865F2?style=flat-square&logo=discord&logoColor=white)](https://discord.gg/sV34vps5hH) [![Twitter/X](https://img.shields.io/badge/Follow-%40BuildSemantica-000000?style=flat-square&logo=x&logoColor=white)](https://x.com/BuildSemantica) [![YouTube](https://img.shields.io/badge/YouTube-Watch%20Demos-FF0000?style=flat-square&logo=youtube&logoColor=white)](https://www.youtube.com/watch?v=QfnNZg4-dZA) [![Changelog](https://img.shields.io/badge/Changelog-View-6E40C9?style=flat-square&logo=keepachangelog&logoColor=white)](CHANGELOG.md)
@@ -56,18 +56,20 @@ pip install semantica
---
Most AI agents run on embeddings, not meaning: similarity scores with no structure, no relationships, and no way to explain why a result came back. Semantica is the semantic/context layer underneath your LLM, vector store, and agent framework: a deterministic infrastructure layer (no LLM required for graph construction, reasoning, or provenance) that turns fragmented enterprise data into a structured, queryable Context Graph and knowledge graph, governed by ontologies and controlled vocabularies (OWL, SHACL, SKOS) so the meaning of your data is explicit, not just its embedding. Decision provenance and audit trails fall out of that structure as a property, not the product itself; in domains a regulator can question, that same structure just happens to double as a straight answer to "why."
Most AI agents act without a trail. They store embeddings, not meaning: context that can't be explained, decisions that can't be audited. In lending, that gap is a compliance exposure, not an inconvenience: an underwriting agent's approval has to survive a regulator's "why" months later.
Semantica sits underneath your LLM, vector store, and agent framework as a deterministic infrastructure layer: no LLM required for graph construction, reasoning, or provenance.
> ⚠️ **System-level explainability, not foundation-model explainability.** Semantica does not expose or reconstruct what happens *inside* the LLM — its internal reasoning or chain-of-thought stays opaque, as it does for any external system. Semantica explains what's *outside* the model: the context and data fed in, the decision produced, its provenance, relevant relationships, applied policies, and the full execution trail.
**Who it's for:**
- **AI/ML platform teams** shipping agents that make consequential decisions and need structured, queryable context, not just a vector index
- **Data platform teams on Databricks or Snowflake** turning tables already in Unity Catalog or a warehouse into a governed, lineage-tracked knowledge graph, without exporting to a third-party SaaS
- **Compliance, risk, and audit teams** who need a straight answer to "why did the AI do that?" in a format a regulator accepts
- **Regulated enterprises** (finance, healthcare, legal, government, defense) that can't ship a black box or send their data to someone else's SaaS to get one
- **AI/ML platform teams** shipping agents that make consequential decisions and need structured, queryable context built from fragmented raw data, not just a vector index
- **Data platform teams on Databricks or Snowflake** who need to turn tables already sitting in Unity Catalog or a Snowflake warehouse into a governed, lineage-tracked knowledge graph, without exporting that data to a third-party SaaS first
- **Compliance, risk, and audit teams** who need a straight answer to "why did the AI do that?" in a format a regulator will actually accept
- **Regulated enterprises** (finance, healthcare, legal, government, defense) that can't ship a black box, and can't send their data to someone else's SaaS to get one
- **Platform and infra engineers** who want the KG, reasoning, and provenance stack self-hosted and swappable, not locked to one vendor's backend
- **Data and knowledge engineers** building a KG from messy, multi-source data, where conflicting facts get flagged and duplicates get merged, not silently overwritten
- **Data and knowledge engineers** building a KG from messy, multi-source data: entities and relationships get extracted, conflicting or contradictory facts are flagged instead of silently overwritten, and duplicates are merged before they turn into noise
**[Quick Start](#quick-start)** &nbsp;·&nbsp; **[Architecture](#architecture)** &nbsp;·&nbsp; **[What You Get](#what-semantica-gives-you)** &nbsp;·&nbsp; **[Why Semantica](#why-semantica)** &nbsp;·&nbsp; **[Decision Intelligence](#decision-intelligence)** &nbsp;·&nbsp; **[Context Graphs](#context-graphs)** &nbsp;·&nbsp; **[Recipe: Audit Trail](#recipe-audit-trail-for-a-regulated-decision)** &nbsp;·&nbsp; **[Module Reference](#module-reference)** &nbsp;·&nbsp; **[Integrations](#integrations)** &nbsp;·&nbsp; **[CLI](#cli)** &nbsp;·&nbsp; **[Performance](#performance)** &nbsp;·&nbsp; **[Install](#installation)**
@@ -81,7 +83,7 @@ Most AI agents run on embeddings, not meaning: similarity scores with no structu
- **Full Auditability:** W3C PROV-O provenance on every fact, with audit trails exportable to JSON, CSV, or RDF
- **Deterministic Reasoning:** Forward chaining, Rete network, Datalog, and SPARQL with fully explainable paths, not black boxes
- **Knowledge Pipeline:** Multi-source ingestion, entity-aware chunking, NER/relation/event extraction, and knowledge graph construction, with semantic deduplication and provenance-preserving merges throughout
- **Enterprise Data Platforms:** Native connectors for Databricks (Unity Catalog + Delta Lake, PAT/OAuth M2M auth, catalog/schema/table/lineage introspection), Snowflake (warehouse/database/schema, key-pair and OAuth auth), and SAP OData (Business Partners, Sales Orders, OAuth2/Basic auth), so data already living in your lakehouse or warehouse becomes graph nodes with provenance, not another export/import hop
- **Enterprise Data Platforms:** Native connectors for Databricks (Unity Catalog + Delta Lake, PAT/OAuth M2M auth, catalog/schema/table/lineage introspection) and Snowflake (warehouse/database/schema, key-pair and OAuth auth), so tables already living in your lakehouse or warehouse become graph nodes with provenance, not another export/import hop
- **Graph Analytics:** Centrality, community detection, link prediction, and shortest-path queries over the graph you just built
- **Polyglot Graph Storage:** Native RDF (embedded Oxigraph, Blazegraph, Apache Jena, Eclipse RDF4J via SPARQL) and Labeled Property Graphs (Neo4j, FalkorDB, Apache AGE, AWS Neptune via Cypher), plus vector stores, all swappable without touching your code
- **Visualization:** Explore any graph, ontology, or timeline in an interactive browser workbench
@@ -139,6 +141,10 @@ compliant = graph.check_decision_rules({"category": "vendor_selection"}) # poli
```bash
semantica doctor
# Python 3.11.9 pass
# semantica 0.6.7 pass
# faiss vector store pass
# Config file pass ~/.semantica/config.yaml
```
**Running in a script or CI?** Progress bars are written only when stdout is an interactive terminal (or a Jupyter notebook), so piping and redirecting stay clean by default. Override with `SEMANTICA_DISABLE_PROGRESS=1` to silence progress everywhere, or `SEMANTICA_FORCE_PROGRESS=1` to keep it when stdout is redirected. `SEMANTICA_DISABLE_PROGRESS` takes precedence.
@@ -163,7 +169,7 @@ Sources → Ingest → Parse → Normalize → Split → Extract → Conflict De
→ Vector Store + Polyglot Graph Store (RDF & LPG) → Export / Visualize / REST · MCP · CLI
```
- **Ingest:** files, web, databases, enterprise data platforms (Databricks, Snowflake, SAP), cloud (Google Drive, Elasticsearch), streams (Kafka, Kinesis), Git, email, MCP
- **Ingest:** files, web, databases, enterprise data platforms (Databricks, Snowflake), cloud (Google Drive, Elasticsearch), streams (Kafka, Kinesis), Git, email, MCP
- **Parse → Normalize → Split:** document parsing, text/entity/date normalization, GraphRAG-native entity-aware chunking
- **Extract → Conflict Detection → Deduplication:** NER, relations, events, triplets; conflicting facts flagged and resolved before they merge
- **Knowledge Graph:** `GraphBuilder` constructs the graph; bi-temporal facts and full graph analytics (centrality, communities, link prediction) run on top of it
@@ -273,7 +279,7 @@ retrieved = ctx.retrieve("who approved the Acme contract?")
## Recipe: Audit Trail for a Regulated Decision
One pattern built on the same Context Graph: record a causally-linked decision chain, attach provenance to every entity, and export a regulator-ready audit trail.
The flagship pattern: record a causally-linked decision chain, attach provenance to every entity, and export a regulator-ready audit trail.
```python
from semantica.context import ContextGraph
@@ -316,7 +322,7 @@ Every module below is independently importable, with working code samples verifi
| Module | What it does |
| --- | --- |
| [`semantica.ingest`](#semanticaingest-multi-source-ingestion) | Files, web, databases, APIs, streams, email, Git, Parquet, Databricks, Snowflake, SAP, MCP |
| [`semantica.ingest`](#semanticaingest-multi-source-ingestion) | Files, web, databases, APIs, streams, email, Git, Parquet, Databricks, Snowflake, MCP |
| [`semantica.semantic_extract`](#semanticasemantic_extract-ner-relations-events-triplets) | NER, relation extraction, event detection, triplet generation |
| [`semantica.kg`](#semanticakg-knowledge-graph-construction--analysis) | Graph construction, centrality, communities, link prediction |
| [`semantica.reasoning`](#semanticareasoning-forward-chaining-rete-datalog-sparql) | Forward chaining, Rete, Datalog, SPARQL, fully explainable |
@@ -345,7 +351,7 @@ Expand any module below for its runnable example.
<summary><b><code>semantica.ingest</code></b>: Multi-Source Ingestion</summary>
<a id="semanticaingest-multi-source-ingestion"></a>
Ingest from files, web, databases, APIs, streams, email, Git repos, Parquet, Databricks, Snowflake, SAP, or MCP servers, all through a unified interface.
Ingest from files, web, databases, APIs, streams, email, Git repos, Parquet, Databricks, Snowflake, or MCP servers, all through a unified interface.
```python
from semantica.ingest import FileIngestor, WebIngestor, ParquetIngestor, DBIngestor
@@ -396,7 +402,7 @@ orders = snowflake.ingest_table("ORDERS", limit=10_000)
> **Security Note:** Never hardcode credentials (`token`, `password`, `private_key`) in production code; pass them via environment variables (e.g., `DATABRICKS_TOKEN`, `SNOWFLAKE_PASSWORD`) or a secrets manager.
**Supported sources:** Local files (PDF, DOCX, PPTX, HTML, TXT, CSV, JSON, YAML, Excel, XML) · Web pages · RSS/Atom feeds · REST APIs · Databases (PostgreSQL, MySQL, SQLite, Oracle, SQL Server) · Parquet datasets · Databricks (Unity Catalog + Delta Lake) · Snowflake · SAP (OData v2/v4) · Git repositories · Email (IMAP/POP3) · Message streams (Kafka, RabbitMQ, Kinesis, Pulsar) · MCP resources · Apache Arrow/Feather/IPC (`ArrowIngestor`)
**Supported sources:** Local files (PDF, DOCX, PPTX, HTML, TXT, CSV, JSON, YAML, Excel, XML) · Web pages · RSS/Atom feeds · REST APIs · Databases (PostgreSQL, MySQL, SQLite, Oracle, SQL Server) · Parquet datasets · Databricks (Unity Catalog + Delta Lake) · Snowflake · Git repositories · Email (IMAP/POP3) · Message streams (Kafka, RabbitMQ, Kinesis, Pulsar) · MCP resources · Apache Arrow/Feather/IPC (`ArrowIngestor`)
DuckDB, Elasticsearch, Google Drive, HuggingFace, MongoDB, and Pandas ingestion also ship (`DuckDBIngestor`, `ElasticIngestor`, `GDriveIngestor`, `HuggingFaceIngestor`, `MongoIngestor`, `PandasIngestor`) but aren't re-exported from the top-level `semantica.ingest` namespace yet — import them directly: `from semantica.ingest.duckdb_ingestor import DuckDBIngestor`.
@@ -1024,7 +1030,7 @@ team = Team(agents=[researcher, analyst], mode="coordinate")
## More Recipes
The audit-trail recipe is [above](#recipe-audit-trail-for-a-regulated-decision). Here are three more common patterns.
The flagship audit-trail recipe is [above](#recipe-audit-trail-for-a-regulated-decision). Here are three more common patterns.
<details>
<summary><b>End-to-End GraphRAG Pipeline</b></summary>
@@ -1141,7 +1147,7 @@ if report.valid:
| **Vector Store** | FAISS · Pinecone · Weaviate · Qdrant · Milvus · PgVector · hybrid + filtered search |
| **Graph Databases (LPG)** | Neo4j · FalkorDB · Apache AGE · AWS Neptune |
| **Triple Stores (RDF)** | Oxigraph (embedded) · Blazegraph · Apache Jena · Eclipse RDF4J · unified `TripletStore` interface · SPARQL query & bulk load |
| **Enterprise Data Platforms** | Databricks (`DatabricksIngestor`: Unity Catalog + Delta Lake, PAT/OAuth M2M, table/query ingestion, catalog/schema/table/lineage introspection) · Snowflake (`SnowflakeIngestor`: warehouse/database/schema, password/key-pair/OAuth auth) · SAP (`SAPIngestor`: OData v2/v4, OAuth2/Basic auth, Business Partners/Sales Orders) |
| **Enterprise Data Platforms** | Databricks (`DatabricksIngestor`: Unity Catalog + Delta Lake, PAT/OAuth M2M, table/query ingestion, catalog/schema/table/lineage introspection) · Snowflake (`SnowflakeIngestor`: warehouse/database/schema, password/key-pair/OAuth auth) |
| **LLM Providers** | **All already supported today:** OpenAI (GPT-4o, o1, o3) · Anthropic (Claude) · Google Gemini · Mistral · Meta Llama · Groq · Cohere · Azure OpenAI · AWS Bedrock · Ollama · DeepSeek · Perplexity · Together AI · Fireworks AI · Replicate · HuggingFace · via `semantica.llms` and LiteLLM |
---
@@ -1513,7 +1519,6 @@ pip install semantica[vectorstore-qdrant] # Qdrant vector store
pip install semantica[vectorstore-pinecone] # Pinecone vector store
pip install semantica[db-snowflake] # Snowflake
pip install semantica[db-databricks] # Databricks (SDK + SQL connector)
pip install semantica[ingest-sap] # SAP OData
pip install semantica[ingest-parquet] # Parquet / PyArrow
pip install semantica[ingest-arrow] # Apache Arrow, Feather, IPC
pip install semantica[viz] # HTML interactive visualization
@@ -1529,20 +1534,6 @@ git clone https://github.com/semantica-agi/semantica.git
cd semantica && pip install -e ".[dev]" && pytest tests/
```
### CI & Deployment
Wiring `semantica` into your own CI is a two-minute job. On GitHub Actions, use the reusable composite action:
```yaml
- uses: semantica-agi/semantica/.github/actions/setup-semantica@main
with:
python-version: '3.11'
```
Copy-paste starting templates for GitHub Actions, GitLab CI, and CircleCI live in [examples/ci/](examples/ci/). The published package itself is verified installable across Ubuntu/macOS/Windows and Python 3.9-3.12 every week by the [Install Matrix workflow](.github/workflows/install-matrix.yml).
Ready-made deployment configs for AWS, GCP, Azure, Fly.io, Railway, Render, Kubernetes, and Helm are in [deploy/](deploy/).
---
## Enterprise
+3 -2
View File
@@ -153,7 +153,7 @@ that attack chain.
- **Risk**: a PR merges without its security/CI checks passing.
**Control**: merges require the `build`, `Analyze Python` (CodeQL), and `security-scan` checks to pass, in strict mode (checks must be re-run against the latest `main`).
- **Risk**: a compromised scanner job reaches secrets or write access.
**Control**: scanning jobs (`CodeQL`, `security-scan.yml`, `defender-for-devops.yml`) run with read-only, least-privilege permissions (typically `contents: read` + `security-events: write` only) and never share a job, environment, or secret scope with the publish job.
**Control**: scanning jobs (`CodeQL`, `security-scan.yml`, `security.yml`, `defender-for-devops.yml`) run with read-only, least-privilege permissions (typically `contents: read` + `security-events: write` only) and never share a job, environment, or secret scope with the publish job.
- **Risk**: secrets are committed accidentally.
**Control**: GitHub secret scanning and push protection are both enabled at the repository level, rejecting pushes that contain recognizable credential patterns before they land in history.
@@ -164,7 +164,8 @@ Every scan below runs continuously in CI, not just at release time:
- **CodeQL** (`security-and-quality` query pack) — Python source: injection, unsafe deserialization, and other code-level vulnerability classes. Runs in `codeql.yml` on every push/PR to `main` and weekly.
- **Bandit** — Python-specific security anti-patterns (hardcoded secrets, unsafe `eval`/`pickle`, weak crypto, etc.); CI fails on any HIGH-severity finding. Runs in `security-scan.yml` on every push/PR to `main` and twice weekly.
- **Semgrep** (`p/security` ruleset) — cross-language static-analysis security patterns. Runs in `security-scan.yml` on every push/PR to `main` and twice weekly.
- **pip-audit** — PyPA-maintained, OSV-backed vulnerability database cross-check against Semantica's pinned dependency tree, including optional LLM-provider extras such as LiteLLM; CI fails on any match. Runs in `security-scan.yml` on every push/PR to `main` and twice weekly, and can be triggered on demand via `workflow_dispatch`.
- **Safety** — known CVEs in Semantica's own installed dependencies, including optional LLM-provider extras such as LiteLLM; CI fails on any match. Runs in `security-scan.yml` on every push/PR to `main` and twice weekly.
- **pip-audit** — independent, PyPA-maintained vulnerability database cross-check against installed dependencies (Safety and pip-audit use different advisory sources, so both run). Runs in `security.yml` weekly.
- **Microsoft Defender for DevOps** (`eslint`, `templateanalyzer`, `terrascan`) — JavaScript/TypeScript lint-security rules and infrastructure-as-code misconfigurations. Runs in `defender-for-devops.yml` on every push/PR to `main` and weekly.
- **Checkov** — Kubernetes, Helm, Dockerfile, GitHub Actions, and secrets-pattern IaC scanning; results upload to the same Security tab as CodeQL. Runs in `defender-for-devops.yml` on every push/PR to `main` and weekly.
- **GitGuardian** — secret-detection check on every pull request, installed as a GitHub App integration (not a repo-local workflow). Runs on every PR.
@@ -0,0 +1,222 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/semantica-agi/semantica/blob/main/cookbook/advanced/09_Semantic_Layer_Construction.ipynb)\n",
"\n",
"# Semantic Layer Construction\n",
"\n",
"## Overview\n",
"\n",
"Build an enterprise semantic layer: construct knowledge graph, generate ontology, create semantic layer, export RDF, and store in triplet store.\n",
"\n",
"\n",
"**Documentation**: [API Reference](https://semantica.readthedocs.io/concepts/)\n",
"\n",
"## Installation\n",
"\n",
"Install Semantica from PyPI:\n",
"\n",
"```bash\n",
"pip install semantica\n",
"# Or with all optional dependencies:\n",
"pip install semantica[all]\n",
"```\n",
"\n",
"## Workflow: Build KG → Generate Ontology → Create Semantic Layer → Export RDF \n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"!pip install -qU semantica\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from semantica.kg import GraphBuilder\n",
"from semantica.ontology import OntologyGenerator\n",
"from semantica.export import RDFExporter\n",
"from semantica.triplet_store import TripletStore\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 1: Build Knowledge Graph\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"builder = GraphBuilder()\n",
"\n",
"entities = [\n",
" {\"id\": \"e1\", \"type\": \"Person\", \"name\": \"Alice\", \"properties\": {\"age\": 30, \"role\": \"Engineer\"}},\n",
" {\"id\": \"e2\", \"type\": \"Person\", \"name\": \"Bob\", \"properties\": {\"age\": 35, \"role\": \"Manager\"}},\n",
" {\"id\": \"e3\", \"type\": \"Organization\", \"name\": \"Tech Corp\", \"properties\": {\"founded\": 2010}},\n",
" {\"id\": \"e4\", \"type\": \"Project\", \"name\": \"Project Alpha\", \"properties\": {\"status\": \"active\"}},\n",
"]\n",
"\n",
"relationships = [\n",
" {\"source\": \"e1\", \"target\": \"e2\", \"type\": \"reports_to\"},\n",
" {\"source\": \"e1\", \"target\": \"e3\", \"type\": \"works_for\"},\n",
" {\"source\": \"e2\", \"target\": \"e3\", \"type\": \"works_for\"},\n",
" {\"source\": \"e1\", \"target\": \"e4\", \"type\": \"works_on\"},\n",
"]\n",
"\n",
"knowledge_graph = builder.build(entities, relationships)\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 2: Generate Ontology\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"generator = OntologyGenerator()\n",
"ontology = generator.generate_from_graph(knowledge_graph)\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 3: Create Semantic Layer\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"def create_mappings(kg, ontology):\n",
" mappings = {\n",
" \"entity_type_mappings\": {},\n",
" \"relationship_type_mappings\": {},\n",
" \"property_mappings\": {}\n",
" }\n",
" \n",
" entity_types = set(e.get(\"type\") for e in entities)\n",
" ontology_classes = ontology.get(\"classes\", [])\n",
" \n",
" for entity_type in entity_types:\n",
" matching_class = next((cls for cls in ontology_classes if cls.get(\"name\") == entity_type), None)\n",
" if matching_class:\n",
" mappings[\"entity_type_mappings\"][entity_type] = matching_class.get(\"uri\", entity_type)\n",
" \n",
" relationship_types = set(r.get(\"type\") for r in relationships)\n",
" ontology_properties = ontology.get(\"properties\", [])\n",
" \n",
" for rel_type in relationship_types:\n",
" matching_prop = next((prop for prop in ontology_properties if prop.get(\"name\") == rel_type), None)\n",
" if matching_prop:\n",
" mappings[\"relationship_type_mappings\"][rel_type] = matching_prop.get(\"uri\", rel_type)\n",
" \n",
" return mappings\n",
"\n",
"mappings = create_mappings(knowledge_graph, ontology)\n",
"\n",
"semantic_layer = {\n",
" \"graph\": knowledge_graph,\n",
" \"ontology\": ontology,\n",
" \"mappings\": mappings,\n",
" \"metadata\": {\n",
" \"version\": \"1.0\",\n",
" \"created_at\": \"2024-01-01\",\n",
" \"description\": \"Enterprise semantic layer\"\n",
" }\n",
"}\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 4: Export RDF\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"exporter = RDFExporter()\n",
"# Export Knowledge Graph\n",
"exporter.export(knowledge_graph, \"knowledge_graph.ttl\", format=\"turtle\")\n",
"print(\"Exported knowledge graph to knowledge_graph.ttl\")\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Summary\n",
"\n",
"Enterprise semantic layer construction:\n",
"- Knowledge Graph Built\n",
"- Ontology Generated\n",
"- Semantic Layer Created with Mappings\n",
"- RDF Export Completed\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": []
}
],
"metadata": {
"kernelspec": {
"display_name": "Python 3",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3",
"version": "3.11.9"
}
},
"nbformat": 4,
"nbformat_minor": 2
}
@@ -0,0 +1,435 @@
{
"nbformat": 4,
"nbformat_minor": 5,
"metadata": {
"kernelspec": {
"display_name": "Python 3",
"language": "python",
"name": "python3"
},
"language_info": {
"name": "python",
"version": "3.10.0"
}
},
"cells": [
{
"cell_type": "markdown",
"id": "cell-0",
"metadata": {},
"source": [
"[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/semantica-agi/semantica/blob/main/cookbook/advanced/13_Manual_Ontology_Snowflake_Mapping.ipynb)\n",
"\n",
"# Manual Ontology + Snowflake Mapping\n",
"\n",
"This notebook answers a specific workflow:\n",
"\n",
"> *\"I want to design the ontology myself — not have AI infer it from my tables — and then map Snowflake data to it explicitly.\"*\n",
"\n",
"### What this notebook demonstrates\n",
"\n",
"| Step | What happens | Who controls it |\n",
"|---|---|---|\n",
"| 1 | Design ontology classes and properties | **You** (Python dict) |\n",
"| 2 | Model n-ary facts with reification | **You** (`AssociativeClassBuilder`) |\n",
"| 3 | Pull rows from Snowflake | Semantica `SnowflakeIngestor` |\n",
"| 4 | Map columns → ontology-aligned graph | **You** (explicit transform) |\n",
"| 5 | Validate + export OWL / SHACL | Semantica `OntologyEngine` |\n",
"| 6 | Load to triplet store and query | Semantica `TripletStore` |\n",
"\n",
"### What this notebook does NOT do\n",
"\n",
"- No LLM-driven ontology generation\n",
"- No schema introspection or table-to-class inference\n",
"- No \"suggest ontology from my data\"\n",
"\n",
"### Standards coverage\n",
"\n",
"| Feature | Status |\n",
"|---|---|\n",
"| OWL 2 (Turtle / RDF-XML) | Supported |\n",
"| SHACL 1.1 shapes | Supported |\n",
"| SPARQL 1.1 | Supported |\n",
"| Reification / n-ary facts | Supported via `AssociativeClassBuilder` |\n",
"| SPARQL 1.2 (reifier annotation, `LATERAL`) | Planned |\n",
"| SHACL 1.2 (`sh:severity` extensions, SHACL-AF) | Planned |"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-1",
"metadata": {},
"outputs": [],
"source": [
"!pip install -qU semantica"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-2",
"metadata": {},
"outputs": [],
"source": [
"import os\n",
"from typing import Any, Dict, List\n",
"\n",
"from semantica.ingest import SnowflakeIngestor\n",
"from semantica.kg.methods import build_kg\n",
"from semantica.ontology import AssociativeClassBuilder, OntologyEngine\n",
"from semantica.triplet_store import TripletStore"
]
},
{
"cell_type": "markdown",
"id": "cell-3",
"metadata": {},
"source": [
"## Step 1: Hand-Design the Ontology in Python\n",
"\n",
"You define every class and property explicitly. Nothing is read from Snowflake at this stage.\n",
"\n",
"**Design decisions that belong to you:**\n",
"- Which classes exist and what they mean\n",
"- Which properties are datatype vs. object properties\n",
"- Domain, range, and cardinality constraints\n",
"- Which properties are required (later enforced by SHACL)\n",
"\n",
"This dict versions with your code. It does not change when your database schema changes."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-4",
"metadata": {},
"outputs": [],
"source": "BASE_URI = \"https://example.com/hr/\"\n\n# Your ontology — designed by you, not inferred by Semantica.\nontology: Dict[str, Any] = {\n \"name\": \"EmploymentDomainOntology\",\n \"uri\": f\"{BASE_URI}EmploymentDomainOntology\",\n \"namespace\": {\"base_uri\": BASE_URI},\n\n # You decide the class taxonomy\n \"classes\": [\n {\"name\": \"Person\", \"uri\": f\"{BASE_URI}Person\"},\n {\"name\": \"Organization\", \"uri\": f\"{BASE_URI}Organization\"},\n {\"name\": \"Role\", \"uri\": f\"{BASE_URI}Role\"},\n # EmploymentEvent is a reification node.\n # It connects Person + Organization + Role and carries salary/date context.\n {\"name\": \"EmploymentEvent\", \"uri\": f\"{BASE_URI}EmploymentEvent\"},\n ],\n\n # Each property carries a full URI so TripletStore stores it as hr:<name>\n # rather than the default urn:property:<name>.\n # This ensures SPARQL queries using PREFIX hr: match what is actually stored.\n \"properties\": [\n # Datatype properties\n {\"name\": \"name\", \"uri\": f\"{BASE_URI}name\", \"type\": \"datatype\", \"domain\": \"Person\", \"range\": \"string\", \"required\": True},\n {\"name\": \"legalName\", \"uri\": f\"{BASE_URI}legalName\", \"type\": \"datatype\", \"domain\": \"Organization\", \"range\": \"string\", \"required\": True},\n {\"name\": \"title\", \"uri\": f\"{BASE_URI}title\", \"type\": \"datatype\", \"domain\": \"Role\", \"range\": \"string\", \"required\": True},\n {\"name\": \"startDate\", \"uri\": f\"{BASE_URI}startDate\", \"type\": \"datatype\", \"domain\": \"EmploymentEvent\", \"range\": \"date\"},\n {\"name\": \"endDate\", \"uri\": f\"{BASE_URI}endDate\", \"type\": \"datatype\", \"domain\": \"EmploymentEvent\", \"range\": \"date\"},\n {\"name\": \"salary\", \"uri\": f\"{BASE_URI}salary\", \"type\": \"datatype\", \"domain\": \"EmploymentEvent\", \"range\": \"decimal\"},\n\n # Object properties — reification spokes (required)\n {\"name\": \"employee\", \"uri\": f\"{BASE_URI}employee\", \"type\": \"object\", \"domain\": \"EmploymentEvent\", \"range\": \"Person\", \"required\": True},\n {\"name\": \"employer\", \"uri\": f\"{BASE_URI}employer\", \"type\": \"object\", \"domain\": \"EmploymentEvent\", \"range\": \"Organization\", \"required\": True},\n {\"name\": \"role\", \"uri\": f\"{BASE_URI}role\", \"type\": \"object\", \"domain\": \"EmploymentEvent\", \"range\": \"Role\", \"required\": True},\n\n # Shortcut edges — direct person→org / person→role without traversing the event node\n {\"name\": \"worksFor\", \"uri\": f\"{BASE_URI}worksFor\", \"type\": \"object\", \"domain\": \"Person\", \"range\": \"Organization\"},\n {\"name\": \"hasRole\", \"uri\": f\"{BASE_URI}hasRole\", \"type\": \"object\", \"domain\": \"Person\", \"range\": \"Role\"},\n ],\n}\n\nontology"
},
{
"cell_type": "markdown",
"id": "cell-5",
"metadata": {},
"source": [
"## Step 2: Reification — Modeling N-Ary Facts\n",
"\n",
"**The problem with binary triples:**\n",
"A simple triple `(Alice, worksFor, Acme)` cannot carry extra context such as salary, start date, or role.\n",
"Standard RDF reification and OWL n-ary patterns solve this by introducing an intermediate node.\n",
"\n",
"Semantica's `AssociativeClassBuilder` is the Pythonic API for this pattern:\n",
"\n",
"```\n",
"EmploymentEvent\n",
" ├── employee → Person (required)\n",
" ├── employer → Organization (required)\n",
" ├── role → Role (required)\n",
" ├── startDate → xsd:date\n",
" ├── endDate → xsd:date\n",
" └── salary → xsd:decimal\n",
"```\n",
"\n",
"**On SPARQL 1.1 vs. SPARQL 1.2:**\n",
"- **SPARQL 1.1 (current):** traverse the event node explicitly — `?event hr:employee ?person ; hr:salary ?salary`\n",
"- **SPARQL 1.2 (planned):** the draft reifier annotation syntax allows attaching context to triples directly, without a separate intermediate node. Semantica will adopt this once the spec is ratified.\n",
"\n",
"**On SHACL 1.1 vs. SHACL 1.2:**\n",
"- **SHACL 1.1 (current):** `sh:NodeShape` + `sh:PropertyShape` constraints are exported for all `required` properties and enforced at load time.\n",
"- **SHACL 1.2 (planned):** `sh:severity` profile extensions and SHACL-AF rules are on the roadmap."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-6",
"metadata": {},
"outputs": [],
"source": "assoc_builder = AssociativeClassBuilder()\n\nemployment_assoc = assoc_builder.create_associative_class(\n name=\"EmploymentEvent\",\n connects=[\"Person\", \"Organization\", \"Role\"],\n temporal=True, # adds startDate / endDate handling\n properties={\n \"startDate\": \"xsd:date\",\n \"endDate\": \"xsd:date\",\n \"salary\": \"xsd:decimal\",\n },\n)\n\nvalidation_result = assoc_builder.validate_associative_class(employment_assoc)\n\n# AssociativeClass is a dataclass — use attribute access, not .get()\nprint(\"AssociativeClass structure:\")\nprint(f\" name: {employment_assoc.name}\")\nprint(f\" connects: {employment_assoc.connects}\")\nprint(f\" temporal: {employment_assoc.temporal}\")\nprint(f\" properties: {list(employment_assoc.properties.keys())}\")\nprint(f\"\\nValidation passed: {validation_result}\")"
},
{
"cell_type": "markdown",
"id": "cell-7",
"metadata": {},
"source": [
"## Step 3: Ingest Snowflake Rows (Extraction Only)\n",
"\n",
"`SnowflakeIngestor` retrieves rows — nothing more. It does **not**:\n",
"- Inspect your table schema\n",
"- Suggest classes or properties\n",
"- Infer relationships from column names\n",
"\n",
"Set `USE_LIVE_SNOWFLAKE=true` plus the env vars below to connect to a real warehouse.\n",
"Otherwise the stub data is used."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-8",
"metadata": {},
"outputs": [],
"source": [
"def fetch_rows_from_snowflake() -> List[Dict[str, Any]]:\n",
" if os.getenv(\"USE_LIVE_SNOWFLAKE\", \"false\").lower() != \"true\":\n",
" return [\n",
" {\n",
" \"EMPLOYEE_ID\": \"E100\",\n",
" \"EMPLOYEE_NAME\": \"Alice Johnson\",\n",
" \"ORG_ID\": \"O10\",\n",
" \"ORG_NAME\": \"Acme Corp\",\n",
" \"ROLE_ID\": \"R7\",\n",
" \"ROLE_TITLE\": \"Senior Engineer\",\n",
" \"START_DATE\": \"2025-01-15\",\n",
" \"END_DATE\": None,\n",
" \"SALARY\": 160000,\n",
" },\n",
" {\n",
" \"EMPLOYEE_ID\": \"E101\",\n",
" \"EMPLOYEE_NAME\": \"Bob Singh\",\n",
" \"ORG_ID\": \"O10\",\n",
" \"ORG_NAME\": \"Acme Corp\",\n",
" \"ROLE_ID\": \"R9\",\n",
" \"ROLE_TITLE\": \"Data Architect\",\n",
" \"START_DATE\": \"2024-09-01\",\n",
" \"END_DATE\": None,\n",
" \"SALARY\": 185000,\n",
" },\n",
" ]\n",
"\n",
" ingestor = SnowflakeIngestor(\n",
" account=os.getenv(\"SNOWFLAKE_ACCOUNT\"),\n",
" user=os.getenv(\"SNOWFLAKE_USER\"),\n",
" password=os.getenv(\"SNOWFLAKE_PASSWORD\"),\n",
" warehouse=os.getenv(\"SNOWFLAKE_WAREHOUSE\"),\n",
" database=os.getenv(\"SNOWFLAKE_DATABASE\"),\n",
" schema=os.getenv(\"SNOWFLAKE_SCHEMA\", \"PUBLIC\"),\n",
" )\n",
" query = (\n",
" \"SELECT EMPLOYEE_ID, EMPLOYEE_NAME, \"\n",
" \"ORG_ID, ORG_NAME, ROLE_ID, ROLE_TITLE, \"\n",
" \"START_DATE, END_DATE, SALARY \"\n",
" \"FROM HR_EMPLOYMENT_FACT\"\n",
" )\n",
" data = ingestor.ingest_query(query)\n",
" ingestor.close()\n",
" return data.data\n",
"\n",
"\n",
"rows = fetch_rows_from_snowflake()\n",
"rows[:2]"
]
},
{
"cell_type": "markdown",
"id": "cell-9",
"metadata": {},
"source": [
"## Step 4: Map Rows to Ontology Concepts Explicitly\n",
"\n",
"This is the semantic transformation layer — the part that makes your ontology real.\n",
"\n",
"Semantica does not guess which column becomes which entity or property.\n",
"Every assignment is code you write and own:\n",
"\n",
"- **Stable node IDs** — deterministic, collision-safe, derived from business keys\n",
"- **Class assignment** — matches what you declared in Step 1\n",
"- **Property routing** — each column value goes to the correct ontology property\n",
"- **Reification wiring** — `EmploymentEvent` is linked to its three participants\n",
"\n",
"When your Snowflake schema changes, only this function needs updating. The ontology stays stable."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-10",
"metadata": {},
"outputs": [],
"source": "def map_rows_to_kg(rows: List[Dict[str, Any]]) -> Dict[str, Any]:\n entities: Dict[str, Dict[str, Any]] = {}\n relationships: List[Dict[str, Any]] = []\n\n for row in rows:\n # Stable, deterministic node IDs derived from business keys\n person_id = f\"person:{row['EMPLOYEE_ID']}\"\n org_id = f\"org:{row['ORG_ID']}\"\n role_id = f\"role:{row['ROLE_ID']}\"\n # Event ID includes all three participants + start date so that\n # a re-hired employee gets a distinct event node, not an overwrite.\n event_id = f\"employment:{row['EMPLOYEE_ID']}:{row['ORG_ID']}:{row['START_DATE']}\"\n\n # Entities — \"type\" must match a class name from Step 1\n entities[person_id] = {\n \"id\": person_id,\n \"type\": \"Person\",\n \"properties\": {\"name\": row[\"EMPLOYEE_NAME\"]},\n }\n entities[org_id] = {\n \"id\": org_id,\n \"type\": \"Organization\",\n \"properties\": {\"legalName\": row[\"ORG_NAME\"]},\n }\n entities[role_id] = {\n \"id\": role_id,\n \"type\": \"Role\",\n \"properties\": {\"title\": row[\"ROLE_TITLE\"]},\n }\n\n # Reification node — filter out None values so TripletStore does not\n # stringify None as the literal \"None\" for open-ended employment.\n event_props = {\n \"startDate\": row[\"START_DATE\"],\n \"endDate\": row[\"END_DATE\"],\n \"salary\": row[\"SALARY\"],\n }\n entities[event_id] = {\n \"id\": event_id,\n \"type\": \"EmploymentEvent\",\n \"properties\": {k: v for k, v in event_props.items() if v is not None},\n }\n\n # Full URIs for relationship types so TripletStore stores hr:<type>\n # instead of the default urn:property:<type>, keeping SPARQL consistent.\n relationships.extend([\n # Shortcut edges — fast SPARQL when context is not needed\n {\"source\": person_id, \"target\": org_id, \"type\": f\"{BASE_URI}worksFor\"},\n {\"source\": person_id, \"target\": role_id, \"type\": f\"{BASE_URI}hasRole\"},\n # Reification spokes — full context via the event node\n {\"source\": event_id, \"target\": person_id, \"type\": f\"{BASE_URI}employee\"},\n {\"source\": event_id, \"target\": org_id, \"type\": f\"{BASE_URI}employer\"},\n {\"source\": event_id, \"target\": role_id, \"type\": f\"{BASE_URI}role\"},\n ])\n\n return build_kg([{\"entities\": list(entities.values()), \"relationships\": relationships}])\n\n\nkg = map_rows_to_kg(rows)\nprint(f\"Entities built: {len(kg.get('entities', []))}\")\nprint(f\"Relationships built: {len(kg.get('relationships', []))}\")\n\nsample = next((e for e in kg[\"entities\"] if e[\"type\"] == \"EmploymentEvent\"), None)\nprint(f\"\\nSample EmploymentEvent node: {sample}\")"
},
{
"cell_type": "markdown",
"id": "cell-11",
"metadata": {},
"source": [
"## Step 5: Validate Ontology and Export OWL + SHACL\n",
"\n",
"`OntologyEngine` validates your ontology dict and serialises it to standards-compliant files.\n",
"\n",
"**Output files:**\n",
"- `employment_manual_ontology.ttl` — OWL 2 Turtle\n",
"- `employment_manual_shapes.ttl` — SHACL 1.1 node and property shapes\n",
"\n",
"**Standards status:**\n",
"\n",
"| Standard | Semantica support |\n",
"|---|---|\n",
"| SPARQL 1.1 | Full |\n",
"| SHACL 1.1 (`sh:NodeShape`, `sh:PropertyShape`, `sh:minCount`, `sh:datatype`, `sh:class`) | Full |\n",
"| SPARQL 1.2 (reifier annotation syntax, `LATERAL`) | Tracked — not yet implemented |\n",
"| SHACL 1.2 (`sh:severity` profiles, SHACL-AF extensions) | Tracked — not yet implemented |"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-12",
"metadata": {},
"outputs": [],
"source": [
"engine = OntologyEngine(base_uri=BASE_URI)\n",
"\n",
"validation = engine.validate(ontology)\n",
"owl_ttl = engine.to_owl(ontology, format=\"turtle\")\n",
"shacl_ttl = engine.to_shacl(ontology, format=\"turtle\")\n",
"\n",
"engine.export_owl(ontology, \"employment_manual_ontology.ttl\", format=\"turtle\")\n",
"engine.export_shacl(ontology, \"employment_manual_shapes.ttl\", format=\"turtle\")\n",
"\n",
"print(f\"Ontology valid: {validation.valid}\")\n",
"print(f\"Ontology consistent: {validation.consistent}\")\n",
"print(f\"OWL output: {len(owl_ttl):,} chars → employment_manual_ontology.ttl\")\n",
"print(f\"SHACL output: {len(shacl_ttl):,} chars → employment_manual_shapes.ttl\")\n",
"\n",
"print(\"\\n--- SHACL shapes (first 20 lines) ---\")\n",
"print(\"\\n\".join(shacl_ttl.splitlines()[:20]))"
]
},
{
"cell_type": "markdown",
"id": "cell-13",
"metadata": {},
"source": [
"## Best-Practice Architecture\n",
"\n",
"```\n",
"┌──────────────────────────────────┐\n",
"│ Ontology as code (Python dict) │ ← versioned alongside your application\n",
"│ + AssociativeClass for n-ary │\n",
"└───────────────┬──────────────────┘\n",
" │ validate + export\n",
" ▼\n",
"┌───────────────────────────────────┐\n",
"│ OWL 2 Turtle │ SHACL 1.1 │ ← standards-compliant artifacts\n",
"└───────────────┬───────────────────┘\n",
" │\n",
" ▼\n",
"┌──────────────────────────────────┐\n",
"│ Snowflake — raw data access │ ← no schema introspection\n",
"└───────────────┬──────────────────┘\n",
" │ explicit mapping layer\n",
" ▼\n",
"┌──────────────────────────────────┐\n",
"│ Ontology-aligned KG │ ← types, IDs, edges match Step 1\n",
"└───────────────┬──────────────────┘\n",
" │ optional\n",
" ▼\n",
"┌──────────────────────────────────┐\n",
"│ Triplet store + SPARQL 1.1 │\n",
"└──────────────────────────────────┘\n",
"```\n",
"\n",
"**Why this split matters:**\n",
"If Semantica inferred the ontology from your Snowflake schema, every schema migration would risk silently changing your semantic model.\n",
"With this pattern, schema changes only touch the mapping function in Step 4 — the ontology remains stable and under your control."
]
},
{
"cell_type": "markdown",
"id": "cell-14",
"metadata": {},
"source": [
"## SPARQL Query Patterns\n",
"\n",
"Two query styles are available because we wrote both shortcut edges and reification spokes.\n",
"\n",
"### Simple lookup — shortcut edge (no context needed)\n",
"\n",
"```sparql\n",
"PREFIX hr: <https://example.com/hr/>\n",
"\n",
"SELECT ?personName ?orgName\n",
"WHERE {\n",
" ?person a hr:Person ;\n",
" hr:name ?personName ;\n",
" hr:worksFor ?org .\n",
" ?org hr:legalName ?orgName .\n",
"}\n",
"```\n",
"\n",
"### Contextual lookup — via reification node (salary, dates, role)\n",
"\n",
"```sparql\n",
"PREFIX hr: <https://example.com/hr/>\n",
"\n",
"SELECT ?personName ?roleTitle ?salary ?startDate\n",
"WHERE {\n",
" ?event a hr:EmploymentEvent ;\n",
" hr:employee ?person ;\n",
" hr:role ?role ;\n",
" hr:salary ?salary ;\n",
" hr:startDate ?startDate .\n",
" ?person hr:name ?personName .\n",
" ?role hr:title ?roleTitle .\n",
"}\n",
"ORDER BY DESC(?salary)\n",
"```\n",
"\n",
"### Future: SPARQL 1.2 reifier syntax\n",
"\n",
"The SPARQL 1.2 draft introduces annotation syntax that lets you attach context directly to triples, without a separate intermediate node.\n",
"Once the spec is ratified Semantica will adopt it, and the contextual query above may be expressible more concisely."
]
},
{
"cell_type": "markdown",
"id": "cell-15",
"metadata": {},
"source": [
"## Step 6 (Optional): Load to Triplet Store and Run SPARQL\n",
"\n",
"Set `STORE_TO_TRIPLET=true` to load the KG into a live triplet store and run the contextual reification query."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-16",
"metadata": {},
"outputs": [],
"source": [
"if os.getenv(\"STORE_TO_TRIPLET\", \"false\").lower() == \"true\":\n",
" store = TripletStore(\n",
" backend=os.getenv(\"TRIPLET_BACKEND\", \"blazegraph\"),\n",
" endpoint=os.getenv(\"TRIPLET_ENDPOINT\", \"http://localhost:9999/blazegraph\"),\n",
" namespace=os.getenv(\"TRIPLET_NAMESPACE\", \"kb\"),\n",
" )\n",
" store_result = store.store(knowledge_graph=kg, ontology=ontology)\n",
" print(\"Store result:\", store_result)\n",
"\n",
" # Contextual reification query — person + role + salary via EmploymentEvent\n",
" query = \"\"\"\n",
" PREFIX hr: <https://example.com/hr/>\n",
"\n",
" SELECT ?personName ?roleTitle ?salary ?startDate\n",
" WHERE {\n",
" ?event a hr:EmploymentEvent ;\n",
" hr:employee ?person ;\n",
" hr:role ?role ;\n",
" hr:salary ?salary ;\n",
" hr:startDate ?startDate .\n",
" ?person hr:name ?personName .\n",
" ?role hr:title ?roleTitle .\n",
" }\n",
" ORDER BY DESC(?salary)\n",
" LIMIT 10\n",
" \"\"\"\n",
" result = store.execute_query(query)\n",
" print(result)\n",
"else:\n",
" print(\"Skipping triplet-store load/query (set STORE_TO_TRIPLET=true to enable)\")"
]
}
]
}
@@ -10,16 +10,15 @@
"\n",
"## Overview\n",
"\n",
"This notebook demonstrates how to build knowledge graphs from extracted entities and relationships using Semantica's graph building modules. You'll learn to use `GraphBuilder` and `EntityResolver`.\n",
"This notebook demonstrates how to build knowledge graphs from entities and relationships using Semantica's graph building modules. You'll learn to use `GraphBuilder` and `EntityResolver`.\n",
"\n",
"**Documentation**: [API Reference](https://semantica.readthedocs.io/reference/kg/)\n",
"\n",
"### Learning Objectives\n",
"\n",
"- Extract entity mentions and relations, and map them into graph records\n",
"- Use `GraphBuilder` to construct a graph whose edges come from the actual extracted relations\n",
"- Use `EntityResolver` to merge duplicate mentions and remap relationship endpoints\n",
"- Use the `semantica.deduplication` module and report the complete deduplicated entity set\n",
"- Use `GraphBuilder` to construct knowledge graphs\n",
"- Use `EntityResolver` to resolve entity conflicts\n",
"**Note**: For deduplication, use the `semantica.deduplication` module.\n",
"\n",
"## Installation\n",
"\n",
@@ -33,217 +32,120 @@
"\n",
"---\n",
"\n",
"## Step 1: Extract Entities and Relations\n",
"## Step 1: Build Knowledge Graph\n",
"\n",
"Extract entity mentions and relations from text. The sample text mentions `Apple Inc.` in two separate sentences, so we can later show how duplicate mentions are resolved into one canonical entity.\n"
"Construct a knowledge graph from entities and relationships.\n"
]
},
{
"cell_type": "code",
"metadata": {},
"source": [
"%pip install semantica\n",
"\n",
"# spaCy models are distributed separately from the spaCy library. This lesson\n",
"# relies on the English model to recognize standalone places such as Cupertino.\n",
"import sys\n",
"import subprocess\n",
"import spacy\n",
"\n",
"try:\n",
" spacy.load(\"en_core_web_sm\")\n",
"except OSError:\n",
" subprocess.check_call([sys.executable, \"-m\", \"spacy\", \"download\", \"en_core_web_sm\"])\n"
],
"execution_count": null,
"outputs": []
"metadata": {},
"outputs": [],
"source": [
"!pip install semantica\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from semantica.kg import GraphBuilder\n",
"from semantica.semantic_extract import NERExtractor, RelationExtractor\n",
"\n",
"text = (\n",
" \"Apple Inc. is headquartered in Cupertino, California. \"\n",
" \"Tim Cook is the CEO of Apple Inc. \"\n",
" \"The company is a technology company.\"\n",
")\n",
"\n",
"builder = GraphBuilder()\n",
"ner_extractor = NERExtractor()\n",
"relation_extractor = RelationExtractor()\n",
"\n",
"mentions = ner_extractor.extract(text)\n",
"relations = relation_extractor.extract(text, mentions)\n",
"text = \"Apple Inc. is a technology company. Tim Cook is the CEO of Apple Inc. Apple Inc. is headquartered in Cupertino, California.\"\n",
"\n",
"print(\"Entity mentions:\")\n",
"for mention in mentions:\n",
" print(f\" {mention.text!r:<13} {mention.label:<7} span=[{mention.start_char}:{mention.end_char}]\")\n",
"\n",
"print(\"\\nExtracted relations:\")\n",
"for rel in relations:\n",
" print(f\" {rel.subject.text!r} --{rel.predicate}--> {rel.object.text!r}\")"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 2: Build the Knowledge Graph\n",
"\n",
"Give every mention a graph ID, then translate each relation's `subject` and `object` into those IDs. Building edges from the actual relation endpoints — rather than guessing endpoints from list positions — is what keeps the graph faithful to the text.\n"
]
},
{
"cell_type": "code",
"metadata": {},
"source": [
"from semantica.kg import GraphBuilder\n",
"entities_list = ner_extractor.extract(text)\n",
"relationships_list = relation_extractor.extract(text, entities_list)\n",
"\n",
"entities = []\n",
"span_to_id = {}\n",
"for i, mention in enumerate(mentions, 1):\n",
" graph_id = f\"e{i}\"\n",
" span_to_id[(mention.start_char, mention.end_char)] = graph_id\n",
"for i, entity in enumerate(entities_list[:5], 1):\n",
" entities.append({\n",
" \"id\": graph_id,\n",
" \"type\": mention.label,\n",
" \"name\": mention.text,\n",
" \"properties\": {},\n",
" \"id\": f\"e{i}\",\n",
" \"type\": entity.label,\n",
" \"name\": entity.text,\n",
" \"properties\": {}\n",
" })\n",
"\n",
"relationships = []\n",
"for rel in relations:\n",
" source_id = span_to_id.get((rel.subject.start_char, rel.subject.end_char))\n",
" target_id = span_to_id.get((rel.object.start_char, rel.object.end_char))\n",
" if source_id is None or target_id is None:\n",
" print(f\"Skipping relation with unmapped endpoint: \"\n",
" f\"{rel.subject.text!r} --{rel.predicate}--> {rel.object.text!r}\")\n",
" continue\n",
"for i, rel in enumerate(relationships_list[:3], 1):\n",
" relationships.append({\n",
" \"source\": source_id,\n",
" \"target\": target_id,\n",
" \"source\": f\"e{1}\",\n",
" \"target\": f\"e{i+1}\",\n",
" \"type\": rel.predicate,\n",
" \"properties\": {},\n",
" \"properties\": {}\n",
" })\n",
"\n",
"builder = GraphBuilder()\n",
"knowledge_graph = builder.build({\"entities\": entities, \"relationships\": relationships})\n",
"knowledge_graph = builder.build(entities, relationships)\n",
"\n",
"id_to_name = {entity[\"id\"]: entity[\"name\"] for entity in entities}\n",
"\n",
"print(f\"Graph entities ({len(knowledge_graph['entities'])}):\")\n",
"for entity in knowledge_graph[\"entities\"]:\n",
" print(f\" {entity['id']}: {entity['name']} ({entity['type']})\")\n",
"\n",
"print(f\"\\nGraph relationships ({len(knowledge_graph['relationships'])}):\")\n",
"for relationship in knowledge_graph[\"relationships\"]:\n",
" print(f\" {id_to_name[relationship['source']]} \"\n",
" f\"--{relationship['type']}--> {id_to_name[relationship['target']]}\")\n",
"\n",
"edges = {\n",
" (id_to_name[r[\"source\"]], r[\"type\"], id_to_name[r[\"target\"]])\n",
" for r in knowledge_graph[\"relationships\"]\n",
"}\n",
"assert (\"Apple Inc.\", \"located_in\", \"Cupertino\") in edges\n",
"assert (\"Tim Cook\", \"works_for\", \"Apple Inc.\") in edges"
],
"execution_count": null,
"outputs": []
"print(f\"Built knowledge graph with {len(knowledge_graph.get('entities', []))} entities\")\n",
"print(f\"Relationships: {len(knowledge_graph.get('relationships', []))}\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 3: Entity Resolution\n",
"## Step 2: Entity Resolution\n",
"\n",
"The graph currently contains two nodes for the same organization. `EntityResolver` merges duplicate mentions into one canonical entity and records which source IDs were merged (`merged_from`), so relationship endpoints can be remapped onto the canonical entity.\n"
"Resolve entity conflicts and duplicates.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from semantica.kg import EntityResolver\n",
"\n",
"entity_resolver = EntityResolver()\n",
"\n",
"resolved_entities = entity_resolver.resolve_entities(entities)\n",
"\n",
"canonical_id = {}\n",
"for entity in resolved_entities:\n",
" for source_id in entity.get(\"merged_from\", [entity[\"id\"]]):\n",
" canonical_id[source_id] = entity[\"id\"]\n",
" if entity.get(\"merged_from\"):\n",
" print(f\"Merged {entity['merged_from']} -> {entity['id']}: {entity['name']}\")\n",
"\n",
"print(f\"\\nMentions in: {len(entities)}, resolved entities out: {len(resolved_entities)}\")\n",
"\n",
"resolved_names = {entity[\"id\"]: entity[\"name\"] for entity in resolved_entities}\n",
"print(\"\\nRelationships remapped onto canonical entities:\")\n",
"for relationship in relationships:\n",
" source = canonical_id[relationship[\"source\"]]\n",
" target = canonical_id[relationship[\"target\"]]\n",
" print(f\" {resolved_names[source]} --{relationship['type']}--> {resolved_names[target]}\")\n",
"\n",
"canonical_entities = {(entity[\"name\"], entity[\"type\"]) for entity in resolved_entities}\n",
"assert canonical_entities == {\n",
" (\"Apple Inc.\", \"ORG\"),\n",
" (\"Tim Cook\", \"PERSON\"),\n",
" (\"Cupertino\", \"GPE\"),\n",
" (\"California\", \"GPE\"),\n",
"}\n",
"assert len(resolved_entities) == 4"
],
"execution_count": null,
"outputs": []
"print(f\"Original entities: {len(entities)}\")\n",
"print(f\"Resolved entities: {len(resolved_entities)}\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 4: Deduplication\n",
"## Step 3: Deduplication\n",
"\n",
"The `semantica.deduplication` module gives finer control over the same problem. Note that `merge_duplicates` returns one `MergeOperation` per duplicate *group* — the complete deduplicated collection is those merged entities plus every entity that was not part of any group.\n"
"Remove duplicate entities from the graph.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from semantica.deduplication import DuplicateDetector, EntityMerger, MergeStrategy\n",
"\n",
"# Detect duplicates\n",
"detector = DuplicateDetector(similarity_threshold=0.8)\n",
"duplicate_groups = detector.detect_duplicate_groups(entities)\n",
"print(f\"Duplicate groups: {len(duplicate_groups)}\")\n",
"for group in duplicate_groups:\n",
" print(f\" {[entity['name'] for entity in group.entities]} \"\n",
" f\"(confidence={group.confidence:.2f})\")\n",
"duplicate_groups = detector.detect_duplicate_groups(knowledge_graph.get('entities', []))\n",
"\n",
"# Merge duplicates\n",
"merger = EntityMerger()\n",
"merge_operations = merger.merge_duplicates(\n",
" entities, strategy=MergeStrategy.KEEP_MOST_COMPLETE\n",
" knowledge_graph.get('entities', []),\n",
" strategy=MergeStrategy.KEEP_MOST_COMPLETE\n",
")\n",
"\n",
"merged_source_ids = {\n",
" entity[\"id\"] for op in merge_operations for entity in op.source_entities\n",
"}\n",
"untouched_entities = [e for e in entities if e[\"id\"] not in merged_source_ids]\n",
"deduplicated_entities = untouched_entities + [\n",
" op.merged_entity for op in merge_operations\n",
"]\n",
"deduplicated_entities = [op.merged_entity for op in merge_operations]\n",
"\n",
"print(f\"\\nMerge operations: {len(merge_operations)}\")\n",
"print(f\"Deduplicated entities ({len(deduplicated_entities)}):\")\n",
"for entity in deduplicated_entities:\n",
" print(f\" {entity['id']}: {entity['name']} ({entity['type']})\")\n",
"\n",
"assert len(merge_operations) == 1\n",
"assert len(deduplicated_entities) == 4"
],
"execution_count": null,
"outputs": []
"print(f\"Original entities: {len(knowledge_graph.get('entities', []))}\")\n",
"print(f\"Deduplicated entities: {len(deduplicated_entities)}\")\n"
]
},
{
"cell_type": "markdown",
@@ -253,10 +155,9 @@
"\n",
"You've learned how to build knowledge graphs:\n",
"\n",
"- **Extraction to graph**: map each mention to a graph ID and build edges from the actual `Relation.subject` / `Relation.object` endpoints\n",
"- **GraphBuilder**: construct knowledge graphs from explicit `{\"entities\": ..., \"relationships\": ...}` input\n",
"- **EntityResolver**: merge duplicate mentions into canonical entities and remap relationship endpoints\n",
"- **Deduplication**: combine `MergeOperation` results with untouched entities to get the complete deduplicated set\n",
"- **GraphBuilder**: Construct knowledge graphs from entities and relationships\n",
"- **EntityResolver**: Resolve entity conflicts and duplicates\n",
"- **Deduplication**: Use `semantica.deduplication` module for removing duplicate entities\n",
"\n",
"Next: Learn how to analyze graphs in the Graph_Analytics notebook.\n"
]
@@ -10,7 +10,7 @@
"\n",
"## Overview\n",
"\n",
"This notebook walks you through creating your first knowledge graph from a simple document. You'll learn the complete end-to-end workflow from ingesting a file to visualizing the resulting knowledge graph — and every step consumes the real output of the step before it.\n",
"This notebook walks you through creating your first knowledge graph from a simple document. You'll learn the complete end-to-end workflow from ingesting a file to visualizing the resulting knowledge graph.\n",
"\n",
"> [!TIP]\n",
"> This is the perfect starting point if you are new to Semantica. No prior knowledge of knowledge graphs is required!\n",
@@ -19,10 +19,10 @@
"\n",
"### 🎯 Learning Objectives\n",
"\n",
"- **Understand the Workflow**: Learn the `File → Parse → Extract → Graph → Visualize` pipeline\n",
"- **Understand the Workflow**: Learn the `File → Parse → Extract → Graph` pipeline\n",
"- **Ingest Data**: Load documents using `FileIngestor`\n",
"- **Parse Content**: Extract text using `DocumentParser`\n",
"- **Extract Knowledge**: Identify entities and relations using `NERExtractor` and `RelationExtractor`\n",
"- **Extract Knowledge**: Identify entities using `NERExtractor`\n",
"- **Build Graph**: Construct a graph using `GraphBuilder`\n",
"- **Visualize**: See your graph come to life with `KGVisualizer`\n",
"\n",
@@ -40,76 +40,71 @@
"\n",
"## 🔄 Simple End-to-End Workflow\n",
"\n",
"The complete workflow consists of five main steps:\n",
"The complete workflow consists of four main steps:\n",
"\n",
"1. **📥 Ingest** - Load data from files or other sources\n",
"2. **📄 Parse** - Extract and structure content from documents\n",
"3. **⛏️ Extract** - Identify entities and relationships\n",
"4. **🕸️ Build Graph** - Construct the knowledge graph\n",
"5. **📊 Visualize** - Render and analyze the graph\n",
"\n",
"Each step is demonstrated in the code cells below, and each cell can be rerun on its own: the sample file is only removed by the optional cleanup cell at the very end.\n",
"Each step is demonstrated in the code cells below.\n",
"\n",
"> [!TIP]\n",
"> **Alternative: Using Semantica Framework**\n",
">\n",
"> \n",
"> For a simpler, high-level approach, you can use the `Semantica` framework class which orchestrates all these steps:\n",
">\n",
"> \n",
"> ```python\n",
"> from semantica.core import Semantica\n",
">\n",
"> \n",
"> framework = Semantica()\n",
"> framework.initialize()\n",
">\n",
"> \n",
"> result = framework.build_knowledge_base(\n",
"> sources=[\"sample_document.txt\"],\n",
"> embeddings=True,\n",
"> graph=True\n",
"> )\n",
">\n",
"> \n",
"> framework.shutdown()\n",
"> ```\n",
">\n",
"> \n",
"> This notebook shows the step-by-step approach for learning. See [Core Module Usage Guide](../../../semantica/core/core_usage.md) for more details.\n",
"\n",
"---\n",
"\n",
"## 📂 Step 1: Ingest a File\n",
"\n",
"In this step, we'll use `FileIngestor` to load a document. The ingestor supports various file formats including PDF, DOCX, TXT, and more. Writing the sample file is idempotent, so this cell can be rerun at any time.\n"
"In this step, we'll use `FileIngestor` to load a document. The ingestor supports various file formats including PDF, DOCX, TXT, and more.\n"
]
},
{
"cell_type": "code",
"metadata": {},
"source": [
"%pip install semantica\n",
"\n",
"# spaCy models are distributed separately from the spaCy library. This lesson\n",
"# relies on the English model to recognize standalone places such as Cupertino.\n",
"import sys\n",
"import subprocess\n",
"import spacy\n",
"\n",
"try:\n",
" spacy.load(\"en_core_web_sm\")\n",
"except OSError:\n",
" subprocess.check_call([sys.executable, \"-m\", \"spacy\", \"download\", \"en_core_web_sm\"])\n"
],
"execution_count": null,
"outputs": []
"metadata": {},
"outputs": [],
"source": [
"!pip install semantica"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from semantica.ingest import FileIngestor\n",
"from pathlib import Path\n",
"\n",
"from semantica.ingest import FileIngestor\n",
"# Initialize the ingestor\n",
"ingestor = FileIngestor()\n",
"\n",
"sample_text = \"\"\"Apple Inc. is headquartered in Cupertino, California.\n",
"In 1976, Steve Jobs founded Apple Inc.\n",
"Tim Cook is the CEO of Apple Inc.\n",
"# Create a sample document for demonstration\n",
"sample_text = \"\"\"\n",
"Apple Inc. is a technology company founded by Steve Jobs, Steve Wozniak, and Ronald Wayne in 1976.\n",
"The company is headquartered in Cupertino, California.\n",
"Tim Cook is the current CEO of Apple Inc.\n",
"Apple designs and manufactures consumer electronics, software, and online services.\n",
"\"\"\"\n",
"\n",
"sample_file = Path(\"sample_document.txt\")\n",
@@ -118,14 +113,12 @@
"print(f\"File: {sample_file}\")\n",
"print(f\"Content length: {len(sample_text)} characters\")\n",
"\n",
"ingestor = FileIngestor()\n",
"# Ingest the file\n",
"file_object = ingestor.ingest_file(sample_file, read_content=True)\n",
"print(f\" File name: {file_object.name}\")\n",
"print(f\" File type: {file_object.file_type}\")\n",
"print(f\" Content available: {file_object.content is not None}\")"
],
"execution_count": null,
"outputs": []
"print(f\" Content available: {file_object.content is not None}\")\n"
]
},
{
"cell_type": "markdown",
@@ -133,58 +126,64 @@
"source": [
"## 📄 Step 2: Parse the Document\n",
"\n",
"After ingesting the file, we need to parse it to extract the text content. `DocumentParser.parse_document()` returns the extracted text under the `\"text\"` key.\n"
"After ingesting the file, we need to parse it to extract the text content. The `DocumentParser` handles various file formats and extracts structured content.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from semantica.parse import DocumentParser\n",
"\n",
"parser = DocumentParser()\n",
"# Parse the document to extract text\n",
"parsed_document = parser.parse_document(str(sample_file))\n",
"\n",
"parsed_content = parsed_document.get(\"text\", \"\")\n",
"assert parsed_content.strip(), \"Parsing produced no text — check the input file\"\n",
"\n",
"print(f\"Parsed content length: {len(parsed_content)} characters\")\n",
"print(f\"Preview: {parsed_content[:120]}...\")"
],
"execution_count": null,
"outputs": []
"parsed_content = parsed_document.get(\"content\", \"\")\n",
"print(f\" Parsed content length: {len(parsed_content) if parsed_content else 0} characters\")\n",
"print(f\" Preview: {parsed_content[:200] if parsed_content else 'N/A'}...\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## ⛏️ Step 3: Extract Entities and Relations\n",
"## ⛏️ Step 3: Extract Entities\n",
"\n",
"Now we'll extract entities and relations from the parsed text. `NERExtractor` identifies people, organizations, locations and dates; `RelationExtractor` finds relations between those mentions. Both operate on the *parsed content from Step 2* — not on a copy of the raw string.\n"
"Now we'll extract entities from the parsed text using Named Entity Recognition (NER). This identifies people, organizations, locations, dates, and other entities in the text.\n",
"\n",
"> [!NOTE]\n",
"> In a real scenario, you would use `NERExtractor` with an LLM or model backend. Here we simulate the output for demonstration purposes.\n"
]
},
{
"cell_type": "code",
"metadata": {},
"source": [
"from semantica.semantic_extract import NERExtractor, RelationExtractor\n",
"\n",
"ner_extractor = NERExtractor()\n",
"relation_extractor = RelationExtractor()\n",
"\n",
"mentions = ner_extractor.extract(parsed_content)\n",
"relations = relation_extractor.extract(parsed_content, mentions)\n",
"\n",
"print(\"Entity mentions:\")\n",
"for mention in mentions:\n",
" print(f\" {mention.text!r:<13} {mention.label:<7} span=[{mention.start_char}:{mention.end_char}]\")\n",
"\n",
"print(\"\\nExtracted relations:\")\n",
"for rel in relations:\n",
" print(f\" {rel.subject.text!r} --{rel.predicate}--> {rel.object.text!r}\")"
],
"execution_count": null,
"outputs": []
"metadata": {},
"outputs": [],
"source": [
"from semantica.semantic_extract import NamedEntityRecognizer, NERExtractor\n",
"\n",
"ner = NamedEntityRecognizer()\n",
"extractor = NERExtractor()\n",
"\n",
"print(f\"\\nText: {parsed_content[:100]}...\")\n",
"\n",
"# Simulated extraction results\n",
"expected_entities = [\n",
" {\"text\": \"Apple Inc.\", \"type\": \"Organization\", \"start\": 0, \"end\": 10},\n",
" {\"text\": \"Steve Jobs\", \"type\": \"Person\", \"start\": 50, \"end\": 60},\n",
" {\"text\": \"Steve Wozniak\", \"type\": \"Person\", \"start\": 62, \"end\": 75},\n",
" {\"text\": \"Ronald Wayne\", \"type\": \"Person\", \"start\": 81, \"end\": 93},\n",
" {\"text\": \"1976\", \"type\": \"Date\", \"start\": 97, \"end\": 101},\n",
" {\"text\": \"Cupertino, California\", \"type\": \"Location\", \"start\": 130, \"end\": 151},\n",
" {\"text\": \"Tim Cook\", \"type\": \"Person\", \"start\": 153, \"end\": 161},\n",
"]\n",
"\n",
"for entity in expected_entities:\n",
" print(f\" - {entity['text']} ({entity['type']})\")\n"
]
},
{
"cell_type": "markdown",
@@ -192,68 +191,58 @@
"source": [
"## 🕸️ Step 4: Build the Knowledge Graph\n",
"\n",
"Using the extracted entities and relations, we construct a knowledge graph with `GraphBuilder`. Every mention gets a graph ID, and each edge is built from the actual `Relation.subject` / `Relation.object` endpoints.\n",
"\n",
"> [!NOTE]\n",
"> The graph will contain one node per *mention*, so `Apple Inc.` appears three times. Merging duplicate mentions into one canonical entity is covered in [07_Building_Knowledge_Graphs.ipynb](./07_Building_Knowledge_Graphs.ipynb).\n"
"Using the extracted entities and relationships, we'll construct a knowledge graph. The graph represents entities as nodes and relationships as edges.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from semantica.kg import GraphBuilder\n",
"\n",
"entities = []\n",
"span_to_id = {}\n",
"for i, mention in enumerate(mentions, 1):\n",
" graph_id = f\"e{i}\"\n",
" span_to_id[(mention.start_char, mention.end_char)] = graph_id\n",
" entities.append({\n",
" \"id\": graph_id,\n",
" \"type\": mention.label,\n",
" \"name\": mention.text,\n",
" \"properties\": {},\n",
" })\n",
"\n",
"relationships = []\n",
"for rel in relations:\n",
" source_id = span_to_id.get((rel.subject.start_char, rel.subject.end_char))\n",
" target_id = span_to_id.get((rel.object.start_char, rel.object.end_char))\n",
" if source_id is None or target_id is None:\n",
" print(f\"Skipping relation with unmapped endpoint: \"\n",
" f\"{rel.subject.text!r} --{rel.predicate}--> {rel.object.text!r}\")\n",
" continue\n",
" relationships.append({\n",
" \"source\": source_id,\n",
" \"target\": target_id,\n",
" \"type\": rel.predicate,\n",
" \"properties\": {},\n",
" })\n",
"import networkx as nx\n",
"\n",
"builder = GraphBuilder()\n",
"knowledge_graph = builder.build({\"entities\": entities, \"relationships\": relationships})\n",
"\n",
"id_to_name = {entity[\"id\"]: entity[\"name\"] for entity in entities}\n",
"# Prepare data for graph construction\n",
"entities_data = [\n",
" {\"id\": f\"entity_{i}\", \"name\": entity[\"text\"], \"type\": entity[\"type\"]}\n",
" for i, entity in enumerate(expected_entities)\n",
"]\n",
"\n",
"print(f\"Nodes (entities): {len(knowledge_graph['entities'])}\")\n",
"for entity in knowledge_graph[\"entities\"]:\n",
" print(f\" {entity['id']}: {entity['name']} ({entity['type']})\")\n",
"relationships_data = [\n",
" {\"source\": \"entity_0\", \"target\": \"entity_1\", \"type\": \"founded_by\"},\n",
" {\"source\": \"entity_0\", \"target\": \"entity_2\", \"type\": \"founded_by\"},\n",
" {\"source\": \"entity_0\", \"target\": \"entity_3\", \"type\": \"founded_by\"},\n",
" {\"source\": \"entity_0\", \"target\": \"entity_4\", \"type\": \"founded_in\"},\n",
" {\"source\": \"entity_0\", \"target\": \"entity_5\", \"type\": \"located_in\"},\n",
" {\"source\": \"entity_6\", \"target\": \"entity_0\", \"type\": \"ceo_of\"},\n",
"]\n",
"\n",
"print(f\"\\nEdges (relationships): {len(knowledge_graph['relationships'])}\")\n",
"for relationship in knowledge_graph[\"relationships\"]:\n",
" print(f\" {id_to_name[relationship['source']]} \"\n",
" f\"--{relationship['type']}--> {id_to_name[relationship['target']]}\")\n",
"# Build the graph using NetworkX\n",
"kg = nx.DiGraph()\n",
"\n",
"edges = {\n",
" (id_to_name[r[\"source\"]], r[\"type\"], id_to_name[r[\"target\"]])\n",
" for r in knowledge_graph[\"relationships\"]\n",
"}\n",
"assert (\"Apple Inc.\", \"located_in\", \"Cupertino\") in edges\n",
"assert (\"Tim Cook\", \"works_for\", \"Apple Inc.\") in edges"
],
"execution_count": null,
"outputs": []
"for entity in entities_data:\n",
" kg.add_node(entity[\"id\"], name=entity[\"name\"], type=entity[\"type\"])\n",
"\n",
"for rel in relationships_data:\n",
" source_name = entities_data[int(rel[\"source\"].split(\"_\")[1])][\"name\"]\n",
" target_name = entities_data[int(rel[\"target\"].split(\"_\")[1])][\"name\"]\n",
" kg.add_edge(rel[\"source\"], rel[\"target\"], type=rel[\"type\"])\n",
"\n",
"print(f\" Nodes (entities): {len(kg.nodes)}\")\n",
"print(f\" Edges (relationships): {len(kg.edges)}\")\n",
"\n",
"for node_id in kg.nodes():\n",
" node_data = kg.nodes[node_id]\n",
" print(f\" Node: {node_data['name']} ({node_data['type']})\")\n",
"\n",
"for source, target, data in kg.edges(data=True):\n",
" source_name = kg.nodes[source]['name']\n",
" target_name = kg.nodes[target]['name']\n",
" print(f\" {source_name} --[{data['type']}]--> {target_name}\")\n"
]
},
{
"cell_type": "markdown",
@@ -261,81 +250,49 @@
"source": [
"## 📊 Step 5: Visualize and Analyze\n",
"\n",
"Finally, we render the knowledge graph with `KGVisualizer` and look at its structure. `visualize_network()` accepts the `GraphBuilder` result directly and can save an interactive HTML file.\n"
"Finally, we'll visualize the knowledge graph and analyze its structure. This helps you understand the relationships and entities in your data.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from semantica.visualization import KGVisualizer\n",
"\n",
"visualizer = KGVisualizer()\n",
"fig = visualizer.visualize_network(\n",
" knowledge_graph, output=\"html\", file_path=\"knowledge_graph.html\"\n",
")\n",
"print(\"Saved interactive visualization to knowledge_graph.html\")\n",
"\n",
"print(f\" Total entities: {len(kg.nodes)}\")\n",
"print(f\" Total relationships: {len(kg.edges)}\")\n",
"\n",
"entity_types = {}\n",
"for entity in knowledge_graph[\"entities\"]:\n",
" entity_types[entity[\"type\"]] = entity_types.get(entity[\"type\"], 0) + 1\n",
"for node_id in kg.nodes():\n",
" entity_type = kg.nodes[node_id]['type']\n",
" entity_types[entity_type] = entity_types.get(entity_type, 0) + 1\n",
"\n",
"print(\"\\nEntities by type:\")\n",
"for entity_type, count in sorted(entity_types.items()):\n",
" print(f\" - {entity_type}: {count}\")\n",
"for etype, count in entity_types.items():\n",
" print(f\" - {etype}: {count}\")\n",
"\n",
"relationship_types = {}\n",
"for relationship in knowledge_graph[\"relationships\"]:\n",
" relationship_types[relationship[\"type\"]] = (\n",
" relationship_types.get(relationship[\"type\"], 0) + 1\n",
" )\n",
"rel_types = {}\n",
"for _, _, data in kg.edges(data=True):\n",
" rel_type = data.get('type', 'unknown')\n",
" rel_types[rel_type] = rel_types.get(rel_type, 0) + 1\n",
"\n",
"print(\"\\nRelationships by type:\")\n",
"for relationship_type, count in sorted(relationship_types.items()):\n",
" print(f\" - {relationship_type}: {count}\")\n",
"for rtype, count in rel_types.items():\n",
" print(f\" - {rtype}: {count}\")\n",
"\n",
"fig"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## 🧹 Optional: Clean Up\n",
"\n",
"Run this cell only when you are done with the notebook. Earlier cells read `sample_document.txt`, so they stay rerunnable until you delete it here.\n"
"# Cleanup\n",
"if sample_file.exists():\n",
" sample_file.unlink()\n"
]
},
{
"cell_type": "code",
"metadata": {},
"source": [
"for path in [sample_file, Path(\"knowledge_graph.html\")]:\n",
" if path.exists():\n",
" path.unlink()\n",
" print(f\"Removed {path}\")"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Summary\n",
"\n",
"You've built your first knowledge graph, end to end:\n",
"\n",
"- **FileIngestor** loaded the sample document\n",
"- **DocumentParser** returned its text under the `\"text\"` key\n",
"- **NERExtractor** / **RelationExtractor** produced real mentions and relations from that text\n",
"- **GraphBuilder** turned them into a graph whose edges come from the actual relation endpoints\n",
"- **KGVisualizer** rendered the result as an interactive network\n",
"\n",
"Next: merge duplicate mentions with `EntityResolver` in [07_Building_Knowledge_Graphs.ipynb](./07_Building_Knowledge_Graphs.ipynb), or explore graph metrics in the Graph Analytics notebook.\n"
]
"outputs": [],
"source": []
}
],
"metadata": {
+1 -2
View File
@@ -497,8 +497,7 @@
"**Next Steps**:\n",
"* Try customizing the `NamespaceManager` to use your organization's URL.\n",
"* Explore `OntologyEvaluator` for deeper quality metrics.\n",
"* Feed the generated ontology into the **Knowledge Graph** module to start reasoning over your data!\n",
"* Put the graph, ontology, and explicit mappings together in [Semantic Layer Basics](./26_Semantic_Layer_Basics.ipynb)."
"* Feed the generated ontology into the **Knowledge Graph** module to start reasoning over your data!"
]
}
],
@@ -1,418 +0,0 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/semantica-agi/semantica/blob/main/cookbook/introduction/26_Semantic_Layer_Basics.ipynb)\n",
"\n",
"# Semantic Layer Basics: Putting the Knowledge Graph, Ontology, and Mappings Together\n",
"\n",
"## Overview\n",
"\n",
"This lesson connects three things you have already met — a knowledge graph, an ontology, and RDF export — into one minimal *semantic layer*: a knowledge graph whose types, relationships, and properties are **explicitly mapped** to ontology terms, so the resulting RDF can be queried with SPARQL against a shared vocabulary.\n",
"\n",
"**Documentation**: [API Reference](https://semantica.readthedocs.io/concepts/)\n",
"\n",
"### 🎯 Learning Objectives\n",
"\n",
"- Build a small knowledge graph with `GraphBuilder`\n",
"- Generate a starter ontology from the graph with `OntologyGenerator`\n",
"- Write **explicit** entity-type, relationship-type, and property mappings to ontology terms\n",
"- Produce ontology-aligned RDF and store it with `TripletStore`\n",
"- Answer a business question with one small SPARQL query\n",
"\n",
"### 📚 Prerequisites\n",
"\n",
"- [07_Building_Knowledge_Graphs.ipynb](./07_Building_Knowledge_Graphs.ipynb) — graphs from entities and relationships\n",
"- [14_Ontology.ipynb](./14_Ontology.ipynb) — ontology generation\n",
"- [20_Triplet_Store.ipynb](./20_Triplet_Store.ipynb) — triplet store backends\n",
"\n",
"> [!NOTE]\n",
"> **Teaching mappings vs. governed mappings.** The mappings in this lesson are a demo: they live in a Python dict and are derived from a generated ontology. A production semantic layer uses governed identifiers, hand-designed ontologies, explicit source mappings, validation (SHACL), provenance, and versioning — that workflow is covered in [Advanced: Manual Ontology + Snowflake Mapping](../advanced/13_Manual_Ontology_Snowflake_Mapping.ipynb).\n",
"\n",
"## Installation\n",
"\n",
"The triplet-store step uses the embedded Oxigraph backend, so install with that extra. Pin at least 0.6.7: earlier releases could generate ontology classes with no URI (#1103), which silently breaks the mappings below instead of failing loudly.\n",
"\n",
"```bash\n",
"pip install \"semantica[tripletstore-oxigraph]>=0.6.7\"\n",
"```\n",
"\n",
"---\n",
"\n",
"## Step 1: Build a Knowledge Graph\n",
"\n",
"Start from a small, explicit set of entities and relationships — two people, an organization, and a project.\n"
]
},
{
"cell_type": "code",
"metadata": {},
"source": [
"!pip install \"semantica[tripletstore-oxigraph]>=0.6.7\"\n"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "code",
"metadata": {},
"source": [
"from semantica.kg import GraphBuilder\n",
"\n",
"entities = [\n",
" {\"id\": \"e1\", \"type\": \"Person\", \"name\": \"Alice\", \"properties\": {\"age\": 30, \"role\": \"Engineer\"}},\n",
" {\"id\": \"e2\", \"type\": \"Person\", \"name\": \"Bob\", \"properties\": {\"age\": 35, \"role\": \"Manager\"}},\n",
" {\"id\": \"e3\", \"type\": \"Organization\", \"name\": \"Tech Corp\", \"properties\": {\"founded\": 2010}},\n",
" {\"id\": \"e4\", \"type\": \"Project\", \"name\": \"Project Alpha\", \"properties\": {\"status\": \"active\"}},\n",
"]\n",
"\n",
"relationships = [\n",
" {\"source\": \"e1\", \"target\": \"e2\", \"type\": \"reports_to\", \"properties\": {}},\n",
" {\"source\": \"e1\", \"target\": \"e3\", \"type\": \"works_for\", \"properties\": {}},\n",
" {\"source\": \"e2\", \"target\": \"e3\", \"type\": \"works_for\", \"properties\": {}},\n",
" {\"source\": \"e1\", \"target\": \"e4\", \"type\": \"works_on\", \"properties\": {}},\n",
"]\n",
"\n",
"builder = GraphBuilder()\n",
"knowledge_graph = builder.build({\"entities\": entities, \"relationships\": relationships})\n",
"\n",
"id_to_name = {entity[\"id\"]: entity[\"name\"] for entity in entities}\n",
"\n",
"print(f\"Entities ({len(knowledge_graph['entities'])}):\")\n",
"for entity in knowledge_graph[\"entities\"]:\n",
" print(f\" {entity['id']}: {entity['name']} ({entity['type']}) {entity['properties']}\")\n",
"\n",
"print(f\"\\nRelationships ({len(knowledge_graph['relationships'])}):\")\n",
"for relationship in knowledge_graph[\"relationships\"]:\n",
" print(f\" {id_to_name[relationship['source']]} \"\n",
" f\"--{relationship['type']}--> {id_to_name[relationship['target']]}\")"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 2: Generate a Starter Ontology\n",
"\n",
"`OntologyGenerator` infers OWL classes and properties from graph records. Because `GraphBuilder` keeps business attributes inside each entity's `properties` dictionary while ontology inference reads record fields, we first create a flat **inference view**. The knowledge graph itself remains unchanged. Two settings matter here:\n",
"\n",
"- `base_uri` puts every generated term in *your* namespace\n",
"- `min_occurrences=1` includes classes that occur only once (the default of 2 would drop `Organization` and `Project` from this tiny demo graph)\n",
"\n",
"Note that the generator normalizes names: the relationship type `works_for` becomes the ontology property `worksFor`. That is exactly why the next step maps terms **explicitly** instead of matching names.\n"
]
},
{
"cell_type": "code",
"metadata": {},
"source": [
"from semantica.ontology import OntologyGenerator\n",
"\n",
"BASE_URI = \"https://example.org/company/\"\n",
"\n",
"# Adapt the property-graph representation to the record shape consumed by\n",
"# OntologyGenerator, so age/role/founded/status become declared properties.\n",
"ontology_input = {\n",
" \"entities\": [\n",
" {\n",
" **{key: value for key, value in entity.items() if key != \"properties\"},\n",
" **entity.get(\"properties\", {}),\n",
" }\n",
" for entity in knowledge_graph[\"entities\"]\n",
" ],\n",
" \"relationships\": knowledge_graph[\"relationships\"],\n",
"}\n",
"\n",
"generator = OntologyGenerator(base_uri=BASE_URI, min_occurrences=1)\n",
"ontology = generator.generate_from_graph(ontology_input)\n",
"\n",
"# OntologyGenerator calls datatype properties `data`; TripletStore's public\n",
"# ontology contract calls them `datatype`. Normalize that boundary explicitly.\n",
"store_ontology = {\n",
" **ontology,\n",
" \"properties\": [\n",
" {**prop, \"type\": \"datatype\" if prop[\"type\"] == \"data\" else prop[\"type\"]}\n",
" for prop in ontology[\"properties\"]\n",
" ],\n",
"}\n",
"\n",
"print(\"Classes:\")\n",
"for ontology_class in ontology[\"classes\"]:\n",
" print(f\" {ontology_class['name']:<14} {ontology_class['uri']}\")\n",
"\n",
"print(\"\\nProperties:\")\n",
"for prop in ontology[\"properties\"]:\n",
" print(f\" {prop['name']:<14} {prop['type']:<7} {prop['uri']} \"\n",
" f\"(domain={prop['domain']}, range={prop['range']})\")\n",
"\n",
"assert len(ontology[\"classes\"]) == 3"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 3: Map the Graph to Ontology Terms\n",
"\n",
"The heart of a semantic layer is the mapping contract: which source type, relationship, and property corresponds to which ontology term.\n",
"\n",
"- **Entity types** and **relationship types**: each generated class/property records the source name it was inferred from (`metadata[\"inferred_from\"]`), so the mapping is read off the ontology itself — no fragile name matching between `works_for` and `worksFor`.\n",
"- **Properties**: the flat inference view makes `name`, `age`, `role`, `founded`, and `status` real generated datatype properties. Every mapping therefore points to a term declared in the ontology — no URI is invented only at mapping time.\n"
]
},
{
"cell_type": "code",
"metadata": {},
"source": [
"entity_type_mappings = {\n",
" ontology_class[\"metadata\"][\"inferred_from\"]: ontology_class[\"uri\"]\n",
" for ontology_class in ontology[\"classes\"]\n",
"}\n",
"\n",
"relationship_type_mappings = {\n",
" prop[\"metadata\"][\"inferred_from\"]: prop[\"uri\"]\n",
" for prop in ontology[\"properties\"]\n",
" if prop[\"type\"] == \"object\"\n",
"}\n",
"\n",
"datatype_property_uris = {\n",
" prop[\"metadata\"][\"inferred_from\"]: prop[\"uri\"]\n",
" for prop in ontology[\"properties\"]\n",
" if prop[\"type\"] != \"object\"\n",
"}\n",
"\n",
"property_mappings = datatype_property_uris\n",
"\n",
"semantic_layer = {\n",
" \"graph\": knowledge_graph,\n",
" \"ontology\": ontology,\n",
" \"mappings\": {\n",
" \"entity_type_mappings\": entity_type_mappings,\n",
" \"relationship_type_mappings\": relationship_type_mappings,\n",
" \"property_mappings\": property_mappings,\n",
" },\n",
"}\n",
"\n",
"for mapping_name, mapping in semantic_layer[\"mappings\"].items():\n",
" print(f\"{mapping_name}:\")\n",
" for source, target in mapping.items():\n",
" print(f\" {source:<12} -> {target}\")\n",
"\n",
"# Every type and relationship in the graph must have an ontology term\n",
"assert set(entity_type_mappings) == {entity[\"type\"] for entity in entities}\n",
"assert set(relationship_type_mappings) == {rel[\"type\"] for rel in relationships}\n",
"assert set(property_mappings) == {\"name\", \"age\", \"role\", \"founded\", \"status\"}\n",
"assert set(property_mappings.values()) <= {prop[\"uri\"] for prop in ontology[\"properties\"]}"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 4: Apply the Mappings\n",
"\n",
"Applying the semantic layer means rewriting the graph so every type, relationship, and property key is an ontology term. This *aligned* graph — not the original one — is what gets exported and stored.\n"
]
},
{
"cell_type": "code",
"metadata": {},
"source": [
"aligned_graph = {\n",
" \"entities\": [\n",
" {\n",
" **entity,\n",
" \"type\": entity_type_mappings[entity[\"type\"]],\n",
" \"properties\": {\n",
" property_mappings[\"name\"]: entity[\"name\"],\n",
" **{\n",
" property_mappings[key]: value\n",
" for key, value in entity[\"properties\"].items()\n",
" },\n",
" },\n",
" }\n",
" for entity in knowledge_graph[\"entities\"]\n",
" ],\n",
" \"relationships\": [\n",
" {**rel, \"type\": relationship_type_mappings[rel[\"type\"]]}\n",
" for rel in knowledge_graph[\"relationships\"]\n",
" ],\n",
"}\n",
"\n",
"print(\"Aligned entity sample:\")\n",
"sample = aligned_graph[\"entities\"][0]\n",
"print(f\" id: {sample['id']}\")\n",
"print(f\" type: {sample['type']}\")\n",
"for key, value in sample[\"properties\"].items():\n",
" print(f\" {key} = {value}\")\n",
"\n",
"print(\"\\nAligned relationship sample:\")\n",
"print(f\" {aligned_graph['relationships'][0]['type']}\")"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 5: Store and Export Complete Ontology-Aligned RDF\n",
"\n",
"`TripletStore.store()` materializes both the ontology declarations and the aligned instance graph. We then read those triples through the store's public API and serialize that complete RDF graph as Turtle. This avoids the compact `RDFExporter` entity projection, which does not include arbitrary entries from an entity's `properties` dictionary.\n"
]
},
{
"cell_type": "code",
"metadata": {},
"source": [
"from rdflib import Graph, Literal, URIRef\n",
"from rdflib.namespace import OWL, RDF\n",
"from semantica.triplet_store import TripletStore\n",
"\n",
"store = TripletStore(backend=\"oxigraph\")\n",
"result = store.store(aligned_graph, store_ontology)\n",
"print(f\"Stored triples: {result['processed']} (failed: {result['failed']})\")\n",
"\n",
"rdf_graph = Graph()\n",
"for triplet in store.get_triplets():\n",
" datatype = triplet.metadata.get(\"datatype\")\n",
" if datatype:\n",
" object_term = Literal(triplet.object, datatype=URIRef(datatype))\n",
" elif triplet.object.startswith((\"http://\", \"https://\", \"urn:\")):\n",
" object_term = URIRef(triplet.object)\n",
" else:\n",
" object_term = Literal(triplet.object)\n",
" rdf_graph.add((URIRef(triplet.subject), URIRef(triplet.predicate), object_term))\n",
"\n",
"rdf_graph.serialize(destination=\"semantic_layer.ttl\", format=\"turtle\")\n",
"turtle = open(\"semantic_layer.ttl\", encoding=\"utf-8\").read()\n",
"print(turtle[:600])\n",
"\n",
"# The exported RDF contains declarations plus mapped instance facts.\n",
"declared_datatype_properties = {\n",
" str(subject) for subject in rdf_graph.subjects(RDF.type, OWL.DatatypeProperty)\n",
"}\n",
"assert result[\"failed\"] == 0\n",
"assert set(property_mappings.values()) <= declared_datatype_properties\n",
"assert (\n",
" URIRef(BASE_URI + \"e1\"),\n",
" URIRef(property_mappings[\"role\"]),\n",
" Literal(\"Engineer\"),\n",
") in rdf_graph\n",
"assert (\n",
" URIRef(BASE_URI + \"e1\"),\n",
" URIRef(relationship_type_mappings[\"works_for\"]),\n",
" URIRef(BASE_URI + \"e3\"),\n",
") in rdf_graph\n",
"print(\"... exported semantic_layer.ttl\")"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 6: Query the Semantic Layer\n",
"\n",
"The embedded Oxigraph backend runs in memory, so there is nothing to start beyond installing the `tripletstore-oxigraph` extra. The organization is constrained by its mapped `name` predicate; the query therefore means *Tech Corp*, rather than accidentally matching employees of every organization.\n"
]
},
{
"cell_type": "code",
"metadata": {},
"source": [
"query = f\"\"\"\n",
"SELECT ?name ?role WHERE {{\n",
" ?person <{BASE_URI}worksFor> ?org .\n",
" ?org <{BASE_URI}name> \"Tech Corp\" .\n",
" ?person <{BASE_URI}name> ?name .\n",
" ?person <{BASE_URI}role> ?role .\n",
"}}\n",
"ORDER BY ?name\n",
"\"\"\"\n",
"query_result = store.execute_query(query)\n",
"\n",
"print(\"\\nWho works for Tech Corp, and in which role?\")\n",
"for binding in query_result.bindings:\n",
" print(f\" {binding['name']['value']} — {binding['role']['value']}\")\n",
"\n",
"assert [(row[\"name\"][\"value\"], row[\"role\"][\"value\"]) for row in query_result.bindings] == [\n",
" (\"Alice\", \"Engineer\"),\n",
" (\"Bob\", \"Manager\"),\n",
"]"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## 🧹 Optional: Clean Up\n"
]
},
{
"cell_type": "code",
"metadata": {},
"source": [
"from pathlib import Path\n",
"\n",
"ttl_file = Path(\"semantic_layer.ttl\")\n",
"if ttl_file.exists():\n",
" ttl_file.unlink()\n",
" print(f\"Removed {ttl_file}\")"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Summary\n",
"\n",
"A minimal semantic layer is a composition, and you have now built each part:\n",
"\n",
"1. **Knowledge graph** — `GraphBuilder` from explicit entities and relationships\n",
"2. **Ontology** — `OntologyGenerator` with your `base_uri`\n",
"3. **Explicit mappings** — entity types, relationship types, and properties, each tied to an ontology term\n",
"4. **Ontology-aligned RDF** — the mappings applied to the graph, materialized with `TripletStore`, and serialized to Turtle from the store's own triples\n",
"5. **Queryable store** — `TripletStore` (embedded Oxigraph) answering a SPARQL question over the shared vocabulary\n",
"\n",
"### Where to go next\n",
"\n",
"The production version of this workflow — hand-designed governed ontologies, explicit source-to-ontology mappings from a warehouse, n-ary modeling, SHACL validation, provenance, and versioning — is covered in [Advanced: Manual Ontology + Snowflake Mapping](../advanced/13_Manual_Ontology_Snowflake_Mapping.ipynb).\n"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "Python 3",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3",
"version": "3.11.9"
}
},
"nbformat": 4,
"nbformat_minor": 2
}
+1 -1
View File
@@ -226,7 +226,7 @@ Pick your goal to see the minimum imports and a working skeleton.
</Tab>
<Tab title="MCP — Claude / Cursor">
Use Semantica from Claude Desktop, Cursor, VS Code, or any MCP-aware tool — no Python code required after setup. 15 tools available instantly.
Use Semantica from Claude Desktop, Cursor, VS Code, or any MCP-aware tool — no Python code required after setup. 12 tools available instantly.
**Step 1 — Install:**
```bash
+2 -2
View File
@@ -53,7 +53,7 @@ python -c "import semantica; print(semantica.__version__)"
- **semantica-server** — Starts the REST API server. Binds to `0.0.0.0:8000`. Use this when another service or application needs programmatic access to Semantica over HTTP.
- **semantica-worker** — Background task processor. Run alongside `semantica-server` when you need async pipeline execution outside the request cycle. Start the server first, then start one or more workers pointing at the same backend.
- **semantica-explorer** — Launches the browser dashboard. Requires `pip install semantica[explorer]`. Use this to explore a saved knowledge graph interactively. See [Explorer Setup](explorer-setup).
- **semantica-mcp** — Runs the MCP server over stdio. Configure it in your MCP client's settings file to expose all 15 tools and 3 resources to Claude Desktop, Cursor, Windsurf, or any MCP-aware client. See [MCP Server](reference/mcp_server).
- **semantica-mcp** — Runs the MCP server over stdio. Configure it in your MCP client's settings file to expose all 12 tools and 3 resources to Claude Desktop, Cursor, Windsurf, or any MCP-aware client. See [MCP Server](reference/mcp_server).
## Usage Examples
@@ -229,6 +229,6 @@ Install the [Microsoft Visual C++ Redistributable](https://aka.ms/vs/17/release/
## Next Steps
- [Explorer Setup](explorer-setup) — Build a graph, save it, and launch the browser dashboard.
- [MCP Server](reference/mcp_server) — All 15 tools and 3 resources exposed over the MCP protocol.
- [MCP Server](reference/mcp_server) — All 12 tools and 3 resources exposed over the MCP protocol.
- [Installation](installation) — Virtual environments, optional extras, and platform-specific notes.
- [Quickstart](quickstart) — End-to-end pipeline walkthrough with working code.
-1
View File
@@ -36,7 +36,6 @@ Essential guides to master the Semantica framework.
- **[Graph Store](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/09_Graph_Store.ipynb)** — Persisting knowledge graphs in Neo4j or FalkorDB. Topics: Neo4j, Cypher, Persistence · *Intermediate*
- **[Ontology](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/14_Ontology.ipynb)** — Defining domain schemas and ontologies to structure your data. Topics: OWL, RDF, Schema Design · *Intermediate*
- **[Seed Data](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/25_Seed_Data.ipynb)** — Bootstrapping a knowledge graph from trusted CSV, JSON, database, and API sources before extraction runs. Topics: SeedDataManager, Foundation Graphs · *Intermediate*
- **[Semantic Layer Basics](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/26_Semantic_Layer_Basics.ipynb)** — Capstone tutorial that combines a knowledge graph, generated ontology, explicit mappings, ontology-aligned RDF, and a SPARQL query. Topics: Semantic Layer, Ontology Mapping, Oxigraph, SPARQL · *Intermediate*
## Advanced Concepts
+6 -6
View File
@@ -16,7 +16,7 @@ icon: "circle-question"
| Python version? | 3.8+ (3.11+ recommended) |
| API key required? | Optional: pattern extraction works with no keys |
| Works with LangChain / LlamaIndex? | Yes: Semantica is a layer on top, not a replacement |
| Production-ready? | Yes: 1,000+ tests, security fixes shipped in every release (see [CHANGELOG](https://github.com/semantica-agi/semantica/blob/main/CHANGELOG.md)) |
| Production-ready? | Yes: 1,000+ tests, v0.5.0 ships with 12 security fixes |
| Latest version? | **v0.6.7** (August 2026) |
| Local LLMs? | Yes: Ollama via LiteLLM, HuggingFaceLLM for air-gapped |
@@ -70,9 +70,9 @@ Yes: MIT licensed, no vendor lock-in, no paywalled features. Some capabilities r
<Accordion title="What's the latest version?" icon="star">
**v0.6.7**: released August 2026.
**v0.5.0**: released May 2026.
Highlights: first-class LangChain integration, SAP OData ingestor, human-editable Markdown round-trip persistence for `ContextGraph`, a structured Action layer for the reasoning engine, and a public `run_shacl_validation` entry point. The 0.6.x line also added first-class CrewAI support and the Semantica RDF vocabulary with deterministic IRIs. See the [CHANGELOG](https://github.com/semantica-agi/semantica/blob/main/CHANGELOG.md) for the full history.
Highlights: Ontology Hub, Distance Intelligence, Parquet/XML ingestion, 12 security fixes, Graph Explorer redesign, NER gateway fix.
```bash
pip install --upgrade semantica
@@ -173,7 +173,7 @@ This includes PyTorch with CUDA, FAISS GPU, and CuPy.
<Accordion title="How does Semantica handle large datasets?" icon="layer-group">
- **Batching**: process documents in configurable chunks to control memory usage
- **Parallel processing**: `PipelineBuilder().set_parallelism(N)` runs independent pipeline steps concurrently
- **Parallel processing**: `Pipeline(workers=N)` runs extraction steps concurrently
- **Delta processing**: update graphs incrementally without full recompute on new data
- **Persistent backends**: swap in-memory NetworkX for Neo4j, FalkorDB, or Apache AGE for large-scale production graphs
@@ -269,13 +269,13 @@ Groq, OpenAI, Anthropic, Google Gemini, Ollama (fully local), DeepSeek, Novita A
<Accordion title="Is Semantica production-ready?" icon="shield-check">
Yes. Every release ships with:
Yes. v0.5.0 ships with:
- 1,000+ passing tests across Python 3.83.12
- `PipelineValidator` and `FailureHandler` with exponential backoff and configurable retry policies
- W3C PROV-O provenance tracking across all modules
- Change management with SHA-256 checksums and full audit trails
- Ongoing security hardening: eval injection, pickle deserialization, SQL injection, XXE, SSRF, ReDoS, and path traversal fixes have all landed across recent releases (see the [CHANGELOG](https://github.com/semantica-agi/semantica/blob/main/CHANGELOG.md) security sections)
- 12 security vulnerability fixes: eval injection, pickle deserialization, SQL injection, XXE, SSRF, ReDoS, path traversal, and more
</Accordion>
+1 -1
View File
@@ -183,7 +183,7 @@ icon: "rocket"
}
```
15 tools available instantly: extract entities, query graph, record decisions, run reasoning, export results.
12 tools available instantly: extract entities, query graph, record decisions, run reasoning, export results.
**Next:** [MCP Server reference →](reference/mcp_server)
</Tab>
+2 -4
View File
@@ -11,7 +11,7 @@ MCP stands for the Model Context Protocol. It is an open standard that allows ex
The Semantica MCP server exposes your knowledge graph as 12 callable tools. By connecting it, any compatible AI client can traverse the graph live, record decisions, run analytics, and export results during a conversation — without you having to write custom tool wrappers.
<Info>
The Semantica MCP server exposes 15 tools and 3 read-only resources. All tools accept and return JSON. No configuration beyond an optional environment variable for graph persistence is required.
The Semantica MCP server exposes 12 tools and 3 read-only resources. All tools accept and return JSON. No configuration beyond an optional environment variable for graph persistence is required.
</Info>
## Architecture & Communication
@@ -132,7 +132,7 @@ docker run --rm -i \
ghcr.io/semantica-agi/semantica-mcp:latest
```
## What the Agent Can Do: The 15 Tools
## What the Agent Can Do: The 12 Tools
Once connected, the LLM can call any of these tools during a conversation. The agent chains them automatically — you do not orchestrate the sequence, you just describe what you want.
@@ -140,8 +140,6 @@ Once connected, the LLM can call any of these tools during a conversation. The a
**Knowledge graph manipulation**`add_entity` adds a node, `add_relationship` adds a directed edge. After extraction, the agent calls these to persist what it found into the live graph.
**Live graph queries and edits**`query_graph` reads the graph without exporting it: fetch one node, walk its neighbours up to five hops, or keyword-search nodes. `update_node` merges properties onto an existing node (for example marking a task node `done`), and `delete_node` archives a node it no longer tracks. When `SEMANTICA_KG_PATH` is set, `update_node` and `delete_node` write their changes back to that file so they survive a restart.
**Decision intelligence**`record_decision` writes a decision as a provenance node with confidence score, reasoning, and decision maker identity. `query_decisions` retrieves past decisions by query or category. `find_precedents` finds the most similar past decisions by semantic similarity. `get_causal_chain` traces decision causality upstream or downstream.
**Reasoning**`run_reasoning` applies forward-chaining IF/THEN rules over a set of facts and returns derived conclusions.
+2 -2
View File
@@ -369,7 +369,7 @@ Semantica was designed for domains where every decision must be explainable and
| `semantica.reasoning` | Forward chaining, Rete, deductive, abductive, SPARQL, Datalog |
| `semantica.ontology` | SHACL, SKOS, alignments, diff/migration, auto-generation, OWL/RDF |
| `semantica.explorer` | FastAPI Knowledge Explorer, Ontology Hub, Distance Intelligence, SHACL Studio |
| `semantica.mcp_server` | MCP stdio server: 15 tools for Claude Desktop, VS Code, Cursor, Windsurf, Cline |
| `semantica.mcp_server` | MCP stdio server: 12 tools for Claude Desktop, VS Code, Cursor, Windsurf, Cline |
| `semantica.vector_store` | FAISS, Pinecone, Weaviate, Qdrant, Milvus, PgVector |
| `semantica.graph_store` | Neo4j, FalkorDB, Apache AGE, Amazon Neptune |
| `semantica.triplet_store` | In-memory and persistent RDF triple store with SPARQL |
@@ -404,7 +404,7 @@ Semantica was designed for domains where every decision must be explainable and
- 1,000+ passing tests with full regression coverage
- `PipelineValidator` catches configuration errors at startup
- `FailureHandler` with exponential backoff and dead-letter queues
- Ongoing security hardening: fixes shipped in every release ([CHANGELOG](https://github.com/semantica-agi/semantica/blob/main/CHANGELOG.md))
- 12 security vulnerabilities fixed in v0.5.0
**Modular by Design** — Import only what you need.
- Use `NERExtractor` without a graph store
+1 -1
View File
@@ -438,7 +438,7 @@ Exposes Semantica as an MCP stdio server for IDE and agent integrations.
python -m semantica.mcp_server
```
**Integrations:** Claude Desktop, VS Code, Cursor, Windsurf, Cline: 15 MCP tools exposed
**Integrations:** Claude Desktop, VS Code, Cursor, Windsurf, Cline: 12 MCP tools exposed
### Seed
+44 -71
View File
@@ -5,7 +5,7 @@ icon: "rocket"
---
<Info>
**v0.6.7**first-class LangChain integration, SAP OData ingestor, human-editable Markdown persistence for `ContextGraph`, and a structured Action layer for the reasoning engine. <a href="https://github.com/semantica-agi/semantica/releases" style={{color:"#10B981",fontWeight:600,textDecoration:"none"}}>What's new →</a>
**v0.5.0**Ontology Hub, Distance Intelligence, Parquet & XML ingestion, 12 security fixes. <a href="https://github.com/semantica-agi/semantica/releases" style={{color:"#10B981",fontWeight:600,textDecoration:"none"}}>What's new →</a>
</Info>
This guide walks you through the end-to-end pipeline for building your first knowledge graph. Start here after installation. An LLM API key is optional: pattern-based extraction works out of the box.
@@ -35,7 +35,7 @@ Verify:
```bash
python -c "import semantica; print(semantica.__version__)"
# 0.6.7
# 0.5.0
```
@@ -62,19 +62,18 @@ sources = ingestor.ingest("data/report.pdf")
```python Web
from semantica.ingest import WebIngestor
ingestor = WebIngestor()
page = ingestor.ingest_url("https://example.com/article")
# WebContent: page.text, page.title, page.html, page.links, page.metadata
ingestor = WebIngestor(max_depth=2)
sources = ingestor.ingest("https://example.com/article")
```
```python Parquet / XML
```python Parquet / XML (v0.5.0)
from semantica.ingest import ParquetIngestor, XMLIngestor
# Single file or Hive-partitioned directory
sources = ParquetIngestor().ingest("data/events.parquet")
# XML; pass an XSD to validate against during ingestion
sources = XMLIngestor().ingest("data/records/", schema_path="schema.xsd")
# XML with XSD schema validation
sources = XMLIngestor(validate_xsd="schema.xsd").ingest("data/records/")
```
</CodeGroup>
@@ -89,24 +88,22 @@ Extract structured text and layout from raw documents.
from semantica.parse import DocumentParser
parser = DocumentParser()
parsed = parser.parse(sources[0].path) # parse() takes a path string
parsed = parser.parse(sources[0])
print(parsed["text"][:200]) # extracted text
print(parsed["metadata"]) # file_path, encoding, size, and format-specific keys
print(parsed.text[:200]) # extracted text
print(parsed.metadata) # title, author, date, source
```
`parse()` returns a `dict` with `text`, `full_text`, and `metadata` keys.
<Tip>
For PDFs with tables, charts, or multi-column layouts, use `DoclingParser` (`pip install semantica[parse-docling]`): it applies advanced layout analysis and returns structured table data alongside text.
For PDFs with tables, charts, or multi-column layouts, use `DoclingParser`: it applies advanced layout analysis and returns structured table data alongside text.
</Tip>
```python
from semantica.parse import DoclingParser
parser = DoclingParser()
parsed = parser.parse(sources[0].path)
print(parsed["tables"]) # structured table data
parsed = parser.parse(sources[0])
print(parsed.tables) # structured table objects
```
</Step>
@@ -120,28 +117,26 @@ Identify named entities and extract typed relationships between them.
```python Pattern-based (fast, no API key)
from semantica.semantic_extract import NERExtractor, RelationExtractor
text = parsed["text"]
ner = NERExtractor(method="pattern")
entities = ner.extract(text)
# Returns: [Entity(text="Apple Inc.", label="ORG", start_char=0, end_char=10, confidence=0.7), ...]
entities = ner.extract(parsed)
# Returns: [{"text": "Apple Inc.", "type": "ORGANIZATION", "confidence": 0.98}, ...]
rel = RelationExtractor(method="pattern")
relationships = rel.extract(text, entities=entities)
# Returns: [Relation(subject=Entity(...), predicate="founded_by", object=Entity(...), confidence=0.7), ...]
rel = RelationExtractor(method="rule")
relationships = rel.extract(parsed, entities=entities)
# Returns: [{"subject": "Steve Jobs", "predicate": "founded", "object": "Apple Inc."}, ...]
```
```python LLM-powered (higher accuracy)
from semantica.semantic_extract import NERExtractor, RelationExtractor
from semantica.llms import Groq
# Reads GROQ_API_KEY from the environment; provider/llm_model select the backend
text = parsed["text"]
llm = Groq(model="llama-3.3-70b-versatile")
ner = NERExtractor(method="llm", provider="groq", llm_model="llama-3.3-70b-versatile")
entities = ner.extract(text)
ner = NERExtractor(method="llm", llm_provider=llm)
entities = ner.extract(parsed)
rel = RelationExtractor(method="llm", provider="groq", llm_model="llama-3.3-70b-versatile")
relationships = rel.extract(text, entities=entities)
rel = RelationExtractor(method="llm", llm_provider=llm)
relationships = rel.extract(parsed, entities=entities)
```
</CodeGroup>
@@ -203,17 +198,16 @@ exporter.export(graph, file_path="graph.nt", format="nt")
from semantica.export import ParquetExporter
exporter = ParquetExporter()
exporter.export(graph, file_path="output/graph")
# Dict input writes one file per key: output/graph_entities.parquet and
# output/graph_relationships.parquet: ready for Spark, BigQuery, Databricks
exporter.export(graph, file_path="output/graph.parquet")
# Writes nodes.parquet + edges.parquet: ready for Spark, BigQuery, Databricks
```
```python ArangoDB
from semantica.export import ArangoAQLExporter
exporter = ArangoAQLExporter()
exporter.export(graph, file_path="graph.aql")
# Writes ready-to-run AQL INSERT statements to graph.aql
aql = exporter.export(graph)
# Returns ready-to-run AQL INSERT statements
```
</CodeGroup>
@@ -278,21 +272,14 @@ relationships = rel.extract(text, entities=entities)
<Accordion title="Multi-source incremental graph build" icon="layer-group">
```python
from semantica.ingest import FileIngestor
from semantica.parse import DocumentParser
from semantica.semantic_extract import NERExtractor, RelationExtractor
from semantica.kg import GraphBuilder
parser = DocumentParser()
ner = NERExtractor(method="pattern")
rel = RelationExtractor(method="pattern")
builder = GraphBuilder(merge_entities=True)
builder = GraphBuilder(merge_entities=True)
all_entities, all_rels = [], []
for source in FileIngestor().ingest("data/reports/"):
text = parser.parse(source.path)["text"]
entities = ner.extract(text)
rels = rel.extract(text, entities=entities)
for doc in parsed_docs:
entities = ner.extract(doc)
rels = rel.extract(doc, entities=entities)
all_entities.extend(entities)
all_rels.extend(rels)
@@ -340,11 +327,10 @@ print(f"Relationships active in 2023: {result_2023['num_relationships']}")
<Accordion title="Persistent graph store: Neo4j, FalkorDB, Apache AGE" icon="database">
```python
from semantica.graph_store import GraphStore
from semantica.graph_store import Neo4jStore
from semantica.kg import GraphBuilder
store = GraphStore(
backend="neo4j",
store = Neo4jStore(
uri="bolt://localhost:7687",
user="neo4j",
password="password",
@@ -372,8 +358,7 @@ graph = builder.build({"entities": entities, "relationships": relationships})
# Retrieve full lineage for any entity
sources = prov.get_all_sources("Apple Inc.")
print(sources[0])
# {"source": "data/report.pdf", "location": None, "timestamp": "...",
# "confidence": 1.0, "metadata": {"confidence": 0.98}}
# {"source": "data/report.pdf", "location": None, "timestamp": "...", "confidence": 0.98}
```
</Accordion>
@@ -387,42 +372,30 @@ print(sources[0])
<Accordion title="No entities extracted" icon="magnifying-glass">
The document likely contains scanned images rather than machine-readable text. `DocumentParser` warns when a PDF has no text layer; switch to `DoclingParser` with OCR enabled:
The document likely contains scanned images rather than machine-readable text. Enable OCR:
```python
from semantica.parse import DoclingParser # pip install semantica[parse-docling]
from semantica.parse import DocumentParser
parser = DoclingParser(enable_ocr=True)
parsed = parser.parse(sources[0].path)
parser = DocumentParser(ocr=True) # enables Tesseract OCR
parsed = parser.parse(sources[0])
```
</Accordion>
<Accordion title="Slow processing on large corpora" icon="gauge">
Enable GPU acceleration and run pipeline steps in parallel:
Enable parallel processing and GPU acceleration:
```bash
pip install semantica[gpu]
```
```python
from semantica.pipeline import PipelineBuilder, ExecutionEngine
from semantica.pipeline import Pipeline
builder = PipelineBuilder()
builder.add_step("ingest", step_type="ingest", source="data/reports/", recursive=True)
builder.add_step("extract", step_type="ner_extract")
builder.add_step("build", step_type="kg_build", merge_entities=True)
pipeline = (
builder
.connect_steps("ingest", "extract")
.connect_steps("extract", "build")
.set_parallelism(8)
.build(name="reports_pipeline")
)
result = ExecutionEngine().execute_pipeline(pipeline)
pipeline = Pipeline(workers=8, batch_size=32)
pipeline.run(sources)
```
</Accordion>
-95
View File
@@ -25,7 +25,6 @@ icon: "brain"
| `DecisionRecorder` | Record decisions with embeddings, causal chains, and metadata |
| `PolicyEngine` | Policy management: `add_policy()`, `check_compliance()`, `get_applicable_policies()` |
| `CausalChainAnalyzer` | Trace how decisions influenced each other: `get_causal_chain(decision_id)` |
| `ErasureCoordinator` | Erase an entity across graph, memory, and vector store, returning an auditable `ErasureReceipt` |
## What You Get
@@ -635,100 +634,6 @@ queried together safely. Vector-store writes are deferred until the in-memory im
commits; adapter synchronization remains best-effort and logs failures.
## ErasureCoordinator
`ContextGraph.purge_node()` is scoped to one graph: the node is removed and a
tombstone is written, but the same content can still be live as an `AgentMemory`
item and as an embedding in the vector store. `ErasureCoordinator` drives the
cascade across every bound store and returns an `ErasureReceipt` recording what
each one reported.
```python
from semantica.context import AgentMemory, ContextGraph, ErasureCoordinator
coordinator = ErasureCoordinator(graph=graph, memory=memory)
receipt = coordinator.erase_entity(
"customer-4471",
reason="GDPR Art. 17 request #882",
)
if not receipt.complete:
# These stores may still hold the entity; handle them out of band.
print(receipt.incomplete_stores)
```
<Warning>
Check the receipt — the call returning is not proof the data is gone. FAISS,
Milvus, and Weaviate expose no delete method, so erasure cannot be completed on
those backends today; the receipt reports `unsupported` rather than a success it
did not achieve.
</Warning>
### Constructor Parameters
| Parameter | Type | Default | Description |
| :--- | :--- | :--- | :--- |
| `graph` | `ContextGraph` | `None` | Anything exposing `purge_node()` |
| `memory` | `AgentMemory` | `None` | Anything exposing `find_by_entity()` and `batch_delete()` |
| `vector_store` | `VectorStore` | `memory.vector_store` | Store holding entity-keyed embeddings; pass `False` to disable the leg |
At least one store is required; a store that is not supplied reports
`not_configured` rather than being silently skipped.
### Methods
| Method | Returns | Description |
| :--- | :--- | :--- |
| `erase_entity(entity_id, reason, at, vector_ids)` | `ErasureReceipt` | Erase one entity from every bound store |
| `erase_entities(entity_ids, reason, at)` | `List[ErasureReceipt]` | One receipt per entity, in order; one failure does not stop the rest |
### Store Statuses
| Status | Meaning |
| :--- | :--- |
| `erased` | Reached, data removed. On the vectors leg this means the store accepted the delete for the ids given — backends offer no portable existence check, so it is not a count of embeddings that were really there |
| `not_found` | Reached, held nothing for this entity |
| `not_configured` | No such store was bound — normal, not a failure |
| `unsupported` | The store cannot delete at all; retrying will not help |
| `failed` | The store was reached and the deletion did not succeed |
### ErasureReceipt
| Member | Type | Description |
| :--- | :--- | :--- |
| `entity_id` | `str` | Entity the erasure was requested for |
| `reason` | `Optional[str]` | Recorded in the receipt and the graph tombstone |
| `erased_at` | `str` | ISO-8601; matches the tombstone's `purged_at` |
| `stores` | `Dict[str, Dict]` | Per-store outcome keyed `vectors`, `memory`, `graph` |
| `complete` | `bool` | `False` when any store reports `unsupported` or `failed` |
| `incomplete_stores` | `List[str]` | Stores that may still hold the entity's data |
| `to_dict()` | `Dict` | Serialized receipt, safe to persist as an audit record |
```python
receipt.to_dict()
# {
# "entity_id": "customer-4471",
# "reason": "GDPR Art. 17 request #882",
# "erased_at": "2026-08-16T09:03:36.813220",
# "complete": False,
# "stores": {
# "vectors": {"status": "unsupported", "backend": "faiss",
# "detail": "backend exposes no delete()/delete_vectors(); ..."},
# "memory": {"status": "erased", "items": 14},
# "graph": {"status": "erased", "nodes": 1, "edges": 3},
# },
# }
```
Erasure runs outward-in — vectors, then memory, then the graph. The tombstone is
the durable attestation that an erasure happened, so it is written last: a crash
mid-cascade leaves the node present and the receipt incomplete, rather than a
tombstone claiming more than actually happened. A store that raises is recorded
as `failed` and the remaining stores are still erased. Erasing the same entity
twice returns a receipt saying there was nothing left to do rather than raising.
## PolicyEngine
`PolicyEngine` manages versioned policies stored in the knowledge graph. Policies are stored as nodes and can be linked to decisions:
+3 -51
View File
@@ -6,7 +6,7 @@ icon: "plug"
**`semantica.mcp_server`** exposes Semantica's knowledge graph, decision intelligence, semantic extraction, and reasoning capabilities as an [MCP (Model Context Protocol)](https://modelcontextprotocol.io) **server over stdio**:
- 15 MCP tools exposed: extract entities, query graph, record decisions, run reasoning, export results
- 12 MCP tools exposed: extract entities, query graph, record decisions, run reasoning, export results
- No Python code required after launch: configure once, use from any MCP-aware client
- Compatible with Claude Desktop, Windsurf, Cline, Continue, VS Code, Roo Code, Cursor
@@ -40,7 +40,7 @@ python -m semantica.mcp_server
## What You Get
- **15 MCP Tools** — Extract entities, extract relations, record decisions, query decisions, find precedents, trace causal chains, add entities, add relationships, run analytics, summarise graph, run reasoning, export graph, query the live graph, update nodes, archive nodes.
- **12 MCP Tools** — Extract entities, extract relations, record decisions, query decisions, find precedents, trace causal chains, add entities, add relationships, run analytics, summarise graph, run reasoning, export graph.
- **3 Readable Resources** — Live graph JSON (`semantica://graph/summary`), decision list, and schema/version info: readable by any MCP client.
- **Zero Infrastructure** — Runs over stdio: no server, no port, no Docker required. One config block to activate in any MCP client.
- **Persistent Graphs** — Point `SEMANTICA_KG_PATH` at a saved graph file to reload it automatically on every server startup.
@@ -159,7 +159,7 @@ The MCP server is included in the base install: no extras required.
## Tools
The MCP server exposes 15 tools that any connected AI assistant can call:
The MCP server exposes 12 tools that any connected AI assistant can call:
| Tool | Category | Description |
| :---- | :-------- | :----------- |
@@ -173,9 +173,6 @@ The MCP server exposes 15 tools that any connected AI assistant can call:
| `add_relationship` | Graph Operations | Add a directed edge between two nodes |
| `get_graph_summary` | Graph Operations | Node count, decision count, graph status |
| `get_graph_analytics` | Graph Operations | PageRank centrality and community detection |
| `query_graph` | Graph Operations | Fetch a node, traverse its neighbours, or keyword-search nodes |
| `update_node` | Graph Operations | Merge properties onto a node and persist to `SEMANTICA_KG_PATH` |
| `delete_node` | Graph Operations | Soft-delete (archive) a node and persist to `SEMANTICA_KG_PATH` |
| `run_reasoning` | Reasoning | Forward-chain IF/THEN rules over facts |
| `export_graph` | Reasoning & Export | Serialise the graph (`turtle`/`ttl`: RDF Turtle aliases, `nt`, `xml`, `json-ld`, `json`) |
@@ -389,51 +386,6 @@ Takes no input parameters.
</Accordion>
<Accordion title="query_graph" icon="magnifying-glass">
Read the live graph in one of three modes, set by `mode`:
- `node` — return a single node by `node_id`.
- `neighbors` (default) — traverse outward and inward from `node_id` up to `depth` hops (clamped to 1-5, default 1). Optional `relationship_types` filters edge types; optional `limit` caps results.
- `search` — keyword match `query` against each node's id and content. Optional `node_type` restricts the scan; `limit` defaults to 50.
**Input:**
```json
{ "mode": "neighbors", "node_id": "apple_inc", "depth": 2 }
```
</Accordion>
<Accordion title="update_node" icon="pen">
Merge a set of properties onto an existing node. The change is applied in memory and, when `SEMANTICA_KG_PATH` is set, written back to that file so it survives a restart. Returns `persisted: false` when no path is configured.
**Input:**
```json
{
"node_id": "task_42",
"properties": { "status": "done", "note": "shipped in v0.6.7" }
}
```
`node_id` and a non-empty `properties` object are required. Updating a missing node returns an error.
</Accordion>
<Accordion title="delete_node" icon="box-archive">
Soft-delete a node: it stays in the graph for history but is marked `status: "archived"`. Persists to `SEMANTICA_KG_PATH` when configured.
**Input:**
```json
{ "node_id": "task_42" }
```
</Accordion>
</AccordionGroup>
### Reasoning
@@ -1,338 +0,0 @@
# Objective Layer for semantica.evals Runner — Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Add per-metric objective support (direction + threshold, or Boolean expectation) to the `evaluate()` runner, overriding evaluator default pass verdicts, backward-compatible when no objective is configured.
**Architecture:** The runner already iterates evaluators and computes per-case status. Objectives are read from `config["<name>"]["objective"]`, validated up front, and applied to each returned metric's `passed` field (and `details`) before aggregation. Error metrics always win over objectives.
**Tech Stack:** Python 3.8+, stdlib only (typing, dataclasses). pytest for tests.
## Global Constraints
- Python >= 3.8: use `typing.Dict/List/Optional/Union`, never builtin generics or `|`.
- Zero new dependencies.
- Do not change the `EvalMetric` shape, the `evaluate()` signature, or the evaluator function signature.
- Existing behavior with no `objective` configured must be byte-for-byte unchanged (all 62 existing tests keep passing).
- Error metrics (`meta` contains `"error"`) always classify the case as `error`, regardless of objective.
- Config errors are programmer errors: raise `ValueError` from `evaluate()` before any evaluator runs (fail-fast).
- Tests go in `tests/evals/`, pytest class style, no new files outside the listed paths.
---
### Task 1: Objective parsing, validation, and re-decision in the runner
**Files:**
- Modify: `semantica/evals/runner.py`
- Test: `tests/evals/test_runner.py`
**Interfaces:**
- Consumes: `EvalMetric` from `.types` (fields: `score`, `passed`, `meta`); `evaluate(cases, evaluators, config=None, target_fn=None)` existing signature.
- Produces: private helpers `_parse_objective(name, eval_config) -> Optional[Dict]` (returns `None` when no objective configured, raises `ValueError` on invalid config) and `_apply_objective(metric, objective) -> bool` (returns the re-decided `passed`). Public `evaluate()` behavior extended as specified.
- [ ] **Step 1: Write the failing tests**
Append a new test class to `tests/evals/test_runner.py`:
```python
class TestObjective:
def test_maximize_with_threshold_pass(self):
# levenshtein similarity 1.0 for identical, objective demands >= 0.5
result = evaluate(
[("apple", "apple")],
evaluators=["levenshtein"],
config={"levenshtein": {"objective": {"direction": "maximize", "threshold": 0.5}}},
)
assert result.cases[0].status == "pass"
assert result.cases[0].metrics["levenshtein"].passed is True
def test_maximize_with_threshold_fail(self):
result = evaluate(
[("apple", "aple")], # similarity < 1.0
evaluators=["levenshtein"],
config={"levenshtein": {"objective": {"direction": "maximize", "threshold": 0.99}}},
)
assert result.cases[0].status == "fail"
assert result.cases[0].metrics["levenshtein"].passed is False
assert "levenshtein" in result.cases[0].details
def test_minimize_with_threshold_pass(self):
# edit distance normalized ~0.2; objective: distance <= 0.5
result = evaluate(
[("night", "nacht")],
evaluators=["levenshtein"],
config={"levenshtein": {"objective": {"direction": "minimize", "threshold": 0.5}}},
)
assert result.cases[0].status == "pass"
assert result.cases[0].metrics["levenshtein"].passed is True
def test_minimize_with_threshold_fail(self):
result = evaluate(
[("night", "nacht")],
evaluators=["levenshtein"],
config={"levenshtein": {"objective": {"direction": "minimize", "threshold": 0.1}}},
)
assert result.cases[0].status == "fail"
def test_expect_true_on_boolean_metric(self):
result = evaluate(
[("ok", "ok")],
evaluators=["exact_match"],
config={"exact_match": {"objective": {"expect": True}}},
)
assert result.cases[0].status == "pass"
def test_expect_false_overrides_passing_metric(self):
# exact_match passes (score 1.0) but expectation is false -> fail
result = evaluate(
[("ok", "ok")],
evaluators=["exact_match"],
config={"exact_match": {"objective": {"expect": False}}},
)
assert result.cases[0].status == "fail"
assert result.cases[0].metrics["exact_match"].passed is False
assert "exact_match" in result.cases[0].details
def test_maximize_without_threshold_is_noop(self):
# identical behavior to no objective: evaluator's own verdict stands
result = evaluate(
[("ok", "no")],
evaluators=["exact_match"],
config={"exact_match": {"objective": {"direction": "maximize"}}},
)
assert result.cases[0].status == "fail"
def test_minimize_without_threshold_raises(self):
with pytest.raises(ValueError):
evaluate(
[("a", "b")],
evaluators=["levenshtein"],
config={"levenshtein": {"objective": {"direction": "minimize"}}},
)
def test_bad_direction_raises(self):
with pytest.raises(ValueError):
evaluate(
[("a", "b")],
evaluators=["levenshtein"],
config={"levenshtein": {"objective": {"direction": "sideways", "threshold": 0.5}}},
)
def test_expect_with_direction_raises(self):
with pytest.raises(ValueError):
evaluate(
[("a", "b")],
evaluators=["levenshtein"],
config={"levenshtein": {"objective": {"expect": True, "direction": "maximize"}}},
)
def test_error_metric_wins_over_objective(self):
result = evaluate(
[("[invalid", "x")],
evaluators=["regex_match"],
config={"regex_match": {"objective": {"direction": "maximize", "threshold": 0.0}}},
)
assert result.cases[0].status == "error"
assert result.errors == 1
assert result.failed == 0
def test_no_objective_unchanged(self):
result = evaluate([("ok", "no")], evaluators=["exact_match"])
assert result.cases[0].status == "fail"
```
- [ ] **Step 2: Run tests to verify they fail**
Run: `python3 -m pytest tests/evals/test_runner.py -q`
Expected: the new `TestObjective` tests fail (objective config ignored → `exact_match` passes under `expect:false` etc.); the pre-existing tests in the file still pass.
- [ ] **Step 3: Implement objective parsing, validation, and re-decision**
In `semantica/evals/runner.py`, add two helpers before `evaluate` and wire them into the evaluator loop.
```python
def _parse_objective(name, eval_config):
"""Return the validated objective dict, or None when not configured.
Raises ValueError for invalid configurations (programmer error).
"""
objective = (eval_config or {}).get("objective")
if objective is None:
return None
direction = objective.get("direction")
threshold = objective.get("threshold")
expect = objective.get("expect")
if expect is not None:
if direction is not None or threshold is not None:
raise ValueError(
f"objective for '{name}': 'expect' cannot be combined with "
"'direction' or 'threshold'"
)
return {"expect": bool(expect)}
if direction == "minimize":
if threshold is None:
raise ValueError(
f"objective for '{name}': 'minimize' requires a 'threshold'"
)
return {"direction": "minimize", "threshold": float(threshold)}
if direction == "maximize":
if threshold is None:
# no bar to re-decide against; treat as absent (evaluator default stands)
return None
return {"direction": "maximize", "threshold": float(threshold)}
raise ValueError(
f"objective for '{name}': 'direction' must be 'maximize' or 'minimize' "
f"(got {direction!r})"
)
def _apply_objective(metric, objective):
"""Return the objective-adjusted pass verdict for a non-error metric."""
if "expect" in objective:
return bool(metric.score) == objective["expect"]
if objective["direction"] == "minimize":
return metric.score <= objective["threshold"]
return metric.score >= objective["threshold"]
```
Then modify the evaluator loop in `evaluate()` so the parsed objective is computed once per case (outside the evaluator loop, since it only depends on merged config), and applied inside the loop:
```python
objective_by_name = {
name: _parse_objective(name, merged.get(name) or {})
for name in evaluators
}
metrics: Dict[str, EvalMetric] = {}
details: Dict[str, Any] = {}
failed, errored = False, False
for name in evaluators:
eval_config = merged.get(name) or {}
try:
metric = get_evaluator(name)(actual, expected, config=eval_config)
objective = objective_by_name.get(name)
if objective is not None and "error" not in metric.meta:
metric = EvalMetric(metric.score, _apply_objective(metric, objective), metric.meta)
metrics[name] = metric
if "error" in metric.meta:
errored = True
details[name] = metric.meta
elif not metric.passed:
failed = True
details[name] = metric.meta
except Exception as exc: # noqa: BLE001
errored = True
metrics[name] = EvalMetric(0.0, False, {"error": str(exc)})
details[name] = {"error": str(exc)}
```
Note: `objective_by_name` is computed once per case (it depends only on merged config), so invalid config raises `ValueError` at the first case — satisfying the fail-fast requirement. `EvalMetric` is a frozen dataclass, so the re-verdict constructs a new instance preserving score/meta.
- [ ] **Step 4: Run tests to verify they pass**
Run: `python3 -m pytest tests/evals/test_runner.py -q`
Expected: all `TestObjective` tests pass; pre-existing tests still pass.
- [ ] **Step 5: Run the full evals suite**
Run: `python3 -m pytest tests/evals -q`
Expected: 62 existing + new tests all pass (no regressions).
- [ ] **Step 6: Commit**
```bash
git add semantica/evals/runner.py tests/evals/test_runner.py
git commit -m "feat(evals): add per-metric objective support to runner"
```
---
### Task 2: Documentation — usage.md and CHANGELOG
**Files:**
- Modify: `semantica/evals/usage.md`
- Modify: `CHANGELOG.md`
**Interfaces:**
- Consumes: the objective config surface implemented in Task 1 (exact keys: `objective.direction`, `objective.threshold`, `objective.expect`; validation rules).
- Produces: docs only.
- [ ] **Step 1: Add objective section to usage.md**
Append a section after the existing "Run the runner over decision records" section:
```markdown
## Set per-evaluator objectives
By default each evaluator decides its own pass/fail. To override that
verdict at the run level, configure an **objective** per evaluator name:
```python
from semantica.evals import evaluate
# Require a minimum similarity (default direction is maximize):
evaluate(
[("apple", "aple")],
evaluators=["levenshtein"],
config={"levenshtein": {"objective": {"direction": "maximize", "threshold": 0.7}}},
)
# Lower is better — override the direction:
evaluate(
[("night", "nacht")],
evaluators=["levenshtein"],
config={"levenshtein": {"objective": {"direction": "minimize", "threshold": 0.5}}},
)
# Boolean expectation on a 0/1 metric:
evaluate(
[("ok", "ok")],
evaluators=["exact_match"],
config={"exact_match": {"objective": {"expect": False}}},
)
```
Rules:
- `maximize` + `threshold`: pass iff `score >= threshold`. `maximize` without
a threshold is a no-op (the evaluator's own verdict stands).
- `minimize` + `threshold`: pass iff `score <= threshold`. `minimize`
**requires** a threshold — omitting it raises `ValueError`.
- `expect` (`true`/`false`): pass iff `bool(score)` matches; cannot be
combined with `direction`/`threshold`.
- A metric whose `meta` contains `"error"` is always an error, never affected
by an objective.
- Invalid objective config raises `ValueError` before any evaluator runs.
```
- [ ] **Step 2: Add CHANGELOG entry**
Under `## [Unreleased]` → `### Added`, insert a new bullet at the top (before the `semantica.evals` module entry), following existing style:
```markdown
- **`semantica.evals` runner gains per-metric objectives** (#1091)
- `evaluate()` now accepts `config={"<evaluator>": {"objective": {"direction": "maximize"|"minimize", "threshold": X}}}` to override the evaluator's default pass verdict with a threshold; `{"objective": {"expect": bool}}` expresses a Boolean expectation
- `minimize` requires a `threshold`; `maximize` without one is a no-op; `expect` cannot be combined with `direction`/`threshold`; invalid config raises `ValueError` before any evaluator runs
- Error metrics are never affected by objectives (error wins over fail)
- Backward compatible: no `objective` key → existing behavior unchanged
- New tests in `tests/evals/test_runner.py::TestObjective`
```
- [ ] **Step 3: Verify docs examples run**
Run the three examples from Step 1 as a Python script (import `evaluate`, run each snippet) to confirm they don't raise unexpectedly. No test output assertion needed beyond "no exception" and sensible status values.
- [ ] **Step 4: Commit**
```bash
git add semantica/evals/usage.md CHANGELOG.md
git commit -m "docs(evals): document per-metric objectives"
```
---
## Self-Review Notes
- **Spec coverage:** §3.1 (config surface) → Task 1 helpers + Task 2 docs; §3.2 (semantics: maximize/minimize/expect) → Task 1 `_apply_objective`; §3.3 (error wins) → Task 1 error branch + `test_error_metric_wins_over_objective`; §3.4 rules 1-3 (validation) → Task 1 `_parse_objective` + 4 validation tests; §3.4 rule 4 → error branch; §3.5 (aggregation unchanged, details on final verdict) → Task 1 loop + `test_expect_false_overrides_passing_metric` asserts `details`; §4 (fail-fast ValueError) → `_parse_objective` at case top; §5 (tests) → Task 1 test class; §6 (compat) → `test_no_objective_unchanged` + full-suite green.
- **Type consistency:** `_parse_objective(name, eval_config) -> Optional[Dict]`, `_apply_objective(metric, objective) -> bool`; `EvalMetric(score, passed, meta)` positional construction preserved everywhere.
- **Backward compat:** objective parsed to `None` for absent config → loop behavior identical to before.
@@ -1,115 +0,0 @@
# Design: Objective layer for `semantica.evals` runner
**Date:** 2026-08-19
**Issue:** semantica-agi/semantica#1091 (assigned to pkupt)
**Base:** PR #1090 (`semantica.evals` module)
## 1. Problem
`semantica.evals` runs named evaluators and aggregates per-case pass/fail, but the pass judgement is hard-coded inside each evaluator — a higher score always means "better". There is no way to express an evaluation objective at the run level:
- apply a threshold the evaluator does not encode (e.g. "F1 must be ≥ 0.7");
- reverse the direction (e.g. "lower edit distance is better");
- express a Boolean expectation (e.g. "this metric should be `false`").
This blocks the domain-specific benchmark harnesses `docs/community-projects.md` says `semantica.evals` supports. Palantir AIP Evals models exactly this: each metric has an **objective** (Boolean expected value, or numeric `maximize`/`minimize` direction with an optional threshold), and a test case passes when **all** its metrics meet their objectives.
## 2. Scope
In scope:
- A per-metric objective configuration consumed by the `evaluate()` runner.
- Runner-level pass/fail re-decision for numeric scores and Boolean metrics.
- Backward-compatible behavior when no objective is configured.
- Tests and docs.
Out of scope:
- Changing the evaluator signature or the `EvalMetric` shape.
- Multi-iteration test cases (AIP Evals has them; Semantica's runner is single-iteration per case).
- Objective-aware aggregation beyond per-case `pass`/`fail` (existing `pass_rate` semantics are kept).
## 3. Design
### 3.1 Configuration surface
Objective is configured per evaluator inside the runner's `config`, under the evaluator name:
```python
config = {
"<evaluator_name>": {
"objective": {
"direction": "maximize" | "minimize",
"threshold": <float>, # optional
}
}
}
```
Boolean-form objective (shorthand): for metrics whose score is Boolean-like (0.0/1.0) or for semantic clarity, `{"objective": {"expect": true}}` / `{"objective": {"expect": false}}` is also supported.
### 3.2 Evaluation semantics
For each metric produced by an evaluator during a case run, if an objective exists for that evaluator name, the runner recomputes the metric's pass verdict:
- **maximize**: pass iff `score >= threshold`. If no `threshold` is given, the objective is treated as absent (evaluator's own verdict stands) — see 3.4 rule 2.
- **minimize**: pass iff `score <= threshold` (threshold required, see 3.4 rule 1).
- **expect**: pass iff `bool(score)` equals `expect` (for Boolean-style metrics).
When an objective is present, the runner **overrides** `metric.passed` with the objective verdict. When absent, `metric.passed` is used unchanged (existing behavior).
The `objective` key is a **reserved runner-level key**: it is consumed by the runner and is passed through to the evaluator function inside `eval_config` (evaluators already ignore unknown config keys via `cfg.get(...)`, so this is harmless); evaluators must not rely on it. The runner re-decision happens on the metric the evaluator returns, so no evaluator change is required.
### 3.3 Interaction with errors
An `EvalMetric` whose `meta` contains `"error"` remains classified as an error regardless of objective (error wins over fail, per the existing contract). Objectives only affect non-error metrics.
### 3.4 Ambiguity rules (explicit decisions)
1. **`minimize` without `threshold`** is rejected at config-validation time with a clear error (`ValueError`), because "lowest is best" has no absolute pass bar without a threshold. (AIP Evals allows direction-only; we require threshold to keep pass/fail well-defined.) — *Chosen for determinism; revisit if a use case demands direction-only minimize.*
2. **`maximize` without `threshold`** behaves like no objective (pass iff evaluator's own `passed`), because the evaluator's default is already "higher is better".
3. **`expect` with a numeric `direction`/`threshold`** is a config error (`ValueError`): pick one form.
4. **Objective on a metric that errors** → the error wins (3.3), objective ignored.
### 3.5 Aggregation
Unchanged:
- Case `status`: `"error"` if any metric errored, else `"fail"` if any failed, else `"pass"`.
- `pass_rate` = passed / total (1.0 on empty).
- `metrics` dict holds the (possibly re-verdict'd) `EvalMetric`; the re-verdict is observable via `metric.passed`.
- `details[name]` is populated when a metric ends up failed **after** objective re-decision (i.e. objective-failed metrics appear in `details`; metrics that pass under objective are not recorded there). This mirrors the existing "record failures in details" behavior applied to the final verdict.
### 3.6 Files
- `semantica/evals/runner.py` — add objective parsing/validation and re-decision inside the evaluator loop.
- `tests/evals/test_runner.py` — new test class(es) for objective semantics.
- `semantica/evals/usage.md` — document the objective config and examples.
- `CHANGELOG.md``[Unreleased]` entry.
No new dependencies; Python ≥ 3.8 (stdlib `typing`).
## 4. Error handling
- Invalid objective config (`direction` not in {maximize, minimize}, both `expect` and `direction`, `minimize` without threshold, non-numeric threshold) → `ValueError` raised at runner config parse, before any evaluator runs. Deterministic, fail-fast.
- These are programmer errors, not per-case data errors — no per-case `error` status involved.
## 5. Testing
New tests in `tests/evals/test_runner.py`:
1. maximize + threshold: score ≥ threshold → pass; below → fail.
2. minimize + threshold: score ≤ threshold → pass; above → fail (e.g. levenshtein on a close pair).
3. minimize without threshold → `ValueError`.
4. expect=true / expect=false on a Boolean metric (exact_match) — pass/fail per expectation.
5. no objective → existing behavior unchanged (evaluator's own verdict).
6. objective + error metric → error wins (status=error, not fail).
7. config error (bad direction) → `ValueError` raised by `evaluate()`.
8. objective turns a passing metric into failing → `details` records it; case status becomes fail.
9. backward-compat: all existing 62 tests keep passing.
## 6. Compatibility
- Public API (`evaluate`, `list_evaluators`, `get_evaluator`, types) unchanged in signature.
- `EvalMetric` shape unchanged (score, passed, meta) — only `passed` may be recomputed by the runner.
- Existing configs (no `objective` key) behave identically.
-36
View File
@@ -1,36 +0,0 @@
# CI templates
Copy-paste starting points for wiring `semantica` into your own project's CI. Each file is a
complete, working config — rename it into your project (see the comment at the top of each file
for the target path) and swap the smoke-test / test step for whatever your project does with
Semantica. Each template installs `semantica` unconditionally and your own project's dependencies
only if a `requirements.txt` is present; if your project uses `pyproject.toml`, Poetry, or Pipenv
instead, adjust the marked install line (each file calls it out inline).
| File | Target path in your repo |
| ---- | ------------------------- |
| [`github-actions.yml`](github-actions.yml) | `.github/workflows/semantica.yml` |
| [`gitlab-ci.yml`](gitlab-ci.yml) | `.gitlab-ci.yml` |
| [`circleci-config.yml`](circleci-config.yml) | `.circleci/config.yml` |
If your own project is hosted on GitHub, you can skip the setup boilerplate entirely and use
Semantica's reusable composite action instead:
```yaml
- uses: semantica-agi/semantica/.github/actions/setup-semantica@main
with:
python-version: '3.11'
# extras: 'explorer,all' # optional
# version: '==0.6.7' # optional, pin an exact release
# cache: 'pip' # optional, only if your repo has a requirements.txt/pyproject.toml/etc.
```
`@main` always tracks this repo's default branch, which is convenient but — like any mutable
ref — can change out from under you between runs. For production CI, pin it to a commit SHA
instead (find one via `git rev-parse` against a tagged release, or the commit history for
[`.github/actions/setup-semantica/`](../../.github/actions/setup-semantica/)) and update the pin
deliberately when you want to pick up changes, the same way this repo's own workflows are pinned
(see [`verify-action-pins.yml`](../../.github/workflows/verify-action-pins.yml)).
It installs Python, installs `semantica`, and verifies the import (pip caching is opt-in via `cache: 'pip'`, since not every caller repo has a requirements file to key the cache on) — see
[`.github/actions/setup-semantica/action.yml`](../../.github/actions/setup-semantica/action.yml).
-40
View File
@@ -1,40 +0,0 @@
# Drop this in as .circleci/config.yml in your own project.
version: 2.1
jobs:
test:
docker:
- image: cimg/python:3.11
steps:
- checkout
# A content-hashed cache key (e.g. `{{ checksum "requirements.txt" }}`)
# is more precise but breaks if that exact file doesn't exist in your
# project - swap in one matched to however you declare dependencies
# once you've adjusted the install step below.
- restore_cache:
keys:
- pip-cache-v1
- run:
name: Install dependencies
command: |
pip install --upgrade pip
pip install semantica
# Install your own project's dependencies however your project
# declares them - adjust this to match, e.g. `pip install -e .`
# for pyproject.toml / setup.cfg, or `poetry install`.
if [ -f requirements.txt ]; then pip install -r requirements.txt; fi
- save_cache:
key: pip-cache-v1
paths:
- ~/.cache/pip
- run:
name: Smoke test
command: python -c "import semantica; print('semantica', semantica.__version__)"
- run:
name: Run tests
command: pytest
workflows:
test:
jobs:
- test
-44
View File
@@ -1,44 +0,0 @@
# Drop this in as .github/workflows/semantica.yml in your own project.
#
# Installs Semantica and runs a smoke import + your test suite. Swap the
# smoke-test step for whatever your project actually does with Semantica
# (build a context graph, run an ingest pipeline, etc.).
#
# Third-party actions below are pinned to a commit SHA rather than a mutable
# tag - a moved tag can silently swap in different code. Update the pin (and
# the trailing "# vX" comment) deliberately when you want a newer version;
# see semantica-agi/semantica's own .github/workflows/verify-action-pins.yml
# for one way to keep pins honest automatically.
name: Semantica
on:
push:
branches: [main]
pull_request:
branches: [main]
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7
with:
python-version: '3.11'
cache: 'pip'
- name: Install dependencies
run: |
python -m pip install --upgrade pip
pip install semantica
# Install your own project's dependencies however your project
# declares them - adjust this to match. Examples:
# pip install -r requirements.txt
# pip install -e . # pyproject.toml / setup.cfg
# pip install -e ".[dev]"
# poetry install
if [ -f requirements.txt ]; then pip install -r requirements.txt; fi
- name: Run tests
run: pytest
-20
View File
@@ -1,20 +0,0 @@
# Drop this in as .gitlab-ci.yml in your own project.
semantica-test:
image: python:3.11-slim
cache:
paths:
- .cache/pip
variables:
PIP_CACHE_DIR: "$CI_PROJECT_DIR/.cache/pip"
script:
- pip install --upgrade pip
- pip install semantica
# Install your own project's dependencies however your project declares
# them - adjust this to match, e.g. `pip install -e .` for pyproject.toml
# / setup.cfg, or `poetry install`.
- if [ -f requirements.txt ]; then pip install -r requirements.txt; fi
- python -c "import semantica; print('semantica', semantica.__version__)"
- pytest
rules:
- if: '$CI_PIPELINE_SOURCE == "merge_request_event"'
- if: '$CI_COMMIT_BRANCH == "main"'
+51 -43
View File
@@ -80,6 +80,7 @@
"integrity": "sha512-QdxmAo/ikZqqRGA8s43ww8lcql6naWRvEz0FFrl6MIlc7Gi6TroXnSdWa5U/kq6fzcpqpHesicQxFZIieZbyIA==",
"dev": true,
"license": "MIT",
"peer": true,
"dependencies": {
"@babel/code-frame": "^7.29.0",
"@babel/generator": "^7.29.6",
@@ -1602,8 +1603,7 @@
"version": "2.0.46",
"resolved": "https://registry.npmjs.org/@types/hammerjs/-/hammerjs-2.0.46.tgz",
"integrity": "sha512-ynRvcq6wvqexJ9brDMS4BnBLzmr0e14d6ZJTEShTBWKymQiHwlAyGu0ZPEFI2Fh1U53F7tN9ufClWM5KvqkKOw==",
"license": "MIT",
"peer": true
"license": "MIT"
},
"node_modules/@types/hast": {
"version": "3.0.5",
@@ -1642,6 +1642,7 @@
"integrity": "sha512-A1sre26ke7HDIuY/M23nd9gfB+nrmhtYyMINbjI1zHJxYteKR6qSMX56FsmjMcDb3SMcjJg5BiRRgOCC/yBD0g==",
"devOptional": true,
"license": "MIT",
"peer": true,
"dependencies": {
"undici-types": "~7.16.0"
}
@@ -1651,6 +1652,7 @@
"resolved": "https://registry.npmjs.org/@types/react/-/react-19.2.14.tgz",
"integrity": "sha512-ilcTH/UniCkMdtexkoCN0bI7pMcJDvmQFPvuPvmEaYA/NSfFTAgdUSLAoVjaRJm7+6PvcM+q1zYOwS4wTYMF9w==",
"license": "MIT",
"peer": true,
"dependencies": {
"csstype": "^3.2.2"
}
@@ -1670,8 +1672,7 @@
"resolved": "https://registry.npmjs.org/@types/trusted-types/-/trusted-types-2.0.7.tgz",
"integrity": "sha512-ScaPdn1dQczgbl0QFTeTOmVHFULt394XJgOQNoyVhZ6r2vLnMLJfBPd53SB52T/3G36VI1/g2MZaX0cwDuXsfw==",
"license": "MIT",
"optional": true,
"peer": true
"optional": true
},
"node_modules/@types/unist": {
"version": "3.0.3",
@@ -1724,6 +1725,7 @@
"integrity": "sha512-/Zb/xaIDfxeJnvishjGdcR4jmr7S+bda8PKNhRGdljDM+elXhlvN0FyPSsMnLmJUrVG9aPO6dof80wjMawsASg==",
"dev": true,
"license": "MIT",
"peer": true,
"dependencies": {
"@typescript-eslint/scope-manager": "8.58.2",
"@typescript-eslint/types": "8.58.2",
@@ -1993,6 +1995,7 @@
"integrity": "sha512-xRQbDb9BnwDafYNn6Vwl839DYVjqXYb1XVGtWAZ1kcDc6iwAL4hg3B1dZlRiuENFeO2H53gFG3in621AdERVAg==",
"dev": true,
"license": "MIT",
"peer": true,
"bin": {
"acorn": "bin/acorn"
},
@@ -2067,9 +2070,9 @@
}
},
"node_modules/baseline-browser-mapping": {
"version": "2.11.20",
"resolved": "https://registry.npmjs.org/baseline-browser-mapping/-/baseline-browser-mapping-2.11.20.tgz",
"integrity": "sha512-H0ulySigv6icDJ1F7SjtdCD6PrhTpdYCmP0CactWy1+ekh0AFd0o1Wn5T8b+hnTmdBx19u9yhL6wvCylXMY7zw==",
"version": "2.10.20",
"resolved": "https://registry.npmjs.org/baseline-browser-mapping/-/baseline-browser-mapping-2.10.20.tgz",
"integrity": "sha512-1AaXxEPfXT+GvTBJFuy4yXVHWJBXa4OdbIebGN/wX5DlsIkU0+wzGnd2lOzokSk51d5LUmqjgBLRLlypLUqInQ==",
"dev": true,
"license": "Apache-2.0",
"bin": {
@@ -2080,9 +2083,9 @@
}
},
"node_modules/brace-expansion": {
"version": "5.0.9",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.9.tgz",
"integrity": "sha512-ScQ4IuvIEF1TMlP7Zt+vjJ//9zlPb2SDcxWxM3bk8s6t6GGdJ7KO1dCcTidOPJKePW30LE/2cT7wCyPho9/Wxg==",
"version": "5.0.8",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.8.tgz",
"integrity": "sha512-JZyDyq3D4AUifKTPOB7DELf6XsB3WdPuNxCtob1vFXPsSXhdAiHBWJ/tJ8HAc9aH84BK+5JFZLNkJKx3G9kzQg==",
"dev": true,
"license": "MIT",
"dependencies": {
@@ -2093,9 +2096,9 @@
}
},
"node_modules/browserslist": {
"version": "4.28.8",
"resolved": "https://registry.npmjs.org/browserslist/-/browserslist-4.28.8.tgz",
"integrity": "sha512-V2NpofLblG64mfOtSgDhOJESZEGogzDMBv/q+W6oc4LXWP/q75eOXoOaaOu1EOadB9U4Bwx/e0yzbvwKH8zalA==",
"version": "4.28.2",
"resolved": "https://registry.npmjs.org/browserslist/-/browserslist-4.28.2.tgz",
"integrity": "sha512-48xSriZYYg+8qXna9kwqjIVzuQxi+KYWp2+5nCYnYKPTr0LvD89Jqk2Or5ogxz0NUMfIjhh2lIUX/LyX9B4oIg==",
"dev": true,
"funding": [
{
@@ -2112,12 +2115,13 @@
}
],
"license": "MIT",
"peer": true,
"dependencies": {
"baseline-browser-mapping": "^2.11.12",
"caniuse-lite": "^1.0.30001809",
"electron-to-chromium": "^1.5.402",
"node-releases": "^2.0.53",
"update-browserslist-db": "^1.3.0"
"baseline-browser-mapping": "^2.10.12",
"caniuse-lite": "^1.0.30001782",
"electron-to-chromium": "^1.5.328",
"node-releases": "^2.0.36",
"update-browserslist-db": "^1.2.3"
},
"bin": {
"browserslist": "cli.js"
@@ -2127,9 +2131,9 @@
}
},
"node_modules/caniuse-lite": {
"version": "1.0.30001810",
"resolved": "https://registry.npmjs.org/caniuse-lite/-/caniuse-lite-1.0.30001810.tgz",
"integrity": "sha512-TITQPUkaz+aVk5GL6NhOdwk1aEaNTSDPsGFWrTuhKGtjTF70jL/Oht2W4c6rXUe5fu7Ie19VIahAXHIIiWWNeg==",
"version": "1.0.30001788",
"resolved": "https://registry.npmjs.org/caniuse-lite/-/caniuse-lite-1.0.30001788.tgz",
"integrity": "sha512-6q8HFp+lOQtcf7wBK+uEenxymVWkGKkjFpCvw5W25cmMwEDU45p1xQFBQv8JDlMMry7eNxyBaR+qxgmTUZkIRQ==",
"dev": true,
"funding": [
{
@@ -2217,8 +2221,7 @@
"version": "2.20.3",
"resolved": "https://registry.npmjs.org/commander/-/commander-2.20.3.tgz",
"integrity": "sha512-GpVkmM8vF2vQUkj2LvZmD35JxeJOLCwJ9cUkugyk2nuhbv3+mJvpLYYt+0+USMxE+oj+ey/lJEnhZw75x/OMcQ==",
"license": "MIT",
"peer": true
"license": "MIT"
},
"node_modules/component-emitter": {
"version": "1.3.1",
@@ -2256,8 +2259,7 @@
"version": "0.0.10",
"resolved": "https://registry.npmjs.org/cssfilter/-/cssfilter-0.0.10.tgz",
"integrity": "sha512-FAaLDaplstoRsDR8XGYH51znUN0UY7nMc6Z9/fvE8EXGwvJE9hu7W2vHwx1+bd6gCYnln9nLbzxFTrcO9YQDZw==",
"license": "MIT",
"peer": true
"license": "MIT"
},
"node_modules/csstype": {
"version": "3.2.3",
@@ -2322,6 +2324,7 @@
"resolved": "https://registry.npmjs.org/d3-selection/-/d3-selection-3.0.0.tgz",
"integrity": "sha512-fmTRWbNMmsmWq6xJV8D19U/gw/bwrHfNXxrIN+HfZgnzqTHp9jOmKMhsTUjXOJnZOdZY9Q28y4yebKzqDKlxlQ==",
"license": "ISC",
"peer": true,
"engines": {
"node": ">=12"
}
@@ -2454,15 +2457,14 @@
"resolved": "https://registry.npmjs.org/dompurify/-/dompurify-3.4.13.tgz",
"integrity": "sha512-2vmYIoqjze2d+kakP8S/nS5shfsl587kzwEjcGlTdiksUVgFHnFCsLYDVj/JNqJVOQZGSYBTmuycv0PodwmnMQ==",
"license": "(MPL-2.0 OR Apache-2.0)",
"peer": true,
"optionalDependencies": {
"@types/trusted-types": "^2.0.7"
}
},
"node_modules/electron-to-chromium": {
"version": "1.5.420",
"resolved": "https://registry.npmjs.org/electron-to-chromium/-/electron-to-chromium-1.5.420.tgz",
"integrity": "sha512-2yD6XreGusOfNV+dUcvipJEXc3n/n7fgr7996aszTG+YY5E4mqM4tOq/3uhP129cazL9YHbVWSpc79ePotWtPA==",
"version": "1.5.340",
"resolved": "https://registry.npmjs.org/electron-to-chromium/-/electron-to-chromium-1.5.340.tgz",
"integrity": "sha512-908qahOGocRMinT2nM3ajCEM99H4iPdv84eagPP3FfZy/1ZGeOy2CZYzjhms81ckOPCXPlW7LkY4XpxD8r1DrA==",
"dev": true,
"license": "ISC"
},
@@ -2537,6 +2539,7 @@
"integrity": "sha512-nuKKvN+oIBO0koN7Tm7dlkmnkc21mtt0QJLwAKzjLq14y6lRTdVG36MZHJ8eQHwdJMwZbQNMlPOYedMq/oVJvQ==",
"dev": true,
"license": "MIT",
"peer": true,
"workspaces": [
"packages/*"
],
@@ -3336,7 +3339,6 @@
"resolved": "https://registry.npmjs.org/marked/-/marked-14.0.0.tgz",
"integrity": "sha512-uIj4+faQ+MgHgwUW1l2PsPglZLOLOT1uErt06dAPtx2kjteLAkbsd/0FiYg/MGS+i7ZKLb7w2WClxHkzOOuryQ==",
"license": "MIT",
"peer": true,
"bin": {
"marked": "bin/marked.js"
},
@@ -4248,9 +4250,9 @@
"license": "MIT"
},
"node_modules/nanoid": {
"version": "3.3.18",
"resolved": "https://registry.npmjs.org/nanoid/-/nanoid-3.3.18.tgz",
"integrity": "sha512-DTg4MJbGMWkfi6VZFdNt2/caMbQy4Ou+Op/hJQvGEWcnVfoA1QA+xzRKAzw9jD6+GVOOeYr/mIcuDSdug6F6+w==",
"version": "3.3.16",
"resolved": "https://registry.npmjs.org/nanoid/-/nanoid-3.3.16.tgz",
"integrity": "sha512-bzlKTyNJ7+LdGIIwy8ijFpIqEQIvafahV7eYykJ8Cvh42EdJeODoJ6gUJXpQJvej1BddH8OqTXZNE/KfbWAu8Q==",
"dev": true,
"funding": [
{
@@ -4274,14 +4276,11 @@
"license": "MIT"
},
"node_modules/node-releases": {
"version": "2.0.54",
"resolved": "https://registry.npmjs.org/node-releases/-/node-releases-2.0.54.tgz",
"integrity": "sha512-YHs7BmmcsdAI5Ozuf8JZo6PT0mv2GIWC9vMfvUC3dp65M8hn7Ux8CPL+2oBI7juNuj9d0ndhTcznq2ODBps9cQ==",
"version": "2.0.37",
"resolved": "https://registry.npmjs.org/node-releases/-/node-releases-2.0.37.tgz",
"integrity": "sha512-1h5gKZCF+pO/o3Iqt5Jp7wc9rH3eJJ0+nh/CIoiRwjRxde/hAHyLPXYN4V3CqKAbiZPSeJFSWHmJsbkicta0Eg==",
"dev": true,
"license": "MIT",
"engines": {
"node": ">=18"
}
"license": "MIT"
},
"node_modules/object-assign": {
"version": "4.1.1",
@@ -4415,6 +4414,7 @@
"integrity": "sha512-QP88BAKvMam/3NxH6vj2o21R6MjxZUAd6nlwAS/pnGvN9IVLocLHxGYIzFhg6fUQ+5th6P4dv4eW9jX3DSIj7A==",
"dev": true,
"license": "MIT",
"peer": true,
"engines": {
"node": ">=12"
},
@@ -4551,6 +4551,7 @@
"resolved": "https://registry.npmjs.org/react/-/react-19.2.5.tgz",
"integrity": "sha512-llUJLzz1zTUBrskt2pwZgLq59AemifIftw4aB7JxOqf1HY2FDaGDxgwpAPVzHU1kdWabH7FauP4i1oEeer2WCA==",
"license": "MIT",
"peer": true,
"engines": {
"node": ">=0.10.0"
}
@@ -4616,6 +4617,7 @@
"resolved": "https://registry.npmjs.org/react-dom/-/react-dom-19.2.5.tgz",
"integrity": "sha512-J5bAZz+DXMMwW/wV3xzKke59Af6CHY7G4uYLN1OvBcKEsWOs4pQExj86BBKamxl/Ik5bx9whOrvBlSDfWzgSag==",
"license": "MIT",
"peer": true,
"dependencies": {
"scheduler": "^0.27.0"
},
@@ -4861,6 +4863,7 @@
"resolved": "https://registry.npmjs.org/sigma/-/sigma-3.0.2.tgz",
"integrity": "sha512-/BUbeOwPGruiBOm0YQQ6ZMcLIZ6tf/W+Jcm7dxZyAX0tK3WP9/sq7/NAWBxPIxVahdGjCJoGwej0Gdrv0DxlQQ==",
"license": "MIT",
"peer": true,
"dependencies": {
"events": "^3.3.0",
"graphology-utils": "^2.5.2"
@@ -4986,6 +4989,7 @@
"integrity": "sha512-X8EX+XV4QR5xCsrgxaED954zTDfY8KqlDtskKEL0cHhyS/P8b4IFOvGDQpsC9Q1XnLq915wEfwwY/zzskCtmhg==",
"dev": true,
"license": "MIT",
"peer": true,
"dependencies": {
"esbuild": "~0.28.0"
},
@@ -5018,6 +5022,7 @@
"integrity": "sha512-jl1vZzPDinLr9eUt3J/t7V6FgNEw9QjvBPdysz9KfQDD41fQrC2Y4vKQdiaUpFT4bXlb1RHhLpp8wtm6M5TgSw==",
"dev": true,
"license": "Apache-2.0",
"peer": true,
"bin": {
"tsc": "bin/tsc",
"tsserver": "bin/tsserver"
@@ -5145,9 +5150,9 @@
}
},
"node_modules/update-browserslist-db": {
"version": "1.3.2",
"resolved": "https://registry.npmjs.org/update-browserslist-db/-/update-browserslist-db-1.3.2.tgz",
"integrity": "sha512-UQ+MSxlhRm1bzjhU+DcuXfjFO1FzNtqhK5+9Yvlp90ItDLk5vT932A0rFu619nf7RVS+Y/VeaUW1jaRDqZ8VJw==",
"version": "1.2.3",
"resolved": "https://registry.npmjs.org/update-browserslist-db/-/update-browserslist-db-1.2.3.tgz",
"integrity": "sha512-Js0m9cx+qOgDxo0eMiFGEueWztz+d4+M3rGlmKPT+T4IS/jP4ylw3Nwpu6cpTTP8R1MAC1kF4VbdLt3ARf209w==",
"dev": true,
"funding": [
{
@@ -5241,6 +5246,7 @@
"resolved": "https://registry.npmjs.org/vis-data/-/vis-data-8.0.3.tgz",
"integrity": "sha512-jhnb6rJNqkKR1Qmlay0VuDXY9ZlvAnYN1udsrP4U+krgZEq7C0yNSKdZqmnCe13mdnf9AdVcdDGFOzy2mpPoqw==",
"license": "(Apache-2.0 OR MIT)",
"peer": true,
"funding": {
"type": "opencollective",
"url": "https://opencollective.com/visjs"
@@ -5295,6 +5301,7 @@
"integrity": "sha512-NTKlcQjlAK7MlQoyb6LgaqHc8sso/pVyUJYWMws3jg21uTJw/LddqIFPcPqP6PzpgbIcZyKI85sFE4HBrQDA8A==",
"dev": true,
"license": "MIT",
"peer": true,
"dependencies": {
"esbuild": "^0.25.0",
"fdir": "^6.4.4",
@@ -5433,6 +5440,7 @@
"integrity": "sha512-rftlrkhHZOcjDwkGlnUtZZkvaPHCsDATp4pGpuOOMDaTdDDXF91wuVDJoWoPsKX/3YPQ5fHuF3STjcYyKr+Qhg==",
"dev": true,
"license": "MIT",
"peer": true,
"funding": {
"url": "https://github.com/sponsors/colinhacks"
}
+1 -1
View File
@@ -9,7 +9,7 @@
"lint": "eslint .",
"preview": "vite preview",
"test:graph-store": "node --test tests/graphStore.multi-edge.test.mjs",
"test:graph-workspace": "node --import tsx --test tests/markdownContentViewer.test.ts tests/graphSceneState.display.test.ts tests/temporalLifecycle.test.ts tests/deterministicExplorerRendering.test.ts tests/smallGraphLayout.test.ts tests/realtimeGraphAttributes.test.ts",
"test:graph-workspace": "node --import tsx --test tests/markdownContentViewer.test.ts tests/graphSceneState.display.test.ts tests/temporalLifecycle.test.ts tests/deterministicExplorerRendering.test.ts",
"test:deterministic-e2e": "node --import tsx --test tests/deterministicExplorerRendering.e2e.ts",
"test:plugin-registry": "node --import tsx --test tests/pluginRegistry.temporal.test.mjs"
},
-1
View File
@@ -86,7 +86,6 @@ export interface EdgeAttributes {
dominantEdgeType?: string;
representativeWeight?: number;
bundleKind?: "parallel" | "bidirectional" | "community";
isSmallGraph?: boolean;
edgeType: string;
@@ -21,6 +21,7 @@ import type Graph from "graphology";
import { batchMergeEdges, batchMergeNodes, graph } from "../../store/graphStore";
import { logEvent } from "../../store/registryStore";
import type { EdgeAttributes, NodeAttributes } from "../../store/graphStore";
import { curveGroupForPair } from "../../store/edgePairKeys.js";
import { InspectorPanel, MetricChip, SurfaceCard } from "../../ui/primitives";
import { lazy, Suspense } from "react";
import { SigmaSceneAdapter } from "./SigmaSceneAdapter";
@@ -41,8 +42,6 @@ import {
import { explorationEffectsShouldLoad, neighborhoodPanelShouldLoad, temporalOverlayShouldLoad } from "./pluginRegistryPredicates";
import { shouldFetchTemporalBounds, shouldFetchTemporalSnapshot } from "./temporalLifecyclePredicates";
import { createTemporalSnapshotGuards, type TemporalSnapshotResponse } from "./temporalSnapshotGuards";
import { SMALL_GRAPH_MAX_NODES } from "./smallGraphLayout";
import { buildRealtimeEdgeAttributes } from "./realtimeGraphAttributes";
import type { LinkPrediction, PathResponse } from "./GraphInspectorPanel";
import type { GraphSceneHandle, GraphSceneRuntime } from "./scene";
import type {
@@ -1057,10 +1056,46 @@ function buildRealtimeNodeAttributes(payload: {
};
}
function synchronizeRealtimeSmallGraphEdges(isSmallGraph: boolean): void {
graph.forEachEdge((edgeId) => {
graph.setEdgeAttribute(edgeId, "isSmallGraph", isSmallGraph);
});
function buildRealtimeEdgeAttributes(payload: {
id: string;
familyId?: string;
source_id: string;
target_id: string;
type?: string;
weight?: number;
properties?: Record<string, unknown>;
}): EdgeAttributes {
const properties = payload.properties || {};
const isInferred = Boolean(properties.inferred);
const isBidirectional = graph.hasDirectedEdge(payload.target_id, payload.source_id);
const baseColor = isInferred ? GRAPH_THEME.palette.accent.path : GRAPH_THEME.palette.muted.edgeStructure;
return {
edgeId: payload.id,
familyId: payload.familyId || payload.id,
sourceId: payload.source_id,
targetId: payload.target_id,
weight: Number(payload.weight ?? 1),
edgeType: payload.type || "related_to",
properties,
size: 1,
baseSize: 1,
color: baseColor,
baseColor,
mutedColor: GRAPH_THEME.palette.muted.edgeOverview,
visualPriority: isInferred ? 0.95 : 0.5,
isBidirectional,
edgeFamily: isInferred ? "path" : isBidirectional ? "bidirectional" : "line",
curveGroup: isBidirectional ? curveGroupForPair(payload.source_id, payload.target_id) : null,
type: "line",
edgeVariant: isInferred ? "pathSignal" : isBidirectional ? "bidirectionalCurve" : "directional",
arrowVisibilityPolicy: isInferred ? "always" : "contextual",
relationshipStrength: isInferred ? 0.95 : 0.52,
isParallelPair: false,
parallelIndex: 0,
parallelCount: 1,
familySize: 1,
};
}
function buildSelectedNodeState(
@@ -1320,7 +1355,6 @@ export function GraphWorkspace({ externalFocusNodeId, externalFocusToken }: Grap
const lastExternalFocusTokenRef = useRef<number | undefined>(undefined);
const pluginRuntimeRef = useRef<GraphSceneRuntime | null>(null);
const appliedGraphSummarySignatureRef = useRef<string | null>(null);
const smallGraphModeRef = useRef(false);
const pluginInteractionStateRef = useRef<GraphInteractionState>({
hoveredNodeId: null,
selectedNodeId: "",
@@ -1348,12 +1382,6 @@ export function GraphWorkspace({ externalFocusNodeId, externalFocusToken }: Grap
}
appliedGraphSummarySignatureRef.current = signature;
smallGraphModeRef.current = Boolean(
graphSummary.layoutReady
&& !graphSummary.hasCoordinates
&& graphSummary.nodeCount > 0
&& graphSummary.nodeCount <= SMALL_GRAPH_MAX_NODES,
);
setGraphReady(true);
setGraphVersion((current) => current + 1);
setIsLayoutRunning(!graphSummary.layoutReady);
@@ -1865,26 +1893,18 @@ export function GraphWorkspace({ externalFocusNodeId, externalFocusToken }: Grap
attributes: buildRealtimeNodeAttributes(payload),
},
]);
if (graph.order > SMALL_GRAPH_MAX_NODES) {
smallGraphModeRef.current = false;
}
synchronizeRealtimeSmallGraphEdges(smallGraphModeRef.current);
logEvent("add-node", `Added node ${payload.label ?? payload.id}${payload.nodeType ? ` (${payload.nodeType})` : ""} via realtime ws`, { nodeId: payload.id, nodeType: payload.nodeType });
setGraphVersion((current) => current + 1);
sceneRef.current?.getRuntime()?.requestRender();
}
if (eventType === "ADD_EDGE") {
const isSmallGraph = smallGraphModeRef.current;
batchMergeEdges([
{
id: String(payload.id),
familyId: payload.familyId ? String(payload.familyId) : String(payload.id),
source: payload.source_id,
target: payload.target_id,
attributes: buildRealtimeEdgeAttributes(payload, {
isBidirectional: graph.hasDirectedEdge(payload.target_id, payload.source_id),
isSmallGraph,
}),
attributes: buildRealtimeEdgeAttributes(payload),
},
]);
logEvent("add-edge", `Added edge ${payload.edgeType ?? payload.id} (${payload.source_id}${payload.target_id}) via realtime ws`, { edgeId: payload.id, edgeType: payload.edgeType, source: payload.source_id, target: payload.target_id });
@@ -1783,7 +1783,6 @@ export function resolveEdgeElementStyle(
const isCommunityBundle = attrs.bundleKind === "community";
const baseSize = Number(attrs.baseSize || attrs.size || 0.9);
const visualPriority = Number(attrs.visualPriority ?? 0);
const isSmallGraphEdge = viewMode === "full" && attrs.isSmallGraph === true;
const isFullBridgeEdge = viewMode === "full" && fullEdgeClass === "bridge";
const isFullBackboneEdge = viewMode === "full" && fullEdgeClass === "backbone";
const shouldCurveBridge = isFullBridgeEdge
@@ -1791,13 +1790,11 @@ export function resolveEdgeElementStyle(
const visibilityPolicy = resolveEdgeVisibilityPolicy(theme, viewMode, zoomTier, isCommunityBundle);
const isContextEdge = isContextEdgeState(state);
const isNonCriticalEdge = isNonCriticalEdgeVariant(edgeVariant);
const belowPriorityThreshold = !isSmallGraphEdge && state === "default"
const belowPriorityThreshold = state === "default"
&& visualPriority < Math.max(tierConfig.edgePriorityThreshold, visibilityPolicy.defaultPriorityThreshold)
&& isNonCriticalEdge;
const hiddenByMutedState = !isSmallGraphEdge
&& (state === "muted" || state === "inactive")
&& visibilityPolicy.hideMuted;
const sampledOut = !isSmallGraphEdge && isNonCriticalEdge
const hiddenByMutedState = (state === "muted" || state === "inactive") && visibilityPolicy.hideMuted;
const sampledOut = isNonCriticalEdge
&& (
(state === "default" && !isContextEdge && shouldSampleOutBackgroundEdge(visibilityPolicy.backgroundSampleRate, visualPriority, edgeId, sourceId, targetId))
|| (
@@ -1840,14 +1837,11 @@ export function resolveEdgeElementStyle(
? resolveEdgeCurvature(theme, state, edgeVariant, attrs, sourceId, targetId)
: 0;
const baseColor = resolveEdgeColor(theme, zoomTier, state, attrs, attrs.color, fullEdgeClass);
const resolvedLodAlpha = resolveEdgeLodAlpha(theme, viewMode, zoomTier, state, attrs, isCommunityBundle, fullEdgeClass);
const lodAlpha = isSmallGraphEdge
? Math.max(resolvedLodAlpha ?? 1, isContextEdge ? 0.62 : 0.46)
: resolvedLodAlpha;
const lodAlpha = resolveEdgeLodAlpha(theme, viewMode, zoomTier, state, attrs, isCommunityBundle, fullEdgeClass);
const color = lodAlpha === null ? baseColor : withAlpha(baseColor, lodAlpha);
const rawSize = Math.max(
baseSize * sizeMultiplier * (isCommunityBundle ? theme.grouped.style.edgeSizeScale : 1),
isSmallGraphEdge ? Math.max(stateConfig.minSize, 0.9) : stateConfig.minSize,
stateConfig.minSize,
);
const interactionMaxSize = (fullEdgeClass === "path" || state === "path")
@@ -1,50 +0,0 @@
import type { EdgeAttributes } from "../../store/graphStore";
import { curveGroupForPair } from "../../store/edgePairKeys.js";
import { GRAPH_THEME } from "./graphTheme";
export type RealtimeEdgePayload = {
id: string;
familyId?: string;
source_id: string;
target_id: string;
type?: string;
weight?: number;
properties?: Record<string, unknown>;
};
export function buildRealtimeEdgeAttributes(
payload: RealtimeEdgePayload,
options: { isBidirectional: boolean; isSmallGraph: boolean },
): EdgeAttributes {
const properties = payload.properties || {};
const isInferred = Boolean(properties.inferred);
const baseColor = isInferred ? GRAPH_THEME.palette.accent.path : GRAPH_THEME.palette.muted.edgeStructure;
return {
edgeId: payload.id,
familyId: payload.familyId || payload.id,
sourceId: payload.source_id,
targetId: payload.target_id,
weight: Number(payload.weight ?? 1),
edgeType: payload.type || "related_to",
properties,
size: 1,
baseSize: 1,
color: baseColor,
baseColor,
mutedColor: GRAPH_THEME.palette.muted.edgeOverview,
visualPriority: isInferred ? 0.95 : 0.5,
isBidirectional: options.isBidirectional,
edgeFamily: isInferred ? "path" : options.isBidirectional ? "bidirectional" : "line",
curveGroup: options.isBidirectional ? curveGroupForPair(payload.source_id, payload.target_id) : null,
type: "line",
edgeVariant: isInferred ? "pathSignal" : options.isBidirectional ? "bidirectionalCurve" : "directional",
arrowVisibilityPolicy: isInferred ? "always" : "contextual",
relationshipStrength: isInferred ? 0.95 : 0.52,
isParallelPair: false,
parallelIndex: 0,
parallelCount: 1,
familySize: 1,
isSmallGraph: options.isSmallGraph,
};
}
@@ -1,135 +0,0 @@
export const SMALL_GRAPH_MAX_NODES = 48;
const PROVIDED_COORDINATE_COVERAGE = 0.92;
const MAX_COMPONENT_RADIUS = 78;
const COMPONENT_GAP = 48;
type LayoutEdge = {
source: string;
target: string;
};
export function shouldUseSmallGraphLayout(nodeCount: number, coordinateCoverage: number): boolean {
return nodeCount > 0
&& nodeCount <= SMALL_GRAPH_MAX_NODES
&& coordinateCoverage < PROVIDED_COORDINATE_COVERAGE;
}
export function resolveGraphLayoutDecision(nodeCount: number, coordinateCoverage: number): {
useProvidedCoordinates: boolean;
useSmallGraphLayout: boolean;
layoutReady: boolean;
} {
const useProvidedCoordinates = coordinateCoverage >= PROVIDED_COORDINATE_COVERAGE;
const useSmallGraphLayout = shouldUseSmallGraphLayout(nodeCount, coordinateCoverage);
return {
useProvidedCoordinates,
useSmallGraphLayout,
layoutReady: useProvidedCoordinates || useSmallGraphLayout,
};
}
export function resolveNodeLayoutPosition(
decision: ReturnType<typeof resolveGraphLayoutDecision>,
provided: { x: number | null; y: number | null },
seeded: { x: number; y: number } | undefined,
): { x: number; y: number } {
if (decision.useProvidedCoordinates) {
return { x: provided.x ?? 0, y: provided.y ?? 0 };
}
if (decision.useSmallGraphLayout) {
return { x: seeded?.x ?? 0, y: seeded?.y ?? 0 };
}
return {
x: provided.x ?? seeded?.x ?? 0,
y: provided.y ?? seeded?.y ?? 0,
};
}
/**
* Produce a compact deterministic layout for small graphs.
*
* ForceAtlas2 is useful for large connected datasets, but it makes tiny graphs
* with several disconnected components look like scattered dots. This layout
* keeps each connected component together and packs components into a centered
* grid so instance relationships remain legible on first render.
*/
export function buildSmallGraphSeedPositions(
nodeIds: string[],
edges: LayoutEdge[],
): Map<string, { x: number; y: number }> {
const ids = [...new Set(nodeIds)].sort((left, right) => left.localeCompare(right));
const adjacency = new Map(ids.map((id) => [id, new Set<string>()]));
edges.forEach(({ source, target }) => {
if (!adjacency.has(source) || !adjacency.has(target) || source === target) {
return;
}
adjacency.get(source)?.add(target);
adjacency.get(target)?.add(source);
});
const visited = new Set<string>();
const components: string[][] = [];
ids.forEach((start) => {
if (visited.has(start)) {
return;
}
const component: string[] = [];
const queue = [start];
visited.add(start);
while (queue.length > 0) {
const current = queue.shift();
if (!current) {
continue;
}
component.push(current);
[...(adjacency.get(current) ?? [])]
.sort((left, right) => left.localeCompare(right))
.forEach((neighbor) => {
if (!visited.has(neighbor)) {
visited.add(neighbor);
queue.push(neighbor);
}
});
}
component.sort((left, right) => {
const degreeDelta = (adjacency.get(right)?.size ?? 0) - (adjacency.get(left)?.size ?? 0);
return degreeDelta || left.localeCompare(right);
});
components.push(component);
});
components.sort((left, right) => right.length - left.length || left[0].localeCompare(right[0]));
const columns = Math.max(1, Math.ceil(Math.sqrt(components.length)));
const rows = Math.max(1, Math.ceil(components.length / columns));
// Adjacent cells must leave room for two maximum-radius components plus a
// readable gap. A smaller row height allows valid 12-node components to
// overlap vertically.
const cellWidth = MAX_COMPONENT_RADIUS * 2 + COMPONENT_GAP;
const cellHeight = MAX_COMPONENT_RADIUS * 2 + COMPONENT_GAP;
const positions = new Map<string, { x: number; y: number }>();
components.forEach((component, componentIndex) => {
const column = componentIndex % columns;
const row = Math.floor(componentIndex / columns);
const centerX = (column - (columns - 1) / 2) * cellWidth;
const centerY = (row - (rows - 1) / 2) * cellHeight;
if (component.length === 1) {
positions.set(component[0], { x: centerX, y: centerY });
return;
}
const radius = Math.min(MAX_COMPONENT_RADIUS, 30 + component.length * 9);
component.forEach((nodeId, nodeIndex) => {
const angle = -Math.PI / 2 + (nodeIndex * Math.PI * 2) / component.length;
positions.set(nodeId, {
x: centerX + Math.cos(angle) * radius,
y: centerY + Math.sin(angle) * radius,
});
});
});
return positions;
}
@@ -16,11 +16,6 @@ import {
} from "./graphTheme";
import { classifyEntityShape } from "./graphEntityShape";
import { createGraphLoadProgress } from "./graphLoading";
import {
buildSmallGraphSeedPositions,
resolveGraphLayoutDecision,
resolveNodeLayoutPosition,
} from "./smallGraphLayout";
import type { GraphLoadProgress, GraphLoadSummary } from "./types";
const SEMANTIC_COLOR_FIELDS = [
@@ -558,19 +553,10 @@ export function useLoadGraph(options: UseLoadGraphOptions = {}) {
: count;
}, 0);
const coordinateCoverage = fetchedNodes.length > 0 ? providedCoordinateCount / fetchedNodes.length : 0;
const {
useProvidedCoordinates,
useSmallGraphLayout,
layoutReady,
} = resolveGraphLayoutDecision(fetchedNodes.length, coordinateCoverage);
const useProvidedCoordinates = coordinateCoverage >= 0.92;
const seededPositions = useProvidedCoordinates
? null
: useSmallGraphLayout
? buildSmallGraphSeedPositions(
fetchedNodes.map((node) => node.id),
fetchedEdges,
)
: buildClusterSeedPositions(
: buildClusterSeedPositions(
draftAttributes.map(({ id, attributes }) => ({
id,
semanticGroup: semanticKeyByNodeId.get(id) ?? structuralColorKey(id, attributes),
@@ -583,9 +569,7 @@ export function useLoadGraph(options: UseLoadGraphOptions = {}) {
const colorIndex = hashString(semanticGroup) % GRAPH_THEME.palette.semantic.length;
const baseColor = GRAPH_THEME.palette.semantic[colorIndex];
const sizeRatio = nodePriorityById.get(id) ?? 0;
const dynamicSize = useSmallGraphLayout
? clamp(5.2, 5.2 + 6.6 * sizeRatio, 11.8)
: clamp(1.8, 1.8 + 8.8 * sizeRatio, 11.8);
const dynamicSize = clamp(1.8, 1.8 + 8.8 * sizeRatio, 11.8);
const hasTemporalBounds = Boolean(attributes.valid_from || attributes.valid_until);
const provenanceCount = getProvenanceCount(attributes.properties ?? {});
const properties = attributes.properties as Record<string, unknown>;
@@ -593,11 +577,12 @@ export function useLoadGraph(options: UseLoadGraphOptions = {}) {
const providedX = readFiniteCoordinate(properties?.x);
const providedY = readFiniteCoordinate(properties?.y);
const seededPosition = seededPositions?.get(id);
const { x, y } = resolveNodeLayoutPosition(
{ useProvidedCoordinates, useSmallGraphLayout, layoutReady },
{ x: providedX, y: providedY },
seededPosition,
);
const x = useProvidedCoordinates
? providedX ?? 0
: providedX ?? seededPosition?.x ?? 0;
const y = useProvidedCoordinates
? providedY ?? 0
: providedY ?? seededPosition?.y ?? 0;
return {
id,
attributes: {
@@ -618,7 +603,6 @@ export function useLoadGraph(options: UseLoadGraphOptions = {}) {
borderSize: 0.72,
entityShape,
...resolveNodeVariantMetadata(baseColor, sizeRatio, hasTemporalBounds, provenanceCount),
...(useSmallGraphLayout ? { labelVisibilityPolicy: "always" as const } : {}),
} as NodeAttributes,
};
});
@@ -675,7 +659,6 @@ export function useLoadGraph(options: UseLoadGraphOptions = {}) {
parallelIndex,
parallelCount,
familySize: familyCounts.get(edge.familyId) ?? 1,
isSmallGraph: useSmallGraphLayout,
...resolveEdgeVariantMetadata(edge, sourcePriority, targetPriority, isBidirectional),
} as EdgeAttributes,
};
@@ -718,7 +701,7 @@ export function useLoadGraph(options: UseLoadGraphOptions = {}) {
loadTimeMs: Math.round(performance.now() - startedAt),
hasCoordinates: useProvidedCoordinates,
layoutSource: useProvidedCoordinates ? "provided" : "runtime",
layoutReady,
layoutReady: useProvidedCoordinates,
} satisfies GraphLoadSummary;
onProgress?.(createGraphLoadProgress({
@@ -407,31 +407,6 @@ test("resolveEdgeElementStyle applies full-graph LOD to directional background e
assert.equal(style.hidden, true);
});
test("resolveEdgeElementStyle keeps small-graph relationships visible in overview", () => {
const style = resolveEdgeElementStyle(
GRAPH_THEME,
"overview",
"inactive",
{
edgeType: "related_to",
weight: 1,
properties: {},
edgeVariant: "directional",
visualPriority: 0.1,
baseSize: 0.5,
isSmallGraph: true,
},
"source",
"target",
"full",
"small-graph-low-priority",
"hidden",
);
assert.equal(style.hidden, false);
assert.ok(Number(style.size ?? 0) >= 0.9);
});
test("classifyFullGraphEdge applies deterministic priority order", () => {
const edgeClass = classifyFullGraphEdge(
"edge-priority",
@@ -1,31 +0,0 @@
import assert from "node:assert/strict";
import test from "node:test";
import { buildRealtimeEdgeAttributes } from "../src/workspaces/GraphWorkspace/realtimeGraphAttributes.ts";
const payload = {
id: "edge-live",
source_id: "source",
target_id: "target",
type: "related_to",
properties: {},
};
test("realtime edges retain the active small-graph visibility marker", () => {
const attributes = buildRealtimeEdgeAttributes(payload, {
isBidirectional: false,
isSmallGraph: true,
});
assert.equal(attributes.isSmallGraph, true);
assert.equal(attributes.edgeVariant, "directional");
});
test("realtime edges do not retain the marker after graph leaves small-graph mode", () => {
const attributes = buildRealtimeEdgeAttributes(payload, {
isBidirectional: false,
isSmallGraph: false,
});
assert.equal(attributes.isSmallGraph, false);
});
-100
View File
@@ -1,100 +0,0 @@
import assert from "node:assert/strict";
import test from "node:test";
import {
SMALL_GRAPH_MAX_NODES,
buildSmallGraphSeedPositions,
resolveGraphLayoutDecision,
resolveNodeLayoutPosition,
shouldUseSmallGraphLayout,
} from "../src/workspaces/GraphWorkspace/smallGraphLayout.ts";
test("small graph layout is selected only when coordinates are not already usable", () => {
assert.equal(shouldUseSmallGraphLayout(12, 0), true);
assert.equal(shouldUseSmallGraphLayout(SMALL_GRAPH_MAX_NODES + 1, 0), false);
assert.equal(shouldUseSmallGraphLayout(12, 0.95), false);
});
test("small graph layout ignores isolated partial coordinates", () => {
const decision = resolveGraphLayoutDecision(12, 1 / 12);
assert.deepEqual(
resolveNodeLayoutPosition(decision, { x: 50_000, y: -50_000 }, { x: 24, y: -18 }),
{ x: 24, y: -18 },
);
assert.deepEqual(
resolveNodeLayoutPosition(decision, { x: 50_000, y: null }, { x: -12, y: 36 }),
{ x: -12, y: 36 },
);
});
test("small graph load is immediately ready and skips runtime stabilization", () => {
assert.deepEqual(resolveGraphLayoutDecision(12, 0), {
useProvidedCoordinates: false,
useSmallGraphLayout: true,
layoutReady: true,
});
assert.deepEqual(resolveGraphLayoutDecision(SMALL_GRAPH_MAX_NODES + 1, 0), {
useProvidedCoordinates: false,
useSmallGraphLayout: false,
layoutReady: false,
});
assert.deepEqual(resolveGraphLayoutDecision(12, 1), {
useProvidedCoordinates: true,
useSmallGraphLayout: false,
layoutReady: true,
});
});
test("small graph layout is deterministic and keeps connected nodes together", () => {
const nodes = ["Apple", "Steve", "Ronald", "Cupertino", "California"];
const edges = [
{ source: "Apple", target: "Steve" },
{ source: "Ronald", target: "Cupertino" },
];
const first = buildSmallGraphSeedPositions(nodes, edges);
const second = buildSmallGraphSeedPositions([...nodes].reverse(), [...edges].reverse());
assert.deepEqual([...first.entries()].sort(), [...second.entries()].sort());
assert.equal(first.size, nodes.length);
const distance = (left: string, right: string) => {
const a = first.get(left);
const b = first.get(right);
assert.ok(a && b);
return Math.hypot(a.x - b.x, a.y - b.y);
};
assert.ok(distance("Apple", "Steve") < distance("Apple", "California"));
assert.ok(distance("Ronald", "Cupertino") < distance("Ronald", "California"));
});
test("small graph layout keeps maximum-radius components separated", () => {
const componentCount = 4;
const nodesPerComponent = 12;
const nodes = Array.from(
{ length: componentCount * nodesPerComponent },
(_, index) => `component-${Math.floor(index / nodesPerComponent)}-node-${index % nodesPerComponent}`,
);
const edges = Array.from({ length: componentCount }).flatMap((_, componentIndex) => {
const prefix = `component-${componentIndex}-node-`;
return Array.from({ length: nodesPerComponent - 1 }, (_unused, nodeIndex) => ({
source: `${prefix}${nodeIndex}`,
target: `${prefix}${nodeIndex + 1}`,
}));
});
const positions = buildSmallGraphSeedPositions(nodes, edges);
for (let leftComponent = 0; leftComponent < componentCount; leftComponent += 1) {
for (let rightComponent = leftComponent + 1; rightComponent < componentCount; rightComponent += 1) {
let closestDistance = Number.POSITIVE_INFINITY;
for (let leftNode = 0; leftNode < nodesPerComponent; leftNode += 1) {
for (let rightNode = 0; rightNode < nodesPerComponent; rightNode += 1) {
const left = positions.get(`component-${leftComponent}-node-${leftNode}`);
const right = positions.get(`component-${rightComponent}-node-${rightNode}`);
assert.ok(left && right);
closestDistance = Math.min(closestDistance, Math.hypot(left.x - right.x, left.y - right.y));
}
}
assert.ok(closestDistance >= 48, `components are only ${closestDistance} units apart`);
}
}
});
+7 -33
View File
@@ -8,9 +8,8 @@ Connects Claude Code, Cursor, Windsurf, Cline, Continue, VS Code (GitHub Copilot
## Quick start
```bash
# From the repo root — no extra install flag needed; the root mcp/ package is
# part of the repository and does not require an external MCP SDK.
pip install -e .
# From the repo root
pip install -e ".[mcp]"
# Test the server (type a JSON-RPC request, press Enter)
python -m mcp
@@ -90,14 +89,7 @@ python -m mcp [--debug]
## Per-tool configuration
### Claude Code (`~/.claude.json` or `.mcp.json`)
Claude Code supports two MCP configuration scopes:
- **User scope**`~/.claude.json` applies across all projects for your user account.
- **Project scope**`.mcp.json` in your project root applies only to that project.
Both files use the same `mcpServers` structure:
### Claude Code (`~/.claude/settings.json`)
```json
{
@@ -105,33 +97,15 @@ Both files use the same `mcpServers` structure:
"semantica": {
"command": "python",
"args": ["-m", "mcp"],
"env": {
"PYTHONPATH": "/path/to/semantica"
}
"cwd": "/path/to/semantica"
}
}
}
```
> **Why `PYTHONPATH`?** The root `mcp/` package is intentionally not included in
> the installed wheel, so `python -m mcp` only works when the repository is on
> Python's import path. Setting `PYTHONPATH` here ensures this works regardless
> of the working directory Claude uses when it launches the server.
Or add it via the CLI (user scope):
Or use the plugin bundle:
```bash
claude mcp add --scope user semantica \
-e PYTHONPATH=/path/to/semantica \
-- python -m mcp
```
Or for project scope (omit `--scope user`):
```bash
claude mcp add semantica \
-e PYTHONPATH=/path/to/semantica \
-- python -m mcp
claude mcp add semantica python -m mcp --cwd /path/to/semantica
```
---
@@ -242,7 +216,7 @@ Add to your Q Developer MCP config:
| Variable | Default | Description |
|---|---|---|
| `SEMANTICA_KG_PATH` | *(in-memory only)* | Path to a JSON file used to **load** the graph on startup and **persist** mutations (record decisions, add entities/relationships) back to disk after each change. When unset the graph lives in memory only and is lost when the server exits. |
| `SEMANTICA_KG_PATH` | *(in-memory)* | Path to persist/load the graph (JSON file) |
---
+7 -35
View File
@@ -16,13 +16,6 @@ log = logging.getLogger("semantica.mcp.session")
_graph: Optional[Any] = None
# Tracks whether the last graph initialisation successfully loaded the
# configured SEMANTICA_KG_PATH file. When True (or no path was configured)
# mutation handlers are allowed to save. When False an existing file failed
# to load; saving would overwrite the original data with an empty graph, so
# persistence is blocked until the process is restarted with a readable file.
_load_ok: bool = True
def get_graph() -> Any:
"""
@@ -31,45 +24,24 @@ def get_graph() -> Any:
The graph is created with advanced_analytics=True so all centrality,
community-detection, and embedding features are available.
"""
global _graph, _load_ok
global _graph
if _graph is None:
from semantica.context import ContextGraph
_graph = ContextGraph(advanced_analytics=True)
_load_ok = True # default: safe to persist
kg_path = os.environ.get("SEMANTICA_KG_PATH", "").strip()
if kg_path and os.path.exists(kg_path):
# Only attempt to load if the file has content. An empty file
# means the path was just created (e.g. a fresh tempfile) and
# should be treated as "start with empty graph" rather than a
# corrupt-file failure.
if os.path.getsize(kg_path) > 0:
try:
_graph.load_from_file(kg_path)
log.info("Graph loaded from %s", kg_path)
except Exception as exc:
log.warning(
"Could not load graph from %s: %s — persistence disabled "
"to protect existing data; restart the server to retry.",
kg_path, exc,
)
_load_ok = False # do not overwrite the original file
try:
_graph.load(kg_path)
log.info("Graph loaded from %s", kg_path)
except Exception as exc:
log.warning("Could not load graph from %s: %s", kg_path, exc)
return _graph
def is_persistence_safe() -> bool:
"""Return True when it is safe to write mutations back to SEMANTICA_KG_PATH.
Returns False after a failed load so that mutation handlers do not
overwrite the original (possibly intact) file with a fresh empty graph.
"""
return _load_ok
def reset_graph() -> None:
"""Reset the singleton (mainly useful in tests)."""
global _graph, _load_ok
global _graph
_graph = None
_load_ok = True
+1 -35
View File
@@ -5,7 +5,6 @@ Decision intelligence tools — record, query, precedents, causal chain, impact.
from __future__ import annotations
import logging
import os
from mcp.schemas import (
ANALYZE_DECISION_IMPACT,
@@ -14,7 +13,7 @@ from mcp.schemas import (
QUERY_DECISIONS,
RECORD_DECISION,
)
from mcp.session import get_graph, is_persistence_safe
from mcp.session import get_graph
log = logging.getLogger("semantica.mcp.tools.decisions")
@@ -38,39 +37,6 @@ def handle_record_decision(args: dict) -> dict:
valid_from=args.get("valid_from"),
valid_until=args.get("valid_until"),
)
# Persist back to disk so the decision survives server restarts.
# Skip when the initial load failed to avoid overwriting original data.
kg_path = os.environ.get("SEMANTICA_KG_PATH", "").strip()
if kg_path:
if not is_persistence_safe():
# Roll back the in-memory mutation so the client-visible state
# matches the persisted state (neither is saved).
if hasattr(graph, "_decisions") and decision_id in graph._decisions:
del graph._decisions[decision_id]
if hasattr(graph, "_decision_index"):
cat = args.get("category", "")
if cat in graph._decision_index:
graph._decision_index[cat].discard(decision_id)
return {
"error": (
"Persistence blocked: the configured SEMANTICA_KG_PATH "
"could not be loaded at startup. Restart the server with "
"a readable graph file to re-enable persistence."
)
}
try:
graph.save_to_file(kg_path)
except Exception as save_exc:
# Atomic write failed. Roll back the in-memory mutation so the
# client-visible and persisted states remain consistent.
if hasattr(graph, "_decisions") and decision_id in graph._decisions:
del graph._decisions[decision_id]
if hasattr(graph, "_decision_index"):
cat = args.get("category", "")
if cat in graph._decision_index:
graph._decision_index[cat].discard(decision_id)
log.exception("save_to_file failed after record_decision; mutation rolled back")
return {"error": f"Mutation rolled back: could not persist graph: {save_exc}"}
return {
"decision_id": decision_id,
"status": "recorded",
+1 -70
View File
@@ -5,10 +5,9 @@ Graph tools — add entities/relationships, search, analytics, summary.
from __future__ import annotations
import logging
import os
from mcp.schemas import ADD_ENTITY, ADD_RELATIONSHIP, EMPTY, GET_ANALYTICS, SEARCH_GRAPH
from mcp.session import get_graph, is_persistence_safe
from mcp.session import get_graph
log = logging.getLogger("semantica.mcp.tools.graph")
@@ -26,35 +25,6 @@ def handle_add_entity(args: dict) -> dict:
node_type=args.get("type", "Entity"),
metadata=args.get("metadata", {}),
)
# Persist back to disk so the entity survives server restarts.
# Skip when the initial load failed to avoid overwriting original data.
kg_path = os.environ.get("SEMANTICA_KG_PATH", "").strip()
if kg_path:
if not is_persistence_safe():
# Roll back: remove the node we just added.
try:
with graph._lock:
graph._drop_node_from_indexes(node_id)
except Exception:
pass
return {
"error": (
"Persistence blocked: the configured SEMANTICA_KG_PATH "
"could not be loaded at startup. Restart the server with "
"a readable graph file to re-enable persistence."
)
}
try:
graph.save_to_file(kg_path)
except Exception as save_exc:
# Roll back: remove the node so in-memory and persisted state agree.
try:
with graph._lock:
graph._drop_node_from_indexes(node_id)
except Exception:
pass
log.exception("save_to_file failed after add_entity; mutation rolled back")
return {"error": f"Mutation rolled back: could not persist graph: {save_exc}"}
return {"status": "added", "id": node_id, "type": args.get("type", "Entity")}
except Exception as exc:
log.exception("add_entity failed")
@@ -76,45 +46,6 @@ def handle_add_relationship(args: dict) -> dict:
edge_type=rel_type,
metadata=args.get("metadata", {}),
)
# Persist back to disk so the relationship survives server restarts.
# Skip when the initial load failed to avoid overwriting original data.
kg_path = os.environ.get("SEMANTICA_KG_PATH", "").strip()
if kg_path:
if not is_persistence_safe():
# Roll back: remove the edge we just added (last matching edge).
try:
with graph._lock:
for edge in reversed(list(graph.edges)):
if (edge.source_id == source
and edge.target_id == target
and edge.edge_type == rel_type):
graph._drop_edge_from_indexes(edge)
break
except Exception:
pass
return {
"error": (
"Persistence blocked: the configured SEMANTICA_KG_PATH "
"could not be loaded at startup. Restart the server with "
"a readable graph file to re-enable persistence."
)
}
try:
graph.save_to_file(kg_path)
except Exception as save_exc:
# Roll back: remove the edge so in-memory and persisted state agree.
try:
with graph._lock:
for edge in reversed(list(graph.edges)):
if (edge.source_id == source
and edge.target_id == target
and edge.edge_type == rel_type):
graph._drop_edge_from_indexes(edge)
break
except Exception:
pass
log.exception("save_to_file failed after add_relationship; mutation rolled back")
return {"error": f"Mutation rolled back: could not persist graph: {save_exc}"}
return {"status": "added", "source": source, "target": target, "type": rel_type}
except Exception as exc:
log.exception("add_relationship failed")
+1 -5
View File
@@ -25,9 +25,5 @@
"mcp"
],
"skills": "./skills",
"agents": [
"./agents/decision-advisor.md",
"./agents/explainability.md",
"./agents/kg-assistant.md"
]
"agents": "./agents"
}
+1 -8
View File
@@ -49,14 +49,7 @@ dependencies = [
"scipy>=1.13.1",
"scikit-learn>=1.7.2",
"umap-learn>=0.5.12",
# thinc (spacy's core dep) dropped Python 3.9 wheels at 8.3.10, and later
# spacy patch releases (3.8.8+) require thinc>=8.3.9-only-on-3.10+ ranges,
# which forces a source build that fails outright on 3.9 (see Install
# Matrix run history). Capping both keeps 3.9 on the last wheel-compatible
# pair; 3.10+ is left unconstrained to always get the latest spacy/thinc.
"spacy>=3.4.0,<3.8.8; python_version < '3.10'",
"spacy>=3.4.0; python_version >= '3.10'",
"thinc<8.3.5; python_version < '3.10'",
"spacy>=3.4.0",
"transformers>=4.20.0",
"torch>=1.13.1",
"sentence-transformers>=2.2.0",
-4
View File
@@ -111,7 +111,6 @@ from .context_graph import ContextEdge, ContextGraph, ContextNode
from .context_retriever import ContextRetriever, RetrievedContext, TemporalGraphRetriever
from .decision_context import DecisionContext
from .entity_linker import EntityLink, EntityLinker, LinkedEntity
from .erasure import ErasureCoordinator, ErasureReceipt
# Decision tracking imports
from .decision_models import (
@@ -146,9 +145,6 @@ __all__ = [
"ContextRetriever",
"RetrievedContext",
"TemporalGraphRetriever",
# Cross-store erasure
"ErasureCoordinator",
"ErasureReceipt",
# Decision tracking models
"Decision",
"DecisionContextModel",
-21
View File
@@ -626,27 +626,6 @@ class AgentMemory:
self.logger.debug(f"Deleted memory item: {memory_id}")
return True
def vector_ids_for(self, memory_id: str) -> List[str]:
"""Return the vector-store ids owned by a memory item.
Read-only view of the ids ``delete_memory()`` would remove for this
item, so a caller that needs to *report* on vector removal can delete
them itself rather than relying on ``delete_memory()``'s best-effort
cascade, which logs a vector-store failure and still returns ``True``.
Mirrors the fallback in ``delete_memory``: an item stored without
tracked vector ids is keyed in the vector store by its own memory id.
Args:
memory_id: Memory identifier.
Returns:
The item's vector ids, or ``[]`` if the item is unknown.
"""
if memory_id not in self.memory_items:
return []
return list(self._vector_ids.get(memory_id, [])) or [memory_id]
def clear_memory(self, **filters) -> int:
"""
Clear memory items matching filters.
+3 -29
View File
@@ -1203,30 +1203,8 @@ class ContextGraph:
"links": links_data,
}
# Write atomically: serialize to a sibling temp file then replace the
# destination in one OS-level rename. This guarantees the destination
# is either the old contents or the new contents — never a partial write
# — so a crash or disk-full error during json.dump cannot corrupt the
# sole persisted copy of the graph.
dest = Path(path)
dest.parent.mkdir(parents=True, exist_ok=True)
fd, tmp_path = tempfile.mkstemp(
dir=dest.parent, prefix=".kg_tmp_", suffix=".json"
)
try:
with os.fdopen(fd, "w", encoding="utf-8") as f:
json.dump(data, f, indent=2, ensure_ascii=False)
f.flush()
os.fsync(f.fileno())
os.replace(tmp_path, dest)
except Exception:
# Clean up the temp file on any failure so we don't litter the
# directory with partial writes.
try:
os.unlink(tmp_path)
except OSError:
pass
raise
with open(path, "w", encoding="utf-8") as f:
json.dump(data, f, indent=2, ensure_ascii=False)
self.logger.info(f"Saved context graph to {path}")
@@ -2662,11 +2640,7 @@ class ContextGraph:
Scope is this graph only. Copies held elsewhere (``AgentMemory``, a
bound vector store, an exported file) are not reached, so this is one
step of an erasure workflow, not the whole of it. Callers who need the
whole workflow -- and a receipt recording which stores it actually
reached -- should drive this through
:class:`~semantica.context.erasure.ErasureCoordinator` rather than
treating a ``True`` here as proof the content is gone.
step of an erasure workflow, not the whole of it.
Args:
node_id: Node to purge.
-76
View File
@@ -239,82 +239,6 @@ print(f"Python importance score: {importance.get('degree', 0)}")
---
## 🧹 Erasing an Entity Everywhere - ErasureCoordinator
`purge_node()` removes an entity from **one graph**. The same content can still be
sitting in agent memory and in your vector store, so purge on its own is one step
of an erasure workflow rather than the whole of it.
`ErasureCoordinator` drives the whole cascade and hands you a receipt saying what
it actually managed to erase.
```python
from semantica.context import AgentMemory, ContextGraph, ErasureCoordinator
coordinator = ErasureCoordinator(graph=knowledge, memory=memory)
receipt = coordinator.erase_entity(
"customer-4471",
reason="GDPR Art. 17 request #882",
)
if receipt.complete:
print("Erased everywhere")
else:
print("Still holding data:", receipt.incomplete_stores)
```
### Always Check the Receipt
The receipt is the point of the feature — **do not treat the call itself as proof
the data is gone**. Each store reports one of five statuses:
| Status | Meaning |
|---|---|
| `erased` | Reached, data removed (on the vectors leg: the store accepted the delete for the ids given) |
| `not_found` | Reached, held nothing for this entity |
| `not_configured` | No such store was bound — normal, not a failure |
| `unsupported` | The store cannot delete at all; retrying will not help |
| `failed` | The store was reached and the deletion did not succeed |
```python
receipt.to_dict()
# {
# "entity_id": "customer-4471",
# "reason": "GDPR Art. 17 request #882",
# "erased_at": "2026-08-16T09:03:36.813220",
# "complete": False,
# "stores": {
# "vectors": {"status": "unsupported", "backend": "faiss",
# "detail": "backend exposes no delete()/delete_vectors(); ..."},
# "memory": {"status": "erased", "items": 14},
# "graph": {"status": "erased", "nodes": 1, "edges": 3},
# },
# }
```
`complete` is `False` when any store reports `unsupported` or `failed`, which is
your signal to handle that store out of band. FAISS, Milvus and Weaviate expose
no delete method today, so erasure genuinely cannot be completed on them — the
coordinator says so rather than reporting a success it did not achieve.
### Good to Know
- **Order is vectors → memory → graph.** The graph tombstone is the durable record
that an erasure happened, so it is written last: a crash mid-cascade leaves the
node present and the receipt incomplete, rather than a tombstone claiming more
than actually happened.
- **A failing store does not abort the rest.** Partial failure is recorded in the
receipt and the remaining stores are still erased.
- **Every store is optional.** `ErasureCoordinator(graph=graph)` is fine; the other
legs report `not_configured`.
- **It is idempotent.** Erasing the same entity twice returns a receipt saying
there was nothing left to do, rather than raising.
- **Batch:** `coordinator.erase_entities([...], reason=...)` returns one receipt per
entity, in order, so one entity's failure does not stop the others.
---
## 🔄 Using Both Together - The Complete Setup
### Your Smart Agent System
+9 -15
View File
@@ -76,11 +76,11 @@ Production Use Cases:
- Insurance: Claim decisions, underwriting assessments
"""
import json
import uuid
from dataclasses import InitVar, dataclass, field
from dataclasses import dataclass, field
from datetime import datetime
from typing import Any, Dict, List, Optional
import json
import uuid
@dataclass
@@ -100,9 +100,8 @@ class Decision:
valid_from: Optional[str] = None
valid_until: Optional[str] = None
metadata: Dict[str, Any] = field(default_factory=dict)
auto_generate_id: InitVar[bool] = True
def __post_init__(self, auto_generate_id: bool) -> None:
def __post_init__(self, auto_generate_id: bool = True):
"""Validate decision data."""
if auto_generate_id and not self.decision_id: # Handle both None and empty string
self.decision_id = str(uuid.uuid4())
@@ -147,9 +146,8 @@ class DecisionContext:
risk_factors: List[str]
cross_system_inputs: Dict[str, Any] = field(default_factory=dict)
metadata: Dict[str, Any] = field(default_factory=dict)
auto_generate_id: InitVar[bool] = True
def __post_init__(self, auto_generate_id: bool) -> None:
def __post_init__(self, auto_generate_id: bool = True):
"""Validate decision context data."""
if auto_generate_id and not self.context_id: # Handle both None and empty string
self.context_id = str(uuid.uuid4())
@@ -186,9 +184,8 @@ class Policy:
created_at: datetime
updated_at: datetime
metadata: Dict[str, Any] = field(default_factory=dict)
auto_generate_id: InitVar[bool] = True
def __post_init__(self, auto_generate_id: bool) -> None:
def __post_init__(self, auto_generate_id: bool = True):
"""Validate policy data."""
if auto_generate_id and not self.policy_id: # Handle both None and empty string
self.policy_id = str(uuid.uuid4())
@@ -230,9 +227,8 @@ class PolicyException:
approval_timestamp: datetime
justification: str
metadata: Dict[str, Any] = field(default_factory=dict)
auto_generate_id: InitVar[bool] = True
def __post_init__(self, auto_generate_id: bool) -> None:
def __post_init__(self, auto_generate_id: bool = True):
"""Validate policy exception data."""
if auto_generate_id and not self.exception_id: # Handle both None and empty string
self.exception_id = str(uuid.uuid4())
@@ -269,9 +265,8 @@ class Precedent:
similarity_score: float
relationship_type: str # "similar_scenario", "same_policy", "exception_precedent"
metadata: Dict[str, Any] = field(default_factory=dict)
auto_generate_id: InitVar[bool] = True
def __post_init__(self, auto_generate_id: bool) -> None:
def __post_init__(self, auto_generate_id: bool = True):
"""Validate precedent data."""
if auto_generate_id and not self.precedent_id: # Handle both None and empty string
self.precedent_id = str(uuid.uuid4())
@@ -310,9 +305,8 @@ class ApprovalChain:
approval_context: str
timestamp: datetime
metadata: Dict[str, Any] = field(default_factory=dict)
auto_generate_id: InitVar[bool] = True
def __post_init__(self, auto_generate_id: bool) -> None:
def __post_init__(self, auto_generate_id: bool = True):
"""Validate approval chain data."""
if auto_generate_id and not self.approval_id: # Handle both None and empty string
self.approval_id = str(uuid.uuid4())
-673
View File
@@ -1,673 +0,0 @@
"""
Cross-store erasure coordination.
``ContextGraph.purge_node()`` is graph-scope by design (#957): it removes the
node and leaves a tombstone, but any copy of the same content held in
``AgentMemory`` or in a bound vector store is untouched. That makes purge one
step of an erasure workflow rather than the whole of it, and leaves the caller
to drive the remaining steps by hand -- with no record of which of them
actually succeeded.
:class:`ErasureCoordinator` drives the cascade across the stores it is given
and returns an :class:`ErasureReceipt` describing what was reached and what was
not. It *composes* the existing public APIs; nothing in ``context_graph.py`` or
``agent_memory.py`` changes, and ``ContextGraph`` keeps its graph-scope
contract.
The property that matters is honest partial reporting. Three vector backends
(FAISS, Milvus, Weaviate) expose no delete at all, so erasure is genuinely not
completable on them today. The receipt says ``unsupported`` for those rather
than reporting a success it did not achieve -- a receipt that reads
"graph: erased, memory: 14 erased, vectors: unsupported on faiss" is
actionable; a bare ``True`` is a compliance liability.
Example:
>>> from semantica.context import ContextGraph, AgentMemory
>>> from semantica.context.erasure import ErasureCoordinator
>>> coordinator = ErasureCoordinator(graph=graph, memory=memory)
>>> receipt = coordinator.erase_entity(
... "customer-4471", reason="GDPR Art. 17 request #882"
... )
>>> receipt.complete
False
>>> receipt.stores["vectors"]["status"]
'unsupported'
"""
from dataclasses import dataclass, field
from datetime import datetime, timezone
from typing import Any, Dict, Iterable, List, Optional, Sequence, Tuple, Union
from ..utils.logging import get_logger
from .context_graph import _normalize_temporal_input
__all__ = [
"ErasureCoordinator",
"ErasureReceipt",
"STATUS_ERASED",
"STATUS_NOT_FOUND",
"STATUS_NOT_CONFIGURED",
"STATUS_UNSUPPORTED",
"STATUS_FAILED",
]
#: The store was reached and the entity's data removed from it. On the vectors
#: leg this means the store accepted the delete for the ids it was given: no
#: backend offers a portable "does this id exist" check, so it is not a count of
#: embeddings that were really there. The memory leg re-queries to confirm and
#: so is the stronger claim of the two.
STATUS_ERASED = "erased"
#: The store was reached and held nothing for this entity.
STATUS_NOT_FOUND = "not_found"
#: No such store was bound to the coordinator. Normal, not a failure.
STATUS_NOT_CONFIGURED = "not_configured"
#: The store exists but cannot delete -- e.g. a vector backend with no delete
#: method. Deliberately distinct from ``failed``: retrying will not help.
STATUS_UNSUPPORTED = "unsupported"
#: The store was reached and the deletion did not succeed.
STATUS_FAILED = "failed"
#: Statuses that leave data behind. A receipt containing any of these is not
#: complete, and the shortfall has to be handled out of band.
_INCOMPLETE_STATUSES = frozenset({STATUS_UNSUPPORTED, STATUS_FAILED})
#: Page size for the memory sweep. See ``_erase_memory`` for why the sweep
#: loops rather than passing one large limit.
_MEMORY_SWEEP_BATCH = 500
logger = get_logger("erasure")
@dataclass
class ErasureReceipt:
"""Auditable record of one entity's erasure across every bound store.
Attributes:
entity_id: The entity the erasure was requested for.
reason: Why it was erased, e.g. an erasure-request reference.
erased_at: ISO-8601 timestamp of the erasure.
stores: Per-store outcome keyed by ``"vectors"``, ``"memory"`` and
``"graph"``, each a dict with at least a ``status`` key drawn from
the ``STATUS_*`` constants in this module.
"""
entity_id: str
reason: Optional[str] = None
erased_at: str = ""
stores: Dict[str, Dict[str, Any]] = field(default_factory=dict)
@property
def complete(self) -> bool:
"""True when no bound store was left holding data.
``not_configured`` and ``not_found`` count as complete -- a store that
was never bound, or that held nothing, leaves no residue. Only
``unsupported`` and ``failed`` mean data survived the erasure.
"""
return not self.incomplete_stores
@property
def incomplete_stores(self) -> List[str]:
"""Names of the stores that may still hold the entity's data."""
return [
name
for name, result in self.stores.items()
if result.get("status") in _INCOMPLETE_STATUSES
]
def to_dict(self) -> Dict[str, Any]:
"""Serialize the receipt, deep-copying the per-store results."""
return {
"entity_id": self.entity_id,
"reason": self.reason,
"erased_at": self.erased_at,
"complete": self.complete,
"stores": {name: dict(result) for name, result in self.stores.items()},
}
class ErasureCoordinator:
"""Drives erasure of an entity across the graph, memory and vector stores.
Every store is optional; a store that is not supplied reports
``not_configured`` rather than being silently skipped, so the receipt still
shows the full shape of the workflow.
Args:
graph: A :class:`~semantica.context.ContextGraph` (or anything exposing
``purge_node``).
memory: An :class:`~semantica.context.AgentMemory` (or anything
exposing ``find_by_entity`` and ``batch_delete``).
vector_store: Vector store holding entity-keyed embeddings. Defaults to
``memory.vector_store`` when a memory is supplied, and stays
overridable for deployments that bind a store the memory does not
own. Pass ``False`` to disable the vector leg entirely.
Note:
Erasure runs outward-in -- vectors, then memory, then the graph. The
graph tombstone is the durable attestation that an erasure happened, so
writing it first would let a crash mid-cascade leave a record claiming
more than actually occurred. Erasing the graph last means a partial
failure leaves the node present and the receipt incomplete, which is
recoverable and honest.
"""
def __init__(
self,
graph: Optional[Any] = None,
memory: Optional[Any] = None,
vector_store: Optional[Any] = None,
):
# `is None` / `is False` rather than truthiness: a real store that
# defines __bool__ or __len__ (an empty one, say) is falsey while being
# a perfectly valid store to erase from.
vector_store_given = vector_store is not None and vector_store is not False
if graph is None and memory is None and not vector_store_given:
raise ValueError(
"ErasureCoordinator needs at least one store to erase from; got "
f"graph=None, memory=None, vector_store={vector_store!r}"
)
self.graph = graph
self.memory = memory
if vector_store is False:
self.vector_store: Optional[Any] = None
elif vector_store is not None:
self.vector_store = vector_store
else:
self.vector_store = getattr(memory, "vector_store", None)
self.logger = logger
def erase_entity(
self,
entity_id: str,
reason: Optional[str] = None,
at: Optional[Union[str, int, float, datetime]] = None,
vector_ids: Optional[Sequence[str]] = None,
) -> ErasureReceipt:
"""Erase one entity from every bound store and return a receipt.
A store that cannot be erased from is recorded in the receipt and the
cascade continues -- partial failure is a result, not an exception.
Aborting on the first failure would leave a half-erased state with no
record of which half.
Args:
entity_id: Entity to erase. Interpreted as a graph node id, an
``entities[].id`` in memory items, and a vector id.
reason: Why it was erased, e.g. an erasure-request reference.
Recorded in the receipt and in the graph tombstone.
at: When the erasure takes effect, used as the receipt's
``erased_at`` and passed to ``purge_node`` so both records
carry the same instant. Accepts anything ``ContextGraph``
accepts -- an ISO string, a ``datetime``, or epoch seconds --
and defaults to now, UTC.
vector_ids: Explicit vector ids to remove, in addition to the
ids owned by the entity's memory items, which are always
included. Defaults to ``[entity_id]``, covering entity-keyed
embeddings written by something other than ``AgentMemory``.
Returns:
An :class:`ErasureReceipt`. Check :attr:`ErasureReceipt.complete`
before treating the erasure as done.
"""
# Resolve the timestamp once and hand the *resolved* value to the graph.
# Passing the caller's `at` through instead would let purge_node take its
# own now() when `at` is None, so the receipt and the tombstone it
# attests to would disagree by however long the cascade took.
erased_at = _normalize_timestamp(at)
stores: Dict[str, Dict[str, Any]] = {}
# Outward-in: vectors, then memory, then the graph last.
#
# The vector leg must also cover the embeddings owned by memory items.
# AgentMemory.delete_memory() deletes an item's vectors best-effort: it
# catches a vector-store failure, logs it, and still returns True, so
# the memory leg cannot tell a full erasure from one that left the
# embedding behind. Deleting those ids here instead puts them behind
# the one leg that reports honestly. Collected before anything is
# deleted, while the items still exist to be enumerated.
stores["vectors"] = self._erase_vectors(
entity_id, self._all_vector_ids(entity_id, vector_ids)
)
stores["memory"] = self._erase_memory(entity_id)
stores["graph"] = self._erase_graph(entity_id, reason, erased_at)
receipt = ErasureReceipt(
entity_id=entity_id,
reason=reason,
erased_at=erased_at,
stores=stores,
)
if receipt.complete:
self.logger.info(
"Erased %r across %d store(s)%s",
entity_id,
len(stores),
f" ({reason})" if reason else "",
)
else:
self.logger.warning(
"Erasure of %r is incomplete; these stores may still hold it: %s",
entity_id,
", ".join(receipt.incomplete_stores),
)
return receipt
def erase_entities(
self,
entity_ids: Iterable[str],
reason: Optional[str] = None,
at: Optional[Union[str, int, float, datetime]] = None,
) -> List[ErasureReceipt]:
"""Erase several entities, returning one receipt per entity.
Each entity is erased independently, so one entity's failure does not
stop the rest. Receipts come back in the order the ids were given.
The timestamp is resolved once for the whole batch so that every
receipt and every graph tombstone record the same instant -- a batch
erasure under a single legal request must not produce tombstones with
diverging ``purged_at`` values.
"""
resolved_at = _normalize_timestamp(at)
return [
self.erase_entity(entity_id, reason=reason, at=resolved_at)
for entity_id in entity_ids
]
# Store legs
def _all_vector_ids(
self, entity_id: str, vector_ids: Optional[Sequence[str]]
) -> List[str]:
"""Caller-supplied vector ids plus the ids owned by memory items.
Best-effort by design: if memory cannot be enumerated here, the memory
leg makes the same call moments later and reports the failure, so the
receipt is still incomplete. Swallowing it there instead would be the
bug this method exists to fix.
Collects vector IDs from ALL memory items before deletion. Must call
find_by_entity with limit=None to get all items, since find_by_entity
doesn't support offset/cursor and we cannot delete while collecting.
"""
ids: List[str] = list(vector_ids) if vector_ids is not None else [entity_id]
if self.memory is None:
return ids
seen_vector_ids = set(ids)
try:
# Get ALL matching memory items in one call (limit=None).
# Pagination with deletion happens in _erase_memory(); here we must
# collect all vector IDs up front before any deletion occurs.
found = self.memory.find_by_entity(entity_id, limit=None)
for item in found:
memory_id = _memory_item_id(item)
if not memory_id:
continue
for vector_id in self.memory.vector_ids_for(memory_id):
if vector_id not in seen_vector_ids:
seen_vector_ids.add(vector_id)
ids.append(vector_id)
except Exception as exc:
self.logger.warning(
"Could not enumerate memory-owned vector ids for %r: %s; "
"the memory leg will report the same failure",
entity_id,
exc,
)
return ids
def _erase_vectors(
self, entity_id: str, vector_ids: Optional[Sequence[str]]
) -> Dict[str, Any]:
"""Remove entity-keyed embeddings from the bound vector store.
``vector_ids`` in the result is the number of ids the store accepted,
not the number of embeddings that existed: backends delete by id and
report success either way, with no portable way to ask what was
actually there. See :data:`STATUS_ERASED`.
"""
if self.vector_store is None:
return {"status": STATUS_NOT_CONFIGURED}
ids = list(vector_ids) if vector_ids is not None else [entity_id]
backend = _vector_backend_name(self.vector_store)
if not ids:
return {"status": STATUS_NOT_FOUND, "backend": backend}
method_name, target = _vector_delete_capability(self.vector_store)
if method_name is None:
# FAISS, Milvus and Weaviate expose no delete at all; FAISS in
# particular cannot remove from a flat index without a rebuild.
self.logger.warning(
"Vector backend %r exposes no delete; %d vector id(s) for %r "
"were not erased",
backend,
len(ids),
entity_id,
)
return {
"status": STATUS_UNSUPPORTED,
"backend": backend,
"vector_ids": len(ids),
"detail": (
"backend exposes no delete()/delete_vectors(); "
"removal requires an index rebuild or an out-of-band process"
),
}
try:
deleted = getattr(target, method_name)(ids)
except NotImplementedError as exc:
# The VectorStore facade declares delete_vectors() unconditionally
# and only fails on the call when its backend cannot delete.
self.logger.warning(
"Vector backend %r cannot delete %d id(s) for %r: %s",
backend,
len(ids),
entity_id,
exc,
)
return {
"status": STATUS_UNSUPPORTED,
"backend": backend,
"vector_ids": len(ids),
"detail": str(exc),
}
except Exception as exc:
self.logger.warning(
"Vector deletion failed for %r on backend %r: %s",
entity_id,
backend,
exc,
exc_info=True,
)
return {
"status": STATUS_FAILED,
"backend": backend,
"vector_ids": len(ids),
"detail": f"{type(exc).__name__}: {exc}",
}
accepted, detail = _interpret_delete_result(deleted)
result: Dict[str, Any] = {
"status": STATUS_ERASED if accepted else STATUS_FAILED,
"backend": backend,
"vector_ids": len(ids),
"via": method_name,
}
# Keep whatever the backend said. Qdrant returns {"status": ...} and
# Pinecone {"deleted": True}, and that detail is the only account of
# the delete anyone gets -- dropping it on the floor would leave the
# receipt less informative than the call it is attesting to.
if detail is not None:
result["backend_result"] = detail
if not accepted:
self.logger.warning(
"Vector backend %r reported no deletion for %r: %s",
backend,
entity_id,
detail,
)
result["detail"] = "store reported the ids were not deleted"
return result
def _erase_memory(self, entity_id: str) -> Dict[str, Any]:
"""Delete every memory item referencing the entity."""
if self.memory is None:
return {"status": STATUS_NOT_CONFIGURED}
deleted = 0
try:
# Sweep in pages until dry rather than passing one large limit:
# ``find_by_entity`` has historically defaulted to ``limit=10`` and
# truncated silently, and a single large number is only correct
# until someone exceeds it. Deleting as we go means the next page
# is the remainder.
while True:
found = self.memory.find_by_entity(entity_id, limit=_MEMORY_SWEEP_BATCH)
if not found:
break
memory_ids = [
memory_id
for memory_id in (_memory_item_id(item) for item in found)
if memory_id
]
if not memory_ids:
self.logger.warning(
"Memory returned %d item(s) for %r with no identifier; "
"cannot delete them",
len(found),
entity_id,
)
return {
"status": STATUS_FAILED,
"items": deleted,
"residual": len(found),
"detail": "memory items carry no 'memory_id'",
}
removed = self.memory.batch_delete(memory_ids)
deleted += removed
if removed == 0:
# No progress: another page would return the same items.
self.logger.warning(
"Memory sweep for %r stalled with %d item(s) remaining",
entity_id,
len(found),
)
return {
"status": STATUS_FAILED,
"items": deleted,
"residual": len(found),
"detail": "batch_delete removed nothing for a non-empty page",
}
if len(found) < _MEMORY_SWEEP_BATCH:
break
# Re-query once rather than trusting the loop's own bookkeeping;
# this is what keeps the leg's `failed` status honest.
residual = self.memory.find_by_entity(entity_id, limit=_MEMORY_SWEEP_BATCH)
except Exception as exc:
self.logger.warning(
"Memory erasure failed for %r after %d item(s): %s",
entity_id,
deleted,
exc,
exc_info=True,
)
return {
"status": STATUS_FAILED,
"items": deleted,
"detail": f"{type(exc).__name__}: {exc}",
}
if residual:
self.logger.warning(
"Memory still holds %d item(s) for %r after erasure",
len(residual),
entity_id,
)
return {
"status": STATUS_FAILED,
"items": deleted,
"residual": len(residual),
"detail": "items referencing the entity survived the sweep",
}
if deleted == 0:
return {"status": STATUS_NOT_FOUND, "items": 0}
return {"status": STATUS_ERASED, "items": deleted}
def _erase_graph(
self,
entity_id: str,
reason: Optional[str],
at: Optional[Union[str, int, float, datetime]],
) -> Dict[str, Any]:
"""Purge the node, and with it every edge that touches it."""
if self.graph is None:
return {"status": STATUS_NOT_CONFIGURED}
try:
# Counted before the purge because the edges are gone afterwards.
edge_count = _incident_edge_count(self.graph, entity_id)
purged = self.graph.purge_node(entity_id, reason=reason, at=at)
except Exception as exc:
self.logger.warning(
"Graph purge failed for %r: %s", entity_id, exc, exc_info=True
)
return {
"status": STATUS_FAILED,
"detail": f"{type(exc).__name__}: {exc}",
}
if not purged:
return {"status": STATUS_NOT_FOUND, "nodes": 0, "edges": 0}
return {"status": STATUS_ERASED, "nodes": 1, "edges": edge_count}
# Helpers
def _normalize_timestamp(at: Optional[Union[str, int, float, datetime]]) -> str:
"""Render ``at`` exactly as the graph tombstone will record it.
Reuses ``ContextGraph``'s own normalizer rather than formatting the value
here, so the receipt and the tombstone written by the same erasure cannot
disagree about when it happened -- an audit record that contradicts the
tombstone it attests to is worse than no record. Normalizing up front also
rejects an unparseable ``at`` before any store is touched, instead of half
way through the cascade.
``None`` resolves to now here rather than being passed along, so the
default path gets one timestamp for both records instead of two ``now()``
calls separated by the length of the cascade.
"""
return _normalize_temporal_input(
at if at is not None else datetime.now(timezone.utc)
)
def _memory_item_id(item: Any) -> Optional[str]:
"""Pull the identifier out of a memory dict as ``find_by_entity`` returns it."""
if not isinstance(item, dict):
return None
memory_id = item.get("memory_id") or item.get("id")
return str(memory_id) if memory_id else None
#: Dict keys a backend uses to report whether a delete succeeded, and the
#: values that mean it did not. Qdrant returns ``{"status": <UpdateStatus>}``
#: and Pinecone ``{"deleted": True}``; neither is a bool, so a bare
#: ``result is False`` check would call every dict a success.
_DELETE_FAILURE_MARKERS = {
"deleted": (False,),
"success": (False,),
"ok": (False,),
"acknowledged": (False,),
"status": ("failed", "error", "failure"),
}
def _interpret_delete_result(result: Any) -> Tuple[bool, Optional[str]]:
"""Decide whether a backend's delete return value reports success.
Returns ``(accepted, detail)``, where ``detail`` is a serializable
rendering of the backend's own response to keep in the receipt (``None``
when there was nothing worth recording).
``None`` counts as accepted: a delete implemented as a void method returns
it on success, and reporting ``failed`` there would be a false alarm --
the opposite of the honesty this module is for, in the other direction.
"""
if result is None:
return True, None
if isinstance(result, bool):
return result, None
if isinstance(result, dict):
rendered = {key: _stringify(value) for key, value in result.items()}
for key, failure_values in _DELETE_FAILURE_MARKERS.items():
if key in result and _is_failure_value(result[key], failure_values):
return False, rendered
return True, rendered
# Anything else (a count, a client response object) is taken at face value;
# there is no cross-backend contract to interpret it against.
return True, _stringify(result)
def _is_failure_value(value: Any, failure_values: Tuple[Any, ...]) -> bool:
"""True when a backend's marker value says the delete did not happen.
Bools are matched by identity so a ``0`` count is not read as ``False``.
String markers are matched as substrings of the rendered value, because a
backend may return an enum whose ``str()`` is ``"UpdateStatus.FAILED"``
rather than a bare ``"failed"``.
"""
for failure in failure_values:
if isinstance(failure, bool):
if value is failure:
return True
elif failure in str(value).lower():
return True
return False
def _stringify(value: Any) -> Any:
"""Render a backend payload value so the receipt stays serializable.
Qdrant's status is an enum, which would make ``to_dict()`` output
unserializable as the audit record it is meant to be.
"""
if isinstance(value, (str, int, float, bool)) or value is None:
return value
return str(value)
def _vector_delete_capability(store: Any) -> Tuple[Optional[str], Any]:
"""Find the delete method to call, and the object to call it on.
Returns ``(None, target)`` when no delete surface exists, which is the
``unsupported`` case.
The ``VectorStore`` facade declares ``delete_vectors()`` for every backend
and only raises ``NotImplementedError`` once called, so probing the facade
alone cannot tell a deletable backend from a delete-less one -- hence the
look at the backend it wraps. Probing rather than calling-and-catching also
keeps a missing method distinguishable from an ``AttributeError`` raised
*inside* a working one, which is exactly where guessing wrong would produce
a false clean bill of health.
"""
target = getattr(store, "_backend_store", None) or store
for name in ("delete_vectors", "delete"):
if callable(getattr(target, name, None)):
return name, target
return None, target
def _vector_backend_name(store: Any) -> str:
"""Best-effort backend label for the receipt."""
backend = getattr(store, "backend", None)
if isinstance(backend, str) and backend:
return backend
inner = getattr(store, "_backend_store", None)
return type(inner if inner is not None else store).__name__
def _incident_edge_count(graph: Any, node_id: str) -> int:
"""Count edges touching ``node_id`` through the graph's public API."""
find_edges = getattr(graph, "find_edges", None)
if not callable(find_edges):
return 0
return sum(
1
for edge in find_edges()
if edge.get("source") == node_id or edge.get("target") == node_id
)
+6 -17
View File
@@ -1,21 +1,10 @@
"""Semantica Evals — evaluation layer for decision intelligence outputs.
"""
Semantica Evals Module
Provides a small library of deterministic and model-backed evaluators plus a
runner for measuring decision records, audit trails, and reasoning output.
Coming Soon
"""
from . import decision_evaluators # noqa: F401 (registers decision_scores)
from . import evaluators # noqa: F401 (registers the generic evaluators)
from .registry import get_evaluator, list_evaluators
from .runner import evaluate
from .types import CaseResult, EvalMetric, EvalSummary
__version__ = "0.1.1"
__status__ = "coming_soon"
__all__ = []
__version__ = "0.1.0"
__all__ = [
"evaluate",
"get_evaluator",
"list_evaluators",
"CaseResult",
"EvalMetric",
"EvalSummary",
]
-90
View File
@@ -1,90 +0,0 @@
"""Decision-specialized evaluator.
``decision_scores`` validates a ``Decision`` (or dict) against field-level and
governance-level checks: expected outcome, confidence bounds, non-empty
required fields, provenance presence, and (when configured) policy compliance
via ``PolicyEngine.check_compliance``.
"""
from typing import Any, Dict, Optional
from .registry import register
from .types import EvalMetric
def _coerce_decision(actual: Any):
"""Return a Decision or None; never raise for dict inputs."""
from semantica.context.decision_models import Decision
if isinstance(actual, Decision):
return actual
if isinstance(actual, dict):
try:
return Decision(**actual)
except (TypeError, ValueError, KeyError):
return None
return None
@register("decision_scores")
def decision_scores(actual, expected=None, config=None, **kwargs):
"""Composite evaluator over a Decision; see module docstring for sub-checks."""
cfg = config or {}
decision = _coerce_decision(actual)
if decision is None:
return EvalMetric(0.0, False, {"error": "input is not a valid Decision or dict"})
checks: Dict[str, bool] = {}
reasons: Dict[str, str] = {}
expected_outcome = cfg.get("expected_outcome", expected)
if expected_outcome is not None:
checks["decision_outcome"] = decision.outcome == expected_outcome
if not checks["decision_outcome"]:
reasons["decision_outcome"] = f"expected {expected_outcome!r}, got {decision.outcome!r}"
lo = cfg.get("min_confidence", 0.0)
hi = cfg.get("max_confidence", 1.0)
checks["decision_confidence"] = lo <= decision.confidence <= hi
if not checks["decision_confidence"]:
reasons["decision_confidence"] = f"{decision.confidence} not in [{lo}, {hi}]"
for field in ("decision_maker", "reasoning", "scenario"):
value = getattr(decision, field, None)
checks[field] = isinstance(value, str) and bool(value.strip())
if not checks[field]:
reasons[field] = f"field {field!r} is empty"
metadata = decision.metadata if isinstance(decision.metadata, dict) else {}
prov = metadata.get(cfg.get("provenance_key", "provenance"))
checks["provenance"] = bool(prov)
if not checks["provenance"]:
reasons["provenance"] = "no provenance record found in metadata"
policy_engine = cfg.get("policy_engine")
policy_id = cfg.get("policy_id")
if policy_engine is not None and policy_id is not None:
try:
compliant = bool(policy_engine.check_compliance(decision, policy_id))
checks["policy"] = compliant == cfg.get("expected_policy_compliant", True)
if not checks["policy"]:
reasons["policy"] = f"compliance={compliant}"
except Exception as exc: # noqa: BLE001
checks["policy"] = False
reasons["policy"] = str(exc)
if cfg.get("causal_chain_exists"):
raise NotImplementedError(
"decision_scores causal_chain_exists is an interface slot reserved for V2"
)
passed_count = sum(checks.values())
total = len(checks)
passed = total > 0 and passed_count == total
meta = dict(checks)
meta["reasons"] = reasons
return EvalMetric(
score=passed_count / total if total else 0.0,
passed=passed,
meta=meta,
)
-181
View File
@@ -1,181 +0,0 @@
"""Generic (non-decision) evaluators for the evals module.
Each evaluator takes ``(actual, expected, config=None, **kwargs)`` and returns
an ``EvalMetric``. Config uses ``min``/``max`` bounds where relevant.
"""
from datetime import datetime
from typing import Any, Dict, List, Optional
from .registry import register
from .types import EvalMetric
def _default_config(config):
return config or {}
@register("exact_match")
def exact_match(actual, expected, config=None, **kwargs):
"""Score 1.0 if ``actual`` equals ``expected`` (scalar or list)."""
matched = actual == expected
return EvalMetric(
score=1.0 if matched else 0.0,
passed=matched,
meta={} if matched else {"reason": f"expected {expected!r}, got {actual!r}"},
)
@register("regex_match")
def regex_match(actual, expected, config=None, **kwargs):
"""Score 1.0 if string ``actual`` matches regex ``expected``."""
import re
try:
matched = re.search(expected, actual) is not None
return EvalMetric(
score=1.0 if matched else 0.0,
passed=matched,
meta={} if matched else {"reason": f"'{actual}' does not match {expected}"},
)
except re.error as exc:
return EvalMetric(0.0, False, {"error": str(exc)})
@register("numeric_range")
def numeric_range(actual, expected=None, config=None, **kwargs):
"""Score 1.0 if number ``actual`` is within inclusive ``[min, max]``."""
cfg = _default_config(config)
lo, hi = cfg.get("min"), cfg.get("max")
passed = lo is not None and hi is not None and lo <= actual <= hi
return EvalMetric(
score=1.0 if passed else 0.0,
passed=passed,
meta={} if passed else {"reason": f"{actual} not in [{lo}, {hi}]"},
)
@register("temporal_range")
def temporal_range(actual, expected=None, config=None, **kwargs):
"""Score 1.0 if datetime ``actual`` is within inclusive ISO-datetime window."""
cfg = _default_config(config)
try:
stamp = datetime.fromisoformat(actual)
lo = datetime.fromisoformat(cfg["min"])
hi = datetime.fromisoformat(cfg["max"])
passed = lo <= stamp <= hi
return EvalMetric(
score=1.0 if passed else 0.0,
passed=passed,
meta={} if passed else {"reason": f"{actual} not in [{cfg['min']}, {cfg['max']}]"},
)
except (KeyError, TypeError, ValueError) as exc:
return EvalMetric(0.0, False, {"error": str(exc)})
@register("length_range")
def length_range(actual, expected=None, config=None, **kwargs):
"""Score 1.0 if length of ``actual`` is within inclusive ``[min, max]``."""
cfg = _default_config(config)
size = len(actual)
lo = cfg.get("min", 0)
hi = cfg.get("max")
passed = hi is not None and lo <= size <= hi
return EvalMetric(
score=1.0 if passed else 0.0,
passed=passed,
meta={} if passed else {"reason": f"length {size} not in [{lo}, {hi}]"},
)
@register("keyword_check")
def keyword_check(actual, expected=None, config=None, **kwargs):
"""Score 1.0 if all required terms appear in ``actual`` (word-boundary matching)."""
cfg = _default_config(config)
required = cfg.get("required") or (expected or [])
import re
tokens = set(re.findall(r"\w+", str(actual).lower()))
missing = [term for term in required if str(term).lower() not in tokens]
passed = not missing
return EvalMetric(
score=1.0 if passed else 0.0,
passed=passed,
meta={} if passed else {"missing": missing},
)
def _levenshtein(a: str, b: str) -> int:
"""Classic Levenshtein edit distance."""
if a == b:
return 0
if not a:
return len(b)
if not b:
return len(a)
prev = list(range(len(b) + 1))
for i, ca in enumerate(a, 1):
cur = [i]
for j, cb in enumerate(b, 1):
cur.append(min(prev[j] + 1, cur[j - 1] + 1, prev[j - 1] + (ca != cb)))
prev = cur
return prev[-1]
@register("levenshtein")
def levenshtein(actual, expected, config=None, **kwargs):
"""Score normalized similarity (1 - distance/max_len) vs ``threshold`` (default 0.8)."""
cfg = _default_config(config)
threshold = cfg.get("threshold", 0.8)
a, b = str(actual), str(expected)
max_len = max(len(a), len(b))
similarity = 1.0 if max_len == 0 else 1.0 - _levenshtein(a, b) / max_len
passed = similarity >= threshold
return EvalMetric(
score=similarity,
passed=passed,
meta={"similarity": similarity},
)
def _tokenize(text: str) -> List[str]:
import re
return re.findall(r"\w+", str(text).lower())
@register("rouge")
def rouge(actual, expected, config=None, **kwargs):
"""ROUGE-1 precision/recall/F1 over tokens; pass on F1 >= ``threshold`` (default 0.0)."""
cfg = _default_config(config)
threshold = cfg.get("threshold", 0.0)
hyp, ref = _tokenize(actual), _tokenize(expected)
from collections import Counter
hyp_c, ref_c = Counter(hyp), Counter(ref)
overlap = sum((hyp_c & ref_c).values())
precision = overlap / len(hyp) if hyp else 0.0
recall = overlap / len(ref) if ref else 0.0
f1 = 0.0 if (precision + recall) == 0 else 2 * precision * recall / (precision + recall)
passed = f1 > 0 and f1 >= threshold
return EvalMetric(
score=f1,
passed=passed,
meta={"precision": precision, "recall": recall, "f1": f1},
)
@register("llm_as_judge")
def llm_as_judge(actual, expected, config=None, **kwargs):
"""Score 1.0 when a caller-supplied ``judge_fn(actual, expected) -> bool`` passes.
The judge resolver stays lazy: no LLM backend is imported unless the caller
provides one in config.
"""
cfg = _default_config(config)
judge_fn = cfg.get("judge_fn")
if judge_fn is None:
return EvalMetric(
0.0, False, {"error": "config['judge_fn'] required (callable(actual, expected) -> bool)"}
)
try:
verdict = bool(judge_fn(actual, expected))
return EvalMetric(score=1.0 if verdict else 0.0, passed=verdict)
except Exception as exc: # noqa: BLE001
return EvalMetric(0.0, False, {"error": str(exc)})
-34
View File
@@ -1,34 +0,0 @@
"""Evaluator registry for the evals module.
Evaluators are plain functions ``fn(actual, expected, config=None, **kwargs)
-> EvalMetric`` registered under a stable string name so the runner and users
can select them by name without importing individual modules.
"""
from typing import Callable, Dict, List
from .types import EvalMetric
EVALUATORS: Dict[str, Callable] = {}
def register(name: str) -> Callable:
"""Decorator registering an evaluator function under ``name``."""
def _register(fn: Callable) -> Callable:
if name in EVALUATORS:
raise ValueError(f"evaluator already registered: {name}")
EVALUATORS[name] = fn
return fn
return _register
def list_evaluators() -> List[str]:
"""Return sorted names of all registered evaluators."""
return sorted(EVALUATORS)
def get_evaluator(name: str) -> Callable:
"""Look up an evaluator by name, raising ValueError with a hint otherwise."""
if name not in EVALUATORS:
raise ValueError(f"unknown evaluator '{name}'. Available: {list_evaluators()}")
return EVALUATORS[name]
-211
View File
@@ -1,211 +0,0 @@
"""Evaluation runner: orchestrates evaluators over a list of cases."""
import math
from typing import Any, Callable, Dict, List, Optional, Tuple, Union
from .registry import get_evaluator
from .types import CaseResult, EvalMetric, EvalSummary
Case = Union[Dict[str, Any], Tuple[Any, Any]]
def _coerce_threshold(name, threshold):
"""Convert ``threshold`` to a finite float, raising ``ValueError`` otherwise.
Accepts any value that ``float()`` accepts (int, float, bool, numeric
strings) as long as the result is finite. Raises ``ValueError`` never
``TypeError`` for non-convertible types, NaN, and infinity so that
all invalid objective config produces the same exception type.
"""
try:
value = float(threshold)
except (TypeError, ValueError) as exc:
raise ValueError(
f"objective for '{name}': 'threshold' must be a finite number "
f"(got {threshold!r})"
) from exc
if not math.isfinite(value):
raise ValueError(
f"objective for '{name}': 'threshold' must be a finite number "
f"(got {threshold!r})"
)
return value
def _parse_objective(name, eval_config):
"""Return the validated objective dict, or None when not configured.
Raises ValueError for invalid configurations (programmer error).
"""
objective = (eval_config or {}).get("objective")
if objective is None:
return None
if not isinstance(objective, dict):
raise ValueError(
f"objective for '{name}': expected a dict, got {type(objective).__name__}"
)
direction = objective.get("direction")
threshold = objective.get("threshold")
expect = objective.get("expect")
if expect is not None:
if not isinstance(expect, bool):
raise ValueError(
f"objective for '{name}': 'expect' must be a bool (got {expect!r})"
)
if direction is not None or threshold is not None:
raise ValueError(
f"objective for '{name}': 'expect' cannot be combined with "
"'direction' or 'threshold'"
)
return {"expect": expect}
if direction == "minimize":
if threshold is None:
raise ValueError(
f"objective for '{name}': 'minimize' requires a 'threshold'"
)
return {"direction": "minimize", "threshold": _coerce_threshold(name, threshold)}
if direction == "maximize":
if threshold is None:
# no bar to re-decide against; treat as absent (evaluator default stands)
return None
return {"direction": "maximize", "threshold": _coerce_threshold(name, threshold)}
raise ValueError(
f"objective for '{name}': 'direction' must be 'maximize' or 'minimize' "
f"(got {direction!r})"
)
def _apply_objective(metric, objective):
"""Return the objective-adjusted pass verdict for a non-error metric."""
if "expect" in objective:
return bool(metric.score) == objective["expect"]
if objective["direction"] == "minimize":
return metric.score <= objective["threshold"]
return metric.score >= objective["threshold"]
def _extract(case: Case, target_fn: Optional[Callable]):
"""Return (case_id, expected, actual, config, per_case_target_fn)."""
if isinstance(case, tuple):
expected, actual = case[0], (case[1] if len(case) > 1 else None)
return str(id(case)), expected, actual, {}, None
case_id = case.get("id") or f"case-{id(case)}"
expected = case.get("expected")
actual = case.get("actual")
config = case.get("config") or {}
per_fn = case.get("target_fn")
return case_id, expected, actual, config, per_fn
def _merge_config(default_config: Dict[str, Any], case_config: Dict[str, Any]) -> Dict[str, Any]:
"""Deep-merge per-case config over the global config (two levels deep).
Level 1 (top-level keys, e.g. evaluator names): merged key-by-key so a
per-case override of one evaluator's settings does not erase the whole
global evaluator entry.
Level 2 (evaluator config keys, e.g. ``"objective"``): also merged
key-by-key so a per-case override that specifies only some objective fields
(e.g. just ``"threshold"``) inherits the rest from the global objective
(e.g. ``"direction"``). Per-case values always take precedence.
Depth-3+ values are replaced wholesale, consistent with the previous
single-level behaviour (no evaluator config currently nests beyond two
levels). Neither the caller's global config nor the case config is
mutated.
"""
merged = dict(default_config)
for key, value in (case_config or {}).items():
if isinstance(value, dict) and isinstance(merged.get(key), dict):
# Merge level-1 dict (evaluator config) key-by-key.
current = dict(merged[key])
for k, v in value.items():
if isinstance(v, dict) and isinstance(current.get(k), dict):
# Merge level-2 dict (e.g. objective sub-dict) key-by-key.
inner = dict(current[k])
inner.update(v)
current[k] = inner
else:
current[k] = v
merged[key] = current
else:
merged[key] = value
return merged
def evaluate(
cases: List[Case],
evaluators: List[str],
config: Optional[Dict[str, Any]] = None,
target_fn: Optional[Callable] = None,
) -> EvalSummary:
"""Run named evaluators over each case and aggregate metrics.
A per-case or top-level ``target_fn`` produces ``actual`` when the case
does not already carry one. Evaluator failures become ``error`` results.
"""
default_config = config or {}
case_results: List[CaseResult] = []
# Validate objective config for every case up front so an invalid objective
# rejects the run before any target_fn or evaluator executes (fail-fast),
# regardless of which case carries it.
pre_resolved = []
for case in cases:
_, _, _, case_config, _ = _extract(case, target_fn)
merged = _merge_config(default_config, case_config)
pre_resolved.append(
{
name: _parse_objective(name, merged.get(name) or {})
for name in evaluators
}
)
for case, objective_by_name in zip(cases, pre_resolved):
case_id, expected, actual, case_config, per_fn = _extract(case, target_fn)
merged = _merge_config(default_config, case_config)
if expected is None:
expected = merged.get("expected")
resolver = per_fn or target_fn
if actual is None and resolver is not None:
try:
actual = resolver(case)
except Exception as exc: # noqa: BLE001
case_results.append(
CaseResult(case_id, "error", {}, {"target_fn": str(exc)})
)
continue
metrics: Dict[str, EvalMetric] = {}
details: Dict[str, Any] = {}
failed, errored = False, False
for name in evaluators:
eval_config = merged.get(name) or {}
try:
metric = get_evaluator(name)(actual, expected, config=eval_config)
objective = objective_by_name.get(name)
if objective is not None and "error" not in metric.meta:
metric = EvalMetric(metric.score, _apply_objective(metric, objective), metric.meta)
metrics[name] = metric
if "error" in metric.meta:
errored = True
details[name] = metric.meta
elif not metric.passed:
failed = True
details[name] = metric.meta
except Exception as exc: # noqa: BLE001
errored = True
metrics[name] = EvalMetric(0.0, False, {"error": str(exc)})
details[name] = {"error": str(exc)}
status = "error" if errored else ("fail" if failed else "pass")
case_results.append(CaseResult(case_id, status, metrics, details))
total = len(case_results)
passed = sum(1 for c in case_results if c.status == "pass")
failed = sum(1 for c in case_results if c.status == "fail")
errors = sum(1 for c in case_results if c.status == "error")
pass_rate = (passed / total) if total else 1.0
return EvalSummary(
total, passed, failed, errors, pass_rate,
cases=case_results,
)
-37
View File
@@ -1,37 +0,0 @@
"""Evals result data models.
Defines the metric and result shapes produced by the evals module.
"""
from dataclasses import dataclass, field
from typing import Any, Dict, List, NamedTuple
@dataclass(frozen=True)
class EvalMetric:
"""One evaluator's numeric score plus pass/fail verdict."""
score: float
passed: bool
meta: Dict[str, Any] = field(default_factory=dict)
class CaseResult(NamedTuple):
"""Evaluation output for a single case."""
case_id: str
status: str
metrics: Dict[str, EvalMetric]
details: Dict[str, Any]
@dataclass
class EvalSummary:
"""Aggregate evaluation output across cases."""
total: int
passed: int
failed: int
errors: int
pass_rate: float
cases: List[CaseResult] = field(default_factory=list)
-176
View File
@@ -1,176 +0,0 @@
# Semantica Evals — Usage
The evals module measures decision intelligence outputs: decision records,
audit trails, and reasoning output — with deterministic and model-backed
evaluators plus a small runner.
## Import
```python
import semantica.evals as evals # through the root lazy proxy
from semantica.evals import evaluate, list_evaluators
```
## Discover evaluators
```python
>>> evals.list_evaluators()
['decision_scores', 'exact_match', 'keyword_check', 'length_range',
'levenshtein', 'llm_as_judge', 'numeric_range', 'regex_match', 'rouge',
'temporal_range']
```
`list_evaluators` returns every name registered by importing the package —
the import wiring runs each evaluator module's `register()` side effects.
## Run the runner over decision records
`evaluate(cases, evaluators, config=None)` accepts a list of cases; each case is
a dict with `expected`, `actual`, optional `config`, and optional `id`. The
`actual` can be a finished `Decision` object or its dict form.
```python
from datetime import datetime
from semantica.context.decision_models import Decision
from semantica.evals import evaluate
decision = Decision(
decision_id="d-1",
category="loan",
scenario="loan-request",
reasoning="vetted by policy",
outcome="approve",
confidence=0.87,
timestamp=datetime.now(),
decision_maker="approver-a",
metadata={"provenance": "workflow:loan/v3"},
)
cases = [
{
"id": "loan-001",
"actual": decision,
"config": {
"decision_scores": {
"expected_outcome": "approve",
"min_confidence": 0.7,
}
},
},
{
"id": "loan-002",
"actual": {
"decision_id": "d-2",
"category": "loan",
"scenario": "loan-request",
"reasoning": "auto",
"outcome": "reject",
"confidence": 0.9,
"timestamp": datetime.now().isoformat(),
"decision_maker": "system",
"metadata": {},
},
"config": {
"decision_scores": {
"expected_outcome": "approve",
"min_confidence": 0.7,
}
},
},
]
summary = evaluate(cases, ["decision_scores"])
```
`evaluate` also runs high-level names like `exact_match`, `keyword_check`, or
`llm_as_judge`; per-case or top-level `config` may carry per-evaluator settings
(e.g. `config={"exact_match": {...}}`).
## Set per-evaluator objectives
By default each evaluator decides its own pass/fail. To override that
verdict at the run level, configure an **objective** per evaluator name:
```python
from semantica.evals import evaluate
# Require a minimum similarity (levenshtein's default bar is >= 0.8; here we set 0.7):
evaluate(
[("apple", "aple")],
evaluators=["levenshtein"],
config={"levenshtein": {"objective": {"direction": "maximize", "threshold": 0.7}}},
)
# Lower is better — override the direction:
evaluate(
[("night", "nacht")],
evaluators=["levenshtein"],
config={"levenshtein": {"objective": {"direction": "minimize", "threshold": 0.7}}},
)
# Boolean expectation — the metric matches (score 1), but we expect it not to:
evaluate(
[("ok", "ok")],
evaluators=["exact_match"],
config={"exact_match": {"objective": {"expect": False}}},
)
```
Rules:
- `maximize` + `threshold`: pass iff `score >= threshold`. `maximize` without
a threshold is a no-op (the evaluator's own verdict stands).
- `minimize` + `threshold`: pass iff `score <= threshold`. `minimize`
**requires** a threshold — omitting it or setting it to `None` raises
`ValueError`.
- `expect` (`true`/`false`): pass iff `bool(score)` matches; cannot be
combined with `direction`/`threshold`. `expect` must be a real boolean
(a string like `"false"` is rejected).
- A metric whose `meta` contains `"error"` is always an error, never affected
by an objective.
- Invalid objective config (non-dict objective, bad `direction`, non-bool
`expect`, missing `minimize` threshold) raises `ValueError` before any
evaluator runs.
## Interpret the summary
```python
>>> summary.total, summary.passed, summary.failed, summary.errors
(2, 1, 1, 0)
>>> summary.pass_rate
0.5
>>> for case in summary.cases:
... print(case.case_id, case.status)
... for name, metric in case.metrics.items():
... print(" ", name, metric.score, metric.passed)
... print(" ", metric.meta.get("reasons"))
loan-001 pass
decision_scores 1.0 True
{}
loan-002 fail
decision_scores 0.667 False
{'decision_outcome': "expected 'approve', got 'reject'",
'provenance': 'no provenance record found in metadata'}
```
`EvalSummary` fields:
- `total` / `passed` / `failed` / `errors` — case counts by status.
- `pass_rate``passed / total` (1.0 on an empty case list).
- `cases` — one `CaseResult` per input case: `case_id`, `status`
(`pass` | `fail` | `error`), `metrics` (name → `EvalMetric` with `score`,
`passed`, `meta`), and `details`.
Evaluator failures do not crash the run; they surface as `status="error"` on
the affected case with the exception text captured in the metric meta.
## Notes
- **`llm_as_judge` needs `config["judge_fn"]`**: a callable
`judge_fn(actual, expected) -> bool` supplied by the caller. Without it the
evaluator fails with `config['judge_fn'] required`.
- **`decision_scores` governance checks are opt-in**: policy compliance is only
evaluated when both `config["policy_engine"]` and `config["policy_id"]` are
provided; otherwise those checks are skipped. The reserved
`causal_chain_exists` slot is not yet implemented.
+1 -99
View File
@@ -233,31 +233,6 @@ async def import_file(
)
#: Aliases kept consistent with `mcp/tools/export.py::_FORMAT_ALIASES` and
#: `RDFExporter._format_aliases` to ensure the two surfaces agree on format names.
#: Maps user-provided format strings to RDFExporter's canonical format names.
_RDF_FORMATS: dict[str, str] = {
"ttl": "turtle",
"turtle": "turtle",
"nt": "ntriples", # RDFExporter canonical is "ntriples", not "nt"
"ntriples": "ntriples",
"n-triples": "ntriples",
"xml": "rdfxml", # RDFExporter canonical is "rdfxml", not "xml"
"rdfxml": "rdfxml",
"rdf-xml": "rdfxml",
"json-ld": "jsonld", # RDFExporter canonical is "jsonld", not "json-ld"
"jsonld": "jsonld",
}
#: Media type and file extension per RDFExporter canonical format name.
_RDF_MEDIA_TYPES: dict[str, tuple[str, str]] = {
"turtle": ("text/turtle", "ttl"),
"ntriples": ("application/n-triples", "nt"),
"rdfxml": ("application/rdf+xml", "rdf"),
"jsonld": ("application/ld+json", "jsonld"),
}
@router.post("/api/export")
async def export_graph(
body: ExportRequest,
@@ -292,81 +267,8 @@ async def export_graph(
content = output.getvalue()
media_type = "text/csv"
extension = "csv"
elif fmt in _RDF_FORMATS:
# Reuses `semantica.export`, the same exporters the MCP `export_graph` tool calls.
# Before this, the Explorer answered 422 for every RDF format while the MCP surface
# offered them, so a graph could be loaded as JSON-LD and never exported back — the
# round trip had to leave the product. See #1131.
try:
from semantica.export import RDFExporter
from semantica.utils.exceptions import ValidationError
except ImportError as exc: # pragma: no cover - optional dependency
raise HTTPException(
status_code=503,
detail=f"RDF export unavailable: {exc}",
) from exc
try:
content = RDFExporter().export_to_rdf(graph_dict, format=_RDF_FORMATS[fmt])
except ValidationError as exc:
# Data validation or serialization failed
raise HTTPException(
status_code=422,
detail=f"RDF export failed: {exc}",
) from exc
except Exception as exc:
# Unexpected error during export
logger.exception("RDF export failed unexpectedly")
raise HTTPException(
status_code=500,
detail=f"RDF export error: {exc}",
) from exc
media_type, extension = _RDF_MEDIA_TYPES[_RDF_FORMATS[fmt]]
elif fmt == "graphml":
# GraphML support using GraphExporter (not GraphMLExporter which doesn't exist)
try:
from semantica.export import GraphExporter
from semantica.utils.exceptions import ValidationError
except ImportError as exc: # pragma: no cover - optional dependency
raise HTTPException(
status_code=503,
detail=f"GraphML export unavailable: {exc}",
) from exc
try:
# GraphExporter.export() writes to file, but we need string content for HTTP response.
# Use a temporary file that is automatically cleaned up.
import tempfile
from pathlib import Path
# Create temp file in a secure directory with automatic cleanup on exception
with tempfile.TemporaryDirectory() as tmpdir:
tmp_path = Path(tmpdir) / "export.graphml"
exporter = GraphExporter(format="graphml")
exporter.export(graph_dict, file_path=tmp_path)
content = tmp_path.read_text(encoding='utf-8')
except ValidationError as exc:
raise HTTPException(
status_code=422,
detail=f"GraphML export failed: {exc}",
) from exc
except Exception as exc:
logger.exception("GraphML export failed unexpectedly")
raise HTTPException(
status_code=500,
detail=f"GraphML export error: {exc}",
) from exc
media_type, extension = "application/xml", "graphml"
else:
raise HTTPException(
status_code=422,
detail=(
f"Unsupported export format '{fmt}'. "
f"Supported: {', '.join(sorted({'json', 'csv', 'graphml'} | set(_RDF_FORMATS)))}"
),
)
raise HTTPException(status_code=422, detail=f"Unsupported export format '{fmt}'")
return Response(
content=content,

Some files were not shown because too many files have changed in this diff Show More