Compare commits

...
Author SHA1 Message Date
Zohaib Hassnain 656120514d docs: update stale latest version claims 2026-09-03 03:49:41 +05:00
Zohaib Hassnain 3c68cd12ad docs(mcp): correct tool count (#1399) 2026-09-03 03:40:52 +05:00
Sameer Kadam 798a7455e4 fix(mcp): complete persistence and setup fixes (#1394)
MCP's stdio transport uses stdout for JSON-RPC framing, so anything else written there corrupts every response after it. The original #1134 bug was progress-tracker output landing on stdout during tool calls that construct a `ContextGraph`, which is exactly what happens on any request that triggers reasoning or extraction. This PR closes out the remaining pieces of that fix: loading now goes through `load_from_file()` instead of the older `load()` path on the root graph, and mutations, `record_decision`, `add_entity`, `add_relationship`, now persist back to `SEMANTICA_KG_PATH` when it's configured, in both MCP server implementations (the root `mcp/` package and the packaged `semantica.mcp_server`), not just one.

Four things came out of review on top of that.

The stdio regression test originally exercised `get_graph_summary`, which doesn't touch the progress tracker at all, so it couldn't have caught the original bug. Swapped it for `run_reasoning`: `Reasoner.infer_with_results()` calls `progress_tracker.start_tracking()` directly, the exact call site that corrupted stdout before, so this is the minimal path that actually proves the fix. The test now spawns a real `python -m mcp` subprocess, sends it a `tools/call` for `run_reasoning`, and asserts every single line on stdout parses as JSON.

Loading a corrupt or unreadable `SEMANTICA_KG_PATH` used to fail silently and fall through to an empty graph, which meant the next mutation would happily save that empty graph over the original file. Both implementations now track whether the initial load actually succeeded. If it didn't, every mutation handler refuses to save and returns an error instead, so a broken file on disk stays broken rather than getting silently replaced with nothing. An empty file is treated differently: that's a fresh destination, not a corrupt one, and starts a normal empty graph without tripping the guard.

`save_to_file` used to `open(path, 'w')` and `json.dump` directly into the destination, so a crash or disk-full error mid-write could leave a truncated file as the only copy of the graph. It now writes to a temp file in the same directory, flushes, fsyncs, and only then `os.replace`s the destination, so the destination is always either the old contents or the new contents, never a partial write. The temp file gets cleaned up if anything fails before the replace.

And since a mutation is applied to the in-memory graph before the save happens, a save failure used to leave the in-memory graph ahead of what's on disk, an entity or decision the client thinks succeeded but that never made it to the file. `record_decision`, `add_entity`, and `add_relationship` all roll back the in-memory mutation now if `save_to_file` raises, so the client-visible state and the persisted state never diverge: either both hold the change or neither does.

104 tests passing across the MCP, persistence, and progress-tracking suites.
2026-09-03 03:35:20 +05:00
Mohd Kaif bd584b7402 Update features list in README
Removed 'Self-Hostable' and 'Auditable' from the features list.
2026-09-02 22:21:41 +05:30
Mohd Kaif a4500f5b20 Merge pull request #1396 from semantica-agi/readme-enterprise-connectors-update
docs: tighten README audience list, add SAP connector mentions, log Salesforce ingestor
2026-09-02 21:50:51 +05:30
KaifAhmad1 bdd12e8ac6 fix: correct JWT auth requirements in changelog, add missing SAP install extra
- CHANGELOG: JWT Bearer requires username too, not just consumer_key + private key
- README: add pip install semantica[ingest-sap] to the install-extras list, which was missing despite SAP appearing in the supported-sources lists

Addresses Qodo review feedback on #1396.
2026-09-02 21:45:24 +05:30
KaifAhmad1 48204d4e02 docs: tighten README audience list, add SAP to connector mentions, log unreleased Salesforce ingestor
- Trim "Who it's for" bullets in README for concision
- Propagate SAP OData connector mentions across README's integration lists (was only in the What's New section)
- Add missing CHANGELOG entry for the unreleased Salesforce ingestor (#1240)
- Remove sample `semantica doctor` output lines from the quickstart snippet
2026-09-02 21:23:29 +05:30
Zohaib Hassnain 1bc873cbbd Merge pull request #1328 from semantica-agi/feat/pinecone-iter-all
feat(vector_store): add pinecone iter_all
2026-09-02 20:07:54 +05:30
KaifAhmad1 4dd88375e1 fix(vector_store): don't short-circuit pinecone iter_all() on an empty page
if not vector_ids: return fired before the continuation token was ever
checked. Pinecone's actual pagination contract is that a scan is only
exhausted when the response carries no pagination token -- a page can
legitimately list zero ids while pagination.next is still set (sparse
or filtered namespaces, eventual-consistency windows on serverless
indexes). This was flagged in review but the fix commit that followed
only addressed the separate repeated-token stall case, not this one.

Reproduced concretely against the unfixed code: a page with data,
followed by an empty page with a live token, followed by a page with
more data -- the last page was silently dropped with no error raised,
exactly the #1083 failure mode (store migrate reporting success after
copying only part of a collection).

Now the empty-page case skips the pointless fetch() call but still
falls through to the same next_token check every other path already
goes through, so a live token continues the scan and only a genuinely
absent token (or one that's stopped advancing) ends it.

Added test_continues_past_an_empty_page_with_a_live_token, the
"empty page + non-None next token" case the original review asked for
and that wasn't otherwise covered.
2026-09-02 19:54:24 +05:30
Mohd Kaif fc899c6966 Merge pull request #1326 from semantica-agi/feat/milvus-iter-all
feat(vector_store): add milvus iter_all
2026-09-02 19:37:54 +05:30
KaifAhmad1 98bd632585 Merge remote-tracking branch 'origin/main' into feat/milvus-iter-all 2026-09-02 19:20:55 +05:30
KevinandSameer Kadam 30a91a3a78 feat(evals): add per-metric objective support to runner (closes #1091) (#1092)
* chore: ignore .worktrees directory

* feat(evals): add eval metric and result models

* feat(evals): add evaluator registry

* feat(evals): add exact/regex/range/length evaluators

* feat(evals): add keyword/levenshtein/rouge/llm-as-judge evaluators

* feat(evals): add decision_scores composite evaluator

* feat(evals): add evaluation runner

* feat(evals): expose public API and module proxy

* fix(evals): resolve __all__ names and repair usage example

* docs(evals): add usage docs and changelog entry

* style(evals): tidy evaluator metadata and wiring comments

* fix(evals): honor expected arg and classify error metrics

* fix(evals): export get_evaluator and fix shared meta default

* fix(evals): guard provenance check against non-dict metadata

* docs: add objective layer design spec for semantica.evals

* docs: refine objective spec for consistency with AIP Evals semantics

* docs: add implementation plan for evals objective layer

* docs: fix plan tests to use module-level pytest import

* feat(evals): add per-metric objective support to runner

* docs(evals): document per-metric objectives

* docs(evals): fix minimize example threshold to demonstrate pass

* fix(evals): validate objective config shape strictly

* docs(evals): clarify objective examples and Boolean semantics

* fix(evals): honor direction-only minimize, fail fast on objectives, deep-merge case config

- minimize without threshold is now a no-op, matching maximize (issue #1091
  requires thresholds to be optional for both directions)
- objective config is parsed for every case before any target_fn/evaluator
  runs, so an invalid per-case objective rejects the run up front
- per-case evaluator config deep-merges over the global config so a case
  that overrides one setting keeps the run-level objective
- regression tests for all three, plus updated docs/CHANGELOG

Addresses 3 of 4 Qodo findings on #1092 (the 4th, 'result models defined
twice', is a false positive: types live in types.py)

* fix: finalize eval objectives review

---------

Co-authored-by: Sameer Kadam <sskadam6305@gmail.com>
2026-09-02 18:47:28 +05:30
Mohd Kaif 23126106a3 Merge pull request #1317 from semantica-agi/feat/weaviate-iter-all
feat(vector_store): add weaviate iter_all
2026-09-02 18:39:57 +05:30
Mohd Kaif 4b001b4c9d Merge branch 'main' into feat/weaviate-iter-all 2026-09-02 18:29:14 +05:30
KaifAhmad1 6c9eb2296d fix(vector_store): don't treat an empty weaviate page as end of scan
iter_all() unconditionally returned on any empty fetch_objects() page,
regardless of pagination mode. That's safe for offset/single_page (an
empty page there is a direct, unambiguous statement about live rows),
but not for cursor mode: `after` has no server-issued continuation
value of its own, it's derived client-side from the last object's uuid,
so an empty page gives nothing to advance it with. If Weaviate's cursor
walks internal storage position rather than strict uuid order, a batch
can in principle land entirely on a gap (e.g. tombstoned objects) with
live data past it -- the same risk already confirmed and fixed for
Qdrant's scroll cursor in #1316. Reproduced concretely against the
pre-fix code: a full page followed by an empty page followed by a page
with real data silently dropped that last page with no error raised.

iter_all() now falls back to offset pagination once when a cursor-mode
page comes back empty, rather than assuming that's the end. Offset
addresses live rows directly by position and has no equivalent gap, so
an empty page there (or in single_page mode) is trustworthy and still
ends the scan immediately.

Also updates test_iter_all_empty_collection_yields_nothing and
test_iter_all_requests_vectors, which needed a second empty page now
that a genuinely empty collection takes two calls (cursor, then the
confirming offset check) to report as such.
2026-09-02 17:43:52 +05:30
KaifAhmad1 1ad17beaf6 fix(vector_store): sync qdrant iter_all() with #1316's stall-guard fix
This branch was forked from an earlier commit of feat/vector-store-iter-all
(#1316), before that PR fixed a false-positive/silent-truncation bug in
QdrantStore.iter_all(): an empty scroll page with a still-advancing cursor
(e.g. a window landing entirely on tombstoned points) was treated as the
end of the collection instead of continuing. Syncing qdrant_store.py,
vector_store.py, and their tests to #1316's current tip (fa967983) so this
branch doesn't reintroduce the already-fixed bug once merged. Content-only
sync of the 4 shared files (verified via diff against origin/feat/vector-store-iter-all)
rather than a full branch merge, to avoid pulling in unrelated main drift
that has landed on that branch since this one diverged.
2026-09-02 17:39:42 +05:30
Zohaib Hassnain 110f6deb1e Merge pull request #1316 from semantica-agi/feat/vector-store-iter-all
feat(vector_store): add iter_all enumeration for cursor-based backends
2026-09-02 17:37:11 +05:30
Zohaib Hassnain b8299b1427 chore: clean it 2026-09-02 17:19:19 +05:30
Zohaib Hassnain fa967983e6 fix(vector_store): dedupe qdrant record conversion, don't abort iter_all on a live cursor with an empty page 2026-09-02 17:19:19 +05:30
Zohaib Hassnain bbd423c50a fix(vector_store): raise on stalled pinecone pagination instead of truncating 2026-09-02 17:19:19 +05:30
Zohaib Hassnain fcdad56893 making it clean 2026-09-02 17:19:19 +05:30
Zohaib Hassnain 6b36379f15 feat(vector_store): add pinecone iter_all 2026-09-02 17:19:19 +05:30
Zohaib Hassnain f2e7b9ed75 fix(vector_store): raise instead of truncating when a qdrant scan cannot advance 2026-09-02 17:19:19 +05:30
Zohaib Hassnain 5e80ebd837 fix(vector_store): drop qdrant migrate wiring, keep iter_all only
VectorStore cannot actually migrate to or from qdrant yet. _init_backend_store constructs QdrantStore without connecting or selecting a collection, so reads raise a Collection not initialized error, and the facade store_vectors dispatches only to add/add_vectors while QdrantStore exposes insert_vectors, so writes raise NotImplementedError.

Both are pre-existing facade gaps that nothing had exposed, since migrate previously only allowed faiss/sqlite/pgvector. Adding qdrant to the allowlist claimed support that does not work end to end, so it is removed along with the dimension inference that only fires for backends missing a .dimension attribute. Tracked separately; this PR keeps just the iter_all primitive.
2026-09-02 17:19:19 +05:30
Zohaib Hassnain d175f894a4 feat(vector_store): add iter_all enumeration and wire up qdrant migration 2026-09-02 17:19:19 +05:30
Mohd Kaif 3d32254b07 Merge pull request #1390 from semantica-agi/fix/security-scan-pip-audit-migration
fix(ci): migrate security-scan from Safety to pip-audit
2026-09-02 16:39:01 +05:30
KaifAhmad1 b5199ae6e3 fix(ci): handle pip-audit skipped dependencies, restore manual trigger, fix stale docs
Addresses review feedback on this PR:

- Guard 2 and the PR-comment JS parser both required every dependency
  in pip-audit's report to carry an array-valued `vulns` field. A
  dependency pip-audit can't resolve/audit is reported instead as
  {"name": ..., "skip_reason": ...} with no `vulns` key at all (see
  pip_audit._format.json.JsonFormat._format_dep) - a normal, documented
  shape, not a malformed one. That made a single unauditable package
  hard-fail the whole job and show "Invalid report structure" in the PR
  comment, reintroducing the same class of scan-unrelated CI break this
  migration was meant to fix for Safety. Both now accept skipped
  entries, treat them as zero vulns, and surface them explicitly (job
  log + PR comment) instead of silently dropping or crashing on them.
  Verified the fixed jq queries and JS parse logic against synthetic
  pip-audit report fixtures covering the normal, skipped, and malformed
  shapes.

- Restored a `workflow_dispatch` trigger on security-scan.yml. Deleting
  security.yml (which had it) left no way to manually run a dependency
  audit on demand.

- Updated SECURITY.md, which still described security.yml as a live
  scanning workflow and Safety as an active scanner after this PR
  deletes both.
2026-09-02 16:22:16 +05:30
Mohd Kaif db48f73755 Merge branch 'main' into fix/security-scan-pip-audit-migration 2026-09-02 16:01:17 +05:30
Zohaib Hassnain 07113d2d2d fix(ci): migrate security scan from Safety to pip audit 2026-09-02 15:21:36 +05:00
Shubham SrivastavaandSameer Kadam 909ccf0ded test: install extractor dispatch mocks per test, not at module scope (#1337)
The module assigned MagicMocks into sys.modules at import time and never
removed them. pytest imports every test module during collection before
running anything, so those mocks were live while later modules were
imported and each bound them into its own globals.

132 tests passed alone and failed in a full-suite run as a result. Full
suite goes from 199 failed / 5506 passed to 67 failed / 5638 passed.

A tearDownModule cannot fix this: collection has already finished by the
time it runs. The extractors resolve 'from .methods import
get_entity_method' lazily inside their methods, so the stand-in only has
to be in sys.modules while a test executes - it is now installed per test
via patch.dict in setUp and removed by addCleanup.

Co-authored-by: Sameer Kadam <sskadam6305@gmail.com>
2026-09-02 15:31:38 +05:30
Guofang.Tang fb69b033be fix(ontology): resolve endpoints in direct property inference (#1229)
The relationship-endpoint fix merged in #1170 covers the main ontology generation pipeline, but the public property-inference path still had the same gap.

`OntologyGenerator.infer_properties()`, and the `PropertyGenerator` it delegates to, fell back to `owl:Thing` for both domain and range when a relationship used entity IDs or aliases instead of explicit `source_type` / `target_type` values. The pipeline resolved those endpoints correctly, but the public API path did not.

This moves the existing alias-building and endpoint-resolution logic out of `OntologyGenerator` and into a shared `relationship_utils` module:

* `build_entity_aliases`
* `get_relationship_endpoint`
* `resolve_relationship_endpoint_type`

`OntologyGenerator` now uses those shared helpers instead of keeping its own copies.

`PropertyGenerator._infer_object_properties()` now also receives the entity list, builds the same alias index, and uses the shared endpoint resolver. This replaces the old fallback:

```python
rel.get("source_type") or self._infer_class_from_entity(...)
```

which could only fall back to `owl:Thing` because `_infer_class_from_entity()` never actually resolved an entity.

There are two small behavior changes from centralizing the logic. `build_entity_aliases()` now converts `entity_type` to `str` before adding it to the alias set, avoiding mixed-type alias values. `resolve_relationship_endpoint_type()` returns `None` rather than `""` when there is no usable explicit type, since an empty string isn't a meaningful endpoint type.

The new `test_public_infer_properties_resolves_id_endpoints` covers the broken public API path directly. It creates entities and ID-based relationships through `infer_classes()` / `infer_properties()` and verifies that the inferred `worksFor` property resolves to `Person` for the domain and `Organization` for the range instead of falling back to `owl:Thing`.

That exercises the same endpoint-resolution behavior already covered by the pipeline tests, but through the public entry point that was still missing it.
2026-09-02 14:26:13 +05:00
Guofang.Tang c10090dc9b fix(ci): fail closed on malformed Safety reports (#1366)
The Security Scan workflow already scans `requirements-ci.txt` directly, but malformed Safety output could still be treated as a clean scan. If the report existed on disk but `vulnerabilities` was missing, `null`, or the wrong type, the workflow could end up counting it as zero findings.

This adds a structural check immediately after the report is written. `vulnerabilities` must be an array; otherwise the step fails closed with a clear error instead of treating a broken report as a successful scan.

There was a related problem in the PR reporting path. The comment step already knew how to render an `Invalid report structure` warning, but that branch was effectively unreachable. In GitHub Actions, a custom `if:` is implicitly gated by `success()` unless it includes a status function such as `always()` or `failure()`. Once the Safety step exited non-zero, the Upload and Comment steps were skipped, so the warning could never be posted.

Fixing that required changing how Safety failures flow through the job rather than just adding another guard. The Safety step now uses `continue-on-error: true`, which lets Bandit and Semgrep continue running and allows the Upload and Comment steps to process the failed or malformed Safety result.

Because `continue-on-error` means the Safety step no longer carries the job's final failure signal itself, the workflow now tracks that state explicitly with `SAFETY_SCAN_STATUS`. It is set to `failed` at the start of the Safety step, before any validation runs, and changes to `passed` only when the report is valid and contains zero vulnerabilities.

That default-failed behavior covers every other exit path: a missing report, malformed `vulnerabilities` field, invalid vulnerability count, Safety failure, or an actual vulnerability finding all leave the status as `failed`.

A final `Enforce Safety Gate` step checks `SAFETY_SCAN_STATUS` and fails the job unless it is exactly `passed`. This keeps the same merge-blocking behavior while still allowing the rest of the security checks and reporting steps to run after a Safety failure.

This is a follow-up to #1356. The overlapping Safety behavior changes and duplicate pip-audit path from that PR were dropped after `main` picked up the canonical fix for the underlying `cuda-toolkit` crash. This change keeps only the report-validation hardening that remains independent of that fix.
2026-09-02 14:12:33 +05:00
Zohaib Hassnain 28c96c2539 fix(ci): update actions/deploy-pages pin to current v5 (v5.0.1) (#1387) 2026-09-02 13:59:42 +05:00
Ahmad Bilal 170b4215a6 fix(vector_store): persist vector_ids and metadata across FAISS index save/load (#1272) (#1314)
`FAISSIndex.save()` previously wrote only the raw FAISS index. `load()` then rebuilt the wrapper with empty `vector_ids` and `metadata`, so that state was never restored.

That made a save/load round trip effectively unusable through the wrapper API: `scan_vectors()` returned no vectors, `count()` returned `0`, and `get_vector(id)` returned `None` for IDs that were present in the underlying FAISS index.

This also affected migration. `semantica store migrate --from faiss` could load a valid FAISS index, see zero vectors through `scan_vectors()`, migrate nothing, and still exit successfully. Since FAISS is a supported migration source, this was a silent data-loss path rather than just a persistence bug.

The fix adds a `.meta.json` sidecar next to the FAISS binary. It stores:

* `vector_ids`
* `metadata`
* `dimension`
* `index_type`

The sidecar is written atomically using a temporary file and rename. `save()` also serializes the metadata before writing the FAISS binary, so a serialization error fails before either persistence artifact is created. That avoids leaving a valid-looking index file behind without the metadata needed to use it correctly.

On load, `dimension` and `index_type` come from the sidecar rather than the caller's arguments. This makes the reconstructed wrapper reflect the index that was actually saved instead of relying on the caller to provide matching values.

There are also explicit checks for incomplete or inconsistent persisted state. If the sidecar is missing, which can happen with indexes written by older versions or when only the FAISS binary was copied, `load()` emits a `RuntimeWarning` instead of silently returning an apparently usable wrapper with no IDs or metadata. If the number of saved vector IDs doesn't match the FAISS index's `ntotal`, `load()` raises `ProcessingError` rather than returning a state where vectors exist in FAISS but can't be reached through `scan_vectors()`.

Metadata serialization changed during review as well. The first version used `json.dumps(..., default=str)`. That avoided failures for values such as `datetime`, `UUID`, and `set`, but it was lossy: those values came back as strings instead of their original Python types.

That was replaced with a tagged encoder/decoder that preserves the supported types across a round trip. It currently handles sets, datetimes, dates, UUIDs, NumPy scalars and arrays, and bytes, with bytes stored as base64.

The decoder also uses an exact-schema check for tagged values. A normal dictionary that happens to contain a reserved tag key alongside other fields is left alone instead of being interpreted as an encoded type.

The final implementation was spread across fourteen commits, mostly following review feedback. Those changes included cleaning up conflict markers from an unfinished stash pop, expanding round-trip and retry coverage, adding the missing-sidecar warning, adding an end-to-end `scan_vectors()` persistence test, replacing lossy metadata serialization with the tagged format, checking FAISS/sidecar count mismatches, adding `bytes` support, and reordering `save()` so metadata serialization happens before the FAISS index is written.
2026-09-02 13:44:42 +05:00
Zohaib Hassnain 5f600a3f36 fix(ci): ignore SFTY-20260723-60537 (CVE-2026-65918) in torchvision, unreachable transitive dep (#1385) 2026-09-02 13:37:03 +05:00
Mohd Kaif 1d18755a4e Delete cookbook/advanced/13_Manual_Ontology_Snowflake_Mapping.ipynb 2026-09-02 13:54:13 +05:30
Mohd Kaif 3acf801273 Merge pull request #1361 from taoche/fix/semantic-layer-basics-intro
docs(cookbook): rewrite Semantic Layer Basics as an introductory workflow
2026-09-02 13:22:42 +05:30
KaifAhmad1 796f181c75 docs(cookbook): fix stale RDFExporter claim, pin oxigraph install version
Step 5 explicitly avoids RDFExporter's compact projection and builds
the Turtle export from the TripletStore's own triples instead, but the
Summary cell still credited RDFExporter -- a leftover from before the
rdflib-based export replaced it. Correct the claim to match the code.

Also pin the install to >=0.6.7: earlier releases could return
ontology classes with an empty uri (#1103), which made entity_type
mappings silently resolve to None instead of raising, so the notebook
would appear to pass while never actually typing its instances.
2026-09-02 13:15:14 +05:30
Mohd Kaif 8e73ed8d4c Merge branch 'main' into fix/semantic-layer-basics-intro 2026-09-02 12:44:19 +05:30
Mohd Kaif 114641b39d Merge pull request #1359 from taoche/fix/cookbook-08-end-to-end
docs(cookbook): make notebook 08 a real rerunnable KG workflow
2026-09-02 12:26:00 +05:30
Mohd Kaif f71b711205 Merge branch 'main' into fix/cookbook-08-end-to-end 2026-09-02 12:11:43 +05:30
Mohd Kaif f9a661a4ed security(deps-dev): bump browserslist from 4.28.2 to 4.28.8 in /explorer (#1382)
Fixes GHSA-73wf-gq98-2v4g (prototype pollution / DoS via unguarded
browserslist-stats.json parsing) and GHSA-c83g-rgw3-j3cx (unbounded
cache growth leading to OOM), both patched upstream in 4.28.7.
2026-09-02 11:49:55 +05:30
Kevin 8e7aaee4f5 fix(vector-store): validate collection schema in MilvusStore.get_collection (#1344)
`get_collection()` attached any collection right after the existence check, with no look at its schema. A collection with an INT64 primary key, or one missing the `metadata` field entirely, would attach without complaint and only fail later, inside `get_vector()` or `get_metadata()`, with an error that gave no hint the real problem was upstream at attach time.

This adds a schema check between the attach and the assignment to `self.collection`, so a mismatch is caught at the point of failure instead of surfacing three calls later as an unrelated-looking error. The check validates against exactly the shape `create_collection()` builds: a `VARCHAR` primary key named `id` with `auto_id=False`, a `FLOAT_VECTOR` field named `vector`, and a `JSON` field named `metadata`. Anything else, wrong dtype, wrong name, a missing field, or an auto-generated id, is rejected before the store ever holds a reference to it.

The auto_id and metadata-dtype checks were added in a second pass after review. A collection with `auto_id=True` still attached cleanly and only broke once the store tried to insert with the explicit ids it always sends, and a `metadata` field that existed but wasn't `JSON`-typed only broke during a later write or metadata filter, for the same reason: schema drift that looked fine at attach time and failed downstream instead of at the source.

Nine tests cover this: the one matching-schema case that should succeed, and each rejection path independently, wrong pk dtype, missing pk, wrong pk name, auto_id pk, missing vector field, wrong vector dtype, missing metadata field, and non-JSON metadata.

Closes #1331.
2026-09-02 00:58:26 +05:00
Zohaib Hassnain af829f5f20 fix(vector_store): don't let iterator close() mask the real scan error, dedupe milvus result shaping, split unavailable/uninitialized messages 2026-09-01 23:09:53 +05:00
Zohaib Hassnain 930e7f9b71 docs(vector_store): note the milvus schema assumption 2026-09-01 23:09:53 +05:00
Zohaib Hassnain e335971dcd feat(vector_store): add milvus iter_all 2026-09-01 23:09:53 +05:00
Zohaib Hassnain 78682076d5 fix(vector_store): raise before yielding on a stalled weaviate cursor, extract v4 dict vectors, dedupe fallback ladder 2026-09-01 23:03:07 +05:00
Zohaib Hassnain 1227947be5 fix(vector_store): dedupe qdrant record conversion, don't abort iter_all on a live cursor with an empty page 2026-09-01 22:51:32 +05:00
Zohaib Hassnain b4a14d87f5 making it clean 2026-09-01 22:51:32 +05:00
Zohaib Hassnain 3bf89e523f fix(vector_store): raise instead of truncating when a qdrant scan cannot advance 2026-09-01 22:51:32 +05:00
Zohaib Hassnain 2b5b62bb8d fix(vector_store): drop qdrant migrate wiring, keep iter_all only
VectorStore cannot actually migrate to or from qdrant yet. _init_backend_store constructs QdrantStore without connecting or selecting a collection, so reads raise a Collection not initialized error, and the facade store_vectors dispatches only to add/add_vectors while QdrantStore exposes insert_vectors, so writes raise NotImplementedError.

Both are pre-existing facade gaps that nothing had exposed, since migrate previously only allowed faiss/sqlite/pgvector. Adding qdrant to the allowlist claimed support that does not work end to end, so it is removed along with the dimension inference that only fires for backends missing a .dimension attribute. Tracked separately; this PR keeps just the iter_all primitive.
2026-09-01 22:51:32 +05:00
Zohaib Hassnain 3a0f3f672a feat(vector_store): add iter_all enumeration and wire up qdrant migration 2026-09-01 22:51:32 +05:00
e040d84d59 feat(context): add ErasureCoordinator for cross-store entity erasure (#1027)
* feat(context): add ErasureCoordinator for cross-store entity erasure

purge_node() is graph-scope by design (#957), so an entity removed from the
graph can survive verbatim as an AgentMemory item and as an embedding. The
changelog names GDPR Article 17 as purge's motivation, and an Article 17
erasure the vector store can still answer queries from is not an erasure --
it is worse than none, because purge_node() returns True and writes a
tombstone attesting the content is gone.

ErasureCoordinator composes the existing public APIs to drive the cascade and
returns an ErasureReceipt recording what each store reported. Nothing in
context_graph.py or agent_memory.py changes behaviorally; ContextGraph keeps
its documented graph-scope contract instead of acquiring references that would
invert the dependency.

Honest partial reporting is the point. Stores report erased / not_found /
not_configured / unsupported / failed, and complete is False when any store
reports unsupported or failed. FAISS, Milvus and Weaviate expose no delete at
all, so erasure genuinely cannot be completed on them today -- the receipt
says so rather than reporting a success it did not achieve.

Erasure runs outward-in (vectors, memory, graph). The tombstone is the durable
attestation, so writing it first would let a crash mid-cascade leave a record
claiming more than happened; erasing the graph last leaves a partial failure
recoverable and honest.

The memory sweep pages until dry and re-queries afterwards rather than
trusting one find_by_entity() call, whose limit=10 default silently truncates
the very check a caller uses to decide the erasure is done. Unsupported vector
backends are detected by probing the wrapped backend, since the VectorStore
facade declares delete_vectors() for every backend and only raises
NotImplementedError once called.

27 tests against real ContextGraph/AgentMemory instances, including the
25-items-on-one-entity regression that fails against a naive single-call
sweep. Full tests/context/ suite: 596 passed.

Closes #1018

* docs(context): document ErasureCoordinator in the context API reference

* test(context): exercise ErasureCoordinator against a real VectorStore

The vector-leg tests asserted the three backend shapes the coordinator
expects -- delete_vectors / delete / neither -- against fakes, which is worth
exactly as much as the assumption that a real store looks like one of them.
VectorStore(backend="inmemory") runs without external services, so it can
hold that assumption to account.

Adds the end-to-end case the receipt actually attests to: a real ContextGraph,
AgentMemory and VectorStore, where the embedding is written by
AgentMemory.store() and has to be gone afterwards. That exercises the memory
leg's own delete_memory() vector cascade rather than the coordinator's model
of it.

The real backend also pins a limit worth knowing before trusting the receipt:
it pops the ids and returns True whether or not they were there, and no
backend offers a portable existence check, so `erased` on the vectors leg
means the store accepted the delete for the ids given -- not that embeddings
were really removed. The memory leg re-queries to confirm and so is the
stronger claim. Documented on STATUS_ERASED, _erase_vectors(), and both status
tables.

tests/context/: 599 passed.

* fix(context): address review findings on ErasureCoordinator

Timestamp drift (high). erase_entity() resolved erased_at up front but passed
the caller's original `at` down to purge_node(), so on the default at=None
path the coordinator and the graph each took their own now() and the receipt
attested to a different instant than the tombstone it points at -- breaking
the invariant this module states most loudly. The resolved value is now what
the graph receives. The existing test passed only because it supplied an
explicit `at`, which hides the drift; the regression test covers at=None,
which is what callers actually use.

Backend delete results. The vectors leg treated anything other than the
literal False as success, but no in-repo backend returns a bool -- Qdrant
returns {"status": <UpdateStatus>} and Pinecone {"deleted": True}, so every
dict read as success and the backend's own account of the delete was thrown
away. Results are now interpreted by shape and the payload is kept in the
receipt as backend_result, stringified so it stays JSON-serializable as an
audit record. Bool markers match by identity so a 0 count isn't read as
False; string markers match as substrings so an enum rendering as
"UpdateStatus.FAILED" isn't read as success.

Falsey vector store. The "at least one store" guard used `not vector_store`,
rejecting a valid store whose __bool__/__len__ makes an empty instance falsey
and then reporting vector_store=None when an object had been passed. It now
separates None (absent) from False (deliberately disabled) from provided, and
echoes what it received.

`at` annotations. Widened to int/float, matching the ContextGraph normalizer
they delegate to, so the coordinator stops advertising less than the API it
wraps.

tests/context/: 608 passed.

* fix(context): report memory-owned vectors that survive erasure (#1018)

A receipt could read complete while an embedding was still in the vector
store. The vector leg deleted `vector_ids` or `[entity_id]`, and the memory
leg relied on `AgentMemory.delete_memory()` to cascade to the vectors each
item owns. That cascade is best-effort: `_delete_vector_ids()` raises when a
backend returns False, `delete_memory()` catches it, logs a warning, and
still returns True. So `batch_delete` counted the item, the residual re-query
found no items, the memory leg reported `erased`, and nothing in the receipt
recorded that the embedding was refused.

Reproduced with a store that deletes the entity-keyed id and refuses the
memory-owned one: `receipt.complete` was True with the embedding still live.
That is the failure mode this module exists to prevent -- a receipt is a
compliance artifact, and one that overstates is worse than none.

Fix by deleting memory-owned vector ids through the coordinator's own vector
leg, which reports honestly, instead of trusting the memory leg's cascade.
The ids are collected before anything is deleted, while the items still exist
to be enumerated, and are unioned with any caller-supplied ids rather than
replacing them.

This needs one addition to AgentMemory: `vector_ids_for(memory_id)`, a
read-only accessor mirroring the fallback in `delete_memory` (an item stored
without tracked ids is keyed by its own memory id). Reaching into
`_vector_ids` from the coordinator would have been the internals-access
pattern this repo keeps getting bitten by. No existing AgentMemory behaviour
changes -- `delete_memory()` still cascades best-effort, so other callers are
unaffected; the coordinator simply no longer depends on that being reliable.
It does mean the vectors are attempted twice, which is a no-op on a working
store and only ever costs a log line.

Note this deviates from the PR's stated "nothing in agent_memory.py changes"
constraint. The constraint could not hold: with `_vector_ids` private and no
portable way to ask a vector store what it still holds, the coordinator had
no way to make the claim truthful without it.

Four tests: the refused-vector case (receipt must be incomplete), that
memory-owned ids reach the store, that explicit `vector_ids` do not displace
them, and the accessor's fallback. The first three were confirmed to fail
against the previous coordinator, on the `receipt.complete` assertion rather
than incidentally. 655 tests pass across tests/context and the agno
integration.

* fix(context): make erasure receipt vector failures honest

* fix(context): optimize erasure pagination handling

* fix(context): use one timestamp for batch erasure

* style: strip trailing whitespace from erasure.py and test file

* docs(changelog): correct test counts to 48 / 738 after review rounds

---------

Co-authored-by: Pravit Ampapathini <pravit.amp@gmail.com>
Co-authored-by: Sameer6305 <sskadam6305@gmail.com>
2026-09-01 23:07:02 +05:30
Mohd Kaif 18fb7c3ec0 Merge branch 'main' into fix/semantic-layer-basics-intro 2026-09-01 21:51:31 +05:30
Mohd Kaif 2eab7ab876 fix(ci): drop --ignore from Safety check, filter accepted CVEs in jq instead (#1371)
* fix(ci): drop --ignore from Safety check, filter accepted CVEs in jq instead

The follow-up to #1370: adding `--ignore SFTY-20260120-40557` to the
`safety check` invocation reintroduced the exact crash #1131/#1157 had
just fixed - "Unhandled exception happened: 'cuda-toolkit'" - but only
once Safety actually has a live vulnerability match to apply the ignore
against (the plain, un-ignored scan against the same requirements-ci.txt
had already succeeded and correctly reported that same match on main,
per the run right before this one).

I couldn't reproduce this locally: my local Safety installation doesn't
surface the live cuda-toolkit CVE match at all (its open-source
vulnerability DB appears to lag CI's), so --ignore never had a real
match to crash on in my testing. That's on me - I should have caught
that my "0 vulnerabilities" local result meant the DB hadn't even seen
the finding yet, not that the fix worked.

Since I can't safely iterate against Safety's own --ignore path without
live-DB access, this moves the "should we still fail on ID X" decision
out of Safety entirely: run the plain scan (the one path an actual CI
run has now proven doesn't crash), then filter the accepted vulnerability
ID out of the report ourselves in jq before counting/printing. Verified
the jq expression directly against a synthetic report shaped like a real
one (id present + one other unrelated id): filters exactly the intended
entry, and - as a bonus - iterating over a null/missing "vulnerabilities"
key with jq now raises inside jq the way the existing guard comment always
assumed it did, rather than silently coming back as 0.

* fix(ci): apply the accepted-CVE exclusion list to the PR comment too

Qodo caught a real gap on this PR: the jq-based exclusion I added only
covers the CI gate (the VULNS count and the failure-path detail print).
The "Comment PR with Security Results" step reads safety-report.json
independently in its own JS, with no filtering at all, so a PR touching
only the accepted cuda-toolkit CVE would still get a comment saying
"Found 1" even though the gate itself correctly treats it as
non-actionable and passes.

Export IGNORED_VULN_IDS via $GITHUB_ENV from the shell step so the JS
step can read the same list, and filter data.vulnerabilities there
before rendering - with a footnote naming what was excluded and why,
so the comment stays transparent about the accepted finding rather than
just silently hiding it.

Verified the JS logic standalone against two synthetic reports: one with
the accepted CVE plus an unrelated real one (shows only the real one,
plus the footnote), and one with only the accepted CVE (shows "No
findings" plus the footnote, rather than misleadingly looking identical
to a clean scan with no explanation).
2026-09-01 19:45:36 +05:30
Mohd Kaif d8822198cf fix(ci): ignore CVE-2025-33228 in cuda-toolkit - unfixable transitive pin, unreachable code path (#1370)
Merging #1357 surfaced a real (not crashed) Safety finding: cuda-toolkit
13.0.3.0 < 13.1.0 is affected by SFTY-20260120-40557 / CVE-2025-33228.

This can't be fixed with a version bump on our end: torch 2.13.0 (the
latest release on PyPI - there is no newer one) hard-pins
`cuda-toolkit[cublas,cudart,cufft,cufile,cupti,curand,cusolver,cusparse,
nvjitlink,nvrtc,nvtx]==13.0.3` on Linux via its own METADATA, not a loose
transitive requirement we control.

The underlying CVE is OS command injection in NVIDIA Nsight Systems'
gfx_hotspot recipe (process_nsys_rep_cli.py), which requires a human to
manually invoke that script with an attacker-supplied string. It isn't
reachable from any Semantica code path, and Nsight Systems isn't even
part of the extras torch requests here (cublas/cudart/cufft/cufile/
cupti/curand/cusolver/cusparse/nvjitlink/nvrtc/nvtx - no Nsight extra
among them).

Ignoring this one vulnerability ID only (not the whole package or a
blanket policy) so CI reflects actionable risk. Re-evaluate once torch
ships a release that pins a patched cuda-toolkit.
2026-09-01 19:14:21 +05:30
Mohd Kaif e6409217dd Merge pull request #1357 from taoche/fix/cookbook-07-graph-mapping
docs(cookbook): correct graph mapping and deduplication in notebook 07
2026-09-01 18:41:18 +05:30
taoche e6c05df33e docs(cookbook): index semantic layer capstone 2026-09-01 19:51:49 +08:00
taoche c20a46f026 docs(cookbook): preserve complete semantic layer RDF 2026-09-01 19:45:40 +08:00
taoche 0dbe9274eb docs(cookbook): install notebook 08 NER model 2026-09-01 19:45:39 +08:00
taoche 3ed31b9182 docs(cookbook): make notebook 07 setup deterministic 2026-09-01 19:45:39 +08:00
Mohd Kaif 7300fb41b1 Merge pull request #1363 from 7487/fix/plugin-manifest-agents-array
fix(plugins): declare agents as an array of file paths in plugin.json
2026-09-01 16:57:52 +05:30
Mohd Kaif 218e5a33f3 Merge branch 'main' into fix/plugin-manifest-agents-array 2026-09-01 16:38:36 +05:30
Mohd Kaif 635f6e52f4 Merge pull request #1332 from semantica-agi/test/backend-facade-contract
Pin facade contract gaps for cloud backends
2026-09-01 15:20:08 +05:30
Mohd Kaif 9240a1b1f7 Merge branch 'main' into test/backend-facade-contract 2026-09-01 15:08:35 +05:30
3254b9be80 Serve the RDF export formats the MCP tool already offers (#1131) (#1157)
* feat(explorer): serve the RDF export formats the MCP tool already offers (#1131)

`POST /api/export` accepted only `json` and `csv` and answered 422 for everything
else, while the MCP `export_graph` tool resolved Turtle, N-Triples, RDF/XML,
JSON-LD, GraphML and Parquet through `semantica.export`. Two surfaces of one
product disagreeing about what the product can do — and for an RDF-native project,
a graph that loads as JSON-LD and cannot be exported as RDF is a one-way door.

The route now reaches the same exporters the MCP tool uses. Nothing is
reimplemented: `RDFExporter.export_to_rdf` and `GraphMLExporter.export` receive the
dict `session.build_graph_dict()` already builds.

The alias table is a copy of `mcp/tools/export.py::_FORMAT_ALIASES` plus the
spellings the issue mentioned (`ntriples`, `rdf-xml`), and a test asserts the two
tables agree — if either drifts, the formats a caller can use would depend on which
door they came through.

Media types and extensions per serializer, so a Turtle export is `text/turtle` and
not `application/json` with a `.json` name.

The 422 message now names what IS supported. The old one said only that the format
was unsupported, which reads as "this format does not exist" rather than "this door
does not open it" — that is what sent me looking through the library.

Parquet is left out on purpose: `ParquetExporter.export` writes a file and returns a
path, so serving it over HTTP is a different shape of change and deserves its own
review.

Tests, in `TestImportExport`: the seven RDF spellings, each **parsed with rdflib**
rather than asserted on strings — a response that merely looks like Turtle is what
lets this class of gap survive a suite. Plus the alias-agreement canary and the
error message. Three mutations (rejecting RDF again, breaking one alias, emptying
the message) each turn the matching tests red.

110 tests in `tests/explorer/test_explorer_api.py` pass.

* fix(explorer): complete RDF export support

* fix(explorer): secure GraphML temporary file handling

* fix(ci): scan declared dependencies with Safety

---------

Co-authored-by: 13g4d0 <13g4d0@users.noreply.github.com>
Co-authored-by: Sameer Kadam <sskadam6305@gmail.com>
2026-09-01 15:01:13 +05:30
Guofang.Tang dc81acaefd fix(ontology): reject normalized class name collisions (#1230)
ClassInferrer.infer_classes() groups entities by their type string, then
normalizes each group name (PascalCase + singularize) when it builds the
ontology class. Two source types that only differ in casing or plurality,
like Person and person, both normalize to the same class name. Nothing
caught that, so the second type's entities silently got treated as
instances of the first type's class, and property inference downstream
picked up whichever properties happened to win.

Added a pass right after entities get grouped by type: normalize every
type name that meets min_occurrences, and if two different source types
land on the same normalized name, raise ValidationError before any class
gets built. The check reuses the exact same min_occurrences filter the
real class-emission loop uses, so it only fires on collisions that would
actually produce duplicate classes, not on types that get filtered out
anyway.

The error carries validation_context with the normalized name mapped to
every source type that collided into it, so whoever's calling this can
see exactly what to rename instead of just getting a generic message.

Added test_class_inference_rejects_normalized_type_collisions covering
Person/person landing on the same class.

Follow-up to #1171.
2026-09-01 14:25:54 +05:00
7487 f44020c742 fix(plugins): declare agents as an array of file paths in plugin.json
Claude Code's plugin schema rejects "agents": "./agents" (a bare
directory string) with:

    Validation errors: agents: Invalid input

so the bundled plugin has never been installable. Unlike "skills",
which accepts a directory string, "agents" must be an array of .md
file paths.

Replaced the string with the explicit list of the three agent files.
Verified with `claude plugin validate plugins` (2.1.231): fails on the
old manifest with the error above, passes after this change.

Added tests/test_plugin_manifest.py to guard the manifest shape: agents
is a non-empty array of existing .md paths that stays in sync with
plugins/agents/, and the skills directory exists.

Fixes #1350
2026-09-01 17:02:52 +08:00
taocheandClaude Fable 5 0e7cef4677 docs(cookbook): move Semantic Layer Basics from Advanced to Introduction
Advanced chapter 09 labeled an introductory composition of
already-taught APIs as an enterprise semantic layer: TripletStore was
imported but never used, property_mappings stayed empty, mappings were
derived by fragile name matching (works_for never matched worksFor),
and the exported RDF was the original graph rather than an
ontology-aligned one.

Replace it with introduction/26_Semantic_Layer_Basics.ipynb, which
demonstrates the minimal semantic-layer composition honestly:

- build a small graph, generate an ontology (min_occurrences=1 so all
  demo classes are inferred, base_uri in the user's namespace)
- explicit entity-type, relationship-type, and property mappings read
  from the ontology's inferred_from metadata instead of name matching
- apply the mappings to produce an ontology-aligned graph, export it
  as Turtle, and note the file exporter's property projection
- store the aligned graph in the embedded Oxigraph TripletStore and
  answer a business question with one SPARQL query
- distinguish teaching mappings from governed production mappings and
  point to Advanced 13 as the production continuation

Advanced 13 gains a positioning note naming the new lesson as its
prerequisite; introduction/14_Ontology links forward to the new
lesson. All cells execute top to bottom (verified with the embedded
Oxigraph backend).

Closes #1325

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 15:56:49 +08:00
taocheandClaude Fable 5 5c366c6b7e docs(cookbook): make notebook 08 a real rerunnable end-to-end workflow
The first-knowledge-graph lesson read the parser output from a key it
never returns (content vs text), then masked the failure with
hard-coded entities, a manually assembled NetworkX graph, and an
uninvoked KGVisualizer; the final cell deleted the sample file, so
rerunning intermediate cells failed. Rework the notebook so every
stage consumes the previous stage's output:

- parse via parsed_document["text"] with an assertion that content
  was actually extracted
- real NERExtractor/RelationExtractor output replaces the simulated
  entities and wrong hard-coded offsets
- GraphBuilder builds the graph from actual relation endpoints via a
  mention-span -> graph-ID map (consistent with notebook 07)
- KGVisualizer.visualize_network renders the graph and saves HTML
- deletion moved to an explicit optional cleanup cell, so parsing and
  downstream cells stay rerunnable

Closes #1289

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 15:45:28 +08:00
taocheandClaude Fable 5 abb65feff0 docs(cookbook): build notebook 07 graph edges from real relation endpoints
The knowledge-graph lesson fabricated relationship endpoints from loop
indices, hid the corruption behind count-only output, and displayed
only merged duplicate groups as the deduplicated result. Rework the
notebook so that:

- graph edges come from Relation.subject/Relation.object mapped
  through a mention-span -> graph-ID table
- the sample text keeps two separate "Apple Inc." mentions without the
  sentence-boundary merge edge case
- entity resolution shows which mentions merged (merged_from) and
  remaps relationship endpoints onto the canonical entity
- deduplication reports merge operations separately from the complete
  deduplicated set (merged + untouched entities)
- each stage prints its transformed records, and lightweight
  assertions pin the expected canonical entities and edges

Closes #1287

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 15:39:25 +08:00
Wei TaoandSameer Kadam 8b125d6476 fix(explorer): keep small Full Graph relationships readable (#1277)
* fix(explorer): keep small full graphs readable

* perf(explorer): avoid redundant realtime edge sync

---------

Co-authored-by: Sameer Kadam <sskadam6305@gmail.com>
2026-09-01 11:11:53 +05:30
Mohd Kaif f3c540cfd2 docs(readme): reposition Semantica as the semantic/context layer (#1348)
Lead the README with Semantica's identity as a semantic/context/knowledge
layer (Context Graph, KG, ontology and vocabulary governance via OWL/SHACL/SKOS),
with decision provenance and audit trails framed as a property of that
structure rather than the flagship pitch.
2026-08-31 22:11:13 +05:30
Sameer Kadam 46b18fbee3 fix(docs): use GraphStore facade in Neo4j quickstart (#1340)
The Neo4j example in the persistent graph store section was passing a raw
Neo4jStore straight into GraphBuilder. GraphBuilder calls add_nodes() and
add_edges() on whatever it's given, and those only exist on the GraphStore
facade, not on Neo4jStore itself. Anyone who copied the example got:

AttributeError: 'Neo4jStore' object has no attribute 'add_nodes'

Swapped the import and construction to GraphStore(backend="neo4j", uri=...,
user=..., password=...), which wraps Neo4jStore internally and actually has
the methods GraphBuilder needs.

Added a test in tests/kg/test_graph_builder_with_graph_store.py that builds
a small graph through GraphBuilder with a mocked GraphStore and checks
add_nodes/add_edges get called. Also kept a test for GraphBuilder without a
graph_store at all, so that path doesn't regress either.

Closes #1135
2026-08-31 20:34:04 +05:00
Mohd Kaif b0679d4f67 fix(ci): stop checkov's suppressed checks from reopening as new alerts (#1346)
* fix(ci): drop unpinnable benchmarks/requirements.txt install

Scorecard flagged this pip install as unpinned-by-hash (#6082). Can't
hash-pin it - benchmarks/requirements.txt doesn't exist in this repo, so
there's nothing to compile a lockfile from. Dropping it instead of
leaving it unpinned: the job already fails on the next real step
(benchmarks/benchmarks_runner.py, also missing), so this line wasn't
doing anything useful to begin with.

* fix(ci): hash-pin the spacy model download in benchmark.yml

Qodo review on this PR: dropping the benchmarks/requirements.txt install
(the previous failure point) let the job actually reach
`python -m spacy download en_core_web_sm`, which fetches an unpinned,
unhashed wheel from spacy-models' GitHub releases - undoing the point of
this PR by exposing a real unpinned-install path instead of a dead one.

Replaced with a hash-pinned direct-URL entry in benchmark-extra.in/.txt
for en_core_web_sm-3.8.0 (matches the spacy==3.8.15 already pinned in
base-deps.txt). uv independently computed the same sha256 I got via a
manual curl+sha256 of the release asset, and a --require-hashes dry-run
install verifies clean.

* fix(ci): stop checkov's suppressed checks from reopening as new alerts

Root cause found, not just worked around: checkov's SARIF exporter
includes every evaluated check as an ordinary result, including ones it
internally marked SKIPPED via the inline # checkov:skip= comments and
checkov.io/skipN annotations already on the Helm chart. It never uses
SARIF's own `suppressions` field and never drops them - so the exact same
already-suppressed finding reopens as a brand-new code scanning alert
number on every single run, forever (#6035/#6036, #6112-6115,
#6128-6131 are all the same 4 findings, manually dismissed 3 times now).

checkov's JSON output *does* correctly record which checks were skipped.
Added .github/scripts/filter_checkov_skipped.py, which cross-references
the JSON's skipped_checks against the SARIF's results (matched by check
ID + the last two path segments, since the two outputs use different path
roots) and drops anything checkov itself already decided to suppress,
before upload. Verified locally against a real checkov+helm run: removed
exactly the 4 known-suppressed helm chart results, left the 2 genuinely
real findings (deploy/gcp/cloudrun-service.yaml, deploy/kubernetes/
deployment.yaml) untouched.
2026-08-31 19:29:24 +05:30
Mohd Kaif d135ad185f fix(ci): drop unpinnable benchmarks/requirements.txt install (#1345)
* fix(ci): drop unpinnable benchmarks/requirements.txt install

Scorecard flagged this pip install as unpinned-by-hash (#6082). Can't
hash-pin it - benchmarks/requirements.txt doesn't exist in this repo, so
there's nothing to compile a lockfile from. Dropping it instead of
leaving it unpinned: the job already fails on the next real step
(benchmarks/benchmarks_runner.py, also missing), so this line wasn't
doing anything useful to begin with.

* fix(ci): hash-pin the spacy model download in benchmark.yml

Qodo review on this PR: dropping the benchmarks/requirements.txt install
(the previous failure point) let the job actually reach
`python -m spacy download en_core_web_sm`, which fetches an unpinned,
unhashed wheel from spacy-models' GitHub releases - undoing the point of
this PR by exposing a real unpinned-install path instead of a dead one.

Replaced with a hash-pinned direct-URL entry in benchmark-extra.in/.txt
for en_core_web_sm-3.8.0 (matches the spacy==3.8.15 already pinned in
base-deps.txt). uv independently computed the same sha256 I got via a
manual curl+sha256 of the release asset, and a --require-hashes dry-run
install verifies clean.
2026-08-31 18:57:28 +05:30
Mohd Kaif 96dbd3f0d4 fix(security): bump checkov to 3.3.16, fix aiohttp CVEs in its lockfile (#1342)
Dependabot flagged 12 aiohttp advisories (1 high, rest moderate/low - CVE
range covering request smuggling, websocket/parser bugs, cookie/redirect
issues) against aiohttp==3.13.5 pinned in checkov.txt. checkov==3.3.1
itself pinned `aiohttp<3.14.0`, which excludes every fixed release;
3.3.16 (latest) relaxes that to `<3.15.0`, so bumping checkov also lets
aiohttp resolve to 3.14.3 (fixes all of them).

Two alerts remain open, both genuinely blocked upstream rather than
something a version bump here can fix:
- asteval: checkov 3.3.16 (latest, still) hard-pins asteval==1.0.6 with
  no range; the fix (1.0.9) is unresolvable without violating checkov's
  own declared dependency - confirmed via `uv pip compile` refusing to
  solve it. Needs checkov itself to bump the pin upstream.
- ecdsa: 0.19.2 is already the latest release; the Minerva timing-attack
  advisory has no patched version, since python-ecdsa's maintainers have
  stated side-channel attacks are out of scope for the project.

Both are checkov's own transitive deps, used only for local static IaC
analysis in defender-for-devops.yml (no network signing/cloud-auth calls
that would actually exercise ecdsa's signing path) - dismissing on
GitHub with that reasoning as a separate step.
2026-08-31 18:21:36 +05:30
Mohd Kaif d4cd44e7f1 fix(docker): split explorer-extra.txt by Python version, fix broken build (#1341)
* fix(docker): split explorer-extra.txt by Python version, fix broken build

main's container-scan.yml has been failing since PR #1338 merged:

  ERROR: In --require-hashes mode, all requirements must have their
  versions pinned with ==. These do not:
      standard-aifc from .../standard_aifc-3.13.0-py3-none-any.whl
      (from audioread==3.1.0->-r explorer-extra.txt (line 30))

Root cause: explorer-extra.txt was compiled with `--python-version 3.11`
but is installed on the Dockerfile's actual python:3.13-slim interpreter.
librosa's audioread dependency needs standard-aifc/standard-sunau only
under `python_version >= "3.13"` (Python 3.13 dropped aifc/sunau from
stdlib) - a file resolved for 3.11 has no hash for those packages at all,
so --require-hashes fails outright once pip resolves against the real
3.13 environment instead of silently under-pinning.

Splits the file in two: explorer-extra-py311.txt (ci.yml, unchanged
resolution) and explorer-extra-py313.txt (Dockerfile, newly compiled for
--python-version 3.13). They aren't interchangeable and shouldn't be
recombined - documented in .github/requirements/README.md, including how
to catch this class of bug before it ships again.

* fix(ci): correct stale -o path in explorer-extra-py311.txt header

Qodo review on this PR: the autogenerated header comment still said
-o .github/requirements/explorer-extra.txt (the pre-rename path), which
would silently regenerate the wrong file if someone copy-pasted it.
2026-08-31 17:52:23 +05:30
Mohd Kaif 832412cc01 Merge pull request #1338 from semantica-agi/fix/scorecard-pinned-dependencies
fix(ci): hash-pin every pip install for Scorecard Pinned-Dependencies
2026-08-31 17:31:14 +05:30
Sameer6305 dfbe98cebd fix(ci): generate Checkov lockfile for Windows 2026-08-31 17:21:59 +05:30
Sameer6305 89c6c8a45d fix(ci): hash-pin Checkov installation 2026-08-31 17:01:16 +05:30
Sameer Kadam 4b24053851 Merge branch 'main' into fix/scorecard-pinned-dependencies 2026-08-31 16:15:42 +05:30
cxzg007andSameer Kadam c111277a1f fix(context): make auto_generate_id a real InitVar in decision models (#1153)
The six decision-model dataclasses (Decision, DecisionContext, Policy,
PolicyException, Precedent, ApprovalChain) accepted `auto_generate_id`
only as a plain `__post_init__` parameter. Since it was neither a field
nor an InitVar, the generated `__init__` never forwarded it, so the
parameter was always its `True` default and the `auto_generate_id=False`
validation branch was unreachable dead code.

Declare `auto_generate_id: InitVar[bool] = True` on each dataclass so the
generated `__init__` forwards it to `__post_init__`, restoring the
required-id contract. Serialization is unaffected because InitVar is not
a real field. Add regression tests covering the auto-generate path, the
required-id error path, and the explicit-id path.

Fixes #1152

Co-authored-by: Sameer Kadam <sskadam6305@gmail.com>
2026-08-31 15:50:40 +05:30
KaifAhmad1 5ed98cefbd fix(ci): hash-pin PEP 517 build isolation deps (setuptools, wheel)
Qodo review on this PR: `pip install --no-deps -e .` / `pip install
--no-deps .` still leaves PEP 517 build isolation on by default, which
fetches [build-system] requires (setuptools==84.0.0, wheel==0.48.0)
completely outside any hash checking - the --require-hashes installs
right next to it didn't cover this at all.

Adds .github/requirements/pep517-build.txt, hash-locked to the exact
pyproject.toml [build-system] requires, and installs it before every
local-source install (Dockerfile, ci.yml, benchmark.yml) with
--no-build-isolation so pip reuses those hash-verified copies instead
of fetching its own.
2026-08-31 15:43:59 +05:30
KaifAhmad1 48737629d9 fix(ci): hash-pin every pip install for Scorecard Pinned-Dependencies
Scorecard's Pinned-Dependencies check requires pip installs to be
hash-verified, not just version-pinned - our existing pkg==X.Y.Z pins
(and even pip install -r requirements-ci.txt, despite that file already
carrying hashes) still scored a 4 because the pin/hash isn't visible on
the command line itself.

Adds .github/requirements/*.txt: hash-locked files generated via
`uv pip compile --generate-hashes` for every pip target that isn't
already requirements-ci.txt, covering standalone CI tooling (build,
wheel, twine, uv, pip-audit, safety/bandit/semgrep/jq, pip/setuptools
bootstrap) and the project's own local-source installs. The latter
(`pip install -e ".[explorer]"`, `pip install -e .`) can't be hash-pinned
directly since there's nothing to hash for a local source tree; split
into `pip install --no-deps -e .` plus a separate hash-pinned install of
the actual fetched dependencies instead.

Also adds --require-hashes to every `-r requirements-ci.txt` install so
hash verification is enforced explicitly rather than only implied by the
file's own content.

Simplifies the Dockerfile in the process: it now installs from the same
pre-generated explorer-extra.txt (copied in at build time) instead of
extracting constraints from requirements-ci.txt at build time, which
also means setuptools gets its CVE-2025-47273 fix as a side effect of
the hash-pinned install rather than a separate upgrade step.

benchmark.yml: pip install -r benchmarks/requirements.txt is left
unpinned - that directory doesn't exist in this repo, so there's nothing
to generate hashes from. Pre-existing breakage, unrelated to this change.
2026-08-31 15:29:40 +05:30
Mohd Kaif 1d62217cc5 Merge pull request #1334 from semantica-agi/fix/container-scan-cves
fix(docker): resolve Trivy-flagged CVEs in the built image
2026-08-31 14:34:32 +05:30
KaifAhmad1 bf292ccbbc fix(docker): address terrascan findings, drop apt-get upgrade
Two terrascan/GHAS findings on the previous commit:
- AC_DOCKER_0052 (no apt-get upgrade in Dockerfiles): dropped it. It also
  wasn't fixing anything - Debian's openssl fix for CVE-2026-14456 is still
  in trixie-proposed-updates, not reachable via a normal upgrade. Pin both
  base images by digest instead (matches #1329's approach) so the docker
  Dependabot ecosystem bumps them once Debian ships a rebuilt image with
  the fix, and document why the QUIC DoS isn't reachable here regardless
  (HTTP-only via uvicorn).
- AC_DOCKER_0010 (pin pip package versions): setuptools was `>=78.1.1`;
  pinned to the exact 84.0.0 already used by pyproject.toml/requirements-ci.txt.

Also fixes two build breaks this introduces on its own: requirements-ci.txt
wasn't in .dockerignore's allowlist or container-scan.yml's path trigger,
so the COPY in the prior commit would have failed the image build outright.
2026-08-31 14:06:28 +05:30
KaifAhmad1 64d942503b fix(docker): extract requirements-ci.txt pins with Python instead of sed
The sed expression to strip requirements-ci.txt's line-continuation
backslash (`[\]$`) is valid POSIX/GNU sed - verified it exits 0 and
strips correctly - but it's easy to misread as broken (a bot reviewer
flagged it as an unterminated bracket expression), and the seemingly
more obvious `\$`/` \$` forms silently fail to match at all rather
than erroring. Swap to a small `re.findall` extraction so there's no
backslash-escaping judgment call left for a reader (bot or human) to
second-guess.
2026-08-31 14:06:02 +05:30
KaifAhmad1 471a420b10 fix(docker): resolve Trivy-flagged CVEs in the built image
Container Security Scan flagged five HIGH-severity findings against
semantica:scan:
- setuptools 70.3.0 (CVE-2025-47273, path traversal) - the base image's
  bundled copy, never touched by our own build. Upgraded explicitly.
- msgpack 1.1.2 (GHSA-6v7p-g79w-8964, OOB read/crash) - `pip install
  ".[explorer]"` re-resolved deps from scratch instead of reusing the
  audited, hash-pinned requirements-ci.txt (which already pins
  msgpack==1.2.1), so it landed on an unpatched transitive version. Now
  installs against a constraints file derived from requirements-ci.txt.
- openssl / libssl3t64 / openssl-provider-legacy (CVE-2026-14456, QUIC
  server DoS) - the Debian fix is still in trixie-proposed-updates, not
  yet promoted to trixie-security, so it can't be pulled via apt today.
  Added an apt upgrade step so the next image rebuild picks it up
  automatically once Debian ships it; documented why this image isn't
  actually exposed to it in the meantime (HTTP-only via uvicorn, no QUIC
  listener).
2026-08-31 14:06:02 +05:30
Mohd KaifandSameer6305 a4aa71ad87 fix(ci): unblock py3.9 install matrix and raise Scorecard pinning/signing (#1329)
* fix(ci): unblock py3.9 install matrix and raise Scorecard pinning/signing

pip install semantica failed on Python 3.9 across all three OSes because
spacy had no upper bound, so pip resolved spacy 3.8.16 whose thinc>=8.3.12
requirement has no cp39 wheels and no working sdist build path. Cap
spacy/thinc for python_version < '3.10' to the last wheel-compatible pair.

Also addresses the two OpenSSF Scorecard findings that were actually
fixable in code:
- Pinned-Dependencies: Dockerfile base images (node:26-alpine,
  python:3.13-slim) were unpinned by digest; pin both, and pin five
  previously-unversioned pip install calls in CI (build, safety, bandit,
  semgrep, jq, pip-audit).
- Signed-Releases: attest-build-provenance only publishes to the GH
  attestations API, which Scorecard doesn't inspect. Sign dist/* with
  Sigstore and attach the .sigstore.json bundles as release assets.

* fix(ci): correct Sigstore artifact inputs

---------

Co-authored-by: Sameer6305 <sskadam6305@gmail.com>
2026-08-31 14:04:28 +05:30
Zohaib Hassnain 73b14c00ba test(vector_store): make contract xfails reachable and cover the backend roster 2026-08-31 13:27:45 +05:00
Zohaib Hassnain bd1ba24b24 fix(vector_store): carry weaviate offset fallback across pages, raise on truncation 2026-08-31 13:22:23 +05:00
Zohaib Hassnain e9a756eac2 test(vector_store): pin facade contract gaps for cloud backends 2026-08-31 12:59:40 +05:00
Zohaib Hassnain e8ff36f088 feat(vector_store): add weaviate iter_all 2026-08-31 03:01:37 +05:00
Zohaib Hassnain ec9e63e16f fix(vector_store): drop qdrant migrate wiring, keep iter_all only
VectorStore cannot actually migrate to or from qdrant yet. _init_backend_store constructs QdrantStore without connecting or selecting a collection, so reads raise a Collection not initialized error, and the facade store_vectors dispatches only to add/add_vectors while QdrantStore exposes insert_vectors, so writes raise NotImplementedError.

Both are pre-existing facade gaps that nothing had exposed, since migrate previously only allowed faiss/sqlite/pgvector. Adding qdrant to the allowlist claimed support that does not work end to end, so it is removed along with the dimension inference that only fires for backends missing a .dimension attribute. Tracked separately; this PR keeps just the iter_all primitive.
2026-08-31 03:00:56 +05:00
Zohaib Hassnain 274d5d1195 feat(vector_store): add iter_all enumeration and wire up qdrant migration 2026-08-31 02:34:26 +05:00
Mohd Kaif fa87a1a9be ci: add npm Dependabot ecosystem and container image scanning (#1286)
* ci: add npm Dependabot ecosystem and container image scanning

- dependabot.yml had no npm ecosystem entry for explorer/, so its
  lockfile was never watched - exactly why the brace-expansion/nanoid
  CVEs fixed in #1280 went undetected. Add it, mirroring the existing
  pip entry's schedule/labels/reviewer conventions.
- New container-scan.yml builds the root Dockerfile's image and scans
  it with Trivy (CRITICAL/HIGH OS+lib CVEs, SARIF to the Security tab)
  and Syft (SPDX SBOM artifact), on push to main, weekly, and manual
  dispatch. Neither the base-image scan nor an SBOM existed before -
  Dependabot's docker entry only bumps the base image tag, it doesn't
  scan built layers.
- Trivy runs report-only for now (no exit-code gate): this is its
  first run against the image, so the CRITICAL/HIGH baseline hasn't
  been triaged yet. Once reviewed, add exit-code: '1' to make it a
  hard gate, same as Safety/Bandit-HIGH in security-scan.yml.

* fix: run Trivy via digest-pinned image, not the aquasecurity/trivy-action wrapper

verify-action-pins.sh failed in CI: the aquasecurity GitHub org has an IP
allow list on its API that 403s the live tag->SHA resolution from
Actions-runner IPs (confirmed reproducible, not transient - resolves fine
from a non-blocked host). Rather than carve a skip exception into the pin
verifier for an org this script already flags as a past tag-repointing
target (see its "LiteLLM/Trivy 2026 incident" comment), pull Trivy as a
sha256-digest-pinned Docker Hub image instead. A digest is immutable and
verifiable independently of GitHub's API entirely, so it sidesteps the
IP block without weakening verification of the one action this repo
already treats as higher-risk. Confirmed the pinned digest
(aquasec/trivy@sha256:62b1e65e...) resolves live against Docker Hub's
registry API.

* fix: match container-scan.yml's push paths to what actually reaches the image

The path filter only watched explorer/package.json and package-lock.json,
but Dockerfile COPYs the whole explorer/ tree plus README.md, LICENSE, and
MANIFEST.in, and .dockerignore controls all of it. A frontend source change
or a README/LICENSE edit would change the built image without triggering a
scan, silently drifting until the next weekly run. Replace the filter with
exactly .dockerignore's opt-in list.
2026-08-30 21:39:30 +05:30
Mohd Kaif 08c78bfb40 fix(security): resolve Scorecard vulnerability and token-permission alerts (#1280)
- Bump explorer's brace-expansion (minimatch dep) 5.0.8 -> 5.0.9 and
  nanoid (postcss dep) 3.3.16 -> 3.3.18, fixing GHSA-rgw5-rvv9-x895 and
  GHSA-2v37-7h3g-55p8 (both DoS via unbounded input, both within the
  existing caret ranges declared by their parents).
- Move codeql.yml and defender-for-devops.yml's security-events: write
  (and codeql.yml's actions: read) from workflow-level down to their
  single job, matching Scorecard's Token-Permissions ideal of a
  read-only top-level default with sensitive scopes granted only where
  used.
2026-08-30 20:58:39 +05:30
Mohd Kaif 56b174781f ci: reusable install action, install-matrix, and release hardening (#1266)
Distribution and trust-signal infrastructure to make pip install semantica
frictionless in downstream CI, and to bring the release pipeline in line
with mature OSS practice.

- .github/actions/setup-semantica: reusable composite action other repos
  can call to install + verify semantica in one step
- install-matrix.yml: verifies the published package installs and imports
  cleanly across Ubuntu/macOS/Windows x Python 3.9-3.12, weekly and on
  release; backs a new README badge
- scorecard.yml: OpenSSF Scorecard analysis, weekly and on push to main,
  backing a new README badge
- release.yml: twine check gate before publish, catching a broken PyPI
  long-description render before it ships
- CITATION.cff: enables GitHub's native "Cite this repository" button
- examples/ci/: copy-paste GitHub Actions, GitLab CI, and CircleCI
  templates for projects adopting semantica
- GROWTH.md: tracked checklist of distribution channels, what's done vs
  outstanding, with guardrails against inflating metrics artificially

Fixes folded in along the way:

- Re-pinned softprops/action-gh-release to the immutable v3.0.3 tag
  instead of the floating v3, after verify-action-pins.sh caught the
  mutable tag had drifted to a newer commit
- setup-semantica now passes extras/version through env vars instead of
  interpolating ${{ inputs.* }} directly into the bash script, closing
  a script-injection vector for callers deriving these from event data
- install-matrix now triggers on the Release workflow's completion
  (workflow_run) instead of release: published, since the GitHub release
  is created before the PyPI upload runs and the old trigger could race
  the publish
- The workflow_run path derives the expected version from the triggering
  tag and passes it into setup-semantica's version input, so pip
  installs and verifies the exact release instead of whatever's latest
  on PyPI at the time
- setup-semantica's pip caching is now opt-in (default disabled), since
  actions/setup-python errors out with cache: 'pip' enabled when the
  caller repo has no requirements.txt/pyproject.toml to key on
- examples/ci/github-actions.yml pins actions/checkout and
  actions/setup-python to verified commit SHAs instead of mutable tags
- examples/ci templates guard the requirements.txt install step with
  -f requirements.txt and call out pyproject.toml/Poetry/Pipenv as
  alternatives, since not every project has a requirements.txt
2026-08-30 17:30:21 +05:00
Zohaib HassnainandSameer Kadam dfda4c561a feat(vector_store): add scan_vectors enumeration and wire up store mi… (#1264)
* feat(vector_store): add scan_vectors enumeration and wire up store migrate

* fix(vector_store): address Qodo finds

* fix(vector_store): make FAISS add_vectors idempotent for retried migrations

---------

Co-authored-by: Sameer Kadam <sskadam6305@gmail.com>
2026-08-30 17:31:57 +05:30
Zohaib HassnainandSameer Kadam ea0dd17bff feat(llms): add Gemini, Ollama, DeepSeek, Novita provider wrappers (#1262)
* feat(llms): add Gemini, Ollama, DeepSeek, Novita provider wrappers

* fix(llms): address Qodo review findings on provider wrappers PR

* fix(llms): address review findings

---------

Co-authored-by: Sameer Kadam <sskadam6305@gmail.com>
2026-08-30 17:00:14 +05:30
Shubham Srivastava f6cd62411b test(integrations): make crewai and langchain test dirs packages (#1252)
Both directories contain a test_degradation.py. Neither had an __init__.py,
so under pytest's default prepend import mode both modules were imported as
plain 'test_degradation' and the second collided with the first:

    import file mismatch:
    imported module 'test_degradation' has this __file__ attribute:
      tests/integrations/crewai/test_degradation.py
    which is not the same as the test file we want to collect:
      tests/integrations/langchain/test_degradation.py

That aborted collection for tests/integrations/, so the langchain
graceful-degradation tests never ran. tests/integrations/__init__.py
already exists, and most directories under tests/ carry one; these two
subpackages were simply missed.

Collection goes from 335 collected, 1 error to 337 collected.

Closes #1251
2026-08-30 15:23:43 +05:00
王林 cac6dfbe45 fix(explorer): surface server error detail in graph loading failures (#1260)
The nodes/edges fetch loops were throwing away the response body whenever the request returned a non-OK status.

Because of that, errors like a `503` caused by a missing `SEMANTICA_API_KEY` only showed up as:

`Fetch failed: 503`

even though the backend was already returning a more useful message in the response `detail`.

This change reads the JSON error body and includes `detail` in the thrown error when it's a string, so `GraphLoadingOverlay` can show the actual backend error to the user.

Closes #1256
2026-08-30 15:16:15 +05:00
Mohd Kaif 80e9737542 Merge pull request #1254 from dex0shubham/fix/1134-progress-stream-stderr
fix(utils): write console progress to stderr instead of stdout
2026-08-30 14:22:17 +05:30
Mohd Kaif 14fb975fa1 Merge branch 'main' into fix/1134-progress-stream-stderr 2026-08-30 14:12:18 +05:30
Mohd Kaif 8b0ac61afd Merge pull request #1255 from semantica-agi/feat/claude-wrapper
feat(llms): add first class Anthropic provider wrapper
2026-08-30 13:43:58 +05:30
KaifAhmad1 9d98eedaa1 fix(llms): replace retired default model and tidy Anthropic wrapper
claude-3-sonnet-20240229 was retired 2025-07-21, so the wrapper's
default model and every copy-paste doc example would fail at
generate() time out of the box. Switch to claude-sonnet-4-6
everywhere (wrapper default, __init__ docstring, docs guide, tests).

Also cleans up leftover docstring typos/spacing from the previous
review pass and adds unavailable-path test coverage for
generate_structured()/generate_typed() to match generate(), clearing
ANTHROPIC_API_KEY in those tests so they don't flake on a runner that
has a real key set.
2026-08-30 13:33:18 +05:30
Zohaib Hassnain 5ffb212a8f Merge branch 'main' into feat/claude-wrapper 2026-08-29 22:15:41 +05:00
Zohaib Hassnain ca7f743dab fix(llms): address Qodo review findings on Anthropic wrapper 2026-08-29 22:15:19 +05:00
Zohaib Hassnain 5cf59fdd88 feat(llms): add first class Anthropic provider wrapper 2026-08-29 22:01:16 +05:00
Mohd Kaif 0384a8de30 Merge pull request #1077 from cxzg007/fix/rete-pattern-matching
fix(reasoning): implement RETE alpha/beta matching with Token model (#300)
2026-08-29 21:57:15 +05:30
KaifAhmad1 f3d5932c24 Merge remote-tracking branch 'origin/main' into fix/rete-pattern-matching
Reconciles this PR's Token-based alpha/beta matching (#300) with the
rule-actions/provenance layer merged separately in #1096. That PR built
bind_reasoner()/execute_matches() action-firing/_executed_activations/
reset_action_history() on top of the still-broken always-True stubs
(via an interim _bindings_for_rule() regex re-extraction), so main and
this branch touched the same propagation code with incompatible shapes.

Kept this branch's Token(facts, bindings) model for alpha/beta
propagation (the actual fix for #300) and layered main's action/
provenance plumbing on top of it, sourcing Match.bindings directly from
Token.bindings instead of re-deriving them with _bindings_for_rule(),
which is now redundant and removed. Also fixes a 2-tuple/3-tuple
unpacking break in test_matches_reasoner_match_rule caused by
Reasoner._match_rule()'s return shape changing upstream, and drops an
unrelated encoding-only .gitignore diff.

Verified: tests/reasoning/ (106 tests) and flake8 --max-line-length=88
both clean on the merged tree.
2026-08-29 21:50:04 +05:30
Mohd Kaif a5aac7e22a Merge pull request #1232 from dex0shubham/test/1167-fastapi-collection-guards
test: guard fastapi-dependent modules so collection succeeds without the explorer extra
2026-08-29 21:20:20 +05:30
dex0shubham 3448ac0689 test(utils): restore both module bindings in the progress fixture
Importing a submodule rebinds it as an attribute of its parent package,
so restoring only the sys.modules entry left
semantica.utils.progress_tracker and
sys.modules['semantica.utils.progress_tracker'] pointing at different
objects for every test that ran afterwards.

Addresses review feedback on #1254.
2026-08-29 15:00:04 +01:00
dex0shubham 70dfbf151c fix(utils): write console progress to stderr instead of stdout
ConsoleProgressDisplay wrote every progress frame to sys.stdout. Progress
is diagnostic output, so stderr is the correct stream for it — tqdm and
most progress renderers default there for the same reason — and stdout
must stay clean for programs that carry a machine-readable protocol on
it. The stdio MCP servers put newline-delimited JSON-RPC on stdout, where
an interleaved progress bar makes a response body unparseable (#1134).

ConsoleProgressDisplay now takes an optional stream, defaulting to
stderr. The stream is resolved per write rather than captured at
construction, so a later rebinding of sys.stderr (pytest capture, for
instance) is honoured. All writes and the four bare flushes route through
it, and the emoji-capability probe now inspects that stream rather than
stdout, so a cp1252 stderr still degrades correctly.

The existing cp1252 tests in tests/deduplication/test_deduplication.py
patched sys.stdout to assert emoji auto-disabling; they now patch the
stream progress is actually written to. Their intent is unchanged.

Closes #1134 (point 1 only; the SEMANTICA_KG_PATH persistence and README
items remain with @akaszubski)
2026-08-29 13:53:48 +01:00
dex0shubham 47446ebdde test(explorer): guard the deterministic-rendering e2e module on fastapi
This module landed after the branch was opened and imports
semantica.explorer.app at module scope, so it reproduced the same
collection error on a clean [dev] install.
2026-08-29 13:31:00 +01:00
dex0shubham c8a591e89e test: scope the security-regression guard to the SPARQL class, guard the new decision-route test
Addresses review feedback on #1232.
2026-08-29 13:27:42 +01:00
dex0shubham ab86127e4e test: guard fastapi-dependent modules so collection succeeds without the explorer extra
Closes #1167
2026-08-29 13:27:42 +01:00
Mohd Kaif 30592b1285 Merge pull request #1245 from HsienW/fix/sliding-window-chunker-non-termination
fix(split): validate sliding window progress
2026-08-29 15:56:29 +05:30
Mohd Kaif aceb69a5bc Merge branch 'main' into fix/sliding-window-chunker-non-termination 2026-08-29 15:44:18 +05:30
Mohd Kaif e13c953bd8 Merge pull request #1249 from semantica-agi/fix/codeql-action-pin
ci: resync github/codeql-action pin to current v4
2026-08-29 14:12:47 +05:30
yzxcj797andSameer Kadam 85d6ccd0a5 fix(memory): find_by_entity returns all matches by default (#1024)
* fix(memory): find_by_entity returns all matches by default (limit=None, not 10)

* Address review: move find_by_entity tests to the AgentMemory area

The regression tests lived in tests/test_seed_manager.py, mixing unrelated
domains. Moved to tests/context/test_agent_memory_find_by_entity.py with a
shared fixture; the unbounded default itself is unchanged and deliberate —
it IS the fix (#1018): an erasure workflow computing what references an
entity cannot paginate, so silently truncating at 10 left live references
behind. Callers that want a page pass an explicit limit.

---------

Co-authored-by: Sameer Kadam <sskadam6305@gmail.com>
2026-08-29 14:09:40 +05:30
Mohd Kaif 19ff5bf200 Merge branch 'main' into fix/codeql-action-pin 2026-08-29 13:27:43 +05:30
Sameer Kadam 4bf525d409 feat(ingest): add production-ready Salesforce ingestor (#1240)
feat(ingest): add Salesforce ingestor

Adds first-class Salesforce ingestion support, following the existing
Connector + Data + Ingestor architecture already used by the
Snowflake and Databricks integrations: SalesforceConnector /
SalesforceData / SalesforceIngestor, exposed lazily from
semantica.ingest so the base install stays unaffected.

SalesforceConnector supports both auth landscapes Salesforce actually
uses in practice: username + password + security token (SOAP login,
on-prem/sandbox), and session_id + instance_url for reusing an
existing authenticated session. Production and sandbox are selected
through domain, credentials can come from environment variables, and
the connector never intentionally puts credential material into logs,
exceptions, or its own repr.

SalesforceIngestor covers ingest_sobject(), ingest_query(),
list_sobjects(), get_sobject_schema(), and export_as_documents(),
against standard sObjects, custom objects (__c), custom metadata
objects (__mdt), platform events (__e), namespaced objects, and
relationship-field traversal (Owner.Name). Pagination follows
nextRecordsUrl/query_more() automatically and stops once a caller's
limit is satisfied rather than continuing to fetch full pages past it.

Dynamically constructed SOQL is validated before it's sent: sObject
names, field names, relationship paths, ORDER BY expressions, and
numeric limits are checked, and WHERE fragments are screened against
common injection primitives after masking quoted string literals so a
value like status = 'union' doesn't false-positive. Raw SOQL passed
directly to ingest_query() stays intentionally caller-controlled,
since that method is documented as the advanced/unvalidated escape
hatch.

Salesforce-specific attributes metadata is stripped from returned
records before they're handed to the rest of the pipeline, while
relationship data, normal field values, and datetime normalization
are preserved. export_as_documents() uses the Salesforce Id as the
stable document identifier and keeps the source record in document
metadata for provenance.

Wired into the unified ingestion API via ingest_salesforce() and
ingest(source_type="salesforce", ...), registered with
MethodRegistry under sobject/query/list_sobjects/schema/documents.
Isolated behind the semantica[db-salesforce] extra
(simple-salesforce>=1.12.0), included in db-all.

JWT Bearer authentication and Bulk API 2.0 are intentionally out of
scope for this first connector; both are documented as deliberate
follow-ups rather than gaps.

fix(ingest): address Salesforce review findings

- limit now validates as a non-negative integer before use; negative,
  string, and float values raise ValidationError instead of silently
  returning an empty result, raising a bare TypeError, or building an
  invalid LIMIT 0 query
- fields is validated as a non-empty list of strings; a bare string
  (e.g. "Id") no longer gets iterated character-by-character into
  nonsense field names, and an empty list no longer builds a
  syntactically invalid SELECT
- the generic connection-failure path now raises with `from None`
  instead of chaining the original exception, so credential or
  request detail from the underlying library can't surface through a
  traceback
- the unified ingest() dispatch no longer coerces a non-dict source
  into None and silently falling back to environment credentials; an
  invalid source now raises
- _validate_order_by rewritten to validate each dot-separated
  component through _validate_field_name, rejecting malformed
  fragments like "Name." or "Owner..Name" that the previous regex let
  through
- CI conflicts from parallel merges resolved; upstream markdown
  dependency changes preserved

test(ingest): add Salesforce JWT coverage

Adds construction and connect() coverage for the JWT Bearer auth path
(consumer_key + privatekey/privatekey_file), the one auth mode that
had no dedicated tests despite handling private key material.
Also removes _SAFE_ORDER_RE, left behind as dead code once
_validate_order_by was rewritten to use _validate_field_name per
component, and fixes a test-isolation leak where an earlier test left
SALESFORCE_AVAILABLE=True behind for a later test that expected it
False when simple-salesforce isn't installed.
2026-08-29 12:55:37 +05:00
Zohaib Hassnain 8858beb6d9 ci: resync github/codeql-action pin to current v4 2026-08-29 12:36:40 +05:00
Alex Smolya d3183d0ab3 feat(explorer): add deterministic rendering E2E example and test (#1037) (#1041)
feat(explorer): add deterministic rendering E2E example and test (#1037)

Adds a deterministic Explorer graph baseline and coverage for the full
build -> persist -> API -> frontend hydration -> canvas rendering path,
so a regression anywhere along that chain shows up in CI instead manually.

examples/explorer_deterministic_rendering_example.py builds the
canonical 4-node, 3-edge graph (Alice -WORKS_AT-> Acme, Bob -KNOWS->
Alice, Acme -LOCATED_IN-> New York) with ContextGraph.add_node()/
add_edge(), persists it with save_to_file() and reloads it with
GraphSession.from_file(), printing the setup prerequisites and the
expected node/edge/label checklist for anyone running it by hand.

tests/explorer/test_explorer_deterministic_rendering_e2e.py covers
graph construction, the serialize/deserialize round trip, GraphSession
loading, and the Explorer API's /api/graph/* responses against the
exact expected nodes, edges, and labels, plus all three auth modes
(unconfigured, API-key required, anonymous opt-in).

fix(explorer): address Qodo review findings for deterministic rendering e2e (#1037)

- configure SEMANTICA_ALLOW_ANONYMOUS=true and document
  SEMANTICA_API_KEY as the alternative in the reproduction
  instructions, so the documented commands don't 503 on a clean
  checkout
- add clean-checkout prerequisites and a visual verification
  checklist to the example
- add edge-label (WORKS_AT, KNOWS, LOCATED_IN), zoom-tier, and
  hover-interaction coverage to the frontend test
- add an explicit auth-enforcement integration test for the
  deterministic graph endpoints

fix(explorer): connect deterministic rendering E2E path

The frontend test built its own node/edge objects directly with
batchMergeNodes()/batchMergeEdges(), bypassing the real loading path
entirely -- it never went through useLoadGraph, never mounted the
canvas, and its fixture didn't even carry the same fields the backend
actually returns (e.g. no color values), so a break in API hydration,
the edge.type -> edgeType mapping, or canvas label rendering could
still pass.

Adds deterministicExplorerRendering.e2e.ts, which mounts the real
Explorer app in Chromium, serves API-shaped /api/graph/nodes and
/api/graph/edges responses through route interception, drives the
app through its actual useLoadGraph hydration path into a real Sigma
canvas, and asserts on captured canvas fillText() calls that
WORKS_AT, KNOWS, and LOCATED_IN are genuinely drawn, both after load
and after Zoom In.

fix(explorer): preserve upstream markdown dependencies
ci(explorer): isolate deterministic backend test dependencies

Wires the new Python test into ci.yml as its own focused step (it
previously only ran manually), installs Playwright's Chromium
browser before the frontend suite, and keeps the deterministic
backend test's dependency install separate from the rest of the
pipeline so it doesn't pull in unrelated optional extras during
collection.

fix(explorer): remove redundant edge label hydration

An earlier commit in this PR added an explicit `label` field to
hydrated edge attributes on the theory that it was needed for edge
labels to render. Review traced through GraphCanvas.tsx's label
resolution (`attrs.edgeType || data.label || ""`, from the earlier
#1009 fix already on main) and found that `edgeType` is set
unconditionally on every edge during hydration, so it always wins the
`||` before `data.label` is ever consulted -- the added field and its
plumbing in useLoadGraph.ts and graphStore.ts never did anything.
Removed both; reran the real Chromium E2E test against the reverted
code and confirmed all three labels still render identically, closing
out the question of whether anything else was actually broken.
2026-08-29 12:25:58 +05:00
Kevin Zhang da642f12fa fix(export): @vocab mints into the shipped ns# namespace (#1236)
fix(export): keep caller data out of the shipped ns# namespace

Every JSON-LD context set @vocab to https://semantica.dev/vocab/,
which 404s, so every bare term in caller data (extracted entity/
relationship types, arbitrary metadata keys) minted under a namespace
the package never ships. The obvious fix, pointing @vocab at
SEMANTICA_NS instead, turned out to be worse than the dead link: since
that namespace is real and populated, every bare term a caller happens
to use now expands into something that looks like official Semantica
vocabulary. An extracted type "ORG" became ns#ORG, a class the
vocabulary never defines. A metadata key "source" attached a plain
string value to sem:source, an owl:ObjectProperty that already exists
in semantica-ns.ttl with a resource-valued range, silently corrupting
its semantics.

@vocab is now removed from all five contexts (four in
json_exporter.py, one in rdf_exporter.py) rather than repointed.
Every document already used explicit semantica: prefixes for its own
terms, so nothing else in the output changes; an unscoped bare term
now simply fails to expand, which is standard JSON-LD behavior for a
context that doesn't know it, instead of being silently claimed by
our namespace.

Two call sites needed to stop handing caller data to @type/bare terms
in the first place:

- Entity nodes are always typed semantica:Entity now, with the
  caller's label carried as a semantica:type string instead of
  minted into @type. This matches how relationship nodes already
  carried their type. sem:type's domain in semantica-ns.ttl opens up
  to cover entities as well as relationships, following the
  sem:confidence precedent, since the property is now legitimately
  emitted for both.
- semantica:metadata gets an explicit @json term definition, so a
  caller's metadata dict travels as one rdf:JSON literal instead of
  having its keys expand as separate predicates. A metadata key can
  no longer collide with a real ontology term no matter what the
  caller names it.

Both JSONExporter and RDFExporter.serialize_to_jsonld got the same
treatment, since they build separate JSON-LD structures for the same
underlying data.

The regression tests assert the negative space this bug lived in: no
context declares @vocab, no caller type label appears as an rdf:type
under ns#, and no caller metadata key appears as a predicate under
ns# at all, only as content inside the single JSON literal.

Closes #1146
2026-08-28 19:53:35 +05:00
Aldrin Joseph 5376f046ca fix(explorer): dedupe temporal snapshot requests and apply latest-wins (#1241)
fix(explorer): dedupe temporal snapshot requests and apply latest-wins

The temporal snapshot effect fetched /api/temporal/snapshot with no
idempotency or ordering guards. Upstream churn (timeline recreation
while bounds settle, play ticks resetting the playhead, drag events)
could re-request the same `at` repeatedly, and with variable network
latency an older position's response could land after a newer one's,
overwriting the active-node count, so the chip visibly lagged the
scrubber.

Add a small stateful guard module (temporalSnapshotGuards.ts) built
around a per-position cache, keyed by the debounced timestamp's
primitive millisecond value rather than the Date object, so upstream
object-identity churn cannot defeat the dedup on its own:

- at most one in-flight request per scrubber position, so identical
  `at` values arriving while a request is pending are dropped instead
  of firing a fresh fetch, breaking the idle/play polling loop;
- successful snapshots are cached per position and re-applied when the
  scrubber returns to it (play wrap-around, back-scrubbing) without a
  network round trip;
- a response is applied only while the scrubber is still on the
  position it was requested for, so an out-of-order response can never
  clobber a newer position's count;
- failed, cancelled, or superseded requests release their position so
  it can be fetched again the next time it's visited, rather than
  stalling it permanently;
- reset() drops all cached and in-flight state when the underlying
  graph summary changes (reload/retry), since snapshots cached against
  the previous graph no longer describe anything real. Keyed on the
  summary query's data identity, which react-query keeps stable
  (staleTime: Infinity plus structural sharing) unless the graph data
  itself was replaced, so reset fires exactly on a real reload and not
  on cosmetic re-renders.

The snapshot effect is wired through the guards end to end: begin()
returns either a fresh sequence number to fetch under or a cached
snapshot to reapply directly; the same shouldApply()/apply() gate
handles both the network and cached-reapply paths so they can't drift
apart; finish() runs from both the fetch's failure branch and its
cleanup function, so a cancelled or failed request is always retryable
on the next visit instead of leaving its position stuck in-flight.

16 unit tests cover dedup, independent positions, revisit re-apply,
play wrap-around, failure retry, stale-sequence protection (a late
response or a late release from a superseded request cannot act on a
newer request's position), reset-on-reload, and cache-bound eviction.

Closes #1128
2026-08-28 19:45:13 +05:00
Guofang.Tang 56d9e9a857 fix(ontology): coalesce normalized property collisions (#1231)
fix(ontology): coalesce normalized property collisions

Different raw property spellings can normalize to the same ontology
name and IRI. works_for and worksFor, for example, both normalize to
worksFor, but property inference emitted a separate definition for
each spelling, so the generated ontology declared two distinct
properties under what would become the same IRI once minted. The same
collapse could also happen across kinds: a relationship type and an
entity attribute that normalize to the same name would previously
produce a data property and an object property sharing one name, with
no signal that anything was wrong.

infer_properties() now runs a coalescing pass after object and data
properties are both inferred. Properties are grouped by (kind, name).
Object properties that collide are merged in occurrence order:
domains and ranges are unioned rather than overwritten, so a property
seen across several source classes keeps every domain instead of
losing all but the first, and occurrence_count is summed across the
merged spellings so downstream confidence/frequency signals stay
correct. Data properties merge domains the same way and reconcile
differing ranges through the existing _get_more_general_type()
widening logic already used elsewhere in this file, rather than a new
implementation.

A name that resolves to both an object property and a data property
is not silently coalesced into either one, since the two kinds mean
different things in the emitted ontology. That case raises a
ValidationError up front, naming every colliding name and which kinds
collided, so the conflict surfaces before an ambiguous ontology is
written rather than after.

Verified beyond the two cases in the new test file: a data property
colliding across two different domain classes correctly unions the
domain instead of keeping only the first class, and three distinct
spellings of the same relationship type collapse into one property
with the occurrence count correctly summed across all three.

Follow-up to #1170 (relationship endpoint types) and #1171 (retained
data properties for normalized class names).
2026-08-28 16:23:46 +05:00
hsien wei dfd668c206 fix(split): validate sliding window progress
- Reject non-positive stride values before chunking and validate temporary overlap overrides before mutating chunker state.

- Restore the original overlap and custom stride with `try/finally` so state remains unchanged after both successful and failed `chunk_with_overlap()` calls.

- Add regression coverage for invalid stride and overlap values, valid boundary cases, and state restoration.
2026-08-28 05:03:45 +08:00
江俊杰 e74d0a274d Merge remote-tracking branch 'upstream/main' into fix/rete-pattern-matching
# Conflicts:
#	CHANGELOG.md
2026-08-25 10:04:54 +08:00
江俊杰 b570794515 perf(reasoning): precompile alpha node condition regex
unify_condition() rebuilt a regex (re.split + concat + re.match) for
every fact tested against every alpha node. Since RETE evaluates many
facts across many alpha nodes, this repeated construction added
significant overhead.

- Extract regex construction into _build_condition_regex() (reused by
  unify_condition and AlphaNode).
- AlphaNode.__init__ now compiles its condition once (no initial
  bindings at alpha time) into self._compiled and reuses it per fact.
- On compile failure, log a WARNING and treat the node as non-matching,
  consistent with the earlier observability fix.
- Add tests for the compiled path and the compile-failure fallback.

Refs #300
2026-08-19 10:24:24 +08:00
江俊杰 a94cec3b36 fix(reasoning): log unify_condition regex errors for observability
Previously unify_condition() silently caught re.error and returned None
with no log context, unlike Reasoner._match_pattern() which logs the
pattern/regex/fact on failure. This made malformed conditions hard to
diagnose in the RETE engine.

- Add a module-level logger ("semantica.rete_engine") for the standalone
  unify_condition() helper.
- On re.error, log a WARNING including the condition pattern, compiled
  regex, and fact string before returning None.
- Also catch unexpected exceptions (noqa BLE001) with the same context,
  mirroring Reasoner._match_pattern behaviour.
- Add tests asserting both error paths log a warning and return None.

Refs #300
2026-08-19 10:24:24 +08:00
江俊杰 f9b1295d14 fix(reasoning): implement RETE alpha/beta matching with Token model (#300)
AlphaNode._matches and BetaNode._can_join were placeholder stubs that
always returned True, so the Rete network fired every rule for every
fact. Add a regex-based unify_condition (reusing Reasoner._match_pattern's
approach) that binds ?vars via named groups and enforces repeated-variable
and cross-condition binding consistency.

Rework propagation around a Token model (facts + bindings) instead of bare
facts: AlphaNode emits single-fact tokens, and BetaNode.join merges left/
right tokens, concatenating facts in condition order and returning a merged
token only when shared variables agree. This fixes a P1 chained-join defect
where rules with three or more conditions lost bindings and accumulated
wrong facts at the third join, and a conflicting third condition could
spuriously fire. Beta nodes now keep both left/right token memories and
join each new token against every token on the opposite side.

Also fix an adjacent bug where beta nodes were never wired into their
inputs' children, blocking propagation. Adds tests/reasoning/test_rete_engine.py
including a TestThreeConditionChain suite (valid match, third-level conflict
suppression, insertion-order independence, complete in-order Match.facts,
multiple left tokens joining one right fact, parity against
Reasoner._match_rule, and reset clearing all token memory).
2026-08-19 10:24:24 +08:00
198 changed files with 37410 additions and 1653 deletions
+3
View File
@@ -18,6 +18,9 @@
.git/**
.github
.github/**
!.github/requirements/
!.github/requirements/explorer-extra-py313.txt
!.github/requirements/pep517-build.txt
.claude
.claude/**
.codex
@@ -0,0 +1,56 @@
name: 'Setup Semantica'
description: 'Install Python, cache pip, and install the semantica package into a workflow'
author: 'Semantica'
inputs:
python-version:
description: 'Python version to set up'
required: false
default: '3.11'
version:
description: 'Version constraint to append to the pip spec, e.g. "==0.6.7" or ">=0.6,<0.7". Leave empty for the latest release.'
required: false
default: ''
extras:
description: 'Comma-separated extras to install, e.g. "explorer,all"'
required: false
default: ''
cache:
description: 'Pip cache mode passed straight to actions/setup-python ("pip" to enable). Left empty (disabled) by default because this action is meant to run standalone in any caller repo, and actions/setup-python errors out if it cannot find a requirements.txt/pyproject.toml/setup.py/poetry.lock to key the cache on. Opt in only when the caller repo has one of those files.'
required: false
default: ''
outputs:
version:
description: 'The installed semantica version'
value: ${{ steps.verify.outputs.version }}
runs:
using: 'composite'
steps:
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7
with:
python-version: ${{ inputs.python-version }}
cache: ${{ inputs.cache }}
- name: Install semantica
shell: bash
env:
SEMANTICA_EXTRAS: ${{ inputs.extras }}
SEMANTICA_VERSION: ${{ inputs.version }}
run: |
python -m pip install --upgrade pip
if [ -n "$SEMANTICA_EXTRAS" ]; then
spec="semantica[$SEMANTICA_EXTRAS]$SEMANTICA_VERSION"
else
spec="semantica$SEMANTICA_VERSION"
fi
python -m pip install -- "$spec"
- name: Verify install
id: verify
shell: bash
run: |
VERSION=$(python -c "import semantica; print(semantica.__version__)")
echo "Installed semantica $VERSION"
echo "version=$VERSION" >> "$GITHUB_OUTPUT"
+23
View File
@@ -101,6 +101,29 @@ updates:
allow:
- dependency-type: "production"
# Explorer frontend (npm)
- package-ecosystem: "npm"
directory: "/explorer"
schedule:
interval: "weekly"
day: "monday"
time: "03:30" # 3:30 AM UTC (9:00 AM IST)
open-pull-requests-limit: 10
reviewers:
- "KaifAhmad1"
assignees:
- "KaifAhmad1"
commit-message:
prefix: "security"
include: "scope"
labels:
- "dependencies"
- "javascript"
- "security"
allow:
- dependency-type: "production"
- dependency-type: "development"
# Docker dependencies (if you use Docker)
- package-ecosystem: "docker"
directory: "/"
+58
View File
@@ -0,0 +1,58 @@
# CI tool requirements
Hash-pinned `pip install` targets for CI/release/Dockerfile steps that install
something other than the project's own audited `requirements-ci.txt` set.
These exist because OpenSSF Scorecard's Pinned-Dependencies check flags any
`pip install` in a workflow or Dockerfile that isn't hash-verified, and
`requirements-ci.txt` alone doesn't cover build/release/security tooling or
the project's own local-source install.
Each `.txt` was generated from the adjacent `.in` (or, for `explorer-extra-py311.txt`,
`explorer-extra-py313.txt`, and `base-deps.txt`, from `pyproject.toml` directly) with:
```
uv pip compile <input> --python-version 3.11 --python-platform linux \
--constraint requirements-ci.txt --generate-hashes -o <output>.txt
```
(`--constraint requirements-ci.txt` is omitted for `bootstrap.txt`,
`build-tools.txt`, `uv-tool.txt`, `twine.txt`, `pip-audit.txt`, and
`security-scan-tools.txt`, since those install standalone tooling with no
version relationship to the project's own dependency tree.)
Regenerate a file the same way after bumping a pinned version, and re-run it
whenever `requirements-ci.txt` changes if the file used `--constraint` (see
each file's own autogenerated header comment for its exact command).
| File | Used by | Installs |
| --- | --- | --- |
| `bootstrap.txt` | security-scan.yml, benchmark.yml | pip, setuptools (upgrade before anything else) |
| `pep517-build.txt` | ci.yml, benchmark.yml, Dockerfile | exact `[build-system] requires` from `pyproject.toml` (setuptools, wheel) - installed with `--no-build-isolation` before any `pip install -e .` / `pip install .`, since `--no-deps` alone doesn't stop pip's PEP 517 build isolation from fetching those two *unhashed* |
| `explorer-extra-py311.txt` | ci.yml | semantica's base deps + the `explorer` extra, resolved for python 3.11 |
| `explorer-extra-py313.txt` | Dockerfile | the same, resolved for python 3.13 (the image's actual interpreter) |
| `pytest-tool.txt` | ci.yml | pytest, for the pre-all-extras deterministic test |
| `uv-tool.txt` | ci.yml | uv, to verify requirements-ci.txt is current |
| `build-tools.txt` | ci.yml, release.yml | build, wheel |
| `twine.txt` | release.yml | twine |
| `pip-audit.txt` | security-scan.yml | pip-audit |
| `security-scan-tools.txt` | security-scan.yml | bandit, semgrep, jq |
| `base-deps.txt` | benchmark.yml | semantica's base deps (no extras) |
| `benchmark-extra.txt` | benchmark.yml | the benchmark-only libs (neo4j, pdfplumber, etc.) |
`explorer-extra-py31{1,3}.txt` and `base-deps.txt` are large (they mirror
most of `requirements-ci.txt`) because semantica's `dependencies` list in
`pyproject.toml` isn't extras-gated - installing the package at all pulls
the full base set. That's expected, not a mistake.
`explorer-extra-py311.txt` and `explorer-extra-py313.txt` are **not**
interchangeable, and can't be collapsed into one file compiled for either
version: `librosa`'s `audioread` dependency needs `standard-aifc` /
`standard-sunau` only under `python_version >= "3.13"` (Python 3.13 dropped
`aifc`/`sunau` from stdlib). A file resolved for 3.11 simply omits those
packages' hashes, so installing it with `--require-hashes` on a real 3.13
interpreter (the Dockerfile's base image) fails outright rather than
silently under-pinning. Any other file shared across a 3.11 and 3.13
consumer would need the same split if it hits a similar stdlib-removal
edge case - check for `ERROR: In --require-hashes mode, all requirements
must have their versions pinned` on the *other* Python version before
assuming one `--python-version` covers every consumer.
File diff suppressed because it is too large Load Diff
+14
View File
@@ -0,0 +1,14 @@
rdflib
neo4j
faiss-cpu
torch
pyarrow
pdfplumber
python-pptx
openpyxl
lxml
python-docx
beautifulsoup4
chardet
langdetect
en-core-web-sm @ https://github.com/explosion/spacy-models/releases/download/en_core_web_sm-3.8.0/en_core_web_sm-3.8.0-py3-none-any.whl
File diff suppressed because it is too large Load Diff
+2
View File
@@ -0,0 +1,2 @@
pip
setuptools
+10
View File
@@ -0,0 +1,10 @@
# This file was autogenerated by uv via the following command:
# uv pip compile .github/requirements/bootstrap.in --generate-hashes --python-version 3.11 --python-platform linux -o .github/requirements/bootstrap.txt
pip==26.2.1 \
--hash=sha256:71138adf1f4ca900cdb7d289c21b7494329f2332b6d85f0e1c42108c0384ed3e \
--hash=sha256:f6ad667e89a1fe78046c8f13232b247200f5258d7828f3f7883d660878e0813f
# via -r .github/requirements/bootstrap.in
setuptools==84.0.0 \
--hash=sha256:51a52592b3b99e102b609654876bd65f19f999935166d1352678931132b0c670 \
--hash=sha256:f4695c21257f0d9b537ec2692c941d02ee143b7cc1276941349a546573b2ef73
# via -r .github/requirements/bootstrap.in
+2
View File
@@ -0,0 +1,2 @@
build==1.6.0
wheel==0.48.0
+20
View File
@@ -0,0 +1,20 @@
# This file was autogenerated by uv via the following command:
# uv pip compile .github/requirements/build-tools.in --generate-hashes --python-version 3.11 --python-platform linux -o .github/requirements/build-tools.txt
build==1.6.0 \
--hash=sha256:bd2c8afc603e7a2e0ce70e2ea85f0a6d02043bafbd307f5bada0f98669eca5af \
--hash=sha256:f7aaf1ebbb79178a02ba248bb524f2176b256017e17e8e4bd4289c7b38cc2bad
# via -r .github/requirements/build-tools.in
packaging==26.3 \
--hash=sha256:94edc256424af38762eb31306eed28beb9f0efc50a8837492c9d6fd6004aed79 \
--hash=sha256:d7193f7c8e4e93f444fde0262bf90af30e16fa0ad0ad44cb553c87339b23cd1c
# via
# build
# wheel
pyproject-hooks==1.2.0 \
--hash=sha256:1e859bd5c40fae9448642dd871adf459e5e2084186e8d2c2a79a824c970da1f8 \
--hash=sha256:9e5c6bfa8dcc30091c74b0cf803c81fdd29d94f01992a7707bc97babb1141913
# via build
wheel==0.48.0 \
--hash=sha256:3217dcc807155e45db462d7ef2431f5ddda0d7273b700d05a67b271ceb1287ab \
--hash=sha256:94800765601e9171bf5d58d066e640662842bcedcbab982b2c90787a2c987322
# via -r .github/requirements/build-tools.in
+1
View File
@@ -0,0 +1 @@
checkov==3.3.16
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+2
View File
@@ -0,0 +1,2 @@
setuptools==84.0.0
wheel==0.48.0
+14
View File
@@ -0,0 +1,14 @@
# This file was autogenerated by uv via the following command:
# uv pip compile .github/requirements/pep517-build.in --generate-hashes --python-version 3.11 --python-platform linux -o .github/requirements/pep517-build.txt
packaging==26.3 \
--hash=sha256:94edc256424af38762eb31306eed28beb9f0efc50a8837492c9d6fd6004aed79 \
--hash=sha256:d7193f7c8e4e93f444fde0262bf90af30e16fa0ad0ad44cb553c87339b23cd1c
# via wheel
setuptools==84.0.0 \
--hash=sha256:51a52592b3b99e102b609654876bd65f19f999935166d1352678931132b0c670 \
--hash=sha256:f4695c21257f0d9b537ec2692c941d02ee143b7cc1276941349a546573b2ef73
# via -r .github/requirements/pep517-build.in
wheel==0.48.0 \
--hash=sha256:3217dcc807155e45db462d7ef2431f5ddda0d7273b700d05a67b271ceb1287ab \
--hash=sha256:94800765601e9171bf5d58d066e640662842bcedcbab982b2c90787a2c987322
# via -r .github/requirements/pep517-build.in
+1
View File
@@ -0,0 +1 @@
pip-audit==2.10.1
+423
View File
@@ -0,0 +1,423 @@
# This file was autogenerated by uv via the following command:
# uv pip compile .github/requirements/pip-audit.in --generate-hashes --python-version 3.11 --python-platform linux -o .github/requirements/pip-audit.txt
boolean-py==5.0 \
--hash=sha256:60cbc4bad079753721d32649545505362c754e121570ada4658b852a3a318d95 \
--hash=sha256:ef28a70bd43115208441b53a045d1549e2f0ec6e3d08a9d142cbc41c1938e8d9
# via license-expression
cachecontrol==0.14.4 \
--hash=sha256:b7ac014ff72ee199b5f8af1de29d60239954f223e948196fa3d84adaffc71d2b \
--hash=sha256:e6220afafa4c22a47dd0badb319f84475d79108100d04e26e8542ef7d3ab05a1
# via pip-audit
certifi==2026.7.22 \
--hash=sha256:62f22742b58a1a33014a2b6b706588a8d7e2a88ae7bd1a6ebe8c992928483775 \
--hash=sha256:741e2c3b351ddf169a738da9f2c048608ff7f2c5cc02f1ebc6b118bb090d5d55
# via requests
charset-normalizer==3.5.1 \
--hash=sha256:00668ebb0609751758682eb0b5857e7c35b9f00e84dfdef062e103244ec94d45 \
--hash=sha256:012a22b88a77ca2e59b98ac5889b0deb604147666032f45e6d6e217634d2550d \
--hash=sha256:01e93745f7f219b703b60ba7afead36cfc4242782be5af484673fc500df12da5 \
--hash=sha256:04368edf83514385ffc3e1cfd4546e595f4f1272dd23ba437a93a9cc3741d47b \
--hash=sha256:0722590aabf9dc6a6c0343d523c05458fa2b5047dbe6302fd526bb570600753f \
--hash=sha256:07ffd07412fc5d5e84cd8952acf9ff7e4ed7a708e69d1bada19d8ba91711353f \
--hash=sha256:09a7bba9f739468c8e78c36a75c33768e53cb1959fc638f510454c14683f00d5 \
--hash=sha256:0b2b1b3fa5670c127b246df1d0c059defd41f689a868a3b9d79df9b1cac42d22 \
--hash=sha256:0c6dfb5ca6723eeed15aa8e564a014d69fcb8812f94eef11fe3631e0508199f5 \
--hash=sha256:0d929fc574b4d6fd9e7c0f5c2ede8716a41911923aa7fa5fce38e0818aa4a1ac \
--hash=sha256:13e3afe97712e8887cd516e960c63f0b93122971e5b5e4b2622fe7701771e838 \
--hash=sha256:15f024313246a4ed976c60f440bb8d257815513a681d212ff74fd46f7d715a90 \
--hash=sha256:195ce897c6153c0700078142cf8efe3e6454ca4cf4357499e4078dfd83396626 \
--hash=sha256:19a3dd5aa73cef1c99687c4fc57db016a9c17104ae1185da88ba566a5d3bebe4 \
--hash=sha256:1d1c7a53a6c2103925cdd6d7229f8c567379f211c869793df679f2e9f738c369 \
--hash=sha256:1f5883d77fd409a261abb5dc8ccbe335720d798b1de4abb3b1d47ccbbc76b53b \
--hash=sha256:21b82d8082f6f5e7f456ef0bd16323d08de1266efbfeb476e64b2a91d1471a4e \
--hash=sha256:252d099029bcbea642f2a06c4ed5046bdf8b5a8150b64afa5e027e88b106e5ee \
--hash=sha256:256dd4d85d9e4dc595e2bc983c980e73f62ddeb3165c58b4c3dfe78c5c8548c1 \
--hash=sha256:26422d45fd13551cf564c58932f7d72b4f58b93b0fcf18c35ba6be12b46bb102 \
--hash=sha256:2679de311c7946dde5d3b6f44941844133ff5c7cb86099c0061ab1e8901c20a8 \
--hash=sha256:29880d17a8eb0b5cfdfd8944b468322928059aa35f1f5fa8ff22b149ec0b42f8 \
--hash=sha256:2bced4061f000f7187254a02ad3433ae17eaf991747ceea2f478422590a5bba9 \
--hash=sha256:2e9cf9253119d8e5d111f05d71626786fd3d6193817316eab1ca088cdb8593cf \
--hash=sha256:2f06b7eae9dbe77fe1d644ca244dad508de8d302870a43f3c559b521270938a0 \
--hash=sha256:2f293479cce755c75f1697e87c409b7ae4c555c7dfecb6e988ad13abba943031 \
--hash=sha256:329fc3ccb63ad22d867d84c2adea759a64079a37ba4a343433b02c7a2816871e \
--hash=sha256:343fb4f2821043bd87095f7b08a1a181febc8e36ac64212143bbfd0a0e1bc235 \
--hash=sha256:3588e376b3ea2eea84976f67273d679f229e24c66dce7b82ae45aef04ff6e072 \
--hash=sha256:35aea775dc2bd5f54cd84a1cd2696cc3207c479cb9cf0bd346f0d343e4300ddb \
--hash=sha256:35fe081843b35aad20ffeccec3eeffbe637b15d14f3fb22cc1b59cd8ec17e93c \
--hash=sha256:36047af20e17097c3bb9476c2b7655f2f7aa51322c0ba58c07695bedf755a950 \
--hash=sha256:3617ac3cfd8b9888f145ad89dd6e692285834b0201c6074a5eeaad3fd4d668c2 \
--hash=sha256:366ec70f5547c640d3ce1985722490f23faf4eb5216a7eeba78277490e78dacb \
--hash=sha256:394fea06235c8543390050ed5f529187074b029fb027213f6c46ac11ab5d950e \
--hash=sha256:3d27167433c0d5f18dc850f07d0b3816221984fecdc405d6c157a6f0b8f8e9e6 \
--hash=sha256:3e5e1224c0a6a90e05843e07adfec669edebec17801c67072f51e59561d63c0b \
--hash=sha256:41876ee62a3dddf48ff1121ad8f0798032aa03f2fd35f21f34a4cab14f18d8d2 \
--hash=sha256:433c5a81eade63b47e522303bad236f59dba55ea6951746f5558355eeed8c75d \
--hash=sha256:4582c27e8c889d64811987b5967fbd3ae0c823fe1fd933b543d55ac20bb475fa \
--hash=sha256:485a0d363cafefcd2538a73c7c838daa2035f09b2c9f9b5e3133f80c6aeb84c2 \
--hash=sha256:494b70049a4d69aec6e8137c13af4cf8db8c9f9820a1392ac293b0dd2987a818 \
--hash=sha256:496846868fea80e479324862fa877f02411f2fd0f83b79ccee2607aa68b2a032 \
--hash=sha256:4abdc5f9ad448c1ecbfae2974b820535d6bc6e7eef63babbab3d81cf46968c71 \
--hash=sha256:4b599739b93b2cbeded49645ae3c8d1405c29ddfbceac1545c87a3f9580a9e96 \
--hash=sha256:4bea7f8ebe90bbd7f0e4a2de42ca6924ba23e3e76418c408ff82f1d46fabd687 \
--hash=sha256:4c4fb141a727957c93edfe5c32a26ceb6b5f6461d67146e2d39f51e16170bea8 \
--hash=sha256:4c9548dc78002099910abaebc0a72ac58b7d30931869e0351c09b507dff4ece3 \
--hash=sha256:4d26f14f041e83dd8edfd61f4cd4fa7285d31798b5bf1f28e70c367ba6c41d61 \
--hash=sha256:4f298bdadb8f0b9e5672877f647d1be9373ef5320c9e2f049795e26cad28b6a9 \
--hash=sha256:52ec005752a56ae79547a05c0139ca2501a0c866390b6115008456b9f0e7cde1 \
--hash=sha256:55261ac0d2941c42f196dd576f543d87a8ee03cd6f5e30dfb4d807b2e3b9121a \
--hash=sha256:56490c595a28b1bb27dfc583e816152a9767721ef58b2c03b13f954d2f707420 \
--hash=sha256:58d3e12c88e0950bca850ae1f7c256055c097639c2edb9eb123af9807d8b15e4 \
--hash=sha256:58d4aa13a59c969dbfdf9e6a9560e242cbfd9e8a8f50c2747714df1a423adf65 \
--hash=sha256:59171c6e45bf07d0d5cab3b0bf81d945035530f6873398b3b531c31184d46663 \
--hash=sha256:5b6d1386bf0096d26d3a863dc0a487a5b4eb9aa93cf5ba69683d29dde6b9d60f \
--hash=sha256:5c0ea61a470e070686aa30892fed79e297d2c8d0ab46b8bcdf027d38c51da591 \
--hash=sha256:5c84bec0ab5ae0c64bfe73a7d2adcb5ce73b467523fc27fd6a28ab2aa6cbe35a \
--hash=sha256:5ca0555312ae2fe82715cada7fac375530c2f3349e1eaa1bcb33d0283ac79a18 \
--hash=sha256:5d8531a6569d025f68e2321e7638fb7978f23db58e5f69f56913837aae03816e \
--hash=sha256:5e2d0e146dcb57034f8b97dc58d2d512cb90aba253960ce449f695fec6a82c6f \
--hash=sha256:5fc45d653ea8c9a20479167e11d4a0f8cb2fa3470737ab6f9c827532313187b7 \
--hash=sha256:6117b84ea48435e5356dc737f5121485c30920ba43375fa7b434fd753df0eac3 \
--hash=sha256:6199d5606e2bbf2b096cf64d03f8b6790c91081d5ac866b8e7bb6422738cc60c \
--hash=sha256:62b55f6722735a6c472f88361cde6640608773d9443cebdbb51abf436a1fcdd3 \
--hash=sha256:687c9ca3035544b113bea2055e180af96fb63c0c476e22a9180f51925186e7b7 \
--hash=sha256:6b7430cf5728e68f6c462254009a6ef4086e1bea43cf2f57aa9c55fb4f50ff96 \
--hash=sha256:6ba32c4d2abf1d2fe7cf27d280f4cca5664233b0f885549c7761719eb977f486 \
--hash=sha256:6c9cdde8becb25a7fde49924511aa2644d6f8081cc8df8e9452724303348d8e3 \
--hash=sha256:6df0ec430f9a831772c23ca5a224cba36517a58a84bb32c32bb59a9fa67c47f6 \
--hash=sha256:6e2912d4babbc65196ac13c2f53468dc57fb8b9c25ef913e8c59ddf7c6dc0e1b \
--hash=sha256:6e5e4d73d588ca5ed09df1b7dcd1b203d1df3c542e3f50d126c947d432b10731 \
--hash=sha256:70055ff39b97c99e7ae40ea3e393fb62aa2e44dbd9b29f8d14f42fb0025c3959 \
--hash=sha256:706bfd38730a5ac7a365793269a00f4e988178cec121391f4248d84ad8c972e9 \
--hash=sha256:7235dc28fc6dd9d832ac7c7bce95367dedb85929f17368a0c2bee1e080b9acbf \
--hash=sha256:774d157f112367ff4abd29019f38f023c24e00e56edc7829c20e358a5a913ad8 \
--hash=sha256:77efcff2b23071c349402ac1066667a3d011f62398d81408c9b88ad991747c9e \
--hash=sha256:789b8982559ae28dad2356519f841655756cdcd96616410590ae0b17454ee64f \
--hash=sha256:7ac76cf9afd34929d76eb7fcb63be476a4853d8a96f0dcf2d0db68a0cbdf9885 \
--hash=sha256:7c0c10730342b0c9b35dd1d619beb8214e520bd96a1f870f452680b238aab3e0 \
--hash=sha256:823f82903d189af463d7df250ef1f7f696f3cee08cc8d91deb565e8d425f6506 \
--hash=sha256:838648accb3a7fd9803fd45c87bce8509648eb0c11bc34e216141300977244f2 \
--hash=sha256:854066be00447fa8de2ccbbe893e2ffc4b123ef16d897af794c1e18bd4a714b0 \
--hash=sha256:85d5855daafc240cc045c026d7a15fd198a09b0fc8ff6f5ecbb5297b509cb11e \
--hash=sha256:85de3134b5379856e323ba37c19c9256d39425f7b76a63af52b09fb4664c2e8f \
--hash=sha256:87e4f41d375c0b9be2fb5251aee4b8a689169e134535aed81bf085c3b647451e \
--hash=sha256:88ca277405c2d3b71c4e1c2ee0e7966e807bcba86a69d11e19ba199d18ae4491 \
--hash=sha256:88e85ab89cb822c1e635f51d6d32e488f94e002e70e2f492bdb8b945543f345a \
--hash=sha256:8ac8c94b6539074e0f40899301273ac8402b9b3e01c7b7ba269ff30340aaaf20 \
--hash=sha256:8fe532b3c966d1fb794e0698e4589d0444017ae77fc0b31edea13c0e35bcc449 \
--hash=sha256:9085f87b0e38a2b92b8923059b4e8789fe40d9279712d15dcc670048d77079af \
--hash=sha256:90b7481fb62fbe172c558bc6fd1c4c98d82004a54a7551f20e11ac9bf0b8708c \
--hash=sha256:92caef967d287a407085d61176fce4012b1dd62daed4eb6d5ceb26d3d2538712 \
--hash=sha256:9362dd90aa7dab48c0054a21187791ccf05473f7dba5d92b8033ae62164675e7 \
--hash=sha256:94d78ecec2605a8d0398b0f365d5f12a63248438516f5dac536a5eff7337df4a \
--hash=sha256:94fbf1c0c6cc0d3d5e50f9a9313a8cdca90dd696d34b381cd1704f8c9e939f20 \
--hash=sha256:950f23cb393f85543777b0433f082cddd25b51ab398eac7971146495679efe5f \
--hash=sha256:96eefc178f8636b9c760c5829345307fd81cfae9ab1e80997dbddeb0f54ee9a3 \
--hash=sha256:96fef3e886d6a9874b14f27fc193fbdc69d5d8035783d86aa4e1cea594e695f9 \
--hash=sha256:977cdbd483a9cff38179bea4fd754289a6f2195c7abd414aba85410b3e66cc5e \
--hash=sha256:978eab16f55b4ab2c2a745be9a0a840bf8f09a7f227d9c76eb30214d078865a5 \
--hash=sha256:994e883d17c559cdfd38c84003c8b27d25424a1077272a17e7cd27bfe0bf57b2 \
--hash=sha256:9ac4444d8d4fd4c4bd08bf451ed3167aa9e7ec6cdb41b648794f1d1103652e36 \
--hash=sha256:9b5db6052055d34d41230fb78d7c439c23dc536a9896f6cb039e8dd92cfc1263 \
--hash=sha256:9d9a0dc7cbe9bec24c3f767c9122c41fe5a1bc43f47cd099d00d393e09769de4 \
--hash=sha256:9dbdd9205662134957cf0c324f639bdc5031c0ca056e2369e238db75187c0f11 \
--hash=sha256:9eea3ab2597a5e65fe65296e2d6a84570845a6b55532d90333d740d48bbc850a \
--hash=sha256:a2028475ba855475b8b4d3cfeb4994269c967aea8b9892dfba907f4263a863a3 \
--hash=sha256:a3a370082ce34d0612f421e15fe011c53bb1feff21a26d06ad4fb244dab5a375 \
--hash=sha256:a545775cfe815855ea32d7c27731d79da358ef2055b4a25830231b1622dd18aa \
--hash=sha256:a5cbd90ecf0fc62e64726917ad083b73001f0563657a87ec3c0b504e277dc90d \
--hash=sha256:a6d095662e73e74f0a49988e0593373e243e3a52e27bfeea0a859e88acf4a0f5 \
--hash=sha256:a6dac12ff6b846103483683f60c5f8fee205121adc58ffd87e90a90a3af69e99 \
--hash=sha256:a951ad59cad9145664a730d3036b40b844e74d2d3683da40111463cd3a83845d \
--hash=sha256:aa1099b956fb795e686d073568f6dc002a0bb89765ea6d5b055dd7d9bf1b116c \
--hash=sha256:aa2bb0b37202dca27175591f761108b5d34096ade1191ffe4808bdf6b1571488 \
--hash=sha256:aae2ee51122d3ae968a3837d97dc24a0aeebb0dea23694422cd172bd30017cd6 \
--hash=sha256:ab743e9bc90c1f73552ec33e10e3331315acd2c397b36065b591b0181de533cc \
--hash=sha256:ac00177c4831ffa650f8609e4bdddd5fe09c03b1c0c47acece7e6ea20421598b \
--hash=sha256:ac13b004224fb341e1e25a1ed5e19d32f57cdb2a403e01f003b46f051a550f6f \
--hash=sha256:acaf604462bf330b0d07e7a07c1d6e4adac79e5fb13e9c5140590542cafacc00 \
--hash=sha256:ae31a1a1db2ee6cc2942fccaf695c934bc7f3db9f2133a3fef1f367cf1a4ab10 \
--hash=sha256:ae4a097991662cd4fff0ddc74e0fe7874f82e00042fa0ea00855645ed0c79598 \
--hash=sha256:aea996a6aba25260827c9ea511d1addfde2da9eb686ac961838509086188b7e6 \
--hash=sha256:b39b69b347e5e47a3b5b8cfc005c68c1ba347474e3960236c4944a8ecd174962 \
--hash=sha256:b54e7e13267d49ffbfe68e25b3cbd774dab38fa37238f71265e91b36146eb21c \
--hash=sha256:b9af956078716df40d985fb0dfeb2c2120c5ca92ba4ff4b388acfd01cdc14d08 \
--hash=sha256:ba2f37ee79e6338845261a3c5b1784e5d1acdff2c0785b284f1b633033d136ab \
--hash=sha256:ba501e667c17d8411f98e67a022d9604ef179aff0e459b7e292c796837c13573 \
--hash=sha256:baf3775a2635e5a11fbd5e4e64ee69c7e86875d224a5c72aca4c141064589a90 \
--hash=sha256:bb57753e36e4855b8ca375069482250a6246372331a3e4f3407eaebb007443f5 \
--hash=sha256:bd6c173f04743d483881bffa1478d5a4624475b8cd1d2194956a75548e191c18 \
--hash=sha256:be47f99644b208bff7766314013f9acf57b056b04191d570d68ad14022cf5b1d \
--hash=sha256:c010f5581d9c612804cc59fcf7b524b707fbcb72828551237ab545bb5c7034af \
--hash=sha256:c1dcc36dcb96abc02236e182d17e0f71430152a6c2c7447421da2d2dc144edea \
--hash=sha256:c428c6c31eb5f4277d7f8eccaf767fbd548ddd5ce3c8b4f4cbbfab3d96b5904c \
--hash=sha256:c658c50ac0c98cd755a2dd50b7977d3bca7df401dcc47fbdfa87db53ef7d4e8b \
--hash=sha256:c71fb0d56c920c269cd3e2e3fe7c610e3f1fdb21a6ce60efa6430ff63676cea6 \
--hash=sha256:c7b742bf31c88566b4bb6335a7f393bb322e580b6bb98df7bd0c25e6e3519ce8 \
--hash=sha256:cc0329df4caaceb950d2f580b5ac716a377f7059624a0bafaeaf8a218c6ed774 \
--hash=sha256:cc5d36d96478aa9c60654bd932525bf32964c62a7281eafdf16d85003a8d6004 \
--hash=sha256:ce854f5f478050ade5a238731c4ca985a7d3b3cb53ff600a9b5c3b689b5f0a7a \
--hash=sha256:ced3fdd71aaa83ce593746c2edb42b7a59cb4c19c8b5c407781c72e493aae55a \
--hash=sha256:cee5dd7c6fb5dd52a0fe2a740f9bc6e3593f5f8b1788bde49de02086f30182b2 \
--hash=sha256:cfa1c0cc3a8f9f53f1243a5a99ac36fd003880199383b37672e86ddda9cb07e2 \
--hash=sha256:d1ee1e296209fdce05b81b663250eefa02213a2da7b41bf26f7829b8ba3545aa \
--hash=sha256:d59b75732e9b6f27388e10c14b0259cc5f2e48c78627d185e6a177b58ad3cffe \
--hash=sha256:d63600d620ad0064c3a748b950ac5ea38a80190e5498532efefa4b7b3f1da1f3 \
--hash=sha256:dd732602a7009217f658d5863d12d79d373a4de0eebc111094bcdd3bb8e0a6cc \
--hash=sha256:e06efa066f7dbadbc84ebc126a97c452a6451dfcf589d89d788484949e1cf795 \
--hash=sha256:e199fb99720074809a7720f1c0b4d919eea8b87e88713e0f8f602f7bef543d9d \
--hash=sha256:e4b018dc5a0eee4676e38fe84a47a427816c590b93b55d9025274ec4d6ffc2dc \
--hash=sha256:e6621fb2a4988d6e53eedc455e5903e2679f3967b8acb3d639f1b63c14a2e893 \
--hash=sha256:e71c909f353863b2b89c83de2ebed71ea6d0df8a6ef65a128193c5e650766bef \
--hash=sha256:e90251c0c7bdd54a100a0dce3c07b7e637278c93af29dbf78ebb89a58c4bac7d \
--hash=sha256:e9fbdce1e47394b09bc9f26ab117dfc8d6491977a11d86f592bb42c779db2fda \
--hash=sha256:eb12fb2ba69ffa05f8695f61c69e591dc4b4a12ac3757ac8af8adb259bf56d17 \
--hash=sha256:eda059b6bc8bc0812d626fd91a7ce01bf583df0a61296eff390fd94141a34e30 \
--hash=sha256:f03ac127268b43ef4fe9e6ab6794a6794b49485a0cc0c1db79876d2f33f75bc7 \
--hash=sha256:f298e218441525d3794428b4c8b8fb8662c6d3ea79925d4807ee6b9a96a3bca5 \
--hash=sha256:f5542f9b941279d82d41eb0aa9f98eba36fe4df5c7086c651df7944935b37182 \
--hash=sha256:f6f7deae3feb4edfa2efaf7c574fe88cbf055038a6abdb40188e4fff66d5699f \
--hash=sha256:f9b1e28d0e8dbfa858abdba91d6b547beaf2df1a59bec6da6faae7b96a4991a9 \
--hash=sha256:f9f8405c2c758532c74fed975dbee57be1f31a6e865c031870c79a6ed3212ada \
--hash=sha256:fa48b1b63d639f9483e0633e092f5851e2348c352f1f9bb6c8182f87884ef876 \
--hash=sha256:fb78f6e7fcd8ad785d28cd577168bc1aaee827b25bb8755638f694794ea98f0a \
--hash=sha256:fbc597639158fd7c14d55e808718848319540f51b0e6746e3eefa59723a4a348 \
--hash=sha256:fce8cbd4997efeb450bd298b54f755dcdff18d496f7a5ddbb4867c6d7c88fdc3 \
--hash=sha256:fd0350afdc3aabd5576f60ea109228bd5538139713c7b094c5cd27c73a98bc6f \
--hash=sha256:fd0a274c0e5f9a21565cd9d3dd749b61f96b7aa1e20a93aa1ba4029518f2e5c0 \
--hash=sha256:fdb8a068947befafba9952162645dc2fecaeb400e64584829ed5e9b2fbe21a7f
# via requests
cyclonedx-python-lib==11.12.0 \
--hash=sha256:0e807521a921a5c3cb8ce1153f8a61d29eedfe76a46aac2796b7c6b573391a54 \
--hash=sha256:16767c4039de90c04e9f03348f8f0ed4b8ff842eaa7eefcad3a95685f970dacf
# via pip-audit
defusedxml==0.7.1 \
--hash=sha256:1bb3032db185915b62d7c6209c5a8792be6a32ab2fedacc84e01b52c51aa3e69 \
--hash=sha256:a352e7e428770286cc899e2542b6cdaedb2b4953ff269a210103ec58f6198a61
# via py-serializable
filelock==3.32.4 \
--hash=sha256:22e58ca3b1ae3b98993b762d7338367ae64fe50252bf78d59da3bfebcdf1cedd \
--hash=sha256:2bde2e4cf732e0153406d8a7bc80620ecf5e621fe0d25e41143c4e3b4733ff30
# via cachecontrol
idna==3.19 \
--hash=sha256:5e0811a4383b21dc5838069f801c4fb62113b7447663d2530d2bd6e77b49bf15 \
--hash=sha256:815e7be7a7806d54abb586dc943addc79e8b2ee16915059658cbeff4b1b43bf4
# via requests
license-expression==30.4.4 \
--hash=sha256:421788fdcadb41f049d2dc934ce666626265aeccefddd25e162a26f23bcbf8a4 \
--hash=sha256:73448f0aacd8d0808895bdc4b2c8e01a8d67646e4188f887375398c761f340fd
# via cyclonedx-python-lib
markdown-it-py==4.2.0 \
--hash=sha256:04a21681d6fbb623de53f6f364d352309d4094dd4194040a10fd51833e418d49 \
--hash=sha256:9f7ebbcd14fe59494226453aed97c1070d83f8d24b6fc3a3bcf9a38092641c4a
# via rich
mdurl==0.1.2 \
--hash=sha256:84008a41e51615a49fc9966191ff91509e3c40b939176e643fd50a5c2196b8f8 \
--hash=sha256:bb413d29f5eea38f31dd4754dd7377d4465116fb207585f97bf925588687c1ba
# via markdown-it-py
msgpack==1.2.2 \
--hash=sha256:06d95f61de7afe4f4ff908a6feebfcb070d0582ac87c9cf3cedf8551cf634516 \
--hash=sha256:0708afbf6a9587f0bfe479a9825c141d14d91e2f6a5c8103cf28bc96f4edb5d9 \
--hash=sha256:0883a1578168929fd1640fbbc4614773f1a130e419a8a817dc2918d9af1b651c \
--hash=sha256:0a652ceeededf71d3fa40c303a02a149d42338d310162367b91c539d4bd6e0a3 \
--hash=sha256:0dd9173c5ebaf5ecc5ca86e7ae1db92934e1d57b856f3dd90698941431f4fd77 \
--hash=sha256:0e3315de5a4b2920ccef48d96b4448025e064a10d0f5a250f6584477d839c8d4 \
--hash=sha256:0e91332144f69bc3018c91232fac26da580ef748fb8eaddd7914d4458001cc4f \
--hash=sha256:0fbc1bed8a535389b41882cfae66376e248cd1680eaa94fd83193c73e1d24986 \
--hash=sha256:11e8c421e117d1c36728b423d0402555cccbf0c6f53e288f0e75b6b12100d70f \
--hash=sha256:1510f24612d4b983dff6935d9273e02c320cfd525727fbcb58836a75f589fdbc \
--hash=sha256:1814f92306ae7862908e9ece7cfd90e0dc87ded3e89b6ae7ffdd1175d6376fdc \
--hash=sha256:1e8cdd1f3e7cc52c751092a9bf740e81e6919ab109cd376ae2d965dad0bbae34 \
--hash=sha256:1f3af0baafd184436501004828bb3df64eeb2fc49dfe9d89abcf604956094563 \
--hash=sha256:1f6b6f8deb07d49090e1808c6ef9cb7d23ca17bef3aa6ed3e5e03df16606e60c \
--hash=sha256:226a62ffe99fe54c5c61d910ec64c3449b7766c3280bd286bf6c94838dde239a \
--hash=sha256:29cc2d5291711a52956a79a51f41c732329df39ad727c886bd8f0b5b9237a808 \
--hash=sha256:336525cc2688e43ea77dfb1a4ce012c8cde561835913801dbfcfdcf4111d8abb \
--hash=sha256:34e83e345194a2a51d8bd447dea9de2104f91e75b247f4735f14f04529f0746b \
--hash=sha256:352ed831042549cca8be23780e1fe7c9177e65ff02bf183509c4b4d33f671782 \
--hash=sha256:3e915d390d7068b257ca8b62f3fc59fad135c8631d1017ab03b0b924b07c5367 \
--hash=sha256:419a45c67a5c04213172a14b1864657e014665b77d7081b107a51707923dd39e \
--hash=sha256:42fd9260416885b4815caca5bdd14dfd5dda6cdade732d6c09104ef8f6228761 \
--hash=sha256:46ec851571d8f1b6e29794ebb9dd36f785008da6d14f57c702e60781d6caf648 \
--hash=sha256:4710d881d8fb047deed2485707409116722af2b992d3fefd73c7667c4e350839 \
--hash=sha256:4955accbd87f27beebef5f3ecc27503aa74cb016fb4f640868e749fd93194a35 \
--hash=sha256:4a4348705be86e029d04e741cf9ed0dfe03e942d7d3b92e838fa80d3aa2c3ebc \
--hash=sha256:4b554d8164ebb526892194f71dcd96ef1fefe0c250087498785d3ffc04a80be3 \
--hash=sha256:4d9a562aec0a92fe536da2e533d313b3d2a6b929157b1dec7ff623446dc0a8ab \
--hash=sha256:51dd39d23cfdea0400ed3ff2d29d1e83bd951d3aea79dc89be5b701a09edfe23 \
--hash=sha256:53679573c75cce5f82359e0bd4e6a97809a6b9a9b7a48fd1ba592f4a82cddc84 \
--hash=sha256:55faa6f8395e23b848c535ad5dcb96b3462f37f5e7f4ac500d500434f7345da7 \
--hash=sha256:58ce37a4a54577115922385d37201d9a44d66d0167dfbbf4770a2e9bf8ea7ba3 \
--hash=sha256:59d5b93efa45fd09f620d0c9ba81cde339a2c9937af3eea42ee9653094ce6640 \
--hash=sha256:6195257a107bf25872ef84aab7295078271eea3ac6413f0506b631f6c9586ed5 \
--hash=sha256:652d1bf13d01bac8fd569def0fe76745e55bcda01e30aa6332d5947ea3788839 \
--hash=sha256:682804bf31e43d46e51a9a33bd575b51e839d715ce6bd5612c055f7b28ad637b \
--hash=sha256:68df2947921d449f6dcfeafd86cb2cdde13327a8b447534bbe4ee5aaf32a5695 \
--hash=sha256:6f53285f20d592ed309ee19e509cc4c77a3bda1db02ad67e8a0949bb227a5a6d \
--hash=sha256:73b0e05c32c3cfc3cd84994908e57430c0ebc6813abf905d3f18ff115d54df3f \
--hash=sha256:77c2e018417dc1d66f235e383877ee885b60ade9d29e494dd581e08af2cb1923 \
--hash=sha256:7826f16edc763e768404f55605ef85dfcf5857e729c1ed29e0d7c180be4fe6d8 \
--hash=sha256:7afa5431f6f3487c584187ca6c8e2a34e9b106529893b3e720eabb068f6ac970 \
--hash=sha256:7d095df2627e5dd59ac7b0c5ad627a671c76e6020171e03cbe4621a61f0562c3 \
--hash=sha256:7fe374ba76eb0ecca13a1703daa8fa85825a6ddddbb52d4c1a732fa524194683 \
--hash=sha256:82b1bdf293267afaadcc608b125e7fc6576bb0785a60c4fa7d07c7ab76ed76ec \
--hash=sha256:86f173a584f72f6164801f31866d22a581f60c991572cf922aed9ab8eb422b77 \
--hash=sha256:8b1415d02e9bf722672af8a90f90813265a0cd0b14163187261e54a5592bc949 \
--hash=sha256:8b2a281b556f120a43e591ea39915741b7ad54d4727b9c4350a0a11692252533 \
--hash=sha256:8c6321a414f8b4a8dc43976b2fa8349156434ca9adedd9a187b796f7e1d3d3fc \
--hash=sha256:8dc4487097571f7311188c3eca2a3e86cd1f1db4c37c7a017bcc3fd38486cbfe \
--hash=sha256:90986cc9aab9d7d1d8f38bcbf65d3f7ac83bdd90c35765db7d691b4829698cba \
--hash=sha256:9352e6cdb510a7b1a5d3ccaccec730e82e50cf3484a3af7bdaab19e23b9589ff \
--hash=sha256:935b1cfad9b908b0fa845010f4271df4c2f04e1cd26e3f18acd61a45f93c9e36 \
--hash=sha256:9b659d77f8726fa5e7038967dda6b68d53cf34472c094cfa5b845454713b90d5 \
--hash=sha256:9bd3d1557c3fe1a095068210708a03e3e4795973392af6f4047060e70abd9a6c \
--hash=sha256:9bf452ff4d4981f25a18e9476e002bcc9263e7928024aa4d7148e25f7be3f929 \
--hash=sha256:9d7fb25b4442fae0cb2590272d06ab4f6caa526ee36a994edb81e946b874813e \
--hash=sha256:9db1ba1c1e6a84245a9dd866265b56b8a1e9461549cc72ed296d8cbfbd32961b \
--hash=sha256:9eb0b0e602064527a045ea28c4f174ed69383587e29cebe28947e3b84106eb2a \
--hash=sha256:9fd7f32e2f0fb334e7ecc5adb5cf0458785bd3a9d9d86f950e1715f101cebce5 \
--hash=sha256:a378e12ccc06d76efde115caf4073b7e5ff3cc18291d1341f9e65fb882e3f754 \
--hash=sha256:a4161eee7799863aee237c35c90427861f7b994416dd81ae829f560b0a81bdcd \
--hash=sha256:a9b4cf3685a135666d27d0d7a73fece74e2fad01d9b508fded89e843512f0e90 \
--hash=sha256:aa1120c653b76d8eafa50423b5eba06b5c9737f8692c74fa3afe03e84b8978ea \
--hash=sha256:b07c03f0da7e5279170df7745ddc732d526c8a198208936ec1a95c11ed2b2d5f \
--hash=sha256:b13b59e66f107cca1ba708dd5307179870ca1b15b19fcee7ccf722e5308d9212 \
--hash=sha256:b542ffc0a5c531eedc40419f291f1bd659aa8d4223408a5b51c88a2796083fd3 \
--hash=sha256:b5c696ae7cd7166b3657261adb855b461ff31f07823fdbae9de8bf80adfccc21 \
--hash=sha256:b68614fba0570349833b7dd999ff0aed4e5cc8d9eb6e3a7d4527be33c65e33d3 \
--hash=sha256:b8dd6c71d20c28d2d0eb0c51e7cccf3584afde3b1364f6629596186c9025bd54 \
--hash=sha256:b9b0c1f2aa7b0026b4bd50718100e8b04175e4f36e160aa852502377b5e572e7 \
--hash=sha256:c522420d78db2431887d45b518e304d86e27b9ad0b30f24e3806a6ad5d8bdbfc \
--hash=sha256:ccfd880988f8438d1c91c77d7edc58e70f4d2012e999167bc154c64c6f06ea6b \
--hash=sha256:cdb6cc6e1127d15879c47a8b3270716243da82d3e7feab1f5946872c75b3d60f \
--hash=sha256:cf66fb38703e61a486b01b56d43bb1f50698fbe99b6bd90feba10f24fab60b3b \
--hash=sha256:d13d07efbf655f9ae7a2352b630c52727b359005b21ba08a507585c9ac8c0896 \
--hash=sha256:d242f3c4ccf55b056e6cf901720dccde58f1df117898f2bbf3bcd6e38ec7c248 \
--hash=sha256:d24b38a825bcca41bb956de50eb98451ef291304a8607fad99e619043d3e79b9 \
--hash=sha256:d3c247d457ae9079974c7ce3c665396754a6d2baff7eaa51332212a8a5a3f13b \
--hash=sha256:d886baa46b2532135e7320067e6a44edb09ba5883a6096b0f9c044533984b8a8 \
--hash=sha256:e05a94a0442de86818a30281c6cc2cb9cc7aa148386fd3541c4d4774b73cb3a9 \
--hash=sha256:e1b99ad34613d5f8477fa5cf99bc4eaeaf27965588007c102370cd9a78fe9de5 \
--hash=sha256:e2eb7ea0ac3911a7aac9d8aaa36d40f216d99455b3274cd3fac38181bcd910cf \
--hash=sha256:e497ee34e8a3342bbde51b27c22d8db05a651df3361dd3daef5b3ab0d66f3e04 \
--hash=sha256:f11e09f10210a91c169e39c7a5a1f9090eaa73ad75555fafad5023c3053c47ba \
--hash=sha256:f466049b8e1ec0854287bbe9a074316826fe0e08dcf707245f98b1ae49e92650 \
--hash=sha256:f80361592c13d7226b4379c8941529b63fe1a9d0e05d2de8f3306b70e522b53f \
--hash=sha256:ffdd2f4950daf7815490f23087963e3420175b9609520b7ff5df64d351159c22
# via cachecontrol
packageurl-python==0.17.6 \
--hash=sha256:1252ce3a102372ca6f86eb968e16f9014c4ba511c5c37d95a7f023e2ca6e5c25 \
--hash=sha256:31a85c2717bc41dd818f3c62908685ff9eebcb68588213745b14a6ee9e7df7c9
# via cyclonedx-python-lib
packaging==26.3 \
--hash=sha256:94edc256424af38762eb31306eed28beb9f0efc50a8837492c9d6fd6004aed79 \
--hash=sha256:d7193f7c8e4e93f444fde0262bf90af30e16fa0ad0ad44cb553c87339b23cd1c
# via
# pip-audit
# pip-requirements-parser
pip==26.2.1 \
--hash=sha256:71138adf1f4ca900cdb7d289c21b7494329f2332b6d85f0e1c42108c0384ed3e \
--hash=sha256:f6ad667e89a1fe78046c8f13232b247200f5258d7828f3f7883d660878e0813f
# via pip-api
pip-api==0.0.34 \
--hash=sha256:8b2d7d7c37f2447373aa2cf8b1f60a2f2b27a84e1e9e0294a3f6ef10eb3ba6bb \
--hash=sha256:9b75e958f14c5a2614bae415f2adf7eeb54d50a2cfbe7e24fd4826471bac3625
# via pip-audit
pip-audit==2.10.1 \
--hash=sha256:1eb4565d19ebe5d48996f4b770b4d2b32887e12cb12cfa637f1a064011b55ffc \
--hash=sha256:99ef3f600a317c1945f1e89e227ef26e1c2d618429b8bd3fa6f4f7c440c4611a
# via -r .github/requirements/pip-audit.in
pip-requirements-parser==32.0.1 \
--hash=sha256:4659bc2a667783e7a15d190f6fccf8b2486685b6dba4c19c3876314769c57526 \
--hash=sha256:b4fa3a7a0be38243123cf9d1f3518da10c51bdb165a2b2985566247f9155a7d3
# via pip-audit
platformdirs==4.11.5 \
--hash=sha256:89f8d42695853b89c7170bd49bc3dc593f98a71e695ede88e06a3b247bc4563b \
--hash=sha256:e8b31f4f8bcbbedef91a6b57a706255e4f148d2a4e01648382a0a47342539173
# via pip-audit
py-serializable==2.1.0 \
--hash=sha256:9d5db56154a867a9b897c0163b33a793c804c80cee984116d02d49e4578fc103 \
--hash=sha256:b56d5d686b5a03ba4f4db5e769dc32336e142fc3bd4d68a8c25579ebb0a67304
# via cyclonedx-python-lib
pygments==2.21.0 \
--hash=sha256:2363c69b61c4a97c838da3b130dcd6468f4848992b21a82f2a63ec34377137d9 \
--hash=sha256:610ca751c9bc2492b38eb9a38a7fbc93edbbb2d7182edaf34e66ae493dee5c8c
# via rich
pyparsing==3.3.2 \
--hash=sha256:850ba148bd908d7e2411587e247a1e4f0327839c40e2e5e6d05a007ecc69911d \
--hash=sha256:c777f4d763f140633dcb6d8a3eda953bf7a214dc4eff598413c070bcdc117cbc
# via pip-requirements-parser
requests==2.34.2 \
--hash=sha256:2a0d60c172f83ac6ab31e4554906c0f3b3588d37b5cb939b1c061f4907e278e0 \
--hash=sha256:f288924cae4e29463698d6d60bc6a4da69c89185ad1e0bcc4104f584e960b9ed
# via
# cachecontrol
# pip-audit
rich==15.0.0 \
--hash=sha256:33bd4ef74232fb73fe9279a257718407f169c09b78a87ad3d296f548e27de0bb \
--hash=sha256:edd07a4824c6b40189fb7ac9bc4c52536e9780fbbfbddf6f1e2502c31b068c36
# via pip-audit
sortedcontainers==2.4.0 \
--hash=sha256:25caa5a06cc30b6b83d11423433f65d1f9d76c4c6a0c90e3379eaa43b9bfdb88 \
--hash=sha256:a163dcaede0f1c021485e957a39245190e74249897e2ae4b2aa38595db237ee0
# via cyclonedx-python-lib
tomli==2.4.1 \
--hash=sha256:01f520d4f53ef97964a240a035ec2a869fe1a37dde002b57ebc4417a27ccd853 \
--hash=sha256:0d85819802132122da43cb86656f8d1f8c6587d54ae7dcaf30e90533028b49fe \
--hash=sha256:136443dbd7e1dee43c68ac2694fde36b2849865fa258d39bf822c10e8068eac5 \
--hash=sha256:1d8591993e228b0c930c4bb0db464bdad97b3289fb981255d6c9a41aedc84b2d \
--hash=sha256:2190f2e9dd7508d2a90ded5ed369255980a1bcdd58e52f7fe24b8162bf9fedbd \
--hash=sha256:2c1c351919aca02858f740c6d33adea0c5deea37f9ecca1cc1ef9e884a619d26 \
--hash=sha256:36d2bd2ad5fb9eaddba5226aa02c8ec3fa4f192631e347b3ed28186d43be6b54 \
--hash=sha256:3d48a93ee1c9b79c04bb38772ee1b64dcf18ff43085896ea460ca8dec96f35f6 \
--hash=sha256:47149d5bd38761ac8be13a84864bf0b7b70bc051806bc3669ab1cbc56216b23c \
--hash=sha256:4ab97e64ccda8756376892c53a72bd1f964e519c77236368527f758fbc36a53a \
--hash=sha256:4b605484e43cdc43f0954ddae319fb75f04cc10dd80d830540060ee7cd0243cd \
--hash=sha256:504aa796fe0569bb43171066009ead363de03675276d2d121ac1a4572397870f \
--hash=sha256:51529d40e3ca50046d7606fa99ce3956a617f9b36380da3b7f0dd3dd28e68cb5 \
--hash=sha256:52c8ef851d9a240f11a88c003eacb03c31fc1c9c4ec64a99a0f922b93874fda9 \
--hash=sha256:559db847dc486944896521f68d8190be1c9e719fced785720d2216fe7022b662 \
--hash=sha256:5a881ab208c0baf688221f8cecc5401bd291d67e38a1ac884d6736cbcd8247e9 \
--hash=sha256:5cb41aa38891e073ee49d55fbc7839cfdb2bc0e600add13874d048c94aadddd1 \
--hash=sha256:5e262d41726bc187e69af7825504c933b6794dc3fbd5945e41a79bb14c31f585 \
--hash=sha256:5ee18d9ebdb417e384b58fe414e8d6af9f4e7a0ae761519fb50f721de398dd4e \
--hash=sha256:7008df2e7655c495dd12d2a4ad038ff878d4ca4b81fccaf82b714e07eae4402c \
--hash=sha256:734e20b57ba95624ecf1841e72b53f6e186355e216e5412de414e3c51e5e3c41 \
--hash=sha256:7c7e1a961a0b2f2472c1ac5b69affa0ae1132c39adcb67aba98568702b9cc23f \
--hash=sha256:7f86fd587c4ed9dd76f318225e7d9b29cfc5a9d43de44e5754db8d1128487085 \
--hash=sha256:7f94b27a62cfad8496c8d2513e1a222dd446f095fca8987fceef261225538a15 \
--hash=sha256:88dceee75c2c63af144e456745e10101eb67361050196b0b6af5d717254dddf7 \
--hash=sha256:8a650c2dbafa08d42e51ba0b62740dae4ecb9338eefa093aa5c78ceb546fcd5c \
--hash=sha256:8d65a2fbf9d2f8352685bc1364177ee3923d6baf5e7f43ea4959d7d8bc326a36 \
--hash=sha256:96481a5786729fd470164b47cdb3e0e58062a496f455ee41b4403be77cb5a076 \
--hash=sha256:a120733b01c45e9a0c34aeef92bf0cf1d56cfe81ed9d47d562f9ed591a9828ac \
--hash=sha256:b1d22e6e9387bf4739fbe23bfa80e93f6b0373a7f1b96c6227c32bef95a4d7a8 \
--hash=sha256:b8c198f8c1805dc42708689ed6864951fd2494f924149d3e4bce7710f8eb5232 \
--hash=sha256:c2541745709bad0264b7d4705ad453b76ccd191e64aa6f0fc66b69a293a45ece \
--hash=sha256:c742f741d58a28940ce01d58f0ab2ea3ced8b12402f162f4d534dfe18ba1cd6a \
--hash=sha256:c7f2c7f2b9ca6bdeef8f0fa897f8e05085923eb091721675170254cbc5b02897 \
--hash=sha256:d312ef37c91508b0ab2cee7da26ec0b3ed2f03ce12bd87a588d771ae15dcf82d \
--hash=sha256:d4d8fe59808a54658fcc0160ecfb1b30f9089906c50b23bcb4c69eddc19ec2b4 \
--hash=sha256:da25dc3563bff5965356133435b757a795a17b17d01dbc0f42fb32447ddfd917 \
--hash=sha256:eab21f45c7f66c13f2a9e0e1535309cee140182a9cdae1e041d02e47291e8396 \
--hash=sha256:eb0dc4e38e6a1fd579e5d50369aa2e10acfc9cace504579b2faabb478e76941a \
--hash=sha256:ec9bfaf3ad2df51ace80688143a6a4ebc09a248f6ff781a9945e51937008fcbc \
--hash=sha256:ede3e6487c5ef5d28634ba3f31f989030ad6af71edfb0055cbbd14189ff240ba \
--hash=sha256:f3c6818a1a86dd6dca7ddcaaf76947d5ba31aecc28cb1b67009a5877c9a64f3f \
--hash=sha256:f758f1b9299d059cc3f6546ae2af89670cb1c4d48ea29c3cacc4fe7de3058257 \
--hash=sha256:f8f0fc26ec2cc2b965b7a3b87cd19c5c6b8c5e5f436b984e85f486d652285c30 \
--hash=sha256:fd0409a3653af6c147209d267a0e4243f0ae46b011aa978b1080359fddc9b6cf \
--hash=sha256:ff18e6a727ee0ab0388507b89d1bc6a22b138d1e2fa56d1ad494586d61d2eae9 \
--hash=sha256:ff2983983d34813c1aeb0fa89091e76c3a22889ee83ab27c5eeb45100560c049
# via pip-audit
tomli-w==1.2.0 \
--hash=sha256:188306098d013b691fcadc011abd66727d3c414c571bb01b1a174ba8c983cf90 \
--hash=sha256:2dd14fac5a47c27be9cd4c976af5a12d87fb1f0b4512f81d69cce3b35ae25021
# via pip-audit
typing-extensions==4.16.0 \
--hash=sha256:481caa481374e813c1b176ada14e97f1f67a4539ce9cfeb3f350d78d6370c2e8 \
--hash=sha256:dc983d19a509c94dba722ee6abd33940f7c05a89e243c47e907eb4db6f1a43e5
# via cyclonedx-python-lib
urllib3==2.7.0 \
--hash=sha256:231e0ec3b63ceb14667c67be60f2f2c40a518cb38b03af60abc813da26505f4c \
--hash=sha256:9fb4c81ebbb1ce9531cce37674bbc6f1360472bc18ca9a553ede278ef7276897
# via requests
+1
View File
@@ -0,0 +1 @@
pytest==9.1.1
+32
View File
@@ -0,0 +1,32 @@
# This file was autogenerated by uv via the following command:
# uv pip compile .github/requirements/pytest-tool.in --generate-hashes --python-version 3.11 --python-platform linux --constraint requirements-ci.txt -o .github/requirements/pytest-tool.txt
iniconfig==2.3.0 \
--hash=sha256:c76315c77db068650d49c5b56314774a7804df16fee4402c1f19d6d15d8c4730 \
--hash=sha256:f631c04d2c48c52b84d0d0549c99ff3859c98df65b3101406327ecc7d53fbf12
# via
# -c requirements-ci.txt
# pytest
packaging==26.3 \
--hash=sha256:94edc256424af38762eb31306eed28beb9f0efc50a8837492c9d6fd6004aed79 \
--hash=sha256:d7193f7c8e4e93f444fde0262bf90af30e16fa0ad0ad44cb553c87339b23cd1c
# via
# -c requirements-ci.txt
# pytest
pluggy==1.6.0 \
--hash=sha256:7dcc130b76258d33b90f61b658791dede3486c3e6bfb003ee5c9bfb396dd22f3 \
--hash=sha256:e920276dd6813095e9377c0bc5566d94c932c33b27a3e3945d8389c374dd4746
# via
# -c requirements-ci.txt
# pytest
pygments==2.20.0 \
--hash=sha256:6757cd03768053ff99f3039c1a36d6c0aa0b263438fcab17520b30a303a82b5f \
--hash=sha256:81a9e26dd42fd28a23a2d169d86d7ac03b46e2f8b59ed4698fb4785f946d0176
# via
# -c requirements-ci.txt
# pytest
pytest==9.1.1 \
--hash=sha256:1088fbde8f2b49d95a549a195707afa7a76a3ce9bcadc26b6d71f0ffda5fe313 \
--hash=sha256:37a86b45efb9a47a61a36449063e8e18d0cab3161329fc099eb21783169c4f0c
# via
# -c requirements-ci.txt
# -r .github/requirements/pytest-tool.in
@@ -0,0 +1,3 @@
bandit==1.9.4
semgrep==1.175.0
jq==1.12.0
File diff suppressed because it is too large Load Diff
+1
View File
@@ -0,0 +1 @@
twine==7.0.0
+470
View File
@@ -0,0 +1,470 @@
# This file was autogenerated by uv via the following command:
# uv pip compile .github/requirements/twine.in --generate-hashes --python-version 3.11 --python-platform linux -o .github/requirements/twine.txt
backports-tarfile==1.2.0 \
--hash=sha256:77e284d754527b01fb1e6fa8a1afe577858ebe4e9dad8919e34c862cb399bc34 \
--hash=sha256:d75e02c268746e1b8144c278978b6e98e85de6ad16f8e4b0844a154557eca991
# via jaraco-context
certifi==2026.7.22 \
--hash=sha256:62f22742b58a1a33014a2b6b706588a8d7e2a88ae7bd1a6ebe8c992928483775 \
--hash=sha256:741e2c3b351ddf169a738da9f2c048608ff7f2c5cc02f1ebc6b118bb090d5d55
# via requests
cffi==2.1.1 \
--hash=sha256:046bfc24911b37851ee1b51aab8bffe713d89c68c6a057b09484ce9fd5f69b4e \
--hash=sha256:06c72bb76605a4b0cd0aad6930b69d4baf7dd5d806cfc409b824191099700e66 \
--hash=sha256:0beceaabe56af686895136a2de78db54ecd8e4046b236b8fd6d6cb61389e9bf2 \
--hash=sha256:154852545011f779917b11c78db2358d095da62a9a172b78ad0a583ee5adc0d0 \
--hash=sha256:194cffa889098ced9976c3fc6340305e43f6303657d298da55366907c05c22d6 \
--hash=sha256:19ee6127ee34de7d83ce3d371ebc5ed91addbdcc39f9ab15ce4eb35a4e534971 \
--hash=sha256:1a18a57b58cfb21fc28d72e876acf10eaed67a1ed96226f92af4df681d571c4c \
--hash=sha256:1aa5645c30469b09530c4ebca77ebf8f17618293c58f8549cb1a543a50236e7d \
--hash=sha256:1dea0e4d7d4f11f619fe8c1d76caf49e24405b4b5743c0e3be16a500ecd930c9 \
--hash=sha256:208f941bb9d18e768138677f0a6d2ce01f590df56043dda1df1535ac57c88517 \
--hash=sha256:210019b6c7cf07f081b4c54635c8cf744377001350e29cc0f81c4377b4797735 \
--hash=sha256:246fa40ce8645a614ff682e0b70f37134e460eaf93a775e0cbe3cca585a67a80 \
--hash=sha256:25792eac27877609e7bb06d42ff88278a6624fff2ba9bbb523c09616b117e80f \
--hash=sha256:27350daa11d4f10c540e6e89dada4c54feb7256ad03e9a4dc075ebad7ba360d1 \
--hash=sha256:28907ab9bfb6aa13184cfc17c6b8e1023c5ab6fd7076d8c20a35e59fe04f8f29 \
--hash=sha256:2ae64be792b8966f2c69538199728b290e34726562896df1e5dc8ffd8d8188e8 \
--hash=sha256:31348097ff5bbe827ccc41795d4dd099d9f0625e7def00ee653c137a490c2a6c \
--hash=sha256:3143d81e29e1e20a9ce10901ec369012947876596f75a222235965f2b7ae832e \
--hash=sha256:3222ba5d678f80a030e6afbcc33dc1ae5cb45facabb61cee2c7016b8432fde48 \
--hash=sha256:3311ed60d36f83378794e1009ac6258bafbf81f7888b4caa7b35a521e3f95813 \
--hash=sha256:334644fbac4eff73d985a17a91226df55d0f394160c4cfb880e084c8f7161cac \
--hash=sha256:34e261f78cb6ceaaa36f42f2613f4380d94d9c759a9c73c769ee6e0247364632 \
--hash=sha256:363e05fa78e15116c3c32c210ee36884fd6b9afa6d440e47112c3bd511d64cb6 \
--hash=sha256:398aff33cee2767e3e781d2554c54bd0dff386bb437581e0d8011fde1a942ec1 \
--hash=sha256:3d22a20b1fb1632cc72c22f95f7b0d2961c3e1c235f245ba4c606c4771035659 \
--hash=sha256:42a494cee34437f05546455144f2b5d9ac09b1face62bcfce597d2e521066688 \
--hash=sha256:42e2f76b9455f5a9a844f770bf3e200ed3da0e15f5df3db9c31fe80b04b3d004 \
--hash=sha256:42f6930c31dc7f50732c9ae793c2786c7b6b044195967bbdde40bb9be81c4cc0 \
--hash=sha256:456a61fa52d579ebf9df2e9552ead5129855dbaff6c1e5a9b1bc408809bdc062 \
--hash=sha256:471cee653ae88de62096552e6d24ccb4a5adb8c8c9f10b5054d0122c15bf2779 \
--hash=sha256:49cbc70e6542d4ccccb936558d1064a8012541e78f821f955cff24e357776c94 \
--hash=sha256:4a7c934f7360e8cd64fe9efadcbd10c7c6364f531e432b9a4bf5ccbc9e0e8b50 \
--hash=sha256:4be96343e422f2dfcd12ab5c9f5aebe03f82f737c6bffeca6830b3875cb44aab \
--hash=sha256:4f42141fc14250de6dde5ee7ea4432be017252d91f19c5ad043c084cea629cac \
--hash=sha256:507a24c282e0f42f8ed737cf048572cbf580468da5555764a8331735e9c736b6 \
--hash=sha256:51b31d1c98274844cfd7838ce00bfc27c7423a4dc00fc0772fc3331c2cc90676 \
--hash=sha256:58acb8ab8e295e6c5ea12f888cbb13cf21511ef2a3303a23f4325c29d17fe5c1 \
--hash=sha256:5a59cc1c4442bc3d5c703bf720b51138d0bfc173618807c9ee2490a7541dd3d9 \
--hash=sha256:5bb4e7ea95dcd6a014a6fef62e62467d67d8e582326443f3d68e71d6320a9fcf \
--hash=sha256:5c58fe613dc5e5336357eff555824a314d8e43282600435c8d1cb6a7a2fedd13 \
--hash=sha256:5e7cecbaadb83884793e05828cee59b210b24583b9c7425d0ba6a754fe22eb4e \
--hash=sha256:616f097f2fe415bc92a247f02e11f634e1f9e9a83d327e3c915c15089c87869e \
--hash=sha256:63bbfd5ded17c4840ac07cd8f1c21ba9d9708141f840b324f422f41b207e3973 \
--hash=sha256:64faea20f4e2613363a1a9b9c7dd73058f3ecd00133a511e72ad7c511658f527 \
--hash=sha256:661c298b4821edebead0c91edd2b00374d67ad7c5a1f7a91d4442633b79d6a72 \
--hash=sha256:68e62fe11f30d5ca8289242866f0a5291402d8529ca2178ab8afc5c9694ae890 \
--hash=sha256:6a8dddef476fab96d066d578fc88526767b836ab5ab21754e1d5bf3879c31c7c \
--hash=sha256:6e192623c49c94421616a5778fba35cf0d5a8d000650c1967ef4448ee5cdd990 \
--hash=sha256:7225e4514edb64eb6740324353e0da0711954fd8d7da4576755b1c6e09b697cd \
--hash=sha256:75f80557d1389eddbd0de2681f6a390a0c5338c31ddaa821381c203fc3fd50d9 \
--hash=sha256:770de9db11e84213beec501cfcaa013b019820ca881e03344dea5844f7876d94 \
--hash=sha256:7750c6449dff7864bb9bb27ddfb0267756189201a3afc911d82b3caacd70dfc3 \
--hash=sha256:7bde5e4cc5c10140859842b9d383af292b22639a4dffb725314baf45968cef80 \
--hash=sha256:7ce713ace7c0e4520535b42b77eaa742c16dab813978064913e5a3cf82973b41 \
--hash=sha256:7da0c5eff80f0197f3b3d1232ec5a682a9325f4ae9016a78f5f5ca35f9ced1f5 \
--hash=sha256:7dbb61fe3a7699468030f71bbe5f8a0e326a151daa91beb11a6fc1f980c55e1c \
--hash=sha256:811bd1e21d32de12efca32393a0ab3f5133b54fce9bd44b8bd77ab07da14bf6a \
--hash=sha256:8ef53b2de9bcb9197d31854256575d59dbac0cba72ac627bb291ef5eceb74be4 \
--hash=sha256:937c0052c05a31ca1daf18de3158eed4dbfcb9cc107adbea227728d647be701e \
--hash=sha256:9d2055050ea716bd38b7f7f1579c275386646b4894c155a3e2f3cd62ed41b7c6 \
--hash=sha256:9f8d177621de5cb38ee3e731eda45d421db093ec0739f46a5594babda7987a98 \
--hash=sha256:a2d7755bef5a12ed488f4ef1f1b69ee9191d7396083b755a5d2295f6edb4768b \
--hash=sha256:a48d62ab9d6f4f98c983223a547af44be6ca3691074c31cecced6facd3ba2dc1 \
--hash=sha256:a4f00aa42f75d6e4595e8866e748cc1705adc0cddfeb2ca86d0d03993d63ba03 \
--hash=sha256:a6e721d4b0e45d5b65e87534470e67b18dcd092c83f68fba09f152b9cbc061af \
--hash=sha256:a730a083190634c65cca36ba5f489531576ebd79bcd5c8e172130f6453127231 \
--hash=sha256:a931079504ecc49efed7744c476a5c343a92fabf66dec2db95edb1b2fdc770e2 \
--hash=sha256:aa9511c62d14da7aacc9b4bf51f3f697a621e83b2d6919008243c3aad168eea3 \
--hash=sha256:ab36d55f9ed2d067327667c2fea18dda018eb628dd6347aa01dda6cf1f5d3836 \
--hash=sha256:ad2c86c495b899d862ea0f4b42891b8713a3bd45dd4105c7fd51c2a72f39f3a5 \
--hash=sha256:aeae0e330c9f6acd681f647d46cefd30c29f93e3392882e792e82080c9691399 \
--hash=sha256:b0431303acaea1089ad4b3e9ce4e6518193def1118d4073ca848635ee4ea2e96 \
--hash=sha256:b5bdfd1c873d4e093aabc0ca84c4ca6dbc4f752afb5c86f146d9742580c9da2e \
--hash=sha256:baed1e86cc735622097354b9d1281406caf42ff42a886d29faa8e8d1630333be \
--hash=sha256:c1453022f490d2459a11819d83ad1d586e9ff65a12ac3e705ffebd46d3685dcf \
--hash=sha256:c26608d2222fb1e94487e4a387d85f13eb55d5ed725cb25a0c589ac4ee60e7bc \
--hash=sha256:c7659f22557c5a0bc4855cd635f55edec690cc008a40768527762cb9fb263455 \
--hash=sha256:c8c69575568085ba0b1b10c0249d779a214aea6f6522e949a0fc9fb0fcb449d0 \
--hash=sha256:c8d2c9fd1f2d16f780d15127abb050d13d1a76c03a4bd87d7e4980e45e511e12 \
--hash=sha256:ca82be1a1d406ecfe1d25dc16cb33488e5a16bf4438c9fb590484ea29d92478b \
--hash=sha256:cc572dace3f60ef98d7b12ff411d20f5362feb31a0439eab0085bbfd349982d7 \
--hash=sha256:d18e5ac0f2f03f4f518d3e23db0f0cad7faa1da8620e9c09461d443bbf6e6692 \
--hash=sha256:d28630f5854ab07ab1fd4aba756de52326c82e6be15d414b12793f1975048b54 \
--hash=sha256:d9c275eaacd24aa73f94ffd6de08fc3f932424d8b6c376f4bed7cde376fe7bc3 \
--hash=sha256:da0e573f9f97159390c89d9f1a9e41908b66d408cc5b58d08cf3847d844c531b \
--hash=sha256:dd31f52ea1086513bb9df30f8fcee9b8918323ae067a3d5b78bc826a000712be \
--hash=sha256:dddad92b554513a31f272570678ba307fb9f618f05e3d4a5eacafff9eae03e1d \
--hash=sha256:df423d40ee8654634421812bc3b196da3f9bd7d32929da813f8394c4348a5358 \
--hash=sha256:df913725b79db7bcf03448f36b7bf8815363417d5b58deecf9305e3e30f0f21a \
--hash=sha256:e0bcb7e0f677f543555d2adff3bf19c05f66cdb4796e5ff602442ab2fe3c4ef7 \
--hash=sha256:e2d65b31f36619cda3999b78b2aa9632e76b78448e7a56fc4240824200e7c4fc \
--hash=sha256:e6e8cff14d6fb0be70a09c0bdc58096f501952d04624ebf867e0e56da2df8960 \
--hash=sha256:f16c709686a78c727bbbf059f92b0bf41c6fc60deec706d2dc19f529175a6125 \
--hash=sha256:f24fb43132a4c6b4cb4eb029492919b2db645be6808d738f244fd146c03c32cb \
--hash=sha256:f53e442b08449d42821fa4a4fba000095af9f62742a500f978a9f557ec44339a \
--hash=sha256:f5cfbc5fe74540d335175b656c725d74d90e3730c626d92575eea35029d9afaa \
--hash=sha256:f81b3b8f3d4e343550fa4baa0e479bba9f2d29ce9c2e9b51d1ce1718d7442fcf \
--hash=sha256:f8ec5e643a9a937f64e1999eb9f75d072263751912dc5cd06d3c85f8f44be7c3 \
--hash=sha256:fb92203a88b3d3053034db775110081c49d28be6551923805e039924093761e4 \
--hash=sha256:fcd22650c908d7b7da162bbfaab594a1227a15d1643a98c68b122ac642fa2264
# via cryptography
charset-normalizer==3.5.1 \
--hash=sha256:00668ebb0609751758682eb0b5857e7c35b9f00e84dfdef062e103244ec94d45 \
--hash=sha256:012a22b88a77ca2e59b98ac5889b0deb604147666032f45e6d6e217634d2550d \
--hash=sha256:01e93745f7f219b703b60ba7afead36cfc4242782be5af484673fc500df12da5 \
--hash=sha256:04368edf83514385ffc3e1cfd4546e595f4f1272dd23ba437a93a9cc3741d47b \
--hash=sha256:0722590aabf9dc6a6c0343d523c05458fa2b5047dbe6302fd526bb570600753f \
--hash=sha256:07ffd07412fc5d5e84cd8952acf9ff7e4ed7a708e69d1bada19d8ba91711353f \
--hash=sha256:09a7bba9f739468c8e78c36a75c33768e53cb1959fc638f510454c14683f00d5 \
--hash=sha256:0b2b1b3fa5670c127b246df1d0c059defd41f689a868a3b9d79df9b1cac42d22 \
--hash=sha256:0c6dfb5ca6723eeed15aa8e564a014d69fcb8812f94eef11fe3631e0508199f5 \
--hash=sha256:0d929fc574b4d6fd9e7c0f5c2ede8716a41911923aa7fa5fce38e0818aa4a1ac \
--hash=sha256:13e3afe97712e8887cd516e960c63f0b93122971e5b5e4b2622fe7701771e838 \
--hash=sha256:15f024313246a4ed976c60f440bb8d257815513a681d212ff74fd46f7d715a90 \
--hash=sha256:195ce897c6153c0700078142cf8efe3e6454ca4cf4357499e4078dfd83396626 \
--hash=sha256:19a3dd5aa73cef1c99687c4fc57db016a9c17104ae1185da88ba566a5d3bebe4 \
--hash=sha256:1d1c7a53a6c2103925cdd6d7229f8c567379f211c869793df679f2e9f738c369 \
--hash=sha256:1f5883d77fd409a261abb5dc8ccbe335720d798b1de4abb3b1d47ccbbc76b53b \
--hash=sha256:21b82d8082f6f5e7f456ef0bd16323d08de1266efbfeb476e64b2a91d1471a4e \
--hash=sha256:252d099029bcbea642f2a06c4ed5046bdf8b5a8150b64afa5e027e88b106e5ee \
--hash=sha256:256dd4d85d9e4dc595e2bc983c980e73f62ddeb3165c58b4c3dfe78c5c8548c1 \
--hash=sha256:26422d45fd13551cf564c58932f7d72b4f58b93b0fcf18c35ba6be12b46bb102 \
--hash=sha256:2679de311c7946dde5d3b6f44941844133ff5c7cb86099c0061ab1e8901c20a8 \
--hash=sha256:29880d17a8eb0b5cfdfd8944b468322928059aa35f1f5fa8ff22b149ec0b42f8 \
--hash=sha256:2bced4061f000f7187254a02ad3433ae17eaf991747ceea2f478422590a5bba9 \
--hash=sha256:2e9cf9253119d8e5d111f05d71626786fd3d6193817316eab1ca088cdb8593cf \
--hash=sha256:2f06b7eae9dbe77fe1d644ca244dad508de8d302870a43f3c559b521270938a0 \
--hash=sha256:2f293479cce755c75f1697e87c409b7ae4c555c7dfecb6e988ad13abba943031 \
--hash=sha256:329fc3ccb63ad22d867d84c2adea759a64079a37ba4a343433b02c7a2816871e \
--hash=sha256:343fb4f2821043bd87095f7b08a1a181febc8e36ac64212143bbfd0a0e1bc235 \
--hash=sha256:3588e376b3ea2eea84976f67273d679f229e24c66dce7b82ae45aef04ff6e072 \
--hash=sha256:35aea775dc2bd5f54cd84a1cd2696cc3207c479cb9cf0bd346f0d343e4300ddb \
--hash=sha256:35fe081843b35aad20ffeccec3eeffbe637b15d14f3fb22cc1b59cd8ec17e93c \
--hash=sha256:36047af20e17097c3bb9476c2b7655f2f7aa51322c0ba58c07695bedf755a950 \
--hash=sha256:3617ac3cfd8b9888f145ad89dd6e692285834b0201c6074a5eeaad3fd4d668c2 \
--hash=sha256:366ec70f5547c640d3ce1985722490f23faf4eb5216a7eeba78277490e78dacb \
--hash=sha256:394fea06235c8543390050ed5f529187074b029fb027213f6c46ac11ab5d950e \
--hash=sha256:3d27167433c0d5f18dc850f07d0b3816221984fecdc405d6c157a6f0b8f8e9e6 \
--hash=sha256:3e5e1224c0a6a90e05843e07adfec669edebec17801c67072f51e59561d63c0b \
--hash=sha256:41876ee62a3dddf48ff1121ad8f0798032aa03f2fd35f21f34a4cab14f18d8d2 \
--hash=sha256:433c5a81eade63b47e522303bad236f59dba55ea6951746f5558355eeed8c75d \
--hash=sha256:4582c27e8c889d64811987b5967fbd3ae0c823fe1fd933b543d55ac20bb475fa \
--hash=sha256:485a0d363cafefcd2538a73c7c838daa2035f09b2c9f9b5e3133f80c6aeb84c2 \
--hash=sha256:494b70049a4d69aec6e8137c13af4cf8db8c9f9820a1392ac293b0dd2987a818 \
--hash=sha256:496846868fea80e479324862fa877f02411f2fd0f83b79ccee2607aa68b2a032 \
--hash=sha256:4abdc5f9ad448c1ecbfae2974b820535d6bc6e7eef63babbab3d81cf46968c71 \
--hash=sha256:4b599739b93b2cbeded49645ae3c8d1405c29ddfbceac1545c87a3f9580a9e96 \
--hash=sha256:4bea7f8ebe90bbd7f0e4a2de42ca6924ba23e3e76418c408ff82f1d46fabd687 \
--hash=sha256:4c4fb141a727957c93edfe5c32a26ceb6b5f6461d67146e2d39f51e16170bea8 \
--hash=sha256:4c9548dc78002099910abaebc0a72ac58b7d30931869e0351c09b507dff4ece3 \
--hash=sha256:4d26f14f041e83dd8edfd61f4cd4fa7285d31798b5bf1f28e70c367ba6c41d61 \
--hash=sha256:4f298bdadb8f0b9e5672877f647d1be9373ef5320c9e2f049795e26cad28b6a9 \
--hash=sha256:52ec005752a56ae79547a05c0139ca2501a0c866390b6115008456b9f0e7cde1 \
--hash=sha256:55261ac0d2941c42f196dd576f543d87a8ee03cd6f5e30dfb4d807b2e3b9121a \
--hash=sha256:56490c595a28b1bb27dfc583e816152a9767721ef58b2c03b13f954d2f707420 \
--hash=sha256:58d3e12c88e0950bca850ae1f7c256055c097639c2edb9eb123af9807d8b15e4 \
--hash=sha256:58d4aa13a59c969dbfdf9e6a9560e242cbfd9e8a8f50c2747714df1a423adf65 \
--hash=sha256:59171c6e45bf07d0d5cab3b0bf81d945035530f6873398b3b531c31184d46663 \
--hash=sha256:5b6d1386bf0096d26d3a863dc0a487a5b4eb9aa93cf5ba69683d29dde6b9d60f \
--hash=sha256:5c0ea61a470e070686aa30892fed79e297d2c8d0ab46b8bcdf027d38c51da591 \
--hash=sha256:5c84bec0ab5ae0c64bfe73a7d2adcb5ce73b467523fc27fd6a28ab2aa6cbe35a \
--hash=sha256:5ca0555312ae2fe82715cada7fac375530c2f3349e1eaa1bcb33d0283ac79a18 \
--hash=sha256:5d8531a6569d025f68e2321e7638fb7978f23db58e5f69f56913837aae03816e \
--hash=sha256:5e2d0e146dcb57034f8b97dc58d2d512cb90aba253960ce449f695fec6a82c6f \
--hash=sha256:5fc45d653ea8c9a20479167e11d4a0f8cb2fa3470737ab6f9c827532313187b7 \
--hash=sha256:6117b84ea48435e5356dc737f5121485c30920ba43375fa7b434fd753df0eac3 \
--hash=sha256:6199d5606e2bbf2b096cf64d03f8b6790c91081d5ac866b8e7bb6422738cc60c \
--hash=sha256:62b55f6722735a6c472f88361cde6640608773d9443cebdbb51abf436a1fcdd3 \
--hash=sha256:687c9ca3035544b113bea2055e180af96fb63c0c476e22a9180f51925186e7b7 \
--hash=sha256:6b7430cf5728e68f6c462254009a6ef4086e1bea43cf2f57aa9c55fb4f50ff96 \
--hash=sha256:6ba32c4d2abf1d2fe7cf27d280f4cca5664233b0f885549c7761719eb977f486 \
--hash=sha256:6c9cdde8becb25a7fde49924511aa2644d6f8081cc8df8e9452724303348d8e3 \
--hash=sha256:6df0ec430f9a831772c23ca5a224cba36517a58a84bb32c32bb59a9fa67c47f6 \
--hash=sha256:6e2912d4babbc65196ac13c2f53468dc57fb8b9c25ef913e8c59ddf7c6dc0e1b \
--hash=sha256:6e5e4d73d588ca5ed09df1b7dcd1b203d1df3c542e3f50d126c947d432b10731 \
--hash=sha256:70055ff39b97c99e7ae40ea3e393fb62aa2e44dbd9b29f8d14f42fb0025c3959 \
--hash=sha256:706bfd38730a5ac7a365793269a00f4e988178cec121391f4248d84ad8c972e9 \
--hash=sha256:7235dc28fc6dd9d832ac7c7bce95367dedb85929f17368a0c2bee1e080b9acbf \
--hash=sha256:774d157f112367ff4abd29019f38f023c24e00e56edc7829c20e358a5a913ad8 \
--hash=sha256:77efcff2b23071c349402ac1066667a3d011f62398d81408c9b88ad991747c9e \
--hash=sha256:789b8982559ae28dad2356519f841655756cdcd96616410590ae0b17454ee64f \
--hash=sha256:7ac76cf9afd34929d76eb7fcb63be476a4853d8a96f0dcf2d0db68a0cbdf9885 \
--hash=sha256:7c0c10730342b0c9b35dd1d619beb8214e520bd96a1f870f452680b238aab3e0 \
--hash=sha256:823f82903d189af463d7df250ef1f7f696f3cee08cc8d91deb565e8d425f6506 \
--hash=sha256:838648accb3a7fd9803fd45c87bce8509648eb0c11bc34e216141300977244f2 \
--hash=sha256:854066be00447fa8de2ccbbe893e2ffc4b123ef16d897af794c1e18bd4a714b0 \
--hash=sha256:85d5855daafc240cc045c026d7a15fd198a09b0fc8ff6f5ecbb5297b509cb11e \
--hash=sha256:85de3134b5379856e323ba37c19c9256d39425f7b76a63af52b09fb4664c2e8f \
--hash=sha256:87e4f41d375c0b9be2fb5251aee4b8a689169e134535aed81bf085c3b647451e \
--hash=sha256:88ca277405c2d3b71c4e1c2ee0e7966e807bcba86a69d11e19ba199d18ae4491 \
--hash=sha256:88e85ab89cb822c1e635f51d6d32e488f94e002e70e2f492bdb8b945543f345a \
--hash=sha256:8ac8c94b6539074e0f40899301273ac8402b9b3e01c7b7ba269ff30340aaaf20 \
--hash=sha256:8fe532b3c966d1fb794e0698e4589d0444017ae77fc0b31edea13c0e35bcc449 \
--hash=sha256:9085f87b0e38a2b92b8923059b4e8789fe40d9279712d15dcc670048d77079af \
--hash=sha256:90b7481fb62fbe172c558bc6fd1c4c98d82004a54a7551f20e11ac9bf0b8708c \
--hash=sha256:92caef967d287a407085d61176fce4012b1dd62daed4eb6d5ceb26d3d2538712 \
--hash=sha256:9362dd90aa7dab48c0054a21187791ccf05473f7dba5d92b8033ae62164675e7 \
--hash=sha256:94d78ecec2605a8d0398b0f365d5f12a63248438516f5dac536a5eff7337df4a \
--hash=sha256:94fbf1c0c6cc0d3d5e50f9a9313a8cdca90dd696d34b381cd1704f8c9e939f20 \
--hash=sha256:950f23cb393f85543777b0433f082cddd25b51ab398eac7971146495679efe5f \
--hash=sha256:96eefc178f8636b9c760c5829345307fd81cfae9ab1e80997dbddeb0f54ee9a3 \
--hash=sha256:96fef3e886d6a9874b14f27fc193fbdc69d5d8035783d86aa4e1cea594e695f9 \
--hash=sha256:977cdbd483a9cff38179bea4fd754289a6f2195c7abd414aba85410b3e66cc5e \
--hash=sha256:978eab16f55b4ab2c2a745be9a0a840bf8f09a7f227d9c76eb30214d078865a5 \
--hash=sha256:994e883d17c559cdfd38c84003c8b27d25424a1077272a17e7cd27bfe0bf57b2 \
--hash=sha256:9ac4444d8d4fd4c4bd08bf451ed3167aa9e7ec6cdb41b648794f1d1103652e36 \
--hash=sha256:9b5db6052055d34d41230fb78d7c439c23dc536a9896f6cb039e8dd92cfc1263 \
--hash=sha256:9d9a0dc7cbe9bec24c3f767c9122c41fe5a1bc43f47cd099d00d393e09769de4 \
--hash=sha256:9dbdd9205662134957cf0c324f639bdc5031c0ca056e2369e238db75187c0f11 \
--hash=sha256:9eea3ab2597a5e65fe65296e2d6a84570845a6b55532d90333d740d48bbc850a \
--hash=sha256:a2028475ba855475b8b4d3cfeb4994269c967aea8b9892dfba907f4263a863a3 \
--hash=sha256:a3a370082ce34d0612f421e15fe011c53bb1feff21a26d06ad4fb244dab5a375 \
--hash=sha256:a545775cfe815855ea32d7c27731d79da358ef2055b4a25830231b1622dd18aa \
--hash=sha256:a5cbd90ecf0fc62e64726917ad083b73001f0563657a87ec3c0b504e277dc90d \
--hash=sha256:a6d095662e73e74f0a49988e0593373e243e3a52e27bfeea0a859e88acf4a0f5 \
--hash=sha256:a6dac12ff6b846103483683f60c5f8fee205121adc58ffd87e90a90a3af69e99 \
--hash=sha256:a951ad59cad9145664a730d3036b40b844e74d2d3683da40111463cd3a83845d \
--hash=sha256:aa1099b956fb795e686d073568f6dc002a0bb89765ea6d5b055dd7d9bf1b116c \
--hash=sha256:aa2bb0b37202dca27175591f761108b5d34096ade1191ffe4808bdf6b1571488 \
--hash=sha256:aae2ee51122d3ae968a3837d97dc24a0aeebb0dea23694422cd172bd30017cd6 \
--hash=sha256:ab743e9bc90c1f73552ec33e10e3331315acd2c397b36065b591b0181de533cc \
--hash=sha256:ac00177c4831ffa650f8609e4bdddd5fe09c03b1c0c47acece7e6ea20421598b \
--hash=sha256:ac13b004224fb341e1e25a1ed5e19d32f57cdb2a403e01f003b46f051a550f6f \
--hash=sha256:acaf604462bf330b0d07e7a07c1d6e4adac79e5fb13e9c5140590542cafacc00 \
--hash=sha256:ae31a1a1db2ee6cc2942fccaf695c934bc7f3db9f2133a3fef1f367cf1a4ab10 \
--hash=sha256:ae4a097991662cd4fff0ddc74e0fe7874f82e00042fa0ea00855645ed0c79598 \
--hash=sha256:aea996a6aba25260827c9ea511d1addfde2da9eb686ac961838509086188b7e6 \
--hash=sha256:b39b69b347e5e47a3b5b8cfc005c68c1ba347474e3960236c4944a8ecd174962 \
--hash=sha256:b54e7e13267d49ffbfe68e25b3cbd774dab38fa37238f71265e91b36146eb21c \
--hash=sha256:b9af956078716df40d985fb0dfeb2c2120c5ca92ba4ff4b388acfd01cdc14d08 \
--hash=sha256:ba2f37ee79e6338845261a3c5b1784e5d1acdff2c0785b284f1b633033d136ab \
--hash=sha256:ba501e667c17d8411f98e67a022d9604ef179aff0e459b7e292c796837c13573 \
--hash=sha256:baf3775a2635e5a11fbd5e4e64ee69c7e86875d224a5c72aca4c141064589a90 \
--hash=sha256:bb57753e36e4855b8ca375069482250a6246372331a3e4f3407eaebb007443f5 \
--hash=sha256:bd6c173f04743d483881bffa1478d5a4624475b8cd1d2194956a75548e191c18 \
--hash=sha256:be47f99644b208bff7766314013f9acf57b056b04191d570d68ad14022cf5b1d \
--hash=sha256:c010f5581d9c612804cc59fcf7b524b707fbcb72828551237ab545bb5c7034af \
--hash=sha256:c1dcc36dcb96abc02236e182d17e0f71430152a6c2c7447421da2d2dc144edea \
--hash=sha256:c428c6c31eb5f4277d7f8eccaf767fbd548ddd5ce3c8b4f4cbbfab3d96b5904c \
--hash=sha256:c658c50ac0c98cd755a2dd50b7977d3bca7df401dcc47fbdfa87db53ef7d4e8b \
--hash=sha256:c71fb0d56c920c269cd3e2e3fe7c610e3f1fdb21a6ce60efa6430ff63676cea6 \
--hash=sha256:c7b742bf31c88566b4bb6335a7f393bb322e580b6bb98df7bd0c25e6e3519ce8 \
--hash=sha256:cc0329df4caaceb950d2f580b5ac716a377f7059624a0bafaeaf8a218c6ed774 \
--hash=sha256:cc5d36d96478aa9c60654bd932525bf32964c62a7281eafdf16d85003a8d6004 \
--hash=sha256:ce854f5f478050ade5a238731c4ca985a7d3b3cb53ff600a9b5c3b689b5f0a7a \
--hash=sha256:ced3fdd71aaa83ce593746c2edb42b7a59cb4c19c8b5c407781c72e493aae55a \
--hash=sha256:cee5dd7c6fb5dd52a0fe2a740f9bc6e3593f5f8b1788bde49de02086f30182b2 \
--hash=sha256:cfa1c0cc3a8f9f53f1243a5a99ac36fd003880199383b37672e86ddda9cb07e2 \
--hash=sha256:d1ee1e296209fdce05b81b663250eefa02213a2da7b41bf26f7829b8ba3545aa \
--hash=sha256:d59b75732e9b6f27388e10c14b0259cc5f2e48c78627d185e6a177b58ad3cffe \
--hash=sha256:d63600d620ad0064c3a748b950ac5ea38a80190e5498532efefa4b7b3f1da1f3 \
--hash=sha256:dd732602a7009217f658d5863d12d79d373a4de0eebc111094bcdd3bb8e0a6cc \
--hash=sha256:e06efa066f7dbadbc84ebc126a97c452a6451dfcf589d89d788484949e1cf795 \
--hash=sha256:e199fb99720074809a7720f1c0b4d919eea8b87e88713e0f8f602f7bef543d9d \
--hash=sha256:e4b018dc5a0eee4676e38fe84a47a427816c590b93b55d9025274ec4d6ffc2dc \
--hash=sha256:e6621fb2a4988d6e53eedc455e5903e2679f3967b8acb3d639f1b63c14a2e893 \
--hash=sha256:e71c909f353863b2b89c83de2ebed71ea6d0df8a6ef65a128193c5e650766bef \
--hash=sha256:e90251c0c7bdd54a100a0dce3c07b7e637278c93af29dbf78ebb89a58c4bac7d \
--hash=sha256:e9fbdce1e47394b09bc9f26ab117dfc8d6491977a11d86f592bb42c779db2fda \
--hash=sha256:eb12fb2ba69ffa05f8695f61c69e591dc4b4a12ac3757ac8af8adb259bf56d17 \
--hash=sha256:eda059b6bc8bc0812d626fd91a7ce01bf583df0a61296eff390fd94141a34e30 \
--hash=sha256:f03ac127268b43ef4fe9e6ab6794a6794b49485a0cc0c1db79876d2f33f75bc7 \
--hash=sha256:f298e218441525d3794428b4c8b8fb8662c6d3ea79925d4807ee6b9a96a3bca5 \
--hash=sha256:f5542f9b941279d82d41eb0aa9f98eba36fe4df5c7086c651df7944935b37182 \
--hash=sha256:f6f7deae3feb4edfa2efaf7c574fe88cbf055038a6abdb40188e4fff66d5699f \
--hash=sha256:f9b1e28d0e8dbfa858abdba91d6b547beaf2df1a59bec6da6faae7b96a4991a9 \
--hash=sha256:f9f8405c2c758532c74fed975dbee57be1f31a6e865c031870c79a6ed3212ada \
--hash=sha256:fa48b1b63d639f9483e0633e092f5851e2348c352f1f9bb6c8182f87884ef876 \
--hash=sha256:fb78f6e7fcd8ad785d28cd577168bc1aaee827b25bb8755638f694794ea98f0a \
--hash=sha256:fbc597639158fd7c14d55e808718848319540f51b0e6746e3eefa59723a4a348 \
--hash=sha256:fce8cbd4997efeb450bd298b54f755dcdff18d496f7a5ddbb4867c6d7c88fdc3 \
--hash=sha256:fd0350afdc3aabd5576f60ea109228bd5538139713c7b094c5cd27c73a98bc6f \
--hash=sha256:fd0a274c0e5f9a21565cd9d3dd749b61f96b7aa1e20a93aa1ba4029518f2e5c0 \
--hash=sha256:fdb8a068947befafba9952162645dc2fecaeb400e64584829ed5e9b2fbe21a7f
# via requests
cryptography==50.0.1 \
--hash=sha256:01f41478cf33fc605a6a089cd56d28b45c6c0b45a1928b61797f2621a04bac71 \
--hash=sha256:05ba322c4da95b262a212c345af888ef2c37c88c0509756ea00a0e6d68850f23 \
--hash=sha256:16c5ecd954b3330ebfb6605eca4fd952da8bef376551d5cc264534e3770a9ee6 \
--hash=sha256:2a93d05e34d5f67fba6f891fe85d929999baa7195e853923ea6d7576c9e68c5e \
--hash=sha256:2b34d76a652ea2b6faf777c35df230c5637842cd904e04f16230c3f9f03e4361 \
--hash=sha256:2ebbfb0f1fed745e91796e3e1080a1440423fdae8ece1b995a1d80883a409054 \
--hash=sha256:30a125032e5642a21ff816e021152bd4e7e94f03eff3f4b7fca41cd22bc3110f \
--hash=sha256:330fbb252391c596f1ae42c5754449dc924e6ad012dca8efe0d703f9f2d12ec6 \
--hash=sha256:359e62deae718bce96170e223fdcb6357e4fbd3bb7a3a75f4430763532560e49 \
--hash=sha256:407fe2b6db00939c05c0e945e9914238f2f0a430974839429dafc82b1ee6bee5 \
--hash=sha256:42be3bb70596b3abe4ac097b75be223e8b3ab614a0e5de068e3dcc54d71d6149 \
--hash=sha256:4c4188f7c0cf655be5c06342b817ed0f9595b69ffa2b12026e5353eed29dea88 \
--hash=sha256:51593d180cf6d179bde5c5d065bed81386b1f381656ae7d042b7ffc87a9895ad \
--hash=sha256:51afcfceb15597cf2635068e4ac9a56b2abde622edde17f37d85fd7b5306497a \
--hash=sha256:53e279950892dc102c6b4e52af03ae5ea92fac572a1ddab78ca73a997f62b69f \
--hash=sha256:55d16b1ef3ee0958d893a977b19777887e546c9954ea81b200c3301a864013f2 \
--hash=sha256:5dd9bda1c12b4162f6ff568eeb5e0ff956c28d14406e875cfe8a63a2d414ff20 \
--hash=sha256:5fe002589592ed749ce77fe0695fcbd3500dd61d7d6db5858a7544c612fa8e45 \
--hash=sha256:5fe939deeb161024a6be98229c953b6591fef1f41214497a78fe793a244c017f \
--hash=sha256:693c99b49bd37d0d096e4334c10232c77248c415b98d35236094cdf96d57258b \
--hash=sha256:76de83fbd91ac49c0feaaa983d0748fd7a53176afac5fb3bf7478d244f0eb527 \
--hash=sha256:79bf008d1f9af6071c797ad133e39915dfee7614f18f18f4db9072eb715064a3 \
--hash=sha256:804728ce710890870f3aaa344b2e161172d258d768ac139d02cfd9092d0d94e6 \
--hash=sha256:8921d58f426793c5f1b47f0b59575780de9a095214958d0eb37d909593db8367 \
--hash=sha256:8df2de9102026855887e4587084f6eabd80ed0f345b8ad8a7ac27ab9bf4723e0 \
--hash=sha256:9cb3cb952cf5a8abd50c782a98a89d71699715e802fe349704b47f2425b42a94 \
--hash=sha256:9dde0a357190eb3b1da1bb9ab750e9c85cba82ca5977aa0836cbb94e92611239 \
--hash=sha256:9ebcdd5519be9b652a46f507817a74591774fc3d6923ac364e4dfa64e36b291b \
--hash=sha256:a0b1a59e3a089064a0ec309e9428c8e3ae4e161419d20ac33600767e83fc658a \
--hash=sha256:a255449073358275b64b67d3f595f268bbef70e72b6edb65e0c70c735bf739c9 \
--hash=sha256:a8f40ea47330e71b594a7e246898f93177c259490c63183dbaf9e571d71ed9a5 \
--hash=sha256:ac02b07824d4d1001bd4367599f839c19cb171924c796e52c23508ac14c2c0cc \
--hash=sha256:aed8db4f6d71c51efb89530e12d9464e7bf2923d46c3205dc794a2a93f8c0648 \
--hash=sha256:b8f852c65863251b9e3a1b8c150ce21e59b522dbb6a7d4bc80e680d38388e986 \
--hash=sha256:be224a65493ec5b74a158ff22a5522ce4a5ca1e543c647a3a4730d4a09e5f959 \
--hash=sha256:ca83d00d9e69cd5eb63f2e69c3a5a59e0cecae5ae14c6ae0b35830fe3b37bad0 \
--hash=sha256:cbf74a81765ee67413503ca6e26dcc4f6f5a519822436cc0a1b97aab6c1b8a17 \
--hash=sha256:d63ae8f6481fec907ac0f588eee8a90aefde112c633131fe540e5711ddbb5a4e \
--hash=sha256:e22dfed744bd4002e909464cb23d2f0b05c6f3113a79ef2e9864a53db737c733 \
--hash=sha256:e2ca8fd1b6b4b82a1c4cb02841d0837e3c12336c2e24b520ab8ab3b969733d8f \
--hash=sha256:e74591e283fe6eb956416c929eb58262a719fe0311fd9054c62c3350ed8760d8 \
--hash=sha256:f74455bb086a85d5e81246412602aaa97ed095e504cd40dd261ef50be42205bf \
--hash=sha256:fb4b9672d389c738b175c4166e78310f8a70358886aacd9173ee03a85ffdc671 \
--hash=sha256:fc3ed7ebd2a8c96f5b166de0ab9b624996bef3b07bbeb19364dfb78222c22c80 \
--hash=sha256:fd3718b960d0b5dd213cdf03f3bcb7000e69dda0de8b956061947ff6bcff5558 \
--hash=sha256:ff838d62ec1bfce4f9ba7fa16f4a7b554cd8d0c299e6be37502161a660c84eef
# via secretstorage
docutils==0.23 \
--hash=sha256:25d013af9bf23bc1c7b2b093dff4208166c53a94786c9e447808335ef1185fea \
--hash=sha256:746f5060322511280a1e50eb76846ed6bf2342984b2ac04dc42caa1a8d78799e
# via readme-renderer
id==1.6.1 \
--hash=sha256:d0732d624fb46fd4e7bc4e5152f00214450953b9e772c182c1c22964def1a069 \
--hash=sha256:f5ec41ed2629a508f5d0988eda142e190c9c6da971100612c4de9ad9f9b237ca
# via twine
idna==3.19 \
--hash=sha256:5e0811a4383b21dc5838069f801c4fb62113b7447663d2530d2bd6e77b49bf15 \
--hash=sha256:815e7be7a7806d54abb586dc943addc79e8b2ee16915059658cbeff4b1b43bf4
# via requests
importlib-metadata==9.0.1 \
--hash=sha256:ab830580bc0ef3db61ce8fae716389e5462b67e033018bab6d8f80ef17172f99 \
--hash=sha256:bba5600596a7e21f3eef53281cf28d6a5195634d2f2b78ff9501a3272c6eaab0
# via keyring
jaraco-classes==3.4.0 \
--hash=sha256:47a024b51d0239c0dd8c8540c6c7f484be3b8fcf0b2d85c13825780d3b3f3acd \
--hash=sha256:f662826b6bed8cace05e7ff873ce0f9283b5c924470fe664fff1c2f00f581790
# via keyring
jaraco-context==6.1.2 \
--hash=sha256:bf8150b79a2d5d91ae48629d8b427a8f7ba0e1097dd6202a9059f29a36379535 \
--hash=sha256:f1a6c9d391e661cc5b8d39861ff077a7dc24dc23833ccee564b234b81c82dfe3
# via keyring
jaraco-functools==4.6.0 \
--hash=sha256:880c577ec9720b3a052d5bc611fb9f2269b3d87902ef42440df443b88e443280 \
--hash=sha256:99e3dc0060c5cbe8fcd1cdb36258e2a65ca40f1566b2033b12abb1bb44dd3c30
# via keyring
jeepney==0.9.0 \
--hash=sha256:97e5714520c16fc0a45695e5365a2e11b81ea79bba796e26f9f1d178cb182683 \
--hash=sha256:cf0e9e845622b81e4a28df94c40345400256ec608d0e55bb8a3feaa9163f5732
# via
# keyring
# secretstorage
keyring==25.7.0 \
--hash=sha256:be4a0b195f149690c166e850609a477c532ddbfbaed96a404d4e43f8d5e2689f \
--hash=sha256:fe01bd85eb3f8fb3dd0405defdeac9a5b4f6f0439edbb3149577f244a2e8245b
# via twine
markdown-it-py==4.2.0 \
--hash=sha256:04a21681d6fbb623de53f6f364d352309d4094dd4194040a10fd51833e418d49 \
--hash=sha256:9f7ebbcd14fe59494226453aed97c1070d83f8d24b6fc3a3bcf9a38092641c4a
# via rich
mdurl==0.1.2 \
--hash=sha256:84008a41e51615a49fc9966191ff91509e3c40b939176e643fd50a5c2196b8f8 \
--hash=sha256:bb413d29f5eea38f31dd4754dd7377d4465116fb207585f97bf925588687c1ba
# via markdown-it-py
more-itertools==11.1.0 \
--hash=sha256:48e8f4d9e7e5878571ecf6f2b4e57634f93cd474cc8cfbd2376f2d11b396e30d \
--hash=sha256:4b65538ae22f6fed0ce4874efd317463a7489796a0939fa66824dd542125a192
# via
# jaraco-classes
# jaraco-functools
nh3==0.3.7 \
--hash=sha256:157ec1eb7a62f3d9a7badb8d82d89aa810e3e24e097eedfa481a25d0c8a99877 \
--hash=sha256:15f5fbf090f5c88d61c820e1fc1fceecb6520cca9fe85649c06b57ef9dc9ff62 \
--hash=sha256:18f4278ecd157d43cb35acd5aae9f35cfa79f546b4922bd86536adc0f6312102 \
--hash=sha256:19f288c938ec6eef1f5d2c6cab47838e71fef8097e1c1233802be5a6230ba086 \
--hash=sha256:4968fe8d2db97c6f047659bf46a449fd8ec377f44ebf3e0a1b96c0d3a333ae32 \
--hash=sha256:5ffdfcb9a686ffb12765376bcfb6b5b55728516d3c0ee317d29982381ded3df8 \
--hash=sha256:614dac4a4c36ad084e78447d16fe898dedd762e354a7ab9cda2984e82f67883d \
--hash=sha256:618e3059caf41ccdf5dcccb3fa9df4cf6e4efe23d1382a8bbfca272a8a4f8bfc \
--hash=sha256:6698a822132beedab80f131c08d8d0ac5a178ddeb488d02ca4b67716ecfac7af \
--hash=sha256:6c3aa50eb26e9228238271db9f983cbc3b006dfbfeca2d4dc34c33ddc6ac5ea5 \
--hash=sha256:6e4280115d44c3b278eef712a86748c1a723105cd79feec46952383117ab4e59 \
--hash=sha256:70f5ac8626e899a4bab0ef74ca2f5bd602f49c7b739e6e5026b4afc6d63dac42 \
--hash=sha256:71860d01c16f4d8c72e334e0674beb2b0899dbd0bf760de18932ef4390303848 \
--hash=sha256:808def0c8c07843e6e50dc84f532457bfa2cfd17417b219a5d9e7c773709331a \
--hash=sha256:874b7d67a067bd29a59223f6270fc30da4edd8e6d87fd219fc93bcbaa662c946 \
--hash=sha256:91a4dab4e94d9fc54b9f67b1adfb23e81fab7ab43f33c3b8c97be9aa38f789ba \
--hash=sha256:94fd6e59553fbb9ffd8ba71bbd5a54e3126ba01799a097ae30d5341d750bc6ac \
--hash=sha256:9b7279d43323a25225df23576af6594a16693f61431170848b8b2ac21ad4f174 \
--hash=sha256:bc42bb1193c1e28a1e74c2cabaca178e118a7103e8832699fef8a2b3e2496493 \
--hash=sha256:be53a4825585f701955cb9baf49f478f56eb81e20294329fe4bc689dd5dd81fa \
--hash=sha256:d56e76bd3cadb09b6b0cef364850811663734b348a25f5f587a2819c495367bd \
--hash=sha256:de2b2aab32ea303405debefdcfc58043d3e635fa3f67b9eb140d2b0e0c0d2563 \
--hash=sha256:e8fd1ab205258b29254f72db377d99e2c96aa7653ef3b015ccab0420b094b506 \
--hash=sha256:eae64328e46a25785535afcb6885b6f182ecaf5ee8c88f8c075422db8aacc65b \
--hash=sha256:f04b7d333b27f13ca439da3cf1c75c2fba34f104969f6ce4ac8e7079699c2f4a \
--hash=sha256:f266d3f1b3647449923a8e406524632220dd5d8b647078dfe45b885d33d10479 \
--hash=sha256:fd4a70efb45d5372174f718878eb7a35c12677626a63b2f103b23b833457dcac
# via readme-renderer
packaging==26.3 \
--hash=sha256:94edc256424af38762eb31306eed28beb9f0efc50a8837492c9d6fd6004aed79 \
--hash=sha256:d7193f7c8e4e93f444fde0262bf90af30e16fa0ad0ad44cb553c87339b23cd1c
# via twine
pycparser==3.0 \
--hash=sha256:600f49d217304a5902ac3c37e1281c9fe94e4d0489de643a9504c5cdfdfc6b29 \
--hash=sha256:b727414169a36b7d524c1c3e31839a521725078d7b2ff038656844266160a992
# via cffi
pygments==2.21.0 \
--hash=sha256:2363c69b61c4a97c838da3b130dcd6468f4848992b21a82f2a63ec34377137d9 \
--hash=sha256:610ca751c9bc2492b38eb9a38a7fbc93edbbb2d7182edaf34e66ae493dee5c8c
# via
# readme-renderer
# rich
readme-renderer==46.0 \
--hash=sha256:af3e964914f6310a33ff67b72a4bdd940bed8d7c3bdecd2d14f40edf284bfe90 \
--hash=sha256:d0dae1f74bb273b534770cb4cccb6bb78735540afdb03c2146f4e19dcd412560
# via twine
requests==2.34.2 \
--hash=sha256:2a0d60c172f83ac6ab31e4554906c0f3b3588d37b5cb939b1c061f4907e278e0 \
--hash=sha256:f288924cae4e29463698d6d60bc6a4da69c89185ad1e0bcc4104f584e960b9ed
# via
# requests-toolbelt
# twine
requests-toolbelt==1.0.0 \
--hash=sha256:7681a0a3d047012b5bdc0ee37d7f8f07ebe76ab08caeccfc3921ce23c88d5bc6 \
--hash=sha256:cccfdd665f0a24fcf4726e690f65639d272bb0637b9b92dfd91a5568ccf6bd06
# via twine
rfc3986==2.0.0 \
--hash=sha256:50b1502b60e289cb37883f3dfd34532b8873c7de9f49bb546641ce9cbd256ebd \
--hash=sha256:97aacf9dbd4bfd829baad6e6309fa6573aaf1be3f6fa735c8ab05e46cecb261c
# via twine
rich==15.0.0 \
--hash=sha256:33bd4ef74232fb73fe9279a257718407f169c09b78a87ad3d296f548e27de0bb \
--hash=sha256:edd07a4824c6b40189fb7ac9bc4c52536e9780fbbfbddf6f1e2502c31b068c36
# via twine
secretstorage==3.5.0 \
--hash=sha256:0ce65888c0725fcb2c5bc0fdb8e5438eece02c523557ea40ce0703c266248137 \
--hash=sha256:f04b8e4689cbce351744d5537bf6b1329c6fc68f91fa666f60a380edddcd11be
# via keyring
twine==7.0.0 \
--hash=sha256:85cdb29c518efef867360ae4acd4b0dfd61c8654a22fca08e6f8539f05022177 \
--hash=sha256:b854164df26db268af05f49aa5c0344b10e27a494343ff05b1e0bad3b135f5a7
# via -r .github/requirements/twine.in
urllib3==2.7.0 \
--hash=sha256:231e0ec3b63ceb14667c67be60f2f2c40a518cb38b03af60abc813da26505f4c \
--hash=sha256:9fb4c81ebbb1ce9531cce37674bbc6f1360472bc18ca9a553ede278ef7276897
# via
# id
# requests
# twine
zipp==4.1.0 \
--hash=sha256:25ad4e16390cd314347dd8f1de67a2ac538ae658ed4ab9db16029c07c188e97f \
--hash=sha256:4cb57381f544315db7688e976e922a2b18cdb513d21cc194eb42232ba2a3e602
# via importlib-metadata
+1
View File
@@ -0,0 +1 @@
uv==0.12.1
+23
View File
@@ -0,0 +1,23 @@
# This file was autogenerated by uv via the following command:
# uv pip compile .github/requirements/uv-tool.in --generate-hashes --python-version 3.11 --python-platform linux -o .github/requirements/uv-tool.txt
uv==0.12.1 \
--hash=sha256:04290ea4001dca31ac8a8324113a4930dccad69ce35dbf6eaae307d54880890d \
--hash=sha256:153ec0959a15397514438aefc1d7cd04235f335dd6bb53ea0f9e6e82c5a49f03 \
--hash=sha256:173ee216f17d89fc39f65339d311a53584fc7de4918d27c0f3c7edafabc6b54d \
--hash=sha256:1de49d9b04438f1ad2f41a1441dbbe19e230b94fca56d632818cfaed69e03bfc \
--hash=sha256:1e8fd95fe98768e29436ad57f9ef7b68dc294b7b9862ef63396af8b15ab85e6c \
--hash=sha256:27211df9b277f440dea438a4e525ba40250fb721ad39b8927eefc2d91f9aea15 \
--hash=sha256:29399e1e73b67ed24abe82bc971aa4eb8419c4de804784290f39cf681f0b51ce \
--hash=sha256:2e9b0b86e180abc5968b979c6e25203b32e85969abb5083ee1e8b88a5aa98a76 \
--hash=sha256:3bd5db002adc763aa8d277f5b44f8d6e3fd82d20f2e51225b0bbdae1badc7259 \
--hash=sha256:41b8fc2335f682312a1ca39a7b4abfd6af800992065c663582ca3e4d51cf9258 \
--hash=sha256:5bd04849dd5346517cc4e57b4b3aa0b01c67c423878260c04f5893a038fe25b6 \
--hash=sha256:6f7e72543264d2420ebb2ddc84696a751af2d6c5910046b7666589118f47292b \
--hash=sha256:71f86410264c69a3e8acd18171897dd8ab1a13350cf40f718e4def5db2b724be \
--hash=sha256:76d87de420213ca92fa403e87023c4c7c6956c6726c6b96d91c42cfe620173a3 \
--hash=sha256:9331dda0dc4990512c232f86e1d3a7b83c13f459777fcc2bd46030911b40eaaa \
--hash=sha256:b255ac23958e45f39f9c7a4cd65890df5ef46f539a3b14de03bd296bbba9cb60 \
--hash=sha256:bd02f2da212e6a983115dc64a6fc94e9256c2d60e056d6b669de0a6025aaec05 \
--hash=sha256:e35e0030480a8c3bf8ecd87ae4a6f6a224009e15e96a6fbb3634ac11ab75d582 \
--hash=sha256:ead7ad064f291a5df358c3ffa8ffab347a32bd5a75a6a068ca22254c2539a829
# via -r .github/requirements/uv-tool.in
+76
View File
@@ -0,0 +1,76 @@
"""Drop checkov-suppressed results from its SARIF output before upload.
checkov's SARIF exporter includes every evaluated check as an ordinary
result, including ones it internally marked SKIPPED via an inline
`# checkov:skip=` comment or a `checkov.io/skipN` resource annotation - it
never uses SARIF's `suppressions` field, and never drops them. checkov's
JSON output *does* correctly record which checks were skipped, so this
cross-references the two: any SARIF result whose (check_id, file) pair
appears in the JSON's skipped_checks is removed before GitHub ever sees it.
Without this, every already-suppressed finding reopens as a brand new code
scanning alert on every run, forever (see #6035/#6036, #6112-6115,
#6128-6131 for the pattern this was chasing before this script existed).
Usage: filter_checkov_skipped.py <json_path> <sarif_in_path> <sarif_out_path>
"""
import json
import sys
def path_suffix(path: str, segments: int = 2) -> str:
"""Last N path segments, normalized to forward slashes, lowercased.
checkov's JSON file_path and SARIF artifactLocation.uri are relative to
different roots (the scanned directory vs. a temp helm-render dir), so
they can't be compared directly - but the last couple of segments
(e.g. "templates/service.yaml") are stable across both and specific
enough in practice to avoid cross-file collisions.
"""
normalized = path.replace("\\", "/").strip("/")
return "/".join(normalized.split("/")[-segments:]).lower()
def main() -> None:
json_path, sarif_in_path, sarif_out_path = sys.argv[1:4]
with open(json_path, encoding="utf-8") as f:
checkov_json = json.load(f)
if isinstance(checkov_json, dict):
checkov_json = [checkov_json]
skipped = set()
for block in checkov_json:
for check in block.get("results", {}).get("skipped_checks", []):
skipped.add((check["check_id"], path_suffix(check["file_path"])))
with open(sarif_in_path, encoding="utf-8") as f:
sarif = json.load(f)
removed = 0
for run in sarif.get("runs", []):
kept = []
for result in run.get("results", []):
rule_id = result.get("ruleId")
locations = result.get("locations") or [{}]
uri = (
locations[0]
.get("physicalLocation", {})
.get("artifactLocation", {})
.get("uri", "")
)
if (rule_id, path_suffix(uri)) in skipped:
removed += 1
continue
kept.append(result)
run["results"] = kept
with open(sarif_out_path, "w", encoding="utf-8") as f:
json.dump(sarif, f)
print(f"Removed {removed} checkov-suppressed result(s) from the SARIF before upload.")
if __name__ == "__main__":
main()
+31 -5
View File
@@ -28,11 +28,37 @@ jobs:
BENCHMARK_REAL_LIBS: "1"
run: |
python -m pip install --upgrade pip
pip install -e .
pip install -r benchmarks/requirements.txt
python -m spacy download en_core_web_sm
pip install rdflib neo4j faiss-cpu torch pyarrow pdfplumber python-pptx openpyxl lxml python-docx beautifulsoup4 chardet langdetect
pip install -r .github/requirements/bootstrap.txt --require-hashes
# --no-deps + a hash-pinned install of the same base dependency set
# (rather than a bare `pip install -e .`) so every fetched package
# is hash-verified (Scorecard Pinned-Dependencies); the local
# editable install itself has nothing to hash.
#
# --no-deps only skips *runtime* dependency resolution - `-e .`
# still does a PEP 517 build, which by default creates an isolated
# build env and fetches [build-system] requires (setuptools,
# wheel) completely outside any hash checking. Install
# pep517-build.txt (pins that exact build-system.requires) first
# and pass --no-build-isolation so pip reuses those hash-verified
# copies instead of fetching its own.
pip install -r .github/requirements/pep517-build.txt --require-hashes
pip install --no-deps --no-build-isolation -e .
pip install -r .github/requirements/base-deps.txt --require-hashes
# NOTE: benchmarks/ does not currently exist in this repo (neither
# requirements.txt nor benchmarks_runner.py below), so this job
# already fails on any real invocation - pre-existing, unrelated to
# this pinning change. The `pip install -r benchmarks/requirements.txt`
# step that used to be here is dropped rather than fixed: there's
# nothing to hash-pin without knowing what that file should
# contain, and an unpinned install here would just re-trip
# Scorecard's Pinned-Dependencies check for no real benefit, since
# the job can't run to completion regardless.
#
# `python -m spacy download en_core_web_sm` fetches an unpinned,
# unhashed wheel from spacy-models' GitHub releases - replaced with
# a hash-pinned direct-URL install of the same 3.8.0 model (matches
# the spacy==3.8.15 pinned in base-deps.txt) via benchmark-extra.txt.
pip install -r .github/requirements/benchmark-extra.txt --require-hashes
- name: Execute Benchmarks (Real Mode)
env:
+43 -6
View File
@@ -33,21 +33,58 @@ jobs:
- name: Install Explorer frontend dependencies
working-directory: explorer
run: npm ci
- name: Install Playwright Chromium
working-directory: explorer
run: npx playwright install --with-deps chromium
- name: Test Explorer frontend
working-directory: explorer
run: |
npm run test:graph-store
npm run test:graph-workspace
npm run test:plugin-registry
npm run test:deterministic-e2e
- name: Build Explorer frontend
working-directory: explorer
run: npm run build
- name: Install Explorer backend test dependencies
run: |
# Run the deterministic backend path before the all-extras CI
# environment is installed. The Explorer extra supplies the
# production API dependencies without importing optional vector
# providers such as Pinecone during test collection.
#
# --no-deps + a separate hash-pinned install (rather than the old
# `pip install -e ".[explorer]" pytest==9.1.1`) so every fetched
# package is hash-verified (Scorecard Pinned-Dependencies); the
# local editable install itself has nothing to hash.
# .github/requirements/explorer-extra-py311.txt is
# `uv pip compile pyproject.toml --extra explorer --python-version 3.11 --constraint requirements-ci.txt --generate-hashes`
# - regenerate it the same way if pyproject.toml's base/explorer
# deps change. Resolved specifically for this job's python 3.11
# (see the Dockerfile's explorer-extra-py313.txt for why this
# can't be shared with python 3.13: audioread needs extra
# standard-aifc/standard-sunau hashes only on 3.13+).
#
# --no-deps only skips *runtime* dependency resolution - `-e .`
# still does a PEP 517 build, which by default creates an isolated
# build env and fetches [build-system] requires (setuptools,
# wheel) completely outside any hash checking. Install
# pep517-build.txt (pins that exact build-system.requires) first
# and pass --no-build-isolation so pip reuses those hash-verified
# copies instead of fetching its own.
pip install -r .github/requirements/pep517-build.txt --require-hashes
pip install --no-deps --no-build-isolation -e .
pip install -r .github/requirements/explorer-extra-py311.txt --require-hashes
pip install -r .github/requirements/pytest-tool.txt --require-hashes
- name: Test deterministic Explorer backend path
run: |
pytest -q tests/explorer/test_explorer_deterministic_rendering_e2e.py
- name: Install pinned Python dependencies
run: |
pip install -r requirements-ci.txt
pip install -r requirements-ci.txt --require-hashes
- name: Verify requirements-ci.txt is up to date
run: |
pip install uv==0.12.1
pip install -r .github/requirements/uv-tool.txt --require-hashes
# Re-resolve with the committed file as a constraint: upstream package
# releases must NOT fail CI (deps only change when pyproject.toml
# changes intentionally). Compare only version lines (pkg==ver),
@@ -58,10 +95,10 @@ jobs:
diff \
<(grep -E '^[a-zA-Z0-9._-]+==' requirements-ci.txt | sed 's/ \\$//') \
<(grep -E '^[a-zA-Z0-9._-]+==' /tmp/requirements-ci-check.txt)
- run: pip install build
# wheel is build-time only (not in requirements-ci.txt) — install the
# same pinned version [build-system] declares so --no-isolation works.
- run: pip install wheel==0.48.0
# build is a dev-time dependency; wheel is build-time only (neither is
# in requirements-ci.txt) — install the same pinned versions
# [build-system] declares so --no-isolation works below.
- run: pip install -r .github/requirements/build-tools.txt --require-hashes
- name: Build package (no isolation — pinned deps)
run: python -m build --no-isolation
- name: Verify Explorer frontend is packaged
+10 -8
View File
@@ -10,13 +10,15 @@ on:
permissions:
contents: read
security-events: write
actions: read
jobs:
analyze:
name: Analyze Python
runs-on: ubuntu-latest
permissions:
contents: read
security-events: write # for github/codeql-action/upload-sarif below
actions: read # for github/codeql-action/init's CodeQL bundle cache lookup
steps:
- name: Checkout repository
@@ -32,7 +34,7 @@ jobs:
# meaningful state carried over from a failed attempt.
- name: Initialize CodeQL (attempt 1)
id: codeql-init-1
uses: github/codeql-action/init@ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd # v4
uses: github/codeql-action/init@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4
continue-on-error: true
with:
languages: python
@@ -42,7 +44,7 @@ jobs:
- name: Initialize CodeQL (attempt 2)
id: codeql-init-2
if: steps.codeql-init-1.outcome == 'failure'
uses: github/codeql-action/init@ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd # v4
uses: github/codeql-action/init@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4
continue-on-error: true
with:
languages: python
@@ -52,17 +54,17 @@ jobs:
- name: Initialize CodeQL (attempt 3)
id: codeql-init-3
if: steps.codeql-init-2.outcome == 'failure'
uses: github/codeql-action/init@ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd # v4
uses: github/codeql-action/init@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4
with:
languages: python
queries: security-and-quality
config-file: .github/codeql/codeql-config.yml
- name: Autobuild
uses: github/codeql-action/autobuild@ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd # v4
uses: github/codeql-action/autobuild@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4
- name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd # v4
uses: github/codeql-action/analyze@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4
with:
category: "/language:python"
upload: false
@@ -72,7 +74,7 @@ jobs:
# Uploads results only when Default Setup is not active.
# If Default Setup is still enabled, this step skips gracefully
# instead of failing the workflow with HTTP 409.
uses: github/codeql-action/upload-sarif@ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd # v4
uses: github/codeql-action/upload-sarif@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4
with:
sarif_file: ${{ steps.codeql.outputs.sarif-output }}
category: "/language:python"
+75
View File
@@ -0,0 +1,75 @@
name: Container Security Scan
on:
push:
branches: [main]
# Mirrors .dockerignore's opt-in list exactly - anything not listed there
# can't reach the build context, so it can't change the built image.
paths:
- 'Dockerfile'
- '.dockerignore'
- 'pyproject.toml'
- 'README.md'
- 'LICENSE'
- 'MANIFEST.in'
- '.github/requirements/explorer-extra-py313.txt'
- '.github/requirements/pep517-build.txt'
- 'semantica/**'
- 'integrations/**'
- 'explorer/**'
- '.github/workflows/container-scan.yml'
schedule:
- cron: '30 2 * * 1' # weekly, catches new CVEs published against the base image between pushes
workflow_dispatch:
permissions:
contents: read
jobs:
scan:
runs-on: ubuntu-latest
permissions:
contents: read
security-events: write # for github/codeql-action/upload-sarif below
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- name: Build image
run: docker build -t semantica:scan .
# Run Trivy as a digest-pinned image rather than the aquasecurity/trivy-action
# marketplace wrapper: the aquasecurity GitHub org has an IP allow list on its
# API that 403s verify-action-pins.sh's live tag->SHA check from Actions-runner
# IPs, and this repo already treats Trivy's action pin as a known past target
# for tag-repointing (see the LiteLLM/Trivy 2026 incident note above). Pulling
# by sha256 digest from Docker Hub is immutable and verifiable independently of
# GitHub's API, so it sidesteps both problems at once instead of carving a skip
# exception into the pin verifier for an org already flagged as higher-risk.
#
# Report-only for now: this is Trivy's first run against this image, so we
# don't yet know the CRITICAL/HIGH baseline. Findings still land in the
# Security tab either way. Once triaged, add `--exit-code 1` (like
# Safety/Bandit-HIGH in security-scan.yml) to make it a hard gate.
- name: Scan image for vulnerabilities (Trivy)
run: |
docker run --rm \
-v /var/run/docker.sock:/var/run/docker.sock \
-v "$PWD:/output" \
aquasec/trivy@sha256:62b1e65e8869bc4b4c6aa4fa2b21595256c7c2f6018a9d9ad61caf87187c1969 \
image --format sarif --output /output/trivy-results.sarif \
--severity CRITICAL,HIGH --ignore-unfixed semantica:scan
- name: Upload Trivy SARIF
if: always()
uses: github/codeql-action/upload-sarif@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4
with:
sarif_file: trivy-results.sarif
category: trivy-container
- name: Generate SBOM (Syft)
if: always()
uses: anchore/sbom-action@3ad7283483fc7af8ff2b4ea19663c2d5ca935e26 # v0.24.2
with:
image: semantica:scan
format: spdx-json
output-file: semantica-sbom.spdx.json
+25 -7
View File
@@ -28,12 +28,14 @@ on:
permissions:
contents: read
security-events: write
jobs:
MSDO:
# currently only windows-latest is supported
runs-on: windows-latest
permissions:
contents: read
security-events: write # for github/codeql-action/upload-sarif below
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
@@ -57,7 +59,7 @@ jobs:
# avoiding the guardian.cmd/checkov exit-code bug in the MSDO wrapper.
tools: eslint,templateanalyzer,terrascan
- name: Upload results to Security tab
uses: github/codeql-action/upload-sarif@ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd # v4
uses: github/codeql-action/upload-sarif@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4
with:
sarif_file: ${{ steps.msdo.outputs.sarifFile }}
@@ -66,7 +68,7 @@ jobs:
python-version: "3.12"
- name: Install Checkov
run: python -m pip install checkov==3.3.1
run: pip install -r .github/requirements/checkov.txt --require-hashes
- name: Run Checkov
shell: pwsh
@@ -74,15 +76,31 @@ jobs:
PYTHONUTF8: "1"
run: |
New-Item -ItemType Directory -Force reports | Out-Null
checkov --directory . --framework kubernetes helm dockerfile github_actions secrets bicep arm --soft-fail --output sarif --output-file-path reports/checkov.sarif
if (-not (Test-Path reports/checkov.sarif)) {
checkov --directory . --framework kubernetes helm dockerfile github_actions secrets bicep arm --soft-fail --output sarif --output json --output-file-path reports
if (-not (Test-Path reports/results_sarif.sarif)) {
$sarif = Get-ChildItem -Path reports -Recurse -Filter *.sarif | Select-Object -First 1
if ($null -eq $sarif) { throw "Checkov did not produce a SARIF file" }
Copy-Item $sarif.FullName reports/checkov.sarif
Copy-Item $sarif.FullName reports/results_sarif.sarif
}
if (-not (Test-Path reports/results_json.json)) {
$json = Get-ChildItem -Path reports -Recurse -Filter *.json | Select-Object -First 1
if ($null -eq $json) { throw "Checkov did not produce a JSON file" }
Copy-Item $json.FullName reports/results_json.json
}
# checkov's SARIF exporter includes checks it internally marked SKIPPED
# (via the inline `# checkov:skip=` comments / `checkov.io/skipN`
# annotations already on the Helm chart) as ordinary un-suppressed
# results - it never uses SARIF's own `suppressions` field, so GitHub
# opens a fresh alert for the same already-suppressed finding on every
# single run (see #6035/#6036, #6112-6115, #6128-6131). checkov's JSON
# output does correctly record the skip, so cross-reference it here
# instead of re-dismissing the same alerts by hand forever.
- name: Filter checkov's own suppressed checks out of the SARIF
run: python .github/scripts/filter_checkov_skipped.py reports/results_json.json reports/results_sarif.sarif reports/checkov.sarif
- name: Upload Checkov results to Security tab
uses: github/codeql-action/upload-sarif@ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd # v4
uses: github/codeql-action/upload-sarif@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4
if: always()
with:
sarif_file: reports/checkov.sarif
+1 -1
View File
@@ -65,4 +65,4 @@ jobs:
- name: Deploy to GitHub Pages
id: deployment
uses: actions/deploy-pages@cd2ce8fcbc39b97be8ca5fce6e763baed58fa128 # v5
uses: actions/deploy-pages@368f82528645a54fb793d4d04e342629a3f51346 # v5
+59
View File
@@ -0,0 +1,59 @@
name: Install Matrix
permissions:
contents: read
on:
schedule:
- cron: '0 6 * * 1' # weekly, catches upstream dependency breakage between releases
workflow_run:
# The Release workflow publishes the GitHub release *before* it uploads to
# PyPI (see release.yml), so triggering on `release: published` would race
# the PyPI upload and could pass by silently installing the prior version.
# workflow_run fires only after the whole Release workflow - including the
# PyPI publish step - has finished.
workflows: ['Release']
types: [completed]
workflow_dispatch:
jobs:
verify-install:
if: github.event_name != 'workflow_run' || github.event.workflow_run.conclusion == 'success'
name: pip install semantica (${{ matrix.os }}, py${{ matrix.python-version }})
runs-on: ${{ matrix.os }}
strategy:
fail-fast: false
matrix:
os: [ubuntu-latest, macos-latest, windows-latest]
python-version: ['3.9', '3.10', '3.11', '3.12']
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- name: Pin expected version for release-triggered runs
id: expected-version
if: github.event_name == 'workflow_run'
shell: bash
env:
EXPECTED_TAG: ${{ github.event.workflow_run.head_branch }}
run: |
expected="${EXPECTED_TAG#v}"
if [ -z "$expected" ]; then
echo "::error::Could not determine a release tag from the triggering workflow run (head_branch was empty)."
exit 1
fi
echo "constraint===$expected" >> "$GITHUB_OUTPUT"
- id: setup-semantica
uses: ./.github/actions/setup-semantica
with:
python-version: ${{ matrix.python-version }}
cache: 'pip'
version: ${{ steps.expected-version.outputs.constraint }}
- name: Smoke test import
shell: bash
run: |
python -c "
import semantica
print('semantica', semantica.__version__, 'installed and importable')
"
+26 -8
View File
@@ -16,7 +16,7 @@ jobs:
cancel-in-progress: false
permissions:
contents: write # for the GitHub Release
id-token: write # for PyPI Trusted Publishing (OIDC) and attestation signing
id-token: write # for PyPI Trusted Publishing (OIDC), attestation signing, and Sigstore
attestations: write # for SLSA build provenance
# If you add another job to this workflow, give it its own explicit
# `permissions:` block rather than relying on the workflow-level default
@@ -39,11 +39,11 @@ jobs:
# Install the pinned dependency set (with hashes) so the sdist/wheel
# build runs against the same versions CI tests against.
- name: Install pinned build dependencies
run: pip install -r requirements-ci.txt
- run: pip install build
# wheel is build-time only (not in requirements-ci.txt) — install the
# same pinned version [build-system] declares so --no-isolation works.
- run: pip install wheel==0.48.0
run: pip install -r requirements-ci.txt --require-hashes
# build is a dev-time dependency; wheel is build-time only (neither is
# in requirements-ci.txt) — install the same pinned versions
# [build-system] declares so --no-isolation works below.
- run: pip install -r .github/requirements/build-tools.txt --require-hashes
- name: Build package (no isolation — pinned deps)
run: python -m build --no-isolation
- name: Verify Explorer frontend is packaged
@@ -63,11 +63,29 @@ jobs:
print("Explorer frontend is packaged")
PY
- name: Verify PyPI long-description will render
run: |
pip install -r .github/requirements/twine.txt --require-hashes
twine check dist/*
- name: Attest build provenance
uses: actions/attest-build-provenance@4d101475d8b20a2381f78447822ac1eab6504dd8 # v4
with:
subject-path: 'dist/*'
- uses: softprops/action-gh-release@3d0d9888cb7fd7b750713d6e236d1fcb99157228 # v3
# attest-build-provenance publishes to the GH attestations API only, which
# OpenSSF Scorecard's Signed-Releases check does not inspect - it looks for
# signature files attached as release assets. Sign here too so
# `dist/*.sigstore.json` bundles ship alongside the wheel/sdist on the
# GitHub Release itself.
- name: Sign artifacts with Sigstore
uses: sigstore/gh-action-sigstore-python@790bc6befb9d733738f18d8f895854b453640ec9 # v3.5.0
with:
files: dist/*
inputs: |
dist/*.whl
dist/*.tar.gz
- uses: softprops/action-gh-release@efb35369e0ad2afab669f228072c1b0d510eae64 # v3.0.3
with:
files: |
dist/*.whl
dist/*.tar.gz
dist/*.sigstore.json
- uses: pypa/gh-action-pypi-publish@dc37677b2e1c63e2034f94d8a5b11f265b73ba33 # release/v1
+45
View File
@@ -0,0 +1,45 @@
name: Scorecard supply-chain security
permissions: read-all
on:
branch_protection_rule:
schedule:
- cron: '30 1 * * 6' # weekly
push:
branches: [main]
jobs:
analysis:
name: Scorecard analysis
runs-on: ubuntu-latest
permissions:
security-events: write # to upload SARIF results
id-token: write # to publish results and get a badge
contents: read
actions: read # to detect GitHub Actions workflows
steps:
- name: Checkout code
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
with:
persist-credentials: false
- name: Run analysis
uses: ossf/scorecard-action@2d1146689b8cda280b9bc96326124645441f03bc # v2.4.4
with:
results_file: results.sarif
results_format: sarif
publish_results: true
- name: Upload artifact
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
with:
name: SARIF file
path: results.sarif
retention-days: 5
- name: Upload to code-scanning
uses: github/codeql-action/upload-sarif@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4
with:
sarif_file: results.sarif
+168 -41
View File
@@ -3,6 +3,7 @@ name: Security Scan
on:
schedule:
- cron: '30 1 * * 1,4' # Mon/Thu 7 AM IST
workflow_dispatch:
push:
branches: [main]
paths-ignore:
@@ -44,46 +45,101 @@ jobs:
- name: Install dependencies
run: |
python -m pip install --upgrade pip
# Install the pinned dependency set FIRST so Safety scans Semantica's
# exact CI/release dependency tree (requirements-ci.txt is generated
# from pyproject.toml extras, so this covers the project's real deps).
pip install -r requirements-ci.txt
# Tooling AFTER the pinned set: installing safety/bandit/semgrep/jq
# first lets the pinned requirements overwrite their transitive deps
# (e.g. rich), which breaks the safety CLI at runtime.
pip install safety bandit semgrep jq
pip install -r .github/requirements/bootstrap.txt --require-hashes
# Install the pinned dependency set FIRST so pip-audit scans
# Semantica's exact CI/release dependency tree (requirements-ci.txt
# is generated from pyproject.toml extras, so this covers the
# project's real deps).
pip install -r requirements-ci.txt --require-hashes
# Tooling AFTER the pinned set: installing it first would let the
# pinned requirements overwrite the tooling's own transitive deps.
pip install -r .github/requirements/pip-audit.txt --require-hashes
pip install -r .github/requirements/security-scan-tools.txt --require-hashes
- name: Run Safety Check (Package Vulnerabilities)
- name: Run pip-audit (Package Vulnerabilities)
continue-on-error: true
run: |
# NOTE: Safety 3.x repurposed --output to select a console format
# (json/text/screen/...), not a file path. Writing JSON to a file
# now requires --save-json; the previous `--output safety-report.json`
# usage was silently invalid and never produced a report.
safety check --save-json safety-report.json || true
# Keep publishing reports and the PR comment even when the audit
# gate fails. The final gate below preserves the failure status.
echo 'AUDIT_SCAN_STATUS=failed' >> "$GITHUB_ENV"
# Guard 1: fail loudly if Safety exited before writing a report at all
# (network error, API auth failure, tool crash). Without this check a
# missing or empty file causes jq to fall back to "0", making a broken
# Same dependency tree Safety used to scan, and the same tool and
# invocation already proven reliable in security.yml.
pip-audit -r requirements-ci.txt --format=json --output=pip-audit-report.json || true
# Guard 1: fail loudly if pip-audit exited before writing a report
# at all (network error, tool crash). Without this check a missing
# or empty file causes jq to fall back to "0", making a broken
# scanner indistinguishable from a clean scan.
if [ ! -s safety-report.json ]; then
echo "::error::Safety scan produced no report (safety-report.json is missing or empty). Treating as failure — check for network errors, API auth failures, or Safety crashes in the logs above."
if [ ! -s pip-audit-report.json ]; then
echo "::error::pip-audit produced no report (pip-audit-report.json is missing or empty). Treating as failure — check for network errors or pip-audit crashes in the logs above."
exit 1
fi
# Guard 2: fail closed when the report doesn't have the shape the
# checks below assume: a non-empty dependencies array, each entry
# either carrying an array-valued vulns field or being a dependency
# pip-audit couldn't resolve/audit, which it reports as
# {"name": ..., "skip_reason": ...} with no vulns field at all
# (see pip_audit._format.json.JsonFormat._format_dep). That's a
# normal, documented report shape, not a malformed one — treating
# it as invalid would fail the whole job over a single unauditable
# package, the same kind of scan-unrelated CI break this migration
# away from Safety was meant to fix.
if ! jq -e '
(.dependencies | type == "array" and length > 0)
and all(.dependencies[]; type == "object" and ((.vulns | type == "array") or (.skip_reason | type == "string")))
' pip-audit-report.json >/dev/null 2>&1; then
echo "::error::pip-audit report has an invalid dependency structure. Expected a non-empty dependencies array where every entry has either a vulns array or a skip_reason. Treating as failure."
exit 1
fi
echo "Checking for package vulnerabilities..."
# No || echo "0" fallback: if jq fails (malformed JSON, missing key,
# vulnerabilities:null) VULNS will be empty or "null" so guard 2 below
# catches it rather than silently treating the broken report as zero.
VULNS=$(jq '.vulnerabilities | length' safety-report.json 2>/dev/null)
# Guard 2 above already confirmed pip-audit-report.json is valid
# JSON with a well-shaped dependencies array, so this count is
# always a plain non-negative integer.
SKIPPED=$(jq '[.dependencies[] | select(has("skip_reason"))] | length' pip-audit-report.json)
if [ "$SKIPPED" -gt 0 ]; then
echo "⚠️ pip-audit could not audit $SKIPPED dependencies (see pip-audit-report.json for skip_reason):"
jq -r '.dependencies[] | select(has("skip_reason")) | " - \(.name): \(.skip_reason)"' pip-audit-report.json
fi
# Guard 2: ensure VULNS is a non-negative integer before the -gt
# Vulnerability IDs reviewed and accepted as non-actionable for this
# project. Empty for now: pip-audit's OSV-backed database doesn't
# currently carry either of the findings Safety used to flag here
# (cuda-toolkit CVE-2025-33228, torchvision CVE-2026-65918), so
# there's nothing to exclude. Left in place so a future finding can
# be added the same way without restructuring this step - see git
# history on this file for the reasoning behind past entries.
IGNORED_VULN_IDS=""
# Exported so the "Comment PR with Security Results" step below can
# apply the same exclusion list to the raw report - it reads
# pip-audit-report.json independently in JS, so without this the PR
# comment would show an accepted finding as live even though this
# gate correctly treats it as non-actionable.
echo "IGNORED_VULN_IDS=$IGNORED_VULN_IDS" >> "$GITHUB_ENV"
# No []? / || echo "0" fallback: if jq fails (malformed JSON) VULNS
# will be empty or "null" so Guard 3 below catches it rather than
# silently treating the broken report as zero.
# `.vulns // []` guards against skipped dependencies, which carry
# no vulns field at all (see the skip_reason handling above) -
# without the fallback, iterating `null[]` raises inside jq and
# this whole computation silently evaluates to empty.
VULNS=$(jq --arg ignored "$IGNORED_VULN_IDS" '
($ignored | split(",") | map(select(length > 0))) as $ignore_list
| [.dependencies[] | (.vulns // [])[] | select(.id as $id | ($ignore_list | index($id)) | not)]
| length
' pip-audit-report.json 2>/dev/null)
# Guard 3: ensure VULNS is a non-negative integer before the -gt
# comparison. "null" (missing/null key) or "" (jq parse failure) would
# cause bash's -gt to throw an arithmetic error and fall through to the
# success branch — the same silent-pass bug as a missing file.
if ! [[ "$VULNS" =~ ^[0-9]+$ ]]; then
echo "::error::Safety report exists but 'vulnerabilities' is missing or non-numeric (got: '${VULNS}'). The report may be malformed or Safety may have written an error-only JSON. Treating as failure."
echo "::error::pip-audit report exists but dependency vulnerabilities are missing or non-numeric (got: '${VULNS}'). The report may be malformed or contain an error-only JSON response. Treating as failure."
exit 1
fi
@@ -92,12 +148,18 @@ jobs:
echo "CI will fail to prevent merging of vulnerable dependencies"
echo ""
echo "Vulnerability details:"
jq -r '.vulnerabilities[] | "- \(.package_name)==\(.analyzed_version): \(.vulnerability_id) (\(.CVE // "no CVE assigned"))"' safety-report.json || true
jq --arg ignored "$IGNORED_VULN_IDS" -r '
($ignored | split(",") | map(select(length > 0))) as $ignore_list
| .dependencies[] as $dependency
| ($dependency.vulns // [])[] | select(.id as $id | ($ignore_list | index($id)) | not)
| "- \($dependency.name)==\($dependency.version): \(.id)"
' pip-audit-report.json || true
exit 1
else
echo "✅ No security vulnerabilities found"
echo "✅ No actionable security vulnerabilities found${IGNORED_VULN_IDS:+ (ignored: $IGNORED_VULN_IDS)}"
echo 'AUDIT_SCAN_STATUS=passed' >> "$GITHUB_ENV"
fi
- name: Run Bandit (Code Security Linter)
run: |
bandit -r semantica/ -f json -o bandit-report.json || true
@@ -135,17 +197,18 @@ jobs:
fi
- name: Upload Security Reports
if: always()
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
with:
name: security-reports
retention-days: 14
path: |
safety-report.json
pip-audit-report.json
bandit-report.json
semgrep-report.json
- name: Comment PR with Security Results
if: github.event_name == 'pull_request'
if: always() && github.event_name == 'pull_request'
uses: actions/github-script@3a2844b7e9c422d3c10d287c895573f7108da1b3 # v9
with:
script: |
@@ -168,6 +231,12 @@ jobs:
}
const items = parse(data);
if (items === null) {
return [
'### ' + title,
'⚠️ Invalid report structure in ' + reportPath + ' — check the job logs.',
].join('\n');
}
if (items.length === 0) {
return [`### ${title}`, `✅ No findings.`].join('\n');
}
@@ -184,14 +253,64 @@ jobs:
return lines.join('\n');
}
const safetySection = renderSection(
'Safety — dependency vulnerabilities',
'safety-report.json',
(data) => (data.vulnerabilities || []).map(
(v) => `- \`${v.package_name}==${v.analyzed_version}\`: ${v.vulnerability_id}` +
(v.CVE ? ` (${v.CVE})` : '') + ` — ${v.advisory || 'no advisory text'}`
)
);
// Mirrors the shell step's own IGNORED_VULN_IDS (passed through
// $GITHUB_ENV) so an accepted, non-actionable CVE that the CI
// gate already excluded doesn't reappear here as a live finding -
// this reads the same raw, unfiltered pip-audit-report.json.
const ignoredVulnIds = (process.env.IGNORED_VULN_IDS || '')
.split(',')
.map((id) => id.trim())
.filter(Boolean);
// A dependency pip-audit couldn't resolve/audit is reported as
// {"name": ..., "skip_reason": ...} with no vulns field at all
// (see pip_audit._format.json.JsonFormat._format_dep) - that's a
// normal report shape, not a malformed one, so it must not be
// treated as an invalid dependency below.
const isSkipped = (dependency) => typeof dependency.skip_reason === 'string';
let skippedDeps = [];
try {
const auditData = JSON.parse(fs.readFileSync('pip-audit-report.json', 'utf8'));
skippedDeps = (auditData.dependencies || []).filter(
(dependency) => dependency && typeof dependency === 'object' && isSkipped(dependency)
);
} catch (e) {
// Unreadable/unparseable report - renderSection's own
// report-missing branch below surfaces this.
}
const pipAuditSection = renderSection(
'pip-audit — dependency vulnerabilities',
'pip-audit-report.json',
(data) => {
if (
!Array.isArray(data.dependencies) ||
data.dependencies.length === 0 ||
data.dependencies.some(
(dependency) =>
!dependency ||
typeof dependency !== 'object' ||
(!Array.isArray(dependency.vulns) && !isSkipped(dependency))
)
) {
return null;
}
return data.dependencies.flatMap((dependency) =>
(dependency.vulns || [])
.filter((vulnerability) => !ignoredVulnIds.includes(vulnerability.id))
.map(
(vulnerability) => `- \`${dependency.name}==${dependency.version}\`: ${vulnerability.id}` +
(vulnerability.fix_versions?.length ? ` (fixed by ${vulnerability.fix_versions.join(', ')})` : '')
)
);
}
) + (ignoredVulnIds.length
? `\n\n_Excluded as accepted, non-actionable findings: ${ignoredVulnIds.join(', ')} — see the workflow file's inline comments for why._`
: '') + (skippedDeps.length
? `\n\n_Could not be audited: ${skippedDeps.map((d) => `\`${d.name}\` (${d.skip_reason})`).join(', ')}_`
: '');
const banditSection = renderSection(
'Bandit — HIGH-severity code issues',
@@ -212,7 +331,7 @@ jobs:
const comment = [
'# 🔒 Security Scan Results',
'',
safetySection,
pipAuditSection,
'',
banditSection,
'',
@@ -222,7 +341,7 @@ jobs:
'',
'*This security scan runs automatically on source-code PRs and bi-weekly (skipped for doc/markdown-only changes).*',
'',
'📊 **Security Policy**: CI fails on Safety vulnerabilities and Bandit HIGH-severity findings. Semgrep findings above are informational and do not block merge.',
'📊 **Security Policy**: CI fails on pip-audit vulnerabilities and Bandit HIGH-severity findings. Semgrep findings above are informational and do not block merge.',
].join('\n');
try {
@@ -237,3 +356,11 @@ jobs:
console.log('⚠️ Could not post security comment:', error.message);
console.log('📋 Security scan results saved to artifacts');
}
- name: Enforce Audit Gate
if: always()
run: |
if [ "${AUDIT_SCAN_STATUS:-failed}" != "passed" ]; then
echo "::error::pip-audit scan failed. See the pip-audit output and uploaded reports above."
exit 1
fi
-42
View File
@@ -1,42 +0,0 @@
name: Security
on:
schedule:
- cron: '0 0 * * 1'
workflow_dispatch:
pull_request:
branches: [main]
paths:
- 'pyproject.toml'
- 'requirements-ci.txt'
- '.github/workflows/security.yml'
permissions:
contents: read
jobs:
audit:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7
with:
python-version: '3.11'
# Upgrade first: actions/setup-python's baked-in setuptools has been
# behind known-vulnerable floors before (e.g. PYSEC-2026-3447 /
# setuptools 75.1.0), so don't trust the preinstalled one.
- run: python -m pip install --upgrade pip setuptools
# Audit the pinned dependency set (requirements-ci.txt is compiled from
# pyproject.toml with --extra all — the same coverage as the [all]
# extra, minus the Linux-only gpu set — so this keeps scan parity with
# CI/release builds without a time-dependent resolution). This is the
# fix for PYSEC-2024-38 (#869): the bare-env job never had fastapi or
# python-multipart installed to look at.
- run: pip install -r requirements-ci.txt
# PR runs gate on findings, since they're scoped to actual
# pyproject.toml changes under review. The schedule/workflow_dispatch
# runs stay non-blocking until a full pass over pre-existing findings
# across the whole [all] tree has been done.
- run: pip install pip-audit
- run: pip-audit -r requirements-ci.txt
continue-on-error: ${{ github.event_name != 'pull_request' }}
BIN
View File
Binary file not shown.
+53
View File
@@ -9,6 +9,37 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## [Unreleased]
### Added
- **Salesforce ingestor** (#1240) by @Sameer6305
- New `SalesforceConnector` / `SalesforceData` / `SalesforceIngestor` (`semantica.ingest`, lazy export), following the same Connector + Data + Ingestor pattern already used for Snowflake/Databricks/SAP
- Auth covers both landscapes Salesforce actually uses: username + password + security token (SOAP login), session_id + instance_url (reusing an existing session), and username + consumer_key + private key (JWT Bearer); production and sandbox are selected via `domain`, and credentials can come from environment variables. Credential material is never intentionally written to logs, exceptions, or `repr()`
- `ingest_sobject()`, `ingest_query()`, `list_sobjects()`, `get_sobject_schema()`, `export_as_documents()` against standard sObjects, custom objects (`__c`), custom metadata (`__mdt`), platform events (`__e`), namespaced objects, and relationship-field traversal (e.g. `Owner.Name`); pagination follows `nextRecordsUrl`/`query_more()` and stops once a caller's `limit` is satisfied
- New `pip install semantica[db-salesforce]` extra (`simple-salesforce>=1.12.0`)
- New `tests/test_salesforce_ingestor.py`
- Docs: `docs/integrations/salesforce.md`
- **`ErasureCoordinator` completes the erasure workflow `purge_node()` only starts — the graph node was removed while the same content survived verbatim in `AgentMemory` and as an embedding** (closes #1018) by @pravit-amp
- New `semantica/context/erasure.py`, exporting `ErasureCoordinator` and `ErasureReceipt` from `semantica.context`. `purge_node()`/`purge_edge()` (#957) are graph-scope by design and their changelog entry documents this gap explicitly; the changelog also names GDPR Article 17 as the motivation, and an Article 17 erasure that removes the node while the content stays retrievable by similarity search is not an erasure — it is worse than not offering one, because `purge_node()` returns `True` and writes a tombstone attesting the content is gone
- The coordinator **composes** the existing public APIs — nothing in `context_graph.py` or `agent_memory.py` changes behaviorally, and `ContextGraph` keeps its documented graph-scope contract rather than acquiring references to `AgentMemory`/`vector_store` that would invert the dependency
- `erase_entity(entity_id, reason=..., at=..., vector_ids=...)` returns an `ErasureReceipt`; `erase_entities([...])` returns one receipt per entity, in order, so one entity's failure does not stop the rest
- **Honest partial reporting is the point.** Each store reports one of five statuses — `erased`, `not_found`, `not_configured` (store never bound; normal), `unsupported` (store cannot delete at all; retrying will not help), `failed` — and `receipt.complete` is `False` when any store reports `unsupported`/`failed`, with `receipt.incomplete_stores` naming them. A receipt reading `graph: erased, memory: 14 erased, vectors: unsupported on faiss` is actionable; a bare `True` is a compliance liability
- **Erasure runs outward-in: vectors → memory → graph.** The graph tombstone is the durable attestation that an erasure happened, so writing it first would let a crash mid-cascade leave a record claiming more than occurred. Erasing the graph last means a partial failure leaves the node present and the receipt incomplete — recoverable and honest; the reverse is neither
- **Partial failure is a result, not an exception**: a store that raises is recorded as `failed` (with the exception type) and the remaining legs still run, rather than aborting into a half-erased state with no record of which half
- **The memory sweep cannot be silently truncated.** `find_by_entity(entity_id, limit=10)` returned `results[:limit]`, so the obvious hand-rolled cascade erases the first ten items and reports success — an erasure check computed from a page already truncated by the very `limit` it was called with. The coordinator sweeps in pages until dry (deleting as it goes, so the next page is the remainder) rather than passing one large number that is only correct until someone exceeds it, then **re-queries once after the sweep** and reports `failed` with the residual count if anything survived. It also stops rather than spinning if `batch_delete` reports no progress on a non-empty page. Note `find_by_entity` returns items keyed `memory_id`, not `id`
- **`unsupported` vector backends are detected by probing, not by calling and catching.** `faiss_store.py`, `milvus_store.py` and `weaviate_store.py` expose no delete at all (FAISS cannot remove from a flat index without a rebuild), while the `VectorStore` facade declares `delete_vectors()` for *every* backend and only raises `NotImplementedError` once called — so probing the facade alone cannot tell a deletable backend from a delete-less one, and the coordinator looks at the backend it wraps. Probing also keeps a missing method distinguishable from an `AttributeError` raised *inside* a working one, which is exactly where guessing wrong produces a false clean bill of health. `NotImplementedError` at call time is still caught and reported as `unsupported`; a store returning `False` is reported as `failed`
- Backends are reached under either supported name — `delete_vectors(ids)` (pinecone/qdrant) or `delete(ids)` (pgvector/sqlite-vec) — and the receipt records which was used
- `vector_store` defaults to `memory.vector_store` when a memory is supplied, stays overridable for deployments binding a store the memory does not own, and accepts `False` to disable the vector leg. Vectors owned by memory items are removed by the memory leg's own `delete_memory()` cascade; the explicit vector leg covers entity-keyed embeddings written by something other than `AgentMemory`
- The receipt's `erased_at` is normalized through `ContextGraph`'s own temporal normalizer, so the receipt and the tombstone written by the same erasure cannot disagree about when it happened; an unparseable `at` is rejected before any store is touched rather than half way through the cascade
- `purge_node()`'s docstring now points at the coordinator, so callers reading the graph-scope caveat find the thing that completes the workflow
- New `tests/context/test_erasure_coordinator.py`: 48 tests against **real** `ContextGraph`/`AgentMemory` instances rather than mocks — the bug lives in the interaction between them, so mocking it away would test nothing. Covers the 25-items-on-one-entity regression that fails against a naive single `find_by_entity()` call, all three vector-backend shapes (`delete_vectors`/`delete`/neither) plus the facade-over-delete-less-backend shape, residual/no-progress/no-identifier memory failures, partial failure continuing the cascade, idempotency, receipt serialization, and `at` normalization
- Full `tests/context/` suite: 738 passed
- **Fixed during review** (Qodo): `erase_entity()` resolved `erased_at` up front but passed the caller's original `at` down to `purge_node()`, so on the default `at=None` path the coordinator and the graph each took their own `now()` and the receipt attested to a different instant than the tombstone it points at — breaking the one invariant this module states most loudly. The resolved timestamp is now passed to the graph. The existing test passed only because it supplied an explicit `at`, which hides the drift; a regression test now covers the `at=None` path that callers actually use
- **Fixed during review** (Qodo): the vectors leg treated any return value other than the literal `False` as success, but no in-repo backend returns a bool — Qdrant returns `{"status": <UpdateStatus>}` and Pinecone `{"deleted": True}`, so every dict was read as a success and the backend's own account of the delete was discarded. Delete results are now interpreted by shape (bool, dict with explicit failure markers, `None` for a void method, anything else at face value) and the backend payload is recorded in the receipt as `backend_result`, stringified so the receipt stays JSON-serializable as the audit record it is meant to be. Bool markers are matched by identity so a `0` count is not read as `False`, and string markers match as substrings so an enum rendering as `"UpdateStatus.FAILED"` is not read as a success
- **Fixed during review** (Qodo): the constructor's "at least one store" guard used `not vector_store`, rejecting a valid store whose `__bool__`/`__len__` makes an empty instance falsey, and reporting `vector_store=None` in the error when an object had been passed; it now distinguishes `None` (absent) from `False` (deliberately disabled) from any other value (provided), and echoes what it actually received
- **Fixed during review** (Qodo): `at` annotations accepted only `str`/`datetime` while the shared `ContextGraph` normalizer they delegate to also takes epoch seconds; widened to `int`/`float` with the docstrings updated, so the coordinator no longer advertises less than the graph API it wraps
- **Known limitation, unchanged by this PR**: erasure still cannot be *completed* on FAISS/Milvus/Weaviate — `delete_vectors()` is declared on the `VectorStore` facade (`vector_store.py:786`) but not implemented across the backend set, under at least three different names. That is worth its own issue; the coordinator ships reporting `unsupported` and starts reporting `erased` for those backends once it is fixed, with no API change here
## [0.6.7] - 2026-08-28
### Added
@@ -128,6 +159,19 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- **Also fixed, on the JSON-LD paths**: the first fix covered the Turtle, N-Triples and RDF/XML serializers, and left both JSON-LD writers interpolating the entity's own text into `f"semantica:entity/{text}"` and the endpoints into `f"semantica:rel/{source}_{target}"`. Three consequences, all live in 0.6.5: an entity whose text contained a space produced an invalid IRI, and a JSON-LD parser dropped that node in full rather than reporting it, so the entity disappeared from the export; every relationship carrying `source`/`target` rather than `source_id`/`target_id` minted the identical `semantica:rel/_`, collapsing all of them onto one node whose types and endpoints merged; and the JSON-LD `@id` disagreed with the Turtle IRI for the same entity, so the two serializations of one knowledge graph were two different graphs. Both JSON-LD writers now use `mint_entity_iri`/`mint_relationship_iri`, and `JSONExporter.export_entities`/`export_relationships` declare the `semantica` prefix their `@context` was already writing `semantica:entities` against — without it a processor reads that as an IRI in the scheme `semantica`, which is the original #1101 defect on a third path
- `tests/export/test_jsonld_iri_minting.py` parses each export with a real JSON-LD processor and asserts the entity survives, the relationships stay distinct, no term expands into the `semantica` scheme, and the JSON-LD `@id` equals the Turtle IRI
- 236 export and ontology tests pass
- **`semantica.evals` runner gains per-metric objectives** (#1091)
- `evaluate()` now accepts `config={"<evaluator>": {"objective": {"direction": "maximize"|"minimize", "threshold": X}}}` to override the evaluator's default pass verdict with a threshold; `{"objective": {"expect": bool}}` expresses a Boolean expectation
- `minimize` requires a `threshold` — omitting it or setting it to `None` raises `ValueError`; `maximize` without a threshold is a no-op (the evaluator's own verdict stands); `expect` cannot be combined with `direction`/`threshold`; invalid config raises `ValueError` before any evaluator runs
- Error metrics are never affected by objectives (error wins over fail)
- Backward compatible: no `objective` key → existing behavior unchanged
- New tests in `tests/evals/test_runner.py::TestObjective`
- **`semantica.evals` is now a fully implemented evaluation module** (was a "Coming Soon" stub in the package layout)
- `evaluate(cases, evaluators, config=None, target_fn=None)` runner with per-case `pass`/`fail`/`error` status and an aggregate `pass_rate`, using a registry of named evaluators (`list_evaluators()`)
- 10 built-in evaluators: `exact_match`, `regex_match`, `numeric_range`, `temporal_range`, `length_range`, `keyword_check`, `levenshtein` (edit-distance similarity), `rouge` (in-house token F1, no new dependencies), `llm_as_judge` (lazy: caller-supplied `judge_fn`), and `decision_scores` (composite over `semantica.context.Decision`)
- `decision_scores` validates field-level (expected outcome, confidence bounds, non-empty maker/reasoning/scenario) and governance-level (provenance record presence; opt-in `PolicyEngine.check_compliance`) checks, coercing dict inputs via `Decision(**actual)` and never crashing on malformed input; an interface slot for causal-chain/embedding checks is reserved and raises `NotImplementedError` (V2)
- `__version__` is `0.1.0`, and the module ships a usage guide at `semantica/evals/usage.md` with worked import/run/interpret examples
- `semantica.evals` is reachable through the root package lazy module proxy (`semantica.evals`)
- 99 unit tests in `tests/evals/` covering every evaluator, registry errors, runner aggregation, decision coercion, and per-metric objectives; `python -m pytest tests/evals -q` → 99 passed
- **First-class CrewAI integration** (#988, closes #962) by @Shindevrp
- New `pip install semantica[crewai]` extra (`crewai>=0.80.0`) — crewai core provides `BaseTool`/`BaseKnowledgeSource`, so `crewai-tools` is intentionally not included, and the extra is intentionally **not** part of the `all` bundle: crewai hard-requires `chromadb~=1.1.0`, which is affected by the unpatched pre-auth code-injection CVE-2026-45829 (see `integrations/crewai/README.md`)
@@ -213,6 +257,15 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
### Fixed
- **RETE engine matched every fact against every rule — `AlphaNode._matches()` and `BetaNode._can_join()` were placeholder stubs that always returned `True`** (closes #300)
- `semantica/reasoning/rete_engine.py` shipped a Rete network whose per-condition alpha test and cross-condition beta join were both `return True` stubs, so `match_patterns()` fired every rule for every fact regardless of predicate, arity, or shared-variable consistency
- New module-level `unify_condition()` reuses the regex-based approach from `Reasoner._match_pattern()`: a condition pattern like `Person(?x)` / `Parent(?x, ?y)` is compiled against a fact's `predicate(arg, ...)` string, `?var` becomes a named capture group, and a variable seen twice within one condition (e.g. `Loves(?x, ?x)`) becomes a backreference, so it only unifies when both positions hold the same value. Returns the bindings dict or `None`
- Reworked propagation to carry partial-match **tokens** instead of bare facts: a new `Token` dataclass bundles the accumulated `facts` with the consistent `bindings`. `AlphaNode` emits a single-fact token per match; `BetaNode.join()` merges a left token with a right token, concatenating their facts in condition order and returning the merged token only when shared variables agree (conflicting values → `None`, no join). Terminal activations carry the full fact list and accumulated bindings through to the emitted match
- This fixes a P1 chained-join defect: rules with three or more conditions (e.g. `Person(?x)`, `Parent(?x, ?y)`, `Located(?y, ?z)`) previously lost bindings and accumulated wrong facts at the third join, and a conflicting third condition could spuriously fire. Beta nodes now keep both `left_tokens` and `right_tokens` memories and join each new token against every token on the opposite side, so deep chains stay binding-consistent and third-level conflicts are correctly suppressed
- Fixed an adjacent network-topology bug surfaced by the above: newly created beta nodes were never appended to their input nodes' `children`, so tokens could not propagate; propagation was reworked to support chained joins and to thread bindings end-to-end
- Reconciled with the rule-actions/provenance layer (#1096) merged after this fix was opened: `execute_matches()` still dedupes and fires `Rule.actions`/legacy `handler` through a bound `Reasoner` via `_make_activation_key`, now sourced from the Token model's own `bindings` instead of the interim `_bindings_for_rule()` regex re-extraction, which is removed as redundant
- New `tests/reasoning/test_rete_engine.py`: `unify_condition` unit cases (single/multi variable, literal args, predicate mismatch, repeated-variable equality), alpha match/reject, beta consistent-join vs conflict-reject, end-to-end rules (single-condition fires only the matching fact; multi-condition join fires only on consistent bindings), and a `TestThreeConditionChain` suite (valid three-condition match, third-level conflict suppression, insertion-order independence, `Match.facts` complete and in condition order, multiple left tokens joining one right fact, parity against `Reasoner._match_rule()`, and `reset()` clearing all token memory)
- **KG provenance tests asserted on generated ID strings instead of stored records, and `kg_provenance.py` was missed by the `utcnow` sweep** (closes #946) by @pravit-amp
- The KG workflow and integration suites checked that a tracker call returned an ID matching a prefix (`assert cent_id.startswith("centrality_")`) without ever reading the record back, so an ID generator that returned a well-formed string and wrote nothing would have passed. Worse, some of those calls named tracker methods that do not exist anywhere in `semantica/` (`track_layer_analysis`, `track_centrality_score`), so the assertions were satisfied with no real interaction behind them
- Those tests now read provenance back through `get_provenance()` and assert on algorithm metadata, and call the methods that actually persist records. Verified by mutation rather than by a green run alone: neutering the manager's storage write (`self.storage.store(...)` → no-op) fails 10 tests
+20
View File
@@ -0,0 +1,20 @@
cff-version: 1.2.0
message: "If you use this software, please cite it as below."
title: "Semantica: Graph-Native Infrastructure for Context and Accountable AI Systems"
type: software
authors:
- name: "Semantica"
repository-code: "https://github.com/semantica-agi/semantica"
url: "https://getsemantica.ai"
license: MIT
version: 0.6.7
date-released: 2026-08-28
keywords:
- knowledge-graph
- context-graph
- ai-agents
- llm
- decision-intelligence
- provenance
- explainability
- graph-rag
+38 -4
View File
@@ -1,5 +1,5 @@
# syntax=docker/dockerfile:1
FROM node:26-alpine AS frontend-builder
FROM node:26-alpine@sha256:2d984a15c9b54fd0aeb608b8e0d0d83529eb34d2966db27a1fb4f1edc3d298a3 AS frontend-builder
WORKDIR /app
COPY explorer/package*.json ./explorer/
@@ -9,7 +9,18 @@ RUN npm ci
COPY explorer/ ./
RUN mkdir -p /app/semantica && npm run build
FROM python:3.13-slim AS runtime
# CVE-2026-14456 (OpenSSL QUIC-server DoS, flagged against this base image's
# openssl/libssl3t64/openssl-provider-legacy): the Debian fix
# (3.5.7-1~deb13u2) is only in trixie-proposed-updates as of this writing,
# not yet promoted to trixie-security, so there's no package to pin here
# today. Deliberately NOT running `apt-get upgrade` to chase it - that
# breaks build reproducibility (terrascan AC_DOCKER_0052) and still
# wouldn't reach a proposed-updates-only package. Once Debian ships the fix
# and rebuilds this tag, the docker Dependabot ecosystem in
# .github/dependabot.yml opens a PR bumping the digest pin above. Also: this
# image only serves plain HTTP via uvicorn and never opens a QUIC listener,
# so the bug isn't reachable here regardless.
FROM python:3.13-slim@sha256:7ce4b6dfe35e55397b7cda544f8a13f191b7ae28dc5aad71fe664dbc9bc2623f AS runtime
ENV PYTHONDONTWRITEBYTECODE=1 \
PYTHONUNBUFFERED=1 \
@@ -22,12 +33,35 @@ WORKDIR /app
RUN groupadd --system semantica \
&& useradd --system --gid semantica --home-dir /app --shell /usr/sbin/nologin semantica
COPY pyproject.toml README.md LICENSE MANIFEST.in ./
COPY pyproject.toml README.md LICENSE MANIFEST.in \
.github/requirements/explorer-extra-py313.txt .github/requirements/pep517-build.txt ./
COPY semantica/ ./semantica/
COPY integrations/ ./integrations/
COPY --from=frontend-builder /app/semantica/static ./semantica/static
RUN pip install --no-cache-dir ".[explorer]" \
# explorer-extra-py313.txt is `uv pip compile pyproject.toml --extra explorer
# --python-version 3.13 --constraint requirements-ci.txt --generate-hashes`
# (see ci.yml's explorer-extra-py311.txt for the CI counterpart, resolved
# for CI's python 3.11 instead - the two aren't interchangeable: audioread
# (via librosa) needs standard-aifc/standard-sunau only on python>=3.13,
# since aifc/sunau left stdlib there, so a 3.11-resolved lockfile is
# missing hashes pip needs on this image's actual 3.13 interpreter and
# --require-hashes fails outright rather than silently under-pinning).
# Every fetched package is hash-verified (Scorecard Pinned-Dependencies)
# and pinned to the same versions CI audited, e.g. msgpack==1.2.1 and
# setuptools==84.0.0 (which also replaces the base image's vulnerable
# 70.3.0, CVE-2025-47273 - nothing else in the tree pulls a newer copy).
# --no-deps on the local package itself: it's our own source tree, not a
# fetch, so there's nothing to hash-pin there - but `pip install .` still
# does a PEP 517 build, which by default creates an *isolated* build env
# and fetches [build-system] requires (setuptools, wheel) completely
# outside any hash checking. pep517-build.txt pins that exact
# build-system.requires; installing it first and passing
# --no-build-isolation makes pip reuse those hash-verified copies instead
# of fetching its own.
RUN pip install --no-cache-dir -r explorer-extra-py313.txt -r pep517-build.txt --require-hashes \
&& pip install --no-cache-dir --no-deps --no-build-isolation . \
&& rm -f explorer-extra-py313.txt pep517-build.txt \
&& chown -R semantica:semantica /app
USER semantica
+131
View File
@@ -0,0 +1,131 @@
# Growth & Distribution Playbook
North star: **10,000 developers who actually use Semantica in real projects**, not a raw PyPI download number. Downloads are a lagging indicator of distribution, not a target to optimize directly.
```
GitHub stars → Website visitors → PyPI installs → Weekly active users → Production deployments → Enterprise customers
```
The last two matter far more than the download count.
## Guardrails — do not do this
- No fake/looping CI jobs that repeatedly `pip install semantica` purely to inflate the graph. It's detectable, it produces zero real users, and it damages credibility with anyone doing diligence (investors, enterprise buyers, security reviewers).
- No package-splitting purely to multiply install counts — only split into `semantica-*` packages when there's a real architectural reason.
- No meaningless Docker pulls or notebook launches with no real content behind them.
- Every item below should get someone from "installed it" to "used it for something real." If a channel can't do that, it's not worth building.
## 30-day priority sprint
Ordered by leverage-to-effort ratio; do these first.
| # | Initiative | Target |
| - | ---------- | ------ |
| 1 | ✅ GitHub Actions example + reusable `setup-semantica` composite action + install-matrix badge | done |
| 2 | Google Colab notebooks | 10 |
| 3 | Docker images (RAG, Graph, Agent, API) | 4-5 |
| 4 | Hugging Face Spaces demos | 3-4 |
| 5 | LangChain integration + example | 1 |
| 6 | LlamaIndex integration + example | 1 |
| 7 | Vector/graph DB integrations (Qdrant, Weaviate, Neo4j) | 3 |
| 8 | MCP server + example | 1 (already have `mcp/` — package as a distributable example) |
| 9 | Production-quality starter repos (FastAPI, Streamlit, Gradio) | 3 |
| 10 | `awesome-rag` / `awesome-llm` / `awesome-knowledge-graph` list submissions | 3+ PRs |
Push everything through: GitHub → Discord (`sV34vps5hH`) → X (`@BuildSemantica`) → GitHub Discussions → Reddit → Hacker News → relevant newsletters.
## Full channel checklist
### CI/CD (highest-intent distribution — installs tied to real pipelines)
- [x] GitHub Actions example in `examples/ci/github-actions.yml`
- [x] Reusable composite GitHub Action — [`.github/actions/setup-semantica`](.github/actions/setup-semantica/action.yml), modeled on `actions/setup-python`; usable by any repo as `uses: semantica-agi/semantica/.github/actions/setup-semantica@main`
- [x] "pip install" status badge in the README, backed by [`.github/workflows/install-matrix.yml`](.github/workflows/install-matrix.yml) — verifies the *published* package installs cleanly on Ubuntu/macOS/Windows across Python 3.9-3.12, weekly + on every release
- [x] GitLab CI template — `examples/ci/gitlab-ci.yml`
- [x] CircleCI template — `examples/ci/circleci-config.yml`
- [ ] Jenkins, Azure DevOps, Bitbucket Pipelines, Buildkite, Travis CI equivalents
### Release pipeline hardening (already had Trusted Publishing/OIDC + SLSA attestation — this rounds it out to match top-tier OSS release practice)
- [x] `twine check` gate in `.github/workflows/release.yml` before publish — catches a broken PyPI long-description render before it goes live instead of after (a malformed README on the live PyPI page is a silent conversion killer)
- [x] `CITATION.cff` (see Academic & research below)
- [x] OpenSSF Scorecard (see Discoverability below)
- [ ] Considered and deliberately skipped: Release Drafter / auto-generated changelogs — this repo hand-curates `CHANGELOG.md` with far more detail (PR numbers, contributors, phase-1 limitations) than a bot would produce. Don't introduce this without checking with maintainers first.
- [ ] Renovate / Dependabot config templates that auto-bump the `semantica` version in downstream repos — real recurring CI runs on real adopters
- [ ] Nightly scheduled workflow template that tests a downstream project against `semantica@latest`
### Containers & dev environments
- [ ] Official Docker images: RAG, Graph, Agent, API, `+Postgres`, `+Neo4j`, `+Qdrant`
- [ ] `docker-compose` examples (repo already has `docker-compose.dev.yml` / `docker-compose.yml` as a base)
- [ ] `.devcontainer/devcontainer.json` for one-click "Reopen in Container"
- [ ] GitHub Codespaces-ready config
- [ ] Gitpod config
- [ ] "Use this template" GitHub repo button so new projects start with `semantica` in `requirements.txt`
### Notebooks & hosted demos
- [ ] 10-20 Google Colab notebooks (Graph RAG, agent memory, entity resolution, semantic search, document intelligence)
- [ ] Kaggle Notebooks/Kernels
- [ ] Binder / mybinder.org config for instant repo launch
- [ ] SageMaker Studio Lab / Databricks Community Edition / Paperspace Gradient examples
- [ ] Hugging Face Spaces (Streamlit/Gradio) demos with `semantica` in `requirements.txt`
- [ ] Public hosted playground (source on GitHub, install visible)
### Framework & data-store integrations
- [x] LangChain integration — `integrations/langchain/` (`SemanticaRetriever`, `SemanticaVectorStore`, `SemanticaKGTool`/`SemanticaDecisionTool`), `pip install semantica[langchain]`, shipped in 0.6.7
- [ ] LlamaIndex integration + example
- [ ] LangGraph example
- [ ] Neo4j integration/example (docs already list it as a supported graph store — turn into a runnable example repo)
- [ ] Vector DB examples: Qdrant, Weaviate, Milvus, Pinecone, Chroma, FAISS, pgvector, OpenSearch/Elasticsearch (FAISS/Pinecone/Weaviate/Qdrant/Milvus/PgVector already supported per `docs/community-projects.md` — package each as a standalone example)
- [ ] LLM provider quickstarts: OpenAI, Anthropic, Gemini, Groq, Ollama, HuggingFace, DeepSeek, LiteLLM (already-supported providers per docs — each gets its own copy-paste quickstart)
- [ ] CrewAI / Agno integration examples (already documented under `docs/integrations/`) — promote as standalone repos, not just docs pages
### Package managers & installers
- [ ] conda-forge feedstock
- [ ] Homebrew formula for the CLI
- [ ] Nix/nixpkgs packaging
- [ ] Chocolatey / Scoop (Windows)
- [ ] Document `uv add semantica` and `poetry add semantica` explicitly alongside `pip install`
### Downstream packages & CLI
- [ ] Genuinely useful `semantica-*` packages only where warranted (e.g. `semantica-rag`, `semantica-connectors`) — each pulls `semantica` as a real dependency
- [ ] Make sure `semantica init / ingest / index / query / serve` CLI flows are the default onboarding path in every tutorial
- [ ] VS Code extension wrapping the CLI (scaffold + run commands from the command palette)
- [ ] JetBrains plugin equivalent
### Templates & starters
- [ ] Cookiecutter templates: `cookiecutter-semantic-rag`, `cookiecutter-ai-agent`, `cookiecutter-enterprise-rag`
- [ ] Starter repos: FastAPI, Streamlit, Gradio, Next.js frontend + Semantica backend
- [ ] Cloud deploy templates: AWS, GCP, Azure, Modal, Railway, Render, Fly.io (repo already has `deploy/azure`, `deploy/gcp`, `deploy/fly`, `deploy/railway`, `deploy/render`, `deploy/kubernetes`, `deploy/helm` — link these prominently from the README/quickstart, they're already-built distribution surface)
- [ ] Terraform / Pulumi / Helm modules published to their respective registries
### Discoverability & curation
- [ ] Submit to `awesome-rag`, `awesome-llm`, `awesome-knowledge-graph`, `awesome-python`
- [ ] Pitch newsletters with engaged Python/AI audiences (Python Weekly, Import AI, TLDR AI, etc.)
- [x] PyPI trove classifiers/keywords and `project.urls` (Homepage/Docs/Repository/Changelog/Bug Tracker) — already complete in `pyproject.toml`
- [ ] Get listed on Papers With Code for any retrieval/graph-RAG benchmark work
- [x] [OpenSSF Scorecard](https://scorecard.dev/viewer/?uri=github.com/semantica-agi/semantica) badge + weekly workflow (`.github/workflows/scorecard.yml`) — a concrete trust signal security/procurement teams check before greenlighting adoption, which gates real (non-CI-bot) install growth at enterprises
### Academic & research
- [x] `CITATION.cff` at repo root — enables GitHub's native "Cite this repository" button, feeds Google Scholar/academic tooling; complements `docs/citation.md` (still needs a real Zenodo DOI to replace the `XXXXXXX` placeholder in both places once one is minted)
- [ ] arXiv paper if there's real architectural novelty to describe
- [ ] Zenodo DOI for citability (`docs/citation.md` already exists — make sure it points to a real DOI)
- [ ] Workshop/tutorial sessions at PyData/ODSC-style events with hands-on install steps
- [ ] University course material / bootcamp adoption outreach
### Content
- [ ] Reproducible benchmark repos (Graph RAG vs vector RAG, retrieval@k, enterprise-scale retrieval) with `pip install semantica && python benchmark.py`
- [ ] 20-30 real-world example applications (RAG, enterprise document intelligence, financial entity graphs, code knowledge graphs, research discovery, agent memory)
- [ ] Blog/tutorial posts on Dev.to, Medium, personal blogs — always with runnable code, not just prose
- [ ] Contribute integrations/PRs to other projects building RAG/agents/knowledge graphs — "I implemented Semantica support" beats "please use Semantica"
## Tracking
Don't just watch the raw PyPI number — use download analytics (e.g. PePy) to separate CI/bot traffic from real installs, and track the funnel above end-to-end where possible (stars → site visits → installs → weekly actives).
+32 -23
View File
@@ -18,15 +18,15 @@
> Ingest your enterprise data, extract what matters, build a Context Graph and knowledge graph (KG), and run graph analytics and causal reasoning over all of it, with full decision provenance baked in. Explainable, traceable, and trustworthy by design.
**Decision Intelligence &nbsp;·&nbsp; Context Management &nbsp;·&nbsp; Deterministic Reasoning &nbsp;·&nbsp; Ontology Management &nbsp;·&nbsp; Knowledge Modeling &nbsp;·&nbsp; End-to-End Traceability**
**Context Management &nbsp;·&nbsp; Knowledge Modeling &nbsp;·&nbsp; Deterministic Reasoning &nbsp;·&nbsp; Ontology Management &nbsp;·&nbsp; Decision Intelligence &nbsp;·&nbsp; End-to-End Traceability**
**Open Source &nbsp;·&nbsp; Self-Hostable &nbsp;·&nbsp; Auditable &nbsp;·&nbsp; Governed &nbsp;·&nbsp; Zero Vendor Lock-In**
**Open Source &nbsp;·&nbsp; Governed &nbsp;·&nbsp; Zero Vendor Lock-In**
**Polyglot Graph Storage &nbsp;·&nbsp; RDF & LPG Support &nbsp;·&nbsp; W3C Standards &nbsp;·&nbsp; Interoperable**
#### Built for High-Stakes, Regulated Domains
[![GitHub Stars](https://img.shields.io/github/stars/semantica-agi/semantica?style=flat-square&color=FFD700&logo=github&logoColor=white&label=Stars)](https://github.com/semantica-agi/semantica) [![GitHub Forks](https://img.shields.io/github/forks/semantica-agi/semantica?style=flat-square&color=6E40C9&logo=github&logoColor=white&label=Forks)](https://github.com/semantica-agi/semantica/network/members) [![Contributors](https://img.shields.io/github/contributors/semantica-agi/semantica?style=flat-square&color=2EA043&logo=github&logoColor=white)](https://github.com/semantica-agi/semantica/graphs/contributors) [![PyPI](https://img.shields.io/pypi/v/semantica.svg?style=flat-square&color=0066CC&logo=pypi&logoColor=white)](https://pypi.org/project/semantica/) [![Total Downloads](https://static.pepy.tech/badge/semantica?style=flat-square)](https://pepy.tech/project/semantica) [![Python 3.8+](https://img.shields.io/badge/python-3.8+-3776AB?style=flat-square&logo=python&logoColor=white)](https://www.python.org/) [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg?style=flat-square)](https://opensource.org/licenses/MIT) [![CI](https://img.shields.io/github/actions/workflow/status/semantica-agi/semantica/ci.yml?style=flat-square&label=CI)](https://github.com/semantica-agi/semantica/actions) [![Ask DeepWiki](https://deepwiki.com/badge.svg)](https://deepwiki.com/semantica-agi/semantica)
[![GitHub Stars](https://img.shields.io/github/stars/semantica-agi/semantica?style=flat-square&color=FFD700&logo=github&logoColor=white&label=Stars)](https://github.com/semantica-agi/semantica) [![GitHub Forks](https://img.shields.io/github/forks/semantica-agi/semantica?style=flat-square&color=6E40C9&logo=github&logoColor=white&label=Forks)](https://github.com/semantica-agi/semantica/network/members) [![Contributors](https://img.shields.io/github/contributors/semantica-agi/semantica?style=flat-square&color=2EA043&logo=github&logoColor=white)](https://github.com/semantica-agi/semantica/graphs/contributors) [![PyPI](https://img.shields.io/pypi/v/semantica.svg?style=flat-square&color=0066CC&logo=pypi&logoColor=white)](https://pypi.org/project/semantica/) [![Total Downloads](https://static.pepy.tech/badge/semantica?style=flat-square)](https://pepy.tech/project/semantica) [![Python 3.8+](https://img.shields.io/badge/python-3.8+-3776AB?style=flat-square&logo=python&logoColor=white)](https://www.python.org/) [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg?style=flat-square)](https://opensource.org/licenses/MIT) [![CI](https://img.shields.io/github/actions/workflow/status/semantica-agi/semantica/ci.yml?style=flat-square&label=CI)](https://github.com/semantica-agi/semantica/actions) [![Install Matrix](https://img.shields.io/github/actions/workflow/status/semantica-agi/semantica/install-matrix.yml?style=flat-square&label=pip%20install)](https://github.com/semantica-agi/semantica/actions/workflows/install-matrix.yml) [![OpenSSF Scorecard](https://api.scorecard.dev/projects/github.com/semantica-agi/semantica/badge?style=flat-square)](https://scorecard.dev/viewer/?uri=github.com/semantica-agi/semantica) [![Ask DeepWiki](https://deepwiki.com/badge.svg)](https://deepwiki.com/semantica-agi/semantica)
[![Website](https://img.shields.io/badge/Website-getsemantica.ai-000000?style=flat-square&logo=googlechrome&logoColor=white)](https://getsemantica.ai/) [![Docs](https://img.shields.io/badge/Docs-docs.getsemantica.ai-0099FF?style=flat-square&logo=readthedocs&logoColor=white)](https://docs.getsemantica.ai/) [![Discord](https://img.shields.io/badge/Discord-Join%20Community-5865F2?style=flat-square&logo=discord&logoColor=white)](https://discord.gg/sV34vps5hH) [![Twitter/X](https://img.shields.io/badge/Follow-%40BuildSemantica-000000?style=flat-square&logo=x&logoColor=white)](https://x.com/BuildSemantica) [![YouTube](https://img.shields.io/badge/YouTube-Watch%20Demos-FF0000?style=flat-square&logo=youtube&logoColor=white)](https://www.youtube.com/watch?v=QfnNZg4-dZA) [![Changelog](https://img.shields.io/badge/Changelog-View-6E40C9?style=flat-square&logo=keepachangelog&logoColor=white)](CHANGELOG.md)
@@ -56,20 +56,18 @@ pip install semantica
---
Most AI agents act without a trail. They store embeddings, not meaning: context that can't be explained, decisions that can't be audited. In lending, that gap is a compliance exposure, not an inconvenience: an underwriting agent's approval has to survive a regulator's "why" months later.
Semantica sits underneath your LLM, vector store, and agent framework as a deterministic infrastructure layer: no LLM required for graph construction, reasoning, or provenance.
Most AI agents run on embeddings, not meaning: similarity scores with no structure, no relationships, and no way to explain why a result came back. Semantica is the semantic/context layer underneath your LLM, vector store, and agent framework: a deterministic infrastructure layer (no LLM required for graph construction, reasoning, or provenance) that turns fragmented enterprise data into a structured, queryable Context Graph and knowledge graph, governed by ontologies and controlled vocabularies (OWL, SHACL, SKOS) so the meaning of your data is explicit, not just its embedding. Decision provenance and audit trails fall out of that structure as a property, not the product itself; in domains a regulator can question, that same structure just happens to double as a straight answer to "why."
> ⚠️ **System-level explainability, not foundation-model explainability.** Semantica does not expose or reconstruct what happens *inside* the LLM — its internal reasoning or chain-of-thought stays opaque, as it does for any external system. Semantica explains what's *outside* the model: the context and data fed in, the decision produced, its provenance, relevant relationships, applied policies, and the full execution trail.
**Who it's for:**
- **AI/ML platform teams** shipping agents that make consequential decisions and need structured, queryable context built from fragmented raw data, not just a vector index
- **Data platform teams on Databricks or Snowflake** who need to turn tables already sitting in Unity Catalog or a Snowflake warehouse into a governed, lineage-tracked knowledge graph, without exporting that data to a third-party SaaS first
- **Compliance, risk, and audit teams** who need a straight answer to "why did the AI do that?" in a format a regulator will actually accept
- **Regulated enterprises** (finance, healthcare, legal, government, defense) that can't ship a black box, and can't send their data to someone else's SaaS to get one
- **AI/ML platform teams** shipping agents that make consequential decisions and need structured, queryable context, not just a vector index
- **Data platform teams on Databricks or Snowflake** turning tables already in Unity Catalog or a warehouse into a governed, lineage-tracked knowledge graph, without exporting to a third-party SaaS
- **Compliance, risk, and audit teams** who need a straight answer to "why did the AI do that?" in a format a regulator accepts
- **Regulated enterprises** (finance, healthcare, legal, government, defense) that can't ship a black box or send their data to someone else's SaaS to get one
- **Platform and infra engineers** who want the KG, reasoning, and provenance stack self-hosted and swappable, not locked to one vendor's backend
- **Data and knowledge engineers** building a KG from messy, multi-source data: entities and relationships get extracted, conflicting or contradictory facts are flagged instead of silently overwritten, and duplicates are merged before they turn into noise
- **Data and knowledge engineers** building a KG from messy, multi-source data, where conflicting facts get flagged and duplicates get merged, not silently overwritten
**[Quick Start](#quick-start)** &nbsp;·&nbsp; **[Architecture](#architecture)** &nbsp;·&nbsp; **[What You Get](#what-semantica-gives-you)** &nbsp;·&nbsp; **[Why Semantica](#why-semantica)** &nbsp;·&nbsp; **[Decision Intelligence](#decision-intelligence)** &nbsp;·&nbsp; **[Context Graphs](#context-graphs)** &nbsp;·&nbsp; **[Recipe: Audit Trail](#recipe-audit-trail-for-a-regulated-decision)** &nbsp;·&nbsp; **[Module Reference](#module-reference)** &nbsp;·&nbsp; **[Integrations](#integrations)** &nbsp;·&nbsp; **[CLI](#cli)** &nbsp;·&nbsp; **[Performance](#performance)** &nbsp;·&nbsp; **[Install](#installation)**
@@ -83,7 +81,7 @@ Semantica sits underneath your LLM, vector store, and agent framework as a deter
- **Full Auditability:** W3C PROV-O provenance on every fact, with audit trails exportable to JSON, CSV, or RDF
- **Deterministic Reasoning:** Forward chaining, Rete network, Datalog, and SPARQL with fully explainable paths, not black boxes
- **Knowledge Pipeline:** Multi-source ingestion, entity-aware chunking, NER/relation/event extraction, and knowledge graph construction, with semantic deduplication and provenance-preserving merges throughout
- **Enterprise Data Platforms:** Native connectors for Databricks (Unity Catalog + Delta Lake, PAT/OAuth M2M auth, catalog/schema/table/lineage introspection) and Snowflake (warehouse/database/schema, key-pair and OAuth auth), so tables already living in your lakehouse or warehouse become graph nodes with provenance, not another export/import hop
- **Enterprise Data Platforms:** Native connectors for Databricks (Unity Catalog + Delta Lake, PAT/OAuth M2M auth, catalog/schema/table/lineage introspection), Snowflake (warehouse/database/schema, key-pair and OAuth auth), and SAP OData (Business Partners, Sales Orders, OAuth2/Basic auth), so data already living in your lakehouse or warehouse becomes graph nodes with provenance, not another export/import hop
- **Graph Analytics:** Centrality, community detection, link prediction, and shortest-path queries over the graph you just built
- **Polyglot Graph Storage:** Native RDF (embedded Oxigraph, Blazegraph, Apache Jena, Eclipse RDF4J via SPARQL) and Labeled Property Graphs (Neo4j, FalkorDB, Apache AGE, AWS Neptune via Cypher), plus vector stores, all swappable without touching your code
- **Visualization:** Explore any graph, ontology, or timeline in an interactive browser workbench
@@ -141,10 +139,6 @@ compliant = graph.check_decision_rules({"category": "vendor_selection"}) # poli
```bash
semantica doctor
# Python 3.11.9 pass
# semantica 0.6.7 pass
# faiss vector store pass
# Config file pass ~/.semantica/config.yaml
```
**Running in a script or CI?** Progress bars are written only when stdout is an interactive terminal (or a Jupyter notebook), so piping and redirecting stay clean by default. Override with `SEMANTICA_DISABLE_PROGRESS=1` to silence progress everywhere, or `SEMANTICA_FORCE_PROGRESS=1` to keep it when stdout is redirected. `SEMANTICA_DISABLE_PROGRESS` takes precedence.
@@ -169,7 +163,7 @@ Sources → Ingest → Parse → Normalize → Split → Extract → Conflict De
→ Vector Store + Polyglot Graph Store (RDF & LPG) → Export / Visualize / REST · MCP · CLI
```
- **Ingest:** files, web, databases, enterprise data platforms (Databricks, Snowflake), cloud (Google Drive, Elasticsearch), streams (Kafka, Kinesis), Git, email, MCP
- **Ingest:** files, web, databases, enterprise data platforms (Databricks, Snowflake, SAP), cloud (Google Drive, Elasticsearch), streams (Kafka, Kinesis), Git, email, MCP
- **Parse → Normalize → Split:** document parsing, text/entity/date normalization, GraphRAG-native entity-aware chunking
- **Extract → Conflict Detection → Deduplication:** NER, relations, events, triplets; conflicting facts flagged and resolved before they merge
- **Knowledge Graph:** `GraphBuilder` constructs the graph; bi-temporal facts and full graph analytics (centrality, communities, link prediction) run on top of it
@@ -279,7 +273,7 @@ retrieved = ctx.retrieve("who approved the Acme contract?")
## Recipe: Audit Trail for a Regulated Decision
The flagship pattern: record a causally-linked decision chain, attach provenance to every entity, and export a regulator-ready audit trail.
One pattern built on the same Context Graph: record a causally-linked decision chain, attach provenance to every entity, and export a regulator-ready audit trail.
```python
from semantica.context import ContextGraph
@@ -322,7 +316,7 @@ Every module below is independently importable, with working code samples verifi
| Module | What it does |
| --- | --- |
| [`semantica.ingest`](#semanticaingest-multi-source-ingestion) | Files, web, databases, APIs, streams, email, Git, Parquet, Databricks, Snowflake, MCP |
| [`semantica.ingest`](#semanticaingest-multi-source-ingestion) | Files, web, databases, APIs, streams, email, Git, Parquet, Databricks, Snowflake, SAP, MCP |
| [`semantica.semantic_extract`](#semanticasemantic_extract-ner-relations-events-triplets) | NER, relation extraction, event detection, triplet generation |
| [`semantica.kg`](#semanticakg-knowledge-graph-construction--analysis) | Graph construction, centrality, communities, link prediction |
| [`semantica.reasoning`](#semanticareasoning-forward-chaining-rete-datalog-sparql) | Forward chaining, Rete, Datalog, SPARQL, fully explainable |
@@ -351,7 +345,7 @@ Expand any module below for its runnable example.
<summary><b><code>semantica.ingest</code></b>: Multi-Source Ingestion</summary>
<a id="semanticaingest-multi-source-ingestion"></a>
Ingest from files, web, databases, APIs, streams, email, Git repos, Parquet, Databricks, Snowflake, or MCP servers, all through a unified interface.
Ingest from files, web, databases, APIs, streams, email, Git repos, Parquet, Databricks, Snowflake, SAP, or MCP servers, all through a unified interface.
```python
from semantica.ingest import FileIngestor, WebIngestor, ParquetIngestor, DBIngestor
@@ -402,7 +396,7 @@ orders = snowflake.ingest_table("ORDERS", limit=10_000)
> **Security Note:** Never hardcode credentials (`token`, `password`, `private_key`) in production code; pass them via environment variables (e.g., `DATABRICKS_TOKEN`, `SNOWFLAKE_PASSWORD`) or a secrets manager.
**Supported sources:** Local files (PDF, DOCX, PPTX, HTML, TXT, CSV, JSON, YAML, Excel, XML) · Web pages · RSS/Atom feeds · REST APIs · Databases (PostgreSQL, MySQL, SQLite, Oracle, SQL Server) · Parquet datasets · Databricks (Unity Catalog + Delta Lake) · Snowflake · Git repositories · Email (IMAP/POP3) · Message streams (Kafka, RabbitMQ, Kinesis, Pulsar) · MCP resources · Apache Arrow/Feather/IPC (`ArrowIngestor`)
**Supported sources:** Local files (PDF, DOCX, PPTX, HTML, TXT, CSV, JSON, YAML, Excel, XML) · Web pages · RSS/Atom feeds · REST APIs · Databases (PostgreSQL, MySQL, SQLite, Oracle, SQL Server) · Parquet datasets · Databricks (Unity Catalog + Delta Lake) · Snowflake · SAP (OData v2/v4) · Git repositories · Email (IMAP/POP3) · Message streams (Kafka, RabbitMQ, Kinesis, Pulsar) · MCP resources · Apache Arrow/Feather/IPC (`ArrowIngestor`)
DuckDB, Elasticsearch, Google Drive, HuggingFace, MongoDB, and Pandas ingestion also ship (`DuckDBIngestor`, `ElasticIngestor`, `GDriveIngestor`, `HuggingFaceIngestor`, `MongoIngestor`, `PandasIngestor`) but aren't re-exported from the top-level `semantica.ingest` namespace yet — import them directly: `from semantica.ingest.duckdb_ingestor import DuckDBIngestor`.
@@ -1030,7 +1024,7 @@ team = Team(agents=[researcher, analyst], mode="coordinate")
## More Recipes
The flagship audit-trail recipe is [above](#recipe-audit-trail-for-a-regulated-decision). Here are three more common patterns.
The audit-trail recipe is [above](#recipe-audit-trail-for-a-regulated-decision). Here are three more common patterns.
<details>
<summary><b>End-to-End GraphRAG Pipeline</b></summary>
@@ -1147,7 +1141,7 @@ if report.valid:
| **Vector Store** | FAISS · Pinecone · Weaviate · Qdrant · Milvus · PgVector · hybrid + filtered search |
| **Graph Databases (LPG)** | Neo4j · FalkorDB · Apache AGE · AWS Neptune |
| **Triple Stores (RDF)** | Oxigraph (embedded) · Blazegraph · Apache Jena · Eclipse RDF4J · unified `TripletStore` interface · SPARQL query & bulk load |
| **Enterprise Data Platforms** | Databricks (`DatabricksIngestor`: Unity Catalog + Delta Lake, PAT/OAuth M2M, table/query ingestion, catalog/schema/table/lineage introspection) · Snowflake (`SnowflakeIngestor`: warehouse/database/schema, password/key-pair/OAuth auth) |
| **Enterprise Data Platforms** | Databricks (`DatabricksIngestor`: Unity Catalog + Delta Lake, PAT/OAuth M2M, table/query ingestion, catalog/schema/table/lineage introspection) · Snowflake (`SnowflakeIngestor`: warehouse/database/schema, password/key-pair/OAuth auth) · SAP (`SAPIngestor`: OData v2/v4, OAuth2/Basic auth, Business Partners/Sales Orders) |
| **LLM Providers** | **All already supported today:** OpenAI (GPT-4o, o1, o3) · Anthropic (Claude) · Google Gemini · Mistral · Meta Llama · Groq · Cohere · Azure OpenAI · AWS Bedrock · Ollama · DeepSeek · Perplexity · Together AI · Fireworks AI · Replicate · HuggingFace · via `semantica.llms` and LiteLLM |
---
@@ -1519,6 +1513,7 @@ pip install semantica[vectorstore-qdrant] # Qdrant vector store
pip install semantica[vectorstore-pinecone] # Pinecone vector store
pip install semantica[db-snowflake] # Snowflake
pip install semantica[db-databricks] # Databricks (SDK + SQL connector)
pip install semantica[ingest-sap] # SAP OData
pip install semantica[ingest-parquet] # Parquet / PyArrow
pip install semantica[ingest-arrow] # Apache Arrow, Feather, IPC
pip install semantica[viz] # HTML interactive visualization
@@ -1534,6 +1529,20 @@ git clone https://github.com/semantica-agi/semantica.git
cd semantica && pip install -e ".[dev]" && pytest tests/
```
### CI & Deployment
Wiring `semantica` into your own CI is a two-minute job. On GitHub Actions, use the reusable composite action:
```yaml
- uses: semantica-agi/semantica/.github/actions/setup-semantica@main
with:
python-version: '3.11'
```
Copy-paste starting templates for GitHub Actions, GitLab CI, and CircleCI live in [examples/ci/](examples/ci/). The published package itself is verified installable across Ubuntu/macOS/Windows and Python 3.9-3.12 every week by the [Install Matrix workflow](.github/workflows/install-matrix.yml).
Ready-made deployment configs for AWS, GCP, Azure, Fly.io, Railway, Render, Kubernetes, and Helm are in [deploy/](deploy/).
---
## Enterprise
+2 -3
View File
@@ -153,7 +153,7 @@ that attack chain.
- **Risk**: a PR merges without its security/CI checks passing.
**Control**: merges require the `build`, `Analyze Python` (CodeQL), and `security-scan` checks to pass, in strict mode (checks must be re-run against the latest `main`).
- **Risk**: a compromised scanner job reaches secrets or write access.
**Control**: scanning jobs (`CodeQL`, `security-scan.yml`, `security.yml`, `defender-for-devops.yml`) run with read-only, least-privilege permissions (typically `contents: read` + `security-events: write` only) and never share a job, environment, or secret scope with the publish job.
**Control**: scanning jobs (`CodeQL`, `security-scan.yml`, `defender-for-devops.yml`) run with read-only, least-privilege permissions (typically `contents: read` + `security-events: write` only) and never share a job, environment, or secret scope with the publish job.
- **Risk**: secrets are committed accidentally.
**Control**: GitHub secret scanning and push protection are both enabled at the repository level, rejecting pushes that contain recognizable credential patterns before they land in history.
@@ -164,8 +164,7 @@ Every scan below runs continuously in CI, not just at release time:
- **CodeQL** (`security-and-quality` query pack) — Python source: injection, unsafe deserialization, and other code-level vulnerability classes. Runs in `codeql.yml` on every push/PR to `main` and weekly.
- **Bandit** — Python-specific security anti-patterns (hardcoded secrets, unsafe `eval`/`pickle`, weak crypto, etc.); CI fails on any HIGH-severity finding. Runs in `security-scan.yml` on every push/PR to `main` and twice weekly.
- **Semgrep** (`p/security` ruleset) — cross-language static-analysis security patterns. Runs in `security-scan.yml` on every push/PR to `main` and twice weekly.
- **Safety** — known CVEs in Semantica's own installed dependencies, including optional LLM-provider extras such as LiteLLM; CI fails on any match. Runs in `security-scan.yml` on every push/PR to `main` and twice weekly.
- **pip-audit** — independent, PyPA-maintained vulnerability database cross-check against installed dependencies (Safety and pip-audit use different advisory sources, so both run). Runs in `security.yml` weekly.
- **pip-audit** — PyPA-maintained, OSV-backed vulnerability database cross-check against Semantica's pinned dependency tree, including optional LLM-provider extras such as LiteLLM; CI fails on any match. Runs in `security-scan.yml` on every push/PR to `main` and twice weekly, and can be triggered on demand via `workflow_dispatch`.
- **Microsoft Defender for DevOps** (`eslint`, `templateanalyzer`, `terrascan`) — JavaScript/TypeScript lint-security rules and infrastructure-as-code misconfigurations. Runs in `defender-for-devops.yml` on every push/PR to `main` and weekly.
- **Checkov** — Kubernetes, Helm, Dockerfile, GitHub Actions, and secrets-pattern IaC scanning; results upload to the same Security tab as CodeQL. Runs in `defender-for-devops.yml` on every push/PR to `main` and weekly.
- **GitGuardian** — secret-detection check on every pull request, installed as a GitHub App integration (not a repo-local workflow). Runs on every PR.
@@ -1,222 +0,0 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/semantica-agi/semantica/blob/main/cookbook/advanced/09_Semantic_Layer_Construction.ipynb)\n",
"\n",
"# Semantic Layer Construction\n",
"\n",
"## Overview\n",
"\n",
"Build an enterprise semantic layer: construct knowledge graph, generate ontology, create semantic layer, export RDF, and store in triplet store.\n",
"\n",
"\n",
"**Documentation**: [API Reference](https://semantica.readthedocs.io/concepts/)\n",
"\n",
"## Installation\n",
"\n",
"Install Semantica from PyPI:\n",
"\n",
"```bash\n",
"pip install semantica\n",
"# Or with all optional dependencies:\n",
"pip install semantica[all]\n",
"```\n",
"\n",
"## Workflow: Build KG → Generate Ontology → Create Semantic Layer → Export RDF \n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"!pip install -qU semantica\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from semantica.kg import GraphBuilder\n",
"from semantica.ontology import OntologyGenerator\n",
"from semantica.export import RDFExporter\n",
"from semantica.triplet_store import TripletStore\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 1: Build Knowledge Graph\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"builder = GraphBuilder()\n",
"\n",
"entities = [\n",
" {\"id\": \"e1\", \"type\": \"Person\", \"name\": \"Alice\", \"properties\": {\"age\": 30, \"role\": \"Engineer\"}},\n",
" {\"id\": \"e2\", \"type\": \"Person\", \"name\": \"Bob\", \"properties\": {\"age\": 35, \"role\": \"Manager\"}},\n",
" {\"id\": \"e3\", \"type\": \"Organization\", \"name\": \"Tech Corp\", \"properties\": {\"founded\": 2010}},\n",
" {\"id\": \"e4\", \"type\": \"Project\", \"name\": \"Project Alpha\", \"properties\": {\"status\": \"active\"}},\n",
"]\n",
"\n",
"relationships = [\n",
" {\"source\": \"e1\", \"target\": \"e2\", \"type\": \"reports_to\"},\n",
" {\"source\": \"e1\", \"target\": \"e3\", \"type\": \"works_for\"},\n",
" {\"source\": \"e2\", \"target\": \"e3\", \"type\": \"works_for\"},\n",
" {\"source\": \"e1\", \"target\": \"e4\", \"type\": \"works_on\"},\n",
"]\n",
"\n",
"knowledge_graph = builder.build(entities, relationships)\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 2: Generate Ontology\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"generator = OntologyGenerator()\n",
"ontology = generator.generate_from_graph(knowledge_graph)\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 3: Create Semantic Layer\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"def create_mappings(kg, ontology):\n",
" mappings = {\n",
" \"entity_type_mappings\": {},\n",
" \"relationship_type_mappings\": {},\n",
" \"property_mappings\": {}\n",
" }\n",
" \n",
" entity_types = set(e.get(\"type\") for e in entities)\n",
" ontology_classes = ontology.get(\"classes\", [])\n",
" \n",
" for entity_type in entity_types:\n",
" matching_class = next((cls for cls in ontology_classes if cls.get(\"name\") == entity_type), None)\n",
" if matching_class:\n",
" mappings[\"entity_type_mappings\"][entity_type] = matching_class.get(\"uri\", entity_type)\n",
" \n",
" relationship_types = set(r.get(\"type\") for r in relationships)\n",
" ontology_properties = ontology.get(\"properties\", [])\n",
" \n",
" for rel_type in relationship_types:\n",
" matching_prop = next((prop for prop in ontology_properties if prop.get(\"name\") == rel_type), None)\n",
" if matching_prop:\n",
" mappings[\"relationship_type_mappings\"][rel_type] = matching_prop.get(\"uri\", rel_type)\n",
" \n",
" return mappings\n",
"\n",
"mappings = create_mappings(knowledge_graph, ontology)\n",
"\n",
"semantic_layer = {\n",
" \"graph\": knowledge_graph,\n",
" \"ontology\": ontology,\n",
" \"mappings\": mappings,\n",
" \"metadata\": {\n",
" \"version\": \"1.0\",\n",
" \"created_at\": \"2024-01-01\",\n",
" \"description\": \"Enterprise semantic layer\"\n",
" }\n",
"}\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 4: Export RDF\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"exporter = RDFExporter()\n",
"# Export Knowledge Graph\n",
"exporter.export(knowledge_graph, \"knowledge_graph.ttl\", format=\"turtle\")\n",
"print(\"Exported knowledge graph to knowledge_graph.ttl\")\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Summary\n",
"\n",
"Enterprise semantic layer construction:\n",
"- Knowledge Graph Built\n",
"- Ontology Generated\n",
"- Semantic Layer Created with Mappings\n",
"- RDF Export Completed\n"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": []
}
],
"metadata": {
"kernelspec": {
"display_name": "Python 3",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3",
"version": "3.11.9"
}
},
"nbformat": 4,
"nbformat_minor": 2
}
@@ -1,435 +0,0 @@
{
"nbformat": 4,
"nbformat_minor": 5,
"metadata": {
"kernelspec": {
"display_name": "Python 3",
"language": "python",
"name": "python3"
},
"language_info": {
"name": "python",
"version": "3.10.0"
}
},
"cells": [
{
"cell_type": "markdown",
"id": "cell-0",
"metadata": {},
"source": [
"[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/semantica-agi/semantica/blob/main/cookbook/advanced/13_Manual_Ontology_Snowflake_Mapping.ipynb)\n",
"\n",
"# Manual Ontology + Snowflake Mapping\n",
"\n",
"This notebook answers a specific workflow:\n",
"\n",
"> *\"I want to design the ontology myself — not have AI infer it from my tables — and then map Snowflake data to it explicitly.\"*\n",
"\n",
"### What this notebook demonstrates\n",
"\n",
"| Step | What happens | Who controls it |\n",
"|---|---|---|\n",
"| 1 | Design ontology classes and properties | **You** (Python dict) |\n",
"| 2 | Model n-ary facts with reification | **You** (`AssociativeClassBuilder`) |\n",
"| 3 | Pull rows from Snowflake | Semantica `SnowflakeIngestor` |\n",
"| 4 | Map columns → ontology-aligned graph | **You** (explicit transform) |\n",
"| 5 | Validate + export OWL / SHACL | Semantica `OntologyEngine` |\n",
"| 6 | Load to triplet store and query | Semantica `TripletStore` |\n",
"\n",
"### What this notebook does NOT do\n",
"\n",
"- No LLM-driven ontology generation\n",
"- No schema introspection or table-to-class inference\n",
"- No \"suggest ontology from my data\"\n",
"\n",
"### Standards coverage\n",
"\n",
"| Feature | Status |\n",
"|---|---|\n",
"| OWL 2 (Turtle / RDF-XML) | Supported |\n",
"| SHACL 1.1 shapes | Supported |\n",
"| SPARQL 1.1 | Supported |\n",
"| Reification / n-ary facts | Supported via `AssociativeClassBuilder` |\n",
"| SPARQL 1.2 (reifier annotation, `LATERAL`) | Planned |\n",
"| SHACL 1.2 (`sh:severity` extensions, SHACL-AF) | Planned |"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-1",
"metadata": {},
"outputs": [],
"source": [
"!pip install -qU semantica"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-2",
"metadata": {},
"outputs": [],
"source": [
"import os\n",
"from typing import Any, Dict, List\n",
"\n",
"from semantica.ingest import SnowflakeIngestor\n",
"from semantica.kg.methods import build_kg\n",
"from semantica.ontology import AssociativeClassBuilder, OntologyEngine\n",
"from semantica.triplet_store import TripletStore"
]
},
{
"cell_type": "markdown",
"id": "cell-3",
"metadata": {},
"source": [
"## Step 1: Hand-Design the Ontology in Python\n",
"\n",
"You define every class and property explicitly. Nothing is read from Snowflake at this stage.\n",
"\n",
"**Design decisions that belong to you:**\n",
"- Which classes exist and what they mean\n",
"- Which properties are datatype vs. object properties\n",
"- Domain, range, and cardinality constraints\n",
"- Which properties are required (later enforced by SHACL)\n",
"\n",
"This dict versions with your code. It does not change when your database schema changes."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-4",
"metadata": {},
"outputs": [],
"source": "BASE_URI = \"https://example.com/hr/\"\n\n# Your ontology — designed by you, not inferred by Semantica.\nontology: Dict[str, Any] = {\n \"name\": \"EmploymentDomainOntology\",\n \"uri\": f\"{BASE_URI}EmploymentDomainOntology\",\n \"namespace\": {\"base_uri\": BASE_URI},\n\n # You decide the class taxonomy\n \"classes\": [\n {\"name\": \"Person\", \"uri\": f\"{BASE_URI}Person\"},\n {\"name\": \"Organization\", \"uri\": f\"{BASE_URI}Organization\"},\n {\"name\": \"Role\", \"uri\": f\"{BASE_URI}Role\"},\n # EmploymentEvent is a reification node.\n # It connects Person + Organization + Role and carries salary/date context.\n {\"name\": \"EmploymentEvent\", \"uri\": f\"{BASE_URI}EmploymentEvent\"},\n ],\n\n # Each property carries a full URI so TripletStore stores it as hr:<name>\n # rather than the default urn:property:<name>.\n # This ensures SPARQL queries using PREFIX hr: match what is actually stored.\n \"properties\": [\n # Datatype properties\n {\"name\": \"name\", \"uri\": f\"{BASE_URI}name\", \"type\": \"datatype\", \"domain\": \"Person\", \"range\": \"string\", \"required\": True},\n {\"name\": \"legalName\", \"uri\": f\"{BASE_URI}legalName\", \"type\": \"datatype\", \"domain\": \"Organization\", \"range\": \"string\", \"required\": True},\n {\"name\": \"title\", \"uri\": f\"{BASE_URI}title\", \"type\": \"datatype\", \"domain\": \"Role\", \"range\": \"string\", \"required\": True},\n {\"name\": \"startDate\", \"uri\": f\"{BASE_URI}startDate\", \"type\": \"datatype\", \"domain\": \"EmploymentEvent\", \"range\": \"date\"},\n {\"name\": \"endDate\", \"uri\": f\"{BASE_URI}endDate\", \"type\": \"datatype\", \"domain\": \"EmploymentEvent\", \"range\": \"date\"},\n {\"name\": \"salary\", \"uri\": f\"{BASE_URI}salary\", \"type\": \"datatype\", \"domain\": \"EmploymentEvent\", \"range\": \"decimal\"},\n\n # Object properties — reification spokes (required)\n {\"name\": \"employee\", \"uri\": f\"{BASE_URI}employee\", \"type\": \"object\", \"domain\": \"EmploymentEvent\", \"range\": \"Person\", \"required\": True},\n {\"name\": \"employer\", \"uri\": f\"{BASE_URI}employer\", \"type\": \"object\", \"domain\": \"EmploymentEvent\", \"range\": \"Organization\", \"required\": True},\n {\"name\": \"role\", \"uri\": f\"{BASE_URI}role\", \"type\": \"object\", \"domain\": \"EmploymentEvent\", \"range\": \"Role\", \"required\": True},\n\n # Shortcut edges — direct person→org / person→role without traversing the event node\n {\"name\": \"worksFor\", \"uri\": f\"{BASE_URI}worksFor\", \"type\": \"object\", \"domain\": \"Person\", \"range\": \"Organization\"},\n {\"name\": \"hasRole\", \"uri\": f\"{BASE_URI}hasRole\", \"type\": \"object\", \"domain\": \"Person\", \"range\": \"Role\"},\n ],\n}\n\nontology"
},
{
"cell_type": "markdown",
"id": "cell-5",
"metadata": {},
"source": [
"## Step 2: Reification — Modeling N-Ary Facts\n",
"\n",
"**The problem with binary triples:**\n",
"A simple triple `(Alice, worksFor, Acme)` cannot carry extra context such as salary, start date, or role.\n",
"Standard RDF reification and OWL n-ary patterns solve this by introducing an intermediate node.\n",
"\n",
"Semantica's `AssociativeClassBuilder` is the Pythonic API for this pattern:\n",
"\n",
"```\n",
"EmploymentEvent\n",
" ├── employee → Person (required)\n",
" ├── employer → Organization (required)\n",
" ├── role → Role (required)\n",
" ├── startDate → xsd:date\n",
" ├── endDate → xsd:date\n",
" └── salary → xsd:decimal\n",
"```\n",
"\n",
"**On SPARQL 1.1 vs. SPARQL 1.2:**\n",
"- **SPARQL 1.1 (current):** traverse the event node explicitly — `?event hr:employee ?person ; hr:salary ?salary`\n",
"- **SPARQL 1.2 (planned):** the draft reifier annotation syntax allows attaching context to triples directly, without a separate intermediate node. Semantica will adopt this once the spec is ratified.\n",
"\n",
"**On SHACL 1.1 vs. SHACL 1.2:**\n",
"- **SHACL 1.1 (current):** `sh:NodeShape` + `sh:PropertyShape` constraints are exported for all `required` properties and enforced at load time.\n",
"- **SHACL 1.2 (planned):** `sh:severity` profile extensions and SHACL-AF rules are on the roadmap."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-6",
"metadata": {},
"outputs": [],
"source": "assoc_builder = AssociativeClassBuilder()\n\nemployment_assoc = assoc_builder.create_associative_class(\n name=\"EmploymentEvent\",\n connects=[\"Person\", \"Organization\", \"Role\"],\n temporal=True, # adds startDate / endDate handling\n properties={\n \"startDate\": \"xsd:date\",\n \"endDate\": \"xsd:date\",\n \"salary\": \"xsd:decimal\",\n },\n)\n\nvalidation_result = assoc_builder.validate_associative_class(employment_assoc)\n\n# AssociativeClass is a dataclass — use attribute access, not .get()\nprint(\"AssociativeClass structure:\")\nprint(f\" name: {employment_assoc.name}\")\nprint(f\" connects: {employment_assoc.connects}\")\nprint(f\" temporal: {employment_assoc.temporal}\")\nprint(f\" properties: {list(employment_assoc.properties.keys())}\")\nprint(f\"\\nValidation passed: {validation_result}\")"
},
{
"cell_type": "markdown",
"id": "cell-7",
"metadata": {},
"source": [
"## Step 3: Ingest Snowflake Rows (Extraction Only)\n",
"\n",
"`SnowflakeIngestor` retrieves rows — nothing more. It does **not**:\n",
"- Inspect your table schema\n",
"- Suggest classes or properties\n",
"- Infer relationships from column names\n",
"\n",
"Set `USE_LIVE_SNOWFLAKE=true` plus the env vars below to connect to a real warehouse.\n",
"Otherwise the stub data is used."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-8",
"metadata": {},
"outputs": [],
"source": [
"def fetch_rows_from_snowflake() -> List[Dict[str, Any]]:\n",
" if os.getenv(\"USE_LIVE_SNOWFLAKE\", \"false\").lower() != \"true\":\n",
" return [\n",
" {\n",
" \"EMPLOYEE_ID\": \"E100\",\n",
" \"EMPLOYEE_NAME\": \"Alice Johnson\",\n",
" \"ORG_ID\": \"O10\",\n",
" \"ORG_NAME\": \"Acme Corp\",\n",
" \"ROLE_ID\": \"R7\",\n",
" \"ROLE_TITLE\": \"Senior Engineer\",\n",
" \"START_DATE\": \"2025-01-15\",\n",
" \"END_DATE\": None,\n",
" \"SALARY\": 160000,\n",
" },\n",
" {\n",
" \"EMPLOYEE_ID\": \"E101\",\n",
" \"EMPLOYEE_NAME\": \"Bob Singh\",\n",
" \"ORG_ID\": \"O10\",\n",
" \"ORG_NAME\": \"Acme Corp\",\n",
" \"ROLE_ID\": \"R9\",\n",
" \"ROLE_TITLE\": \"Data Architect\",\n",
" \"START_DATE\": \"2024-09-01\",\n",
" \"END_DATE\": None,\n",
" \"SALARY\": 185000,\n",
" },\n",
" ]\n",
"\n",
" ingestor = SnowflakeIngestor(\n",
" account=os.getenv(\"SNOWFLAKE_ACCOUNT\"),\n",
" user=os.getenv(\"SNOWFLAKE_USER\"),\n",
" password=os.getenv(\"SNOWFLAKE_PASSWORD\"),\n",
" warehouse=os.getenv(\"SNOWFLAKE_WAREHOUSE\"),\n",
" database=os.getenv(\"SNOWFLAKE_DATABASE\"),\n",
" schema=os.getenv(\"SNOWFLAKE_SCHEMA\", \"PUBLIC\"),\n",
" )\n",
" query = (\n",
" \"SELECT EMPLOYEE_ID, EMPLOYEE_NAME, \"\n",
" \"ORG_ID, ORG_NAME, ROLE_ID, ROLE_TITLE, \"\n",
" \"START_DATE, END_DATE, SALARY \"\n",
" \"FROM HR_EMPLOYMENT_FACT\"\n",
" )\n",
" data = ingestor.ingest_query(query)\n",
" ingestor.close()\n",
" return data.data\n",
"\n",
"\n",
"rows = fetch_rows_from_snowflake()\n",
"rows[:2]"
]
},
{
"cell_type": "markdown",
"id": "cell-9",
"metadata": {},
"source": [
"## Step 4: Map Rows to Ontology Concepts Explicitly\n",
"\n",
"This is the semantic transformation layer — the part that makes your ontology real.\n",
"\n",
"Semantica does not guess which column becomes which entity or property.\n",
"Every assignment is code you write and own:\n",
"\n",
"- **Stable node IDs** — deterministic, collision-safe, derived from business keys\n",
"- **Class assignment** — matches what you declared in Step 1\n",
"- **Property routing** — each column value goes to the correct ontology property\n",
"- **Reification wiring** — `EmploymentEvent` is linked to its three participants\n",
"\n",
"When your Snowflake schema changes, only this function needs updating. The ontology stays stable."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-10",
"metadata": {},
"outputs": [],
"source": "def map_rows_to_kg(rows: List[Dict[str, Any]]) -> Dict[str, Any]:\n entities: Dict[str, Dict[str, Any]] = {}\n relationships: List[Dict[str, Any]] = []\n\n for row in rows:\n # Stable, deterministic node IDs derived from business keys\n person_id = f\"person:{row['EMPLOYEE_ID']}\"\n org_id = f\"org:{row['ORG_ID']}\"\n role_id = f\"role:{row['ROLE_ID']}\"\n # Event ID includes all three participants + start date so that\n # a re-hired employee gets a distinct event node, not an overwrite.\n event_id = f\"employment:{row['EMPLOYEE_ID']}:{row['ORG_ID']}:{row['START_DATE']}\"\n\n # Entities — \"type\" must match a class name from Step 1\n entities[person_id] = {\n \"id\": person_id,\n \"type\": \"Person\",\n \"properties\": {\"name\": row[\"EMPLOYEE_NAME\"]},\n }\n entities[org_id] = {\n \"id\": org_id,\n \"type\": \"Organization\",\n \"properties\": {\"legalName\": row[\"ORG_NAME\"]},\n }\n entities[role_id] = {\n \"id\": role_id,\n \"type\": \"Role\",\n \"properties\": {\"title\": row[\"ROLE_TITLE\"]},\n }\n\n # Reification node — filter out None values so TripletStore does not\n # stringify None as the literal \"None\" for open-ended employment.\n event_props = {\n \"startDate\": row[\"START_DATE\"],\n \"endDate\": row[\"END_DATE\"],\n \"salary\": row[\"SALARY\"],\n }\n entities[event_id] = {\n \"id\": event_id,\n \"type\": \"EmploymentEvent\",\n \"properties\": {k: v for k, v in event_props.items() if v is not None},\n }\n\n # Full URIs for relationship types so TripletStore stores hr:<type>\n # instead of the default urn:property:<type>, keeping SPARQL consistent.\n relationships.extend([\n # Shortcut edges — fast SPARQL when context is not needed\n {\"source\": person_id, \"target\": org_id, \"type\": f\"{BASE_URI}worksFor\"},\n {\"source\": person_id, \"target\": role_id, \"type\": f\"{BASE_URI}hasRole\"},\n # Reification spokes — full context via the event node\n {\"source\": event_id, \"target\": person_id, \"type\": f\"{BASE_URI}employee\"},\n {\"source\": event_id, \"target\": org_id, \"type\": f\"{BASE_URI}employer\"},\n {\"source\": event_id, \"target\": role_id, \"type\": f\"{BASE_URI}role\"},\n ])\n\n return build_kg([{\"entities\": list(entities.values()), \"relationships\": relationships}])\n\n\nkg = map_rows_to_kg(rows)\nprint(f\"Entities built: {len(kg.get('entities', []))}\")\nprint(f\"Relationships built: {len(kg.get('relationships', []))}\")\n\nsample = next((e for e in kg[\"entities\"] if e[\"type\"] == \"EmploymentEvent\"), None)\nprint(f\"\\nSample EmploymentEvent node: {sample}\")"
},
{
"cell_type": "markdown",
"id": "cell-11",
"metadata": {},
"source": [
"## Step 5: Validate Ontology and Export OWL + SHACL\n",
"\n",
"`OntologyEngine` validates your ontology dict and serialises it to standards-compliant files.\n",
"\n",
"**Output files:**\n",
"- `employment_manual_ontology.ttl` — OWL 2 Turtle\n",
"- `employment_manual_shapes.ttl` — SHACL 1.1 node and property shapes\n",
"\n",
"**Standards status:**\n",
"\n",
"| Standard | Semantica support |\n",
"|---|---|\n",
"| SPARQL 1.1 | Full |\n",
"| SHACL 1.1 (`sh:NodeShape`, `sh:PropertyShape`, `sh:minCount`, `sh:datatype`, `sh:class`) | Full |\n",
"| SPARQL 1.2 (reifier annotation syntax, `LATERAL`) | Tracked — not yet implemented |\n",
"| SHACL 1.2 (`sh:severity` profiles, SHACL-AF extensions) | Tracked — not yet implemented |"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-12",
"metadata": {},
"outputs": [],
"source": [
"engine = OntologyEngine(base_uri=BASE_URI)\n",
"\n",
"validation = engine.validate(ontology)\n",
"owl_ttl = engine.to_owl(ontology, format=\"turtle\")\n",
"shacl_ttl = engine.to_shacl(ontology, format=\"turtle\")\n",
"\n",
"engine.export_owl(ontology, \"employment_manual_ontology.ttl\", format=\"turtle\")\n",
"engine.export_shacl(ontology, \"employment_manual_shapes.ttl\", format=\"turtle\")\n",
"\n",
"print(f\"Ontology valid: {validation.valid}\")\n",
"print(f\"Ontology consistent: {validation.consistent}\")\n",
"print(f\"OWL output: {len(owl_ttl):,} chars → employment_manual_ontology.ttl\")\n",
"print(f\"SHACL output: {len(shacl_ttl):,} chars → employment_manual_shapes.ttl\")\n",
"\n",
"print(\"\\n--- SHACL shapes (first 20 lines) ---\")\n",
"print(\"\\n\".join(shacl_ttl.splitlines()[:20]))"
]
},
{
"cell_type": "markdown",
"id": "cell-13",
"metadata": {},
"source": [
"## Best-Practice Architecture\n",
"\n",
"```\n",
"┌──────────────────────────────────┐\n",
"│ Ontology as code (Python dict) │ ← versioned alongside your application\n",
"│ + AssociativeClass for n-ary │\n",
"└───────────────┬──────────────────┘\n",
" │ validate + export\n",
" ▼\n",
"┌───────────────────────────────────┐\n",
"│ OWL 2 Turtle │ SHACL 1.1 │ ← standards-compliant artifacts\n",
"└───────────────┬───────────────────┘\n",
" │\n",
" ▼\n",
"┌──────────────────────────────────┐\n",
"│ Snowflake — raw data access │ ← no schema introspection\n",
"└───────────────┬──────────────────┘\n",
" │ explicit mapping layer\n",
" ▼\n",
"┌──────────────────────────────────┐\n",
"│ Ontology-aligned KG │ ← types, IDs, edges match Step 1\n",
"└───────────────┬──────────────────┘\n",
" │ optional\n",
" ▼\n",
"┌──────────────────────────────────┐\n",
"│ Triplet store + SPARQL 1.1 │\n",
"└──────────────────────────────────┘\n",
"```\n",
"\n",
"**Why this split matters:**\n",
"If Semantica inferred the ontology from your Snowflake schema, every schema migration would risk silently changing your semantic model.\n",
"With this pattern, schema changes only touch the mapping function in Step 4 — the ontology remains stable and under your control."
]
},
{
"cell_type": "markdown",
"id": "cell-14",
"metadata": {},
"source": [
"## SPARQL Query Patterns\n",
"\n",
"Two query styles are available because we wrote both shortcut edges and reification spokes.\n",
"\n",
"### Simple lookup — shortcut edge (no context needed)\n",
"\n",
"```sparql\n",
"PREFIX hr: <https://example.com/hr/>\n",
"\n",
"SELECT ?personName ?orgName\n",
"WHERE {\n",
" ?person a hr:Person ;\n",
" hr:name ?personName ;\n",
" hr:worksFor ?org .\n",
" ?org hr:legalName ?orgName .\n",
"}\n",
"```\n",
"\n",
"### Contextual lookup — via reification node (salary, dates, role)\n",
"\n",
"```sparql\n",
"PREFIX hr: <https://example.com/hr/>\n",
"\n",
"SELECT ?personName ?roleTitle ?salary ?startDate\n",
"WHERE {\n",
" ?event a hr:EmploymentEvent ;\n",
" hr:employee ?person ;\n",
" hr:role ?role ;\n",
" hr:salary ?salary ;\n",
" hr:startDate ?startDate .\n",
" ?person hr:name ?personName .\n",
" ?role hr:title ?roleTitle .\n",
"}\n",
"ORDER BY DESC(?salary)\n",
"```\n",
"\n",
"### Future: SPARQL 1.2 reifier syntax\n",
"\n",
"The SPARQL 1.2 draft introduces annotation syntax that lets you attach context directly to triples, without a separate intermediate node.\n",
"Once the spec is ratified Semantica will adopt it, and the contextual query above may be expressible more concisely."
]
},
{
"cell_type": "markdown",
"id": "cell-15",
"metadata": {},
"source": [
"## Step 6 (Optional): Load to Triplet Store and Run SPARQL\n",
"\n",
"Set `STORE_TO_TRIPLET=true` to load the KG into a live triplet store and run the contextual reification query."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "cell-16",
"metadata": {},
"outputs": [],
"source": [
"if os.getenv(\"STORE_TO_TRIPLET\", \"false\").lower() == \"true\":\n",
" store = TripletStore(\n",
" backend=os.getenv(\"TRIPLET_BACKEND\", \"blazegraph\"),\n",
" endpoint=os.getenv(\"TRIPLET_ENDPOINT\", \"http://localhost:9999/blazegraph\"),\n",
" namespace=os.getenv(\"TRIPLET_NAMESPACE\", \"kb\"),\n",
" )\n",
" store_result = store.store(knowledge_graph=kg, ontology=ontology)\n",
" print(\"Store result:\", store_result)\n",
"\n",
" # Contextual reification query — person + role + salary via EmploymentEvent\n",
" query = \"\"\"\n",
" PREFIX hr: <https://example.com/hr/>\n",
"\n",
" SELECT ?personName ?roleTitle ?salary ?startDate\n",
" WHERE {\n",
" ?event a hr:EmploymentEvent ;\n",
" hr:employee ?person ;\n",
" hr:role ?role ;\n",
" hr:salary ?salary ;\n",
" hr:startDate ?startDate .\n",
" ?person hr:name ?personName .\n",
" ?role hr:title ?roleTitle .\n",
" }\n",
" ORDER BY DESC(?salary)\n",
" LIMIT 10\n",
" \"\"\"\n",
" result = store.execute_query(query)\n",
" print(result)\n",
"else:\n",
" print(\"Skipping triplet-store load/query (set STORE_TO_TRIPLET=true to enable)\")"
]
}
]
}
@@ -10,15 +10,16 @@
"\n",
"## Overview\n",
"\n",
"This notebook demonstrates how to build knowledge graphs from entities and relationships using Semantica's graph building modules. You'll learn to use `GraphBuilder` and `EntityResolver`.\n",
"This notebook demonstrates how to build knowledge graphs from extracted entities and relationships using Semantica's graph building modules. You'll learn to use `GraphBuilder` and `EntityResolver`.\n",
"\n",
"**Documentation**: [API Reference](https://semantica.readthedocs.io/reference/kg/)\n",
"\n",
"### Learning Objectives\n",
"\n",
"- Use `GraphBuilder` to construct knowledge graphs\n",
"- Use `EntityResolver` to resolve entity conflicts\n",
"**Note**: For deduplication, use the `semantica.deduplication` module.\n",
"- Extract entity mentions and relations, and map them into graph records\n",
"- Use `GraphBuilder` to construct a graph whose edges come from the actual extracted relations\n",
"- Use `EntityResolver` to merge duplicate mentions and remap relationship endpoints\n",
"- Use the `semantica.deduplication` module and report the complete deduplicated entity set\n",
"\n",
"## Installation\n",
"\n",
@@ -32,120 +33,217 @@
"\n",
"---\n",
"\n",
"## Step 1: Build Knowledge Graph\n",
"## Step 1: Extract Entities and Relations\n",
"\n",
"Construct a knowledge graph from entities and relationships.\n"
"Extract entity mentions and relations from text. The sample text mentions `Apple Inc.` in two separate sentences, so we can later show how duplicate mentions are resolved into one canonical entity.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"!pip install semantica\n"
]
"%pip install semantica\n",
"\n",
"# spaCy models are distributed separately from the spaCy library. This lesson\n",
"# relies on the English model to recognize standalone places such as Cupertino.\n",
"import sys\n",
"import subprocess\n",
"import spacy\n",
"\n",
"try:\n",
" spacy.load(\"en_core_web_sm\")\n",
"except OSError:\n",
" subprocess.check_call([sys.executable, \"-m\", \"spacy\", \"download\", \"en_core_web_sm\"])\n"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from semantica.kg import GraphBuilder\n",
"from semantica.semantic_extract import NERExtractor, RelationExtractor\n",
"\n",
"builder = GraphBuilder()\n",
"text = (\n",
" \"Apple Inc. is headquartered in Cupertino, California. \"\n",
" \"Tim Cook is the CEO of Apple Inc. \"\n",
" \"The company is a technology company.\"\n",
")\n",
"\n",
"ner_extractor = NERExtractor()\n",
"relation_extractor = RelationExtractor()\n",
"\n",
"text = \"Apple Inc. is a technology company. Tim Cook is the CEO of Apple Inc. Apple Inc. is headquartered in Cupertino, California.\"\n",
"mentions = ner_extractor.extract(text)\n",
"relations = relation_extractor.extract(text, mentions)\n",
"\n",
"entities_list = ner_extractor.extract(text)\n",
"relationships_list = relation_extractor.extract(text, entities_list)\n",
"print(\"Entity mentions:\")\n",
"for mention in mentions:\n",
" print(f\" {mention.text!r:<13} {mention.label:<7} span=[{mention.start_char}:{mention.end_char}]\")\n",
"\n",
"entities = []\n",
"for i, entity in enumerate(entities_list[:5], 1):\n",
" entities.append({\n",
" \"id\": f\"e{i}\",\n",
" \"type\": entity.label,\n",
" \"name\": entity.text,\n",
" \"properties\": {}\n",
" })\n",
"\n",
"relationships = []\n",
"for i, rel in enumerate(relationships_list[:3], 1):\n",
" relationships.append({\n",
" \"source\": f\"e{1}\",\n",
" \"target\": f\"e{i+1}\",\n",
" \"type\": rel.predicate,\n",
" \"properties\": {}\n",
" })\n",
"\n",
"knowledge_graph = builder.build(entities, relationships)\n",
"\n",
"print(f\"Built knowledge graph with {len(knowledge_graph.get('entities', []))} entities\")\n",
"print(f\"Relationships: {len(knowledge_graph.get('relationships', []))}\")"
]
"print(\"\\nExtracted relations:\")\n",
"for rel in relations:\n",
" print(f\" {rel.subject.text!r} --{rel.predicate}--> {rel.object.text!r}\")"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 2: Entity Resolution\n",
"## Step 2: Build the Knowledge Graph\n",
"\n",
"Resolve entity conflicts and duplicates.\n"
"Give every mention a graph ID, then translate each relation's `subject` and `object` into those IDs. Building edges from the actual relation endpoints — rather than guessing endpoints from list positions — is what keeps the graph faithful to the text.\n"
]
},
{
"cell_type": "code",
"metadata": {},
"source": [
"from semantica.kg import GraphBuilder\n",
"\n",
"entities = []\n",
"span_to_id = {}\n",
"for i, mention in enumerate(mentions, 1):\n",
" graph_id = f\"e{i}\"\n",
" span_to_id[(mention.start_char, mention.end_char)] = graph_id\n",
" entities.append({\n",
" \"id\": graph_id,\n",
" \"type\": mention.label,\n",
" \"name\": mention.text,\n",
" \"properties\": {},\n",
" })\n",
"\n",
"relationships = []\n",
"for rel in relations:\n",
" source_id = span_to_id.get((rel.subject.start_char, rel.subject.end_char))\n",
" target_id = span_to_id.get((rel.object.start_char, rel.object.end_char))\n",
" if source_id is None or target_id is None:\n",
" print(f\"Skipping relation with unmapped endpoint: \"\n",
" f\"{rel.subject.text!r} --{rel.predicate}--> {rel.object.text!r}\")\n",
" continue\n",
" relationships.append({\n",
" \"source\": source_id,\n",
" \"target\": target_id,\n",
" \"type\": rel.predicate,\n",
" \"properties\": {},\n",
" })\n",
"\n",
"builder = GraphBuilder()\n",
"knowledge_graph = builder.build({\"entities\": entities, \"relationships\": relationships})\n",
"\n",
"id_to_name = {entity[\"id\"]: entity[\"name\"] for entity in entities}\n",
"\n",
"print(f\"Graph entities ({len(knowledge_graph['entities'])}):\")\n",
"for entity in knowledge_graph[\"entities\"]:\n",
" print(f\" {entity['id']}: {entity['name']} ({entity['type']})\")\n",
"\n",
"print(f\"\\nGraph relationships ({len(knowledge_graph['relationships'])}):\")\n",
"for relationship in knowledge_graph[\"relationships\"]:\n",
" print(f\" {id_to_name[relationship['source']]} \"\n",
" f\"--{relationship['type']}--> {id_to_name[relationship['target']]}\")\n",
"\n",
"edges = {\n",
" (id_to_name[r[\"source\"]], r[\"type\"], id_to_name[r[\"target\"]])\n",
" for r in knowledge_graph[\"relationships\"]\n",
"}\n",
"assert (\"Apple Inc.\", \"located_in\", \"Cupertino\") in edges\n",
"assert (\"Tim Cook\", \"works_for\", \"Apple Inc.\") in edges"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 3: Entity Resolution\n",
"\n",
"The graph currently contains two nodes for the same organization. `EntityResolver` merges duplicate mentions into one canonical entity and records which source IDs were merged (`merged_from`), so relationship endpoints can be remapped onto the canonical entity.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from semantica.kg import EntityResolver\n",
"\n",
"entity_resolver = EntityResolver()\n",
"\n",
"resolved_entities = entity_resolver.resolve_entities(entities)\n",
"\n",
"print(f\"Original entities: {len(entities)}\")\n",
"print(f\"Resolved entities: {len(resolved_entities)}\")"
]
"canonical_id = {}\n",
"for entity in resolved_entities:\n",
" for source_id in entity.get(\"merged_from\", [entity[\"id\"]]):\n",
" canonical_id[source_id] = entity[\"id\"]\n",
" if entity.get(\"merged_from\"):\n",
" print(f\"Merged {entity['merged_from']} -> {entity['id']}: {entity['name']}\")\n",
"\n",
"print(f\"\\nMentions in: {len(entities)}, resolved entities out: {len(resolved_entities)}\")\n",
"\n",
"resolved_names = {entity[\"id\"]: entity[\"name\"] for entity in resolved_entities}\n",
"print(\"\\nRelationships remapped onto canonical entities:\")\n",
"for relationship in relationships:\n",
" source = canonical_id[relationship[\"source\"]]\n",
" target = canonical_id[relationship[\"target\"]]\n",
" print(f\" {resolved_names[source]} --{relationship['type']}--> {resolved_names[target]}\")\n",
"\n",
"canonical_entities = {(entity[\"name\"], entity[\"type\"]) for entity in resolved_entities}\n",
"assert canonical_entities == {\n",
" (\"Apple Inc.\", \"ORG\"),\n",
" (\"Tim Cook\", \"PERSON\"),\n",
" (\"Cupertino\", \"GPE\"),\n",
" (\"California\", \"GPE\"),\n",
"}\n",
"assert len(resolved_entities) == 4"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 3: Deduplication\n",
"## Step 4: Deduplication\n",
"\n",
"Remove duplicate entities from the graph.\n"
"The `semantica.deduplication` module gives finer control over the same problem. Note that `merge_duplicates` returns one `MergeOperation` per duplicate *group* — the complete deduplicated collection is those merged entities plus every entity that was not part of any group.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from semantica.deduplication import DuplicateDetector, EntityMerger, MergeStrategy\n",
"\n",
"# Detect duplicates\n",
"detector = DuplicateDetector(similarity_threshold=0.8)\n",
"duplicate_groups = detector.detect_duplicate_groups(knowledge_graph.get('entities', []))\n",
"duplicate_groups = detector.detect_duplicate_groups(entities)\n",
"print(f\"Duplicate groups: {len(duplicate_groups)}\")\n",
"for group in duplicate_groups:\n",
" print(f\" {[entity['name'] for entity in group.entities]} \"\n",
" f\"(confidence={group.confidence:.2f})\")\n",
"\n",
"# Merge duplicates\n",
"merger = EntityMerger()\n",
"merge_operations = merger.merge_duplicates(\n",
" knowledge_graph.get('entities', []),\n",
" strategy=MergeStrategy.KEEP_MOST_COMPLETE\n",
" entities, strategy=MergeStrategy.KEEP_MOST_COMPLETE\n",
")\n",
"\n",
"deduplicated_entities = [op.merged_entity for op in merge_operations]\n",
"merged_source_ids = {\n",
" entity[\"id\"] for op in merge_operations for entity in op.source_entities\n",
"}\n",
"untouched_entities = [e for e in entities if e[\"id\"] not in merged_source_ids]\n",
"deduplicated_entities = untouched_entities + [\n",
" op.merged_entity for op in merge_operations\n",
"]\n",
"\n",
"print(f\"Original entities: {len(knowledge_graph.get('entities', []))}\")\n",
"print(f\"Deduplicated entities: {len(deduplicated_entities)}\")\n"
]
"print(f\"\\nMerge operations: {len(merge_operations)}\")\n",
"print(f\"Deduplicated entities ({len(deduplicated_entities)}):\")\n",
"for entity in deduplicated_entities:\n",
" print(f\" {entity['id']}: {entity['name']} ({entity['type']})\")\n",
"\n",
"assert len(merge_operations) == 1\n",
"assert len(deduplicated_entities) == 4"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
@@ -155,9 +253,10 @@
"\n",
"You've learned how to build knowledge graphs:\n",
"\n",
"- **GraphBuilder**: Construct knowledge graphs from entities and relationships\n",
"- **EntityResolver**: Resolve entity conflicts and duplicates\n",
"- **Deduplication**: Use `semantica.deduplication` module for removing duplicate entities\n",
"- **Extraction to graph**: map each mention to a graph ID and build edges from the actual `Relation.subject` / `Relation.object` endpoints\n",
"- **GraphBuilder**: construct knowledge graphs from explicit `{\"entities\": ..., \"relationships\": ...}` input\n",
"- **EntityResolver**: merge duplicate mentions into canonical entities and remap relationship endpoints\n",
"- **Deduplication**: combine `MergeOperation` results with untouched entities to get the complete deduplicated set\n",
"\n",
"Next: Learn how to analyze graphs in the Graph_Analytics notebook.\n"
]
@@ -10,7 +10,7 @@
"\n",
"## Overview\n",
"\n",
"This notebook walks you through creating your first knowledge graph from a simple document. You'll learn the complete end-to-end workflow from ingesting a file to visualizing the resulting knowledge graph.\n",
"This notebook walks you through creating your first knowledge graph from a simple document. You'll learn the complete end-to-end workflow from ingesting a file to visualizing the resulting knowledge graph — and every step consumes the real output of the step before it.\n",
"\n",
"> [!TIP]\n",
"> This is the perfect starting point if you are new to Semantica. No prior knowledge of knowledge graphs is required!\n",
@@ -19,10 +19,10 @@
"\n",
"### 🎯 Learning Objectives\n",
"\n",
"- **Understand the Workflow**: Learn the `File → Parse → Extract → Graph` pipeline\n",
"- **Understand the Workflow**: Learn the `File → Parse → Extract → Graph → Visualize` pipeline\n",
"- **Ingest Data**: Load documents using `FileIngestor`\n",
"- **Parse Content**: Extract text using `DocumentParser`\n",
"- **Extract Knowledge**: Identify entities using `NERExtractor`\n",
"- **Extract Knowledge**: Identify entities and relations using `NERExtractor` and `RelationExtractor`\n",
"- **Build Graph**: Construct a graph using `GraphBuilder`\n",
"- **Visualize**: See your graph come to life with `KGVisualizer`\n",
"\n",
@@ -40,71 +40,76 @@
"\n",
"## 🔄 Simple End-to-End Workflow\n",
"\n",
"The complete workflow consists of four main steps:\n",
"The complete workflow consists of five main steps:\n",
"\n",
"1. **📥 Ingest** - Load data from files or other sources\n",
"2. **📄 Parse** - Extract and structure content from documents\n",
"3. **⛏️ Extract** - Identify entities and relationships\n",
"4. **🕸️ Build Graph** - Construct the knowledge graph\n",
"5. **📊 Visualize** - Render and analyze the graph\n",
"\n",
"Each step is demonstrated in the code cells below.\n",
"Each step is demonstrated in the code cells below, and each cell can be rerun on its own: the sample file is only removed by the optional cleanup cell at the very end.\n",
"\n",
"> [!TIP]\n",
"> **Alternative: Using Semantica Framework**\n",
"> \n",
">\n",
"> For a simpler, high-level approach, you can use the `Semantica` framework class which orchestrates all these steps:\n",
"> \n",
">\n",
"> ```python\n",
"> from semantica.core import Semantica\n",
"> \n",
">\n",
"> framework = Semantica()\n",
"> framework.initialize()\n",
"> \n",
">\n",
"> result = framework.build_knowledge_base(\n",
"> sources=[\"sample_document.txt\"],\n",
"> embeddings=True,\n",
"> graph=True\n",
"> )\n",
"> \n",
">\n",
"> framework.shutdown()\n",
"> ```\n",
"> \n",
">\n",
"> This notebook shows the step-by-step approach for learning. See [Core Module Usage Guide](../../../semantica/core/core_usage.md) for more details.\n",
"\n",
"---\n",
"\n",
"## 📂 Step 1: Ingest a File\n",
"\n",
"In this step, we'll use `FileIngestor` to load a document. The ingestor supports various file formats including PDF, DOCX, TXT, and more.\n"
"In this step, we'll use `FileIngestor` to load a document. The ingestor supports various file formats including PDF, DOCX, TXT, and more. Writing the sample file is idempotent, so this cell can be rerun at any time.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"!pip install semantica"
]
"%pip install semantica\n",
"\n",
"# spaCy models are distributed separately from the spaCy library. This lesson\n",
"# relies on the English model to recognize standalone places such as Cupertino.\n",
"import sys\n",
"import subprocess\n",
"import spacy\n",
"\n",
"try:\n",
" spacy.load(\"en_core_web_sm\")\n",
"except OSError:\n",
" subprocess.check_call([sys.executable, \"-m\", \"spacy\", \"download\", \"en_core_web_sm\"])\n"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from semantica.ingest import FileIngestor\n",
"from pathlib import Path\n",
"\n",
"# Initialize the ingestor\n",
"ingestor = FileIngestor()\n",
"from semantica.ingest import FileIngestor\n",
"\n",
"# Create a sample document for demonstration\n",
"sample_text = \"\"\"\n",
"Apple Inc. is a technology company founded by Steve Jobs, Steve Wozniak, and Ronald Wayne in 1976.\n",
"The company is headquartered in Cupertino, California.\n",
"Tim Cook is the current CEO of Apple Inc.\n",
"Apple designs and manufactures consumer electronics, software, and online services.\n",
"sample_text = \"\"\"Apple Inc. is headquartered in Cupertino, California.\n",
"In 1976, Steve Jobs founded Apple Inc.\n",
"Tim Cook is the CEO of Apple Inc.\n",
"\"\"\"\n",
"\n",
"sample_file = Path(\"sample_document.txt\")\n",
@@ -113,12 +118,14 @@
"print(f\"File: {sample_file}\")\n",
"print(f\"Content length: {len(sample_text)} characters\")\n",
"\n",
"# Ingest the file\n",
"ingestor = FileIngestor()\n",
"file_object = ingestor.ingest_file(sample_file, read_content=True)\n",
"print(f\" File name: {file_object.name}\")\n",
"print(f\" File type: {file_object.file_type}\")\n",
"print(f\" Content available: {file_object.content is not None}\")\n"
]
"print(f\" Content available: {file_object.content is not None}\")"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
@@ -126,64 +133,58 @@
"source": [
"## 📄 Step 2: Parse the Document\n",
"\n",
"After ingesting the file, we need to parse it to extract the text content. The `DocumentParser` handles various file formats and extracts structured content.\n"
"After ingesting the file, we need to parse it to extract the text content. `DocumentParser.parse_document()` returns the extracted text under the `\"text\"` key.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from semantica.parse import DocumentParser\n",
"\n",
"parser = DocumentParser()\n",
"# Parse the document to extract text\n",
"parsed_document = parser.parse_document(str(sample_file))\n",
"parsed_content = parsed_document.get(\"content\", \"\")\n",
"print(f\" Parsed content length: {len(parsed_content) if parsed_content else 0} characters\")\n",
"print(f\" Preview: {parsed_content[:200] if parsed_content else 'N/A'}...\")"
]
"\n",
"parsed_content = parsed_document.get(\"text\", \"\")\n",
"assert parsed_content.strip(), \"Parsing produced no text — check the input file\"\n",
"\n",
"print(f\"Parsed content length: {len(parsed_content)} characters\")\n",
"print(f\"Preview: {parsed_content[:120]}...\")"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## ⛏️ Step 3: Extract Entities\n",
"## ⛏️ Step 3: Extract Entities and Relations\n",
"\n",
"Now we'll extract entities from the parsed text using Named Entity Recognition (NER). This identifies people, organizations, locations, dates, and other entities in the text.\n",
"\n",
"> [!NOTE]\n",
"> In a real scenario, you would use `NERExtractor` with an LLM or model backend. Here we simulate the output for demonstration purposes.\n"
"Now we'll extract entities and relations from the parsed text. `NERExtractor` identifies people, organizations, locations and dates; `RelationExtractor` finds relations between those mentions. Both operate on the *parsed content from Step 2* — not on a copy of the raw string.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from semantica.semantic_extract import NamedEntityRecognizer, NERExtractor\n",
"from semantica.semantic_extract import NERExtractor, RelationExtractor\n",
"\n",
"ner = NamedEntityRecognizer()\n",
"extractor = NERExtractor()\n",
"ner_extractor = NERExtractor()\n",
"relation_extractor = RelationExtractor()\n",
"\n",
"print(f\"\\nText: {parsed_content[:100]}...\")\n",
"mentions = ner_extractor.extract(parsed_content)\n",
"relations = relation_extractor.extract(parsed_content, mentions)\n",
"\n",
"# Simulated extraction results\n",
"expected_entities = [\n",
" {\"text\": \"Apple Inc.\", \"type\": \"Organization\", \"start\": 0, \"end\": 10},\n",
" {\"text\": \"Steve Jobs\", \"type\": \"Person\", \"start\": 50, \"end\": 60},\n",
" {\"text\": \"Steve Wozniak\", \"type\": \"Person\", \"start\": 62, \"end\": 75},\n",
" {\"text\": \"Ronald Wayne\", \"type\": \"Person\", \"start\": 81, \"end\": 93},\n",
" {\"text\": \"1976\", \"type\": \"Date\", \"start\": 97, \"end\": 101},\n",
" {\"text\": \"Cupertino, California\", \"type\": \"Location\", \"start\": 130, \"end\": 151},\n",
" {\"text\": \"Tim Cook\", \"type\": \"Person\", \"start\": 153, \"end\": 161},\n",
"]\n",
"print(\"Entity mentions:\")\n",
"for mention in mentions:\n",
" print(f\" {mention.text!r:<13} {mention.label:<7} span=[{mention.start_char}:{mention.end_char}]\")\n",
"\n",
"for entity in expected_entities:\n",
" print(f\" - {entity['text']} ({entity['type']})\")\n"
]
"print(\"\\nExtracted relations:\")\n",
"for rel in relations:\n",
" print(f\" {rel.subject.text!r} --{rel.predicate}--> {rel.object.text!r}\")"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
@@ -191,58 +192,68 @@
"source": [
"## 🕸️ Step 4: Build the Knowledge Graph\n",
"\n",
"Using the extracted entities and relationships, we'll construct a knowledge graph. The graph represents entities as nodes and relationships as edges.\n"
"Using the extracted entities and relations, we construct a knowledge graph with `GraphBuilder`. Every mention gets a graph ID, and each edge is built from the actual `Relation.subject` / `Relation.object` endpoints.\n",
"\n",
"> [!NOTE]\n",
"> The graph will contain one node per *mention*, so `Apple Inc.` appears three times. Merging duplicate mentions into one canonical entity is covered in [07_Building_Knowledge_Graphs.ipynb](./07_Building_Knowledge_Graphs.ipynb).\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from semantica.kg import GraphBuilder\n",
"import networkx as nx\n",
"\n",
"entities = []\n",
"span_to_id = {}\n",
"for i, mention in enumerate(mentions, 1):\n",
" graph_id = f\"e{i}\"\n",
" span_to_id[(mention.start_char, mention.end_char)] = graph_id\n",
" entities.append({\n",
" \"id\": graph_id,\n",
" \"type\": mention.label,\n",
" \"name\": mention.text,\n",
" \"properties\": {},\n",
" })\n",
"\n",
"relationships = []\n",
"for rel in relations:\n",
" source_id = span_to_id.get((rel.subject.start_char, rel.subject.end_char))\n",
" target_id = span_to_id.get((rel.object.start_char, rel.object.end_char))\n",
" if source_id is None or target_id is None:\n",
" print(f\"Skipping relation with unmapped endpoint: \"\n",
" f\"{rel.subject.text!r} --{rel.predicate}--> {rel.object.text!r}\")\n",
" continue\n",
" relationships.append({\n",
" \"source\": source_id,\n",
" \"target\": target_id,\n",
" \"type\": rel.predicate,\n",
" \"properties\": {},\n",
" })\n",
"\n",
"builder = GraphBuilder()\n",
"knowledge_graph = builder.build({\"entities\": entities, \"relationships\": relationships})\n",
"\n",
"# Prepare data for graph construction\n",
"entities_data = [\n",
" {\"id\": f\"entity_{i}\", \"name\": entity[\"text\"], \"type\": entity[\"type\"]}\n",
" for i, entity in enumerate(expected_entities)\n",
"]\n",
"id_to_name = {entity[\"id\"]: entity[\"name\"] for entity in entities}\n",
"\n",
"relationships_data = [\n",
" {\"source\": \"entity_0\", \"target\": \"entity_1\", \"type\": \"founded_by\"},\n",
" {\"source\": \"entity_0\", \"target\": \"entity_2\", \"type\": \"founded_by\"},\n",
" {\"source\": \"entity_0\", \"target\": \"entity_3\", \"type\": \"founded_by\"},\n",
" {\"source\": \"entity_0\", \"target\": \"entity_4\", \"type\": \"founded_in\"},\n",
" {\"source\": \"entity_0\", \"target\": \"entity_5\", \"type\": \"located_in\"},\n",
" {\"source\": \"entity_6\", \"target\": \"entity_0\", \"type\": \"ceo_of\"},\n",
"]\n",
"print(f\"Nodes (entities): {len(knowledge_graph['entities'])}\")\n",
"for entity in knowledge_graph[\"entities\"]:\n",
" print(f\" {entity['id']}: {entity['name']} ({entity['type']})\")\n",
"\n",
"# Build the graph using NetworkX\n",
"kg = nx.DiGraph()\n",
"print(f\"\\nEdges (relationships): {len(knowledge_graph['relationships'])}\")\n",
"for relationship in knowledge_graph[\"relationships\"]:\n",
" print(f\" {id_to_name[relationship['source']]} \"\n",
" f\"--{relationship['type']}--> {id_to_name[relationship['target']]}\")\n",
"\n",
"for entity in entities_data:\n",
" kg.add_node(entity[\"id\"], name=entity[\"name\"], type=entity[\"type\"])\n",
"\n",
"for rel in relationships_data:\n",
" source_name = entities_data[int(rel[\"source\"].split(\"_\")[1])][\"name\"]\n",
" target_name = entities_data[int(rel[\"target\"].split(\"_\")[1])][\"name\"]\n",
" kg.add_edge(rel[\"source\"], rel[\"target\"], type=rel[\"type\"])\n",
"\n",
"print(f\" Nodes (entities): {len(kg.nodes)}\")\n",
"print(f\" Edges (relationships): {len(kg.edges)}\")\n",
"\n",
"for node_id in kg.nodes():\n",
" node_data = kg.nodes[node_id]\n",
" print(f\" Node: {node_data['name']} ({node_data['type']})\")\n",
"\n",
"for source, target, data in kg.edges(data=True):\n",
" source_name = kg.nodes[source]['name']\n",
" target_name = kg.nodes[target]['name']\n",
" print(f\" {source_name} --[{data['type']}]--> {target_name}\")\n"
]
"edges = {\n",
" (id_to_name[r[\"source\"]], r[\"type\"], id_to_name[r[\"target\"]])\n",
" for r in knowledge_graph[\"relationships\"]\n",
"}\n",
"assert (\"Apple Inc.\", \"located_in\", \"Cupertino\") in edges\n",
"assert (\"Tim Cook\", \"works_for\", \"Apple Inc.\") in edges"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
@@ -250,49 +261,81 @@
"source": [
"## 📊 Step 5: Visualize and Analyze\n",
"\n",
"Finally, we'll visualize the knowledge graph and analyze its structure. This helps you understand the relationships and entities in your data.\n"
"Finally, we render the knowledge graph with `KGVisualizer` and look at its structure. `visualize_network()` accepts the `GraphBuilder` result directly and can save an interactive HTML file.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from semantica.visualization import KGVisualizer\n",
"\n",
"visualizer = KGVisualizer()\n",
"\n",
"print(f\" Total entities: {len(kg.nodes)}\")\n",
"print(f\" Total relationships: {len(kg.edges)}\")\n",
"fig = visualizer.visualize_network(\n",
" knowledge_graph, output=\"html\", file_path=\"knowledge_graph.html\"\n",
")\n",
"print(\"Saved interactive visualization to knowledge_graph.html\")\n",
"\n",
"entity_types = {}\n",
"for node_id in kg.nodes():\n",
" entity_type = kg.nodes[node_id]['type']\n",
" entity_types[entity_type] = entity_types.get(entity_type, 0) + 1\n",
"for entity in knowledge_graph[\"entities\"]:\n",
" entity_types[entity[\"type\"]] = entity_types.get(entity[\"type\"], 0) + 1\n",
"\n",
"for etype, count in entity_types.items():\n",
" print(f\" - {etype}: {count}\")\n",
"print(\"\\nEntities by type:\")\n",
"for entity_type, count in sorted(entity_types.items()):\n",
" print(f\" - {entity_type}: {count}\")\n",
"\n",
"rel_types = {}\n",
"for _, _, data in kg.edges(data=True):\n",
" rel_type = data.get('type', 'unknown')\n",
" rel_types[rel_type] = rel_types.get(rel_type, 0) + 1\n",
"relationship_types = {}\n",
"for relationship in knowledge_graph[\"relationships\"]:\n",
" relationship_types[relationship[\"type\"]] = (\n",
" relationship_types.get(relationship[\"type\"], 0) + 1\n",
" )\n",
"\n",
"for rtype, count in rel_types.items():\n",
" print(f\" - {rtype}: {count}\")\n",
"print(\"\\nRelationships by type:\")\n",
"for relationship_type, count in sorted(relationship_types.items()):\n",
" print(f\" - {relationship_type}: {count}\")\n",
"\n",
"# Cleanup\n",
"if sample_file.exists():\n",
" sample_file.unlink()\n"
"fig"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## 🧹 Optional: Clean Up\n",
"\n",
"Run this cell only when you are done with the notebook. Earlier cells read `sample_document.txt`, so they stay rerunnable until you delete it here.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": []
"source": [
"for path in [sample_file, Path(\"knowledge_graph.html\")]:\n",
" if path.exists():\n",
" path.unlink()\n",
" print(f\"Removed {path}\")"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Summary\n",
"\n",
"You've built your first knowledge graph, end to end:\n",
"\n",
"- **FileIngestor** loaded the sample document\n",
"- **DocumentParser** returned its text under the `\"text\"` key\n",
"- **NERExtractor** / **RelationExtractor** produced real mentions and relations from that text\n",
"- **GraphBuilder** turned them into a graph whose edges come from the actual relation endpoints\n",
"- **KGVisualizer** rendered the result as an interactive network\n",
"\n",
"Next: merge duplicate mentions with `EntityResolver` in [07_Building_Knowledge_Graphs.ipynb](./07_Building_Knowledge_Graphs.ipynb), or explore graph metrics in the Graph Analytics notebook.\n"
]
}
],
"metadata": {
+2 -1
View File
@@ -497,7 +497,8 @@
"**Next Steps**:\n",
"* Try customizing the `NamespaceManager` to use your organization's URL.\n",
"* Explore `OntologyEvaluator` for deeper quality metrics.\n",
"* Feed the generated ontology into the **Knowledge Graph** module to start reasoning over your data!"
"* Feed the generated ontology into the **Knowledge Graph** module to start reasoning over your data!\n",
"* Put the graph, ontology, and explicit mappings together in [Semantic Layer Basics](./26_Semantic_Layer_Basics.ipynb)."
]
}
],
@@ -0,0 +1,418 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/semantica-agi/semantica/blob/main/cookbook/introduction/26_Semantic_Layer_Basics.ipynb)\n",
"\n",
"# Semantic Layer Basics: Putting the Knowledge Graph, Ontology, and Mappings Together\n",
"\n",
"## Overview\n",
"\n",
"This lesson connects three things you have already met — a knowledge graph, an ontology, and RDF export — into one minimal *semantic layer*: a knowledge graph whose types, relationships, and properties are **explicitly mapped** to ontology terms, so the resulting RDF can be queried with SPARQL against a shared vocabulary.\n",
"\n",
"**Documentation**: [API Reference](https://semantica.readthedocs.io/concepts/)\n",
"\n",
"### 🎯 Learning Objectives\n",
"\n",
"- Build a small knowledge graph with `GraphBuilder`\n",
"- Generate a starter ontology from the graph with `OntologyGenerator`\n",
"- Write **explicit** entity-type, relationship-type, and property mappings to ontology terms\n",
"- Produce ontology-aligned RDF and store it with `TripletStore`\n",
"- Answer a business question with one small SPARQL query\n",
"\n",
"### 📚 Prerequisites\n",
"\n",
"- [07_Building_Knowledge_Graphs.ipynb](./07_Building_Knowledge_Graphs.ipynb) — graphs from entities and relationships\n",
"- [14_Ontology.ipynb](./14_Ontology.ipynb) — ontology generation\n",
"- [20_Triplet_Store.ipynb](./20_Triplet_Store.ipynb) — triplet store backends\n",
"\n",
"> [!NOTE]\n",
"> **Teaching mappings vs. governed mappings.** The mappings in this lesson are a demo: they live in a Python dict and are derived from a generated ontology. A production semantic layer uses governed identifiers, hand-designed ontologies, explicit source mappings, validation (SHACL), provenance, and versioning — that workflow is covered in [Advanced: Manual Ontology + Snowflake Mapping](../advanced/13_Manual_Ontology_Snowflake_Mapping.ipynb).\n",
"\n",
"## Installation\n",
"\n",
"The triplet-store step uses the embedded Oxigraph backend, so install with that extra. Pin at least 0.6.7: earlier releases could generate ontology classes with no URI (#1103), which silently breaks the mappings below instead of failing loudly.\n",
"\n",
"```bash\n",
"pip install \"semantica[tripletstore-oxigraph]>=0.6.7\"\n",
"```\n",
"\n",
"---\n",
"\n",
"## Step 1: Build a Knowledge Graph\n",
"\n",
"Start from a small, explicit set of entities and relationships — two people, an organization, and a project.\n"
]
},
{
"cell_type": "code",
"metadata": {},
"source": [
"!pip install \"semantica[tripletstore-oxigraph]>=0.6.7\"\n"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "code",
"metadata": {},
"source": [
"from semantica.kg import GraphBuilder\n",
"\n",
"entities = [\n",
" {\"id\": \"e1\", \"type\": \"Person\", \"name\": \"Alice\", \"properties\": {\"age\": 30, \"role\": \"Engineer\"}},\n",
" {\"id\": \"e2\", \"type\": \"Person\", \"name\": \"Bob\", \"properties\": {\"age\": 35, \"role\": \"Manager\"}},\n",
" {\"id\": \"e3\", \"type\": \"Organization\", \"name\": \"Tech Corp\", \"properties\": {\"founded\": 2010}},\n",
" {\"id\": \"e4\", \"type\": \"Project\", \"name\": \"Project Alpha\", \"properties\": {\"status\": \"active\"}},\n",
"]\n",
"\n",
"relationships = [\n",
" {\"source\": \"e1\", \"target\": \"e2\", \"type\": \"reports_to\", \"properties\": {}},\n",
" {\"source\": \"e1\", \"target\": \"e3\", \"type\": \"works_for\", \"properties\": {}},\n",
" {\"source\": \"e2\", \"target\": \"e3\", \"type\": \"works_for\", \"properties\": {}},\n",
" {\"source\": \"e1\", \"target\": \"e4\", \"type\": \"works_on\", \"properties\": {}},\n",
"]\n",
"\n",
"builder = GraphBuilder()\n",
"knowledge_graph = builder.build({\"entities\": entities, \"relationships\": relationships})\n",
"\n",
"id_to_name = {entity[\"id\"]: entity[\"name\"] for entity in entities}\n",
"\n",
"print(f\"Entities ({len(knowledge_graph['entities'])}):\")\n",
"for entity in knowledge_graph[\"entities\"]:\n",
" print(f\" {entity['id']}: {entity['name']} ({entity['type']}) {entity['properties']}\")\n",
"\n",
"print(f\"\\nRelationships ({len(knowledge_graph['relationships'])}):\")\n",
"for relationship in knowledge_graph[\"relationships\"]:\n",
" print(f\" {id_to_name[relationship['source']]} \"\n",
" f\"--{relationship['type']}--> {id_to_name[relationship['target']]}\")"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 2: Generate a Starter Ontology\n",
"\n",
"`OntologyGenerator` infers OWL classes and properties from graph records. Because `GraphBuilder` keeps business attributes inside each entity's `properties` dictionary while ontology inference reads record fields, we first create a flat **inference view**. The knowledge graph itself remains unchanged. Two settings matter here:\n",
"\n",
"- `base_uri` puts every generated term in *your* namespace\n",
"- `min_occurrences=1` includes classes that occur only once (the default of 2 would drop `Organization` and `Project` from this tiny demo graph)\n",
"\n",
"Note that the generator normalizes names: the relationship type `works_for` becomes the ontology property `worksFor`. That is exactly why the next step maps terms **explicitly** instead of matching names.\n"
]
},
{
"cell_type": "code",
"metadata": {},
"source": [
"from semantica.ontology import OntologyGenerator\n",
"\n",
"BASE_URI = \"https://example.org/company/\"\n",
"\n",
"# Adapt the property-graph representation to the record shape consumed by\n",
"# OntologyGenerator, so age/role/founded/status become declared properties.\n",
"ontology_input = {\n",
" \"entities\": [\n",
" {\n",
" **{key: value for key, value in entity.items() if key != \"properties\"},\n",
" **entity.get(\"properties\", {}),\n",
" }\n",
" for entity in knowledge_graph[\"entities\"]\n",
" ],\n",
" \"relationships\": knowledge_graph[\"relationships\"],\n",
"}\n",
"\n",
"generator = OntologyGenerator(base_uri=BASE_URI, min_occurrences=1)\n",
"ontology = generator.generate_from_graph(ontology_input)\n",
"\n",
"# OntologyGenerator calls datatype properties `data`; TripletStore's public\n",
"# ontology contract calls them `datatype`. Normalize that boundary explicitly.\n",
"store_ontology = {\n",
" **ontology,\n",
" \"properties\": [\n",
" {**prop, \"type\": \"datatype\" if prop[\"type\"] == \"data\" else prop[\"type\"]}\n",
" for prop in ontology[\"properties\"]\n",
" ],\n",
"}\n",
"\n",
"print(\"Classes:\")\n",
"for ontology_class in ontology[\"classes\"]:\n",
" print(f\" {ontology_class['name']:<14} {ontology_class['uri']}\")\n",
"\n",
"print(\"\\nProperties:\")\n",
"for prop in ontology[\"properties\"]:\n",
" print(f\" {prop['name']:<14} {prop['type']:<7} {prop['uri']} \"\n",
" f\"(domain={prop['domain']}, range={prop['range']})\")\n",
"\n",
"assert len(ontology[\"classes\"]) == 3"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 3: Map the Graph to Ontology Terms\n",
"\n",
"The heart of a semantic layer is the mapping contract: which source type, relationship, and property corresponds to which ontology term.\n",
"\n",
"- **Entity types** and **relationship types**: each generated class/property records the source name it was inferred from (`metadata[\"inferred_from\"]`), so the mapping is read off the ontology itself — no fragile name matching between `works_for` and `worksFor`.\n",
"- **Properties**: the flat inference view makes `name`, `age`, `role`, `founded`, and `status` real generated datatype properties. Every mapping therefore points to a term declared in the ontology — no URI is invented only at mapping time.\n"
]
},
{
"cell_type": "code",
"metadata": {},
"source": [
"entity_type_mappings = {\n",
" ontology_class[\"metadata\"][\"inferred_from\"]: ontology_class[\"uri\"]\n",
" for ontology_class in ontology[\"classes\"]\n",
"}\n",
"\n",
"relationship_type_mappings = {\n",
" prop[\"metadata\"][\"inferred_from\"]: prop[\"uri\"]\n",
" for prop in ontology[\"properties\"]\n",
" if prop[\"type\"] == \"object\"\n",
"}\n",
"\n",
"datatype_property_uris = {\n",
" prop[\"metadata\"][\"inferred_from\"]: prop[\"uri\"]\n",
" for prop in ontology[\"properties\"]\n",
" if prop[\"type\"] != \"object\"\n",
"}\n",
"\n",
"property_mappings = datatype_property_uris\n",
"\n",
"semantic_layer = {\n",
" \"graph\": knowledge_graph,\n",
" \"ontology\": ontology,\n",
" \"mappings\": {\n",
" \"entity_type_mappings\": entity_type_mappings,\n",
" \"relationship_type_mappings\": relationship_type_mappings,\n",
" \"property_mappings\": property_mappings,\n",
" },\n",
"}\n",
"\n",
"for mapping_name, mapping in semantic_layer[\"mappings\"].items():\n",
" print(f\"{mapping_name}:\")\n",
" for source, target in mapping.items():\n",
" print(f\" {source:<12} -> {target}\")\n",
"\n",
"# Every type and relationship in the graph must have an ontology term\n",
"assert set(entity_type_mappings) == {entity[\"type\"] for entity in entities}\n",
"assert set(relationship_type_mappings) == {rel[\"type\"] for rel in relationships}\n",
"assert set(property_mappings) == {\"name\", \"age\", \"role\", \"founded\", \"status\"}\n",
"assert set(property_mappings.values()) <= {prop[\"uri\"] for prop in ontology[\"properties\"]}"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 4: Apply the Mappings\n",
"\n",
"Applying the semantic layer means rewriting the graph so every type, relationship, and property key is an ontology term. This *aligned* graph — not the original one — is what gets exported and stored.\n"
]
},
{
"cell_type": "code",
"metadata": {},
"source": [
"aligned_graph = {\n",
" \"entities\": [\n",
" {\n",
" **entity,\n",
" \"type\": entity_type_mappings[entity[\"type\"]],\n",
" \"properties\": {\n",
" property_mappings[\"name\"]: entity[\"name\"],\n",
" **{\n",
" property_mappings[key]: value\n",
" for key, value in entity[\"properties\"].items()\n",
" },\n",
" },\n",
" }\n",
" for entity in knowledge_graph[\"entities\"]\n",
" ],\n",
" \"relationships\": [\n",
" {**rel, \"type\": relationship_type_mappings[rel[\"type\"]]}\n",
" for rel in knowledge_graph[\"relationships\"]\n",
" ],\n",
"}\n",
"\n",
"print(\"Aligned entity sample:\")\n",
"sample = aligned_graph[\"entities\"][0]\n",
"print(f\" id: {sample['id']}\")\n",
"print(f\" type: {sample['type']}\")\n",
"for key, value in sample[\"properties\"].items():\n",
" print(f\" {key} = {value}\")\n",
"\n",
"print(\"\\nAligned relationship sample:\")\n",
"print(f\" {aligned_graph['relationships'][0]['type']}\")"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 5: Store and Export Complete Ontology-Aligned RDF\n",
"\n",
"`TripletStore.store()` materializes both the ontology declarations and the aligned instance graph. We then read those triples through the store's public API and serialize that complete RDF graph as Turtle. This avoids the compact `RDFExporter` entity projection, which does not include arbitrary entries from an entity's `properties` dictionary.\n"
]
},
{
"cell_type": "code",
"metadata": {},
"source": [
"from rdflib import Graph, Literal, URIRef\n",
"from rdflib.namespace import OWL, RDF\n",
"from semantica.triplet_store import TripletStore\n",
"\n",
"store = TripletStore(backend=\"oxigraph\")\n",
"result = store.store(aligned_graph, store_ontology)\n",
"print(f\"Stored triples: {result['processed']} (failed: {result['failed']})\")\n",
"\n",
"rdf_graph = Graph()\n",
"for triplet in store.get_triplets():\n",
" datatype = triplet.metadata.get(\"datatype\")\n",
" if datatype:\n",
" object_term = Literal(triplet.object, datatype=URIRef(datatype))\n",
" elif triplet.object.startswith((\"http://\", \"https://\", \"urn:\")):\n",
" object_term = URIRef(triplet.object)\n",
" else:\n",
" object_term = Literal(triplet.object)\n",
" rdf_graph.add((URIRef(triplet.subject), URIRef(triplet.predicate), object_term))\n",
"\n",
"rdf_graph.serialize(destination=\"semantic_layer.ttl\", format=\"turtle\")\n",
"turtle = open(\"semantic_layer.ttl\", encoding=\"utf-8\").read()\n",
"print(turtle[:600])\n",
"\n",
"# The exported RDF contains declarations plus mapped instance facts.\n",
"declared_datatype_properties = {\n",
" str(subject) for subject in rdf_graph.subjects(RDF.type, OWL.DatatypeProperty)\n",
"}\n",
"assert result[\"failed\"] == 0\n",
"assert set(property_mappings.values()) <= declared_datatype_properties\n",
"assert (\n",
" URIRef(BASE_URI + \"e1\"),\n",
" URIRef(property_mappings[\"role\"]),\n",
" Literal(\"Engineer\"),\n",
") in rdf_graph\n",
"assert (\n",
" URIRef(BASE_URI + \"e1\"),\n",
" URIRef(relationship_type_mappings[\"works_for\"]),\n",
" URIRef(BASE_URI + \"e3\"),\n",
") in rdf_graph\n",
"print(\"... exported semantic_layer.ttl\")"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Step 6: Query the Semantic Layer\n",
"\n",
"The embedded Oxigraph backend runs in memory, so there is nothing to start beyond installing the `tripletstore-oxigraph` extra. The organization is constrained by its mapped `name` predicate; the query therefore means *Tech Corp*, rather than accidentally matching employees of every organization.\n"
]
},
{
"cell_type": "code",
"metadata": {},
"source": [
"query = f\"\"\"\n",
"SELECT ?name ?role WHERE {{\n",
" ?person <{BASE_URI}worksFor> ?org .\n",
" ?org <{BASE_URI}name> \"Tech Corp\" .\n",
" ?person <{BASE_URI}name> ?name .\n",
" ?person <{BASE_URI}role> ?role .\n",
"}}\n",
"ORDER BY ?name\n",
"\"\"\"\n",
"query_result = store.execute_query(query)\n",
"\n",
"print(\"\\nWho works for Tech Corp, and in which role?\")\n",
"for binding in query_result.bindings:\n",
" print(f\" {binding['name']['value']} — {binding['role']['value']}\")\n",
"\n",
"assert [(row[\"name\"][\"value\"], row[\"role\"][\"value\"]) for row in query_result.bindings] == [\n",
" (\"Alice\", \"Engineer\"),\n",
" (\"Bob\", \"Manager\"),\n",
"]"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## 🧹 Optional: Clean Up\n"
]
},
{
"cell_type": "code",
"metadata": {},
"source": [
"from pathlib import Path\n",
"\n",
"ttl_file = Path(\"semantic_layer.ttl\")\n",
"if ttl_file.exists():\n",
" ttl_file.unlink()\n",
" print(f\"Removed {ttl_file}\")"
],
"execution_count": null,
"outputs": []
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Summary\n",
"\n",
"A minimal semantic layer is a composition, and you have now built each part:\n",
"\n",
"1. **Knowledge graph** — `GraphBuilder` from explicit entities and relationships\n",
"2. **Ontology** — `OntologyGenerator` with your `base_uri`\n",
"3. **Explicit mappings** — entity types, relationship types, and properties, each tied to an ontology term\n",
"4. **Ontology-aligned RDF** — the mappings applied to the graph, materialized with `TripletStore`, and serialized to Turtle from the store's own triples\n",
"5. **Queryable store** — `TripletStore` (embedded Oxigraph) answering a SPARQL question over the shared vocabulary\n",
"\n",
"### Where to go next\n",
"\n",
"The production version of this workflow — hand-designed governed ontologies, explicit source-to-ontology mappings from a warehouse, n-ary modeling, SHACL validation, provenance, and versioning — is covered in [Advanced: Manual Ontology + Snowflake Mapping](../advanced/13_Manual_Ontology_Snowflake_Mapping.ipynb).\n"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "Python 3",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3",
"version": "3.11.9"
}
},
"nbformat": 4,
"nbformat_minor": 2
}
+1 -1
View File
@@ -226,7 +226,7 @@ Pick your goal to see the minimum imports and a working skeleton.
</Tab>
<Tab title="MCP — Claude / Cursor">
Use Semantica from Claude Desktop, Cursor, VS Code, or any MCP-aware tool — no Python code required after setup. 12 tools available instantly.
Use Semantica from Claude Desktop, Cursor, VS Code, or any MCP-aware tool — no Python code required after setup. 15 tools available instantly.
**Step 1 — Install:**
```bash
+2 -2
View File
@@ -53,7 +53,7 @@ python -c "import semantica; print(semantica.__version__)"
- **semantica-server** — Starts the REST API server. Binds to `0.0.0.0:8000`. Use this when another service or application needs programmatic access to Semantica over HTTP.
- **semantica-worker** — Background task processor. Run alongside `semantica-server` when you need async pipeline execution outside the request cycle. Start the server first, then start one or more workers pointing at the same backend.
- **semantica-explorer** — Launches the browser dashboard. Requires `pip install semantica[explorer]`. Use this to explore a saved knowledge graph interactively. See [Explorer Setup](explorer-setup).
- **semantica-mcp** — Runs the MCP server over stdio. Configure it in your MCP client's settings file to expose all 12 tools and 3 resources to Claude Desktop, Cursor, Windsurf, or any MCP-aware client. See [MCP Server](reference/mcp_server).
- **semantica-mcp** — Runs the MCP server over stdio. Configure it in your MCP client's settings file to expose all 15 tools and 3 resources to Claude Desktop, Cursor, Windsurf, or any MCP-aware client. See [MCP Server](reference/mcp_server).
## Usage Examples
@@ -229,6 +229,6 @@ Install the [Microsoft Visual C++ Redistributable](https://aka.ms/vs/17/release/
## Next Steps
- [Explorer Setup](explorer-setup) — Build a graph, save it, and launch the browser dashboard.
- [MCP Server](reference/mcp_server) — All 12 tools and 3 resources exposed over the MCP protocol.
- [MCP Server](reference/mcp_server) — All 15 tools and 3 resources exposed over the MCP protocol.
- [Installation](installation) — Virtual environments, optional extras, and platform-specific notes.
- [Quickstart](quickstart) — End-to-end pipeline walkthrough with working code.
+1
View File
@@ -36,6 +36,7 @@ Essential guides to master the Semantica framework.
- **[Graph Store](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/09_Graph_Store.ipynb)** — Persisting knowledge graphs in Neo4j or FalkorDB. Topics: Neo4j, Cypher, Persistence · *Intermediate*
- **[Ontology](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/14_Ontology.ipynb)** — Defining domain schemas and ontologies to structure your data. Topics: OWL, RDF, Schema Design · *Intermediate*
- **[Seed Data](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/25_Seed_Data.ipynb)** — Bootstrapping a knowledge graph from trusted CSV, JSON, database, and API sources before extraction runs. Topics: SeedDataManager, Foundation Graphs · *Intermediate*
- **[Semantic Layer Basics](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/26_Semantic_Layer_Basics.ipynb)** — Capstone tutorial that combines a knowledge graph, generated ontology, explicit mappings, ontology-aligned RDF, and a SPARQL query. Topics: Semantic Layer, Ontology Mapping, Oxigraph, SPARQL · *Intermediate*
## Advanced Concepts
+2 -1
View File
@@ -106,7 +106,8 @@
"integrations/langchain",
"integrations/docling",
"integrations/snowflake",
"integrations/databricks"
"integrations/databricks",
"integrations/salesforce"
]
},
{
+5 -5
View File
@@ -16,7 +16,7 @@ icon: "circle-question"
| Python version? | 3.8+ (3.11+ recommended) |
| API key required? | Optional: pattern extraction works with no keys |
| Works with LangChain / LlamaIndex? | Yes: Semantica is a layer on top, not a replacement |
| Production-ready? | Yes: 1,000+ tests, v0.5.0 ships with 12 security fixes |
| Production-ready? | Yes: 1,000+ tests, security fixes shipped in every release (see [CHANGELOG](https://github.com/semantica-agi/semantica/blob/main/CHANGELOG.md)) |
| Latest version? | **v0.6.7** (August 2026) |
| Local LLMs? | Yes: Ollama via LiteLLM, HuggingFaceLLM for air-gapped |
@@ -70,9 +70,9 @@ Yes: MIT licensed, no vendor lock-in, no paywalled features. Some capabilities r
<Accordion title="What's the latest version?" icon="star">
**v0.5.0**: released May 2026.
**v0.6.7**: released August 2026.
Highlights: Ontology Hub, Distance Intelligence, Parquet/XML ingestion, 12 security fixes, Graph Explorer redesign, NER gateway fix.
Highlights: first-class LangChain integration, SAP OData ingestor, human-editable Markdown round-trip persistence for `ContextGraph`, a structured Action layer for the reasoning engine, and a public `run_shacl_validation` entry point. The 0.6.x line also added first-class CrewAI support and the Semantica RDF vocabulary with deterministic IRIs. See the [CHANGELOG](https://github.com/semantica-agi/semantica/blob/main/CHANGELOG.md) for the full history.
```bash
pip install --upgrade semantica
@@ -269,13 +269,13 @@ Groq, OpenAI, Anthropic, Google Gemini, Ollama (fully local), DeepSeek, Novita A
<Accordion title="Is Semantica production-ready?" icon="shield-check">
Yes. v0.5.0 ships with:
Yes. Every release ships with:
- 1,000+ passing tests across Python 3.83.12
- `PipelineValidator` and `FailureHandler` with exponential backoff and configurable retry policies
- W3C PROV-O provenance tracking across all modules
- Change management with SHA-256 checksums and full audit trails
- 12 security vulnerability fixes: eval injection, pickle deserialization, SQL injection, XXE, SSRF, ReDoS, path traversal, and more
- Ongoing security hardening: eval injection, pickle deserialization, SQL injection, XXE, SSRF, ReDoS, and path traversal fixes have all landed across recent releases (see the [CHANGELOG](https://github.com/semantica-agi/semantica/blob/main/CHANGELOG.md) security sections)
</Accordion>
+1 -1
View File
@@ -183,7 +183,7 @@ icon: "rocket"
}
```
12 tools available instantly: extract entities, query graph, record decisions, run reasoning, export results.
15 tools available instantly: extract entities, query graph, record decisions, run reasoning, export results.
**Next:** [MCP Server reference →](reference/mcp_server)
</Tab>
+152 -24
View File
@@ -28,6 +28,7 @@ The `semantica.llms` module provides a unified interface for connecting to Large
## When To Use / When Not To Use
**Use LLM integrations for:**
- Text generation, summarization, and question-answering tasks
- Complex reasoning that requires natural language understanding
- Structured data extraction from unstructured text
@@ -35,6 +36,7 @@ The `semantica.llms` module provides a unified interface for connecting to Large
- Tasks where context, ambiguity, or domain knowledge matter
**Deterministic tools may be better for:**
- Pattern matching that regular expressions can handle
- Simple rule-based classification with clear criteria
- Mathematical calculations or statistical analysis
@@ -42,6 +44,7 @@ The `semantica.llms` module provides a unified interface for connecting to Large
- Data transformations with known logic
**A full LLM may be unnecessary for:**
- Simple keyword search or exact string matching
- Deterministic workflows with predefined decision trees
- High-frequency, low-latency operations where inference overhead matters
@@ -59,7 +62,7 @@ Four factors drive provider selection, each optimized for different use cases:
**Accuracy** matters most in high-stakes decisions: clinical contraindication checks, credit committee reasoning, and legal document analysis. Frontier models like Claude or GPT-4 available through `LiteLLM` provide the strongest reasoning capabilities.
**Data residency** constraints eliminate cloud providers for classified or HIPAA-regulated workloads. `HuggingFaceLLM` with local model paths enables fully air-gapped deployments without network calls.
**Data residency** constraints eliminate cloud providers for classified or HIPAA-regulated workloads. `HuggingFaceLLM` with local model paths, or `Ollama` pointed at a local server, both enable fully air-gapped deployments without network calls.
**Cost at scale** favors high-throughput providers like Novita AI for bulk extraction pipelines processing thousands of documents per hour where per-token costs accumulate quickly.
@@ -143,6 +146,131 @@ risk_data = oai.generate_structured(
The default model `gpt-3.5-turbo` is fine for classification and light extraction. Switch to `gpt-4o` for complex multi-step regulatory reasoning or document understanding.
## Anthropic — Complex Reasoning and Structured Extraction
**Anthropic** provides the Claude model family, built with an emphasis on careful, instruction-following behavior and strong performance on multi-step reasoning, long-document analysis, and code-related tasks. Claude models tend to be more cautious about ambiguous instructions than other providers. That matters when the cost of a confidently wrong answer is high.
The `Anthropic` provider wraps the Claude API. Reach for it when the task involves reasoning through several dependent steps (not just single-turn extraction), when you're processing long source documents that need to stay in context, or when you need schema-validated structured output rather than best-effort JSON.
Install with `pip install "semantica[llm-anthropic]"` (or just `pip install anthropic`) before using this provider.
```python
from semantica.llms import Anthropic
claude = Anthropic(model="claude-sonnet-4-6", api_key="YOUR_ANTHROPIC_KEY")
# api_key falls back to the ANTHROPIC_API_KEY environment variable
# is_available() only confirms a client was constructed from some key.
# It does not validate the key or check network reachability - an
# invalid or expired key still passes this check and fails at generate().
if not claude.is_available():
raise RuntimeError("Anthropic provider not configured - set ANTHROPIC_API_KEY")
# Plain generation - multi-step reasoning over a contract clause
verdict = claude.generate(
"A vendor contract has a 30-day termination-for-convenience clause "
"but a 90-day data-return obligation that survives termination. "
"If the customer terminates on day 1, when must vendor-held data "
"be returned? Answer with the date basis only.",
temperature=0.1,
)
print(verdict)
# "Day 120 from termination notice. The 90-day return period runs from
# the termination date (day 30), not from the notice date."
# Structured, schema-validated output
from pydantic import BaseModel
class ContractRisk(BaseModel):
clause: str
risk_level: str
days_to_deadline: int
risk = claude.generate_typed(
"Extract the termination clause risk from: vendor contract, "
"30-day termination for convenience, 90-day post-termination "
"data return obligation.",
schema=ContractRisk,
)
print(risk.risk_level, risk.days_to_deadline)
# "medium" 90
```
Model selection follows the same tier structure as the other providers: a Haiku model for high-volume classification where cost matters more than depth, a Sonnet model as the default for most extraction and reasoning tasks, an Opus model when a task genuinely needs the deepest reasoning available and latency/cost are secondary. Check Anthropic's docs for the current model identifiers, since they're versioned and change over time.
## Gemini — Long Context and Multimodal Input
**Gemini** is Google's model family, with a context window large enough to hold entire codebases or long regulatory filings in a single call, and native support for image and document input alongside text. Reach for it when a task needs to reference a large amount of source material at once, or when the input isn't plain text.
The `Gemini` provider tries the newer `google-genai` SDK first and falls back to the older `google-generativeai` package if that's what's installed. Install with `pip install "semantica[llm-gemini]"` (or `pip install google-genai`) before using this provider.
```python
from semantica.llms import Gemini
gemini = Gemini(model="gemini-pro", api_key="YOUR_GEMINI_KEY")
# api_key falls back to the GEMINI_API_KEY environment variable
if not gemini.is_available():
raise RuntimeError("Gemini provider not configured - set GEMINI_API_KEY")
response = gemini.generate(
"Summarize the key obligations in a standard NDA in three bullet points."
)
print(response)
data = gemini.generate_structured(
"Extract the party names and effective date from: "
"This Agreement is entered into between Acme Corp and Globex LLC, "
"effective January 1, 2026."
)
print(data)
```
## Ollama — Local, Air-Gapped Inference
**Ollama** runs models entirely on your own machine, with no API key and no outbound network call. It's the right choice for air-gapped environments, offline development, or any workload where the source data can't leave the local network.
Unlike the other providers here, `Ollama` takes a `base_url` instead of an `api_key`. It talks to a local Ollama server over HTTP. Start the server with `ollama serve` and pull a model with `ollama pull llama2` before using this provider. Install the Python client with `pip install "semantica[llm-ollama]"` (or `pip install ollama`).
```python
from semantica.llms import Ollama
llm = Ollama(model="llama2", base_url="http://localhost:11434")
if not llm.is_available():
raise RuntimeError("Ollama provider not configured - is 'ollama serve' running?")
response = llm.generate("Explain the difference between a hash map and a tree map.")
print(response)
```
`is_available()` for Ollama does a real connectivity check (it calls the server's `list()` endpoint), unlike the API-key-based providers above, so a `False` here usually means the server isn't running rather than a missing credential.
## DeepSeek — Budget Reasoning at Scale
**DeepSeek** exposes an OpenAI-compatible API at a fraction of the cost of the larger US providers, with reasoning quality that holds up well for extraction and classification work. It's a reasonable default when you're processing a large volume of documents and don't need the deepest reasoning tier.
Install with `pip install "semantica[llm-deepseek]"` (or `pip install openai`, since DeepSeek is accessed through the OpenAI client pointed at a different base URL).
```python
from semantica.llms import DeepSeek
llm = DeepSeek(model="deepseek-chat", api_key="YOUR_DEEPSEEK_KEY")
# api_key falls back to the DEEPSEEK_API_KEY environment variable
if not llm.is_available():
raise RuntimeError("DeepSeek provider not configured - set DEEPSEEK_API_KEY")
response = llm.generate("List three risks of using a floating IP in a Kubernetes ingress.")
print(response)
data = llm.generate_structured(
"Extract the CVE ID and affected product from: "
"CVE-2024-3400 affects PAN-OS GlobalProtect gateways."
)
print(data)
```
## LiteLLM — One Interface, 100+ Providers
**LiteLLM** is a universal adapter that provides a single interface to over 100 different LLM providers, including Anthropic Claude, Azure OpenAI, AWS Bedrock, Google Vertex AI, and local Ollama instances. It acts as a translation layer, converting your unified API calls into provider-specific requests, enabling easy switching between providers without code changes.
@@ -306,30 +434,32 @@ for t in triplets:
## Novita AI — Cost-Efficient Bulk Extraction
Novita AI exposes an OpenAI-compatible API and is available as a built-in provider for the extraction layer. It is accessed differently from the `semantica.llms` classes — through `create_provider` from `semantica.semantic_extract.providers` — making it the right choice for high-volume NER pipelines where per-call cost matters.
**Novita AI** exposes an OpenAI-compatible API at low per-call cost, making it a reasonable choice for high-volume NER pipelines where cost matters more than getting the single best answer.
Install with `pip install "semantica[llm-novita]"` (or `pip install openai`, since Novita is accessed through the OpenAI client pointed at a different base URL).
```python
from semantica.llms import Novita
llm = Novita(model="deepseek/deepseek-v3.2", api_key="YOUR_NOVITA_KEY")
# api_key falls back to the NOVITA_API_KEY environment variable
if not llm.is_available():
raise RuntimeError("Novita provider not configured - set NOVITA_API_KEY")
response = llm.generate("Summarize the Basel III leverage ratio requirement.")
data = llm.generate_structured(
"Extract drug names and dosages from: "
"Patient received warfarin 5mg daily, aspirin 75mg daily, metformin 500mg twice daily."
)
```
Novita is also reachable as a provider name string for the NER interface, without going through the `Novita` class directly:
```python
from semantica.semantic_extract.providers import create_provider
from semantica.semantic_extract import NamedEntityRecognizer
# create_provider pools instances — same key reuses the same object
provider = create_provider(
"novita",
api_key="YOUR_NOVITA_KEY", # or set NOVITA_API_KEY env var
model="deepseek/deepseek-v3.2", # default model
)
if provider.is_available():
# Plain generation
response = provider.generate("Summarise the Basel III leverage ratio requirement.")
# Structured extraction — returns parsed dict
data = provider.generate_structured(
"Extract drug names and dosages from: "
"Patient received warfarin 5mg daily, aspirin 75mg daily, metformin 500mg twice daily."
)
# Use Novita through the NER interface — provider name as string
ner = NamedEntityRecognizer(
methods=["llm"],
provider="novita",
@@ -339,11 +469,9 @@ entities = ner.extract_entities(
"CVE-2024-3400 is exploited by UNC3886 targeting PAN-OS GlobalProtect."
)
for e in entities:
print("{} ({}) conf={:.2f}".format(e.text, e.label, e.confidence))
print("{} ({}) conf={:.2f}".format(e.text, e.label, e.confidence))
```
Novita requires the `openai` Python client under the hood — install with `pip install "semantica[llm-openai]"` or `pip install openai`.
## Domain Examples
<Tabs>
+4 -2
View File
@@ -11,7 +11,7 @@ MCP stands for the Model Context Protocol. It is an open standard that allows ex
The Semantica MCP server exposes your knowledge graph as 12 callable tools. By connecting it, any compatible AI client can traverse the graph live, record decisions, run analytics, and export results during a conversation — without you having to write custom tool wrappers.
<Info>
The Semantica MCP server exposes 12 tools and 3 read-only resources. All tools accept and return JSON. No configuration beyond an optional environment variable for graph persistence is required.
The Semantica MCP server exposes 15 tools and 3 read-only resources. All tools accept and return JSON. No configuration beyond an optional environment variable for graph persistence is required.
</Info>
## Architecture & Communication
@@ -132,7 +132,7 @@ docker run --rm -i \
ghcr.io/semantica-agi/semantica-mcp:latest
```
## What the Agent Can Do: The 12 Tools
## What the Agent Can Do: The 15 Tools
Once connected, the LLM can call any of these tools during a conversation. The agent chains them automatically — you do not orchestrate the sequence, you just describe what you want.
@@ -140,6 +140,8 @@ Once connected, the LLM can call any of these tools during a conversation. The a
**Knowledge graph manipulation**`add_entity` adds a node, `add_relationship` adds a directed edge. After extraction, the agent calls these to persist what it found into the live graph.
**Live graph queries and edits**`query_graph` reads the graph without exporting it: fetch one node, walk its neighbours up to five hops, or keyword-search nodes. `update_node` merges properties onto an existing node (for example marking a task node `done`), and `delete_node` archives a node it no longer tracks. When `SEMANTICA_KG_PATH` is set, `update_node` and `delete_node` write their changes back to that file so they survive a restart.
**Decision intelligence**`record_decision` writes a decision as a provenance node with confidence score, reasoning, and decision maker identity. `query_decisions` retrieves past decisions by query or category. `find_precedents` finds the most similar past decisions by semantic similarity. `get_causal_chain` traces decision causality upstream or downstream.
**Reasoning**`run_reasoning` applies forward-chaining IF/THEN rules over a set of facts and returns derived conclusions.
+2 -2
View File
@@ -369,7 +369,7 @@ Semantica was designed for domains where every decision must be explainable and
| `semantica.reasoning` | Forward chaining, Rete, deductive, abductive, SPARQL, Datalog |
| `semantica.ontology` | SHACL, SKOS, alignments, diff/migration, auto-generation, OWL/RDF |
| `semantica.explorer` | FastAPI Knowledge Explorer, Ontology Hub, Distance Intelligence, SHACL Studio |
| `semantica.mcp_server` | MCP stdio server: 12 tools for Claude Desktop, VS Code, Cursor, Windsurf, Cline |
| `semantica.mcp_server` | MCP stdio server: 15 tools for Claude Desktop, VS Code, Cursor, Windsurf, Cline |
| `semantica.vector_store` | FAISS, Pinecone, Weaviate, Qdrant, Milvus, PgVector |
| `semantica.graph_store` | Neo4j, FalkorDB, Apache AGE, Amazon Neptune |
| `semantica.triplet_store` | In-memory and persistent RDF triple store with SPARQL |
@@ -404,7 +404,7 @@ Semantica was designed for domains where every decision must be explainable and
- 1,000+ passing tests with full regression coverage
- `PipelineValidator` catches configuration errors at startup
- `FailureHandler` with exponential backoff and dead-letter queues
- 12 security vulnerabilities fixed in v0.5.0
- Ongoing security hardening: fixes shipped in every release ([CHANGELOG](https://github.com/semantica-agi/semantica/blob/main/CHANGELOG.md))
**Modular by Design** — Import only what you need.
- Use `NERExtractor` without a graph store
+376
View File
@@ -0,0 +1,376 @@
---
title: "Salesforce Integration"
description: "Ingest CRM records from Salesforce sObjects and SOQL queries into Semantica's KG pipeline."
icon: "cloud"
---
> Extract Accounts, Contacts, Opportunities, and custom objects from Salesforce into Semantica with username/password/security-token, JWT bearer, or session-based authentication.
## Installation
```bash
# Install with Salesforce support
pip install "semantica[db-salesforce]"
# Or install the connector separately
pip install simple-salesforce>=1.12.0
```
## Basic Usage
```python
from semantica.ingest import SalesforceIngestor
import os
ingestor = SalesforceIngestor(
username=os.getenv("SALESFORCE_USERNAME"),
password=os.getenv("SALESFORCE_PASSWORD"),
security_token=os.getenv("SALESFORCE_SECURITY_TOKEN"),
domain=os.getenv("SALESFORCE_DOMAIN", "login"), # "test" for sandbox
)
data = ingestor.ingest_sobject("Account", fields=["Id", "Name", "Industry"], limit=1000)
print(f"Retrieved {data.row_count} of {data.total_size} matching records")
print(f"Columns: {data.columns}")
```
<Tip>
Use environment variables (or a `.env` file with `python-dotenv`) to keep credentials out of source code. `SalesforceIngestor()` with no arguments reads from `SALESFORCE_*` environment variables automatically.
</Tip>
## Authentication Methods
<Tabs>
<Tab title="Username / Password / Security Token">
```python
import os
from semantica.ingest import SalesforceIngestor
ingestor = SalesforceIngestor(
username=os.getenv("SALESFORCE_USERNAME"),
password=os.getenv("SALESFORCE_PASSWORD"),
security_token=os.getenv("SALESFORCE_SECURITY_TOKEN"),
domain="login", # production; use "test" for sandbox
)
```
Set the required environment variables before running:
```bash
export SALESFORCE_USERNAME="your-username@example.com"
export SALESFORCE_PASSWORD="your-password"
export SALESFORCE_SECURITY_TOKEN="your-security-token"
```
The standard server-side flow. The security token is appended to the
password during Salesforce SOAP login. Generate or reset it under
**Settings → My Personal Information → Reset My Security Token**.
</Tab>
<Tab title="JWT Bearer (Recommended for CI/CD)">
```python
import os
from semantica.ingest import SalesforceIngestor
ingestor = SalesforceIngestor(
username=os.getenv("SALESFORCE_USERNAME"),
consumer_key=os.getenv("SALESFORCE_CONSUMER_KEY"),
privatekey_file=os.getenv("SALESFORCE_PRIVATE_KEY_FILE"),
domain="login", # or "test" for sandbox
)
```
```bash
export SALESFORCE_USERNAME="your-username@example.com"
export SALESFORCE_CONSUMER_KEY="your-connected-app-consumer-key"
export SALESFORCE_PRIVATE_KEY_FILE="/path/to/server.key"
```
The JWT bearer flow authenticates with a signed token — no password
is transmitted. Ideal for server-to-server integrations and CI/CD
pipelines. Requires a Salesforce connected app configured with
**Use digital signatures** and the pre-authorised user listed under
**Manage → Profiles / Permission Sets**.
If you prefer to pass the key material as a string instead of a file
path, use `SALESFORCE_PRIVATE_KEY` (the PEM contents) in place of
`SALESFORCE_PRIVATE_KEY_FILE`.
</Tab>
<Tab title="Session ID + Instance URL">
```python
ingestor = SalesforceIngestor(
session_id=os.getenv("SALESFORCE_SESSION_ID"),
instance_url=os.getenv("SALESFORCE_INSTANCE_URL"),
)
```
Use this when your environment already manages the OAuth token
lifecycle (e.g. a connected app obtaining tokens via the web-server
or device flow). Pass the access token as `session_id` and the full
instance URL (e.g. `https://myorg.my.salesforce.com`) as
`instance_url`.
</Tab>
<Tab title="Sandbox">
```python
import os
from semantica.ingest import SalesforceIngestor
ingestor = SalesforceIngestor(
username=os.getenv("SALESFORCE_USERNAME"),
password=os.getenv("SALESFORCE_PASSWORD"),
security_token=os.getenv("SALESFORCE_SECURITY_TOKEN"),
domain="test", # routes to test.salesforce.com
)
```
```bash
export SALESFORCE_USERNAME="your-sandbox-username@example.com.sandbox"
export SALESFORCE_PASSWORD="your-password"
export SALESFORCE_SECURITY_TOKEN="your-security-token"
export SALESFORCE_DOMAIN="test"
```
Replace `domain="login"` with `domain="test"` (or set
`SALESFORCE_DOMAIN=test` in your environment) to connect to a
developer or full sandbox.
</Tab>
</Tabs>
### Environment variables
All constructor parameters have environment-variable fallbacks:
| Variable | Parameter | Default |
|---|---|---|
| `SALESFORCE_USERNAME` | `username` | — |
| `SALESFORCE_PASSWORD` | `password` | — |
| `SALESFORCE_SECURITY_TOKEN` | `security_token` | — |
| `SALESFORCE_DOMAIN` | `domain` | `"login"` |
| `SALESFORCE_INSTANCE_URL` | `instance_url` | — |
| `SALESFORCE_SESSION_ID` | `session_id` | — |
| `SALESFORCE_CONSUMER_KEY` | `consumer_key` | — |
| `SALESFORCE_PRIVATE_KEY_FILE` | `privatekey_file` | — |
| `SALESFORCE_PRIVATE_KEY` | `privatekey` | — |
| `SALESFORCE_API_VERSION` | `api_version` | library default (`59.0`) |
## Object Ingestion
### Ingest a standard object
```python
data = ingestor.ingest_sobject(
"Account",
fields=["Id", "Name", "Industry", "AnnualRevenue", "BillingCity"],
where="Type = 'Customer' AND AnnualRevenue > 1000000",
order_by="Name ASC",
limit=5000,
)
print(f"Retrieved {data.row_count} of {data.total_size} matching records")
```
<Note>
`data.row_count` is the number of records in `data.data` (i.e. what was actually returned after any `limit`). `data.total_size` is Salesforce's `totalSize` — the number of records matching the query *before* the limit. Compare them to know whether you got all results.
</Note>
### Ingest a custom object
Custom objects end with `__c` in their API name:
```python
data = ingestor.ingest_sobject(
"My_Custom_Object__c",
fields=["Id", "Name", "Custom_Field__c"],
)
```
Relationship traversal fields (`Owner.Name`) are also supported:
```python
data = ingestor.ingest_sobject(
"Contact",
fields=["Id", "Name", "Email", "Account.Name", "Owner.Name"],
limit=10000,
)
```
### Let Semantica choose the fields
When `fields` is omitted, all selectable fields are fetched via `describe()`
(one extra API call). Compound address and geolocation fields (`type=address`,
`type=location`) are automatically excluded — select their components
(`BillingStreet`, `BillingCity`, `Location__Latitude__s`, …) individually if
you need them.
```python
data = ingestor.ingest_sobject("Opportunity")
```
## Raw SOQL Ingestion
Pass any valid SOQL query verbatim — pagination is handled automatically:
```python
data = ingestor.ingest_query("""
SELECT Id, Name, StageName, Amount, CloseDate,
Account.Name, Owner.Name
FROM Opportunity
WHERE IsClosed = false
ORDER BY CloseDate ASC
""")
print(f"Open opportunities: {data.row_count}")
```
The query is passed to the Salesforce REST API unchanged. The caller is
responsible for SOQL correctness and safety.
<Warning>
`ingest_query` does not validate or sanitise the SOQL string. Use
`ingest_sobject` (which validates sObject names, field names, and WHERE/ORDER
BY fragments) when building queries from application-controlled inputs.
</Warning>
## Document Export
Convert ingested records to the Semantica document format for use with
`GraphBuilder`:
```python
documents = ingestor.export_as_documents(
data,
id_field="Id", # default; Salesforce 18-char record Id
text_fields=["Name", "Description"], # omit to join all string fields
)
print(f"Created {len(documents)} documents")
# Each document:
# {
# "id": "001xx000003GYk2AAG",
# "text": "Acme Corp Enterprise software company",
# "metadata": {
# "source": "salesforce",
# "sobject": "Account",
# "instance_url": "https://myorg.my.salesforce.com",
# "row_data": { ... full cleaned record ... }
# }
# }
```
Feed the documents directly into `GraphBuilder`:
```python
from semantica.kg import GraphBuilder
builder = GraphBuilder()
kg = builder.build(documents)
```
## Object and Schema Discovery
```python
# List all accessible sObjects
sobject_names = ingestor.list_sobjects()
print(sobject_names[:10]) # ["Account", "Case", "Contact", ...]
# Inspect fields for a specific sObject
schema = ingestor.get_sobject_schema("Account")
for field in schema["fields"]:
print(f"{field['name']}: {field['type']} (nillable={field['nillable']})")
```
## Context Manager
Prefer the context manager for long-running jobs — it opens one connection on
entry and closes it on exit, so every ingestion call inside the `with` block
reuses the same authenticated session:
```python
with SalesforceIngestor(
username=os.getenv("SALESFORCE_USERNAME"),
password=os.getenv("SALESFORCE_PASSWORD"),
security_token=os.getenv("SALESFORCE_SECURITY_TOKEN"),
) as sf:
accounts = sf.ingest_sobject("Account", limit=10000)
contacts = sf.ingest_sobject("Contact", limit=10000)
sobjects = sf.list_sobjects()
```
## Convenience Function
Use `ingest_salesforce()` for one-liner ingestion:
```python
from semantica.ingest import ingest_salesforce
# Fetch records
data = ingest_salesforce(
method="sobject",
sobject_name="Account",
fields=["Id", "Name", "Industry"],
limit=500,
)
# Execute raw SOQL (credentials from environment variables)
data = ingest_salesforce(
method="query",
soql="SELECT Id, Name FROM Contact WHERE IsActive = true",
)
# Ingest + export to documents in one step
docs = ingest_salesforce(
method="documents",
sobject_name="Account",
text_fields=["Name", "Description"],
limit=1000,
)
# List accessible sObjects
sobject_names = ingest_salesforce(method="list_sobjects")
```
Or use the unified `ingest()` dispatcher:
```python
from semantica.ingest import ingest
result = ingest(
None,
source_type="salesforce",
method="sobject",
sobject_name="Account",
fields=["Id", "Name"],
limit=500,
)
data = result["data"] # SalesforceData
```
## Troubleshooting
```python
import os
from semantica.ingest import SalesforceConnector
connector = SalesforceConnector(
username=os.getenv("SALESFORCE_USERNAME"),
password=os.getenv("SALESFORCE_PASSWORD"),
security_token=os.getenv("SALESFORCE_SECURITY_TOKEN"),
)
if not connector.test_connection():
print("Connection failed: check username, password, security token, and domain")
```
Common causes of authentication failures:
- **Wrong domain**: production orgs use `domain="login"`; sandboxes use `domain="test"`.
- **Stale security token**: reset it under **Settings → Reset My Security Token**. The new token is emailed to you.
- **IP restriction**: your org's trusted IP ranges may block the originating IP. Check **Setup → Network Access**.
- **API access disabled**: ensure the connected profile has the **API Enabled** permission.
## See Also
- [Ingest Module](../reference/ingest) — Full `SalesforceIngestor` API and all other ingestors.
- [Snowflake Integration](snowflake) — Relational warehouse connector with a similar design.
- [Databricks Integration](databricks) — Lakehouse connector.
- [Installation](../installation) — All optional dependency extras.
- [Knowledge Graph](../reference/kg) — Build a KG from ingested Salesforce data.
+1 -1
View File
@@ -438,7 +438,7 @@ Exposes Semantica as an MCP stdio server for IDE and agent integrations.
python -m semantica.mcp_server
```
**Integrations:** Claude Desktop, VS Code, Cursor, Windsurf, Cline: 12 MCP tools exposed
**Integrations:** Claude Desktop, VS Code, Cursor, Windsurf, Cline: 15 MCP tools exposed
### Seed
+5 -4
View File
@@ -5,7 +5,7 @@ icon: "rocket"
---
<Info>
**v0.5.0**Ontology Hub, Distance Intelligence, Parquet & XML ingestion, 12 security fixes. <a href="https://github.com/semantica-agi/semantica/releases" style={{color:"#10B981",fontWeight:600,textDecoration:"none"}}>What's new →</a>
**v0.6.7**first-class LangChain integration, SAP OData ingestor, human-editable Markdown persistence for `ContextGraph`, and a structured Action layer for the reasoning engine. <a href="https://github.com/semantica-agi/semantica/releases" style={{color:"#10B981",fontWeight:600,textDecoration:"none"}}>What's new →</a>
</Info>
This guide walks you through the end-to-end pipeline for building your first knowledge graph. Start here after installation. An LLM API key is optional: pattern-based extraction works out of the box.
@@ -35,7 +35,7 @@ Verify:
```bash
python -c "import semantica; print(semantica.__version__)"
# 0.5.0
# 0.6.7
```
@@ -327,10 +327,11 @@ print(f"Relationships active in 2023: {result_2023['num_relationships']}")
<Accordion title="Persistent graph store: Neo4j, FalkorDB, Apache AGE" icon="database">
```python
from semantica.graph_store import Neo4jStore
from semantica.graph_store import GraphStore
from semantica.kg import GraphBuilder
store = Neo4jStore(
store = GraphStore(
backend="neo4j",
uri="bolt://localhost:7687",
user="neo4j",
password="password",
+95
View File
@@ -25,6 +25,7 @@ icon: "brain"
| `DecisionRecorder` | Record decisions with embeddings, causal chains, and metadata |
| `PolicyEngine` | Policy management: `add_policy()`, `check_compliance()`, `get_applicable_policies()` |
| `CausalChainAnalyzer` | Trace how decisions influenced each other: `get_causal_chain(decision_id)` |
| `ErasureCoordinator` | Erase an entity across graph, memory, and vector store, returning an auditable `ErasureReceipt` |
## What You Get
@@ -634,6 +635,100 @@ queried together safely. Vector-store writes are deferred until the in-memory im
commits; adapter synchronization remains best-effort and logs failures.
## ErasureCoordinator
`ContextGraph.purge_node()` is scoped to one graph: the node is removed and a
tombstone is written, but the same content can still be live as an `AgentMemory`
item and as an embedding in the vector store. `ErasureCoordinator` drives the
cascade across every bound store and returns an `ErasureReceipt` recording what
each one reported.
```python
from semantica.context import AgentMemory, ContextGraph, ErasureCoordinator
coordinator = ErasureCoordinator(graph=graph, memory=memory)
receipt = coordinator.erase_entity(
"customer-4471",
reason="GDPR Art. 17 request #882",
)
if not receipt.complete:
# These stores may still hold the entity; handle them out of band.
print(receipt.incomplete_stores)
```
<Warning>
Check the receipt — the call returning is not proof the data is gone. FAISS,
Milvus, and Weaviate expose no delete method, so erasure cannot be completed on
those backends today; the receipt reports `unsupported` rather than a success it
did not achieve.
</Warning>
### Constructor Parameters
| Parameter | Type | Default | Description |
| :--- | :--- | :--- | :--- |
| `graph` | `ContextGraph` | `None` | Anything exposing `purge_node()` |
| `memory` | `AgentMemory` | `None` | Anything exposing `find_by_entity()` and `batch_delete()` |
| `vector_store` | `VectorStore` | `memory.vector_store` | Store holding entity-keyed embeddings; pass `False` to disable the leg |
At least one store is required; a store that is not supplied reports
`not_configured` rather than being silently skipped.
### Methods
| Method | Returns | Description |
| :--- | :--- | :--- |
| `erase_entity(entity_id, reason, at, vector_ids)` | `ErasureReceipt` | Erase one entity from every bound store |
| `erase_entities(entity_ids, reason, at)` | `List[ErasureReceipt]` | One receipt per entity, in order; one failure does not stop the rest |
### Store Statuses
| Status | Meaning |
| :--- | :--- |
| `erased` | Reached, data removed. On the vectors leg this means the store accepted the delete for the ids given — backends offer no portable existence check, so it is not a count of embeddings that were really there |
| `not_found` | Reached, held nothing for this entity |
| `not_configured` | No such store was bound — normal, not a failure |
| `unsupported` | The store cannot delete at all; retrying will not help |
| `failed` | The store was reached and the deletion did not succeed |
### ErasureReceipt
| Member | Type | Description |
| :--- | :--- | :--- |
| `entity_id` | `str` | Entity the erasure was requested for |
| `reason` | `Optional[str]` | Recorded in the receipt and the graph tombstone |
| `erased_at` | `str` | ISO-8601; matches the tombstone's `purged_at` |
| `stores` | `Dict[str, Dict]` | Per-store outcome keyed `vectors`, `memory`, `graph` |
| `complete` | `bool` | `False` when any store reports `unsupported` or `failed` |
| `incomplete_stores` | `List[str]` | Stores that may still hold the entity's data |
| `to_dict()` | `Dict` | Serialized receipt, safe to persist as an audit record |
```python
receipt.to_dict()
# {
# "entity_id": "customer-4471",
# "reason": "GDPR Art. 17 request #882",
# "erased_at": "2026-08-16T09:03:36.813220",
# "complete": False,
# "stores": {
# "vectors": {"status": "unsupported", "backend": "faiss",
# "detail": "backend exposes no delete()/delete_vectors(); ..."},
# "memory": {"status": "erased", "items": 14},
# "graph": {"status": "erased", "nodes": 1, "edges": 3},
# },
# }
```
Erasure runs outward-in — vectors, then memory, then the graph. The tombstone is
the durable attestation that an erasure happened, so it is written last: a crash
mid-cascade leaves the node present and the receipt incomplete, rather than a
tombstone claiming more than actually happened. A store that raises is recorded
as `failed` and the remaining stores are still erased. Erasing the same entity
twice returns a receipt saying there was nothing left to do rather than raising.
## PolicyEngine
`PolicyEngine` manages versioned policies stored in the knowledge graph. Policies are stored as nodes and can be linked to decisions:
+51 -3
View File
@@ -6,7 +6,7 @@ icon: "plug"
**`semantica.mcp_server`** exposes Semantica's knowledge graph, decision intelligence, semantic extraction, and reasoning capabilities as an [MCP (Model Context Protocol)](https://modelcontextprotocol.io) **server over stdio**:
- 12 MCP tools exposed: extract entities, query graph, record decisions, run reasoning, export results
- 15 MCP tools exposed: extract entities, query graph, record decisions, run reasoning, export results
- No Python code required after launch: configure once, use from any MCP-aware client
- Compatible with Claude Desktop, Windsurf, Cline, Continue, VS Code, Roo Code, Cursor
@@ -40,7 +40,7 @@ python -m semantica.mcp_server
## What You Get
- **12 MCP Tools** — Extract entities, extract relations, record decisions, query decisions, find precedents, trace causal chains, add entities, add relationships, run analytics, summarise graph, run reasoning, export graph.
- **15 MCP Tools** — Extract entities, extract relations, record decisions, query decisions, find precedents, trace causal chains, add entities, add relationships, run analytics, summarise graph, run reasoning, export graph, query the live graph, update nodes, archive nodes.
- **3 Readable Resources** — Live graph JSON (`semantica://graph/summary`), decision list, and schema/version info: readable by any MCP client.
- **Zero Infrastructure** — Runs over stdio: no server, no port, no Docker required. One config block to activate in any MCP client.
- **Persistent Graphs** — Point `SEMANTICA_KG_PATH` at a saved graph file to reload it automatically on every server startup.
@@ -159,7 +159,7 @@ The MCP server is included in the base install: no extras required.
## Tools
The MCP server exposes 12 tools that any connected AI assistant can call:
The MCP server exposes 15 tools that any connected AI assistant can call:
| Tool | Category | Description |
| :---- | :-------- | :----------- |
@@ -173,6 +173,9 @@ The MCP server exposes 12 tools that any connected AI assistant can call:
| `add_relationship` | Graph Operations | Add a directed edge between two nodes |
| `get_graph_summary` | Graph Operations | Node count, decision count, graph status |
| `get_graph_analytics` | Graph Operations | PageRank centrality and community detection |
| `query_graph` | Graph Operations | Fetch a node, traverse its neighbours, or keyword-search nodes |
| `update_node` | Graph Operations | Merge properties onto a node and persist to `SEMANTICA_KG_PATH` |
| `delete_node` | Graph Operations | Soft-delete (archive) a node and persist to `SEMANTICA_KG_PATH` |
| `run_reasoning` | Reasoning | Forward-chain IF/THEN rules over facts |
| `export_graph` | Reasoning & Export | Serialise the graph (`turtle`/`ttl`: RDF Turtle aliases, `nt`, `xml`, `json-ld`, `json`) |
@@ -386,6 +389,51 @@ Takes no input parameters.
</Accordion>
<Accordion title="query_graph" icon="magnifying-glass">
Read the live graph in one of three modes, set by `mode`:
- `node` — return a single node by `node_id`.
- `neighbors` (default) — traverse outward and inward from `node_id` up to `depth` hops (clamped to 1-5, default 1). Optional `relationship_types` filters edge types; optional `limit` caps results.
- `search` — keyword match `query` against each node's id and content. Optional `node_type` restricts the scan; `limit` defaults to 50.
**Input:**
```json
{ "mode": "neighbors", "node_id": "apple_inc", "depth": 2 }
```
</Accordion>
<Accordion title="update_node" icon="pen">
Merge a set of properties onto an existing node. The change is applied in memory and, when `SEMANTICA_KG_PATH` is set, written back to that file so it survives a restart. Returns `persisted: false` when no path is configured.
**Input:**
```json
{
"node_id": "task_42",
"properties": { "status": "done", "note": "shipped in v0.6.7" }
}
```
`node_id` and a non-empty `properties` object are required. Updating a missing node returns an error.
</Accordion>
<Accordion title="delete_node" icon="box-archive">
Soft-delete a node: it stays in the graph for history but is marked `status: "archived"`. Persists to `SEMANTICA_KG_PATH` when configured.
**Input:**
```json
{ "node_id": "task_42" }
```
</Accordion>
</AccordionGroup>
### Reasoning
@@ -0,0 +1,338 @@
# Objective Layer for semantica.evals Runner — Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Add per-metric objective support (direction + threshold, or Boolean expectation) to the `evaluate()` runner, overriding evaluator default pass verdicts, backward-compatible when no objective is configured.
**Architecture:** The runner already iterates evaluators and computes per-case status. Objectives are read from `config["<name>"]["objective"]`, validated up front, and applied to each returned metric's `passed` field (and `details`) before aggregation. Error metrics always win over objectives.
**Tech Stack:** Python 3.8+, stdlib only (typing, dataclasses). pytest for tests.
## Global Constraints
- Python >= 3.8: use `typing.Dict/List/Optional/Union`, never builtin generics or `|`.
- Zero new dependencies.
- Do not change the `EvalMetric` shape, the `evaluate()` signature, or the evaluator function signature.
- Existing behavior with no `objective` configured must be byte-for-byte unchanged (all 62 existing tests keep passing).
- Error metrics (`meta` contains `"error"`) always classify the case as `error`, regardless of objective.
- Config errors are programmer errors: raise `ValueError` from `evaluate()` before any evaluator runs (fail-fast).
- Tests go in `tests/evals/`, pytest class style, no new files outside the listed paths.
---
### Task 1: Objective parsing, validation, and re-decision in the runner
**Files:**
- Modify: `semantica/evals/runner.py`
- Test: `tests/evals/test_runner.py`
**Interfaces:**
- Consumes: `EvalMetric` from `.types` (fields: `score`, `passed`, `meta`); `evaluate(cases, evaluators, config=None, target_fn=None)` existing signature.
- Produces: private helpers `_parse_objective(name, eval_config) -> Optional[Dict]` (returns `None` when no objective configured, raises `ValueError` on invalid config) and `_apply_objective(metric, objective) -> bool` (returns the re-decided `passed`). Public `evaluate()` behavior extended as specified.
- [ ] **Step 1: Write the failing tests**
Append a new test class to `tests/evals/test_runner.py`:
```python
class TestObjective:
def test_maximize_with_threshold_pass(self):
# levenshtein similarity 1.0 for identical, objective demands >= 0.5
result = evaluate(
[("apple", "apple")],
evaluators=["levenshtein"],
config={"levenshtein": {"objective": {"direction": "maximize", "threshold": 0.5}}},
)
assert result.cases[0].status == "pass"
assert result.cases[0].metrics["levenshtein"].passed is True
def test_maximize_with_threshold_fail(self):
result = evaluate(
[("apple", "aple")], # similarity < 1.0
evaluators=["levenshtein"],
config={"levenshtein": {"objective": {"direction": "maximize", "threshold": 0.99}}},
)
assert result.cases[0].status == "fail"
assert result.cases[0].metrics["levenshtein"].passed is False
assert "levenshtein" in result.cases[0].details
def test_minimize_with_threshold_pass(self):
# edit distance normalized ~0.2; objective: distance <= 0.5
result = evaluate(
[("night", "nacht")],
evaluators=["levenshtein"],
config={"levenshtein": {"objective": {"direction": "minimize", "threshold": 0.5}}},
)
assert result.cases[0].status == "pass"
assert result.cases[0].metrics["levenshtein"].passed is True
def test_minimize_with_threshold_fail(self):
result = evaluate(
[("night", "nacht")],
evaluators=["levenshtein"],
config={"levenshtein": {"objective": {"direction": "minimize", "threshold": 0.1}}},
)
assert result.cases[0].status == "fail"
def test_expect_true_on_boolean_metric(self):
result = evaluate(
[("ok", "ok")],
evaluators=["exact_match"],
config={"exact_match": {"objective": {"expect": True}}},
)
assert result.cases[0].status == "pass"
def test_expect_false_overrides_passing_metric(self):
# exact_match passes (score 1.0) but expectation is false -> fail
result = evaluate(
[("ok", "ok")],
evaluators=["exact_match"],
config={"exact_match": {"objective": {"expect": False}}},
)
assert result.cases[0].status == "fail"
assert result.cases[0].metrics["exact_match"].passed is False
assert "exact_match" in result.cases[0].details
def test_maximize_without_threshold_is_noop(self):
# identical behavior to no objective: evaluator's own verdict stands
result = evaluate(
[("ok", "no")],
evaluators=["exact_match"],
config={"exact_match": {"objective": {"direction": "maximize"}}},
)
assert result.cases[0].status == "fail"
def test_minimize_without_threshold_raises(self):
with pytest.raises(ValueError):
evaluate(
[("a", "b")],
evaluators=["levenshtein"],
config={"levenshtein": {"objective": {"direction": "minimize"}}},
)
def test_bad_direction_raises(self):
with pytest.raises(ValueError):
evaluate(
[("a", "b")],
evaluators=["levenshtein"],
config={"levenshtein": {"objective": {"direction": "sideways", "threshold": 0.5}}},
)
def test_expect_with_direction_raises(self):
with pytest.raises(ValueError):
evaluate(
[("a", "b")],
evaluators=["levenshtein"],
config={"levenshtein": {"objective": {"expect": True, "direction": "maximize"}}},
)
def test_error_metric_wins_over_objective(self):
result = evaluate(
[("[invalid", "x")],
evaluators=["regex_match"],
config={"regex_match": {"objective": {"direction": "maximize", "threshold": 0.0}}},
)
assert result.cases[0].status == "error"
assert result.errors == 1
assert result.failed == 0
def test_no_objective_unchanged(self):
result = evaluate([("ok", "no")], evaluators=["exact_match"])
assert result.cases[0].status == "fail"
```
- [ ] **Step 2: Run tests to verify they fail**
Run: `python3 -m pytest tests/evals/test_runner.py -q`
Expected: the new `TestObjective` tests fail (objective config ignored → `exact_match` passes under `expect:false` etc.); the pre-existing tests in the file still pass.
- [ ] **Step 3: Implement objective parsing, validation, and re-decision**
In `semantica/evals/runner.py`, add two helpers before `evaluate` and wire them into the evaluator loop.
```python
def _parse_objective(name, eval_config):
"""Return the validated objective dict, or None when not configured.
Raises ValueError for invalid configurations (programmer error).
"""
objective = (eval_config or {}).get("objective")
if objective is None:
return None
direction = objective.get("direction")
threshold = objective.get("threshold")
expect = objective.get("expect")
if expect is not None:
if direction is not None or threshold is not None:
raise ValueError(
f"objective for '{name}': 'expect' cannot be combined with "
"'direction' or 'threshold'"
)
return {"expect": bool(expect)}
if direction == "minimize":
if threshold is None:
raise ValueError(
f"objective for '{name}': 'minimize' requires a 'threshold'"
)
return {"direction": "minimize", "threshold": float(threshold)}
if direction == "maximize":
if threshold is None:
# no bar to re-decide against; treat as absent (evaluator default stands)
return None
return {"direction": "maximize", "threshold": float(threshold)}
raise ValueError(
f"objective for '{name}': 'direction' must be 'maximize' or 'minimize' "
f"(got {direction!r})"
)
def _apply_objective(metric, objective):
"""Return the objective-adjusted pass verdict for a non-error metric."""
if "expect" in objective:
return bool(metric.score) == objective["expect"]
if objective["direction"] == "minimize":
return metric.score <= objective["threshold"]
return metric.score >= objective["threshold"]
```
Then modify the evaluator loop in `evaluate()` so the parsed objective is computed once per case (outside the evaluator loop, since it only depends on merged config), and applied inside the loop:
```python
objective_by_name = {
name: _parse_objective(name, merged.get(name) or {})
for name in evaluators
}
metrics: Dict[str, EvalMetric] = {}
details: Dict[str, Any] = {}
failed, errored = False, False
for name in evaluators:
eval_config = merged.get(name) or {}
try:
metric = get_evaluator(name)(actual, expected, config=eval_config)
objective = objective_by_name.get(name)
if objective is not None and "error" not in metric.meta:
metric = EvalMetric(metric.score, _apply_objective(metric, objective), metric.meta)
metrics[name] = metric
if "error" in metric.meta:
errored = True
details[name] = metric.meta
elif not metric.passed:
failed = True
details[name] = metric.meta
except Exception as exc: # noqa: BLE001
errored = True
metrics[name] = EvalMetric(0.0, False, {"error": str(exc)})
details[name] = {"error": str(exc)}
```
Note: `objective_by_name` is computed once per case (it depends only on merged config), so invalid config raises `ValueError` at the first case — satisfying the fail-fast requirement. `EvalMetric` is a frozen dataclass, so the re-verdict constructs a new instance preserving score/meta.
- [ ] **Step 4: Run tests to verify they pass**
Run: `python3 -m pytest tests/evals/test_runner.py -q`
Expected: all `TestObjective` tests pass; pre-existing tests still pass.
- [ ] **Step 5: Run the full evals suite**
Run: `python3 -m pytest tests/evals -q`
Expected: 62 existing + new tests all pass (no regressions).
- [ ] **Step 6: Commit**
```bash
git add semantica/evals/runner.py tests/evals/test_runner.py
git commit -m "feat(evals): add per-metric objective support to runner"
```
---
### Task 2: Documentation — usage.md and CHANGELOG
**Files:**
- Modify: `semantica/evals/usage.md`
- Modify: `CHANGELOG.md`
**Interfaces:**
- Consumes: the objective config surface implemented in Task 1 (exact keys: `objective.direction`, `objective.threshold`, `objective.expect`; validation rules).
- Produces: docs only.
- [ ] **Step 1: Add objective section to usage.md**
Append a section after the existing "Run the runner over decision records" section:
```markdown
## Set per-evaluator objectives
By default each evaluator decides its own pass/fail. To override that
verdict at the run level, configure an **objective** per evaluator name:
```python
from semantica.evals import evaluate
# Require a minimum similarity (default direction is maximize):
evaluate(
[("apple", "aple")],
evaluators=["levenshtein"],
config={"levenshtein": {"objective": {"direction": "maximize", "threshold": 0.7}}},
)
# Lower is better — override the direction:
evaluate(
[("night", "nacht")],
evaluators=["levenshtein"],
config={"levenshtein": {"objective": {"direction": "minimize", "threshold": 0.5}}},
)
# Boolean expectation on a 0/1 metric:
evaluate(
[("ok", "ok")],
evaluators=["exact_match"],
config={"exact_match": {"objective": {"expect": False}}},
)
```
Rules:
- `maximize` + `threshold`: pass iff `score >= threshold`. `maximize` without
a threshold is a no-op (the evaluator's own verdict stands).
- `minimize` + `threshold`: pass iff `score <= threshold`. `minimize`
**requires** a threshold — omitting it raises `ValueError`.
- `expect` (`true`/`false`): pass iff `bool(score)` matches; cannot be
combined with `direction`/`threshold`.
- A metric whose `meta` contains `"error"` is always an error, never affected
by an objective.
- Invalid objective config raises `ValueError` before any evaluator runs.
```
- [ ] **Step 2: Add CHANGELOG entry**
Under `## [Unreleased]` → `### Added`, insert a new bullet at the top (before the `semantica.evals` module entry), following existing style:
```markdown
- **`semantica.evals` runner gains per-metric objectives** (#1091)
- `evaluate()` now accepts `config={"<evaluator>": {"objective": {"direction": "maximize"|"minimize", "threshold": X}}}` to override the evaluator's default pass verdict with a threshold; `{"objective": {"expect": bool}}` expresses a Boolean expectation
- `minimize` requires a `threshold`; `maximize` without one is a no-op; `expect` cannot be combined with `direction`/`threshold`; invalid config raises `ValueError` before any evaluator runs
- Error metrics are never affected by objectives (error wins over fail)
- Backward compatible: no `objective` key → existing behavior unchanged
- New tests in `tests/evals/test_runner.py::TestObjective`
```
- [ ] **Step 3: Verify docs examples run**
Run the three examples from Step 1 as a Python script (import `evaluate`, run each snippet) to confirm they don't raise unexpectedly. No test output assertion needed beyond "no exception" and sensible status values.
- [ ] **Step 4: Commit**
```bash
git add semantica/evals/usage.md CHANGELOG.md
git commit -m "docs(evals): document per-metric objectives"
```
---
## Self-Review Notes
- **Spec coverage:** §3.1 (config surface) → Task 1 helpers + Task 2 docs; §3.2 (semantics: maximize/minimize/expect) → Task 1 `_apply_objective`; §3.3 (error wins) → Task 1 error branch + `test_error_metric_wins_over_objective`; §3.4 rules 1-3 (validation) → Task 1 `_parse_objective` + 4 validation tests; §3.4 rule 4 → error branch; §3.5 (aggregation unchanged, details on final verdict) → Task 1 loop + `test_expect_false_overrides_passing_metric` asserts `details`; §4 (fail-fast ValueError) → `_parse_objective` at case top; §5 (tests) → Task 1 test class; §6 (compat) → `test_no_objective_unchanged` + full-suite green.
- **Type consistency:** `_parse_objective(name, eval_config) -> Optional[Dict]`, `_apply_objective(metric, objective) -> bool`; `EvalMetric(score, passed, meta)` positional construction preserved everywhere.
- **Backward compat:** objective parsed to `None` for absent config → loop behavior identical to before.
@@ -0,0 +1,115 @@
# Design: Objective layer for `semantica.evals` runner
**Date:** 2026-08-19
**Issue:** semantica-agi/semantica#1091 (assigned to pkupt)
**Base:** PR #1090 (`semantica.evals` module)
## 1. Problem
`semantica.evals` runs named evaluators and aggregates per-case pass/fail, but the pass judgement is hard-coded inside each evaluator — a higher score always means "better". There is no way to express an evaluation objective at the run level:
- apply a threshold the evaluator does not encode (e.g. "F1 must be ≥ 0.7");
- reverse the direction (e.g. "lower edit distance is better");
- express a Boolean expectation (e.g. "this metric should be `false`").
This blocks the domain-specific benchmark harnesses `docs/community-projects.md` says `semantica.evals` supports. Palantir AIP Evals models exactly this: each metric has an **objective** (Boolean expected value, or numeric `maximize`/`minimize` direction with an optional threshold), and a test case passes when **all** its metrics meet their objectives.
## 2. Scope
In scope:
- A per-metric objective configuration consumed by the `evaluate()` runner.
- Runner-level pass/fail re-decision for numeric scores and Boolean metrics.
- Backward-compatible behavior when no objective is configured.
- Tests and docs.
Out of scope:
- Changing the evaluator signature or the `EvalMetric` shape.
- Multi-iteration test cases (AIP Evals has them; Semantica's runner is single-iteration per case).
- Objective-aware aggregation beyond per-case `pass`/`fail` (existing `pass_rate` semantics are kept).
## 3. Design
### 3.1 Configuration surface
Objective is configured per evaluator inside the runner's `config`, under the evaluator name:
```python
config = {
"<evaluator_name>": {
"objective": {
"direction": "maximize" | "minimize",
"threshold": <float>, # optional
}
}
}
```
Boolean-form objective (shorthand): for metrics whose score is Boolean-like (0.0/1.0) or for semantic clarity, `{"objective": {"expect": true}}` / `{"objective": {"expect": false}}` is also supported.
### 3.2 Evaluation semantics
For each metric produced by an evaluator during a case run, if an objective exists for that evaluator name, the runner recomputes the metric's pass verdict:
- **maximize**: pass iff `score >= threshold`. If no `threshold` is given, the objective is treated as absent (evaluator's own verdict stands) — see 3.4 rule 2.
- **minimize**: pass iff `score <= threshold` (threshold required, see 3.4 rule 1).
- **expect**: pass iff `bool(score)` equals `expect` (for Boolean-style metrics).
When an objective is present, the runner **overrides** `metric.passed` with the objective verdict. When absent, `metric.passed` is used unchanged (existing behavior).
The `objective` key is a **reserved runner-level key**: it is consumed by the runner and is passed through to the evaluator function inside `eval_config` (evaluators already ignore unknown config keys via `cfg.get(...)`, so this is harmless); evaluators must not rely on it. The runner re-decision happens on the metric the evaluator returns, so no evaluator change is required.
### 3.3 Interaction with errors
An `EvalMetric` whose `meta` contains `"error"` remains classified as an error regardless of objective (error wins over fail, per the existing contract). Objectives only affect non-error metrics.
### 3.4 Ambiguity rules (explicit decisions)
1. **`minimize` without `threshold`** is rejected at config-validation time with a clear error (`ValueError`), because "lowest is best" has no absolute pass bar without a threshold. (AIP Evals allows direction-only; we require threshold to keep pass/fail well-defined.) — *Chosen for determinism; revisit if a use case demands direction-only minimize.*
2. **`maximize` without `threshold`** behaves like no objective (pass iff evaluator's own `passed`), because the evaluator's default is already "higher is better".
3. **`expect` with a numeric `direction`/`threshold`** is a config error (`ValueError`): pick one form.
4. **Objective on a metric that errors** → the error wins (3.3), objective ignored.
### 3.5 Aggregation
Unchanged:
- Case `status`: `"error"` if any metric errored, else `"fail"` if any failed, else `"pass"`.
- `pass_rate` = passed / total (1.0 on empty).
- `metrics` dict holds the (possibly re-verdict'd) `EvalMetric`; the re-verdict is observable via `metric.passed`.
- `details[name]` is populated when a metric ends up failed **after** objective re-decision (i.e. objective-failed metrics appear in `details`; metrics that pass under objective are not recorded there). This mirrors the existing "record failures in details" behavior applied to the final verdict.
### 3.6 Files
- `semantica/evals/runner.py` — add objective parsing/validation and re-decision inside the evaluator loop.
- `tests/evals/test_runner.py` — new test class(es) for objective semantics.
- `semantica/evals/usage.md` — document the objective config and examples.
- `CHANGELOG.md``[Unreleased]` entry.
No new dependencies; Python ≥ 3.8 (stdlib `typing`).
## 4. Error handling
- Invalid objective config (`direction` not in {maximize, minimize}, both `expect` and `direction`, `minimize` without threshold, non-numeric threshold) → `ValueError` raised at runner config parse, before any evaluator runs. Deterministic, fail-fast.
- These are programmer errors, not per-case data errors — no per-case `error` status involved.
## 5. Testing
New tests in `tests/evals/test_runner.py`:
1. maximize + threshold: score ≥ threshold → pass; below → fail.
2. minimize + threshold: score ≤ threshold → pass; above → fail (e.g. levenshtein on a close pair).
3. minimize without threshold → `ValueError`.
4. expect=true / expect=false on a Boolean metric (exact_match) — pass/fail per expectation.
5. no objective → existing behavior unchanged (evaluator's own verdict).
6. objective + error metric → error wins (status=error, not fail).
7. config error (bad direction) → `ValueError` raised by `evaluate()`.
8. objective turns a passing metric into failing → `details` records it; case status becomes fail.
9. backward-compat: all existing 62 tests keep passing.
## 6. Compatibility
- Public API (`evaluate`, `list_evaluators`, `get_evaluator`, types) unchanged in signature.
- `EvalMetric` shape unchanged (score, passed, meta) — only `passed` may be recomputed by the runner.
- Existing configs (no `objective` key) behave identically.
+36
View File
@@ -0,0 +1,36 @@
# CI templates
Copy-paste starting points for wiring `semantica` into your own project's CI. Each file is a
complete, working config — rename it into your project (see the comment at the top of each file
for the target path) and swap the smoke-test / test step for whatever your project does with
Semantica. Each template installs `semantica` unconditionally and your own project's dependencies
only if a `requirements.txt` is present; if your project uses `pyproject.toml`, Poetry, or Pipenv
instead, adjust the marked install line (each file calls it out inline).
| File | Target path in your repo |
| ---- | ------------------------- |
| [`github-actions.yml`](github-actions.yml) | `.github/workflows/semantica.yml` |
| [`gitlab-ci.yml`](gitlab-ci.yml) | `.gitlab-ci.yml` |
| [`circleci-config.yml`](circleci-config.yml) | `.circleci/config.yml` |
If your own project is hosted on GitHub, you can skip the setup boilerplate entirely and use
Semantica's reusable composite action instead:
```yaml
- uses: semantica-agi/semantica/.github/actions/setup-semantica@main
with:
python-version: '3.11'
# extras: 'explorer,all' # optional
# version: '==0.6.7' # optional, pin an exact release
# cache: 'pip' # optional, only if your repo has a requirements.txt/pyproject.toml/etc.
```
`@main` always tracks this repo's default branch, which is convenient but — like any mutable
ref — can change out from under you between runs. For production CI, pin it to a commit SHA
instead (find one via `git rev-parse` against a tagged release, or the commit history for
[`.github/actions/setup-semantica/`](../../.github/actions/setup-semantica/)) and update the pin
deliberately when you want to pick up changes, the same way this repo's own workflows are pinned
(see [`verify-action-pins.yml`](../../.github/workflows/verify-action-pins.yml)).
It installs Python, installs `semantica`, and verifies the import (pip caching is opt-in via `cache: 'pip'`, since not every caller repo has a requirements file to key the cache on) — see
[`.github/actions/setup-semantica/action.yml`](../../.github/actions/setup-semantica/action.yml).
+40
View File
@@ -0,0 +1,40 @@
# Drop this in as .circleci/config.yml in your own project.
version: 2.1
jobs:
test:
docker:
- image: cimg/python:3.11
steps:
- checkout
# A content-hashed cache key (e.g. `{{ checksum "requirements.txt" }}`)
# is more precise but breaks if that exact file doesn't exist in your
# project - swap in one matched to however you declare dependencies
# once you've adjusted the install step below.
- restore_cache:
keys:
- pip-cache-v1
- run:
name: Install dependencies
command: |
pip install --upgrade pip
pip install semantica
# Install your own project's dependencies however your project
# declares them - adjust this to match, e.g. `pip install -e .`
# for pyproject.toml / setup.cfg, or `poetry install`.
if [ -f requirements.txt ]; then pip install -r requirements.txt; fi
- save_cache:
key: pip-cache-v1
paths:
- ~/.cache/pip
- run:
name: Smoke test
command: python -c "import semantica; print('semantica', semantica.__version__)"
- run:
name: Run tests
command: pytest
workflows:
test:
jobs:
- test
+44
View File
@@ -0,0 +1,44 @@
# Drop this in as .github/workflows/semantica.yml in your own project.
#
# Installs Semantica and runs a smoke import + your test suite. Swap the
# smoke-test step for whatever your project actually does with Semantica
# (build a context graph, run an ingest pipeline, etc.).
#
# Third-party actions below are pinned to a commit SHA rather than a mutable
# tag - a moved tag can silently swap in different code. Update the pin (and
# the trailing "# vX" comment) deliberately when you want a newer version;
# see semantica-agi/semantica's own .github/workflows/verify-action-pins.yml
# for one way to keep pins honest automatically.
name: Semantica
on:
push:
branches: [main]
pull_request:
branches: [main]
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7
with:
python-version: '3.11'
cache: 'pip'
- name: Install dependencies
run: |
python -m pip install --upgrade pip
pip install semantica
# Install your own project's dependencies however your project
# declares them - adjust this to match. Examples:
# pip install -r requirements.txt
# pip install -e . # pyproject.toml / setup.cfg
# pip install -e ".[dev]"
# poetry install
if [ -f requirements.txt ]; then pip install -r requirements.txt; fi
- name: Run tests
run: pytest
+20
View File
@@ -0,0 +1,20 @@
# Drop this in as .gitlab-ci.yml in your own project.
semantica-test:
image: python:3.11-slim
cache:
paths:
- .cache/pip
variables:
PIP_CACHE_DIR: "$CI_PROJECT_DIR/.cache/pip"
script:
- pip install --upgrade pip
- pip install semantica
# Install your own project's dependencies however your project declares
# them - adjust this to match, e.g. `pip install -e .` for pyproject.toml
# / setup.cfg, or `poetry install`.
- if [ -f requirements.txt ]; then pip install -r requirements.txt; fi
- python -c "import semantica; print('semantica', semantica.__version__)"
- pytest
rules:
- if: '$CI_PIPELINE_SOURCE == "merge_request_event"'
- if: '$CI_COMMIT_BRANCH == "main"'
@@ -0,0 +1,195 @@
"""
Deterministic Explorer Rendering E2E Example.
Demonstrates building, serializing, and reloading a deterministic 4-node,
3-edge knowledge graph baseline for visual inspection in Semantica Explorer (#1037).
Graph topology:
Alice (Person, #63E6FF) --WORKS_AT--> Acme (Organization, #A78BFA)
Bob (Person, #63E6FF) --KNOWS--> Alice (Person, #63E6FF)
Acme (Organization, #A78BFA) --LOCATED_IN--> New York (Location, #34D399)
Clean Checkout Prerequisites:
1. Python backend dependencies:
pip install -e ".[explorer]"
2. Frontend workspace dependencies:
cd explorer && npm install && cd ..
Usage:
# 1. Generate the deterministic graph baseline:
python examples/explorer_deterministic_rendering_example.py
# 2. Launch Explorer with local dev authentication (Option A - Dev mode):
# Terminal 1 (Backend API):
SEMANTICA_ALLOW_ANONYMOUS=true python -m semantica.explorer --graph explorer_e2e_test_graph.json --port 8000 --no-browser
# Terminal 2 (Frontend UI):
cd explorer && npm run dev
# Open http://localhost:5173
# 2. Launch Explorer (Option B - Standalone CLI server):
SEMANTICA_ALLOW_ANONYMOUS=true python -m semantica.explorer --graph explorer_e2e_test_graph.json --port 8000
# Open http://localhost:8000
# Secure authentication alternative:
export SEMANTICA_API_KEY="your-secret-api-key"
python -m semantica.explorer --graph explorer_e2e_test_graph.json --port 8000
# Send HTTP header: X-API-Key: your-secret-api-key
Verification Checklist:
- Exactly 4 nodes visible on canvas:
* Alice (Person, #63E6FF)
* Bob (Person, #63E6FF)
* Acme (Organization, #A78BFA)
* New York (Location, #34D399)
- Exactly 3 directed edges with canonical relationship labels:
* Alice -> Acme (WORKS_AT)
* Bob -> Alice (KNOWS)
* Acme -> New York (LOCATED_IN)
- Zoom behavior:
* Zoom in to Inspection tier (ratio <= 0.5): directional arrows and node labels scale clearly.
* Zoom out to Overview tier (ratio > 1.2): layout remains stable and non-colliding.
- Hover & Selection interactions:
* Hover over 'Alice': node halo triggers; incident edges (WORKS_AT, KNOWS) highlight in local context.
* Click an edge: Inspector panel confirms edgeType ('WORKS_AT', 'KNOWS', or 'LOCATED_IN').
"""
from __future__ import annotations
import json
from pathlib import Path
from semantica.context.context_graph import ContextGraph
from semantica.explorer.session import GraphSession
def build_deterministic_graph() -> ContextGraph:
"""Build the exact 4-node, 3-edge graph specified in #1037."""
graph = ContextGraph(advanced_analytics=False)
# 1. Add exactly 4 nodes
graph.add_node(
"alice",
node_type="Person",
content="Alice",
color="#63E6FF",
)
graph.add_node(
"bob",
node_type="Person",
content="Bob",
color="#63E6FF",
)
graph.add_node(
"acme",
node_type="Organization",
content="Acme",
color="#A78BFA",
)
graph.add_node(
"new_york",
node_type="Location",
content="New York",
color="#34D399",
)
# 2. Add exactly 3 directed edges
graph.add_edge("alice", "acme", edge_type="WORKS_AT", weight=1.0)
graph.add_edge("bob", "alice", edge_type="KNOWS", weight=1.0)
graph.add_edge("acme", "new_york", edge_type="LOCATED_IN", weight=1.0)
return graph
def main() -> None:
print("=" * 75)
print("Semantica Explorer Deterministic Graph Generator (#1037)")
print("=" * 75)
print("1. Building deterministic ContextGraph...")
graph = build_deterministic_graph()
print(
f" ✓ Graph built with {len(graph.nodes)} nodes "
f"and {len(graph.edges)} edges."
)
output_path = Path("explorer_e2e_test_graph.json").resolve()
print(f"2. Persisting graph to '{output_path.name}'...")
graph.save_to_file(str(output_path))
print(f" ✓ Graph saved to {output_path}")
# Verify JSON format
with open(output_path, "r", encoding="utf-8") as f:
data = json.load(f)
assert len(data.get("nodes", [])) == 4
assert len(data.get("edges", [])) == 3
print("3. Verifying reload via GraphSession.from_file()...")
session = GraphSession.from_file(str(output_path))
stats = session.get_stats()
nodes, total_nodes = session.get_nodes()
edges, total_edges = session.get_edges()
assert stats["node_count"] == 4
assert stats["edge_count"] == 3
assert total_nodes == 4
assert total_edges == 3
print(
f" ✓ Graph reloaded successfully without mutation "
f"(nodes: {total_nodes}, edges: {total_edges}).\n"
)
print("=" * 75)
print("Clean Checkout Prerequisites:")
print("=" * 75)
print(" pip install -e '.[explorer]'")
print(" cd explorer && npm install && cd ..\n")
print("=" * 75)
print("Reproduction instructions to view in Semantica Explorer:")
print("=" * 75)
print("Option A (Frontend dev server + API backend — recommended for development):")
print(
f" 1. Backend: SEMANTICA_ALLOW_ANONYMOUS=true python -m semantica.explorer "
f"--graph {output_path} --port 8000 --no-browser"
)
print(" 2. Frontend: cd explorer && npm run dev")
print(" 3. Open http://localhost:5173 to inspect the graph canvas.\n")
print("Option B (Standalone Explorer CLI server):")
print(
f" SEMANTICA_ALLOW_ANONYMOUS=true python -m semantica.explorer "
f"--graph {output_path} --port 8000"
)
print(" Open http://localhost:8000\n")
print("Secure Authentication Alternative:")
print(" export SEMANTICA_API_KEY='your-secret-api-key'")
print(
f" python -m semantica.explorer --graph {output_path} --port 8000"
)
print(" Send header: 'X-API-Key: your-secret-api-key'\n")
print("=" * 75)
print("Verification Checklist:")
print("=" * 75)
print(" 1. Nodes (4 total):")
print(" - Alice (Person, #63E6FF)")
print(" - Bob (Person, #63E6FF)")
print(" - Acme (Organization, #A78BFA)")
print(" - New York (Location, #34D399)")
print(" 2. Directed Edges & Canonical Labels (3 total):")
print(" - Alice -> Acme [WORKS_AT]")
print(" - Bob -> Alice [KNOWS]")
print(" - Acme -> New York [LOCATED_IN]")
print(" 3. Zoom Interactions:")
print(" - Inspection tier (zoom in): directional arrows & labels remain legible.")
print(" - Overview tier (zoom out): nodes and edges maintain layout integrity.")
print(" 4. Hover & Selection Interactions:")
print(" - Hover Alice: node halo triggers and incident edges (WORKS_AT, KNOWS) highlight.")
print(" - Click edge: Inspector panel displays edgeType label ('WORKS_AT', 'KNOWS', 'LOCATED_IN').")
print("=" * 75)
if __name__ == "__main__":
main()
+43 -51
View File
@@ -80,7 +80,6 @@
"integrity": "sha512-QdxmAo/ikZqqRGA8s43ww8lcql6naWRvEz0FFrl6MIlc7Gi6TroXnSdWa5U/kq6fzcpqpHesicQxFZIieZbyIA==",
"dev": true,
"license": "MIT",
"peer": true,
"dependencies": {
"@babel/code-frame": "^7.29.0",
"@babel/generator": "^7.29.6",
@@ -1603,7 +1602,8 @@
"version": "2.0.46",
"resolved": "https://registry.npmjs.org/@types/hammerjs/-/hammerjs-2.0.46.tgz",
"integrity": "sha512-ynRvcq6wvqexJ9brDMS4BnBLzmr0e14d6ZJTEShTBWKymQiHwlAyGu0ZPEFI2Fh1U53F7tN9ufClWM5KvqkKOw==",
"license": "MIT"
"license": "MIT",
"peer": true
},
"node_modules/@types/hast": {
"version": "3.0.5",
@@ -1642,7 +1642,6 @@
"integrity": "sha512-A1sre26ke7HDIuY/M23nd9gfB+nrmhtYyMINbjI1zHJxYteKR6qSMX56FsmjMcDb3SMcjJg5BiRRgOCC/yBD0g==",
"devOptional": true,
"license": "MIT",
"peer": true,
"dependencies": {
"undici-types": "~7.16.0"
}
@@ -1652,7 +1651,6 @@
"resolved": "https://registry.npmjs.org/@types/react/-/react-19.2.14.tgz",
"integrity": "sha512-ilcTH/UniCkMdtexkoCN0bI7pMcJDvmQFPvuPvmEaYA/NSfFTAgdUSLAoVjaRJm7+6PvcM+q1zYOwS4wTYMF9w==",
"license": "MIT",
"peer": true,
"dependencies": {
"csstype": "^3.2.2"
}
@@ -1672,7 +1670,8 @@
"resolved": "https://registry.npmjs.org/@types/trusted-types/-/trusted-types-2.0.7.tgz",
"integrity": "sha512-ScaPdn1dQczgbl0QFTeTOmVHFULt394XJgOQNoyVhZ6r2vLnMLJfBPd53SB52T/3G36VI1/g2MZaX0cwDuXsfw==",
"license": "MIT",
"optional": true
"optional": true,
"peer": true
},
"node_modules/@types/unist": {
"version": "3.0.3",
@@ -1725,7 +1724,6 @@
"integrity": "sha512-/Zb/xaIDfxeJnvishjGdcR4jmr7S+bda8PKNhRGdljDM+elXhlvN0FyPSsMnLmJUrVG9aPO6dof80wjMawsASg==",
"dev": true,
"license": "MIT",
"peer": true,
"dependencies": {
"@typescript-eslint/scope-manager": "8.58.2",
"@typescript-eslint/types": "8.58.2",
@@ -1995,7 +1993,6 @@
"integrity": "sha512-xRQbDb9BnwDafYNn6Vwl839DYVjqXYb1XVGtWAZ1kcDc6iwAL4hg3B1dZlRiuENFeO2H53gFG3in621AdERVAg==",
"dev": true,
"license": "MIT",
"peer": true,
"bin": {
"acorn": "bin/acorn"
},
@@ -2070,9 +2067,9 @@
}
},
"node_modules/baseline-browser-mapping": {
"version": "2.10.20",
"resolved": "https://registry.npmjs.org/baseline-browser-mapping/-/baseline-browser-mapping-2.10.20.tgz",
"integrity": "sha512-1AaXxEPfXT+GvTBJFuy4yXVHWJBXa4OdbIebGN/wX5DlsIkU0+wzGnd2lOzokSk51d5LUmqjgBLRLlypLUqInQ==",
"version": "2.11.20",
"resolved": "https://registry.npmjs.org/baseline-browser-mapping/-/baseline-browser-mapping-2.11.20.tgz",
"integrity": "sha512-H0ulySigv6icDJ1F7SjtdCD6PrhTpdYCmP0CactWy1+ekh0AFd0o1Wn5T8b+hnTmdBx19u9yhL6wvCylXMY7zw==",
"dev": true,
"license": "Apache-2.0",
"bin": {
@@ -2083,9 +2080,9 @@
}
},
"node_modules/brace-expansion": {
"version": "5.0.8",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.8.tgz",
"integrity": "sha512-JZyDyq3D4AUifKTPOB7DELf6XsB3WdPuNxCtob1vFXPsSXhdAiHBWJ/tJ8HAc9aH84BK+5JFZLNkJKx3G9kzQg==",
"version": "5.0.9",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.9.tgz",
"integrity": "sha512-ScQ4IuvIEF1TMlP7Zt+vjJ//9zlPb2SDcxWxM3bk8s6t6GGdJ7KO1dCcTidOPJKePW30LE/2cT7wCyPho9/Wxg==",
"dev": true,
"license": "MIT",
"dependencies": {
@@ -2096,9 +2093,9 @@
}
},
"node_modules/browserslist": {
"version": "4.28.2",
"resolved": "https://registry.npmjs.org/browserslist/-/browserslist-4.28.2.tgz",
"integrity": "sha512-48xSriZYYg+8qXna9kwqjIVzuQxi+KYWp2+5nCYnYKPTr0LvD89Jqk2Or5ogxz0NUMfIjhh2lIUX/LyX9B4oIg==",
"version": "4.28.8",
"resolved": "https://registry.npmjs.org/browserslist/-/browserslist-4.28.8.tgz",
"integrity": "sha512-V2NpofLblG64mfOtSgDhOJESZEGogzDMBv/q+W6oc4LXWP/q75eOXoOaaOu1EOadB9U4Bwx/e0yzbvwKH8zalA==",
"dev": true,
"funding": [
{
@@ -2115,13 +2112,12 @@
}
],
"license": "MIT",
"peer": true,
"dependencies": {
"baseline-browser-mapping": "^2.10.12",
"caniuse-lite": "^1.0.30001782",
"electron-to-chromium": "^1.5.328",
"node-releases": "^2.0.36",
"update-browserslist-db": "^1.2.3"
"baseline-browser-mapping": "^2.11.12",
"caniuse-lite": "^1.0.30001809",
"electron-to-chromium": "^1.5.402",
"node-releases": "^2.0.53",
"update-browserslist-db": "^1.3.0"
},
"bin": {
"browserslist": "cli.js"
@@ -2131,9 +2127,9 @@
}
},
"node_modules/caniuse-lite": {
"version": "1.0.30001788",
"resolved": "https://registry.npmjs.org/caniuse-lite/-/caniuse-lite-1.0.30001788.tgz",
"integrity": "sha512-6q8HFp+lOQtcf7wBK+uEenxymVWkGKkjFpCvw5W25cmMwEDU45p1xQFBQv8JDlMMry7eNxyBaR+qxgmTUZkIRQ==",
"version": "1.0.30001810",
"resolved": "https://registry.npmjs.org/caniuse-lite/-/caniuse-lite-1.0.30001810.tgz",
"integrity": "sha512-TITQPUkaz+aVk5GL6NhOdwk1aEaNTSDPsGFWrTuhKGtjTF70jL/Oht2W4c6rXUe5fu7Ie19VIahAXHIIiWWNeg==",
"dev": true,
"funding": [
{
@@ -2221,7 +2217,8 @@
"version": "2.20.3",
"resolved": "https://registry.npmjs.org/commander/-/commander-2.20.3.tgz",
"integrity": "sha512-GpVkmM8vF2vQUkj2LvZmD35JxeJOLCwJ9cUkugyk2nuhbv3+mJvpLYYt+0+USMxE+oj+ey/lJEnhZw75x/OMcQ==",
"license": "MIT"
"license": "MIT",
"peer": true
},
"node_modules/component-emitter": {
"version": "1.3.1",
@@ -2259,7 +2256,8 @@
"version": "0.0.10",
"resolved": "https://registry.npmjs.org/cssfilter/-/cssfilter-0.0.10.tgz",
"integrity": "sha512-FAaLDaplstoRsDR8XGYH51znUN0UY7nMc6Z9/fvE8EXGwvJE9hu7W2vHwx1+bd6gCYnln9nLbzxFTrcO9YQDZw==",
"license": "MIT"
"license": "MIT",
"peer": true
},
"node_modules/csstype": {
"version": "3.2.3",
@@ -2324,7 +2322,6 @@
"resolved": "https://registry.npmjs.org/d3-selection/-/d3-selection-3.0.0.tgz",
"integrity": "sha512-fmTRWbNMmsmWq6xJV8D19U/gw/bwrHfNXxrIN+HfZgnzqTHp9jOmKMhsTUjXOJnZOdZY9Q28y4yebKzqDKlxlQ==",
"license": "ISC",
"peer": true,
"engines": {
"node": ">=12"
}
@@ -2457,14 +2454,15 @@
"resolved": "https://registry.npmjs.org/dompurify/-/dompurify-3.4.13.tgz",
"integrity": "sha512-2vmYIoqjze2d+kakP8S/nS5shfsl587kzwEjcGlTdiksUVgFHnFCsLYDVj/JNqJVOQZGSYBTmuycv0PodwmnMQ==",
"license": "(MPL-2.0 OR Apache-2.0)",
"peer": true,
"optionalDependencies": {
"@types/trusted-types": "^2.0.7"
}
},
"node_modules/electron-to-chromium": {
"version": "1.5.340",
"resolved": "https://registry.npmjs.org/electron-to-chromium/-/electron-to-chromium-1.5.340.tgz",
"integrity": "sha512-908qahOGocRMinT2nM3ajCEM99H4iPdv84eagPP3FfZy/1ZGeOy2CZYzjhms81ckOPCXPlW7LkY4XpxD8r1DrA==",
"version": "1.5.420",
"resolved": "https://registry.npmjs.org/electron-to-chromium/-/electron-to-chromium-1.5.420.tgz",
"integrity": "sha512-2yD6XreGusOfNV+dUcvipJEXc3n/n7fgr7996aszTG+YY5E4mqM4tOq/3uhP129cazL9YHbVWSpc79ePotWtPA==",
"dev": true,
"license": "ISC"
},
@@ -2539,7 +2537,6 @@
"integrity": "sha512-nuKKvN+oIBO0koN7Tm7dlkmnkc21mtt0QJLwAKzjLq14y6lRTdVG36MZHJ8eQHwdJMwZbQNMlPOYedMq/oVJvQ==",
"dev": true,
"license": "MIT",
"peer": true,
"workspaces": [
"packages/*"
],
@@ -3339,6 +3336,7 @@
"resolved": "https://registry.npmjs.org/marked/-/marked-14.0.0.tgz",
"integrity": "sha512-uIj4+faQ+MgHgwUW1l2PsPglZLOLOT1uErt06dAPtx2kjteLAkbsd/0FiYg/MGS+i7ZKLb7w2WClxHkzOOuryQ==",
"license": "MIT",
"peer": true,
"bin": {
"marked": "bin/marked.js"
},
@@ -4250,9 +4248,9 @@
"license": "MIT"
},
"node_modules/nanoid": {
"version": "3.3.16",
"resolved": "https://registry.npmjs.org/nanoid/-/nanoid-3.3.16.tgz",
"integrity": "sha512-bzlKTyNJ7+LdGIIwy8ijFpIqEQIvafahV7eYykJ8Cvh42EdJeODoJ6gUJXpQJvej1BddH8OqTXZNE/KfbWAu8Q==",
"version": "3.3.18",
"resolved": "https://registry.npmjs.org/nanoid/-/nanoid-3.3.18.tgz",
"integrity": "sha512-DTg4MJbGMWkfi6VZFdNt2/caMbQy4Ou+Op/hJQvGEWcnVfoA1QA+xzRKAzw9jD6+GVOOeYr/mIcuDSdug6F6+w==",
"dev": true,
"funding": [
{
@@ -4276,11 +4274,14 @@
"license": "MIT"
},
"node_modules/node-releases": {
"version": "2.0.37",
"resolved": "https://registry.npmjs.org/node-releases/-/node-releases-2.0.37.tgz",
"integrity": "sha512-1h5gKZCF+pO/o3Iqt5Jp7wc9rH3eJJ0+nh/CIoiRwjRxde/hAHyLPXYN4V3CqKAbiZPSeJFSWHmJsbkicta0Eg==",
"version": "2.0.54",
"resolved": "https://registry.npmjs.org/node-releases/-/node-releases-2.0.54.tgz",
"integrity": "sha512-YHs7BmmcsdAI5Ozuf8JZo6PT0mv2GIWC9vMfvUC3dp65M8hn7Ux8CPL+2oBI7juNuj9d0ndhTcznq2ODBps9cQ==",
"dev": true,
"license": "MIT"
"license": "MIT",
"engines": {
"node": ">=18"
}
},
"node_modules/object-assign": {
"version": "4.1.1",
@@ -4414,7 +4415,6 @@
"integrity": "sha512-QP88BAKvMam/3NxH6vj2o21R6MjxZUAd6nlwAS/pnGvN9IVLocLHxGYIzFhg6fUQ+5th6P4dv4eW9jX3DSIj7A==",
"dev": true,
"license": "MIT",
"peer": true,
"engines": {
"node": ">=12"
},
@@ -4551,7 +4551,6 @@
"resolved": "https://registry.npmjs.org/react/-/react-19.2.5.tgz",
"integrity": "sha512-llUJLzz1zTUBrskt2pwZgLq59AemifIftw4aB7JxOqf1HY2FDaGDxgwpAPVzHU1kdWabH7FauP4i1oEeer2WCA==",
"license": "MIT",
"peer": true,
"engines": {
"node": ">=0.10.0"
}
@@ -4617,7 +4616,6 @@
"resolved": "https://registry.npmjs.org/react-dom/-/react-dom-19.2.5.tgz",
"integrity": "sha512-J5bAZz+DXMMwW/wV3xzKke59Af6CHY7G4uYLN1OvBcKEsWOs4pQExj86BBKamxl/Ik5bx9whOrvBlSDfWzgSag==",
"license": "MIT",
"peer": true,
"dependencies": {
"scheduler": "^0.27.0"
},
@@ -4863,7 +4861,6 @@
"resolved": "https://registry.npmjs.org/sigma/-/sigma-3.0.2.tgz",
"integrity": "sha512-/BUbeOwPGruiBOm0YQQ6ZMcLIZ6tf/W+Jcm7dxZyAX0tK3WP9/sq7/NAWBxPIxVahdGjCJoGwej0Gdrv0DxlQQ==",
"license": "MIT",
"peer": true,
"dependencies": {
"events": "^3.3.0",
"graphology-utils": "^2.5.2"
@@ -4989,7 +4986,6 @@
"integrity": "sha512-X8EX+XV4QR5xCsrgxaED954zTDfY8KqlDtskKEL0cHhyS/P8b4IFOvGDQpsC9Q1XnLq915wEfwwY/zzskCtmhg==",
"dev": true,
"license": "MIT",
"peer": true,
"dependencies": {
"esbuild": "~0.28.0"
},
@@ -5022,7 +5018,6 @@
"integrity": "sha512-jl1vZzPDinLr9eUt3J/t7V6FgNEw9QjvBPdysz9KfQDD41fQrC2Y4vKQdiaUpFT4bXlb1RHhLpp8wtm6M5TgSw==",
"dev": true,
"license": "Apache-2.0",
"peer": true,
"bin": {
"tsc": "bin/tsc",
"tsserver": "bin/tsserver"
@@ -5150,9 +5145,9 @@
}
},
"node_modules/update-browserslist-db": {
"version": "1.2.3",
"resolved": "https://registry.npmjs.org/update-browserslist-db/-/update-browserslist-db-1.2.3.tgz",
"integrity": "sha512-Js0m9cx+qOgDxo0eMiFGEueWztz+d4+M3rGlmKPT+T4IS/jP4ylw3Nwpu6cpTTP8R1MAC1kF4VbdLt3ARf209w==",
"version": "1.3.2",
"resolved": "https://registry.npmjs.org/update-browserslist-db/-/update-browserslist-db-1.3.2.tgz",
"integrity": "sha512-UQ+MSxlhRm1bzjhU+DcuXfjFO1FzNtqhK5+9Yvlp90ItDLk5vT932A0rFu619nf7RVS+Y/VeaUW1jaRDqZ8VJw==",
"dev": true,
"funding": [
{
@@ -5246,7 +5241,6 @@
"resolved": "https://registry.npmjs.org/vis-data/-/vis-data-8.0.3.tgz",
"integrity": "sha512-jhnb6rJNqkKR1Qmlay0VuDXY9ZlvAnYN1udsrP4U+krgZEq7C0yNSKdZqmnCe13mdnf9AdVcdDGFOzy2mpPoqw==",
"license": "(Apache-2.0 OR MIT)",
"peer": true,
"funding": {
"type": "opencollective",
"url": "https://opencollective.com/visjs"
@@ -5301,7 +5295,6 @@
"integrity": "sha512-NTKlcQjlAK7MlQoyb6LgaqHc8sso/pVyUJYWMws3jg21uTJw/LddqIFPcPqP6PzpgbIcZyKI85sFE4HBrQDA8A==",
"dev": true,
"license": "MIT",
"peer": true,
"dependencies": {
"esbuild": "^0.25.0",
"fdir": "^6.4.4",
@@ -5440,7 +5433,6 @@
"integrity": "sha512-rftlrkhHZOcjDwkGlnUtZZkvaPHCsDATp4pGpuOOMDaTdDDXF91wuVDJoWoPsKX/3YPQ5fHuF3STjcYyKr+Qhg==",
"dev": true,
"license": "MIT",
"peer": true,
"funding": {
"url": "https://github.com/sponsors/colinhacks"
}
+2 -1
View File
@@ -9,7 +9,8 @@
"lint": "eslint .",
"preview": "vite preview",
"test:graph-store": "node --test tests/graphStore.multi-edge.test.mjs",
"test:graph-workspace": "node --import tsx --test tests/markdownContentViewer.test.ts tests/graphSceneState.display.test.ts tests/temporalLifecycle.test.ts",
"test:graph-workspace": "node --import tsx --test tests/markdownContentViewer.test.ts tests/graphSceneState.display.test.ts tests/temporalLifecycle.test.ts tests/deterministicExplorerRendering.test.ts tests/smallGraphLayout.test.ts tests/realtimeGraphAttributes.test.ts",
"test:deterministic-e2e": "node --import tsx --test tests/deterministicExplorerRendering.e2e.ts",
"test:plugin-registry": "node --import tsx --test tests/pluginRegistry.temporal.test.mjs"
},
"dependencies": {
+1
View File
@@ -86,6 +86,7 @@ export interface EdgeAttributes {
dominantEdgeType?: string;
representativeWeight?: number;
bundleKind?: "parallel" | "bidirectional" | "community";
isSmallGraph?: boolean;
edgeType: string;
@@ -21,7 +21,6 @@ import type Graph from "graphology";
import { batchMergeEdges, batchMergeNodes, graph } from "../../store/graphStore";
import { logEvent } from "../../store/registryStore";
import type { EdgeAttributes, NodeAttributes } from "../../store/graphStore";
import { curveGroupForPair } from "../../store/edgePairKeys.js";
import { InspectorPanel, MetricChip, SurfaceCard } from "../../ui/primitives";
import { lazy, Suspense } from "react";
import { SigmaSceneAdapter } from "./SigmaSceneAdapter";
@@ -41,6 +40,9 @@ import {
} from "./plugins";
import { explorationEffectsShouldLoad, neighborhoodPanelShouldLoad, temporalOverlayShouldLoad } from "./pluginRegistryPredicates";
import { shouldFetchTemporalBounds, shouldFetchTemporalSnapshot } from "./temporalLifecyclePredicates";
import { createTemporalSnapshotGuards, type TemporalSnapshotResponse } from "./temporalSnapshotGuards";
import { SMALL_GRAPH_MAX_NODES } from "./smallGraphLayout";
import { buildRealtimeEdgeAttributes } from "./realtimeGraphAttributes";
import type { LinkPrediction, PathResponse } from "./GraphInspectorPanel";
import type { GraphSceneHandle, GraphSceneRuntime } from "./scene";
import type {
@@ -1055,46 +1057,10 @@ function buildRealtimeNodeAttributes(payload: {
};
}
function buildRealtimeEdgeAttributes(payload: {
id: string;
familyId?: string;
source_id: string;
target_id: string;
type?: string;
weight?: number;
properties?: Record<string, unknown>;
}): EdgeAttributes {
const properties = payload.properties || {};
const isInferred = Boolean(properties.inferred);
const isBidirectional = graph.hasDirectedEdge(payload.target_id, payload.source_id);
const baseColor = isInferred ? GRAPH_THEME.palette.accent.path : GRAPH_THEME.palette.muted.edgeStructure;
return {
edgeId: payload.id,
familyId: payload.familyId || payload.id,
sourceId: payload.source_id,
targetId: payload.target_id,
weight: Number(payload.weight ?? 1),
edgeType: payload.type || "related_to",
properties,
size: 1,
baseSize: 1,
color: baseColor,
baseColor,
mutedColor: GRAPH_THEME.palette.muted.edgeOverview,
visualPriority: isInferred ? 0.95 : 0.5,
isBidirectional,
edgeFamily: isInferred ? "path" : isBidirectional ? "bidirectional" : "line",
curveGroup: isBidirectional ? curveGroupForPair(payload.source_id, payload.target_id) : null,
type: "line",
edgeVariant: isInferred ? "pathSignal" : isBidirectional ? "bidirectionalCurve" : "directional",
arrowVisibilityPolicy: isInferred ? "always" : "contextual",
relationshipStrength: isInferred ? 0.95 : 0.52,
isParallelPair: false,
parallelIndex: 0,
parallelCount: 1,
familySize: 1,
};
function synchronizeRealtimeSmallGraphEdges(isSmallGraph: boolean): void {
graph.forEachEdge((edgeId) => {
graph.setEdgeAttribute(edgeId, "isSmallGraph", isSmallGraph);
});
}
function buildSelectedNodeState(
@@ -1354,6 +1320,7 @@ export function GraphWorkspace({ externalFocusNodeId, externalFocusToken }: Grap
const lastExternalFocusTokenRef = useRef<number | undefined>(undefined);
const pluginRuntimeRef = useRef<GraphSceneRuntime | null>(null);
const appliedGraphSummarySignatureRef = useRef<string | null>(null);
const smallGraphModeRef = useRef(false);
const pluginInteractionStateRef = useRef<GraphInteractionState>({
hoveredNodeId: null,
selectedNodeId: "",
@@ -1381,6 +1348,12 @@ export function GraphWorkspace({ externalFocusNodeId, externalFocusToken }: Grap
}
appliedGraphSummarySignatureRef.current = signature;
smallGraphModeRef.current = Boolean(
graphSummary.layoutReady
&& !graphSummary.hasCoordinates
&& graphSummary.nodeCount > 0
&& graphSummary.nodeCount <= SMALL_GRAPH_MAX_NODES,
);
setGraphReady(true);
setGraphVersion((current) => current + 1);
setIsLayoutRunning(!graphSummary.layoutReady);
@@ -1479,6 +1452,23 @@ export function GraphWorkspace({ externalFocusNodeId, externalFocusToken }: Grap
summary?.edgeCount,
]);
// Guards the snapshot lifecycle: at most one in-flight request per scrubber
// position (identical-`at` polls are deduplicated, breaking the idle/play
// polling loop), applied snapshots are cached and re-applied on revisit, and
// a response applies only while the scrubber is still on its position
// (out-of-order responses cannot clobber the active-node count).
const temporalSnapshotGuardsRef = useRef<ReturnType<typeof createTemporalSnapshotGuards> | null>(null);
if (temporalSnapshotGuardsRef.current === null) {
temporalSnapshotGuardsRef.current = createTemporalSnapshotGuards();
}
const temporalSnapshotGuards = temporalSnapshotGuardsRef.current;
// A new graph summary means the graph data was replaced (reload/retry);
// snapshots cached against the previous graph are stale, so reset all state.
useEffect(() => {
temporalSnapshotGuards.reset();
}, [summary]);
useEffect(() => {
if (!canFetchTemporalSnapshot) {
return;
@@ -1488,37 +1478,67 @@ export function GraphWorkspace({ externalFocusNodeId, externalFocusToken }: Grap
return;
}
const atMs = debouncedTime.getTime();
const { seq, cached } = temporalSnapshotGuards.begin(atMs);
if (seq === null) {
// An identical request is already in flight: one request per position.
return;
}
let cancelled = false;
const applyData = (data: TemporalSnapshotResponse) => {
const nextActiveIds = new Set(data.active_node_ids);
requestAnimationFrame(() => {
if (cancelled) return;
if (!temporalSnapshotGuards.shouldApply(atMs, seq)) {
// The scrubber moved on (or this request was superseded): release the
// position so a return to it refetches instead of stalling.
temporalSnapshotGuards.finish(atMs, seq);
return;
}
const previous = prevActiveIdsRef.current;
previous.forEach((id) => {
if (!nextActiveIds.has(id) && graph.hasNode(id)) {
graph.setNodeAttribute(id, "hidden", true);
}
});
nextActiveIds.forEach((id) => {
if (graph.hasNode(id)) {
graph.setNodeAttribute(id, "hidden", false);
}
});
prevActiveIdsRef.current = nextActiveIds;
setActiveNodeCount(data.active_node_count);
setGraphVersion((current) => current + 1);
sceneRef.current?.getRuntime()?.requestRender();
temporalSnapshotGuards.apply(atMs, seq, data);
});
};
if (cached) {
// Returning to a position whose snapshot was already applied: re-apply
// the cached result without a network request.
applyData(cached);
return;
}
const applySnapshot = async () => {
try {
const at = debouncedTime.toISOString();
const response = await fetch(`/api/temporal/snapshot?at=${encodeURIComponent(at)}`);
if (!response.ok || cancelled) return;
const data: { active_node_ids: string[]; active_node_count: number } = await response.json();
if (!response.ok) {
// A failed request must be retryable if the scrubber returns.
if (!cancelled) temporalSnapshotGuards.finish(atMs, seq);
return;
}
if (cancelled) return;
const nextActiveIds = new Set(data.active_node_ids);
requestAnimationFrame(() => {
if (cancelled) return;
const previous = prevActiveIdsRef.current;
previous.forEach((id) => {
if (!nextActiveIds.has(id) && graph.hasNode(id)) {
graph.setNodeAttribute(id, "hidden", true);
}
});
nextActiveIds.forEach((id) => {
if (graph.hasNode(id)) {
graph.setNodeAttribute(id, "hidden", false);
}
});
prevActiveIdsRef.current = nextActiveIds;
setActiveNodeCount(data.active_node_count);
setGraphVersion((current) => current + 1);
sceneRef.current?.getRuntime()?.requestRender();
});
const data: TemporalSnapshotResponse = await response.json();
if (cancelled) return;
applyData(data);
} catch (fetchError) {
temporalSnapshotGuards.finish(atMs, seq);
if (!cancelled) {
console.error("[Temporal] Snapshot fetch failed", fetchError);
}
@@ -1528,6 +1548,8 @@ export function GraphWorkspace({ externalFocusNodeId, externalFocusToken }: Grap
applySnapshot();
return () => {
cancelled = true;
// A cancelled request must be retryable when its position is revisited.
temporalSnapshotGuards.finish(atMs, seq);
};
}, [
canFetchTemporalSnapshot,
@@ -1843,18 +1865,26 @@ export function GraphWorkspace({ externalFocusNodeId, externalFocusToken }: Grap
attributes: buildRealtimeNodeAttributes(payload),
},
]);
if (graph.order > SMALL_GRAPH_MAX_NODES) {
smallGraphModeRef.current = false;
}
synchronizeRealtimeSmallGraphEdges(smallGraphModeRef.current);
logEvent("add-node", `Added node ${payload.label ?? payload.id}${payload.nodeType ? ` (${payload.nodeType})` : ""} via realtime ws`, { nodeId: payload.id, nodeType: payload.nodeType });
setGraphVersion((current) => current + 1);
sceneRef.current?.getRuntime()?.requestRender();
}
if (eventType === "ADD_EDGE") {
const isSmallGraph = smallGraphModeRef.current;
batchMergeEdges([
{
id: String(payload.id),
familyId: payload.familyId ? String(payload.familyId) : String(payload.id),
source: payload.source_id,
target: payload.target_id,
attributes: buildRealtimeEdgeAttributes(payload),
attributes: buildRealtimeEdgeAttributes(payload, {
isBidirectional: graph.hasDirectedEdge(payload.target_id, payload.source_id),
isSmallGraph,
}),
},
]);
logEvent("add-edge", `Added edge ${payload.edgeType ?? payload.id} (${payload.source_id}${payload.target_id}) via realtime ws`, { edgeId: payload.id, edgeType: payload.edgeType, source: payload.source_id, target: payload.target_id });
@@ -1783,6 +1783,7 @@ export function resolveEdgeElementStyle(
const isCommunityBundle = attrs.bundleKind === "community";
const baseSize = Number(attrs.baseSize || attrs.size || 0.9);
const visualPriority = Number(attrs.visualPriority ?? 0);
const isSmallGraphEdge = viewMode === "full" && attrs.isSmallGraph === true;
const isFullBridgeEdge = viewMode === "full" && fullEdgeClass === "bridge";
const isFullBackboneEdge = viewMode === "full" && fullEdgeClass === "backbone";
const shouldCurveBridge = isFullBridgeEdge
@@ -1790,11 +1791,13 @@ export function resolveEdgeElementStyle(
const visibilityPolicy = resolveEdgeVisibilityPolicy(theme, viewMode, zoomTier, isCommunityBundle);
const isContextEdge = isContextEdgeState(state);
const isNonCriticalEdge = isNonCriticalEdgeVariant(edgeVariant);
const belowPriorityThreshold = state === "default"
const belowPriorityThreshold = !isSmallGraphEdge && state === "default"
&& visualPriority < Math.max(tierConfig.edgePriorityThreshold, visibilityPolicy.defaultPriorityThreshold)
&& isNonCriticalEdge;
const hiddenByMutedState = (state === "muted" || state === "inactive") && visibilityPolicy.hideMuted;
const sampledOut = isNonCriticalEdge
const hiddenByMutedState = !isSmallGraphEdge
&& (state === "muted" || state === "inactive")
&& visibilityPolicy.hideMuted;
const sampledOut = !isSmallGraphEdge && isNonCriticalEdge
&& (
(state === "default" && !isContextEdge && shouldSampleOutBackgroundEdge(visibilityPolicy.backgroundSampleRate, visualPriority, edgeId, sourceId, targetId))
|| (
@@ -1837,11 +1840,14 @@ export function resolveEdgeElementStyle(
? resolveEdgeCurvature(theme, state, edgeVariant, attrs, sourceId, targetId)
: 0;
const baseColor = resolveEdgeColor(theme, zoomTier, state, attrs, attrs.color, fullEdgeClass);
const lodAlpha = resolveEdgeLodAlpha(theme, viewMode, zoomTier, state, attrs, isCommunityBundle, fullEdgeClass);
const resolvedLodAlpha = resolveEdgeLodAlpha(theme, viewMode, zoomTier, state, attrs, isCommunityBundle, fullEdgeClass);
const lodAlpha = isSmallGraphEdge
? Math.max(resolvedLodAlpha ?? 1, isContextEdge ? 0.62 : 0.46)
: resolvedLodAlpha;
const color = lodAlpha === null ? baseColor : withAlpha(baseColor, lodAlpha);
const rawSize = Math.max(
baseSize * sizeMultiplier * (isCommunityBundle ? theme.grouped.style.edgeSizeScale : 1),
stateConfig.minSize,
isSmallGraphEdge ? Math.max(stateConfig.minSize, 0.9) : stateConfig.minSize,
);
const interactionMaxSize = (fullEdgeClass === "path" || state === "path")
@@ -0,0 +1,50 @@
import type { EdgeAttributes } from "../../store/graphStore";
import { curveGroupForPair } from "../../store/edgePairKeys.js";
import { GRAPH_THEME } from "./graphTheme";
export type RealtimeEdgePayload = {
id: string;
familyId?: string;
source_id: string;
target_id: string;
type?: string;
weight?: number;
properties?: Record<string, unknown>;
};
export function buildRealtimeEdgeAttributes(
payload: RealtimeEdgePayload,
options: { isBidirectional: boolean; isSmallGraph: boolean },
): EdgeAttributes {
const properties = payload.properties || {};
const isInferred = Boolean(properties.inferred);
const baseColor = isInferred ? GRAPH_THEME.palette.accent.path : GRAPH_THEME.palette.muted.edgeStructure;
return {
edgeId: payload.id,
familyId: payload.familyId || payload.id,
sourceId: payload.source_id,
targetId: payload.target_id,
weight: Number(payload.weight ?? 1),
edgeType: payload.type || "related_to",
properties,
size: 1,
baseSize: 1,
color: baseColor,
baseColor,
mutedColor: GRAPH_THEME.palette.muted.edgeOverview,
visualPriority: isInferred ? 0.95 : 0.5,
isBidirectional: options.isBidirectional,
edgeFamily: isInferred ? "path" : options.isBidirectional ? "bidirectional" : "line",
curveGroup: options.isBidirectional ? curveGroupForPair(payload.source_id, payload.target_id) : null,
type: "line",
edgeVariant: isInferred ? "pathSignal" : options.isBidirectional ? "bidirectionalCurve" : "directional",
arrowVisibilityPolicy: isInferred ? "always" : "contextual",
relationshipStrength: isInferred ? 0.95 : 0.52,
isParallelPair: false,
parallelIndex: 0,
parallelCount: 1,
familySize: 1,
isSmallGraph: options.isSmallGraph,
};
}
@@ -0,0 +1,135 @@
export const SMALL_GRAPH_MAX_NODES = 48;
const PROVIDED_COORDINATE_COVERAGE = 0.92;
const MAX_COMPONENT_RADIUS = 78;
const COMPONENT_GAP = 48;
type LayoutEdge = {
source: string;
target: string;
};
export function shouldUseSmallGraphLayout(nodeCount: number, coordinateCoverage: number): boolean {
return nodeCount > 0
&& nodeCount <= SMALL_GRAPH_MAX_NODES
&& coordinateCoverage < PROVIDED_COORDINATE_COVERAGE;
}
export function resolveGraphLayoutDecision(nodeCount: number, coordinateCoverage: number): {
useProvidedCoordinates: boolean;
useSmallGraphLayout: boolean;
layoutReady: boolean;
} {
const useProvidedCoordinates = coordinateCoverage >= PROVIDED_COORDINATE_COVERAGE;
const useSmallGraphLayout = shouldUseSmallGraphLayout(nodeCount, coordinateCoverage);
return {
useProvidedCoordinates,
useSmallGraphLayout,
layoutReady: useProvidedCoordinates || useSmallGraphLayout,
};
}
export function resolveNodeLayoutPosition(
decision: ReturnType<typeof resolveGraphLayoutDecision>,
provided: { x: number | null; y: number | null },
seeded: { x: number; y: number } | undefined,
): { x: number; y: number } {
if (decision.useProvidedCoordinates) {
return { x: provided.x ?? 0, y: provided.y ?? 0 };
}
if (decision.useSmallGraphLayout) {
return { x: seeded?.x ?? 0, y: seeded?.y ?? 0 };
}
return {
x: provided.x ?? seeded?.x ?? 0,
y: provided.y ?? seeded?.y ?? 0,
};
}
/**
* Produce a compact deterministic layout for small graphs.
*
* ForceAtlas2 is useful for large connected datasets, but it makes tiny graphs
* with several disconnected components look like scattered dots. This layout
* keeps each connected component together and packs components into a centered
* grid so instance relationships remain legible on first render.
*/
export function buildSmallGraphSeedPositions(
nodeIds: string[],
edges: LayoutEdge[],
): Map<string, { x: number; y: number }> {
const ids = [...new Set(nodeIds)].sort((left, right) => left.localeCompare(right));
const adjacency = new Map(ids.map((id) => [id, new Set<string>()]));
edges.forEach(({ source, target }) => {
if (!adjacency.has(source) || !adjacency.has(target) || source === target) {
return;
}
adjacency.get(source)?.add(target);
adjacency.get(target)?.add(source);
});
const visited = new Set<string>();
const components: string[][] = [];
ids.forEach((start) => {
if (visited.has(start)) {
return;
}
const component: string[] = [];
const queue = [start];
visited.add(start);
while (queue.length > 0) {
const current = queue.shift();
if (!current) {
continue;
}
component.push(current);
[...(adjacency.get(current) ?? [])]
.sort((left, right) => left.localeCompare(right))
.forEach((neighbor) => {
if (!visited.has(neighbor)) {
visited.add(neighbor);
queue.push(neighbor);
}
});
}
component.sort((left, right) => {
const degreeDelta = (adjacency.get(right)?.size ?? 0) - (adjacency.get(left)?.size ?? 0);
return degreeDelta || left.localeCompare(right);
});
components.push(component);
});
components.sort((left, right) => right.length - left.length || left[0].localeCompare(right[0]));
const columns = Math.max(1, Math.ceil(Math.sqrt(components.length)));
const rows = Math.max(1, Math.ceil(components.length / columns));
// Adjacent cells must leave room for two maximum-radius components plus a
// readable gap. A smaller row height allows valid 12-node components to
// overlap vertically.
const cellWidth = MAX_COMPONENT_RADIUS * 2 + COMPONENT_GAP;
const cellHeight = MAX_COMPONENT_RADIUS * 2 + COMPONENT_GAP;
const positions = new Map<string, { x: number; y: number }>();
components.forEach((component, componentIndex) => {
const column = componentIndex % columns;
const row = Math.floor(componentIndex / columns);
const centerX = (column - (columns - 1) / 2) * cellWidth;
const centerY = (row - (rows - 1) / 2) * cellHeight;
if (component.length === 1) {
positions.set(component[0], { x: centerX, y: centerY });
return;
}
const radius = Math.min(MAX_COMPONENT_RADIUS, 30 + component.length * 9);
component.forEach((nodeId, nodeIndex) => {
const angle = -Math.PI / 2 + (nodeIndex * Math.PI * 2) / component.length;
positions.set(nodeId, {
x: centerX + Math.cos(angle) * radius,
y: centerY + Math.sin(angle) * radius,
});
});
});
return positions;
}
@@ -0,0 +1,113 @@
/**
* Guards for the temporal snapshot fetch/apply lifecycle.
*
* The snapshot effect previously fetched /api/temporal/snapshot with no
* idempotency or ordering protection. Upstream churn (timeline recreation
* while bounds settle, play ticks resetting the playhead, drag events) could
* re-request the same `at` repeatedly, and responses could arrive after the
* scrubber had moved on.
*
* The guards enforce:
* - at most one in-flight request per scrubber position (identical `at`
* values are deduplicated while a request is pending, breaking the
* idle/play polling loop);
* - successful snapshots are cached per position and re-applied when the
* scrubber returns (play wrap-around, back-scrubbing) without a refetch;
* - a response is applied only while the scrubber is still on its position,
* so out-of-order responses cannot clobber a newer position's count;
* - failed, cancelled, or superseded requests release their position so it
* can be fetched again on the next visit;
* - `reset()` drops all state when the underlying graph data is replaced
* (reload/retry), because cached snapshots describe the previous graph.
*
* `createTemporalSnapshotGuards()` is stateful by design.
*/
export interface TemporalSnapshotResponse {
active_node_ids: string[];
active_node_count: number;
}
export interface TemporalSnapshotRequest {
/** null when the request was deduplicated because one is already in flight. */
seq: number | null;
/** The snapshot previously applied for this position, when revisiting it. */
cached: TemporalSnapshotResponse | null;
}
export interface TemporalSnapshotGuards {
/** Begin (or dedupe) a request for `atMs`; marks it as the current position. */
begin(atMs: number): TemporalSnapshotRequest;
/** True when the response for `atMs`/`seq` may be applied (scrubber still on `atMs`). */
shouldApply(atMs: number, seq: number): boolean;
/** Record a successful application and cache its snapshot for revisits. */
apply(atMs: number, seq: number, data: TemporalSnapshotResponse): void;
/** Release a position whose request failed, was cancelled, or was superseded. */
finish(atMs: number, seq: number): void;
/** Drop all state; call when the underlying graph data is replaced (reload). */
reset(): void;
}
interface SnapshotEntry {
seq: number;
/** null while the request is in flight (or before the first success). */
data: TemporalSnapshotResponse | null;
}
/** Upper bound on cached positions so long scrubbing sessions stay bounded. */
const MAX_CACHED_POSITIONS = 256;
export function createTemporalSnapshotGuards(): TemporalSnapshotGuards {
const entries = new Map<number, SnapshotEntry>();
let latestRequestSeq = 0;
let currentAtMs: number | null = null;
const evictOldest = () => {
while (entries.size > MAX_CACHED_POSITIONS) {
const oldestAtMs = entries.keys().next().value;
if (oldestAtMs === undefined) return;
entries.delete(oldestAtMs);
}
};
return {
begin(atMs) {
const existing = entries.get(atMs);
if (existing && existing.data === null) {
// Identical request already in flight: dedupe, but the scrubber is here now.
currentAtMs = atMs;
return { seq: null, cached: null };
}
latestRequestSeq += 1;
const seq = latestRequestSeq;
entries.set(atMs, { seq, data: existing?.data ?? null });
currentAtMs = atMs;
evictOldest();
return { seq, cached: existing?.data ?? null };
},
shouldApply(atMs, seq) {
return atMs === currentAtMs && entries.get(atMs)?.seq === seq;
},
apply(atMs, seq, data) {
const entry = entries.get(atMs);
if (entry && entry.seq === seq) {
entry.data = data;
}
},
finish(atMs, seq) {
const entry = entries.get(atMs);
if (entry && entry.seq === seq && entry.data === null) {
entries.delete(atMs);
}
},
reset() {
entries.clear();
latestRequestSeq = 0;
currentAtMs = null;
},
};
}
@@ -16,6 +16,11 @@ import {
} from "./graphTheme";
import { classifyEntityShape } from "./graphEntityShape";
import { createGraphLoadProgress } from "./graphLoading";
import {
buildSmallGraphSeedPositions,
resolveGraphLayoutDecision,
resolveNodeLayoutPosition,
} from "./smallGraphLayout";
import type { GraphLoadProgress, GraphLoadSummary } from "./types";
const SEMANTIC_COLOR_FIELDS = [
@@ -318,6 +323,20 @@ interface EdgeListResponse {
const PAGE_LIMIT = 1000;
/** Surface the server's `detail` message (e.g. auth/setup guidance) on non-OK responses. */
async function fetchErrorDetail(response: Response): Promise<string> {
try {
const body: unknown = await response.json();
const detail = (body as { detail?: unknown } | null)?.detail;
if (typeof detail === "string" && detail.trim()) {
return `${detail.trim()}`;
}
} catch {
// Non-JSON or unreadable body: fall back to the status-only message.
}
return "";
}
async function fetchAllNodes(
signal: AbortSignal,
onProgress?: (progress: GraphLoadProgress) => void,
@@ -335,7 +354,7 @@ async function fetchAllNodes(
const response = await fetch(url.toString(), { signal });
if (!response.ok) {
throw new Error(`Fetch failed: ${response.status}`);
throw new Error(`Fetch failed: ${response.status}${await fetchErrorDetail(response)}`);
}
const data: NodeListResponse = await response.json();
@@ -390,7 +409,7 @@ async function fetchAllEdges(
const response = await fetch(url.toString(), { signal });
if (!response.ok) {
throw new Error(`Fetch failed: ${response.status}`);
throw new Error(`Fetch failed: ${response.status}${await fetchErrorDetail(response)}`);
}
const data: EdgeListResponse = await response.json();
@@ -539,10 +558,19 @@ export function useLoadGraph(options: UseLoadGraphOptions = {}) {
: count;
}, 0);
const coordinateCoverage = fetchedNodes.length > 0 ? providedCoordinateCount / fetchedNodes.length : 0;
const useProvidedCoordinates = coordinateCoverage >= 0.92;
const {
useProvidedCoordinates,
useSmallGraphLayout,
layoutReady,
} = resolveGraphLayoutDecision(fetchedNodes.length, coordinateCoverage);
const seededPositions = useProvidedCoordinates
? null
: buildClusterSeedPositions(
: useSmallGraphLayout
? buildSmallGraphSeedPositions(
fetchedNodes.map((node) => node.id),
fetchedEdges,
)
: buildClusterSeedPositions(
draftAttributes.map(({ id, attributes }) => ({
id,
semanticGroup: semanticKeyByNodeId.get(id) ?? structuralColorKey(id, attributes),
@@ -555,7 +583,9 @@ export function useLoadGraph(options: UseLoadGraphOptions = {}) {
const colorIndex = hashString(semanticGroup) % GRAPH_THEME.palette.semantic.length;
const baseColor = GRAPH_THEME.palette.semantic[colorIndex];
const sizeRatio = nodePriorityById.get(id) ?? 0;
const dynamicSize = clamp(1.8, 1.8 + 8.8 * sizeRatio, 11.8);
const dynamicSize = useSmallGraphLayout
? clamp(5.2, 5.2 + 6.6 * sizeRatio, 11.8)
: clamp(1.8, 1.8 + 8.8 * sizeRatio, 11.8);
const hasTemporalBounds = Boolean(attributes.valid_from || attributes.valid_until);
const provenanceCount = getProvenanceCount(attributes.properties ?? {});
const properties = attributes.properties as Record<string, unknown>;
@@ -563,12 +593,11 @@ export function useLoadGraph(options: UseLoadGraphOptions = {}) {
const providedX = readFiniteCoordinate(properties?.x);
const providedY = readFiniteCoordinate(properties?.y);
const seededPosition = seededPositions?.get(id);
const x = useProvidedCoordinates
? providedX ?? 0
: providedX ?? seededPosition?.x ?? 0;
const y = useProvidedCoordinates
? providedY ?? 0
: providedY ?? seededPosition?.y ?? 0;
const { x, y } = resolveNodeLayoutPosition(
{ useProvidedCoordinates, useSmallGraphLayout, layoutReady },
{ x: providedX, y: providedY },
seededPosition,
);
return {
id,
attributes: {
@@ -589,6 +618,7 @@ export function useLoadGraph(options: UseLoadGraphOptions = {}) {
borderSize: 0.72,
entityShape,
...resolveNodeVariantMetadata(baseColor, sizeRatio, hasTemporalBounds, provenanceCount),
...(useSmallGraphLayout ? { labelVisibilityPolicy: "always" as const } : {}),
} as NodeAttributes,
};
});
@@ -645,6 +675,7 @@ export function useLoadGraph(options: UseLoadGraphOptions = {}) {
parallelIndex,
parallelCount,
familySize: familyCounts.get(edge.familyId) ?? 1,
isSmallGraph: useSmallGraphLayout,
...resolveEdgeVariantMetadata(edge, sourcePriority, targetPriority, isBidirectional),
} as EdgeAttributes,
};
@@ -687,7 +718,7 @@ export function useLoadGraph(options: UseLoadGraphOptions = {}) {
loadTimeMs: Math.round(performance.now() - startedAt),
hasCoordinates: useProvidedCoordinates,
layoutSource: useProvidedCoordinates ? "provided" : "runtime",
layoutReady: useProvidedCoordinates,
layoutReady,
} satisfies GraphLoadSummary;
onProgress?.(createGraphLoadProgress({
@@ -0,0 +1,103 @@
import assert from "node:assert/strict";
import { spawn, type ChildProcess } from "node:child_process";
import { existsSync } from "node:fs";
import { setTimeout as delay } from "node:timers/promises";
import test from "node:test";
import { chromium, type Page } from "playwright";
const PORT = 4173;
const BASE_URL = `http://127.0.0.1:${PORT}`;
const nodes = [
{ id: "alice", type: "Person", content: "Alice", properties: {} },
{ id: "bob", type: "Person", content: "Bob", properties: {} },
{ id: "acme", type: "Organization", content: "Acme", properties: {} },
{ id: "new_york", type: "Location", content: "New York", properties: {} },
];
const edges = [
{ id: "edge_alice_acme", familyId: "edge_alice_acme", source: "alice", target: "acme", type: "WORKS_AT", weight: 1, properties: {} },
{ id: "edge_bob_alice", familyId: "edge_bob_alice", source: "bob", target: "alice", type: "KNOWS", weight: 1, properties: {} },
{ id: "edge_acme_new_york", familyId: "edge_acme_new_york", source: "acme", target: "new_york", type: "LOCATED_IN", weight: 1, properties: {} },
];
let server: ChildProcess | undefined;
async function startVite(): Promise<void> {
server = spawn("npm", ["run", "dev", "--", "--host", "127.0.0.1", "--port", String(PORT)], {
cwd: process.cwd(),
stdio: "ignore",
});
for (let attempt = 0; attempt < 50; attempt += 1) {
try {
const response = await fetch(BASE_URL);
if (response.ok) return;
} catch {
// Vite is still starting.
}
await delay(100);
}
throw new Error("Vite did not become ready");
}
async function installApiFixture(page: Page): Promise<void> {
await page.route("**/api/graph/**", async (route) => {
const pathname = new URL(route.request().url()).pathname;
if (pathname === "/api/graph/stats") {
await route.fulfill({ json: { node_count: 4, edge_count: 3 } });
} else if (pathname === "/api/graph/nodes") {
await route.fulfill({ json: { nodes, total: nodes.length, skip: 0, limit: 1000, next_cursor: null } });
} else if (pathname === "/api/graph/edges") {
await route.fulfill({ json: { edges, total: edges.length, skip: 0, limit: 1000, next_cursor: null } });
} else {
await route.continue();
}
});
}
test("real Explorer loading path hydrates and renders API edge labels", async (t) => {
await startVite();
t.after(async () => {
server?.kill();
});
const browser = await chromium.launch({
headless: true,
executablePath: process.env.CHROMIUM_PATH || (existsSync("/usr/bin/chromium") ? "/usr/bin/chromium" : undefined),
});
t.after(() => browser.close());
const page = await browser.newPage({ viewport: { width: 1440, height: 1000 } });
await page.addInitScript(() => {
const captured = (window as Window & { __capturedCanvasText?: string[] }).__capturedCanvasText = [];
const originalFillText = CanvasRenderingContext2D.prototype.fillText;
CanvasRenderingContext2D.prototype.fillText = function (text: string, ...args: [number, number, number?, number?]) {
captured.push(String(text));
return originalFillText.call(this, text, ...args);
};
});
await installApiFixture(page);
await page.goto(BASE_URL);
await page.getByRole("button", { name: /Open Semantica Explorer/ }).click();
await page.locator("canvas").nth(0).waitFor({ state: "attached" });
await page.waitForFunction(() => document.querySelectorAll("canvas").length >= 2);
await page.waitForFunction(() => {
const labels = (window as Window & { __capturedCanvasText?: string[] }).__capturedCanvasText ?? [];
return ["WORKS_AT", "KNOWS", "LOCATED_IN"].every((label) => labels.includes(label));
}, undefined, { timeout: 10_000 });
const capturedLabels = await page.evaluate(() => (window as Window & { __capturedCanvasText?: string[] }).__capturedCanvasText ?? []);
for (const label of ["WORKS_AT", "KNOWS", "LOCATED_IN"]) {
assert.ok(capturedLabels.includes(label), `Expected rendered edge label ${label}`);
}
assert.ok(capturedLabels.includes("Alice"));
await page.getByRole("button", { name: "Zoom In" }).click();
await page.waitForTimeout(250);
const labelsAfterZoom = await page.evaluate(() => (window as Window & { __capturedCanvasText?: string[] }).__capturedCanvasText ?? []);
for (const label of ["WORKS_AT", "KNOWS", "LOCATED_IN"]) {
assert.ok(labelsAfterZoom.includes(label), `Expected edge label ${label} after zoom`);
}
});
@@ -0,0 +1,508 @@
import test from "node:test";
import assert from "node:assert/strict";
import {
batchMergeEdges,
batchMergeNodes,
clearGraph,
graph,
} from "../src/store/graphStore.ts";
import {
buildStructuralDistanceSnapshot,
classifyFullGraphEdge,
resolveDisplayGraph,
resolveEdgeElementStyle,
resolveEdgeVisualState,
resolveNodeElementStyle,
resolveNodeVisualState,
shouldForceNodeLabel,
} from "../src/workspaces/GraphWorkspace/graphSceneState.ts";
import { GRAPH_THEME, type GraphZoomTier } from "../src/workspaces/GraphWorkspace/graphTheme.ts";
test.beforeEach(() => {
clearGraph();
});
test.after(() => {
clearGraph();
});
/**
* Loads the canonical 4-node, 3-edge deterministic test graph (Semantica #1037).
*
* Graph structure:
* Alice (Person) --WORKS_AT--> Acme (Organization)
* Bob (Person) --KNOWS--> Alice (Person)
* Acme (Organization) --LOCATED_IN--> New York (Location)
*/
function loadDeterministicTestGraph() {
batchMergeNodes([
{
id: "alice",
attributes: {
label: "Alice",
content: "Alice",
x: 0,
y: 0,
size: 8,
color: "#63E6FF",
baseColor: "#63E6FF",
nodeType: "Person",
semanticGroup: "Person",
properties: {},
},
},
{
id: "bob",
attributes: {
label: "Bob",
content: "Bob",
x: -50,
y: 0,
size: 8,
color: "#63E6FF",
baseColor: "#63E6FF",
nodeType: "Person",
semanticGroup: "Person",
properties: {},
},
},
{
id: "acme",
attributes: {
label: "Acme",
content: "Acme",
x: 50,
y: 0,
size: 8,
color: "#A78BFA",
baseColor: "#A78BFA",
nodeType: "Organization",
semanticGroup: "Organization",
properties: {},
},
},
{
id: "new_york",
attributes: {
label: "New York",
content: "New York",
x: 100,
y: 0,
size: 8,
color: "#34D399",
baseColor: "#34D399",
nodeType: "Location",
semanticGroup: "Location",
properties: {},
},
},
]);
batchMergeEdges([
{
id: "edge_alice_acme",
source: "alice",
target: "acme",
attributes: {
edgeId: "edge_alice_acme",
edgeType: "WORKS_AT",
weight: 1.0,
visualPriority: 0.8,
baseSize: 0.8,
properties: {},
},
},
{
id: "edge_bob_alice",
source: "bob",
target: "alice",
attributes: {
edgeId: "edge_bob_alice",
edgeType: "KNOWS",
weight: 1.0,
visualPriority: 0.8,
baseSize: 0.8,
properties: {},
},
},
{
id: "edge_acme_new_york",
source: "acme",
target: "new_york",
attributes: {
edgeId: "edge_acme_new_york",
edgeType: "LOCATED_IN",
weight: 1.0,
visualPriority: 0.8,
baseSize: 0.8,
properties: {},
},
},
]);
}
test("deterministic graph contains exactly 4 nodes and 3 edges in store", () => {
loadDeterministicTestGraph();
assert.equal(graph.order, 4, "Expected exactly 4 nodes");
assert.equal(graph.size, 3, "Expected exactly 3 edges");
// Verify node identities and labels
const alice = graph.getNodeAttributes("alice");
const bob = graph.getNodeAttributes("bob");
const acme = graph.getNodeAttributes("acme");
const newYork = graph.getNodeAttributes("new_york");
assert.equal(alice.label, "Alice");
assert.equal(alice.nodeType, "Person");
assert.equal(alice.color, "#63E6FF");
assert.equal(bob.label, "Bob");
assert.equal(bob.nodeType, "Person");
assert.equal(bob.color, "#63E6FF");
assert.equal(acme.label, "Acme");
assert.equal(acme.nodeType, "Organization");
assert.equal(acme.color, "#A78BFA");
assert.equal(newYork.label, "New York");
assert.equal(newYork.nodeType, "Location");
assert.equal(newYork.color, "#34D399");
// Verify edge connectivity and canonical edgeType labels
const edgeAliceAcme = graph.getEdgeAttributes("edge_alice_acme");
const edgeBobAlice = graph.getEdgeAttributes("edge_bob_alice");
const edgeAcmeNewYork = graph.getEdgeAttributes("edge_acme_new_york");
assert.equal(edgeAliceAcme.edgeType, "WORKS_AT");
assert.equal(graph.source("edge_alice_acme"), "alice");
assert.equal(graph.target("edge_alice_acme"), "acme");
assert.equal(edgeBobAlice.edgeType, "KNOWS");
assert.equal(graph.source("edge_bob_alice"), "bob");
assert.equal(graph.target("edge_bob_alice"), "alice");
assert.equal(edgeAcmeNewYork.edgeType, "LOCATED_IN");
assert.equal(graph.source("edge_acme_new_york"), "acme");
assert.equal(graph.target("edge_acme_new_york"), "new_york");
});
test("display graph resolution preserves all 4 nodes and 3 edges in full view", () => {
loadDeterministicTestGraph();
const { graph: displayGraph } = resolveDisplayGraph("", [], [], "full", { aggregationEnabled: false });
assert.equal(displayGraph.order, 4);
assert.equal(displayGraph.size, 3);
assert.ok(displayGraph.hasNode("alice"));
assert.ok(displayGraph.hasNode("bob"));
assert.ok(displayGraph.hasNode("acme"));
assert.ok(displayGraph.hasNode("new_york"));
assert.ok(displayGraph.hasEdge("edge_alice_acme"));
assert.ok(displayGraph.hasEdge("edge_bob_alice"));
assert.ok(displayGraph.hasEdge("edge_acme_new_york"));
});
test("structural distance calculation resolves correct hop counts across the 3-edge chain", () => {
loadDeterministicTestGraph();
// From Bob: Bob (0) -> Alice (1) -> Acme (2) -> New York (3)
const distances = buildStructuralDistanceSnapshot(graph, "bob", 3);
assert.equal(distances.bob, 0);
assert.equal(distances.alice, 1);
assert.equal(distances.acme, 2);
assert.equal(distances.new_york, 3);
});
test("edge rendering and canonical edge labels remain legible across zoom tiers and inspection modes", () => {
loadDeterministicTestGraph();
const canonicalEdges = [
{ id: "edge_alice_acme", source: "alice", target: "acme", label: "WORKS_AT" },
{ id: "edge_bob_alice", source: "bob", target: "alice", label: "KNOWS" },
{ id: "edge_acme_new_york", source: "acme", target: "new_york", label: "LOCATED_IN" },
];
// 1. Edge attributes preserve canonical edgeType labels in graph store:
for (const item of canonicalEdges) {
const attrs = graph.getEdgeAttributes(item.id);
assert.equal(attrs.edgeType, item.label, `Edge ${item.id} must have edgeType ${item.label}`);
assert.equal(graph.source(item.id), item.source);
assert.equal(graph.target(item.id), item.target);
}
// 2. In active context / neighbor state across all zoom tiers (overview, structure, inspection):
const allTiers: GraphZoomTier[] = ["overview", "structure", "inspection"];
for (const tier of allTiers) {
for (const item of canonicalEdges) {
const attrs = graph.getEdgeAttributes(item.id);
const contextStyle = resolveEdgeElementStyle(
GRAPH_THEME,
tier,
"neighbor",
attrs,
item.source,
item.target,
"full",
item.id,
);
assert.equal(
contextStyle.hidden,
false,
`Edge ${item.id} (${item.label}) in context state 'neighbor' must be visible in zoom tier '${tier}'`,
);
assert.ok(
contextStyle.size !== undefined && contextStyle.size > 0,
`Edge ${item.id} (${item.label}) must have positive render size in zoom tier '${tier}'`,
);
}
}
// 3. In selected state in inspection zoom tier (close examination of edge details and label):
for (const item of canonicalEdges) {
const attrs = graph.getEdgeAttributes(item.id);
const selectedStyle = resolveEdgeElementStyle(
GRAPH_THEME,
"inspection",
"selected",
attrs,
item.source,
item.target,
"full",
item.id,
"selected",
);
assert.equal(
selectedStyle.hidden,
false,
`Selected edge ${item.id} (${item.label}) must be visible in inspection zoom tier`,
);
assert.ok(
selectedStyle.size !== undefined && selectedStyle.size > 0,
`Selected edge ${item.id} (${item.label}) must have positive render size`,
);
}
// 4. Verify inspection zoom tier camera and arrow rendering settings
assert.equal(GRAPH_THEME.zoomTiers.inspection.showContextualArrows, true);
assert.equal(GRAPH_THEME.zoomTiers.inspection.showCurves, true);
});
test("node hover interaction preserves edge visibility and highlights canonical incident edge types", () => {
loadDeterministicTestGraph();
// Scenario 1: Hover Alice
// Incident edges: Alice -> Acme (WORKS_AT) and Bob -> Alice (KNOWS)
const aliceAttrs = graph.getNodeAttributes("alice");
const aliceVisual = resolveNodeVisualState("alice", "structure", "alice", "", "", new Set(), new Set(), new Set());
assert.equal(aliceVisual, "hovered");
const aliceStyle = resolveNodeElementStyle(GRAPH_THEME, "structure", "hovered", aliceAttrs, "Alice");
assert.equal(aliceStyle.forceLabel, true, "Hovered Alice must force-render label");
assert.equal(aliceStyle.label, "Alice");
assert.equal(aliceStyle.showHalo, true, "Hovered Alice must show interactive halo");
const aliceIncidentEdges = new Set(["edge_alice_acme", "edge_bob_alice"]);
// Edge Alice -> Acme (WORKS_AT) under Alice hover
const aliceAcmeAttrs = graph.getEdgeAttributes("edge_alice_acme");
assert.equal(aliceAcmeAttrs.edgeType, "WORKS_AT");
const aliceAcmeState = resolveEdgeVisualState(
"edge_alice_acme",
"alice",
"acme",
"structure",
"alice",
"",
"",
new Set(),
new Set(),
aliceIncidentEdges,
);
assert.equal(aliceAcmeState, "hovered");
const aliceAcmeStyle = resolveEdgeElementStyle(
GRAPH_THEME,
"structure",
"hovered",
aliceAcmeAttrs,
"alice",
"acme",
"full",
"edge_alice_acme",
);
assert.equal(aliceAcmeStyle.hidden, false, "Incident edge WORKS_AT must remain visible on hover");
assert.ok(aliceAcmeStyle.size !== undefined && aliceAcmeStyle.size > 0);
// Edge Bob -> Alice (KNOWS) under Alice hover
const bobAliceAttrs = graph.getEdgeAttributes("edge_bob_alice");
assert.equal(bobAliceAttrs.edgeType, "KNOWS");
const bobAliceState = resolveEdgeVisualState(
"edge_bob_alice",
"bob",
"alice",
"structure",
"alice",
"",
"",
new Set(),
new Set(),
aliceIncidentEdges,
);
assert.equal(bobAliceState, "hovered");
const bobAliceStyle = resolveEdgeElementStyle(
GRAPH_THEME,
"structure",
"hovered",
bobAliceAttrs,
"bob",
"alice",
"full",
"edge_bob_alice",
);
assert.equal(bobAliceStyle.hidden, false, "Incident edge KNOWS must remain visible on hover");
// Non-incident edge Acme -> New York (LOCATED_IN) under Alice hover
const acmeNyAttrs = graph.getEdgeAttributes("edge_acme_new_york");
assert.equal(acmeNyAttrs.edgeType, "LOCATED_IN");
const acmeNyState = resolveEdgeVisualState(
"edge_acme_new_york",
"acme",
"new_york",
"structure",
"alice",
"",
"",
new Set(),
new Set(),
aliceIncidentEdges,
);
assert.equal(acmeNyState, "muted");
// Scenario 2: Hover Acme
// Incident edges: Alice -> Acme (WORKS_AT) and Acme -> New York (LOCATED_IN)
const acmeAttrs = graph.getNodeAttributes("acme");
const acmeStyle = resolveNodeElementStyle(GRAPH_THEME, "structure", "hovered", acmeAttrs, "Acme");
assert.equal(acmeStyle.forceLabel, true);
assert.equal(acmeStyle.label, "Acme");
const acmeIncidentEdges = new Set(["edge_alice_acme", "edge_acme_new_york"]);
const acmeNyHoverState = resolveEdgeVisualState(
"edge_acme_new_york",
"acme",
"new_york",
"structure",
"acme",
"",
"",
new Set(),
new Set(),
acmeIncidentEdges,
);
assert.equal(acmeNyHoverState, "hovered");
const acmeNyHoverStyle = resolveEdgeElementStyle(
GRAPH_THEME,
"structure",
"hovered",
acmeNyAttrs,
"acme",
"new_york",
"full",
"edge_acme_new_york",
);
assert.equal(acmeNyHoverStyle.hidden, false, "Incident edge LOCATED_IN must remain visible on hover");
// Scenario 3: Hover Bob
// Incident edge: Bob -> Alice (KNOWS)
const bobAttrs = graph.getNodeAttributes("bob");
const bobStyle = resolveNodeElementStyle(GRAPH_THEME, "structure", "hovered", bobAttrs, "Bob");
assert.equal(bobStyle.forceLabel, true);
assert.equal(bobStyle.label, "Bob");
const bobIncidentEdges = new Set(["edge_bob_alice"]);
const bobAliceHoverState = resolveEdgeVisualState(
"edge_bob_alice",
"bob",
"alice",
"structure",
"bob",
"",
"",
new Set(),
new Set(),
bobIncidentEdges,
);
assert.equal(bobAliceHoverState, "hovered");
});
test("edge selection maintains canonical edge type labels and active visual state", () => {
loadDeterministicTestGraph();
const edgeCases = [
{ id: "edge_alice_acme", source: "alice", target: "acme", label: "WORKS_AT" },
{ id: "edge_bob_alice", source: "bob", target: "alice", label: "KNOWS" },
{ id: "edge_acme_new_york", source: "acme", target: "new_york", label: "LOCATED_IN" },
];
for (const { id, source, target, label } of edgeCases) {
const attrs = graph.getEdgeAttributes(id);
assert.equal(attrs.edgeType, label);
const visualState = resolveEdgeVisualState(
id,
source,
target,
"inspection",
null,
"",
id, // selected edge
new Set(),
new Set(),
);
assert.equal(visualState, "selected", `Selected edge ${id} must resolve to 'selected' state`);
const style = resolveEdgeElementStyle(
GRAPH_THEME,
"inspection",
"selected",
attrs,
source,
target,
"full",
id,
"selected",
);
assert.equal(style.hidden, false, `Selected edge ${id} (${label}) must not be hidden`);
assert.ok(
style.size !== undefined && style.size > 0,
`Selected edge ${id} (${label}) must have positive render size`,
);
}
});
test("node labels remain forced visible during hover, selection, and inspection zoom tier", () => {
loadDeterministicTestGraph();
const nodes = ["alice", "bob", "acme", "new_york"];
for (const nid of nodes) {
const attrs = graph.getNodeAttributes(nid);
// Hover state forces label visibility
const hoverForcesLabel = shouldForceNodeLabel(GRAPH_THEME, "structure", "hovered", attrs, 0);
assert.equal(hoverForcesLabel, true, `Node ${nid} label must force visible on hover`);
// Selected state forces label visibility
const selectForcesLabel = shouldForceNodeLabel(GRAPH_THEME, "structure", "selected", attrs, 0);
assert.equal(selectForcesLabel, true, `Node ${nid} label must force visible on selection`);
// Resolved style emits actual string label
const style = resolveNodeElementStyle(GRAPH_THEME, "inspection", "hovered", attrs, attrs.label);
assert.equal(style.forceLabel, true);
assert.equal(style.label, attrs.label);
}
});
@@ -407,6 +407,31 @@ test("resolveEdgeElementStyle applies full-graph LOD to directional background e
assert.equal(style.hidden, true);
});
test("resolveEdgeElementStyle keeps small-graph relationships visible in overview", () => {
const style = resolveEdgeElementStyle(
GRAPH_THEME,
"overview",
"inactive",
{
edgeType: "related_to",
weight: 1,
properties: {},
edgeVariant: "directional",
visualPriority: 0.1,
baseSize: 0.5,
isSmallGraph: true,
},
"source",
"target",
"full",
"small-graph-low-priority",
"hidden",
);
assert.equal(style.hidden, false);
assert.ok(Number(style.size ?? 0) >= 0.9);
});
test("classifyFullGraphEdge applies deterministic priority order", () => {
const edgeClass = classifyFullGraphEdge(
"edge-priority",
@@ -0,0 +1,31 @@
import assert from "node:assert/strict";
import test from "node:test";
import { buildRealtimeEdgeAttributes } from "../src/workspaces/GraphWorkspace/realtimeGraphAttributes.ts";
const payload = {
id: "edge-live",
source_id: "source",
target_id: "target",
type: "related_to",
properties: {},
};
test("realtime edges retain the active small-graph visibility marker", () => {
const attributes = buildRealtimeEdgeAttributes(payload, {
isBidirectional: false,
isSmallGraph: true,
});
assert.equal(attributes.isSmallGraph, true);
assert.equal(attributes.edgeVariant, "directional");
});
test("realtime edges do not retain the marker after graph leaves small-graph mode", () => {
const attributes = buildRealtimeEdgeAttributes(payload, {
isBidirectional: false,
isSmallGraph: false,
});
assert.equal(attributes.isSmallGraph, false);
});
+100
View File
@@ -0,0 +1,100 @@
import assert from "node:assert/strict";
import test from "node:test";
import {
SMALL_GRAPH_MAX_NODES,
buildSmallGraphSeedPositions,
resolveGraphLayoutDecision,
resolveNodeLayoutPosition,
shouldUseSmallGraphLayout,
} from "../src/workspaces/GraphWorkspace/smallGraphLayout.ts";
test("small graph layout is selected only when coordinates are not already usable", () => {
assert.equal(shouldUseSmallGraphLayout(12, 0), true);
assert.equal(shouldUseSmallGraphLayout(SMALL_GRAPH_MAX_NODES + 1, 0), false);
assert.equal(shouldUseSmallGraphLayout(12, 0.95), false);
});
test("small graph layout ignores isolated partial coordinates", () => {
const decision = resolveGraphLayoutDecision(12, 1 / 12);
assert.deepEqual(
resolveNodeLayoutPosition(decision, { x: 50_000, y: -50_000 }, { x: 24, y: -18 }),
{ x: 24, y: -18 },
);
assert.deepEqual(
resolveNodeLayoutPosition(decision, { x: 50_000, y: null }, { x: -12, y: 36 }),
{ x: -12, y: 36 },
);
});
test("small graph load is immediately ready and skips runtime stabilization", () => {
assert.deepEqual(resolveGraphLayoutDecision(12, 0), {
useProvidedCoordinates: false,
useSmallGraphLayout: true,
layoutReady: true,
});
assert.deepEqual(resolveGraphLayoutDecision(SMALL_GRAPH_MAX_NODES + 1, 0), {
useProvidedCoordinates: false,
useSmallGraphLayout: false,
layoutReady: false,
});
assert.deepEqual(resolveGraphLayoutDecision(12, 1), {
useProvidedCoordinates: true,
useSmallGraphLayout: false,
layoutReady: true,
});
});
test("small graph layout is deterministic and keeps connected nodes together", () => {
const nodes = ["Apple", "Steve", "Ronald", "Cupertino", "California"];
const edges = [
{ source: "Apple", target: "Steve" },
{ source: "Ronald", target: "Cupertino" },
];
const first = buildSmallGraphSeedPositions(nodes, edges);
const second = buildSmallGraphSeedPositions([...nodes].reverse(), [...edges].reverse());
assert.deepEqual([...first.entries()].sort(), [...second.entries()].sort());
assert.equal(first.size, nodes.length);
const distance = (left: string, right: string) => {
const a = first.get(left);
const b = first.get(right);
assert.ok(a && b);
return Math.hypot(a.x - b.x, a.y - b.y);
};
assert.ok(distance("Apple", "Steve") < distance("Apple", "California"));
assert.ok(distance("Ronald", "Cupertino") < distance("Ronald", "California"));
});
test("small graph layout keeps maximum-radius components separated", () => {
const componentCount = 4;
const nodesPerComponent = 12;
const nodes = Array.from(
{ length: componentCount * nodesPerComponent },
(_, index) => `component-${Math.floor(index / nodesPerComponent)}-node-${index % nodesPerComponent}`,
);
const edges = Array.from({ length: componentCount }).flatMap((_, componentIndex) => {
const prefix = `component-${componentIndex}-node-`;
return Array.from({ length: nodesPerComponent - 1 }, (_unused, nodeIndex) => ({
source: `${prefix}${nodeIndex}`,
target: `${prefix}${nodeIndex + 1}`,
}));
});
const positions = buildSmallGraphSeedPositions(nodes, edges);
for (let leftComponent = 0; leftComponent < componentCount; leftComponent += 1) {
for (let rightComponent = leftComponent + 1; rightComponent < componentCount; rightComponent += 1) {
let closestDistance = Number.POSITIVE_INFINITY;
for (let leftNode = 0; leftNode < nodesPerComponent; leftNode += 1) {
for (let rightNode = 0; rightNode < nodesPerComponent; rightNode += 1) {
const left = positions.get(`component-${leftComponent}-node-${leftNode}`);
const right = positions.get(`component-${rightComponent}-node-${rightNode}`);
assert.ok(left && right);
closestDistance = Math.min(closestDistance, Math.hypot(left.x - right.x, left.y - right.y));
}
}
assert.ok(closestDistance >= 48, `components are only ${closestDistance} units apart`);
}
}
});
@@ -0,0 +1,150 @@
import test from "node:test";
import assert from "node:assert/strict";
import { createTemporalSnapshotGuards } from "../src/workspaces/GraphWorkspace/temporalSnapshotGuards.ts";
const POSITION_1 = new Date("2023-07-02T00:00:00Z").getTime();
const POSITION_2 = new Date("2024-01-02T00:00:00Z").getTime();
const POSITION_3 = new Date("2024-07-02T00:00:00Z").getTime();
const SNAPSHOT = { active_node_ids: ["n1", "n2"], active_node_count: 2 };
// ── begin: one request per scrubber position ─────────────────────────────────
test("begin: a new position returns a fresh request sequence", () => {
const guards = createTemporalSnapshotGuards();
assert.deepEqual(guards.begin(POSITION_1), { seq: 1, cached: null });
});
test("begin: an identical in-flight request is deduplicated (no duplicate fetch)", () => {
const guards = createTemporalSnapshotGuards();
guards.begin(POSITION_1);
assert.deepEqual(guards.begin(POSITION_1), { seq: null, cached: null });
});
test("begin: distinct positions request independently", () => {
const guards = createTemporalSnapshotGuards();
assert.equal(guards.begin(POSITION_1).seq, 1);
assert.equal(guards.begin(POSITION_2).seq, 2);
});
test("begin: revisiting an applied position returns its cached snapshot", () => {
const guards = createTemporalSnapshotGuards();
const { seq } = guards.begin(POSITION_1);
guards.apply(POSITION_1, seq, SNAPSHOT);
const revisit = guards.begin(POSITION_1);
assert.equal(revisit.seq, 2);
assert.deepEqual(revisit.cached, SNAPSHOT);
});
test("begin: a failed position (finished) can be requested again", () => {
const guards = createTemporalSnapshotGuards();
const { seq } = guards.begin(POSITION_1);
guards.finish(POSITION_1, seq);
const retry = guards.begin(POSITION_1);
assert.equal(retry.seq, 2);
assert.equal(retry.cached, null);
});
test("finish: does not clear a position whose snapshot was already applied", () => {
const guards = createTemporalSnapshotGuards();
const { seq } = guards.begin(POSITION_1);
guards.apply(POSITION_1, seq, SNAPSHOT);
guards.finish(POSITION_1, seq);
assert.deepEqual(guards.begin(POSITION_1).cached, SNAPSHOT);
});
test("finish: a stale sequence cannot release a newer request's position", () => {
const guards = createTemporalSnapshotGuards();
const first = guards.begin(POSITION_1);
guards.finish(POSITION_1, first.seq);
guards.begin(POSITION_1); // seq 2, in flight again
guards.finish(POSITION_1, first.seq); // stale seq: must not release seq 2
assert.deepEqual(guards.begin(POSITION_1), { seq: null, cached: null });
});
// ── shouldApply: applied only while the scrubber is on that position ─────────
test("shouldApply: the current position's response is applied", () => {
const guards = createTemporalSnapshotGuards();
const { seq } = guards.begin(POSITION_1);
assert.equal(guards.shouldApply(POSITION_1, seq), true);
});
test("shouldApply: a response for a position the scrubber left is discarded", () => {
const guards = createTemporalSnapshotGuards();
const { seq: seq1 } = guards.begin(POSITION_1);
guards.begin(POSITION_2);
assert.equal(guards.shouldApply(POSITION_1, seq1), false);
assert.equal(guards.shouldApply(POSITION_2, 2), true);
});
test("shouldApply: a late response for the position the scrubber returned to is applied", () => {
const guards = createTemporalSnapshotGuards();
const { seq: seq1 } = guards.begin(POSITION_1);
const { seq: seq2 } = guards.begin(POSITION_2);
guards.begin(POSITION_1); // back to 1: deduplicated, no new request
assert.equal(guards.shouldApply(POSITION_1, seq1), true);
assert.equal(guards.shouldApply(POSITION_2, seq2), false);
});
test("shouldApply: an unknown sequence is discarded", () => {
const guards = createTemporalSnapshotGuards();
guards.begin(POSITION_1);
assert.equal(guards.shouldApply(POSITION_1, 99), false);
});
test("shouldApply: after a reset no pre-reset response applies", () => {
const guards = createTemporalSnapshotGuards();
const { seq } = guards.begin(POSITION_1);
guards.reset();
assert.equal(guards.shouldApply(POSITION_1, seq), false);
});
// ── apply: caching for revisits ─────────────────────────────────────────────
test("apply: stores the snapshot so a revisit re-applies it without a request", () => {
const guards = createTemporalSnapshotGuards();
const { seq } = guards.begin(POSITION_1);
guards.apply(POSITION_1, seq, SNAPSHOT);
guards.begin(POSITION_2);
assert.deepEqual(guards.begin(POSITION_1).cached, SNAPSHOT);
});
test("apply: play wrap-around re-applies the wrapped-to position's snapshot", () => {
const guards = createTemporalSnapshotGuards();
const { seq } = guards.begin(POSITION_1);
guards.apply(POSITION_1, seq, SNAPSHOT);
guards.begin(POSITION_2);
guards.begin(POSITION_3);
const wrap = guards.begin(POSITION_1);
assert.deepEqual(wrap.cached, SNAPSHOT);
assert.equal(guards.shouldApply(POSITION_1, wrap.seq), true);
});
// ── reset: graph reload ─────────────────────────────────────────────────────
test("reset: clears requested and cached state so positions refetch", () => {
const guards = createTemporalSnapshotGuards();
const { seq } = guards.begin(POSITION_1);
guards.apply(POSITION_1, seq, SNAPSHOT);
guards.reset();
const fresh = guards.begin(POSITION_1);
assert.equal(fresh.seq, 1);
assert.equal(fresh.cached, null);
});
// ── cache bound ─────────────────────────────────────────────────────────────
test("cache: oldest positions are evicted when the cache is full", () => {
const guards = createTemporalSnapshotGuards();
const count = 300;
for (let i = 0; i < count; i++) {
const { seq } = guards.begin(POSITION_1 + i * 1000);
guards.apply(POSITION_1 + i * 1000, seq, SNAPSHOT);
}
const oldest = guards.begin(POSITION_1);
assert.equal(oldest.cached, null); // evicted: must refetch on revisit
const newest = guards.begin(POSITION_1 + (count - 1) * 1000);
assert.deepEqual(newest.cached, SNAPSHOT); // still cached
});
+33 -7
View File
@@ -8,8 +8,9 @@ Connects Claude Code, Cursor, Windsurf, Cline, Continue, VS Code (GitHub Copilot
## Quick start
```bash
# From the repo root
pip install -e ".[mcp]"
# From the repo root — no extra install flag needed; the root mcp/ package is
# part of the repository and does not require an external MCP SDK.
pip install -e .
# Test the server (type a JSON-RPC request, press Enter)
python -m mcp
@@ -89,7 +90,14 @@ python -m mcp [--debug]
## Per-tool configuration
### Claude Code (`~/.claude/settings.json`)
### Claude Code (`~/.claude.json` or `.mcp.json`)
Claude Code supports two MCP configuration scopes:
- **User scope**`~/.claude.json` applies across all projects for your user account.
- **Project scope**`.mcp.json` in your project root applies only to that project.
Both files use the same `mcpServers` structure:
```json
{
@@ -97,15 +105,33 @@ python -m mcp [--debug]
"semantica": {
"command": "python",
"args": ["-m", "mcp"],
"cwd": "/path/to/semantica"
"env": {
"PYTHONPATH": "/path/to/semantica"
}
}
}
}
```
Or use the plugin bundle:
> **Why `PYTHONPATH`?** The root `mcp/` package is intentionally not included in
> the installed wheel, so `python -m mcp` only works when the repository is on
> Python's import path. Setting `PYTHONPATH` here ensures this works regardless
> of the working directory Claude uses when it launches the server.
Or add it via the CLI (user scope):
```bash
claude mcp add semantica python -m mcp --cwd /path/to/semantica
claude mcp add --scope user semantica \
-e PYTHONPATH=/path/to/semantica \
-- python -m mcp
```
Or for project scope (omit `--scope user`):
```bash
claude mcp add semantica \
-e PYTHONPATH=/path/to/semantica \
-- python -m mcp
```
---
@@ -216,7 +242,7 @@ Add to your Q Developer MCP config:
| Variable | Default | Description |
|---|---|---|
| `SEMANTICA_KG_PATH` | *(in-memory)* | Path to persist/load the graph (JSON file) |
| `SEMANTICA_KG_PATH` | *(in-memory only)* | Path to a JSON file used to **load** the graph on startup and **persist** mutations (record decisions, add entities/relationships) back to disk after each change. When unset the graph lives in memory only and is lost when the server exits. |
---
+35 -7
View File
@@ -16,6 +16,13 @@ log = logging.getLogger("semantica.mcp.session")
_graph: Optional[Any] = None
# Tracks whether the last graph initialisation successfully loaded the
# configured SEMANTICA_KG_PATH file. When True (or no path was configured)
# mutation handlers are allowed to save. When False an existing file failed
# to load; saving would overwrite the original data with an empty graph, so
# persistence is blocked until the process is restarted with a readable file.
_load_ok: bool = True
def get_graph() -> Any:
"""
@@ -24,24 +31,45 @@ def get_graph() -> Any:
The graph is created with advanced_analytics=True so all centrality,
community-detection, and embedding features are available.
"""
global _graph
global _graph, _load_ok
if _graph is None:
from semantica.context import ContextGraph
_graph = ContextGraph(advanced_analytics=True)
_load_ok = True # default: safe to persist
kg_path = os.environ.get("SEMANTICA_KG_PATH", "").strip()
if kg_path and os.path.exists(kg_path):
try:
_graph.load(kg_path)
log.info("Graph loaded from %s", kg_path)
except Exception as exc:
log.warning("Could not load graph from %s: %s", kg_path, exc)
# Only attempt to load if the file has content. An empty file
# means the path was just created (e.g. a fresh tempfile) and
# should be treated as "start with empty graph" rather than a
# corrupt-file failure.
if os.path.getsize(kg_path) > 0:
try:
_graph.load_from_file(kg_path)
log.info("Graph loaded from %s", kg_path)
except Exception as exc:
log.warning(
"Could not load graph from %s: %s — persistence disabled "
"to protect existing data; restart the server to retry.",
kg_path, exc,
)
_load_ok = False # do not overwrite the original file
return _graph
def is_persistence_safe() -> bool:
"""Return True when it is safe to write mutations back to SEMANTICA_KG_PATH.
Returns False after a failed load so that mutation handlers do not
overwrite the original (possibly intact) file with a fresh empty graph.
"""
return _load_ok
def reset_graph() -> None:
"""Reset the singleton (mainly useful in tests)."""
global _graph
global _graph, _load_ok
_graph = None
_load_ok = True
+35 -1
View File
@@ -5,6 +5,7 @@ Decision intelligence tools — record, query, precedents, causal chain, impact.
from __future__ import annotations
import logging
import os
from mcp.schemas import (
ANALYZE_DECISION_IMPACT,
@@ -13,7 +14,7 @@ from mcp.schemas import (
QUERY_DECISIONS,
RECORD_DECISION,
)
from mcp.session import get_graph
from mcp.session import get_graph, is_persistence_safe
log = logging.getLogger("semantica.mcp.tools.decisions")
@@ -37,6 +38,39 @@ def handle_record_decision(args: dict) -> dict:
valid_from=args.get("valid_from"),
valid_until=args.get("valid_until"),
)
# Persist back to disk so the decision survives server restarts.
# Skip when the initial load failed to avoid overwriting original data.
kg_path = os.environ.get("SEMANTICA_KG_PATH", "").strip()
if kg_path:
if not is_persistence_safe():
# Roll back the in-memory mutation so the client-visible state
# matches the persisted state (neither is saved).
if hasattr(graph, "_decisions") and decision_id in graph._decisions:
del graph._decisions[decision_id]
if hasattr(graph, "_decision_index"):
cat = args.get("category", "")
if cat in graph._decision_index:
graph._decision_index[cat].discard(decision_id)
return {
"error": (
"Persistence blocked: the configured SEMANTICA_KG_PATH "
"could not be loaded at startup. Restart the server with "
"a readable graph file to re-enable persistence."
)
}
try:
graph.save_to_file(kg_path)
except Exception as save_exc:
# Atomic write failed. Roll back the in-memory mutation so the
# client-visible and persisted states remain consistent.
if hasattr(graph, "_decisions") and decision_id in graph._decisions:
del graph._decisions[decision_id]
if hasattr(graph, "_decision_index"):
cat = args.get("category", "")
if cat in graph._decision_index:
graph._decision_index[cat].discard(decision_id)
log.exception("save_to_file failed after record_decision; mutation rolled back")
return {"error": f"Mutation rolled back: could not persist graph: {save_exc}"}
return {
"decision_id": decision_id,
"status": "recorded",
+70 -1
View File
@@ -5,9 +5,10 @@ Graph tools — add entities/relationships, search, analytics, summary.
from __future__ import annotations
import logging
import os
from mcp.schemas import ADD_ENTITY, ADD_RELATIONSHIP, EMPTY, GET_ANALYTICS, SEARCH_GRAPH
from mcp.session import get_graph
from mcp.session import get_graph, is_persistence_safe
log = logging.getLogger("semantica.mcp.tools.graph")
@@ -25,6 +26,35 @@ def handle_add_entity(args: dict) -> dict:
node_type=args.get("type", "Entity"),
metadata=args.get("metadata", {}),
)
# Persist back to disk so the entity survives server restarts.
# Skip when the initial load failed to avoid overwriting original data.
kg_path = os.environ.get("SEMANTICA_KG_PATH", "").strip()
if kg_path:
if not is_persistence_safe():
# Roll back: remove the node we just added.
try:
with graph._lock:
graph._drop_node_from_indexes(node_id)
except Exception:
pass
return {
"error": (
"Persistence blocked: the configured SEMANTICA_KG_PATH "
"could not be loaded at startup. Restart the server with "
"a readable graph file to re-enable persistence."
)
}
try:
graph.save_to_file(kg_path)
except Exception as save_exc:
# Roll back: remove the node so in-memory and persisted state agree.
try:
with graph._lock:
graph._drop_node_from_indexes(node_id)
except Exception:
pass
log.exception("save_to_file failed after add_entity; mutation rolled back")
return {"error": f"Mutation rolled back: could not persist graph: {save_exc}"}
return {"status": "added", "id": node_id, "type": args.get("type", "Entity")}
except Exception as exc:
log.exception("add_entity failed")
@@ -46,6 +76,45 @@ def handle_add_relationship(args: dict) -> dict:
edge_type=rel_type,
metadata=args.get("metadata", {}),
)
# Persist back to disk so the relationship survives server restarts.
# Skip when the initial load failed to avoid overwriting original data.
kg_path = os.environ.get("SEMANTICA_KG_PATH", "").strip()
if kg_path:
if not is_persistence_safe():
# Roll back: remove the edge we just added (last matching edge).
try:
with graph._lock:
for edge in reversed(list(graph.edges)):
if (edge.source_id == source
and edge.target_id == target
and edge.edge_type == rel_type):
graph._drop_edge_from_indexes(edge)
break
except Exception:
pass
return {
"error": (
"Persistence blocked: the configured SEMANTICA_KG_PATH "
"could not be loaded at startup. Restart the server with "
"a readable graph file to re-enable persistence."
)
}
try:
graph.save_to_file(kg_path)
except Exception as save_exc:
# Roll back: remove the edge so in-memory and persisted state agree.
try:
with graph._lock:
for edge in reversed(list(graph.edges)):
if (edge.source_id == source
and edge.target_id == target
and edge.edge_type == rel_type):
graph._drop_edge_from_indexes(edge)
break
except Exception:
pass
log.exception("save_to_file failed after add_relationship; mutation rolled back")
return {"error": f"Mutation rolled back: could not persist graph: {save_exc}"}
return {"status": "added", "source": source, "target": target, "type": rel_type}
except Exception as exc:
log.exception("add_relationship failed")
+5 -1
View File
@@ -25,5 +25,9 @@
"mcp"
],
"skills": "./skills",
"agents": "./agents"
"agents": [
"./agents/decision-advisor.md",
"./agents/explainability.md",
"./agents/kg-assistant.md"
]
}
+12 -3
View File
@@ -49,7 +49,14 @@ dependencies = [
"scipy>=1.13.1",
"scikit-learn>=1.7.2",
"umap-learn>=0.5.12",
"spacy>=3.4.0",
# thinc (spacy's core dep) dropped Python 3.9 wheels at 8.3.10, and later
# spacy patch releases (3.8.8+) require thinc>=8.3.9-only-on-3.10+ ranges,
# which forces a source build that fails outright on 3.9 (see Install
# Matrix run history). Capping both keeps 3.9 on the last wheel-compatible
# pair; 3.10+ is left unconstrained to always get the latest spacy/thinc.
"spacy>=3.4.0,<3.8.8; python_version < '3.10'",
"spacy>=3.4.0; python_version >= '3.10'",
"thinc<8.3.5; python_version < '3.10'",
"transformers>=4.20.0",
"torch>=1.13.1",
"sentence-transformers>=2.2.0",
@@ -107,11 +114,12 @@ llm-gemini = ["google-genai>=0.1.0"]
llm-anthropic = ["anthropic>=0.122.0"]
llm-ollama = ["ollama>=0.1.0"]
llm-deepseek = ["openai>=1.0.0"]
llm-novita = ["openai>=1.0.0"]
llm-litellm = ["litellm>=1.83.9"]
llm-instructor = ["instructor>=1.15.3"]
llm-all = [
"semantica[llm-openai,llm-groq,llm-gemini,llm-anthropic,llm-ollama,llm-deepseek,llm-litellm,llm-instructor]"
"semantica[llm-openai,llm-groq,llm-gemini,llm-anthropic,llm-ollama,llm-deepseek,llm-novita,llm-litellm,llm-instructor]"
]
# ---- Document Parsing ----
@@ -124,12 +132,13 @@ shacl = ["pyshacl>=0.25.0"]
db-snowflake = ["snowflake-connector-python>=4.6.0", "cryptography>=49.0.0"]
db-databricks = ["databricks-sdk>=0.60.0", "databricks-sql-connector>=4.0.0"]
db-arrow = ["pyarrow>=24.0.0"]
db-salesforce = ["simple-salesforce>=1.12.0"]
ingest-parquet = ["pyarrow>=24.0.0"]
ingest-arrow = ["pyarrow>=24.0.0"]
ingest-sap = ["requests>=2.28.0"]
db-all = [
"semantica[db-snowflake,db-databricks,db-arrow]"
"semantica[db-snowflake,db-databricks,db-salesforce,db-arrow]"
]
# ---- Embedding / Models ----
+114 -9
View File
@@ -3714,19 +3714,61 @@ def store_stats(cli_ctx: CLIContext, backend: str, fmt: str, local_json: bool) -
_run_with_error_handling(_action)
_MIGRATE_SUPPORTED_BACKENDS = {"faiss", "sqlite", "pgvector"}
_MIGRATE_BATCH_SIZE = 500
def _migrate_backend_config(vs_cfg: Dict[str, Any], backend: str) -> Dict[str, Any]:
"""Resolve per-backend config out of the vector_store config section.
Supports both a per-backend nested shape (``vector_store.faiss.dimension``)
and the common flat single-backend shape (``vector_store.backend`` +
sibling keys), since either can appear depending on how many backends a
user has configured.
"""
nested = vs_cfg.get(backend)
if isinstance(nested, dict):
return dict(nested)
if vs_cfg.get("backend") == backend:
return {k: v for k, v in vs_cfg.items() if k != "backend"}
return {}
def _require_faiss_index_path(cfg: Dict[str, Any], role: str) -> str:
"""FAISS has no server to hold state between commands: a fresh FAISSStore
starts empty and nothing outside the process persists it, so migration
needs an explicit on-disk index to read from or write to."""
index_path = cfg.get("index_path")
if not index_path:
raise click.ClickException(
f"faiss as migration {role} requires 'index_path' in the vector_store "
f"config (vector_store.faiss.index_path or vector_store.index_path "
f"when faiss is the configured backend)."
)
return index_path
@store.command("migrate")
@click.option("--from", "from_backend", required=True)
@click.option("--to", "to_backend", required=True)
@click.option("--namespace", default=None)
@click.option("--dry-run", "local_dry", is_flag=True, default=False)
@click.option("--json", "local_json", is_flag=True, default=False)
@click.pass_obj
def store_migrate(cli_ctx: CLIContext, from_backend: str, to_backend: str,
namespace: Optional[str], local_dry: bool) -> None:
namespace: Optional[str], local_dry: bool, local_json: bool) -> None:
"""Migrate data between backends.
Direct migration is only wired up between faiss, sqlite, and pgvector -
these are the backends whose storage contract supports paging through
every stored vector. Migrating to or from qdrant, pinecone, milvus, or
weaviate still needs the export/reindex workaround below, since each of
those needs its own enumeration design (Qdrant scroll, Pinecone list,
etc.) that hasn't been built yet.
\b
Example:
semantica store migrate --from faiss --to qdrant --namespace production --dry-run
semantica store migrate --from faiss --to sqlite --namespace production --dry-run
"""
cli_ctx = _require_ctx(cli_ctx)
@@ -3734,13 +3776,76 @@ def store_migrate(cli_ctx: CLIContext, from_backend: str, to_backend: str,
if _is_dry(cli_ctx, local_dry):
_dry(cli_ctx, "migrate", from_backend=from_backend, to_backend=to_backend)
return
raise click.ClickException(
f"Direct backend migration ({from_backend}{to_backend}) is not yet supported "
"by the vector store layer. To migrate, export your data first:\n"
" semantica export --format parquet --output dump.parquet\n"
f" semantica embed index dump.parquet --store {to_backend}"
+ (f" --namespace {namespace}" if namespace else "")
)
if from_backend not in _MIGRATE_SUPPORTED_BACKENDS or to_backend not in _MIGRATE_SUPPORTED_BACKENDS:
raise click.ClickException(
f"Direct backend migration ({from_backend}{to_backend}) is only supported "
f"between {', '.join(sorted(_MIGRATE_SUPPORTED_BACKENDS))}. To migrate involving "
"another backend, export your data first:\n"
" semantica export --format parquet --output dump.parquet\n"
f" semantica embed index dump.parquet --store {to_backend}"
+ (f" --namespace {namespace}" if namespace else "")
)
from .vector_store import VectorStore
vs_cfg = cli_ctx.config.to_dict().get("vector_store", {}) or {}
source_cfg = _migrate_backend_config(vs_cfg, from_backend)
dest_cfg = _migrate_backend_config(vs_cfg, to_backend)
source_index_path = None
if from_backend == "faiss":
source_index_path = _require_faiss_index_path(source_cfg, "source")
dest_index_path = None
if to_backend == "faiss":
dest_index_path = _require_faiss_index_path(dest_cfg, "destination")
source = VectorStore(backend=from_backend, config=source_cfg)
if source_index_path:
source._backend_store.load_index(source_index_path)
source_dimension = getattr(source._backend_store, "dimension", None)
if source_dimension and "dimension" not in dest_cfg:
dest_cfg["dimension"] = source_dimension
dest = VectorStore(backend=to_backend, config=dest_cfg)
if dest_index_path and Path(dest_index_path).exists():
dest._backend_store.load_index(dest_index_path)
migrated = 0
vectors_batch: List[Any] = []
metadata_batch: List[Dict[str, Any]] = []
ids_batch: List[str] = []
def _flush() -> None:
nonlocal migrated
if not vectors_batch:
return
dest.store_vectors(list(vectors_batch), list(metadata_batch), ids=list(ids_batch))
migrated += len(vectors_batch)
vectors_batch.clear()
metadata_batch.clear()
ids_batch.clear()
for item in source.iter_vectors(batch_size=_MIGRATE_BATCH_SIZE):
meta = dict(item.get("metadata") or {})
if namespace and "namespace" not in meta:
meta["namespace"] = namespace
vectors_batch.append(item["vector"])
metadata_batch.append(meta)
ids_batch.append(item["id"])
if len(vectors_batch) >= _MIGRATE_BATCH_SIZE:
_flush()
_flush()
if dest_index_path and migrated:
dest._backend_store.save_index(dest_index_path)
result = {"from": from_backend, "to": to_backend, "migrated": migrated}
if _is_json(cli_ctx, local_json):
_jecho(result)
else:
_ok(cli_ctx, f"Migrated {migrated} vectors from {from_backend} to {to_backend}")
_run_with_error_handling(_action)
+4
View File
@@ -111,6 +111,7 @@ from .context_graph import ContextEdge, ContextGraph, ContextNode
from .context_retriever import ContextRetriever, RetrievedContext, TemporalGraphRetriever
from .decision_context import DecisionContext
from .entity_linker import EntityLink, EntityLinker, LinkedEntity
from .erasure import ErasureCoordinator, ErasureReceipt
# Decision tracking imports
from .decision_models import (
@@ -145,6 +146,9 @@ __all__ = [
"ContextRetriever",
"RetrievedContext",
"TemporalGraphRetriever",
# Cross-store erasure
"ErasureCoordinator",
"ErasureReceipt",
# Decision tracking models
"Decision",
"DecisionContextModel",
+31 -4
View File
@@ -626,6 +626,27 @@ class AgentMemory:
self.logger.debug(f"Deleted memory item: {memory_id}")
return True
def vector_ids_for(self, memory_id: str) -> List[str]:
"""Return the vector-store ids owned by a memory item.
Read-only view of the ids ``delete_memory()`` would remove for this
item, so a caller that needs to *report* on vector removal can delete
them itself rather than relying on ``delete_memory()``'s best-effort
cascade, which logs a vector-store failure and still returns ``True``.
Mirrors the fallback in ``delete_memory``: an item stored without
tracked vector ids is keyed in the vector store by its own memory id.
Args:
memory_id: Memory identifier.
Returns:
The item's vector ids, or ``[]`` if the item is unknown.
"""
if memory_id not in self.memory_items:
return []
return list(self._vector_ids.get(memory_id, [])) or [memory_id]
def clear_memory(self, **filters) -> int:
"""
Clear memory items matching filters.
@@ -1286,13 +1307,19 @@ class AgentMemory:
"""
return self.retrieve(content, max_results=limit, **kwargs)
def find_by_entity(self, entity_id: str, limit: int = 10) -> List[Dict[str, Any]]:
def find_by_entity(
self, entity_id: str, limit: Optional[int] = None
) -> List[Dict[str, Any]]:
"""
Find by entity.
Args:
entity_id: Entity ID to search for
limit: Maximum results (default: 10)
limit: Maximum results. None (the default) returns ALL matches.
The previous default of 10 silently truncated results an
erasure workflow computing "what references this entity"
from a truncated page would leave the remainder live
(#1018). Callers that want pagination pass an explicit limit.
Returns:
List of memory dicts containing the entity
@@ -1308,9 +1335,9 @@ class AgentMemory:
if mem_dict:
results.append(mem_dict)
break
if len(results) >= limit:
if limit is not None and len(results) >= limit:
break
return results[:limit]
return results if limit is None else results[:limit]
def find_by_relationship(
self, relationship_type: str, limit: int = 10
+29 -3
View File
@@ -1203,8 +1203,30 @@ class ContextGraph:
"links": links_data,
}
with open(path, "w", encoding="utf-8") as f:
json.dump(data, f, indent=2, ensure_ascii=False)
# Write atomically: serialize to a sibling temp file then replace the
# destination in one OS-level rename. This guarantees the destination
# is either the old contents or the new contents — never a partial write
# — so a crash or disk-full error during json.dump cannot corrupt the
# sole persisted copy of the graph.
dest = Path(path)
dest.parent.mkdir(parents=True, exist_ok=True)
fd, tmp_path = tempfile.mkstemp(
dir=dest.parent, prefix=".kg_tmp_", suffix=".json"
)
try:
with os.fdopen(fd, "w", encoding="utf-8") as f:
json.dump(data, f, indent=2, ensure_ascii=False)
f.flush()
os.fsync(f.fileno())
os.replace(tmp_path, dest)
except Exception:
# Clean up the temp file on any failure so we don't litter the
# directory with partial writes.
try:
os.unlink(tmp_path)
except OSError:
pass
raise
self.logger.info(f"Saved context graph to {path}")
@@ -2640,7 +2662,11 @@ class ContextGraph:
Scope is this graph only. Copies held elsewhere (``AgentMemory``, a
bound vector store, an exported file) are not reached, so this is one
step of an erasure workflow, not the whole of it.
step of an erasure workflow, not the whole of it. Callers who need the
whole workflow -- and a receipt recording which stores it actually
reached -- should drive this through
:class:`~semantica.context.erasure.ErasureCoordinator` rather than
treating a ``True`` here as proof the content is gone.
Args:
node_id: Node to purge.
+76
View File
@@ -239,6 +239,82 @@ print(f"Python importance score: {importance.get('degree', 0)}")
---
## 🧹 Erasing an Entity Everywhere - ErasureCoordinator
`purge_node()` removes an entity from **one graph**. The same content can still be
sitting in agent memory and in your vector store, so purge on its own is one step
of an erasure workflow rather than the whole of it.
`ErasureCoordinator` drives the whole cascade and hands you a receipt saying what
it actually managed to erase.
```python
from semantica.context import AgentMemory, ContextGraph, ErasureCoordinator
coordinator = ErasureCoordinator(graph=knowledge, memory=memory)
receipt = coordinator.erase_entity(
"customer-4471",
reason="GDPR Art. 17 request #882",
)
if receipt.complete:
print("Erased everywhere")
else:
print("Still holding data:", receipt.incomplete_stores)
```
### Always Check the Receipt
The receipt is the point of the feature — **do not treat the call itself as proof
the data is gone**. Each store reports one of five statuses:
| Status | Meaning |
|---|---|
| `erased` | Reached, data removed (on the vectors leg: the store accepted the delete for the ids given) |
| `not_found` | Reached, held nothing for this entity |
| `not_configured` | No such store was bound — normal, not a failure |
| `unsupported` | The store cannot delete at all; retrying will not help |
| `failed` | The store was reached and the deletion did not succeed |
```python
receipt.to_dict()
# {
# "entity_id": "customer-4471",
# "reason": "GDPR Art. 17 request #882",
# "erased_at": "2026-08-16T09:03:36.813220",
# "complete": False,
# "stores": {
# "vectors": {"status": "unsupported", "backend": "faiss",
# "detail": "backend exposes no delete()/delete_vectors(); ..."},
# "memory": {"status": "erased", "items": 14},
# "graph": {"status": "erased", "nodes": 1, "edges": 3},
# },
# }
```
`complete` is `False` when any store reports `unsupported` or `failed`, which is
your signal to handle that store out of band. FAISS, Milvus and Weaviate expose
no delete method today, so erasure genuinely cannot be completed on them — the
coordinator says so rather than reporting a success it did not achieve.
### Good to Know
- **Order is vectors → memory → graph.** The graph tombstone is the durable record
that an erasure happened, so it is written last: a crash mid-cascade leaves the
node present and the receipt incomplete, rather than a tombstone claiming more
than actually happened.
- **A failing store does not abort the rest.** Partial failure is recorded in the
receipt and the remaining stores are still erased.
- **Every store is optional.** `ErasureCoordinator(graph=graph)` is fine; the other
legs report `not_configured`.
- **It is idempotent.** Erasing the same entity twice returns a receipt saying
there was nothing left to do, rather than raising.
- **Batch:** `coordinator.erase_entities([...], reason=...)` returns one receipt per
entity, in order, so one entity's failure does not stop the others.
---
## 🔄 Using Both Together - The Complete Setup
### Your Smart Agent System
+15 -9
View File
@@ -76,11 +76,11 @@ Production Use Cases:
- Insurance: Claim decisions, underwriting assessments
"""
from dataclasses import dataclass, field
from datetime import datetime
from typing import Any, Dict, List, Optional
import json
import uuid
from dataclasses import InitVar, dataclass, field
from datetime import datetime
from typing import Any, Dict, List, Optional
@dataclass
@@ -100,8 +100,9 @@ class Decision:
valid_from: Optional[str] = None
valid_until: Optional[str] = None
metadata: Dict[str, Any] = field(default_factory=dict)
auto_generate_id: InitVar[bool] = True
def __post_init__(self, auto_generate_id: bool = True):
def __post_init__(self, auto_generate_id: bool) -> None:
"""Validate decision data."""
if auto_generate_id and not self.decision_id: # Handle both None and empty string
self.decision_id = str(uuid.uuid4())
@@ -146,8 +147,9 @@ class DecisionContext:
risk_factors: List[str]
cross_system_inputs: Dict[str, Any] = field(default_factory=dict)
metadata: Dict[str, Any] = field(default_factory=dict)
auto_generate_id: InitVar[bool] = True
def __post_init__(self, auto_generate_id: bool = True):
def __post_init__(self, auto_generate_id: bool) -> None:
"""Validate decision context data."""
if auto_generate_id and not self.context_id: # Handle both None and empty string
self.context_id = str(uuid.uuid4())
@@ -184,8 +186,9 @@ class Policy:
created_at: datetime
updated_at: datetime
metadata: Dict[str, Any] = field(default_factory=dict)
auto_generate_id: InitVar[bool] = True
def __post_init__(self, auto_generate_id: bool = True):
def __post_init__(self, auto_generate_id: bool) -> None:
"""Validate policy data."""
if auto_generate_id and not self.policy_id: # Handle both None and empty string
self.policy_id = str(uuid.uuid4())
@@ -227,8 +230,9 @@ class PolicyException:
approval_timestamp: datetime
justification: str
metadata: Dict[str, Any] = field(default_factory=dict)
auto_generate_id: InitVar[bool] = True
def __post_init__(self, auto_generate_id: bool = True):
def __post_init__(self, auto_generate_id: bool) -> None:
"""Validate policy exception data."""
if auto_generate_id and not self.exception_id: # Handle both None and empty string
self.exception_id = str(uuid.uuid4())
@@ -265,8 +269,9 @@ class Precedent:
similarity_score: float
relationship_type: str # "similar_scenario", "same_policy", "exception_precedent"
metadata: Dict[str, Any] = field(default_factory=dict)
auto_generate_id: InitVar[bool] = True
def __post_init__(self, auto_generate_id: bool = True):
def __post_init__(self, auto_generate_id: bool) -> None:
"""Validate precedent data."""
if auto_generate_id and not self.precedent_id: # Handle both None and empty string
self.precedent_id = str(uuid.uuid4())
@@ -305,8 +310,9 @@ class ApprovalChain:
approval_context: str
timestamp: datetime
metadata: Dict[str, Any] = field(default_factory=dict)
auto_generate_id: InitVar[bool] = True
def __post_init__(self, auto_generate_id: bool = True):
def __post_init__(self, auto_generate_id: bool) -> None:
"""Validate approval chain data."""
if auto_generate_id and not self.approval_id: # Handle both None and empty string
self.approval_id = str(uuid.uuid4())

Some files were not shown because too many files have changed in this diff Show More