* fix(pipeline): resolve broken import and missing run() in PipelineWithProvenance
Fix two bugs in pipeline_provenance.py:
1. Wrong import path: `from .pipeline import Pipeline` fails because
`semantica/pipeline/pipeline.py` does not exist. Pipeline lives in
`pipeline_builder.py`. Fixed to `from .pipeline_builder import Pipeline`.
2. Pipeline dataclass has no run() method. PipelineWithProvenance.run()
now delegates to ExecutionEngine.execute_pipeline(), which is the
intended execution path for built pipelines.
Additional changes:
- Constructor now accepts a built Pipeline instance (breaking the previous
unusable API that tried to instantiate a dataclass with **config).
- Replace deprecated datetime.utcnow() with datetime.now(timezone.utc).
- Add test suite covering import, instantiation, execution, attribute
delegation, and provenance graceful degradation.
Fixes#858
* test: address Qodo review findings
- Remove redundant test_import_succeeds (module-level import already
guards against import regression at collection time).
- Fix test_provenance_disabled_when_import_fails to deterministically
simulate ImportError via sys.modules patch and assert provenance is
actually toggled off (runner.provenance is False).
* fix(pipeline): update provenance callers for Pipeline API
---------
Co-authored-by: Sameer Kadam <sskadam6305@gmail.com>
Co-authored-by: Russell Jurney <russell.jurney@gmail.com>
CORSMiddleware doesn't cover WebSocket handshakes at all (Starlette's CORS
support only wraps HTTP), so under SEMANTICA_ALLOW_ANONYMOUS=true --
the mode docker-compose.dev.yml ships -- is_valid_api_key's anonymous
bypass accepted a /ws/graph-updates connection from any origin. Loopback
binding isn't a boundary against a browser: any page the operator has
open can still reach ws://localhost:8000/ws/graph-updates directly, and
ConnectionManager.broadcast sends every graph_mutation to every
connected socket with no per-connection scoping. Combined with
/api/import accepting multipart/form-data (a CORS-safelisted content
type that skips preflight), a hostile page could write to the graph
over REST and read the result back over the unauthenticated WebSocket
-- demonstrated end-to-end in the report with a real client.
Not affected: any deployment with SEMANTICA_API_KEY configured -- the
handshake already rejects without a valid key in that mode. This is an
anonymous-mode-only, development-configuration exposure.
Fix: check the handshake's Origin header against
app.state.explorer_settings['allowed_origins'], the same list
CORSMiddleware already enforces for HTTP, before the key check. A
missing Origin (native/CLI clients, which never set the header --
only browsers do) is still allowed through, since the browser is the
only threat this closes.
4 new tests in test_explorer_auth.py: hostile Origin rejected under
anonymous mode; hostile Origin rejected even with a correct key
(Origin is checked before the key, so a leaked key alone can't
hijack the socket); an allowlisted Origin still connects under
anonymous mode; a missing Origin still connects under anonymous mode
(native clients keep working). Full explorer suite: 226 passed.
Co-authored-by: Sameer Kadam <sskadam6305@gmail.com>
Qodo's re-review confirmed the multi-IP fallback fix but kept the proxy
finding open: logging-and-falling-back when a proxy applies still let
the DNS-pinning protection be silently skipped under proxy
configuration, rather than enforcing a clear policy either way.
Implemented Qodo's preferred option: proxies are now disabled outright
for this SSRF-sensitive fetcher via session.trust_env = False, so
HTTP_PROXY/HTTPS_PROXY/NO_PROXY env vars are never consulted in the
first place (a configured proxy would perform its own DNS resolution
of the target host outside this process's control, reopening the
DNS check-then-use race pinning exists to close). The adapter also
keeps a fail-closed backstop: if a proxy is somehow still configured
despite trust_env=False (e.g. set explicitly by future code), it now
raises a clear 502 instead of silently connecting through the proxy
unpinned.
_validate_fetch_url's destination classification (blocking private/
internal targets) is unaffected either way — it runs before any of
this and doesn't depend on proxy configuration.
4 new tests: trust_env is disabled on every pinned session; an
HTTP_PROXY env var pointed at an address that would fail if contacted
is confirmed genuinely unused (real local-server fetch still succeeds
directly); and the fail-closed backstop actually raises when a proxy
is forced onto the session. Full explorer + triplet_store suite: 572
passed.
Four findings from PR #916's automated review, all addressed:
- CodeQL (HIGH): the test HTTPS server's SSLContext allowed TLSv1/TLSv1.1
by not setting a minimum version. Added
ssl_ctx.minimum_version = ssl.TLSVersion.TLSv1_2.
- github-code-quality: unused `cryptography` local in
_make_self_signed_cert — importorskip's return value was never used.
- Qodo (reliability): _validate_fetch_url() only returned the first
validated IP, and _make_pinned_session() pinned to just that one
address, so a fetch would fail outright if the first-returned A/AAAA
record happened to be unreachable even though a later one would work.
_validate_fetch_url() now returns every validated IP (deduplicated,
in resolution order); _make_pinned_session() takes the full list and
falls back through each one via a custom Connection._new_conn
override, matching the fallback behavior a normal DNS-resolving
connection would already get for free. Verified with a real test:
pin to an unreachable loopback address followed by a real one, confirm
the fetch still succeeds by falling back; and a real test confirming
it still raises (rather than silently re-resolving the hostname) when
every pinned address is unreachable.
- Qodo (security): when an HTTP(S) proxy applies, the adapter falls back
to the unpinned path rather than pinning. This is a real, but
architecturally unavoidable, limitation from the client side: for a
forward proxy, the *proxy* performs its own DNS resolution of the
target host on the application's behalf, a resolution this process
has no visibility into or control over — there's no client-side pin
that closes that race. _validate_fetch_url's destination
classification still fully applies either way; only the secondary
DNS-pinning hardening doesn't extend through a proxy. Added an info
log when this fallback path is taken so it's observable rather than
silent, and expanded the code comment to make the reasoning explicit
for the next reader/reviewer rather than looking like an oversight.
Tests: 3 new tests in test_ontology_dns_pinning.py (multi-IP fallback
success, all-unreachable failure, deduplicated multi-record resolution).
Full explorer + triplet_store suite: 569 passed.
Two follow-up hardening items flagged as secondary/deferred during
GHSA-8c7v-62gr-hj6g and GHSA-8vgg-8mr4-r236's fixes:
1. DNS check-then-use (TOCTOU) window in the ontology URL fetcher.
_validate_fetch_url() resolved and validated a hostname once, but
_fetch_url_sync() then let requests resolve the same hostname again
independently at connect time — a low-TTL or rebinding DNS answer
could differ between the two lookups, reopening the SSRF window the
validation exists to close.
_validate_fetch_url() now returns the validated IP, and a new
_make_pinned_session() builds a per-hop requests.Session whose
connection pool is pinned directly to that IP (bypassing DNS
resolution for the connection entirely), while explicitly restoring
the real hostname as the outgoing HTTP Host header and, for HTTPS,
the TLS SNI server_hostname/assert_hostname — so the connection
reaches the validated IP but still presents (and is verified
against) the real hostname's identity, keeping virtual hosting and
certificate validation correct.
Note: an earlier version of this fix set `_dns_host` post-construction
assuming it was decoupled from `host`, matching some other urllib3
releases; in the installed version (2.7.0), `host` is a property
that reads/writes `_dns_host` directly, so that approach silently
changed the Host header too. Verified with a real (non-mocked) local
HTTP server, a real local HTTPS server with a self-signed cert
(proving SNI/cert-hostname verification checks the real hostname,
not the pinned IP), and a negative control confirming a hostname/cert
mismatch is still correctly rejected — not silently bypassed.
2. Pre-wrapped object IRIs skipped full validation in
_format_object_for_sparql/_format_object_for_ntriples (Blazegraph,
RDF4J). A triplet object already wrapped in `<...>` only had its
inner content checked for a literal space or `>`, not run through
sparql_escaping.validate_uri() like the unwrapped-object branch —
flagged by automated review during GHSA-8vgg-8mr4-r236's fix. Both
branches now validate identically.
Tests: tests/explorer/test_ontology_dns_pinning.py (6 tests, including
2 real local-server end-to-end checks and 2 real-TLS checks with a
generated self-signed cert, gracefully skipped if `cryptography` isn't
installed); updated tests/explorer/test_ontology_ssrf.py for the new
per-hop session construction; 4 new tests in
tests/triplet_store/test_sparql_injection.py for the object-IRI fix.
Full explorer + triplet_store suite: 566 passed.
Two follow-up fixes to the initial ReDoS patch (CodeQL py/polynomial-redos
#1897), raised during code review:
--- Fix 1: _PREFIX_DECL regression — inline prologues and CRLF (#review-1) ---
The first ReDoS fix replaced the ambiguous trailing \s* with [ \t]*(?:\n|$),
but that introduced a behavioral regression:
* Inline prologues — PREFIX ex: <...> SELECT ... on a single line were no
longer stripped because the mandatory (?:\n|$) anchor never matched when
non-whitespace content followed the IRI on the same line.
* CRLF line endings — PREFIX ex: <...>\r\n failed because \r is not in
[ \t]* and the anchor expected a bare \n.
Root cause: the end-of-line anchor was unnecessary; the only thing needed
to eliminate backtracking ambiguity is ensuring the IRI body character class
and the trailing whitespace quantifier are disjoint.
Fix: change the IRI body from <[^>]*> to <[^>\r\n]*>, which:
- excludes CR and LF from the IRI match (semantically correct — SPARQL
IRIs cannot span line boundaries)
- makes [^>\r\n]* and the trailing [ \t]* have zero character overlap,
eliminating all backtracking ambiguity without any end-of-line anchor
No anchor is used, so both inline prologues and CRLF/LF endings work
naturally. ReDoS payloads (base< + !< x 10,000) still complete in <1 ms.
--- Fix 2: oversized-query length guard obscured error (#review-2) ---
The initial patch placed the _SPARQL_MAX_QUERY_LEN guard inside
_is_read_only_query(), which caused execute_sparql() to return the same
generic 'Only SELECT' error for both genuinely disallowed query types and
oversized inputs. Clients could not distinguish the two rejection reasons.
Fix: move the length check out of _is_read_only_query() and into
execute_sparql() as an explicit early gate, alongside the other resource
limits (_SPARQL_MAX_ROWS, _SPARQL_MAX_GRAPH_NODES). Oversized queries now
return a specific message naming the limit, the received length, and the
remediation step. _is_read_only_query() is documented to be length-agnostic.
_SPARQL_MAX_QUERY_LEN is relocated to the resource-limits block with the
other constants.
--- Tests added ---
tests/test_security_regression.py:
- test_inline_prefix_before_select_allowed (Fix 1 regression)
- test_crlf_line_endings_with_prefix (Fix 1 regression)
- test_crlf_multiple_prefixes_then_select (Fix 1 regression)
- test_inline_prefix_before_insert_still_blocked (Fix 1 security check)
- test_long_valid_query_not_rejected_by_is_read_only (Fix 2 separation)
tests/explorer/test_sparql_route.py:
- test_oversized_query_returns_distinct_length_error (Fix 2 error message)
- test_oversized_query_never_touches_the_graph (Fix 2 short-circuit)
- test_query_exactly_at_length_limit_is_accepted (Fix 2 boundary)
All 82 tests pass.
The _PREFIX_DECL pattern used \s* as a trailing quantifier after
<[^>]*>. On inputs that start with ase< but contain no closing >
(e.g. ase<!<<!<<!<...), the regex engine explores exponentially many
ways to split the match between [^>]* and \s*, causing polynomial
backtracking against user-controlled SPARQL query input.
Fix:
- Replace ^\s* / \s+ / \s* with ^[ \t]* / [ \t]+ / [ \t]*
so the leading/internal whitespace quantifiers only match horizontal
whitespace (no overlap with the <[^>]*> IRI part).
- Replace the ambiguous trailing \s* with [ \t]*(?:\n|$), which
matches only horizontal whitespace followed by a hard line boundary.
[^>]* and [ \t]* have disjoint character sets, eliminating the
backtracking ambiguity entirely.
- Add _SPARQL_MAX_QUERY_LEN = 10_000 guard at the top of
_is_read_only_query as defence-in-depth: rejects oversized input
before any regex work, bounding worst-case cost even if a future
pattern change reintroduces ambiguity.
Verified: ReDoS payload ase< + !< x 5000 completes in <1 ms.
Normal PREFIX/BASE stripping and read-only query detection unchanged.
Fixes: CodeQL py/polynomial-redos alert #1897
CWE: CWE-1333, CWE-730, CWE-400
* security: sanitize Cypher labels/relationship types/property keys (GHSA-482h-hw99-h62p)
Node labels and property keys passed to create_node/create_relationship
were interpolated directly into Cypher strings in the Neptune, Neo4j, and
FalkorDB graph stores. Property values are parameterized, but labels and
keys can't be bound as parameters, and nothing validated them, so a
document-derived entity type or property name could close the current
Cypher token early and append arbitrary statements (e.g. DETACH DELETE),
running with the application's database credentials.
- New shared semantica/graph_store/query_sanitize.py: sanitize_identifier()
generalizes age_store.py's existing _sanitize_label/_sanitize_rel_type
(the only backend that already validated this) into a helper the other
backends can import without an import cycle with graph_store.py/methods.py.
- Applied at every label/relationship-type/property-key interpolation site
in amazon_neptune.py, neo4j_store.py, falkordb_store.py, graph_store.py
(degree_centrality's own query builder), and methods.py
(update_relationship's own query builder) — create_node, create_nodes,
create_relationship, get_nodes, get_relationships, get_neighbors,
shortest_path, update_node, create_index, and all relationship-type
filters.
- depth/max_depth path-length parameters are also cast to int before
interpolation as defense-in-depth (they're already typed int, but
Python doesn't enforce that at runtime).
Added tests/graph_store/test_cypher_injection.py (12 tests covering the
sanitizer directly and reproducing the advisory's injection payload
against Neptune/Neo4j/FalkorDB create_node/create_relationship — asserts
the malicious query is never built or sent), plus regression tests for
graph_store.py's degree_centrality and methods.py's update_relationship.
Full graph_store test suite (224 tests) passes with no regressions.
* fix(graph-store): prevent depth-based Cypher injection
* test(graph-store): tighten injection regression assertions
* docs(changelog): add PR #910 (GHSA-482h Cypher injection) entry
---------
Co-authored-by: Sameer6305 <sskadam6305@gmail.com>
Two issues in the last round of commits:
1. explorer/auth.py added a new, opt-in APIKeyAuthMiddleware
(EXPLORER_API_KEY) and wired it into create_app(), but in doing so
removed the Depends(require_auth) dependency from every router and
deleted the /ws/graph-updates handshake check entirely. The new
middleware also fails OPEN (allows all requests) when its key is
unset, the opposite of require_auth's fail-closed design. Since
GHSA-j4mq-hprp-987v (the unauthenticated-Explorer-API advisory) is
already merged into main via require_auth, this would have reverted
a merged Critical fix the moment this branch merges. Removed
explorer/auth.py, restored the per-router dependencies and the
WebSocket auth check. Kept auth.py's one genuine improvement (adding
X-API-Key to the CORS allow_headers list) by folding it into the
existing CORS middleware config.
2. sparql.py's new _is_read_only_query() hardening (comment/PREFIX
stripping + forbidden-keyword scan) used `#[^\n]*` to strip SPARQL
comments, but a bare '#' also appears inside standard RDF namespace
IRIs (e.g. ".../1999/02/22-rdf-syntax-ns#") — the regex struck
everything after that '#' as a "comment", corrupting the query and
rejecting any legitimate SELECT using rdf:/rdfs:-style PREFIX
declarations. Confirmed by the fact the new hardening's own inlined
test copy failed against two of its own cases. Fixed by only
treating '#' as a comment-start at line-start or after whitespace,
which distinguishes ".../ns#" (preceded by a word character) from an
actual comment (preceded by whitespace/newline in every realistic
case, including the attacker's own comment-hiding PoC). Also fixed
the companion PREFIX/BASE regex, which required a prefix-name token
between the keyword and the IRI even for bare `BASE <...>`
declarations (which have none).
tests/test_security_regression.py's SPARQL section now imports the real
_is_read_only_query instead of maintaining a parallel inlined copy that
had silently drifted from — and shared the same bug as — the real
implementation; removed its TestAPIKeyAuth class (tested the now-deleted
auth.py) since equivalent, more thorough coverage already exists in
tests/explorer/test_explorer_auth.py. Updated tests/explorer/test_sparql_route.py's
multi-statement-injection test to reflect that the keyword scan now
catches "SELECT ... ; DROP ALL" itself rather than relying on rdflib's
parser, and added a new test confirming the parser still catches
multi-statement syntax that doesn't contain any forbidden keyword.
Full explorer/vector_store/security-regression/age_store suite: 543
passed (the only failures are 6 pre-existing, unrelated Pinecone-client
mocking issues).
The previous rework of the redirect loop closed the response on each
redirect hop but dropped the try/finally around the success path, so the
terminal response (the one actually read and returned) was left
unclosed, leaking the connection back to the pool unclosed under load.
Triplet.subject and Triplet.predicate (and, in some builders, .object)
were interpolated directly into SPARQL update/query strings in the
Blazegraph and RDF4J stores, and into a SELECT filter in the Jena store.
A subject containing '>' closes the '<...>' IRI token early, so the rest
of the value is parsed as more SPARQL. Entity names are document text in
the normal ingest pipeline, so anyone whose content gets processed could
append operations like CLEAR ALL, running with the application's store
credentials.
Applied the existing sparql_escaping.validate_uri (already used by
anzo_store.py, the one backend that was already hardened) at every
subject/predicate/object interpolation site:
- blazegraph_store.py: _build_insert_data, _triplets_to_rdf (unreachable
dead code today but same fix applied for consistency/future-proofing),
bulk_load's graph option, get_triplets's filter, delete_triplet.
- rdf4j_store.py: _triplets_to_ntriples, get_triplets's filter,
delete_triplet. (add_triplets's graph option was already validated.)
- jena_store.py: get_triplets's filter — the only vulnerable site;
add_triplets/delete_triplet already use rdflib's native Python API
(Graph.add/.remove with URIRef) rather than building query strings, so
they were never exploitable this way.
Added tests/triplet_store/test_sparql_injection.py (12 tests) reproducing
the advisory's own injection payload against all three backends' write
and read paths, asserting the malicious query is never built or sent.
Full triplet_store suite (330 tests) passes with no regressions.
Note: while adding read-path test coverage, found that jena_store.py's
get_triplets() WHERE-clause filter syntax is malformed SPARQL (missing a
FILTER()/separator before the equality conditions) — a pre-existing
correctness bug unrelated to this fix, worth a separate follow-up.
* security: require API-key auth on all Explorer API routes (GHSA-j4mq-hprp-987v)
Every Explorer route (bulk import/export, delete, LLM-backed ontology
generation, SPARQL, etc.) was mounted with no authentication, and both
server entrypoints bind 0.0.0.0 by default. Anyone reaching the port got
full read/write/delete on the graph.
- Add require_auth dependency (explorer/dependencies.py): checks
X-API-Key against SEMANTICA_API_KEY, fails closed with 503 if
unconfigured (not silently anonymous), 401 on wrong/missing key.
SEMANTICA_ALLOW_ANONYMOUS=true opts out explicitly for local dev.
- Wire dependencies=[Depends(require_auth)] into all 11 API routers in
both explorer/app.py and server.py. /health, /api/info, static assets,
and the SPA catch-all stay public.
- /ws/graph-updates handshake now checks the same key via header or
?api_key= query param (browsers can't set custom WS headers) before
accepting the connection.
- Default bind changed from 0.0.0.0 to 127.0.0.1 in server.py's main()
and cli.py's `server start`; the CLI warns if a non-loopback host is
passed explicitly without a key configured.
- Startup logging reports the resolved auth mode in both app factories.
- Document/generate SEMANTICA_API_KEY in the deploy recipes that expose
a public endpoint by default: docker-compose, Railway, Fly, Render.
Added tests/explorer/test_explorer_auth.py covering fail-closed default,
wrong/missing/correct key, anonymous opt-in, public-route exemptions, and
the WS handshake. Added tests/explorer/conftest.py defaulting the
pre-existing ~200 explorer tests to SEMANTICA_ALLOW_ANONYMOUS=true so
they keep exercising route logic without needing a key.
* fix CORS
---------
Co-authored-by: Zohaib Hassnain <109234410+ZohaibHassan16@users.noreply.github.com>
* fix(ingest): add SSRF protection for WebIngestor and RESTIngestor
Block non-http(s) schemes and private/loopback/link-local targets before
outbound requests, with allow_private_ips opt-in for trusted deployments.
* fix(ingest): fail closed on SSRF DNS resolution errors
* fix(ingest): validate SSRF targets on every HTTP redirect hop
* fix(ingest): avoid blocking on SSRF DNS executor shutdown
* fix(ingest): parse allow_private_ips without truthy-string pitfalls
* docs(ingest): clarify robots.txt SSRF/validation comment
---------
Co-authored-by: Pravit Ampapathini
- vector_store.save(): use v.tolist() instead of list(v) so numpy float32
vectors round-trip through JSON instead of raising TypeError.
- ontology._fetch_url_sync(): resolve relative Location headers via urljoin
before re-validating (previously any relative redirect was rejected
outright), and close every response instead of leaking the connection
across redirect hops.
- sparql.execute_sparql(): move _build_rdflib_graph inside the handler's
error handling so the graph-size cap returns a clean SparqlResponse
error instead of an unhandled 500.
- add regression tests for all three.
* test(deduplication): cover ClusterBuilder, MergeStrategyManager, and batch paths
Add focused coverage for union-find clustering, property merge rules,
merge_duplicates, embedding similarity, and incremental detection (#866).
* test(deduplication): assert unrelated clusters without conditional skip
Make cluster-separation coverage fail closed by using mocked pairs and
unconditional assertions for distinct Apple vs Microsoft cluster IDs.
* test(deduplication): tighten update_clusters attachment assertions
Require the incremental path to place the new near-duplicate in the
same rebuilt cluster instead of accepting a vacuous cluster-count check.
* test(deduplication): strengthen incremental detect_duplicates wrapper checks
Assert real DuplicateCandidate matches, score threshold, and new×existing
routing instead of only checking that the wrapper returns a list.
* test(deduplication): verify metadata provenance behavior
Assert that preserve_provenance writes metadata.provenance fields and add a disabled-path test so regressions do not pass through merge_entities metadata alone.
---------
Co-authored-by: Pravit Ampapathini <pravitampapathini@users.noreply.github.com>