* fix(security): sanitize node_id in Content-Disposition to prevent header injection (CWE-113)
* fix(security): cap link prediction at 10k nodes with semaphore to prevent OOM DoS (CWE-770)
* fix(security): sanitize imported node IDs to prevent stored header injection chain (CWE-20)
* test(security): add self-contained PoC runner with real measured output
* test(security): add regression tests for header injection, DoS cap, import sanitization
* fix(security): comprehensive fix for header injection, DoS, and import ID sanitization
* fix: move semaphore to wrap entire data-load+scoring region, use node-specific edge queries (Qodo #2, #3)
* fix: sanitize edge source/target IDs to match sanitized node IDs (Qodo #4)
* fix: scope 999_999 check to predict_links function via AST (Qodo #1)
* fix: add explicit None guard to _sanitize_import_node_id
* fix(security): close import-sanitizer bypass, enforce link-prediction cap before the expensive scan
Follow-up to the fixes in this PR, found in review:
- export_import.py's "properties" in raw_node fast path stored the id
verbatim, completely skipping _sanitize_import_node_id() -- a node
payload of {"id": "<crlf>", "properties": {}} (the shape this app's
own /api/export produces) bypassed the VULN-3 fix entirely. That
branch now sanitizes id before storing.
- The link-prediction 10k-node cap checked `total` only after calling
session.get_nodes()/get_edges(), which normalize the graph's entire
matching set before applying `limit` -- so the DoS guard ran after
the expensive work it exists to prevent had already happened, on
every request regardless of graph size. Added
GraphSession.get_raw_counts(), an O(1) check against the raw
len(graph.nodes)/len(graph.edges), and moved the size check ahead of
the normalizing calls (also added an edge-count cap).
- 5 of the existing regression tests asserted that literal words like
"Set-Cookie"/"Content-Type" disappear from the sanitized value -- the
sanitizer strips \r\n\x00"\ , not letters, so those assertions failed
against this PR's own fix as submitted. Corrected to assert on the
actual security property (no \r/\n survives), and added end-to-end
tests that exercise the real /api/import -> /api/provenance/report
route chain so the properties-key bypass has regression coverage.
Full explorer suite: 241 passed. tests/test_security_regression_pr2.py: 30 passed.
---------
Co-authored-by: Zohaib Hassnain <109234410+ZohaibHassan16@users.noreply.github.com>
Co-authored-by: Mohd Kaif <98801504+KaifAhmad1@users.noreply.github.com>
Co-authored-by: KaifAhmad1 <kaifahmad087@gmail.com>