Files
semantica/SECURITY.md
Mohd KaifandSameer6305 b59211ea7f security: SHA-pin all Actions, harden release pipeline, add pin verification (#824)
* security: SHA-pin all Actions, harden release pipeline, add pin verification

Hardens the CI/CD supply chain against the LiteLLM/Trivy-style attack (a
compromised third-party Action with a mutable tag stealing a long-lived
publishing token) and closes several related gaps found in an audit of the
actual repository state.

- Pin every third-party GitHub Action across all workflows to a full commit
  SHA (tag kept as a trailing comment); add verify-action-pins.yml, a CI
  check that confirms via the GitHub API that each pin still matches its
  tag, on every workflow change, push to main, and weekly.
- Scope release.yml permissions to the job level (workflow defaults to
  contents: read); add a concurrency group so simultaneous tag pushes can't
  race the publish job.
- Add SLSA build provenance attestation (actions/attest-build-provenance)
  for every released wheel.
- Fix a latent bug in security-scan.yml: the PR-comment step was missing
  pull-requests: write and silently failing; add bounded artifact retention
  for uploaded scan reports.
- Group github-actions Dependabot updates to cut review noise.
- Document the resulting posture in SECURITY.md for auditors/regulated
  adopters, including what's enforced and what a fork needs to reconfigure
  for itself (environment/branch protection, Trusted Publishing trust).

Also (via GitHub API, not in this diff): created a protected `pypi`
environment with a required reviewer restricted to v* tags, and enabled
branch protection on main (required review, required status checks, no
force-push/deletion).

* fix: harden verify-action-pins per PR #824 bot review

Addresses real findings from the automated review on #824:

- The script previously only matched uses: lines that already contained a
  40-hex SHA, so a newly added mutable-tag action (e.g. some/action@v1)
  would never be scanned at all and the check would pass silently. It now
  matches every uses: line and hard-fails on any ref that isn't a full
  commit SHA.
- A tag that fails to resolve via the GitHub API (rate limit, deleted tag)
  previously only logged a warning and continued; that's now a hard
  failure too, since an unverifiable pin is exactly the failure mode this
  check exists to catch.
- verify-action-pins.yml only triggered on .github/workflows/** changes,
  so an edit to the verifier script itself wouldn't run the check that
  verifies it. Added the script path to both trigger filters.

The reviewer's claim that slash-containing tag comments (release/v1) break
the API lookup did not reproduce - tested directly against
pypa/gh-action-pypi-publish@release/v1 and GitHub's commits API resolves
multi-segment refs natively - so no change was needed there.

Verified with a synthetic test workflow containing a mutable-tag action,
a correctly-pinned SHA, and a deliberately mismatched SHA: the updated
script now catches the first and third cases and passes the second. Also
re-ran against the real workflow tree (40/40 pins still verify clean).

* fix: repair broken Safety scan and PR comment formatting

The "Comment PR with Security Results" step was producing garbled output
(literal \n characters instead of newlines, "undefined:" labels) because:

- Every line in the JS comment builder used \n (escaped backslash-n)
  inside template literals, which JS renders as the literal two-character
  string \n, not a newline.
- The Semgrep section read issue.rule_id, but Semgrep's JSON field is
  check_id - hence "undefined: <path>" for every entry.

Rewrote the comment builder to construct each section as an array of
lines joined with a real '\n', with correct field names, and collapsed
long finding lists into a <details> block instead of a flat list.
Verified by extracting the exact script and running it under node against
synthetic fixtures matching each tool's real JSON schema (found/clean/
missing-report paths all render correctly).

While tracing the "undefined" and always-empty Safety section, found the
Safety step itself was silently broken:

- `safety check --json --output safety-report.json` is invalid in
  Safety 3.x: --output now selects a console format (json/text/screen),
  not a file path. The command errored on every run (swallowed by
  `|| true`), so safety-report.json was never created and the PR comment
  always fell back to a generic "scan completed" message. Switched to
  `--save-json`, which is the correct flag for writing a JSON report to
  disk, and confirmed against the real safety 3.8.1 CLI locally.
- Even with a report, the code read vuln.package - the real field is
  package_name.
- The job never installed Semantica's own dependencies before scanning,
  so `safety check` (which defaults to scanning the environment) was
  auditing the scanner tools' own dependencies, not Semantica's. Added
  `pip install -e ".[llm-litellm]"` so the project's actual dependency
  tree - including the LiteLLM extra this whole hardening effort is
  about - is what gets scanned.

Also updated the corresponding SECURITY.md bullet to describe what Safety
actually covers now.

* fix: remove unused pypdf2 dependency (CVE-2023-36464)

Now that the Safety scan step actually runs (see previous commit), it
correctly failed this PR's checks on CVE-2023-36464 in pypdf2==3.0.1 - a
real, pre-existing vulnerability that was invisible until the scan was
fixed.

PyPDF2 is not a patchable dependency here: the project is discontinued
(merged into `pypdf`), 3.0.1 is its final release, and there is no fixed
version to upgrade to. Grepping the repo for `import PyPDF2` / `from
PyPDF2` turns up nothing - it was never actually imported anywhere. Its
only presence outside pyproject.toml was in docstrings describing a
"PyPDF2.PdfReader() fallback" for PDF parsing that was never implemented
in code; pdfplumber is the library actually used. Removed the dependency
and corrected the stale docstrings in parse/__init__.py, parse/methods.py,
parse/pdf_parser.py, and ingest/email_ingestor.py accordingly.

* fix: suppress Bandit B324 false positives on non-cryptographic MD5 use

Same pattern as the previous pypdf2 commit: fixing the Safety scan
surfaced this PR's own Bandit HIGH-severity gate actually blocking on 10
pre-existing findings, all Bandit B324 ("Use of weak MD5 hash for
security").

Checked each of the 10 call sites: every one uses hashlib.md5() to build
a short deterministic cache key, entity ID, or IRI suffix from already-
non-secret input (query text, entity text/type, class/property names) -
none are used for passwords, tokens, or integrity verification of
untrusted data. This is exactly the case Bandit's own message points at
("Consider usedforsecurity=False").

Did not use usedforsecurity=False itself: that keyword argument was
added to hashlib in Python 3.9, and pyproject.toml declares
`requires-python = ">=3.8"` - adding it unconditionally risks a TypeError
on 3.8. Used a targeted `# nosec B324` comment with a one-line
justification instead, which suppresses only this specific check and
carries no runtime behavior change on any supported Python version.

Verified locally: bandit -r semantica/ -ll now reports 0 HIGH-severity
findings (was 10).

* docs: add CHANGELOG entry for #824 CI/CD supply-chain hardening

Covers the SHA-pinning + verify-action-pins.yml enforcement, release.yml
hardening (job-scoped permissions, concurrency, SLSA provenance), the
pypi environment/branch protection GitHub-side config, the
security-scan.yml Safety/comment-formatting fixes, and the two
vulnerabilities those fixes surfaced (pypdf2 CVE-2023-36464 removal,
Bandit B324 suppression).

* fix: close two remaining gaps missed by upstream bot-review fixes

verify-action-pins.sh:
- Quoted uses: lines (e.g. uses: owner/action@SHA) were not matched
  by the existing regex, so a SHA-pinned action written with quotes would
  silently skip verification. Updated the main ERE to accept an optional
  leading/trailing single or double quote around the owner/action@ref
  value, and excluded quote chars from the inner character classes so the
  ref is still extracted cleanly.
- The grep input glob only covered *.yml. GitHub also treats *.yaml as a
  valid workflow extension. Added *.yaml to the glob and a 2>/dev/null
  guard so the command doesn't fail when no *.yaml files exist.

security-scan.yml (on top of Kaif's --save-json fix in 67c7ec2a):
- Kaif's fix kept the '|| echo 0' fallback on the VULNS= line, so all
  five scanner-failure modes (file missing, empty file, malformed JSON,
  valid JSON with no 'vulnerabilities' key, vulnerabilities: null) still
  silently produce VULNS=0 or VULNS=null and pass the merge-blocker check.
- Added guard 1: '[ ! -s safety-report.json ]' fails loudly if Safety
  crashed before writing a report (covers missing and empty-file cases).
- Dropped the '|| echo 0' fallback and added guard 2: '[[ ! VULNS =~
  ^[0-9]+$ ]]' fails loudly on non-integer VULNS (covers malformed JSON,
  missing key, and null cases). Both guards emit ::error:: annotations.
- Verified with a 7-case simulation: all 5 failure modes now exit 1;
  genuine zero-vuln and real-vuln cases still behave correctly.

* fix: correct bash [[ =~ ]] quoting that broke verify-action-pins.sh in CI

The regex for matching uses: lines was embedded directly inline in a
[[ =~ ]] test with literal \" and \' escape sequences. Bash's conditional-
expression parser interprets these as shell syntax rather than regex
literals, producing:

  syntax error in conditional expression: unexpected token ')'

at line 27 on every CI run.

Fix: move the regex into a USES_PATTERN variable using safe single-quote
shell-string concatenation so the [[ =~ ]] parser receives an unquoted
variable reference ($USES_PATTERN) rather than a literal pattern containing
bash-special characters. The regex semantics are identical: optional
leading/trailing quote around owner/action@ref, quote chars excluded from
capture groups.

Verified in real bash 5.2.21 (Git for Windows):
  No syntax error on the real 40-pin workflow tree (Checked 40)
  Unquoted SHA pin:      MATCH, correct repo+ref extracted
  Double-quoted SHA pin: MATCH, correct repo+ref extracted
  Single-quoted SHA pin: MATCH, correct repo+ref extracted
  .yaml extension file:  MATCH, correct repo+ref extracted
  ./local-action:        NO MATCH (correct)
  docker://:             NO MATCH (correct)

* docs: add 3 missing items to fork-reconfiguration checklist in SECURITY.md

The checklist covered Trusted Publishing trust, protected environment,
branch protection, and Dependabot github-actions entry. Three non-forking
controls described elsewhere in SECURITY.md were omitted:

- GitHub secret scanning and push protection (repo settings, not copied
  on fork)
- GitGuardian (GitHub App installation scoped to this specific repo,
  requires separate install on any fork)
- CodeQL Default Setup vs Advanced Setup state (repo setting that affects
  whether the upload-sarif step in codeql.yml does anything)

Added as items 5, 6, 7 matching the existing numbered bullet style.

* fix: update github/codeql-action pins to v4 tip (SHA drift caught by verify check)

verify-action-pins caught that github/codeql-action@v4 tag was re-pointed
upstream:

  old: f205ea1c3313d32999d8d6a48b4f6530d4437b38
  new: d1ba80a13dd99fba24a470575428917156a28b43

Updated all 8 occurrences across codeql.yml (init x3, autobuild, analyze,
upload-sarif) and defender-for-devops.yml (upload-sarif x2). Tag comment
# v4 unchanged — the tag itself hasn't changed, only what commit it points to.

---------

Co-authored-by: Sameer6305 <sskadam6305@gmail.com>
2026-08-03 19:17:10 +05:30

14 KiB

Security Policy

Supported Versions

We actively support the following versions of Semantica with security updates:

Version Supported
0.2.3
0.2.2
0.2.1
0.2.0
0.1.1
0.1.0
< 0.1.0

Reporting a Vulnerability

We take security vulnerabilities seriously. If you discover a security vulnerability, please follow these steps:

1. Do NOT create a public GitHub issue

Security vulnerabilities should be reported privately to prevent potential exploitation.

2. Report Security Issue

Create a GitHub Security Advisory or contact us via the security email listed in SUPPORT.md.

Include the following information:

  • Type of vulnerability (e.g., XSS, SQL injection, authentication bypass)
  • Affected component (module, function, or file)
  • Steps to reproduce (detailed description or proof-of-concept code)
  • Potential impact (what could an attacker do?)
  • Suggested fix (if you have one)
  • Your contact information (for follow-up questions)

3. Response Timeline

  • Initial Response: Within 24 hours for critical issues; within 48 hours for non-critical issues
  • Status Update: Within 7 days
  • Resolution: Depends on severity and complexity

4. Disclosure Policy

  • We will acknowledge receipt of your report within 48 hours
  • We will provide regular updates on the status of the vulnerability
  • Once fixed, we will credit you (if desired) in the security advisory
  • We will coordinate public disclosure with you

Security Update Process

  1. Assessment: We assess the severity using CVSS scoring
  2. Fix Development: We develop and test a fix
  3. Release: We release a security update
  4. Advisory: We publish a security advisory on GitHub
  5. Communication: We notify users through appropriate channels

Severity Levels

Critical

  • Remote code execution
  • Authentication bypass
  • Data breach or exposure
  • Response Time: Immediate (within 24 hours)

High

  • Privilege escalation
  • Significant data leakage
  • Denial of service
  • Response Time: Within 7 days

Medium

  • Information disclosure
  • Cross-site scripting (XSS)
  • CSRF vulnerabilities
  • Response Time: Within 30 days

Low

  • Minor information leakage
  • Best practice violations
  • Response Time: Next release cycle

Known Security Considerations

Dependencies

We regularly update dependencies to address security vulnerabilities. However, you should:

  • Keep your dependencies up to date
  • Review security advisories for our dependencies
  • Use tools like pip-audit or safety to check for known vulnerabilities

API Keys and Credentials

  • Never commit API keys or credentials to the repository
  • Use environment variables or secure configuration management
  • Rotate keys regularly
  • Use least-privilege access principles

Data Handling

  • Be cautious when processing untrusted data
  • Validate and sanitize all inputs
  • Use parameterized queries for database operations
  • Implement rate limiting for public APIs

Network Security

  • Use HTTPS for all network communications
  • Validate SSL/TLS certificates
  • Be cautious with external API calls
  • Implement proper authentication and authorization

CI/CD Supply-Chain Security

Semantica's build and release pipeline is explicitly hardened against CI/CD supply-chain attacks — the class of attack behind the March 2026 LiteLLM/Trivy incident, where a compromised third-party Action with a mutable tag was used to steal a long-lived publishing token, after which malicious packages were pushed straight to PyPI without ever touching the source repository. Every control below maps directly to closing one step of that attack chain.

Immutable build inputs

  • Risk: a tag (@v4, @release/v1) is re-pointed by a compromised upstream maintainer or account, silently changing what every consumer's CI runs. Control: every third-party GitHub Action in every workflow is pinned to a full 40-character commit SHA, with the human-readable tag kept only as a trailing comment (e.g. actions/checkout@3d3c42e... # v7).
  • Risk: a SHA pin drifts out of sync with its own comment over time, or is mistyped. Control: verify-action-pins.yml fails closed on any uses: reference that isn't a full commit SHA (catching a newly added mutable tag, not just auditing existing pins), resolves every pinned tag via the GitHub API on each workflow change, on every push to main, and weekly, and fails if the SHA no longer matches the tag it claims to be — an API lookup that can't be resolved is treated as a failure, not a silent skip.
  • Risk: manually re-pinning ~15 actions across 8 workflow files on every upstream release is error-prone. Control: Dependabot (github-actions ecosystem) opens a grouped PR that bumps the SHA and the tag comment together whenever an action releases — pins never require hand-editing.

Publishing pipeline (highest-privilege path)

  • Risk: a long-lived PYPI_TOKEN sitting in repo/org secrets is exfiltrated by any compromised step. Control: PyPI publishing uses Trusted Publishing (OIDC) (id-token: write) — there is no long-lived PyPI credential anywhere in this repository to steal.
  • Risk: a compromised CI run publishes to PyPI with no human in the loop. Control: the publish job runs only inside a protected pypi GitHub Environment with a required human reviewer — every release needs manual approval in the Actions UI before it runs.
  • Risk: the release job could be triggered from an arbitrary branch/ref. Control: the pypi environment's deployment-branch policy is restricted to v* tags only.
  • Risk: a scanner or unrelated job inherits publish-level credentials. Control: release.yml sets permissions: contents: read at the workflow level; contents: write / id-token: write / attestations: write are granted only to the release job, never workflow-wide.
  • Risk: two tag pushes race through the publish pipeline simultaneously. Control: concurrency: group: release-${{ github.ref }} serializes releases per tag.
  • Risk: a consumer can't verify a wheel on PyPI actually came from this repo's CI. Control: SLSA build provenance is attested for every release via actions/attest-build-provenance, producing a signed, verifiable record of the exact commit and workflow run that produced the artifact (checkable with gh attestation verify).

Repository controls

  • Risk: unreviewed or force-pushed changes land on main. Control: main requires 1 approving PR review (stale approvals dismissed on new pushes), resolved conversations, and blocks force-pushes and branch deletion.
  • Risk: a PR merges without its security/CI checks passing. Control: merges require the build, Analyze Python (CodeQL), and security-scan checks to pass, in strict mode (checks must be re-run against the latest main).
  • Risk: a compromised scanner job reaches secrets or write access. Control: scanning jobs (CodeQL, security-scan.yml, security.yml, defender-for-devops.yml) run with read-only, least-privilege permissions (typically contents: read + security-events: write only) and never share a job, environment, or secret scope with the publish job.
  • Risk: secrets are committed accidentally. Control: GitHub secret scanning and push protection are both enabled at the repository level, rejecting pushes that contain recognizable credential patterns before they land in history.

Automated Security Scanning

Every scan below runs continuously in CI, not just at release time:

  • CodeQL (security-and-quality query pack) — Python source: injection, unsafe deserialization, and other code-level vulnerability classes. Runs in codeql.yml on every push/PR to main and weekly.
  • Bandit — Python-specific security anti-patterns (hardcoded secrets, unsafe eval/pickle, weak crypto, etc.); CI fails on any HIGH-severity finding. Runs in security-scan.yml on every push/PR to main and twice weekly.
  • Semgrep (p/security ruleset) — cross-language static-analysis security patterns. Runs in security-scan.yml on every push/PR to main and twice weekly.
  • Safety — known CVEs in Semantica's own installed dependencies, including optional LLM-provider extras such as LiteLLM; CI fails on any match. Runs in security-scan.yml on every push/PR to main and twice weekly.
  • pip-audit — independent, PyPA-maintained vulnerability database cross-check against installed dependencies (Safety and pip-audit use different advisory sources, so both run). Runs in security.yml weekly.
  • Microsoft Defender for DevOps (eslint, templateanalyzer, terrascan) — JavaScript/TypeScript lint-security rules and infrastructure-as-code misconfigurations. Runs in defender-for-devops.yml on every push/PR to main and weekly.
  • Checkov — Kubernetes, Helm, Dockerfile, GitHub Actions, and secrets-pattern IaC scanning; results upload to the same Security tab as CodeQL. Runs in defender-for-devops.yml on every push/PR to main and weekly.
  • GitGuardian — secret-detection check on every pull request, installed as a GitHub App integration (not a repo-local workflow). Runs on every PR.
  • GitHub secret scanning + push protection — blocks known credential patterns before they're pushed, and continuously scans existing history. Platform-level, continuous.
  • Dependabot — version/security PRs for Python, Docker, and GitHub Actions dependencies, grouped where relevant to reduce review noise. Configured in .github/dependabot.yml, runs weekly for security-relevant packages and monthly for docs dependencies.
  • verify-action-pins.yml — enforces that every Action reference is a full commit SHA (failing on a newly introduced mutable tag) and confirms each SHA still matches the tag it claims to be. Runs on every workflow change, every push to main, and weekly.

All SARIF-producing scanners (CodeQL, Checkov, Microsoft Defender) publish findings to the repository's Security → Code scanning alerts tab, giving a single audit trail across tools rather than scattered per-tool reports.

Adopting this posture in a fork or downstream deployment

Teams standing up their own instance of Semantica, or forking it for an internal/regulated deployment, can reuse this posture directly:

  1. Keep Dependabot's github-actions ecosystem entry — it is what keeps SHA pins current without manual maintenance.
  2. Re-run verify-action-pins.yml after re-pointing the repository's Actions at your own mirrors, if you do so.
  3. If you publish your own PyPI package from a fork, configure your own Trusted Publishing trust relationship on PyPI (Trusted Publishing is scoped to a specific owner/repo + workflow filename) and your own protected environment with your own required reviewers — these are not transferable from this repository.
  4. Branch protection, environment protection, and repository secret scanning are repository settings, not workflow files — cloning or forking the repo does not copy them. They must be re-applied via the GitHub UI or API on the new repository.
  5. GitHub secret scanning and push protection are repository settings that don't carry over to a fork either — re-enable both under the new repository's Security settings, not just Dependabot.
  6. GitGuardian runs as a GitHub App installation scoped to this specific repository, not a workflow file — a fork gets no secret-detection coverage from it until the app is installed separately on the new repo.
  7. CodeQL's upload-sarif step in codeql.yml only runs meaningfully if Default Setup is not already enabled for the repository (it's designed to skip gracefully otherwise) — check whether Default Setup or Advanced Setup is active on the new repository and adjust expectations for where CodeQL findings show up accordingly.

Dependency Security Policy

Regular Updates

  • We monitor security advisories for all dependencies
  • We update dependencies regularly in our development branch
  • Critical security updates are backported to supported versions

Reporting Dependency Vulnerabilities

If you discover a vulnerability in one of our dependencies:

  1. Check if it's already reported upstream
  2. Report to us if it affects Semantica specifically
  3. We will coordinate with upstream maintainers if needed

Security Scanning

We use automated tools to scan for vulnerabilities:

  • Dependabot: Automated dependency updates and security alerts
  • GitHub Security Advisories: Vulnerability tracking
  • Manual Reviews: Regular security audits

Best Practices for Users

  1. Keep Semantica Updated: Always use the latest stable version
  2. Review Dependencies: Regularly update your project dependencies
  3. Secure Configuration: Use secure defaults and proper configuration
  4. Monitor Logs: Watch for suspicious activity
  5. Report Issues: Don't hesitate to report potential security issues

Security Acknowledgments

We appreciate responsible disclosure. Security researchers who help us improve the security of Semantica will be:

  • Credited in security advisories (if desired)
  • Listed in our security acknowledgments
  • Recognized for their contribution

Contact

For security-related questions or concerns:

  • Private Reporting: Please do not report vulnerabilities in public issues.
  • GitHub Security Advisories: Report vulnerability

Additional Resources


Thank you for helping keep Semantica and its users safe!