mirror of
https://github.com/semantica-agi/semantica.git
synced 2026-08-29 04:26:20 +00:00
* security: SHA-pin all Actions, harden release pipeline, add pin verification
Hardens the CI/CD supply chain against the LiteLLM/Trivy-style attack (a
compromised third-party Action with a mutable tag stealing a long-lived
publishing token) and closes several related gaps found in an audit of the
actual repository state.
- Pin every third-party GitHub Action across all workflows to a full commit
SHA (tag kept as a trailing comment); add verify-action-pins.yml, a CI
check that confirms via the GitHub API that each pin still matches its
tag, on every workflow change, push to main, and weekly.
- Scope release.yml permissions to the job level (workflow defaults to
contents: read); add a concurrency group so simultaneous tag pushes can't
race the publish job.
- Add SLSA build provenance attestation (actions/attest-build-provenance)
for every released wheel.
- Fix a latent bug in security-scan.yml: the PR-comment step was missing
pull-requests: write and silently failing; add bounded artifact retention
for uploaded scan reports.
- Group github-actions Dependabot updates to cut review noise.
- Document the resulting posture in SECURITY.md for auditors/regulated
adopters, including what's enforced and what a fork needs to reconfigure
for itself (environment/branch protection, Trusted Publishing trust).
Also (via GitHub API, not in this diff): created a protected `pypi`
environment with a required reviewer restricted to v* tags, and enabled
branch protection on main (required review, required status checks, no
force-push/deletion).
* fix: harden verify-action-pins per PR #824 bot review
Addresses real findings from the automated review on #824:
- The script previously only matched uses: lines that already contained a
40-hex SHA, so a newly added mutable-tag action (e.g. some/action@v1)
would never be scanned at all and the check would pass silently. It now
matches every uses: line and hard-fails on any ref that isn't a full
commit SHA.
- A tag that fails to resolve via the GitHub API (rate limit, deleted tag)
previously only logged a warning and continued; that's now a hard
failure too, since an unverifiable pin is exactly the failure mode this
check exists to catch.
- verify-action-pins.yml only triggered on .github/workflows/** changes,
so an edit to the verifier script itself wouldn't run the check that
verifies it. Added the script path to both trigger filters.
The reviewer's claim that slash-containing tag comments (release/v1) break
the API lookup did not reproduce - tested directly against
pypa/gh-action-pypi-publish@release/v1 and GitHub's commits API resolves
multi-segment refs natively - so no change was needed there.
Verified with a synthetic test workflow containing a mutable-tag action,
a correctly-pinned SHA, and a deliberately mismatched SHA: the updated
script now catches the first and third cases and passes the second. Also
re-ran against the real workflow tree (40/40 pins still verify clean).
* fix: repair broken Safety scan and PR comment formatting
The "Comment PR with Security Results" step was producing garbled output
(literal \n characters instead of newlines, "undefined:" labels) because:
- Every line in the JS comment builder used \n (escaped backslash-n)
inside template literals, which JS renders as the literal two-character
string \n, not a newline.
- The Semgrep section read issue.rule_id, but Semgrep's JSON field is
check_id - hence "undefined: <path>" for every entry.
Rewrote the comment builder to construct each section as an array of
lines joined with a real '\n', with correct field names, and collapsed
long finding lists into a <details> block instead of a flat list.
Verified by extracting the exact script and running it under node against
synthetic fixtures matching each tool's real JSON schema (found/clean/
missing-report paths all render correctly).
While tracing the "undefined" and always-empty Safety section, found the
Safety step itself was silently broken:
- `safety check --json --output safety-report.json` is invalid in
Safety 3.x: --output now selects a console format (json/text/screen),
not a file path. The command errored on every run (swallowed by
`|| true`), so safety-report.json was never created and the PR comment
always fell back to a generic "scan completed" message. Switched to
`--save-json`, which is the correct flag for writing a JSON report to
disk, and confirmed against the real safety 3.8.1 CLI locally.
- Even with a report, the code read vuln.package - the real field is
package_name.
- The job never installed Semantica's own dependencies before scanning,
so `safety check` (which defaults to scanning the environment) was
auditing the scanner tools' own dependencies, not Semantica's. Added
`pip install -e ".[llm-litellm]"` so the project's actual dependency
tree - including the LiteLLM extra this whole hardening effort is
about - is what gets scanned.
Also updated the corresponding SECURITY.md bullet to describe what Safety
actually covers now.
* fix: remove unused pypdf2 dependency (CVE-2023-36464)
Now that the Safety scan step actually runs (see previous commit), it
correctly failed this PR's checks on CVE-2023-36464 in pypdf2==3.0.1 - a
real, pre-existing vulnerability that was invisible until the scan was
fixed.
PyPDF2 is not a patchable dependency here: the project is discontinued
(merged into `pypdf`), 3.0.1 is its final release, and there is no fixed
version to upgrade to. Grepping the repo for `import PyPDF2` / `from
PyPDF2` turns up nothing - it was never actually imported anywhere. Its
only presence outside pyproject.toml was in docstrings describing a
"PyPDF2.PdfReader() fallback" for PDF parsing that was never implemented
in code; pdfplumber is the library actually used. Removed the dependency
and corrected the stale docstrings in parse/__init__.py, parse/methods.py,
parse/pdf_parser.py, and ingest/email_ingestor.py accordingly.
* fix: suppress Bandit B324 false positives on non-cryptographic MD5 use
Same pattern as the previous pypdf2 commit: fixing the Safety scan
surfaced this PR's own Bandit HIGH-severity gate actually blocking on 10
pre-existing findings, all Bandit B324 ("Use of weak MD5 hash for
security").
Checked each of the 10 call sites: every one uses hashlib.md5() to build
a short deterministic cache key, entity ID, or IRI suffix from already-
non-secret input (query text, entity text/type, class/property names) -
none are used for passwords, tokens, or integrity verification of
untrusted data. This is exactly the case Bandit's own message points at
("Consider usedforsecurity=False").
Did not use usedforsecurity=False itself: that keyword argument was
added to hashlib in Python 3.9, and pyproject.toml declares
`requires-python = ">=3.8"` - adding it unconditionally risks a TypeError
on 3.8. Used a targeted `# nosec B324` comment with a one-line
justification instead, which suppresses only this specific check and
carries no runtime behavior change on any supported Python version.
Verified locally: bandit -r semantica/ -ll now reports 0 HIGH-severity
findings (was 10).
* docs: add CHANGELOG entry for #824 CI/CD supply-chain hardening
Covers the SHA-pinning + verify-action-pins.yml enforcement, release.yml
hardening (job-scoped permissions, concurrency, SLSA provenance), the
pypi environment/branch protection GitHub-side config, the
security-scan.yml Safety/comment-formatting fixes, and the two
vulnerabilities those fixes surfaced (pypdf2 CVE-2023-36464 removal,
Bandit B324 suppression).
* fix: close two remaining gaps missed by upstream bot-review fixes
verify-action-pins.sh:
- Quoted uses: lines (e.g. uses: owner/action@SHA) were not matched
by the existing regex, so a SHA-pinned action written with quotes would
silently skip verification. Updated the main ERE to accept an optional
leading/trailing single or double quote around the owner/action@ref
value, and excluded quote chars from the inner character classes so the
ref is still extracted cleanly.
- The grep input glob only covered *.yml. GitHub also treats *.yaml as a
valid workflow extension. Added *.yaml to the glob and a 2>/dev/null
guard so the command doesn't fail when no *.yaml files exist.
security-scan.yml (on top of Kaif's --save-json fix in 67c7ec2a):
- Kaif's fix kept the '|| echo 0' fallback on the VULNS= line, so all
five scanner-failure modes (file missing, empty file, malformed JSON,
valid JSON with no 'vulnerabilities' key, vulnerabilities: null) still
silently produce VULNS=0 or VULNS=null and pass the merge-blocker check.
- Added guard 1: '[ ! -s safety-report.json ]' fails loudly if Safety
crashed before writing a report (covers missing and empty-file cases).
- Dropped the '|| echo 0' fallback and added guard 2: '[[ ! VULNS =~
^[0-9]+$ ]]' fails loudly on non-integer VULNS (covers malformed JSON,
missing key, and null cases). Both guards emit ::error:: annotations.
- Verified with a 7-case simulation: all 5 failure modes now exit 1;
genuine zero-vuln and real-vuln cases still behave correctly.
* fix: correct bash [[ =~ ]] quoting that broke verify-action-pins.sh in CI
The regex for matching uses: lines was embedded directly inline in a
[[ =~ ]] test with literal \" and \' escape sequences. Bash's conditional-
expression parser interprets these as shell syntax rather than regex
literals, producing:
syntax error in conditional expression: unexpected token ')'
at line 27 on every CI run.
Fix: move the regex into a USES_PATTERN variable using safe single-quote
shell-string concatenation so the [[ =~ ]] parser receives an unquoted
variable reference ($USES_PATTERN) rather than a literal pattern containing
bash-special characters. The regex semantics are identical: optional
leading/trailing quote around owner/action@ref, quote chars excluded from
capture groups.
Verified in real bash 5.2.21 (Git for Windows):
No syntax error on the real 40-pin workflow tree (Checked 40)
Unquoted SHA pin: MATCH, correct repo+ref extracted
Double-quoted SHA pin: MATCH, correct repo+ref extracted
Single-quoted SHA pin: MATCH, correct repo+ref extracted
.yaml extension file: MATCH, correct repo+ref extracted
./local-action: NO MATCH (correct)
docker://: NO MATCH (correct)
* docs: add 3 missing items to fork-reconfiguration checklist in SECURITY.md
The checklist covered Trusted Publishing trust, protected environment,
branch protection, and Dependabot github-actions entry. Three non-forking
controls described elsewhere in SECURITY.md were omitted:
- GitHub secret scanning and push protection (repo settings, not copied
on fork)
- GitGuardian (GitHub App installation scoped to this specific repo,
requires separate install on any fork)
- CodeQL Default Setup vs Advanced Setup state (repo setting that affects
whether the upload-sarif step in codeql.yml does anything)
Added as items 5, 6, 7 matching the existing numbered bullet style.
* fix: update github/codeql-action pins to v4 tip (SHA drift caught by verify check)
verify-action-pins caught that github/codeql-action@v4 tag was re-pointed
upstream:
old: f205ea1c3313d32999d8d6a48b4f6530d4437b38
new: d1ba80a13dd99fba24a470575428917156a28b43
Updated all 8 occurrences across codeql.yml (init x3, autobuild, analyze,
upload-sarif) and defender-for-devops.yml (upload-sarif x2). Tag comment
# v4 unchanged — the tag itself hasn't changed, only what commit it points to.
---------
Co-authored-by: Sameer6305 <sskadam6305@gmail.com>
237 lines
9.7 KiB
YAML
237 lines
9.7 KiB
YAML
name: Security Scan
|
|
|
|
on:
|
|
schedule:
|
|
- cron: '30 1 * * 1,4' # Mon/Thu 7 AM IST
|
|
push:
|
|
branches: [main]
|
|
paths-ignore:
|
|
- 'docs/**'
|
|
- 'mkdocs.yml'
|
|
- 'requirements-docs.txt'
|
|
- '**/*.md'
|
|
pull_request:
|
|
branches: [main]
|
|
paths-ignore:
|
|
- 'docs/**'
|
|
- 'mkdocs.yml'
|
|
- 'requirements-docs.txt'
|
|
- '**/*.md'
|
|
|
|
permissions:
|
|
contents: read
|
|
|
|
jobs:
|
|
security-scan:
|
|
runs-on: ubuntu-latest
|
|
permissions:
|
|
contents: read
|
|
security-events: write
|
|
actions: read
|
|
# Needed for the "Comment PR with Security Results" step below. Safe on
|
|
# pull_request (not pull_request_target): GitHub always forces a
|
|
# read-only token for PRs from forks regardless of this permission.
|
|
pull-requests: write
|
|
|
|
steps:
|
|
- name: Checkout repository
|
|
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
|
|
|
- name: Set up Python
|
|
uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7
|
|
with:
|
|
python-version: '3.11'
|
|
|
|
- name: Install dependencies
|
|
run: |
|
|
python -m pip install --upgrade pip
|
|
pip install safety bandit semgrep jq
|
|
# Install the project itself (core deps + the LiteLLM provider extra)
|
|
# so Safety scans Semantica's actual dependency tree, not just the
|
|
# scanner tools' own dependencies.
|
|
pip install -e ".[llm-litellm]"
|
|
|
|
- name: Run Safety Check (Package Vulnerabilities)
|
|
run: |
|
|
# NOTE: Safety 3.x repurposed --output to select a console format
|
|
# (json/text/screen/...), not a file path. Writing JSON to a file
|
|
# now requires --save-json; the previous `--output safety-report.json`
|
|
# usage was silently invalid and never produced a report.
|
|
safety check --save-json safety-report.json || true
|
|
|
|
# Guard 1: fail loudly if Safety exited before writing a report at all
|
|
# (network error, API auth failure, tool crash). Without this check a
|
|
# missing or empty file causes jq to fall back to "0", making a broken
|
|
# scanner indistinguishable from a clean scan.
|
|
if [ ! -s safety-report.json ]; then
|
|
echo "::error::Safety scan produced no report (safety-report.json is missing or empty). Treating as failure — check for network errors, API auth failures, or Safety crashes in the logs above."
|
|
exit 1
|
|
fi
|
|
|
|
echo "Checking for package vulnerabilities..."
|
|
|
|
# No || echo "0" fallback: if jq fails (malformed JSON, missing key,
|
|
# vulnerabilities:null) VULNS will be empty or "null" so guard 2 below
|
|
# catches it rather than silently treating the broken report as zero.
|
|
VULNS=$(jq '.vulnerabilities | length' safety-report.json 2>/dev/null)
|
|
|
|
# Guard 2: ensure VULNS is a non-negative integer before the -gt
|
|
# comparison. "null" (missing/null key) or "" (jq parse failure) would
|
|
# cause bash's -gt to throw an arithmetic error and fall through to the
|
|
# success branch — the same silent-pass bug as a missing file.
|
|
if ! [[ "$VULNS" =~ ^[0-9]+$ ]]; then
|
|
echo "::error::Safety report exists but 'vulnerabilities' is missing or non-numeric (got: '${VULNS}'). The report may be malformed or Safety may have written an error-only JSON. Treating as failure."
|
|
exit 1
|
|
fi
|
|
|
|
if [ "$VULNS" -gt 0 ]; then
|
|
echo "❌ Security vulnerabilities found: $VULNS"
|
|
echo "CI will fail to prevent merging of vulnerable dependencies"
|
|
echo ""
|
|
echo "Vulnerability details:"
|
|
jq -r '.vulnerabilities[] | "- \(.package_name)==\(.analyzed_version): \(.vulnerability_id) (\(.CVE // "no CVE assigned"))"' safety-report.json || true
|
|
exit 1
|
|
else
|
|
echo "✅ No security vulnerabilities found"
|
|
fi
|
|
|
|
- name: Run Bandit (Code Security Linter)
|
|
run: |
|
|
bandit -r semantica/ -f json -o bandit-report.json || true
|
|
echo "Checking for HIGH severity security issues..."
|
|
|
|
# Count HIGH severity issues
|
|
HIGH_ISSUES=$(bandit -r semantica/ -f json -ll 2>/dev/null | jq -r '.results[]? | select(.issue_severity == "HIGH") | .test_name' 2>/dev/null | wc -l || echo "0")
|
|
|
|
if [ "$HIGH_ISSUES" -gt 0 ]; then
|
|
echo "❌ HIGH severity security issues found: $HIGH_ISSUES"
|
|
echo "CI will fail to prevent merging of high-risk code"
|
|
echo ""
|
|
echo "High severity issues:"
|
|
bandit -r semantica/ -ll | grep "Severity: High" -A 5 -B 1 || true
|
|
exit 1
|
|
else
|
|
echo "✅ No HIGH severity security issues found"
|
|
fi
|
|
|
|
- name: Run Semgrep (Static Analysis)
|
|
run: |
|
|
echo "Running Semgrep static analysis..."
|
|
semgrep --config=auto --json --output=semgrep-report.json semantica/ || true
|
|
|
|
# Run security-focused rules
|
|
echo "Checking for security patterns..."
|
|
SECURITY_ISSUES=$(semgrep --config=p/security --json semantica/ 2>/dev/null | jq '.results | length' 2>/dev/null || echo "0")
|
|
|
|
if [ "$SECURITY_ISSUES" -gt 0 ]; then
|
|
echo "⚠️ Security patterns found: $SECURITY_ISSUES"
|
|
echo "Review these findings for potential improvements"
|
|
semgrep --config=p/security semantica/ || true
|
|
else
|
|
echo "✅ No security patterns found"
|
|
fi
|
|
|
|
- name: Upload Security Reports
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
|
|
with:
|
|
name: security-reports
|
|
retention-days: 14
|
|
path: |
|
|
safety-report.json
|
|
bandit-report.json
|
|
semgrep-report.json
|
|
|
|
- name: Comment PR with Security Results
|
|
if: github.event_name == 'pull_request'
|
|
uses: actions/github-script@3a2844b7e9c422d3c10d287c895573f7108da1b3 # v9
|
|
with:
|
|
script: |
|
|
const fs = require('fs');
|
|
|
|
// Renders one tool's findings as a section. `items` is already
|
|
// the list of pre-formatted "- `thing` in `where`" strings; this
|
|
// just handles the found/not-found/report-missing framing and
|
|
// collapses long lists into a <details> block so the comment
|
|
// doesn't turn into a wall of text.
|
|
function renderSection(title, reportPath, parse) {
|
|
let data;
|
|
try {
|
|
data = JSON.parse(fs.readFileSync(reportPath, 'utf8'));
|
|
} catch (e) {
|
|
return [
|
|
`### ${title}`,
|
|
`⚠️ No report found at \`${reportPath}\` — the scan may have failed before producing output. Check the job logs.`,
|
|
].join('\n');
|
|
}
|
|
|
|
const items = parse(data);
|
|
if (items.length === 0) {
|
|
return [`### ${title}`, `✅ No findings.`].join('\n');
|
|
}
|
|
|
|
const lines = [`### ${title}`, `Found **${items.length}**.`, ''];
|
|
const shown = items.slice(0, 15);
|
|
if (items.length > 15) {
|
|
lines.push('<details>', '<summary>Show all findings</summary>', '');
|
|
lines.push(...items);
|
|
lines.push('', '</details>');
|
|
} else {
|
|
lines.push(...shown);
|
|
}
|
|
return lines.join('\n');
|
|
}
|
|
|
|
const safetySection = renderSection(
|
|
'Safety — dependency vulnerabilities',
|
|
'safety-report.json',
|
|
(data) => (data.vulnerabilities || []).map(
|
|
(v) => `- \`${v.package_name}==${v.analyzed_version}\`: ${v.vulnerability_id}` +
|
|
(v.CVE ? ` (${v.CVE})` : '') + ` — ${v.advisory || 'no advisory text'}`
|
|
)
|
|
);
|
|
|
|
const banditSection = renderSection(
|
|
'Bandit — HIGH-severity code issues',
|
|
'bandit-report.json',
|
|
(data) => (data.results || [])
|
|
.filter((issue) => issue.issue_severity === 'HIGH')
|
|
.map((issue) => `- \`${issue.test_name}\` in \`${issue.filename}:${issue.line_number}\``)
|
|
);
|
|
|
|
const semgrepSection = renderSection(
|
|
'Semgrep — static analysis patterns',
|
|
'semgrep-report.json',
|
|
(data) => (data.results || []).map(
|
|
(issue) => `- \`${issue.check_id}\` in \`${issue.path}:${issue.start?.line ?? '?'}\``
|
|
)
|
|
);
|
|
|
|
const comment = [
|
|
'# 🔒 Security Scan Results',
|
|
'',
|
|
safetySection,
|
|
'',
|
|
banditSection,
|
|
'',
|
|
semgrepSection,
|
|
'',
|
|
'---',
|
|
'',
|
|
'*This security scan runs automatically on source-code PRs and bi-weekly (skipped for doc/markdown-only changes).*',
|
|
'',
|
|
'📊 **Security Policy**: CI fails on Safety vulnerabilities and Bandit HIGH-severity findings. Semgrep findings above are informational and do not block merge.',
|
|
].join('\n');
|
|
|
|
try {
|
|
await github.rest.issues.createComment({
|
|
issue_number: context.issue.number,
|
|
owner: context.repo.owner,
|
|
repo: context.repo.repo,
|
|
body: comment,
|
|
});
|
|
console.log('✅ Security comment posted successfully');
|
|
} catch (error) {
|
|
console.log('⚠️ Could not post security comment:', error.message);
|
|
console.log('📋 Security scan results saved to artifacts');
|
|
}
|