The first-knowledge-graph lesson read the parser output from a key it
never returns (content vs text), then masked the failure with
hard-coded entities, a manually assembled NetworkX graph, and an
uninvoked KGVisualizer; the final cell deleted the sample file, so
rerunning intermediate cells failed. Rework the notebook so every
stage consumes the previous stage's output:
- parse via parsed_document["text"] with an assertion that content
was actually extracted
- real NERExtractor/RelationExtractor output replaces the simulated
entities and wrong hard-coded offsets
- GraphBuilder builds the graph from actual relation endpoints via a
mention-span -> graph-ID map (consistent with notebook 07)
- KGVisualizer.visualize_network renders the graph and saves HTML
- deletion moved to an explicit optional cleanup cell, so parsing and
downstream cells stay rerunnable
Closes#1289
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Lead the README with Semantica's identity as a semantic/context/knowledge
layer (Context Graph, KG, ontology and vocabulary governance via OWL/SHACL/SKOS),
with decision provenance and audit trails framed as a property of that
structure rather than the flagship pitch.
The Neo4j example in the persistent graph store section was passing a raw
Neo4jStore straight into GraphBuilder. GraphBuilder calls add_nodes() and
add_edges() on whatever it's given, and those only exist on the GraphStore
facade, not on Neo4jStore itself. Anyone who copied the example got:
AttributeError: 'Neo4jStore' object has no attribute 'add_nodes'
Swapped the import and construction to GraphStore(backend="neo4j", uri=...,
user=..., password=...), which wraps Neo4jStore internally and actually has
the methods GraphBuilder needs.
Added a test in tests/kg/test_graph_builder_with_graph_store.py that builds
a small graph through GraphBuilder with a mocked GraphStore and checks
add_nodes/add_edges get called. Also kept a test for GraphBuilder without a
graph_store at all, so that path doesn't regress either.
Closes#1135
* fix(ci): drop unpinnable benchmarks/requirements.txt install
Scorecard flagged this pip install as unpinned-by-hash (#6082). Can't
hash-pin it - benchmarks/requirements.txt doesn't exist in this repo, so
there's nothing to compile a lockfile from. Dropping it instead of
leaving it unpinned: the job already fails on the next real step
(benchmarks/benchmarks_runner.py, also missing), so this line wasn't
doing anything useful to begin with.
* fix(ci): hash-pin the spacy model download in benchmark.yml
Qodo review on this PR: dropping the benchmarks/requirements.txt install
(the previous failure point) let the job actually reach
`python -m spacy download en_core_web_sm`, which fetches an unpinned,
unhashed wheel from spacy-models' GitHub releases - undoing the point of
this PR by exposing a real unpinned-install path instead of a dead one.
Replaced with a hash-pinned direct-URL entry in benchmark-extra.in/.txt
for en_core_web_sm-3.8.0 (matches the spacy==3.8.15 already pinned in
base-deps.txt). uv independently computed the same sha256 I got via a
manual curl+sha256 of the release asset, and a --require-hashes dry-run
install verifies clean.
* fix(ci): stop checkov's suppressed checks from reopening as new alerts
Root cause found, not just worked around: checkov's SARIF exporter
includes every evaluated check as an ordinary result, including ones it
internally marked SKIPPED via the inline # checkov:skip= comments and
checkov.io/skipN annotations already on the Helm chart. It never uses
SARIF's own `suppressions` field and never drops them - so the exact same
already-suppressed finding reopens as a brand-new code scanning alert
number on every single run, forever (#6035/#6036, #6112-6115,
#6128-6131 are all the same 4 findings, manually dismissed 3 times now).
checkov's JSON output *does* correctly record which checks were skipped.
Added .github/scripts/filter_checkov_skipped.py, which cross-references
the JSON's skipped_checks against the SARIF's results (matched by check
ID + the last two path segments, since the two outputs use different path
roots) and drops anything checkov itself already decided to suppress,
before upload. Verified locally against a real checkov+helm run: removed
exactly the 4 known-suppressed helm chart results, left the 2 genuinely
real findings (deploy/gcp/cloudrun-service.yaml, deploy/kubernetes/
deployment.yaml) untouched.
* fix(ci): drop unpinnable benchmarks/requirements.txt install
Scorecard flagged this pip install as unpinned-by-hash (#6082). Can't
hash-pin it - benchmarks/requirements.txt doesn't exist in this repo, so
there's nothing to compile a lockfile from. Dropping it instead of
leaving it unpinned: the job already fails on the next real step
(benchmarks/benchmarks_runner.py, also missing), so this line wasn't
doing anything useful to begin with.
* fix(ci): hash-pin the spacy model download in benchmark.yml
Qodo review on this PR: dropping the benchmarks/requirements.txt install
(the previous failure point) let the job actually reach
`python -m spacy download en_core_web_sm`, which fetches an unpinned,
unhashed wheel from spacy-models' GitHub releases - undoing the point of
this PR by exposing a real unpinned-install path instead of a dead one.
Replaced with a hash-pinned direct-URL entry in benchmark-extra.in/.txt
for en_core_web_sm-3.8.0 (matches the spacy==3.8.15 already pinned in
base-deps.txt). uv independently computed the same sha256 I got via a
manual curl+sha256 of the release asset, and a --require-hashes dry-run
install verifies clean.
Dependabot flagged 12 aiohttp advisories (1 high, rest moderate/low - CVE
range covering request smuggling, websocket/parser bugs, cookie/redirect
issues) against aiohttp==3.13.5 pinned in checkov.txt. checkov==3.3.1
itself pinned `aiohttp<3.14.0`, which excludes every fixed release;
3.3.16 (latest) relaxes that to `<3.15.0`, so bumping checkov also lets
aiohttp resolve to 3.14.3 (fixes all of them).
Two alerts remain open, both genuinely blocked upstream rather than
something a version bump here can fix:
- asteval: checkov 3.3.16 (latest, still) hard-pins asteval==1.0.6 with
no range; the fix (1.0.9) is unresolvable without violating checkov's
own declared dependency - confirmed via `uv pip compile` refusing to
solve it. Needs checkov itself to bump the pin upstream.
- ecdsa: 0.19.2 is already the latest release; the Minerva timing-attack
advisory has no patched version, since python-ecdsa's maintainers have
stated side-channel attacks are out of scope for the project.
Both are checkov's own transitive deps, used only for local static IaC
analysis in defender-for-devops.yml (no network signing/cloud-auth calls
that would actually exercise ecdsa's signing path) - dismissing on
GitHub with that reasoning as a separate step.
* fix(docker): split explorer-extra.txt by Python version, fix broken build
main's container-scan.yml has been failing since PR #1338 merged:
ERROR: In --require-hashes mode, all requirements must have their
versions pinned with ==. These do not:
standard-aifc from .../standard_aifc-3.13.0-py3-none-any.whl
(from audioread==3.1.0->-r explorer-extra.txt (line 30))
Root cause: explorer-extra.txt was compiled with `--python-version 3.11`
but is installed on the Dockerfile's actual python:3.13-slim interpreter.
librosa's audioread dependency needs standard-aifc/standard-sunau only
under `python_version >= "3.13"` (Python 3.13 dropped aifc/sunau from
stdlib) - a file resolved for 3.11 has no hash for those packages at all,
so --require-hashes fails outright once pip resolves against the real
3.13 environment instead of silently under-pinning.
Splits the file in two: explorer-extra-py311.txt (ci.yml, unchanged
resolution) and explorer-extra-py313.txt (Dockerfile, newly compiled for
--python-version 3.13). They aren't interchangeable and shouldn't be
recombined - documented in .github/requirements/README.md, including how
to catch this class of bug before it ships again.
* fix(ci): correct stale -o path in explorer-extra-py311.txt header
Qodo review on this PR: the autogenerated header comment still said
-o .github/requirements/explorer-extra.txt (the pre-rename path), which
would silently regenerate the wrong file if someone copy-pasted it.
The six decision-model dataclasses (Decision, DecisionContext, Policy,
PolicyException, Precedent, ApprovalChain) accepted `auto_generate_id`
only as a plain `__post_init__` parameter. Since it was neither a field
nor an InitVar, the generated `__init__` never forwarded it, so the
parameter was always its `True` default and the `auto_generate_id=False`
validation branch was unreachable dead code.
Declare `auto_generate_id: InitVar[bool] = True` on each dataclass so the
generated `__init__` forwards it to `__post_init__`, restoring the
required-id contract. Serialization is unaffected because InitVar is not
a real field. Add regression tests covering the auto-generate path, the
required-id error path, and the explicit-id path.
Fixes#1152
Co-authored-by: Sameer Kadam <sskadam6305@gmail.com>
Qodo review on this PR: `pip install --no-deps -e .` / `pip install
--no-deps .` still leaves PEP 517 build isolation on by default, which
fetches [build-system] requires (setuptools==84.0.0, wheel==0.48.0)
completely outside any hash checking - the --require-hashes installs
right next to it didn't cover this at all.
Adds .github/requirements/pep517-build.txt, hash-locked to the exact
pyproject.toml [build-system] requires, and installs it before every
local-source install (Dockerfile, ci.yml, benchmark.yml) with
--no-build-isolation so pip reuses those hash-verified copies instead
of fetching its own.
Scorecard's Pinned-Dependencies check requires pip installs to be
hash-verified, not just version-pinned - our existing pkg==X.Y.Z pins
(and even pip install -r requirements-ci.txt, despite that file already
carrying hashes) still scored a 4 because the pin/hash isn't visible on
the command line itself.
Adds .github/requirements/*.txt: hash-locked files generated via
`uv pip compile --generate-hashes` for every pip target that isn't
already requirements-ci.txt, covering standalone CI tooling (build,
wheel, twine, uv, pip-audit, safety/bandit/semgrep/jq, pip/setuptools
bootstrap) and the project's own local-source installs. The latter
(`pip install -e ".[explorer]"`, `pip install -e .`) can't be hash-pinned
directly since there's nothing to hash for a local source tree; split
into `pip install --no-deps -e .` plus a separate hash-pinned install of
the actual fetched dependencies instead.
Also adds --require-hashes to every `-r requirements-ci.txt` install so
hash verification is enforced explicitly rather than only implied by the
file's own content.
Simplifies the Dockerfile in the process: it now installs from the same
pre-generated explorer-extra.txt (copied in at build time) instead of
extracting constraints from requirements-ci.txt at build time, which
also means setuptools gets its CVE-2025-47273 fix as a side effect of
the hash-pinned install rather than a separate upgrade step.
benchmark.yml: pip install -r benchmarks/requirements.txt is left
unpinned - that directory doesn't exist in this repo, so there's nothing
to generate hashes from. Pre-existing breakage, unrelated to this change.
Two terrascan/GHAS findings on the previous commit:
- AC_DOCKER_0052 (no apt-get upgrade in Dockerfiles): dropped it. It also
wasn't fixing anything - Debian's openssl fix for CVE-2026-14456 is still
in trixie-proposed-updates, not reachable via a normal upgrade. Pin both
base images by digest instead (matches #1329's approach) so the docker
Dependabot ecosystem bumps them once Debian ships a rebuilt image with
the fix, and document why the QUIC DoS isn't reachable here regardless
(HTTP-only via uvicorn).
- AC_DOCKER_0010 (pin pip package versions): setuptools was `>=78.1.1`;
pinned to the exact 84.0.0 already used by pyproject.toml/requirements-ci.txt.
Also fixes two build breaks this introduces on its own: requirements-ci.txt
wasn't in .dockerignore's allowlist or container-scan.yml's path trigger,
so the COPY in the prior commit would have failed the image build outright.
The sed expression to strip requirements-ci.txt's line-continuation
backslash (`[\]$`) is valid POSIX/GNU sed - verified it exits 0 and
strips correctly - but it's easy to misread as broken (a bot reviewer
flagged it as an unterminated bracket expression), and the seemingly
more obvious `\$`/` \$` forms silently fail to match at all rather
than erroring. Swap to a small `re.findall` extraction so there's no
backslash-escaping judgment call left for a reader (bot or human) to
second-guess.
Container Security Scan flagged five HIGH-severity findings against
semantica:scan:
- setuptools 70.3.0 (CVE-2025-47273, path traversal) - the base image's
bundled copy, never touched by our own build. Upgraded explicitly.
- msgpack 1.1.2 (GHSA-6v7p-g79w-8964, OOB read/crash) - `pip install
".[explorer]"` re-resolved deps from scratch instead of reusing the
audited, hash-pinned requirements-ci.txt (which already pins
msgpack==1.2.1), so it landed on an unpatched transitive version. Now
installs against a constraints file derived from requirements-ci.txt.
- openssl / libssl3t64 / openssl-provider-legacy (CVE-2026-14456, QUIC
server DoS) - the Debian fix is still in trixie-proposed-updates, not
yet promoted to trixie-security, so it can't be pulled via apt today.
Added an apt upgrade step so the next image rebuild picks it up
automatically once Debian ships it; documented why this image isn't
actually exposed to it in the meantime (HTTP-only via uvicorn, no QUIC
listener).
* fix(ci): unblock py3.9 install matrix and raise Scorecard pinning/signing
pip install semantica failed on Python 3.9 across all three OSes because
spacy had no upper bound, so pip resolved spacy 3.8.16 whose thinc>=8.3.12
requirement has no cp39 wheels and no working sdist build path. Cap
spacy/thinc for python_version < '3.10' to the last wheel-compatible pair.
Also addresses the two OpenSSF Scorecard findings that were actually
fixable in code:
- Pinned-Dependencies: Dockerfile base images (node:26-alpine,
python:3.13-slim) were unpinned by digest; pin both, and pin five
previously-unversioned pip install calls in CI (build, safety, bandit,
semgrep, jq, pip-audit).
- Signed-Releases: attest-build-provenance only publishes to the GH
attestations API, which Scorecard doesn't inspect. Sign dist/* with
Sigstore and attach the .sigstore.json bundles as release assets.
* fix(ci): correct Sigstore artifact inputs
---------
Co-authored-by: Sameer6305 <sskadam6305@gmail.com>
* ci: add npm Dependabot ecosystem and container image scanning
- dependabot.yml had no npm ecosystem entry for explorer/, so its
lockfile was never watched - exactly why the brace-expansion/nanoid
CVEs fixed in #1280 went undetected. Add it, mirroring the existing
pip entry's schedule/labels/reviewer conventions.
- New container-scan.yml builds the root Dockerfile's image and scans
it with Trivy (CRITICAL/HIGH OS+lib CVEs, SARIF to the Security tab)
and Syft (SPDX SBOM artifact), on push to main, weekly, and manual
dispatch. Neither the base-image scan nor an SBOM existed before -
Dependabot's docker entry only bumps the base image tag, it doesn't
scan built layers.
- Trivy runs report-only for now (no exit-code gate): this is its
first run against the image, so the CRITICAL/HIGH baseline hasn't
been triaged yet. Once reviewed, add exit-code: '1' to make it a
hard gate, same as Safety/Bandit-HIGH in security-scan.yml.
* fix: run Trivy via digest-pinned image, not the aquasecurity/trivy-action wrapper
verify-action-pins.sh failed in CI: the aquasecurity GitHub org has an IP
allow list on its API that 403s the live tag->SHA resolution from
Actions-runner IPs (confirmed reproducible, not transient - resolves fine
from a non-blocked host). Rather than carve a skip exception into the pin
verifier for an org this script already flags as a past tag-repointing
target (see its "LiteLLM/Trivy 2026 incident" comment), pull Trivy as a
sha256-digest-pinned Docker Hub image instead. A digest is immutable and
verifiable independently of GitHub's API entirely, so it sidesteps the
IP block without weakening verification of the one action this repo
already treats as higher-risk. Confirmed the pinned digest
(aquasec/trivy@sha256:62b1e65e...) resolves live against Docker Hub's
registry API.
* fix: match container-scan.yml's push paths to what actually reaches the image
The path filter only watched explorer/package.json and package-lock.json,
but Dockerfile COPYs the whole explorer/ tree plus README.md, LICENSE, and
MANIFEST.in, and .dockerignore controls all of it. A frontend source change
or a README/LICENSE edit would change the built image without triggering a
scan, silently drifting until the next weekly run. Replace the filter with
exactly .dockerignore's opt-in list.
- Bump explorer's brace-expansion (minimatch dep) 5.0.8 -> 5.0.9 and
nanoid (postcss dep) 3.3.16 -> 3.3.18, fixing GHSA-rgw5-rvv9-x895 and
GHSA-2v37-7h3g-55p8 (both DoS via unbounded input, both within the
existing caret ranges declared by their parents).
- Move codeql.yml and defender-for-devops.yml's security-events: write
(and codeql.yml's actions: read) from workflow-level down to their
single job, matching Scorecard's Token-Permissions ideal of a
read-only top-level default with sensitive scopes granted only where
used.
Distribution and trust-signal infrastructure to make pip install semantica
frictionless in downstream CI, and to bring the release pipeline in line
with mature OSS practice.
- .github/actions/setup-semantica: reusable composite action other repos
can call to install + verify semantica in one step
- install-matrix.yml: verifies the published package installs and imports
cleanly across Ubuntu/macOS/Windows x Python 3.9-3.12, weekly and on
release; backs a new README badge
- scorecard.yml: OpenSSF Scorecard analysis, weekly and on push to main,
backing a new README badge
- release.yml: twine check gate before publish, catching a broken PyPI
long-description render before it ships
- CITATION.cff: enables GitHub's native "Cite this repository" button
- examples/ci/: copy-paste GitHub Actions, GitLab CI, and CircleCI
templates for projects adopting semantica
- GROWTH.md: tracked checklist of distribution channels, what's done vs
outstanding, with guardrails against inflating metrics artificially
Fixes folded in along the way:
- Re-pinned softprops/action-gh-release to the immutable v3.0.3 tag
instead of the floating v3, after verify-action-pins.sh caught the
mutable tag had drifted to a newer commit
- setup-semantica now passes extras/version through env vars instead of
interpolating ${{ inputs.* }} directly into the bash script, closing
a script-injection vector for callers deriving these from event data
- install-matrix now triggers on the Release workflow's completion
(workflow_run) instead of release: published, since the GitHub release
is created before the PyPI upload runs and the old trigger could race
the publish
- The workflow_run path derives the expected version from the triggering
tag and passes it into setup-semantica's version input, so pip
installs and verifies the exact release instead of whatever's latest
on PyPI at the time
- setup-semantica's pip caching is now opt-in (default disabled), since
actions/setup-python errors out with cache: 'pip' enabled when the
caller repo has no requirements.txt/pyproject.toml to key on
- examples/ci/github-actions.yml pins actions/checkout and
actions/setup-python to verified commit SHAs instead of mutable tags
- examples/ci templates guard the requirements.txt install step with
-f requirements.txt and call out pyproject.toml/Poetry/Pipenv as
alternatives, since not every project has a requirements.txt
Both directories contain a test_degradation.py. Neither had an __init__.py,
so under pytest's default prepend import mode both modules were imported as
plain 'test_degradation' and the second collided with the first:
import file mismatch:
imported module 'test_degradation' has this __file__ attribute:
tests/integrations/crewai/test_degradation.py
which is not the same as the test file we want to collect:
tests/integrations/langchain/test_degradation.py
That aborted collection for tests/integrations/, so the langchain
graceful-degradation tests never ran. tests/integrations/__init__.py
already exists, and most directories under tests/ carry one; these two
subpackages were simply missed.
Collection goes from 335 collected, 1 error to 337 collected.
Closes#1251
The nodes/edges fetch loops were throwing away the response body whenever the request returned a non-OK status.
Because of that, errors like a `503` caused by a missing `SEMANTICA_API_KEY` only showed up as:
`Fetch failed: 503`
even though the backend was already returning a more useful message in the response `detail`.
This change reads the JSON error body and includes `detail` in the thrown error when it's a string, so `GraphLoadingOverlay` can show the actual backend error to the user.
Closes#1256
claude-3-sonnet-20240229 was retired 2025-07-21, so the wrapper's
default model and every copy-paste doc example would fail at
generate() time out of the box. Switch to claude-sonnet-4-6
everywhere (wrapper default, __init__ docstring, docs guide, tests).
Also cleans up leftover docstring typos/spacing from the previous
review pass and adds unavailable-path test coverage for
generate_structured()/generate_typed() to match generate(), clearing
ANTHROPIC_API_KEY in those tests so they don't flake on a runner that
has a real key set.
Reconciles this PR's Token-based alpha/beta matching (#300) with the
rule-actions/provenance layer merged separately in #1096. That PR built
bind_reasoner()/execute_matches() action-firing/_executed_activations/
reset_action_history() on top of the still-broken always-True stubs
(via an interim _bindings_for_rule() regex re-extraction), so main and
this branch touched the same propagation code with incompatible shapes.
Kept this branch's Token(facts, bindings) model for alpha/beta
propagation (the actual fix for #300) and layered main's action/
provenance plumbing on top of it, sourcing Match.bindings directly from
Token.bindings instead of re-deriving them with _bindings_for_rule(),
which is now redundant and removed. Also fixes a 2-tuple/3-tuple
unpacking break in test_matches_reasoner_match_rule caused by
Reasoner._match_rule()'s return shape changing upstream, and drops an
unrelated encoding-only .gitignore diff.
Verified: tests/reasoning/ (106 tests) and flake8 --max-line-length=88
both clean on the merged tree.
Importing a submodule rebinds it as an attribute of its parent package,
so restoring only the sys.modules entry left
semantica.utils.progress_tracker and
sys.modules['semantica.utils.progress_tracker'] pointing at different
objects for every test that ran afterwards.
Addresses review feedback on #1254.
ConsoleProgressDisplay wrote every progress frame to sys.stdout. Progress
is diagnostic output, so stderr is the correct stream for it — tqdm and
most progress renderers default there for the same reason — and stdout
must stay clean for programs that carry a machine-readable protocol on
it. The stdio MCP servers put newline-delimited JSON-RPC on stdout, where
an interleaved progress bar makes a response body unparseable (#1134).
ConsoleProgressDisplay now takes an optional stream, defaulting to
stderr. The stream is resolved per write rather than captured at
construction, so a later rebinding of sys.stderr (pytest capture, for
instance) is honoured. All writes and the four bare flushes route through
it, and the emoji-capability probe now inspects that stream rather than
stdout, so a cp1252 stderr still degrades correctly.
The existing cp1252 tests in tests/deduplication/test_deduplication.py
patched sys.stdout to assert emoji auto-disabling; they now patch the
stream progress is actually written to. Their intent is unchanged.
Closes#1134 (point 1 only; the SEMANTICA_KG_PATH persistence and README
items remain with @akaszubski)
This module landed after the branch was opened and imports
semantica.explorer.app at module scope, so it reproduced the same
collection error on a clean [dev] install.
* fix(memory): find_by_entity returns all matches by default (limit=None, not 10)
* Address review: move find_by_entity tests to the AgentMemory area
The regression tests lived in tests/test_seed_manager.py, mixing unrelated
domains. Moved to tests/context/test_agent_memory_find_by_entity.py with a
shared fixture; the unbounded default itself is unchanged and deliberate —
it IS the fix (#1018): an erasure workflow computing what references an
entity cannot paginate, so silently truncating at 10 left live references
behind. Callers that want a page pass an explicit limit.
---------
Co-authored-by: Sameer Kadam <sskadam6305@gmail.com>
feat(ingest): add Salesforce ingestor
Adds first-class Salesforce ingestion support, following the existing
Connector + Data + Ingestor architecture already used by the
Snowflake and Databricks integrations: SalesforceConnector /
SalesforceData / SalesforceIngestor, exposed lazily from
semantica.ingest so the base install stays unaffected.
SalesforceConnector supports both auth landscapes Salesforce actually
uses in practice: username + password + security token (SOAP login,
on-prem/sandbox), and session_id + instance_url for reusing an
existing authenticated session. Production and sandbox are selected
through domain, credentials can come from environment variables, and
the connector never intentionally puts credential material into logs,
exceptions, or its own repr.
SalesforceIngestor covers ingest_sobject(), ingest_query(),
list_sobjects(), get_sobject_schema(), and export_as_documents(),
against standard sObjects, custom objects (__c), custom metadata
objects (__mdt), platform events (__e), namespaced objects, and
relationship-field traversal (Owner.Name). Pagination follows
nextRecordsUrl/query_more() automatically and stops once a caller's
limit is satisfied rather than continuing to fetch full pages past it.
Dynamically constructed SOQL is validated before it's sent: sObject
names, field names, relationship paths, ORDER BY expressions, and
numeric limits are checked, and WHERE fragments are screened against
common injection primitives after masking quoted string literals so a
value like status = 'union' doesn't false-positive. Raw SOQL passed
directly to ingest_query() stays intentionally caller-controlled,
since that method is documented as the advanced/unvalidated escape
hatch.
Salesforce-specific attributes metadata is stripped from returned
records before they're handed to the rest of the pipeline, while
relationship data, normal field values, and datetime normalization
are preserved. export_as_documents() uses the Salesforce Id as the
stable document identifier and keeps the source record in document
metadata for provenance.
Wired into the unified ingestion API via ingest_salesforce() and
ingest(source_type="salesforce", ...), registered with
MethodRegistry under sobject/query/list_sobjects/schema/documents.
Isolated behind the semantica[db-salesforce] extra
(simple-salesforce>=1.12.0), included in db-all.
JWT Bearer authentication and Bulk API 2.0 are intentionally out of
scope for this first connector; both are documented as deliberate
follow-ups rather than gaps.
fix(ingest): address Salesforce review findings
- limit now validates as a non-negative integer before use; negative,
string, and float values raise ValidationError instead of silently
returning an empty result, raising a bare TypeError, or building an
invalid LIMIT 0 query
- fields is validated as a non-empty list of strings; a bare string
(e.g. "Id") no longer gets iterated character-by-character into
nonsense field names, and an empty list no longer builds a
syntactically invalid SELECT
- the generic connection-failure path now raises with `from None`
instead of chaining the original exception, so credential or
request detail from the underlying library can't surface through a
traceback
- the unified ingest() dispatch no longer coerces a non-dict source
into None and silently falling back to environment credentials; an
invalid source now raises
- _validate_order_by rewritten to validate each dot-separated
component through _validate_field_name, rejecting malformed
fragments like "Name." or "Owner..Name" that the previous regex let
through
- CI conflicts from parallel merges resolved; upstream markdown
dependency changes preserved
test(ingest): add Salesforce JWT coverage
Adds construction and connect() coverage for the JWT Bearer auth path
(consumer_key + privatekey/privatekey_file), the one auth mode that
had no dedicated tests despite handling private key material.
Also removes _SAFE_ORDER_RE, left behind as dead code once
_validate_order_by was rewritten to use _validate_field_name per
component, and fixes a test-isolation leak where an earlier test left
SALESFORCE_AVAILABLE=True behind for a later test that expected it
False when simple-salesforce isn't installed.