feat(ingest): add Salesforce ingestor Adds first-class Salesforce ingestion support, following the existing Connector + Data + Ingestor architecture already used by the Snowflake and Databricks integrations: SalesforceConnector / SalesforceData / SalesforceIngestor, exposed lazily from semantica.ingest so the base install stays unaffected. SalesforceConnector supports both auth landscapes Salesforce actually uses in practice: username + password + security token (SOAP login, on-prem/sandbox), and session_id + instance_url for reusing an existing authenticated session. Production and sandbox are selected through domain, credentials can come from environment variables, and the connector never intentionally puts credential material into logs, exceptions, or its own repr. SalesforceIngestor covers ingest_sobject(), ingest_query(), list_sobjects(), get_sobject_schema(), and export_as_documents(), against standard sObjects, custom objects (__c), custom metadata objects (__mdt), platform events (__e), namespaced objects, and relationship-field traversal (Owner.Name). Pagination follows nextRecordsUrl/query_more() automatically and stops once a caller's limit is satisfied rather than continuing to fetch full pages past it. Dynamically constructed SOQL is validated before it's sent: sObject names, field names, relationship paths, ORDER BY expressions, and numeric limits are checked, and WHERE fragments are screened against common injection primitives after masking quoted string literals so a value like status = 'union' doesn't false-positive. Raw SOQL passed directly to ingest_query() stays intentionally caller-controlled, since that method is documented as the advanced/unvalidated escape hatch. Salesforce-specific attributes metadata is stripped from returned records before they're handed to the rest of the pipeline, while relationship data, normal field values, and datetime normalization are preserved. export_as_documents() uses the Salesforce Id as the stable document identifier and keeps the source record in document metadata for provenance. Wired into the unified ingestion API via ingest_salesforce() and ingest(source_type="salesforce", ...), registered with MethodRegistry under sobject/query/list_sobjects/schema/documents. Isolated behind the semantica[db-salesforce] extra (simple-salesforce>=1.12.0), included in db-all. JWT Bearer authentication and Bulk API 2.0 are intentionally out of scope for this first connector; both are documented as deliberate follow-ups rather than gaps. fix(ingest): address Salesforce review findings - limit now validates as a non-negative integer before use; negative, string, and float values raise ValidationError instead of silently returning an empty result, raising a bare TypeError, or building an invalid LIMIT 0 query - fields is validated as a non-empty list of strings; a bare string (e.g. "Id") no longer gets iterated character-by-character into nonsense field names, and an empty list no longer builds a syntactically invalid SELECT - the generic connection-failure path now raises with `from None` instead of chaining the original exception, so credential or request detail from the underlying library can't surface through a traceback - the unified ingest() dispatch no longer coerces a non-dict source into None and silently falling back to environment credentials; an invalid source now raises - _validate_order_by rewritten to validate each dot-separated component through _validate_field_name, rejecting malformed fragments like "Name." or "Owner..Name" that the previous regex let through - CI conflicts from parallel merges resolved; upstream markdown dependency changes preserved test(ingest): add Salesforce JWT coverage Adds construction and connect() coverage for the JWT Bearer auth path (consumer_key + privatekey/privatekey_file), the one auth mode that had no dedicated tests despite handling private key material. Also removes _SAFE_ORDER_RE, left behind as dead code once _validate_order_by was rewritten to use _validate_field_name per component, and fixes a test-isolation leak where an earlier test left SALESFORCE_AVAILABLE=True behind for a later test that expected it False when simple-salesforce isn't installed.
11 KiB
title, description, icon
| title | description | icon |
|---|---|---|
| Salesforce Integration | Ingest CRM records from Salesforce sObjects and SOQL queries into Semantica's KG pipeline. | cloud |
Extract Accounts, Contacts, Opportunities, and custom objects from Salesforce into Semantica with username/password/security-token, JWT bearer, or session-based authentication.
Installation
# Install with Salesforce support
pip install "semantica[db-salesforce]"
# Or install the connector separately
pip install simple-salesforce>=1.12.0
Basic Usage
from semantica.ingest import SalesforceIngestor
import os
ingestor = SalesforceIngestor(
username=os.getenv("SALESFORCE_USERNAME"),
password=os.getenv("SALESFORCE_PASSWORD"),
security_token=os.getenv("SALESFORCE_SECURITY_TOKEN"),
domain=os.getenv("SALESFORCE_DOMAIN", "login"), # "test" for sandbox
)
data = ingestor.ingest_sobject("Account", fields=["Id", "Name", "Industry"], limit=1000)
print(f"Retrieved {data.row_count} of {data.total_size} matching records")
print(f"Columns: {data.columns}")
Authentication Methods
```python import os from semantica.ingest import SalesforceIngestoringestor = SalesforceIngestor(
username=os.getenv("SALESFORCE_USERNAME"),
password=os.getenv("SALESFORCE_PASSWORD"),
security_token=os.getenv("SALESFORCE_SECURITY_TOKEN"),
domain="login", # production; use "test" for sandbox
)
```
Set the required environment variables before running:
```bash
export SALESFORCE_USERNAME="your-username@example.com"
export SALESFORCE_PASSWORD="your-password"
export SALESFORCE_SECURITY_TOKEN="your-security-token"
```
The standard server-side flow. The security token is appended to the
password during Salesforce SOAP login. Generate or reset it under
**Settings → My Personal Information → Reset My Security Token**.
ingestor = SalesforceIngestor(
username=os.getenv("SALESFORCE_USERNAME"),
consumer_key=os.getenv("SALESFORCE_CONSUMER_KEY"),
privatekey_file=os.getenv("SALESFORCE_PRIVATE_KEY_FILE"),
domain="login", # or "test" for sandbox
)
```
```bash
export SALESFORCE_USERNAME="your-username@example.com"
export SALESFORCE_CONSUMER_KEY="your-connected-app-consumer-key"
export SALESFORCE_PRIVATE_KEY_FILE="/path/to/server.key"
```
The JWT bearer flow authenticates with a signed token — no password
is transmitted. Ideal for server-to-server integrations and CI/CD
pipelines. Requires a Salesforce connected app configured with
**Use digital signatures** and the pre-authorised user listed under
**Manage → Profiles / Permission Sets**.
If you prefer to pass the key material as a string instead of a file
path, use `SALESFORCE_PRIVATE_KEY` (the PEM contents) in place of
`SALESFORCE_PRIVATE_KEY_FILE`.
ingestor = SalesforceIngestor(
username=os.getenv("SALESFORCE_USERNAME"),
password=os.getenv("SALESFORCE_PASSWORD"),
security_token=os.getenv("SALESFORCE_SECURITY_TOKEN"),
domain="test", # routes to test.salesforce.com
)
```
```bash
export SALESFORCE_USERNAME="your-sandbox-username@example.com.sandbox"
export SALESFORCE_PASSWORD="your-password"
export SALESFORCE_SECURITY_TOKEN="your-security-token"
export SALESFORCE_DOMAIN="test"
```
Replace `domain="login"` with `domain="test"` (or set
`SALESFORCE_DOMAIN=test` in your environment) to connect to a
developer or full sandbox.
Environment variables
All constructor parameters have environment-variable fallbacks:
| Variable | Parameter | Default |
|---|---|---|
SALESFORCE_USERNAME |
username |
— |
SALESFORCE_PASSWORD |
password |
— |
SALESFORCE_SECURITY_TOKEN |
security_token |
— |
SALESFORCE_DOMAIN |
domain |
"login" |
SALESFORCE_INSTANCE_URL |
instance_url |
— |
SALESFORCE_SESSION_ID |
session_id |
— |
SALESFORCE_CONSUMER_KEY |
consumer_key |
— |
SALESFORCE_PRIVATE_KEY_FILE |
privatekey_file |
— |
SALESFORCE_PRIVATE_KEY |
privatekey |
— |
SALESFORCE_API_VERSION |
api_version |
library default (59.0) |
Object Ingestion
Ingest a standard object
data = ingestor.ingest_sobject(
"Account",
fields=["Id", "Name", "Industry", "AnnualRevenue", "BillingCity"],
where="Type = 'Customer' AND AnnualRevenue > 1000000",
order_by="Name ASC",
limit=5000,
)
print(f"Retrieved {data.row_count} of {data.total_size} matching records")
Ingest a custom object
Custom objects end with __c in their API name:
data = ingestor.ingest_sobject(
"My_Custom_Object__c",
fields=["Id", "Name", "Custom_Field__c"],
)
Relationship traversal fields (Owner.Name) are also supported:
data = ingestor.ingest_sobject(
"Contact",
fields=["Id", "Name", "Email", "Account.Name", "Owner.Name"],
limit=10000,
)
Let Semantica choose the fields
When fields is omitted, all selectable fields are fetched via describe()
(one extra API call). Compound address and geolocation fields (type=address,
type=location) are automatically excluded — select their components
(BillingStreet, BillingCity, Location__Latitude__s, …) individually if
you need them.
data = ingestor.ingest_sobject("Opportunity")
Raw SOQL Ingestion
Pass any valid SOQL query verbatim — pagination is handled automatically:
data = ingestor.ingest_query("""
SELECT Id, Name, StageName, Amount, CloseDate,
Account.Name, Owner.Name
FROM Opportunity
WHERE IsClosed = false
ORDER BY CloseDate ASC
""")
print(f"Open opportunities: {data.row_count}")
The query is passed to the Salesforce REST API unchanged. The caller is responsible for SOQL correctness and safety.
`ingest_query` does not validate or sanitise the SOQL string. Use `ingest_sobject` (which validates sObject names, field names, and WHERE/ORDER BY fragments) when building queries from application-controlled inputs.Document Export
Convert ingested records to the Semantica document format for use with
GraphBuilder:
documents = ingestor.export_as_documents(
data,
id_field="Id", # default; Salesforce 18-char record Id
text_fields=["Name", "Description"], # omit to join all string fields
)
print(f"Created {len(documents)} documents")
# Each document:
# {
# "id": "001xx000003GYk2AAG",
# "text": "Acme Corp Enterprise software company",
# "metadata": {
# "source": "salesforce",
# "sobject": "Account",
# "instance_url": "https://myorg.my.salesforce.com",
# "row_data": { ... full cleaned record ... }
# }
# }
Feed the documents directly into GraphBuilder:
from semantica.kg import GraphBuilder
builder = GraphBuilder()
kg = builder.build(documents)
Object and Schema Discovery
# List all accessible sObjects
sobject_names = ingestor.list_sobjects()
print(sobject_names[:10]) # ["Account", "Case", "Contact", ...]
# Inspect fields for a specific sObject
schema = ingestor.get_sobject_schema("Account")
for field in schema["fields"]:
print(f"{field['name']}: {field['type']} (nillable={field['nillable']})")
Context Manager
Prefer the context manager for long-running jobs — it opens one connection on
entry and closes it on exit, so every ingestion call inside the with block
reuses the same authenticated session:
with SalesforceIngestor(
username=os.getenv("SALESFORCE_USERNAME"),
password=os.getenv("SALESFORCE_PASSWORD"),
security_token=os.getenv("SALESFORCE_SECURITY_TOKEN"),
) as sf:
accounts = sf.ingest_sobject("Account", limit=10000)
contacts = sf.ingest_sobject("Contact", limit=10000)
sobjects = sf.list_sobjects()
Convenience Function
Use ingest_salesforce() for one-liner ingestion:
from semantica.ingest import ingest_salesforce
# Fetch records
data = ingest_salesforce(
method="sobject",
sobject_name="Account",
fields=["Id", "Name", "Industry"],
limit=500,
)
# Execute raw SOQL (credentials from environment variables)
data = ingest_salesforce(
method="query",
soql="SELECT Id, Name FROM Contact WHERE IsActive = true",
)
# Ingest + export to documents in one step
docs = ingest_salesforce(
method="documents",
sobject_name="Account",
text_fields=["Name", "Description"],
limit=1000,
)
# List accessible sObjects
sobject_names = ingest_salesforce(method="list_sobjects")
Or use the unified ingest() dispatcher:
from semantica.ingest import ingest
result = ingest(
None,
source_type="salesforce",
method="sobject",
sobject_name="Account",
fields=["Id", "Name"],
limit=500,
)
data = result["data"] # SalesforceData
Troubleshooting
import os
from semantica.ingest import SalesforceConnector
connector = SalesforceConnector(
username=os.getenv("SALESFORCE_USERNAME"),
password=os.getenv("SALESFORCE_PASSWORD"),
security_token=os.getenv("SALESFORCE_SECURITY_TOKEN"),
)
if not connector.test_connection():
print("Connection failed: check username, password, security token, and domain")
Common causes of authentication failures:
- Wrong domain: production orgs use
domain="login"; sandboxes usedomain="test". - Stale security token: reset it under Settings → Reset My Security Token. The new token is emailed to you.
- IP restriction: your org's trusted IP ranges may block the originating IP. Check Setup → Network Access.
- API access disabled: ensure the connected profile has the API Enabled permission.
See Also
- Ingest Module — Full
SalesforceIngestorAPI and all other ingestors. - Snowflake Integration — Relational warehouse connector with a similar design.
- Databricks Integration — Lakehouse connector.
- Installation — All optional dependency extras.
- Knowledge Graph — Build a KG from ingested Salesforce data.