mirror of
https://github.com/semantica-agi/semantica.git
synced 2026-08-29 04:26:20 +00:00
Adds DatabricksIngestor to semantica/ingest/, mirroring SnowflakeIngestor's structure and public API shape: table/query ingestion via databricks-sql-connector, Unity Catalog metadata and lineage via databricks-sdk, and export-as-documents for KG construction. Closes #747
4.0 KiB
4.0 KiB
title, description, icon
| title | description | icon |
|---|---|---|
| Snowflake Integration | Ingest structured data from Snowflake tables and queries into Semantica's KG pipeline. | snowflake |
Extract data from Snowflake into Semantica with password, key-pair, OAuth, and SSO authentication.
Installation
# Install with Snowflake support
pip install "semantica[db-snowflake]"
# Or install the connector separately
pip install snowflake-connector-python
Basic Usage
from semantica.ingest import SnowflakeIngestor
import os
ingestor = SnowflakeIngestor(
account=os.getenv("SNOWFLAKE_ACCOUNT"),
user=os.getenv("SNOWFLAKE_USER"),
password=os.getenv("SNOWFLAKE_PASSWORD"),
warehouse=os.getenv("SNOWFLAKE_WAREHOUSE"),
database=os.getenv("SNOWFLAKE_DATABASE"),
schema=os.getenv("SNOWFLAKE_SCHEMA"),
)
data = ingestor.ingest_table("CUSTOMERS")
print(f"Retrieved {data.row_count} rows: columns: {data.columns}")
Authentication Methods
```python ingestor = SnowflakeIngestor( account="myaccount", user="myuser", password="mypassword", warehouse="COMPUTE_WH", ) ``` ```python ingestor = SnowflakeIngestor( account="myaccount", user="myuser", private_key_path="/path/to/rsa_key.p8", warehouse="COMPUTE_WH", ) ``` Preferred for production: no password stored in config. ```python ingestor = SnowflakeIngestor( account="myaccount", user="myuser", authenticator="oauth", token="your_oauth_token", warehouse="COMPUTE_WH", ) ``` ```python ingestor = SnowflakeIngestor( account="myaccount", user="myuser", authenticator="externalbrowser", warehouse="COMPUTE_WH", ) ```Querying
Ingest a table with filters
data = ingestor.ingest_table(
"CUSTOMERS",
where="COUNTRY = 'USA' AND CREATED_DATE > '2024-01-01'",
order_by="CREATED_DATE DESC",
limit=10000,
)
Custom SQL
data = ingestor.ingest_query("""
SELECT CUSTOMER_ID, SUM(AMOUNT) AS TOTAL_AMOUNT
FROM SALES
WHERE DATE >= '2024-01-01'
GROUP BY CUSTOMER_ID
""")
Schema introspection
schema = ingestor.get_table_schema("CUSTOMERS")
for column in schema["columns"]:
print(f"{column['name']}: {column['type']}")
Export as Semantica Documents
documents = ingestor.export_as_documents(
data,
id_field="CUSTOMER_ID",
text_fields=["NAME", "EMAIL", "NOTES"],
)
print(f"Created {len(documents)} documents for processing")
Batch Processing Large Tables
PAGE_SIZE = 5000
for page in range(total_pages):
data = ingestor.ingest_table(
"LARGE_TABLE",
limit=PAGE_SIZE,
offset=page * PAGE_SIZE,
)
process_batch(data)
Or use the built-in batch_size parameter:
data = ingestor.ingest_query(
"SELECT * FROM LARGE_TABLE",
batch_size=5000,
)
Troubleshooting
from semantica.ingest import SnowflakeConnector
connector = SnowflakeConnector(account="myaccount", user="myuser", password="mypassword")
if not connector.test_connection():
print("Connection failed: check credentials and account identifier")
See Also
- Ingest Module — Full SnowflakeIngestor and all other ingestors.
- Databricks Integration — Companion connector for a Snowflake + Databricks hybrid estate.
- Pipeline — Use Snowflake ingestion as a pipeline step.
- Installation — All optional dependency extras.
- Knowledge Graph — Build a KG from ingested Snowflake data.