mirror of
https://github.com/semantica-agi/semantica.git
synced 2026-09-04 04:01:07 +00:00
Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
4e8dfb72e9 |
@@ -26,7 +26,7 @@ each file's own autogenerated header comment for its exact command).
|
|||||||
|
|
||||||
| File | Used by | Installs |
|
| File | Used by | Installs |
|
||||||
| --- | --- | --- |
|
| --- | --- | --- |
|
||||||
| `bootstrap.txt` | security-scan.yml, benchmark.yml | pip, setuptools (upgrade before anything else) |
|
| `bootstrap.txt` | security.yml, security-scan.yml, benchmark.yml | pip, setuptools (upgrade before anything else) |
|
||||||
| `pep517-build.txt` | ci.yml, benchmark.yml, Dockerfile | exact `[build-system] requires` from `pyproject.toml` (setuptools, wheel) - installed with `--no-build-isolation` before any `pip install -e .` / `pip install .`, since `--no-deps` alone doesn't stop pip's PEP 517 build isolation from fetching those two *unhashed* |
|
| `pep517-build.txt` | ci.yml, benchmark.yml, Dockerfile | exact `[build-system] requires` from `pyproject.toml` (setuptools, wheel) - installed with `--no-build-isolation` before any `pip install -e .` / `pip install .`, since `--no-deps` alone doesn't stop pip's PEP 517 build isolation from fetching those two *unhashed* |
|
||||||
| `explorer-extra-py311.txt` | ci.yml | semantica's base deps + the `explorer` extra, resolved for python 3.11 |
|
| `explorer-extra-py311.txt` | ci.yml | semantica's base deps + the `explorer` extra, resolved for python 3.11 |
|
||||||
| `explorer-extra-py313.txt` | Dockerfile | the same, resolved for python 3.13 (the image's actual interpreter) |
|
| `explorer-extra-py313.txt` | Dockerfile | the same, resolved for python 3.13 (the image's actual interpreter) |
|
||||||
@@ -34,8 +34,8 @@ each file's own autogenerated header comment for its exact command).
|
|||||||
| `uv-tool.txt` | ci.yml | uv, to verify requirements-ci.txt is current |
|
| `uv-tool.txt` | ci.yml | uv, to verify requirements-ci.txt is current |
|
||||||
| `build-tools.txt` | ci.yml, release.yml | build, wheel |
|
| `build-tools.txt` | ci.yml, release.yml | build, wheel |
|
||||||
| `twine.txt` | release.yml | twine |
|
| `twine.txt` | release.yml | twine |
|
||||||
| `pip-audit.txt` | security-scan.yml | pip-audit |
|
| `pip-audit.txt` | security.yml | pip-audit |
|
||||||
| `security-scan-tools.txt` | security-scan.yml | bandit, semgrep, jq |
|
| `security-scan-tools.txt` | security-scan.yml | safety, bandit, semgrep, jq |
|
||||||
| `base-deps.txt` | benchmark.yml | semantica's base deps (no extras) |
|
| `base-deps.txt` | benchmark.yml | semantica's base deps (no extras) |
|
||||||
| `benchmark-extra.txt` | benchmark.yml | the benchmark-only libs (neo4j, pdfplumber, etc.) |
|
| `benchmark-extra.txt` | benchmark.yml | the benchmark-only libs (neo4j, pdfplumber, etc.) |
|
||||||
|
|
||||||
|
|||||||
@@ -11,4 +11,3 @@ python-docx
|
|||||||
beautifulsoup4
|
beautifulsoup4
|
||||||
chardet
|
chardet
|
||||||
langdetect
|
langdetect
|
||||||
en-core-web-sm @ https://github.com/explosion/spacy-models/releases/download/en_core_web_sm-3.8.0/en_core_web_sm-3.8.0-py3-none-any.whl
|
|
||||||
|
|||||||
@@ -408,9 +408,6 @@ cuda-toolkit==13.0.3.0 \
|
|||||||
# via
|
# via
|
||||||
# -c requirements-ci.txt
|
# -c requirements-ci.txt
|
||||||
# torch
|
# torch
|
||||||
en-core-web-sm @ https://github.com/explosion/spacy-models/releases/download/en_core_web_sm-3.8.0/en_core_web_sm-3.8.0-py3-none-any.whl \
|
|
||||||
--hash=sha256:1932429db727d4bff3deed6b34cfc05df17794f4a52eeb26cf8928f7c1a0fb85
|
|
||||||
# via -r .github/requirements/benchmark-extra.in
|
|
||||||
et-xmlfile==2.0.0 \
|
et-xmlfile==2.0.0 \
|
||||||
--hash=sha256:7a91720bc756843502c3b7504c77b8fe44217c85c537d85037f0f536151b2caa \
|
--hash=sha256:7a91720bc756843502c3b7504c77b8fe44217c85c537d85037f0f536151b2caa \
|
||||||
--hash=sha256:dab3f4764309081ce75662649be815c4c9081e88f0837825f90fd28317d4da54
|
--hash=sha256:dab3f4764309081ce75662649be815c4c9081e88f0837825f90fd28317d4da54
|
||||||
|
|||||||
@@ -1 +1 @@
|
|||||||
checkov==3.3.16
|
checkov==3.3.1
|
||||||
|
|||||||
+200
-194
@@ -8,126 +8,127 @@ aiohappyeyeballs==2.7.1 \
|
|||||||
--hash=sha256:065665c041c42a5938ed220bdcd7230f22527fbec085e1853d2402c8a3615d9d \
|
--hash=sha256:065665c041c42a5938ed220bdcd7230f22527fbec085e1853d2402c8a3615d9d \
|
||||||
--hash=sha256:9243213661e29250eb41368e5daa826fc017156c3b8a11440826b2e3ed376472
|
--hash=sha256:9243213661e29250eb41368e5daa826fc017156c3b8a11440826b2e3ed376472
|
||||||
# via aiohttp
|
# via aiohttp
|
||||||
aiohttp==3.14.3 \
|
aiohttp==3.13.5 \
|
||||||
--hash=sha256:03cd2bde3d7f085b64e549c985f4bb928cad7e8ecf5323bfca320db548d81b39 \
|
--hash=sha256:019a67772e034a0e6b9b17c13d0a8fe56ad9fb150fc724b7f3ffd3724288d9e5 \
|
||||||
--hash=sha256:041badb8f84396357c4d3ad26de6afd7a32b112f43d3c63045c0c8278cfd2043 \
|
--hash=sha256:02222e7e233295f40e011c1b00e3b0bd451f22cf853a0304c3595633ee47da4b \
|
||||||
--hash=sha256:0a5ff2dfbb9ce645fa5b8ef3e02c6c0b9cc3f6030ff863d0c51fffc50cb5541b \
|
--hash=sha256:023ecba036ddd840b0b19bf195bfae970083fd7024ce1ac22e9bba90464620e9 \
|
||||||
--hash=sha256:0fdea2281997af69da84c77ffa6f5938a0285f21fb3887c249d67419ca865b3d \
|
--hash=sha256:02e048037a6501a5ec1f6fc9736135aec6eb8a004ce48838cb951c515f32c80b \
|
||||||
--hash=sha256:11fb37ef075669eee52ab1928fbf6e1741fada40409fa309ebde9607a962aebf \
|
--hash=sha256:0494a01ca9584eea1e5fbd6d748e61ecff218c51b576ee1999c23db7066417d8 \
|
||||||
--hash=sha256:134ac5ddcf61c6fad984b9a5727d83492ada43d63471db20fb73042c13fca62f \
|
--hash=sha256:0f7a18f258d124cd678c5fe072fe4432a4d5232b0657fca7c1847f599233c83a \
|
||||||
--hash=sha256:152516815ef926786a0b6ae2b8f1fd2e0c71582dee0b435636865316fd4891b7 \
|
--hash=sha256:10a75acfcf794edf9d8db50e5a7ec5fc818b2a8d3f591ce93bc7b1210df016d2 \
|
||||||
--hash=sha256:1576145bdceeb92382d899751e12743a3a5b8e460a841e3e50543859e54864dc \
|
--hash=sha256:110e448e02c729bcebb18c60b9214a87ba33bac4a9fa5e9a5f139938b56c6cb1 \
|
||||||
--hash=sha256:16100ad3ab8d649fdfbee87602d9d2dcdca9df0b9eda8a1b5fdc0d41f96da559 \
|
--hash=sha256:147b4f501d0292077f29d5268c16bb7c864a1f054d7001c4c1812c0421ea1ed0 \
|
||||||
--hash=sha256:16ea7e24c309fb7c0bbd505d149abe4fe4dccfb8db911db7dbec0921bc889a6f \
|
--hash=sha256:157826e2fa245d2ef46c83ea8a5faf77ca19355d278d425c29fda0beb3318037 \
|
||||||
--hash=sha256:18c441d0a8fca6de8d1f546849b9f0ab20d435993e2c5b59562b2fae6be2f929 \
|
--hash=sha256:15c933ad7920b7d9a20de151efcd05a6e38302cbf0e10c9b2acb9a42210a2416 \
|
||||||
--hash=sha256:18cb43369747b2ae007bd2655fb8e63a099c2ff1d207962943636dac989b3147 \
|
--hash=sha256:178c7b5e62b454c2bc790786e6058c3cc968613b4419251b478c153a4aec32b1 \
|
||||||
--hash=sha256:1b59533861b70a2185c8f4f350f791f39d64358ef6944ce71c5240c9ec0982c9 \
|
--hash=sha256:18a2f6c1182c51baa1d28d68fea51513cb2a76612f038853c0ad3c145423d3d9 \
|
||||||
--hash=sha256:1c5281acc88b92396f88c7e1e2748f8466689df22b80170e4f51efa712fb47a8 \
|
--hash=sha256:1efb06900858bb618ff5cee184ae2de5828896c448403d51fb633f09e109be0a \
|
||||||
--hash=sha256:1c5ec8fb1bcc31a8466f74aaf26c345d5c386fa4bd08a3f0eb9c7a4a3fe8b5bf \
|
--hash=sha256:20058e23909b9e65f9da62b396b77dfa95965cbe840f8def6e572538b1d32e36 \
|
||||||
--hash=sha256:1caa7b0d05f3e3a36f87788c59e970a7ee1cefcfcbb924a9f138c4a6551c9cb7 \
|
--hash=sha256:206b7b3ef96e4ce211754f0cd003feb28b7d81f0ad26b8d077a5d5161436067f \
|
||||||
--hash=sha256:21c016079415ed3fd676963e9793700a566d85dbbd6bfc564b9b2d209147dcc8 \
|
--hash=sha256:20ae0ff08b1f2c8788d6fb85afcb798654ae6ba0b747575f8562de738078457b \
|
||||||
--hash=sha256:2498f0fe69ead802f9675beca44a7c21c62fdaa4ec5145ea1c3ad6edbee29f85 \
|
--hash=sha256:2294172ce08a82fb7c7273485895de1fa1186cc8294cfeb6aef4af42ad261174 \
|
||||||
--hash=sha256:25bd2708db6bdf6a6630dd37bdcdfcb47c4434d22ac69c64665b802910140b30 \
|
--hash=sha256:241a94f7de7c0c3b616627aaad530fe2cb620084a8b144d3be7b6ecfe95bae3b \
|
||||||
--hash=sha256:270d3dace9ca2f10f0da5d8ebe519b7a310fc6112ed916e32df5866df0888553 \
|
--hash=sha256:26d2f8546f1dfa75efa50c3488215a903c0168d253b75fba4210f57ab77a0fb8 \
|
||||||
--hash=sha256:2e1161602f45a54de2ce0905243a95f58cb42dcd378402f3697f5e0b21e9d2e7 \
|
--hash=sha256:2837fb92951564d6339cedae4a7231692aa9f73cbc4fb2e04263b96844e03b4e \
|
||||||
--hash=sha256:2e9878ae68e4a5f1c0abe4dd497dbc3d51946f5837b56759e2a02e78fa90ef86 \
|
--hash=sha256:2994be9f6e51046c4f864598fd9abeb4fba6e88f0b2152422c9666dcd4aea9c6 \
|
||||||
--hash=sha256:30402d03a7c0ff52bce290b57e564e9079fd9d0cb545c8aba73f86a103162d2e \
|
--hash=sha256:2d6d44a5b48132053c2f6cd5c8cb14bc67e99a63594e336b0f2af81e94d5530c \
|
||||||
--hash=sha256:33a2d7c28d33797a2e99923dffa63f83d908a19b6bf26cfe80fa790aa5e1a75a \
|
--hash=sha256:31cebae8b26f8a615d2b546fee45d5ffb76852ae6450e2a03f42c9102260d6fe \
|
||||||
--hash=sha256:362a3fd481769cac1a824514bcd86fda51c65e8fe6e051099e008fddde6db17c \
|
--hash=sha256:327cc432fdf1356fb4fbc6fe833ad4e9f6aacb71a8acaa5f1855e4b25910e4a9 \
|
||||||
--hash=sha256:38901a84da3ce22249f6e860bf8f90d141bcab7da090cc398f8bb58c0e44b7da \
|
--hash=sha256:329f292ed14d38a6c4c435e465f48bebb47479fd676a0411936cc371643225cc \
|
||||||
--hash=sha256:39aded8c7f3b935b54aab1d8d73c70ec0ee2d3ec3b943e0e86611bc150ba47f5 \
|
--hash=sha256:330f5da04c987f1d5bdb8ae189137c77139f36bd1cb23779ca1a354a4b027800 \
|
||||||
--hash=sha256:3a26434dafe408229ff3403458ca58de24fb51936504decac49ce6755f77e59d \
|
--hash=sha256:33add2463dde55c4f2d9635c6ab33ce154e5ecf322bd26d09af95c5f81cfa286 \
|
||||||
--hash=sha256:3ae5b3a59436d089b5395d910121a390feed4d00578eb95a0fd1a329fe963100 \
|
--hash=sha256:347542f0ea3f95b2a955ee6656461fa1c776e401ac50ebce055a6c38454a0adf \
|
||||||
--hash=sha256:3d4f72af88ac2474bb5bca640030320e3d38a0163a1d7533500e87be458eef71 \
|
--hash=sha256:39380e12bd1f2fdab4285b6e055ad48efbaed5c836433b142ed4f5b9be71036a \
|
||||||
--hash=sha256:3f42e9b78301f11c8f861746175d8b9c1ccef713fcad9eab396e2f6db8ed4a22 \
|
--hash=sha256:3a807cabd5115fb55af198b98178997a5e0e57dead43eb74a93d9c07d6d4a7dc \
|
||||||
--hash=sha256:42a67efc36300d052fb4508a53e8b6901b9284b599ae63945c377569c5fcc1e1 \
|
--hash=sha256:3b13560160d07e047a93f23aaa30718606493036253d5430887514715b67c9d9 \
|
||||||
--hash=sha256:48d67b87db6279c044760787eb01f6413032c2e6f3ba1cafaa492b1c8e578479 \
|
--hash=sha256:3df334e39d4c2f899a914f1dba283c1aadc311790733f705182998c6f7cae665 \
|
||||||
--hash=sha256:498c6c623134f8e09a3c4e60bcd607a0b4590dd7dbf08dd40851b27cbb520ccb \
|
--hash=sha256:4bb6bf5811620003614076bdc807ef3b5e38244f9d25ca5fe888eaccea2a9832 \
|
||||||
--hash=sha256:49f7325beb0f85ef4aef5f48f490269575f83e6e2acad00a1d80b807eb027062 \
|
--hash=sha256:4beac52e9fe46d6abf98b0176a88154b742e878fdf209d2248e99fcdf73cd297 \
|
||||||
--hash=sha256:4e3ac92d90e92773b2362d506068e9a948192bd553e743c5b2429e28527c8661 \
|
--hash=sha256:4e704c52438f66fdd89588346183d898bb42167cf88f8b7ff1c0f9fc957c348f \
|
||||||
--hash=sha256:530125ee1163c4219af35dc3aa1206e541e7b31b6efc1a3f93b70a136f65d427 \
|
--hash=sha256:4eac02d9af4813ee289cd63a361576da36dba57f5a1ab36377bc2600db0cbb73 \
|
||||||
--hash=sha256:5373dc80ad1aa2fb9ad95c83f24eef418bbda3a61375f128e5b0192e4f3f9b32 \
|
--hash=sha256:53fc049ed6390d05423ba33103ded7281fe897cf97878f369a527070bd95795b \
|
||||||
--hash=sha256:53e5179d8abb5710f8e83ba207c41c8d1261fcffd4616500e15ca2b7a33be10a \
|
--hash=sha256:55b3bdd3292283295774ab585160c4004f4f2f203946997f49aac032c84649e9 \
|
||||||
--hash=sha256:53e7b4ce82b54a8bcc71b3b67a5cbd177ca1d7f592cbc92cd38b7349f73482db \
|
--hash=sha256:57653eac22c6a4c13eb22ecf4d673d64a12f266e72785ab1c8b8e5940d0e8090 \
|
||||||
--hash=sha256:543906c127fb1d929b95076db19b83fa2d46751006ff1e23b093aa5ac4d8db42 \
|
--hash=sha256:60869c7ac4aaabe7110f26499f3e6e5696eae98144735b12a9c3d9eae2b51a49 \
|
||||||
--hash=sha256:54cfcdee2770dac994417cbb0ee1f3eb0e7cb6b30c79bf44f2c02ff79ec5124a \
|
--hash=sha256:636bc362f0c5bbc7372bc3ae49737f9e3030dbce469f0f422c8f38079780363d \
|
||||||
--hash=sha256:55bdcc472aafe2de4a253045cc128007a64f1e0264fb675791e132ea5edaa3bd \
|
--hash=sha256:676e5651705ad5d8a70aeb8eb6936c436d8ebbd56e63436cb7dd9bb36d2a9a46 \
|
||||||
--hash=sha256:56f355e79f71aef2a85c80305cc915f894b170dba76de5fe84f6351939b83c06 \
|
--hash=sha256:69f571de7500e0557801c0b51f4780482c0ec5fe2ac851af5a92cfce1af1cb83 \
|
||||||
--hash=sha256:5895ef58c4620afe02fa16044f023dc4dafec08158f9d08874a46a7dbc0341b8 \
|
--hash=sha256:6a7cbeb06d1070f1d14895eeeed4dac5913b22d7b456f2eb969f11f4b3993796 \
|
||||||
--hash=sha256:5bcb6ff3fdab1258a192679ff1a05d44f59626430aa05cd1a9d2447423599228 \
|
--hash=sha256:6cf81fe010b8c17b09495cbd15c1d35afbc8fb405c0c9cf4738e5ae3af1d65be \
|
||||||
--hash=sha256:5f08ec777f35ee70720233b8b9811d3bb5d728137f30ac91b7457709c3261ac0 \
|
--hash=sha256:6e27ea05d184afac78aabbac667450c75e54e35f62238d44463131bd3f96753d \
|
||||||
--hash=sha256:614c61d478b83953e261d02bb2df750f17227cd33ef8002945bf5aebbde21919 \
|
--hash=sha256:6f1cbf0c7926d315c3c26c2da41fd2b5d2fe01ac0e157b78caefc51a782196cf \
|
||||||
--hash=sha256:617105e2c3018ee38d0c8ce5ee3c84f621a6d8b9f723202aacaff28449ca91ee \
|
--hash=sha256:6f497a6876aa4b1a102b04996ce4c1170c7040d83faa9387dd921c16e30d5c83 \
|
||||||
--hash=sha256:6debfa7312ff9d4c124dc71d72e9a0a4b9e0879e48ba6fcb42bef5c3300289e2 \
|
--hash=sha256:756c3c304d394977519824449600adaf2be0ccee76d206ee339c5e76b70ded25 \
|
||||||
--hash=sha256:7041d52c3a7fa20c9e8c182b534704abb19502c8bdcbde7ab23bfda6f642394f \
|
--hash=sha256:77dfa48c9f8013271011e51c00f8ada19851f013cde2c48fca1ba5e0caf5bb06 \
|
||||||
--hash=sha256:70c987b27534f9ae1a723f47ae921571d616da21d3208282bf4c52af5164ac43 \
|
--hash=sha256:7996023b2ed59489ae4762256c8516df9820f751cf2c5da8ed2fb20ee50abab3 \
|
||||||
--hash=sha256:74ab5b6a9fb13e873e5a90946588baecaf488745e1db1a4a5c433f971f035098 \
|
--hash=sha256:7ab7229b6f9b5c1ba4910d6c41a9eb11f543eadb3f384df1b4c293f4e73d44d6 \
|
||||||
--hash=sha256:78253b573e6ffab5028924fc98bc281aae05445969982a10864bc360dea2016c \
|
--hash=sha256:7becdf835feff2f4f335d7477f121af787e3504b48b449ff737afb35869ba7bb \
|
||||||
--hash=sha256:7a75aa63cbf9b21cfaf60dc2657e19df2c2867d91707d653fee171ffeedd1371 \
|
--hash=sha256:7c35b0bf0b48a70b4cb4fc5d7bed9b932532728e124874355de1a0af8ec4bc88 \
|
||||||
--hash=sha256:8800c996b01c2772a783e3e46f3e1abd5823029adca0df54231960de9bfefa5b \
|
--hash=sha256:7c4b6668b2b2b9027f209ddf647f2a4407784b5d88b8be4efcc72036f365baf9 \
|
||||||
--hash=sha256:89176250f686cb9853c0fb7ead90e639e915b84a6f43eedc2a4e7ec21f1037f0 \
|
--hash=sha256:7e5dc4311bd5ac493886c63cbf76ab579dbe4641268e7c74e48e774c74b6f2be \
|
||||||
--hash=sha256:8a5fd34f7f7410d1730d5c2ba873cacb2eed3fede366feb268a70ba22581ed8f \
|
--hash=sha256:888e78eb5ca55a615d285c3c09a7a91b42e9dd6fc699b166ebd5dee87c9ccf14 \
|
||||||
--hash=sha256:8b3b60de05f3dcb6f6a00f818bb2ec781cee4de0645f59ccaf99b1d1823b6100 \
|
--hash=sha256:898703aa2667e3c5ca4c54ca36cd73f58b7a38ef87a5606414799ebce4d3fd3a \
|
||||||
--hash=sha256:8f2f1c4c032c7cedd7d8da6f54c97b70266c6570c3108d3fdffee7188bb70529 \
|
--hash=sha256:8b14eb3262fad0dc2f89c1a43b13727e709504972186ff6a99a3ecaa77102b6c \
|
||||||
--hash=sha256:9491196535a88924a60afd5b5f434b5b203b6cc616250878dbdb223a8f7844bc \
|
--hash=sha256:8bd3ec6376e68a41f9f95f5ed170e2fcf22d4eb27a1f8cb361d0508f6e0557f3 \
|
||||||
--hash=sha256:9aa6e61fdf20105c4144e755bd586008ff450791d67b1c8146fdc15959c4d51c \
|
--hash=sha256:8cf20a8d6868cb15a73cab329ffc07291ba8c22b1b88176026106ae39aa6df0f \
|
||||||
--hash=sha256:9d9edccfe496b476db5f398d97b865e9a6752bcf8aec4eef8390ce20fb64bb41 \
|
--hash=sha256:8f14c50708bb156b3a3ca7230b3d820199d56a48e3af76fa21c2d6087190fe3d \
|
||||||
--hash=sha256:9fc7b5bfec6573f3ae844f457fdde5adeb713f8b8e4a81ad64fc207b49383716 \
|
--hash=sha256:8f546a4dc1e6a5edbb9fd1fd6ad18134550e096a5a43f4ad74acfbd834fc6670 \
|
||||||
--hash=sha256:a0dc483c00da8b673abbb367eb6f8d8f4bcec30eb58529ea13cb42e7fd2dfa33 \
|
--hash=sha256:912d4b6af530ddb1338a66229dac3a25ff11d4448be3ec3d6340583995f56031 \
|
||||||
--hash=sha256:a3a8296e7ab5c295f53f1041487cb088e1480775aafbf7fe545d93b770a0f96f \
|
--hash=sha256:9277145d36a01653863899c665243871434694bcc3431922c3b35c978061bdb8 \
|
||||||
--hash=sha256:a3e22975f905b89a55a488c2a08f2fdb2186175349e917d48985cc468a3d4c6e \
|
--hash=sha256:95d14ca7abefde230f7639ec136ade282655431fd5db03c343b19dda72dd1643 \
|
||||||
--hash=sha256:a4af35c443e0b1a1bd6a8af3f3485d7fda15c142751a00f3ff8090f0b93346fa \
|
--hash=sha256:999802d5fa0389f58decd24b537c54aa63c01c3219ce17d1214cbda3c2b22d2d \
|
||||||
--hash=sha256:a94dbaae5ae27bd849c93570669bff91e0510f33a80805738e3de72a7be0447b \
|
--hash=sha256:9a0f4474b6ea6818b41f82172d799e4b3d29e22c2c520ce4357856fced9af2f8 \
|
||||||
--hash=sha256:ac74facc01463f138b0da5580329cfcc82818dea5656e83ddcd11268fc12ff80 \
|
--hash=sha256:9b16c653d38eb1a611cc898c41e76859ca27f119d25b53c12875fd0474ae31a8 \
|
||||||
--hash=sha256:ad4c8b7488d745d2ca4838ebd8ae5ba9b56341d30b1da43640e4ce87f9f49646 \
|
--hash=sha256:9d98cc980ecc96be6eb4c1994ce35d28d8b1f5e5208a23b421187d1209dbb7d1 \
|
||||||
--hash=sha256:b014a6ed7cf912e787149fdc529166d3ceabac23f26efeea3158c9aba2354e7e \
|
--hash=sha256:9efcc0f11d850cefcafdd9275b9576ad3bfb539bed96807663b32ad99c4d4b88 \
|
||||||
--hash=sha256:b20032766aedf6261c7a566585a40867d092ac03a0d81592d5370ef9b054f99b \
|
--hash=sha256:a2567b72e1ffc3ab25510db43f355b29eeada56c0a622e58dcdb19530eb0a3cb \
|
||||||
--hash=sha256:b2466434105a4e03113c36ec775cc2ebe6676b62eae326fa670bb607ef788c1c \
|
--hash=sha256:a5029cc80718bbd545123cd8fe5d15025eccaaaace5d0eeec6bd556ad6163d61 \
|
||||||
--hash=sha256:b304db572b4368edd8dda8a2274f73156fe15558fca4a917cb8a09fc47af5963 \
|
--hash=sha256:a60eaa2d440cd4707696b52e40ed3e2b0f73f65be07fd0ef23b6b539c9c0b0b4 \
|
||||||
--hash=sha256:ba59d59aba08ac02fc03b0c8983ccd5ee39a199d0552ce9e6d2b4845b34d59ae \
|
--hash=sha256:a79a6d399cef33a11b6f004c67bb07741d91f2be01b8d712d52c75711b1e07c7 \
|
||||||
--hash=sha256:bd52f811e65f6fb634b1047159657c98f52b407f8efec907bcfc09da9a4c0a25 \
|
--hash=sha256:a84792f8631bf5a94e52d9cc881c0b824ab42717165a5579c760b830d9392ac9 \
|
||||||
--hash=sha256:bdd0e2834dce1a26c1bbe26464861e16bbe217042cbff619247c11594472518c \
|
--hash=sha256:a8a4d3427e8de1312ddf309cc482186466c79895b3a139fed3259fc01dfa9a5b \
|
||||||
--hash=sha256:c23ec8ee9d5ab2f5421f9c7fffce208435607af27fd46d4a44e031954352838f \
|
--hash=sha256:a8aca50daa9493e9e13c0f566201a9006f080e7c50e5e90d0b06f53146a54500 \
|
||||||
--hash=sha256:c39846c3aad97a8530c89d7a3869a8f8e9e3762c6ac0504481e5c80948f7e807 \
|
--hash=sha256:aa6d0d932e0f39c02b80744273cd5c388a2d9bc07760a03164f229c8e02662f6 \
|
||||||
--hash=sha256:c3c200cf9757edd785051dc699c7ecbec22110dbfcb3fefc7a9f9695eda8ea7a \
|
--hash=sha256:ab2899f9fa2f9f741896ebb6fa07c4c883bfa5c7f2ddd8cf2aafa86fa981b2d2 \
|
||||||
--hash=sha256:c7d3a97c678d34fc5b59da671ee9cd630096ddc643e7b5a30d54a2a6f3574d3f \
|
--hash=sha256:af545c2cffdb0967a96b6249e6f5f7b0d92cdfd267f9d5238d5b9ca63e8edb10 \
|
||||||
--hash=sha256:c8653fd547c93a61aadc612007790f5555cdd18946fa48cf45e26d8ea4ea473d \
|
--hash=sha256:b18f31b80d5a33661e08c89e202edabf1986e9b49c42b4504371daeaa11b47c1 \
|
||||||
--hash=sha256:cc7cb243a68167172f48c1fd43cee91ec4b1d40cefd190edd43369d1a6bc9c82 \
|
--hash=sha256:b20df693de16f42b2472a9c485e1c948ee55524786a0a34345511afdd22246f3 \
|
||||||
--hash=sha256:ccd4893707b3e2a13e39c90d43cf80edf2e4d0457935bcc103bf2346214c3f15 \
|
--hash=sha256:b38765950832f7d728297689ad78f5f2cf79ff82487131c4d26fe6ceecdc5f8e \
|
||||||
--hash=sha256:cd817772b2fcf2b8c0905795318485f9ec16eae60b29feb7f4c77085311637f0 \
|
--hash=sha256:b6f6cd1560c5fa427e3b6074bb24d2c64e225afbb7165008903bd42e4e33e28a \
|
||||||
--hash=sha256:cda5fd5c95ad7a125a2e8464acc78b98b94c475a3780d6aa0aa157c93f470f4d \
|
--hash=sha256:bace460460ed20614fa6bc8cb09966c0b8517b8c58ad8046828c6078d25333b5 \
|
||||||
--hash=sha256:cef89a58e628c4efcac3275c2d68083f82426dcdc89c1492a6f654f9f7ea6ab9 \
|
--hash=sha256:bca9ef7517fd7874a1a08970ae88f497bf5c984610caa0bf40bd7e8450852b95 \
|
||||||
--hash=sha256:d1558173930a5a8d3069cee5c92fc91c87c4dbcb099debbb3622053717145a19 \
|
--hash=sha256:c180f480207a9b2475f2b8d8bd7204e47aec952d084b2a2be58a782ffcf96074 \
|
||||||
--hash=sha256:d6088ec9894113802bddb3c09e974929aed2c7b3a8c456219b8aab4481f1a239 \
|
--hash=sha256:c2b2355dc094e5f7d45a7bb262fe7207aa0460b37a0d87027dcf21b5d890e7d5 \
|
||||||
--hash=sha256:d6218d92e450824e9b4881f44e8c09f1853b490f9a64130801024a4793b1b3b0 \
|
--hash=sha256:c564dd5f09ddc9d8f2c2d0a301cd30a79a2cc1b46dd1a73bef8f0038863d016b \
|
||||||
--hash=sha256:d77640cc618c1d99fc4f8589c0f24a730adfa54eb1e57ef7bf0c8dfb78da898c \
|
--hash=sha256:c632ce9c0b534fbe25b52c974515ed674937c5b99f549a92127c85f771a78772 \
|
||||||
--hash=sha256:d7d2deec16eeedf55f2c7cf75b521ea3856a5177e123844f8fd0f114ce252cb5 \
|
--hash=sha256:c719f65bebcdf6716f10e9eff80d27567f7892d8988c06de12bbbd39307c6e3a \
|
||||||
--hash=sha256:db332af25642007330fca8be5c4d194caf2bea7a7fc84415aff3497af5dfee6b \
|
--hash=sha256:c86969d012e51b8e415a8c6ce96f7857d6a87d6207303ab02d5d11ef0cad2274 \
|
||||||
--hash=sha256:dd54d0e8717de95939766febac482ac0474d8ac3b048115f9f2b1d23a16e7db4 \
|
--hash=sha256:c974fb66180e58709b6fc402846f13791240d180b74de81d23913abe48e96d94 \
|
||||||
--hash=sha256:ddcac3c6b382e81f1dd0499199d4136b877beb4cb5ef770bbbfba56c4b8f55d2 \
|
--hash=sha256:c9883051c6972f58bfc4ebb2116345ee2aa151178e99c3f2b2bbe2af712abd13 \
|
||||||
--hash=sha256:df82f3787c940c94986b34222d59c9e38843fba85139f36e85255a82ad5355a9 \
|
--hash=sha256:ca9ac61ac6db4eb6c2a0cd1d0f7e1357647b638ccc92f7e9d8d133e71ed3c6ac \
|
||||||
--hash=sha256:dfa68deb2a443bdaa3ea5297b0699c1464f08aef3812b486d1348eee61b07dc0 \
|
--hash=sha256:cb979826071c0986a5f08333a36104153478ce6018c58cba7f9caddaf63d5d67 \
|
||||||
--hash=sha256:dff9461ec275f22135650d5ba4b4931a11f3958df7dfbb8db630000d4dee0883 \
|
--hash=sha256:cd3db5927bf9167d5a6157ddb2f036f6b6b0ad001ac82355d43e97a4bde76d76 \
|
||||||
--hash=sha256:e1e74298bab6ee0d6e749ed4fd1901c7e604bdda32c03d787a2cc71c46d0433d \
|
--hash=sha256:d147004fede1b12f6013a6dbb2a26a986a671a03c6ea740ddc76500e5f1c399f \
|
||||||
--hash=sha256:e2667f0bbe7eb6c74eae5e9691441ad186e5845ca3cff63230fc09c4e7514f5d \
|
--hash=sha256:d3a4834f221061624b8887090637db9ad4f61752001eae37d56c52fddade2dc8 \
|
||||||
--hash=sha256:e3be98a7c30b8c25d573dafba7171d66dfb05ee6a9070fc46535464ff97700a6 \
|
--hash=sha256:d9010032a0b9710f58012a1e9c222528763d860ba2ee1422c03473eab47703e7 \
|
||||||
--hash=sha256:e568e14940c09955aa51f4e645b6daa18a581c5dcfcd73744dcc86a856e3ced3 \
|
--hash=sha256:d97f93fdae594d886c5a866636397e2bcab146fd7a132fd6bb9ce182224452f8 \
|
||||||
--hash=sha256:e72ee89e28d907a18f46959b4eb0bb06701cc7f8cf4366e00029e2ccfaaf5924 \
|
--hash=sha256:df23d57718f24badef8656c49743e11a89fd6f5358fa8a7b96e728fda2abf7d3 \
|
||||||
--hash=sha256:e92eb8acc45eb6a9f4935071a77edf5b85cc6f8dfad5cd99e97653c26593cdde \
|
--hash=sha256:df6104c009713d3a89621096f3e3e88cc323fd269dbd7c20afe18535094320be \
|
||||||
--hash=sha256:ea05e1f97ceea523942d9b2a7d7c0359d781d683d6b043f5943a602b14da4787 \
|
--hash=sha256:e5e5f7debc7a57af53fdf5c5009f9391d9f4c12867049d509bf7bb164a6e295b \
|
||||||
--hash=sha256:eac645b09bcfdf73df7536331f0678c1086ea250981118ddb5199e17ccef72bb \
|
--hash=sha256:e7d2f8616f0ff60bd332022279011776c3ac0faa0f1b463f7bb12326fbc97a1c \
|
||||||
--hash=sha256:eb0495d778817619273c108784292be161a924b9f5ae5cbbc70a2caa6838250b \
|
--hash=sha256:e999f0c88a458c836d5fb521814e92ed2172c649200336a6df514987c1488258 \
|
||||||
--hash=sha256:ebe8e504f058fe91223351cecd2d9d6946c9d241bb0250d898ffbdf584cc72b0 \
|
--hash=sha256:eb4639f32fd4a9904ab8fb45bf3383ba71137f3d9d4ba25b3b3f3109977c5b8c \
|
||||||
--hash=sha256:ed099d105449c4f9e84f24af203cd131349d4761d8813fa7e02c32e7128cd910 \
|
--hash=sha256:ec707059ee75732b1ba130ed5f9580fe10ff75180c812bc267ded039db5128c6 \
|
||||||
--hash=sha256:f0f177d1b195b9e06376cfd7d308d8a1b920909a609d03ac82a8c73bbb16d3b9 \
|
--hash=sha256:ecc26751323224cf8186efcf7fbcbc30f4e1d8c7970659daf25ad995e4032a56 \
|
||||||
--hash=sha256:f3d2669fe7dec7fc359ecdb5984b29b50d85d5d00f8c1cb61de4f4a24ee42627 \
|
--hash=sha256:ee5e86776273de1795947d17bddd6bb19e0365fd2af4289c0d2c5454b6b1d36b \
|
||||||
--hash=sha256:f4e05329faa0ea1a404b37de4f034fd2c2defcca06a68dc6745e4e56c88e8a48 \
|
--hash=sha256:f1162a1492032c82f14271e831c8f4b49f2b6078f4f5fc74de2c912fa225d51d \
|
||||||
--hash=sha256:f53bcd52f585e1ac3e590d61434eb61f9a88c38df041b4ea126d97144344a77b \
|
--hash=sha256:f34ecee82858e41dd217734f0c41a532bd066bcaab636ad830f03a30b2a96f2a \
|
||||||
--hash=sha256:f55119f7bf25f49ed210f6096090715da24f2943c62102448915fde3c62877ce \
|
--hash=sha256:f85c6f327bf0b8c29da7d93b1cabb6363fb5e4e160a32fa241ed2dce21b73162 \
|
||||||
--hash=sha256:f631fe87a6f30df5fbe6d79640b25e4cffb38c31c7fb6f10871517b84b0f8c1a \
|
--hash=sha256:f92995dfec9420bb69ae629abf422e516923ba79ba4403bc750d94fb4a6c68c1 \
|
||||||
--hash=sha256:f8fb78a83c9e5f741ca3a68cfb455c1f5bb83b4e7249a3848b3cd78d0a8563b0 \
|
--hash=sha256:fb0540c854ac9c0c5ad495908fdfd3e332d553ec731698c0e29b1877ba0d2ec6 \
|
||||||
--hash=sha256:fa9467a8113aa69d3d7c55a70ef0b7c636010a40993f3df9d9d0d73b3eb7ef24 \
|
--hash=sha256:fceedde51fbd67ee2bcc8c0b33d0126cc8b51ef3bbde2f86662bd6d5a6f10ec5 \
|
||||||
--hash=sha256:fd51ebf9d3a00c074df4ede271023f4d2dba289bcc740b88191872716014e3c5
|
--hash=sha256:fe6970addfea9e5e081401bcbadf865d2b6da045472f58af08427e108d618540 \
|
||||||
|
--hash=sha256:fee86b7c4bd29bdaf0d53d14739b08a106fdda809ca5fe032a15f52fae5fe254
|
||||||
# via checkov
|
# via checkov
|
||||||
aiomultiprocess==0.9.1 \
|
aiomultiprocess==0.9.1 \
|
||||||
--hash=sha256:3a7b3bb3c38dbfb4d9d1194ece5934b6d32cf0280e8edbe64a7d215bba1322c6 \
|
--hash=sha256:3a7b3bb3c38dbfb4d9d1194ece5934b6d32cf0280e8edbe64a7d215bba1322c6 \
|
||||||
@@ -156,9 +157,9 @@ attrs==26.1.0 \
|
|||||||
# aiohttp
|
# aiohttp
|
||||||
# jsonschema
|
# jsonschema
|
||||||
# referencing
|
# referencing
|
||||||
bc-detect-secrets==1.5.50 \
|
bc-detect-secrets==1.5.47 \
|
||||||
--hash=sha256:016ce9e79f692adbabcbef4a7293db427352911c3f81404a136e6f5c3e54a7f2 \
|
--hash=sha256:46f88c710b0fd8c5f2e54b361d793b5e1469197884da73cfc6f488b614366fc3 \
|
||||||
--hash=sha256:99037375d9cb49ed07e5bb12722f4bbb76fb8acaff6f367eef9df7559e7642b3
|
--hash=sha256:a9be28a2e564f2b19731991df39e63ae6372cc84d828ee24e50c094cbb4c154c
|
||||||
# via checkov
|
# via checkov
|
||||||
bc-jsonpath-ng==1.6.1 \
|
bc-jsonpath-ng==1.6.1 \
|
||||||
--hash=sha256:2c85bb1d194376808fe1fc49558dd484e39024b15c719995e22de811e6ba4dc8 \
|
--hash=sha256:2c85bb1d194376808fe1fc49558dd484e39024b15c719995e22de811e6ba4dc8 \
|
||||||
@@ -483,9 +484,9 @@ charset-normalizer==3.5.1 \
|
|||||||
# via
|
# via
|
||||||
# checkov
|
# checkov
|
||||||
# requests
|
# requests
|
||||||
checkov==3.3.16 \
|
checkov==3.3.1 \
|
||||||
--hash=sha256:43e5383418a8b52d39747e2daaec4da4a6b7db3e2f64ab6c93f8b89c044c7665 \
|
--hash=sha256:1e781a58de8310ec99756205a7991adcfe66524a52642c6474e9d86a7cc9c635 \
|
||||||
--hash=sha256:6f7f611f45c765af9b6acd43e603a438153d86007d02dffa7903a1125f9b5089
|
--hash=sha256:aafc571cc937ddaa0714df30f2b9d79302a07cb9a41b3e0168c7eecb8172db14
|
||||||
# via -r .github/requirements/checkov.in
|
# via -r .github/requirements/checkov.in
|
||||||
click==8.5.0 \
|
click==8.5.0 \
|
||||||
--hash=sha256:255bc9599cf7748b4b1a446ccc735421bd08a2ae529a8b88597d3de5664ee360 \
|
--hash=sha256:255bc9599cf7748b4b1a446ccc735421bd08a2ae529a8b88597d3de5664ee360 \
|
||||||
@@ -983,73 +984,79 @@ networkx==2.6.3 \
|
|||||||
--hash=sha256:80b6b89c77d1dfb64a4c7854981b60aeea6360ac02c6d4e4913319e0a313abef \
|
--hash=sha256:80b6b89c77d1dfb64a4c7854981b60aeea6360ac02c6d4e4913319e0a313abef \
|
||||||
--hash=sha256:c0946ed31d71f1b732b5aaa6da5a0388a345019af232ce2f49c766e2d6795c51
|
--hash=sha256:c0946ed31d71f1b732b5aaa6da5a0388a345019af232ce2f49c766e2d6795c51
|
||||||
# via checkov
|
# via checkov
|
||||||
numpy==2.5.2 \
|
numpy==2.4.6 \
|
||||||
--hash=sha256:0090ccdd57ec2703e9b49d0bf554767370581c1dd0a6b2bb2b2d9def317d042a \
|
--hash=sha256:001fbb8e08d942dd57599e781f2472269ee7f2755fae407b4f67b2f0b17da3f1 \
|
||||||
--hash=sha256:078f9b027b478c9379b9677babbf0f8b8f1ecfada27636d7b9a93990c638739f \
|
--hash=sha256:0280e0356c0829a18d9de1cb7eee50ec22ca639878d7240307ca0943d73cd2c4 \
|
||||||
--hash=sha256:07d4e89f3a9ab0a9ba24264ccdb642b3dd951b2281e8883a5481a4aa79cc31a7 \
|
--hash=sha256:043191bfa8eab18c776647b62723ac9dddece59743b13f49b2016094129c2b3f \
|
||||||
--hash=sha256:0a4035ae1129ff8777f08bfbd44f1e5d8e9c049ce0c2dd78fc0d92c13e7251c0 \
|
--hash=sha256:06ca2f61ec4385a07a6977c55ba998a4466c123642b4a32694d3128fce18c079 \
|
||||||
--hash=sha256:0aadf13b60048d501e05fa699efaf7734e2494f3498a4c2a5521d822640324f3 \
|
--hash=sha256:0a041d3d761dc3c35cc56ce0351506a02bcbc25f7b169f652435141a17db9096 \
|
||||||
--hash=sha256:14e373cfc6387177e8409dac3c7159be8eb05cd77096cd7c950268b86f62831c \
|
--hash=sha256:0ab0a9c4ffb1a6d95ef519fe4247dba8eb6b18ad93999f76b7f657039acabd47 \
|
||||||
--hash=sha256:1ab3d4a901f844ea836c3e80bf463c6a27d7f3c14e8e292fcf28d348b25b9bce \
|
--hash=sha256:0c9136e14ed34a9e343a31c533d78a9813a69a3148332bce5e9821cb2f996e66 \
|
||||||
--hash=sha256:24b9dc2e3d84aa58523798805194e23e736f3f6ce2d1a5b92583ae734e6dbda8 \
|
--hash=sha256:110f8b71aacb688ec69062bb7f6938a0f8acb01b7c1c4beb453c65b6d234584d \
|
||||||
--hash=sha256:27650bb0e7140fa3d37b9923b4803645e0b125d190f326eecfd3f4dad8e8ade1 \
|
--hash=sha256:112b06a867b235ef466ed3508ddf0238050df9c727cafb5301ac385b899189a1 \
|
||||||
--hash=sha256:28ac63476ec7651484215ee7fa15a1f78b57c14621f01e392afe17b9a1390ce4 \
|
--hash=sha256:17f9ade344e7d9b464a084d69bcf18fc691cb1db67c62ed80820bf4926d78f0e \
|
||||||
--hash=sha256:29b86ff8a6cc556b47ec6b64b194815cc80e6bf5eedcc6cddfd65318cb0b4eee \
|
--hash=sha256:1e254a00cdf42b1e4d5b3d68d33af63268d41340d8885df2ab6470f2e1500147 \
|
||||||
--hash=sha256:29d81e97f668489cba8ebfd796b9bdd453525d35dd9e162e2daec94bf3fc7740 \
|
--hash=sha256:1e978ec1e8bd0e0e4de6bb75de9d30cbb74db6b6a2bb727618613703ca0167dd \
|
||||||
--hash=sha256:2cc779226e476d1e1f08c74068c419e60f41a9e0e069c92f6671d31d5c985e98 \
|
--hash=sha256:25c692919ac5a01f170a3bfcd62d745b24fd095c353d50812637d6fcab442e75 \
|
||||||
--hash=sha256:2ffa7bacab3e2ee1b19ed31766bb60bb380b68c23f051e199c5cc598afd68710 \
|
--hash=sha256:260a5d70215b61ab4fadf5c7baacd64821842975eea312125ed3c39a6391b063 \
|
||||||
--hash=sha256:318b9a4c845dbea06708a29c84ee429cc3065048db34cdb799047643492050ee \
|
--hash=sha256:2803abfebfc990042cd494d8ce2d5f82e9d847af6d35ec486923aa19dbad5e73 \
|
||||||
--hash=sha256:34c319e2963be042673fb46570501b2f06c41924e17e3563d58646b4380dfb68 \
|
--hash=sha256:29a287e0cf63ff528da061de6b9f64a4618da591ca1046aafc54062e40ca7eab \
|
||||||
--hash=sha256:3a2f061cebd9e3d23bdcfaaded5e2293a4c6a5b60fa42df85d410a725ce621bf \
|
--hash=sha256:29cb7f67d10b479ff07c17d33e39f78c07f71c40ef30d63c153d340e96cd3fb4 \
|
||||||
--hash=sha256:3cdec01fa790a186d430433fdd4d4ffb70eed6f0eeb4bf05c8dbe2dce0a9bcb8 \
|
--hash=sha256:3213d622a0283a39a93d188f3cf72b26862df52fbb4ca3697f51705016523d41 \
|
||||||
--hash=sha256:3e4c367352d3747784248a227fbec218e193b56f7e6692e3b64fc805478ecfdf \
|
--hash=sha256:33111801a01c12a8a1e3721f0a9232f8cfc8ae2c6b7098167e6f623c6073f402 \
|
||||||
--hash=sha256:40f4d451aed46a8046a1aae41c4e55fb3612273df9c502480135e1501576a34b \
|
--hash=sha256:357cc07a6d7b0b182ff02249616a03742827ebb1277546b5c7cd7f7620a45698 \
|
||||||
--hash=sha256:44ef9675d908e65f9953063837c3277730f3f4437615a4cdab67b366cabaf884 \
|
--hash=sha256:38efbc8de75c7a0fc1ac190162d892787f3f47b57cc291231aafee36b80982b7 \
|
||||||
--hash=sha256:4bbd96c833ecc8cc069ce518078fc8c60cb9cbfb0fea5b7a803ad65035596d03 \
|
--hash=sha256:4081eb135ac24158bd51cdfbef16f1c64df7063b1143f24731387137c092bec8 \
|
||||||
--hash=sha256:4ec954036759bcee3aa484f8603bd9c14f3e776293b85578b8734c2d72777c69 \
|
--hash=sha256:40fdc1ae7125e518ea98e53e69a4ebc27e1fd50510c47b7ea130cf21e5e1d42b \
|
||||||
--hash=sha256:4f9744f9fbdcea0bc552e8f19e1f141f811a3f9bc2be2cc6e86d982cab23e3f4 \
|
--hash=sha256:4cfe66903cc32a9921a6733d96b19bb6abf310397581bbad89c228f5abaf0ee8 \
|
||||||
--hash=sha256:50a68f4bacd8a2b33d8da3d2269d0d78500f86ea582e4786dc10f5ef2c2c6842 \
|
--hash=sha256:511dbaf848decaaaf4b4ca48032619fb3138710c4bf7da7617765edad1ef96b0 \
|
||||||
--hash=sha256:50e500dc868e9313530ce12ba470fe50ff3afe3d62993ed6eff652dacd555b65 \
|
--hash=sha256:55cced7c52e981362f708ad635198e97a752dfba412cc03c23bbf3bd8d5cd662 \
|
||||||
--hash=sha256:52c808f96484f5571a5cc863775ce50247c17dfb3b0361f8ed6b4b0456f80080 \
|
--hash=sha256:56b39e5e0622a09a25bf5baf62f4bcf0cb8a41ae6e2819cf49bbc5a74c083f91 \
|
||||||
--hash=sha256:5f8e00be2ec6f45f4e8a41a527f68d44a7d96fee92a650e4d8b1326f77f61e6e \
|
--hash=sha256:5dbbdb29840ca3d91ee0fece42fc29278886d908280bfec0a5846c6f901a3eb0 \
|
||||||
--hash=sha256:60e902ac295855348a5ca2ea4c89108989a9f5fddfad3dfc0a8f36b10358567e \
|
--hash=sha256:5f9fb9157b4ce2971008323afe46053787b526ef624fea915b261468a8421a0f \
|
||||||
--hash=sha256:65f188481f1669e26f62b701e8205d19e460fa4a9b52a1414ba382330e4a3414 \
|
--hash=sha256:6180d8b35af935aed8ece3a85e0a43f87393ae0ac87c8d2c8bd2c993f7270ef3 \
|
||||||
--hash=sha256:6950c4b7dd562453090548ba7f5da7e59f57f85663f15d5dcc60e249192f7e59 \
|
--hash=sha256:68a5124b13fa6cc2086764a20005d30bc0548146f7f5322f02fce212ca14317f \
|
||||||
--hash=sha256:6a9bb119fb8dd21ba30b3f0e555b7e2b081bd9883af21ec9c1c633d161cda3a8 \
|
--hash=sha256:68bb27509ac1b9a3443094260f6326150663b06abe40b73a2f81160623da5b67 \
|
||||||
--hash=sha256:6b588cc8f902d6bff201c19fd00c43ab8545671e3554d014e12e14139e5e8617 \
|
--hash=sha256:6f41ae150c4e32db4f3310cdaf64b1593a03dbabe29eec77fc9b50fe64061df6 \
|
||||||
--hash=sha256:6df895598c0edcb41030126c89e0f353b07d93238116143b7405e937359736c4 \
|
--hash=sha256:7265a2f3d436e54ef9f2b52b5c937e6be778781bd97a590319d7348f1c1ca997 \
|
||||||
--hash=sha256:6e8172ddfcf5cf74b811d372b570b83c60bd2de87a6fbfbebdadb4a9bd9c6cbb \
|
--hash=sha256:72fbe16c6fac95aedf5937fa873445cec2110be35d8a4e9433d7501fd98dae6b \
|
||||||
--hash=sha256:7354826bc6f8f69402e9b7fe28d15fcd34feebd74f856f111585c5b0c9fb0251 \
|
--hash=sha256:7d92c3819208a60205a12a245c91ad70cb0a85336659b19b834205573ac8456e \
|
||||||
--hash=sha256:7587f53dfbd5edc0f7b87c6217b4c6d2d1f2ef9c3da70bc1315e7db5f8d7ec9d \
|
--hash=sha256:8155154c7c691289fe18f510b5d4657c68c67989f293f0535a91360392ff6538 \
|
||||||
--hash=sha256:77843ca236b777e67f8d6b3660ea116e499612703a0ecd7093f316201eb9d8e2 \
|
--hash=sha256:81a1cca95ed5bb92aa8b10dd2cdc9a0d3853a50fad926c28b5d7e8ea54389627 \
|
||||||
--hash=sha256:7999d4ddb0c4025018373fd787510d46e04c769467af22869707b3c1cfd459ab \
|
--hash=sha256:89cd468399cfd2504718f0ba50e410dca55a170b61a02ad92bb18c8a65186e93 \
|
||||||
--hash=sha256:85aaccb24182c25df891ad0ec333585967e115269d5f1b17f2c9ae005bc96657 \
|
--hash=sha256:8ad03c0965fb3c692200e74d458ca28c1dbb4ce96f9a479a8aa041ad5fabca02 \
|
||||||
--hash=sha256:8e4cb9a754c8a0c62eaa88273a5fba3391f4a610d1dee893c0755da31c083f15 \
|
--hash=sha256:90f9849678c75fe7afa2d348ac842c168b0a4d3d61919687216dfc547976d853 \
|
||||||
--hash=sha256:8ee9c4eeb8454b3660a8b53493563c3e121c2fc94fbd72b848ef814ed7b676a9 \
|
--hash=sha256:948424b06129ce883307e8cff868c31396d8dc7630a59c61d70d98dbe70f222c \
|
||||||
--hash=sha256:9a0731745a72a184490a582fb4af2533512bd071ace67785b5fdffc0ae58dce8 \
|
--hash=sha256:9cd5ffd25db4e7ba6a375693b3fc0fc1791ec636c17db3720da19bde7180ec43 \
|
||||||
--hash=sha256:9e9413326d726c2545bfa65d2c0876871e8d8386e77f992c1d426e180bbd4323 \
|
--hash=sha256:a0df0043bdb289bde1f62da130d20df23d58b45429f752bc7a8fc5325a225ecd \
|
||||||
--hash=sha256:a610dc7e3c52edd39c2bc2375ff9c3fd59cb3ad00e4472d36f83bc1457145788 \
|
--hash=sha256:a2c306dea656c12c68f51f4cea133cbe78ca7435eb28c735eac1d3ebe73be6e8 \
|
||||||
--hash=sha256:a839318485284a6fb31be4f8f2c91c8f2cb22f4543c4a8903f12b0671ffe07cc \
|
--hash=sha256:a7830bab239b79cda9c08c2da014761cafb48da6150e1da17ac06283f43b6089 \
|
||||||
--hash=sha256:afb3f0632d6b2e3ba04dbce8d1e48d321b369138b73830b5ca371a0e8d479d56 \
|
--hash=sha256:a7c711e21628b52034bb5ab8d1bce291f752fcc5e92accc615778acee1ff4778 \
|
||||||
--hash=sha256:b879fb674276e331513fb136b78dbc6bd3c848309e0d841cfd63be3896c4cfc1 \
|
--hash=sha256:aaf159caa35993cb1f56fb9b8e4610d35758e7ca005412eb1daa856a78c9c4b1 \
|
||||||
--hash=sha256:b9727f472d2f3888053b8a75ab0cb94745a9de224bb5846dbadc0092101bc71d \
|
--hash=sha256:ae506e6902902557576a26ff33eda8695e7ecb3cb36c3b573a0765dee114ebdb \
|
||||||
--hash=sha256:ba0a474801b8dc67b66bf465548abc90e82b44d2611b5770f33008dcabffe8ec \
|
--hash=sha256:b507f5c4c1d508876d1819b6bf9a49d365b96320b5d4993426b33a23ca4b8261 \
|
||||||
--hash=sha256:bd68ece1553d2023c09a4226d9e41c586ad2d20594d1a456186c33513d2cb3f2 \
|
--hash=sha256:bf162abab1c1a736333192707cef898e735a5ca00f38f27eeedf44b39d9e85eb \
|
||||||
--hash=sha256:c081cbe16ba1ab53078e5ff29013621e33c509eedab055775d956427712c236e \
|
--hash=sha256:c1a2af6c6ef86344a6b0db6b97834208bf598db514f2b155042439b62605601a \
|
||||||
--hash=sha256:c1f017dc0875c9209d219f97feceb7d54c2661bb243deb4114478e1295808af7 \
|
--hash=sha256:c2d37ab77531417474168eb79d6d80b14f821a966818505d03013d0833edb7a8 \
|
||||||
--hash=sha256:cebc2d6dbb605a7703d59751dea4bd6b0ab127a5a4338a6f432df1936fef8b26 \
|
--hash=sha256:c4fc99836233ea196540b17ab0983aff60ed07941751930f5f4d05bc3b3b7359 \
|
||||||
--hash=sha256:cf7de32f486e4ac9e2d93b810f9e9ac72a728dd46a32a0bb403222f27f653514 \
|
--hash=sha256:d581b735e177fdcdce6fed8e7e8880a3fb6ee4e3653a3ac6af01c6f4c03effc5 \
|
||||||
--hash=sha256:d482d171c406ae88c5b19cad3b6a1c4c5209f886ab74bc44c2c865c23f52d860 \
|
--hash=sha256:d6da64deb6b8ed903e7560180a92f2d804ee1ba5eeb849ac2748b8c1aba1f6d7 \
|
||||||
--hash=sha256:d6a48072864e3324e194a8fbb3c657bcc5b5c869dbc64c9537b1d5c862572c0a \
|
--hash=sha256:d8e8286dd7cea7895157318d1b91cdacac64c479f3cbc8dce548331728484751 \
|
||||||
--hash=sha256:d787cf769c3baeb5f6235e778edb52c08dfa923789b5958f28e6450f96107cb1 \
|
--hash=sha256:ddea102b48f9e339f3948bf22040944184627a30fdf7f858667673b9c5f033c8 \
|
||||||
--hash=sha256:dc649493697006bc90614a5f0bbc8cb3cb1866715c474e473694968d7e6b99ab \
|
--hash=sha256:dfa20cc6ca228e6b155b11da03825975ce66aea520985dbbddf0f2a5a495c605 \
|
||||||
--hash=sha256:ddf47472af2e4280d79bac82304f5e80150211f1b9e614b760061d5fdfbb6eba \
|
--hash=sha256:e3e5193ef5a3dc73bceee50f7fdc2c90dbb76c42df8d8fae3d1067a583df579e \
|
||||||
--hash=sha256:e5651f3f87add730ee6608d915009e19c911fba0cb000c7e3ea994b7d768eb12 \
|
--hash=sha256:e3eeb0aabd6bd5ce64faae67e9935203a6991b4bc2a485a767fbafb2c5125f45 \
|
||||||
--hash=sha256:e79aba74ffaf5f78a050d777c184cddf8fdffabab38acf5f3ef1fecbc17895d6 \
|
--hash=sha256:e5805d5a22fd19c8ccff10a9561f9df94436b0545619ea579db2d3c35294bce2 \
|
||||||
--hash=sha256:eaa088384c46f519dacb93b7ec483a6d6b19a4a2085ae4f25ab9b1c43d387d1e \
|
--hash=sha256:e85b752a1e912b70eaad4fafbd4d1238007ab221de2009b9a2f5ae7461239895 \
|
||||||
--hash=sha256:eaca7ff36f0f52e2111ec71f169d8fd3e889e7ddc0d2592e0d703fd8d3ce8fac \
|
--hash=sha256:eaf7fa2de5c0be8ae6ff8e9bea2ccd725e980541244521d8d4b5f3354a27babe \
|
||||||
--hash=sha256:f06571a052127dc1b4e8b83029b4d1b20daa2b64a31cdd181fc6bc774e9000eb \
|
--hash=sha256:ebfb099f8dcf083deef3ac1ca4c1503f387cf76296fcb3816b66f5ecb5f54fdb \
|
||||||
--hash=sha256:fd0d703772bba096843785bd38371e31bb4a0c1151497ad5739d182114a73f7f
|
--hash=sha256:ece3d2cfe132e7d51f44a832b303895e6f2d499c5e74dfbdb06ee246147a304a \
|
||||||
|
--hash=sha256:ed9749eef4cbd126da3dc1d6bcb3a57f5eb7ac6a6484146bdbf743f552dfc577 \
|
||||||
|
--hash=sha256:ede83e07a75dd06bc501566c1eca2afc0d61677c1472ac9ad93fdee6e638a48d \
|
||||||
|
--hash=sha256:ef4aea96ce4d3b074422cb4f2f64e216bf9e213004bb58ecfdf50ea02ea8eb9a \
|
||||||
|
--hash=sha256:f3a3570c4a2a16746ac2c31a7c7c7b0c186b95ce902e33db6f28094ed7387dda \
|
||||||
|
--hash=sha256:f407cb6b8e9d6d8c626bc73c945db1706035af8fd632295547bf1c9e46d092d6 \
|
||||||
|
--hash=sha256:f74a575920ab21fe304421a3fc28793d82e299cae9eccb37084e9fc7f3617c20
|
||||||
# via rustworkx
|
# via rustworkx
|
||||||
orjson==3.12.0 \
|
orjson==3.12.0 \
|
||||||
--hash=sha256:010811c1b69773450a01cef97727a67b223242f350b77d4ca000e59a9ef2155a \
|
--hash=sha256:010811c1b69773450a01cef97727a67b223242f350b77d4ca000e59a9ef2155a \
|
||||||
@@ -1943,7 +1950,6 @@ typing-extensions==4.16.0 \
|
|||||||
--hash=sha256:481caa481374e813c1b176ada14e97f1f67a4539ce9cfeb3f350d78d6370c2e8 \
|
--hash=sha256:481caa481374e813c1b176ada14e97f1f67a4539ce9cfeb3f350d78d6370c2e8 \
|
||||||
--hash=sha256:dc983d19a509c94dba722ee6abd33940f7c05a89e243c47e907eb4db6f1a43e5
|
--hash=sha256:dc983d19a509c94dba722ee6abd33940f7c05a89e243c47e907eb4db6f1a43e5
|
||||||
# via
|
# via
|
||||||
# aiohttp
|
|
||||||
# aiosignal
|
# aiosignal
|
||||||
# beautifulsoup4
|
# beautifulsoup4
|
||||||
# checkov
|
# checkov
|
||||||
|
|||||||
@@ -1,3 +1,4 @@
|
|||||||
|
safety==3.8.1
|
||||||
bandit==1.9.4
|
bandit==1.9.4
|
||||||
semgrep==1.175.0
|
semgrep==1.175.0
|
||||||
jq==1.12.0
|
jq==1.12.0
|
||||||
|
|||||||
@@ -1,5 +1,9 @@
|
|||||||
# This file was autogenerated by uv via the following command:
|
# This file was autogenerated by uv via the following command:
|
||||||
# uv pip compile .github/requirements/security-scan-tools.in --generate-hashes --python-version 3.11 --python-platform linux -o .github/requirements/security-scan-tools.txt
|
# uv pip compile .github/requirements/security-scan-tools.in --generate-hashes --python-version 3.11 --python-platform linux -o .github/requirements/security-scan-tools.txt
|
||||||
|
annotated-doc==0.0.5 \
|
||||||
|
--hash=sha256:117bac03a25ede5df5440e855b32d556049ca169ead221505badf432fed4b101 \
|
||||||
|
--hash=sha256:c7e58ce09192557605d8bbd92836d7e1d520ac9580096042c0bfd197efacf1bb
|
||||||
|
# via typer
|
||||||
annotated-types==0.8.0 \
|
annotated-types==0.8.0 \
|
||||||
--hash=sha256:13b2beaad985e05e2d6407ee4c4f35590b11f8d693a258a561055cac8f64cab7 \
|
--hash=sha256:13b2beaad985e05e2d6407ee4c4f35590b11f8d693a258a561055cac8f64cab7 \
|
||||||
--hash=sha256:f072f4d804ea359e4eaf198b1af7a8b0943881a87f31bb764f8bf219bb9419e0
|
--hash=sha256:f072f4d804ea359e4eaf198b1af7a8b0943881a87f31bb764f8bf219bb9419e0
|
||||||
@@ -20,6 +24,10 @@ attrs==26.1.0 \
|
|||||||
# jsonschema
|
# jsonschema
|
||||||
# referencing
|
# referencing
|
||||||
# semgrep
|
# semgrep
|
||||||
|
authlib==1.8.0 \
|
||||||
|
--hash=sha256:88aebbd9af6757e14e912d5dc007ae1dc1f3e27e3b2152ce7c552ee2c3b3c121 \
|
||||||
|
--hash=sha256:f3ecd5f1da737262fb53bf1a4d95c4ea1ad9dd509316587a255c99ab1838a4f0
|
||||||
|
# via safety
|
||||||
bandit==1.9.4 \
|
bandit==1.9.4 \
|
||||||
--hash=sha256:b589e5de2afe70bd4d53fa0c1da6199f4085af666fde00e8a034f152a52cd628 \
|
--hash=sha256:b589e5de2afe70bd4d53fa0c1da6199f4085af666fde00e8a034f152a52cd628 \
|
||||||
--hash=sha256:f89ffa663767f5a0585ea075f01020207e966a9c0f2b9ef56a57c7963a3f6f8e
|
--hash=sha256:f89ffa663767f5a0585ea075f01020207e966a9c0f2b9ef56a57c7963a3f6f8e
|
||||||
@@ -42,6 +50,7 @@ certifi==2026.7.22 \
|
|||||||
# httpcore
|
# httpcore
|
||||||
# httpx
|
# httpx
|
||||||
# requests
|
# requests
|
||||||
|
# safety
|
||||||
cffi==2.1.1 \
|
cffi==2.1.1 \
|
||||||
--hash=sha256:046bfc24911b37851ee1b51aab8bffe713d89c68c6a057b09484ce9fd5f69b4e \
|
--hash=sha256:046bfc24911b37851ee1b51aab8bffe713d89c68c6a057b09484ce9fd5f69b4e \
|
||||||
--hash=sha256:06c72bb76605a4b0cd0aad6930b69d4baf7dd5d806cfc409b824191099700e66 \
|
--hash=sha256:06c72bb76605a4b0cd0aad6930b69d4baf7dd5d806cfc409b824191099700e66 \
|
||||||
@@ -323,12 +332,19 @@ click==8.4.2 \
|
|||||||
--hash=sha256:e6f9f66136c816745b9d65817da91d61d957fb16e02e4dcd0552553c5a197b76
|
--hash=sha256:e6f9f66136c816745b9d65817da91d61d957fb16e02e4dcd0552553c5a197b76
|
||||||
# via
|
# via
|
||||||
# click-option-group
|
# click-option-group
|
||||||
|
# nltk
|
||||||
|
# safety
|
||||||
# semgrep
|
# semgrep
|
||||||
|
# typer
|
||||||
# uvicorn
|
# uvicorn
|
||||||
click-option-group==0.5.9 \
|
click-option-group==0.5.9 \
|
||||||
--hash=sha256:ad2599248bd373e2e19bec5407967c3eec1d0d4fc4a5e77b08a0481e75991080 \
|
--hash=sha256:ad2599248bd373e2e19bec5407967c3eec1d0d4fc4a5e77b08a0481e75991080 \
|
||||||
--hash=sha256:f94ed2bc4cf69052e0f29592bd1e771a1789bd7bfc482dd0bc482134aff95823
|
--hash=sha256:f94ed2bc4cf69052e0f29592bd1e771a1789bd7bfc482dd0bc482134aff95823
|
||||||
# via semgrep
|
# via semgrep
|
||||||
|
cloudpickle==3.1.2 \
|
||||||
|
--hash=sha256:7fda9eb655c9c230dab534f1983763de5835249750e85fbcef43aaa30a9a2414 \
|
||||||
|
--hash=sha256:9acb47f6afd73f60dc1df93bb801b472f05ff42fa6c84167d25cb206be1fbf4a
|
||||||
|
# via joblib
|
||||||
colorama==0.4.6 \
|
colorama==0.4.6 \
|
||||||
--hash=sha256:08695f5cb7ed6e0531a20572697297273c47b8cae5a63ffc6d6ed5c201be6e44 \
|
--hash=sha256:08695f5cb7ed6e0531a20572697297273c47b8cae5a63ffc6d6ed5c201be6e44 \
|
||||||
--hash=sha256:4f1d9991f5acc0ca119f9d443620b77f9d6b33703e51011c16baf57afb285fc6
|
--hash=sha256:4f1d9991f5acc0ca119f9d443620b77f9d6b33703e51011c16baf57afb285fc6
|
||||||
@@ -380,7 +396,20 @@ cryptography==50.0.1 \
|
|||||||
--hash=sha256:fc3ed7ebd2a8c96f5b166de0ab9b624996bef3b07bbeb19364dfb78222c22c80 \
|
--hash=sha256:fc3ed7ebd2a8c96f5b166de0ab9b624996bef3b07bbeb19364dfb78222c22c80 \
|
||||||
--hash=sha256:fd3718b960d0b5dd213cdf03f3bcb7000e69dda0de8b956061947ff6bcff5558 \
|
--hash=sha256:fd3718b960d0b5dd213cdf03f3bcb7000e69dda0de8b956061947ff6bcff5558 \
|
||||||
--hash=sha256:ff838d62ec1bfce4f9ba7fa16f4a7b554cd8d0c299e6be37502161a660c84eef
|
--hash=sha256:ff838d62ec1bfce4f9ba7fa16f4a7b554cd8d0c299e6be37502161a660c84eef
|
||||||
# via pyjwt
|
# via
|
||||||
|
# authlib
|
||||||
|
# joserfc
|
||||||
|
# pyjwt
|
||||||
|
defusedxml==0.7.1 \
|
||||||
|
--hash=sha256:1bb3032db185915b62d7c6209c5a8792be6a32ab2fedacc84e01b52c51aa3e69 \
|
||||||
|
--hash=sha256:a352e7e428770286cc899e2542b6cdaedb2b4953ff269a210103ec58f6198a61
|
||||||
|
# via nltk
|
||||||
|
dparse==0.6.4 \
|
||||||
|
--hash=sha256:90b29c39e3edc36c6284c82c4132648eaf28a01863eb3c231c2512196132201a \
|
||||||
|
--hash=sha256:fbab4d50d54d0e739fbb4dedfc3d92771003a5b9aa8545ca7a7045e3b174af57
|
||||||
|
# via
|
||||||
|
# safety
|
||||||
|
# safety-schemas
|
||||||
exceptiongroup==1.2.2 \
|
exceptiongroup==1.2.2 \
|
||||||
--hash=sha256:3111b9d131c238bec2f8f516e123e14ba243563fb135d3fe885990585aa7795b \
|
--hash=sha256:3111b9d131c238bec2f8f516e123e14ba243563fb135d3fe885990585aa7795b \
|
||||||
--hash=sha256:47c2edf7c6738fafb49fd34290706d1a1a2f4d1c6df275526b62cbb4aa5393cc
|
--hash=sha256:47c2edf7c6738fafb49fd34290706d1a1a2f4d1c6df275526b62cbb4aa5393cc
|
||||||
@@ -389,6 +418,10 @@ face==26.0.1 \
|
|||||||
--hash=sha256:8183d94bc248baaea855a9f8445f97a22a9988908e60abddccc6e251da77c4c6 \
|
--hash=sha256:8183d94bc248baaea855a9f8445f97a22a9988908e60abddccc6e251da77c4c6 \
|
||||||
--hash=sha256:ab0a83c37c9789dce658a67a9a80eafaa113c9ec37c5a9d950ff5480542a062d
|
--hash=sha256:ab0a83c37c9789dce658a67a9a80eafaa113c9ec37c5a9d950ff5480542a062d
|
||||||
# via glom
|
# via glom
|
||||||
|
filelock==3.32.4 \
|
||||||
|
--hash=sha256:22e58ca3b1ae3b98993b762d7338367ae64fe50252bf78d59da3bfebcdf1cedd \
|
||||||
|
--hash=sha256:2bde2e4cf732e0153406d8a7bc80620ecf5e621fe0d25e41143c4e3b4733ff30
|
||||||
|
# via safety
|
||||||
glom==25.12.0 \
|
glom==25.12.0 \
|
||||||
--hash=sha256:1ae7da88be3693df40ad27bdf57a765a55c075c86c971bcddd67927403eb0069 \
|
--hash=sha256:1ae7da88be3693df40ad27bdf57a765a55c075c86c971bcddd67927403eb0069 \
|
||||||
--hash=sha256:b9f21e77f71a6576a43864e85066b8cc3f0f778d0d50961563f8981377a6dcb1
|
--hash=sha256:b9f21e77f71a6576a43864e85066b8cc3f0f778d0d50961563f8981377a6dcb1
|
||||||
@@ -410,7 +443,9 @@ httpcore==1.0.9 \
|
|||||||
httpx==0.28.1 \
|
httpx==0.28.1 \
|
||||||
--hash=sha256:75e98c5f16b0f35b567856f597f06ff2270a374470a5c2392242528e3e3e42fc \
|
--hash=sha256:75e98c5f16b0f35b567856f597f06ff2270a374470a5c2392242528e3e3e42fc \
|
||||||
--hash=sha256:d909fcccc110f8c7faf814ca82a9a4d816bc5a6dbfea25d6591d6985b8ba59ad
|
--hash=sha256:d909fcccc110f8c7faf814ca82a9a4d816bc5a6dbfea25d6591d6985b8ba59ad
|
||||||
# via mcp
|
# via
|
||||||
|
# mcp
|
||||||
|
# safety
|
||||||
httpx-sse==0.4.3 \
|
httpx-sse==0.4.3 \
|
||||||
--hash=sha256:0ac1c9fe3c0afad2e0ebb25a934a59f4c7823b60792691f779fad2c5568830fc \
|
--hash=sha256:0ac1c9fe3c0afad2e0ebb25a934a59f4c7823b60792691f779fad2c5568830fc \
|
||||||
--hash=sha256:9b1ed0127459a66014aec3c56bebd93da3c1bc8bb6618c8082039a44889a755d
|
--hash=sha256:9b1ed0127459a66014aec3c56bebd93da3c1bc8bb6618c8082039a44889a755d
|
||||||
@@ -426,6 +461,18 @@ importlib-metadata==8.7.1 \
|
|||||||
--hash=sha256:49fef1ae6440c182052f407c8d34a68f72efc36db9ca90dc0113398f2fdde8bb \
|
--hash=sha256:49fef1ae6440c182052f407c8d34a68f72efc36db9ca90dc0113398f2fdde8bb \
|
||||||
--hash=sha256:5a1f80bf1daa489495071efbb095d75a634cf28a8bc299581244063b53176151
|
--hash=sha256:5a1f80bf1daa489495071efbb095d75a634cf28a8bc299581244063b53176151
|
||||||
# via opentelemetry-api
|
# via opentelemetry-api
|
||||||
|
jinja2==3.1.6 \
|
||||||
|
--hash=sha256:0137fb05990d35f1275a587e9aee6d56da821fc83491a0fb838183be43f66d6d \
|
||||||
|
--hash=sha256:85ece4451f492d0c13c5dd7c13a64681a86afae63a5f347908daf103ce6d2f67
|
||||||
|
# via safety
|
||||||
|
joblib==1.6.0 \
|
||||||
|
--hash=sha256:2ccc96785b12046c08fd6d55839c12857831b54a3c1673ffadd2f04bfc4eda03 \
|
||||||
|
--hash=sha256:3dbbf9f6e4b592a2357b854608e980fe6390d131d7a82f011a377ef2ebef7aba
|
||||||
|
# via nltk
|
||||||
|
joserfc==1.7.5 \
|
||||||
|
--hash=sha256:add2c2c84e8373b084d526a8b53daba5d7a513a118cd2dcd9fc9f979d0922159 \
|
||||||
|
--hash=sha256:d5ff536e658e17664f8c1b1ab60dc4aa62aa973fcef1edd33cc44bda45d6f5ea
|
||||||
|
# via authlib
|
||||||
jq==1.12.0 \
|
jq==1.12.0 \
|
||||||
--hash=sha256:02112ca560f90c6b1ea31829bb7777fbc5b1f1d13f78b2c6ce5cefa8233cee7e \
|
--hash=sha256:02112ca560f90c6b1ea31829bb7777fbc5b1f1d13f78b2c6ce5cefa8233cee7e \
|
||||||
--hash=sha256:067ea0d3ee2cd7f7ba9c5d5c1925b9b0f83e1869c97a65ef11d8d76bd91ece6e \
|
--hash=sha256:067ea0d3ee2cd7f7ba9c5d5c1925b9b0f83e1869c97a65ef11d8d76bd91ece6e \
|
||||||
@@ -502,6 +549,101 @@ markdown-it-py==4.2.0 \
|
|||||||
--hash=sha256:04a21681d6fbb623de53f6f364d352309d4094dd4194040a10fd51833e418d49 \
|
--hash=sha256:04a21681d6fbb623de53f6f364d352309d4094dd4194040a10fd51833e418d49 \
|
||||||
--hash=sha256:9f7ebbcd14fe59494226453aed97c1070d83f8d24b6fc3a3bcf9a38092641c4a
|
--hash=sha256:9f7ebbcd14fe59494226453aed97c1070d83f8d24b6fc3a3bcf9a38092641c4a
|
||||||
# via rich
|
# via rich
|
||||||
|
markupsafe==3.0.3 \
|
||||||
|
--hash=sha256:0303439a41979d9e74d18ff5e2dd8c43ed6c6001fd40e5bf2e43f7bd9bbc523f \
|
||||||
|
--hash=sha256:068f375c472b3e7acbe2d5318dea141359e6900156b5b2ba06a30b169086b91a \
|
||||||
|
--hash=sha256:0bf2a864d67e76e5c9a34dc26ec616a66b9888e25e7b9460e1c76d3293bd9dbf \
|
||||||
|
--hash=sha256:0db14f5dafddbb6d9208827849fad01f1a2609380add406671a26386cdf15a19 \
|
||||||
|
--hash=sha256:0eb9ff8191e8498cca014656ae6b8d61f39da5f95b488805da4bb029cccbfbaf \
|
||||||
|
--hash=sha256:0f4b68347f8c5eab4a13419215bdfd7f8c9b19f2b25520968adfad23eb0ce60c \
|
||||||
|
--hash=sha256:1085e7fbddd3be5f89cc898938f42c0b3c711fdcb37d75221de2666af647c175 \
|
||||||
|
--hash=sha256:116bb52f642a37c115f517494ea5feb03889e04df47eeff5b130b1808ce7c219 \
|
||||||
|
--hash=sha256:12c63dfb4a98206f045aa9563db46507995f7ef6d83b2f68eda65c307c6829eb \
|
||||||
|
--hash=sha256:133a43e73a802c5562be9bbcd03d090aa5a1fe899db609c29e8c8d815c5f6de6 \
|
||||||
|
--hash=sha256:1353ef0c1b138e1907ae78e2f6c63ff67501122006b0f9abad68fda5f4ffc6ab \
|
||||||
|
--hash=sha256:15d939a21d546304880945ca1ecb8a039db6b4dc49b2c5a400387cdae6a62e26 \
|
||||||
|
--hash=sha256:177b5253b2834fe3678cb4a5f0059808258584c559193998be2601324fdeafb1 \
|
||||||
|
--hash=sha256:1872df69a4de6aead3491198eaf13810b565bdbeec3ae2dc8780f14458ec73ce \
|
||||||
|
--hash=sha256:1b4b79e8ebf6b55351f0d91fe80f893b4743f104bff22e90697db1590e47a218 \
|
||||||
|
--hash=sha256:1b52b4fb9df4eb9ae465f8d0c228a00624de2334f216f178a995ccdcf82c4634 \
|
||||||
|
--hash=sha256:1ba88449deb3de88bd40044603fafffb7bc2b055d626a330323a9ed736661695 \
|
||||||
|
--hash=sha256:1cc7ea17a6824959616c525620e387f6dd30fec8cb44f649e31712db02123dad \
|
||||||
|
--hash=sha256:218551f6df4868a8d527e3062d0fb968682fe92054e89978594c28e642c43a73 \
|
||||||
|
--hash=sha256:26a5784ded40c9e318cfc2bdb30fe164bdb8665ded9cd64d500a34fb42067b1c \
|
||||||
|
--hash=sha256:2713baf880df847f2bece4230d4d094280f4e67b1e813eec43b4c0e144a34ffe \
|
||||||
|
--hash=sha256:2a15a08b17dd94c53a1da0438822d70ebcd13f8c3a95abe3a9ef9f11a94830aa \
|
||||||
|
--hash=sha256:2f981d352f04553a7171b8e44369f2af4055f888dfb147d55e42d29e29e74559 \
|
||||||
|
--hash=sha256:32001d6a8fc98c8cb5c947787c5d08b0a50663d139f1305bac5885d98d9b40fa \
|
||||||
|
--hash=sha256:3524b778fe5cfb3452a09d31e7b5adefeea8c5be1d43c4f810ba09f2ceb29d37 \
|
||||||
|
--hash=sha256:3537e01efc9d4dccdf77221fb1cb3b8e1a38d5428920e0657ce299b20324d758 \
|
||||||
|
--hash=sha256:35add3b638a5d900e807944a078b51922212fb3dedb01633a8defc4b01a3c85f \
|
||||||
|
--hash=sha256:38664109c14ffc9e7437e86b4dceb442b0096dfe3541d7864d9cbe1da4cf36c8 \
|
||||||
|
--hash=sha256:3a7e8ae81ae39e62a41ec302f972ba6ae23a5c5396c8e60113e9066ef893da0d \
|
||||||
|
--hash=sha256:3b562dd9e9ea93f13d53989d23a7e775fdfd1066c33494ff43f5418bc8c58a5c \
|
||||||
|
--hash=sha256:457a69a9577064c05a97c41f4e65148652db078a3a509039e64d3467b9e7ef97 \
|
||||||
|
--hash=sha256:4bd4cd07944443f5a265608cc6aab442e4f74dff8088b0dfc8238647b8f6ae9a \
|
||||||
|
--hash=sha256:4e885a3d1efa2eadc93c894a21770e4bc67899e3543680313b09f139e149ab19 \
|
||||||
|
--hash=sha256:4faffd047e07c38848ce017e8725090413cd80cbc23d86e55c587bf979e579c9 \
|
||||||
|
--hash=sha256:509fa21c6deb7a7a273d629cf5ec029bc209d1a51178615ddf718f5918992ab9 \
|
||||||
|
--hash=sha256:5678211cb9333a6468fb8d8be0305520aa073f50d17f089b5b4b477ea6e67fdc \
|
||||||
|
--hash=sha256:591ae9f2a647529ca990bc681daebdd52c8791ff06c2bfa05b65163e28102ef2 \
|
||||||
|
--hash=sha256:5a7d5dc5140555cf21a6fefbdbf8723f06fcd2f63ef108f2854de715e4422cb4 \
|
||||||
|
--hash=sha256:69c0b73548bc525c8cb9a251cddf1931d1db4d2258e9599c28c07ef3580ef354 \
|
||||||
|
--hash=sha256:6b5420a1d9450023228968e7e6a9ce57f65d148ab56d2313fcd589eee96a7a50 \
|
||||||
|
--hash=sha256:722695808f4b6457b320fdc131280796bdceb04ab50fe1795cd540799ebe1698 \
|
||||||
|
--hash=sha256:729586769a26dbceff69f7a7dbbf59ab6572b99d94576a5592625d5b411576b9 \
|
||||||
|
--hash=sha256:77f0643abe7495da77fb436f50f8dab76dbc6e5fd25d39589a0f1fe6548bfa2b \
|
||||||
|
--hash=sha256:795e7751525cae078558e679d646ae45574b47ed6e7771863fcc079a6171a0fc \
|
||||||
|
--hash=sha256:7be7b61bb172e1ed687f1754f8e7484f1c8019780f6f6b0786e76bb01c2ae115 \
|
||||||
|
--hash=sha256:7c3fb7d25180895632e5d3148dbdc29ea38ccb7fd210aa27acbd1201a1902c6e \
|
||||||
|
--hash=sha256:7e68f88e5b8799aa49c85cd116c932a1ac15caaa3f5db09087854d218359e485 \
|
||||||
|
--hash=sha256:83891d0e9fb81a825d9a6d61e3f07550ca70a076484292a70fde82c4b807286f \
|
||||||
|
--hash=sha256:8485f406a96febb5140bfeca44a73e3ce5116b2501ac54fe953e488fb1d03b12 \
|
||||||
|
--hash=sha256:8709b08f4a89aa7586de0aadc8da56180242ee0ada3999749b183aa23df95025 \
|
||||||
|
--hash=sha256:8f71bc33915be5186016f675cd83a1e08523649b0e33efdb898db577ef5bb009 \
|
||||||
|
--hash=sha256:915c04ba3851909ce68ccc2b8e2cd691618c4dc4c4232fb7982bca3f41fd8c3d \
|
||||||
|
--hash=sha256:949b8d66bc381ee8b007cd945914c721d9aba8e27f71959d750a46f7c282b20b \
|
||||||
|
--hash=sha256:94c6f0bb423f739146aec64595853541634bde58b2135f27f61c1ffd1cd4d16a \
|
||||||
|
--hash=sha256:9a1abfdc021a164803f4d485104931fb8f8c1efd55bc6b748d2f5774e78b62c5 \
|
||||||
|
--hash=sha256:9b79b7a16f7fedff2495d684f2b59b0457c3b493778c9eed31111be64d58279f \
|
||||||
|
--hash=sha256:a320721ab5a1aba0a233739394eb907f8c8da5c98c9181d1161e77a0c8e36f2d \
|
||||||
|
--hash=sha256:a4afe79fb3de0b7097d81da19090f4df4f8d3a2b3adaa8764138aac2e44f3af1 \
|
||||||
|
--hash=sha256:ad2cf8aa28b8c020ab2fc8287b0f823d0a7d8630784c31e9ee5edea20f406287 \
|
||||||
|
--hash=sha256:b8512a91625c9b3da6f127803b166b629725e68af71f8184ae7e7d54686a56d6 \
|
||||||
|
--hash=sha256:bc51efed119bc9cfdf792cdeaa4d67e8f6fcccab66ed4bfdd6bde3e59bfcbb2f \
|
||||||
|
--hash=sha256:bdc919ead48f234740ad807933cdf545180bfbe9342c2bb451556db2ed958581 \
|
||||||
|
--hash=sha256:bdd37121970bfd8be76c5fb069c7751683bdf373db1ed6c010162b2a130248ed \
|
||||||
|
--hash=sha256:be8813b57049a7dc738189df53d69395eba14fb99345e0a5994914a3864c8a4b \
|
||||||
|
--hash=sha256:c0c0b3ade1c0b13b936d7970b1d37a57acde9199dc2aecc4c336773e1d86049c \
|
||||||
|
--hash=sha256:c47a551199eb8eb2121d4f0f15ae0f923d31350ab9280078d1e5f12b249e0026 \
|
||||||
|
--hash=sha256:c4ffb7ebf07cfe8931028e3e4c85f0357459a3f9f9490886198848f4fa002ec8 \
|
||||||
|
--hash=sha256:ccfcd093f13f0f0b7fdd0f198b90053bf7b2f02a3927a30e63f3ccc9df56b676 \
|
||||||
|
--hash=sha256:d2ee202e79d8ed691ceebae8e0486bd9a2cd4794cec4824e1c99b6f5009502f6 \
|
||||||
|
--hash=sha256:d53197da72cc091b024dd97249dfc7794d6a56530370992a5e1a08983ad9230e \
|
||||||
|
--hash=sha256:d6dd0be5b5b189d31db7cda48b91d7e0a9795f31430b7f271219ab30f1d3ac9d \
|
||||||
|
--hash=sha256:d88b440e37a16e651bda4c7c2b930eb586fd15ca7406cb39e211fcff3bf3017d \
|
||||||
|
--hash=sha256:de8a88e63464af587c950061a5e6a67d3632e36df62b986892331d4620a35c01 \
|
||||||
|
--hash=sha256:df2449253ef108a379b8b5d6b43f4b1a8e81a061d6537becd5582fba5f9196d7 \
|
||||||
|
--hash=sha256:e1c1493fb6e50ab01d20a22826e57520f1284df32f2d8601fdd90b6304601419 \
|
||||||
|
--hash=sha256:e1cf1972137e83c5d4c136c43ced9ac51d0e124706ee1c8aa8532c1287fa8795 \
|
||||||
|
--hash=sha256:e2103a929dfa2fcaf9bb4e7c091983a49c9ac3b19c9061b6d5427dd7d14d81a1 \
|
||||||
|
--hash=sha256:e56b7d45a839a697b5eb268c82a71bd8c7f6c94d6fd50c3d577fa39a9f1409f5 \
|
||||||
|
--hash=sha256:e8afc3f2ccfa24215f8cb28dcf43f0113ac3c37c2f0f0806d8c70e4228c5cf4d \
|
||||||
|
--hash=sha256:e8fc20152abba6b83724d7ff268c249fa196d8259ff481f3b1476383f8f24e42 \
|
||||||
|
--hash=sha256:eaa9599de571d72e2daf60164784109f19978b327a3910d3e9de8c97b5b70cfe \
|
||||||
|
--hash=sha256:ec15a59cf5af7be74194f7ab02d0f59a62bdcf1a537677ce67a2537c9b87fcda \
|
||||||
|
--hash=sha256:f190daf01f13c72eac4efd5c430a8de82489d9cff23c364c3ea822545032993e \
|
||||||
|
--hash=sha256:f34c41761022dd093b4b6896d4810782ffbabe30f2d443ff5f083e0cbbb8c737 \
|
||||||
|
--hash=sha256:f3e98bb3798ead92273dc0e5fd0f31ade220f59a266ffd8a4f6065e0a3ce0523 \
|
||||||
|
--hash=sha256:f42d0984e947b8adf7dd6dde396e720934d12c506ce84eea8476409563607591 \
|
||||||
|
--hash=sha256:f71a396b3bf33ecaa1626c255855702aca4d3d9fea5e051b41ac59a9c1c41edc \
|
||||||
|
--hash=sha256:f9e130248f4462aaa8e2552d547f36ddadbeaa573879158d721bbd33dfe4743a \
|
||||||
|
--hash=sha256:fed51ac40f757d41b7c48425901843666a6677e3e8eb0abcff09e4ba6e664f50
|
||||||
|
# via jinja2
|
||||||
|
marshmallow==4.3.1 \
|
||||||
|
--hash=sha256:e65accfbe277546df92ed7996a678c90e063e9a7c2a2f5e03f7d0b90e3768c42 \
|
||||||
|
--hash=sha256:fb6b8048af08d4ab061610d5b7d3696a7e4c95337dbda880edb9f95812cabc20
|
||||||
|
# via safety
|
||||||
mcp==1.29.0 \
|
mcp==1.29.0 \
|
||||||
--hash=sha256:52d01f334de1868cc3bb2d6604931126a67631f99a6c5d3b82ba47290315ec36 \
|
--hash=sha256:52d01f334de1868cc3bb2d6604931126a67631f99a6c5d3b82ba47290315ec36 \
|
||||||
--hash=sha256:f5a075bb611f23d6f4d080c6a1699fa62772eebc562ba9e66b306ddde1c755f7
|
--hash=sha256:f5a075bb611f23d6f4d080c6a1699fa62772eebc562ba9e66b306ddde1c755f7
|
||||||
@@ -510,6 +652,10 @@ mdurl==0.1.2 \
|
|||||||
--hash=sha256:84008a41e51615a49fc9966191ff91509e3c40b939176e643fd50a5c2196b8f8 \
|
--hash=sha256:84008a41e51615a49fc9966191ff91509e3c40b939176e643fd50a5c2196b8f8 \
|
||||||
--hash=sha256:bb413d29f5eea38f31dd4754dd7377d4465116fb207585f97bf925588687c1ba
|
--hash=sha256:bb413d29f5eea38f31dd4754dd7377d4465116fb207585f97bf925588687c1ba
|
||||||
# via markdown-it-py
|
# via markdown-it-py
|
||||||
|
nltk==3.10.3 \
|
||||||
|
--hash=sha256:bb9327a461c3811c2fa4900e03840401f2126adfb30c0072827c433bd2444ea4 \
|
||||||
|
--hash=sha256:ff9598a8e20518ee0d557745890cc4435b9578489e2dcbc69c4f81fa060caf7c
|
||||||
|
# via safety
|
||||||
opentelemetry-api==1.37.0 \
|
opentelemetry-api==1.37.0 \
|
||||||
--hash=sha256:540735b120355bd5112738ea53621f8d5edb35ebcd6fe21ada3ab1c61d1cd9a7 \
|
--hash=sha256:540735b120355bd5112738ea53621f8d5edb35ebcd6fe21ada3ab1c61d1cd9a7 \
|
||||||
--hash=sha256:accf2024d3e89faec14302213bc39550ec0f4095d1cf5ca688e1bfb1c8612f47
|
--hash=sha256:accf2024d3e89faec14302213bc39550ec0f4095d1cf5ca688e1bfb1c8612f47
|
||||||
@@ -570,7 +716,10 @@ packaging==26.3 \
|
|||||||
--hash=sha256:94edc256424af38762eb31306eed28beb9f0efc50a8837492c9d6fd6004aed79 \
|
--hash=sha256:94edc256424af38762eb31306eed28beb9f0efc50a8837492c9d6fd6004aed79 \
|
||||||
--hash=sha256:d7193f7c8e4e93f444fde0262bf90af30e16fa0ad0ad44cb553c87339b23cd1c
|
--hash=sha256:d7193f7c8e4e93f444fde0262bf90af30e16fa0ad0ad44cb553c87339b23cd1c
|
||||||
# via
|
# via
|
||||||
|
# dparse
|
||||||
# opentelemetry-instrumentation
|
# opentelemetry-instrumentation
|
||||||
|
# safety
|
||||||
|
# safety-schemas
|
||||||
# semgrep
|
# semgrep
|
||||||
peewee==3.19.0 \
|
peewee==3.19.0 \
|
||||||
--hash=sha256:de220b94766e6008c466e00ce4ba5299b9a832117d9eb36d45d0062f3cfd7417 \
|
--hash=sha256:de220b94766e6008c466e00ce4ba5299b9a832117d9eb36d45d0062f3cfd7417 \
|
||||||
@@ -600,6 +749,8 @@ pydantic==2.13.5 \
|
|||||||
# via
|
# via
|
||||||
# mcp
|
# mcp
|
||||||
# pydantic-settings
|
# pydantic-settings
|
||||||
|
# safety
|
||||||
|
# safety-schemas
|
||||||
pydantic-core==2.46.5 \
|
pydantic-core==2.46.5 \
|
||||||
--hash=sha256:013d6f3483d81e02e7c328831808f336c8596ee33b4bd4026b9ffb1e960b8942 \
|
--hash=sha256:013d6f3483d81e02e7c328831808f336c8596ee33b4bd4026b9ffb1e960b8942 \
|
||||||
--hash=sha256:03b9666e41e35d8909852ba191a0607520f81b74eaf12ccf8737005dbb313821 \
|
--hash=sha256:03b9666e41e35d8909852ba191a0607520f81b74eaf12ccf8737005dbb313821 \
|
||||||
@@ -825,6 +976,122 @@ referencing==0.37.0 \
|
|||||||
# via
|
# via
|
||||||
# jsonschema
|
# jsonschema
|
||||||
# jsonschema-specifications
|
# jsonschema-specifications
|
||||||
|
regex==2026.8.31 \
|
||||||
|
--hash=sha256:0087dfa879bf01c5eb290848c7de22f717d8d4218a997080e63ae4813bc55104 \
|
||||||
|
--hash=sha256:026a7cd6c20a2a5bf3249a4a1c7f076af86b17188e2ffd17722e2ed24f433f9a \
|
||||||
|
--hash=sha256:073b9cb8c44e197a4d1d8b819a3329f6b20866d83d2700f78b9d33e1f1a75116 \
|
||||||
|
--hash=sha256:0abb98dd76a3ffe3b401fe93aadac135ecd6ba4a71d7b4be4a333de8d691e834 \
|
||||||
|
--hash=sha256:0bb6121dbf90c7de42610459398a81cbb90bc870e2cc003248f3f2b65d45f2b6 \
|
||||||
|
--hash=sha256:0ec77a1ce2350c74fe3821d1c6555107d41f6969c369f4ee197a10cec97632ec \
|
||||||
|
--hash=sha256:0ee80c5d20a62ae819f39a4f5b0c7f1dbbeb28186de6138840eb8c138e96f99e \
|
||||||
|
--hash=sha256:13f036b42889e8cad5f1ee2eadb48c656b2f44c5944035e0f697cb6ef81757ba \
|
||||||
|
--hash=sha256:15e9e862c6e905ef66ea5f019deb5ac5fdeebf8fc134ea4c7b5d5c2eb7bdcdd8 \
|
||||||
|
--hash=sha256:18ac65e72e8454343df30ca1d8a4ad604d3419b96e0ef8e2dc3a69642bb557b4 \
|
||||||
|
--hash=sha256:18c7e0348286f5073867d339d7cab60ed200b77b48d7a9be4edbcdc2c996a62b \
|
||||||
|
--hash=sha256:1930ade186f2b519fe9c4bdfd3a77410e469bd91423a995888b91f3beb12679b \
|
||||||
|
--hash=sha256:1e74e38c5a9ed3a70a0e0a89498eb664211b97c162d77b1131f37636779f36b4 \
|
||||||
|
--hash=sha256:222c906a555bdbd5322f15778bb2b4f238c26e1d52c9445f1e50f5e4452909b3 \
|
||||||
|
--hash=sha256:241c614ab811e29f2e67e2828404dd10a2dc675ec2c75a6017ec310fd09117b9 \
|
||||||
|
--hash=sha256:26a6ddc85198558b0c74b856f6440132d6f97248c22589bf52cf13df2fa44fdc \
|
||||||
|
--hash=sha256:2c5f4fc5463ac732ed49cb87ffdf2eab3d909a0df4100211ce4be3af1ad729cb \
|
||||||
|
--hash=sha256:2d28ad9d016ac681843b059ddca376b9ff833ec218c938035d925c8af44c6de7 \
|
||||||
|
--hash=sha256:34c8d36a5f70c16e3f406ae1c93a47ea4b2a40e29b02639cf41915b6fea5ce26 \
|
||||||
|
--hash=sha256:360c916117c988b120ba05aa106cd3c1aa7c0f4575a2db0d605d502b4ee334f4 \
|
||||||
|
--hash=sha256:38179404d70581402831c2c0de0c8ec3483d272beab2244095cb09b4eeb30ef7 \
|
||||||
|
--hash=sha256:3b3a020f2a43e9016624047ecc15cd0d472c11dfbe4d12fe030f574570467f35 \
|
||||||
|
--hash=sha256:3e139e792b016a614b9af4a43e036b259a8d32f751e9b5bda77b4af652ad8a17 \
|
||||||
|
--hash=sha256:40f4cdf6d38663cf8f56a52edde25ca6dbfb857f5a7d49cd7de3e0e1a0883bf4 \
|
||||||
|
--hash=sha256:4301de5a58a28fe95b6a865d3b97b5cea073bb4c6ad743211c32b004f32d5096 \
|
||||||
|
--hash=sha256:43581e1f0c1f624cb7e2e8195c443f6e3004fc376bd12d644cdc8e613c973323 \
|
||||||
|
--hash=sha256:453e9ffb310eede3f35303d7fb2e891382c98888d54f162e5a2e0174d1b75331 \
|
||||||
|
--hash=sha256:45537c0d48a84dd0f840ea7c308445ad1e83a04d28d6fc394d71ad24f9f55d2b \
|
||||||
|
--hash=sha256:45b0450d6ae52e2dfcdb5e58987b829ed5fc01b709fc5ff09a1e81ab13c5262a \
|
||||||
|
--hash=sha256:4c3ac1eec883a1d0fbba167e90bb1beb72289e765966b464f9b333090dfcae2e \
|
||||||
|
--hash=sha256:50a8677cca3d4df536776380161744d41ea5001f99cc2c4638e6b0625839fa61 \
|
||||||
|
--hash=sha256:520b14582a59f43ba9ba595938349e70238009f8deb8c35d5bbfe33e44fd0ba9 \
|
||||||
|
--hash=sha256:52f03cd8f259d8fb482a9e142ad17c8d1c931a69a7a932922f2222df05875d59 \
|
||||||
|
--hash=sha256:56f7516b00f720231b26fdcd41ac13cceab7a8c1c903b1ab98e173b0962a771d \
|
||||||
|
--hash=sha256:66df1812cf0fd5f0f59e4341c54247a15397354ee01231e1c2620b08032f3361 \
|
||||||
|
--hash=sha256:69c42c35758cf46c31d976d63c79fbbcb114fe192aa4c721c734204d0e3d7555 \
|
||||||
|
--hash=sha256:69fbc60c1c34790037cfd350dd1600436fdfea9ca221761c614fc5e633c7cabd \
|
||||||
|
--hash=sha256:6d5537087013e5ce841b9d0f19a564f18f33fa79489a7e8865f5a38ba2a4de7d \
|
||||||
|
--hash=sha256:6d5c9841dd924437e34d43bdbecbb31bc1a01c57bd974af8e1a0a98b0a7a731c \
|
||||||
|
--hash=sha256:6fcbf68a10dd6a564c737147e013e5dea6180c032e3c363629cf4d0f9d258752 \
|
||||||
|
--hash=sha256:7010dae7e7064ee091703cafce0143693e56931bb3d21a82483bb96ad8a37751 \
|
||||||
|
--hash=sha256:722c2dba81c28494dae77f06c0fd33f0ad215e1b7cc6e2b0f3bad36656413f84 \
|
||||||
|
--hash=sha256:75b888caf9469df3826876ae0e2f92f37e7bbad0455cfa028852d99815af9dd0 \
|
||||||
|
--hash=sha256:75cc2d43987040df8655c25b47c1d452c7d59b28df108d7b2c19a003d021601f \
|
||||||
|
--hash=sha256:79c7b6bd11620dc722a94e160965fa0e64124ca8841afaf9683d8fa659431cf5 \
|
||||||
|
--hash=sha256:7aa0688964b66ac50e2bf3b04b9e88bdab58fa5ea8130b403d72668df6f54cb9 \
|
||||||
|
--hash=sha256:7c06a4cbe33f8ad72c3bd9590630c07e55c7a7c581253d287b6ca645e2879051 \
|
||||||
|
--hash=sha256:7daf31011e73c16f8b824bc6a6992f0de8a9ae13133001d757668c852bcc6502 \
|
||||||
|
--hash=sha256:81391983ff052f922baebb0955a3be455d5731351b3a93e0638a8150bd44b8b5 \
|
||||||
|
--hash=sha256:8231dfdbb4baf59d35a10fc1115846bdcc43b30ab6ec8809ec807bfeea48a119 \
|
||||||
|
--hash=sha256:861a12bd9e8d3f26a9a36cc1b3426edacc70395b2e4f37c1402f40345e9c06db \
|
||||||
|
--hash=sha256:868d9113a744f2bfffa31197cadcda5b7fc3951a8621dd5899f9c0e4208ca196 \
|
||||||
|
--hash=sha256:897c2e301226fdfaf1a0c68219607718c40699df82dff09fd366b489b4c6e6d8 \
|
||||||
|
--hash=sha256:8b6bcc66372b493faa2b6153cd16a44db3bfa316411f81c4ba5d0ffa693244df \
|
||||||
|
--hash=sha256:8b7f1bdf1f36555fa0317f4f6cbbd5312f886edf9f2a41c8c298ffb9ad9f4a1a \
|
||||||
|
--hash=sha256:8d3e98b55372aa36b1e046a56a10f13cf0ef782ad6c86dbd64f3897c7e7a7a02 \
|
||||||
|
--hash=sha256:91a478b9a76b7f2b4cc704ec5f438041012ae7914716f8de0d56c11c9706203f \
|
||||||
|
--hash=sha256:9350fd448a6442ae27853ab9d4b8d5a0bcb6d7774923a4fdfddd104c4458b35f \
|
||||||
|
--hash=sha256:95c25f91b7c3f8121946e175a731eccf097dfeff065ab1204dbaad1ebf8ada6e \
|
||||||
|
--hash=sha256:976c265b3a42b806cf58afd3c5a64417e1bbd804289bf4abd38ea7395623531d \
|
||||||
|
--hash=sha256:98183eb943ebcd2e89fd9fcb4103bfafc5369cff9479561a5c96de2fe90cae68 \
|
||||||
|
--hash=sha256:98381539ee2dd88794f3ce6e40166f59b93e6e3ee9cd27dea9f2dd6b857f3dbc \
|
||||||
|
--hash=sha256:9a991b561615498877b042b13a788cc2f33c99087a9540627c397037c58ae795 \
|
||||||
|
--hash=sha256:9acbc6901bea11ad2f21d32b0790cbe2cb0194b521ea239231e1ee9627efd585 \
|
||||||
|
--hash=sha256:9b9e48a4ae2378c7bb29df0cbe2426cf0929ddbbae5819225c1fe133e6bb368d \
|
||||||
|
--hash=sha256:9fe2540d8da1bbf12f7c1b909a9ae47c2b343fa2a2084280c21ead1c9fb0e6f7 \
|
||||||
|
--hash=sha256:a1c9cd392daa08d3a3d5b663443a08071f4efbc1476f902142d51a229c60e852 \
|
||||||
|
--hash=sha256:a54f6b1b418e40b908ff9b9dd3e5fa638a2bd1bbe6e24180dc097c92b1deed0f \
|
||||||
|
--hash=sha256:a55bfb3914b760d5103d313a1053d301b2776f4677eb7f4d09f6420c625d97dd \
|
||||||
|
--hash=sha256:a679703a46574dcfbbae42acbc538d37653fa78dd2a3826f27c2dab386ea194d \
|
||||||
|
--hash=sha256:a75efe8109ebfaa5574aff49882fe471287ecb7959d96d29660cec937e5af1ce \
|
||||||
|
--hash=sha256:aac83eab8d47e3c290b9d30a34f94e3d888b7dd42f7cc45b8d204154cec3017b \
|
||||||
|
--hash=sha256:abd6b935adcd6c19733f20080a85972c6199cc9599dd8d16c9bbd1bbada569d8 \
|
||||||
|
--hash=sha256:aea17d86e7581e589fb8c43b70dc5f6588b1897390442536697a551bc66e2fd6 \
|
||||||
|
--hash=sha256:b40aee7f8df89d239943a932bfb53809f6b2c2ad53c049ee329100a54d3e1cfd \
|
||||||
|
--hash=sha256:b94165c6b98404ca40838852febd60df4fa6380dc0898f28dedaf5fca638e7ca \
|
||||||
|
--hash=sha256:bb1ca9e722c7270fb4267abee42cf8cfa97bc8e361b73839a50f00fcd2b76636 \
|
||||||
|
--hash=sha256:bb392c55059edb1bda593ee12218f5198a337535ff5e52f806c224c57b98716b \
|
||||||
|
--hash=sha256:bc00f39b7201fca5a15f12580f9dfb84b226323ad24043ec71b1132b5dbab711 \
|
||||||
|
--hash=sha256:bdbc6e87c9868ab2e7f29eed32b04583420df1b9b19e718f212e140c01f8b026 \
|
||||||
|
--hash=sha256:c01865f6a72c776064e4f58030e59f925e5fef32066aab3cb1a97be191f7bdd1 \
|
||||||
|
--hash=sha256:c72238cc48cd020f415e9dd3cba6c6b1af559d613358d282f7957cf61f0bcf6b \
|
||||||
|
--hash=sha256:c7ffcdf6fe74cedd4e36a9de2fb072b526a978e9b2d4fd2431edca96d80a67cd \
|
||||||
|
--hash=sha256:c9ba0b56ca6547e238323452178e5d9889886c99cdd17a4333d026f3c84471c5 \
|
||||||
|
--hash=sha256:c9c7a13d018f4f84503986564a543c2f7657a4bec4895f2c2cc584fb09d7429b \
|
||||||
|
--hash=sha256:caa959da9bb21394131eaf5c57698b47926ebada98c6796cfb4e754a52de001f \
|
||||||
|
--hash=sha256:cf427a3bebc873a2601601fc5e8453d1396b52d694ad65788fa2b22fe7b0f920 \
|
||||||
|
--hash=sha256:cf6c32d2a6bdaac692915ab81f28b62525d937abeac80149260db2c904a5df97 \
|
||||||
|
--hash=sha256:d27a3bdd19aa00974ac53ba14faea80ecef412f2d957c0071a869d7baea820f4 \
|
||||||
|
--hash=sha256:d59beef8054a851b2a3f42f56f94770981973699ab4c7f0b5f6984c26205b76c \
|
||||||
|
--hash=sha256:d84db4aaf4b5c5c4d512ce06420850c909865fa7d6223081dc8e9dbde7a83754 \
|
||||||
|
--hash=sha256:d9759f4cc91880cfafdb11b7b2bc83e34f2f16d103fd94f936d804cbfdb9c1aa \
|
||||||
|
--hash=sha256:dacc364aa1c06cb3fffb1705ff313cb3622c94d8c248f29e57bac2acadd77bf7 \
|
||||||
|
--hash=sha256:dbed5cea80c5a67c3f95f16d011d68174eb81a5efccf87a3ad0822b79d74baae \
|
||||||
|
--hash=sha256:deab998bd9314f7e93f519d3f62f1fd9e83a2db654f579cadac3968fbc1b5976 \
|
||||||
|
--hash=sha256:def853717c37661f59942c76ad06e060630f6e297257bcfb6f203d2daf497d41 \
|
||||||
|
--hash=sha256:dfc722cb60e40e6fefa483a7583baa4af55ac87babb5ecfc8989e54e5e182d1d \
|
||||||
|
--hash=sha256:e169081d7ae955f4bd1a590a7ec29f1032eae6889539cf7047bd0f7b09daedc9 \
|
||||||
|
--hash=sha256:e5578ad134fa81286622faff397650cfa2249f640af783b8c2abbae1c70dacdd \
|
||||||
|
--hash=sha256:e67af1dcebc0663cd90253cfb4653f991d0995160ec9ca3132924d7956e17c6e \
|
||||||
|
--hash=sha256:ebe363e5c252dc9011b0380c9b0b8ef559573dcc325ec8f3165129d21af10b63 \
|
||||||
|
--hash=sha256:ec9a66ed2ed23611dcfaa87a860f1511a56ded56f01dd161eeebddb6e25590c3 \
|
||||||
|
--hash=sha256:ed723dc78dd6f676f38083bd86194dbe91befd8c3ecb9cd2f47147bfe7d26dd1 \
|
||||||
|
--hash=sha256:ed865d560365bb3797e4e05dcbd83fb7a045893cc54f0d72588f90eb05c68fee \
|
||||||
|
--hash=sha256:efefb4c85414b6e4be19a53f90d58b573f551b7e4d1dc1e566f7030b6ca4fa8f \
|
||||||
|
--hash=sha256:f078f774d094ea32302163419141fda36176b954069956296406ae1cf4b00222 \
|
||||||
|
--hash=sha256:f2ecb87363dd9e13fa9def0a5c7a61ef5ccc952c08b99672e6f95fdb2463ccd9 \
|
||||||
|
--hash=sha256:f59d36c5356ca6ff79b1a91ef39845c0dd71eeee6b98d71cd0972307eba77260 \
|
||||||
|
--hash=sha256:f696d058d233923b7259d2d963f92b9cf2906063820f27cbd4085529d78861c3 \
|
||||||
|
--hash=sha256:f69c363342b81fce87f2e9dafd05ec041b67ee3b74c08ee9d2be5aeab8d484da \
|
||||||
|
--hash=sha256:f8b784a28492f4020dc90ef6b6d0bb3ca591cb1331de6362968308ed5243b550 \
|
||||||
|
--hash=sha256:f9594423bace86d47d080ae92329315b977fe6466ac998e36a88563c9c6d0259 \
|
||||||
|
--hash=sha256:fb7df717e6c9f2b59aebdf558242da87b2b5cd5961b9469efe8f01762dfe4cc1 \
|
||||||
|
--hash=sha256:ff7cc959f3535028c03c201bbe6703ce1cb5051164f08bca9f814e04333fbb48
|
||||||
|
# via nltk
|
||||||
requests==2.34.2 \
|
requests==2.34.2 \
|
||||||
--hash=sha256:2a0d60c172f83ac6ab31e4554906c0f3b3588d37b5cb939b1c061f4907e278e0 \
|
--hash=sha256:2a0d60c172f83ac6ab31e4554906c0f3b3588d37b5cb939b1c061f4907e278e0 \
|
||||||
--hash=sha256:f288924cae4e29463698d6d60bc6a4da69c89185ad1e0bcc4104f584e960b9ed
|
--hash=sha256:f288924cae4e29463698d6d60bc6a4da69c89185ad1e0bcc4104f584e960b9ed
|
||||||
@@ -837,6 +1104,7 @@ rich==15.0.0 \
|
|||||||
# via
|
# via
|
||||||
# bandit
|
# bandit
|
||||||
# semgrep
|
# semgrep
|
||||||
|
# typer
|
||||||
rpds-py==2026.6.3 \
|
rpds-py==2026.6.3 \
|
||||||
--hash=sha256:0be972be84cfcaf46c8c6edf690ca0f154ac17babf1f6a955a51579b34ad2dc5 \
|
--hash=sha256:0be972be84cfcaf46c8c6edf690ca0f154ac17babf1f6a955a51579b34ad2dc5 \
|
||||||
--hash=sha256:127565fead0a10943b282957bd5447804ff3160ad79f2ad2635e6d249e380680 \
|
--hash=sha256:127565fead0a10943b282957bd5447804ff3160ad79f2ad2635e6d249e380680 \
|
||||||
@@ -960,7 +1228,10 @@ rpds-py==2026.6.3 \
|
|||||||
ruamel-yaml==0.19.1 \
|
ruamel-yaml==0.19.1 \
|
||||||
--hash=sha256:27592957fedf6e0b62f281e96effd28043345e0e66001f97683aa9a40c667c93 \
|
--hash=sha256:27592957fedf6e0b62f281e96effd28043345e0e66001f97683aa9a40c667c93 \
|
||||||
--hash=sha256:53eb66cd27849eff968ebf8f0bf61f46cdac2da1d1f3576dd4ccee9b25c31993
|
--hash=sha256:53eb66cd27849eff968ebf8f0bf61f46cdac2da1d1f3576dd4ccee9b25c31993
|
||||||
# via semgrep
|
# via
|
||||||
|
# safety
|
||||||
|
# safety-schemas
|
||||||
|
# semgrep
|
||||||
ruamel-yaml-clib==0.2.15 \
|
ruamel-yaml-clib==0.2.15 \
|
||||||
--hash=sha256:014181cdec565c8745b7cbc4de3bf2cc8ced05183d986e6d1200168e5bb59490 \
|
--hash=sha256:014181cdec565c8745b7cbc4de3bf2cc8ced05183d986e6d1200168e5bb59490 \
|
||||||
--hash=sha256:04d21dc9c57d9608225da28285900762befbb0165ae48482c15d8d4989d4af14 \
|
--hash=sha256:04d21dc9c57d9608225da28285900762befbb0165ae48482c15d8d4989d4af14 \
|
||||||
@@ -1024,6 +1295,14 @@ ruamel-yaml-clib==0.2.15 \
|
|||||||
--hash=sha256:fd4c928ddf6bce586285daa6d90680b9c291cfd045fc40aad34e445d57b1bf51 \
|
--hash=sha256:fd4c928ddf6bce586285daa6d90680b9c291cfd045fc40aad34e445d57b1bf51 \
|
||||||
--hash=sha256:fe239bdfdae2302e93bd6e8264bd9b71290218fff7084a9db250b55caaccf43f
|
--hash=sha256:fe239bdfdae2302e93bd6e8264bd9b71290218fff7084a9db250b55caaccf43f
|
||||||
# via semgrep
|
# via semgrep
|
||||||
|
safety==3.8.1 \
|
||||||
|
--hash=sha256:953c1c3c60c873f53a6cc250b2a9c4b38bb6ef45f0625990e43f20bff916c965 \
|
||||||
|
--hash=sha256:e646123b976bbb6707cfaacae8c926e2f886b744a60e0f410e8610a3a4eaf7be
|
||||||
|
# via -r .github/requirements/security-scan-tools.in
|
||||||
|
safety-schemas==0.0.16 \
|
||||||
|
--hash=sha256:3bb04d11bd4b5cc79f9fa183c658a6a8cf827a9ceec443a5ffa6eed38a50a24e \
|
||||||
|
--hash=sha256:6760515d3fd1e6535b251cd73014bd431d12fe0bfb8b6e8880a9379b5ab7aa44
|
||||||
|
# via safety
|
||||||
semantic-version==2.10.0 \
|
semantic-version==2.10.0 \
|
||||||
--hash=sha256:bdabb6d336998cbb378d4b9db3a4b56a1e3235701dc05ea2690d9a997ed5041c \
|
--hash=sha256:bdabb6d336998cbb378d4b9db3a4b56a1e3235701dc05ea2690d9a997ed5041c \
|
||||||
--hash=sha256:de78a3b8e0feda74cabc54aab2da702113e33ac9d9eb9d2389bcf1f58b7d9177
|
--hash=sha256:de78a3b8e0feda74cabc54aab2da702113e33ac9d9eb9d2389bcf1f58b7d9177
|
||||||
@@ -1038,6 +1317,10 @@ semgrep==1.175.0 \
|
|||||||
--hash=sha256:e8b14c91558f765b9dd155a99b0071bfe64f61577cda8eb4964132155232c1af \
|
--hash=sha256:e8b14c91558f765b9dd155a99b0071bfe64f61577cda8eb4964132155232c1af \
|
||||||
--hash=sha256:e8ecd7ee8ef1033c9635111c6e162e778834d61416968c5c1c5e7b7fba35c34e
|
--hash=sha256:e8ecd7ee8ef1033c9635111c6e162e778834d61416968c5c1c5e7b7fba35c34e
|
||||||
# via -r .github/requirements/security-scan-tools.in
|
# via -r .github/requirements/security-scan-tools.in
|
||||||
|
shellingham==1.5.4 \
|
||||||
|
--hash=sha256:7ecfff8f2fd72616f7481040475a65b2bf8af90a56c89140852d1120324e8686 \
|
||||||
|
--hash=sha256:8dbca0739d487e5bd35ab3ca4b36e11c4078f3a234bfce294b0a0291363404de
|
||||||
|
# via typer
|
||||||
sse-starlette==3.4.8 \
|
sse-starlette==3.4.8 \
|
||||||
--hash=sha256:6e82314c786709a3cd9520f2285cf9fff90e181e598e8a357b0cf80f66afba0d \
|
--hash=sha256:6e82314c786709a3cd9520f2285cf9fff90e181e598e8a357b0cf80f66afba0d \
|
||||||
--hash=sha256:ed89ffbb75cbf78a5fe2f2109cd584792ee7f9dfac96f791db546df8f15f3f9c
|
--hash=sha256:ed89ffbb75cbf78a5fe2f2109cd584792ee7f9dfac96f791db546df8f15f3f9c
|
||||||
@@ -1052,6 +1335,10 @@ stevedore==5.9.1 \
|
|||||||
--hash=sha256:5c8ff3a9f336cc1a06ac0f597bc79d11a2f950bfd32e290ca56b5a301fafafbf \
|
--hash=sha256:5c8ff3a9f336cc1a06ac0f597bc79d11a2f950bfd32e290ca56b5a301fafafbf \
|
||||||
--hash=sha256:e97a2667923efda926e8713fde6a73616df68210a3cbc6f02b48967b676fd8bf
|
--hash=sha256:e97a2667923efda926e8713fde6a73616df68210a3cbc6f02b48967b676fd8bf
|
||||||
# via bandit
|
# via bandit
|
||||||
|
tenacity==9.1.4 \
|
||||||
|
--hash=sha256:6095a360c919085f28c6527de529e76a06ad89b23659fa881ae0649b867a9d55 \
|
||||||
|
--hash=sha256:adb31d4c263f2bd041081ab33b498309a57c77f9acf2db65aadf0898179cf93a
|
||||||
|
# via safety
|
||||||
tomli==2.4.1 \
|
tomli==2.4.1 \
|
||||||
--hash=sha256:01f520d4f53ef97964a240a035ec2a869fe1a37dde002b57ebc4417a27ccd853 \
|
--hash=sha256:01f520d4f53ef97964a240a035ec2a869fe1a37dde002b57ebc4417a27ccd853 \
|
||||||
--hash=sha256:0d85819802132122da43cb86656f8d1f8c6587d54ae7dcaf30e90533028b49fe \
|
--hash=sha256:0d85819802132122da43cb86656f8d1f8c6587d54ae7dcaf30e90533028b49fe \
|
||||||
@@ -1101,6 +1388,22 @@ tomli==2.4.1 \
|
|||||||
--hash=sha256:ff18e6a727ee0ab0388507b89d1bc6a22b138d1e2fa56d1ad494586d61d2eae9 \
|
--hash=sha256:ff18e6a727ee0ab0388507b89d1bc6a22b138d1e2fa56d1ad494586d61d2eae9 \
|
||||||
--hash=sha256:ff2983983d34813c1aeb0fa89091e76c3a22889ee83ab27c5eeb45100560c049
|
--hash=sha256:ff2983983d34813c1aeb0fa89091e76c3a22889ee83ab27c5eeb45100560c049
|
||||||
# via semgrep
|
# via semgrep
|
||||||
|
tomlkit==0.15.1 \
|
||||||
|
--hash=sha256:177a05aece5a8ca5266fd3c448abb47b8d352f09d477d3ca8332db4d89b24304 \
|
||||||
|
--hash=sha256:e25bbf38843005246210a12982776f27f99cb9be67160e14434d0c0d21ee1e97
|
||||||
|
# via safety
|
||||||
|
tqdm==4.70.0 \
|
||||||
|
--hash=sha256:55b0b0dbd97462d06ebee91e4dac24ed4d4702be82b24f07e6c1d27e08cea220 \
|
||||||
|
--hash=sha256:7f585706bfddbdebf89daac705b2dfcc16890130727d3197ca62c732b4310953
|
||||||
|
# via nltk
|
||||||
|
truststore==0.10.4 \
|
||||||
|
--hash=sha256:9d91bd436463ad5e4ee4aba766628dd6cd7010cf3e2461756b3303710eebc301 \
|
||||||
|
--hash=sha256:adaeaecf1cbb5f4de3b1959b42d41f6fab57b2b1666adb59e89cb0b53361d981
|
||||||
|
# via safety
|
||||||
|
typer==0.25.1 \
|
||||||
|
--hash=sha256:75caa44ed46a03fb2dab8808753ffacdbfea88495e74c85a28c5eefcf5f39c89 \
|
||||||
|
--hash=sha256:9616eb8853a09ffeabab1698952f33c6f29ffdbceb4eaeecf571880e8d7664cc
|
||||||
|
# via safety
|
||||||
typing-extensions==4.16.0 \
|
typing-extensions==4.16.0 \
|
||||||
--hash=sha256:481caa481374e813c1b176ada14e97f1f67a4539ce9cfeb3f350d78d6370c2e8 \
|
--hash=sha256:481caa481374e813c1b176ada14e97f1f67a4539ce9cfeb3f350d78d6370c2e8 \
|
||||||
--hash=sha256:dc983d19a509c94dba722ee6abd33940f7c05a89e243c47e907eb4db6f1a43e5
|
--hash=sha256:dc983d19a509c94dba722ee6abd33940f7c05a89e243c47e907eb4db6f1a43e5
|
||||||
@@ -1114,6 +1417,8 @@ typing-extensions==4.16.0 \
|
|||||||
# pydantic
|
# pydantic
|
||||||
# pydantic-core
|
# pydantic-core
|
||||||
# referencing
|
# referencing
|
||||||
|
# safety
|
||||||
|
# safety-schemas
|
||||||
# semgrep
|
# semgrep
|
||||||
# starlette
|
# starlette
|
||||||
# typing-inspection
|
# typing-inspection
|
||||||
|
|||||||
@@ -1,76 +0,0 @@
|
|||||||
"""Drop checkov-suppressed results from its SARIF output before upload.
|
|
||||||
|
|
||||||
checkov's SARIF exporter includes every evaluated check as an ordinary
|
|
||||||
result, including ones it internally marked SKIPPED via an inline
|
|
||||||
`# checkov:skip=` comment or a `checkov.io/skipN` resource annotation - it
|
|
||||||
never uses SARIF's `suppressions` field, and never drops them. checkov's
|
|
||||||
JSON output *does* correctly record which checks were skipped, so this
|
|
||||||
cross-references the two: any SARIF result whose (check_id, file) pair
|
|
||||||
appears in the JSON's skipped_checks is removed before GitHub ever sees it.
|
|
||||||
|
|
||||||
Without this, every already-suppressed finding reopens as a brand new code
|
|
||||||
scanning alert on every run, forever (see #6035/#6036, #6112-6115,
|
|
||||||
#6128-6131 for the pattern this was chasing before this script existed).
|
|
||||||
|
|
||||||
Usage: filter_checkov_skipped.py <json_path> <sarif_in_path> <sarif_out_path>
|
|
||||||
"""
|
|
||||||
|
|
||||||
import json
|
|
||||||
import sys
|
|
||||||
|
|
||||||
|
|
||||||
def path_suffix(path: str, segments: int = 2) -> str:
|
|
||||||
"""Last N path segments, normalized to forward slashes, lowercased.
|
|
||||||
|
|
||||||
checkov's JSON file_path and SARIF artifactLocation.uri are relative to
|
|
||||||
different roots (the scanned directory vs. a temp helm-render dir), so
|
|
||||||
they can't be compared directly - but the last couple of segments
|
|
||||||
(e.g. "templates/service.yaml") are stable across both and specific
|
|
||||||
enough in practice to avoid cross-file collisions.
|
|
||||||
"""
|
|
||||||
normalized = path.replace("\\", "/").strip("/")
|
|
||||||
return "/".join(normalized.split("/")[-segments:]).lower()
|
|
||||||
|
|
||||||
|
|
||||||
def main() -> None:
|
|
||||||
json_path, sarif_in_path, sarif_out_path = sys.argv[1:4]
|
|
||||||
|
|
||||||
with open(json_path, encoding="utf-8") as f:
|
|
||||||
checkov_json = json.load(f)
|
|
||||||
if isinstance(checkov_json, dict):
|
|
||||||
checkov_json = [checkov_json]
|
|
||||||
|
|
||||||
skipped = set()
|
|
||||||
for block in checkov_json:
|
|
||||||
for check in block.get("results", {}).get("skipped_checks", []):
|
|
||||||
skipped.add((check["check_id"], path_suffix(check["file_path"])))
|
|
||||||
|
|
||||||
with open(sarif_in_path, encoding="utf-8") as f:
|
|
||||||
sarif = json.load(f)
|
|
||||||
|
|
||||||
removed = 0
|
|
||||||
for run in sarif.get("runs", []):
|
|
||||||
kept = []
|
|
||||||
for result in run.get("results", []):
|
|
||||||
rule_id = result.get("ruleId")
|
|
||||||
locations = result.get("locations") or [{}]
|
|
||||||
uri = (
|
|
||||||
locations[0]
|
|
||||||
.get("physicalLocation", {})
|
|
||||||
.get("artifactLocation", {})
|
|
||||||
.get("uri", "")
|
|
||||||
)
|
|
||||||
if (rule_id, path_suffix(uri)) in skipped:
|
|
||||||
removed += 1
|
|
||||||
continue
|
|
||||||
kept.append(result)
|
|
||||||
run["results"] = kept
|
|
||||||
|
|
||||||
with open(sarif_out_path, "w", encoding="utf-8") as f:
|
|
||||||
json.dump(sarif, f)
|
|
||||||
|
|
||||||
print(f"Removed {removed} checkov-suppressed result(s) from the SARIF before upload.")
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
|
||||||
main()
|
|
||||||
@@ -44,20 +44,12 @@ jobs:
|
|||||||
pip install -r .github/requirements/pep517-build.txt --require-hashes
|
pip install -r .github/requirements/pep517-build.txt --require-hashes
|
||||||
pip install --no-deps --no-build-isolation -e .
|
pip install --no-deps --no-build-isolation -e .
|
||||||
pip install -r .github/requirements/base-deps.txt --require-hashes
|
pip install -r .github/requirements/base-deps.txt --require-hashes
|
||||||
# NOTE: benchmarks/ does not currently exist in this repo (neither
|
# NOTE: benchmarks/ does not currently exist in this repo, so this
|
||||||
# requirements.txt nor benchmarks_runner.py below), so this job
|
# step and the run below it fail on any real invocation - pre-existing,
|
||||||
# already fails on any real invocation - pre-existing, unrelated to
|
# unrelated to this pinning change. Left as-is since there's nothing
|
||||||
# this pinning change. The `pip install -r benchmarks/requirements.txt`
|
# to hash without knowing what belongs there.
|
||||||
# step that used to be here is dropped rather than fixed: there's
|
pip install -r benchmarks/requirements.txt
|
||||||
# nothing to hash-pin without knowing what that file should
|
python -m spacy download en_core_web_sm
|
||||||
# contain, and an unpinned install here would just re-trip
|
|
||||||
# Scorecard's Pinned-Dependencies check for no real benefit, since
|
|
||||||
# the job can't run to completion regardless.
|
|
||||||
#
|
|
||||||
# `python -m spacy download en_core_web_sm` fetches an unpinned,
|
|
||||||
# unhashed wheel from spacy-models' GitHub releases - replaced with
|
|
||||||
# a hash-pinned direct-URL install of the same 3.8.0 model (matches
|
|
||||||
# the spacy==3.8.15 pinned in base-deps.txt) via benchmark-extra.txt.
|
|
||||||
pip install -r .github/requirements/benchmark-extra.txt --require-hashes
|
pip install -r .github/requirements/benchmark-extra.txt --require-hashes
|
||||||
|
|
||||||
- name: Execute Benchmarks (Real Mode)
|
- name: Execute Benchmarks (Real Mode)
|
||||||
|
|||||||
@@ -12,63 +12,13 @@ on:
|
|||||||
- '**/*.md'
|
- '**/*.md'
|
||||||
pull_request:
|
pull_request:
|
||||||
branches: [main]
|
branches: [main]
|
||||||
|
paths-ignore:
|
||||||
|
- 'docs/**'
|
||||||
|
- 'docs_check.py'
|
||||||
|
- '**/*.md'
|
||||||
|
|
||||||
jobs:
|
jobs:
|
||||||
# Detect whether this PR touches any source files (non-docs/non-markdown).
|
|
||||||
# The result drives the `build` job's `if:` condition so that:
|
|
||||||
# - docs-only PRs: `build` is skipped (satisfies the required check).
|
|
||||||
# - code PRs: `build` runs exactly as before.
|
|
||||||
# Push events (to main) keep their own paths-ignore above and never reach
|
|
||||||
# this job, so the push optimization is unaffected.
|
|
||||||
changes:
|
|
||||||
runs-on: ubuntu-latest
|
|
||||||
# Only needed for pull_request events; push events are pre-filtered above.
|
|
||||||
if: github.event_name == 'pull_request'
|
|
||||||
outputs:
|
|
||||||
src: ${{ steps.filter.outputs.src }}
|
|
||||||
steps:
|
|
||||||
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
|
||||||
with:
|
|
||||||
# Fetch enough history to compute the merge base against the PR base.
|
|
||||||
fetch-depth: 0
|
|
||||||
- name: Check for source changes
|
|
||||||
id: filter
|
|
||||||
run: |
|
|
||||||
# List files changed in this PR relative to the true merge base.
|
|
||||||
# Using three-dot merge-base diff so changes on the base branch that
|
|
||||||
# are not part of this PR do not appear in the file list.
|
|
||||||
# If every changed file matches docs/** or *.md (any depth) or
|
|
||||||
# docs_check.py, this is a docs-only PR and src=false; otherwise
|
|
||||||
# src=true.
|
|
||||||
BASE="${{ github.event.pull_request.base.sha }}"
|
|
||||||
HEAD="${{ github.event.pull_request.head.sha }}"
|
|
||||||
MERGE_BASE=$(git merge-base "$BASE" "$HEAD")
|
|
||||||
CHANGED=$(git diff --name-only "$MERGE_BASE" "$HEAD")
|
|
||||||
echo "Changed files:"
|
|
||||||
echo "$CHANGED"
|
|
||||||
NON_DOCS=$(echo "$CHANGED" | grep -Ev '^(docs/|docs_check\.py|.*\.md$)' || true)
|
|
||||||
if [ -n "$NON_DOCS" ]; then
|
|
||||||
echo "src=true" >> "$GITHUB_OUTPUT"
|
|
||||||
else
|
|
||||||
echo "src=false" >> "$GITHUB_OUTPUT"
|
|
||||||
fi
|
|
||||||
|
|
||||||
build:
|
build:
|
||||||
needs: [changes]
|
|
||||||
# For pull_request events:
|
|
||||||
# - skip only when changes ran successfully and explicitly set src=false
|
|
||||||
# (i.e. a confirmed docs-only PR).
|
|
||||||
# - run when changes succeeded with src=true (source changes present).
|
|
||||||
# - run when changes failed or was cancelled (fail-closed: missing output
|
|
||||||
# must not silently skip the build).
|
|
||||||
# For push/non-PR events: changes is skipped; always() prevents the build
|
|
||||||
# from being skipped due to a skipped needs dependency.
|
|
||||||
if: >-
|
|
||||||
always() && (
|
|
||||||
github.event_name != 'pull_request' ||
|
|
||||||
needs.changes.result != 'success' ||
|
|
||||||
needs.changes.outputs.src == 'true'
|
|
||||||
)
|
|
||||||
runs-on: ubuntu-latest
|
runs-on: ubuntu-latest
|
||||||
steps:
|
steps:
|
||||||
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
||||||
|
|||||||
@@ -76,28 +76,12 @@ jobs:
|
|||||||
PYTHONUTF8: "1"
|
PYTHONUTF8: "1"
|
||||||
run: |
|
run: |
|
||||||
New-Item -ItemType Directory -Force reports | Out-Null
|
New-Item -ItemType Directory -Force reports | Out-Null
|
||||||
checkov --directory . --framework kubernetes helm dockerfile github_actions secrets bicep arm --soft-fail --output sarif --output json --output-file-path reports
|
checkov --directory . --framework kubernetes helm dockerfile github_actions secrets bicep arm --soft-fail --output sarif --output-file-path reports/checkov.sarif
|
||||||
if (-not (Test-Path reports/results_sarif.sarif)) {
|
if (-not (Test-Path reports/checkov.sarif)) {
|
||||||
$sarif = Get-ChildItem -Path reports -Recurse -Filter *.sarif | Select-Object -First 1
|
$sarif = Get-ChildItem -Path reports -Recurse -Filter *.sarif | Select-Object -First 1
|
||||||
if ($null -eq $sarif) { throw "Checkov did not produce a SARIF file" }
|
if ($null -eq $sarif) { throw "Checkov did not produce a SARIF file" }
|
||||||
Copy-Item $sarif.FullName reports/results_sarif.sarif
|
Copy-Item $sarif.FullName reports/checkov.sarif
|
||||||
}
|
}
|
||||||
if (-not (Test-Path reports/results_json.json)) {
|
|
||||||
$json = Get-ChildItem -Path reports -Recurse -Filter *.json | Select-Object -First 1
|
|
||||||
if ($null -eq $json) { throw "Checkov did not produce a JSON file" }
|
|
||||||
Copy-Item $json.FullName reports/results_json.json
|
|
||||||
}
|
|
||||||
|
|
||||||
# checkov's SARIF exporter includes checks it internally marked SKIPPED
|
|
||||||
# (via the inline `# checkov:skip=` comments / `checkov.io/skipN`
|
|
||||||
# annotations already on the Helm chart) as ordinary un-suppressed
|
|
||||||
# results - it never uses SARIF's own `suppressions` field, so GitHub
|
|
||||||
# opens a fresh alert for the same already-suppressed finding on every
|
|
||||||
# single run (see #6035/#6036, #6112-6115, #6128-6131). checkov's JSON
|
|
||||||
# output does correctly record the skip, so cross-reference it here
|
|
||||||
# instead of re-dismissing the same alerts by hand forever.
|
|
||||||
- name: Filter checkov's own suppressed checks out of the SARIF
|
|
||||||
run: python .github/scripts/filter_checkov_skipped.py reports/results_json.json reports/results_sarif.sarif reports/checkov.sarif
|
|
||||||
|
|
||||||
- name: Upload Checkov results to Security tab
|
- name: Upload Checkov results to Security tab
|
||||||
uses: github/codeql-action/upload-sarif@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4
|
uses: github/codeql-action/upload-sarif@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4
|
||||||
|
|||||||
@@ -65,4 +65,4 @@ jobs:
|
|||||||
|
|
||||||
- name: Deploy to GitHub Pages
|
- name: Deploy to GitHub Pages
|
||||||
id: deployment
|
id: deployment
|
||||||
uses: actions/deploy-pages@368f82528645a54fb793d4d04e342629a3f51346 # v5
|
uses: actions/deploy-pages@cd2ce8fcbc39b97be8ca5fce6e763baed58fa128 # v5
|
||||||
|
|||||||
@@ -3,7 +3,6 @@ name: Security Scan
|
|||||||
on:
|
on:
|
||||||
schedule:
|
schedule:
|
||||||
- cron: '30 1 * * 1,4' # Mon/Thu 7 AM IST
|
- cron: '30 1 * * 1,4' # Mon/Thu 7 AM IST
|
||||||
workflow_dispatch:
|
|
||||||
push:
|
push:
|
||||||
branches: [main]
|
branches: [main]
|
||||||
paths-ignore:
|
paths-ignore:
|
||||||
@@ -13,65 +12,17 @@ on:
|
|||||||
- '**/*.md'
|
- '**/*.md'
|
||||||
pull_request:
|
pull_request:
|
||||||
branches: [main]
|
branches: [main]
|
||||||
|
paths-ignore:
|
||||||
|
- 'docs/**'
|
||||||
|
- 'mkdocs.yml'
|
||||||
|
- 'requirements-docs.txt'
|
||||||
|
- '**/*.md'
|
||||||
|
|
||||||
permissions:
|
permissions:
|
||||||
contents: read
|
contents: read
|
||||||
|
|
||||||
jobs:
|
jobs:
|
||||||
# Detect whether this PR touches any source files (non-docs/non-markdown).
|
|
||||||
# The result drives the `security-scan` job's `if:` condition so that:
|
|
||||||
# - docs-only PRs: `security-scan` is skipped (satisfies the required check).
|
|
||||||
# - code PRs: the full scan runs exactly as before.
|
|
||||||
# Schedule and workflow_dispatch runs always skip this job and run the scan
|
|
||||||
# unconditionally (the security-scan job's if: accounts for that below).
|
|
||||||
# Push events (to main) keep their own paths-ignore above.
|
|
||||||
changes:
|
|
||||||
runs-on: ubuntu-latest
|
|
||||||
if: github.event_name == 'pull_request'
|
|
||||||
outputs:
|
|
||||||
src: ${{ steps.filter.outputs.src }}
|
|
||||||
steps:
|
|
||||||
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
|
||||||
with:
|
|
||||||
fetch-depth: 0
|
|
||||||
- name: Check for source changes
|
|
||||||
id: filter
|
|
||||||
run: |
|
|
||||||
# List files changed in this PR relative to the true merge base.
|
|
||||||
# Using three-dot merge-base diff so changes on the base branch that
|
|
||||||
# are not part of this PR do not appear in the file list.
|
|
||||||
# If every changed file matches the docs/markdown paths-ignore list
|
|
||||||
# (at any directory depth), this is a docs-only PR and src=false;
|
|
||||||
# otherwise src=true.
|
|
||||||
BASE="${{ github.event.pull_request.base.sha }}"
|
|
||||||
HEAD="${{ github.event.pull_request.head.sha }}"
|
|
||||||
MERGE_BASE=$(git merge-base "$BASE" "$HEAD")
|
|
||||||
CHANGED=$(git diff --name-only "$MERGE_BASE" "$HEAD")
|
|
||||||
echo "Changed files:"
|
|
||||||
echo "$CHANGED"
|
|
||||||
NON_DOCS=$(echo "$CHANGED" | grep -Ev '^(docs/|mkdocs\.yml$|requirements-docs\.txt$|.*\.md$)' || true)
|
|
||||||
if [ -n "$NON_DOCS" ]; then
|
|
||||||
echo "src=true" >> "$GITHUB_OUTPUT"
|
|
||||||
else
|
|
||||||
echo "src=false" >> "$GITHUB_OUTPUT"
|
|
||||||
fi
|
|
||||||
|
|
||||||
security-scan:
|
security-scan:
|
||||||
# For pull_request events:
|
|
||||||
# - skip only when changes ran successfully and explicitly set src=false
|
|
||||||
# (i.e. a confirmed docs-only PR).
|
|
||||||
# - run when changes succeeded with src=true (source changes present).
|
|
||||||
# - run when changes failed or was cancelled (fail-closed: missing output
|
|
||||||
# must not silently skip the security scan).
|
|
||||||
# For schedule/workflow_dispatch/push: changes is skipped; always() ensures
|
|
||||||
# the scan still runs unconditionally for those triggers.
|
|
||||||
needs: [changes]
|
|
||||||
if: >-
|
|
||||||
always() && (
|
|
||||||
github.event_name != 'pull_request' ||
|
|
||||||
needs.changes.result != 'success' ||
|
|
||||||
needs.changes.outputs.src == 'true'
|
|
||||||
)
|
|
||||||
runs-on: ubuntu-latest
|
runs-on: ubuntu-latest
|
||||||
permissions:
|
permissions:
|
||||||
contents: read
|
contents: read
|
||||||
@@ -94,100 +45,45 @@ jobs:
|
|||||||
- name: Install dependencies
|
- name: Install dependencies
|
||||||
run: |
|
run: |
|
||||||
pip install -r .github/requirements/bootstrap.txt --require-hashes
|
pip install -r .github/requirements/bootstrap.txt --require-hashes
|
||||||
# Install the pinned dependency set FIRST so pip-audit scans
|
# Install the pinned dependency set FIRST so Safety scans Semantica's
|
||||||
# Semantica's exact CI/release dependency tree (requirements-ci.txt
|
# exact CI/release dependency tree (requirements-ci.txt is generated
|
||||||
# is generated from pyproject.toml extras, so this covers the
|
# from pyproject.toml extras, so this covers the project's real deps).
|
||||||
# project's real deps).
|
|
||||||
pip install -r requirements-ci.txt --require-hashes
|
pip install -r requirements-ci.txt --require-hashes
|
||||||
# Tooling AFTER the pinned set: installing it first would let the
|
# Tooling AFTER the pinned set: installing safety/bandit/semgrep/jq
|
||||||
# pinned requirements overwrite the tooling's own transitive deps.
|
# first lets the pinned requirements overwrite their transitive deps
|
||||||
pip install -r .github/requirements/pip-audit.txt --require-hashes
|
# (e.g. rich), which breaks the safety CLI at runtime.
|
||||||
pip install -r .github/requirements/security-scan-tools.txt --require-hashes
|
pip install -r .github/requirements/security-scan-tools.txt --require-hashes
|
||||||
|
|
||||||
- name: Run pip-audit (Package Vulnerabilities)
|
- name: Run Safety Check (Package Vulnerabilities)
|
||||||
continue-on-error: true
|
|
||||||
run: |
|
run: |
|
||||||
# Keep publishing reports and the PR comment even when the audit
|
# NOTE: Safety 3.x repurposed --output to select a console format
|
||||||
# gate fails. The final gate below preserves the failure status.
|
# (json/text/screen/...), not a file path. Writing JSON to a file
|
||||||
echo 'AUDIT_SCAN_STATUS=failed' >> "$GITHUB_ENV"
|
# now requires --save-json; the previous `--output safety-report.json`
|
||||||
|
# usage was silently invalid and never produced a report.
|
||||||
|
safety check --save-json safety-report.json || true
|
||||||
|
|
||||||
# Same dependency tree Safety used to scan, and the same tool and
|
# Guard 1: fail loudly if Safety exited before writing a report at all
|
||||||
# invocation already proven reliable in security.yml.
|
# (network error, API auth failure, tool crash). Without this check a
|
||||||
pip-audit -r requirements-ci.txt --format=json --output=pip-audit-report.json || true
|
# missing or empty file causes jq to fall back to "0", making a broken
|
||||||
|
|
||||||
# Guard 1: fail loudly if pip-audit exited before writing a report
|
|
||||||
# at all (network error, tool crash). Without this check a missing
|
|
||||||
# or empty file causes jq to fall back to "0", making a broken
|
|
||||||
# scanner indistinguishable from a clean scan.
|
# scanner indistinguishable from a clean scan.
|
||||||
if [ ! -s pip-audit-report.json ]; then
|
if [ ! -s safety-report.json ]; then
|
||||||
echo "::error::pip-audit produced no report (pip-audit-report.json is missing or empty). Treating as failure — check for network errors or pip-audit crashes in the logs above."
|
echo "::error::Safety scan produced no report (safety-report.json is missing or empty). Treating as failure — check for network errors, API auth failures, or Safety crashes in the logs above."
|
||||||
exit 1
|
|
||||||
fi
|
|
||||||
|
|
||||||
# Guard 2: fail closed when the report doesn't have the shape the
|
|
||||||
# checks below assume: a non-empty dependencies array, each entry
|
|
||||||
# either carrying an array-valued vulns field or being a dependency
|
|
||||||
# pip-audit couldn't resolve/audit, which it reports as
|
|
||||||
# {"name": ..., "skip_reason": ...} with no vulns field at all
|
|
||||||
# (see pip_audit._format.json.JsonFormat._format_dep). That's a
|
|
||||||
# normal, documented report shape, not a malformed one — treating
|
|
||||||
# it as invalid would fail the whole job over a single unauditable
|
|
||||||
# package, the same kind of scan-unrelated CI break this migration
|
|
||||||
# away from Safety was meant to fix.
|
|
||||||
if ! jq -e '
|
|
||||||
(.dependencies | type == "array" and length > 0)
|
|
||||||
and all(.dependencies[]; type == "object" and ((.vulns | type == "array") or (.skip_reason | type == "string")))
|
|
||||||
' pip-audit-report.json >/dev/null 2>&1; then
|
|
||||||
echo "::error::pip-audit report has an invalid dependency structure. Expected a non-empty dependencies array where every entry has either a vulns array or a skip_reason. Treating as failure."
|
|
||||||
exit 1
|
exit 1
|
||||||
fi
|
fi
|
||||||
|
|
||||||
echo "Checking for package vulnerabilities..."
|
echo "Checking for package vulnerabilities..."
|
||||||
|
|
||||||
# Guard 2 above already confirmed pip-audit-report.json is valid
|
# No || echo "0" fallback: if jq fails (malformed JSON, missing key,
|
||||||
# JSON with a well-shaped dependencies array, so this count is
|
# vulnerabilities:null) VULNS will be empty or "null" so guard 2 below
|
||||||
# always a plain non-negative integer.
|
# catches it rather than silently treating the broken report as zero.
|
||||||
SKIPPED=$(jq '[.dependencies[] | select(has("skip_reason"))] | length' pip-audit-report.json)
|
VULNS=$(jq '.vulnerabilities | length' safety-report.json 2>/dev/null)
|
||||||
if [ "$SKIPPED" -gt 0 ]; then
|
|
||||||
echo "⚠️ pip-audit could not audit $SKIPPED dependencies (see pip-audit-report.json for skip_reason):"
|
|
||||||
jq -r '.dependencies[] | select(has("skip_reason")) | " - \(.name): \(.skip_reason)"' pip-audit-report.json
|
|
||||||
fi
|
|
||||||
|
|
||||||
# Vulnerability IDs reviewed and accepted as non-actionable for this
|
# Guard 2: ensure VULNS is a non-negative integer before the -gt
|
||||||
# project. Empty for now: pip-audit's OSV-backed database doesn't
|
|
||||||
# currently carry either of the findings Safety used to flag here
|
|
||||||
# (cuda-toolkit CVE-2025-33228, torchvision CVE-2026-65918), so
|
|
||||||
# there's nothing to exclude. Left in place so a future finding can
|
|
||||||
# be added the same way without restructuring this step - see git
|
|
||||||
# history on this file for the reasoning behind past entries.
|
|
||||||
IGNORED_VULN_IDS=""
|
|
||||||
|
|
||||||
# Exported so the "Comment PR with Security Results" step below can
|
|
||||||
# apply the same exclusion list to the raw report - it reads
|
|
||||||
# pip-audit-report.json independently in JS, so without this the PR
|
|
||||||
# comment would show an accepted finding as live even though this
|
|
||||||
# gate correctly treats it as non-actionable.
|
|
||||||
echo "IGNORED_VULN_IDS=$IGNORED_VULN_IDS" >> "$GITHUB_ENV"
|
|
||||||
|
|
||||||
# No []? / || echo "0" fallback: if jq fails (malformed JSON) VULNS
|
|
||||||
# will be empty or "null" so Guard 3 below catches it rather than
|
|
||||||
# silently treating the broken report as zero.
|
|
||||||
# `.vulns // []` guards against skipped dependencies, which carry
|
|
||||||
# no vulns field at all (see the skip_reason handling above) -
|
|
||||||
# without the fallback, iterating `null[]` raises inside jq and
|
|
||||||
# this whole computation silently evaluates to empty.
|
|
||||||
VULNS=$(jq --arg ignored "$IGNORED_VULN_IDS" '
|
|
||||||
($ignored | split(",") | map(select(length > 0))) as $ignore_list
|
|
||||||
| [.dependencies[] | (.vulns // [])[] | select(.id as $id | ($ignore_list | index($id)) | not)]
|
|
||||||
| length
|
|
||||||
' pip-audit-report.json 2>/dev/null)
|
|
||||||
|
|
||||||
# Guard 3: ensure VULNS is a non-negative integer before the -gt
|
|
||||||
# comparison. "null" (missing/null key) or "" (jq parse failure) would
|
# comparison. "null" (missing/null key) or "" (jq parse failure) would
|
||||||
# cause bash's -gt to throw an arithmetic error and fall through to the
|
# cause bash's -gt to throw an arithmetic error and fall through to the
|
||||||
# success branch — the same silent-pass bug as a missing file.
|
# success branch — the same silent-pass bug as a missing file.
|
||||||
if ! [[ "$VULNS" =~ ^[0-9]+$ ]]; then
|
if ! [[ "$VULNS" =~ ^[0-9]+$ ]]; then
|
||||||
echo "::error::pip-audit report exists but dependency vulnerabilities are missing or non-numeric (got: '${VULNS}'). The report may be malformed or contain an error-only JSON response. Treating as failure."
|
echo "::error::Safety report exists but 'vulnerabilities' is missing or non-numeric (got: '${VULNS}'). The report may be malformed or Safety may have written an error-only JSON. Treating as failure."
|
||||||
exit 1
|
exit 1
|
||||||
fi
|
fi
|
||||||
|
|
||||||
@@ -196,18 +92,12 @@ jobs:
|
|||||||
echo "CI will fail to prevent merging of vulnerable dependencies"
|
echo "CI will fail to prevent merging of vulnerable dependencies"
|
||||||
echo ""
|
echo ""
|
||||||
echo "Vulnerability details:"
|
echo "Vulnerability details:"
|
||||||
jq --arg ignored "$IGNORED_VULN_IDS" -r '
|
jq -r '.vulnerabilities[] | "- \(.package_name)==\(.analyzed_version): \(.vulnerability_id) (\(.CVE // "no CVE assigned"))"' safety-report.json || true
|
||||||
($ignored | split(",") | map(select(length > 0))) as $ignore_list
|
|
||||||
| .dependencies[] as $dependency
|
|
||||||
| ($dependency.vulns // [])[] | select(.id as $id | ($ignore_list | index($id)) | not)
|
|
||||||
| "- \($dependency.name)==\($dependency.version): \(.id)"
|
|
||||||
' pip-audit-report.json || true
|
|
||||||
exit 1
|
exit 1
|
||||||
else
|
else
|
||||||
echo "✅ No actionable security vulnerabilities found${IGNORED_VULN_IDS:+ (ignored: $IGNORED_VULN_IDS)}"
|
echo "✅ No security vulnerabilities found"
|
||||||
echo 'AUDIT_SCAN_STATUS=passed' >> "$GITHUB_ENV"
|
|
||||||
fi
|
fi
|
||||||
|
|
||||||
- name: Run Bandit (Code Security Linter)
|
- name: Run Bandit (Code Security Linter)
|
||||||
run: |
|
run: |
|
||||||
bandit -r semantica/ -f json -o bandit-report.json || true
|
bandit -r semantica/ -f json -o bandit-report.json || true
|
||||||
@@ -245,18 +135,17 @@ jobs:
|
|||||||
fi
|
fi
|
||||||
|
|
||||||
- name: Upload Security Reports
|
- name: Upload Security Reports
|
||||||
if: always()
|
|
||||||
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
|
||||||
with:
|
with:
|
||||||
name: security-reports
|
name: security-reports
|
||||||
retention-days: 14
|
retention-days: 14
|
||||||
path: |
|
path: |
|
||||||
pip-audit-report.json
|
safety-report.json
|
||||||
bandit-report.json
|
bandit-report.json
|
||||||
semgrep-report.json
|
semgrep-report.json
|
||||||
|
|
||||||
- name: Comment PR with Security Results
|
- name: Comment PR with Security Results
|
||||||
if: always() && github.event_name == 'pull_request'
|
if: github.event_name == 'pull_request'
|
||||||
uses: actions/github-script@3a2844b7e9c422d3c10d287c895573f7108da1b3 # v9
|
uses: actions/github-script@3a2844b7e9c422d3c10d287c895573f7108da1b3 # v9
|
||||||
with:
|
with:
|
||||||
script: |
|
script: |
|
||||||
@@ -279,12 +168,6 @@ jobs:
|
|||||||
}
|
}
|
||||||
|
|
||||||
const items = parse(data);
|
const items = parse(data);
|
||||||
if (items === null) {
|
|
||||||
return [
|
|
||||||
'### ' + title,
|
|
||||||
'⚠️ Invalid report structure in ' + reportPath + ' — check the job logs.',
|
|
||||||
].join('\n');
|
|
||||||
}
|
|
||||||
if (items.length === 0) {
|
if (items.length === 0) {
|
||||||
return [`### ${title}`, `✅ No findings.`].join('\n');
|
return [`### ${title}`, `✅ No findings.`].join('\n');
|
||||||
}
|
}
|
||||||
@@ -301,64 +184,14 @@ jobs:
|
|||||||
return lines.join('\n');
|
return lines.join('\n');
|
||||||
}
|
}
|
||||||
|
|
||||||
// Mirrors the shell step's own IGNORED_VULN_IDS (passed through
|
const safetySection = renderSection(
|
||||||
// $GITHUB_ENV) so an accepted, non-actionable CVE that the CI
|
'Safety — dependency vulnerabilities',
|
||||||
// gate already excluded doesn't reappear here as a live finding -
|
'safety-report.json',
|
||||||
// this reads the same raw, unfiltered pip-audit-report.json.
|
(data) => (data.vulnerabilities || []).map(
|
||||||
const ignoredVulnIds = (process.env.IGNORED_VULN_IDS || '')
|
(v) => `- \`${v.package_name}==${v.analyzed_version}\`: ${v.vulnerability_id}` +
|
||||||
.split(',')
|
(v.CVE ? ` (${v.CVE})` : '') + ` — ${v.advisory || 'no advisory text'}`
|
||||||
.map((id) => id.trim())
|
)
|
||||||
.filter(Boolean);
|
);
|
||||||
|
|
||||||
// A dependency pip-audit couldn't resolve/audit is reported as
|
|
||||||
// {"name": ..., "skip_reason": ...} with no vulns field at all
|
|
||||||
// (see pip_audit._format.json.JsonFormat._format_dep) - that's a
|
|
||||||
// normal report shape, not a malformed one, so it must not be
|
|
||||||
// treated as an invalid dependency below.
|
|
||||||
const isSkipped = (dependency) => typeof dependency.skip_reason === 'string';
|
|
||||||
|
|
||||||
let skippedDeps = [];
|
|
||||||
try {
|
|
||||||
const auditData = JSON.parse(fs.readFileSync('pip-audit-report.json', 'utf8'));
|
|
||||||
skippedDeps = (auditData.dependencies || []).filter(
|
|
||||||
(dependency) => dependency && typeof dependency === 'object' && isSkipped(dependency)
|
|
||||||
);
|
|
||||||
} catch (e) {
|
|
||||||
// Unreadable/unparseable report - renderSection's own
|
|
||||||
// report-missing branch below surfaces this.
|
|
||||||
}
|
|
||||||
|
|
||||||
const pipAuditSection = renderSection(
|
|
||||||
'pip-audit — dependency vulnerabilities',
|
|
||||||
'pip-audit-report.json',
|
|
||||||
(data) => {
|
|
||||||
if (
|
|
||||||
!Array.isArray(data.dependencies) ||
|
|
||||||
data.dependencies.length === 0 ||
|
|
||||||
data.dependencies.some(
|
|
||||||
(dependency) =>
|
|
||||||
!dependency ||
|
|
||||||
typeof dependency !== 'object' ||
|
|
||||||
(!Array.isArray(dependency.vulns) && !isSkipped(dependency))
|
|
||||||
)
|
|
||||||
) {
|
|
||||||
return null;
|
|
||||||
}
|
|
||||||
|
|
||||||
return data.dependencies.flatMap((dependency) =>
|
|
||||||
(dependency.vulns || [])
|
|
||||||
.filter((vulnerability) => !ignoredVulnIds.includes(vulnerability.id))
|
|
||||||
.map(
|
|
||||||
(vulnerability) => `- \`${dependency.name}==${dependency.version}\`: ${vulnerability.id}` +
|
|
||||||
(vulnerability.fix_versions?.length ? ` (fixed by ${vulnerability.fix_versions.join(', ')})` : '')
|
|
||||||
)
|
|
||||||
);
|
|
||||||
}
|
|
||||||
) + (ignoredVulnIds.length
|
|
||||||
? `\n\n_Excluded as accepted, non-actionable findings: ${ignoredVulnIds.join(', ')} — see the workflow file's inline comments for why._`
|
|
||||||
: '') + (skippedDeps.length
|
|
||||||
? `\n\n_Could not be audited: ${skippedDeps.map((d) => `\`${d.name}\` (${d.skip_reason})`).join(', ')}_`
|
|
||||||
: '');
|
|
||||||
|
|
||||||
const banditSection = renderSection(
|
const banditSection = renderSection(
|
||||||
'Bandit — HIGH-severity code issues',
|
'Bandit — HIGH-severity code issues',
|
||||||
@@ -379,7 +212,7 @@ jobs:
|
|||||||
const comment = [
|
const comment = [
|
||||||
'# 🔒 Security Scan Results',
|
'# 🔒 Security Scan Results',
|
||||||
'',
|
'',
|
||||||
pipAuditSection,
|
safetySection,
|
||||||
'',
|
'',
|
||||||
banditSection,
|
banditSection,
|
||||||
'',
|
'',
|
||||||
@@ -389,7 +222,7 @@ jobs:
|
|||||||
'',
|
'',
|
||||||
'*This security scan runs automatically on source-code PRs and bi-weekly (skipped for doc/markdown-only changes).*',
|
'*This security scan runs automatically on source-code PRs and bi-weekly (skipped for doc/markdown-only changes).*',
|
||||||
'',
|
'',
|
||||||
'📊 **Security Policy**: CI fails on pip-audit vulnerabilities and Bandit HIGH-severity findings. Semgrep findings above are informational and do not block merge.',
|
'📊 **Security Policy**: CI fails on Safety vulnerabilities and Bandit HIGH-severity findings. Semgrep findings above are informational and do not block merge.',
|
||||||
].join('\n');
|
].join('\n');
|
||||||
|
|
||||||
try {
|
try {
|
||||||
@@ -404,11 +237,3 @@ jobs:
|
|||||||
console.log('⚠️ Could not post security comment:', error.message);
|
console.log('⚠️ Could not post security comment:', error.message);
|
||||||
console.log('📋 Security scan results saved to artifacts');
|
console.log('📋 Security scan results saved to artifacts');
|
||||||
}
|
}
|
||||||
|
|
||||||
- name: Enforce Audit Gate
|
|
||||||
if: always()
|
|
||||||
run: |
|
|
||||||
if [ "${AUDIT_SCAN_STATUS:-failed}" != "passed" ]; then
|
|
||||||
echo "::error::pip-audit scan failed. See the pip-audit output and uploaded reports above."
|
|
||||||
exit 1
|
|
||||||
fi
|
|
||||||
|
|||||||
@@ -0,0 +1,42 @@
|
|||||||
|
name: Security
|
||||||
|
|
||||||
|
on:
|
||||||
|
schedule:
|
||||||
|
- cron: '0 0 * * 1'
|
||||||
|
workflow_dispatch:
|
||||||
|
pull_request:
|
||||||
|
branches: [main]
|
||||||
|
paths:
|
||||||
|
- 'pyproject.toml'
|
||||||
|
- 'requirements-ci.txt'
|
||||||
|
- '.github/workflows/security.yml'
|
||||||
|
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
audit:
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
steps:
|
||||||
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
|
||||||
|
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7
|
||||||
|
with:
|
||||||
|
python-version: '3.11'
|
||||||
|
# Upgrade first: actions/setup-python's baked-in setuptools has been
|
||||||
|
# behind known-vulnerable floors before (e.g. PYSEC-2026-3447 /
|
||||||
|
# setuptools 75.1.0), so don't trust the preinstalled one.
|
||||||
|
- run: pip install -r .github/requirements/bootstrap.txt --require-hashes
|
||||||
|
# Audit the pinned dependency set (requirements-ci.txt is compiled from
|
||||||
|
# pyproject.toml with --extra all — the same coverage as the [all]
|
||||||
|
# extra, minus the Linux-only gpu set — so this keeps scan parity with
|
||||||
|
# CI/release builds without a time-dependent resolution). This is the
|
||||||
|
# fix for PYSEC-2024-38 (#869): the bare-env job never had fastapi or
|
||||||
|
# python-multipart installed to look at.
|
||||||
|
- run: pip install -r requirements-ci.txt --require-hashes
|
||||||
|
# PR runs gate on findings, since they're scoped to actual
|
||||||
|
# pyproject.toml changes under review. The schedule/workflow_dispatch
|
||||||
|
# runs stay non-blocking until a full pass over pre-existing findings
|
||||||
|
# across the whole [all] tree has been done.
|
||||||
|
- run: pip install -r .github/requirements/pip-audit.txt --require-hashes
|
||||||
|
- run: pip-audit -r requirements-ci.txt
|
||||||
|
continue-on-error: ${{ github.event_name != 'pull_request' }}
|
||||||
BIN
Binary file not shown.
@@ -9,37 +9,6 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|||||||
|
|
||||||
## [Unreleased]
|
## [Unreleased]
|
||||||
|
|
||||||
### Added
|
|
||||||
|
|
||||||
- **Salesforce ingestor** (#1240) by @Sameer6305
|
|
||||||
- New `SalesforceConnector` / `SalesforceData` / `SalesforceIngestor` (`semantica.ingest`, lazy export), following the same Connector + Data + Ingestor pattern already used for Snowflake/Databricks/SAP
|
|
||||||
- Auth covers both landscapes Salesforce actually uses: username + password + security token (SOAP login), session_id + instance_url (reusing an existing session), and username + consumer_key + private key (JWT Bearer); production and sandbox are selected via `domain`, and credentials can come from environment variables. Credential material is never intentionally written to logs, exceptions, or `repr()`
|
|
||||||
- `ingest_sobject()`, `ingest_query()`, `list_sobjects()`, `get_sobject_schema()`, `export_as_documents()` against standard sObjects, custom objects (`__c`), custom metadata (`__mdt`), platform events (`__e`), namespaced objects, and relationship-field traversal (e.g. `Owner.Name`); pagination follows `nextRecordsUrl`/`query_more()` and stops once a caller's `limit` is satisfied
|
|
||||||
- New `pip install semantica[db-salesforce]` extra (`simple-salesforce>=1.12.0`)
|
|
||||||
- New `tests/test_salesforce_ingestor.py`
|
|
||||||
- Docs: `docs/integrations/salesforce.md`
|
|
||||||
|
|
||||||
- **`ErasureCoordinator` completes the erasure workflow `purge_node()` only starts — the graph node was removed while the same content survived verbatim in `AgentMemory` and as an embedding** (closes #1018) by @pravit-amp
|
|
||||||
- New `semantica/context/erasure.py`, exporting `ErasureCoordinator` and `ErasureReceipt` from `semantica.context`. `purge_node()`/`purge_edge()` (#957) are graph-scope by design and their changelog entry documents this gap explicitly; the changelog also names GDPR Article 17 as the motivation, and an Article 17 erasure that removes the node while the content stays retrievable by similarity search is not an erasure — it is worse than not offering one, because `purge_node()` returns `True` and writes a tombstone attesting the content is gone
|
|
||||||
- The coordinator **composes** the existing public APIs — nothing in `context_graph.py` or `agent_memory.py` changes behaviorally, and `ContextGraph` keeps its documented graph-scope contract rather than acquiring references to `AgentMemory`/`vector_store` that would invert the dependency
|
|
||||||
- `erase_entity(entity_id, reason=..., at=..., vector_ids=...)` returns an `ErasureReceipt`; `erase_entities([...])` returns one receipt per entity, in order, so one entity's failure does not stop the rest
|
|
||||||
- **Honest partial reporting is the point.** Each store reports one of five statuses — `erased`, `not_found`, `not_configured` (store never bound; normal), `unsupported` (store cannot delete at all; retrying will not help), `failed` — and `receipt.complete` is `False` when any store reports `unsupported`/`failed`, with `receipt.incomplete_stores` naming them. A receipt reading `graph: erased, memory: 14 erased, vectors: unsupported on faiss` is actionable; a bare `True` is a compliance liability
|
|
||||||
- **Erasure runs outward-in: vectors → memory → graph.** The graph tombstone is the durable attestation that an erasure happened, so writing it first would let a crash mid-cascade leave a record claiming more than occurred. Erasing the graph last means a partial failure leaves the node present and the receipt incomplete — recoverable and honest; the reverse is neither
|
|
||||||
- **Partial failure is a result, not an exception**: a store that raises is recorded as `failed` (with the exception type) and the remaining legs still run, rather than aborting into a half-erased state with no record of which half
|
|
||||||
- **The memory sweep cannot be silently truncated.** `find_by_entity(entity_id, limit=10)` returned `results[:limit]`, so the obvious hand-rolled cascade erases the first ten items and reports success — an erasure check computed from a page already truncated by the very `limit` it was called with. The coordinator sweeps in pages until dry (deleting as it goes, so the next page is the remainder) rather than passing one large number that is only correct until someone exceeds it, then **re-queries once after the sweep** and reports `failed` with the residual count if anything survived. It also stops rather than spinning if `batch_delete` reports no progress on a non-empty page. Note `find_by_entity` returns items keyed `memory_id`, not `id`
|
|
||||||
- **`unsupported` vector backends are detected by probing, not by calling and catching.** `faiss_store.py`, `milvus_store.py` and `weaviate_store.py` expose no delete at all (FAISS cannot remove from a flat index without a rebuild), while the `VectorStore` facade declares `delete_vectors()` for *every* backend and only raises `NotImplementedError` once called — so probing the facade alone cannot tell a deletable backend from a delete-less one, and the coordinator looks at the backend it wraps. Probing also keeps a missing method distinguishable from an `AttributeError` raised *inside* a working one, which is exactly where guessing wrong produces a false clean bill of health. `NotImplementedError` at call time is still caught and reported as `unsupported`; a store returning `False` is reported as `failed`
|
|
||||||
- Backends are reached under either supported name — `delete_vectors(ids)` (pinecone/qdrant) or `delete(ids)` (pgvector/sqlite-vec) — and the receipt records which was used
|
|
||||||
- `vector_store` defaults to `memory.vector_store` when a memory is supplied, stays overridable for deployments binding a store the memory does not own, and accepts `False` to disable the vector leg. Vectors owned by memory items are removed by the memory leg's own `delete_memory()` cascade; the explicit vector leg covers entity-keyed embeddings written by something other than `AgentMemory`
|
|
||||||
- The receipt's `erased_at` is normalized through `ContextGraph`'s own temporal normalizer, so the receipt and the tombstone written by the same erasure cannot disagree about when it happened; an unparseable `at` is rejected before any store is touched rather than half way through the cascade
|
|
||||||
- `purge_node()`'s docstring now points at the coordinator, so callers reading the graph-scope caveat find the thing that completes the workflow
|
|
||||||
- New `tests/context/test_erasure_coordinator.py`: 48 tests against **real** `ContextGraph`/`AgentMemory` instances rather than mocks — the bug lives in the interaction between them, so mocking it away would test nothing. Covers the 25-items-on-one-entity regression that fails against a naive single `find_by_entity()` call, all three vector-backend shapes (`delete_vectors`/`delete`/neither) plus the facade-over-delete-less-backend shape, residual/no-progress/no-identifier memory failures, partial failure continuing the cascade, idempotency, receipt serialization, and `at` normalization
|
|
||||||
- Full `tests/context/` suite: 738 passed
|
|
||||||
- **Fixed during review** (Qodo): `erase_entity()` resolved `erased_at` up front but passed the caller's original `at` down to `purge_node()`, so on the default `at=None` path the coordinator and the graph each took their own `now()` and the receipt attested to a different instant than the tombstone it points at — breaking the one invariant this module states most loudly. The resolved timestamp is now passed to the graph. The existing test passed only because it supplied an explicit `at`, which hides the drift; a regression test now covers the `at=None` path that callers actually use
|
|
||||||
- **Fixed during review** (Qodo): the vectors leg treated any return value other than the literal `False` as success, but no in-repo backend returns a bool — Qdrant returns `{"status": <UpdateStatus>}` and Pinecone `{"deleted": True}`, so every dict was read as a success and the backend's own account of the delete was discarded. Delete results are now interpreted by shape (bool, dict with explicit failure markers, `None` for a void method, anything else at face value) and the backend payload is recorded in the receipt as `backend_result`, stringified so the receipt stays JSON-serializable as the audit record it is meant to be. Bool markers are matched by identity so a `0` count is not read as `False`, and string markers match as substrings so an enum rendering as `"UpdateStatus.FAILED"` is not read as a success
|
|
||||||
- **Fixed during review** (Qodo): the constructor's "at least one store" guard used `not vector_store`, rejecting a valid store whose `__bool__`/`__len__` makes an empty instance falsey, and reporting `vector_store=None` in the error when an object had been passed; it now distinguishes `None` (absent) from `False` (deliberately disabled) from any other value (provided), and echoes what it actually received
|
|
||||||
- **Fixed during review** (Qodo): `at` annotations accepted only `str`/`datetime` while the shared `ContextGraph` normalizer they delegate to also takes epoch seconds; widened to `int`/`float` with the docstrings updated, so the coordinator no longer advertises less than the graph API it wraps
|
|
||||||
- **Known limitation, unchanged by this PR**: erasure still cannot be *completed* on FAISS/Milvus/Weaviate — `delete_vectors()` is declared on the `VectorStore` facade (`vector_store.py:786`) but not implemented across the backend set, under at least three different names. That is worth its own issue; the coordinator ships reporting `unsupported` and starts reporting `erased` for those backends once it is fixed, with no API change here
|
|
||||||
|
|
||||||
## [0.6.7] - 2026-08-28
|
## [0.6.7] - 2026-08-28
|
||||||
|
|
||||||
### Added
|
### Added
|
||||||
@@ -159,19 +128,6 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|||||||
- **Also fixed, on the JSON-LD paths**: the first fix covered the Turtle, N-Triples and RDF/XML serializers, and left both JSON-LD writers interpolating the entity's own text into `f"semantica:entity/{text}"` and the endpoints into `f"semantica:rel/{source}_{target}"`. Three consequences, all live in 0.6.5: an entity whose text contained a space produced an invalid IRI, and a JSON-LD parser dropped that node in full rather than reporting it, so the entity disappeared from the export; every relationship carrying `source`/`target` rather than `source_id`/`target_id` minted the identical `semantica:rel/_`, collapsing all of them onto one node whose types and endpoints merged; and the JSON-LD `@id` disagreed with the Turtle IRI for the same entity, so the two serializations of one knowledge graph were two different graphs. Both JSON-LD writers now use `mint_entity_iri`/`mint_relationship_iri`, and `JSONExporter.export_entities`/`export_relationships` declare the `semantica` prefix their `@context` was already writing `semantica:entities` against — without it a processor reads that as an IRI in the scheme `semantica`, which is the original #1101 defect on a third path
|
- **Also fixed, on the JSON-LD paths**: the first fix covered the Turtle, N-Triples and RDF/XML serializers, and left both JSON-LD writers interpolating the entity's own text into `f"semantica:entity/{text}"` and the endpoints into `f"semantica:rel/{source}_{target}"`. Three consequences, all live in 0.6.5: an entity whose text contained a space produced an invalid IRI, and a JSON-LD parser dropped that node in full rather than reporting it, so the entity disappeared from the export; every relationship carrying `source`/`target` rather than `source_id`/`target_id` minted the identical `semantica:rel/_`, collapsing all of them onto one node whose types and endpoints merged; and the JSON-LD `@id` disagreed with the Turtle IRI for the same entity, so the two serializations of one knowledge graph were two different graphs. Both JSON-LD writers now use `mint_entity_iri`/`mint_relationship_iri`, and `JSONExporter.export_entities`/`export_relationships` declare the `semantica` prefix their `@context` was already writing `semantica:entities` against — without it a processor reads that as an IRI in the scheme `semantica`, which is the original #1101 defect on a third path
|
||||||
- `tests/export/test_jsonld_iri_minting.py` parses each export with a real JSON-LD processor and asserts the entity survives, the relationships stay distinct, no term expands into the `semantica` scheme, and the JSON-LD `@id` equals the Turtle IRI
|
- `tests/export/test_jsonld_iri_minting.py` parses each export with a real JSON-LD processor and asserts the entity survives, the relationships stay distinct, no term expands into the `semantica` scheme, and the JSON-LD `@id` equals the Turtle IRI
|
||||||
- 236 export and ontology tests pass
|
- 236 export and ontology tests pass
|
||||||
- **`semantica.evals` runner gains per-metric objectives** (#1091)
|
|
||||||
- `evaluate()` now accepts `config={"<evaluator>": {"objective": {"direction": "maximize"|"minimize", "threshold": X}}}` to override the evaluator's default pass verdict with a threshold; `{"objective": {"expect": bool}}` expresses a Boolean expectation
|
|
||||||
- `minimize` requires a `threshold` — omitting it or setting it to `None` raises `ValueError`; `maximize` without a threshold is a no-op (the evaluator's own verdict stands); `expect` cannot be combined with `direction`/`threshold`; invalid config raises `ValueError` before any evaluator runs
|
|
||||||
- Error metrics are never affected by objectives (error wins over fail)
|
|
||||||
- Backward compatible: no `objective` key → existing behavior unchanged
|
|
||||||
- New tests in `tests/evals/test_runner.py::TestObjective`
|
|
||||||
- **`semantica.evals` is now a fully implemented evaluation module** (was a "Coming Soon" stub in the package layout)
|
|
||||||
- `evaluate(cases, evaluators, config=None, target_fn=None)` runner with per-case `pass`/`fail`/`error` status and an aggregate `pass_rate`, using a registry of named evaluators (`list_evaluators()`)
|
|
||||||
- 10 built-in evaluators: `exact_match`, `regex_match`, `numeric_range`, `temporal_range`, `length_range`, `keyword_check`, `levenshtein` (edit-distance similarity), `rouge` (in-house token F1, no new dependencies), `llm_as_judge` (lazy: caller-supplied `judge_fn`), and `decision_scores` (composite over `semantica.context.Decision`)
|
|
||||||
- `decision_scores` validates field-level (expected outcome, confidence bounds, non-empty maker/reasoning/scenario) and governance-level (provenance record presence; opt-in `PolicyEngine.check_compliance`) checks, coercing dict inputs via `Decision(**actual)` and never crashing on malformed input; an interface slot for causal-chain/embedding checks is reserved and raises `NotImplementedError` (V2)
|
|
||||||
- `__version__` is `0.1.0`, and the module ships a usage guide at `semantica/evals/usage.md` with worked import/run/interpret examples
|
|
||||||
- `semantica.evals` is reachable through the root package lazy module proxy (`semantica.evals`)
|
|
||||||
- 99 unit tests in `tests/evals/` covering every evaluator, registry errors, runner aggregation, decision coercion, and per-metric objectives; `python -m pytest tests/evals -q` → 99 passed
|
|
||||||
|
|
||||||
- **First-class CrewAI integration** (#988, closes #962) by @Shindevrp
|
- **First-class CrewAI integration** (#988, closes #962) by @Shindevrp
|
||||||
- New `pip install semantica[crewai]` extra (`crewai>=0.80.0`) — crewai core provides `BaseTool`/`BaseKnowledgeSource`, so `crewai-tools` is intentionally not included, and the extra is intentionally **not** part of the `all` bundle: crewai hard-requires `chromadb~=1.1.0`, which is affected by the unpatched pre-auth code-injection CVE-2026-45829 (see `integrations/crewai/README.md`)
|
- New `pip install semantica[crewai]` extra (`crewai>=0.80.0`) — crewai core provides `BaseTool`/`BaseKnowledgeSource`, so `crewai-tools` is intentionally not included, and the extra is intentionally **not** part of the `all` bundle: crewai hard-requires `chromadb~=1.1.0`, which is affected by the unpatched pre-auth code-injection CVE-2026-45829 (see `integrations/crewai/README.md`)
|
||||||
|
|||||||
+1
-1
@@ -20,7 +20,7 @@ RUN mkdir -p /app/semantica && npm run build
|
|||||||
# .github/dependabot.yml opens a PR bumping the digest pin above. Also: this
|
# .github/dependabot.yml opens a PR bumping the digest pin above. Also: this
|
||||||
# image only serves plain HTTP via uvicorn and never opens a QUIC listener,
|
# image only serves plain HTTP via uvicorn and never opens a QUIC listener,
|
||||||
# so the bug isn't reachable here regardless.
|
# so the bug isn't reachable here regardless.
|
||||||
FROM python:3.13-slim@sha256:7ce4b6dfe35e55397b7cda544f8a13f191b7ae28dc5aad71fe664dbc9bc2623f AS runtime
|
FROM python:3.14-slim@sha256:cae66f2ef0ec51a9891263eeee7f987dacf0a9879e8aa9353d5606e0530619a5 AS runtime
|
||||||
|
|
||||||
ENV PYTHONDONTWRITEBYTECODE=1 \
|
ENV PYTHONDONTWRITEBYTECODE=1 \
|
||||||
PYTHONUNBUFFERED=1 \
|
PYTHONUNBUFFERED=1 \
|
||||||
|
|||||||
@@ -18,9 +18,9 @@
|
|||||||
|
|
||||||
> Ingest your enterprise data, extract what matters, build a Context Graph and knowledge graph (KG), and run graph analytics and causal reasoning over all of it, with full decision provenance baked in. Explainable, traceable, and trustworthy by design.
|
> Ingest your enterprise data, extract what matters, build a Context Graph and knowledge graph (KG), and run graph analytics and causal reasoning over all of it, with full decision provenance baked in. Explainable, traceable, and trustworthy by design.
|
||||||
|
|
||||||
**Context Management · Knowledge Modeling · Deterministic Reasoning · Ontology Management · Decision Intelligence · End-to-End Traceability**
|
**Decision Intelligence · Context Management · Deterministic Reasoning · Ontology Management · Knowledge Modeling · End-to-End Traceability**
|
||||||
|
|
||||||
**Open Source · Governed · Zero Vendor Lock-In**
|
**Open Source · Self-Hostable · Auditable · Governed · Zero Vendor Lock-In**
|
||||||
|
|
||||||
**Polyglot Graph Storage · RDF & LPG Support · W3C Standards · Interoperable**
|
**Polyglot Graph Storage · RDF & LPG Support · W3C Standards · Interoperable**
|
||||||
|
|
||||||
@@ -56,18 +56,20 @@ pip install semantica
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
Most AI agents run on embeddings, not meaning: similarity scores with no structure, no relationships, and no way to explain why a result came back. Semantica is the semantic/context layer underneath your LLM, vector store, and agent framework: a deterministic infrastructure layer (no LLM required for graph construction, reasoning, or provenance) that turns fragmented enterprise data into a structured, queryable Context Graph and knowledge graph, governed by ontologies and controlled vocabularies (OWL, SHACL, SKOS) so the meaning of your data is explicit, not just its embedding. Decision provenance and audit trails fall out of that structure as a property, not the product itself; in domains a regulator can question, that same structure just happens to double as a straight answer to "why."
|
Most AI agents act without a trail. They store embeddings, not meaning: context that can't be explained, decisions that can't be audited. In lending, that gap is a compliance exposure, not an inconvenience: an underwriting agent's approval has to survive a regulator's "why" months later.
|
||||||
|
|
||||||
|
Semantica sits underneath your LLM, vector store, and agent framework as a deterministic infrastructure layer: no LLM required for graph construction, reasoning, or provenance.
|
||||||
|
|
||||||
> ⚠️ **System-level explainability, not foundation-model explainability.** Semantica does not expose or reconstruct what happens *inside* the LLM — its internal reasoning or chain-of-thought stays opaque, as it does for any external system. Semantica explains what's *outside* the model: the context and data fed in, the decision produced, its provenance, relevant relationships, applied policies, and the full execution trail.
|
> ⚠️ **System-level explainability, not foundation-model explainability.** Semantica does not expose or reconstruct what happens *inside* the LLM — its internal reasoning or chain-of-thought stays opaque, as it does for any external system. Semantica explains what's *outside* the model: the context and data fed in, the decision produced, its provenance, relevant relationships, applied policies, and the full execution trail.
|
||||||
|
|
||||||
**Who it's for:**
|
**Who it's for:**
|
||||||
|
|
||||||
- **AI/ML platform teams** shipping agents that make consequential decisions and need structured, queryable context, not just a vector index
|
- **AI/ML platform teams** shipping agents that make consequential decisions and need structured, queryable context built from fragmented raw data, not just a vector index
|
||||||
- **Data platform teams on Databricks or Snowflake** turning tables already in Unity Catalog or a warehouse into a governed, lineage-tracked knowledge graph, without exporting to a third-party SaaS
|
- **Data platform teams on Databricks or Snowflake** who need to turn tables already sitting in Unity Catalog or a Snowflake warehouse into a governed, lineage-tracked knowledge graph, without exporting that data to a third-party SaaS first
|
||||||
- **Compliance, risk, and audit teams** who need a straight answer to "why did the AI do that?" in a format a regulator accepts
|
- **Compliance, risk, and audit teams** who need a straight answer to "why did the AI do that?" in a format a regulator will actually accept
|
||||||
- **Regulated enterprises** (finance, healthcare, legal, government, defense) that can't ship a black box or send their data to someone else's SaaS to get one
|
- **Regulated enterprises** (finance, healthcare, legal, government, defense) that can't ship a black box, and can't send their data to someone else's SaaS to get one
|
||||||
- **Platform and infra engineers** who want the KG, reasoning, and provenance stack self-hosted and swappable, not locked to one vendor's backend
|
- **Platform and infra engineers** who want the KG, reasoning, and provenance stack self-hosted and swappable, not locked to one vendor's backend
|
||||||
- **Data and knowledge engineers** building a KG from messy, multi-source data, where conflicting facts get flagged and duplicates get merged, not silently overwritten
|
- **Data and knowledge engineers** building a KG from messy, multi-source data: entities and relationships get extracted, conflicting or contradictory facts are flagged instead of silently overwritten, and duplicates are merged before they turn into noise
|
||||||
|
|
||||||
**[Quick Start](#quick-start)** · **[Architecture](#architecture)** · **[What You Get](#what-semantica-gives-you)** · **[Why Semantica](#why-semantica)** · **[Decision Intelligence](#decision-intelligence)** · **[Context Graphs](#context-graphs)** · **[Recipe: Audit Trail](#recipe-audit-trail-for-a-regulated-decision)** · **[Module Reference](#module-reference)** · **[Integrations](#integrations)** · **[CLI](#cli)** · **[Performance](#performance)** · **[Install](#installation)**
|
**[Quick Start](#quick-start)** · **[Architecture](#architecture)** · **[What You Get](#what-semantica-gives-you)** · **[Why Semantica](#why-semantica)** · **[Decision Intelligence](#decision-intelligence)** · **[Context Graphs](#context-graphs)** · **[Recipe: Audit Trail](#recipe-audit-trail-for-a-regulated-decision)** · **[Module Reference](#module-reference)** · **[Integrations](#integrations)** · **[CLI](#cli)** · **[Performance](#performance)** · **[Install](#installation)**
|
||||||
|
|
||||||
@@ -81,7 +83,7 @@ Most AI agents run on embeddings, not meaning: similarity scores with no structu
|
|||||||
- **Full Auditability:** W3C PROV-O provenance on every fact, with audit trails exportable to JSON, CSV, or RDF
|
- **Full Auditability:** W3C PROV-O provenance on every fact, with audit trails exportable to JSON, CSV, or RDF
|
||||||
- **Deterministic Reasoning:** Forward chaining, Rete network, Datalog, and SPARQL with fully explainable paths, not black boxes
|
- **Deterministic Reasoning:** Forward chaining, Rete network, Datalog, and SPARQL with fully explainable paths, not black boxes
|
||||||
- **Knowledge Pipeline:** Multi-source ingestion, entity-aware chunking, NER/relation/event extraction, and knowledge graph construction, with semantic deduplication and provenance-preserving merges throughout
|
- **Knowledge Pipeline:** Multi-source ingestion, entity-aware chunking, NER/relation/event extraction, and knowledge graph construction, with semantic deduplication and provenance-preserving merges throughout
|
||||||
- **Enterprise Data Platforms:** Native connectors for Databricks (Unity Catalog + Delta Lake, PAT/OAuth M2M auth, catalog/schema/table/lineage introspection), Snowflake (warehouse/database/schema, key-pair and OAuth auth), and SAP OData (Business Partners, Sales Orders, OAuth2/Basic auth), so data already living in your lakehouse or warehouse becomes graph nodes with provenance, not another export/import hop
|
- **Enterprise Data Platforms:** Native connectors for Databricks (Unity Catalog + Delta Lake, PAT/OAuth M2M auth, catalog/schema/table/lineage introspection) and Snowflake (warehouse/database/schema, key-pair and OAuth auth), so tables already living in your lakehouse or warehouse become graph nodes with provenance, not another export/import hop
|
||||||
- **Graph Analytics:** Centrality, community detection, link prediction, and shortest-path queries over the graph you just built
|
- **Graph Analytics:** Centrality, community detection, link prediction, and shortest-path queries over the graph you just built
|
||||||
- **Polyglot Graph Storage:** Native RDF (embedded Oxigraph, Blazegraph, Apache Jena, Eclipse RDF4J via SPARQL) and Labeled Property Graphs (Neo4j, FalkorDB, Apache AGE, AWS Neptune via Cypher), plus vector stores, all swappable without touching your code
|
- **Polyglot Graph Storage:** Native RDF (embedded Oxigraph, Blazegraph, Apache Jena, Eclipse RDF4J via SPARQL) and Labeled Property Graphs (Neo4j, FalkorDB, Apache AGE, AWS Neptune via Cypher), plus vector stores, all swappable without touching your code
|
||||||
- **Visualization:** Explore any graph, ontology, or timeline in an interactive browser workbench
|
- **Visualization:** Explore any graph, ontology, or timeline in an interactive browser workbench
|
||||||
@@ -139,6 +141,10 @@ compliant = graph.check_decision_rules({"category": "vendor_selection"}) # poli
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
semantica doctor
|
semantica doctor
|
||||||
|
# Python 3.11.9 pass
|
||||||
|
# semantica 0.6.7 pass
|
||||||
|
# faiss vector store pass
|
||||||
|
# Config file pass ~/.semantica/config.yaml
|
||||||
```
|
```
|
||||||
|
|
||||||
**Running in a script or CI?** Progress bars are written only when stdout is an interactive terminal (or a Jupyter notebook), so piping and redirecting stay clean by default. Override with `SEMANTICA_DISABLE_PROGRESS=1` to silence progress everywhere, or `SEMANTICA_FORCE_PROGRESS=1` to keep it when stdout is redirected. `SEMANTICA_DISABLE_PROGRESS` takes precedence.
|
**Running in a script or CI?** Progress bars are written only when stdout is an interactive terminal (or a Jupyter notebook), so piping and redirecting stay clean by default. Override with `SEMANTICA_DISABLE_PROGRESS=1` to silence progress everywhere, or `SEMANTICA_FORCE_PROGRESS=1` to keep it when stdout is redirected. `SEMANTICA_DISABLE_PROGRESS` takes precedence.
|
||||||
@@ -163,7 +169,7 @@ Sources → Ingest → Parse → Normalize → Split → Extract → Conflict De
|
|||||||
→ Vector Store + Polyglot Graph Store (RDF & LPG) → Export / Visualize / REST · MCP · CLI
|
→ Vector Store + Polyglot Graph Store (RDF & LPG) → Export / Visualize / REST · MCP · CLI
|
||||||
```
|
```
|
||||||
|
|
||||||
- **Ingest:** files, web, databases, enterprise data platforms (Databricks, Snowflake, SAP), cloud (Google Drive, Elasticsearch), streams (Kafka, Kinesis), Git, email, MCP
|
- **Ingest:** files, web, databases, enterprise data platforms (Databricks, Snowflake), cloud (Google Drive, Elasticsearch), streams (Kafka, Kinesis), Git, email, MCP
|
||||||
- **Parse → Normalize → Split:** document parsing, text/entity/date normalization, GraphRAG-native entity-aware chunking
|
- **Parse → Normalize → Split:** document parsing, text/entity/date normalization, GraphRAG-native entity-aware chunking
|
||||||
- **Extract → Conflict Detection → Deduplication:** NER, relations, events, triplets; conflicting facts flagged and resolved before they merge
|
- **Extract → Conflict Detection → Deduplication:** NER, relations, events, triplets; conflicting facts flagged and resolved before they merge
|
||||||
- **Knowledge Graph:** `GraphBuilder` constructs the graph; bi-temporal facts and full graph analytics (centrality, communities, link prediction) run on top of it
|
- **Knowledge Graph:** `GraphBuilder` constructs the graph; bi-temporal facts and full graph analytics (centrality, communities, link prediction) run on top of it
|
||||||
@@ -273,7 +279,7 @@ retrieved = ctx.retrieve("who approved the Acme contract?")
|
|||||||
|
|
||||||
## Recipe: Audit Trail for a Regulated Decision
|
## Recipe: Audit Trail for a Regulated Decision
|
||||||
|
|
||||||
One pattern built on the same Context Graph: record a causally-linked decision chain, attach provenance to every entity, and export a regulator-ready audit trail.
|
The flagship pattern: record a causally-linked decision chain, attach provenance to every entity, and export a regulator-ready audit trail.
|
||||||
|
|
||||||
```python
|
```python
|
||||||
from semantica.context import ContextGraph
|
from semantica.context import ContextGraph
|
||||||
@@ -316,7 +322,7 @@ Every module below is independently importable, with working code samples verifi
|
|||||||
|
|
||||||
| Module | What it does |
|
| Module | What it does |
|
||||||
| --- | --- |
|
| --- | --- |
|
||||||
| [`semantica.ingest`](#semanticaingest-multi-source-ingestion) | Files, web, databases, APIs, streams, email, Git, Parquet, Databricks, Snowflake, SAP, MCP |
|
| [`semantica.ingest`](#semanticaingest-multi-source-ingestion) | Files, web, databases, APIs, streams, email, Git, Parquet, Databricks, Snowflake, MCP |
|
||||||
| [`semantica.semantic_extract`](#semanticasemantic_extract-ner-relations-events-triplets) | NER, relation extraction, event detection, triplet generation |
|
| [`semantica.semantic_extract`](#semanticasemantic_extract-ner-relations-events-triplets) | NER, relation extraction, event detection, triplet generation |
|
||||||
| [`semantica.kg`](#semanticakg-knowledge-graph-construction--analysis) | Graph construction, centrality, communities, link prediction |
|
| [`semantica.kg`](#semanticakg-knowledge-graph-construction--analysis) | Graph construction, centrality, communities, link prediction |
|
||||||
| [`semantica.reasoning`](#semanticareasoning-forward-chaining-rete-datalog-sparql) | Forward chaining, Rete, Datalog, SPARQL, fully explainable |
|
| [`semantica.reasoning`](#semanticareasoning-forward-chaining-rete-datalog-sparql) | Forward chaining, Rete, Datalog, SPARQL, fully explainable |
|
||||||
@@ -345,7 +351,7 @@ Expand any module below for its runnable example.
|
|||||||
<summary><b><code>semantica.ingest</code></b>: Multi-Source Ingestion</summary>
|
<summary><b><code>semantica.ingest</code></b>: Multi-Source Ingestion</summary>
|
||||||
<a id="semanticaingest-multi-source-ingestion"></a>
|
<a id="semanticaingest-multi-source-ingestion"></a>
|
||||||
|
|
||||||
Ingest from files, web, databases, APIs, streams, email, Git repos, Parquet, Databricks, Snowflake, SAP, or MCP servers, all through a unified interface.
|
Ingest from files, web, databases, APIs, streams, email, Git repos, Parquet, Databricks, Snowflake, or MCP servers, all through a unified interface.
|
||||||
|
|
||||||
```python
|
```python
|
||||||
from semantica.ingest import FileIngestor, WebIngestor, ParquetIngestor, DBIngestor
|
from semantica.ingest import FileIngestor, WebIngestor, ParquetIngestor, DBIngestor
|
||||||
@@ -396,7 +402,7 @@ orders = snowflake.ingest_table("ORDERS", limit=10_000)
|
|||||||
|
|
||||||
> **Security Note:** Never hardcode credentials (`token`, `password`, `private_key`) in production code; pass them via environment variables (e.g., `DATABRICKS_TOKEN`, `SNOWFLAKE_PASSWORD`) or a secrets manager.
|
> **Security Note:** Never hardcode credentials (`token`, `password`, `private_key`) in production code; pass them via environment variables (e.g., `DATABRICKS_TOKEN`, `SNOWFLAKE_PASSWORD`) or a secrets manager.
|
||||||
|
|
||||||
**Supported sources:** Local files (PDF, DOCX, PPTX, HTML, TXT, CSV, JSON, YAML, Excel, XML) · Web pages · RSS/Atom feeds · REST APIs · Databases (PostgreSQL, MySQL, SQLite, Oracle, SQL Server) · Parquet datasets · Databricks (Unity Catalog + Delta Lake) · Snowflake · SAP (OData v2/v4) · Git repositories · Email (IMAP/POP3) · Message streams (Kafka, RabbitMQ, Kinesis, Pulsar) · MCP resources · Apache Arrow/Feather/IPC (`ArrowIngestor`)
|
**Supported sources:** Local files (PDF, DOCX, PPTX, HTML, TXT, CSV, JSON, YAML, Excel, XML) · Web pages · RSS/Atom feeds · REST APIs · Databases (PostgreSQL, MySQL, SQLite, Oracle, SQL Server) · Parquet datasets · Databricks (Unity Catalog + Delta Lake) · Snowflake · Git repositories · Email (IMAP/POP3) · Message streams (Kafka, RabbitMQ, Kinesis, Pulsar) · MCP resources · Apache Arrow/Feather/IPC (`ArrowIngestor`)
|
||||||
|
|
||||||
DuckDB, Elasticsearch, Google Drive, HuggingFace, MongoDB, and Pandas ingestion also ship (`DuckDBIngestor`, `ElasticIngestor`, `GDriveIngestor`, `HuggingFaceIngestor`, `MongoIngestor`, `PandasIngestor`) but aren't re-exported from the top-level `semantica.ingest` namespace yet — import them directly: `from semantica.ingest.duckdb_ingestor import DuckDBIngestor`.
|
DuckDB, Elasticsearch, Google Drive, HuggingFace, MongoDB, and Pandas ingestion also ship (`DuckDBIngestor`, `ElasticIngestor`, `GDriveIngestor`, `HuggingFaceIngestor`, `MongoIngestor`, `PandasIngestor`) but aren't re-exported from the top-level `semantica.ingest` namespace yet — import them directly: `from semantica.ingest.duckdb_ingestor import DuckDBIngestor`.
|
||||||
|
|
||||||
@@ -1024,7 +1030,7 @@ team = Team(agents=[researcher, analyst], mode="coordinate")
|
|||||||
|
|
||||||
## More Recipes
|
## More Recipes
|
||||||
|
|
||||||
The audit-trail recipe is [above](#recipe-audit-trail-for-a-regulated-decision). Here are three more common patterns.
|
The flagship audit-trail recipe is [above](#recipe-audit-trail-for-a-regulated-decision). Here are three more common patterns.
|
||||||
|
|
||||||
<details>
|
<details>
|
||||||
<summary><b>End-to-End GraphRAG Pipeline</b></summary>
|
<summary><b>End-to-End GraphRAG Pipeline</b></summary>
|
||||||
@@ -1141,7 +1147,7 @@ if report.valid:
|
|||||||
| **Vector Store** | FAISS · Pinecone · Weaviate · Qdrant · Milvus · PgVector · hybrid + filtered search |
|
| **Vector Store** | FAISS · Pinecone · Weaviate · Qdrant · Milvus · PgVector · hybrid + filtered search |
|
||||||
| **Graph Databases (LPG)** | Neo4j · FalkorDB · Apache AGE · AWS Neptune |
|
| **Graph Databases (LPG)** | Neo4j · FalkorDB · Apache AGE · AWS Neptune |
|
||||||
| **Triple Stores (RDF)** | Oxigraph (embedded) · Blazegraph · Apache Jena · Eclipse RDF4J · unified `TripletStore` interface · SPARQL query & bulk load |
|
| **Triple Stores (RDF)** | Oxigraph (embedded) · Blazegraph · Apache Jena · Eclipse RDF4J · unified `TripletStore` interface · SPARQL query & bulk load |
|
||||||
| **Enterprise Data Platforms** | Databricks (`DatabricksIngestor`: Unity Catalog + Delta Lake, PAT/OAuth M2M, table/query ingestion, catalog/schema/table/lineage introspection) · Snowflake (`SnowflakeIngestor`: warehouse/database/schema, password/key-pair/OAuth auth) · SAP (`SAPIngestor`: OData v2/v4, OAuth2/Basic auth, Business Partners/Sales Orders) |
|
| **Enterprise Data Platforms** | Databricks (`DatabricksIngestor`: Unity Catalog + Delta Lake, PAT/OAuth M2M, table/query ingestion, catalog/schema/table/lineage introspection) · Snowflake (`SnowflakeIngestor`: warehouse/database/schema, password/key-pair/OAuth auth) |
|
||||||
| **LLM Providers** | **All already supported today:** OpenAI (GPT-4o, o1, o3) · Anthropic (Claude) · Google Gemini · Mistral · Meta Llama · Groq · Cohere · Azure OpenAI · AWS Bedrock · Ollama · DeepSeek · Perplexity · Together AI · Fireworks AI · Replicate · HuggingFace · via `semantica.llms` and LiteLLM |
|
| **LLM Providers** | **All already supported today:** OpenAI (GPT-4o, o1, o3) · Anthropic (Claude) · Google Gemini · Mistral · Meta Llama · Groq · Cohere · Azure OpenAI · AWS Bedrock · Ollama · DeepSeek · Perplexity · Together AI · Fireworks AI · Replicate · HuggingFace · via `semantica.llms` and LiteLLM |
|
||||||
|
|
||||||
---
|
---
|
||||||
@@ -1513,7 +1519,6 @@ pip install semantica[vectorstore-qdrant] # Qdrant vector store
|
|||||||
pip install semantica[vectorstore-pinecone] # Pinecone vector store
|
pip install semantica[vectorstore-pinecone] # Pinecone vector store
|
||||||
pip install semantica[db-snowflake] # Snowflake
|
pip install semantica[db-snowflake] # Snowflake
|
||||||
pip install semantica[db-databricks] # Databricks (SDK + SQL connector)
|
pip install semantica[db-databricks] # Databricks (SDK + SQL connector)
|
||||||
pip install semantica[ingest-sap] # SAP OData
|
|
||||||
pip install semantica[ingest-parquet] # Parquet / PyArrow
|
pip install semantica[ingest-parquet] # Parquet / PyArrow
|
||||||
pip install semantica[ingest-arrow] # Apache Arrow, Feather, IPC
|
pip install semantica[ingest-arrow] # Apache Arrow, Feather, IPC
|
||||||
pip install semantica[viz] # HTML interactive visualization
|
pip install semantica[viz] # HTML interactive visualization
|
||||||
|
|||||||
+3
-2
@@ -153,7 +153,7 @@ that attack chain.
|
|||||||
- **Risk**: a PR merges without its security/CI checks passing.
|
- **Risk**: a PR merges without its security/CI checks passing.
|
||||||
**Control**: merges require the `build`, `Analyze Python` (CodeQL), and `security-scan` checks to pass, in strict mode (checks must be re-run against the latest `main`).
|
**Control**: merges require the `build`, `Analyze Python` (CodeQL), and `security-scan` checks to pass, in strict mode (checks must be re-run against the latest `main`).
|
||||||
- **Risk**: a compromised scanner job reaches secrets or write access.
|
- **Risk**: a compromised scanner job reaches secrets or write access.
|
||||||
**Control**: scanning jobs (`CodeQL`, `security-scan.yml`, `defender-for-devops.yml`) run with read-only, least-privilege permissions (typically `contents: read` + `security-events: write` only) and never share a job, environment, or secret scope with the publish job.
|
**Control**: scanning jobs (`CodeQL`, `security-scan.yml`, `security.yml`, `defender-for-devops.yml`) run with read-only, least-privilege permissions (typically `contents: read` + `security-events: write` only) and never share a job, environment, or secret scope with the publish job.
|
||||||
- **Risk**: secrets are committed accidentally.
|
- **Risk**: secrets are committed accidentally.
|
||||||
**Control**: GitHub secret scanning and push protection are both enabled at the repository level, rejecting pushes that contain recognizable credential patterns before they land in history.
|
**Control**: GitHub secret scanning and push protection are both enabled at the repository level, rejecting pushes that contain recognizable credential patterns before they land in history.
|
||||||
|
|
||||||
@@ -164,7 +164,8 @@ Every scan below runs continuously in CI, not just at release time:
|
|||||||
- **CodeQL** (`security-and-quality` query pack) — Python source: injection, unsafe deserialization, and other code-level vulnerability classes. Runs in `codeql.yml` on every push/PR to `main` and weekly.
|
- **CodeQL** (`security-and-quality` query pack) — Python source: injection, unsafe deserialization, and other code-level vulnerability classes. Runs in `codeql.yml` on every push/PR to `main` and weekly.
|
||||||
- **Bandit** — Python-specific security anti-patterns (hardcoded secrets, unsafe `eval`/`pickle`, weak crypto, etc.); CI fails on any HIGH-severity finding. Runs in `security-scan.yml` on every push/PR to `main` and twice weekly.
|
- **Bandit** — Python-specific security anti-patterns (hardcoded secrets, unsafe `eval`/`pickle`, weak crypto, etc.); CI fails on any HIGH-severity finding. Runs in `security-scan.yml` on every push/PR to `main` and twice weekly.
|
||||||
- **Semgrep** (`p/security` ruleset) — cross-language static-analysis security patterns. Runs in `security-scan.yml` on every push/PR to `main` and twice weekly.
|
- **Semgrep** (`p/security` ruleset) — cross-language static-analysis security patterns. Runs in `security-scan.yml` on every push/PR to `main` and twice weekly.
|
||||||
- **pip-audit** — PyPA-maintained, OSV-backed vulnerability database cross-check against Semantica's pinned dependency tree, including optional LLM-provider extras such as LiteLLM; CI fails on any match. Runs in `security-scan.yml` on every push/PR to `main` and twice weekly, and can be triggered on demand via `workflow_dispatch`.
|
- **Safety** — known CVEs in Semantica's own installed dependencies, including optional LLM-provider extras such as LiteLLM; CI fails on any match. Runs in `security-scan.yml` on every push/PR to `main` and twice weekly.
|
||||||
|
- **pip-audit** — independent, PyPA-maintained vulnerability database cross-check against installed dependencies (Safety and pip-audit use different advisory sources, so both run). Runs in `security.yml` weekly.
|
||||||
- **Microsoft Defender for DevOps** (`eslint`, `templateanalyzer`, `terrascan`) — JavaScript/TypeScript lint-security rules and infrastructure-as-code misconfigurations. Runs in `defender-for-devops.yml` on every push/PR to `main` and weekly.
|
- **Microsoft Defender for DevOps** (`eslint`, `templateanalyzer`, `terrascan`) — JavaScript/TypeScript lint-security rules and infrastructure-as-code misconfigurations. Runs in `defender-for-devops.yml` on every push/PR to `main` and weekly.
|
||||||
- **Checkov** — Kubernetes, Helm, Dockerfile, GitHub Actions, and secrets-pattern IaC scanning; results upload to the same Security tab as CodeQL. Runs in `defender-for-devops.yml` on every push/PR to `main` and weekly.
|
- **Checkov** — Kubernetes, Helm, Dockerfile, GitHub Actions, and secrets-pattern IaC scanning; results upload to the same Security tab as CodeQL. Runs in `defender-for-devops.yml` on every push/PR to `main` and weekly.
|
||||||
- **GitGuardian** — secret-detection check on every pull request, installed as a GitHub App integration (not a repo-local workflow). Runs on every PR.
|
- **GitGuardian** — secret-detection check on every pull request, installed as a GitHub App integration (not a repo-local workflow). Runs on every PR.
|
||||||
|
|||||||
@@ -0,0 +1,222 @@
|
|||||||
|
{
|
||||||
|
"cells": [
|
||||||
|
{
|
||||||
|
"cell_type": "markdown",
|
||||||
|
"metadata": {},
|
||||||
|
"source": [
|
||||||
|
"[](https://colab.research.google.com/github/semantica-agi/semantica/blob/main/cookbook/advanced/09_Semantic_Layer_Construction.ipynb)\n",
|
||||||
|
"\n",
|
||||||
|
"# Semantic Layer Construction\n",
|
||||||
|
"\n",
|
||||||
|
"## Overview\n",
|
||||||
|
"\n",
|
||||||
|
"Build an enterprise semantic layer: construct knowledge graph, generate ontology, create semantic layer, export RDF, and store in triplet store.\n",
|
||||||
|
"\n",
|
||||||
|
"\n",
|
||||||
|
"**Documentation**: [API Reference](https://semantica.readthedocs.io/concepts/)\n",
|
||||||
|
"\n",
|
||||||
|
"## Installation\n",
|
||||||
|
"\n",
|
||||||
|
"Install Semantica from PyPI:\n",
|
||||||
|
"\n",
|
||||||
|
"```bash\n",
|
||||||
|
"pip install semantica\n",
|
||||||
|
"# Or with all optional dependencies:\n",
|
||||||
|
"pip install semantica[all]\n",
|
||||||
|
"```\n",
|
||||||
|
"\n",
|
||||||
|
"## Workflow: Build KG → Generate Ontology → Create Semantic Layer → Export RDF \n"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "code",
|
||||||
|
"execution_count": null,
|
||||||
|
"metadata": {},
|
||||||
|
"outputs": [],
|
||||||
|
"source": [
|
||||||
|
"!pip install -qU semantica\n"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "code",
|
||||||
|
"execution_count": null,
|
||||||
|
"metadata": {},
|
||||||
|
"outputs": [],
|
||||||
|
"source": [
|
||||||
|
"from semantica.kg import GraphBuilder\n",
|
||||||
|
"from semantica.ontology import OntologyGenerator\n",
|
||||||
|
"from semantica.export import RDFExporter\n",
|
||||||
|
"from semantica.triplet_store import TripletStore\n"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "markdown",
|
||||||
|
"metadata": {},
|
||||||
|
"source": [
|
||||||
|
"## Step 1: Build Knowledge Graph\n"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "code",
|
||||||
|
"execution_count": null,
|
||||||
|
"metadata": {},
|
||||||
|
"outputs": [],
|
||||||
|
"source": [
|
||||||
|
"builder = GraphBuilder()\n",
|
||||||
|
"\n",
|
||||||
|
"entities = [\n",
|
||||||
|
" {\"id\": \"e1\", \"type\": \"Person\", \"name\": \"Alice\", \"properties\": {\"age\": 30, \"role\": \"Engineer\"}},\n",
|
||||||
|
" {\"id\": \"e2\", \"type\": \"Person\", \"name\": \"Bob\", \"properties\": {\"age\": 35, \"role\": \"Manager\"}},\n",
|
||||||
|
" {\"id\": \"e3\", \"type\": \"Organization\", \"name\": \"Tech Corp\", \"properties\": {\"founded\": 2010}},\n",
|
||||||
|
" {\"id\": \"e4\", \"type\": \"Project\", \"name\": \"Project Alpha\", \"properties\": {\"status\": \"active\"}},\n",
|
||||||
|
"]\n",
|
||||||
|
"\n",
|
||||||
|
"relationships = [\n",
|
||||||
|
" {\"source\": \"e1\", \"target\": \"e2\", \"type\": \"reports_to\"},\n",
|
||||||
|
" {\"source\": \"e1\", \"target\": \"e3\", \"type\": \"works_for\"},\n",
|
||||||
|
" {\"source\": \"e2\", \"target\": \"e3\", \"type\": \"works_for\"},\n",
|
||||||
|
" {\"source\": \"e1\", \"target\": \"e4\", \"type\": \"works_on\"},\n",
|
||||||
|
"]\n",
|
||||||
|
"\n",
|
||||||
|
"knowledge_graph = builder.build(entities, relationships)\n"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "markdown",
|
||||||
|
"metadata": {},
|
||||||
|
"source": [
|
||||||
|
"## Step 2: Generate Ontology\n"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "code",
|
||||||
|
"execution_count": null,
|
||||||
|
"metadata": {},
|
||||||
|
"outputs": [],
|
||||||
|
"source": [
|
||||||
|
"generator = OntologyGenerator()\n",
|
||||||
|
"ontology = generator.generate_from_graph(knowledge_graph)\n"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "markdown",
|
||||||
|
"metadata": {},
|
||||||
|
"source": [
|
||||||
|
"## Step 3: Create Semantic Layer\n"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "code",
|
||||||
|
"execution_count": null,
|
||||||
|
"metadata": {},
|
||||||
|
"outputs": [],
|
||||||
|
"source": [
|
||||||
|
"def create_mappings(kg, ontology):\n",
|
||||||
|
" mappings = {\n",
|
||||||
|
" \"entity_type_mappings\": {},\n",
|
||||||
|
" \"relationship_type_mappings\": {},\n",
|
||||||
|
" \"property_mappings\": {}\n",
|
||||||
|
" }\n",
|
||||||
|
" \n",
|
||||||
|
" entity_types = set(e.get(\"type\") for e in entities)\n",
|
||||||
|
" ontology_classes = ontology.get(\"classes\", [])\n",
|
||||||
|
" \n",
|
||||||
|
" for entity_type in entity_types:\n",
|
||||||
|
" matching_class = next((cls for cls in ontology_classes if cls.get(\"name\") == entity_type), None)\n",
|
||||||
|
" if matching_class:\n",
|
||||||
|
" mappings[\"entity_type_mappings\"][entity_type] = matching_class.get(\"uri\", entity_type)\n",
|
||||||
|
" \n",
|
||||||
|
" relationship_types = set(r.get(\"type\") for r in relationships)\n",
|
||||||
|
" ontology_properties = ontology.get(\"properties\", [])\n",
|
||||||
|
" \n",
|
||||||
|
" for rel_type in relationship_types:\n",
|
||||||
|
" matching_prop = next((prop for prop in ontology_properties if prop.get(\"name\") == rel_type), None)\n",
|
||||||
|
" if matching_prop:\n",
|
||||||
|
" mappings[\"relationship_type_mappings\"][rel_type] = matching_prop.get(\"uri\", rel_type)\n",
|
||||||
|
" \n",
|
||||||
|
" return mappings\n",
|
||||||
|
"\n",
|
||||||
|
"mappings = create_mappings(knowledge_graph, ontology)\n",
|
||||||
|
"\n",
|
||||||
|
"semantic_layer = {\n",
|
||||||
|
" \"graph\": knowledge_graph,\n",
|
||||||
|
" \"ontology\": ontology,\n",
|
||||||
|
" \"mappings\": mappings,\n",
|
||||||
|
" \"metadata\": {\n",
|
||||||
|
" \"version\": \"1.0\",\n",
|
||||||
|
" \"created_at\": \"2024-01-01\",\n",
|
||||||
|
" \"description\": \"Enterprise semantic layer\"\n",
|
||||||
|
" }\n",
|
||||||
|
"}\n"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "markdown",
|
||||||
|
"metadata": {},
|
||||||
|
"source": [
|
||||||
|
"## Step 4: Export RDF\n"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "code",
|
||||||
|
"execution_count": null,
|
||||||
|
"metadata": {},
|
||||||
|
"outputs": [],
|
||||||
|
"source": [
|
||||||
|
"exporter = RDFExporter()\n",
|
||||||
|
"# Export Knowledge Graph\n",
|
||||||
|
"exporter.export(knowledge_graph, \"knowledge_graph.ttl\", format=\"turtle\")\n",
|
||||||
|
"print(\"Exported knowledge graph to knowledge_graph.ttl\")\n"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "markdown",
|
||||||
|
"metadata": {},
|
||||||
|
"source": [
|
||||||
|
"## Summary\n",
|
||||||
|
"\n",
|
||||||
|
"Enterprise semantic layer construction:\n",
|
||||||
|
"- Knowledge Graph Built\n",
|
||||||
|
"- Ontology Generated\n",
|
||||||
|
"- Semantic Layer Created with Mappings\n",
|
||||||
|
"- RDF Export Completed\n"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "markdown",
|
||||||
|
"metadata": {},
|
||||||
|
"source": []
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "markdown",
|
||||||
|
"metadata": {},
|
||||||
|
"source": []
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "markdown",
|
||||||
|
"metadata": {},
|
||||||
|
"source": []
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"metadata": {
|
||||||
|
"kernelspec": {
|
||||||
|
"display_name": "Python 3",
|
||||||
|
"language": "python",
|
||||||
|
"name": "python3"
|
||||||
|
},
|
||||||
|
"language_info": {
|
||||||
|
"codemirror_mode": {
|
||||||
|
"name": "ipython",
|
||||||
|
"version": 3
|
||||||
|
},
|
||||||
|
"file_extension": ".py",
|
||||||
|
"mimetype": "text/x-python",
|
||||||
|
"name": "python",
|
||||||
|
"nbconvert_exporter": "python",
|
||||||
|
"pygments_lexer": "ipython3",
|
||||||
|
"version": "3.11.9"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"nbformat": 4,
|
||||||
|
"nbformat_minor": 2
|
||||||
|
}
|
||||||
@@ -0,0 +1,435 @@
|
|||||||
|
{
|
||||||
|
"nbformat": 4,
|
||||||
|
"nbformat_minor": 5,
|
||||||
|
"metadata": {
|
||||||
|
"kernelspec": {
|
||||||
|
"display_name": "Python 3",
|
||||||
|
"language": "python",
|
||||||
|
"name": "python3"
|
||||||
|
},
|
||||||
|
"language_info": {
|
||||||
|
"name": "python",
|
||||||
|
"version": "3.10.0"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"cells": [
|
||||||
|
{
|
||||||
|
"cell_type": "markdown",
|
||||||
|
"id": "cell-0",
|
||||||
|
"metadata": {},
|
||||||
|
"source": [
|
||||||
|
"[](https://colab.research.google.com/github/semantica-agi/semantica/blob/main/cookbook/advanced/13_Manual_Ontology_Snowflake_Mapping.ipynb)\n",
|
||||||
|
"\n",
|
||||||
|
"# Manual Ontology + Snowflake Mapping\n",
|
||||||
|
"\n",
|
||||||
|
"This notebook answers a specific workflow:\n",
|
||||||
|
"\n",
|
||||||
|
"> *\"I want to design the ontology myself — not have AI infer it from my tables — and then map Snowflake data to it explicitly.\"*\n",
|
||||||
|
"\n",
|
||||||
|
"### What this notebook demonstrates\n",
|
||||||
|
"\n",
|
||||||
|
"| Step | What happens | Who controls it |\n",
|
||||||
|
"|---|---|---|\n",
|
||||||
|
"| 1 | Design ontology classes and properties | **You** (Python dict) |\n",
|
||||||
|
"| 2 | Model n-ary facts with reification | **You** (`AssociativeClassBuilder`) |\n",
|
||||||
|
"| 3 | Pull rows from Snowflake | Semantica `SnowflakeIngestor` |\n",
|
||||||
|
"| 4 | Map columns → ontology-aligned graph | **You** (explicit transform) |\n",
|
||||||
|
"| 5 | Validate + export OWL / SHACL | Semantica `OntologyEngine` |\n",
|
||||||
|
"| 6 | Load to triplet store and query | Semantica `TripletStore` |\n",
|
||||||
|
"\n",
|
||||||
|
"### What this notebook does NOT do\n",
|
||||||
|
"\n",
|
||||||
|
"- No LLM-driven ontology generation\n",
|
||||||
|
"- No schema introspection or table-to-class inference\n",
|
||||||
|
"- No \"suggest ontology from my data\"\n",
|
||||||
|
"\n",
|
||||||
|
"### Standards coverage\n",
|
||||||
|
"\n",
|
||||||
|
"| Feature | Status |\n",
|
||||||
|
"|---|---|\n",
|
||||||
|
"| OWL 2 (Turtle / RDF-XML) | Supported |\n",
|
||||||
|
"| SHACL 1.1 shapes | Supported |\n",
|
||||||
|
"| SPARQL 1.1 | Supported |\n",
|
||||||
|
"| Reification / n-ary facts | Supported via `AssociativeClassBuilder` |\n",
|
||||||
|
"| SPARQL 1.2 (reifier annotation, `LATERAL`) | Planned |\n",
|
||||||
|
"| SHACL 1.2 (`sh:severity` extensions, SHACL-AF) | Planned |"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "code",
|
||||||
|
"execution_count": null,
|
||||||
|
"id": "cell-1",
|
||||||
|
"metadata": {},
|
||||||
|
"outputs": [],
|
||||||
|
"source": [
|
||||||
|
"!pip install -qU semantica"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "code",
|
||||||
|
"execution_count": null,
|
||||||
|
"id": "cell-2",
|
||||||
|
"metadata": {},
|
||||||
|
"outputs": [],
|
||||||
|
"source": [
|
||||||
|
"import os\n",
|
||||||
|
"from typing import Any, Dict, List\n",
|
||||||
|
"\n",
|
||||||
|
"from semantica.ingest import SnowflakeIngestor\n",
|
||||||
|
"from semantica.kg.methods import build_kg\n",
|
||||||
|
"from semantica.ontology import AssociativeClassBuilder, OntologyEngine\n",
|
||||||
|
"from semantica.triplet_store import TripletStore"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "markdown",
|
||||||
|
"id": "cell-3",
|
||||||
|
"metadata": {},
|
||||||
|
"source": [
|
||||||
|
"## Step 1: Hand-Design the Ontology in Python\n",
|
||||||
|
"\n",
|
||||||
|
"You define every class and property explicitly. Nothing is read from Snowflake at this stage.\n",
|
||||||
|
"\n",
|
||||||
|
"**Design decisions that belong to you:**\n",
|
||||||
|
"- Which classes exist and what they mean\n",
|
||||||
|
"- Which properties are datatype vs. object properties\n",
|
||||||
|
"- Domain, range, and cardinality constraints\n",
|
||||||
|
"- Which properties are required (later enforced by SHACL)\n",
|
||||||
|
"\n",
|
||||||
|
"This dict versions with your code. It does not change when your database schema changes."
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "code",
|
||||||
|
"execution_count": null,
|
||||||
|
"id": "cell-4",
|
||||||
|
"metadata": {},
|
||||||
|
"outputs": [],
|
||||||
|
"source": "BASE_URI = \"https://example.com/hr/\"\n\n# Your ontology — designed by you, not inferred by Semantica.\nontology: Dict[str, Any] = {\n \"name\": \"EmploymentDomainOntology\",\n \"uri\": f\"{BASE_URI}EmploymentDomainOntology\",\n \"namespace\": {\"base_uri\": BASE_URI},\n\n # You decide the class taxonomy\n \"classes\": [\n {\"name\": \"Person\", \"uri\": f\"{BASE_URI}Person\"},\n {\"name\": \"Organization\", \"uri\": f\"{BASE_URI}Organization\"},\n {\"name\": \"Role\", \"uri\": f\"{BASE_URI}Role\"},\n # EmploymentEvent is a reification node.\n # It connects Person + Organization + Role and carries salary/date context.\n {\"name\": \"EmploymentEvent\", \"uri\": f\"{BASE_URI}EmploymentEvent\"},\n ],\n\n # Each property carries a full URI so TripletStore stores it as hr:<name>\n # rather than the default urn:property:<name>.\n # This ensures SPARQL queries using PREFIX hr: match what is actually stored.\n \"properties\": [\n # Datatype properties\n {\"name\": \"name\", \"uri\": f\"{BASE_URI}name\", \"type\": \"datatype\", \"domain\": \"Person\", \"range\": \"string\", \"required\": True},\n {\"name\": \"legalName\", \"uri\": f\"{BASE_URI}legalName\", \"type\": \"datatype\", \"domain\": \"Organization\", \"range\": \"string\", \"required\": True},\n {\"name\": \"title\", \"uri\": f\"{BASE_URI}title\", \"type\": \"datatype\", \"domain\": \"Role\", \"range\": \"string\", \"required\": True},\n {\"name\": \"startDate\", \"uri\": f\"{BASE_URI}startDate\", \"type\": \"datatype\", \"domain\": \"EmploymentEvent\", \"range\": \"date\"},\n {\"name\": \"endDate\", \"uri\": f\"{BASE_URI}endDate\", \"type\": \"datatype\", \"domain\": \"EmploymentEvent\", \"range\": \"date\"},\n {\"name\": \"salary\", \"uri\": f\"{BASE_URI}salary\", \"type\": \"datatype\", \"domain\": \"EmploymentEvent\", \"range\": \"decimal\"},\n\n # Object properties — reification spokes (required)\n {\"name\": \"employee\", \"uri\": f\"{BASE_URI}employee\", \"type\": \"object\", \"domain\": \"EmploymentEvent\", \"range\": \"Person\", \"required\": True},\n {\"name\": \"employer\", \"uri\": f\"{BASE_URI}employer\", \"type\": \"object\", \"domain\": \"EmploymentEvent\", \"range\": \"Organization\", \"required\": True},\n {\"name\": \"role\", \"uri\": f\"{BASE_URI}role\", \"type\": \"object\", \"domain\": \"EmploymentEvent\", \"range\": \"Role\", \"required\": True},\n\n # Shortcut edges — direct person→org / person→role without traversing the event node\n {\"name\": \"worksFor\", \"uri\": f\"{BASE_URI}worksFor\", \"type\": \"object\", \"domain\": \"Person\", \"range\": \"Organization\"},\n {\"name\": \"hasRole\", \"uri\": f\"{BASE_URI}hasRole\", \"type\": \"object\", \"domain\": \"Person\", \"range\": \"Role\"},\n ],\n}\n\nontology"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "markdown",
|
||||||
|
"id": "cell-5",
|
||||||
|
"metadata": {},
|
||||||
|
"source": [
|
||||||
|
"## Step 2: Reification — Modeling N-Ary Facts\n",
|
||||||
|
"\n",
|
||||||
|
"**The problem with binary triples:**\n",
|
||||||
|
"A simple triple `(Alice, worksFor, Acme)` cannot carry extra context such as salary, start date, or role.\n",
|
||||||
|
"Standard RDF reification and OWL n-ary patterns solve this by introducing an intermediate node.\n",
|
||||||
|
"\n",
|
||||||
|
"Semantica's `AssociativeClassBuilder` is the Pythonic API for this pattern:\n",
|
||||||
|
"\n",
|
||||||
|
"```\n",
|
||||||
|
"EmploymentEvent\n",
|
||||||
|
" ├── employee → Person (required)\n",
|
||||||
|
" ├── employer → Organization (required)\n",
|
||||||
|
" ├── role → Role (required)\n",
|
||||||
|
" ├── startDate → xsd:date\n",
|
||||||
|
" ├── endDate → xsd:date\n",
|
||||||
|
" └── salary → xsd:decimal\n",
|
||||||
|
"```\n",
|
||||||
|
"\n",
|
||||||
|
"**On SPARQL 1.1 vs. SPARQL 1.2:**\n",
|
||||||
|
"- **SPARQL 1.1 (current):** traverse the event node explicitly — `?event hr:employee ?person ; hr:salary ?salary`\n",
|
||||||
|
"- **SPARQL 1.2 (planned):** the draft reifier annotation syntax allows attaching context to triples directly, without a separate intermediate node. Semantica will adopt this once the spec is ratified.\n",
|
||||||
|
"\n",
|
||||||
|
"**On SHACL 1.1 vs. SHACL 1.2:**\n",
|
||||||
|
"- **SHACL 1.1 (current):** `sh:NodeShape` + `sh:PropertyShape` constraints are exported for all `required` properties and enforced at load time.\n",
|
||||||
|
"- **SHACL 1.2 (planned):** `sh:severity` profile extensions and SHACL-AF rules are on the roadmap."
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "code",
|
||||||
|
"execution_count": null,
|
||||||
|
"id": "cell-6",
|
||||||
|
"metadata": {},
|
||||||
|
"outputs": [],
|
||||||
|
"source": "assoc_builder = AssociativeClassBuilder()\n\nemployment_assoc = assoc_builder.create_associative_class(\n name=\"EmploymentEvent\",\n connects=[\"Person\", \"Organization\", \"Role\"],\n temporal=True, # adds startDate / endDate handling\n properties={\n \"startDate\": \"xsd:date\",\n \"endDate\": \"xsd:date\",\n \"salary\": \"xsd:decimal\",\n },\n)\n\nvalidation_result = assoc_builder.validate_associative_class(employment_assoc)\n\n# AssociativeClass is a dataclass — use attribute access, not .get()\nprint(\"AssociativeClass structure:\")\nprint(f\" name: {employment_assoc.name}\")\nprint(f\" connects: {employment_assoc.connects}\")\nprint(f\" temporal: {employment_assoc.temporal}\")\nprint(f\" properties: {list(employment_assoc.properties.keys())}\")\nprint(f\"\\nValidation passed: {validation_result}\")"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "markdown",
|
||||||
|
"id": "cell-7",
|
||||||
|
"metadata": {},
|
||||||
|
"source": [
|
||||||
|
"## Step 3: Ingest Snowflake Rows (Extraction Only)\n",
|
||||||
|
"\n",
|
||||||
|
"`SnowflakeIngestor` retrieves rows — nothing more. It does **not**:\n",
|
||||||
|
"- Inspect your table schema\n",
|
||||||
|
"- Suggest classes or properties\n",
|
||||||
|
"- Infer relationships from column names\n",
|
||||||
|
"\n",
|
||||||
|
"Set `USE_LIVE_SNOWFLAKE=true` plus the env vars below to connect to a real warehouse.\n",
|
||||||
|
"Otherwise the stub data is used."
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "code",
|
||||||
|
"execution_count": null,
|
||||||
|
"id": "cell-8",
|
||||||
|
"metadata": {},
|
||||||
|
"outputs": [],
|
||||||
|
"source": [
|
||||||
|
"def fetch_rows_from_snowflake() -> List[Dict[str, Any]]:\n",
|
||||||
|
" if os.getenv(\"USE_LIVE_SNOWFLAKE\", \"false\").lower() != \"true\":\n",
|
||||||
|
" return [\n",
|
||||||
|
" {\n",
|
||||||
|
" \"EMPLOYEE_ID\": \"E100\",\n",
|
||||||
|
" \"EMPLOYEE_NAME\": \"Alice Johnson\",\n",
|
||||||
|
" \"ORG_ID\": \"O10\",\n",
|
||||||
|
" \"ORG_NAME\": \"Acme Corp\",\n",
|
||||||
|
" \"ROLE_ID\": \"R7\",\n",
|
||||||
|
" \"ROLE_TITLE\": \"Senior Engineer\",\n",
|
||||||
|
" \"START_DATE\": \"2025-01-15\",\n",
|
||||||
|
" \"END_DATE\": None,\n",
|
||||||
|
" \"SALARY\": 160000,\n",
|
||||||
|
" },\n",
|
||||||
|
" {\n",
|
||||||
|
" \"EMPLOYEE_ID\": \"E101\",\n",
|
||||||
|
" \"EMPLOYEE_NAME\": \"Bob Singh\",\n",
|
||||||
|
" \"ORG_ID\": \"O10\",\n",
|
||||||
|
" \"ORG_NAME\": \"Acme Corp\",\n",
|
||||||
|
" \"ROLE_ID\": \"R9\",\n",
|
||||||
|
" \"ROLE_TITLE\": \"Data Architect\",\n",
|
||||||
|
" \"START_DATE\": \"2024-09-01\",\n",
|
||||||
|
" \"END_DATE\": None,\n",
|
||||||
|
" \"SALARY\": 185000,\n",
|
||||||
|
" },\n",
|
||||||
|
" ]\n",
|
||||||
|
"\n",
|
||||||
|
" ingestor = SnowflakeIngestor(\n",
|
||||||
|
" account=os.getenv(\"SNOWFLAKE_ACCOUNT\"),\n",
|
||||||
|
" user=os.getenv(\"SNOWFLAKE_USER\"),\n",
|
||||||
|
" password=os.getenv(\"SNOWFLAKE_PASSWORD\"),\n",
|
||||||
|
" warehouse=os.getenv(\"SNOWFLAKE_WAREHOUSE\"),\n",
|
||||||
|
" database=os.getenv(\"SNOWFLAKE_DATABASE\"),\n",
|
||||||
|
" schema=os.getenv(\"SNOWFLAKE_SCHEMA\", \"PUBLIC\"),\n",
|
||||||
|
" )\n",
|
||||||
|
" query = (\n",
|
||||||
|
" \"SELECT EMPLOYEE_ID, EMPLOYEE_NAME, \"\n",
|
||||||
|
" \"ORG_ID, ORG_NAME, ROLE_ID, ROLE_TITLE, \"\n",
|
||||||
|
" \"START_DATE, END_DATE, SALARY \"\n",
|
||||||
|
" \"FROM HR_EMPLOYMENT_FACT\"\n",
|
||||||
|
" )\n",
|
||||||
|
" data = ingestor.ingest_query(query)\n",
|
||||||
|
" ingestor.close()\n",
|
||||||
|
" return data.data\n",
|
||||||
|
"\n",
|
||||||
|
"\n",
|
||||||
|
"rows = fetch_rows_from_snowflake()\n",
|
||||||
|
"rows[:2]"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "markdown",
|
||||||
|
"id": "cell-9",
|
||||||
|
"metadata": {},
|
||||||
|
"source": [
|
||||||
|
"## Step 4: Map Rows to Ontology Concepts Explicitly\n",
|
||||||
|
"\n",
|
||||||
|
"This is the semantic transformation layer — the part that makes your ontology real.\n",
|
||||||
|
"\n",
|
||||||
|
"Semantica does not guess which column becomes which entity or property.\n",
|
||||||
|
"Every assignment is code you write and own:\n",
|
||||||
|
"\n",
|
||||||
|
"- **Stable node IDs** — deterministic, collision-safe, derived from business keys\n",
|
||||||
|
"- **Class assignment** — matches what you declared in Step 1\n",
|
||||||
|
"- **Property routing** — each column value goes to the correct ontology property\n",
|
||||||
|
"- **Reification wiring** — `EmploymentEvent` is linked to its three participants\n",
|
||||||
|
"\n",
|
||||||
|
"When your Snowflake schema changes, only this function needs updating. The ontology stays stable."
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "code",
|
||||||
|
"execution_count": null,
|
||||||
|
"id": "cell-10",
|
||||||
|
"metadata": {},
|
||||||
|
"outputs": [],
|
||||||
|
"source": "def map_rows_to_kg(rows: List[Dict[str, Any]]) -> Dict[str, Any]:\n entities: Dict[str, Dict[str, Any]] = {}\n relationships: List[Dict[str, Any]] = []\n\n for row in rows:\n # Stable, deterministic node IDs derived from business keys\n person_id = f\"person:{row['EMPLOYEE_ID']}\"\n org_id = f\"org:{row['ORG_ID']}\"\n role_id = f\"role:{row['ROLE_ID']}\"\n # Event ID includes all three participants + start date so that\n # a re-hired employee gets a distinct event node, not an overwrite.\n event_id = f\"employment:{row['EMPLOYEE_ID']}:{row['ORG_ID']}:{row['START_DATE']}\"\n\n # Entities — \"type\" must match a class name from Step 1\n entities[person_id] = {\n \"id\": person_id,\n \"type\": \"Person\",\n \"properties\": {\"name\": row[\"EMPLOYEE_NAME\"]},\n }\n entities[org_id] = {\n \"id\": org_id,\n \"type\": \"Organization\",\n \"properties\": {\"legalName\": row[\"ORG_NAME\"]},\n }\n entities[role_id] = {\n \"id\": role_id,\n \"type\": \"Role\",\n \"properties\": {\"title\": row[\"ROLE_TITLE\"]},\n }\n\n # Reification node — filter out None values so TripletStore does not\n # stringify None as the literal \"None\" for open-ended employment.\n event_props = {\n \"startDate\": row[\"START_DATE\"],\n \"endDate\": row[\"END_DATE\"],\n \"salary\": row[\"SALARY\"],\n }\n entities[event_id] = {\n \"id\": event_id,\n \"type\": \"EmploymentEvent\",\n \"properties\": {k: v for k, v in event_props.items() if v is not None},\n }\n\n # Full URIs for relationship types so TripletStore stores hr:<type>\n # instead of the default urn:property:<type>, keeping SPARQL consistent.\n relationships.extend([\n # Shortcut edges — fast SPARQL when context is not needed\n {\"source\": person_id, \"target\": org_id, \"type\": f\"{BASE_URI}worksFor\"},\n {\"source\": person_id, \"target\": role_id, \"type\": f\"{BASE_URI}hasRole\"},\n # Reification spokes — full context via the event node\n {\"source\": event_id, \"target\": person_id, \"type\": f\"{BASE_URI}employee\"},\n {\"source\": event_id, \"target\": org_id, \"type\": f\"{BASE_URI}employer\"},\n {\"source\": event_id, \"target\": role_id, \"type\": f\"{BASE_URI}role\"},\n ])\n\n return build_kg([{\"entities\": list(entities.values()), \"relationships\": relationships}])\n\n\nkg = map_rows_to_kg(rows)\nprint(f\"Entities built: {len(kg.get('entities', []))}\")\nprint(f\"Relationships built: {len(kg.get('relationships', []))}\")\n\nsample = next((e for e in kg[\"entities\"] if e[\"type\"] == \"EmploymentEvent\"), None)\nprint(f\"\\nSample EmploymentEvent node: {sample}\")"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "markdown",
|
||||||
|
"id": "cell-11",
|
||||||
|
"metadata": {},
|
||||||
|
"source": [
|
||||||
|
"## Step 5: Validate Ontology and Export OWL + SHACL\n",
|
||||||
|
"\n",
|
||||||
|
"`OntologyEngine` validates your ontology dict and serialises it to standards-compliant files.\n",
|
||||||
|
"\n",
|
||||||
|
"**Output files:**\n",
|
||||||
|
"- `employment_manual_ontology.ttl` — OWL 2 Turtle\n",
|
||||||
|
"- `employment_manual_shapes.ttl` — SHACL 1.1 node and property shapes\n",
|
||||||
|
"\n",
|
||||||
|
"**Standards status:**\n",
|
||||||
|
"\n",
|
||||||
|
"| Standard | Semantica support |\n",
|
||||||
|
"|---|---|\n",
|
||||||
|
"| SPARQL 1.1 | Full |\n",
|
||||||
|
"| SHACL 1.1 (`sh:NodeShape`, `sh:PropertyShape`, `sh:minCount`, `sh:datatype`, `sh:class`) | Full |\n",
|
||||||
|
"| SPARQL 1.2 (reifier annotation syntax, `LATERAL`) | Tracked — not yet implemented |\n",
|
||||||
|
"| SHACL 1.2 (`sh:severity` profiles, SHACL-AF extensions) | Tracked — not yet implemented |"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "code",
|
||||||
|
"execution_count": null,
|
||||||
|
"id": "cell-12",
|
||||||
|
"metadata": {},
|
||||||
|
"outputs": [],
|
||||||
|
"source": [
|
||||||
|
"engine = OntologyEngine(base_uri=BASE_URI)\n",
|
||||||
|
"\n",
|
||||||
|
"validation = engine.validate(ontology)\n",
|
||||||
|
"owl_ttl = engine.to_owl(ontology, format=\"turtle\")\n",
|
||||||
|
"shacl_ttl = engine.to_shacl(ontology, format=\"turtle\")\n",
|
||||||
|
"\n",
|
||||||
|
"engine.export_owl(ontology, \"employment_manual_ontology.ttl\", format=\"turtle\")\n",
|
||||||
|
"engine.export_shacl(ontology, \"employment_manual_shapes.ttl\", format=\"turtle\")\n",
|
||||||
|
"\n",
|
||||||
|
"print(f\"Ontology valid: {validation.valid}\")\n",
|
||||||
|
"print(f\"Ontology consistent: {validation.consistent}\")\n",
|
||||||
|
"print(f\"OWL output: {len(owl_ttl):,} chars → employment_manual_ontology.ttl\")\n",
|
||||||
|
"print(f\"SHACL output: {len(shacl_ttl):,} chars → employment_manual_shapes.ttl\")\n",
|
||||||
|
"\n",
|
||||||
|
"print(\"\\n--- SHACL shapes (first 20 lines) ---\")\n",
|
||||||
|
"print(\"\\n\".join(shacl_ttl.splitlines()[:20]))"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "markdown",
|
||||||
|
"id": "cell-13",
|
||||||
|
"metadata": {},
|
||||||
|
"source": [
|
||||||
|
"## Best-Practice Architecture\n",
|
||||||
|
"\n",
|
||||||
|
"```\n",
|
||||||
|
"┌──────────────────────────────────┐\n",
|
||||||
|
"│ Ontology as code (Python dict) │ ← versioned alongside your application\n",
|
||||||
|
"│ + AssociativeClass for n-ary │\n",
|
||||||
|
"└───────────────┬──────────────────┘\n",
|
||||||
|
" │ validate + export\n",
|
||||||
|
" ▼\n",
|
||||||
|
"┌───────────────────────────────────┐\n",
|
||||||
|
"│ OWL 2 Turtle │ SHACL 1.1 │ ← standards-compliant artifacts\n",
|
||||||
|
"└───────────────┬───────────────────┘\n",
|
||||||
|
" │\n",
|
||||||
|
" ▼\n",
|
||||||
|
"┌──────────────────────────────────┐\n",
|
||||||
|
"│ Snowflake — raw data access │ ← no schema introspection\n",
|
||||||
|
"└───────────────┬──────────────────┘\n",
|
||||||
|
" │ explicit mapping layer\n",
|
||||||
|
" ▼\n",
|
||||||
|
"┌──────────────────────────────────┐\n",
|
||||||
|
"│ Ontology-aligned KG │ ← types, IDs, edges match Step 1\n",
|
||||||
|
"└───────────────┬──────────────────┘\n",
|
||||||
|
" │ optional\n",
|
||||||
|
" ▼\n",
|
||||||
|
"┌──────────────────────────────────┐\n",
|
||||||
|
"│ Triplet store + SPARQL 1.1 │\n",
|
||||||
|
"└──────────────────────────────────┘\n",
|
||||||
|
"```\n",
|
||||||
|
"\n",
|
||||||
|
"**Why this split matters:**\n",
|
||||||
|
"If Semantica inferred the ontology from your Snowflake schema, every schema migration would risk silently changing your semantic model.\n",
|
||||||
|
"With this pattern, schema changes only touch the mapping function in Step 4 — the ontology remains stable and under your control."
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "markdown",
|
||||||
|
"id": "cell-14",
|
||||||
|
"metadata": {},
|
||||||
|
"source": [
|
||||||
|
"## SPARQL Query Patterns\n",
|
||||||
|
"\n",
|
||||||
|
"Two query styles are available because we wrote both shortcut edges and reification spokes.\n",
|
||||||
|
"\n",
|
||||||
|
"### Simple lookup — shortcut edge (no context needed)\n",
|
||||||
|
"\n",
|
||||||
|
"```sparql\n",
|
||||||
|
"PREFIX hr: <https://example.com/hr/>\n",
|
||||||
|
"\n",
|
||||||
|
"SELECT ?personName ?orgName\n",
|
||||||
|
"WHERE {\n",
|
||||||
|
" ?person a hr:Person ;\n",
|
||||||
|
" hr:name ?personName ;\n",
|
||||||
|
" hr:worksFor ?org .\n",
|
||||||
|
" ?org hr:legalName ?orgName .\n",
|
||||||
|
"}\n",
|
||||||
|
"```\n",
|
||||||
|
"\n",
|
||||||
|
"### Contextual lookup — via reification node (salary, dates, role)\n",
|
||||||
|
"\n",
|
||||||
|
"```sparql\n",
|
||||||
|
"PREFIX hr: <https://example.com/hr/>\n",
|
||||||
|
"\n",
|
||||||
|
"SELECT ?personName ?roleTitle ?salary ?startDate\n",
|
||||||
|
"WHERE {\n",
|
||||||
|
" ?event a hr:EmploymentEvent ;\n",
|
||||||
|
" hr:employee ?person ;\n",
|
||||||
|
" hr:role ?role ;\n",
|
||||||
|
" hr:salary ?salary ;\n",
|
||||||
|
" hr:startDate ?startDate .\n",
|
||||||
|
" ?person hr:name ?personName .\n",
|
||||||
|
" ?role hr:title ?roleTitle .\n",
|
||||||
|
"}\n",
|
||||||
|
"ORDER BY DESC(?salary)\n",
|
||||||
|
"```\n",
|
||||||
|
"\n",
|
||||||
|
"### Future: SPARQL 1.2 reifier syntax\n",
|
||||||
|
"\n",
|
||||||
|
"The SPARQL 1.2 draft introduces annotation syntax that lets you attach context directly to triples, without a separate intermediate node.\n",
|
||||||
|
"Once the spec is ratified Semantica will adopt it, and the contextual query above may be expressible more concisely."
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "markdown",
|
||||||
|
"id": "cell-15",
|
||||||
|
"metadata": {},
|
||||||
|
"source": [
|
||||||
|
"## Step 6 (Optional): Load to Triplet Store and Run SPARQL\n",
|
||||||
|
"\n",
|
||||||
|
"Set `STORE_TO_TRIPLET=true` to load the KG into a live triplet store and run the contextual reification query."
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"cell_type": "code",
|
||||||
|
"execution_count": null,
|
||||||
|
"id": "cell-16",
|
||||||
|
"metadata": {},
|
||||||
|
"outputs": [],
|
||||||
|
"source": [
|
||||||
|
"if os.getenv(\"STORE_TO_TRIPLET\", \"false\").lower() == \"true\":\n",
|
||||||
|
" store = TripletStore(\n",
|
||||||
|
" backend=os.getenv(\"TRIPLET_BACKEND\", \"blazegraph\"),\n",
|
||||||
|
" endpoint=os.getenv(\"TRIPLET_ENDPOINT\", \"http://localhost:9999/blazegraph\"),\n",
|
||||||
|
" namespace=os.getenv(\"TRIPLET_NAMESPACE\", \"kb\"),\n",
|
||||||
|
" )\n",
|
||||||
|
" store_result = store.store(knowledge_graph=kg, ontology=ontology)\n",
|
||||||
|
" print(\"Store result:\", store_result)\n",
|
||||||
|
"\n",
|
||||||
|
" # Contextual reification query — person + role + salary via EmploymentEvent\n",
|
||||||
|
" query = \"\"\"\n",
|
||||||
|
" PREFIX hr: <https://example.com/hr/>\n",
|
||||||
|
"\n",
|
||||||
|
" SELECT ?personName ?roleTitle ?salary ?startDate\n",
|
||||||
|
" WHERE {\n",
|
||||||
|
" ?event a hr:EmploymentEvent ;\n",
|
||||||
|
" hr:employee ?person ;\n",
|
||||||
|
" hr:role ?role ;\n",
|
||||||
|
" hr:salary ?salary ;\n",
|
||||||
|
" hr:startDate ?startDate .\n",
|
||||||
|
" ?person hr:name ?personName .\n",
|
||||||
|
" ?role hr:title ?roleTitle .\n",
|
||||||
|
" }\n",
|
||||||
|
" ORDER BY DESC(?salary)\n",
|
||||||
|
" LIMIT 10\n",
|
||||||
|
" \"\"\"\n",
|
||||||
|
" result = store.execute_query(query)\n",
|
||||||
|
" print(result)\n",
|
||||||
|
"else:\n",
|
||||||
|
" print(\"Skipping triplet-store load/query (set STORE_TO_TRIPLET=true to enable)\")"
|
||||||
|
]
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
@@ -10,16 +10,15 @@
|
|||||||
"\n",
|
"\n",
|
||||||
"## Overview\n",
|
"## Overview\n",
|
||||||
"\n",
|
"\n",
|
||||||
"This notebook demonstrates how to build knowledge graphs from extracted entities and relationships using Semantica's graph building modules. You'll learn to use `GraphBuilder` and `EntityResolver`.\n",
|
"This notebook demonstrates how to build knowledge graphs from entities and relationships using Semantica's graph building modules. You'll learn to use `GraphBuilder` and `EntityResolver`.\n",
|
||||||
"\n",
|
"\n",
|
||||||
"**Documentation**: [API Reference](https://semantica.readthedocs.io/reference/kg/)\n",
|
"**Documentation**: [API Reference](https://semantica.readthedocs.io/reference/kg/)\n",
|
||||||
"\n",
|
"\n",
|
||||||
"### Learning Objectives\n",
|
"### Learning Objectives\n",
|
||||||
"\n",
|
"\n",
|
||||||
"- Extract entity mentions and relations, and map them into graph records\n",
|
"- Use `GraphBuilder` to construct knowledge graphs\n",
|
||||||
"- Use `GraphBuilder` to construct a graph whose edges come from the actual extracted relations\n",
|
"- Use `EntityResolver` to resolve entity conflicts\n",
|
||||||
"- Use `EntityResolver` to merge duplicate mentions and remap relationship endpoints\n",
|
"**Note**: For deduplication, use the `semantica.deduplication` module.\n",
|
||||||
"- Use the `semantica.deduplication` module and report the complete deduplicated entity set\n",
|
|
||||||
"\n",
|
"\n",
|
||||||
"## Installation\n",
|
"## Installation\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -33,217 +32,120 @@
|
|||||||
"\n",
|
"\n",
|
||||||
"---\n",
|
"---\n",
|
||||||
"\n",
|
"\n",
|
||||||
"## Step 1: Extract Entities and Relations\n",
|
"## Step 1: Build Knowledge Graph\n",
|
||||||
"\n",
|
"\n",
|
||||||
"Extract entity mentions and relations from text. The sample text mentions `Apple Inc.` in two separate sentences, so we can later show how duplicate mentions are resolved into one canonical entity.\n"
|
"Construct a knowledge graph from entities and relationships.\n"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"metadata": {},
|
|
||||||
"source": [
|
|
||||||
"%pip install semantica\n",
|
|
||||||
"\n",
|
|
||||||
"# spaCy models are distributed separately from the spaCy library. This lesson\n",
|
|
||||||
"# relies on the English model to recognize standalone places such as Cupertino.\n",
|
|
||||||
"import sys\n",
|
|
||||||
"import subprocess\n",
|
|
||||||
"import spacy\n",
|
|
||||||
"\n",
|
|
||||||
"try:\n",
|
|
||||||
" spacy.load(\"en_core_web_sm\")\n",
|
|
||||||
"except OSError:\n",
|
|
||||||
" subprocess.check_call([sys.executable, \"-m\", \"spacy\", \"download\", \"en_core_web_sm\"])\n"
|
|
||||||
],
|
|
||||||
"execution_count": null,
|
"execution_count": null,
|
||||||
"outputs": []
|
"metadata": {},
|
||||||
|
"outputs": [],
|
||||||
|
"source": [
|
||||||
|
"!pip install semantica\n"
|
||||||
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
|
"execution_count": null,
|
||||||
"metadata": {},
|
"metadata": {},
|
||||||
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
|
"from semantica.kg import GraphBuilder\n",
|
||||||
"from semantica.semantic_extract import NERExtractor, RelationExtractor\n",
|
"from semantica.semantic_extract import NERExtractor, RelationExtractor\n",
|
||||||
"\n",
|
"\n",
|
||||||
"text = (\n",
|
"builder = GraphBuilder()\n",
|
||||||
" \"Apple Inc. is headquartered in Cupertino, California. \"\n",
|
|
||||||
" \"Tim Cook is the CEO of Apple Inc. \"\n",
|
|
||||||
" \"The company is a technology company.\"\n",
|
|
||||||
")\n",
|
|
||||||
"\n",
|
|
||||||
"ner_extractor = NERExtractor()\n",
|
"ner_extractor = NERExtractor()\n",
|
||||||
"relation_extractor = RelationExtractor()\n",
|
"relation_extractor = RelationExtractor()\n",
|
||||||
"\n",
|
"\n",
|
||||||
"mentions = ner_extractor.extract(text)\n",
|
"text = \"Apple Inc. is a technology company. Tim Cook is the CEO of Apple Inc. Apple Inc. is headquartered in Cupertino, California.\"\n",
|
||||||
"relations = relation_extractor.extract(text, mentions)\n",
|
|
||||||
"\n",
|
"\n",
|
||||||
"print(\"Entity mentions:\")\n",
|
"entities_list = ner_extractor.extract(text)\n",
|
||||||
"for mention in mentions:\n",
|
"relationships_list = relation_extractor.extract(text, entities_list)\n",
|
||||||
" print(f\" {mention.text!r:<13} {mention.label:<7} span=[{mention.start_char}:{mention.end_char}]\")\n",
|
|
||||||
"\n",
|
|
||||||
"print(\"\\nExtracted relations:\")\n",
|
|
||||||
"for rel in relations:\n",
|
|
||||||
" print(f\" {rel.subject.text!r} --{rel.predicate}--> {rel.object.text!r}\")"
|
|
||||||
],
|
|
||||||
"execution_count": null,
|
|
||||||
"outputs": []
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"cell_type": "markdown",
|
|
||||||
"metadata": {},
|
|
||||||
"source": [
|
|
||||||
"## Step 2: Build the Knowledge Graph\n",
|
|
||||||
"\n",
|
|
||||||
"Give every mention a graph ID, then translate each relation's `subject` and `object` into those IDs. Building edges from the actual relation endpoints — rather than guessing endpoints from list positions — is what keeps the graph faithful to the text.\n"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"cell_type": "code",
|
|
||||||
"metadata": {},
|
|
||||||
"source": [
|
|
||||||
"from semantica.kg import GraphBuilder\n",
|
|
||||||
"\n",
|
"\n",
|
||||||
"entities = []\n",
|
"entities = []\n",
|
||||||
"span_to_id = {}\n",
|
"for i, entity in enumerate(entities_list[:5], 1):\n",
|
||||||
"for i, mention in enumerate(mentions, 1):\n",
|
|
||||||
" graph_id = f\"e{i}\"\n",
|
|
||||||
" span_to_id[(mention.start_char, mention.end_char)] = graph_id\n",
|
|
||||||
" entities.append({\n",
|
" entities.append({\n",
|
||||||
" \"id\": graph_id,\n",
|
" \"id\": f\"e{i}\",\n",
|
||||||
" \"type\": mention.label,\n",
|
" \"type\": entity.label,\n",
|
||||||
" \"name\": mention.text,\n",
|
" \"name\": entity.text,\n",
|
||||||
" \"properties\": {},\n",
|
" \"properties\": {}\n",
|
||||||
" })\n",
|
" })\n",
|
||||||
"\n",
|
"\n",
|
||||||
"relationships = []\n",
|
"relationships = []\n",
|
||||||
"for rel in relations:\n",
|
"for i, rel in enumerate(relationships_list[:3], 1):\n",
|
||||||
" source_id = span_to_id.get((rel.subject.start_char, rel.subject.end_char))\n",
|
|
||||||
" target_id = span_to_id.get((rel.object.start_char, rel.object.end_char))\n",
|
|
||||||
" if source_id is None or target_id is None:\n",
|
|
||||||
" print(f\"Skipping relation with unmapped endpoint: \"\n",
|
|
||||||
" f\"{rel.subject.text!r} --{rel.predicate}--> {rel.object.text!r}\")\n",
|
|
||||||
" continue\n",
|
|
||||||
" relationships.append({\n",
|
" relationships.append({\n",
|
||||||
" \"source\": source_id,\n",
|
" \"source\": f\"e{1}\",\n",
|
||||||
" \"target\": target_id,\n",
|
" \"target\": f\"e{i+1}\",\n",
|
||||||
" \"type\": rel.predicate,\n",
|
" \"type\": rel.predicate,\n",
|
||||||
" \"properties\": {},\n",
|
" \"properties\": {}\n",
|
||||||
" })\n",
|
" })\n",
|
||||||
"\n",
|
"\n",
|
||||||
"builder = GraphBuilder()\n",
|
"knowledge_graph = builder.build(entities, relationships)\n",
|
||||||
"knowledge_graph = builder.build({\"entities\": entities, \"relationships\": relationships})\n",
|
|
||||||
"\n",
|
"\n",
|
||||||
"id_to_name = {entity[\"id\"]: entity[\"name\"] for entity in entities}\n",
|
"print(f\"Built knowledge graph with {len(knowledge_graph.get('entities', []))} entities\")\n",
|
||||||
"\n",
|
"print(f\"Relationships: {len(knowledge_graph.get('relationships', []))}\")"
|
||||||
"print(f\"Graph entities ({len(knowledge_graph['entities'])}):\")\n",
|
]
|
||||||
"for entity in knowledge_graph[\"entities\"]:\n",
|
|
||||||
" print(f\" {entity['id']}: {entity['name']} ({entity['type']})\")\n",
|
|
||||||
"\n",
|
|
||||||
"print(f\"\\nGraph relationships ({len(knowledge_graph['relationships'])}):\")\n",
|
|
||||||
"for relationship in knowledge_graph[\"relationships\"]:\n",
|
|
||||||
" print(f\" {id_to_name[relationship['source']]} \"\n",
|
|
||||||
" f\"--{relationship['type']}--> {id_to_name[relationship['target']]}\")\n",
|
|
||||||
"\n",
|
|
||||||
"edges = {\n",
|
|
||||||
" (id_to_name[r[\"source\"]], r[\"type\"], id_to_name[r[\"target\"]])\n",
|
|
||||||
" for r in knowledge_graph[\"relationships\"]\n",
|
|
||||||
"}\n",
|
|
||||||
"assert (\"Apple Inc.\", \"located_in\", \"Cupertino\") in edges\n",
|
|
||||||
"assert (\"Tim Cook\", \"works_for\", \"Apple Inc.\") in edges"
|
|
||||||
],
|
|
||||||
"execution_count": null,
|
|
||||||
"outputs": []
|
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"metadata": {},
|
"metadata": {},
|
||||||
"source": [
|
"source": [
|
||||||
"## Step 3: Entity Resolution\n",
|
"## Step 2: Entity Resolution\n",
|
||||||
"\n",
|
"\n",
|
||||||
"The graph currently contains two nodes for the same organization. `EntityResolver` merges duplicate mentions into one canonical entity and records which source IDs were merged (`merged_from`), so relationship endpoints can be remapped onto the canonical entity.\n"
|
"Resolve entity conflicts and duplicates.\n"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
|
"execution_count": null,
|
||||||
"metadata": {},
|
"metadata": {},
|
||||||
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"from semantica.kg import EntityResolver\n",
|
"from semantica.kg import EntityResolver\n",
|
||||||
"\n",
|
"\n",
|
||||||
"entity_resolver = EntityResolver()\n",
|
"entity_resolver = EntityResolver()\n",
|
||||||
|
"\n",
|
||||||
"resolved_entities = entity_resolver.resolve_entities(entities)\n",
|
"resolved_entities = entity_resolver.resolve_entities(entities)\n",
|
||||||
"\n",
|
"\n",
|
||||||
"canonical_id = {}\n",
|
"print(f\"Original entities: {len(entities)}\")\n",
|
||||||
"for entity in resolved_entities:\n",
|
"print(f\"Resolved entities: {len(resolved_entities)}\")"
|
||||||
" for source_id in entity.get(\"merged_from\", [entity[\"id\"]]):\n",
|
]
|
||||||
" canonical_id[source_id] = entity[\"id\"]\n",
|
|
||||||
" if entity.get(\"merged_from\"):\n",
|
|
||||||
" print(f\"Merged {entity['merged_from']} -> {entity['id']}: {entity['name']}\")\n",
|
|
||||||
"\n",
|
|
||||||
"print(f\"\\nMentions in: {len(entities)}, resolved entities out: {len(resolved_entities)}\")\n",
|
|
||||||
"\n",
|
|
||||||
"resolved_names = {entity[\"id\"]: entity[\"name\"] for entity in resolved_entities}\n",
|
|
||||||
"print(\"\\nRelationships remapped onto canonical entities:\")\n",
|
|
||||||
"for relationship in relationships:\n",
|
|
||||||
" source = canonical_id[relationship[\"source\"]]\n",
|
|
||||||
" target = canonical_id[relationship[\"target\"]]\n",
|
|
||||||
" print(f\" {resolved_names[source]} --{relationship['type']}--> {resolved_names[target]}\")\n",
|
|
||||||
"\n",
|
|
||||||
"canonical_entities = {(entity[\"name\"], entity[\"type\"]) for entity in resolved_entities}\n",
|
|
||||||
"assert canonical_entities == {\n",
|
|
||||||
" (\"Apple Inc.\", \"ORG\"),\n",
|
|
||||||
" (\"Tim Cook\", \"PERSON\"),\n",
|
|
||||||
" (\"Cupertino\", \"GPE\"),\n",
|
|
||||||
" (\"California\", \"GPE\"),\n",
|
|
||||||
"}\n",
|
|
||||||
"assert len(resolved_entities) == 4"
|
|
||||||
],
|
|
||||||
"execution_count": null,
|
|
||||||
"outputs": []
|
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"metadata": {},
|
"metadata": {},
|
||||||
"source": [
|
"source": [
|
||||||
"## Step 4: Deduplication\n",
|
"## Step 3: Deduplication\n",
|
||||||
"\n",
|
"\n",
|
||||||
"The `semantica.deduplication` module gives finer control over the same problem. Note that `merge_duplicates` returns one `MergeOperation` per duplicate *group* — the complete deduplicated collection is those merged entities plus every entity that was not part of any group.\n"
|
"Remove duplicate entities from the graph.\n"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
|
"execution_count": null,
|
||||||
"metadata": {},
|
"metadata": {},
|
||||||
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"from semantica.deduplication import DuplicateDetector, EntityMerger, MergeStrategy\n",
|
"from semantica.deduplication import DuplicateDetector, EntityMerger, MergeStrategy\n",
|
||||||
"\n",
|
"\n",
|
||||||
|
"# Detect duplicates\n",
|
||||||
"detector = DuplicateDetector(similarity_threshold=0.8)\n",
|
"detector = DuplicateDetector(similarity_threshold=0.8)\n",
|
||||||
"duplicate_groups = detector.detect_duplicate_groups(entities)\n",
|
"duplicate_groups = detector.detect_duplicate_groups(knowledge_graph.get('entities', []))\n",
|
||||||
"print(f\"Duplicate groups: {len(duplicate_groups)}\")\n",
|
|
||||||
"for group in duplicate_groups:\n",
|
|
||||||
" print(f\" {[entity['name'] for entity in group.entities]} \"\n",
|
|
||||||
" f\"(confidence={group.confidence:.2f})\")\n",
|
|
||||||
"\n",
|
"\n",
|
||||||
|
"# Merge duplicates\n",
|
||||||
"merger = EntityMerger()\n",
|
"merger = EntityMerger()\n",
|
||||||
"merge_operations = merger.merge_duplicates(\n",
|
"merge_operations = merger.merge_duplicates(\n",
|
||||||
" entities, strategy=MergeStrategy.KEEP_MOST_COMPLETE\n",
|
" knowledge_graph.get('entities', []),\n",
|
||||||
|
" strategy=MergeStrategy.KEEP_MOST_COMPLETE\n",
|
||||||
")\n",
|
")\n",
|
||||||
"\n",
|
"\n",
|
||||||
"merged_source_ids = {\n",
|
"deduplicated_entities = [op.merged_entity for op in merge_operations]\n",
|
||||||
" entity[\"id\"] for op in merge_operations for entity in op.source_entities\n",
|
|
||||||
"}\n",
|
|
||||||
"untouched_entities = [e for e in entities if e[\"id\"] not in merged_source_ids]\n",
|
|
||||||
"deduplicated_entities = untouched_entities + [\n",
|
|
||||||
" op.merged_entity for op in merge_operations\n",
|
|
||||||
"]\n",
|
|
||||||
"\n",
|
"\n",
|
||||||
"print(f\"\\nMerge operations: {len(merge_operations)}\")\n",
|
"print(f\"Original entities: {len(knowledge_graph.get('entities', []))}\")\n",
|
||||||
"print(f\"Deduplicated entities ({len(deduplicated_entities)}):\")\n",
|
"print(f\"Deduplicated entities: {len(deduplicated_entities)}\")\n"
|
||||||
"for entity in deduplicated_entities:\n",
|
]
|
||||||
" print(f\" {entity['id']}: {entity['name']} ({entity['type']})\")\n",
|
|
||||||
"\n",
|
|
||||||
"assert len(merge_operations) == 1\n",
|
|
||||||
"assert len(deduplicated_entities) == 4"
|
|
||||||
],
|
|
||||||
"execution_count": null,
|
|
||||||
"outputs": []
|
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
@@ -253,10 +155,9 @@
|
|||||||
"\n",
|
"\n",
|
||||||
"You've learned how to build knowledge graphs:\n",
|
"You've learned how to build knowledge graphs:\n",
|
||||||
"\n",
|
"\n",
|
||||||
"- **Extraction to graph**: map each mention to a graph ID and build edges from the actual `Relation.subject` / `Relation.object` endpoints\n",
|
"- **GraphBuilder**: Construct knowledge graphs from entities and relationships\n",
|
||||||
"- **GraphBuilder**: construct knowledge graphs from explicit `{\"entities\": ..., \"relationships\": ...}` input\n",
|
"- **EntityResolver**: Resolve entity conflicts and duplicates\n",
|
||||||
"- **EntityResolver**: merge duplicate mentions into canonical entities and remap relationship endpoints\n",
|
"- **Deduplication**: Use `semantica.deduplication` module for removing duplicate entities\n",
|
||||||
"- **Deduplication**: combine `MergeOperation` results with untouched entities to get the complete deduplicated set\n",
|
|
||||||
"\n",
|
"\n",
|
||||||
"Next: Learn how to analyze graphs in the Graph_Analytics notebook.\n"
|
"Next: Learn how to analyze graphs in the Graph_Analytics notebook.\n"
|
||||||
]
|
]
|
||||||
|
|||||||
@@ -10,7 +10,7 @@
|
|||||||
"\n",
|
"\n",
|
||||||
"## Overview\n",
|
"## Overview\n",
|
||||||
"\n",
|
"\n",
|
||||||
"This notebook walks you through creating your first knowledge graph from a simple document. You'll learn the complete end-to-end workflow from ingesting a file to visualizing the resulting knowledge graph — and every step consumes the real output of the step before it.\n",
|
"This notebook walks you through creating your first knowledge graph from a simple document. You'll learn the complete end-to-end workflow from ingesting a file to visualizing the resulting knowledge graph.\n",
|
||||||
"\n",
|
"\n",
|
||||||
"> [!TIP]\n",
|
"> [!TIP]\n",
|
||||||
"> This is the perfect starting point if you are new to Semantica. No prior knowledge of knowledge graphs is required!\n",
|
"> This is the perfect starting point if you are new to Semantica. No prior knowledge of knowledge graphs is required!\n",
|
||||||
@@ -19,10 +19,10 @@
|
|||||||
"\n",
|
"\n",
|
||||||
"### 🎯 Learning Objectives\n",
|
"### 🎯 Learning Objectives\n",
|
||||||
"\n",
|
"\n",
|
||||||
"- **Understand the Workflow**: Learn the `File → Parse → Extract → Graph → Visualize` pipeline\n",
|
"- **Understand the Workflow**: Learn the `File → Parse → Extract → Graph` pipeline\n",
|
||||||
"- **Ingest Data**: Load documents using `FileIngestor`\n",
|
"- **Ingest Data**: Load documents using `FileIngestor`\n",
|
||||||
"- **Parse Content**: Extract text using `DocumentParser`\n",
|
"- **Parse Content**: Extract text using `DocumentParser`\n",
|
||||||
"- **Extract Knowledge**: Identify entities and relations using `NERExtractor` and `RelationExtractor`\n",
|
"- **Extract Knowledge**: Identify entities using `NERExtractor`\n",
|
||||||
"- **Build Graph**: Construct a graph using `GraphBuilder`\n",
|
"- **Build Graph**: Construct a graph using `GraphBuilder`\n",
|
||||||
"- **Visualize**: See your graph come to life with `KGVisualizer`\n",
|
"- **Visualize**: See your graph come to life with `KGVisualizer`\n",
|
||||||
"\n",
|
"\n",
|
||||||
@@ -40,76 +40,71 @@
|
|||||||
"\n",
|
"\n",
|
||||||
"## 🔄 Simple End-to-End Workflow\n",
|
"## 🔄 Simple End-to-End Workflow\n",
|
||||||
"\n",
|
"\n",
|
||||||
"The complete workflow consists of five main steps:\n",
|
"The complete workflow consists of four main steps:\n",
|
||||||
"\n",
|
"\n",
|
||||||
"1. **📥 Ingest** - Load data from files or other sources\n",
|
"1. **📥 Ingest** - Load data from files or other sources\n",
|
||||||
"2. **📄 Parse** - Extract and structure content from documents\n",
|
"2. **📄 Parse** - Extract and structure content from documents\n",
|
||||||
"3. **⛏️ Extract** - Identify entities and relationships\n",
|
"3. **⛏️ Extract** - Identify entities and relationships\n",
|
||||||
"4. **🕸️ Build Graph** - Construct the knowledge graph\n",
|
"4. **🕸️ Build Graph** - Construct the knowledge graph\n",
|
||||||
"5. **📊 Visualize** - Render and analyze the graph\n",
|
|
||||||
"\n",
|
"\n",
|
||||||
"Each step is demonstrated in the code cells below, and each cell can be rerun on its own: the sample file is only removed by the optional cleanup cell at the very end.\n",
|
"Each step is demonstrated in the code cells below.\n",
|
||||||
"\n",
|
"\n",
|
||||||
"> [!TIP]\n",
|
"> [!TIP]\n",
|
||||||
"> **Alternative: Using Semantica Framework**\n",
|
"> **Alternative: Using Semantica Framework**\n",
|
||||||
">\n",
|
"> \n",
|
||||||
"> For a simpler, high-level approach, you can use the `Semantica` framework class which orchestrates all these steps:\n",
|
"> For a simpler, high-level approach, you can use the `Semantica` framework class which orchestrates all these steps:\n",
|
||||||
">\n",
|
"> \n",
|
||||||
"> ```python\n",
|
"> ```python\n",
|
||||||
"> from semantica.core import Semantica\n",
|
"> from semantica.core import Semantica\n",
|
||||||
">\n",
|
"> \n",
|
||||||
"> framework = Semantica()\n",
|
"> framework = Semantica()\n",
|
||||||
"> framework.initialize()\n",
|
"> framework.initialize()\n",
|
||||||
">\n",
|
"> \n",
|
||||||
"> result = framework.build_knowledge_base(\n",
|
"> result = framework.build_knowledge_base(\n",
|
||||||
"> sources=[\"sample_document.txt\"],\n",
|
"> sources=[\"sample_document.txt\"],\n",
|
||||||
"> embeddings=True,\n",
|
"> embeddings=True,\n",
|
||||||
"> graph=True\n",
|
"> graph=True\n",
|
||||||
"> )\n",
|
"> )\n",
|
||||||
">\n",
|
"> \n",
|
||||||
"> framework.shutdown()\n",
|
"> framework.shutdown()\n",
|
||||||
"> ```\n",
|
"> ```\n",
|
||||||
">\n",
|
"> \n",
|
||||||
"> This notebook shows the step-by-step approach for learning. See [Core Module Usage Guide](../../../semantica/core/core_usage.md) for more details.\n",
|
"> This notebook shows the step-by-step approach for learning. See [Core Module Usage Guide](../../../semantica/core/core_usage.md) for more details.\n",
|
||||||
"\n",
|
"\n",
|
||||||
"---\n",
|
"---\n",
|
||||||
"\n",
|
"\n",
|
||||||
"## 📂 Step 1: Ingest a File\n",
|
"## 📂 Step 1: Ingest a File\n",
|
||||||
"\n",
|
"\n",
|
||||||
"In this step, we'll use `FileIngestor` to load a document. The ingestor supports various file formats including PDF, DOCX, TXT, and more. Writing the sample file is idempotent, so this cell can be rerun at any time.\n"
|
"In this step, we'll use `FileIngestor` to load a document. The ingestor supports various file formats including PDF, DOCX, TXT, and more.\n"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"metadata": {},
|
|
||||||
"source": [
|
|
||||||
"%pip install semantica\n",
|
|
||||||
"\n",
|
|
||||||
"# spaCy models are distributed separately from the spaCy library. This lesson\n",
|
|
||||||
"# relies on the English model to recognize standalone places such as Cupertino.\n",
|
|
||||||
"import sys\n",
|
|
||||||
"import subprocess\n",
|
|
||||||
"import spacy\n",
|
|
||||||
"\n",
|
|
||||||
"try:\n",
|
|
||||||
" spacy.load(\"en_core_web_sm\")\n",
|
|
||||||
"except OSError:\n",
|
|
||||||
" subprocess.check_call([sys.executable, \"-m\", \"spacy\", \"download\", \"en_core_web_sm\"])\n"
|
|
||||||
],
|
|
||||||
"execution_count": null,
|
"execution_count": null,
|
||||||
"outputs": []
|
"metadata": {},
|
||||||
|
"outputs": [],
|
||||||
|
"source": [
|
||||||
|
"!pip install semantica"
|
||||||
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
|
"execution_count": null,
|
||||||
"metadata": {},
|
"metadata": {},
|
||||||
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
|
"from semantica.ingest import FileIngestor\n",
|
||||||
"from pathlib import Path\n",
|
"from pathlib import Path\n",
|
||||||
"\n",
|
"\n",
|
||||||
"from semantica.ingest import FileIngestor\n",
|
"# Initialize the ingestor\n",
|
||||||
|
"ingestor = FileIngestor()\n",
|
||||||
"\n",
|
"\n",
|
||||||
"sample_text = \"\"\"Apple Inc. is headquartered in Cupertino, California.\n",
|
"# Create a sample document for demonstration\n",
|
||||||
"In 1976, Steve Jobs founded Apple Inc.\n",
|
"sample_text = \"\"\"\n",
|
||||||
"Tim Cook is the CEO of Apple Inc.\n",
|
"Apple Inc. is a technology company founded by Steve Jobs, Steve Wozniak, and Ronald Wayne in 1976.\n",
|
||||||
|
"The company is headquartered in Cupertino, California.\n",
|
||||||
|
"Tim Cook is the current CEO of Apple Inc.\n",
|
||||||
|
"Apple designs and manufactures consumer electronics, software, and online services.\n",
|
||||||
"\"\"\"\n",
|
"\"\"\"\n",
|
||||||
"\n",
|
"\n",
|
||||||
"sample_file = Path(\"sample_document.txt\")\n",
|
"sample_file = Path(\"sample_document.txt\")\n",
|
||||||
@@ -118,14 +113,12 @@
|
|||||||
"print(f\"File: {sample_file}\")\n",
|
"print(f\"File: {sample_file}\")\n",
|
||||||
"print(f\"Content length: {len(sample_text)} characters\")\n",
|
"print(f\"Content length: {len(sample_text)} characters\")\n",
|
||||||
"\n",
|
"\n",
|
||||||
"ingestor = FileIngestor()\n",
|
"# Ingest the file\n",
|
||||||
"file_object = ingestor.ingest_file(sample_file, read_content=True)\n",
|
"file_object = ingestor.ingest_file(sample_file, read_content=True)\n",
|
||||||
"print(f\" File name: {file_object.name}\")\n",
|
"print(f\" File name: {file_object.name}\")\n",
|
||||||
"print(f\" File type: {file_object.file_type}\")\n",
|
"print(f\" File type: {file_object.file_type}\")\n",
|
||||||
"print(f\" Content available: {file_object.content is not None}\")"
|
"print(f\" Content available: {file_object.content is not None}\")\n"
|
||||||
],
|
]
|
||||||
"execution_count": null,
|
|
||||||
"outputs": []
|
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
@@ -133,58 +126,64 @@
|
|||||||
"source": [
|
"source": [
|
||||||
"## 📄 Step 2: Parse the Document\n",
|
"## 📄 Step 2: Parse the Document\n",
|
||||||
"\n",
|
"\n",
|
||||||
"After ingesting the file, we need to parse it to extract the text content. `DocumentParser.parse_document()` returns the extracted text under the `\"text\"` key.\n"
|
"After ingesting the file, we need to parse it to extract the text content. The `DocumentParser` handles various file formats and extracts structured content.\n"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
|
"execution_count": null,
|
||||||
"metadata": {},
|
"metadata": {},
|
||||||
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"from semantica.parse import DocumentParser\n",
|
"from semantica.parse import DocumentParser\n",
|
||||||
"\n",
|
"\n",
|
||||||
"parser = DocumentParser()\n",
|
"parser = DocumentParser()\n",
|
||||||
|
"# Parse the document to extract text\n",
|
||||||
"parsed_document = parser.parse_document(str(sample_file))\n",
|
"parsed_document = parser.parse_document(str(sample_file))\n",
|
||||||
"\n",
|
"parsed_content = parsed_document.get(\"content\", \"\")\n",
|
||||||
"parsed_content = parsed_document.get(\"text\", \"\")\n",
|
"print(f\" Parsed content length: {len(parsed_content) if parsed_content else 0} characters\")\n",
|
||||||
"assert parsed_content.strip(), \"Parsing produced no text — check the input file\"\n",
|
"print(f\" Preview: {parsed_content[:200] if parsed_content else 'N/A'}...\")"
|
||||||
"\n",
|
]
|
||||||
"print(f\"Parsed content length: {len(parsed_content)} characters\")\n",
|
|
||||||
"print(f\"Preview: {parsed_content[:120]}...\")"
|
|
||||||
],
|
|
||||||
"execution_count": null,
|
|
||||||
"outputs": []
|
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"metadata": {},
|
"metadata": {},
|
||||||
"source": [
|
"source": [
|
||||||
"## ⛏️ Step 3: Extract Entities and Relations\n",
|
"## ⛏️ Step 3: Extract Entities\n",
|
||||||
"\n",
|
"\n",
|
||||||
"Now we'll extract entities and relations from the parsed text. `NERExtractor` identifies people, organizations, locations and dates; `RelationExtractor` finds relations between those mentions. Both operate on the *parsed content from Step 2* — not on a copy of the raw string.\n"
|
"Now we'll extract entities from the parsed text using Named Entity Recognition (NER). This identifies people, organizations, locations, dates, and other entities in the text.\n",
|
||||||
|
"\n",
|
||||||
|
"> [!NOTE]\n",
|
||||||
|
"> In a real scenario, you would use `NERExtractor` with an LLM or model backend. Here we simulate the output for demonstration purposes.\n"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"metadata": {},
|
|
||||||
"source": [
|
|
||||||
"from semantica.semantic_extract import NERExtractor, RelationExtractor\n",
|
|
||||||
"\n",
|
|
||||||
"ner_extractor = NERExtractor()\n",
|
|
||||||
"relation_extractor = RelationExtractor()\n",
|
|
||||||
"\n",
|
|
||||||
"mentions = ner_extractor.extract(parsed_content)\n",
|
|
||||||
"relations = relation_extractor.extract(parsed_content, mentions)\n",
|
|
||||||
"\n",
|
|
||||||
"print(\"Entity mentions:\")\n",
|
|
||||||
"for mention in mentions:\n",
|
|
||||||
" print(f\" {mention.text!r:<13} {mention.label:<7} span=[{mention.start_char}:{mention.end_char}]\")\n",
|
|
||||||
"\n",
|
|
||||||
"print(\"\\nExtracted relations:\")\n",
|
|
||||||
"for rel in relations:\n",
|
|
||||||
" print(f\" {rel.subject.text!r} --{rel.predicate}--> {rel.object.text!r}\")"
|
|
||||||
],
|
|
||||||
"execution_count": null,
|
"execution_count": null,
|
||||||
"outputs": []
|
"metadata": {},
|
||||||
|
"outputs": [],
|
||||||
|
"source": [
|
||||||
|
"from semantica.semantic_extract import NamedEntityRecognizer, NERExtractor\n",
|
||||||
|
"\n",
|
||||||
|
"ner = NamedEntityRecognizer()\n",
|
||||||
|
"extractor = NERExtractor()\n",
|
||||||
|
"\n",
|
||||||
|
"print(f\"\\nText: {parsed_content[:100]}...\")\n",
|
||||||
|
"\n",
|
||||||
|
"# Simulated extraction results\n",
|
||||||
|
"expected_entities = [\n",
|
||||||
|
" {\"text\": \"Apple Inc.\", \"type\": \"Organization\", \"start\": 0, \"end\": 10},\n",
|
||||||
|
" {\"text\": \"Steve Jobs\", \"type\": \"Person\", \"start\": 50, \"end\": 60},\n",
|
||||||
|
" {\"text\": \"Steve Wozniak\", \"type\": \"Person\", \"start\": 62, \"end\": 75},\n",
|
||||||
|
" {\"text\": \"Ronald Wayne\", \"type\": \"Person\", \"start\": 81, \"end\": 93},\n",
|
||||||
|
" {\"text\": \"1976\", \"type\": \"Date\", \"start\": 97, \"end\": 101},\n",
|
||||||
|
" {\"text\": \"Cupertino, California\", \"type\": \"Location\", \"start\": 130, \"end\": 151},\n",
|
||||||
|
" {\"text\": \"Tim Cook\", \"type\": \"Person\", \"start\": 153, \"end\": 161},\n",
|
||||||
|
"]\n",
|
||||||
|
"\n",
|
||||||
|
"for entity in expected_entities:\n",
|
||||||
|
" print(f\" - {entity['text']} ({entity['type']})\")\n"
|
||||||
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
@@ -192,68 +191,58 @@
|
|||||||
"source": [
|
"source": [
|
||||||
"## 🕸️ Step 4: Build the Knowledge Graph\n",
|
"## 🕸️ Step 4: Build the Knowledge Graph\n",
|
||||||
"\n",
|
"\n",
|
||||||
"Using the extracted entities and relations, we construct a knowledge graph with `GraphBuilder`. Every mention gets a graph ID, and each edge is built from the actual `Relation.subject` / `Relation.object` endpoints.\n",
|
"Using the extracted entities and relationships, we'll construct a knowledge graph. The graph represents entities as nodes and relationships as edges.\n"
|
||||||
"\n",
|
|
||||||
"> [!NOTE]\n",
|
|
||||||
"> The graph will contain one node per *mention*, so `Apple Inc.` appears three times. Merging duplicate mentions into one canonical entity is covered in [07_Building_Knowledge_Graphs.ipynb](./07_Building_Knowledge_Graphs.ipynb).\n"
|
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
|
"execution_count": null,
|
||||||
"metadata": {},
|
"metadata": {},
|
||||||
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"from semantica.kg import GraphBuilder\n",
|
"from semantica.kg import GraphBuilder\n",
|
||||||
"\n",
|
"import networkx as nx\n",
|
||||||
"entities = []\n",
|
|
||||||
"span_to_id = {}\n",
|
|
||||||
"for i, mention in enumerate(mentions, 1):\n",
|
|
||||||
" graph_id = f\"e{i}\"\n",
|
|
||||||
" span_to_id[(mention.start_char, mention.end_char)] = graph_id\n",
|
|
||||||
" entities.append({\n",
|
|
||||||
" \"id\": graph_id,\n",
|
|
||||||
" \"type\": mention.label,\n",
|
|
||||||
" \"name\": mention.text,\n",
|
|
||||||
" \"properties\": {},\n",
|
|
||||||
" })\n",
|
|
||||||
"\n",
|
|
||||||
"relationships = []\n",
|
|
||||||
"for rel in relations:\n",
|
|
||||||
" source_id = span_to_id.get((rel.subject.start_char, rel.subject.end_char))\n",
|
|
||||||
" target_id = span_to_id.get((rel.object.start_char, rel.object.end_char))\n",
|
|
||||||
" if source_id is None or target_id is None:\n",
|
|
||||||
" print(f\"Skipping relation with unmapped endpoint: \"\n",
|
|
||||||
" f\"{rel.subject.text!r} --{rel.predicate}--> {rel.object.text!r}\")\n",
|
|
||||||
" continue\n",
|
|
||||||
" relationships.append({\n",
|
|
||||||
" \"source\": source_id,\n",
|
|
||||||
" \"target\": target_id,\n",
|
|
||||||
" \"type\": rel.predicate,\n",
|
|
||||||
" \"properties\": {},\n",
|
|
||||||
" })\n",
|
|
||||||
"\n",
|
"\n",
|
||||||
"builder = GraphBuilder()\n",
|
"builder = GraphBuilder()\n",
|
||||||
"knowledge_graph = builder.build({\"entities\": entities, \"relationships\": relationships})\n",
|
|
||||||
"\n",
|
"\n",
|
||||||
"id_to_name = {entity[\"id\"]: entity[\"name\"] for entity in entities}\n",
|
"# Prepare data for graph construction\n",
|
||||||
|
"entities_data = [\n",
|
||||||
|
" {\"id\": f\"entity_{i}\", \"name\": entity[\"text\"], \"type\": entity[\"type\"]}\n",
|
||||||
|
" for i, entity in enumerate(expected_entities)\n",
|
||||||
|
"]\n",
|
||||||
"\n",
|
"\n",
|
||||||
"print(f\"Nodes (entities): {len(knowledge_graph['entities'])}\")\n",
|
"relationships_data = [\n",
|
||||||
"for entity in knowledge_graph[\"entities\"]:\n",
|
" {\"source\": \"entity_0\", \"target\": \"entity_1\", \"type\": \"founded_by\"},\n",
|
||||||
" print(f\" {entity['id']}: {entity['name']} ({entity['type']})\")\n",
|
" {\"source\": \"entity_0\", \"target\": \"entity_2\", \"type\": \"founded_by\"},\n",
|
||||||
|
" {\"source\": \"entity_0\", \"target\": \"entity_3\", \"type\": \"founded_by\"},\n",
|
||||||
|
" {\"source\": \"entity_0\", \"target\": \"entity_4\", \"type\": \"founded_in\"},\n",
|
||||||
|
" {\"source\": \"entity_0\", \"target\": \"entity_5\", \"type\": \"located_in\"},\n",
|
||||||
|
" {\"source\": \"entity_6\", \"target\": \"entity_0\", \"type\": \"ceo_of\"},\n",
|
||||||
|
"]\n",
|
||||||
"\n",
|
"\n",
|
||||||
"print(f\"\\nEdges (relationships): {len(knowledge_graph['relationships'])}\")\n",
|
"# Build the graph using NetworkX\n",
|
||||||
"for relationship in knowledge_graph[\"relationships\"]:\n",
|
"kg = nx.DiGraph()\n",
|
||||||
" print(f\" {id_to_name[relationship['source']]} \"\n",
|
|
||||||
" f\"--{relationship['type']}--> {id_to_name[relationship['target']]}\")\n",
|
|
||||||
"\n",
|
"\n",
|
||||||
"edges = {\n",
|
"for entity in entities_data:\n",
|
||||||
" (id_to_name[r[\"source\"]], r[\"type\"], id_to_name[r[\"target\"]])\n",
|
" kg.add_node(entity[\"id\"], name=entity[\"name\"], type=entity[\"type\"])\n",
|
||||||
" for r in knowledge_graph[\"relationships\"]\n",
|
"\n",
|
||||||
"}\n",
|
"for rel in relationships_data:\n",
|
||||||
"assert (\"Apple Inc.\", \"located_in\", \"Cupertino\") in edges\n",
|
" source_name = entities_data[int(rel[\"source\"].split(\"_\")[1])][\"name\"]\n",
|
||||||
"assert (\"Tim Cook\", \"works_for\", \"Apple Inc.\") in edges"
|
" target_name = entities_data[int(rel[\"target\"].split(\"_\")[1])][\"name\"]\n",
|
||||||
],
|
" kg.add_edge(rel[\"source\"], rel[\"target\"], type=rel[\"type\"])\n",
|
||||||
"execution_count": null,
|
"\n",
|
||||||
"outputs": []
|
"print(f\" Nodes (entities): {len(kg.nodes)}\")\n",
|
||||||
|
"print(f\" Edges (relationships): {len(kg.edges)}\")\n",
|
||||||
|
"\n",
|
||||||
|
"for node_id in kg.nodes():\n",
|
||||||
|
" node_data = kg.nodes[node_id]\n",
|
||||||
|
" print(f\" Node: {node_data['name']} ({node_data['type']})\")\n",
|
||||||
|
"\n",
|
||||||
|
"for source, target, data in kg.edges(data=True):\n",
|
||||||
|
" source_name = kg.nodes[source]['name']\n",
|
||||||
|
" target_name = kg.nodes[target]['name']\n",
|
||||||
|
" print(f\" {source_name} --[{data['type']}]--> {target_name}\")\n"
|
||||||
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
@@ -261,81 +250,49 @@
|
|||||||
"source": [
|
"source": [
|
||||||
"## 📊 Step 5: Visualize and Analyze\n",
|
"## 📊 Step 5: Visualize and Analyze\n",
|
||||||
"\n",
|
"\n",
|
||||||
"Finally, we render the knowledge graph with `KGVisualizer` and look at its structure. `visualize_network()` accepts the `GraphBuilder` result directly and can save an interactive HTML file.\n"
|
"Finally, we'll visualize the knowledge graph and analyze its structure. This helps you understand the relationships and entities in your data.\n"
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
|
"execution_count": null,
|
||||||
"metadata": {},
|
"metadata": {},
|
||||||
|
"outputs": [],
|
||||||
"source": [
|
"source": [
|
||||||
"from semantica.visualization import KGVisualizer\n",
|
"from semantica.visualization import KGVisualizer\n",
|
||||||
"\n",
|
"\n",
|
||||||
"visualizer = KGVisualizer()\n",
|
"visualizer = KGVisualizer()\n",
|
||||||
"fig = visualizer.visualize_network(\n",
|
"\n",
|
||||||
" knowledge_graph, output=\"html\", file_path=\"knowledge_graph.html\"\n",
|
"print(f\" Total entities: {len(kg.nodes)}\")\n",
|
||||||
")\n",
|
"print(f\" Total relationships: {len(kg.edges)}\")\n",
|
||||||
"print(\"Saved interactive visualization to knowledge_graph.html\")\n",
|
|
||||||
"\n",
|
"\n",
|
||||||
"entity_types = {}\n",
|
"entity_types = {}\n",
|
||||||
"for entity in knowledge_graph[\"entities\"]:\n",
|
"for node_id in kg.nodes():\n",
|
||||||
" entity_types[entity[\"type\"]] = entity_types.get(entity[\"type\"], 0) + 1\n",
|
" entity_type = kg.nodes[node_id]['type']\n",
|
||||||
|
" entity_types[entity_type] = entity_types.get(entity_type, 0) + 1\n",
|
||||||
"\n",
|
"\n",
|
||||||
"print(\"\\nEntities by type:\")\n",
|
"for etype, count in entity_types.items():\n",
|
||||||
"for entity_type, count in sorted(entity_types.items()):\n",
|
" print(f\" - {etype}: {count}\")\n",
|
||||||
" print(f\" - {entity_type}: {count}\")\n",
|
|
||||||
"\n",
|
"\n",
|
||||||
"relationship_types = {}\n",
|
"rel_types = {}\n",
|
||||||
"for relationship in knowledge_graph[\"relationships\"]:\n",
|
"for _, _, data in kg.edges(data=True):\n",
|
||||||
" relationship_types[relationship[\"type\"]] = (\n",
|
" rel_type = data.get('type', 'unknown')\n",
|
||||||
" relationship_types.get(relationship[\"type\"], 0) + 1\n",
|
" rel_types[rel_type] = rel_types.get(rel_type, 0) + 1\n",
|
||||||
" )\n",
|
|
||||||
"\n",
|
"\n",
|
||||||
"print(\"\\nRelationships by type:\")\n",
|
"for rtype, count in rel_types.items():\n",
|
||||||
"for relationship_type, count in sorted(relationship_types.items()):\n",
|
" print(f\" - {rtype}: {count}\")\n",
|
||||||
" print(f\" - {relationship_type}: {count}\")\n",
|
|
||||||
"\n",
|
"\n",
|
||||||
"fig"
|
"# Cleanup\n",
|
||||||
],
|
"if sample_file.exists():\n",
|
||||||
"execution_count": null,
|
" sample_file.unlink()\n"
|
||||||
"outputs": []
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"cell_type": "markdown",
|
|
||||||
"metadata": {},
|
|
||||||
"source": [
|
|
||||||
"## 🧹 Optional: Clean Up\n",
|
|
||||||
"\n",
|
|
||||||
"Run this cell only when you are done with the notebook. Earlier cells read `sample_document.txt`, so they stay rerunnable until you delete it here.\n"
|
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"metadata": {},
|
|
||||||
"source": [
|
|
||||||
"for path in [sample_file, Path(\"knowledge_graph.html\")]:\n",
|
|
||||||
" if path.exists():\n",
|
|
||||||
" path.unlink()\n",
|
|
||||||
" print(f\"Removed {path}\")"
|
|
||||||
],
|
|
||||||
"execution_count": null,
|
"execution_count": null,
|
||||||
"outputs": []
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"cell_type": "markdown",
|
|
||||||
"metadata": {},
|
"metadata": {},
|
||||||
"source": [
|
"outputs": [],
|
||||||
"## Summary\n",
|
"source": []
|
||||||
"\n",
|
|
||||||
"You've built your first knowledge graph, end to end:\n",
|
|
||||||
"\n",
|
|
||||||
"- **FileIngestor** loaded the sample document\n",
|
|
||||||
"- **DocumentParser** returned its text under the `\"text\"` key\n",
|
|
||||||
"- **NERExtractor** / **RelationExtractor** produced real mentions and relations from that text\n",
|
|
||||||
"- **GraphBuilder** turned them into a graph whose edges come from the actual relation endpoints\n",
|
|
||||||
"- **KGVisualizer** rendered the result as an interactive network\n",
|
|
||||||
"\n",
|
|
||||||
"Next: merge duplicate mentions with `EntityResolver` in [07_Building_Knowledge_Graphs.ipynb](./07_Building_Knowledge_Graphs.ipynb), or explore graph metrics in the Graph Analytics notebook.\n"
|
|
||||||
]
|
|
||||||
}
|
}
|
||||||
],
|
],
|
||||||
"metadata": {
|
"metadata": {
|
||||||
|
|||||||
@@ -497,8 +497,7 @@
|
|||||||
"**Next Steps**:\n",
|
"**Next Steps**:\n",
|
||||||
"* Try customizing the `NamespaceManager` to use your organization's URL.\n",
|
"* Try customizing the `NamespaceManager` to use your organization's URL.\n",
|
||||||
"* Explore `OntologyEvaluator` for deeper quality metrics.\n",
|
"* Explore `OntologyEvaluator` for deeper quality metrics.\n",
|
||||||
"* Feed the generated ontology into the **Knowledge Graph** module to start reasoning over your data!\n",
|
"* Feed the generated ontology into the **Knowledge Graph** module to start reasoning over your data!"
|
||||||
"* Put the graph, ontology, and explicit mappings together in [Semantic Layer Basics](./26_Semantic_Layer_Basics.ipynb)."
|
|
||||||
]
|
]
|
||||||
}
|
}
|
||||||
],
|
],
|
||||||
|
|||||||
@@ -1,418 +0,0 @@
|
|||||||
{
|
|
||||||
"cells": [
|
|
||||||
{
|
|
||||||
"cell_type": "markdown",
|
|
||||||
"metadata": {},
|
|
||||||
"source": [
|
|
||||||
"[](https://colab.research.google.com/github/semantica-agi/semantica/blob/main/cookbook/introduction/26_Semantic_Layer_Basics.ipynb)\n",
|
|
||||||
"\n",
|
|
||||||
"# Semantic Layer Basics: Putting the Knowledge Graph, Ontology, and Mappings Together\n",
|
|
||||||
"\n",
|
|
||||||
"## Overview\n",
|
|
||||||
"\n",
|
|
||||||
"This lesson connects three things you have already met — a knowledge graph, an ontology, and RDF export — into one minimal *semantic layer*: a knowledge graph whose types, relationships, and properties are **explicitly mapped** to ontology terms, so the resulting RDF can be queried with SPARQL against a shared vocabulary.\n",
|
|
||||||
"\n",
|
|
||||||
"**Documentation**: [API Reference](https://semantica.readthedocs.io/concepts/)\n",
|
|
||||||
"\n",
|
|
||||||
"### 🎯 Learning Objectives\n",
|
|
||||||
"\n",
|
|
||||||
"- Build a small knowledge graph with `GraphBuilder`\n",
|
|
||||||
"- Generate a starter ontology from the graph with `OntologyGenerator`\n",
|
|
||||||
"- Write **explicit** entity-type, relationship-type, and property mappings to ontology terms\n",
|
|
||||||
"- Produce ontology-aligned RDF and store it with `TripletStore`\n",
|
|
||||||
"- Answer a business question with one small SPARQL query\n",
|
|
||||||
"\n",
|
|
||||||
"### 📚 Prerequisites\n",
|
|
||||||
"\n",
|
|
||||||
"- [07_Building_Knowledge_Graphs.ipynb](./07_Building_Knowledge_Graphs.ipynb) — graphs from entities and relationships\n",
|
|
||||||
"- [14_Ontology.ipynb](./14_Ontology.ipynb) — ontology generation\n",
|
|
||||||
"- [20_Triplet_Store.ipynb](./20_Triplet_Store.ipynb) — triplet store backends\n",
|
|
||||||
"\n",
|
|
||||||
"> [!NOTE]\n",
|
|
||||||
"> **Teaching mappings vs. governed mappings.** The mappings in this lesson are a demo: they live in a Python dict and are derived from a generated ontology. A production semantic layer uses governed identifiers, hand-designed ontologies, explicit source mappings, validation (SHACL), provenance, and versioning — that workflow is covered in [Advanced: Manual Ontology + Snowflake Mapping](../advanced/13_Manual_Ontology_Snowflake_Mapping.ipynb).\n",
|
|
||||||
"\n",
|
|
||||||
"## Installation\n",
|
|
||||||
"\n",
|
|
||||||
"The triplet-store step uses the embedded Oxigraph backend, so install with that extra. Pin at least 0.6.7: earlier releases could generate ontology classes with no URI (#1103), which silently breaks the mappings below instead of failing loudly.\n",
|
|
||||||
"\n",
|
|
||||||
"```bash\n",
|
|
||||||
"pip install \"semantica[tripletstore-oxigraph]>=0.6.7\"\n",
|
|
||||||
"```\n",
|
|
||||||
"\n",
|
|
||||||
"---\n",
|
|
||||||
"\n",
|
|
||||||
"## Step 1: Build a Knowledge Graph\n",
|
|
||||||
"\n",
|
|
||||||
"Start from a small, explicit set of entities and relationships — two people, an organization, and a project.\n"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"cell_type": "code",
|
|
||||||
"metadata": {},
|
|
||||||
"source": [
|
|
||||||
"!pip install \"semantica[tripletstore-oxigraph]>=0.6.7\"\n"
|
|
||||||
],
|
|
||||||
"execution_count": null,
|
|
||||||
"outputs": []
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"cell_type": "code",
|
|
||||||
"metadata": {},
|
|
||||||
"source": [
|
|
||||||
"from semantica.kg import GraphBuilder\n",
|
|
||||||
"\n",
|
|
||||||
"entities = [\n",
|
|
||||||
" {\"id\": \"e1\", \"type\": \"Person\", \"name\": \"Alice\", \"properties\": {\"age\": 30, \"role\": \"Engineer\"}},\n",
|
|
||||||
" {\"id\": \"e2\", \"type\": \"Person\", \"name\": \"Bob\", \"properties\": {\"age\": 35, \"role\": \"Manager\"}},\n",
|
|
||||||
" {\"id\": \"e3\", \"type\": \"Organization\", \"name\": \"Tech Corp\", \"properties\": {\"founded\": 2010}},\n",
|
|
||||||
" {\"id\": \"e4\", \"type\": \"Project\", \"name\": \"Project Alpha\", \"properties\": {\"status\": \"active\"}},\n",
|
|
||||||
"]\n",
|
|
||||||
"\n",
|
|
||||||
"relationships = [\n",
|
|
||||||
" {\"source\": \"e1\", \"target\": \"e2\", \"type\": \"reports_to\", \"properties\": {}},\n",
|
|
||||||
" {\"source\": \"e1\", \"target\": \"e3\", \"type\": \"works_for\", \"properties\": {}},\n",
|
|
||||||
" {\"source\": \"e2\", \"target\": \"e3\", \"type\": \"works_for\", \"properties\": {}},\n",
|
|
||||||
" {\"source\": \"e1\", \"target\": \"e4\", \"type\": \"works_on\", \"properties\": {}},\n",
|
|
||||||
"]\n",
|
|
||||||
"\n",
|
|
||||||
"builder = GraphBuilder()\n",
|
|
||||||
"knowledge_graph = builder.build({\"entities\": entities, \"relationships\": relationships})\n",
|
|
||||||
"\n",
|
|
||||||
"id_to_name = {entity[\"id\"]: entity[\"name\"] for entity in entities}\n",
|
|
||||||
"\n",
|
|
||||||
"print(f\"Entities ({len(knowledge_graph['entities'])}):\")\n",
|
|
||||||
"for entity in knowledge_graph[\"entities\"]:\n",
|
|
||||||
" print(f\" {entity['id']}: {entity['name']} ({entity['type']}) {entity['properties']}\")\n",
|
|
||||||
"\n",
|
|
||||||
"print(f\"\\nRelationships ({len(knowledge_graph['relationships'])}):\")\n",
|
|
||||||
"for relationship in knowledge_graph[\"relationships\"]:\n",
|
|
||||||
" print(f\" {id_to_name[relationship['source']]} \"\n",
|
|
||||||
" f\"--{relationship['type']}--> {id_to_name[relationship['target']]}\")"
|
|
||||||
],
|
|
||||||
"execution_count": null,
|
|
||||||
"outputs": []
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"cell_type": "markdown",
|
|
||||||
"metadata": {},
|
|
||||||
"source": [
|
|
||||||
"## Step 2: Generate a Starter Ontology\n",
|
|
||||||
"\n",
|
|
||||||
"`OntologyGenerator` infers OWL classes and properties from graph records. Because `GraphBuilder` keeps business attributes inside each entity's `properties` dictionary while ontology inference reads record fields, we first create a flat **inference view**. The knowledge graph itself remains unchanged. Two settings matter here:\n",
|
|
||||||
"\n",
|
|
||||||
"- `base_uri` puts every generated term in *your* namespace\n",
|
|
||||||
"- `min_occurrences=1` includes classes that occur only once (the default of 2 would drop `Organization` and `Project` from this tiny demo graph)\n",
|
|
||||||
"\n",
|
|
||||||
"Note that the generator normalizes names: the relationship type `works_for` becomes the ontology property `worksFor`. That is exactly why the next step maps terms **explicitly** instead of matching names.\n"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"cell_type": "code",
|
|
||||||
"metadata": {},
|
|
||||||
"source": [
|
|
||||||
"from semantica.ontology import OntologyGenerator\n",
|
|
||||||
"\n",
|
|
||||||
"BASE_URI = \"https://example.org/company/\"\n",
|
|
||||||
"\n",
|
|
||||||
"# Adapt the property-graph representation to the record shape consumed by\n",
|
|
||||||
"# OntologyGenerator, so age/role/founded/status become declared properties.\n",
|
|
||||||
"ontology_input = {\n",
|
|
||||||
" \"entities\": [\n",
|
|
||||||
" {\n",
|
|
||||||
" **{key: value for key, value in entity.items() if key != \"properties\"},\n",
|
|
||||||
" **entity.get(\"properties\", {}),\n",
|
|
||||||
" }\n",
|
|
||||||
" for entity in knowledge_graph[\"entities\"]\n",
|
|
||||||
" ],\n",
|
|
||||||
" \"relationships\": knowledge_graph[\"relationships\"],\n",
|
|
||||||
"}\n",
|
|
||||||
"\n",
|
|
||||||
"generator = OntologyGenerator(base_uri=BASE_URI, min_occurrences=1)\n",
|
|
||||||
"ontology = generator.generate_from_graph(ontology_input)\n",
|
|
||||||
"\n",
|
|
||||||
"# OntologyGenerator calls datatype properties `data`; TripletStore's public\n",
|
|
||||||
"# ontology contract calls them `datatype`. Normalize that boundary explicitly.\n",
|
|
||||||
"store_ontology = {\n",
|
|
||||||
" **ontology,\n",
|
|
||||||
" \"properties\": [\n",
|
|
||||||
" {**prop, \"type\": \"datatype\" if prop[\"type\"] == \"data\" else prop[\"type\"]}\n",
|
|
||||||
" for prop in ontology[\"properties\"]\n",
|
|
||||||
" ],\n",
|
|
||||||
"}\n",
|
|
||||||
"\n",
|
|
||||||
"print(\"Classes:\")\n",
|
|
||||||
"for ontology_class in ontology[\"classes\"]:\n",
|
|
||||||
" print(f\" {ontology_class['name']:<14} {ontology_class['uri']}\")\n",
|
|
||||||
"\n",
|
|
||||||
"print(\"\\nProperties:\")\n",
|
|
||||||
"for prop in ontology[\"properties\"]:\n",
|
|
||||||
" print(f\" {prop['name']:<14} {prop['type']:<7} {prop['uri']} \"\n",
|
|
||||||
" f\"(domain={prop['domain']}, range={prop['range']})\")\n",
|
|
||||||
"\n",
|
|
||||||
"assert len(ontology[\"classes\"]) == 3"
|
|
||||||
],
|
|
||||||
"execution_count": null,
|
|
||||||
"outputs": []
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"cell_type": "markdown",
|
|
||||||
"metadata": {},
|
|
||||||
"source": [
|
|
||||||
"## Step 3: Map the Graph to Ontology Terms\n",
|
|
||||||
"\n",
|
|
||||||
"The heart of a semantic layer is the mapping contract: which source type, relationship, and property corresponds to which ontology term.\n",
|
|
||||||
"\n",
|
|
||||||
"- **Entity types** and **relationship types**: each generated class/property records the source name it was inferred from (`metadata[\"inferred_from\"]`), so the mapping is read off the ontology itself — no fragile name matching between `works_for` and `worksFor`.\n",
|
|
||||||
"- **Properties**: the flat inference view makes `name`, `age`, `role`, `founded`, and `status` real generated datatype properties. Every mapping therefore points to a term declared in the ontology — no URI is invented only at mapping time.\n"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"cell_type": "code",
|
|
||||||
"metadata": {},
|
|
||||||
"source": [
|
|
||||||
"entity_type_mappings = {\n",
|
|
||||||
" ontology_class[\"metadata\"][\"inferred_from\"]: ontology_class[\"uri\"]\n",
|
|
||||||
" for ontology_class in ontology[\"classes\"]\n",
|
|
||||||
"}\n",
|
|
||||||
"\n",
|
|
||||||
"relationship_type_mappings = {\n",
|
|
||||||
" prop[\"metadata\"][\"inferred_from\"]: prop[\"uri\"]\n",
|
|
||||||
" for prop in ontology[\"properties\"]\n",
|
|
||||||
" if prop[\"type\"] == \"object\"\n",
|
|
||||||
"}\n",
|
|
||||||
"\n",
|
|
||||||
"datatype_property_uris = {\n",
|
|
||||||
" prop[\"metadata\"][\"inferred_from\"]: prop[\"uri\"]\n",
|
|
||||||
" for prop in ontology[\"properties\"]\n",
|
|
||||||
" if prop[\"type\"] != \"object\"\n",
|
|
||||||
"}\n",
|
|
||||||
"\n",
|
|
||||||
"property_mappings = datatype_property_uris\n",
|
|
||||||
"\n",
|
|
||||||
"semantic_layer = {\n",
|
|
||||||
" \"graph\": knowledge_graph,\n",
|
|
||||||
" \"ontology\": ontology,\n",
|
|
||||||
" \"mappings\": {\n",
|
|
||||||
" \"entity_type_mappings\": entity_type_mappings,\n",
|
|
||||||
" \"relationship_type_mappings\": relationship_type_mappings,\n",
|
|
||||||
" \"property_mappings\": property_mappings,\n",
|
|
||||||
" },\n",
|
|
||||||
"}\n",
|
|
||||||
"\n",
|
|
||||||
"for mapping_name, mapping in semantic_layer[\"mappings\"].items():\n",
|
|
||||||
" print(f\"{mapping_name}:\")\n",
|
|
||||||
" for source, target in mapping.items():\n",
|
|
||||||
" print(f\" {source:<12} -> {target}\")\n",
|
|
||||||
"\n",
|
|
||||||
"# Every type and relationship in the graph must have an ontology term\n",
|
|
||||||
"assert set(entity_type_mappings) == {entity[\"type\"] for entity in entities}\n",
|
|
||||||
"assert set(relationship_type_mappings) == {rel[\"type\"] for rel in relationships}\n",
|
|
||||||
"assert set(property_mappings) == {\"name\", \"age\", \"role\", \"founded\", \"status\"}\n",
|
|
||||||
"assert set(property_mappings.values()) <= {prop[\"uri\"] for prop in ontology[\"properties\"]}"
|
|
||||||
],
|
|
||||||
"execution_count": null,
|
|
||||||
"outputs": []
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"cell_type": "markdown",
|
|
||||||
"metadata": {},
|
|
||||||
"source": [
|
|
||||||
"## Step 4: Apply the Mappings\n",
|
|
||||||
"\n",
|
|
||||||
"Applying the semantic layer means rewriting the graph so every type, relationship, and property key is an ontology term. This *aligned* graph — not the original one — is what gets exported and stored.\n"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"cell_type": "code",
|
|
||||||
"metadata": {},
|
|
||||||
"source": [
|
|
||||||
"aligned_graph = {\n",
|
|
||||||
" \"entities\": [\n",
|
|
||||||
" {\n",
|
|
||||||
" **entity,\n",
|
|
||||||
" \"type\": entity_type_mappings[entity[\"type\"]],\n",
|
|
||||||
" \"properties\": {\n",
|
|
||||||
" property_mappings[\"name\"]: entity[\"name\"],\n",
|
|
||||||
" **{\n",
|
|
||||||
" property_mappings[key]: value\n",
|
|
||||||
" for key, value in entity[\"properties\"].items()\n",
|
|
||||||
" },\n",
|
|
||||||
" },\n",
|
|
||||||
" }\n",
|
|
||||||
" for entity in knowledge_graph[\"entities\"]\n",
|
|
||||||
" ],\n",
|
|
||||||
" \"relationships\": [\n",
|
|
||||||
" {**rel, \"type\": relationship_type_mappings[rel[\"type\"]]}\n",
|
|
||||||
" for rel in knowledge_graph[\"relationships\"]\n",
|
|
||||||
" ],\n",
|
|
||||||
"}\n",
|
|
||||||
"\n",
|
|
||||||
"print(\"Aligned entity sample:\")\n",
|
|
||||||
"sample = aligned_graph[\"entities\"][0]\n",
|
|
||||||
"print(f\" id: {sample['id']}\")\n",
|
|
||||||
"print(f\" type: {sample['type']}\")\n",
|
|
||||||
"for key, value in sample[\"properties\"].items():\n",
|
|
||||||
" print(f\" {key} = {value}\")\n",
|
|
||||||
"\n",
|
|
||||||
"print(\"\\nAligned relationship sample:\")\n",
|
|
||||||
"print(f\" {aligned_graph['relationships'][0]['type']}\")"
|
|
||||||
],
|
|
||||||
"execution_count": null,
|
|
||||||
"outputs": []
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"cell_type": "markdown",
|
|
||||||
"metadata": {},
|
|
||||||
"source": [
|
|
||||||
"## Step 5: Store and Export Complete Ontology-Aligned RDF\n",
|
|
||||||
"\n",
|
|
||||||
"`TripletStore.store()` materializes both the ontology declarations and the aligned instance graph. We then read those triples through the store's public API and serialize that complete RDF graph as Turtle. This avoids the compact `RDFExporter` entity projection, which does not include arbitrary entries from an entity's `properties` dictionary.\n"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"cell_type": "code",
|
|
||||||
"metadata": {},
|
|
||||||
"source": [
|
|
||||||
"from rdflib import Graph, Literal, URIRef\n",
|
|
||||||
"from rdflib.namespace import OWL, RDF\n",
|
|
||||||
"from semantica.triplet_store import TripletStore\n",
|
|
||||||
"\n",
|
|
||||||
"store = TripletStore(backend=\"oxigraph\")\n",
|
|
||||||
"result = store.store(aligned_graph, store_ontology)\n",
|
|
||||||
"print(f\"Stored triples: {result['processed']} (failed: {result['failed']})\")\n",
|
|
||||||
"\n",
|
|
||||||
"rdf_graph = Graph()\n",
|
|
||||||
"for triplet in store.get_triplets():\n",
|
|
||||||
" datatype = triplet.metadata.get(\"datatype\")\n",
|
|
||||||
" if datatype:\n",
|
|
||||||
" object_term = Literal(triplet.object, datatype=URIRef(datatype))\n",
|
|
||||||
" elif triplet.object.startswith((\"http://\", \"https://\", \"urn:\")):\n",
|
|
||||||
" object_term = URIRef(triplet.object)\n",
|
|
||||||
" else:\n",
|
|
||||||
" object_term = Literal(triplet.object)\n",
|
|
||||||
" rdf_graph.add((URIRef(triplet.subject), URIRef(triplet.predicate), object_term))\n",
|
|
||||||
"\n",
|
|
||||||
"rdf_graph.serialize(destination=\"semantic_layer.ttl\", format=\"turtle\")\n",
|
|
||||||
"turtle = open(\"semantic_layer.ttl\", encoding=\"utf-8\").read()\n",
|
|
||||||
"print(turtle[:600])\n",
|
|
||||||
"\n",
|
|
||||||
"# The exported RDF contains declarations plus mapped instance facts.\n",
|
|
||||||
"declared_datatype_properties = {\n",
|
|
||||||
" str(subject) for subject in rdf_graph.subjects(RDF.type, OWL.DatatypeProperty)\n",
|
|
||||||
"}\n",
|
|
||||||
"assert result[\"failed\"] == 0\n",
|
|
||||||
"assert set(property_mappings.values()) <= declared_datatype_properties\n",
|
|
||||||
"assert (\n",
|
|
||||||
" URIRef(BASE_URI + \"e1\"),\n",
|
|
||||||
" URIRef(property_mappings[\"role\"]),\n",
|
|
||||||
" Literal(\"Engineer\"),\n",
|
|
||||||
") in rdf_graph\n",
|
|
||||||
"assert (\n",
|
|
||||||
" URIRef(BASE_URI + \"e1\"),\n",
|
|
||||||
" URIRef(relationship_type_mappings[\"works_for\"]),\n",
|
|
||||||
" URIRef(BASE_URI + \"e3\"),\n",
|
|
||||||
") in rdf_graph\n",
|
|
||||||
"print(\"... exported semantic_layer.ttl\")"
|
|
||||||
],
|
|
||||||
"execution_count": null,
|
|
||||||
"outputs": []
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"cell_type": "markdown",
|
|
||||||
"metadata": {},
|
|
||||||
"source": [
|
|
||||||
"## Step 6: Query the Semantic Layer\n",
|
|
||||||
"\n",
|
|
||||||
"The embedded Oxigraph backend runs in memory, so there is nothing to start beyond installing the `tripletstore-oxigraph` extra. The organization is constrained by its mapped `name` predicate; the query therefore means *Tech Corp*, rather than accidentally matching employees of every organization.\n"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"cell_type": "code",
|
|
||||||
"metadata": {},
|
|
||||||
"source": [
|
|
||||||
"query = f\"\"\"\n",
|
|
||||||
"SELECT ?name ?role WHERE {{\n",
|
|
||||||
" ?person <{BASE_URI}worksFor> ?org .\n",
|
|
||||||
" ?org <{BASE_URI}name> \"Tech Corp\" .\n",
|
|
||||||
" ?person <{BASE_URI}name> ?name .\n",
|
|
||||||
" ?person <{BASE_URI}role> ?role .\n",
|
|
||||||
"}}\n",
|
|
||||||
"ORDER BY ?name\n",
|
|
||||||
"\"\"\"\n",
|
|
||||||
"query_result = store.execute_query(query)\n",
|
|
||||||
"\n",
|
|
||||||
"print(\"\\nWho works for Tech Corp, and in which role?\")\n",
|
|
||||||
"for binding in query_result.bindings:\n",
|
|
||||||
" print(f\" {binding['name']['value']} — {binding['role']['value']}\")\n",
|
|
||||||
"\n",
|
|
||||||
"assert [(row[\"name\"][\"value\"], row[\"role\"][\"value\"]) for row in query_result.bindings] == [\n",
|
|
||||||
" (\"Alice\", \"Engineer\"),\n",
|
|
||||||
" (\"Bob\", \"Manager\"),\n",
|
|
||||||
"]"
|
|
||||||
],
|
|
||||||
"execution_count": null,
|
|
||||||
"outputs": []
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"cell_type": "markdown",
|
|
||||||
"metadata": {},
|
|
||||||
"source": [
|
|
||||||
"## 🧹 Optional: Clean Up\n"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"cell_type": "code",
|
|
||||||
"metadata": {},
|
|
||||||
"source": [
|
|
||||||
"from pathlib import Path\n",
|
|
||||||
"\n",
|
|
||||||
"ttl_file = Path(\"semantic_layer.ttl\")\n",
|
|
||||||
"if ttl_file.exists():\n",
|
|
||||||
" ttl_file.unlink()\n",
|
|
||||||
" print(f\"Removed {ttl_file}\")"
|
|
||||||
],
|
|
||||||
"execution_count": null,
|
|
||||||
"outputs": []
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"cell_type": "markdown",
|
|
||||||
"metadata": {},
|
|
||||||
"source": [
|
|
||||||
"## Summary\n",
|
|
||||||
"\n",
|
|
||||||
"A minimal semantic layer is a composition, and you have now built each part:\n",
|
|
||||||
"\n",
|
|
||||||
"1. **Knowledge graph** — `GraphBuilder` from explicit entities and relationships\n",
|
|
||||||
"2. **Ontology** — `OntologyGenerator` with your `base_uri`\n",
|
|
||||||
"3. **Explicit mappings** — entity types, relationship types, and properties, each tied to an ontology term\n",
|
|
||||||
"4. **Ontology-aligned RDF** — the mappings applied to the graph, materialized with `TripletStore`, and serialized to Turtle from the store's own triples\n",
|
|
||||||
"5. **Queryable store** — `TripletStore` (embedded Oxigraph) answering a SPARQL question over the shared vocabulary\n",
|
|
||||||
"\n",
|
|
||||||
"### Where to go next\n",
|
|
||||||
"\n",
|
|
||||||
"The production version of this workflow — hand-designed governed ontologies, explicit source-to-ontology mappings from a warehouse, n-ary modeling, SHACL validation, provenance, and versioning — is covered in [Advanced: Manual Ontology + Snowflake Mapping](../advanced/13_Manual_Ontology_Snowflake_Mapping.ipynb).\n"
|
|
||||||
]
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"metadata": {
|
|
||||||
"kernelspec": {
|
|
||||||
"display_name": "Python 3",
|
|
||||||
"language": "python",
|
|
||||||
"name": "python3"
|
|
||||||
},
|
|
||||||
"language_info": {
|
|
||||||
"codemirror_mode": {
|
|
||||||
"name": "ipython",
|
|
||||||
"version": 3
|
|
||||||
},
|
|
||||||
"file_extension": ".py",
|
|
||||||
"mimetype": "text/x-python",
|
|
||||||
"name": "python",
|
|
||||||
"nbconvert_exporter": "python",
|
|
||||||
"pygments_lexer": "ipython3",
|
|
||||||
"version": "3.11.9"
|
|
||||||
}
|
|
||||||
},
|
|
||||||
"nbformat": 4,
|
|
||||||
"nbformat_minor": 2
|
|
||||||
}
|
|
||||||
@@ -185,7 +185,7 @@ Centralized `ConfigManager` with environment variable overrides. No magic defaul
|
|||||||
| **Deduplication v2** | `blocking_v2`, `hybrid_v2`, `semantic_v2`: up to 7x faster than v1 |
|
| **Deduplication v2** | `blocking_v2`, `hybrid_v2`, `semantic_v2`: up to 7x faster than v1 |
|
||||||
| **Indexed search** | Explorer search at 0.004ms on 118k nodes (v0.5.0) |
|
| **Indexed search** | Explorer search at 0.004ms on 118k nodes (v0.5.0) |
|
||||||
|
|
||||||
- [Modules](/modules) — Full module documentation with code examples.
|
- [Modules](modules) — Full module documentation with code examples.
|
||||||
- [Learning More](/learning-more) — Configuration reference, performance guide, and troubleshooting.
|
- [Learning More](learning-more) — Configuration reference, performance guide, and troubleshooting.
|
||||||
- [Pipeline Reference](/reference/pipeline) — Pipeline orchestration, workers, and retry policies.
|
- [Pipeline Reference](reference/pipeline) — Pipeline orchestration, workers, and retry policies.
|
||||||
- [Core Reference](/reference/core) — Framework lifecycle, plugin registry, and configuration.
|
- [Core Reference](reference/core) — Framework lifecycle, plugin registry, and configuration.
|
||||||
|
|||||||
+183
-4
@@ -1,9 +1,14 @@
|
|||||||
/* ============================================================
|
/* ============================================================
|
||||||
SEMANTICA DOCS — DESIGN SYSTEM
|
SEMANTICA DOCS — PREMIUM DESIGN SYSTEM
|
||||||
Dark-first (#080C10 bg, #10B981 emerald accent)
|
Dark-first (#080C10 bg, #10B981 emerald accent)
|
||||||
Minimal, static styling — no decorative motion.
|
|
||||||
============================================================ */
|
============================================================ */
|
||||||
|
|
||||||
|
/* ── Keyframes ─────────────────────────────────────────────── */
|
||||||
|
@keyframes pageFadeIn {
|
||||||
|
from { opacity: 0; transform: translateY(6px); }
|
||||||
|
to { opacity: 1; transform: translateY(0); }
|
||||||
|
}
|
||||||
|
|
||||||
/* ── Global ─────────────────────────────────────────────────── */
|
/* ── Global ─────────────────────────────────────────────────── */
|
||||||
html {
|
html {
|
||||||
scroll-behavior: smooth;
|
scroll-behavior: smooth;
|
||||||
@@ -24,7 +29,16 @@ html {
|
|||||||
}
|
}
|
||||||
::-webkit-scrollbar-thumb:hover { background: rgba(16, 185, 129, 0.4); }
|
::-webkit-scrollbar-thumb:hover { background: rgba(16, 185, 129, 0.4); }
|
||||||
|
|
||||||
/* ── Focus rings (accessibility — kept) ─────────────────────── */
|
/* ── Page entrance ──────────────────────────────────────────── */
|
||||||
|
main,
|
||||||
|
article,
|
||||||
|
[class*="content-area"],
|
||||||
|
[class*="ContentArea"],
|
||||||
|
[class*="prose"] {
|
||||||
|
animation: pageFadeIn 0.35s ease both;
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ── Focus rings ─────────────────────────────────────────────── */
|
||||||
*:focus-visible {
|
*:focus-visible {
|
||||||
outline: 2px solid rgba(16, 185, 129, 0.55) !important;
|
outline: 2px solid rgba(16, 185, 129, 0.55) !important;
|
||||||
outline-offset: 3px !important;
|
outline-offset: 3px !important;
|
||||||
@@ -45,7 +59,7 @@ h1::after {
|
|||||||
left: 0;
|
left: 0;
|
||||||
width: 44px;
|
width: 44px;
|
||||||
height: 2px;
|
height: 2px;
|
||||||
background: #10B981;
|
background: linear-gradient(90deg, #10B981 0%, transparent 100%);
|
||||||
border-radius: 1px;
|
border-radius: 1px;
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -57,6 +71,9 @@ article a,
|
|||||||
[class*="prose"] a {
|
[class*="prose"] a {
|
||||||
text-decoration-color: rgba(16, 185, 129, 0.35);
|
text-decoration-color: rgba(16, 185, 129, 0.35);
|
||||||
text-underline-offset: 3px;
|
text-underline-offset: 3px;
|
||||||
|
transition:
|
||||||
|
text-decoration-color 0.15s ease,
|
||||||
|
color 0.15s ease;
|
||||||
}
|
}
|
||||||
|
|
||||||
article a:hover,
|
article a:hover,
|
||||||
@@ -72,6 +89,14 @@ blockquote {
|
|||||||
padding: 0.9rem 1.2rem !important;
|
padding: 0.9rem 1.2rem !important;
|
||||||
font-style: italic;
|
font-style: italic;
|
||||||
color: rgba(255, 255, 255, 0.68) !important;
|
color: rgba(255, 255, 255, 0.68) !important;
|
||||||
|
transition:
|
||||||
|
border-color 0.2s ease,
|
||||||
|
background-color 0.2s ease !important;
|
||||||
|
}
|
||||||
|
|
||||||
|
blockquote:hover {
|
||||||
|
border-left-color: rgba(16, 185, 129, 0.65) !important;
|
||||||
|
background: rgba(16, 185, 129, 0.07) !important;
|
||||||
}
|
}
|
||||||
|
|
||||||
/* ── HR / Divider ────────────────────────────────────────────── */
|
/* ── HR / Divider ────────────────────────────────────────────── */
|
||||||
@@ -98,11 +123,165 @@ table thead th {
|
|||||||
border-bottom: 1px solid rgba(16, 185, 129, 0.18) !important;
|
border-bottom: 1px solid rgba(16, 185, 129, 0.18) !important;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
table tbody tr {
|
||||||
|
transition: background-color 0.15s ease;
|
||||||
|
cursor: default;
|
||||||
|
}
|
||||||
|
|
||||||
|
table tbody tr:hover {
|
||||||
|
background-color: rgba(16, 185, 129, 0.06) !important;
|
||||||
|
}
|
||||||
|
|
||||||
|
table tbody tr:hover td {
|
||||||
|
background-color: transparent !important;
|
||||||
|
}
|
||||||
|
|
||||||
|
table td,
|
||||||
|
table th {
|
||||||
|
transition: background-color 0.15s ease;
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ── CODE BLOCKS ─────────────────────────────────────────────── */
|
||||||
|
pre,
|
||||||
|
[class*="codeblock"],
|
||||||
|
[class*="code-group"],
|
||||||
|
[class*="CodeBlock"],
|
||||||
|
[data-rehype-pretty-code-fragment] {
|
||||||
|
transition:
|
||||||
|
box-shadow 0.25s cubic-bezier(0.4, 0, 0.2, 1),
|
||||||
|
border-color 0.25s cubic-bezier(0.4, 0, 0.2, 1),
|
||||||
|
transform 0.25s cubic-bezier(0.4, 0, 0.2, 1) !important;
|
||||||
|
}
|
||||||
|
|
||||||
|
pre:hover,
|
||||||
|
[class*="codeblock"]:hover,
|
||||||
|
[class*="CodeBlock"]:hover,
|
||||||
|
[data-rehype-pretty-code-fragment]:hover {
|
||||||
|
transform: translateY(-1px) !important;
|
||||||
|
box-shadow:
|
||||||
|
0 0 0 1px rgba(16, 185, 129, 0.18),
|
||||||
|
0 2px 12px rgba(16, 185, 129, 0.06),
|
||||||
|
0 8px 32px rgba(0, 0, 0, 0.2) !important;
|
||||||
|
border-color: rgba(16, 185, 129, 0.2) !important;
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ── CARDS ───────────────────────────────────────────────────── */
|
||||||
|
[class*="card"],
|
||||||
|
[class*="Card"],
|
||||||
|
[data-card],
|
||||||
|
.group\/card {
|
||||||
|
transition:
|
||||||
|
transform 0.22s ease,
|
||||||
|
box-shadow 0.22s ease,
|
||||||
|
border-color 0.22s ease !important;
|
||||||
|
}
|
||||||
|
|
||||||
|
[class*="card"]:hover,
|
||||||
|
[class*="Card"]:hover,
|
||||||
|
[data-card]:hover,
|
||||||
|
.group\/card:hover {
|
||||||
|
transform: translateY(-3px) !important;
|
||||||
|
box-shadow:
|
||||||
|
0 8px 28px rgba(0, 0, 0, 0.18),
|
||||||
|
0 0 0 1px rgba(16, 185, 129, 0.22) !important;
|
||||||
|
border-color: rgba(16, 185, 129, 0.28) !important;
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ── CALLOUTS / ADMONITIONS ──────────────────────────────────── */
|
||||||
|
[class*="callout"],
|
||||||
|
[class*="Callout"],
|
||||||
|
[class*="admonition"] {
|
||||||
|
transition:
|
||||||
|
box-shadow 0.2s ease,
|
||||||
|
border-color 0.2s ease !important;
|
||||||
|
}
|
||||||
|
|
||||||
|
[class*="callout"]:hover,
|
||||||
|
[class*="Callout"]:hover,
|
||||||
|
[class*="admonition"]:hover {
|
||||||
|
box-shadow: 0 2px 16px rgba(16, 185, 129, 0.08) !important;
|
||||||
|
border-color: rgba(16, 185, 129, 0.35) !important;
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ── STEPS ───────────────────────────────────────────────────── */
|
||||||
|
[class*="step"],
|
||||||
|
[class*="Step"] {
|
||||||
|
transition: background-color 0.15s ease !important;
|
||||||
|
}
|
||||||
|
|
||||||
|
[class*="step"]:hover,
|
||||||
|
[class*="Step"]:hover {
|
||||||
|
background-color: rgba(16, 185, 129, 0.04) !important;
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ── INLINE CODE ─────────────────────────────────────────────── */
|
||||||
|
:not(pre) > code {
|
||||||
|
transition:
|
||||||
|
background-color 0.15s ease,
|
||||||
|
color 0.15s ease !important;
|
||||||
|
cursor: text;
|
||||||
|
}
|
||||||
|
|
||||||
|
:not(pre) > code:hover {
|
||||||
|
background-color: rgba(16, 185, 129, 0.16) !important;
|
||||||
|
}
|
||||||
|
|
||||||
/* ── NAVIGATION / SIDEBAR ────────────────────────────────────── */
|
/* ── NAVIGATION / SIDEBAR ────────────────────────────────────── */
|
||||||
nav a,
|
nav a,
|
||||||
[class*="sidebar"] a,
|
[class*="sidebar"] a,
|
||||||
[class*="Sidebar"] a {
|
[class*="Sidebar"] a {
|
||||||
|
transition: color 0.15s ease !important;
|
||||||
text-decoration: none;
|
text-decoration: none;
|
||||||
|
position: relative;
|
||||||
|
}
|
||||||
|
|
||||||
|
nav a::after,
|
||||||
|
[class*="sidebar"] a::after,
|
||||||
|
[class*="Sidebar"] a::after {
|
||||||
|
content: "";
|
||||||
|
position: absolute;
|
||||||
|
bottom: -1px;
|
||||||
|
left: 0;
|
||||||
|
width: 0;
|
||||||
|
height: 1px;
|
||||||
|
background: #10B981;
|
||||||
|
transition: width 0.2s ease;
|
||||||
|
}
|
||||||
|
|
||||||
|
nav a:hover::after,
|
||||||
|
[class*="sidebar"] a:hover::after,
|
||||||
|
[class*="Sidebar"] a:hover::after {
|
||||||
|
width: 100%;
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ── TEXT / LIST ITEMS ───────────────────────────────────────── */
|
||||||
|
ul > li,
|
||||||
|
ol > li {
|
||||||
|
border-radius: 3px;
|
||||||
|
transition: background-color 0.12s ease;
|
||||||
|
}
|
||||||
|
|
||||||
|
ul > li:hover,
|
||||||
|
ol > li:hover {
|
||||||
|
background-color: rgba(16, 185, 129, 0.04);
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ── PRIMARY BUTTON / CTA ────────────────────────────────────── */
|
||||||
|
button[class*="primary"],
|
||||||
|
a[class*="primary"],
|
||||||
|
[class*="btn-primary"],
|
||||||
|
[class*="ButtonPrimary"] {
|
||||||
|
transition:
|
||||||
|
box-shadow 0.2s ease,
|
||||||
|
transform 0.2s ease !important;
|
||||||
|
}
|
||||||
|
|
||||||
|
button[class*="primary"]:hover,
|
||||||
|
a[class*="primary"]:hover,
|
||||||
|
[class*="btn-primary"]:hover,
|
||||||
|
[class*="ButtonPrimary"]:hover {
|
||||||
|
box-shadow: 0 0 22px rgba(16, 185, 129, 0.28) !important;
|
||||||
|
transform: translateY(-1px) !important;
|
||||||
}
|
}
|
||||||
|
|
||||||
/* ── HIDE THEME TOGGLE ───────────────────────────────────────── */
|
/* ── HIDE THEME TOGGLE ───────────────────────────────────────── */
|
||||||
|
|||||||
+14
-14
@@ -5,7 +5,7 @@ icon: "compass"
|
|||||||
---
|
---
|
||||||
|
|
||||||
<Info>
|
<Info>
|
||||||
Every module works independently — import only what you need. This page maps developer goals to starting points. The [Module Reference](/modules) covers every module in depth.
|
Every module works independently — import only what you need. This page maps developer goals to starting points. The [Module Reference](modules) covers every module in depth.
|
||||||
</Info>
|
</Info>
|
||||||
|
|
||||||
## Quick Reference
|
## Quick Reference
|
||||||
@@ -89,7 +89,7 @@ Pick your goal to see the minimum imports and a working skeleton.
|
|||||||
Pass `method="pattern"` to `NERExtractor` for zero-cost, zero-API-key extraction. Switch to `method="llm"` with any of the supported providers for higher recall.
|
Pass `method="pattern"` to `NERExtractor` for zero-cost, zero-API-key extraction. Switch to `method="llm"` with any of the supported providers for higher recall.
|
||||||
</Tip>
|
</Tip>
|
||||||
|
|
||||||
**Next:** [Quickstart →](/quickstart) — full pipeline with visualization and export.
|
**Next:** [Quickstart →](quickstart) — full pipeline with visualization and export.
|
||||||
</Tab>
|
</Tab>
|
||||||
|
|
||||||
<Tab title="Build GraphRAG">
|
<Tab title="Build GraphRAG">
|
||||||
@@ -122,7 +122,7 @@ Pick your goal to see the minimum imports and a working skeleton.
|
|||||||
print(result["reasoning_path"]) # multi-hop trace
|
print(result["reasoning_path"]) # multi-hop trace
|
||||||
```
|
```
|
||||||
|
|
||||||
**Next:** [Context module reference →](/reference/context)
|
**Next:** [Context module reference →](reference/context)
|
||||||
</Tab>
|
</Tab>
|
||||||
|
|
||||||
<Tab title="Add Agent Memory">
|
<Tab title="Add Agent Memory">
|
||||||
@@ -163,7 +163,7 @@ Pick your goal to see the minimum imports and a working skeleton.
|
|||||||
`decision_tracking=True` is required. Without it, `record_decision()` raises `RuntimeError`.
|
`decision_tracking=True` is required. Without it, `record_decision()` raises `RuntimeError`.
|
||||||
</Note>
|
</Note>
|
||||||
|
|
||||||
**Next:** [Context module reference →](/reference/context)
|
**Next:** [Context module reference →](reference/context)
|
||||||
</Tab>
|
</Tab>
|
||||||
|
|
||||||
<Tab title="Track Provenance">
|
<Tab title="Track Provenance">
|
||||||
@@ -195,7 +195,7 @@ Pick your goal to see the minimum imports and a working skeleton.
|
|||||||
diff = manager.diff("v1.0", "v1.1")
|
diff = manager.diff("v1.0", "v1.1")
|
||||||
```
|
```
|
||||||
|
|
||||||
**Next:** [Provenance reference →](/reference/provenance) · [Change Management reference →](/reference/change_management)
|
**Next:** [Provenance reference →](reference/provenance) · [Change Management reference →](reference/change_management)
|
||||||
</Tab>
|
</Tab>
|
||||||
|
|
||||||
<Tab title="Export">
|
<Tab title="Export">
|
||||||
@@ -222,11 +222,11 @@ Pick your goal to see the minimum imports and a working skeleton.
|
|||||||
|
|
||||||
**Formats:** Turtle · JSON-LD · N-Triples · RDF/XML · Parquet · Cypher · Arrow · OWL · CSV · ArangoDB AQL
|
**Formats:** Turtle · JSON-LD · N-Triples · RDF/XML · Parquet · Cypher · Arrow · OWL · CSV · ArangoDB AQL
|
||||||
|
|
||||||
**Next:** [Export module reference →](/reference/export)
|
**Next:** [Export module reference →](reference/export)
|
||||||
</Tab>
|
</Tab>
|
||||||
|
|
||||||
<Tab title="MCP — Claude / Cursor">
|
<Tab title="MCP — Claude / Cursor">
|
||||||
Use Semantica from Claude Desktop, Cursor, VS Code, or any MCP-aware tool — no Python code required after setup. 15 tools available instantly.
|
Use Semantica from Claude Desktop, Cursor, VS Code, or any MCP-aware tool — no Python code required after setup. 12 tools available instantly.
|
||||||
|
|
||||||
**Step 1 — Install:**
|
**Step 1 — Install:**
|
||||||
```bash
|
```bash
|
||||||
@@ -268,7 +268,7 @@ Pick your goal to see the minimum imports and a working skeleton.
|
|||||||
Set `SEMANTICA_KG_PATH` to persist your graph across restarts. Without it, all data is lost when the server process exits.
|
Set `SEMANTICA_KG_PATH` to persist your graph across restarts. Without it, all data is lost when the server process exits.
|
||||||
</Warning>
|
</Warning>
|
||||||
|
|
||||||
**Next:** [MCP Server reference →](/reference/mcp_server)
|
**Next:** [MCP Server reference →](reference/mcp_server)
|
||||||
</Tab>
|
</Tab>
|
||||||
</Tabs>
|
</Tabs>
|
||||||
|
|
||||||
@@ -283,11 +283,11 @@ Pick your goal to see the minimum imports and a working skeleton.
|
|||||||
|
|
||||||
Use **both together** via `AgentContext` (GraphRAG) to get grounded LLM responses where every claim traces back to a source node.
|
Use **both together** via `AgentContext` (GraphRAG) to get grounded LLM responses where every claim traces back to a source node.
|
||||||
|
|
||||||
See also: [Core Concepts](/concepts)
|
See also: [Core Concepts](concepts)
|
||||||
</Accordion>
|
</Accordion>
|
||||||
|
|
||||||
<Accordion title="I just want to run something quickly." icon="rocket">
|
<Accordion title="I just want to run something quickly." icon="rocket">
|
||||||
Start with the [Quickstart](/quickstart). It builds a complete pipeline (ingest → parse → extract → graph → visualize → export) with no API key required.
|
Start with the [Quickstart](quickstart). It builds a complete pipeline (ingest → parse → extract → graph → visualize → export) with no API key required.
|
||||||
</Accordion>
|
</Accordion>
|
||||||
|
|
||||||
<Accordion title="I'm adding Semantica to an existing agent — what's the minimum?" icon="plug">
|
<Accordion title="I'm adding Semantica to an existing agent — what's the minimum?" icon="plug">
|
||||||
@@ -304,7 +304,7 @@ Pick your goal to see the minimum imports and a working skeleton.
|
|||||||
)
|
)
|
||||||
```
|
```
|
||||||
|
|
||||||
[Context module reference →](/reference/context)
|
[Context module reference →](reference/context)
|
||||||
</Accordion>
|
</Accordion>
|
||||||
|
|
||||||
<Accordion title="I need a compliance-ready pipeline — what's the minimum stack?" icon="shield-check">
|
<Accordion title="I need a compliance-ready pipeline — what's the minimum stack?" icon="shield-check">
|
||||||
@@ -322,6 +322,6 @@ Pick your goal to see the minimum imports and a working skeleton.
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
- [Quickstart](/quickstart) — Full pipeline in 5 minutes.
|
- [Quickstart](quickstart) — Full pipeline in 5 minutes.
|
||||||
- [Module Reference](/modules) — Every module with examples and common chains.
|
- [Module Reference](modules) — Every module with examples and common chains.
|
||||||
- [API Reference](/reference/context) — Complete class and method documentation.
|
- [API Reference](reference/context) — Complete class and method documentation.
|
||||||
|
|||||||
+2
-2
@@ -48,5 +48,5 @@ Published research using Semantica? [Let us know](https://github.com/semantica-a
|
|||||||
|
|
||||||
## See Also
|
## See Also
|
||||||
|
|
||||||
- [License](/project-license) — MIT License details.
|
- [License](project-license) — MIT License details.
|
||||||
- [Community](/community) — Connect with the Semantica community.
|
- [Community](community) — Connect with the Semantica community.
|
||||||
|
|||||||
+9
-9
@@ -24,7 +24,7 @@ After installation the following commands are available:
|
|||||||
| `semantica-mcp` | `semantica.mcp_server:main` | MCP server (stdio) for Claude Desktop, Cursor, Windsurf, and other MCP clients |
|
| `semantica-mcp` | `semantica.mcp_server:main` | MCP server (stdio) for Claude Desktop, Cursor, Windsurf, and other MCP clients |
|
||||||
|
|
||||||
<Note>
|
<Note>
|
||||||
`semantica-explorer` requires `pip install semantica[explorer]`. Running it without that extra will immediately print an error and exit. See [Explorer Setup](/explorer-setup) for the full walkthrough.
|
`semantica-explorer` requires `pip install semantica[explorer]`. Running it without that extra will immediately print an error and exit. See [Explorer Setup](explorer-setup) for the full walkthrough.
|
||||||
</Note>
|
</Note>
|
||||||
|
|
||||||
|
|
||||||
@@ -52,8 +52,8 @@ python -c "import semantica; print(semantica.__version__)"
|
|||||||
- **semantica** — The general-purpose CLI. Use it for one-off pipeline runs, entity extraction, and graph operations from a shell script or CI job.
|
- **semantica** — The general-purpose CLI. Use it for one-off pipeline runs, entity extraction, and graph operations from a shell script or CI job.
|
||||||
- **semantica-server** — Starts the REST API server. Binds to `0.0.0.0:8000`. Use this when another service or application needs programmatic access to Semantica over HTTP.
|
- **semantica-server** — Starts the REST API server. Binds to `0.0.0.0:8000`. Use this when another service or application needs programmatic access to Semantica over HTTP.
|
||||||
- **semantica-worker** — Background task processor. Run alongside `semantica-server` when you need async pipeline execution outside the request cycle. Start the server first, then start one or more workers pointing at the same backend.
|
- **semantica-worker** — Background task processor. Run alongside `semantica-server` when you need async pipeline execution outside the request cycle. Start the server first, then start one or more workers pointing at the same backend.
|
||||||
- **semantica-explorer** — Launches the browser dashboard. Requires `pip install semantica[explorer]`. Use this to explore a saved knowledge graph interactively. See [Explorer Setup](/explorer-setup).
|
- **semantica-explorer** — Launches the browser dashboard. Requires `pip install semantica[explorer]`. Use this to explore a saved knowledge graph interactively. See [Explorer Setup](explorer-setup).
|
||||||
- **semantica-mcp** — Runs the MCP server over stdio. Configure it in your MCP client's settings file to expose all 15 tools and 3 resources to Claude Desktop, Cursor, Windsurf, or any MCP-aware client. See [MCP Server](/reference/mcp_server).
|
- **semantica-mcp** — Runs the MCP server over stdio. Configure it in your MCP client's settings file to expose all 12 tools and 3 resources to Claude Desktop, Cursor, Windsurf, or any MCP-aware client. See [MCP Server](reference/mcp_server).
|
||||||
|
|
||||||
|
|
||||||
## Usage Examples
|
## Usage Examples
|
||||||
@@ -116,7 +116,7 @@ python -c "import semantica; print(semantica.__version__)"
|
|||||||
echo '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"test","version":"1.0"}}}' | semantica-mcp
|
echo '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"test","version":"1.0"}}}' | semantica-mcp
|
||||||
```
|
```
|
||||||
|
|
||||||
You should receive a JSON-RPC response. See [MCP Server](/reference/mcp_server) for the full list of tools and resources.
|
You should receive a JSON-RPC response. See [MCP Server](reference/mcp_server) for the full list of tools and resources.
|
||||||
</Tab>
|
</Tab>
|
||||||
<Tab title="Explorer">
|
<Tab title="Explorer">
|
||||||
```bash
|
```bash
|
||||||
@@ -124,7 +124,7 @@ python -c "import semantica; print(semantica.__version__)"
|
|||||||
semantica-explorer --graph my_graph.json
|
semantica-explorer --graph my_graph.json
|
||||||
```
|
```
|
||||||
|
|
||||||
See [Explorer Setup](/explorer-setup) for the full walkthrough including how to build and save a graph file.
|
See [Explorer Setup](explorer-setup) for the full walkthrough including how to build and save a graph file.
|
||||||
</Tab>
|
</Tab>
|
||||||
<Tab title="Python module form">
|
<Tab title="Python module form">
|
||||||
Every command also runs as a Python module: useful when the script directory is not on `PATH`:
|
Every command also runs as a Python module: useful when the script directory is not on `PATH`:
|
||||||
@@ -228,7 +228,7 @@ Install the [Microsoft Visual C++ Redistributable](https://aka.ms/vs/17/release/
|
|||||||
|
|
||||||
## Next Steps
|
## Next Steps
|
||||||
|
|
||||||
- [Explorer Setup](/explorer-setup) — Build a graph, save it, and launch the browser dashboard.
|
- [Explorer Setup](explorer-setup) — Build a graph, save it, and launch the browser dashboard.
|
||||||
- [MCP Server](/reference/mcp_server) — All 15 tools and 3 resources exposed over the MCP protocol.
|
- [MCP Server](reference/mcp_server) — All 12 tools and 3 resources exposed over the MCP protocol.
|
||||||
- [Installation](/installation) — Virtual environments, optional extras, and platform-specific notes.
|
- [Installation](installation) — Virtual environments, optional extras, and platform-specific notes.
|
||||||
- [Quickstart](/quickstart) — End-to-end pipeline walkthrough with working code.
|
- [Quickstart](quickstart) — End-to-end pipeline walkthrough with working code.
|
||||||
|
|||||||
@@ -109,12 +109,12 @@ def my_ingestor(source):
|
|||||||
method_registry.register("file", "my_format", my_ingestor)
|
method_registry.register("file", "my_format", my_ingestor)
|
||||||
```
|
```
|
||||||
|
|
||||||
See [Architecture](/architecture#extension-points) for the full extension guide.
|
See [Architecture](architecture#extension-points) for the full extension guide.
|
||||||
|
|
||||||
|
|
||||||
## How to Contribute
|
## How to Contribute
|
||||||
|
|
||||||
- [Contributing Guide](/contributing-guide) — Submit code, documentation, tests, or cookbook notebooks.
|
- [Contributing Guide](contributing-guide) — Submit code, documentation, tests, or cookbook notebooks.
|
||||||
- [GitHub Issues](https://github.com/semantica-agi/semantica/issues) — Report bugs, request features, or propose integrations.
|
- [GitHub Issues](https://github.com/semantica-agi/semantica/issues) — Report bugs, request features, or propose integrations.
|
||||||
- [Discord](https://discord.gg/sV34vps5hH) — Share what you're building with the community.
|
- [Discord](https://discord.gg/sV34vps5hH) — Share what you're building with the community.
|
||||||
- [GitHub Discussions](https://github.com/semantica-agi/semantica/discussions) — Long-form questions, design discussions, and ideas.
|
- [GitHub Discussions](https://github.com/semantica-agi/semantica/discussions) — Long-form questions, design discussions, and ideas.
|
||||||
|
|||||||
+5
-5
@@ -55,7 +55,7 @@ There's no single right way to contribute. Pick the path that fits your skills a
|
|||||||
- Review open pull requests
|
- Review open pull requests
|
||||||
- Share your Semantica projects in GitHub Discussions
|
- Share your Semantica projects in GitHub Discussions
|
||||||
|
|
||||||
See the [Contributing Guide](/contributing-guide) for the full development workflow.
|
See the [Contributing Guide](contributing-guide) for the full development workflow.
|
||||||
|
|
||||||
|
|
||||||
## Stay Connected
|
## Stay Connected
|
||||||
@@ -68,7 +68,7 @@ See the [Contributing Guide](/contributing-guide) for the full development workf
|
|||||||
|
|
||||||
## See Also
|
## See Also
|
||||||
|
|
||||||
- [Contributing Guide](/contributing-guide) — Step-by-step guide for submitting PRs and setting up your dev environment.
|
- [Contributing Guide](contributing-guide) — Step-by-step guide for submitting PRs and setting up your dev environment.
|
||||||
- [Community Projects](/community-projects) — Projects and integrations built by the community.
|
- [Community Projects](community-projects) — Projects and integrations built by the community.
|
||||||
- [FAQ](/faq) — Common questions answered.
|
- [FAQ](faq) — Common questions answered.
|
||||||
- [Governance](/governance) — How the project is run and decisions are made.
|
- [Governance](governance) — How the project is run and decisions are made.
|
||||||
|
|||||||
+87
-118
@@ -5,7 +5,7 @@ icon: "book-open"
|
|||||||
---
|
---
|
||||||
|
|
||||||
<Info>
|
<Info>
|
||||||
New here? Start with [Getting Started](/getting-started) for hands-on examples, then return here for deeper understanding.
|
New here? Start with [Getting Started](getting-started) for hands-on examples, then return here for deeper understanding.
|
||||||
</Info>
|
</Info>
|
||||||
|
|
||||||
Semantica transforms unstructured data: documents, web pages, reports, databases: into **knowledge graphs**: structured representations that AI systems can query, reason about, and trace back to sources.
|
Semantica transforms unstructured data: documents, web pages, reports, databases: into **knowledge graphs**: structured representations that AI systems can query, reason about, and trace back to sources.
|
||||||
@@ -38,19 +38,18 @@ This structure makes knowledge **searchable**, **connectable**, **queryable**, a
|
|||||||
Scanning text to find and classify real-world entities:
|
Scanning text to find and classify real-world entities:
|
||||||
|
|
||||||
```python
|
```python
|
||||||
# "Apple Inc. was founded by Steve Jobs in 1976 in Cupertino."
|
# Input: "Apple Inc. was founded by Steve Jobs in 1976 in Cupertino."
|
||||||
[
|
{
|
||||||
Entity(text="Apple Inc.", label="ORG", start_char=0, end_char=10, confidence=0.98),
|
"entities": [
|
||||||
Entity(text="Steve Jobs", label="PERSON", start_char=25, end_char=35, confidence=0.99),
|
{"text": "Apple Inc.", "type": "ORGANIZATION", "confidence": 0.98},
|
||||||
Entity(text="1976", label="DATE", start_char=39, end_char=43, confidence=0.95),
|
{"text": "Steve Jobs", "type": "PERSON", "confidence": 0.99},
|
||||||
Entity(text="Cupertino", label="GPE", start_char=47, end_char=56, confidence=0.97),
|
{"text": "1976", "type": "DATE", "confidence": 0.95},
|
||||||
]
|
{"text": "Cupertino", "type": "LOCATION", "confidence": 0.97}
|
||||||
|
]
|
||||||
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
`NERExtractor(method=...).extract(text)` returns a list of `Entity` objects, each
|
Each entity gets a type, confidence score, and a link to its source document. Three extraction methods are available:
|
||||||
with a `label`, character offsets (`start_char` / `end_char`), a `confidence`
|
|
||||||
score, and a `metadata` dict recording the extraction method. Three methods are
|
|
||||||
available:
|
|
||||||
|
|
||||||
| Method | Speed | Accuracy | Requirements |
|
| Method | Speed | Accuracy | Requirements |
|
||||||
| :------ | :----- | :-------- | :------------ |
|
| :------ | :----- | :-------- | :------------ |
|
||||||
@@ -63,19 +62,15 @@ available:
|
|||||||
Finding how entities connect to each other:
|
Finding how entities connect to each other:
|
||||||
|
|
||||||
```python
|
```python
|
||||||
jobs = Entity(text="Steve Jobs", label="PERSON", start_char=25, end_char=35)
|
{
|
||||||
apple = Entity(text="Apple Inc.", label="ORG", start_char=0, end_char=10)
|
"relationships": [
|
||||||
|
{"subject": "Steve Jobs", "predicate": "founded", "object": "Apple Inc.", "confidence": 0.92},
|
||||||
[
|
{"subject": "Apple Inc.", "predicate": "located_in", "object": "Cupertino", "confidence": 0.89}
|
||||||
Relation(subject=jobs, predicate="founded", object=apple, confidence=0.92),
|
]
|
||||||
Relation(subject=apple, predicate="located_in", object=Entity(text="Cupertino", label="GPE", start_char=47, end_char=56), confidence=0.89),
|
}
|
||||||
]
|
|
||||||
```
|
```
|
||||||
|
|
||||||
`RelationExtractor(method=...).extract(text, entities=entities)` returns a list of
|
Relationships can be extracted via rule-based methods, ML models, or LLMs: each producing typed triplets with confidence scores and source attribution.
|
||||||
`Relation` objects: typed subject-predicate-object triples (the endpoints are
|
|
||||||
`Entity` objects) with confidence scores and source attribution. Extraction runs
|
|
||||||
via pattern rules, ML models, or LLMs.
|
|
||||||
|
|
||||||
|
|
||||||
## Knowledge Graph vs. Vector Store
|
## Knowledge Graph vs. Vector Store
|
||||||
@@ -99,10 +94,9 @@ Both store information for AI retrieval: but they're built for different jobs.
|
|||||||
```python
|
```python
|
||||||
from semantica.kg import GraphBuilder, PathFinder
|
from semantica.kg import GraphBuilder, PathFinder
|
||||||
|
|
||||||
graph = GraphBuilder(merge_entities=True).build(
|
graph = GraphBuilder(merge_entities=True).build(entities=entities, relationships=rels)
|
||||||
{"entities": entities, "relationships": rels}
|
finder = PathFinder()
|
||||||
)
|
path = finder.dijkstra_shortest_path(graph, "Steve Jobs", "Tim Cook")
|
||||||
path = PathFinder().dijkstra_shortest_path(graph, "Steve Jobs", "Tim Cook")
|
|
||||||
```
|
```
|
||||||
</Tab>
|
</Tab>
|
||||||
|
|
||||||
@@ -146,16 +140,8 @@ Both store information for AI retrieval: but they're built for different jobs.
|
|||||||
context = AgentContext(
|
context = AgentContext(
|
||||||
vector_store=VectorStore(backend="faiss", dimension=768),
|
vector_store=VectorStore(backend="faiss", dimension=768),
|
||||||
knowledge_graph=ContextGraph(advanced_analytics=True),
|
knowledge_graph=ContextGraph(advanced_analytics=True),
|
||||||
graph_expansion=True,
|
|
||||||
)
|
)
|
||||||
|
result = context.query("Who founded Apple?", mode="graphrag")
|
||||||
# store() extracts entities and populates the graph + vector index
|
|
||||||
context.store([{"content": "Steve Jobs co-founded Apple Inc. in 1976."}])
|
|
||||||
|
|
||||||
# retrieve() blends vector similarity with graph traversal
|
|
||||||
results = context.retrieve("Who founded Apple?", use_graph=True, expand_graph=True)
|
|
||||||
for r in results:
|
|
||||||
print(r["score"], r["content"], r["source"])
|
|
||||||
```
|
```
|
||||||
</Tab>
|
</Tab>
|
||||||
</Tabs>
|
</Tabs>
|
||||||
@@ -217,7 +203,7 @@ ontology = {
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
Semantica can auto-generate ontologies from your knowledge graph or import existing OWL/RDF/Turtle ontologies. The **Ontology Hub** (v0.5.0) adds a visual editor, SHACL Studio, alignment authoring, and a live health dashboard. See the [Ontology reference](/reference/ontology) for the full 6-stage generation pipeline.
|
Semantica can auto-generate ontologies from your knowledge graph or import existing OWL/RDF/Turtle ontologies. The **Ontology Hub** (v0.5.0) adds a visual editor, SHACL Studio, alignment authoring, and a live health dashboard. See the [Ontology reference](reference/ontology) for the full 6-stage generation pipeline.
|
||||||
|
|
||||||
|
|
||||||
## Reasoning & Inference
|
## Reasoning & Inference
|
||||||
@@ -235,80 +221,70 @@ Inferred: Steve Jobs has a connection to Cupertino
|
|||||||
Applies IF/THEN rules repeatedly until no new facts can be derived. Best for alert systems, compliance checks, and trigger-based workflows.
|
Applies IF/THEN rules repeatedly until no new facts can be derived. Best for alert systems, compliance checks, and trigger-based workflows.
|
||||||
|
|
||||||
```python
|
```python
|
||||||
from semantica.reasoning import Reasoner
|
from semantica.reasoning import Reasoner, Rule, Fact, RuleType
|
||||||
|
|
||||||
engine = Reasoner()
|
engine = Reasoner()
|
||||||
engine.add_fact("Manager(Alice)")
|
engine.add_fact(Fact(subject="Alice", predicate="is_a", obj="Manager"))
|
||||||
engine.add_rule("IF Manager(?x) THEN HasAuthority(?x)")
|
engine.add_rule(Rule(
|
||||||
|
rule_type=RuleType.FORWARD_CHAIN,
|
||||||
results = engine.forward_chain() # list of InferenceResult
|
conditions=[{"subject": "?x", "predicate": "is_a", "object": "Manager"}],
|
||||||
for r in results:
|
conclusion={"subject": "?x", "predicate": "has_authority", "object": "true"}
|
||||||
print(r.conclusion) # "HasAuthority(Alice)"
|
))
|
||||||
|
result = engine.infer()
|
||||||
```
|
```
|
||||||
</Tab>
|
</Tab>
|
||||||
<Tab title="Rete Network">
|
<Tab title="Rete Network">
|
||||||
Efficient pattern matching for large rule sets: the Rete algorithm avoids re-evaluating rules whose preconditions haven't changed. Best for thousands of rules over millions of facts.
|
Efficient pattern matching for large rule sets: the Rete algorithm avoids re-evaluating rules whose preconditions haven't changed. Best for thousands of rules over millions of facts.
|
||||||
|
|
||||||
```python
|
```python
|
||||||
from semantica.reasoning import ReteEngine, Rule, Fact
|
from semantica.reasoning import ReteEngine
|
||||||
|
|
||||||
engine = ReteEngine()
|
engine = ReteEngine()
|
||||||
engine.build_network([
|
engine.load_rules("rules/domain_rules.json")
|
||||||
Rule(rule_id="r1", name="manager_authority",
|
results = engine.run(kg)
|
||||||
conditions=["Manager(?x)"], conclusion="HasAuthority(?x)"),
|
|
||||||
])
|
|
||||||
engine.add_fact(Fact(fact_id="f1", predicate="Manager", arguments=["Alice"]))
|
|
||||||
|
|
||||||
matches = engine.match_patterns()
|
|
||||||
results = engine.execute_matches(matches) # ["HasAuthority(?x)"]
|
|
||||||
```
|
```
|
||||||
</Tab>
|
</Tab>
|
||||||
<Tab title="LLM Reasoning">
|
<Tab title="Deductive & Abductive">
|
||||||
`GraphReasoner` answers open-ended questions over a knowledge graph with an
|
**Deductive**: classical syllogistic reasoning from premises to guaranteed conclusions.
|
||||||
LLM, returning a natural-language answer grounded in the graph's facts. Best
|
|
||||||
for exploratory and investigative questions that fixed rules can't anticipate.
|
**Abductive**: infers the most likely explanation for observed evidence. Best for diagnostic and investigative use cases.
|
||||||
|
|
||||||
```python
|
```python
|
||||||
from semantica.reasoning import GraphReasoner
|
from semantica.reasoning import GraphReasoner
|
||||||
|
|
||||||
reasoner = GraphReasoner(provider="openai", model="gpt-4o-mini")
|
graph_reasoner = GraphReasoner(kg)
|
||||||
answer = reasoner.reason(kg, "Which suppliers are indirectly exposed to the Acme outage?")
|
graph_reasoner.add_rule({"if": [{"subject": "?a", "predicate": "parent_of", "object": "?b"}], "then": {"subject": "?a", "predicate": "ancestor_of", "object": "?b"}})
|
||||||
|
inferences = graph_reasoner.infer(kg)
|
||||||
```
|
```
|
||||||
</Tab>
|
</Tab>
|
||||||
<Tab title="Datalog (v0.4.0)">
|
<Tab title="Datalog (v0.4.0)">
|
||||||
Recursive Horn clause rules with fixpoint semantics: handles transitive closure and recursive relationships that forward chaining cannot express.
|
Recursive Horn clause rules with fixpoint semantics: handles transitive closure and recursive relationships that forward chaining cannot express.
|
||||||
|
|
||||||
```python
|
```python
|
||||||
from semantica.reasoning import DatalogReasoner
|
from semantica.reasoning import DatalogReasoner, DatalogFact, DatalogRule
|
||||||
|
|
||||||
reasoner = DatalogReasoner()
|
reasoner = DatalogReasoner()
|
||||||
reasoner.add_fact("parent(alice, bob)")
|
reasoner.add_fact(DatalogFact("parent", ("alice", "bob")))
|
||||||
reasoner.add_fact("parent(bob, charlie)")
|
reasoner.add_rule(DatalogRule("ancestor(?X, ?Y) :- parent(?X, ?Y)."))
|
||||||
reasoner.add_rule("ancestor(X, Y) :- parent(X, Y).")
|
reasoner.evaluate()
|
||||||
reasoner.add_rule("ancestor(X, Z) :- parent(X, Y), ancestor(Y, Z).")
|
results = reasoner.query("ancestor(alice, ?Z)")
|
||||||
|
|
||||||
reasoner.derive_all()
|
|
||||||
results = reasoner.query("ancestor(alice, ?Z)") # {"Z": "bob"} and {"Z": "charlie"}, order not guaranteed
|
|
||||||
```
|
```
|
||||||
</Tab>
|
</Tab>
|
||||||
<Tab title="Engine Comparison">
|
<Tab title="Engine Comparison">
|
||||||
|
|
||||||
| Engine | Class | Best For |
|
| Engine | Description | Best For |
|
||||||
| :------ | :----- | :-------- |
|
| :------ | :----------- | :-------- |
|
||||||
| Forward chaining | `Reasoner` | Alert systems, compliance checks |
|
| Forward chaining | Applies rules until fixpoint | Alert systems, compliance checks |
|
||||||
| Rete network | `ReteEngine` | Large rule sets, high fact throughput |
|
| Rete network | Efficient pattern matching | Large rule sets, high fact throughput |
|
||||||
| SPARQL expansion | `SPARQLReasoner` | Semantic web, ontology reasoning over RDF |
|
| Deductive | Classical syllogistic reasoning | Mathematical and logical inference |
|
||||||
| Datalog (v0.4.0) | `DatalogReasoner` | Transitive closure, graph reachability |
|
| Abductive | Most likely explanation | Diagnostics, investigation |
|
||||||
| Temporal | `TemporalReasoningEngine` | Allen interval algebra, time-aware inference |
|
| SPARQL | Query-based inference over RDF | Semantic web, ontology reasoning |
|
||||||
| LLM over the graph | `GraphReasoner` | Open-ended, investigative questions |
|
| Datalog (v0.4.0) | Recursive Horn clause rules | Transitive closure, graph reachability |
|
||||||
|
|
||||||
</Tab>
|
</Tab>
|
||||||
</Tabs>
|
</Tabs>
|
||||||
|
|
||||||
`Reasoner.forward_chain()` returns `InferenceResult` objects that carry the rule
|
All engines produce **explainable inference paths**: not black-box conclusions. Every derived fact includes the rules and premises that produced it.
|
||||||
applied (`rule_used`) and the premises it fired on, and `ExplanationGenerator`
|
|
||||||
turns one into a step-by-step natural-language justification: reasoning here is
|
|
||||||
**not** a black box.
|
|
||||||
|
|
||||||
|
|
||||||
## Temporal Intelligence
|
## Temporal Intelligence
|
||||||
@@ -337,18 +313,13 @@ Explore the semantic neighborhood of any entity in your graph: useful for unders
|
|||||||
```python
|
```python
|
||||||
from semantica.kg import SimilarityCalculator
|
from semantica.kg import SimilarityCalculator
|
||||||
|
|
||||||
calc = SimilarityCalculator(method="cosine") # "cosine" | "euclidean" | "manhattan" | "correlation"
|
calc = SimilarityCalculator()
|
||||||
|
scores = calc.calculate_similarity(entity_a, entity_b)
|
||||||
# Similarity for every unique pair of node embeddings: {(node_a, node_b): score}
|
|
||||||
pairs = calc.pairwise_similarity({"apple": vec_apple, "google": vec_google, "nest": vec_nest})
|
|
||||||
|
|
||||||
# Or rank a set of embeddings by closeness to one query vector
|
|
||||||
nearest = calc.find_most_similar(embeddings, query_embedding, top_k=10)
|
|
||||||
```
|
```
|
||||||
|
|
||||||
**Features:** N×N semantic distance matrices, ego-mode visualization, distance band classification (`direct` / `near` / `mid-range` / `distant`), embedding cache optimization for large graphs.
|
**Features:** N×N semantic distance matrices, ego-mode visualization, distance band classification (`near` / `mid` / `far`), embedding cache optimization for large graphs.
|
||||||
|
|
||||||
The [Visualization module](/reference/visualization) renders distance matrices as interactive heatmaps and ego-mode neighborhood graphs. The [Explorer](/reference/explorer) embeds distance intelligence directly in the browser dashboard.
|
The [Visualization module](reference/visualization) renders distance matrices as interactive heatmaps and ego-mode neighborhood graphs. The [Explorer](reference/explorer) embeds distance intelligence directly in the browser dashboard.
|
||||||
|
|
||||||
|
|
||||||
## Deduplication & Entity Resolution
|
## Deduplication & Entity Resolution
|
||||||
@@ -370,11 +341,11 @@ Real-world data contains the same entity under many names: "Apple", "Apple Inc."
|
|||||||
```python
|
```python
|
||||||
from semantica.deduplication import DuplicateDetector, EntityMerger
|
from semantica.deduplication import DuplicateDetector, EntityMerger
|
||||||
|
|
||||||
detector = DuplicateDetector(similarity_threshold=0.85)
|
detector = DuplicateDetector(similarity_threshold=0.85)
|
||||||
candidates = detector.detect_duplicates(entities)
|
duplicates = detector.detect_duplicates(entities)
|
||||||
|
|
||||||
merger = EntityMerger()
|
merger = EntityMerger()
|
||||||
operations = merger.merge_duplicates(entities, strategy="keep_most_complete")
|
deduplicated_entities = merger.merge_duplicates(entities)
|
||||||
```
|
```
|
||||||
</Tab>
|
</Tab>
|
||||||
</Tabs>
|
</Tabs>
|
||||||
@@ -390,21 +361,19 @@ Every fact in Semantica links back to:
|
|||||||
- The **reasoning steps** that produced any inferred fact
|
- The **reasoning steps** that produced any inferred fact
|
||||||
|
|
||||||
<Note>
|
<Note>
|
||||||
This is W3C PROV-O compliant lineage: suitable for regulated industries that require audit trails (HIPAA, SOX, GDPR, FDA 21 CFR Part 11). `ProvenanceManager.export_prov(format="turtle")` serialises the recorded lineage as PROV-O RDF.
|
This is W3C PROV-O compliant lineage: suitable for regulated industries that require audit trails (HIPAA, SOX, GDPR, FDA 21 CFR Part 11). Use `RDFExporter(include_provenance=True)` to embed provenance inline in any RDF export.
|
||||||
</Note>
|
</Note>
|
||||||
|
|
||||||
```python
|
```python
|
||||||
from semantica.provenance import ProvenanceManager
|
from semantica.provenance import ProvenanceManager
|
||||||
|
|
||||||
prov = ProvenanceManager()
|
prov = ProvenanceManager()
|
||||||
prov.track_entity("apple_inc", source="report.pdf",
|
lineage = prov.get_entity_lineage("apple_inc")
|
||||||
metadata={"extractor": "NamedEntityRecognizer", "confidence": 0.98})
|
|
||||||
|
|
||||||
record = prov.get_provenance("apple_inc") # dict; use get_lineage() for the full chain
|
print(f"Source: {lineage.source_document}")
|
||||||
print(record["source_document"])
|
print(f"Method: {lineage.extraction_method}")
|
||||||
print(record["timestamp"])
|
print(f"Extracted: {lineage.timestamp}")
|
||||||
print(record["checksum"])
|
print(f"Checksum: {lineage.checksum}")
|
||||||
print(record["metadata"]) # extractor, confidence, and any custom keys
|
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|
||||||
@@ -444,7 +413,7 @@ When multiple sources disagree on the same fact, Semantica flags and resolves th
|
|||||||
- **Majority vote**: aggregate across all sources with ≥ 2 agreeing
|
- **Majority vote**: aggregate across all sources with ≥ 2 agreeing
|
||||||
- **Manual review**: flag for human arbitration; continue pipeline without blocking
|
- **Manual review**: flag for human arbitration; continue pipeline without blocking
|
||||||
|
|
||||||
See the [Conflicts reference](/reference/conflicts) for `ConflictResolver`, `SourceTracker`, and `InvestigationGuideGenerator`.
|
See the [Conflicts reference](reference/conflicts) for `ConflictResolver`, `SourceTracker`, and `InvestigationGuideGenerator`.
|
||||||
|
|
||||||
|
|
||||||
## Custom Plugin Development
|
## Custom Plugin Development
|
||||||
@@ -487,32 +456,32 @@ Semantica is designed for extension. Any component: ingestor, extractor, graph b
|
|||||||
**Extension points available:** ingestors, parsers, normalizers, extractors, reasoning engines, export formats, vector store backends, graph store backends, visualization renderers.
|
**Extension points available:** ingestors, parsers, normalizers, extractors, reasoning engines, export formats, vector store backends, graph store backends, visualization renderers.
|
||||||
|
|
||||||
</Accordion>
|
</Accordion>
|
||||||
<Accordion title="MethodRegistry: swap a built-in graph operation for your own">
|
<Accordion title="MethodRegistry: add domain-specific graph operations">
|
||||||
|
|
||||||
`method_registry` lets you register an alternative implementation for a
|
`MethodRegistry` lets you register custom methods on knowledge graph objects by name: useful for adding domain-specific graph operations without subclassing.
|
||||||
knowledge-graph task (`build`, `analyze`, `centrality`, `resolve`, …) under a
|
|
||||||
name, then select it wherever that task runs.
|
|
||||||
|
|
||||||
```python
|
```python
|
||||||
from semantica.kg import method_registry
|
from semantica.kg import MethodRegistry
|
||||||
from semantica.kg.methods import calculate_centrality
|
|
||||||
|
|
||||||
def fast_centrality(graph, **kwargs):
|
registry = MethodRegistry()
|
||||||
"""Custom centrality implementation."""
|
|
||||||
|
def find_supply_chain_hops(graph, source_node, max_hops=3):
|
||||||
|
"""Custom BFS traversal for supply chain graphs."""
|
||||||
...
|
...
|
||||||
|
|
||||||
# register(task, name, func)
|
# Register under a string key
|
||||||
method_registry.register("centrality", "fast_centrality", fast_centrality)
|
registry.register("supply_chain_hops", find_supply_chain_hops)
|
||||||
|
|
||||||
# The task wrappers consult method_registry, so the name is now selectable:
|
# Call by name on any graph object
|
||||||
scores = calculate_centrality(kg, method="fast_centrality")
|
result = registry.call("supply_chain_hops", kg, source_node="Supplier_A", max_hops=5)
|
||||||
|
|
||||||
print(method_registry.list_all("centrality")) # {"centrality": ["fast_centrality", ...]}
|
# List all registered methods
|
||||||
|
print(registry.list_methods()) # ["supply_chain_hops", ...]
|
||||||
```
|
```
|
||||||
|
|
||||||
</Accordion>
|
</Accordion>
|
||||||
</AccordionGroup>
|
</AccordionGroup>
|
||||||
|
|
||||||
- [Quickstart Tutorial](/quickstart) — Build a full pipeline with code.
|
- [Quickstart Tutorial](quickstart) — Build a full pipeline with code.
|
||||||
- [Modules Guide](/modules) — Every module explained with examples.
|
- [Modules Guide](modules) — Every module explained with examples.
|
||||||
- [API Reference](/reference/context) — Complete technical reference.
|
- [API Reference](reference/context) — Complete technical reference.
|
||||||
|
|||||||
@@ -85,5 +85,5 @@ All contributors are expected to follow the [Contributor Covenant Code of Conduc
|
|||||||
- [GitHub Discussions](https://github.com/semantica-agi/semantica/discussions)
|
- [GitHub Discussions](https://github.com/semantica-agi/semantica/discussions)
|
||||||
- [Discord](https://discord.gg/sV34vps5hH)
|
- [Discord](https://discord.gg/sV34vps5hH)
|
||||||
|
|
||||||
- [Community](/community) — Community guidelines and values.
|
- [Community](community) — Community guidelines and values.
|
||||||
- [Governance](/governance) — How decisions are made and the project is run.
|
- [Governance](governance) — How decisions are made and the project is run.
|
||||||
|
|||||||
+1
-2
@@ -8,7 +8,7 @@ icon: "flask"
|
|||||||
**Where to start:**
|
**Where to start:**
|
||||||
- **New to Semantica**: begin with [Core Tutorials](#core-tutorials)
|
- **New to Semantica**: begin with [Core Tutorials](#core-tutorials)
|
||||||
- **Building an application**: see [Advanced Concepts](#advanced-concepts)
|
- **Building an application**: see [Advanced Concepts](#advanced-concepts)
|
||||||
- **Need installation help**: see the [Installation Guide](/installation)
|
- **Need installation help**: see the [Installation Guide](installation)
|
||||||
</Tip>
|
</Tip>
|
||||||
|
|
||||||
<Note>
|
<Note>
|
||||||
@@ -36,7 +36,6 @@ Essential guides to master the Semantica framework.
|
|||||||
- **[Graph Store](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/09_Graph_Store.ipynb)** — Persisting knowledge graphs in Neo4j or FalkorDB. Topics: Neo4j, Cypher, Persistence · *Intermediate*
|
- **[Graph Store](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/09_Graph_Store.ipynb)** — Persisting knowledge graphs in Neo4j or FalkorDB. Topics: Neo4j, Cypher, Persistence · *Intermediate*
|
||||||
- **[Ontology](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/14_Ontology.ipynb)** — Defining domain schemas and ontologies to structure your data. Topics: OWL, RDF, Schema Design · *Intermediate*
|
- **[Ontology](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/14_Ontology.ipynb)** — Defining domain schemas and ontologies to structure your data. Topics: OWL, RDF, Schema Design · *Intermediate*
|
||||||
- **[Seed Data](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/25_Seed_Data.ipynb)** — Bootstrapping a knowledge graph from trusted CSV, JSON, database, and API sources before extraction runs. Topics: SeedDataManager, Foundation Graphs · *Intermediate*
|
- **[Seed Data](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/25_Seed_Data.ipynb)** — Bootstrapping a knowledge graph from trusted CSV, JSON, database, and API sources before extraction runs. Topics: SeedDataManager, Foundation Graphs · *Intermediate*
|
||||||
- **[Semantic Layer Basics](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/26_Semantic_Layer_Basics.ipynb)** — Capstone tutorial that combines a knowledge graph, generated ontology, explicit mappings, ontology-aligned RDF, and a SPARQL query. Topics: Semantic Layer, Ontology Mapping, Oxigraph, SPARQL · *Intermediate*
|
|
||||||
|
|
||||||
|
|
||||||
## Advanced Concepts
|
## Advanced Concepts
|
||||||
|
|||||||
+27
-28
@@ -121,23 +121,6 @@
|
|||||||
"pages": [
|
"pages": [
|
||||||
"vector_stores/pgvector"
|
"vector_stores/pgvector"
|
||||||
]
|
]
|
||||||
},
|
|
||||||
{
|
|
||||||
"group": "FAQ",
|
|
||||||
"pages": [
|
|
||||||
"faq"
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"group": "Community",
|
|
||||||
"pages": [
|
|
||||||
"community",
|
|
||||||
"community-projects",
|
|
||||||
"contributing-guide",
|
|
||||||
"governance",
|
|
||||||
"citation",
|
|
||||||
"project-license"
|
|
||||||
]
|
|
||||||
}
|
}
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
@@ -184,17 +167,7 @@
|
|||||||
"guides/policy-engine",
|
"guides/policy-engine",
|
||||||
"guides/visualization",
|
"guides/visualization",
|
||||||
"guides/distance-intelligence",
|
"guides/distance-intelligence",
|
||||||
"guides/graph-analytics"
|
"guides/graph-analytics",
|
||||||
]
|
|
||||||
}
|
|
||||||
]
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"tab": "API Reference",
|
|
||||||
"groups": [
|
|
||||||
{
|
|
||||||
"group": "Context & Intelligence",
|
|
||||||
"pages": [
|
|
||||||
"reference/context",
|
"reference/context",
|
||||||
"reference/kg",
|
"reference/kg",
|
||||||
"reference/temporal",
|
"reference/temporal",
|
||||||
@@ -263,6 +236,32 @@
|
|||||||
]
|
]
|
||||||
}
|
}
|
||||||
]
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"tab": "FAQ",
|
||||||
|
"groups": [
|
||||||
|
{
|
||||||
|
"group": "FAQ",
|
||||||
|
"pages": [
|
||||||
|
"faq"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"group": "Community",
|
||||||
|
"pages": [
|
||||||
|
"community",
|
||||||
|
"community-projects",
|
||||||
|
"contributing-guide",
|
||||||
|
"governance",
|
||||||
|
"citation",
|
||||||
|
"project-license"
|
||||||
|
]
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"tab": "Changelog",
|
||||||
|
"href": "https://github.com/semantica-agi/semantica/releases"
|
||||||
}
|
}
|
||||||
]
|
]
|
||||||
},
|
},
|
||||||
|
|||||||
@@ -6,7 +6,7 @@ icon: "map"
|
|||||||
|
|
||||||
**`semantica-explorer`** is an **interactive browser dashboard** for knowledge graph exploration. You give it a graph file, it starts a local server, and opens a browser tab where you can search nodes, find paths, inspect provenance, and run analytics: no code required after launch.
|
**`semantica-explorer`** is an **interactive browser dashboard** for knowledge graph exploration. You give it a graph file, it starts a local server, and opens a browser tab where you can search nodes, find paths, inspect provenance, and run analytics: no code required after launch.
|
||||||
|
|
||||||
This page covers everything needed to go from zero to a running Explorer. For the full REST API reference and endpoint catalogue, see [Explorer Reference](/reference/explorer).
|
This page covers everything needed to go from zero to a running Explorer. For the full REST API reference and endpoint catalogue, see [Explorer Reference](reference/explorer).
|
||||||
|
|
||||||
|
|
||||||
## Prerequisites
|
## Prerequisites
|
||||||
@@ -27,7 +27,7 @@ Verify:
|
|||||||
semantica-explorer --help
|
semantica-explorer --help
|
||||||
```
|
```
|
||||||
|
|
||||||
You should see the usage message with the four available flags. If you see `command not found`, activate your virtual environment first. See [CLI Setup](/cli-setup#troubleshooting) for PATH help.
|
You should see the usage message with the four available flags. If you see `command not found`, activate your virtual environment first. See [CLI Setup](cli-setup#troubleshooting) for PATH help.
|
||||||
|
|
||||||
|
|
||||||
## Minimal End-to-End Example
|
## Minimal End-to-End Example
|
||||||
@@ -264,7 +264,7 @@ Once running, Explorer exposes a REST API and dashboard for:
|
|||||||
|
|
||||||
The full endpoint catalogue is documented in the Swagger UI at `/docs` and in the reference page below.
|
The full endpoint catalogue is documented in the Swagger UI at `/docs` and in the reference page below.
|
||||||
|
|
||||||
- [Explorer Reference](/reference/explorer) — Every REST endpoint, WebSocket events, analytics, and all supported flags.
|
- [Explorer Reference](reference/explorer) — Every REST endpoint, WebSocket events, analytics, and all supported flags.
|
||||||
- [CLI Setup](/cli-setup) — All five Semantica executables and when to use each one.
|
- [CLI Setup](cli-setup) — All five Semantica executables and when to use each one.
|
||||||
- [Context Module](/reference/context) — Full documentation for ContextGraph: build, query, save, and load.
|
- [Context Module](reference/context) — Full documentation for ContextGraph: build, query, save, and load.
|
||||||
- [Quickstart](/quickstart) — End-to-end pipeline: ingest → extract → build graph → export.
|
- [Quickstart](quickstart) — End-to-end pipeline: ingest → extract → build graph → export.
|
||||||
|
|||||||
+8
-8
@@ -16,7 +16,7 @@ icon: "circle-question"
|
|||||||
| Python version? | 3.8+ (3.11+ recommended) |
|
| Python version? | 3.8+ (3.11+ recommended) |
|
||||||
| API key required? | Optional: pattern extraction works with no keys |
|
| API key required? | Optional: pattern extraction works with no keys |
|
||||||
| Works with LangChain / LlamaIndex? | Yes: Semantica is a layer on top, not a replacement |
|
| Works with LangChain / LlamaIndex? | Yes: Semantica is a layer on top, not a replacement |
|
||||||
| Production-ready? | Yes: 1,000+ tests, security fixes shipped in every release (see [CHANGELOG](https://github.com/semantica-agi/semantica/blob/main/CHANGELOG.md)) |
|
| Production-ready? | Yes: 1,000+ tests, v0.5.0 ships with 12 security fixes |
|
||||||
| Latest version? | **v0.6.7** (August 2026) |
|
| Latest version? | **v0.6.7** (August 2026) |
|
||||||
| Local LLMs? | Yes: Ollama via LiteLLM, HuggingFaceLLM for air-gapped |
|
| Local LLMs? | Yes: Ollama via LiteLLM, HuggingFaceLLM for air-gapped |
|
||||||
|
|
||||||
@@ -70,9 +70,9 @@ Yes: MIT licensed, no vendor lock-in, no paywalled features. Some capabilities r
|
|||||||
|
|
||||||
<Accordion title="What's the latest version?" icon="star">
|
<Accordion title="What's the latest version?" icon="star">
|
||||||
|
|
||||||
**v0.6.7**: released August 2026.
|
**v0.5.0**: released May 2026.
|
||||||
|
|
||||||
Highlights: first-class LangChain integration, SAP OData ingestor, human-editable Markdown round-trip persistence for `ContextGraph`, a structured Action layer for the reasoning engine, and a public `run_shacl_validation` entry point. The 0.6.x line also added first-class CrewAI support and the Semantica RDF vocabulary with deterministic IRIs. See the [CHANGELOG](https://github.com/semantica-agi/semantica/blob/main/CHANGELOG.md) for the full history.
|
Highlights: Ontology Hub, Distance Intelligence, Parquet/XML ingestion, 12 security fixes, Graph Explorer redesign, NER gateway fix.
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
pip install --upgrade semantica
|
pip install --upgrade semantica
|
||||||
@@ -93,7 +93,7 @@ pip install --upgrade semantica
|
|||||||
pip install semantica
|
pip install semantica
|
||||||
```
|
```
|
||||||
|
|
||||||
See [Installation](/installation) for virtual environment setup, optional extras (`[gpu]`, `[all]`, provider-specific), and platform-specific troubleshooting.
|
See [Installation](installation) for virtual environment setup, optional extras (`[gpu]`, `[all]`, provider-specific), and platform-specific troubleshooting.
|
||||||
|
|
||||||
</Accordion>
|
</Accordion>
|
||||||
|
|
||||||
@@ -173,7 +173,7 @@ This includes PyTorch with CUDA, FAISS GPU, and CuPy.
|
|||||||
<Accordion title="How does Semantica handle large datasets?" icon="layer-group">
|
<Accordion title="How does Semantica handle large datasets?" icon="layer-group">
|
||||||
|
|
||||||
- **Batching**: process documents in configurable chunks to control memory usage
|
- **Batching**: process documents in configurable chunks to control memory usage
|
||||||
- **Parallel processing**: the `semantica.pipeline` module can run independent, parallel-safe steps in the same dependency layer concurrently (see the [Pipeline guide](/guides/pipeline))
|
- **Parallel processing**: `Pipeline(workers=N)` runs extraction steps concurrently
|
||||||
- **Delta processing**: update graphs incrementally without full recompute on new data
|
- **Delta processing**: update graphs incrementally without full recompute on new data
|
||||||
- **Persistent backends**: swap in-memory NetworkX for Neo4j, FalkorDB, or Apache AGE for large-scale production graphs
|
- **Persistent backends**: swap in-memory NetworkX for Neo4j, FalkorDB, or Apache AGE for large-scale production graphs
|
||||||
|
|
||||||
@@ -269,13 +269,13 @@ Groq, OpenAI, Anthropic, Google Gemini, Ollama (fully local), DeepSeek, Novita A
|
|||||||
|
|
||||||
<Accordion title="Is Semantica production-ready?" icon="shield-check">
|
<Accordion title="Is Semantica production-ready?" icon="shield-check">
|
||||||
|
|
||||||
Yes. Every release ships with:
|
Yes. v0.5.0 ships with:
|
||||||
|
|
||||||
- 1,000+ passing tests across Python 3.8–3.12
|
- 1,000+ passing tests across Python 3.8–3.12
|
||||||
- `PipelineValidator` and `FailureHandler` with exponential backoff and configurable retry policies
|
- `PipelineValidator` and `FailureHandler` with exponential backoff and configurable retry policies
|
||||||
- W3C PROV-O provenance tracking across all modules
|
- W3C PROV-O provenance tracking across all modules
|
||||||
- Change management with SHA-256 checksums and full audit trails
|
- Change management with SHA-256 checksums and full audit trails
|
||||||
- Ongoing security hardening: eval injection, pickle deserialization, SQL injection, XXE, SSRF, ReDoS, and path traversal fixes have all landed across recent releases (see the [CHANGELOG](https://github.com/semantica-agi/semantica/blob/main/CHANGELOG.md) security sections)
|
- 12 security vulnerability fixes: eval injection, pickle deserialization, SQL injection, XXE, SSRF, ReDoS, path traversal, and more
|
||||||
|
|
||||||
</Accordion>
|
</Accordion>
|
||||||
|
|
||||||
@@ -350,4 +350,4 @@ set PYTHONIOENCODING=utf-8
|
|||||||
|
|
||||||
- [Discord](https://discord.gg/sV34vps5hH) — Community chat and live support.
|
- [Discord](https://discord.gg/sV34vps5hH) — Community chat and live support.
|
||||||
- [GitHub Issues](https://github.com/semantica-agi/semantica/issues) — Bug reports and feature requests.
|
- [GitHub Issues](https://github.com/semantica-agi/semantica/issues) — Bug reports and feature requests.
|
||||||
- [Contributing](/contributing-guide) — Help improve Semantica.
|
- [Contributing](contributing-guide) — Help improve Semantica.
|
||||||
|
|||||||
+38
-46
@@ -5,7 +5,7 @@ icon: "rocket"
|
|||||||
---
|
---
|
||||||
|
|
||||||
<Tip>
|
<Tip>
|
||||||
Already installed? Jump straight to [Quickstart](/quickstart). Need setup help first? See [Installation](/installation).
|
Already installed? Jump straight to [Quickstart](quickstart). Need setup help first? See [Installation](installation).
|
||||||
</Tip>
|
</Tip>
|
||||||
|
|
||||||
## What You Can Build
|
## What You Can Build
|
||||||
@@ -52,15 +52,15 @@ icon: "rocket"
|
|||||||
|
|
||||||
| Track | You want to... | Start with |
|
| Track | You want to... | Start with |
|
||||||
| :----- | :-------------- | :--------- |
|
| :----- | :-------------- | :--------- |
|
||||||
| **Knowledge Graph** | Turn documents into structured, queryable graphs | [Quickstart → Step 1](/quickstart) |
|
| **Knowledge Graph** | Turn documents into structured, queryable graphs | [Quickstart → Step 1](quickstart) |
|
||||||
| **Agent Context** | Give your AI agent persistent memory and decision tracking | [Context reference](/reference/context) |
|
| **Agent Context** | Give your AI agent persistent memory and decision tracking | [Context reference](reference/context) |
|
||||||
| **GraphRAG** | Ground LLM answers in structured knowledge | [Concepts → GraphRAG](/concepts#graphrag) |
|
| **GraphRAG** | Ground LLM answers in structured knowledge | [Concepts → GraphRAG](concepts#graphrag) |
|
||||||
| **MCP Integration** | Use Semantica from Claude Desktop or VS Code | [MCP Server](/reference/mcp_server) |
|
| **MCP Integration** | Use Semantica from Claude Desktop or VS Code | [MCP Server](reference/mcp_server) |
|
||||||
|
|
||||||
</Step>
|
</Step>
|
||||||
|
|
||||||
<Step title="Run the pipeline">
|
<Step title="Run the pipeline">
|
||||||
The full 6-step pipeline: ingest, parse, extract, build, visualize, export: is in the [Quickstart](/quickstart). Takes under 5 minutes with pattern-based extraction (no API key required).
|
The full 6-step pipeline: ingest, parse, extract, build, visualize, export: is in the [Quickstart](quickstart). Takes under 5 minutes with pattern-based extraction (no API key required).
|
||||||
|
|
||||||
<Note>
|
<Note>
|
||||||
An LLM API key is **optional** for the quickstart. Pattern-based extraction works out of the box: upgrade to LLM extraction for higher accuracy when you're ready.
|
An LLM API key is **optional** for the quickstart. Pattern-based extraction works out of the box: upgrade to LLM extraction for higher accuracy when you're ready.
|
||||||
@@ -84,13 +84,13 @@ icon: "rocket"
|
|||||||
# 1. Ingest
|
# 1. Ingest
|
||||||
sources = FileIngestor().ingest("data/report.pdf")
|
sources = FileIngestor().ingest("data/report.pdf")
|
||||||
|
|
||||||
# 2. Parse (extract_text returns a plain string for any supported format)
|
# 2. Parse
|
||||||
text = DocumentParser().extract_text(sources[0].path)
|
parsed = DocumentParser().parse(sources[0])
|
||||||
|
|
||||||
# 3. Extract (extractors take text, return Entity / Relation objects)
|
# 3. Extract
|
||||||
ner = NERExtractor(method="pattern") # no API key needed
|
ner = NERExtractor(method="pattern") # no API key needed
|
||||||
entities = ner.extract(text)
|
entities = ner.extract(parsed)
|
||||||
relationships = RelationExtractor(method="pattern").extract(text, entities=entities)
|
relationships = RelationExtractor().extract(parsed, entities=entities)
|
||||||
|
|
||||||
# 4. Build
|
# 4. Build
|
||||||
graph = GraphBuilder(merge_entities=True).build(
|
graph = GraphBuilder(merge_entities=True).build(
|
||||||
@@ -99,7 +99,7 @@ icon: "rocket"
|
|||||||
print(f"{len(graph['entities'])} nodes, {len(graph['relationships'])} edges")
|
print(f"{len(graph['entities'])} nodes, {len(graph['relationships'])} edges")
|
||||||
```
|
```
|
||||||
|
|
||||||
**Next:** [Full pipeline walkthrough →](/quickstart)
|
**Next:** [Full pipeline walkthrough →](quickstart)
|
||||||
</Tab>
|
</Tab>
|
||||||
|
|
||||||
<Tab title="Agent Context">
|
<Tab title="Agent Context">
|
||||||
@@ -131,7 +131,7 @@ icon: "rocket"
|
|||||||
precedents = context.find_precedents("model selection", limit=5)
|
precedents = context.find_precedents("model selection", limit=5)
|
||||||
```
|
```
|
||||||
|
|
||||||
**Next:** [Context module reference →](/reference/context)
|
**Next:** [Context module reference →](reference/context)
|
||||||
</Tab>
|
</Tab>
|
||||||
|
|
||||||
<Tab title="GraphRAG">
|
<Tab title="GraphRAG">
|
||||||
@@ -144,32 +144,24 @@ icon: "rocket"
|
|||||||
context = AgentContext(
|
context = AgentContext(
|
||||||
vector_store=VectorStore(backend="faiss", dimension=768),
|
vector_store=VectorStore(backend="faiss", dimension=768),
|
||||||
knowledge_graph=ContextGraph(advanced_analytics=True),
|
knowledge_graph=ContextGraph(advanced_analytics=True),
|
||||||
graph_expansion=True, # blend graph traversal into retrieval
|
|
||||||
max_expansion_hops=3, # how far to walk from the seed nodes
|
|
||||||
)
|
)
|
||||||
|
|
||||||
# store() runs extraction and populates both the vector index and the graph
|
# Load your knowledge graph
|
||||||
context.store([
|
context.load_graph("company_kg.json")
|
||||||
{"content": "Steve Wozniak co-founded Apple with Steve Jobs in 1976."},
|
|
||||||
{"content": "Tony Fadell led the iPod team at Apple, then founded Nest."},
|
|
||||||
])
|
|
||||||
|
|
||||||
# GraphRAG retrieval: seed from vector matches, expand along graph edges
|
# Multi-hop GraphRAG query
|
||||||
results = context.retrieve(
|
result = context.query(
|
||||||
"What companies were founded by people who worked at Apple?",
|
"What companies were founded by people who worked at Apple?",
|
||||||
use_graph=True,
|
mode="graphrag",
|
||||||
expand_graph=True,
|
reasoning=True,
|
||||||
)
|
)
|
||||||
for r in results:
|
|
||||||
print(f"[{r['score']:.3f}] {r['content'][:70]} (source: {r['source']})")
|
# Every claim links back to a source node
|
||||||
|
for claim in result.claims:
|
||||||
|
print(f"{claim.text} → source: {claim.source_node}")
|
||||||
```
|
```
|
||||||
|
|
||||||
Each result carries `content`, `score`, `source`, and `metadata`. For a
|
**Next:** [GraphRAG concepts →](concepts#graphrag)
|
||||||
grounded natural-language answer plus an auditable traversal, use
|
|
||||||
`context.query_with_reasoning(query, llm_provider=...)` — it returns
|
|
||||||
`response`, `reasoning_path`, `sources`, and `confidence`.
|
|
||||||
|
|
||||||
**Next:** [GraphRAG concepts →](/concepts#graphrag)
|
|
||||||
</Tab>
|
</Tab>
|
||||||
|
|
||||||
<Tab title="MCP Integration">
|
<Tab title="MCP Integration">
|
||||||
@@ -191,9 +183,9 @@ icon: "rocket"
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
15 tools available instantly: extract entities, query graph, record decisions, run reasoning, export results.
|
12 tools available instantly: extract entities, query graph, record decisions, run reasoning, export results.
|
||||||
|
|
||||||
**Next:** [MCP Server reference →](/reference/mcp_server)
|
**Next:** [MCP Server reference →](reference/mcp_server)
|
||||||
</Tab>
|
</Tab>
|
||||||
</Tabs>
|
</Tabs>
|
||||||
|
|
||||||
@@ -202,29 +194,29 @@ icon: "rocket"
|
|||||||
|
|
||||||
Semantica uses a modular, layered architecture: import only what you need.
|
Semantica uses a modular, layered architecture: import only what you need.
|
||||||
|
|
||||||
- **[Input Layer](/reference/ingest)** — Load and prepare data from any source. Modules: `ingest`, `parse`, `split`, `normalize`
|
- **[Input Layer](reference/ingest)** — Load and prepare data from any source. Modules: `ingest`, `parse`, `split`, `normalize`
|
||||||
- **[Semantic Layer](/reference/semantic_extract)** — Extract meaning from raw text. Modules: `semantic_extract`, `kg`, `ontology`, `reasoning`
|
- **[Semantic Layer](reference/semantic_extract)** — Extract meaning from raw text. Modules: `semantic_extract`, `kg`, `ontology`, `reasoning`
|
||||||
- **[Storage Layer](/reference/vector_store)** — Persist knowledge for retrieval. Modules: `embeddings`, `vector_store`, `graph_store`, `triplet_store`
|
- **[Storage Layer](reference/vector_store)** — Persist knowledge for retrieval. Modules: `embeddings`, `vector_store`, `graph_store`, `triplet_store`
|
||||||
- **[Quality Layer](/reference/deduplication)** — Validate and deduplicate. Modules: `deduplication`, `conflicts`
|
- **[Quality Layer](reference/deduplication)** — Validate and deduplicate. Modules: `deduplication`, `conflicts`
|
||||||
- **[Context Layer](/reference/context)** — Track decisions and lineage. Modules: `context`, `provenance`, `change_management`
|
- **[Context Layer](reference/context)** — Track decisions and lineage. Modules: `context`, `provenance`, `change_management`
|
||||||
- **[Output Layer](/reference/export)** — Deliver results downstream. Modules: `export`, `visualization`, `pipeline`, `explorer`
|
- **[Output Layer](reference/export)** — Deliver results downstream. Modules: `export`, `visualization`, `pipeline`, `explorer`
|
||||||
|
|
||||||
|
|
||||||
## Which Module Do I Need?
|
## Which Module Do I Need?
|
||||||
|
|
||||||
See the [Choose the Right Module](/choose-your-module) guide — it maps 35+ developer goals to the right starting point across all 27 modules, with working code for the most common paths.
|
See the [Choose the Right Module](choose-your-module) guide — it maps 35+ developer goals to the right starting point across all 27 modules, with working code for the most common paths.
|
||||||
|
|
||||||
|
|
||||||
## Next Steps
|
## Next Steps
|
||||||
|
|
||||||
- [Core Concepts](/concepts) — Knowledge graphs, ontologies, and reasoning explained in depth.
|
- [Core Concepts](concepts) — Knowledge graphs, ontologies, and reasoning explained in depth.
|
||||||
- [Quickstart Tutorial](/quickstart) — Full 6-step pipeline walkthrough with working code.
|
- [Quickstart Tutorial](quickstart) — Full 6-step pipeline walkthrough with working code.
|
||||||
- [Module Reference](/modules) — Every module, class, and common chain explained.
|
- [Module Reference](modules) — Every module, class, and common chain explained.
|
||||||
- [API Reference](/reference/context) — Complete module documentation for every class and method.
|
- [API Reference](reference/context) — Complete module documentation for every class and method.
|
||||||
|
|
||||||
|
|
||||||
## Help
|
## Help
|
||||||
|
|
||||||
- [Discord](https://discord.gg/sV34vps5hH) — Ask questions, share projects, get community support.
|
- [Discord](https://discord.gg/sV34vps5hH) — Ask questions, share projects, get community support.
|
||||||
- [GitHub Issues](https://github.com/semantica-agi/semantica/issues) — Report bugs or request features.
|
- [GitHub Issues](https://github.com/semantica-agi/semantica/issues) — Report bugs or request features.
|
||||||
- [FAQ](/faq) — Common questions answered.
|
- [FAQ](faq) — Common questions answered.
|
||||||
|
|||||||
+4
-4
@@ -214,7 +214,7 @@ A vulnerability in XML parsers that allows attackers to read arbitrary files or
|
|||||||
|
|
||||||
## See Also
|
## See Also
|
||||||
|
|
||||||
- [Core Concepts](/concepts) — Deeper explanation of key ideas with code examples.
|
- [Core Concepts](concepts) — Deeper explanation of key ideas with code examples.
|
||||||
- [Getting Started](/getting-started) — First working examples: no prior graph experience required.
|
- [Getting Started](getting-started) — First working examples: no prior graph experience required.
|
||||||
- [Modules Guide](/modules) — All 27 modules explained with code and pipeline chains.
|
- [Modules Guide](modules) — All 27 modules explained with code and pipeline chains.
|
||||||
- [API Reference](/reference/context) — Complete technical reference for every class and method.
|
- [API Reference](reference/context) — Complete technical reference for every class and method.
|
||||||
|
|||||||
+3
-3
@@ -74,10 +74,10 @@ Semantica follows **Semantic Versioning** (`MAJOR.MINOR.PATCH`):
|
|||||||
|
|
||||||
## License
|
## License
|
||||||
|
|
||||||
MIT License: see [LICENSE](https://github.com/semantica-agi/semantica/blob/main/LICENSE) and the [License page](/project-license).
|
MIT License: see [LICENSE](https://github.com/semantica-agi/semantica/blob/main/LICENSE) and the [License page](project-license).
|
||||||
|
|
||||||
|
|
||||||
## See Also
|
## See Also
|
||||||
|
|
||||||
- [Contributing](/contributing-guide) — How to submit changes.
|
- [Contributing](contributing-guide) — How to submit changes.
|
||||||
- [Community](/community) — Community guidelines and channels.
|
- [Community](community) — Community guidelines and channels.
|
||||||
|
|||||||
@@ -46,7 +46,7 @@ Agent Memory provides persistent storage and intelligent retrieval of informatio
|
|||||||
- Simple retrieval tasks where relationships between entities don't matter
|
- Simple retrieval tasks where relationships between entities don't matter
|
||||||
|
|
||||||
<Info>
|
<Info>
|
||||||
This guide covers the memory layer. For graph-enriched traversal and entity linking, see [Context Graphs](/guides/context-graphs). For decision accountability — recording, auditing, and causally tracing what the agent chose — see [Decision Intelligence](/guides/decision-intelligence).
|
This guide covers the memory layer. For graph-enriched traversal and entity linking, see [Context Graphs](context-graphs). For decision accountability — recording, auditing, and causally tracing what the agent chose — see [Decision Intelligence](decision-intelligence).
|
||||||
</Info>
|
</Info>
|
||||||
|
|
||||||
## Setting Up a Persistent Memory Context
|
## Setting Up a Persistent Memory Context
|
||||||
@@ -657,10 +657,10 @@ print("Total memories: {}".format(s.get("total_items", 0)))
|
|||||||
|
|
||||||
## Related Guides
|
## Related Guides
|
||||||
|
|
||||||
- [Context Graphs](/guides/context-graphs) — How the underlying `ContextGraph` stores entity nodes and decision nodes; temporal interval reasoning; deduplication before node insertion; ontology from graph.
|
- [Context Graphs](context-graphs) — How the underlying `ContextGraph` stores entity nodes and decision nodes; temporal interval reasoning; deduplication before node insertion; ontology from graph.
|
||||||
- [Decision Intelligence](/guides/decision-intelligence) — Recording decisions as graph nodes with causal chains and policy gating.
|
- [Decision Intelligence](decision-intelligence) — Recording decisions as graph nodes with causal chains and policy gating.
|
||||||
- [Multi-Agent Systems](/guides/multi-agent) — Coordinating multiple agents through a shared `AgentContext` and save/load handoffs.
|
- [Multi-Agent Systems](multi-agent) — Coordinating multiple agents through a shared `AgentContext` and save/load handoffs.
|
||||||
- [LLM Integrations](/guides/llm-integrations) — Configuring the LLM provider passed to `query_with_reasoning()`.
|
- [LLM Integrations](llm-integrations) — Configuring the LLM provider passed to `query_with_reasoning()`.
|
||||||
- [Deduplication Guide](deduplication) — Full reference for `DuplicateDetector`, `EntityMerger`, similarity methods, and cluster strategies.
|
- [Deduplication Guide](deduplication) — Full reference for `DuplicateDetector`, `EntityMerger`, similarity methods, and cluster strategies.
|
||||||
- [Ontology Management](ontology) — Generate and validate OWL ontologies from the knowledge graph; export to Turtle, OWL/XML, JSON-LD.
|
- [Ontology Management](ontology) — Generate and validate OWL ontologies from the knowledge graph; export to Turtle, OWL/XML, JSON-LD.
|
||||||
- [Context Module Reference](../reference/context) — Full API: `AgentContext`, `AgentMemory`, `MemoryItem`, `ContextRetriever`.
|
- [Context Module Reference](../reference/context) — Full API: `AgentContext`, `AgentMemory`, `MemoryItem`, `ContextRetriever`.
|
||||||
|
|||||||
@@ -496,8 +496,8 @@ print("Model v1.1 verified and approved for production.")
|
|||||||
|
|
||||||
## Related Guides
|
## Related Guides
|
||||||
|
|
||||||
- [Context Graphs](/guides/context-graphs) — `ContextGraph.to_dict()` feeds `create_snapshot()`
|
- [Context Graphs](context-graphs) — `ContextGraph.to_dict()` feeds `create_snapshot()`
|
||||||
- [Ontology Management](ontology) — pair ontology versioning with graph versioning for a complete schema + data audit trail
|
- [Ontology Management](ontology) — pair ontology versioning with graph versioning for a complete schema + data audit trail
|
||||||
- [SHACL Validation](/guides/shacl-validation) — validate graph data at each version gate before snapshotting
|
- [SHACL Validation](shacl-validation) — validate graph data at each version gate before snapshotting
|
||||||
- [Provenance](provenance) — combine change management with W3C PROV-O lineage for a full audit trail
|
- [Provenance](provenance) — combine change management with W3C PROV-O lineage for a full audit trail
|
||||||
- [Visualization](visualization) — `TemporalVisualizer.visualize_snapshot_comparison()` and `visualize_metrics_evolution()` render version diffs as interactive charts
|
- [Visualization](visualization) — `TemporalVisualizer.visualize_snapshot_comparison()` and `visualize_metrics_evolution()` render version diffs as interactive charts
|
||||||
|
|||||||
@@ -69,7 +69,7 @@ flowchart TD
|
|||||||
2. **Conflict Detection** — Call `detect_entity_conflicts()` to surface all property disagreements at once, or `detect_value_conflicts()` to target a specific property.
|
2. **Conflict Detection** — Call `detect_entity_conflicts()` to surface all property disagreements at once, or `detect_value_conflicts()` to target a specific property.
|
||||||
3. **Resolution** — For each conflict, apply a strategy (`CREDIBILITY_WEIGHTED`, `MOST_RECENT`, `VOTING`, etc.) or route it for expert review (`EXPERT_REVIEW`).
|
3. **Resolution** — For each conflict, apply a strategy (`CREDIBILITY_WEIGHTED`, `MOST_RECENT`, `VOTING`, etc.) or route it for expert review (`EXPERT_REVIEW`).
|
||||||
4. **Persist Canonical Values** — Write resolved values back to your canonical entities or graph store. See [Persisting resolved values](#persisting-resolved-values).
|
4. **Persist Canonical Values** — Write resolved values back to your canonical entities or graph store. See [Persisting resolved values](#persisting-resolved-values).
|
||||||
5. **SHACL Validation** — Enforce structural constraints on the resolved graph to confirm it satisfies your ontology. See [SHACL Validation](/guides/shacl-validation).
|
5. **SHACL Validation** — Enforce structural constraints on the resolved graph to confirm it satisfies your ontology. See [SHACL Validation](shacl-validation).
|
||||||
|
|
||||||
## Quick Start: A Beginner Example
|
## Quick Start: A Beginner Example
|
||||||
|
|
||||||
@@ -698,6 +698,6 @@ Calling `set_resolution_rule()` for every entity-property pair just to apply the
|
|||||||
|
|
||||||
- [Deduplication](deduplication) — remove duplicate nodes before running conflict detection
|
- [Deduplication](deduplication) — remove duplicate nodes before running conflict detection
|
||||||
- [Provenance](provenance) — track which source each resolved value came from, and verify the audit trail cryptographically
|
- [Provenance](provenance) — track which source each resolved value came from, and verify the audit trail cryptographically
|
||||||
- [SHACL Validation](/guides/shacl-validation) — enforce structural constraints after conflicts are resolved
|
- [SHACL Validation](shacl-validation) — enforce structural constraints after conflicts are resolved
|
||||||
- [Change Management](/guides/change-management) — snapshot the graph before and after conflict resolution runs
|
- [Change Management](change-management) — snapshot the graph before and after conflict resolution runs
|
||||||
- [Ontology Management](ontology) — align entity types to a shared vocabulary to reduce type conflicts at the schema level
|
- [Ontology Management](ontology) — align entity types to a shared vocabulary to reduce type conflicts at the schema level
|
||||||
|
|||||||
@@ -50,7 +50,7 @@ A context graph is a property graph that stores entities as **nodes** and relati
|
|||||||
- Cases where setup complexity exceeds the relationship complexity
|
- Cases where setup complexity exceeds the relationship complexity
|
||||||
|
|
||||||
<Info>
|
<Info>
|
||||||
ContextGraph is an **in-memory data structure**. All nodes, edges, and metadata are stored in Python dictionaries and lists. For standalone graphs, persist state with `save_to_file()`. When using `AgentContext`, call `AgentContext.save()` instead — it saves the graph, the FAISS vector index, and memory in one step. For analytical operations on top of a populated graph — centrality rankings, community detection, node embeddings, link prediction — see the [Graph Analytics guide](/guides/graph-analytics). For recording and querying decisions stored as nodes, see the [Decision Intelligence guide](/guides/decision-intelligence).
|
ContextGraph is an **in-memory data structure**. All nodes, edges, and metadata are stored in Python dictionaries and lists. For standalone graphs, persist state with `save_to_file()`. When using `AgentContext`, call `AgentContext.save()` instead — it saves the graph, the FAISS vector index, and memory in one step. For analytical operations on top of a populated graph — centrality rankings, community detection, node embeddings, link prediction — see the [Graph Analytics guide](graph-analytics). For recording and querying decisions stored as nodes, see the [Decision Intelligence guide](decision-intelligence).
|
||||||
</Info>
|
</Info>
|
||||||
|
|
||||||
## Constructing the Graph
|
## Constructing the Graph
|
||||||
@@ -704,8 +704,8 @@ for n in stress_reach:
|
|||||||
|
|
||||||
## Related Guides
|
## Related Guides
|
||||||
|
|
||||||
- [Graph Analytics](/guides/graph-analytics) — centrality rankings, community detection, node embeddings, and link prediction on a populated `ContextGraph`
|
- [Graph Analytics](graph-analytics) — centrality rankings, community detection, node embeddings, and link prediction on a populated `ContextGraph`
|
||||||
- [Decision Intelligence](/guides/decision-intelligence) — recording decisions as typed nodes, causal chain analysis, precedent search, and policy enforcement
|
- [Decision Intelligence](decision-intelligence) — recording decisions as typed nodes, causal chain analysis, precedent search, and policy enforcement
|
||||||
- [Ingest](ingest) — loading data from PDFs, APIs, databases, STIX bundles, and RSS feeds into the graph
|
- [Ingest](ingest) — loading data from PDFs, APIs, databases, STIX bundles, and RSS feeds into the graph
|
||||||
- [Deduplication](deduplication) — detecting and merging near-duplicate nodes before insertion to prevent graph fragmentation
|
- [Deduplication](deduplication) — detecting and merging near-duplicate nodes before insertion to prevent graph fragmentation
|
||||||
- [Reasoning](reasoning) — temporal interval algebra (Allen relations), forward/backward chaining, and SPARQL over the knowledge graph
|
- [Reasoning](reasoning) — temporal interval algebra (Allen relations), forward/backward chaining, and SPARQL over the knowledge graph
|
||||||
|
|||||||
@@ -638,8 +638,8 @@ results = context.find_precedents("APT29 infrastructure attribution", limit=5)
|
|||||||
|
|
||||||
## Related Guides
|
## Related Guides
|
||||||
|
|
||||||
- [Context Graphs](/guides/context-graphs) — how `ContextGraph` stores decision nodes and causal edges
|
- [Context Graphs](context-graphs) — how `ContextGraph` stores decision nodes and causal edges
|
||||||
- [Distance Intelligence](/guides/distance-intelligence) — `trace_decision_causality()` annotates causal chains with confidence decay and distance bands
|
- [Distance Intelligence](distance-intelligence) — `trace_decision_causality()` annotates causal chains with confidence decay and distance bands
|
||||||
- [Provenance](provenance) — W3C PROV-O audit trail that wraps decision records in standards-compliant provenance
|
- [Provenance](provenance) — W3C PROV-O audit trail that wraps decision records in standards-compliant provenance
|
||||||
- [MCP Server](/guides/mcp-server) — expose decision recording and precedent search to LLM agents via the `record_decision` and `find_precedents` tools
|
- [MCP Server](mcp-server) — expose decision recording and precedent search to LLM agents via the `record_decision` and `find_precedents` tools
|
||||||
- [Change Management](/guides/change-management) — checkpoint decision state with `flush_checkpoint()` for versioned snapshots
|
- [Change Management](change-management) — checkpoint decision state with `flush_checkpoint()` for versioned snapshots
|
||||||
|
|||||||
@@ -612,7 +612,7 @@ The similarity threshold controls sensitivity. Start at 0.7 and examine false po
|
|||||||
## Related Guides
|
## Related Guides
|
||||||
|
|
||||||
- [Ingest Anything](ingest) — multi-source ingestion creates the duplicates this module resolves
|
- [Ingest Anything](ingest) — multi-source ingestion creates the duplicates this module resolves
|
||||||
- [Context Graphs](/guides/context-graphs) — store deduplicated entities directly in the knowledge graph
|
- [Context Graphs](context-graphs) — store deduplicated entities directly in the knowledge graph
|
||||||
- [Conflict Resolution](/guides/conflict-resolution) — after merging, reconcile disagreeing property values on the canonical entity
|
- [Conflict Resolution](conflict-resolution) — after merging, reconcile disagreeing property values on the canonical entity
|
||||||
- [Provenance](provenance) — track merge lineage so every canonical entity traces back to its original sources
|
- [Provenance](provenance) — track merge lineage so every canonical entity traces back to its original sources
|
||||||
- [Pipeline](pipeline) — chain ingest, deduplicate, and store as a `PipelineBuilder` workflow
|
- [Pipeline](pipeline) — chain ingest, deduplicate, and store as a `PipelineBuilder` workflow
|
||||||
|
|||||||
@@ -557,8 +557,8 @@ for chain in chains:
|
|||||||
|
|
||||||
## Related Guides
|
## Related Guides
|
||||||
|
|
||||||
- [Context Graphs](/guides/context-graphs) — `ContextGraph` node and edge model; `add_edge(weight=...)` feeds confidence decay
|
- [Context Graphs](context-graphs) — `ContextGraph` node and edge model; `add_edge(weight=...)` feeds confidence decay
|
||||||
- [Graph Analytics](/guides/graph-analytics) — centrality, community detection, Node2Vec embeddings, link prediction
|
- [Graph Analytics](graph-analytics) — centrality, community detection, Node2Vec embeddings, link prediction
|
||||||
- [Agent Memory](/guides/agent-memory) — proximity-blended retrieval (`proximity_weight`) integrates distance intelligence into memory search
|
- [Agent Memory](agent-memory) — proximity-blended retrieval (`proximity_weight`) integrates distance intelligence into memory search
|
||||||
- [Decision Intelligence](/guides/decision-intelligence) — `trace_decision_causality()` for causal chains with distance annotations
|
- [Decision Intelligence](decision-intelligence) — `trace_decision_causality()` for causal chains with distance annotations
|
||||||
- [Reasoning & Rules](reasoning) — `TemporalReasoningEngine` for Allen interval algebra over time-bounded graph nodes
|
- [Reasoning & Rules](reasoning) — `TemporalReasoningEngine` for Allen interval algebra over time-bounded graph nodes
|
||||||
|
|||||||
@@ -443,8 +443,8 @@ For semantic reasoning and ontology work, OWL/XML is the format — it is the on
|
|||||||
|
|
||||||
## Related Guides
|
## Related Guides
|
||||||
|
|
||||||
- [Context Graphs](/guides/context-graphs) — the `ContextGraph` object whose `to_dict()` feeds all exports
|
- [Context Graphs](context-graphs) — the `ContextGraph` object whose `to_dict()` feeds all exports
|
||||||
- [Ontology Management](ontology) — export OWL ontologies generated from your graph
|
- [Ontology Management](ontology) — export OWL ontologies generated from your graph
|
||||||
- [Reasoning & Rules](reasoning) — reasoning results can be exported as RDF triples
|
- [Reasoning & Rules](reasoning) — reasoning results can be exported as RDF triples
|
||||||
- [Change Management](/guides/change-management) — snapshot a graph before exporting to prove the export was made from a verified state
|
- [Change Management](change-management) — snapshot a graph before exporting to prove the export was made from a verified state
|
||||||
- [Pipeline](pipeline) — chain ingest, extract, and export in a single `PipelineBuilder`
|
- [Pipeline](pipeline) — chain ingest, extract, and export in a single `PipelineBuilder`
|
||||||
|
|||||||
@@ -310,7 +310,7 @@ for node1, node2, score in predictions:
|
|||||||
A score above 0.8 is worth analyst review — these aren't random; they're edges the topology of the existing graph strongly implies. Scores below 0.5 are noise. The sweet spot for human review is 0.6–0.8: plausible but not yet confirmed.
|
A score above 0.8 is worth analyst review — these aren't random; they're edges the topology of the existing graph strongly implies. Scores below 0.5 are noise. The sweet spot for human review is 0.6–0.8: plausible but not yet confirmed.
|
||||||
|
|
||||||
<Info>
|
<Info>
|
||||||
Link prediction is also available on `Decision` nodes through `DecisionQuery.predict_decision_relationships(decision_id, top_k)`. See the [Decision Intelligence guide](/guides/decision-intelligence) for how to surface causal relationships between past decisions.
|
Link prediction is also available on `Decision` nodes through `DecisionQuery.predict_decision_relationships(decision_id, top_k)`. See the [Decision Intelligence guide](decision-intelligence) for how to surface causal relationships between past decisions.
|
||||||
</Info>
|
</Info>
|
||||||
|
|
||||||
## Understanding Your Decision History
|
## Understanding Your Decision History
|
||||||
@@ -538,7 +538,7 @@ print(f"\n{len(result['communities'])} exposure clusters "
|
|||||||
|
|
||||||
## Related Guides
|
## Related Guides
|
||||||
|
|
||||||
- [Context Graphs](/guides/context-graphs) — building and querying the underlying `ContextGraph`
|
- [Context Graphs](context-graphs) — building and querying the underlying `ContextGraph`
|
||||||
- [Visualization](visualization) — render centrality rankings and community clusters as interactive dashboards
|
- [Visualization](visualization) — render centrality rankings and community clusters as interactive dashboards
|
||||||
- [Decision Intelligence](/guides/decision-intelligence) — link prediction and structural similarity applied to decision nodes
|
- [Decision Intelligence](decision-intelligence) — link prediction and structural similarity applied to decision nodes
|
||||||
- [GraphRAG](/guides/graphrag) — using analytics results to ground LLM generation in the most contextually relevant subgraph
|
- [GraphRAG](graphrag) — using analytics results to ground LLM generation in the most contextually relevant subgraph
|
||||||
|
|||||||
@@ -576,9 +576,9 @@ The vector search and graph traversal run independently, then their scores are f
|
|||||||
|
|
||||||
## Related Guides
|
## Related Guides
|
||||||
|
|
||||||
- [Semantic Extraction](/guides/semantic-extraction) — build the graph from raw unstructured text
|
- [Semantic Extraction](semantic-extraction) — build the graph from raw unstructured text
|
||||||
- [Agent Memory](/guides/agent-memory) — store, retrieve, and persist agent memories
|
- [Agent Memory](agent-memory) — store, retrieve, and persist agent memories
|
||||||
- [Context Graphs](/guides/context-graphs) — build and traverse the knowledge graph directly
|
- [Context Graphs](context-graphs) — build and traverse the knowledge graph directly
|
||||||
- [Reasoning](reasoning) — derive new facts and run inference rules over the graph
|
- [Reasoning](reasoning) — derive new facts and run inference rules over the graph
|
||||||
- [Decision Intelligence](/guides/decision-intelligence) — causal chains, policy enforcement, decision tracking
|
- [Decision Intelligence](decision-intelligence) — causal chains, policy enforcement, decision tracking
|
||||||
- [LLM Integrations](/guides/llm-integrations) — connect Groq, OpenAI, Anthropic, HuggingFace, and 100+ more
|
- [LLM Integrations](llm-integrations) — connect Groq, OpenAI, Anthropic, HuggingFace, and 100+ more
|
||||||
|
|||||||
@@ -951,8 +951,8 @@ print(f"Compliance graph: {graph.stats()['node_count']} nodes, "
|
|||||||
## Related Guides
|
## Related Guides
|
||||||
|
|
||||||
- [Pipeline](pipeline) — chain ingest steps with `PipelineBuilder` for automated, retryable, parallelised workflows
|
- [Pipeline](pipeline) — chain ingest steps with `PipelineBuilder` for automated, retryable, parallelised workflows
|
||||||
- [Context Graphs](/guides/context-graphs) — storing and querying the entities you ingest as a typed property graph
|
- [Context Graphs](context-graphs) — storing and querying the entities you ingest as a typed property graph
|
||||||
- [Semantic Extraction](/guides/semantic-extraction) — NER, relation extraction, and triplet extraction from ingested text
|
- [Semantic Extraction](semantic-extraction) — NER, relation extraction, and triplet extraction from ingested text
|
||||||
- [Provenance](provenance) — tracking the origin document, confidence score, and ingestion timestamp for every extracted entity
|
- [Provenance](provenance) — tracking the origin document, confidence score, and ingestion timestamp for every extracted entity
|
||||||
- [Databricks Integration](../integrations/databricks) — Unity Catalog setup, PAT/OAuth M2M authentication, and lineage introspection
|
- [Databricks Integration](../integrations/databricks) — Unity Catalog setup, PAT/OAuth M2M authentication, and lineage introspection
|
||||||
- [Snowflake Integration](../integrations/snowflake) — warehouse setup and password/key-pair/OAuth authentication
|
- [Snowflake Integration](../integrations/snowflake) — warehouse setup and password/key-pair/OAuth authentication
|
||||||
|
|||||||
@@ -719,7 +719,7 @@ for src in best["sources"]:
|
|||||||
|
|
||||||
## Related Guides
|
## Related Guides
|
||||||
|
|
||||||
- [Agent Memory](/guides/agent-memory) — using `query_with_reasoning()` with any LLM provider for graph-grounded retrieval
|
- [Agent Memory](agent-memory) — using `query_with_reasoning()` with any LLM provider for graph-grounded retrieval
|
||||||
- [Multi-Agent Systems](/guides/multi-agent) — wiring different LLM providers to different agent tiers in a shared-graph pipeline
|
- [Multi-Agent Systems](multi-agent) — wiring different LLM providers to different agent tiers in a shared-graph pipeline
|
||||||
- [Semantic Extraction](/guides/semantic-extraction) — LLM-powered NER, relation extraction, event detection, and triplet extraction
|
- [Semantic Extraction](semantic-extraction) — LLM-powered NER, relation extraction, event detection, and triplet extraction
|
||||||
- [GraphRAG](/guides/graphrag) — multi-hop graph reasoning with `query_with_reasoning()`
|
- [GraphRAG](graphrag) — multi-hop graph reasoning with `query_with_reasoning()`
|
||||||
|
|||||||
@@ -11,7 +11,7 @@ MCP stands for the Model Context Protocol. It is an open standard that allows ex
|
|||||||
The Semantica MCP server exposes your knowledge graph as 12 callable tools. By connecting it, any compatible AI client can traverse the graph live, record decisions, run analytics, and export results during a conversation — without you having to write custom tool wrappers.
|
The Semantica MCP server exposes your knowledge graph as 12 callable tools. By connecting it, any compatible AI client can traverse the graph live, record decisions, run analytics, and export results during a conversation — without you having to write custom tool wrappers.
|
||||||
|
|
||||||
<Info>
|
<Info>
|
||||||
The Semantica MCP server exposes 15 tools and 3 read-only resources. All tools accept and return JSON. No configuration beyond an optional environment variable for graph persistence is required.
|
The Semantica MCP server exposes 12 tools and 3 read-only resources. All tools accept and return JSON. No configuration beyond an optional environment variable for graph persistence is required.
|
||||||
</Info>
|
</Info>
|
||||||
|
|
||||||
## Architecture & Communication
|
## Architecture & Communication
|
||||||
@@ -132,7 +132,7 @@ docker run --rm -i \
|
|||||||
ghcr.io/semantica-agi/semantica-mcp:latest
|
ghcr.io/semantica-agi/semantica-mcp:latest
|
||||||
```
|
```
|
||||||
|
|
||||||
## What the Agent Can Do: The 15 Tools
|
## What the Agent Can Do: The 12 Tools
|
||||||
|
|
||||||
Once connected, the LLM can call any of these tools during a conversation. The agent chains them automatically — you do not orchestrate the sequence, you just describe what you want.
|
Once connected, the LLM can call any of these tools during a conversation. The agent chains them automatically — you do not orchestrate the sequence, you just describe what you want.
|
||||||
|
|
||||||
@@ -140,8 +140,6 @@ Once connected, the LLM can call any of these tools during a conversation. The a
|
|||||||
|
|
||||||
**Knowledge graph manipulation** — `add_entity` adds a node, `add_relationship` adds a directed edge. After extraction, the agent calls these to persist what it found into the live graph.
|
**Knowledge graph manipulation** — `add_entity` adds a node, `add_relationship` adds a directed edge. After extraction, the agent calls these to persist what it found into the live graph.
|
||||||
|
|
||||||
**Live graph queries and edits** — `query_graph` reads the graph without exporting it: fetch one node, walk its neighbours up to five hops, or keyword-search nodes. `update_node` merges properties onto an existing node (for example marking a task node `done`), and `delete_node` archives a node it no longer tracks. When `SEMANTICA_KG_PATH` is set, `update_node` and `delete_node` write their changes back to that file so they survive a restart.
|
|
||||||
|
|
||||||
**Decision intelligence** — `record_decision` writes a decision as a provenance node with confidence score, reasoning, and decision maker identity. `query_decisions` retrieves past decisions by query or category. `find_precedents` finds the most similar past decisions by semantic similarity. `get_causal_chain` traces decision causality upstream or downstream.
|
**Decision intelligence** — `record_decision` writes a decision as a provenance node with confidence score, reasoning, and decision maker identity. `query_decisions` retrieves past decisions by query or category. `find_precedents` finds the most similar past decisions by semantic similarity. `get_causal_chain` traces decision causality upstream or downstream.
|
||||||
|
|
||||||
**Reasoning** — `run_reasoning` applies forward-chaining IF/THEN rules over a set of facts and returns derived conclusions.
|
**Reasoning** — `run_reasoning` applies forward-chaining IF/THEN rules over a set of facts and returns derived conclusions.
|
||||||
@@ -343,7 +341,7 @@ The result is a fully auditable credit decision trail with precedent links, read
|
|||||||
## Related Guides
|
## Related Guides
|
||||||
|
|
||||||
- [Reasoning & Rules](reasoning) — the engine behind the `run_reasoning` tool
|
- [Reasoning & Rules](reasoning) — the engine behind the `run_reasoning` tool
|
||||||
- [Decision Intelligence](/guides/decision-intelligence) — how decisions are stored as causal graph nodes
|
- [Decision Intelligence](decision-intelligence) — how decisions are stored as causal graph nodes
|
||||||
- [Context Graphs](/guides/context-graphs) — the graph that `add_entity` and `add_relationship` write to
|
- [Context Graphs](context-graphs) — the graph that `add_entity` and `add_relationship` write to
|
||||||
- [Export & Serialization](export) — all export formats available via `export_graph`
|
- [Export & Serialization](export) — all export formats available via `export_graph`
|
||||||
- [Ontology Management](ontology) — generate OWL ontologies from the graph built via MCP
|
- [Ontology Management](ontology) — generate OWL ontologies from the graph built via MCP
|
||||||
|
|||||||
@@ -55,7 +55,7 @@ Semantica coordinates agents through shared context (memory and knowledge graphs
|
|||||||
Semantica coordinates multiple agents through a shared `ContextGraph` — agents read and write to the same graph, or hand off serialized state via `save()` and `load()`, with no message broker required. Use this pattern when splitting work across ingestion, enrichment, reasoning, and reporting roles that must share a single evidence base.
|
Semantica coordinates multiple agents through a shared `ContextGraph` — agents read and write to the same graph, or hand off serialized state via `save()` and `load()`, with no message broker required. Use this pattern when splitting work across ingestion, enrichment, reasoning, and reporting roles that must share a single evidence base.
|
||||||
|
|
||||||
<Info>
|
<Info>
|
||||||
This guide covers multi-agent coordination. For the memory layer each agent uses internally, see [Agent Memory](/guides/agent-memory). For graph traversal and entity linking, see [Context Graphs](/guides/context-graphs). For decision recording and precedent matching, see [Decision Intelligence](/guides/decision-intelligence).
|
This guide covers multi-agent coordination. For the memory layer each agent uses internally, see [Agent Memory](agent-memory). For graph traversal and entity linking, see [Context Graphs](context-graphs). For decision recording and precedent matching, see [Decision Intelligence](decision-intelligence).
|
||||||
</Info>
|
</Info>
|
||||||
|
|
||||||
## The Three Coordination Patterns
|
## The Three Coordination Patterns
|
||||||
@@ -679,7 +679,7 @@ context.retrieve("...", user_id="analyst-jsmith")
|
|||||||
|
|
||||||
## Related Guides
|
## Related Guides
|
||||||
|
|
||||||
- [Agent Memory](/guides/agent-memory) — memory storage, retrieval, persistence, and the working memory window each agent uses internally
|
- [Agent Memory](agent-memory) — memory storage, retrieval, persistence, and the working memory window each agent uses internally
|
||||||
- [Context Graphs](/guides/context-graphs) — build and traverse the shared `ContextGraph` directly; temporal interval reasoning; entity deduplication before node insertion
|
- [Context Graphs](context-graphs) — build and traverse the shared `ContextGraph` directly; temporal interval reasoning; entity deduplication before node insertion
|
||||||
- [Decision Intelligence](/guides/decision-intelligence) — record and trace decisions across agent handoffs with causal chain analysis
|
- [Decision Intelligence](decision-intelligence) — record and trace decisions across agent handoffs with causal chain analysis
|
||||||
- [LLM Integrations](/guides/llm-integrations) — configure the LLM provider passed to `query_with_reasoning()` in each agent
|
- [LLM Integrations](llm-integrations) — configure the LLM provider passed to `query_with_reasoning()` in each agent
|
||||||
|
|||||||
@@ -297,7 +297,7 @@ export_rdf(ontology, "cyber_threat.jsonld", format="jsonld")
|
|||||||
export_rdf(ontology, "cyber_threat.nt", format="ntriples")
|
export_rdf(ontology, "cyber_threat.nt", format="ntriples")
|
||||||
```
|
```
|
||||||
|
|
||||||
The exported Turtle file is the input to Semantica's SHACL validation pipeline. See the [SHACL Validation](/guides/shacl-validation) guide for how to generate constraint shapes from this ontology and run them against live graph data.
|
The exported Turtle file is the input to Semantica's SHACL validation pipeline. See the [SHACL Validation](shacl-validation) guide for how to generate constraint shapes from this ontology and run them against live graph data.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -503,8 +503,8 @@ else:
|
|||||||
|
|
||||||
## Related Guides
|
## Related Guides
|
||||||
|
|
||||||
- [SHACL Validation](/guides/shacl-validation) — generate W3C SHACL constraint shapes from your ontology and validate live graph data against them
|
- [SHACL Validation](shacl-validation) — generate W3C SHACL constraint shapes from your ontology and validate live graph data against them
|
||||||
- [Reasoning & Rules](reasoning) — apply forward/backward-chaining rules over your ontology to derive new facts
|
- [Reasoning & Rules](reasoning) — apply forward/backward-chaining rules over your ontology to derive new facts
|
||||||
- [Export & Serialization](export) — export graphs to RDF, GraphML, CSV, and Neo4j Cypher
|
- [Export & Serialization](export) — export graphs to RDF, GraphML, CSV, and Neo4j Cypher
|
||||||
- [Semantic Extraction](/guides/semantic-extraction) — extract entities and relationships that feed ontology generation
|
- [Semantic Extraction](semantic-extraction) — extract entities and relationships that feed ontology generation
|
||||||
- [Context Graphs](/guides/context-graphs) — the knowledge graph that ontology generation reads from
|
- [Context Graphs](context-graphs) — the knowledge graph that ontology generation reads from
|
||||||
|
|||||||
@@ -717,6 +717,6 @@ print(f"Compliance delta update: {result.output}")
|
|||||||
## Related Guides
|
## Related Guides
|
||||||
|
|
||||||
- [Ingest](ingest) — all source types for the ingest step: PDFs, APIs, databases, RSS feeds, STIX directories, and streams
|
- [Ingest](ingest) — all source types for the ingest step: PDFs, APIs, databases, RSS feeds, STIX directories, and streams
|
||||||
- [Semantic Extraction](/guides/semantic-extraction) — NER, relation extraction, triplet extraction, and event detection for the extract step
|
- [Semantic Extraction](semantic-extraction) — NER, relation extraction, triplet extraction, and event detection for the extract step
|
||||||
- [Context Graphs](/guides/context-graphs) — building and querying the `ContextGraph` that the store step populates
|
- [Context Graphs](context-graphs) — building and querying the `ContextGraph` that the store step populates
|
||||||
- [Provenance](provenance) — tracking the origin document, confidence score, and pipeline run ID for every extracted entity
|
- [Provenance](provenance) — tracking the origin document, confidence score, and pipeline run ID for every extracted entity
|
||||||
|
|||||||
@@ -662,9 +662,9 @@ print("Policy updated to v2.4.0")
|
|||||||
|
|
||||||
## Related Guides
|
## Related Guides
|
||||||
|
|
||||||
- [Decision Intelligence](/guides/decision-intelligence) — `record_decision()`, causal chains, and precedent search — the decisions that `check_compliance()` evaluates
|
- [Decision Intelligence](decision-intelligence) — `record_decision()`, causal chains, and precedent search — the decisions that `check_compliance()` evaluates
|
||||||
- [Reasoning & Rules](reasoning) — complement policy rules with formal inference for logical conflict detection
|
- [Reasoning & Rules](reasoning) — complement policy rules with formal inference for logical conflict detection
|
||||||
- [SHACL Validation](/guides/shacl-validation) — enforce structural constraints on policy nodes themselves
|
- [SHACL Validation](shacl-validation) — enforce structural constraints on policy nodes themselves
|
||||||
- [Change Management](/guides/change-management) — version-snapshot the policy graph alongside the knowledge graph
|
- [Change Management](change-management) — version-snapshot the policy graph alongside the knowledge graph
|
||||||
- [Provenance](provenance) — W3C PROV-O lineage for every policy decision and exception
|
- [Provenance](provenance) — W3C PROV-O lineage for every policy decision and exception
|
||||||
- [MCP Server](/guides/mcp-server) — expose `record_decision` and `find_precedents` as MCP tools for AI agents
|
- [MCP Server](mcp-server) — expose `record_decision` and `find_precedents` as MCP tools for AI agents
|
||||||
|
|||||||
@@ -659,7 +659,7 @@ Note: the banking example above passes `agent_id="credit_data_service_v2"` to `t
|
|||||||
|
|
||||||
## Related Guides
|
## Related Guides
|
||||||
|
|
||||||
- [Semantic Extraction](/guides/semantic-extraction) — the NER and relation extraction pipeline that auto-generates provenance entries for every extracted entity
|
- [Semantic Extraction](semantic-extraction) — the NER and relation extraction pipeline that auto-generates provenance entries for every extracted entity
|
||||||
- [Conflict Resolution](/guides/conflict-resolution) — provenance property sources feed directly into conflict detection; every resolved value is traceable to its source
|
- [Conflict Resolution](conflict-resolution) — provenance property sources feed directly into conflict detection; every resolved value is traceable to its source
|
||||||
- [Deduplication](deduplication) — merge operations are recorded in merge history; pair with provenance for a complete lineage from source to canonical entity
|
- [Deduplication](deduplication) — merge operations are recorded in merge history; pair with provenance for a complete lineage from source to canonical entity
|
||||||
- [Provenance Reference](../reference/provenance) — full storage backend API, `InMemoryStorage`, `SQLiteStorage`, and `ProvenanceEntry` schema
|
- [Provenance Reference](../reference/provenance) — full storage backend API, `InMemoryStorage`, `SQLiteStorage`, and `ProvenanceEntry` schema
|
||||||
|
|||||||
@@ -838,9 +838,9 @@ if proof:
|
|||||||
|
|
||||||
## Related Guides
|
## Related Guides
|
||||||
|
|
||||||
- [Semantic Extraction](/guides/semantic-extraction) — extract the entities and relationships that populate the graph facts you reason over
|
- [Semantic Extraction](semantic-extraction) — extract the entities and relationships that populate the graph facts you reason over
|
||||||
- [GraphRAG](/guides/graphrag) — retrieve graph-grounded context for LLM responses
|
- [GraphRAG](graphrag) — retrieve graph-grounded context for LLM responses
|
||||||
- [Ontology Management](ontology) — generate OWL ontologies to give your rules formal semantics
|
- [Ontology Management](ontology) — generate OWL ontologies to give your rules formal semantics
|
||||||
- [Decision Intelligence](/guides/decision-intelligence) — record and trace inferred decisions through the full causal chain
|
- [Decision Intelligence](decision-intelligence) — record and trace inferred decisions through the full causal chain
|
||||||
- [Context Graphs](/guides/context-graphs) — the knowledge graph that reasoning operates over
|
- [Context Graphs](context-graphs) — the knowledge graph that reasoning operates over
|
||||||
- [MCP Server](/guides/mcp-server) — expose `run_reasoning` as a tool for Claude and other agents
|
- [MCP Server](mcp-server) — expose `run_reasoning` as a tool for Claude and other agents
|
||||||
|
|||||||
@@ -71,7 +71,7 @@ This pipeline transforms documents like "APT29 deployed HAMMERTOSS malware targe
|
|||||||
`semantica.semantic_extract` turns unstructured text into structured graph-ready output: it identifies named entities, extracts relationships between them, detects time-anchored events, resolves coreferences, and serialises everything as RDF triplets. Use it to populate a `ContextGraph` from raw documents — intelligence reports, clinical notes, regulatory filings, or any free-text corpus.
|
`semantica.semantic_extract` turns unstructured text into structured graph-ready output: it identifies named entities, extracts relationships between them, detects time-anchored events, resolves coreferences, and serialises everything as RDF triplets. Use it to populate a `ContextGraph` from raw documents — intelligence reports, clinical notes, regulatory filings, or any free-text corpus.
|
||||||
|
|
||||||
<Info>
|
<Info>
|
||||||
Extracted entities and relationships feed into `ContextGraph` via `AgentContext.store()`. For how they are attributed back to source documents, see the [Provenance Guide](provenance). For how the populated graph is queried and traversed, see [Context Graphs](/guides/context-graphs).
|
Extracted entities and relationships feed into `ContextGraph` via `AgentContext.store()`. For how they are attributed back to source documents, see the [Provenance Guide](provenance). For how the populated graph is queried and traversed, see [Context Graphs](context-graphs).
|
||||||
</Info>
|
</Info>
|
||||||
|
|
||||||
## Step 1 — Named Entity Recognition: who and what is in the text
|
## Step 1 — Named Entity Recognition: who and what is in the text
|
||||||
@@ -664,8 +664,8 @@ The fallback behaviour is automatic: if the primary method returns an empty list
|
|||||||
## Related Guides
|
## Related Guides
|
||||||
|
|
||||||
- [Provenance Guide](provenance) — track every extracted entity and chunk back to its source document
|
- [Provenance Guide](provenance) — track every extracted entity and chunk back to its source document
|
||||||
- [Agent Memory Guide](/guides/agent-memory) — store extracted knowledge as searchable agent memories with graph enrichment
|
- [Agent Memory Guide](agent-memory) — store extracted knowledge as searchable agent memories with graph enrichment
|
||||||
- [Context Graphs Guide](/guides/context-graphs) — how extracted entities populate `ContextGraph` nodes and edges
|
- [Context Graphs Guide](context-graphs) — how extracted entities populate `ContextGraph` nodes and edges
|
||||||
- [GraphRAG Guide](/guides/graphrag) — retrieve facts from the populated graph to ground LLM responses
|
- [GraphRAG Guide](graphrag) — retrieve facts from the populated graph to ground LLM responses
|
||||||
- [Reasoning Guide](reasoning) — derive new facts, run SPARQL queries, and apply inference rules over the extracted graph
|
- [Reasoning Guide](reasoning) — derive new facts, run SPARQL queries, and apply inference rules over the extracted graph
|
||||||
- [Semantic Extract Reference](../reference/semantic_extract) — full API for all extractor classes, providers, and validators
|
- [Semantic Extract Reference](../reference/semantic_extract) — full API for all extractor classes, providers, and validators
|
||||||
|
|||||||
@@ -740,5 +740,5 @@ def validate_before_publish(data_graph_str: str, ontology: dict) -> None:
|
|||||||
- [Ontology Management](ontology) — generate the OWL ontology that SHACL shapes are derived from
|
- [Ontology Management](ontology) — generate the OWL ontology that SHACL shapes are derived from
|
||||||
- [Reasoning & Rules](reasoning) — complement SHACL structural constraints with logical inference rules
|
- [Reasoning & Rules](reasoning) — complement SHACL structural constraints with logical inference rules
|
||||||
- [Export & Serialization](export) — serialize graph data to Turtle/RDF/XML for `run_shacl_validation` input
|
- [Export & Serialization](export) — serialize graph data to Turtle/RDF/XML for `run_shacl_validation` input
|
||||||
- [Conflict Resolution](/guides/conflict-resolution) — detect and resolve data conflicts before SHACL validation
|
- [Conflict Resolution](conflict-resolution) — detect and resolve data conflicts before SHACL validation
|
||||||
- [Change Management](/guides/change-management) — version-gate SHACL shapes alongside ontology versions
|
- [Change Management](change-management) — version-gate SHACL shapes alongside ontology versions
|
||||||
|
|||||||
@@ -614,8 +614,8 @@ fig.write_html("out.html") # manual export
|
|||||||
|
|
||||||
## Related Guides
|
## Related Guides
|
||||||
|
|
||||||
- [Context Graphs](/guides/context-graphs) — `graph.to_dict()` is the primary input for `KGVisualizer`
|
- [Context Graphs](context-graphs) — `graph.to_dict()` is the primary input for `KGVisualizer`
|
||||||
- [Ontology Management](ontology) — `OntologyVisualizer` renders ontologies produced by `OntologyGenerator`
|
- [Ontology Management](ontology) — `OntologyVisualizer` renders ontologies produced by `OntologyGenerator`
|
||||||
- [Change Management](/guides/change-management) — `TemporalVersionManager` snapshots feed `visualize_metrics_evolution()` and `visualize_snapshot_comparison()`
|
- [Change Management](change-management) — `TemporalVersionManager` snapshots feed `visualize_metrics_evolution()` and `visualize_snapshot_comparison()`
|
||||||
- [Graph Analytics](/guides/graph-analytics) — centrality scores, community dicts, and connectivity results that feed the `AnalyticsVisualizer`
|
- [Graph Analytics](graph-analytics) — centrality scores, community dicts, and connectivity results that feed the `AnalyticsVisualizer`
|
||||||
- [Export & Serialization](export) — export the same graph to GraphML, GEXF, or DOT for Gephi and Graphviz
|
- [Export & Serialization](export) — export the same graph to GraphML, GEXF, or DOT for Gephi and Graphviz
|
||||||
|
|||||||
+14
-14
@@ -185,8 +185,8 @@ decision_id = context.record_decision(
|
|||||||
|
|
||||||
</CodeGroup>
|
</CodeGroup>
|
||||||
|
|
||||||
- [Full Quickstart](/quickstart) — Step-by-step pipeline walkthrough
|
- [Full Quickstart](quickstart) — Step-by-step pipeline walkthrough
|
||||||
- [Cookbook](/cookbook) — 40+ real-world Jupyter notebooks
|
- [Cookbook](cookbook) — 40+ real-world Jupyter notebooks
|
||||||
- [Join Discord](https://discord.gg/sV34vps5hH) — Community chat and support
|
- [Join Discord](https://discord.gg/sV34vps5hH) — Community chat and support
|
||||||
|
|
||||||
|
|
||||||
@@ -195,7 +195,7 @@ decision_id = context.record_decision(
|
|||||||
Semantica was designed for domains where every decision must be explainable and every fact must be traceable.
|
Semantica was designed for domains where every decision must be explainable and every fact must be traceable.
|
||||||
|
|
||||||
<Warning>
|
<Warning>
|
||||||
**This is system-level explainability, not foundation-model explainability.** Semantica does not expose, reconstruct, or explain what happens *inside* the LLM/foundation model — its internal reasoning or chain-of-thought stays opaque, as it does for any external system. What Semantica explains is *outside* the model: the context and data fed in, the decision produced, its provenance, the relevant relationships, the policies applied, and the full execution trail. See [Core Concepts](/concepts) for the full scope note.
|
**This is system-level explainability, not foundation-model explainability.** Semantica does not expose, reconstruct, or explain what happens *inside* the LLM/foundation model — its internal reasoning or chain-of-thought stays opaque, as it does for any external system. What Semantica explains is *outside* the model: the context and data fed in, the decision produced, its provenance, the relevant relationships, the policies applied, and the full execution trail. See [Core Concepts](concepts) for the full scope note.
|
||||||
</Warning>
|
</Warning>
|
||||||
|
|
||||||
**Healthcare & Life Sciences**
|
**Healthcare & Life Sciences**
|
||||||
@@ -242,35 +242,35 @@ Semantica was designed for domains where every decision must be explainable and
|
|||||||
```bash
|
```bash
|
||||||
pip install semantica
|
pip install semantica
|
||||||
```
|
```
|
||||||
See [Installation](/installation) for optional extras (`[all]`, `[neo4j]`, `[pinecone]`) and environment setup.
|
See [Installation](installation) for optional extras (`[all]`, `[neo4j]`, `[pinecone]`) and environment setup.
|
||||||
</Step>
|
</Step>
|
||||||
<Step title="Run the Quickstart">
|
<Step title="Run the Quickstart">
|
||||||
Build a complete knowledge graph pipeline in [5 minutes](/quickstart):
|
Build a complete knowledge graph pipeline in [5 minutes](quickstart):
|
||||||
- Ingest documents from any source
|
- Ingest documents from any source
|
||||||
- Extract entities and relationships
|
- Extract entities and relationships
|
||||||
- Build and query the graph
|
- Build and query the graph
|
||||||
- Record and trace a decision
|
- Record and trace a decision
|
||||||
</Step>
|
</Step>
|
||||||
<Step title="Learn the mental model">
|
<Step title="Learn the mental model">
|
||||||
[Core Concepts](/concepts) covers:
|
[Core Concepts](concepts) covers:
|
||||||
- Knowledge graphs vs. vector stores: when to use each
|
- Knowledge graphs vs. vector stores: when to use each
|
||||||
- What GraphRAG is and how Semantica implements it
|
- What GraphRAG is and how Semantica implements it
|
||||||
- How provenance and decision tracking work together
|
- How provenance and decision tracking work together
|
||||||
- The accountability layer architecture
|
- The accountability layer architecture
|
||||||
</Step>
|
</Step>
|
||||||
<Step title="Go deep on any module">
|
<Step title="Go deep on any module">
|
||||||
Every module has a dedicated [reference page](/reference/context) with:
|
Every module has a dedicated [reference page](reference/context) with:
|
||||||
- Full class and method documentation
|
- Full class and method documentation
|
||||||
- Parameter tables with types and defaults
|
- Parameter tables with types and defaults
|
||||||
- Runnable code examples for each feature
|
- Runnable code examples for each feature
|
||||||
</Step>
|
</Step>
|
||||||
</Steps>
|
</Steps>
|
||||||
|
|
||||||
- [Installation](/installation) — Get Semantica installed in under a minute
|
- [Installation](installation) — Get Semantica installed in under a minute
|
||||||
- [Quickstart](/quickstart) — Build a complete knowledge graph pipeline in 5 minutes
|
- [Quickstart](quickstart) — Build a complete knowledge graph pipeline in 5 minutes
|
||||||
- [Core Concepts](/concepts) — The mental model behind the API
|
- [Core Concepts](concepts) — The mental model behind the API
|
||||||
- [API Reference](/reference/context) — Exact module, class, and method details
|
- [API Reference](reference/context) — Exact module, class, and method details
|
||||||
- [Cookbook](/cookbook) — Domain notebooks for real-world use cases
|
- [Cookbook](cookbook) — Domain notebooks for real-world use cases
|
||||||
- [Changelog](https://github.com/semantica-agi/semantica/releases) — Release history
|
- [Changelog](https://github.com/semantica-agi/semantica/releases) — Release history
|
||||||
|
|
||||||
|
|
||||||
@@ -369,7 +369,7 @@ Semantica was designed for domains where every decision must be explainable and
|
|||||||
| `semantica.reasoning` | Forward chaining, Rete, deductive, abductive, SPARQL, Datalog |
|
| `semantica.reasoning` | Forward chaining, Rete, deductive, abductive, SPARQL, Datalog |
|
||||||
| `semantica.ontology` | SHACL, SKOS, alignments, diff/migration, auto-generation, OWL/RDF |
|
| `semantica.ontology` | SHACL, SKOS, alignments, diff/migration, auto-generation, OWL/RDF |
|
||||||
| `semantica.explorer` | FastAPI Knowledge Explorer, Ontology Hub, Distance Intelligence, SHACL Studio |
|
| `semantica.explorer` | FastAPI Knowledge Explorer, Ontology Hub, Distance Intelligence, SHACL Studio |
|
||||||
| `semantica.mcp_server` | MCP stdio server: 15 tools for Claude Desktop, VS Code, Cursor, Windsurf, Cline |
|
| `semantica.mcp_server` | MCP stdio server: 12 tools for Claude Desktop, VS Code, Cursor, Windsurf, Cline |
|
||||||
| `semantica.vector_store` | FAISS, Pinecone, Weaviate, Qdrant, Milvus, PgVector |
|
| `semantica.vector_store` | FAISS, Pinecone, Weaviate, Qdrant, Milvus, PgVector |
|
||||||
| `semantica.graph_store` | Neo4j, FalkorDB, Apache AGE, Amazon Neptune |
|
| `semantica.graph_store` | Neo4j, FalkorDB, Apache AGE, Amazon Neptune |
|
||||||
| `semantica.triplet_store` | In-memory and persistent RDF triple store with SPARQL |
|
| `semantica.triplet_store` | In-memory and persistent RDF triple store with SPARQL |
|
||||||
@@ -404,7 +404,7 @@ Semantica was designed for domains where every decision must be explainable and
|
|||||||
- 1,000+ passing tests with full regression coverage
|
- 1,000+ passing tests with full regression coverage
|
||||||
- `PipelineValidator` catches configuration errors at startup
|
- `PipelineValidator` catches configuration errors at startup
|
||||||
- `FailureHandler` with exponential backoff and dead-letter queues
|
- `FailureHandler` with exponential backoff and dead-letter queues
|
||||||
- Ongoing security hardening: fixes shipped in every release ([CHANGELOG](https://github.com/semantica-agi/semantica/blob/main/CHANGELOG.md))
|
- 12 security vulnerabilities fixed in v0.5.0
|
||||||
|
|
||||||
**Modular by Design** — Import only what you need.
|
**Modular by Design** — Import only what you need.
|
||||||
- Use `NERExtractor` without a graph store
|
- Use `NERExtractor` without a graph store
|
||||||
|
|||||||
@@ -183,6 +183,6 @@ Install the [Microsoft Visual C++ Redistributable](https://aka.ms/vs/17/release/
|
|||||||
|
|
||||||
## Next Steps
|
## Next Steps
|
||||||
|
|
||||||
- [Getting Started](/getting-started) — Understand what Semantica does before you build.
|
- [Getting Started](getting-started) — Understand what Semantica does before you build.
|
||||||
- [Build the Pipeline](/quickstart) — Follow the end-to-end workflow with code.
|
- [Build the Pipeline](quickstart) — Follow the end-to-end workflow with code.
|
||||||
- [Browse Examples](/cookbook) — See notebook examples organized by use case.
|
- [Browse Examples](cookbook) — See notebook examples organized by use case.
|
||||||
|
|||||||
@@ -193,7 +193,7 @@ if not connector.test_connection():
|
|||||||
## See Also
|
## See Also
|
||||||
|
|
||||||
- [Ingest Module](../reference/ingest) — Full DatabricksIngestor and all other ingestors.
|
- [Ingest Module](../reference/ingest) — Full DatabricksIngestor and all other ingestors.
|
||||||
- [Snowflake Integration](/integrations/snowflake) — Companion connector for a Snowflake + Databricks hybrid estate.
|
- [Snowflake Integration](snowflake) — Companion connector for a Snowflake + Databricks hybrid estate.
|
||||||
- [Pipeline](../reference/pipeline) — Use Databricks ingestion as a pipeline step.
|
- [Pipeline](../reference/pipeline) — Use Databricks ingestion as a pipeline step.
|
||||||
- [Installation](../installation) — All optional dependency extras.
|
- [Installation](../installation) — All optional dependency extras.
|
||||||
- [Knowledge Graph](../reference/kg) — Build a KG from ingested Databricks data.
|
- [Knowledge Graph](../reference/kg) — Build a KG from ingested Databricks data.
|
||||||
|
|||||||
@@ -370,7 +370,7 @@ Common causes of authentication failures:
|
|||||||
## See Also
|
## See Also
|
||||||
|
|
||||||
- [Ingest Module](../reference/ingest) — Full `SalesforceIngestor` API and all other ingestors.
|
- [Ingest Module](../reference/ingest) — Full `SalesforceIngestor` API and all other ingestors.
|
||||||
- [Snowflake Integration](/integrations/snowflake) — Relational warehouse connector with a similar design.
|
- [Snowflake Integration](snowflake) — Relational warehouse connector with a similar design.
|
||||||
- [Databricks Integration](/integrations/databricks) — Lakehouse connector.
|
- [Databricks Integration](databricks) — Lakehouse connector.
|
||||||
- [Installation](../installation) — All optional dependency extras.
|
- [Installation](../installation) — All optional dependency extras.
|
||||||
- [Knowledge Graph](../reference/kg) — Build a KG from ingested Salesforce data.
|
- [Knowledge Graph](../reference/kg) — Build a KG from ingested Salesforce data.
|
||||||
|
|||||||
@@ -172,7 +172,7 @@ if not connector.test_connection():
|
|||||||
## See Also
|
## See Also
|
||||||
|
|
||||||
- [Ingest Module](../reference/ingest) — Full SnowflakeIngestor and all other ingestors.
|
- [Ingest Module](../reference/ingest) — Full SnowflakeIngestor and all other ingestors.
|
||||||
- [Databricks Integration](/integrations/databricks) — Companion connector for a Snowflake + Databricks hybrid estate.
|
- [Databricks Integration](databricks) — Companion connector for a Snowflake + Databricks hybrid estate.
|
||||||
- [Pipeline](../reference/pipeline) — Use Snowflake ingestion as a pipeline step.
|
- [Pipeline](../reference/pipeline) — Use Snowflake ingestion as a pipeline step.
|
||||||
- [Installation](../installation) — All optional dependency extras.
|
- [Installation](../installation) — All optional dependency extras.
|
||||||
- [Knowledge Graph](../reference/kg) — Build a KG from ingested Snowflake data.
|
- [Knowledge Graph](../reference/kg) — Build a KG from ingested Snowflake data.
|
||||||
|
|||||||
+13
-13
@@ -9,9 +9,9 @@ Whether you're running your first pipeline or deploying Semantica in production,
|
|||||||
|
|
||||||
## Learning Paths
|
## Learning Paths
|
||||||
|
|
||||||
- **Beginner (1–2 hrs)** — New to Semantica and knowledge graphs. [Start with Installation →](/installation)
|
- **Beginner (1–2 hrs)** — New to Semantica and knowledge graphs. [Start with Installation →](installation)
|
||||||
- **Intermediate (4–6 hrs)** — Comfortable with basics, building real applications. [Start with Modules →](/modules)
|
- **Intermediate (4–6 hrs)** — Comfortable with basics, building real applications. [Start with Modules →](modules)
|
||||||
- **Advanced (8+ hrs)** — Enterprise deployments, customization, and extension. [Start with Architecture →](/architecture)
|
- **Advanced (8+ hrs)** — Enterprise deployments, customization, and extension. [Start with Architecture →](architecture)
|
||||||
|
|
||||||
<Tabs>
|
<Tabs>
|
||||||
<Tab title="Beginner (1–2 hrs)">
|
<Tab title="Beginner (1–2 hrs)">
|
||||||
@@ -19,16 +19,16 @@ Whether you're running your first pipeline or deploying Semantica in production,
|
|||||||
|
|
||||||
<Steps>
|
<Steps>
|
||||||
<Step title="Set up your environment">
|
<Step title="Set up your environment">
|
||||||
[Installation Guide](/installation): virtual environments, optional extras, platform-specific fixes.
|
[Installation Guide](installation): virtual environments, optional extras, platform-specific fixes.
|
||||||
</Step>
|
</Step>
|
||||||
<Step title="Understand the core ideas">
|
<Step title="Understand the core ideas">
|
||||||
[Core Concepts](/concepts): what knowledge graphs are, how embeddings work, what extraction does.
|
[Core Concepts](concepts): what knowledge graphs are, how embeddings work, what extraction does.
|
||||||
</Step>
|
</Step>
|
||||||
<Step title="Run your first example">
|
<Step title="Run your first example">
|
||||||
[Getting Started](/getting-started): 5-minute code walkthrough with pattern-based extraction (no API key needed).
|
[Getting Started](getting-started): 5-minute code walkthrough with pattern-based extraction (no API key needed).
|
||||||
</Step>
|
</Step>
|
||||||
<Step title="Build your first knowledge graph">
|
<Step title="Build your first knowledge graph">
|
||||||
[Quickstart Tutorial](/quickstart): full 6-step pipeline from ingestion to visualization.
|
[Quickstart Tutorial](quickstart): full 6-step pipeline from ingestion to visualization.
|
||||||
</Step>
|
</Step>
|
||||||
<Step title="Explore interactively">
|
<Step title="Explore interactively">
|
||||||
[Welcome to Semantica notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/01_Welcome_to_Semantica.ipynb): Jupyter walkthrough of every module.
|
[Welcome to Semantica notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/01_Welcome_to_Semantica.ipynb): Jupyter walkthrough of every module.
|
||||||
@@ -40,13 +40,13 @@ Whether you're running your first pipeline or deploying Semantica in production,
|
|||||||
|
|
||||||
<Steps>
|
<Steps>
|
||||||
<Step title="Learn every module">
|
<Step title="Learn every module">
|
||||||
[Modules Guide](/modules): all 27 modules with code examples and common pipeline chains.
|
[Modules Guide](modules): all 27 modules with code examples and common pipeline chains.
|
||||||
</Step>
|
</Step>
|
||||||
<Step title="Build production knowledge graphs">
|
<Step title="Build production knowledge graphs">
|
||||||
[Building Knowledge Graphs notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/07_Building_Knowledge_Graphs.ipynb): multi-source, deduplication, conflict resolution.
|
[Building Knowledge Graphs notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/07_Building_Knowledge_Graphs.ipynb): multi-source, deduplication, conflict resolution.
|
||||||
</Step>
|
</Step>
|
||||||
<Step title="Add semantic search">
|
<Step title="Add semantic search">
|
||||||
[Embedding Generation notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/12_Embedding_Generation.ipynb): generating embeddings, provider and model switching, dimensions. Then [Vector Store notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/13_Vector_Store.ipynb): storing and searching vectors for retrieval.
|
[Embeddings notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/introduction/09_Embeddings.ipynb): providers, pooling strategies, vector stores.
|
||||||
</Step>
|
</Step>
|
||||||
<Step title="Multi-source integration">
|
<Step title="Multi-source integration">
|
||||||
[Multi-Source Data Integration notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/advanced/06_Multi_Source_Data_Integration.ipynb) for multi-source patterns.
|
[Multi-Source Data Integration notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/advanced/06_Multi_Source_Data_Integration.ipynb) for multi-source patterns.
|
||||||
@@ -58,7 +58,7 @@ Whether you're running your first pipeline or deploying Semantica in production,
|
|||||||
|
|
||||||
<Steps>
|
<Steps>
|
||||||
<Step title="Understand the architecture">
|
<Step title="Understand the architecture">
|
||||||
[Architecture Guide](/architecture): four-layer design, extension points, and design decisions.
|
[Architecture Guide](architecture): four-layer design, extension points, and design decisions.
|
||||||
</Step>
|
</Step>
|
||||||
<Step title="Temporal intelligence">
|
<Step title="Temporal intelligence">
|
||||||
[Temporal Graphs notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/advanced/10_Temporal_Knowledge_Graphs.ipynb): `valid_from`/`valid_until`, Allen interval algebra, point-in-time queries.
|
[Temporal Graphs notebook](https://github.com/semantica-agi/semantica/blob/main/cookbook/advanced/10_Temporal_Knowledge_Graphs.ipynb): `valid_from`/`valid_until`, Allen interval algebra, point-in-time queries.
|
||||||
@@ -236,6 +236,6 @@ The `blocking_v2`, `hybrid_v2`, and `semantic_v2` strategies reduce O(n²) compa
|
|||||||
- **Graph exports**: encrypt sensitive exports at rest; use the v0.5.0 SSRF-safe `base_url` validation when configuring custom LLM gateways
|
- **Graph exports**: encrypt sensitive exports at rest; use the v0.5.0 SSRF-safe `base_url` validation when configuring custom LLM gateways
|
||||||
- **XML ingestion**: always use `XMLIngestor` (v0.5.0), which uses the XXE-safe lxml backend; never parse untrusted XML with the standard library parser
|
- **XML ingestion**: always use `XMLIngestor` (v0.5.0), which uses the XXE-safe lxml backend; never parse untrusted XML with the standard library parser
|
||||||
|
|
||||||
- [Cookbook](/cookbook) — Interactive Jupyter notebooks from beginner to advanced.
|
- [Cookbook](cookbook) — Interactive Jupyter notebooks from beginner to advanced.
|
||||||
- [FAQ](/faq) — Common questions answered.
|
- [FAQ](faq) — Common questions answered.
|
||||||
- [API Reference](/reference/core) — Complete technical documentation.
|
- [API Reference](reference/core) — Complete technical documentation.
|
||||||
|
|||||||
+32
-32
@@ -9,7 +9,7 @@ icon: "puzzle-piece"
|
|||||||
</Info>
|
</Info>
|
||||||
|
|
||||||
<Tip>
|
<Tip>
|
||||||
Not sure which module to use? The [Choose the Right Module](/choose-your-module) guide maps 35+ developer goals to modules with code examples — start there if you're orienting for the first time.
|
Not sure which module to use? The [Choose the Right Module](choose-your-module) guide maps 35+ developer goals to modules with code examples — start there if you're orienting for the first time.
|
||||||
</Tip>
|
</Tip>
|
||||||
|
|
||||||
Semantica is organized into **27 modules** across six logical layers. Each module is independently importable: you never pay for what you don't use.
|
Semantica is organized into **27 modules** across six logical layers. Each module is independently importable: you never pay for what you don't use.
|
||||||
@@ -438,7 +438,7 @@ Exposes Semantica as an MCP stdio server for IDE and agent integrations.
|
|||||||
python -m semantica.mcp_server
|
python -m semantica.mcp_server
|
||||||
```
|
```
|
||||||
|
|
||||||
**Integrations:** Claude Desktop, VS Code, Cursor, Windsurf, Cline: 15 MCP tools exposed
|
**Integrations:** Claude Desktop, VS Code, Cursor, Windsurf, Cline: 12 MCP tools exposed
|
||||||
|
|
||||||
### Seed
|
### Seed
|
||||||
|
|
||||||
@@ -680,34 +680,34 @@ versioner.create_snapshot(kg, "2024-Q1", author="user@example.com", description=
|
|||||||
|
|
||||||
| Module | Purpose | Key Classes |
|
| Module | Purpose | Key Classes |
|
||||||
| :------ | :------- | :----------- |
|
| :------ | :------- | :----------- |
|
||||||
| [ingest](/reference/ingest) | Data ingestion | `FileIngestor`, `WebIngestor`, `ParquetIngestor`, `XMLIngestor` |
|
| [ingest](reference/ingest) | Data ingestion | `FileIngestor`, `WebIngestor`, `ParquetIngestor`, `XMLIngestor` |
|
||||||
| [parse](/reference/parse) | Document parsing | `DocumentParser`, `DoclingParser` |
|
| [parse](reference/parse) | Document parsing | `DocumentParser`, `DoclingParser` |
|
||||||
| [split](/reference/split) | Text chunking | `TextSplitter` |
|
| [split](reference/split) | Text chunking | `TextSplitter` |
|
||||||
| [normalize](/reference/normalize) | Data cleaning | `TextNormalizer`, `EntityNormalizer`, `LanguageDetector` |
|
| [normalize](reference/normalize) | Data cleaning | `TextNormalizer`, `EntityNormalizer`, `LanguageDetector` |
|
||||||
| [semantic_extract](/reference/semantic_extract) | NER & relation extraction | `NERExtractor`, `RelationExtractor`, `TripletExtractor`, `SemanticAnalyzer`, `SemanticNetworkExtractor`, `ExtractionValidator` |
|
| [semantic_extract](reference/semantic_extract) | NER & relation extraction | `NERExtractor`, `RelationExtractor`, `TripletExtractor`, `SemanticAnalyzer`, `SemanticNetworkExtractor`, `ExtractionValidator` |
|
||||||
| [kg](/reference/kg) | Graph construction | `GraphBuilder`, `TemporalGraphQuery`, `SimilarityCalculator` |
|
| [kg](reference/kg) | Graph construction | `GraphBuilder`, `TemporalGraphQuery`, `SimilarityCalculator` |
|
||||||
| [ontology](/reference/ontology) | Schema management | `OntologyGenerator`, `SHACLGenerator` |
|
| [ontology](reference/ontology) | Schema management | `OntologyGenerator`, `SHACLGenerator` |
|
||||||
| [reasoning](/reference/reasoning) | Logical inference | `Reasoner`, `DatalogReasoner` |
|
| [reasoning](reference/reasoning) | Logical inference | `Reasoner`, `DatalogReasoner` |
|
||||||
| [embeddings](/reference/embeddings) | Vector embeddings | `EmbeddingGenerator` |
|
| [embeddings](reference/embeddings) | Vector embeddings | `EmbeddingGenerator` |
|
||||||
| [vector_store](/reference/vector_store) | Vector database | `VectorStore` |
|
| [vector_store](reference/vector_store) | Vector database | `VectorStore` |
|
||||||
| [graph_store](/reference/graph_store) | Graph database | `GraphStore` |
|
| [graph_store](reference/graph_store) | Graph database | `GraphStore` |
|
||||||
| [triplet_store](/reference/triplet_store) | RDF triple store | `TripletStore` |
|
| [triplet_store](reference/triplet_store) | RDF triple store | `TripletStore` |
|
||||||
| [deduplication](/reference/deduplication) | Entity resolution | `EntityResolver`, `DuplicateDetector`, `ClusterBuilder`, `MergeStrategyManager` |
|
| [deduplication](reference/deduplication) | Entity resolution | `EntityResolver`, `DuplicateDetector`, `ClusterBuilder`, `MergeStrategyManager` |
|
||||||
| [conflicts](/reference/conflicts) | Conflict resolution | `ConflictDetector` |
|
| [conflicts](reference/conflicts) | Conflict resolution | `ConflictDetector` |
|
||||||
| [context](/reference/context) | Agent context & decisions | `AgentContext`, `ContextGraph` |
|
| [context](reference/context) | Agent context & decisions | `AgentContext`, `ContextGraph` |
|
||||||
| [provenance](/reference/provenance) | W3C PROV-O lineage | `ProvenanceManager` |
|
| [provenance](reference/provenance) | W3C PROV-O lineage | `ProvenanceManager` |
|
||||||
| [change_management](/reference/change_management) | Version control | `TemporalVersionManager` |
|
| [change_management](reference/change_management) | Version control | `TemporalVersionManager` |
|
||||||
| [export](/reference/export) | Data export | `RDFExporter`, `ParquetExporter` |
|
| [export](reference/export) | Data export | `RDFExporter`, `ParquetExporter` |
|
||||||
| [visualization](/reference/visualization) | Graph visualization | `KGVisualizer` |
|
| [visualization](reference/visualization) | Graph visualization | `KGVisualizer` |
|
||||||
| [pipeline](/reference/pipeline) | Workflow orchestration | `Pipeline`, `PipelineBuilder` |
|
| [pipeline](reference/pipeline) | Workflow orchestration | `Pipeline`, `PipelineBuilder` |
|
||||||
| [explorer](/reference/explorer) | Knowledge Explorer UI | `semantica-explorer --graph <file>` |
|
| [explorer](reference/explorer) | Knowledge Explorer UI | `semantica-explorer --graph <file>` |
|
||||||
| [llms](/reference/llms) | LLM providers | `Groq`, `OpenAI`, `create_provider` |
|
| [llms](reference/llms) | LLM providers | `Groq`, `OpenAI`, `create_provider` |
|
||||||
| [mcp_server](/reference/mcp_server) | MCP stdio server | `python -m semantica.mcp_server` |
|
| [mcp_server](reference/mcp_server) | MCP stdio server | `python -m semantica.mcp_server` |
|
||||||
| [seed](/reference/seed) | KG bootstrapping from structured sources | `SeedManager` |
|
| [seed](reference/seed) | KG bootstrapping from structured sources | `SeedManager` |
|
||||||
| [evals](/reference/evals) | Quality evaluation | `KGEvaluator`, `ExtractionEvaluator`, `PipelineEvaluator`, `RegressionTracker` |
|
| [evals](reference/evals) | Quality evaluation | `KGEvaluator`, `ExtractionEvaluator`, `PipelineEvaluator`, `RegressionTracker` |
|
||||||
| [core](/reference/core) | Base classes & registry | `Semantica`, `ConfigManager`, `PluginRegistry`, `LifecycleManager` |
|
| [core](reference/core) | Base classes & registry | `Semantica`, `ConfigManager`, `PluginRegistry`, `LifecycleManager` |
|
||||||
| [utils](/reference/utils) | Shared utilities | `helpers`, `validators` |
|
| [utils](reference/utils) | Shared utilities | `helpers`, `validators` |
|
||||||
|
|
||||||
- [Getting Started](/getting-started) — Your first knowledge graph in 5 minutes.
|
- [Getting Started](getting-started) — Your first knowledge graph in 5 minutes.
|
||||||
- [Cookbook](/cookbook) — 40+ domain notebooks with real-world examples.
|
- [Cookbook](cookbook) — 40+ domain notebooks with real-world examples.
|
||||||
- [API Reference](/reference/context) — Full technical documentation.
|
- [API Reference](reference/context) — Full technical documentation.
|
||||||
|
|||||||
@@ -76,5 +76,5 @@ By contributing to Semantica, you agree that your contributions will be licensed
|
|||||||
|
|
||||||
## See Also
|
## See Also
|
||||||
|
|
||||||
- [Contributing](/contributing-guide) — How to contribute to the project.
|
- [Contributing](contributing-guide) — How to contribute to the project.
|
||||||
- [Citation](/citation) — How to cite Semantica in research.
|
- [Citation](citation) — How to cite Semantica in research.
|
||||||
|
|||||||
+66
-92
@@ -5,7 +5,7 @@ icon: "rocket"
|
|||||||
---
|
---
|
||||||
|
|
||||||
<Info>
|
<Info>
|
||||||
**v0.6.7** — first-class LangChain integration, SAP OData ingestor, human-editable Markdown persistence for `ContextGraph`, and a structured Action layer for the reasoning engine. <a href="https://github.com/semantica-agi/semantica/releases" style={{color:"#10B981",fontWeight:600,textDecoration:"none"}}>What's new →</a>
|
**v0.5.0** — Ontology Hub, Distance Intelligence, Parquet & XML ingestion, 12 security fixes. <a href="https://github.com/semantica-agi/semantica/releases" style={{color:"#10B981",fontWeight:600,textDecoration:"none"}}>What's new →</a>
|
||||||
</Info>
|
</Info>
|
||||||
|
|
||||||
This guide walks you through the end-to-end pipeline for building your first knowledge graph. Start here after installation. An LLM API key is optional: pattern-based extraction works out of the box.
|
This guide walks you through the end-to-end pipeline for building your first knowledge graph. Start here after installation. An LLM API key is optional: pattern-based extraction works out of the box.
|
||||||
@@ -35,7 +35,7 @@ Verify:
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
python -c "import semantica; print(semantica.__version__)"
|
python -c "import semantica; print(semantica.__version__)"
|
||||||
# 0.6.7
|
# 0.5.0
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|
||||||
@@ -47,24 +47,36 @@ python -c "import semantica; print(semantica.__version__)"
|
|||||||
|
|
||||||
<Step title="Ingest">
|
<Step title="Ingest">
|
||||||
|
|
||||||
Load a document from a file or directory. The rest of this walkthrough follows
|
Load a document from a file, directory, URL, or database.
|
||||||
the file path; other sources are shown afterwards.
|
|
||||||
|
|
||||||
```python
|
<CodeGroup>
|
||||||
|
|
||||||
|
```python File
|
||||||
from semantica.ingest import FileIngestor
|
from semantica.ingest import FileIngestor
|
||||||
|
|
||||||
ingestor = FileIngestor()
|
ingestor = FileIngestor()
|
||||||
sources = ingestor.ingest("data/report.pdf")
|
sources = ingestor.ingest("data/report.pdf")
|
||||||
# Also accepts a directory, .docx, .html, .json, .csv, .xlsx, .pptx, .parquet, .xml
|
# Also accepts: .docx, .html, .json, .csv, .xlsx, .pptx, .parquet, .xml
|
||||||
```
|
```
|
||||||
|
|
||||||
<Tip>
|
```python Web
|
||||||
**Other sources.** `WebIngestor().ingest_url(url)` returns a `WebContent` whose
|
from semantica.ingest import WebIngestor
|
||||||
`.text` you can feed straight into the Extract step (no parsing needed).
|
|
||||||
`ParquetIngestor().ingest(path)` and `XMLIngestor().ingest(path, schema_path=...)`
|
ingestor = WebIngestor(max_depth=2)
|
||||||
return structured records rather than documents; build a graph from those with
|
sources = ingestor.ingest("https://example.com/article")
|
||||||
`GraphBuilder().build({"entities": [...], "relationships": [...]})` directly.
|
```
|
||||||
</Tip>
|
|
||||||
|
```python Parquet / XML (v0.5.0)
|
||||||
|
from semantica.ingest import ParquetIngestor, XMLIngestor
|
||||||
|
|
||||||
|
# Single file or Hive-partitioned directory
|
||||||
|
sources = ParquetIngestor().ingest("data/events.parquet")
|
||||||
|
|
||||||
|
# XML with XSD schema validation
|
||||||
|
sources = XMLIngestor(validate_xsd="schema.xsd").ingest("data/records/")
|
||||||
|
```
|
||||||
|
|
||||||
|
</CodeGroup>
|
||||||
|
|
||||||
</Step>
|
</Step>
|
||||||
|
|
||||||
@@ -76,26 +88,22 @@ Extract structured text and layout from raw documents.
|
|||||||
from semantica.parse import DocumentParser
|
from semantica.parse import DocumentParser
|
||||||
|
|
||||||
parser = DocumentParser()
|
parser = DocumentParser()
|
||||||
parsed = parser.parse(sources[0].path) # parse() takes a path string
|
parsed = parser.parse(sources[0])
|
||||||
|
|
||||||
print(parsed["full_text"][:200]) # extracted text
|
print(parsed.text[:200]) # extracted text
|
||||||
print(parsed["metadata"]) # document properties (fields vary by format)
|
print(parsed.metadata) # title, author, date, source
|
||||||
```
|
```
|
||||||
|
|
||||||
`parse()` returns a `dict`. `full_text` and `metadata` are present for every
|
|
||||||
format; other keys depend on the parser (`pages` for PDF, `tables` and
|
|
||||||
`paragraphs` for DOCX, `tables` for `DoclingParser`).
|
|
||||||
|
|
||||||
<Tip>
|
<Tip>
|
||||||
For PDFs with tables, charts, or multi-column layouts, use `DoclingParser` (`pip install semantica[parse-docling]`): it applies advanced layout analysis and returns structured table data alongside text.
|
For PDFs with tables, charts, or multi-column layouts, use `DoclingParser`: it applies advanced layout analysis and returns structured table data alongside text.
|
||||||
</Tip>
|
</Tip>
|
||||||
|
|
||||||
```python
|
```python
|
||||||
from semantica.parse import DoclingParser
|
from semantica.parse import DoclingParser
|
||||||
|
|
||||||
parser = DoclingParser()
|
parser = DoclingParser()
|
||||||
parsed = parser.parse(sources[0].path)
|
parsed = parser.parse(sources[0])
|
||||||
print(parsed["tables"]) # structured table data
|
print(parsed.tables) # structured table objects
|
||||||
```
|
```
|
||||||
|
|
||||||
</Step>
|
</Step>
|
||||||
@@ -109,28 +117,26 @@ Identify named entities and extract typed relationships between them.
|
|||||||
```python Pattern-based (fast, no API key)
|
```python Pattern-based (fast, no API key)
|
||||||
from semantica.semantic_extract import NERExtractor, RelationExtractor
|
from semantica.semantic_extract import NERExtractor, RelationExtractor
|
||||||
|
|
||||||
text = parsed["full_text"]
|
|
||||||
|
|
||||||
ner = NERExtractor(method="pattern")
|
ner = NERExtractor(method="pattern")
|
||||||
entities = ner.extract(text)
|
entities = ner.extract(parsed)
|
||||||
# Returns: [Entity(text="Apple Inc.", label="ORG", start_char=0, end_char=10, confidence=0.7), ...]
|
# Returns: [{"text": "Apple Inc.", "type": "ORGANIZATION", "confidence": 0.98}, ...]
|
||||||
|
|
||||||
rel = RelationExtractor(method="pattern")
|
rel = RelationExtractor(method="rule")
|
||||||
relationships = rel.extract(text, entities=entities)
|
relationships = rel.extract(parsed, entities=entities)
|
||||||
# Returns: [Relation(subject=Entity(...), predicate="founded_by", object=Entity(...), confidence=0.7), ...]
|
# Returns: [{"subject": "Steve Jobs", "predicate": "founded", "object": "Apple Inc."}, ...]
|
||||||
```
|
```
|
||||||
|
|
||||||
```python LLM-powered (higher accuracy)
|
```python LLM-powered (higher accuracy)
|
||||||
from semantica.semantic_extract import NERExtractor, RelationExtractor
|
from semantica.semantic_extract import NERExtractor, RelationExtractor
|
||||||
|
from semantica.llms import Groq
|
||||||
|
|
||||||
# Reads GROQ_API_KEY from the environment; provider/llm_model select the backend
|
llm = Groq(model="llama-3.3-70b-versatile")
|
||||||
text = parsed["full_text"]
|
|
||||||
|
|
||||||
ner = NERExtractor(method="llm", provider="groq", llm_model="llama-3.3-70b-versatile")
|
ner = NERExtractor(method="llm", llm_provider=llm)
|
||||||
entities = ner.extract(text)
|
entities = ner.extract(parsed)
|
||||||
|
|
||||||
rel = RelationExtractor(method="llm", provider="groq", llm_model="llama-3.3-70b-versatile")
|
rel = RelationExtractor(method="llm", llm_provider=llm)
|
||||||
relationships = rel.extract(text, entities=entities)
|
relationships = rel.extract(parsed, entities=entities)
|
||||||
```
|
```
|
||||||
|
|
||||||
</CodeGroup>
|
</CodeGroup>
|
||||||
@@ -192,17 +198,16 @@ exporter.export(graph, file_path="graph.nt", format="nt")
|
|||||||
from semantica.export import ParquetExporter
|
from semantica.export import ParquetExporter
|
||||||
|
|
||||||
exporter = ParquetExporter()
|
exporter = ParquetExporter()
|
||||||
exporter.export(graph, file_path="output/graph")
|
exporter.export(graph, file_path="output/graph.parquet")
|
||||||
# Dict input writes one file per key: output/graph_entities.parquet and
|
# Writes nodes.parquet + edges.parquet: ready for Spark, BigQuery, Databricks
|
||||||
# output/graph_relationships.parquet: ready for Spark, BigQuery, Databricks
|
|
||||||
```
|
```
|
||||||
|
|
||||||
```python ArangoDB
|
```python ArangoDB
|
||||||
from semantica.export import ArangoAQLExporter
|
from semantica.export import ArangoAQLExporter
|
||||||
|
|
||||||
exporter = ArangoAQLExporter()
|
exporter = ArangoAQLExporter()
|
||||||
exporter.export(graph, file_path="graph.aql")
|
aql = exporter.export(graph)
|
||||||
# Writes ready-to-run AQL INSERT statements to graph.aql
|
# Returns ready-to-run AQL INSERT statements
|
||||||
```
|
```
|
||||||
|
|
||||||
</CodeGroup>
|
</CodeGroup>
|
||||||
@@ -267,21 +272,14 @@ relationships = rel.extract(text, entities=entities)
|
|||||||
<Accordion title="Multi-source incremental graph build" icon="layer-group">
|
<Accordion title="Multi-source incremental graph build" icon="layer-group">
|
||||||
|
|
||||||
```python
|
```python
|
||||||
from semantica.ingest import FileIngestor
|
|
||||||
from semantica.parse import DocumentParser
|
|
||||||
from semantica.semantic_extract import NERExtractor, RelationExtractor
|
|
||||||
from semantica.kg import GraphBuilder
|
from semantica.kg import GraphBuilder
|
||||||
|
|
||||||
parser = DocumentParser()
|
builder = GraphBuilder(merge_entities=True)
|
||||||
ner = NERExtractor(method="pattern")
|
|
||||||
rel = RelationExtractor(method="pattern")
|
|
||||||
builder = GraphBuilder(merge_entities=True)
|
|
||||||
|
|
||||||
all_entities, all_rels = [], []
|
all_entities, all_rels = [], []
|
||||||
for source in FileIngestor().ingest("data/reports/"):
|
|
||||||
text = parser.parse(source.path)["full_text"]
|
for doc in parsed_docs:
|
||||||
entities = ner.extract(text)
|
entities = ner.extract(doc)
|
||||||
rels = rel.extract(text, entities=entities)
|
rels = rel.extract(doc, entities=entities)
|
||||||
all_entities.extend(entities)
|
all_entities.extend(entities)
|
||||||
all_rels.extend(rels)
|
all_rels.extend(rels)
|
||||||
|
|
||||||
@@ -329,11 +327,10 @@ print(f"Relationships active in 2023: {result_2023['num_relationships']}")
|
|||||||
<Accordion title="Persistent graph store: Neo4j, FalkorDB, Apache AGE" icon="database">
|
<Accordion title="Persistent graph store: Neo4j, FalkorDB, Apache AGE" icon="database">
|
||||||
|
|
||||||
```python
|
```python
|
||||||
from semantica.graph_store import GraphStore
|
from semantica.graph_store import Neo4jStore
|
||||||
from semantica.kg import GraphBuilder
|
from semantica.kg import GraphBuilder
|
||||||
|
|
||||||
store = GraphStore(
|
store = Neo4jStore(
|
||||||
backend="neo4j",
|
|
||||||
uri="bolt://localhost:7687",
|
uri="bolt://localhost:7687",
|
||||||
user="neo4j",
|
user="neo4j",
|
||||||
password="password",
|
password="password",
|
||||||
@@ -361,8 +358,7 @@ graph = builder.build({"entities": entities, "relationships": relationships})
|
|||||||
# Retrieve full lineage for any entity
|
# Retrieve full lineage for any entity
|
||||||
sources = prov.get_all_sources("Apple Inc.")
|
sources = prov.get_all_sources("Apple Inc.")
|
||||||
print(sources[0])
|
print(sources[0])
|
||||||
# {"source": "data/report.pdf", "location": None, "timestamp": "...",
|
# {"source": "data/report.pdf", "location": None, "timestamp": "...", "confidence": 0.98}
|
||||||
# "confidence": 1.0, "metadata": {"confidence": 0.98}}
|
|
||||||
```
|
```
|
||||||
|
|
||||||
</Accordion>
|
</Accordion>
|
||||||
@@ -376,54 +372,32 @@ print(sources[0])
|
|||||||
|
|
||||||
<Accordion title="No entities extracted" icon="magnifying-glass">
|
<Accordion title="No entities extracted" icon="magnifying-glass">
|
||||||
|
|
||||||
The document likely contains scanned images rather than machine-readable text. `DocumentParser` warns when a PDF has no text layer; switch to `DoclingParser` with OCR enabled:
|
The document likely contains scanned images rather than machine-readable text. Enable OCR:
|
||||||
|
|
||||||
```python
|
```python
|
||||||
from semantica.parse import DoclingParser # pip install semantica[parse-docling]
|
from semantica.parse import DocumentParser
|
||||||
|
|
||||||
parser = DoclingParser(enable_ocr=True)
|
parser = DocumentParser(ocr=True) # enables Tesseract OCR
|
||||||
parsed = parser.parse(sources[0].path)
|
parsed = parser.parse(sources[0])
|
||||||
```
|
```
|
||||||
|
|
||||||
</Accordion>
|
</Accordion>
|
||||||
|
|
||||||
<Accordion title="Slow processing on large corpora" icon="gauge">
|
<Accordion title="Slow processing on large corpora" icon="gauge">
|
||||||
|
|
||||||
Install the GPU extras so embedding and ML inference run on CUDA:
|
Enable parallel processing and GPU acceleration:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
pip install semantica[gpu]
|
pip install semantica[gpu]
|
||||||
```
|
```
|
||||||
|
|
||||||
Scan the directory for paths first (no file contents are read), then handle one
|
|
||||||
document at a time and write to a persistent graph backend instead of the
|
|
||||||
in-memory graph:
|
|
||||||
|
|
||||||
```python
|
```python
|
||||||
from semantica.ingest import FileIngestor
|
from semantica.pipeline import Pipeline
|
||||||
from semantica.parse import DocumentParser
|
|
||||||
from semantica.semantic_extract import NERExtractor, RelationExtractor
|
|
||||||
from semantica.graph_store import GraphStore
|
|
||||||
from semantica.kg import GraphBuilder
|
|
||||||
|
|
||||||
ingestor = FileIngestor()
|
pipeline = Pipeline(workers=8, batch_size=32)
|
||||||
parser = DocumentParser()
|
pipeline.run(sources)
|
||||||
ner = NERExtractor(method="pattern")
|
|
||||||
rel = RelationExtractor(method="pattern")
|
|
||||||
store = GraphStore(backend="neo4j", uri="bolt://localhost:7687",
|
|
||||||
user="neo4j", password="password")
|
|
||||||
builder = GraphBuilder(merge_entities=True, graph_store=store)
|
|
||||||
|
|
||||||
for info in ingestor.scan_directory("data/reports/", recursive=True):
|
|
||||||
text = parser.parse(info["path"])["full_text"] # one document loaded at a time
|
|
||||||
entities = ner.extract(text)
|
|
||||||
rels = rel.extract(text, entities=entities)
|
|
||||||
builder.build({"entities": entities, "relationships": rels})
|
|
||||||
```
|
```
|
||||||
|
|
||||||
For multi-step orchestration with configurable parallelism, see the
|
|
||||||
[Pipeline guide](/guides/pipeline).
|
|
||||||
|
|
||||||
</Accordion>
|
</Accordion>
|
||||||
|
|
||||||
<Accordion title="Memory errors on large graphs" icon="memory">
|
<Accordion title="Memory errors on large graphs" icon="memory">
|
||||||
@@ -454,7 +428,7 @@ pip install --upgrade semantica
|
|||||||
|
|
||||||
## Next Steps
|
## Next Steps
|
||||||
|
|
||||||
- [Core Concepts](/concepts) — Knowledge graphs, ontologies, reasoning engines: the mental model behind Semantica.
|
- [Core Concepts](concepts) — Knowledge graphs, ontologies, reasoning engines: the mental model behind Semantica.
|
||||||
- [Module Reference](/modules) — Every module explained with key classes and common chains.
|
- [Module Reference](modules) — Every module explained with key classes and common chains.
|
||||||
- [API Reference](/reference/context) — Complete documentation for every module, class, and parameter.
|
- [API Reference](reference/context) — Complete documentation for every module, class, and parameter.
|
||||||
- [Cookbook](/cookbook) — 40+ interactive Jupyter notebooks with real-world datasets.
|
- [Cookbook](cookbook) — 40+ interactive Jupyter notebooks with real-world datasets.
|
||||||
|
|||||||
@@ -351,6 +351,6 @@ for record in history:
|
|||||||
</AccordionGroup>
|
</AccordionGroup>
|
||||||
|
|
||||||
- [Provenance](provenance) — W3C PROV-O lineage tracking.
|
- [Provenance](provenance) — W3C PROV-O lineage tracking.
|
||||||
- [Knowledge Graph](/reference/kg) — The graph being versioned.
|
- [Knowledge Graph](kg) — The graph being versioned.
|
||||||
- [Export](export) — Export versioned snapshots.
|
- [Export](export) — Export versioned snapshots.
|
||||||
- [Conflicts](/reference/conflicts) — Detect conflicts introduced between versions.
|
- [Conflicts](conflicts) — Detect conflicts introduced between versions.
|
||||||
|
|||||||
@@ -453,4 +453,4 @@ class InvestigationStep:
|
|||||||
- [Deduplication](deduplication) — Resolve duplicate entities before conflict detection.
|
- [Deduplication](deduplication) — Resolve duplicate entities before conflict detection.
|
||||||
- [Ontology](ontology) — Logical conflicts use SHACL shapes and ontology axioms.
|
- [Ontology](ontology) — Logical conflicts use SHACL shapes and ontology axioms.
|
||||||
- [Provenance](provenance) — Track which source each conflicting fact came from.
|
- [Provenance](provenance) — Track which source each conflicting fact came from.
|
||||||
- [Knowledge Graph](/reference/kg) — The graph being checked for conflicts.
|
- [Knowledge Graph](kg) — The graph being checked for conflicts.
|
||||||
|
|||||||
@@ -25,7 +25,6 @@ icon: "brain"
|
|||||||
| `DecisionRecorder` | Record decisions with embeddings, causal chains, and metadata |
|
| `DecisionRecorder` | Record decisions with embeddings, causal chains, and metadata |
|
||||||
| `PolicyEngine` | Policy management: `add_policy()`, `check_compliance()`, `get_applicable_policies()` |
|
| `PolicyEngine` | Policy management: `add_policy()`, `check_compliance()`, `get_applicable_policies()` |
|
||||||
| `CausalChainAnalyzer` | Trace how decisions influenced each other: `get_causal_chain(decision_id)` |
|
| `CausalChainAnalyzer` | Trace how decisions influenced each other: `get_causal_chain(decision_id)` |
|
||||||
| `ErasureCoordinator` | Erase an entity across graph, memory, and vector store, returning an auditable `ErasureReceipt` |
|
|
||||||
|
|
||||||
|
|
||||||
## What You Get
|
## What You Get
|
||||||
@@ -449,7 +448,7 @@ print("Nodes: {}, Edges: {}".format(stats["node_count"], stats["edge_count"]))
|
|||||||
`ContextGraph` exposes a full Distance Intelligence API for exploring semantic neighborhoods and blending proximity into retrieval.
|
`ContextGraph` exposes a full Distance Intelligence API for exploring semantic neighborhoods and blending proximity into retrieval.
|
||||||
|
|
||||||
<Info>
|
<Info>
|
||||||
Full Distance Intelligence reference — distance matrices, API endpoints, embedding cache, Explorer UI — is covered in the dedicated [Distance Intelligence](/reference/distance) page. This section documents the context-layer API.
|
Full Distance Intelligence reference — distance matrices, API endpoints, embedding cache, Explorer UI — is covered in the dedicated [Distance Intelligence](distance) page. This section documents the context-layer API.
|
||||||
</Info>
|
</Info>
|
||||||
|
|
||||||
### Neighbors with Distance Metadata
|
### Neighbors with Distance Metadata
|
||||||
@@ -635,100 +634,6 @@ queried together safely. Vector-store writes are deferred until the in-memory im
|
|||||||
commits; adapter synchronization remains best-effort and logs failures.
|
commits; adapter synchronization remains best-effort and logs failures.
|
||||||
|
|
||||||
|
|
||||||
## ErasureCoordinator
|
|
||||||
|
|
||||||
`ContextGraph.purge_node()` is scoped to one graph: the node is removed and a
|
|
||||||
tombstone is written, but the same content can still be live as an `AgentMemory`
|
|
||||||
item and as an embedding in the vector store. `ErasureCoordinator` drives the
|
|
||||||
cascade across every bound store and returns an `ErasureReceipt` recording what
|
|
||||||
each one reported.
|
|
||||||
|
|
||||||
```python
|
|
||||||
from semantica.context import AgentMemory, ContextGraph, ErasureCoordinator
|
|
||||||
|
|
||||||
coordinator = ErasureCoordinator(graph=graph, memory=memory)
|
|
||||||
|
|
||||||
receipt = coordinator.erase_entity(
|
|
||||||
"customer-4471",
|
|
||||||
reason="GDPR Art. 17 request #882",
|
|
||||||
)
|
|
||||||
|
|
||||||
if not receipt.complete:
|
|
||||||
# These stores may still hold the entity; handle them out of band.
|
|
||||||
print(receipt.incomplete_stores)
|
|
||||||
```
|
|
||||||
|
|
||||||
<Warning>
|
|
||||||
Check the receipt — the call returning is not proof the data is gone. FAISS,
|
|
||||||
Milvus, and Weaviate expose no delete method, so erasure cannot be completed on
|
|
||||||
those backends today; the receipt reports `unsupported` rather than a success it
|
|
||||||
did not achieve.
|
|
||||||
</Warning>
|
|
||||||
|
|
||||||
### Constructor Parameters
|
|
||||||
|
|
||||||
| Parameter | Type | Default | Description |
|
|
||||||
| :--- | :--- | :--- | :--- |
|
|
||||||
| `graph` | `ContextGraph` | `None` | Anything exposing `purge_node()` |
|
|
||||||
| `memory` | `AgentMemory` | `None` | Anything exposing `find_by_entity()` and `batch_delete()` |
|
|
||||||
| `vector_store` | `VectorStore` | `memory.vector_store` | Store holding entity-keyed embeddings; pass `False` to disable the leg |
|
|
||||||
|
|
||||||
At least one store is required; a store that is not supplied reports
|
|
||||||
`not_configured` rather than being silently skipped.
|
|
||||||
|
|
||||||
### Methods
|
|
||||||
|
|
||||||
| Method | Returns | Description |
|
|
||||||
| :--- | :--- | :--- |
|
|
||||||
| `erase_entity(entity_id, reason, at, vector_ids)` | `ErasureReceipt` | Erase one entity from every bound store |
|
|
||||||
| `erase_entities(entity_ids, reason, at)` | `List[ErasureReceipt]` | One receipt per entity, in order; one failure does not stop the rest |
|
|
||||||
|
|
||||||
### Store Statuses
|
|
||||||
|
|
||||||
| Status | Meaning |
|
|
||||||
| :--- | :--- |
|
|
||||||
| `erased` | Reached, data removed. On the vectors leg this means the store accepted the delete for the ids given — backends offer no portable existence check, so it is not a count of embeddings that were really there |
|
|
||||||
| `not_found` | Reached, held nothing for this entity |
|
|
||||||
| `not_configured` | No such store was bound — normal, not a failure |
|
|
||||||
| `unsupported` | The store cannot delete at all; retrying will not help |
|
|
||||||
| `failed` | The store was reached and the deletion did not succeed |
|
|
||||||
|
|
||||||
### ErasureReceipt
|
|
||||||
|
|
||||||
| Member | Type | Description |
|
|
||||||
| :--- | :--- | :--- |
|
|
||||||
| `entity_id` | `str` | Entity the erasure was requested for |
|
|
||||||
| `reason` | `Optional[str]` | Recorded in the receipt and the graph tombstone |
|
|
||||||
| `erased_at` | `str` | ISO-8601; matches the tombstone's `purged_at` |
|
|
||||||
| `stores` | `Dict[str, Dict]` | Per-store outcome keyed `vectors`, `memory`, `graph` |
|
|
||||||
| `complete` | `bool` | `False` when any store reports `unsupported` or `failed` |
|
|
||||||
| `incomplete_stores` | `List[str]` | Stores that may still hold the entity's data |
|
|
||||||
| `to_dict()` | `Dict` | Serialized receipt, safe to persist as an audit record |
|
|
||||||
|
|
||||||
```python
|
|
||||||
receipt.to_dict()
|
|
||||||
# {
|
|
||||||
# "entity_id": "customer-4471",
|
|
||||||
# "reason": "GDPR Art. 17 request #882",
|
|
||||||
# "erased_at": "2026-08-16T09:03:36.813220",
|
|
||||||
# "complete": False,
|
|
||||||
# "stores": {
|
|
||||||
# "vectors": {"status": "unsupported", "backend": "faiss",
|
|
||||||
# "detail": "backend exposes no delete()/delete_vectors(); ..."},
|
|
||||||
# "memory": {"status": "erased", "items": 14},
|
|
||||||
# "graph": {"status": "erased", "nodes": 1, "edges": 3},
|
|
||||||
# },
|
|
||||||
# }
|
|
||||||
```
|
|
||||||
|
|
||||||
Erasure runs outward-in — vectors, then memory, then the graph. The tombstone is
|
|
||||||
the durable attestation that an erasure happened, so it is written last: a crash
|
|
||||||
mid-cascade leaves the node present and the receipt incomplete, rather than a
|
|
||||||
tombstone claiming more than actually happened. A store that raises is recorded
|
|
||||||
as `failed` and the remaining stores are still erased. Erasing the same entity
|
|
||||||
twice returns a receipt saying there was nothing left to do rather than raising.
|
|
||||||
|
|
||||||
|
|
||||||
## PolicyEngine
|
## PolicyEngine
|
||||||
|
|
||||||
`PolicyEngine` manages versioned policies stored in the knowledge graph. Policies are stored as nodes and can be linked to decisions:
|
`PolicyEngine` manages versioned policies stored in the knowledge graph. Policies are stored as nodes and can be linked to decisions:
|
||||||
@@ -1087,8 +992,8 @@ class EntityLink:
|
|||||||
</Tab>
|
</Tab>
|
||||||
</Tabs>
|
</Tabs>
|
||||||
|
|
||||||
- [Vector Store](/reference/vector_store) — Embedding storage backend for memory retrieval.
|
- [Vector Store](vector_store) — Embedding storage backend for memory retrieval.
|
||||||
- [Knowledge Graph](/reference/kg) — Graph algorithms and analytics used inside ContextGraph.
|
- [Knowledge Graph](kg) — Graph algorithms and analytics used inside ContextGraph.
|
||||||
- [Reasoning](reasoning) — Logical inference layered on top of context.
|
- [Reasoning](reasoning) — Logical inference layered on top of context.
|
||||||
- [Provenance](provenance) — W3C PROV-O lineage for every stored fact.
|
- [Provenance](provenance) — W3C PROV-O lineage for every stored fact.
|
||||||
|
|
||||||
|
|||||||
@@ -227,6 +227,6 @@ result = build_knowledge_base(sources=["doc.pdf"], method="fast")
|
|||||||
</Tip>
|
</Tip>
|
||||||
|
|
||||||
- [Pipeline](pipeline) — Pipeline execution and step orchestration.
|
- [Pipeline](pipeline) — Pipeline execution and step orchestration.
|
||||||
- [Utils](/reference/utils) — Shared utilities used by Core internally.
|
- [Utils](utils) — Shared utilities used by Core internally.
|
||||||
- [Getting Started](../getting-started) — Learn the basics before using Core.
|
- [Getting Started](../getting-started) — Learn the basics before using Core.
|
||||||
- [LLMs](/reference/llms) — Configure LLM providers via ConfigManager.
|
- [LLMs](llms) — Configure LLM providers via ConfigManager.
|
||||||
|
|||||||
@@ -437,7 +437,7 @@ result = calculate_similarity(entity_a, entity_b, method="drug_name")
|
|||||||
</Tab>
|
</Tab>
|
||||||
</Tabs>
|
</Tabs>
|
||||||
|
|
||||||
- [Conflicts](/reference/conflicts) — Detect value conflicts between non-duplicate entities.
|
- [Conflicts](conflicts) — Detect value conflicts between non-duplicate entities.
|
||||||
- [Knowledge Graph](/reference/kg) — GraphBuilder uses deduplication during construction.
|
- [Knowledge Graph](kg) — GraphBuilder uses deduplication during construction.
|
||||||
- [Normalize](/reference/normalize) — Normalize entity names before deduplication.
|
- [Normalize](normalize) — Normalize entity names before deduplication.
|
||||||
- [Provenance](provenance) — Track merged entity lineage.
|
- [Provenance](provenance) — Track merged entity lineage.
|
||||||
|
|||||||
@@ -607,7 +607,9 @@ The Knowledge Explorer embeds Distance Intelligence directly in the browser dash
|
|||||||
The 10× cache improvement applies when the graph is unchanged between requests. In write-heavy pipelines where nodes are added continuously, cache hit rates will be lower. Use `force_refresh=False` (default) for read-heavy Explorer usage and `force_refresh=True` for batch pipeline contexts.
|
The 10× cache improvement applies when the graph is unchanged between requests. In write-heavy pipelines where nodes are added continuously, cache hit rates will be lower. Use `force_refresh=False` (default) for read-heavy Explorer usage and `force_refresh=True` for batch pipeline contexts.
|
||||||
</Note>
|
</Note>
|
||||||
|
|
||||||
- [Context Module](/reference/context) — `ContextGraph.get_neighbors()` and proximity-blended retrieval.
|
- [Context Module](context) — `ContextGraph.get_neighbors()` and proximity-blended retrieval.
|
||||||
- [Knowledge Graph Module](/reference/kg) — `NodeEmbedder`, `SimilarityCalculator`, and graph analytics.
|
- [Knowledge Graph Module](kg) — `NodeEmbedder`, `SimilarityCalculator`, and graph analytics.
|
||||||
- [Visualization](visualization) — Programmatic distance heatmaps and ego-mode graph renders.
|
- [Visualization](visualization) — Programmatic distance heatmaps and ego-mode graph renders.
|
||||||
- [Explorer](/reference/explorer) — Knowledge Explorer with built-in Distance Intelligence dashboard.
|
- [Explorer](explorer) — Knowledge Explorer with built-in Distance Intelligence dashboard.
|
||||||
|
|
||||||
|
- [Distance Intelligence](https://github.com/semantica-agi/semantica/blob/main/cookbook/advanced/12_Distance_Intelligence.ipynb) — Semantic neighborhoods and distance matrices · Advanced
|
||||||
|
|||||||
@@ -619,7 +619,7 @@ providers = check_available_providers()
|
|||||||
# → {"sentence_transformers": True, "fastembed": True, "openai": False}
|
# → {"sentence_transformers": True, "fastembed": True, "openai": False}
|
||||||
```
|
```
|
||||||
|
|
||||||
- [Vector Store](/reference/vector_store) — Store and search the generated embeddings.
|
- [Vector Store](vector_store) — Store and search the generated embeddings.
|
||||||
- [Split](/reference/split) — Chunk text before embedding for better retrieval quality.
|
- [Split](split) — Chunk text before embedding for better retrieval quality.
|
||||||
- [KG Module](/reference/kg) — Distance Intelligence uses graph embeddings for semantic neighbourhoods.
|
- [KG Module](kg) — Distance Intelligence uses graph embeddings for semantic neighbourhoods.
|
||||||
- [Deduplication](deduplication) — Semantic deduplication uses embedding distance for entity resolution.
|
- [Deduplication](deduplication) — Semantic deduplication uses embedding distance for entity resolution.
|
||||||
|
|||||||
+46
-206
@@ -1,224 +1,64 @@
|
|||||||
---
|
---
|
||||||
title: "Evals Module"
|
title: "Evals Module"
|
||||||
description: "Score decision records, audit trails, and reasoning output with deterministic and model-backed evaluators plus a small run harness."
|
description: "Evaluation framework for measuring Knowledge Graph quality, extraction accuracy, and pipeline performance: coming soon."
|
||||||
icon: "chart-line"
|
icon: "chart-line"
|
||||||
---
|
---
|
||||||
|
|
||||||
`semantica.evals` measures the quality of decision intelligence outputs. It takes
|
**`semantica.evals`** is planned as a comprehensive evaluation framework for measuring **extraction accuracy, graph quality, and pipeline performance**.
|
||||||
the decisions, audit trails, and reasoning text your pipeline produces and scores
|
|
||||||
them against expectations you define, returning a structured summary you can log,
|
|
||||||
assert on in tests, or track across runs.
|
|
||||||
|
|
||||||
- A registry of named evaluators, from exact string matching to ROUGE overlap and
|
<Warning>
|
||||||
LLM-as-judge
|
**`semantica.evals` is not yet implemented.** The module is a placeholder with `__all__ = []`. No classes or functions are available for import. This page describes the planned API only.
|
||||||
- `decision_scores`, a composite evaluator for `Decision` objects that checks
|
</Warning>
|
||||||
outcome, confidence bounds, required fields, provenance, and (optionally)
|
|
||||||
policy compliance
|
|
||||||
- A `evaluate()` runner that applies several evaluators to a list of cases and
|
|
||||||
aggregates pass / fail / error counts
|
|
||||||
- Per-evaluator **objectives** that let you override an evaluator's built-in
|
|
||||||
verdict at the run level
|
|
||||||
|
|
||||||
<Note>
|
## Planned Features
|
||||||
The module is versioned separately from the package: `semantica.evals.__version__`
|
|
||||||
is `"0.1.0"`. The public surface described here is stable, but expect additive
|
|
||||||
changes (new evaluators, new objective options) before it reaches 1.0.
|
|
||||||
</Note>
|
|
||||||
|
|
||||||
## Public API
|
When released, `semantica.evals` will provide:
|
||||||
|
|
||||||
| Name | Kind | Role |
|
| Planned Class | Role |
|
||||||
| :--- | :--- | :--- |
|
|
||||||
| `evaluate(cases, evaluators, config=None, target_fn=None)` | function | Run named evaluators over each case, return an `EvalSummary` |
|
|
||||||
| `list_evaluators()` | function | Sorted names of every registered evaluator |
|
|
||||||
| `get_evaluator(name)` | function | Look up a single evaluator function by name |
|
|
||||||
| `EvalMetric` | dataclass (frozen) | One evaluator's result: `score`, `passed`, `meta` |
|
|
||||||
| `CaseResult` | namedtuple | One case's result: `case_id`, `status`, `metrics`, `details` |
|
|
||||||
| `EvalSummary` | dataclass | Aggregate across cases: `total`, `passed`, `failed`, `errors`, `pass_rate`, `cases` |
|
|
||||||
|
|
||||||
```python
|
|
||||||
import semantica.evals as evals
|
|
||||||
from semantica.evals import evaluate, list_evaluators, get_evaluator
|
|
||||||
```
|
|
||||||
|
|
||||||
## Built-in evaluators
|
|
||||||
|
|
||||||
Every evaluator is a plain function `fn(actual, expected, config=None) -> EvalMetric`
|
|
||||||
registered under a stable name. `list_evaluators()` returns the current set:
|
|
||||||
|
|
||||||
```python
|
|
||||||
>>> list_evaluators()
|
|
||||||
['decision_scores', 'exact_match', 'keyword_check', 'length_range',
|
|
||||||
'levenshtein', 'llm_as_judge', 'numeric_range', 'regex_match', 'rouge',
|
|
||||||
'temporal_range']
|
|
||||||
```
|
|
||||||
|
|
||||||
| Name | Passes when | Relevant `config` keys |
|
|
||||||
| :--- | :--- | :--- |
|
|
||||||
| `exact_match` | `actual == expected` | none |
|
|
||||||
| `regex_match` | `re.search(expected, actual)` matches | none |
|
|
||||||
| `keyword_check` | every required term appears in `actual` (word-boundary) | `required` (falls back to `expected`) |
|
|
||||||
| `numeric_range` | `min <= actual <= max` | `min`, `max` (both required) |
|
|
||||||
| `temporal_range` | ISO datetime `actual` falls in `[min, max]` | `min`, `max` as ISO strings (both required) |
|
|
||||||
| `length_range` | `min <= len(actual) <= max` | `min` (default 0), `max` (required) |
|
|
||||||
| `levenshtein` | normalized similarity `>= threshold` | `threshold` (default 0.8) |
|
|
||||||
| `rouge` | ROUGE-1 F1 `> 0` and `>= threshold` | `threshold` (default 0.0) |
|
|
||||||
| `llm_as_judge` | caller-supplied `judge_fn(actual, expected)` returns truthy | `judge_fn` (required callable) |
|
|
||||||
| `decision_scores` | all configured sub-checks on a `Decision` pass | see below |
|
|
||||||
|
|
||||||
An evaluator that cannot run (bad regex, unparseable datetime, no `judge_fn`) returns an
|
|
||||||
`EvalMetric` with an `"error"` key in `meta` rather than raising. Evaluators that
|
|
||||||
require numeric bounds (`numeric_range`, `length_range`) instead return a failing
|
|
||||||
metric with a `"reason"` key when the bound is missing — they do not raise and do
|
|
||||||
not set `"error"`.
|
|
||||||
|
|
||||||
### `decision_scores`
|
|
||||||
|
|
||||||
`decision_scores` accepts a `Decision` (from `semantica.context.decision_models`)
|
|
||||||
or its dict form and runs a set of field-level and governance checks. The score is
|
|
||||||
the fraction of checks that passed; `passed` is `True` only when all of them did.
|
|
||||||
|
|
||||||
| Sub-check | Controlled by |
|
|
||||||
| :--- | :--- |
|
| :--- | :--- |
|
||||||
| Outcome matches | `expected_outcome` in config, or the case's `expected`; **skipped** when neither is set |
|
| `KGEvaluator` | Completeness, consistency, schema compliance, coverage, and orphan node detection |
|
||||||
| Confidence in range | `min_confidence` (default 0.0), `max_confidence` (default 1.0); always run |
|
| `ExtractionEvaluator` | NER precision / recall / F1 and relation extraction metrics against gold datasets |
|
||||||
| `decision_maker`, `reasoning`, `scenario` non-empty | always run |
|
| `PipelineBenchmark` | Throughput (docs/sec), per-step latency, peak memory, and error rate |
|
||||||
| Provenance present in metadata | `provenance_key` (default `"provenance"`); always run |
|
| `RegressionTracker` | Record runs and compare metrics across commits or config changes |
|
||||||
| Policy compliance | `policy_engine` and `policy_id` both set; skipped otherwise |
|
| `EvalReport` | Structured report: `{scores, regressions, recommendations}` |
|
||||||
|
| `DeduplicationEvaluator` | Merge precision, false positive / false negative rates |
|
||||||
|
| `ReasoningEvaluator` | Inference accuracy, rule coverage, and derivation depth |
|
||||||
|
|
||||||
Passing `causal_chain_exists` in config raises `NotImplementedError`. That key is a
|
## Current Workaround
|
||||||
reserved slot for a future release.
|
|
||||||
|
|
||||||
## Running an evaluation
|
Until `semantica.evals` ships, use `semantica.ontology.OntologyEvaluator` for ontology quality metrics:
|
||||||
|
|
||||||
`evaluate()` takes a list of cases and a list of evaluator names. A case is either
|
|
||||||
a `(expected, actual)` tuple or a dict:
|
|
||||||
|
|
||||||
```python
|
```python
|
||||||
{
|
from semantica.ontology import OntologyEvaluator
|
||||||
"id": "loan-001", # optional, generated if absent
|
|
||||||
"expected": ..., # optional; some evaluators read it, some don't
|
evaluator = OntologyEvaluator()
|
||||||
"actual": ..., # the value under test
|
|
||||||
"config": {...}, # optional, per-evaluator settings for this case
|
# evaluate_ontology takes the ontology dict only
|
||||||
"target_fn": callable, # optional, called with the case to produce `actual`
|
result = evaluator.evaluate_ontology(ontology)
|
||||||
}
|
|
||||||
|
print("Coverage: ", result.coverage_score)
|
||||||
|
print("Completeness:", result.completeness_score)
|
||||||
|
print("Gaps: ", result.gaps)
|
||||||
|
print("Suggestions: ", result.suggestions)
|
||||||
|
|
||||||
|
# Full report with class granularity and relation completeness
|
||||||
|
report = evaluator.generate_report(ontology)
|
||||||
|
print("Coverage score: ", report["evaluation"]["coverage_score"])
|
||||||
|
print("Completeness score:", report["evaluation"]["completeness_score"])
|
||||||
|
print("Relation coverage: ", report["relation_completeness"]["relation_coverage"])
|
||||||
```
|
```
|
||||||
|
|
||||||
If `actual` is missing, the runner calls the case's `target_fn` (or the
|
`EvaluationResult` fields returned by `evaluate_ontology()`:
|
||||||
`target_fn` passed to `evaluate()`) to produce it. Per-case `config` is deep-merged
|
|
||||||
over the top-level `config`, so a case can override one evaluator's settings
|
|
||||||
without discarding the rest.
|
|
||||||
|
|
||||||
```python
|
| Field | Type | Description |
|
||||||
from datetime import datetime
|
| :----- | :---- | :----------- |
|
||||||
|
| `coverage_score` | `float` | Fraction of competency questions answerable by the ontology |
|
||||||
|
| `completeness_score` | `float` | Average of class and property completeness scores |
|
||||||
|
| `gaps` | `List[str]` | Identified gaps in coverage |
|
||||||
|
| `suggestions` | `List[str]` | Improvement suggestions |
|
||||||
|
| `metrics` | `dict` | Detailed sub-metrics |
|
||||||
|
|
||||||
from semantica.context.decision_models import Decision
|
- [Semantic Extract](semantic_extract) — Extraction module.
|
||||||
from semantica.evals import evaluate
|
- [Knowledge Graph](kg) — Graph quality assessment.
|
||||||
|
- [Pipeline](pipeline) — Pipeline performance metrics.
|
||||||
decision = Decision(
|
- [Ontology Evaluator](ontology) — Available now for ontology quality metrics.
|
||||||
decision_id="d-1",
|
|
||||||
category="loan",
|
|
||||||
scenario="loan-request",
|
|
||||||
reasoning="vetted against lending policy v3",
|
|
||||||
outcome="approve",
|
|
||||||
confidence=0.87,
|
|
||||||
timestamp=datetime.now(),
|
|
||||||
decision_maker="approver-a",
|
|
||||||
metadata={"provenance": "workflow:loan/v3"},
|
|
||||||
)
|
|
||||||
|
|
||||||
cases = [
|
|
||||||
{
|
|
||||||
"id": "loan-001",
|
|
||||||
"actual": decision,
|
|
||||||
"config": {
|
|
||||||
"decision_scores": {
|
|
||||||
"expected_outcome": "approve",
|
|
||||||
"min_confidence": 0.7,
|
|
||||||
}
|
|
||||||
},
|
|
||||||
},
|
|
||||||
]
|
|
||||||
|
|
||||||
summary = evaluate(cases, ["decision_scores"])
|
|
||||||
print(summary.pass_rate) # 1.0
|
|
||||||
```
|
|
||||||
|
|
||||||
Evaluators run independently per case. If one raises, that case's `status` becomes
|
|
||||||
`"error"` and the exception text is captured in the metric's `meta`; the rest of
|
|
||||||
the run continues.
|
|
||||||
|
|
||||||
## Objectives
|
|
||||||
|
|
||||||
By default each evaluator decides its own pass / fail. An **objective** overrides
|
|
||||||
that verdict at the run level, keyed by evaluator name under `config`:
|
|
||||||
|
|
||||||
```python
|
|
||||||
# Raise levenshtein's bar from its default 0.8 to 0.9
|
|
||||||
evaluate(
|
|
||||||
[("apple", "aple")],
|
|
||||||
evaluators=["levenshtein"],
|
|
||||||
config={"levenshtein": {"objective": {"direction": "maximize", "threshold": 0.9}}},
|
|
||||||
)
|
|
||||||
|
|
||||||
# Lower is better
|
|
||||||
evaluate(
|
|
||||||
[("night", "nacht")],
|
|
||||||
evaluators=["levenshtein"],
|
|
||||||
config={"levenshtein": {"objective": {"direction": "minimize", "threshold": 0.5}}},
|
|
||||||
)
|
|
||||||
|
|
||||||
# Expect the metric NOT to match
|
|
||||||
evaluate(
|
|
||||||
[("ok", "ok")],
|
|
||||||
evaluators=["exact_match"],
|
|
||||||
config={"exact_match": {"objective": {"expect": False}}},
|
|
||||||
)
|
|
||||||
```
|
|
||||||
|
|
||||||
Rules:
|
|
||||||
|
|
||||||
- `maximize` with `threshold`: pass iff `score >= threshold`. `maximize` with no
|
|
||||||
threshold is a no-op and the evaluator's own verdict stands.
|
|
||||||
- `minimize` with `threshold`: pass iff `score <= threshold`. `minimize`
|
|
||||||
**requires** a threshold; omitting it raises `ValueError`.
|
|
||||||
- `expect` (`True` / `False`): pass iff `bool(score)` equals it. Cannot be combined
|
|
||||||
with `direction` or `threshold`, and must be a real boolean.
|
|
||||||
- A metric that already carries an `"error"` in its `meta` is unaffected by any
|
|
||||||
objective.
|
|
||||||
- Invalid objective config is validated for every case before any evaluator runs,
|
|
||||||
so a bad objective fails the whole run up front rather than partway through.
|
|
||||||
|
|
||||||
## Reading the summary
|
|
||||||
|
|
||||||
```python
|
|
||||||
summary = evaluate(cases, ["decision_scores"])
|
|
||||||
|
|
||||||
summary.total, summary.passed, summary.failed, summary.errors
|
|
||||||
summary.pass_rate # passed / total, or 1.0 for an empty case list
|
|
||||||
|
|
||||||
for case in summary.cases:
|
|
||||||
print(case.case_id, case.status) # status: "pass" | "fail" | "error"
|
|
||||||
for name, metric in case.metrics.items():
|
|
||||||
print(name, metric.score, metric.passed)
|
|
||||||
print(metric.meta.get("reasons", {})) # per-sub-check failure reasons
|
|
||||||
```
|
|
||||||
|
|
||||||
`EvalMetric` is frozen (`score: float`, `passed: bool`, `meta: dict`). `CaseResult`
|
|
||||||
is a namedtuple, and `EvalSummary` is a plain dataclass, so all three are
|
|
||||||
straightforward to serialize for logging or regression tracking.
|
|
||||||
|
|
||||||
## Notes
|
|
||||||
|
|
||||||
- `llm_as_judge` needs `config["judge_fn"]`, a callable
|
|
||||||
`judge_fn(actual, expected) -> bool` you supply. No LLM backend is imported
|
|
||||||
unless you pass one in.
|
|
||||||
- `decision_scores` governance checks are opt-in: policy compliance is only
|
|
||||||
evaluated when both `policy_engine` and `policy_id` are present.
|
|
||||||
|
|
||||||
## See also
|
|
||||||
|
|
||||||
- [Decision Intelligence](/guides/decision-intelligence) — producing the `Decision` records this module scores
|
|
||||||
- [Reasoning](/reference/reasoning) — inference output that reasoning-text evaluators can measure
|
|
||||||
- [Policy Engine](/guides/policy-engine) — the `policy_engine` used by `decision_scores`
|
|
||||||
- [Ontology Evaluator](/reference/ontology) — separate tooling for ontology quality metrics
|
|
||||||
|
|||||||
@@ -403,7 +403,7 @@ Semantic neighborhood requires node embeddings stored in node properties (keys `
|
|||||||
**Session state lost after restart**
|
**Session state lost after restart**
|
||||||
Session state is in-memory only. Use `POST /api/export` to save a JSON snapshot before shutting down.
|
Session state is in-memory only. Use `POST /api/export` to save a JSON snapshot before shutting down.
|
||||||
|
|
||||||
- [Context](/reference/context) — Build and save the ContextGraph that Explorer loads.
|
- [Context](context) — Build and save the ContextGraph that Explorer loads.
|
||||||
- [Ontology](ontology) — Programmatic ontology management and SHACL generation.
|
- [Ontology](ontology) — Programmatic ontology management and SHACL generation.
|
||||||
- [Visualization](visualization) — Programmatic graph rendering without the Explorer server.
|
- [Visualization](visualization) — Programmatic graph rendering without the Explorer server.
|
||||||
- [Export](export) — Export to RDF, Parquet, and other formats without launching a server.
|
- [Export](export) — Export to RDF, Parquet, and other formats without launching a server.
|
||||||
|
|||||||
@@ -394,7 +394,7 @@ The `export_csv` convenience function delegates to `CSVExporter.export()`. For p
|
|||||||
**Match your export format to your consumer.** Neo4j → `cypher`; ArangoDB → `aql`; Gephi/yEd → `graphml` or `gexf`; semantic web tools → `turtle` or `json-ld`; analytics pipelines → `parquet`; zero-copy IPC → `arrow`.
|
**Match your export format to your consumer.** Neo4j → `cypher`; ArangoDB → `aql`; Gephi/yEd → `graphml` or `gexf`; semantic web tools → `turtle` or `json-ld`; analytics pipelines → `parquet`; zero-copy IPC → `arrow`.
|
||||||
</Tip>
|
</Tip>
|
||||||
|
|
||||||
- [Triplet Store](/reference/triplet_store) — Store RDF exports in a SPARQL-queryable backend.
|
- [Triplet Store](triplet_store) — Store RDF exports in a SPARQL-queryable backend.
|
||||||
- [Ontology](ontology) — Export OWL ontologies.
|
- [Ontology](ontology) — Export OWL ontologies.
|
||||||
- [Provenance](provenance) — Include provenance metadata in RDF exports.
|
- [Provenance](provenance) — Include provenance metadata in RDF exports.
|
||||||
- [Pipeline](pipeline) — Add export as a final pipeline step.
|
- [Pipeline](pipeline) — Add export as a final pipeline step.
|
||||||
|
|||||||
@@ -503,7 +503,7 @@ stats = store.get_stats()
|
|||||||
</Tab>
|
</Tab>
|
||||||
</Tabs>
|
</Tabs>
|
||||||
|
|
||||||
- [KG Module](/reference/kg) — Build the graph before persisting it.
|
- [KG Module](kg) — Build the graph before persisting it.
|
||||||
- [Triplet Store](/reference/triplet_store) — RDF triple store for semantic web and SPARQL queries.
|
- [Triplet Store](triplet_store) — RDF triple store for semantic web and SPARQL queries.
|
||||||
- [Visualization](visualization) — Visualize graphs stored in any backend.
|
- [Visualization](visualization) — Visualize graphs stored in any backend.
|
||||||
- [Context](/reference/context) — AgentContext uses GraphStore for memory retrieval.
|
- [Context](context) — AgentContext uses GraphStore for memory retrieval.
|
||||||
|
|||||||
@@ -646,7 +646,7 @@ from semantica.ingest import ingest_file
|
|||||||
result = ingest_file("source_path", method="my_format")
|
result = ingest_file("source_path", method="my_format")
|
||||||
```
|
```
|
||||||
|
|
||||||
- [Parse](/reference/parse) — Parse raw sources into structured text and tables.
|
- [Parse](parse) — Parse raw sources into structured text and tables.
|
||||||
- [Pipeline](pipeline) — Orchestrate ingest as the first pipeline step.
|
- [Pipeline](pipeline) — Orchestrate ingest as the first pipeline step.
|
||||||
- [Snowflake Integration](../integrations/snowflake) — Snowflake-specific setup and authentication guide.
|
- [Snowflake Integration](../integrations/snowflake) — Snowflake-specific setup and authentication guide.
|
||||||
- [Databricks Integration](../integrations/databricks) — Databricks Unity Catalog setup, authentication, and lineage guide.
|
- [Databricks Integration](../integrations/databricks) — Databricks Unity Catalog setup, authentication, and lineage guide.
|
||||||
|
|||||||
@@ -75,10 +75,10 @@ kg = builder.build({"entities": entities, "relationships": relationships})
|
|||||||
## Temporal Knowledge Graphs (v0.4.0+)
|
## Temporal Knowledge Graphs (v0.4.0+)
|
||||||
|
|
||||||
<Info>
|
<Info>
|
||||||
Full temporal reference including `BiTemporalFact`, `TemporalReasoningEngine`, Allen interval algebra, and `TemporalNormalizer` is covered in the dedicated [Temporal Intelligence](/reference/temporal) page. This section documents the KG-layer temporal API.
|
Full temporal reference including `BiTemporalFact`, `TemporalReasoningEngine`, Allen interval algebra, and `TemporalNormalizer` is covered in the dedicated [Temporal Intelligence](temporal) page. This section documents the KG-layer temporal API.
|
||||||
</Info>
|
</Info>
|
||||||
|
|
||||||
The temporal stack — see the [Temporal Intelligence](/reference/temporal) page for the full reference.
|
The temporal stack — see the [Temporal Intelligence](temporal) page for the full reference.
|
||||||
|
|
||||||
### Building a Temporal Graph
|
### Building a Temporal Graph
|
||||||
|
|
||||||
@@ -264,7 +264,7 @@ versioner.verify_checksum(past_kg)
|
|||||||
```
|
```
|
||||||
|
|
||||||
<Tip>
|
<Tip>
|
||||||
See the [Temporal Intelligence](/reference/temporal) reference for the full class API, domain examples (personnel changes, policy evolution, financial timelines), and configuration options.
|
See the [Temporal Intelligence](temporal) reference for the full class API, domain examples (personnel changes, policy evolution, financial timelines), and configuration options.
|
||||||
</Tip>
|
</Tip>
|
||||||
|
|
||||||
|
|
||||||
@@ -475,10 +475,10 @@ kg:
|
|||||||
default_validity: infinite
|
default_validity: infinite
|
||||||
```
|
```
|
||||||
|
|
||||||
- [Graph Store](/reference/graph_store) — Persist graphs in Neo4j, FalkorDB, or Apache AGE.
|
- [Graph Store](graph_store) — Persist graphs in Neo4j, FalkorDB, or Apache AGE.
|
||||||
- [Semantic Extract](/reference/semantic_extract) — Source of entities and relationships fed to GraphBuilder.
|
- [Semantic Extract](semantic_extract) — Source of entities and relationships fed to GraphBuilder.
|
||||||
- [Visualization](visualization) — Visualize knowledge graphs interactively.
|
- [Visualization](visualization) — Visualize knowledge graphs interactively.
|
||||||
- [Conflicts](/reference/conflicts) — Conflict detection and resolution.
|
- [Conflicts](conflicts) — Conflict detection and resolution.
|
||||||
|
|
||||||
### Cookbooks
|
### Cookbooks
|
||||||
|
|
||||||
|
|||||||
@@ -439,7 +439,7 @@ extractor = NERExtractor(
|
|||||||
)
|
)
|
||||||
```
|
```
|
||||||
|
|
||||||
- [Semantic Extract](/reference/semantic_extract) — Use LLMs for NER and relation extraction.
|
- [Semantic Extract](semantic_extract) — Use LLMs for NER and relation extraction.
|
||||||
- [Agno Integration](../integrations/agno) — LLM providers in Agno multi-agent teams.
|
- [Agno Integration](../integrations/agno) — LLM providers in Agno multi-agent teams.
|
||||||
- [Reasoning](reasoning) — LLM-backed deductive and abductive reasoning.
|
- [Reasoning](reasoning) — LLM-backed deductive and abductive reasoning.
|
||||||
- [Context](/reference/context) — GraphRAG uses LLMs for reasoning over knowledge graphs.
|
- [Context](context) — GraphRAG uses LLMs for reasoning over knowledge graphs.
|
||||||
|
|||||||
@@ -6,7 +6,7 @@ icon: "plug"
|
|||||||
|
|
||||||
**`semantica.mcp_server`** exposes Semantica's knowledge graph, decision intelligence, semantic extraction, and reasoning capabilities as an [MCP (Model Context Protocol)](https://modelcontextprotocol.io) **server over stdio**:
|
**`semantica.mcp_server`** exposes Semantica's knowledge graph, decision intelligence, semantic extraction, and reasoning capabilities as an [MCP (Model Context Protocol)](https://modelcontextprotocol.io) **server over stdio**:
|
||||||
|
|
||||||
- 15 MCP tools exposed: extract entities, query graph, record decisions, run reasoning, export results
|
- 12 MCP tools exposed: extract entities, query graph, record decisions, run reasoning, export results
|
||||||
- No Python code required after launch: configure once, use from any MCP-aware client
|
- No Python code required after launch: configure once, use from any MCP-aware client
|
||||||
- Compatible with Claude Desktop, Windsurf, Cline, Continue, VS Code, Roo Code, Cursor
|
- Compatible with Claude Desktop, Windsurf, Cline, Continue, VS Code, Roo Code, Cursor
|
||||||
|
|
||||||
@@ -40,12 +40,12 @@ python -m semantica.mcp_server
|
|||||||
|
|
||||||
## What You Get
|
## What You Get
|
||||||
|
|
||||||
- **15 MCP Tools** — Extract entities, extract relations, record decisions, query decisions, find precedents, trace causal chains, add entities, add relationships, run analytics, summarise graph, run reasoning, export graph, query the live graph, update nodes, archive nodes.
|
- **12 MCP Tools** — Extract entities, extract relations, record decisions, query decisions, find precedents, trace causal chains, add entities, add relationships, run analytics, summarise graph, run reasoning, export graph.
|
||||||
- **3 Readable Resources** — Live graph JSON (`semantica://graph/summary`), decision list, and schema/version info: readable by any MCP client.
|
- **3 Readable Resources** — Live graph JSON (`semantica://graph/summary`), decision list, and schema/version info: readable by any MCP client.
|
||||||
- **Zero Infrastructure** — Runs over stdio: no server, no port, no Docker required. One config block to activate in any MCP client.
|
- **Zero Infrastructure** — Runs over stdio: no server, no port, no Docker required. One config block to activate in any MCP client.
|
||||||
- **Persistent Graphs** — Point `SEMANTICA_KG_PATH` at a saved graph file to reload it automatically on every server startup.
|
- **Persistent Graphs** — Point `SEMANTICA_KG_PATH` at a saved graph file to reload it automatically on every server startup.
|
||||||
- **Decision Intelligence** — Record decisions, find precedents via hybrid similarity search, and trace causal chains across agent runs.
|
- **Decision Intelligence** — Record decisions, find precedents via hybrid similarity search, and trace causal chains across agent runs.
|
||||||
- **REST Alternative** — The [Explorer](/reference/explorer) module offers a full HTTP API and browser dashboard if you prefer programmatic access.
|
- **REST Alternative** — The [Explorer](explorer) module offers a full HTTP API and browser dashboard if you prefer programmatic access.
|
||||||
|
|
||||||
## Installation
|
## Installation
|
||||||
|
|
||||||
@@ -159,7 +159,7 @@ The MCP server is included in the base install: no extras required.
|
|||||||
|
|
||||||
## Tools
|
## Tools
|
||||||
|
|
||||||
The MCP server exposes 15 tools that any connected AI assistant can call:
|
The MCP server exposes 12 tools that any connected AI assistant can call:
|
||||||
|
|
||||||
| Tool | Category | Description |
|
| Tool | Category | Description |
|
||||||
| :---- | :-------- | :----------- |
|
| :---- | :-------- | :----------- |
|
||||||
@@ -173,9 +173,6 @@ The MCP server exposes 15 tools that any connected AI assistant can call:
|
|||||||
| `add_relationship` | Graph Operations | Add a directed edge between two nodes |
|
| `add_relationship` | Graph Operations | Add a directed edge between two nodes |
|
||||||
| `get_graph_summary` | Graph Operations | Node count, decision count, graph status |
|
| `get_graph_summary` | Graph Operations | Node count, decision count, graph status |
|
||||||
| `get_graph_analytics` | Graph Operations | PageRank centrality and community detection |
|
| `get_graph_analytics` | Graph Operations | PageRank centrality and community detection |
|
||||||
| `query_graph` | Graph Operations | Fetch a node, traverse its neighbours, or keyword-search nodes |
|
|
||||||
| `update_node` | Graph Operations | Merge properties onto a node and persist to `SEMANTICA_KG_PATH` |
|
|
||||||
| `delete_node` | Graph Operations | Soft-delete (archive) a node and persist to `SEMANTICA_KG_PATH` |
|
|
||||||
| `run_reasoning` | Reasoning | Forward-chain IF/THEN rules over facts |
|
| `run_reasoning` | Reasoning | Forward-chain IF/THEN rules over facts |
|
||||||
| `export_graph` | Reasoning & Export | Serialise the graph (`turtle`/`ttl`: RDF Turtle aliases, `nt`, `xml`, `json-ld`, `json`) |
|
| `export_graph` | Reasoning & Export | Serialise the graph (`turtle`/`ttl`: RDF Turtle aliases, `nt`, `xml`, `json-ld`, `json`) |
|
||||||
|
|
||||||
@@ -389,51 +386,6 @@ Takes no input parameters.
|
|||||||
|
|
||||||
</Accordion>
|
</Accordion>
|
||||||
|
|
||||||
<Accordion title="query_graph" icon="magnifying-glass">
|
|
||||||
|
|
||||||
Read the live graph in one of three modes, set by `mode`:
|
|
||||||
|
|
||||||
- `node` — return a single node by `node_id`.
|
|
||||||
- `neighbors` (default) — traverse outward and inward from `node_id` up to `depth` hops (clamped to 1-5, default 1). Optional `relationship_types` filters edge types; optional `limit` caps results.
|
|
||||||
- `search` — keyword match `query` against each node's id and content. Optional `node_type` restricts the scan; `limit` defaults to 50.
|
|
||||||
|
|
||||||
**Input:**
|
|
||||||
|
|
||||||
```json
|
|
||||||
{ "mode": "neighbors", "node_id": "apple_inc", "depth": 2 }
|
|
||||||
```
|
|
||||||
|
|
||||||
</Accordion>
|
|
||||||
|
|
||||||
<Accordion title="update_node" icon="pen">
|
|
||||||
|
|
||||||
Merge a set of properties onto an existing node. The change is applied in memory and, when `SEMANTICA_KG_PATH` is set, written back to that file so it survives a restart. Returns `persisted: false` when no path is configured.
|
|
||||||
|
|
||||||
**Input:**
|
|
||||||
|
|
||||||
```json
|
|
||||||
{
|
|
||||||
"node_id": "task_42",
|
|
||||||
"properties": { "status": "done", "note": "shipped in v0.6.7" }
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
`node_id` and a non-empty `properties` object are required. Updating a missing node returns an error.
|
|
||||||
|
|
||||||
</Accordion>
|
|
||||||
|
|
||||||
<Accordion title="delete_node" icon="box-archive">
|
|
||||||
|
|
||||||
Soft-delete a node: it stays in the graph for history but is marked `status: "archived"`. Persists to `SEMANTICA_KG_PATH` when configured.
|
|
||||||
|
|
||||||
**Input:**
|
|
||||||
|
|
||||||
```json
|
|
||||||
{ "node_id": "task_42" }
|
|
||||||
```
|
|
||||||
|
|
||||||
</Accordion>
|
|
||||||
|
|
||||||
</AccordionGroup>
|
</AccordionGroup>
|
||||||
|
|
||||||
### Reasoning
|
### Reasoning
|
||||||
@@ -493,7 +445,7 @@ The MCP server exposes three readable resources:
|
|||||||
| `semantica://decisions/list` | All recorded decisions (up to 50) |
|
| `semantica://decisions/list` | All recorded decisions (up to 50) |
|
||||||
| `semantica://schema/info` | Server version and available tools |
|
| `semantica://schema/info` | Server version and available tools |
|
||||||
|
|
||||||
- [Context](/reference/context) — The ContextGraph that the MCP server operates on.
|
- [Context](context) — The ContextGraph that the MCP server operates on.
|
||||||
- [Semantic Extract](/reference/semantic_extract) — NER and relation extraction powering the MCP tools.
|
- [Semantic Extract](semantic_extract) — NER and relation extraction powering the MCP tools.
|
||||||
- [Reasoning](reasoning) — Forward-chaining engine behind run_reasoning.
|
- [Reasoning](reasoning) — Forward-chaining engine behind run_reasoning.
|
||||||
- [Agno Integration](../integrations/agno) — Use Semantica inside Agno multi-agent teams.
|
- [Agno Integration](../integrations/agno) — Use Semantica inside Agno multi-agent teams.
|
||||||
|
|||||||
@@ -584,7 +584,7 @@ normalized = normalize_text("Apple Inc.", method="expand_suffixes")
|
|||||||
# → "Apple Incorporated"
|
# → "Apple Incorporated"
|
||||||
```
|
```
|
||||||
|
|
||||||
- [Parse](/reference/parse) — Parse documents before normalization.
|
- [Parse](parse) — Parse documents before normalization.
|
||||||
- [Split](/reference/split) — Chunk normalized text for embedding.
|
- [Split](split) — Chunk normalized text for embedding.
|
||||||
- [Deduplication](deduplication) — Resolve duplicate entities after normalization.
|
- [Deduplication](deduplication) — Resolve duplicate entities after normalization.
|
||||||
- [Pipeline](pipeline) — Include normalization as a named pipeline step.
|
- [Pipeline](pipeline) — Include normalization as a named pipeline step.
|
||||||
|
|||||||
@@ -287,6 +287,6 @@ ontology_data = ingest_ontology("schema.jsonld") # JSON-LD
|
|||||||
</Note>
|
</Note>
|
||||||
|
|
||||||
- [Reasoning](reasoning) — Apply inference rules over ontology axioms.
|
- [Reasoning](reasoning) — Apply inference rules over ontology axioms.
|
||||||
- [Knowledge Graph](/reference/kg) — The graph being modeled by the ontology.
|
- [Knowledge Graph](kg) — The graph being modeled by the ontology.
|
||||||
- [Export](export) — Export ontologies as RDF, OWL, or JSON-LD.
|
- [Export](export) — Export ontologies as RDF, OWL, or JSON-LD.
|
||||||
- [Conflicts](/reference/conflicts) — Detect ontology constraint violations.
|
- [Conflicts](conflicts) — Detect ontology constraint violations.
|
||||||
|
|||||||
@@ -298,6 +298,6 @@ for source in sources:
|
|||||||
</Note>
|
</Note>
|
||||||
|
|
||||||
- [Ingest](ingest) — Load files before parsing.
|
- [Ingest](ingest) — Load files before parsing.
|
||||||
- [Split](/reference/split) — Chunk parsed text for embedding and extraction.
|
- [Split](split) — Chunk parsed text for embedding and extraction.
|
||||||
- [Docling Integration](../integrations/docling) — Full Docling integration setup guide.
|
- [Docling Integration](../integrations/docling) — Full Docling integration setup guide.
|
||||||
- [Semantic Extract](/reference/semantic_extract) — Extract entities and relations from parsed text.
|
- [Semantic Extract](semantic_extract) — Extract entities and relations from parsed text.
|
||||||
|
|||||||
@@ -497,7 +497,7 @@ result = engine.execute_pipeline(
|
|||||||
|
|
||||||
## SPARQL CONSTRUCT Template Steps
|
## SPARQL CONSTRUCT Template Steps
|
||||||
|
|
||||||
Use the `"construct_template"` step type to render and execute a [SPARQL CONSTRUCT template](/reference/triplet_store#sparql-construct-templates) as part of a pipeline. `store_backend` and `construct_template_registry` are execution-time resources, not step config — pass them to `execute_pipeline()`, the same way `delta_mode` steps receive `version_manager` and `triplet_store`:
|
Use the `"construct_template"` step type to render and execute a [SPARQL CONSTRUCT template](triplet_store#sparql-construct-templates) as part of a pipeline. `store_backend` and `construct_template_registry` are execution-time resources, not step config — pass them to `execute_pipeline()`, the same way `delta_mode` steps receive `version_manager` and `triplet_store`:
|
||||||
|
|
||||||
```python
|
```python
|
||||||
from semantica.pipeline import PipelineBuilder, ExecutionEngine
|
from semantica.pipeline import PipelineBuilder, ExecutionEngine
|
||||||
@@ -589,6 +589,6 @@ StepStatus.SKIPPED # Skipped due to FailureHandler "skip" strategy
|
|||||||
</AccordionGroup>
|
</AccordionGroup>
|
||||||
|
|
||||||
- [Ingest](ingest) — First step in most pipelines.
|
- [Ingest](ingest) — First step in most pipelines.
|
||||||
- [Semantic Extract](/reference/semantic_extract) — Core extraction step.
|
- [Semantic Extract](semantic_extract) — Core extraction step.
|
||||||
- [Knowledge Graph](/reference/kg) — Graph construction step.
|
- [Knowledge Graph](kg) — Graph construction step.
|
||||||
- [Export](export) — Final output step.
|
- [Export](export) — Final output step.
|
||||||
|
|||||||
@@ -522,7 +522,7 @@ Provenance tracking in Semantica produces the following audit artifacts:
|
|||||||
`ProvenanceManager` does not include built-in Turtle or JSON-LD serialization. Use `entry.to_dict()` and `get_lineage()` to retrieve provenance data, then serialize with your preferred RDF library if W3C PROV-O RDF output is required.
|
`ProvenanceManager` does not include built-in Turtle or JSON-LD serialization. Use `entry.to_dict()` and `get_lineage()` to retrieve provenance data, then serialize with your preferred RDF library if W3C PROV-O RDF output is required.
|
||||||
</Note>
|
</Note>
|
||||||
|
|
||||||
- [Change Management](/reference/change_management) — Version control and snapshot audit trails.
|
- [Change Management](change_management) — Version control and snapshot audit trails.
|
||||||
- [Ingest](ingest) — Provenance begins at the ingestion stage.
|
- [Ingest](ingest) — Provenance begins at the ingestion stage.
|
||||||
- [Export](export) — Include provenance metadata in RDF exports.
|
- [Export](export) — Include provenance metadata in RDF exports.
|
||||||
- [Context](/reference/context) — Decision provenance via AgentContext.
|
- [Context](context) — Decision provenance via AgentContext.
|
||||||
|
|||||||
@@ -482,7 +482,7 @@ step.confidence # float
|
|||||||
`GraphReasoner` requires a configured LLM provider. If the provider fails to initialize, `reason()` returns an error string instead of raising. Check `reasoner.provider is not None` before calling if you need to surface failures explicitly.
|
`GraphReasoner` requires a configured LLM provider. If the provider fails to initialize, `reason()` returns an error string instead of raising. Check `reasoner.provider is not None` before calling if you need to surface failures explicitly.
|
||||||
</Warning>
|
</Warning>
|
||||||
|
|
||||||
- [Knowledge Graph](/reference/kg) — The knowledge graph being reasoned over.
|
- [Knowledge Graph](kg) — The knowledge graph being reasoned over.
|
||||||
- [Ontology](ontology) — Ontology axioms and SHACL constraints for logical reasoning.
|
- [Ontology](ontology) — Ontology axioms and SHACL constraints for logical reasoning.
|
||||||
- [Triplet Store](/reference/triplet_store) — RDF backend for SPARQL-based reasoning.
|
- [Triplet Store](triplet_store) — RDF backend for SPARQL-based reasoning.
|
||||||
- [Context](/reference/context) — Reasoning integrated into agent decision intelligence.
|
- [Context](context) — Reasoning integrated into agent decision intelligence.
|
||||||
|
|||||||
@@ -322,6 +322,6 @@ export SEMANTICA_SEED_MERGE_STRATEGY=seed_first
|
|||||||
</Tip>
|
</Tip>
|
||||||
|
|
||||||
- [Ingest](ingest) — Load unstructured data alongside seed data.
|
- [Ingest](ingest) — Load unstructured data alongside seed data.
|
||||||
- [Knowledge Graph](/reference/kg) — The target graph that seed data populates.
|
- [Knowledge Graph](kg) — The target graph that seed data populates.
|
||||||
- [Deduplication](deduplication) — Handle duplicates during seed-extracted merge.
|
- [Deduplication](deduplication) — Handle duplicates during seed-extracted merge.
|
||||||
- [Pipeline](pipeline) — Incorporate seed loading as a named pipeline step.
|
- [Pipeline](pipeline) — Incorporate seed loading as a named pipeline step.
|
||||||
|
|||||||
@@ -410,7 +410,7 @@ triplets = trip.extract(text)
|
|||||||
| `ml` | Fast | Free | High | Limited |
|
| `ml` | Fast | Free | High | Limited |
|
||||||
| `llm` | Medium | API cost | Highest | Yes (schema) |
|
| `llm` | Medium | API cost | Highest | Yes (schema) |
|
||||||
|
|
||||||
- [LLM Providers](/reference/llms) — Configure which LLM is used for extraction.
|
- [LLM Providers](llms) — Configure which LLM is used for extraction.
|
||||||
- [Knowledge Graph](/reference/kg) — Build graphs from extracted entities and relationships.
|
- [Knowledge Graph](kg) — Build graphs from extracted entities and relationships.
|
||||||
- [Parse Module](/reference/parse) — Parse documents before extraction.
|
- [Parse Module](parse) — Parse documents before extraction.
|
||||||
- [Deduplication](deduplication) — Resolve duplicate entities after extraction.
|
- [Deduplication](deduplication) — Resolve duplicate entities after extraction.
|
||||||
|
|||||||
@@ -373,7 +373,7 @@ for chunk in chunks:
|
|||||||
|
|
||||||
For the full pipeline orchestration API, see the [Pipeline reference](pipeline).
|
For the full pipeline orchestration API, see the [Pipeline reference](pipeline).
|
||||||
|
|
||||||
- [Parse](/reference/parse) — Parse documents before chunking: produces sections and metadata.
|
- [Parse](parse) — Parse documents before chunking: produces sections and metadata.
|
||||||
- [Embeddings](/reference/embeddings) — Embed chunks for vector search and semantic chunking.
|
- [Embeddings](embeddings) — Embed chunks for vector search and semantic chunking.
|
||||||
- [Semantic Extract](/reference/semantic_extract) — Extract entities and relations from individual chunks.
|
- [Semantic Extract](semantic_extract) — Extract entities and relations from individual chunks.
|
||||||
- [Pipeline](pipeline) — Integrate splitting as a named pipeline step.
|
- [Pipeline](pipeline) — Integrate splitting as a named pipeline step.
|
||||||
|
|||||||
@@ -874,8 +874,8 @@ kg:
|
|||||||
engine: allen # allen | point_in_time_only
|
engine: allen # allen | point_in_time_only
|
||||||
```
|
```
|
||||||
|
|
||||||
- [Knowledge Graph Module](/reference/kg) — Core graph construction, `GraphBuilder`, analytics.
|
- [Knowledge Graph Module](kg) — Core graph construction, `GraphBuilder`, analytics.
|
||||||
- [Context Module](/reference/context) — Decision temporal windows and `find_active_nodes()`.
|
- [Context Module](context) — Decision temporal windows and `find_active_nodes()`.
|
||||||
- [Provenance](provenance) — W3C PROV-O lineage stamped alongside temporal metadata.
|
- [Provenance](provenance) — W3C PROV-O lineage stamped alongside temporal metadata.
|
||||||
- [Export](export) — OWL, Turtle, JSON-LD, and Parquet export with temporal annotations.
|
- [Export](export) — OWL, Turtle, JSON-LD, and Parquet export with temporal annotations.
|
||||||
|
|
||||||
|
|||||||
@@ -564,4 +564,4 @@ for row in result.bindings:
|
|||||||
- [Export](export) — Export knowledge graphs to RDF formats.
|
- [Export](export) — Export knowledge graphs to RDF formats.
|
||||||
- [Ontology](ontology) — Load OWL ontologies and store as RDF triples.
|
- [Ontology](ontology) — Load OWL ontologies and store as RDF triples.
|
||||||
- [Reasoning](reasoning) — SPARQL-based property chain inference.
|
- [Reasoning](reasoning) — SPARQL-based property chain inference.
|
||||||
- [Graph Store](/reference/graph_store) — Property graph alternative for Cypher queries.
|
- [Graph Store](graph_store) — Property graph alternative for Cypher queries.
|
||||||
|
|||||||
@@ -222,5 +222,5 @@ from semantica.utils import read_json_file
|
|||||||
config = read_json_file("config.json")
|
config = read_json_file("config.json")
|
||||||
```
|
```
|
||||||
|
|
||||||
- [Core](/reference/core) — Framework orchestration that uses Utils internally.
|
- [Core](core) — Framework orchestration that uses Utils internally.
|
||||||
- [Pipeline](pipeline) — Uses ProgressTracker for per-step tracking.
|
- [Pipeline](pipeline) — Uses ProgressTracker for per-step tracking.
|
||||||
|
|||||||
@@ -588,7 +588,7 @@ store.create_index(index_type="pq", metric="L2", m=8)
|
|||||||
</Tab>
|
</Tab>
|
||||||
</Tabs>
|
</Tabs>
|
||||||
|
|
||||||
- [Embeddings](/reference/embeddings) — Generate the vectors stored here.
|
- [Embeddings](embeddings) — Generate the vectors stored here.
|
||||||
- [Context](/reference/context) — AgentContext uses VectorStore for memory retrieval.
|
- [Context](context) — AgentContext uses VectorStore for memory retrieval.
|
||||||
- [Split](/reference/split) — Chunk documents before embedding and storing.
|
- [Split](split) — Chunk documents before embedding and storing.
|
||||||
- [Ingest](ingest) — Ingest documents before embedding and storing.
|
- [Ingest](ingest) — Ingest documents before embedding and storing.
|
||||||
|
|||||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user