mirror of
https://github.com/semantica-agi/semantica.git
synced 2026-08-29 04:26:20 +00:00
main
20
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1a220da477 | fix(export): address the review findings on the metadata pass-through | ||
|
|
d06434ae31 |
Merge upstream/main into metadata-passthrough
#1123 through #1127 landed while this was open, and #1125 rewrote the same four entity loops this branch extends. Confidence is now normalised through normalize_confidence, which returns None for a value that has no xsd:decimal form, so the clause can be absent. Resolved by folding that into the clause list this branch already builds: the Turtle path assembles its predicate-object clauses and then terminates the last one, which is what makes a variable-length list work at all, and an omitted confidence is simply one clause fewer. RDF/XML and JSON-LD take the upstream conditional as written, with the metadata call after it. |
||
|
|
eb7427d12c |
fix(export): carry metadata through every RDF serialization (#1154)
convert_kg_to_rdf copies metadata into the RDF-ready dictionary at
rdf_exporter.py:302 and no serializer has ever read it back out. Turtle,
N-Triples, RDF/XML and RDFExporter's JSON-LD each write an entity's id,
type, text and confidence and nothing else, so an entity keeps its
confidence score and loses what produced it. JSONExporter's json-ld path
keeps the same fields, which is how one knowledge graph exported two ways
carried the user's data through one exporter and none through the other.
Measured on
|
||
|
|
8f6948f85d |
fix(export): close the four review gaps in the default-graph change
All four are in the branch that recognises an already-converted document, which has to survive every shape JSON-LD allows rather than the one shape Semantica happens to produce. A knowledge graph carrying a context of its own took the already-JSON-LD branch and skipped its own conversion, leaving entity ids, relationship endpoints, types and confidences as raw keys. The entities/relationships test now runs first, and a converted document never has those keys, so the double-conversion guard is unaffected. A context that is a URL or an array cannot be merged key by key, and was being dropped in favour of Semantica's defaults, silently changing how every term expands. Both are kept as an array now, the caller's winning, which is the same precedence the dictionary branch already used. An explicit null is left alone on purpose: in an array it resets the active context and would take the semantica prefix with it. @graph may be a single node object as well as an array. list() on a dictionary yields its keys, so an object-valued graph was replaced by a list of strings. A caller may hand us a document that is deliberately a named graph. That name is theirs to keep, so it is no longer flattened; it is nested one level and the export's own provenance goes beside it, in the default graph, where a plain reader can see it. Four tests, one per case, all failing before this commit. |
||
|
|
60eb595d62 |
fix(export): keep JSON-LD payloads in the default graph
A JSON-LD document with a top-level @id and a top-level @graph is a named graph. Its members become quads named by that @id, and the default graph is left empty. rdflib.Graph.parse() keeps the default graph and discards the rest without reporting anything, so every consumer that loads an export the ordinary way saw the document header and none of the data. _convert_to_jsonld wrote the payload into @graph and then stamped a document @id beside it, which named every list export and every generic-dict export. export_knowledge_graph made it worse: it converted the graph to JSON-LD and handed the finished document back to export(), which converted it a second time. The converted document no longer carries entities/relationships keys, so the second pass treated it as opaque and buried the whole knowledge graph inside @graph, under a name that is a wall-clock timestamp. A two-entity, one-relationship graph exported to JSON-LD parsed as 2 triples with Graph() and 21 quads with Dataset(). The 19 missing triples were the entire knowledge graph. The document node now goes inside @graph when the payload lives there, and is the document itself otherwise, so no export names its own graph by accident. An already-converted document is merged rather than nested, which also stops the export carrying two document nodes and two @context blocks. Semantica's reader has the mirror of this bug (#1129), so these exports could not be read back by Semantica either. |
||
|
|
71cffb15e9 |
fix(core): consume the fallback flag at the call site, not in the helper
Review finding, reproduced. `call_custom_method(..., **kwargs)` builds a
fresh dict from the unpacking, so popping `fallback_on_custom_error`
inside the helper left the caller's own kwargs untouched. On the fallback
path the flag was then forwarded straight into the default
implementation, which is exactly the case the flag exists for.
Instrumenting the default exporter shows it arriving:
config handed to the default exporter: {'fallback_on_custom_error': True}
Most defaults take **kwargs and ignore it, which is why nothing failed
loudly, but any default with a fixed signature raises TypeError on it.
The helper's docstring promised the flag was never forwarded, so the
promise was false rather than merely untidy.
All 58 sites now pop the flag from their own bag and pass it explicitly.
One site in normalize/methods.py names its bag `**context` rather than
`**kwargs`, and is handled too.
3 further tests: the flag reaches neither the default implementation nor
a successful custom method, and a per-module guard that every call site
has a matching pop, since a site that forgets one reintroduces the leak
silently.
Failure set across the six affected modules is unchanged against
upstream/main: 37 pre-existing, none new.
|
||
|
|
d7ee22cf1f |
fix(export): keep the full predicate on the reified relationship
Review finding, reproduced. The reified node reduced the relationship type to its last fragment or path component, so https://a.example/ns#employs and https://b.example/ns#employs both became semantica:type "employs". The temporal node no longer said which predicate it described, and it disagreed with the direct triple written beside it, which carries the full IRI. The full predicate is written instead. I had flagged the local-name form as a deliberate simplification in the PR description; the collision case shows it was the wrong call. 2 further tests. |
||
|
|
efdfa39c15 |
fix(export): address review findings on the confidence typing fix
1. An absurd magnitude expanded instead of being rejected. xsd:decimal has no exponent notation, so the value has to be written out in full, and "1e100000000" is eleven characters that expand to a hundred million digits. "1e100000" already produced a 100,001 character string here. The export path continues past validation errors, so one malformed field could exhaust memory. Values beyond MAX_CONFIDENCE_EXPONENT are now omitted like any other unusable value. 1e-9 still round-trips. 2. Decimal keeps the sign of zero, so 0.0 and -0.0 serialised as "0" and "-0", which are two distinct RDF terms. That is exactly the duplicate this PR exists to remove, so zero is normalised. 4 further tests. |
||
|
|
66e3333e41 |
fix(ontology): address review findings on the SHACL namespace fix
Four findings from the automated review, all reproduced first. 1. The fix only reached Turtle. `_uri` was the single place I corrected, and JSON-LD and N-Triples build sh:targetClass, sh:path and sh:class straight from graph.base_uri, so two of the three formats went on emitting shapes that match nothing. That is the defect this PR claims to close, still live wherever the output is not Turtle. All three serializers now resolve through one `_term_iri`, and the pySHACL violation test runs against each of them. 2. Classes and properties shared one name-keyed index built with setdefault, so a property named after a class was permanently mapped to the class IRI and its sh:path validated the wrong predicate. The index is now split into class_iris and property_iris, and each call site says which it wants. 3. OntologyEngine.to_shacl forwarded target_namespace and attach_domainless_properties through generate(**options), which never reads them, so both were silently dropped on the public path. They are now named parameters passed to the constructor, and documented. 4. The opt-in attachment logged at debug. It broadens constraint generation, so it warns. 7 further tests, including the target-namespace and real-violation checks parametrised across Turtle, N-Triples and JSON-LD. |
||
|
|
9ca83d397f |
fix(export): address review findings on the ontology schema fix
Four findings from the automated review, all reproduced first. 1. The name fallback minted invalid IRIs. `_term_iri` pasted a raw name onto the ontology base, so a class named "Customer Account" produced <https://example.org/onto/Customer Account>. rdflib only warns about the space, Oxigraph rejects it with "Invalid IRI code point". That is the same class of defect this PR set out to fix, introduced by the fix itself. Local names are now percent-encoded. 2. `improve_coherence` raised AttributeError. It lives on OntologyOptimizer, which holds no namespace manager, so the URI fallback I added there crashed on any ontology carrying a class without a URI. It now mints from the ontology's own base through a shared module-level helper. 3. `owl:Thing` was treated as an absolute IRI. It matches the generic scheme grammar, so `_is_absolute_iri` accepted it and domains and ranges came out as the term <owl:Thing> rather than <http://www.w3.org/2002/07/owl#Thing>. This is the live path: stage 4 of the generator assigns ["owl:Thing"] to object properties with no inferred endpoints. Absoluteness is now decided on a real scheme, and the well-known prefixes expand. 4. Unusable property entries were dropped in silence. Non-dictionary entries and definitions carrying no type are now named in a warning. 6 further tests, including a strict-parser check through Oxigraph, which is what catches the space that rdflib waves through. |
||
|
|
d5dc4eabac |
fix(core): let a registered custom method refuse (#1108)
Every module supporting custom methods wrapped the registered callable in
a bare `except Exception`, logged a warning, and carried on into the
built-in implementation:
try:
return custom_method(data, file_path, format=format, **kwargs)
except Exception as e:
logger.warning(f"Custom method {method} failed: {e}, falling back to default")
That makes a registered method advisory. It can add behaviour, but it
cannot decline. For a gate, a validator or a policy check, declining is
the entire purpose: raising is how such a method says "do not produce
this output". Catching the exception and running the default produces
exactly the output the method was registered to prevent, and the only
trace is a warning.
Demonstrated with a verifier that rejects invalid RDF and deletes the
file. The fallback wrote it straight back.
`call_custom_method` in utils/custom_methods.py now holds the policy in
one place: an exception from a registered method propagates. Callers who
relied on the old behaviour can pass `fallback_on_custom_error=True`,
which restores warn-and-continue for that call and is consumed by the
policy rather than forwarded to the method.
The swallow was in six modules, not only the one the issue was filed
against, so all 58 sites are converted: export 13, ingest 13, normalize
13, parse 12, embeddings 4, kg 3. The rewrite is mechanical and uniform.
Sentinel comparison is by identity, so a custom method returning None, 0,
"" or an empty list is not mistaken for a failure.
13 tests in tests/utils/test_custom_method_can_refuse.py, including the
issue's own demonstration and a guard that no module still carries the
swallow. Across the six affected modules the failure set is identical to
upstream/main: 37 pre-existing failures before and after, none new, with
869 passing against 856 on the baseline.
|
||
|
|
f60ca6a529 |
fix(export): give the OWL-Time interval a subject the graph can reach (#1106)
include_temporal=True emitted a well formed OWL-Time interval hanging off
a relationship IRI that appears nowhere else in the graph. A relationship
is written as a single triple, <e1> <employs> <e2>, so there is no node
for the time to attach to:
<...#rel_0_0940a860> time:hasTime <...#rel_0_0940a860__valid_interval> .
Counting inbound arcs to that subject gives zero. The timestamps parse,
they validate, and no query can reach them from the relationship they
describe, which is the only thing they are for.
The JSON-LD path already reifies relationships as sem:Relationship with
sem:source, sem:target and sem:type, and the vocabulary declares all four
terms. Turtle now emits the same shape when it has temporal data to
attach, so the two serializations describe relationships the same way and
the interval has a reachable subject.
The direct triple is unchanged, and nothing is reified when a
relationship carries no temporal data, so default output is untouched.
7 tests in tests/export/test_owl_time_reachability.py, including a SPARQL
walk from the edge to its validity interval, which is what the dangling
node made impossible, and a check that every emitted term is declared in
the shipped vocabulary. Export and ontology suites pass at 228 tests.
|
||
|
|
05c21af117 |
fix(export): write confidence as one typed decimal on every path (#1100, #1102)
#1100 — the four serializers rendered the same confidence four different ways. Turtle wrote it bare, which the Turtle grammar reads as xsd:decimal. N-Triples typed it xsd:float. RDF/XML wrote a plain literal with no datatype. JSON-LD wrote a native JSON number, which expands to xsd:double. For confidence 0.9 that is four distinct RDF terms, so a FILTER matches at most one of them, and merging two exports of one graph gives an entity two different confidence values. N-Triples also omitted the triple entirely when confidence was absent, while the other three wrote the 1.0 default, so the two serializations differed in the number of triples as well as in their datatype. `normalize_confidence` now produces one canonical lexical form and every path writes it with CONFIDENCE_DATATYPE. xsd:decimal is the choice because it is what the Turtle path already produced, so the most used output is unchanged, and because it is exact: xsd:float is 32 bit binary and cannot represent 0.9 at all. Values that arrive in exponent notation are reformatted, since 1e-05 is not a valid xsd:decimal. #1102 — the Turtle path interpolated the value with no type check, so a confidence of "high" produced `semantica:confidence high .` and made the entire document unparseable. One bad field cost the whole export. A value that cannot be a decimal is now omitted with a warning naming the entity, rather than written as something the vocabulary contradicts. Numeric strings are still accepted. Booleans are not, since bool subclasses int and True would otherwise become a confidence of 1. sem:confidence in the shipped vocabulary declared no rdfs:range, deliberately, because declaring one would have contradicted three of the four exporters. It now declares xsd:decimal, and a drift guard asserts the vocabulary and the serializers agree. 20 tests in tests/export/test_confidence_literal_typing.py, comparing the parsed graphs of all four formats rather than their text. Export and ontology suites pass at 240 tests. |
||
|
|
981c9d9208 |
fix(ontology): target the namespace the data uses, and stop inventing constraints (#1104, #1105)
#1104 — SHACLGenerator used one namespace for two jobs. `base_uri` says where the shape resources live, and it was also used to expand every sh:targetClass and sh:path. With the default "https://semantica.dev/shapes/" that made shapes target <https://semantica.dev/shapes/Person>, while data carries the ontology's own class IRI or the semantica:ns# vocabulary. The shapes matched nothing. That failure is silent. A shape with no focus nodes is vacuously satisfied, so pySHACL reports conforms=True on data that plainly breaks the stated constraints. The shipped validator agrees the file is fine. The two namespaces are now separate. `target_namespace` resolves in this order: an explicit argument, the ontology's declared namespace, the namespace of any absolute IRI a term already carries, the ontology URI, and finally the vocabulary namespace the package ships rather than the shapes namespace. Every class and property name is indexed to the IRI it expands to, and `_uri` resolves through that index, so shapes always name the terms the data uses. #1105 — a property with no declared domain was attached to every node shape. That states a constraint the ontology does not, and with minCount 1 it makes every instance of every class invalid. Such a property is now left unattached, with a warning naming it. Passing attach_domainless_properties=True restores the old behaviour. tests/ontology/test_shacl_target_namespace.py adds 17 tests that validate real data through pySHACL rather than reading the shapes text, so a shape that targets nothing cannot pass by being ignored. They cover a generated ontology, one that declares only a namespace, and one that carries only class URIs. tests/ontology/test_ontology_advanced.py::test_no_domain_property_attaches_to_all_shapes asserted the #1105 behaviour, so it pinned the defect in place. It is now two tests: the old expectation against the explicit opt-in, and the new default. Export and ontology suites pass at 239 tests. |
||
|
|
c30ec14858 |
fix(export): read the ontology shape the generator actually emits (#1103)
OWLExporter read `object_properties` and `data_properties`, while
OntologyGenerator emits one combined `properties` list tagged with
type/@type. Every generated property was therefore dropped, and a
generated ontology exported as classes alone.
Class IRIs were worse. ClassInferrer writes `"uri": None` when it is
given no namespace manager, so the stage 3 guard `if "uri" not in cls`
never fired: the key is present, only its value is missing. The exporter
then interpolated the empty string into `<>`, which is a relative IRI
that resolves against the parser's base. Under rdflib that base is the
current working directory, so a two-class ontology parsed as one subject
carrying two rdfs:label values, and the identity of that subject changed
with the directory the export ran from. Oxigraph rejects the same file
outright with "No scheme found in an absolute IRI".
Changes:
- Accept both dict shapes. `_split_properties` classifies the combined
`properties` list by type/@type and merges it with any explicit
`object_properties` and `data_properties`.
- Resolve class and property IRIs through `_term_iri`, falling back from
uri to iri to id to a name joined onto the ontology base. A term with
none of those is skipped with a warning rather than emitted as `<>`.
- Resolve domain and range references through the class index, so a bare
name such as "Person" lands on the IRI that class was exported under
instead of staying relative.
- Resolve data property ranges properly. "string", "xsd:string" and a
full IRI now all give one well formed datatype. The previous
`rdfs:range xsd:{range}` produced `xsd:xsd:string` for generator output,
which no parser accepts. Turtle keeps the compact xsd: form the module
already used.
- Fix the two `not in` guards in the generator so a present-but-None uri
is minted, and mint an absolute IRI rather than assigning a bare name.
- Escape XML text and attribute values, which were interpolated raw, so a
label containing & or < no longer breaks the document.
Turtle and RDF/XML now serialise the same 25 triples for the same
ontology, and both are accepted by rdflib and by Oxigraph.
10 regression tests in tests/export/test_owl_exporter_generator_schema.py,
driven by a real OntologyGenerator run and asserting on the parsed graph
rather than on serialised text. All 10 fail on the parent commit. The
export and ontology suites pass at 231 tests.
|
||
|
|
e03212cd66 |
fix(provenance): compare timestamp ranges by instant, not by spelling
Review finding on #1121, and correct: with new entries carrying +00:00 and entries written earlier carrying nothing, query_recorded_between() and audit_log() compared ISO strings directly, which orders by how a timestamp is spelled rather than when it happened. Two consequences, both introduced by the offset this PR adds: - An inclusive naive bound naming a stored offset-bearing timestamp sorts below it, because the stored value is the longer string, so the record it names is excluded from its own range. - A bound in another offset lands wherever its digits fall. "2026-08-19T19:45:00+05:30" is 14:15Z, before an entry at 14:19Z, but string comparison puts it after. Both paths now compare instants, through a new to_utc_datetime() helper that reads a missing offset as UTC. That is what the naive values actually were: provenance stamped with datetime.utcnow(), so reading them as UTC keeps a stored naive value and the same instant written with an offset comparing equal instead of ordering by representation. It is also the read side the remaining 147 call sites will need whenever the rest of the package is converted. A bound that cannot be read as a timestamp keeps the historical string comparison rather than raising on a call that used to work. Five new tests cover the inclusive naive bound, the other-offset bound, legacy and offset-bearing entries ordered together, audit_log's since filter, and the unreadable-bound fallback. The first two fail with manager.py reverted; the rest are guards. 569 provenance, export and ontology tests pass, and the full-suite failure set is unchanged at 329, all from optional dependencies missing locally. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
83c04a57d6 |
fix(export,provenance): write timestamps with an explicit UTC offset (#1114)
semantica/export/ stamped every value with datetime.now().isoformat(), which reads the machine's local clock. semantica/provenance/ stamped its own with datetime.utcnow().isoformat(), which reads UTC. Both return a naive datetime and both serialize identically, so once the value is out of the process nothing distinguishes them: the same string means two different instants depending on which module wrote it. In RDF the consequence is silent rather than loud. Under XSD 1.1 a value with no timezone compared against one with a timezone is indeterminate whenever the two fall inside the 14-hour window, SPARQL turns an indeterminate comparison into an error, and FILTER discards errors as non-matches. Loading a Semantica-stamped export into Oxigraph next to two correctly stamped ones and asking which were written before a given instant returns the other two and drops ours, with no error anywhere. prov:generatedAtTime, prov:startedAtTime, prov:endedAtTime and prov:atTime all carry values written this way, so an audit trail cannot be ordered against timestamps from any other system. Adds utc_now()/utc_now_iso() to semantica/utils/helpers.py, exported from semantica.utils, and uses them at all 29 call sites in export/ (json_exporter, yaml_exporter, report_generator, export_provenance) and provenance/ (manager, schemas, bridge_axiom). Values now read 2026-08-19T14:19:04.229937+00:00: one unambiguous instant, comparable against any correctly stamped value, and valid xsd:dateTimeStamp. sem:exportedAt's range in the vocabulary that landed with #1109 is tightened from xsd:dateTime to xsd:dateTimeStamp accordingly. Its comment had to explain why the weaker range was necessary; that reason is gone. datetime.utcnow() is also deprecated as of Python 3.12 and scheduled for removal. Constructing a ProvenanceEntry under -W error::DeprecationWarning on 3.13 raised; it no longer does. Two new test modules cover offset presence on every export and provenance path, PROV-O literals valid as xsd:dateTimeStamp, comparison against a timezone-aware instant without TypeError, the Oxigraph filter that dropped the naive value, the declared range matching what the exporter writes, and the document @id remaining a valid IRI with +00:00 in it. The filter test picks a bound inside the indeterminate window on purpose: a bound years away is determinate even for a naive value, and the test would pass without the fix. 13 of the 14 fail with this commit's semantica/export, semantica/provenance and vocabulary reverted. The remaining 147 naive call sites, in context/, vector_store/, seed/ and elsewhere, are deliberately untouched: those timestamps are compared against values parsed back from previously stored naive strings, so converting the write side alone would raise TypeError on existing data. That sweep needs a read-side migration and belongs in its own change. No new failures across the suite: 329 pre-existing failures before and after, all from optional dependencies missing in the local environment. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
75b026c6dd |
fix(export): mint JSON-LD @ids the same way the RDF serializers do (#1101)
The #1101 fix covered serialize_to_turtle, serialize_to_ntriples and serialize_to_rdfxml. Both JSON-LD writers were left interpolating the entity's own text into f"semantica:entity/{text}" and the endpoints into f"semantica:rel/{source}_{target}". Three consequences, all reproducible on 0.6.5 through the public API: - An entity whose text contains a space, which is most organisation and person names an extractor produces, mints an invalid IRI. A JSON-LD parser drops that node in full and says nothing, so the entity is simply missing from the export: rdflib reads 6 triples for {"text": "AcmeCorp"} and 2 for {"text": "Acme Corp"}. - serialize_to_jsonld resolved endpoints from source_id/target_id only, while the rest of the module accepts source/target too. Every relationship carrying the second form minted the identical "semantica:rel/_", so all of them collapsed onto one node and their types and endpoints merged into a graph nobody wrote. - The JSON-LD @id and the Turtle IRI for one entity disagreed (ns#entity/Acme Corp vs ns#entity_a73cb4563ee2e72c), so the two serializations of one knowledge graph were two different graphs. Both writers now use mint_entity_iri/mint_relationship_iri, resolving endpoints both ways and passing the list index the RDF paths pass, so one knowledge graph carries one node identity whichever serializer wrote it. JSONExporter.export_entities and export_relationships also declare the semantica prefix their @context was already writing "semantica:entities" against. Without the declaration a processor reads that as an IRI in the scheme semantica rather than the namespace expansion, which is the original #1101 defect on a third path: rdflib returns the predicate literally as semantica:entities. tests/export/test_jsonld_iri_minting.py parses each export with a real JSON-LD processor rather than asserting on the JSON text, and covers all seven claims above. Each test fails on the parent commit. 236 export and ontology tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e55c03bd39 |
fix: resolve temporal endpoints both ways, and stop declaring a range the exporters contradict
Both from review on #1109. The temporal fallback minted from source_id only, while the main serializer accepts source_id or source. Relationships using the second form therefore hashed two empty strings, and once the IRI became deterministic that turned a latent problem into an active one: unrelated relationships at the same list index collided on the same IRI across exports, so their temporal data aliased when loaded together. Endpoints are now resolved the way serialize_to_turtle resolves them, before minting. The vocabulary declared sem:confidence with range xsd:decimal, which the N-Triples serializer contradicts by typing the same value xsd:float. Neither is safe to declare while the two serializers disagree, since the Turtle path writes the value bare and the Turtle grammar reads that as xsd:decimal. The range is dropped with the reasoning recorded on the term and a pointer to #1100, which tracks the disagreement itself. Extends the drift guard rather than only fixing the instance: a new test asserts that any range this vocabulary declares matches the datatype the serializers actually emit, so the class of contradiction that review caught fails the build next time. 228 export and ontology tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e1092ac507 |
feat(ontology): declare the Semantica vocabulary, and mint entity IRIs deterministically
Closes #1107, closes #1101. Every RDF export mints terms in https://semantica.dev/ns#, and nothing declared what those terms meant. The namespace returns 404 and no vocabulary shipped with the package, so a consumer receiving an export could not tell semantica:text from a typo of it: in the open world an undeclared IRI is unknown rather than wrong, and every RDF tool treats the two alike. Closed-world checking is what separates them, and it needs a document to check against. semantica/ontology/vocabulary/semantica-ns.ttl declares the fourteen terms the exporters actually emit, drawn from the emitting call sites rather than from what a vocabulary ought to contain. It ships inside the package so it loads without a network round trip, and is the same document intended to be served at the namespace IRI once hosting and content negotiation are sorted. tests/ontology/test_vocabulary.py ties the document to the code: every term the serializers can write must be declared, so adding a term to an exporter without declaring it fails the build rather than shipping an undeclared IRI. The vocabulary alone would not have made those IRIs resolve, because the fallback path minted them from Python's builtin hash(). That is randomised per process, so the same entity received a different IRI on every run and exports could not be diffed, deduplicated against an earlier load, or joined to a provenance record written by an earlier process. Minting now uses SHA-256 and writes a full IRI in the declared namespace rather than semantica:entity_N, which inside angle brackets is an IRI in the scheme "semantica" rather than the prefix expansion, and so never joined with anything written through the prefix. The same applies to the default entity and relationship types in the Turtle path. 134 export tests and 91 ontology tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |