Changelog
All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
v2.0.4 — 2026-10-08#
Patch release on the v2.x line under the support policy: one BrowserDb durability hardening fix, plus new compatibility and crash tests. File format is unchanged (v7). No API changes. Native and WASI builds behave exactly as in v2.0.3.
Fixed#
BrowserDbno longer drops a dirty page it cannot read from the IndexedDB flush (#470).checkpoint(),importGraph()and every write collected dirty pages withread_page_raw(id).ok(), so a page that failed to read was left out of the flush while the call still reported success. The read error is now returned. No current code path leaves a dirty id without a page, so this guards against a future bug rather than fixing observed data loss. Affected releases: v0.20.0 through v2.0.3,BrowserDbonly. The browser unit tests compile again on wasm32 (a duplicatepage_cache_capacityandtempfile-baseddir_synctests).
Tests#
- Golden-file compatibility corpus (#391).
tests/golden/holds v7 files written by minigraf 2.0.3, each with a JSON manifest of its CRC32,tx_countand expected query results. They cover one checkpoint, several checkpoints, indexes rebuilt on open, a WAL left by a crash, a WAL whose entries were already checkpointed, and same-transaction multi-values (#371).tests/golden_corpus_test.rsopens a copy of each file and checks its manifest after open, after a checkpoint and reopen, and after one more write. Golden files are never regenerated, only added; a newgolden-corpusjob inpolicy.ymlfails a pull request that changes one. The generator is a standalone crate pinned tominigraf = "=2.0.3"(tests/golden/gen/v7/). - SIGKILL crash test checks committed data (#384).
tests/crash_kill_test.rsused to check only that the file reopened and one attribute query ran, so it missed #370. The child now runs a seeded workload (batches over several entities and attributes, batched retracts, re-assertions after a retract) and logs each transaction afterexecutereturns; it checkpoints after every transaction, every third, never (WAL replay), automatically every 2 writes, or every third withSyncMode::Normal. After aSIGKILLat a random point, the reopened file must hold exactly the model after the last logged transaction or the one in flight, through a full scan, attribute (AEVT), attribute+value (AVET) and entity-bound (EAVT) queries and:as-of; again after a second reopen and after a checkpoint plus reopen. On v2.x the workload avoids the known issues #371 (two values of one attribute in one transaction) and #435 (asserting a triple that is already live returns one row per assertion); the v3 version covers both. Per-PR CI runs 5 rounds; the nightly crash-kill workflow runs 100 per OS.MINIGRAF_CRASH_KILL_SEEDreplays a failure.
Known issues#
- Same-transaction multi-valued facts can read back as one value (#371, #287), a fact retracted and re-asserted in one
WriteTransactionis lost at commit (#477), a later assertion does not replace an earlier valid-time window (#435), andsave()is not crash-atomic (#374); all are fixed in v3.0.0. All v2.x known issues, with affected versions and workarounds, are listed in the pinned issue #421.
v2.0.3 — 2026-10-06#
Patch release on the v2.x line with two data-integrity fixes, under the support policy. File format is unchanged (v7). No API changes. Upgrading is recommended for every v2.x user.
Fixed#
- Queries returned wrong results with more than 65,535 uncheckpointed facts (#445). Each fact not yet checkpointed was indexed by its position in memory, stored as a 16-bit number that stopped counting at 65,535. Every later fact was indexed as that same position, so entity-bound queries (
[:e :attr ?v]) and attribute scans ([?e :attr ?v]) silently lost or duplicated facts. This affected in-memory databases with more than 65,535 facts and file databases with more than 65,535 facts written since the last checkpoint. The stored data was always intact, and a checkpoint made file databases correct again. Positions are now 32-bit and checked. Affected releases: v0.7.0 through v2.0.2. The file format is unchanged. - Reopening next to an already-checkpointed WAL rewound the transaction counter (#447). The WAL normally disappears after a checkpoint. It survives one only if the process crashes between the checkpoint's commit and the WAL delete, if the delete fails (WAL-006), or with
wal_checkpoint_threshold(usize::MAX). When such a WAL held only entries already in the file, replay skipped them all and then reset the transaction counter from the facts it had loaded, which were none, so it reset to 0. The next transactions reusedtx_countvalues already in the file,:as-ofanswers mixed old and new transactions, and the next checkpoint made that history permanent. The counter now never moves backwards on replay. Affected releases: v0.5.0 through v2.0.2.
Known issues#
- Same-transaction multi-valued facts can read back as one value (#371, #287); the fix needs file format v8 and ships in v3.0.0.
save()is not crash-atomic (#374). All v2.x known issues, with affected versions and workarounds, are listed in the pinned issue #421.
v2.0.2 — 2026-09-27#
Final planned release on the v2.x line. File format is unchanged (v7). No breaking API changes; one new user-facing error code (API-010). v2.x gets data-integrity and security fixes for 12 months after v3.0.0 ships (support policy).
Performance#
-
Bound-entity point queries no longer pay for other attributes' history (#323).
[:e :attr ?v]now range-scans only(e, :attr)in the EAVT index instead of every record the entity has ever written. Reading a rarely changed attribute of a heavily rewritten entity no longer slows down as that entity's history grows: with 2,000 retract/reassert cycles on another attribute of the same entity, it drops from 4.58 ms to 19.0 µs. Reading the heavily rewritten attribute itself, and attribute scans ([?e :attr ?v]), are about 1.35–1.4× faster (4.64 ms → 3.36 ms), from cheaper net-assert grouping and removing a redundant dedup pass. The file format is unchanged. -
Attribute scans are bounded to exactly the queried attribute (#381).
[?e :attr ?v]computed the end of its AEVT range by incrementing the attribute's last byte. For attributes whose last byte is0x7For0xBF— including about 1 in 64 non-ASCII characters, such as:丿or:ÿ— the result was not valid UTF-8, so the scan ran to the end of the whole AEVT index and read every later fact before filtering. It also read facts for prefix siblings (:ab,:a/bwhen scanning:a). The range now ends atattribute + "\0", which covers exactly one attribute. Scanning 100:丿facts in a checkpointed file with 40,000 facts on neighbouring attributes drops from 45.4 ms to 251 µs. Results are unchanged. -
Checkpoints no longer re-encode the whole index (#315).
checkpoint()decoded and re-serialised every entry of all four covering indexes and re-read every page of the file for its checksum, so checkpointing after one new fact cost almost as much as after thousands. Index leaves that receive no new entries are now copied verbatim, only the leaves that do are decoded and re-split, and unchanged fact pages are no longer re-hashed. Leaves are still bounds-checked and the leaf chain is checked for cycles, so a damaged index fails the checkpoint instead of being copied forward. Checkpoint after one new fact drops from 246 ms to 26.3 ms on a 100k-fact file and from 23.9 ms to 4.0 ms at 10k. Cost still grows with graph size, because the index pages are copied on every checkpoint; the file format is unchanged. Thecheckpoint()andwal_checkpoint_thresholddocs now state that a checkpoint is not a durability boundary: the WAL is already crash-durable.
Fixed#
- A power loss could lose a new database or WAL file, or bring back a deleted WAL (#389). On Linux and other POSIX systems, a created or deleted file survives a power loss only once its parent directory is fsynced; Minigraf fsynced file contents but never a directory. Creating the
.graphfile or the<db>.walsidecar, and deleting the WAL after a checkpoint, now fsync the parent directory after the file operation. A resurrected WAL was already harmless, since replay skips entries at or below the last checkpointed transaction, but a lost.graphor WAL lost committed data. Windows needs no directory sync and is unchanged. Process kills were never affected, because they leave the OS page cache intact. - Recursive rules failed with a literal start when the recursion goes through another rule (#297).
(query [:find ?y :where (chain :a ?y)])over(chain ?x ?y) :- (link ?x ?mid) (chain ?mid ?y)failed withUnbound variable in rule head: ?mid(INT-022). The magic-sets rewrite, which runs when a rule argument is given as a literal, dropped earlier rule calls from the rules that pass bindings to the next call, so a variable bound only by such a call (here?mid, bound bylink) was never bound. #300 fixed the case where a fact pattern binds it. Those calls are now kept, and results match the same query without a literal start. or-joinfailed with INT-031 when no rows reached it (#405). A valid query such as[?item :order-item/product ?p] (or-join [?item] ...)returnedor-join variable ?item is not bound in the incoming scopeon an empty database, or whenever the clauses before theor-joinmatched nothing, but worked once matching data existed. A safety check read the bound variables from the incoming rows, so with no rows every join variable looked unbound.or-joinnow returns an empty result when no rows reach it, likeorandnot-joinalready did. This applies to queries and rule bodies.- A query with
$slotbind tokens run throughexecute()reported an internal error (#407). Such a query failed with[INT-025] internal: unsubstituted :valid-at bind slot reached the executorwhen it had a:valid-atslot, and with slots in other positions it could run and quietly return nothing.Minigraf::execute(),WriteTransaction::execute()and the REPL now reject it before execution with the new user error API-010, which names the slots and points toprepare(). Prepared queries are unchanged. - Bound-entity queries could drop rows when one
WriteTransactionwrote the same attribute in several valid-time windows (#323). Both facts share atx_count, and the selective lookup path de-duplicated on(entity, attribute, tx_count, asserted), so one window's value was silently lost, while the same query via a full scan returned both. That de-duplication has been removed, and bound-entity queries now always match a full scan. - The INT-054 negative-cycle error named its two predicates in a random order (#410). Registering
(rule [(p ?x) (not (q ?x))])and then(rule [(q ?x) (not (p ?x))])reportedpredicate 'q' is involved in a negative cycle through 'p'on some runs and the reverse on others, because stratification walked the rule dependency graph inHashMaporder. It now walks rules in sorted order, so the same rules always give the same message. Which rules are accepted or rejected is unchanged.
Tests#
- Crash at every point inside
save()(#374, #390). A fault-injection test makes a checkpoint fail at each page write and each sync in turn, on a file already written by several checkpoints, then reopens what reached storage. At all 47 crash points the file reopens with either the old header and exactly the previously checkpointed facts (oldlast_checkpointed_tx_count, so WAL replay re-applies the in-flight transaction) or the new header and all facts, and every fact is found by both entity and attribute lookups. 45 of the 47 reopens take the index-rebuild path. This checks the v2.x behaviour described in the known-issues list (#421); it models a process kill, not a power loss.
CI#
- New
Policyworkflow (#395). An MSRV job runscargo check --all-featuresandcargo test --libon Rust 1.89, resolving dependencies to 1.89-compatible versions, and fails ifrust-versioninCargo.tomldrifts from the tested toolchain.cargo-semver-checkscompares the public API with the latest release tag and blocks PRs tomain; on major-version branches such asv3it reports breaks without failing.cargo-deny(deny.toml) allows only permissive licenses compatible with MIT OR Apache-2.0, rejects duplicate crate versions and wildcard requirements, and requires every crate to come from crates.io.
Documentation#
-
Unsafe WAL recovery advice removed from
docs/ERROR_REFERENCE.md. The WAL-001, WAL-002 and STG-011 resolutions said a.walfile could be deleted with no data loss, or that a database could be rebuilt from the WAL alone. Both are wrong: transactions committed since the last checkpoint exist only in the WAL, and the WAL holds nothing older than that checkpoint. The resolutions now say to copy both files first, to checkpoint with the version that wrote the WAL, and what is lost if the WAL is deleted. -
CLAUDE.md now matches the milestones (v2.0.2 final v2.x release, v3.0.0 format v8 and data integrity, v3.1.0 features) and documents
error.rs,magic_sets.rs,fault_inject.rsandsrc/browser/. ROADMAP lists #405 and #407 in the v2.0.2 scope. -
README: MSRV (1.89) stated, v2.x known issue (#371) shown near the top, Maven coordinates corrected to
io.github.project-minigraf, Android listed on Maven Central, and binding download locations point to the binding repos. -
.github/SECURITY.mdlists 2.x as the supported line. About 50 broken wiki links indocs/ERROR_REFERENCE.mdnow use full wiki URLs. -
Stability and support policy (#397). PHILOSOPHY.md §7 now states a format policy instead of "frozen for decades": what counts as a format change, that v(N+1) always reads v(N), and that older formats are dropped only in a major release. A new support policy gives v2.x data-integrity and security fixes for 12 months after v3.0.0 ships. "Production-ready", "12–15 months to production" and the check mark on "never losing data" are replaced with claims that hold today. SECURITY.md links the policy.
-
Support tiers (#400). The Rust crate and Python are Tier 1 (fully tested, released with every core release). Node.js, browser WASM, WASI, Java/JVM, Android, Swift and C are Tier 2 (experimental, best-effort releases). Tier 1 platforms are Linux ext4/xfs, macOS APFS and Windows NTFS on local disk. The README platform table shows each binding's tier.
-
Known-issues process (#399). Issues labelled
data-integrity,corruptionordurabilitystay open until their fix is in a published release (CONTRIBUTING.md). A newknown-issuelabel and a pinned issue (#421) list each bug in the current release with affected versions, workaround and fix version; the README links it from a new "Known issues" section. #287 was reopened under this rule. The bug report template asks for the affected version range, binding and filesystem.
Known issues#
- Same-transaction multi-valued facts can read back as one value (#371, #287); the fix needs file format v8 and ships in v3.0.0. Workaround: write or retract each value of a multi-valued attribute in its own call.
save()is not crash-atomic (#374), and indexes damaged by the pre-v2.0.1 rebuild bug are not repaired (#373). All v2.x known issues, with affected versions and workarounds, are listed in the pinned issue #421.
Notes#
- v3.0.0 drops read support for file formats v1–v6. v3.0.0 reads format v7 and migrates it to v8. Every v1.x and v2.x release migrates v1–v6 files to v7 on open, so open any older file once with v2.x before upgrading to v3.0.0.
- On v2.x, reading a heavily rewritten attribute still costs time proportional to its history, because v7 index keys carry neither the value nor the assert/retract flag; the structural fix needs the v8 keys and is tracked for v3.0.0 in #379.
- Checkpoints on v2.x still copy every index page, so their cost grows with graph size. Copy-on-write index pages on the v3 branch, which also make
save()crash-atomic, are needed for checkpoints proportional to the change alone (#374).
v2.0.1 — 2026-09-25#
Patch release on the v2.x line. File format is unchanged (v7); no API changes.
Fixed#
- Entity-bound lookups could silently return nothing after an index rebuild on open (#370). When the index checksum does not match on open (for example after a process was killed during
save()), Minigraf rebuilds all four indexes from the fact pages. The rebuild computed each fact's location by re-packing all facts contiguously, but a file written by several checkpoints has partially filled pages that a contiguous re-pack does not reproduce. From the second checkpoint onward, rebuilt index entries pointed at the wrong page/slot: entity-bound queries ([:some/ident ?a ?v]) returned[]or errored while attribute-driven queries still worked, andcheckpoint()copied the bad entries forward. The rebuild now reads each fact's real on-disk location (#372).
Known issues#
- Files whose indexes were already rebuilt with wrong locations by v2.0.0 are not repaired by this release. A public integrity check and rebuild-indexes-from-fact-pages operation is tracked in #373; a crash-atomic
save()in #374. - Two values of the same attribute for one entity written in a single
transact(or retracted in a singleretract) can read back as one value, differently per query path (#371, #287). The fix changes the index key layout (file format v8) and ships in v3.0.0. Workaround on v2.x: write or retract each value of a multi-valued attribute in its own call.
v2.0.0 — 2026-08-26#
This was originally slated as v1.3.0, matching the GitHub milestone name under
which most of this work was tracked. It ships as v2.0.0 instead: this
project's CHANGELOG header (above) commits to Semantic Versioning, and under
strict SemVer any backward-incompatible change to the public API past 1.0.0
requires a major version bump, not a minor one. The breaking changes below —
a public method return-type change, two OpenOptions struct-literal breaks
plus #[non_exhaustive] closing that gap for good, Windows lock-semantics
changes, and an MSRV bump — are exactly that. GitHub milestones v1.4.0,
v1.5.0, v1.6.0, and the pre-existing untitled "2.0" (format v8 + window
completeness + FFI expansion) were renumbered to v2.1.0, v2.2.0, v2.3.0, and
v3.0.0 respectively to keep the sequence consistent.
Breaking-ish changes#
-
Adding
OpenOptions::allow_unlockedbroke struct-literal construction. At the time this field landed,OpenOptionshad all-public fields and was not#[non_exhaustive], so under Rust's semver rules any downstream code building it asOpenOptions { wal_checkpoint_threshold: .., page_cache_size: .., .. }without..Default::default()no longer compiled. Callers who spread..Default::default(), or who use the chainable builder methods, were unaffected. This branch had to update two such literals in its own tree — theDefaultimpl itself and one test — and only the latter is a "downstream literal" in the sense that matters to a consumer.This was the first post-1.0 field addition to
OpenOptions. The earliermax_derived_factsandmax_resultsfields landed in v0.19.0, before 1.0, so they set no precedent under the stability guarantee in PHILOSOPHY.md §7. Combined with thesynchronousfield below, it's whyOpenOptionsis now#[non_exhaustive]— see that entry. -
OpenOptionsgains a second post-1.0 field:synchronous(#302). Same semver consequence asallow_unlockedabove — a struct-literalOpenOptions { .. }built without..Default::default()no longer compiles. One in-tree literal (src/db.rs, a test) needed updating; every other in-tree construction already used the spread form or the chainable builder. -
OpenOptionsis now#[non_exhaustive]. This is a stronger break than the two field additions above: it disallows any struct-literal construction from outside this crate, including the..Default::default()spread form that made those additions non-breaking for spread-form callers. Construct viaOpenOptions::new()/OpenOptions::default()plus the chainable builder methods instead. Landed alongside theallow_unlocked/synchronousfields specifically so this is the last time a newOpenOptionsfield forces a major-version bump. A newwal_checkpoint_threshold(n)builder method was added in the same change — the field had none before, which is why every in-tree construction site (tests, benches, one example) used a struct literal for it and had to be converted to the builder form. -
Minimum supported Rust version is now 1.89 (was effectively 1.85, implied by
edition = "2024"). Required forstd::fs::File::try_lock, stabilised in 1.89. Declared asrust-versioninCargo.toml, so cargo reports a clear error rather than a confusing one. -
On Windows, an open database can no longer be read through any other file handle — including another handle in the same process.
LockFileExlocks are mandatory rather than advisory and exclude every handle but the one holding them, so copying or backing up an open.graphfile now fails on Windows where it previously succeeded, and so does opening it yourself for a raw read while the database is live. The error is os error 33, "another process has locked a portion of the file", even when the other process is you. Unix locks remain advisory and are unaffected. Close the database before reading the file directly — which also avoids a torn read. -
open()now fails on filesystems that cannot lock at all rather than proceeding unprotected. SetOpenOptions::allow_unlocked(true)to accept the risk on single-writer deployments. -
This detection does not cover NFSv3 exports mounted with
-o nolock(#334). Linux's NFS client servesflock/OFD locks for anolockmount out of its own local, in-kernel lock table instead of contacting the server, sotry_lockreports an ordinary local success and looks identical to a mount where locking genuinely works —allow_unlockedis never consulted because the code never learns the mount can't really lock. Two separate client hosts writing the samenolockexport can each believe they hold the lock with no cross-host coordination. There is no way to detect this from the lock call's result, sonolockNFSv3 exports are an unsupported deployment for multi-writer use — the same category as running mixed Minigraf versions against one file, below. Seedocs/ERROR_REFERENCE.mdSTG-027. -
A database genuinely locked by another process now reports its error later than before — about 375ms later.
FileBackend::open_withretries a cross-processWouldBlockup to 10 times with backoff doubling from 5ms and capped at 50ms before concluding another process holds the lock. This exists becauseforktransiently duplicates the lock-holding descriptor into a child untilexecvecloses its close-on-exec descriptors, so any process that spawns subprocesses — the normal pattern for the Python and Node bindings — could see a spuriousWouldBlockfrom its own child. SQLite solves the same problem withbusy_timeout; this is the same idea, bounded so a real conflict still fails, just not instantly. A same-process conflict (this process already holds the file open) is unaffected: it fails immediately, since waiting could never help. -
A refused open on a filesystem that cannot lock may leave a 0-byte
.graphfile behind. The file must exist before it can be locked, so the create-then-lock ordering now creates before it can know the lock will be refused. Harmless — the next open seesis_newand initialises normally — but worth knowing if you inspect the filesystem after a failed open. -
Minigraf's,WriteTransaction's, andPreparedQuery's public methods now returnResult<T, MinigrafError>instead ofanyhow::Result<T>(#277).MinigrafErrorimplementsstd::error::Error(so?still works againstanyhow::Result/Box<dyn Error>callers) and adds.category() -> ErrorCategoryand.code() -> &'static strfor structured error matching. This lands the foundation for issue #277; every error currently surfaces as the genericINT-000code until each error category's follow-up PR wires up its realdocs/ERROR_REFERENCE.mdcodes. Every error'sDisplay/to_string()output now also carries a[CODE]prefix (e.g.[INT-000] unclassified internal error: ...) that it did not have before — a behavior change beyond the type signature. Anything that string-matches error output is affected, including the interactive REPL's printed error messages and all 7 language-binding repos (Python, Node, WASM, Java, Android, Swift, C). -
src/query/datalog/parser.rsnow carries realPRS-0xxcodes instead of falling back toINT-000(#358, #277 2/6). All 79 documentedPRS-0xxcodes are wired up viabail_coded!/err_coded!at every parser call site that maps onto one; a parse error's.code()is now e.g."PRS-005"("Unclosed list") instead of the generic"INT-000", and itsDisplay/to_string()prefix changes to match.parser.rs's internal signatures changed fromResult<_, String>toanyhow::Result<_>(crate-internal only — the publicdb.execute()/db.prepare()surface is unchanged, stillResult<T, MinigrafError>). Two documented codes, PRS-028 ("window expression cannot be empty") and PRS-049 ("unexpected end of fact vector"), are wired up but unreachable through the public API — thehas_over/len() >= 4conditions that guard their call sites already guarantee the "bad" case can't occur, so no input reaches them. This is a pre-existing property of the parser logic, not something this PR changed. Two static message strings changed wording to match theirdocs/ERROR_REFERENCE.md"Error text" verbatim (previously drifted):"Retract argument must be a vector"→"...must be a vector of facts"(PRS-044). Anything that string-matches parser error output — including the interactive REPL and all 7 language-binding repos — is affected by both the code prefix and this wording change. -
Storage/file-format errors now surface their real
STG-0xxcode (#359, part of #277's rollout): all 27 documentedSTG-0xxcodes are wired up acrosssrc/storage/{mod,persistent_facts,btree,btree_v6,packed_pages}.rsandsrc/storage/backend/file.rs— header validation (STG-001–STG-009), header re-read failures (STG-010), on-disk B+tree/packed-page corruption (STG-011–STG-015), a poisoned backend mutex (STG-016), page-count arithmetic overflow guards (STG-017–STG-024), and the three lock conflict errors (STG-025–STG-027).docs/ERROR_REFERENCE.md's STG section's "Error text" lines were rewritten from one-off example values to canonical{}-placeholder templates so they byte-for-byte match theREGISTRYentries the sync test checks against.FileBackend::open_withno longer collapses a more specific header-parse failure (e.g. a bad magic number) into the genericSTG-010— it now propagates aCodedErroralready present in the chain as-is, only falling back toSTG-010for a genuinely uncoded (e.g. raw I/O) failure. Not covered by a dedicated regression test:STG-011's call site is guarded by adebug_assert!that always fires first in a debug/test build, so it's release-build-only and unreachable undercargo test; the 8 arithmetic-overflow guards (STG-017–STG-024) andSTG-027(a filesystem that refuses locking outright — no CI runner provides one) are covered by the registry↔doc sync test and code review but have no dedicated trigger test, consistent with their own "should not occur under normal operation" documentation. Twobail!sites with no matching documentedSTGcode ("Header checksum mismatch", reached on a genuine on-disk corruption of an otherwise-valid header, andinto_backend's multiple-owners precondition) were deliberately left uncoded (INT-000) rather than assigned an undocumented code — left for the final API+INT audit PR. -
src/wal.rs's errors now carry realWAL-0xxcodes instead of falling back toINT-000(#360, toward #277): a bad.walmagic number isWAL-001, an unsupported WAL version isWAL-002, a fact whose serialised size exceeds the WAL's per-entry limit (~4080 bytes) isWAL-003, a fact whose serialised size exceedsu32::MAXisWAL-004(practically unreachable — WAL-003's limit fires first), a WAL header'snum_factsexceeding the platformusizeisWAL-005(unreachable on any 64-bit target), and a failure to delete the sidecar.walfile after a successful checkpoint isWAL-006.docs/ERROR_REFERENCE.md's WAL-002/003/004/006 "Error text" entries were rewritten from a concrete example value to the canonical{}-placeholder template form so the registry↔doc sync test can compare them byte-for-byte; the concrete example now lives in each entry's prose instead. -
Query executor errors now carry real
QRY-00Ncodes (#361, part of #277) instead of the genericINT-000fallback:QRY-001invalid entity,QRY-002attribute must be a keyword,QRY-003cannot transact a pseudo-attribute,QRY-004invalid value,QRY-005transaction failed,QRY-006retraction failed,QRY-007unknown predicate,QRY-008functions lock poisoned,QRY-009rules lock poisoned — the full QRY category (9 of 9 documented codes; seedocs/ERROR_REFERENCE.md). Anything that was matching onMinigrafError::code() == "INT-000"for one of these conditions now sees the specific code instead.QRY-001/004/005/006'sERROR_REFERENCE.md"Error text" entries were also rewritten from a concrete example (e.g.`Invalid entity: "not-a-uuid"`) to the canonical{}-placeholder template form (`Invalid entity: {}`) they share withREGISTRY, per the design spec's migration convention; the worked example moved to each entry's "Example" section, unchanged.QRY-001..006(the transact/retract validation codes) are wired intoDatalogExecutor::execute_transact/execute_retract, but that code path is not reachable throughMinigraf::execute()today — the public write path routestransact/retractthroughMinigraf::materialize_transaction/materialize_retraction(src/db.rs) instead, which raise their own uncoded (stillINT-000) errors, two of which — "attribute must be a keyword" and "cannot transact a pseudo-attribute" — are already earmarked asAPI-003/API-004for the API+INT category PR (#277 step 6). -
The core binary size budget is raised from 1MB to ~1.2MB (
binary-size.yml'sSIZE_LIMIT_BYTES, and PHILOSOPHY.md §4's stated target), toward #277. The project was already at the 1MB wire before #277 started (measured ~1034 KiB pre-rollout on a controlled local build); #360 (WAL, 6 codes) and #361 (QRY, 9 codes) together added only ~9 KiB, and even that was enough to fail CI. All existing size levers —opt-level = "z",lto = true,codegen-units = 1,strip = "symbols", and the non-genericbail_coded!/err_coded!/format_templatedesign (no per-call-site monomorphization, unlike the genericbuild_btreefixed for the same reason in v0.x) — were already in place;cargo bloatshows no single dominant symbol to cut, just diffuse growth across ~2000 small functions. Extrapolating to the full 130-code rollout (PRS: 79, STG: 27, WAL: 6, QRY: 9, API+INT: 9) puts the total growth at roughly 60-70 KiB over the pre-#277 baseline, which no further micro-optimization of already-maxed settings can absorb. ~1.2MB keeps meaningful headroom above that projection without abandoning the philosophy's small-binary principle. -
Every remaining
bail!/anyhow!call site in the crate now carries a structured code — nothing falls through toINT-000by default anymore (#362, #277 step 6/6, the final sub-PR of the umbrella issue). Two parts:Part 1 —
src/db.rs's API-layer call sites now carry their realAPI-0xxcodes:API-001write lock poisoned (execute/begin_write/checkpoint),API-002unexpected command variant in the write path (structurally unreachable — kept for defensive completeness, no test possible),API-003/API-004attribute-must-be-keyword / cannot-transact-a-pseudo-attribute at thematerialize_transaction/materialize_retractionlayer (mirroringQRY-002/QRY-003),API-005/API-006/API-007db.prepare()rejectingtransact/retract/rule,API-008function registry lock poisoned (register_aggregate/register_predicate),API-009WAL not initialized (also structurally unreachable —wal_write_stamped_batchalways populateswaltwo lines above theok_or_else). Regression tests assert the exact code for every reachable one, including two new lock-poisoning tests that deliberately panic while holdingwrite_lock/functions(viacatch_unwind) to exerciseAPI-001/API-008for real rather than by inspection.db.rs's own "invalid entity"/"invalid value" checks inmaterialize_transaction/materialize_retractionhave noAPI-0xxcounterpart indocs/ERROR_REFERENCE.md(onlyQRY-001/QRY-004do, and those cover a different, unreachable-from-db.rscall path — see the QRY entry above) and aWriteTransaction-already-active reentrancy check has no documented code at all, so all three became newINT-0xxcodes (INT-001,INT-002,INT-003) — part of the audit below.Part 2 — every other uncoded
bail!/anyhow!site crate-wide (~250 of them, across 18 files:temporal.rs,parser.rs's long-tail internal-error/syntax branches the PRS migration left uncoded,evaluator.rs,executor.rs's remaining sites,prepared.rs,functions.rs,rules.rs,stratification.rs,graph/storage.rs,storage/mod.rs,storage/persistent_facts.rs,storage/packed_pages.rs,storage/btree.rs,storage/btree_v6.rs,storage/backend/{file,memory}.rs,storage/cache.rs,browser/buffer.rs) — assigned one of 52 newINT-0xxcodes (INT-004–INT-055;docs/ERROR_REFERENCE.md's old "Appendix: Internal Errors" bullet list is gone, replaced byINT-006's formal entry, which covers the same 8 "internal parser error: expected X token" strings via a{}placeholder). Mechanically-repeated call sites share one parameterised code rather than getting one each — mirroring the patternQRY-005/STG-016/WAL-006already established for "wrap an arbitrary underlying cause" and "one poisoned-lock code per resource": e.g.INT-011covers all ten "unexpected end of<clause>" parser messages,INT-018covers all four "variable not bound by any outer/earlier clause" messages,INT-050covers all thirteen lock-poisoned sites across four files ({}names the specific lock), andINT-048/INT-049are broad "storage: arithmetic overflow"/"storage: internal invariant violation" buckets covering dozens of B+tree/packed-page bounds-check and overflow-guard sites that were never meaningfully distinct from one another. A handful of genuinely user-reachable conditions that the original five category PRs simply never got to — the recursion/ derived-facts/result-limit family inevaluator.rs(INT-020) and the runtime aggregate type-mismatch family infunctions.rs(INT-041–INT-044) among them — areINT-0xxrather thanQRY-0xxpurely because this PR's scope is "assign every remaining site anINTcode," not "reopen the closed PRS/QRY/STG/WAL category PRs to extend their code ranges"; two existing regression tests intests/error_codes_qry_test.rsthat had locked inINT-000for exactly these two families before this PR (recursive_rule_derived_facts_limit_returns_int000,aggregate_type_mismatch_returns_int000) were updated in place to assert their new codes instead, and two similarly-named unit tests insrc/query/datalog/evaluator.rswere fixed the same way. -
The
REGISTRY↔docs/ERROR_REFERENCE.mdsync test is now a full bidirectional equality check (#362, #277 step 6/6), replacing the subset-only check the foundation PR (#357) introduced. Renamedregistry_is_a_subset_of_error_reference_doc→registry_matches_error_reference_doc_bidirectionally: it now also asserts every documented### CODEsection with an "Error text" line has a matchingREGISTRYrow, not just the reverse. This is only safe now that Part 2 above means no documented code can be missing a registry entry because its call site hasn't been migrated yet — the last such gap is closed. The final tally: 122 pre-#362ErrorCodevariants (PRS×79,QRY×9,STG×27,WAL×6,INT-000×1) plus 9 newAPI-0xxand 55 newINT-0xx(INT-001–INT-055) = 186 total.
Bug fixes#
-
A
.graphfile is no longer bricked when its holder dies in a container (#317):FileLockrecorded the holder's PID in a.graph.locksidecar and refused to open when that PID equalled our own. Every container's main process is PID 1 in its own namespace, so a container killed with SIGKILL left the sidecar holding1, and the replacement container — also PID 1 — read its own PID back and refused. The refusal was permanent; no code path could clear it, and an operator had to delete the file by hand. Thepid == our_pidrefusal was itself the fix for #304, so the two defects were entangled.A PID cannot serve as a liveness token: it is not unique across PID namespaces and is recycled within one. The sidecar is gone, and the lock now lives in the kernel via
std::fs::File::try_lockon the.graphfile itself. The kernel releases it whenever the process exits, however it exits, so a crashed holder leaves nothing behind. Becauseflockand OFD locks attach to the open file description rather than the process, a second open within one process is also denied, so #304 stays fixed with no PID bookkeeping at all.A leftover
.graph.lockfrom a version before 1.3 is ignored and never deleted — it no longer has any effect, and a still-running old process may still depend on it. Running mixed versions against one file is not supported; upgrade all writers together.Reported and diagnosed by @ocasazza in #317, who supplied the reproduction, identified the PID-namespace mechanism, and correctly traced the refusal back to the #304 fix that caused it. That analysis is what made this straightforward to verify. One of the tests from their proposed patch is the ancestor of the regression test that now covers this.
-
Keywords may contain
?::person/alive?now lexes as a single keyword. Previously?terminated the keyword, so[:e :alive? true]lexed as:alivefollowed by a stray?symbol and was rejected as a four-element fact — reportingOptional 4th element of a fact must be a map, which named the value rather than the keyword.?is a constituent character in EDN keywords and predicate-style names (:artist/dead?) are idiomatic. Query variables are unaffected: a?not preceded by:still begins a symbol.tests/grammar/grammar.pestupdated to match. -
The parser no longer silently discards trailing input after a complete form (#305).
parse_ednreturned successfully as soon as it parsed one complete top-level value, without checking whether any tokens remained — so[1 2 3] garbage ) ] } 12345parsed as[1 2 3]with everything after it dropped, no error, no indication anything was ignored.parse_ednnow errors if tokens remain once the top-level value is parsed. Separately,(query [...])and(rule [...])accepted and silently ignored any arguments after their vector, the same defect(transact ...)/(retract ...)already closed for #303 — so(query [:find ...] :max-results 5)parsed as the query alone, discarding:max-results 5(the correct place for that clause is inside the vector:[:find ... :max-results 5]) rather than rejecting the malformed call. Both now reject extra trailing arguments with the same "unexpected trailing argument(s)" wordingtransact/retractalready use.
Performance#
-
BrowserDbno longer allocates a redundant LRU page cache (#275):open_in_memory(),open(), andimport_graph()now pass page cache capacity0toPersistentFactStorage::newinstead of256. The backingBrowserBufferBackendis already aHashMap-resident, fully-loaded page store, so everyPageCachehit was a second RAM-to-RAM copy plus LRU bookkeeping for no benefit.This required fixing
PageCacheitself: capacity0previously meant unbounded (the eviction loop'scapacity > 0guard skipped eviction but insertion still ran unconditionally), which would have made this change strictly worse — an ever-growing duplicate of every page ever read, never reclaimed.PageCache::new(0)now means the cache is disabled outright:get_or_loadalways reads through to the backend andput_dirtyis a no-op. No existing caller passed0before this change, so no other behaviour is affected. -
not/not-joinclauses are pushed to the earliest position where their variables are bound (#248):optimizer::plan()now acceptsNot/NotJoinalongsidePattern/Exprand interleaves each one right after the pattern that completes its required variables, instead of always applying it as a global post-filter after every pattern (and afteror/or-join). This shrinks the intermediate binding set flowing into any patterns that follow it in the plan. A clause whose variables are only ever bound by anor/or-join(not by any pattern in the same clause list) is deferred to the post-filter exactly as before —or/or-joinisn't visible to the planner, so pushing it early would run the clause before its dependency exists. Benchmark:negation/not_pushdown_selectivity(10k facts, three joins after thenot) shows the effect scaling with the excluded fraction, from ~2% at 0% excluded up to ~55% faster at 100% excluded. -
or/or-joinbranches can now short-circuit (#250): branches are still sorted cheapest-first (landed earlier), but once the branches evaluated so far already produce a match for every incoming binding, remaining (costlier) branches are skipped instead of always being evaluated. This is always safe foror-join— branches are projected down to already-bound outer variables before merging, so a skipped branch could only have produced a row identical to one already kept. For plainor, it's only safe when the branches evaluated so far introduce no variable beyond what's already bound in the incoming scope (a pure filter) — a branch that binds a genuinely new variable can legitimately contribute a distinct row for an already-"covered" key, so that case still evaluates every branch, exactly as before. Benchmark:disjunction/or_short_circuitanddisjunction/or_join_short_circuit(10k facts, both branches fully covering) show ~11% and ~27% improvement respectively. -
WAL writes no longer pay a full
fsync(), and can skip the per-write flush entirely (#302):WalWriter::append_entrycalledFile::sync_all()(fsync) after every write, flushing inode metadata the WAL — a pure-append file — never needed alongside the data. It now callsFile::sync_data()(fdatasync) instead: same durability guarantee, one fewer metadata flush, for every existing caller, no other change required. A newOpenOptions::synchronousfield (SyncMode::Full, the default, matches this fdatasync behavior;SyncMode::Normalskips the per-write flush altogether) lets bulk loaders/migrations that can safely re-run from a checkpoint watermark trade per-write durability for throughput —checkpoint()still fsyncs the main file unconditionally in both modes, so it remains the hard durability boundary regardless ofsynchronous. Separately:begin_write()+ Nexecute()+ onecommit()already collapsed N facts into a single WAL write / single fsync before this change; that existing batching pattern is now documented (README, wiki Performance Tuning page) since it was easy to miss.
v1.2.3 — 2026-08-10#
No changes to the core crate. Released so the language bindings have a version to republish on, carrying a fix that lives in a binding rather than here.
v1.2.2 made a second handle on an already-open file an error (#304). The
UniFFI bindings and the C API already had a way to release a handle on demand
(destroy() and minigraf_close respectively), but the Node binding did not,
and JavaScript has no deterministic destructor — so reopening a .graph file
in one process became impossible there. minigraf-node gained a close()
method (project-minigraf/minigraf-node#1); this release is what lets it ship.
Binding versions track the core version, so all seven bindings republish at 1.2.3.
v1.2.2 — 2026-08-10#
Drop-in replacement for v1.2.1. No file-format changes, no public API changes.
Bug fixes#
-
Two handles on one file in the same process (#304):
FileLock::acquiretreated a lock file whose holder PID equalled our own as stale, deleted it, and opened anyway — on the theory that it could only be a leaked handle. The far more likely cause is a handle that is open and in use right now elsewhere in the process, so the "self-heal" silently produced two liveFileBackends on one file. Each caches its ownheader.page_count, allocates new pages from that count, and bounds-checksread_pageagainst it, so the two page tables diverge. That — not an allocation off-by-one — is the source of the intermittentPage N out of bounds (total pages: M). It also corrupted the lock itself: whichever handle dropped first ranFileLock::dropand removed the lock file that by then belonged to the survivor, leaving a live handle unlocked and admitting other processes too. A lock held by our own PID is now refused with an error naming the same-process case; a lock held by a different, dead process is still reclaimed as before. Behaviour change: callers that relied on reopening a path they already hold open now get an error instead of silent corruption, and should reuse the existing handle (Minigrafis cheap to clone and all clones share one database). -
String comparison in predicates:
[(< ?a ?b)],>,<=and>=now order two strings lexicographically instead of failing. Previously every ordering predicate routed throughto_float_pair, which errors on non-numeric operands, so a string comparison silently matched no rows — and because both directions returned nothing, a range filter dropped the entire relation rather than partitioning it. This most often bit ISO-8601 timestamps stored as strings. The predicate path now agrees withvalue_lt/value_cmp, whichmin,maxand:order-byalready use. Mixed-type comparison (string vs number) remains an error, and NaN comparisons remain false.
v1.2.1 — 2026-06-28#
Drop-in replacement for v1.2.0. No file-format changes, no public API changes.
Bug fixes#
- Magic sets
fbadornment: queries with a literal keyword in value position (e.g.(reports-to ?emp :alice)) now return correct results instead of empty results (#298). Root cause: seed fact used an ephemeral random UUID as entity; the 1-arg guard bound the variable to that UUID rather than the keyword value. Fixed with a deterministic sentinel entity (Uuid::new_v5) and a 2-arg guard that binds from value position. - Magic sets recursive rules:
(reach :a ?y)with intermediate variable?midand keyword entity binding now evaluates correctly (#297).
v1.2.0 — 2026-06-26#
Drop-in replacement for v1.1.x. No file-format changes, no public API breaking changes. Upgrading requires no code changes.
Features#
- Magic Sets rewriting — recursive Datalog queries with bound arguments are now automatically rewritten top-down via magic sets, propagating bound values into recursive rules to avoid full-relation scans (#289)
- Adornment classification, seed fact generation, magic guard injection, SCC-aware propagation rule generation, full
rewrite()wired intoexecute_query_with_rules - Limitation: mutual recursion through negation is not rewritten (documented in ROADMAP §9.6)
- Adornment classification, seed fact generation, magic guard injection, SCC-aware propagation rule generation, full
- Per-query complexity limits —
:max-derived-factsand:max-resultsclauses added toquery; global default raised to 1,000,000 derived facts (#288, #290)
Bug fixes#
selective_fact_fetch: includeassertedflag in dedup key — previously retracted facts could shadow live facts under certain access patterns (#285, #286)selective_fact_fetch: restore per-pattern entity priority lost in a prior refactor (#283)- v5→v6 migration: fix hang caused by using
header.page_countas the B-tree start page instead of the correct offset (#272)
Performance#
- Eliminate backend mutex hold on cache hits in
CommittedFactLoaderImpl::resolve— read path no longer acquires the write lock when the page is cached (#279) - Pre-build
MutexStorageBackendinCommittedFactLoaderImplto eliminate oneArc::cloneperresolve()call (#280, #281)
Infrastructure#
- Split Java, Android, Swift, and C bindings into independent repos under
project-minigraforg — completes the #231 follow-up work.minigraf-java: https://github.com/project-minigraf/minigraf-javaminigraf-android: https://github.com/project-minigraf/minigraf-androidminigraf-swift: https://github.com/project-minigraf/minigraf-swiftminigraf-c: https://github.com/project-minigraf/minigraf-c
minigraf-ffiyanked from crates.io — UniFFI layer inlined into each binding repo's private shim; use the per-language packages insteadcascade.ymlupdated to dispatchcore-releaseto all 7 binding repos on every version tag- All binding repos migrated to OIDC trusted publishing —
NPM_TOKENandCARGO_REGISTRY_TOKENsecrets removed; crates.io usesrust-lang/crates-io-auth-action@v1, npm usessetup-node@v6OIDC - Python, Node, and WASM binding repos now include a
preparejob that pins theminigrafdependency version inCargo.tomland commits before building — consistent with Java, Android, Swift, and C
Documentation#
- Add
docs/ERROR_REFERENCE.md: full inventory of user-facing errors (PRS/QRY/STG/WAL/API categories, 113 entries) with cause, resolution steps, and bad-input examples; docs-only reference codes PRS-001…API-009 (#192)
v1.1.1 — 2026-05-17#
Patch release. Fixes cargo-dist Windows build failure that prevented REPL binaries and crates.io publish from completing for v1.1.0. No code changes to the library itself.
Build#
- Exclude
fuzzcrate from workspace so cargo-dist can build on Windows (MSVC linker incompatible with#![no_main]libFuzzer targets) — fixes #263
v1.1.0 — 2026-05-17#
Drop-in replacement for v1.0.0. No file-format changes, no public API changes, no query surface changes. Upgrading requires no code changes.
Performance#
- Hash-join replaces nested-loop join for multi-clause queries — O(N) instead of O(N²) for large fact sets (#202, #203, #204)
- Selective B+Tree fact fetch: queries with bound entity/attribute skip full-scan and read only relevant index pages (#208)
- Predicate push-down into join ordering (#207)
- Cost-based clause ordering for
not/orrules (#206, #205) - SIMD crossover analysis and benchmarking infrastructure (#229)
Bug fixes#
- Fixed critical correctness and durability bugs found during deep audit (#225): fact visibility edge cases, WAL entry ordering, checkpoint atomicity
- Fixed read-only handle
Droptriggering unnecessary checkpoint, modifying the file on close (#226) QueryResult::TransactedandRetractedreverted to tuple variants — struct-variant form introduced post-1.0 was a breaking change (#261)- Added
Minigraf::current_tx_count() -> u64as additive API to expose the:as-ofmonotonic counter without breaking existing pattern matches
Reliability & testing#
- WAL fault injection harness (
FaultInjectingBackend) with 9 crash-recovery tests (#209, #210, #214) - Storage/migration resilience: migration matrix, index corruption recovery, concurrency stress tests (#215, #216, #217)
- Property-based query correctness tests against a reference evaluator (proptest, #212)
- Datalog parser/evaluator fuzz targets with seed corpus (#213)
- Per-module branch coverage gates in CI (#219)
- Long-haul smoke suite: 500 entities × 10 write/read/checkpoint cycles (#220)
- XTDB and Datomic semantic compatibility tests (#221)
Internal#
- Workspace-wide clippy lint enforcement — 336 violations fixed (#232)
- Grammar conformance test harness: pest shadow grammar + EDN corpus (#233)
- CI: codecov-action v3→v5, actions-rs→dtolnay migration
Reliability Hardening — 2026-05-17#
Summary#
Six PRs hardening the v1.0.0 codebase: WAL fault injection, storage/migration resilience, query correctness (property-based + coverage gates), long-haul smoke testing, and XTDB/Datomic semantic compatibility. No API changes, no file-format changes, no new runtime dependencies.
PR #254 — #209 + #210 + #214: WAL Fault Injection
FaultInjectingBackend(#[cfg(test)]-only wrapper aroundStorageBackend) — configurable write-fail, flush-fail, read-fault injection- 9 new WAL fault tests: write-fail propagation, flush-fail without data corruption, read-fault on WAL replay, CRC corruption discard, checkpoint atomicity under backend failure, partial checkpoint recovery via WAL replay, multi-writer serialisation, concurrent write+checkpoint, backend error propagation as
Errnot panic
PR #257 — #215 + #216 + #217: Storage & Migration Resilience
tests/migration_matrix_test.rs— 5 migration tests: v7 round-trip, v3 empty migrate, corrupt magic returnsErr, unsupported version returnsErr, WAL replay idempotenttests/index_corruption_test.rs— 5 corruption-resilience tests: checksum mismatch triggers index rebuild, btree leaf/internal corrupt pages returnErrwithout panic, root pointer mismatch handled, non-critical corruption still serves queries on good data- 5 new concurrency stress tests in
tests/concurrency_test.rs: stress readers during writer, failed write then success, rollback after partial work, open/write/checkpoint/query loop per thread, nightly stress loop (#[ignore])
PR #256 — #212 + #213 + #219: Property-Based Testing & Coverage Gates
tests/property_test.rs(proptest,cfg(not(wasm32))): 3 property tests — EAV model correctness vs naive reference evaluator, bi-temporal monotonicity, retract visibility invariant.github/workflows/coverage-gates.yml— per-module branch coverage thresholds; CI fails if coverage drops below gate
PR #258 — #220: Long-Haul Smoke Suite
tests/smoke_test.rs(#[ignore]nightly):smoke_large_graph_10_cycles— 500 entities × 10 attributes × 10 update cycles; 7 invariants: active count (333), retracted count, fact count bounds, temporal snapshot integrity, prepared query consistency, recursive rule transitive closure, WAL checkpoint round-trip.github/workflows/smoke.yml— nightly 5am UTC, 15-min timeout, runs--include-ignored
PR #259 — #221: XTDB & Datomic Compatibility Corpus
tests/xtdb_compat_test.rs— 10 semantic ports (Apache 2.0): EAV, tx-time:as-of, valid-time:valid-at, retraction (current + historical), Datalog join, negation, recursive rules, parameterised queries, combined bi-temporaltests/datomic_compat_test.rs— 9 independently written semantic ports: datom model, multi-entity attribute, tx-time:as-of, retract-entity, multi-variable:find, ground-value binding, parameterised query (prepared), named reusable rules, predicate expression filter
Tests#
935 tests passing (943 total, 8 ignored: 6 or+neg-cycle stratification doc tests, 1 nightly concurrency stress, 1 nightly smoke).
Optimizer & Benchmarks — 2026-05-16#
Summary#
Three optimisation PRs extending the v1.0.0 performance work. No API changes, no file-format changes, no new runtime dependencies. Test count unchanged at 850.
PR #249 — #207 + #206: Predicate Push-Down & Mixed Rule Optimization
optimizer::plan()extended to acceptExpr(predicate/arithmetic) clauses; they are interleaved at the earliest position where all their variables are bound, minimising intermediate binding setsStratifiedEvaluatorgains mixed-rule path: when a rule stratum contains both positive-only rules and rules withnot/not-join, the evaluator now evaluates all positive rules first and applies negation filters in a second pass within the same stratum — eliminates a class of ordering-dependent bugs in cross-stratum negation
PR #251 — #205: Cost-Based not/or Ordering
optimizer::plan()assigns selectivity estimates tonot/not-joinandor/or-joinclauses; they are sorted after their ground-variable producers but before unconstrained pattern scansnot/not-joinblocks placed after the patterns that produce their join variables (was: end of plan regardless);or/or-joinblocks placed by estimated output cardinality
PR #253 — #229: SIMD Benchmarking & Crossover Analysis
benches/simd_helpers.rs:valid_time_filter_simd,as_of_filter_simd,sum_simd_i64— portable SIMD kernels usingwide::i64x4/u64x4- Criterion benchmark groups:
simd_temporal,simd_as_of,simd_aggregate— scalar vs. SIMD crossover analysis at 1K–1M facts - Analysis result: SIMD crossover occurs at ~8K–16K facts per query; scalar path preferred below that threshold; SIMD integration deferred to post-1.0 backlog pending real-workload profiling
Tests#
850 tests passing (844 passing + 6 ignored: confirmed or+neg-cycle stratification bug, deferred to post-1.0 backlog). All optimizer PRs are optimisations; no new integration tests added.
Performance — 2026-05-15#
Summary#
Four O(N²) query-engine bottlenecks eliminated. No API changes, no file-format changes, no new dependencies.
PR #246 — #208: B+Tree Selective Lookup
get_facts_by_entity,get_facts_by_attribute,get_facts_by_entity_attributepromoted from#[cfg(test)]to production insrc/graph/storage.rs- New
selective_fact_fetchhelper inexecutor.rs: inspects query patterns for bound entity literals and bound attribute strings; calls index-driven fetches instead ofget_all_facts()when ≤4 distinct lookups detected as_ofqueries continue to use full scan (required for correctness)- New benchmark groups:
btree_lookup/entity_point,btree_lookup/attribute_scan
PR #247 — #202 + #203 + #204: Hash-Join Cluster
- #202 (
executor.rs):not/not-joinbodies pre-computed once intoHashSet<Vec<(String, Value)>>keyed on join variables; O(1) probe per outer binding replaces O(N) re-scan.normalize_valuehandleskeyword→Refrepresentation asymmetry in value position. - #203 (
executor.rs):or/or-joinbranches now evaluated from a single empty seed; branch results hash-joined back onto incoming bindings on shared user-visible variables (__-prefixed metadata keys excluded).or-joinprojects tojoin_varsbefore joining. - #204 (
matcher.rs):join_with_patterndetects join variable (entity position first, value position second), buildsHashMap<Value, Vec<Bindings>>once, probes per existing binding O(1).normalize_join_valuehandleskeyword→Ref. Falls back to nested-loop when no join variable found.
Tests#
850 tests passing (844 passing + 6 ignored: confirmed or+neg-cycle stratification bug, deferred to post-1.0 backlog).
v1.0.0 — Cross-Platform Release (2026-05-01)#
Milestone#
This is the v1.0.0 release. The public Rust API and the .graph file format are now stable
and committed to semantic versioning. File format stability is guaranteed from this release.
Cross-platform summary#
All cross-platform targets have shipped:
- 8.1a — Browser WASM (
BrowserDb,IndexedDbBackend,@minigraf/browseron npm) — v0.20.0 - 8.1b — Server-side WASM (
wasm32-wasip1/ WASI, Wasmtime/Wasmer CI) — v0.20.0 - 8.2 — Mobile bindings (Android
.aaron GitHub Packages, iOS.xcframeworkvia SPM, UniFFI) — v0.21.0 - 8.3a — Python (
minigrafon PyPI, pre-built wheels) — v0.22.0 - 8.3b — Java/JVM (
io.github.adityamukho:minigraf-jvmon Maven Central, fat JAR) — v0.23.0 - 8.3c — C FFI (
minigraf.h+ platform tarballs on GitHub Releases) — v0.24.0 - 8.3d — Node.js (
minigrafon npm, pre-built.nodebinaries) — v0.25.0
Also in this release#
pkg/renamed tominigraf-wasm/,swift/renamed tominigraf-swift/— consistent top-level naming across all workspace packages (issue #179)@minigraf/browsernow published to npm on every tagged release (issue #179)@minigraf/wasipublished to npm on every tagged release (issue #178) — WASI binary packaged for Node.js WASI consumers- Per-platform READMEs added:
minigraf-wasm/,minigraf-node/,minigraf-ffi/python/,minigraf-c/,minigraf-ffi/java/
Tests#
795 tests passing (788 passing + 7 ignored: confirmed or+neg-cycle stratification bug,
deferred to post-1.0 backlog).
v0.25.0 — 2026-04-26#
Added#
- Node.js bindings published to npm as
minigraf. Install withnpm install minigraf. No build step required — prebuilt.nodebinaries for Linux x86_64/aarch64, macOS universal2, Windows x86_64. API:new MiniGrafDb(path),MiniGrafDb.inMemory(),.execute(datalog),.checkpoint(). Full TypeScript definitions included.
v0.24.0 — C Bindings (2026-04-26)#
Added#
- C bindings distributed as GitHub Releases tarballs.
Download
minigraf-c-v0.24.0-<platform>.tar.gz(Linux/macOS) or.zip(Windows) from the release page. Each archive contains the prebuilt shared library plusminigraf.h. API:minigraf_open,minigraf_open_in_memory,minigraf_execute,minigraf_string_free,minigraf_checkpoint,minigraf_close,minigraf_last_error. Memory contract mirrors SQLite:minigraf_executereturns a heap-allocated JSON string; callminigraf_string_freeto release it. minigraf-c/: new workspace crate (cdylib+staticlib) —Cargo.toml,src/lib.rsminigraf-c/cbindgen.toml: cbindgen 0.29.2 configurationminigraf-c/include/minigraf.h: committed stable header (cbindgen-generated).github/workflows/c-ci.yml: PR test matrix on 4 platforms + header drift check.github/workflows/c-release.yml: release workflow — builds + packages platform tarballs, uploads to GitHub Releases
795 tests.
v0.23.0 — Java Desktop JVM Bindings (2026-04-25)#
Added#
- Java desktop JVM bindings published to Maven Central as
io.github.adityamukho:minigraf-jvm:0.23.0. Add to Gradle:implementation("io.github.adityamukho:minigraf-jvm:0.23.0"). Fat JAR with embedded natives for Linux x86_64/aarch64, macOS universal2, Windows x86_64. API:MiniGrafDb.open(path),MiniGrafDb.openInMemory(),.execute(datalog),.checkpoint(). minigraf-ffi/java/: Gradle 8.11 project —build.gradle.kts,settings.gradle.kts,NativeLoader.kt(runtime native extraction from JAR resources), and Gradle wrapperminigraf-ffi/java/src/test/kotlin/.../BasicTest.kt: JUnit 5 suite (in-memory, transact/query, error handling, file-backed persistence).github/workflows/java-ci.yml: PR test matrix on 4 platforms (Linux x86_64, Linux aarch64, macOS universal2, Windows x86_64).github/workflows/java-release.yml: release workflow — cross-compiles natives on 4 platforms, assembles fat JAR, publishes to Maven Central via Sonatype OSSRH
795 tests.
v0.22.0 — Python Bindings (2026-04-25)#
Added#
- Python bindings published to PyPI as
minigraf. Install withpip install minigraf. API:MiniGrafDb.open(path),MiniGrafDb.open_in_memory(),.execute(datalog),.checkpoint(). Pre-built wheels for Linux x86_64/aarch64, macOS universal2, Windows x86_64.
v0.21.1 — Patch: mobile/WASM docs (2026-04-19)#
Changed#
src/lib.rs: added Feature Flags section and WebAssembly targets subsection to crate-level docs — browser feature,wasm32-unknown-unknowntarget switcher note, and WASI build commandREADME.md: updated "For Mobile Apps" section — replaced the cross-platform placeholder with current state, added Kotlin/Swift quick-start snippets and link to wiki integration guide- Wiki
Use-Cases.md: replaced Integration placeholder with full Android (Gradle setup, Kotlin API, error handling, threading) and iOS (SPM setup, Swift API, error handling, async) integration guides
795 tests.
v0.21.0 — Android/iOS Mobile Bindings (2026-04-19)#
Added#
minigraf-fficrate: UniFFI 0.31 bindings exposingMiniGrafDb(open, openInMemory, execute, checkpoint) andMiniGrafError(Parse, Query, Storage, Other) to Kotlin and Swift- Android
.aarrelease artifact, published to GitHub Packages (io.github.adityamukho:minigraf-android) - iOS
.xcframeworkrelease artifact, distributed via Swift Package Manager (Package.swiftat repo root) mobile.ymlCI workflow: cross-compiles Android targets withcargo-ndk, generates Kotlin/Swift UniFFI bindings, assembles AAR with Gradle, assembles xcframework withxcodebuild, and publishes both on every tagdocs-checkCI job inrust.ymlandrelease.yml— gates releases oncargo doc --all-featurespassing cleanly
Fixed#
release.yml: addeddocs-checktohostjob'sneedsandifconditionwasm-release.yml/mobile.yml: retry loops extended from 20 to 40 attempts;inputs.tag || github.ref_nameordering correctedminigraf-ffi/android/gradlew: removed inner double-quotes fromDEFAULT_JVM_OPTSand replaced xargs/sed eval block with directexec— fixes "Could not find main class" and garbled usage outputminigraf-ffi/android/build.gradle.kts: addedandroid { publishing { singleVariant("release") } }— fixes AGP 8.x "SoftwareComponent 'release' not found"mobile.ymlPackage.swift commit: pushes to unprotectedswift-releasesbranch and moves tag viagh api -F force=true— avoids branch-protection blocks and string/boolean type mismatch
795 tests.
v0.20.1 — Patch: docs.rs browser module visibility (2026-04-19)#
Fixed#
browsermodule now appears on docs.rs: addeddocsrsto thecfggate anddoc(cfg(...))badge annotation (src/lib.rs)
v0.20.0 — WebAssembly Support (2026-04-18)#
Added#
- Browser WASM (
wasm32-unknown-unknown+wasm-bindgen):BrowserDbpublic API:open_in_memory(),execute(),checkpoint(),export_graph(),import_graph()BrowserBufferBackend— in-memoryStorageBackendover a flat page buffer, identical byte layout to the native.graphformatIndexedDbBackend— page-granular IndexedDB storage (one 4 KB entry per page); only dirty pages written on checkpointwasm-packbuild workflow (wasm32-unknown-unknown --features browser) generatingminigraf-wasm/with JS glue and TypeScript definitionswasm-bindgen-testbrowser integration tests (Chrome + Firefox viawasm-pack test)
- Server-side WASM (
wasm32-wasip1/ WASI):FileBackendverified under WASI's capability-based filesystem (no backend changes needed)- CI workflow (
wasm-wasi.yml) builds, unit-tests, and smoke-tests under Wasmtime and Wasmer on every push/PR - Thread-dependent tests gated with
#[cfg(not(target_os = "wasi"))]
- Cross-platform compatibility tests (issue #150):
tests/cross_platform_compat_test.rs— native round-trip (raw page byte copy) and fixture-readability teststests/fixtures/compat.graph— committed v7 binary fixture containing:alice :name "Alice"and:alice :age 30examples/generate_compat_fixture.rs— reproducible fixture generator (native only; no-op on wasm32)native_fixture_readable_by_browser_dbwasm-bindgen-test — loads native fixture viaBrowserDb::import_graph, verifies both facts
- Release workflow: WASM artifacts (WASI binary + browser tarball) built and attached on every tag;
cargo publishto crates.io on release
795 tests.
v0.19.0 — Publish Preparation (2026-04-08)#
Changed (breaking — internal visibility only)#
Minigraf::repl()factory method replaces directRepl::new(FactStorage)constructor — users calldb.repl().run()instead- All internal types narrowed to
pub(crate):FactStorage,PersistentFactStorage,FileHeader,StorageBackend,DatalogExecutor,PatternMatcher,Fact,TxId,VALID_TIME_FOREVER,Wal, and all related internals Minigraf::inner_fact_storage()removed (was unused)
Added#
Minigraf::repl(&self) -> Repl<'_>— constructs an interactive REPL session;Replnow borrows&Minigraffor lifetime safety- Full rustdoc on all public API items with
# Examplesdoctests [package.metadata.docs.rs]inCargo.toml— docs.rs builds withall-features = true#![warn(missing_docs)]— enforces documentation coverage going forward- crates.io and docs.rs badges in
README.md - Installation section in
README.md(cargo add minigraf/[dependencies]block) - macOS and Windows added to CI test matrix (
rust.yml) - Strict
cargo clippy -- -D warningsstep inrust-clippy.yml
Fixed#
- Bare
.unwrap()in library code replaced with.expect("lock poisoned")(RwLock operations incache.rs,evaluator.rs) and.expect("WAL not initialized")(db.rs) FileHeader::to_bytesnow takesselfby value (clippywrong_self_convention)- Broken intra-doc link
[Repl::run]indb.rsfixed to[crate::repl::Repl::run]
788 tests.
v0.18.0 — Prepared Statements (2026-04-04)#
Added#
Minigraf::prepare(query_str) -> Result<PreparedQuery>— parse and plan a query once, returning aPreparedQuerythat can be executed many times with different bind valuesPreparedQuery::execute(bindings: &[(&str, BindValue)]) -> Result<QueryResult>— substitute named$slottokens and run against the current fact store state; plan is reused on each callBindValueenum —Entity(Uuid),Val(Value),TxCount(u64),Timestamp(i64),AnyValidTime; each variant is permitted only in the appropriate bind-slot position$identifierbind slot tokens in parser — accepted in entity position, value position,:as-of, and:valid-at; attribute position is intentionally rejected at prepare timeEdnValue::BindSlot(String),AsOf::Slot(String),ValidAt::Slot(String),Expr::Slot(String)AST variants (parse-only; panic at runtime if unsubstituted)BindValueandPreparedQueryre-exported fromlib.rs(public API surface)tests/prepared_statements_test.rs— 17 integration tests covering all slot positions, combined temporal + entity parameterisation, plan reuse, and all error paths
Internal#
src/query/datalog/prepared.rs— new module:prepare_query(), substitution logic, 19 unit tests; manualDebugimpl forPreparedQuery(avoidsFactStorage: Debugbound)- Panic guards (no slot-name interpolation) in
executor.rs(4 sites) andstorage.rs(1 site) for unsubstituted slot variants; CodeQL-safe (no user-controlled string in panic message)
Unchanged#
db.execute(str)string API — no breaking change- Executor, optimizer, matcher — no changes required
v0.17.0 — User-Defined Functions (2026-04-02)#
Added#
Minigraf::register_aggregate(name, init, step, finalise)— register a custom aggregate function usable in both:findgrouping and:over(window) clausesMinigraf::register_predicate(name, f)— register a single-argument filter predicate usable in[(name? ?var)]:whereclausesFunctionRegistry::register_aggregate_desc/register_predicate_desc(internal API)WindowFunc::Udf(String)andUnaryOp::Udf(String)AST variants for runtime-resolved functionsUdfOps,AggImpl,PredicateDesctypes infunctions.rs
Changed#
AggregateDescnow usesAggImpldiscriminator instead ofwindow_compatible+window_opsapply_expr_clausesnow returnsResult<Vec<Binding>>and accepts&FunctionRegistryeval_expracceptsOption<&FunctionRegistry>for UDF predicate resolutionWindowSpec::func_name()now returnsStringinstead of&'static str- Parser emits
Udfvariants for unknown names instead of erroring (runtime validation)
Test count: 727 tests#
v0.16.0 — Window Functions (2026-04-02)#
Added#
- Window functions in Datalog
:findclause:(sum ?v :over (...)),(count ?v :over (...)),(min ?v :over (...)),(max ?v :over (...)),(avg ?v :over (...)),(rank :over (...)),(row-number :over (...))with unbounded-preceding (cumulative from partition start to current row) frame :partition-by ?varoptional clause: absent means whole result set is one partition:order-by ?varrequired in every:overclause;:descoptional (default ascending)FunctionRegistry(src/query/datalog/functions.rs): string-keyed registry of aggregate descriptors; all built-in aggregates migrated into it;window_ops(init/step/finalise) on window-compatible entries;is_builtinflag separates built-ins from future UDFs- Mixed queries: regular aggregates and window functions may coexist in the same
:findclause; aggregates collapse rows first, windows annotate over collapsed rows AggregateDesc,AggState,WindowOpstypes infunctions.rsWindowFunc,Order,WindowSpec,FindSpec::Windowtypes intypes.rstests/window_functions_test.rs: 12 integration tests (cumulative sum, running count/min/avg, rank with ties, row-number, partition-by, desc ordering, mixed aggregate+window, single-row and empty-result edge cases, lag/lead parse rejection)
Changed#
FindSpec::Aggregate { func }: type offuncchanged fromAggFuncenum toString; dispatch goes throughFunctionRegistry— internal change, no public API impactAggFuncenum removed fromtypes.rs; all aggregate dispatch centralised infunctions.rsapply_aggregationandapply_agg_funcremoved fromexecutor.rs; replaced byapply_post_processing+ helpers
Total#
707 tests (unit + integration + doc)
v0.15.0 — Temporal Metadata Bindings (2026-04-01)#
Added#
- Temporal pseudo-attributes:
:db/valid-from,:db/valid-to,:db/tx-count,:db/tx-id, and:db/valid-atare now first-class bindable values in Datalog:wherepatterns PseudoAttrenum andAttributeSpecwrapper type intypes.rs— clean type-safe representation for real vs. pseudo attributes inPatternparse_query_patterninparser.rs— detects:db/*keywords in the attribute position; rejects them in entity/value positions (parse error)PatternMatcher::from_slice_with_valid_atconstructor — passes query-levelvalid_atinto the matcher- Hard-error guard in executor: per-fact pseudo-attrs (
:db/valid-from,:db/valid-to,:db/tx-count,:db/tx-id) require:any-valid-time; error message tells user exactly what to add :db/valid-atbinds the effective query timestamp: explicit:valid-at <ts>→Value::Integer(ts), no:valid-at→Value::Integer(now),:any-valid-time→Value::Null:any-valid-timenow accepted as a standalone top-level query keyword (previously required:valid-at :any-valid-timeform)tests/temporal_metadata_test.rs: 16 new integration tests covering time-interval range queries, time-point lookups, tx-time correlation,:db/valid-atsemantics, and all parse/runtime error guards
Total#
647 tests (438 unit + 209 integration)
v0.14.0 — Tests + Error Coverage (2026-03-31)#
Added#
tests/production_patterns_test.rs: 8 cross-feature integration tests combining not+as-of, not-join+count, count+not, count+valid-at, recursion+not, or+count, or+sum, count+as-of-sequencetests/error_handling_test.rs: 8 integration-level error-path tests covering runtime type errors (sum/string, sum/mixed, max/boolean), stratification errors (negative cycles), and parse safety errors (not-join unbound join var, or mismatched vars, aggregate unbound var)- Stream 3: ~109 unit tests for parser-unreachable branches and aggregation/arithmetic edge cases in
executor.rsandevaluator.rs cargo-llvm-covbranch coverage command documented inCONTRIBUTING.md- CI coverage enforcement:
cargo-tarpaulin --fail-under 75gates every PR; Codecov 75% threshold with 2% drop tolerance;fail_ci_if_error: true - Nightly
cargo-llvm-cov --branchworkflow: uploads LCOV to Codecov (branch-coverageflag) and attaches HTML artifact (30-day retention); also triggerable viaworkflow_dispatch - Codecov badge added to
README.md
Coverage#
- Branch coverage:
executor.rs~85.71% (from ~75%),evaluator.rs~89.29% (from ~73%) - Remaining uncovered branches: NaN-check defensive code not reachable via public API
- Total: 617 tests (424 unit + 187 integration + 6 doc)
Known Issues#
or-with-negative-cycle: stratification does not currently detect negative cycles insideorbranches. Tracked via#[ignore]intests/error_handling_test.rs::or_negative_cycle_rejected.
[0.13.1] — 2026-03-27#
Performance#
filter_facts_for_querysnapshot fix — function now returnsArc<[Fact]>instead of a throwawayFactStorage, eliminating the O(N) four-BTreeMap index rebuild that occurred on every non-rules query call.execute_querypath constructs zeroFactStorageobjects.execute_query_with_rulesstill convertsArc<[Fact]>back toFactStorageforStratifiedEvaluator(deferred).- ~62–65% speedup on non-rules queries at 10K facts:
query/point_entity/10k22 ms → 8.6 ms;aggregation/count_scale/10k28 ms → 9.7 ms. - Evaluator loop:
accumulated_factscomputed once per iteration (was 4 separateget_asserted_facts()calls).
Added#
PatternMatcher::from_slice(Arc<[Fact]>)constructor — creates a matcher from an immutable fact snapshot without index reconstruction.
Technical#
apply_or_clausesandevaluate_not_joinsignatures updated to acceptArc<[Fact]>instead of&FactStorage.- 6 new tests: 4 in
matcher.rs(unit), 2 inexecutor.rs(unit).
Tests#
- Total: 568 tests passing (390 unit + 172 integration + 6 doc)
[0.13.0] — 2026-03-26#
Added#
- Disjunction (
or/or-join): queries and rule bodies can now use(or branch1 branch2 ...)and(or-join [?v...] branch1 branch2 ...)where-clauses. Branches support all other clause types includingnot,not-join,Expr, and nestedor/or-join.(and ...)groups multiple clauses into a single branch. match_patterns_seededonPatternMatcherfor seeded branch evaluation.evaluate_branchandapply_or_clausesaspub(crate)helpers inexecutor.rs.
Technical#
WhereClauseenum gainsOr(Vec<Vec<WhereClause>>)andOrJoin { join_vars, branches }variants.DependencyGraph::from_rulesrefactored with recursivecollect_clause_depshelper;Or/OrJoinbranches contribute positive dependency edges.- Rules with
or/or-joinin their bodies route to themixed_rulespath inStratifiedEvaluator.
[0.12.0] - 2026-03-25#
Added#
BinOpenum (14 variants:Lt,Gt,Lte,Gte,Eq,Neq,Add,Sub,Mul,Div,StartsWith,EndsWith,Contains,Matches) intypes.rsUnaryOpenum (5 variants:StringQ,IntegerQ,FloatQ,BooleanQ,NilQ) intypes.rsExprenum (Var,Lit,BinOp,UnaryOp) — composable expression AST intypes.rsWhereClause::Expr { expr: Expr, binding: Option<String> }variant —None= filter,Some(var)= arithmetic bindingparse_expr_arg/parse_expr/parse_expr_clauseinparser.rs; dispatch at all 4 clause sites (query:where, rule body,notbody,not-joinbody)- Parse-time regex validation for
matches?patterns viaregex-lite; invalid patterns are rejected with a clear error check_expr_safety+check_expr_safety_with_boundinparser.rs— forward-pass safety check; recurses intonot/not-joinbodies; unboundExpr::Varreferences are rejected at parse timeouter_vars_from_clauseupdated forWhereClause::Expr— binding variable contributes to scope for subsequent clauseseval_expr,eval_binop,is_truthy,apply_expr_clausesinexecutor.rs— evaluate expression trees against a binding; type mismatches and div/0 silently drop the rowapply_expr_clauses_in_evaluatorinevaluator.rs— sibling helper for rule body andnot-joinevaluation pathsnot_body_matchesinexecutor.rsupdated to seed with outer binding for expr-onlynotbodiestests/predicate_expr_test.rs— 28 integration tests covering all operators, silent-drop semantics, integer division, NaN, int/float promotion, string predicates, regex, expr innotbody, expr in rule body, bi-temporal + expr, arithmetic into aggregate
Semantics#
- Comparison operators (
<,>,<=,>=) require both operands to be numeric (IntegerorFloat); type mismatch → row dropped =/!=use structural equality onValue— type mismatch returnsfalse/true, not an error- Integer
+Floatpromotes toFloat; integer division truncates; division by zero → row dropped; NaN result → row dropped is_truthy:Boolean(true)→ true; non-zeroIntegerorFloat→ true; everything else (includingKeyword,Ref,Null, zero, empty string,Boolean(false),-0.0) → falsematches?pattern compiled at eval time viaregex-lite; pattern must be a string literal validated at parse time
Tests#
- Added
tests/predicate_expr_test.rs(28 integration tests) - Total: 527 tests passing (365 unit + 156 integration + 6 doc)
[0.11.0] - 2026-03-25#
Added#
- Aggregation in
:findclause:count,count-distinct,sum,sum-distinct,min,max :withgrouping clause — variables that participate in grouping but are excluded from output rowsAggFuncenum andFindSpecenum insrc/query/datalog/types.rs;DatalogQuery.findmigrated fromVec<String>toVec<FindSpec>;DatalogQuery.with_vars: Vec<String>field addedapply_aggregationpost-processing step inexecutor.rs— runs after binding collection when any aggregate is presentextract_variableshelper inexecutor.rs— non-aggregate extraction path (replaces inline loops)apply_agg_funcandvalue_type_namehelpers inexecutor.rsparse_aggregatehelper inparser.rs;:findarm extended to acceptEdnValue::List(aggregate expressions);:withkeyword arm added- Parse-time validation: aggregate variables must be bound in
:where;:withwithout any aggregate is rejected tests/aggregation_test.rs— 24 integration tests covering all aggregates,:with, rules, negation, temporal filters
Semantics#
count/count-distinctwith no grouping vars on zero bindings →[[0]](SQL behavior)- All other aggregates on zero bindings → empty result set
- All aggregates skip
Value::Nullsilently (SQL behavior) - Type mismatches (e.g.
sumonString) fail fast with a runtime error min/maxon mixedInteger/Floatis a runtime error:with ?vadds?vto the grouping key without adding it to output columns
Tests#
- Added
tests/aggregation_test.rs(24 integration tests) - Total: 461 tests passing (327 unit + 128 integration + 6 doc)
[0.10.0] - 2026-03-24#
Added#
src/query/datalog/stratification.rs—DependencyGraphandstratify(): analyse rule dependency graphs at registration time; programs with negative cycles are rejected with a clear errorWhereClause::Not(Vec<WhereClause>)andWhereClause::NotJoin { join_vars, clauses }variants intypes.rs; all exhaustive matches updated(not clause…)in:whereand rule bodies — stratified negation where all body variables must be pre-bound by outer clauses(not-join [?v…] clause…)— existentially-quantified negation with explicit join-variable declaration; body variables not injoin_varsare fresh/unbound- Safety check at parse time: every
notbody variable must be bound by an outer clause; everyjoin_varsvariable innot-joinmust be bound by an outer clause - Nesting constraint:
not-joincannot appear insidenotor anothernot-join— rejected at parse time StratifiedEvaluatorinevaluator.rs: stratifies rules, runs positive rules first, then appliesnot/not-joinfilters per binding for mixed rulesevaluate_not_joinfree function inevaluator.rs: builds partial binding fromjoin_vars, convertsPatternandRuleInvocationbody clauses to patterns, runsPatternMatcher; returnstrueif body is satisfiable (reject outer binding)rule_invocation_to_patternextracted aspub(super)free function fromRecursiveEvaluator- Two not-post-filter sites in
executor.rsnow handle bothNotandNotJoinviaevaluate_not_join tests/negation_test.rs— 10 integration tests fornot: basic absence, multi-clause, rule body, time-travel, negative cycle rejectiontests/not_join_test.rs— 14 integration tests fornot-join: basic exclusion, multiple join vars, multi-clause body, rule body,:as-of,:valid-at, negative cycle at registration,not+not-joincoexistence,RuleInvocationin body end-to-end
Changed#
Rule.bodychanged fromVec<EdnValue>toVec<WhereClause>to support negation clauses alongside patternsexecutor.rsexecute_query_with_rulesnow delegates toStratifiedEvaluatorinstead ofRecursiveEvaluatordirectlyrules.rsregister_rulerunsstratify()after each registration; returnsErron negative cycle (rules are not registered on error)
[0.9.0] - 2026-03-23#
Added#
src/storage/btree_v6.rs— proper on-disk B+tree for all four covering indexes (EAVT, AEVT, AVET, VAET); each B+tree node is one 4KB page (internal + leaf), withbuild_btreefor bulk-load andrange_scanfor leaf-chain traversalOnDiskIndexReaderstruct +CommittedIndexReadertrait — page-cache-backed index lookup replacing the full in-memory BTreeMap; index memory usage is now O(cache_pages), not O(facts)MutexStorageBackend<B>adapter — holds backend mutex only for the duration of a singleread_pagecall on a cache miss; cache-warm pages require no lock, enabling concurrent range scans to proceed in paralleltests/btree_v6_test.rs— 8 integration tests covering B+tree insert/range-scan, multi-page leaf chains, concurrent scan correctness with Barrier-synchronised threads, and v5→v6 migration roundtriptest_concurrent_range_scans_correctnessunit test inbtree_v6.rs— verifies all 8 concurrent threads return identical non-empty scan resultsbench_concurrent_btree_scanCriterion benchmark — measures wall-clock latency at 2/4/8 concurrent EAVT range scans; results updated inBENCHMARKS.mdFileHeaderv6 (80 bytes): addsfact_page_count u64field at bytes 72–80; automatic v5→v6 migration on first checkpoint
Changed#
FORMAT_VERSIONbumped 5→6; v5 databases auto-migrated on first saveBENCHMARKS.mdupdated with v6 open/memory improvements, concurrent B+tree scan results, heaptrack v6 numbers, and a "How to read these numbers" methodology sectionREADME.mdandBENCHMARKS.md: performance table updated to reflect v6 open-time reduction (~2.4×) and peak-heap reduction (~21%)
Fixed#
- Concurrent B+tree range scans no longer serialise on cache-warm pages —
4→8 threadscaling ratio improved from ~2.2× to ~1.9×
[0.8.0] - 2026-03-22#
Added#
BENCHMARKS.md— full Criterion benchmark results at 1K/10K/100K/1M facts with machine spec, HTML report references, and heaptrack memory profilesexamples/memory_profile.rs— heaptrack profiling binary; accepts fact count as positional argCargo.tomlmetadata:repository,keywords,categories,readme,documentationfields- Memory profile table in
README.md"Performance" section
Changed#
README.mdPerformance section now links toBENCHMARKS.mdfor full benchmark details- README status text updated to reflect the completed benchmark and publish-preparation work
- crates.io publication deferred to a later preparation release (API cleanup; file format v6 now complete)
Removed#
- Dead
clapdependency from[dependencies]—clapwas listed but never imported in library or binary code
[0.7.1] - 2026-03-22#
Fixed#
- Retraction semantics in Datalog queries:
filter_facts_for_queryStep 2 now computes the net view per(entity, attribute, value)triple vianet_asserted_facts(). Previously, retracted facts continued to appear in query results because the original assertion record remained in the append-only log. Now, for each EAV triple in the tx window, only the record with the highesttx_countis considered — if it is a retraction, the triple is excluded from results. - Oversized facts are now rejected early in
db.rs(check_fact_sizes) before any WAL write, using theMAX_FACT_BYTESconstant (4 080 bytes) exported frompacked_pages.rs. Previously, oversized facts could cause a panic deep in the page-packing path.
Added#
net_asserted_facts(facts: Vec<Fact>) -> Vec<Fact>helper insrc/graph/storage.rs: groups facts by EAV triple, keeps the record with the highesttx_count, and discards the triple if that record is a retraction. Used by bothexecutor.rsandstorage.rs.check_fact_sizes(facts: &[Fact])insrc/db.rs: validates all facts againstMAX_FACT_BYTESand returns a descriptiveErrbefore writing to the WAL.MAX_FACT_BYTES: usizeconstant insrc/storage/packed_pages.rs:PAGE_SIZE - PACKED_HEADER_SIZE - 4= 4 080 bytes.tests/retraction_test.rs— 7 integration tests covering: assert/retract with no:as-of, as-of snapshot before/after retraction boundary, re-assert after retract,:any-valid-timewith retraction, recursive rule retraction visibility at and before the retraction boundary.tests/edge_cases_test.rs— 4 integration tests covering: oversized-fact file-backed error path,MAX_FACT_BYTESexact boundary (accepted),MAX_FACT_BYTES + 1(rejected), in-memory database has no size limit.
[0.7.0] - 2026-03-22#
Added#
- Packed fact pages (
page_type = 0x02): ~25 facts per 4KB page, ~25× disk space reduction vs v4 - LRU page cache (
src/storage/cache.rs): configurable capacity (default 256 pages = 1MB) OpenOptions::page_cache_size(usize)— tune page cache capacityCommittedFactReadertrait: index-driven fact resolution via page cache (no startup load-all)- File format v5:
fact_page_formatheader field; auto-migration from v4 on first open - Page-based CRC32 checksum (v5): streams raw committed pages instead of all facts
Changed#
PersistentFactStorage::new()takespage_cache_capacity: usizeas second argument- Committed facts no longer loaded into
Vec<Fact>at startup; only pending facts held in memory FactStorage::get_facts_by_entity,get_facts_by_attributeuse EAVT/AEVT index range scans
Fixed#
- v4 databases auto-migrated to v5 packed format on first open (no data loss)
[0.6.0] - 2026-03-21#
Added#
- Four Datomic-style covering indexes (EAVT, AEVT, AVET, VAET) with bi-temporal keys (
valid_from,valid_toin all key tuples) FactRef { page_id: u64, slot_index: u16 }— forward-compatible disk location pointer (slot_index=0 in 6.1)- Canonical value encoding (
encode_value) with sort-order-preserving byte representation - B+tree page serialization for index persistence (
src/storage/btree.rs) FileHeaderv4 (72 bytes): addseavt_root_page,aevt_root_page,avet_root_page,vaet_root_page(4×8=32 bytes),index_checksum(u32), replacing thereservedfield- CRC32 sync check on open: index mismatch triggers automatic rebuild
FactStorage::replace_indexes()andindex_counts()for index lifecycle management- Query optimizer (
src/query/datalog/optimizer.rs):IndexHintenum,select_index(),plan()with selectivity-based join reordering - Join reordering skipped under
wasmfeature flag Cargo.toml[features]section withdefault = []andwasm = []- 6 integration tests in
tests/index_test.rsfor save/reload, bi-temporal, recursive rules regression
Changed#
FactStorageinternal structure:FactData { facts, indexes }under singleArc<RwLock<FactData>>for consistent snapshotsPersistentFactStorage::save()writes index B+tree pages and updates header checksumPersistentFactStorage::load()performs sync check and fast-path index loadexecutor::execute_query()now callsoptimizer::plan()before pattern matching- File format version bumped 3→4; automatic v1/v2/v3→v4 migration on first save
FORMAT_VERSIONconstant updated to 4
Fixed#
- NaN values in
Value::Floatnow canonicalize to a single bit pattern in index encoding (deterministic sort order)
[0.5.0] - 2026-03-21#
Added#
- Write-ahead log (WAL): fact-level sidecar
<db>.walwith CRC32-protected binary entries WriteTransactionAPI:begin_write()/commit()/rollback()for explicit ACID transactions- Crash recovery: WAL entries replayed on open; corrupt/partial entries discarded at first bad CRC32
- Checkpoint:
checkpoint()flushes WAL facts to.graphand deletes the WAL; auto-checkpoint on configurable threshold FileHeaderv3:last_checkpointed_tx_countfield (repurposes unusededge_countslot)FactStoragehelpers:get_all_facts(),restore_tx_counter(),allocate_tx_count()OpenOptionsbuilder:OpenOptions::new().path("db.graph").open()orMinigraf::in_memory()--file <path>CLI flag for the REPL binary- 41 new tests covering WAL, crash recovery, transactions, and checkpoint
Changed#
src/minigraf.rsreplaced bysrc/db.rs—Minigraf,OpenOptions,WriteTransactionpublic API- File format version bumped 2→3; automatic v1/v2→v3 migration on first checkpoint
- REPL version string now tracks
CARGO_PKG_VERSIONautomatically
Fixed#
- WAL-before-apply ordering: facts are now applied to in-memory state only after the WAL entry is fsynced, ensuring crash safety for both implicit (
execute()) and explicit (WriteTransaction) write paths
[0.4.0] - 2026-03-21#
Added#
- Bi-temporal support: every fact now carries transaction time (
tx_id,tx_count) and valid time (valid_from,valid_to) :as-of Nquery modifier for transaction time travel (counter or ISO 8601 timestamp):valid-at "date"query modifier for valid time point-in-time queries:valid-at :any-valid-timeto disable valid time filtering(transact {:valid-from ... :valid-to ...} [...])syntax for specifying valid time- Per-fact valid time override in transact (4-element fact vectors with metadata map)
- File format version 2 with automatic migration from version 1
Changed#
- Breaking behaviour: queries without
:valid-atnow return only currently valid facts (valid_from <= now < valid_to). Existing databases from that era are unaffected because all migrated facts havevalid_to = MAX. FactStorage::transact()now accepts an optionalTransactOptionsparameter
Fixed#
PersistentFactStorage::load()previously discarded originaltx_idwhen loading facts from disk, making time-travel queries on persisted databases incorrect
[0.3.0] - 2026-03-10#
Added#
- Datalog core implementation with recursive rules
- Entity-Attribute-Value (EAV) data model
- Pattern matching with variable unification
- Semi-naive evaluation for recursive rules
- Transitive closure support with cycle handling
- Rule registry for rule management
- Persistent storage with postcard serialization
- REPL with multi-line command support and comments
- 123 comprehensive tests (94 unit + 26 integration + 3 doc)
Changed#
- Replaced GQL-inspired syntax with Datalog EDN syntax
- Data model changed from property graph to EAV triples
- Query executor rewritten for Datalog pattern matching
[0.2.0] - 2026-02-01#
Added#
- Persistent storage backend with
.graphfile format - StorageBackend trait for platform abstraction
- FileBackend implementation (4KB pages, cross-platform)
- MemoryBackend for testing
- PersistentGraphStorage layer for serialization
- Embedded API (
Minigraf::open(),Minigraf::execute()) - Auto-save on drop
Changed#
- Graph storage now supports persistence
[0.1.0] - 2026-01-15#
Added#
- Initial release
- In-memory property graph implementation
- Basic graph operations (nodes, edges, properties)
- Interactive REPL
- Thread-safe storage with
Arc<RwLock<>>