CandidX: typed blocks with stable hashes under unused schema extensions
TLDR: Does what the Generic Value hashing used in ICRC3 does, but better. Would be happy if Dfinity takes it from here and maintains the packages or adds the function to existing ones. There is TS, Rust, Ocaml, Motoko native lib, Web assembly version that can be easily added to the Motoko compiler.
ICRC-3 addresses a real problem: Candid’s schema evolution lets an older client
discard fields it does not know, which prevents it from verifying the complete
block. Its fixed generic Value type preserves that information. The
ICRC-3 specification explains this rationale.
I want the compiler to check my application’s block schema while keeping
that property. A generic Value type does not enforce that an amount field
exists and contains a Nat, or catch a misspelled map key. I also want to
avoid carrying application field names as strings in every block. My approach
is typed records and arrays, canonical bytes derived from their used data,
and verification of the complete bytes before typed decoding. This is the
encoding foundation for a separate block protocol; CandidX itself supplies
neither archive endpoints nor certification.
CandidX a Candid-style appendix, four language packages, Wasm integration, shared fuzzing, and timing/heap measurements. The packages are version 1.0.0 under MIT. The appendix proposes upstream adoption; the profile is usable independently.
The identity we want
These should produce the same bytes and application hash:
record { a = "abc" } : record { a : text }
record { a = "abc"; b = null } : record { a : text; b : opt text }
Their canonical bytes are 4449444c016c016171010003616263.
Adding absent optional fields or unselected cases to existing variants should
preserve old hashes anywhere: nested records, arrays, and finite recursive
values. No application schema ID belongs in the preimage. Actual new data,
different scalar widths or field IDs, and extra present option wrappers change
the identity.
Authors construct ordinary typed records rather than manually assembling
string-keyed generic blocks. This uses a different hashing construction from
the existing ICRC-3 Value hash; it is a separate typed-block convention.
Why ordinary Candid is not enough
The same typed record produced three ordinary byte sequences in Motoko,
Rust, and JavaScript. Their type-table discovery orders and numbering differed
despite identical value payloads. Equivalent host representations and type
sharing can also change the table.
The Candid specification allows this freedom and value-dependent variant narrowing. There is some explicit performance history: the Rust serializer disables equivalent recursive-type deduplication because repeated unrolling is expensive. Other differences follow from an encoding relation that never
required canonical ordering. Reordering a header alone also leaves padded
integers, NaNs, and variant indices to address.
Proposed encoding
CandidX emits ordinary DIDL bytes, with a finite wire type derived from the
used data:
- Omit absent optional record fields and unselected variant cases; retain
explicitnullandreservedfields. - Preserve every present option wrapper, including
Some(None). - Join array element types using their actual fields and cases, preserve
element order, and emit every payload under the joined type. - Treat Blob and
[Nat8]identically, including empty binary data. Other
empty vectors derivevec empty. - Preserve scalar kinds and widths, exact UTF-8, numeric Candid field IDs,
signed zero, infinities, and subnormals. Canonicalize NaNs per float width. - Intern equivalent derived types, number them by a defined traversal, and
emit shortest LEB128 integers.
The type table remains necessary for decoding. It describes used data rather
than the full declared historical schema. Actual function/service and opaque
reference values are outside this data profile; unused declarations are accepted.
The complete rules and explicit resource bounds are in the appendix.
An application choosing SHA-256 computes SHA256(canonical_candid_bytes).
The libraries supply bytes, with no hash API or proposed Motoko SHA-256 builtin.
The isolated Motoko compiler patch provides to_candid_canonical(block).
Rust uses existing CandidType values, TypeScript uses existing SDK IDL types,
and OCaml uses typed witnesses with native records. Those paths derive used
types and emit bytes directly. Only the temporary native Motoko library uses
CandidX.canonicalize(to_candid(block)).
A canonical header depends on the values, so the direct encoders visit them
twice internally while performing one encoding operation. They create no
intermediate ordinary message to decode and re-encode. The Motoko runtime
reads the original native heap; source and IR interpreters have matching OCaml
adapters. A separate import-free Wasm ABI accepts typed memory nodes, and
Prim.canonicalizeCandid(blob) handles inbound ordinary messages.
Verification
Three complete fixed-suite runs passed, including native-library checks under Motoko 0.16.3. Each run checks 93 direct typed cases, 476 inbound normalization cases per backend, and 870 native optimization regressions in three fresh instances. The 61 frozen goldens remain unchanged.
Three 500-case seeded runs cover 3,390 matching typed records, 1,695 unused-schema-extension pairs, and 16,992 messages through each normalizer and Wasm normalization embedding. Canonical bytes and external hashes agree. Independent saved-data replay checks 36,453 messages, including complete normalized-data preservation.
The round-trip assertion compares decoded values with independently
constructed original typed values. Re-encoding is a secondary control.
Every encoder’s actual bytes are also decoded through Rust and TypeScript;
native Motoko checks its own recovery. Mutation controls verify that these
comparisons reject changed options, array order/length, variant tags/payloads,
large integers, text, and finite float values, including signed zero. NaNs
compare by class because their bit patterns are intentionally canonicalized.
Tests disable ordinary encoders/parsers on direct paths; compiler call-chain
checks verify the native heap interface. Extracted Cargo/npm packages and an
isolated OCaml installation have real public-API consumer tests. Wasm reuses
Rust and is counted as an integration path, not an independent codec.
Review found and fixed defects including variant joins, lazy Motoko big-integer
initialization, swallowed serializer errors, changing source shapes, spoofed
JavaScript typed-array metadata, and unpaired UTF-16 surrogates. The fixes
retain regressions. This is testing and review evidence, not a formal proof.
Rust and Motoko use their existing ordinary decoders. SDK 6.1.0 has three
reproducible limitations: a leading text BOM, omitted optional tuple elements,
and some recursive optional fields. CandidX’s optional nonmutating
decodeCompatible adapter recovers all tested original values; vanilla SDK
limitations are recorded separately. OCaml has no standalone ordinary binary
decoder in the audited repositories, so its outputs are checked using the
independent model and other language decoders.
An older schema can discard new present data. Verify complete original bytes
before typed projection; transporting each canonical block as a blob
preserves them. This addresses the information-preservation concern in the ICRC-3 specification
Performance and memory
| Workload | Rust direct / ordinary | TypeScript direct / ordinary | OCaml direct, µs | Native Motoko fallback / ordinary | Compiled Motoko direct / ordinary |
|---|---|---|---|---|---|
| small | 1.67x | 0.39x | 13.01 | 34.71x | 2.53x |
| records-128 | 4.36x | 0.52x | 283.79 | 13.28x | 1.40x |
| blob-64k | 1.04x | 0.35x | 25.88 | 33.37x | 1.11x |
| records-1024 | 5.03x | 0.57x | 2201.62 | 12.78x | 1.38x |
These are three sequential local runs, seven trials per phase, against frozen
sources. Ordinary encode/decode, direct encoding, preencoded normalization,
encode-and-normalize, and decoding canonical input are measured separately.
Hashing is outside the timings. Absolute times across runtimes are not a
language ranking. Wasmtime fuel is not IC instructions or cycles.
Direct encoding does not automatically mean faster encoding. The final Rust
descriptor-cache optimization improved record workloads by 8–10%, but its
direct record path still takes about twice the pipeline time in the isolated
capture. It requests 57–61% fewer allocation bytes on those blocks.
The actual compiled Motoko runtime uses 1.31–2.62 times ordinary encoding
instructions for the four block workloads. For 1,024 transactions, direct
encoding allocates 188,848 bytes versus 458,376 for encode-and-normalize, a
59% reduction. Numeric vectors and multibyte text are more expensive:
12.40x and 14.00x ordinary instructions in separate probes. These are local
canister measurements through PocketIC, not cycle-cost predictions.
Repeated direct-encoding batches reclaimed temporary objects and retained
outputs still decoded correctly. The numeric-vector batch hit its instruction
cap before reclamation and is recorded separately. Post-GC occupied heap is
not exact live heap because of collector partitions and bookkeeping.
V8 record workloads show lower occupied heap and ArrayBuffer use, but the
64 KiB binary workload uses more V8 heap while fewer ArrayBuffer bytes.
OCaml allocation and standalone Wasm arena high-water are reported separately.
The standalone typed ABI includes host marshalling costs and is not the same
memory path as the native Motoko heap visitor. Full ranges, allocation methods,
source fingerprints, and historical baselines are retained in the repository.
Adoption
The proposed sequence is to review the appendix and conformance vectors,
then offer opt-in integrations to Candid/spec and Rust, Motoko, and icp-js-core.
Existing ordinary encoding and heap/stable-storage behavior stay unchanged.
I would appreciate review of the identity rules, especially option presence,
primitive widths, NaNs, empty binary data, and array joins. Would this be a
useful typed-block convention, and which upstream APIs would make adoption
easiest? The local cx contract is fixed: incompatible alternatives need a
different profile so historical block bytes remain stable.