CandidX typed blocks with stable hashes under unused schema extensions

CandidX: typed blocks with stable hashes under unused schema extensions

TLDR: Does what the Generic Value hashing used in ICRC3 does, but better. Would be happy if Dfinity takes it from here and maintains the packages or adds the function to existing ones. There is TS, Rust, Ocaml, Motoko native lib, Web assembly version that can be easily added to the Motoko compiler.


ICRC-3 addresses a real problem: Candid’s schema evolution lets an older client
discard fields it does not know, which prevents it from verifying the complete
block. Its fixed generic Value type preserves that information. The
ICRC-3 specification explains this rationale.

I want the compiler to check my application’s block schema while keeping
that property. A generic Value type does not enforce that an amount field
exists and contains a Nat, or catch a misspelled map key. I also want to
avoid carrying application field names as strings in every block. My approach
is typed records and arrays, canonical bytes derived from their used data,
and verification of the complete bytes before typed decoding. This is the
encoding foundation for a separate block protocol; CandidX itself supplies
neither archive endpoints nor certification.

CandidX a Candid-style appendix, four language packages, Wasm integration, shared fuzzing, and timing/heap measurements. The packages are version 1.0.0 under MIT. The appendix proposes upstream adoption; the profile is usable independently.

The identity we want

These should produce the same bytes and application hash:

record { a = "abc" } : record { a : text }
record { a = "abc"; b = null } : record { a : text; b : opt text }

Their canonical bytes are 4449444c016c016171010003616263.
Adding absent optional fields or unselected cases to existing variants should
preserve old hashes anywhere: nested records, arrays, and finite recursive
values. No application schema ID belongs in the preimage. Actual new data,
different scalar widths or field IDs, and extra present option wrappers change
the identity.

Authors construct ordinary typed records rather than manually assembling
string-keyed generic blocks. This uses a different hashing construction from
the existing ICRC-3 Value hash; it is a separate typed-block convention.

Why ordinary Candid is not enough

The same typed record produced three ordinary byte sequences in Motoko,
Rust, and JavaScript. Their type-table discovery orders and numbering differed
despite identical value payloads. Equivalent host representations and type
sharing can also change the table.

The Candid specification allows this freedom and value-dependent variant narrowing. There is some explicit performance history: the Rust serializer disables equivalent recursive-type deduplication because repeated unrolling is expensive. Other differences follow from an encoding relation that never
required canonical ordering. Reordering a header alone also leaves padded
integers, NaNs, and variant indices to address.

Proposed encoding

CandidX emits ordinary DIDL bytes, with a finite wire type derived from the
used data:

  • Omit absent optional record fields and unselected variant cases; retain
    explicit null and reserved fields.
  • Preserve every present option wrapper, including Some(None).
  • Join array element types using their actual fields and cases, preserve
    element order, and emit every payload under the joined type.
  • Treat Blob and [Nat8] identically, including empty binary data. Other
    empty vectors derive vec empty.
  • Preserve scalar kinds and widths, exact UTF-8, numeric Candid field IDs,
    signed zero, infinities, and subnormals. Canonicalize NaNs per float width.
  • Intern equivalent derived types, number them by a defined traversal, and
    emit shortest LEB128 integers.

The type table remains necessary for decoding. It describes used data rather
than the full declared historical schema. Actual function/service and opaque
reference values are outside this data profile; unused declarations are accepted.
The complete rules and explicit resource bounds are in the appendix.

An application choosing SHA-256 computes SHA256(canonical_candid_bytes).
The libraries supply bytes, with no hash API or proposed Motoko SHA-256 builtin.
The isolated Motoko compiler patch provides to_candid_canonical(block).
Rust uses existing CandidType values, TypeScript uses existing SDK IDL types,
and OCaml uses typed witnesses with native records. Those paths derive used
types and emit bytes directly. Only the temporary native Motoko library uses
CandidX.canonicalize(to_candid(block)).

A canonical header depends on the values, so the direct encoders visit them
twice internally while performing one encoding operation. They create no
intermediate ordinary message to decode and re-encode. The Motoko runtime
reads the original native heap; source and IR interpreters have matching OCaml
adapters. A separate import-free Wasm ABI accepts typed memory nodes, and
Prim.canonicalizeCandid(blob) handles inbound ordinary messages.

Verification

Three complete fixed-suite runs passed, including native-library checks under Motoko 0.16.3. Each run checks 93 direct typed cases, 476 inbound normalization cases per backend, and 870 native optimization regressions in three fresh instances. The 61 frozen goldens remain unchanged.

Three 500-case seeded runs cover 3,390 matching typed records, 1,695 unused-schema-extension pairs, and 16,992 messages through each normalizer and Wasm normalization embedding. Canonical bytes and external hashes agree. Independent saved-data replay checks 36,453 messages, including complete normalized-data preservation.

The round-trip assertion compares decoded values with independently
constructed original typed values. Re-encoding is a secondary control.
Every encoder’s actual bytes are also decoded through Rust and TypeScript;
native Motoko checks its own recovery. Mutation controls verify that these
comparisons reject changed options, array order/length, variant tags/payloads,
large integers, text, and finite float values, including signed zero. NaNs
compare by class because their bit patterns are intentionally canonicalized.

Tests disable ordinary encoders/parsers on direct paths; compiler call-chain
checks verify the native heap interface. Extracted Cargo/npm packages and an
isolated OCaml installation have real public-API consumer tests. Wasm reuses
Rust and is counted as an integration path, not an independent codec.

Review found and fixed defects including variant joins, lazy Motoko big-integer
initialization, swallowed serializer errors, changing source shapes, spoofed
JavaScript typed-array metadata, and unpaired UTF-16 surrogates. The fixes
retain regressions. This is testing and review evidence, not a formal proof.

Rust and Motoko use their existing ordinary decoders. SDK 6.1.0 has three
reproducible limitations: a leading text BOM, omitted optional tuple elements,
and some recursive optional fields. CandidX’s optional nonmutating
decodeCompatible adapter recovers all tested original values; vanilla SDK
limitations are recorded separately. OCaml has no standalone ordinary binary
decoder in the audited repositories, so its outputs are checked using the
independent model and other language decoders.

An older schema can discard new present data. Verify complete original bytes
before typed projection; transporting each canonical block as a blob
preserves them. This addresses the information-preservation concern in the ICRC-3 specification

Performance and memory

Workload Rust direct / ordinary TypeScript direct / ordinary OCaml direct, µs Native Motoko fallback / ordinary Compiled Motoko direct / ordinary
small 1.67x 0.39x 13.01 34.71x 2.53x
records-128 4.36x 0.52x 283.79 13.28x 1.40x
blob-64k 1.04x 0.35x 25.88 33.37x 1.11x
records-1024 5.03x 0.57x 2201.62 12.78x 1.38x

These are three sequential local runs, seven trials per phase, against frozen
sources. Ordinary encode/decode, direct encoding, preencoded normalization,
encode-and-normalize, and decoding canonical input are measured separately.
Hashing is outside the timings. Absolute times across runtimes are not a
language ranking. Wasmtime fuel is not IC instructions or cycles.

Direct encoding does not automatically mean faster encoding. The final Rust
descriptor-cache optimization improved record workloads by 8–10%, but its
direct record path still takes about twice the pipeline time in the isolated
capture. It requests 57–61% fewer allocation bytes on those blocks.

The actual compiled Motoko runtime uses 1.31–2.62 times ordinary encoding
instructions for the four block workloads. For 1,024 transactions, direct
encoding allocates 188,848 bytes versus 458,376 for encode-and-normalize, a
59% reduction. Numeric vectors and multibyte text are more expensive:
12.40x and 14.00x ordinary instructions in separate probes. These are local
canister measurements through PocketIC, not cycle-cost predictions.

Repeated direct-encoding batches reclaimed temporary objects and retained
outputs still decoded correctly. The numeric-vector batch hit its instruction
cap before reclamation and is recorded separately. Post-GC occupied heap is
not exact live heap because of collector partitions and bookkeeping.

V8 record workloads show lower occupied heap and ArrayBuffer use, but the
64 KiB binary workload uses more V8 heap while fewer ArrayBuffer bytes.
OCaml allocation and standalone Wasm arena high-water are reported separately.
The standalone typed ABI includes host marshalling costs and is not the same
memory path as the native Motoko heap visitor. Full ranges, allocation methods,
source fingerprints, and historical baselines are retained in the repository.

Adoption

The proposed sequence is to review the appendix and conformance vectors,
then offer opt-in integrations to Candid/spec and Rust, Motoko, and icp-js-core.
Existing ordinary encoding and heap/stable-storage behavior stay unchanged.

I would appreciate review of the identity rules, especially option presence,
primitive widths, NaNs, empty binary data, and array joins. Would this be a
useful typed-block convention, and which upstream APIs would make adoption
easiest? The local cx contract is fixed: incompatible alternatives need a
different profile so historical block bytes remain stable.

One thing I’d like to understand better: ICRC-3’s Value hash reuses the IC’s representation-independent hash (the one used for request IDs), so it commits to field names and stays encoding-agnostic. Since Candid only puts 32-bit idx hashes of field names on the wire, a CandidX block isn’t self-describing. How do you expect explorers, indexers, and auditors to get the schema needed to interpret blocks, and should the schema be committed to somewhere?

Also, since the block hash only commits to field IDs, and idx-hash collisions are easy to construct, two schemas could assign different names and meanings to the same ID without changing the hash. Is that acceptable for audit purposes, or do you see a way to bind names into the identity? Candid’s field ID is hash(name) = Σ name[i] · 223^(k−i) mod 2^32, a simple polynomial with no cryptographic properties

Thanks for the input Quint!

Previously I’ve heard this in defence of the Generic Value (GV) and that workflow has been taken care of in CandidX

https://forum.dfinity.org/t/toward-a-cleaner-icrc-3-one-format-for-icrc-1-2-clear-rules-for-the-future/54934/18

Explorers get it from the published in canister Candid metadata. It’s actually easier to display a block, because it doesn’t need special treatment, same thing the IC dashboard does for any canister method it can do here as well. The way I envision it, if one uses CandidX to hash their blocks, the blocks will appear same way they appear in the ledger’s get_transactions. https://dashboard.internetcomputer.org/canister/mxzaz-hqaaa-aaaar-qaada-cai#get_transactions

CandidX is encoder-independent, while its canonical format and value model are Candid-specific.

Interesting point.

record { credit_rarxagrx : nat }
record { debit_frudenvf : nat }

Both names produce ID 3993915879.
Notes: Adding named fields does not change existing field IDs. Adding or reordering fields leaves old IDs unchanged; variant tags follow the same rule. A new name that collides with an existing ID in the same record or variant is rejected.

Bitcoin blocks are also not self-describing, so is that really a requirement?
The canister metadata containing the interface description is governed by the controllers and signed by the IC. If they go malicious, they can do a lot more than change the metadata.

It doesn’t seem to be a problem for sending transactions with update calls, so why would it be a problem for reading them in a log? I mean someone could turn icrc1_transfer args from

type TransferArg = record {
  to : Account;
  fee : opt nat;
  memo : opt blob;
  from_subaccount : opt blob;
  created_at_time : opt nat64;
  amount : nat;
};

into

type TransferArg = record {
  get_rarxagrx : Account;
  free_frudenvf : opt nat;
  donuts_aarxagrx : nat;
};

and it will produce a valid encoded transaction. Btw CandidX could be also used to hash arguments and sign them additionally.
Provided that controllers aren’t malicious and don’t break the rules when changing the schema - never delete fields, only add optional fields or variants, I don’t think being self-describing will buy us anything. If you have a concrete attack scenario, it would be great to hear it.

If there is a good reason for it, we can make CandidX self describing - containing the field names of used fields in the header - once, not everywhere like GV, should be compact enough and not lose properties. What I don’t like about GV is:

  • easy to make typing errors, typecheck won’t catch them
  • hard to catch errors
  • need to write encoders and decoders from typed to GV and back
  • if an error slips in you need to keep it around for backwards compatibility
  • slower to encode
  • much bigger in size
  • painful to write code that reads it - or it use a helper that’s slow
  • explorers need special GV readers to display it well

It barely works for something like icrc ledgers, but if you have an app that will make use of many canisters and many different kinds of blocks and clients which read and process them, it will not be very nice.

Here’s a concrete attack scenario, using your own collision pair: a controller upgrades the schema, replacing credit_rarxagrx : nat with debit_frudenvf : nat. Same ID, same type, so every historical block hash still verifies and Candid compatibility checks see no change, but every past credit now renders as a debit. Your same-record collision rule doesn’t catch it, since the old name is replaced, not kept alongside. Under the Value hash this rename breaks verification.

That’s the property I care about for audit: a certified log should stop even the operator from reinterpreting history, not just from rewriting it. Canister metadata doesn’t help here, since it only certifies the current schema.

I agree with much of your criticism of GV ergonomics tho!

Ok, but if the controller goes malicious, they can do hundreds of other more dangerous things.

That is if explorers don’t decide to store previous schemas and don’t have any compatibility checks. The explorer has the schema, so it could detect if a field named ‘credit’ dissapeared and put a red warning there.

For the block, it`s all the same __3993915879__ : 1234
If someone changes the documentation around how to decode Bitcoin blocks, it will be the same attack.

I think it’s better to offload the burden to explorers, instead of messing up ergonomics and performance.