The RAM requirement makes sense to me: canisters need to live in memory for fast execution and the “always on” model of computation (orthogonal persistence) but the SSD requirement is what I don’t understand.
The node hardware specs require ~32 TB of raw NVMe (5 × 6.4 TB in Gen-2, a similar total in the main Gen-1 configurations).
Why do IC nodes require ~32 TB of NVMe storage?
From my understanding, each node stores the full replicated state of its subnet, which is on the order of 1–2 TiB today. Even if the node keeps checkpoints of that state, plus the GuestOS, consensus artifacts and catch-up packages, I struggle to get to 32 TB.
What choice motivated this number? Which proportion of those 32TB is used is practice? In theory, could a smaller configuration work for a node, and if so how small?
Without looking into it further, my guess is that it needs it for blockchain consensus. SSD writes and retrieves data too slow to be able to be used for updating the blockchain.
Are you sure those figures aren’t only indicative of what is accessible to the Cloud Engine users? I’m asking about the specs of the nodes. I think I’ve seen somewhere that the same node could be used in several Cloud Engines at the same time (which actually could be an answer to my question)
Interesting question, I checked out some things quickly, thought the answer was the RAID, but it’s not.
RAID 0 (being used) is striping: it spreads data across all five drives for speed. It doesn’t mirror or keep a backup copy, so most of the 32tb are available.
The storage system needs additional space for:
Checkpoints: older certified versions of the state must remain available while newer ones are created.
State changes: the IC stores changed memory pages in additional files and periodically merges them. Old and new files can coexist during that process.
Recovery and synchronization: checkpoints let restarted or replacement nodes catch up safely.
Dfinity’s April 2025 storage roadmap distinguishes two steps: 1 → 2 TiB State synchronization speed and occasional hashing of the entire state needed testing. Beyond 2 TiB Retained old states and storage files could fill the physical disks under worst-case workloads. Reducing that overhead involves performance tradeoffs and protocol changes.
One reason for having a lot of SSDs with RAID0 is to get more speed, lower latency, better UX. Looks like security & availability is another one, where you model it to handle worst-case workloads. You can’t make that very optimistic, or something can make the whole subnet go down. The roadmap points out software upgrades can improve it a bit
Yes there are VM’s for CE only (mainnet has only full spec capacity nodes, not part of them) but a node is a node if they are able to run Guest OS just like any other node, so in a 4 node config you could still run somewhat limited canisters.
There was a thread that I can’t find anymore trying different ways to implement blob storage and using the excess capacity was mentioned.
In principle, the two limiting factors are how well we can even handle (hashing, checkpointing, etc) large states, and how much disk we use.
For most of the lifetime of the IC, the handling was the actual bottleneck, and a lot of of work was necessary to get to the current limit of 2TiB. The links above go into details of these things, like the introduction of log-structured merge trees (LSMT).
At this point, the disk use starts to become relevant, but mainly in a worst-case sense. For example, the files on disk for a checkpoint might be larger than the state size due to how we handle overlay files, and there might be multiple checkpoints on disk at any point, as required by the consensus algorithm. In practice however, we don’t use much more disk than the subnet size. We do want to keep enough disk space for the worst case though, because running out of disk might have very bad consequences for the subnet.
We do plan to improve these things very soon. That should help at the very least to allow larger states on small engine disks. It will also help to increase the 2TiB mainnet limit as well, but there we also need to improve the handling of large states on top of the pure disk space calculations.