# 6/8/22’s Multi-Continent IC Outage - Are Boundary Nodes the top attack vector for the IC?

**URL:** <https://forum.dfinity.org/t/6-8-22-s-multi-continent-ic-outage-are-boundary-nodes-the-top-attack-vector-for-the-ic/13624>\
**Category:** Developers\
**Tags:** Boundary-nodes, Vulnerability\
**Created:** [June 8, 2022, 11:23pm UTC](https://forum.dfinity.org/t/6-8-22-s-multi-continent-ic-outage-are-boundary-nodes-the-top-attack-vector-for-the-ic/13624 "2022-06-08T23:23:57Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![justmythoughts](https://sea1.discourse-cdn.com/flex023/user_avatar/forum.dfinity.org/justmythoughts/32/5226_2.png) [@justmythoughts](https://forum.dfinity.org/u/justmythoughts)\
**Post date:** [June 8, 2022, 11:23pm UTC](https://forum.dfinity.org/t/6-8-22-s-multi-continent-ic-outage-are-boundary-nodes-the-top-attack-vector-for-the-ic/13624/1 "2022-06-08T23:23:57Z")

</div>

Today’s [multi-continent outage](https://forum.dfinity.org/t/is-the-network-down-for-anyone-else/13613) on the IC has inspired me ask more questions about the resiliency of the IC’s Boundary Nodes.

If the failure of a **single** Boundary Node, that was not under attack of any kind can disrupt users in half of the US, Latin America, and Western Europe (and that is just what was reported), then…

I would argue that reinforcing the boundary node infrastructure is a P1 issue (all hands on deck), and more important than any other work happening on the IC at this very moment (yes, even Bitcoin Integration and the SNS).

> [@Is the network down for anyone else?](https://forum.dfinity.org/t/is-the-network-down-for-anyone-else/13613/20):
>
> The IC itself was actually performing as expected and was not down, however one of the Boundary Nodes (in US-east) which acts as a gateway for the IC experienced intermittent issues that prevented failover. As such users in that geo-area who were DNS routed to that impacted boundary node experienced their requests being dropped by the BN. Advanced users in the impacted area were technically able to reach the IC by talking to other boundary nodes. The method of targeting specific boundary nodes …

> [@Is the network down for anyone else?](https://forum.dfinity.org/t/is-the-network-down-for-anyone-else/13613/10):
>
> We identified and rolled out a fix, that affected access from certain geographies. namely, Latam Western Europe Southeast US we checked access from the above locations and confirmed the issue and its fix. @nicopoggi@oss @Chloros88

These statements don’t seem to be consistent - or if they are, I’m a bit disappointed in the resilience of the Boundary Node infrastructure.

Here’s a few questions that come to mind:

- How do all IC requests coming from users in Minneapolis to Latin America to Western Europe go down because a single Boundary Node goes down in US-east?

- What are the regions that boundary nodes are running in?

- How many Boundary Nodes are being run in total? 1-2 in each continent? [https://dashboard.internetcomputer.org](https://dashboard.internetcomputer.org) shows only 1 boundary node in all of North America.  

- Why should developers feel like their applications and the IC is “infinitely scalable” and “more resilient than AWS” if there is only one boundary node in the US?

- Are there plans to scale out boundary nodes (centralized through DFINITY or through external parties)?

- Where are the Boundary Nodes being run? What is the infrastructure that they are running on? AWS? GCP? Independent data centers? DFINITY owned hardware?

Tagging DFINITY team members that have spoken before regarding Boundary Nodes for visibility:  
@Daniel-Bloom @faraz.shaikh @Jan @rrkapitz @yotam

---

<div class="post-metadata">

**Author:** ![martin\_DFN1](https://sea1.discourse-cdn.com/flex023/user_avatar/forum.dfinity.org/martin_dfn1/32/4617_2.png) [@martin\_DFN1](https://forum.dfinity.org/u/martin_DFN1)\
**Post date:** [June 9, 2022, 3:45am UTC](https://forum.dfinity.org/t/6-8-22-s-multi-continent-ic-outage-are-boundary-nodes-the-top-attack-vector-for-the-ic/13624/2 "2022-06-09T03:45:47Z")

</div>

Your eyes do not deceive you. This should not happen. We agree. We can describe the state of the world, but you should know that we have been diving deep into BNs and their future. We believe we have a good plan (many engineering managers and directors and even Jan and Dom studied the issue together for hours in Zürich last week), but we are working out the kinks before publishing it. Boundary nodes are a work in progress and we welcome feedback like this even while we are working on improving things.

TL;DR: Operations and availability of a single boundary node in no way dictate the availability or operations of the IC. Boundary infra-structure is stateless and fault-tolerant.

Today’s incident was the fallout of an issue with the boundary node fail-over mechanism.  
Two recently added boundary nodes refused to failover, after a misconfigured deployment.  
We understand why failover wasn’t triggered and have our work cut out.

Some quick answers to your question. Please expect a detailed postmortem once we have it in a form ready for public consumption:

Q How do all IC requests coming from users in Minneapolis to Latin America to Western Europe go down because a single Boundary Node (BN) goes down in US-east?  
A. Every region is backed by a pool (3+) of boundary nodes. Each pool has a backup pool of last resort, and eventually, there is a boundary node of last resort. You are right in questioning why requests went to a single node. In today’s incident, flaws in heartbeat detection logic prevented us from failing over to the next BN. The fix deployed was to trigger a manual failover by taking the offending boundary node offline.

Q What are the regions that boundary nodes are running in?  
A. Region is an “optional” abstraction, i.e. division into region doesn’t divide the availability and fault tolerance of boundary nodes. i.e. you can access _any_ boundary node at any time from anywhere. DNS resolution of the ic0.app can actually be controlled by the end-user to point to the desired boundary node. The dashboard should give you a list of boundary nodes irrespective of their parent regions; if you set the name resolution of ic0.app to the IP of one of the boundary nodes you will always go to the same BN.

Q How many Boundary Nodes are being run in total? 1-2 in each continent? shows only 1 boundary node in all of North America.  
A. Please see the answer above on regions. We have 20+ boundary nodes (20 are active and more are on standby). There is ongoing work to reduce the onboarding overhead for new boundary nodes such that community members can run boundary nodes.

Q Why should developers feel like their applications and the IC is “infinitely scalable” and “more resilient than AWS” if there is only one boundary node in the US?  
There are many boundary nodes – please see the answer above.  
Dfinity is invested and committed to the reliability, availability and scalability of the boundary nodes infrastructure.  
The final goal is community-owned infrastructure.

Q Are there plans to scale out boundary nodes (centralized through DFINITY or through external parties)?  
A See answer above

Q Where are the Boundary Nodes being run? What is the infrastructure that they are running on? AWS? GCP? Independent data centers? Dfinity-owned hardware?  
A Independent data center + Dfinity owned data-centers and hardware

Cheers,  
The Boundary Nodes team

---

<div class="post-metadata">

**Author:** ![justmythoughts](https://sea1.discourse-cdn.com/flex023/user_avatar/forum.dfinity.org/justmythoughts/32/5226_2.png) [@justmythoughts](https://forum.dfinity.org/u/justmythoughts)\
**Post date:** [June 9, 2022, 5:45am UTC](https://forum.dfinity.org/t/6-8-22-s-multi-continent-ic-outage-are-boundary-nodes-the-top-attack-vector-for-the-ic/13624/3 "2022-06-09T05:45:52Z")

</div>

@martin_DFN1

Thank you for the prompt response.

Looking forward to reading the postmortem - glad to hear yesterday’s issue was an implementation bug and not a design flaw.

Also, very interested to hear more about the future of boundary nodes on the IC. Many developers are wondering how they can scale up and balance load, as well as throttle and rate-limit principals to prevent DDOS/cycle drain attacks.

I believe the boundary nodes have an incredibly important part to play in the future of the IC, and I look forward to DFINITY revealing more about what that part will look like.

---

<div class="post-metadata">

**Author:** ![diegop](https://sea1.discourse-cdn.com/flex023/user_avatar/forum.dfinity.org/diegop/32/569_2.png) [@diegop](https://forum.dfinity.org/u/diegop)\
**Post date:** [June 9, 2022, 6:27am UTC](https://forum.dfinity.org/t/6-8-22-s-multi-continent-ic-outage-are-boundary-nodes-the-top-attack-vector-for-the-ic/13624/4 "2022-06-09T06:27:35Z")

</div>

For visibility @martin_DFN1 is a senior engineering manager at Dfinity and leads boundary node team.

---

<div class="post-metadata">

**Author:** ![justmythoughts](https://sea1.discourse-cdn.com/flex023/user_avatar/forum.dfinity.org/justmythoughts/32/5226_2.png) [@justmythoughts](https://forum.dfinity.org/u/justmythoughts)\
**Post date:** [June 9, 2022, 6:38pm UTC](https://forum.dfinity.org/t/6-8-22-s-multi-continent-ic-outage-are-boundary-nodes-the-top-attack-vector-for-the-ic/13624/5 "2022-06-09T18:38:58Z")

</div>

@martin_DFN1

> [@martin\_DFN1](#):
>
> We have 20+ boundary nodes (20 are active and more are on standby).

Why does [https://dashboard.internetcomputer.org/](https://dashboard.internetcomputer.org/) (as of the timestamp of this post) then show a count of 13 boundary nodes?

 ![image](https://us1.discourse-cdn.com/flex023/uploads/dfn/original/2X/f/f5b4ab1f90bc6b90be15cf494c930725a8806a66.jpeg)

Is the public facing dashboard out of date?

I assumed all of the numbers coming from the Internet Computer Dashboard were powered by a live internal API and not hardcoded.

---

<div class="post-metadata">

**Author:** ![martin\_DFN1](https://sea1.discourse-cdn.com/flex023/user_avatar/forum.dfinity.org/martin_dfn1/32/4617_2.png) [@martin\_DFN1](https://forum.dfinity.org/u/martin_DFN1)\
**Post date:** [June 9, 2022, 7:11pm UTC](https://forum.dfinity.org/t/6-8-22-s-multi-continent-ic-outage-are-boundary-nodes-the-top-attack-vector-for-the-ic/13624/6 "2022-06-09T19:11:58Z")

</div>

There is a visionary plan under Proposal [35671](https://dashboard.internetcomputer.org/proposal/35671)

We are working on a concrete roadmap, that is too early to share at the moment. We have to balance the powers that node providers have with the security interests of the IC. The crypto people are working on how this can be achieved; and we have to work with all the other teams’ schedules. It’s like threading a needle while changing the engines on a flying airplane.

---

<div class="post-metadata">

**Author:** ![martin\_DFN1](https://sea1.discourse-cdn.com/flex023/user_avatar/forum.dfinity.org/martin_dfn1/32/4617_2.png) [@martin\_DFN1](https://forum.dfinity.org/u/martin_DFN1)\
**Post date:** [June 9, 2022, 7:13pm UTC](https://forum.dfinity.org/t/6-8-22-s-multi-continent-ic-outage-are-boundary-nodes-the-top-attack-vector-for-the-ic/13624/7 "2022-06-09T19:13:45Z")

</div>

Ooh, good question. There is a lot of deployment going on at the moment and while we have many machines sitting around it wasn’t quite clear at the time of writing how many are actually processing requests. Once we have the new BN-VMs fully tested there will be more appearing as quickly as finance allows and we have Ops capacity.

---

<div class="post-metadata">

**Author:** ![jzxchiang](https://sea1.discourse-cdn.com/flex023/user_avatar/forum.dfinity.org/jzxchiang/32/2592_2.png) [@jzxchiang](https://forum.dfinity.org/u/jzxchiang)\
**Post date:** [June 9, 2022, 9:51pm UTC](https://forum.dfinity.org/t/6-8-22-s-multi-continent-ic-outage-are-boundary-nodes-the-top-attack-vector-for-the-ic/13624/8 "2022-06-09T21:51:25Z")

</div>

Thanks for raising this issue. One boundary node serving all of North America… I live in NA and it’s impressive my requests to mainnet get resolved as quickly as they do, given that a single node is handling all of that.

Once we add more BNs, I’m hoping we can also increase the per-subnet, per-BN [rate limits](https://forum.dfinity.org/t/how-would-internet-identity-handle-a-denial-of-service-attack/12791/7) we apply right now.

Looking forward to more updates on this front, including decentralization of BNs.

---

<div class="post-metadata">

**Author:** ![martin\_DFN1](https://sea1.discourse-cdn.com/flex023/user_avatar/forum.dfinity.org/martin_dfn1/32/4617_2.png) [@martin\_DFN1](https://forum.dfinity.org/u/martin_DFN1)\
**Post date:** [June 14, 2022, 8:05pm UTC](https://forum.dfinity.org/t/6-8-22-s-multi-continent-ic-outage-are-boundary-nodes-the-top-attack-vector-for-the-ic/13624/9 "2022-06-14T20:05:18Z")

</div>

It’s actually more like three nodes serving the NA US. And if those fail or slow down the traffic is directed to other nodes. The problem this post is about, occurred because one of the BNs started to fail in a component that affects processing but not the heartbeat that the failover is based upon. A heartbeat should quickly execute the main components and deliver a healthy/sick verdict. So in this case the failover mechanism kept routing to the failed node We’ll improve the heartbeat functionality.

---

<div class="post-metadata">

**Author:** ![diegop](https://sea1.discourse-cdn.com/flex023/user_avatar/forum.dfinity.org/diegop/32/569_2.png) [@diegop](https://forum.dfinity.org/u/diegop)\
**Post date:** [June 17, 2022, 12:45am UTC](https://forum.dfinity.org/t/6-8-22-s-multi-continent-ic-outage-are-boundary-nodes-the-top-attack-vector-for-the-ic/13624/10 "2022-06-17T00:45:55Z")

</div>

**Update:**

@martin_DFN1 has posted an incident retrospective here:

> [@North Atlantic Region boundary node outage Incident Retrospective - Wednesday Jun 8, 2022](https://forum.dfinity.org/t/north-atlantic-region-boundary-node-outage-incident-retrospective-wednesday-jun-8-2022/13856):
>
> Summary On June 8, 2022 two users on the developer forum reported being unable to access the IC ( “Internal Server Error. Failed to fetch response: TypeError: Failed to fetch”). DFINITY engineers looked at the logs of the boundary node (BN) machines and found that the Marseille BN was down, and on other machines the filesystems containing the logs were close to overflowing. DFINITY Engineers reproduced the errors by filling the file system. The overflow was caused by a Journalbeat bug [JB-2362…](https://github.com/elastic/beats/issues/23627)
