Notice history
Aug 2026
No notices reported this month
Jul 2026
- ResolvedUTCResolvedUTC
All users now have access to TingleFlow, and our systems remain stable.
We sincerely apologize for the length of this outage. It was a highly unusual incident due to both its complexity and the nature of the underlying issue. Over the coming days, we will conduct a thorough review to understand exactly what went wrong, identify areas for improvement, and help prevent similar incidents from happening again.
- UpdateUTCUpdateUTC
We have begun allowing traffic back into TingleFlow incrementally.
Some users will regain access before others as we carefully monitor system stability and performance.
Users may regain access in the following priority order:
TingleFlow+ Subscribers (Creators)
TingleFlow+ Subscribers (Viewers)
Creators (users with two or more uploaded videos)
Viewers
To further reduce system load, access will also be rolled out gradually by geographic region. This means users within the same category may regain access at different times.
Regional rollout order:
North America
South America
Africa
Europe
Asia
Oceania
This phased rollout helps us closely monitor system performance and ensure the platform remains stable as users regain access. We appreciate everyone's patience while we complete the recovery.
- UpdateUTCUpdateUTC
It appears that the fix has been successful, and our systems are returning to normal. We expect to begin allowing users back onto TingleFlow incrementally within the next few hours. To help ensure a stable recovery and avoid overloading our infrastructure, we will gradually restore access using DNS routing. We will share more information on which users will gain access first when we're ready to start letting users access TingleFlow again.
- MonitoringUTCMonitoringUTC
We believe that we have deployed an appropriate fix. We will ensure it is working as expected, then we will be able to let traffic in incrementally if all goes well.
- UpdateUTCUpdateUTC
We have identified the exact root cause of the issue and have began working on a solution. Our team is actively working to restore all services, and we'll continue to provide updates throughout the day.
- UpdateUTCUpdateUTC
We are continuing to work on a fix for this incident.
- IdentifiedUTCIdentifiedUTC
We believe we have identified the internal issue that contributed to the TingleFlow outage. We are currently carrying out the required fixes, testing, and maintenance work to safely restore services. Thank you for your continued patience and understanding while we work to bring TingleFlow back online.
- UpdateUTCUpdateUTC
TingleFlow has released a statement regarding this outage. You can read it here: https://blog.tingleflow.com/tingleflow-addresses-unplanned-service-outage/
- UpdateUTCUpdateUTC
We believe we have narrowed down the issue and are very close to identifying it.
- UpdateUTCUpdateUTC
We are continuing to investigate the root cause of this outage. We will provide another update when available. We believe we are close to identifying the root cause.
- UpdateUTCUpdateUTC
We are continuing to investigate the root cause of this outage. We will provide another update when available.
- UpdateUTCUpdateUTC
The engineering team has dug deeper into debug logs and operating system-level metrics to better understand the issue. This data showed that Consul KV writes were getting blocked for long periods of time, resulting in what is known as “contention.”
The cause of the contention is not currently obvious, though one theory is that the shift from 64 to 128 CPU Core servers early in the outage may have made the problem worse. After reviewing the HTOP data and performance debugging data shown in the screenshots below, we have concluded that it is ideal to return to 64 Core servers, similar to those used before the outage.
We have begun preparing the hardware for this transition. Consul is installed, operating system configurations have been triple checked, and the machines have been readied for service in as detailed a manner as possible. We will then transition the Consul cluster back to 64 CPU Core servers.
This will be our fourth attempt at diagnosing the root cause of the incident.
- UpdateUTCUpdateUTC
We are continuing to look for the root cause of this issue. We deeply apologise for the inconvenience caused during this incident.
- UpdateUTCUpdateUTC
At this point, it is now clear to us that overall Consul usage is not the only contributing factor to the performance degradation. Given this realization, we will again pivot. Instead of looking at Consul from the perspective of the TingleFlow services that depend on it, we will start looking at Consul internals for clues.
- UpdateUTCUpdateUTC
We are currently looking into whether hardware failure may be a part of the issue. However, faster hardware hasn't helped and, as we now know, likely hurt stability. Resetting Consul’s internal state hasn't helped either. There is no user traffic coming in, yet Consul is still slow. We have again leveraged iptables to let traffic back into the cluster slowly. We are looking into whether the cluster simply getting pushed back into an unhealthy state by the sheer volume of thousands of containers trying to reconnect has caused this. This was our third attempt at diagnosing the root cause of the incident, no cause has yet been identified.
We have ruled out:
DDoS/Cyber attack
Frontend issues
Authentication issues
Our next step is to reduce Consul usage and then systematically reintroduce it.
TingleFlow also currently reports 0 users online, for the first time.
- UpdateUTCUpdateUTC
The reset went smoothly, and initially, the metrics looked good. When we removed the iptables block, the service discovery and health check load from the internal services returned as expected. However, Consul performance began to degrade again, and now we are back to where we started: 50th percentile on KV write operations was back at 4 seconds. Services and systems that depended on Consul are now starting to mark themselves “unhealthy.” The system has fell back into the problematic state. There is clearly something about our load on Consul that is causing problems, and over 12 hours into the incident, we still are unsure as to what it is.
- UpdateUTCUpdateUTC
We expected that restoring from a snapshot taken when the system was healthy would bring the cluster into a healthy state, but we had one additional concern. Even though TingleFlow did not have any user-generated traffic flowing through the system at this point, internal TingleFlow services and systems were still live and reaching out to Consul to learn the location of their dependencies and to update their health information. These reads and writes were generating a significant load on the cluster. We were worried that this load might immediately push the cluster back into an unhealthy state even if the cluster reset was successful. To address this concern, we configured iptables on the cluster to block access. This would allow us to bring the cluster back up in a controlled way and help us understand if the load we were putting on Consul independent of user traffic was part of the problem.
- UpdateUTCUpdateUTC
Based on the severity of this incident, we have changed the status to Critical. We cannot yet provide an ETA but we encourage users to stay informed here.
- InvestigatingUTCInvestigatingUTC
We are currently investigating this incident. We will provide updates when they become available.