TingleFlow - Notice history

Notice history

Aug 2026

No notices reported this month

Jul 2026

Issues with accessing TingleFlow
  • Resolved
    UTC
    Resolved

    All users now have access to TingleFlow, and our systems remain stable.

    We sincerely apologize for the length of this outage. It was a highly unusual incident due to both its complexity and the nature of the underlying issue. Over the coming days, we will conduct a thorough review to understand exactly what went wrong, identify areas for improvement, and help prevent similar incidents from happening again.

  • Update
    UTC
    Update

    We have begun allowing traffic back into TingleFlow incrementally.

    Some users will regain access before others as we carefully monitor system stability and performance.

    Users may regain access in the following priority order:

    1. TingleFlow+ Subscribers (Creators)

    2. TingleFlow+ Subscribers (Viewers)

    3. Creators (users with two or more uploaded videos)

    4. Viewers

    To further reduce system load, access will also be rolled out gradually by geographic region. This means users within the same category may regain access at different times.

    Regional rollout order:

    1. North America

    2. South America

    3. Africa

    4. Europe

    5. Asia

    6. Oceania

    This phased rollout helps us closely monitor system performance and ensure the platform remains stable as users regain access. We appreciate everyone's patience while we complete the recovery.

  • Update
    UTC
    Update

    It appears that the fix has been successful, and our systems are returning to normal. We expect to begin allowing users back onto TingleFlow incrementally within the next few hours. To help ensure a stable recovery and avoid overloading our infrastructure, we will gradually restore access using DNS routing. We will share more information on which users will gain access first when we're ready to start letting users access TingleFlow again.

  • Monitoring
    UTC
    Monitoring

    We believe that we have deployed an appropriate fix. We will ensure it is working as expected, then we will be able to let traffic in incrementally if all goes well.

  • Update
    UTC
    Update

    We have identified the exact root cause of the issue and have began working on a solution. Our team is actively working to restore all services, and we'll continue to provide updates throughout the day.

  • Update
    UTC
    Update

    We are continuing to work on a fix for this incident.

  • Identified
    UTC
    Identified

    We believe we have identified the internal issue that contributed to the TingleFlow outage. We are currently carrying out the required fixes, testing, and maintenance work to safely restore services. Thank you for your continued patience and understanding while we work to bring TingleFlow back online.

  • Update
    UTC
    Update

    TingleFlow has released a statement regarding this outage. You can read it here: https://blog.tingleflow.com/tingleflow-addresses-unplanned-service-outage/

  • Update
    UTC
    Update

    We believe we have narrowed down the issue and are very close to identifying it.

  • Update
    UTC
    Update

    We are continuing to investigate the root cause of this outage. We will provide another update when available. We believe we are close to identifying the root cause.

  • Update
    UTC
    Update

    We are continuing to investigate the root cause of this outage. We will provide another update when available.

  • Update
    UTC
    Update

    The engineering team has dug deeper into debug logs and operating system-level metrics to better understand the issue. This data showed that Consul KV writes were getting blocked for long periods of time, resulting in what is known as “contention.”

    The cause of the contention is not currently obvious, though one theory is that the shift from 64 to 128 CPU Core servers early in the outage may have made the problem worse. After reviewing the HTOP data and performance debugging data shown in the screenshots below, we have concluded that it is ideal to return to 64 Core servers, similar to those used before the outage.

    We have begun preparing the hardware for this transition. Consul is installed, operating system configurations have been triple checked, and the machines have been readied for service in as detailed a manner as possible. We will then transition the Consul cluster back to 64 CPU Core servers.

    This will be our fourth attempt at diagnosing the root cause of the incident.

  • Update
    UTC
    Update

    We are continuing to look for the root cause of this issue. We deeply apologise for the inconvenience caused during this incident.

  • Update
    UTC
    Update

    At this point, it is now clear to us that overall Consul usage is not the only contributing factor to the performance degradation. Given this realization, we will again pivot. Instead of looking at Consul from the perspective of the TingleFlow services that depend on it, we will start looking at Consul internals for clues.

  • Update
    UTC
    Update

    We are currently looking into whether hardware failure may be a part of the issue. However, faster hardware hasn't helped and, as we now know, likely hurt stability. Resetting Consul’s internal state hasn't helped either. There is no user traffic coming in, yet Consul is still slow. We have again leveraged iptables to let traffic back into the cluster slowly. We are looking into whether the cluster simply getting pushed back into an unhealthy state by the sheer volume of thousands of containers trying to reconnect has caused this. This was our third attempt at diagnosing the root cause of the incident, no cause has yet been identified.

    We have ruled out:

    • DDoS/Cyber attack

    • Frontend issues

    • Authentication issues

    Our next step is to reduce Consul usage and then systematically reintroduce it.

    TingleFlow also currently reports 0 users online, for the first time.

  • Update
    UTC
    Update

    The reset went smoothly, and initially, the metrics looked good. When we removed the iptables block, the service discovery and health check load from the internal services returned as expected. However, Consul performance began to degrade again, and now we are back to where we started: 50th percentile on KV write operations was back at 4 seconds. Services and systems that depended on Consul are now starting to mark themselves “unhealthy.” The system has fell back into the problematic state. There is clearly something about our load on Consul that is causing problems, and over 12 hours into the incident, we still are unsure as to what it is.

  • Update
    UTC
    Update

    We expected that restoring from a snapshot taken when the system was healthy would bring the cluster into a healthy state, but we had one additional concern. Even though TingleFlow did not have any user-generated traffic flowing through the system at this point, internal TingleFlow services and systems were still live and reaching out to Consul to learn the location of their dependencies and to update their health information. These reads and writes were generating a significant load on the cluster. We were worried that this load might immediately push the cluster back into an unhealthy state even if the cluster reset was successful. To address this concern, we configured iptables on the cluster to block access. This would allow us to bring the cluster back up in a controlled way and help us understand if the load we were putting on Consul independent of user traffic was part of the problem.

  • Update
    UTC
    Update

    Based on the severity of this incident, we have changed the status to Critical. We cannot yet provide an ETA but we encourage users to stay informed here.

  • Investigating
    UTC
    Investigating

    We are currently investigating this incident. We will provide updates when they become available.

Jun 2026 to Aug 2026

Next