We counted the public incidents at five hosting providers this month. The one with fifty is not the unreliable one
Cloudflare posted at least 50 incidents to its public status page between August 3 and August 23, 2026. Netlify posted two in the whole month. Reading that as a reliability ranking gets it backwards: Netlify's two were rated critical and major, Cloudflare's fifty included exactly one major, and the Cloudflare incident that left customer data unreachable for roughly 32 hours was labelled minor. We pulled the raw incident feeds from Cloudflare, GitHub, Vercel, Netlify and OpenAI, read every published root-cause analysis from July and August (GitHub's 7 hour 47 minute outage, Azure's five-hour fibre maintenance bug, Google Cloud's 15 hours at 44 degrees), and worked out what a small business can actually do with any of it.
Every hosting provider, CDN and platform your website depends on runs a public status page. Almost nobody reads them, and the people who do usually read them wrong.
This month gave us a good excuse to look properly. July and August 2026 were unusually bad for the infrastructure underneath the web: AWS, Azure, Google Cloud, Cloudflare and GitHub all had incidents worth reading about, several with unusually candid post-mortems. So we pulled the machine-readable incident feeds that most status pages publish, counted everything, and read every root-cause analysis we could find.
The counting produced a result that looks like an indictment and is not one. Here is the raw tally, taken from each provider’s own public incident API on the morning of August 23, 2026:
| Provider | Incidents posted in August | Highest severity used | Critical or major |
|---|---|---|---|
| Cloudflare | 50 or more, Aug 3 to Aug 23 | major | 1 |
| GitHub | 18 | critical | 5 |
| OpenAI | 17 | minor | 0 |
| Vercel | 6 | major | 2 |
| Netlify | 2 | critical | 2 |
One caveat on the first row, and it makes the point rather than undermining it. These feeds return the 50 most recent incidents. For GitHub, Vercel, Netlify and OpenAI, fifty entries reach back well past August 1, so those counts are complete. For Cloudflare, fifty entries only reach back to August 3. Fifty is a floor, not a total: Cloudflare posted at least fifty incidents in twenty-one days, and we cannot see how many more there were because the cap cut us off.
So Cloudflare posted at least twenty-five times as many incidents as Netlify. If you stop there you conclude Cloudflare is having a catastrophic month and Netlify is running beautifully. Both halves of that are wrong, and understanding why is the whole point of this post.
Incident count measures how much a provider tells you
Cloudflare’s fifty entries include things like “Network Performance Issues in Mumbai”, “Elevated Errors in Fuzhou and Foshan”, “Workers AI GLM 5.2 is unavailable” and “Cloudflare Lists Bulk Actions failing”. Cloudflare operates in hundreds of cities and posts an incident when one of them has a bad afternoon. Forty-two of the fifty were rated minor, seven were rated none, and exactly one was rated major.
Netlify posted two incidents. One, on August 3, was elevated build errors lasting about an hour. The other, on August 18, was six minutes long: between 20:02 and 20:08 UTC, visitors to Netlify-hosted sites got error pages instead of content. Netlify rated that six-minute event critical, which is the correct label for “the sites we host were down”, and wrote a short, clear summary of it.
So one provider posts a regional latency blip in Mumbai and one posts only the things that broke the product. Neither approach is dishonest. But it means the incident count is a measure of disclosure policy, not of reliability, and comparing two providers on it is close to meaningless. The generalisation worth keeping: a provider with a busy status page is telling you more, not failing more, and a provider with a quiet one has not proven anything.
Severity labels describe blast radius, not your afternoon
The second trap is worse, because it fools people who do read the status page.
The single most consequential Cloudflare incident this month was rated minor. On August 7, R2, Cloudflare’s object storage service, had an availability problem in its Eastern North America region. Here is the timeline from Cloudflare’s own updates:
- 14:52 to 17:02 UTC, August 7. Writes to a small set of R2 buckets in ENAM fail.
- 18:42 UTC. Cloudflare opens the incident.
- 23:28 UTC, August 8. Availability restored for nearly all impacted buckets. A number of objects uploaded through multipart uploads during that two-hour window on August 7 “remain unavailable.”
- 02:35 UTC, August 9. Everything restored.
For a customer whose files were in one of those buckets, that is roughly 32 hours during which data they had uploaded could not be read back. Not slow. Not degraded. Unreachable, with no way to know when or whether it was coming back. If those were customer invoices, product images or uploaded documents, that is an outage in every sense a business cares about.
It was labelled minor because the blast radius was small: a subset of buckets in one region. That is a perfectly coherent thing for a severity scale to mean, and it is not the thing most people assume it means. Severity on a status page answers “how many customers?”, not “how bad for the ones affected?” If you are one of the affected, the label is not about you.
The one incident Cloudflare rated major in three weeks is instructive in the opposite direction. On August 12, IP addresses used by Cloudflare’s Email Security service got listed by Spamhaus, and downstream mail systems that filter against Spamhaus started rejecting mail that had passed through the service. No server was down. No code was broken. Customers just stopped receiving email for a little over two hours. It got the higher label because it hit everyone using the product, and because “email silently stops arriving” is about the worst failure mode there is for a business. We have written separately about how much of email deliverability rests on reputation lists nobody controls; this is that, at provider scale.
What actually broke, in the providers’ own words
The July and August post-mortems are worth reading because several of them are unusually detailed, and because the same failure shape shows up again and again.
GitHub, August 17, 7 hours 47 minutes. From 13:28 to 21:15 UTC, GitHub.com had roughly 20% error rates on web and API traffic, and about 50% on archive and raw-content downloads. GitHub’s write-up traces it to a misconfigured autoscaling policy: an Istio service-mesh sidecar hit its concurrency limit and failed to scale because the policy watched the host service’s limits and not the sidecar’s. One failure cascaded, four HAProxy nodes exhausted their flow limits, and the authentication gateway degraded. Then it got worse on its own: a latent retry bug in VS Code amplified traffic roughly tenfold, and the Copilot Token Service went from a normal 7,000 to 9,000 requests per second to 70,000 to 100,000. Recovery required deliberately blocking token requests with a 403 and ramping traffic back up site by site.
GitHub, August 6 to 7, 9 hours. A routine deployment to an internal Actions service exposed a capacity weakness. As pods were replaced, the remaining capacity saturated and crashed, cascading across clusters. At peak, 71% of workflow runs hit infrastructure failures and 75% of the rest were delayed more than five minutes. A second phase followed, in which runners kept getting assigned invalid jobs and retrying them, which stopped them picking up real work.
Microsoft Azure, July 23, about 5 hours. Routine fibre maintenance in the West US region. A bug in the request conversion system, in Microsoft’s own words, “incorrectly marked additional devices as part of the maintenance event and caused a set of IP routes to be removed from more devices than intended.” The datacentre lost its connection to the wide-area network and 27 Azure services went offline from 14:44 to 19:41 UTC.
Google Cloud, July, about 15 hours. An electrical fault on the utility grid upstream of a datacentre in the europe-west4-a zone. The backup power system for one side of the facility failed to take the load, the chiller controller went offline during the voltage event, and temperatures inside reached 44°C. VMware Engine, NetApp Volumes and Bare Metal Solutions were affected.
AWS CloudFront, July 16, several hours. From 09:45 UTC, elevated 5xx errors for CloudFront customers using VPC Origins. AWS attributed it to an internal constraint on the fleet managing connections to private VPC origins; once that constraint was reached, the system distributing routing configuration to network processors stopped loading updated configuration correctly. Customers using other origin types were unaffected. Hugging Face, the UK National Lottery and Fallout 76 were among those that went down.
Read together, four of those five are the same story: a small, local limit was reached, and the system’s own recovery behaviour turned it into a large, global failure. A sidecar concurrency cap. A pod-replacement capacity dip. A maintenance script’s device list. A connection-fleet constraint. None of them was the interesting part. The interesting part was retries, cascades and automation amplifying a small fault into hours of downtime. The Google Cloud one is the outlier and the oldest failure mode there is: the power went out and the backup did not work.
The lesson the Google Cloud incident teaches that the others do not
There is one finding from the July Google Cloud outage that we think is the single most useful thing in this whole post, and it has nothing to do with counting incidents.
The standard advice from every cloud provider is to distribute workloads across multiple availability zones for resilience. The europe-west4-a incident showed that for certain specialised services, a “zone” can have a single-datacentre dependency that the customer has no way to see. Forrester’s Biswajeet Mahapatra put it plainly to The Register: customers “are rarely given visibility into whether a particular managed service has a single-datacentre dependency.”
In other words, you can do everything the architecture guide says, pay for redundancy, and still have a hidden single point of failure because the redundancy you bought is not redundant for the specific service you used. That is not a Google problem. It is structural to managed cloud services, and it applies to every provider.
What a small business should actually take from this
Most of what gets written about cloud outages is aimed at people with a platform team and a multi-region budget. Here is the version for a business running a website, a booking system or a small SaaS product, in rough order of value for effort.
1. Subscribe to the status pages you actually depend on. This costs five minutes and is the single highest-return item here. Every status page in the table above has an email or RSS subscription, and most support webhooks. The thing you are buying is not prevention, it is the ten minutes between “the site is broken and I have no idea why” and “the site is broken because Vercel is having an incident, and I should stop debugging.” Make a list of what a page load actually touches: DNS, CDN, hosting, database, payment processor, email sender. Subscribe to all of them.
2. Know where your data physically is, and ask whether that is one place. The R2 incident was regional. The Google Cloud incident was one datacentre inside one zone. If everything you have lives in one region, which for most small deployments it does, you should at least know that, and know it is a deliberate cost decision rather than an accident. For most small businesses, single-region is the correct choice. It only becomes a problem when nobody ever made it as a choice.
3. Have backups that are not at your provider. This is the item people skip, and the R2 incident is the argument for it. Data was not lost; it was unreachable, for a day and a half, with an unhelpful “minor” label attached. A copy of your database and your uploaded files somewhere else, taken automatically, is the difference between an inconvenient day and an existential one. It also protects you against the more common disasters, which are a deleted table and a lapsed credit card, not a datacentre fire.
4. Do not architect around the outage you just read about. Every one of the incidents above would have been survived by a different, more expensive architecture, and each of those architectures has its own failure modes. Multi-region costs real money and adds complexity that itself causes downtime. The honest maths for a business doing, say, $2,000 of revenue a day: a five-hour outage costs a few hundred dollars of deferred sales. Multi-region hosting costs more than that every month, permanently, and buys you nothing on the far more likely failure, which is a bug in your own code. Spend the money on backups and monitoring first.
5. Treat SLA credits as fiction. A typical 99.9% SLA permits about 43 minutes of downtime a month, and the remedy for breaching it is a credit against your bill, usually a percentage, usually capped, usually requiring you to file a claim within a fixed window. On a $20 a month plan, a full month’s credit is $20. If downtime genuinely costs your business meaningful money, the SLA is not the instrument that protects you; a tested backup and a plan for what staff do while the site is down is.
6. Write the two-paragraph plan. Not a disaster recovery binder. Two paragraphs: who notices, who they call, what you tell customers, and what the business does while it is down. If you take payments online, the important line is what happens to orders placed during the outage. Most small businesses discover they need this at 4pm on a Friday.
The part nobody wants to say
Cloud infrastructure spending hit $143 billion in Q2 2026, up 43% year on year, the fastest growth in eight years, and AWS, Azure and Google Cloud together now take 67% of it, up from 63% three quarters earlier, according to Synergy Research. Concentration is increasing, not decreasing, and it is increasing because AI workloads are pulling everyone onto the same handful of platforms.
The practical consequence is that the reliability of your small business website is now substantially outside your control, and will be more so next year. That is not an argument for running your own servers, which almost always goes worse. It is an argument for being clear-eyed about what you are buying: you are renting somebody else’s operational excellence, which is genuinely better than yours would be, and accepting that when it fails you have no lever to pull except waiting.
Given that, the useful work is not preventing the outage. It is shortening the time between the outage starting and you knowing what it is, and making sure your data exists somewhere the outage cannot reach. Both of those are cheap. Everything else on the list is expensive and mostly not worth it.
If you want a second opinion on what your site actually depends on, send us the URL. Mapping it out takes about twenty minutes and usually turns up one dependency the owner did not know was there.
Sources
Incident counts and timelines were taken from each provider’s public incident API on August 23, 2026:
- Cloudflare System Status, incident feed and the August 7 to 9 R2 and August 12 Email Security incident updates
- GitHub Status, including the published root-cause analyses for the August 6 Actions incident and the August 17 GitHub.com incident
- Vercel Status and Netlify Status, including Netlify’s August 18 summary
- OpenAI Status
Post-mortems and reporting:
- Microsoft fiber foul-up cut off Azure California for almost five hours, The Register, July 24, 2026
- Google Cloud outage shows it’s still hard to understand hyperscalers’ real resilience regimes, The Register, July 21, 2026
- AWS CloudFront outage serves errors instead of websites, The Register, July 16, 2026
- Enterprise cloud infrastructure uptake shows no sign of slowing, The Register, August 1, 2026, reporting Synergy Research figures for Q2 2026