181 comments

udave
Hope it survive the next few years of codification of all of humanities unfinished cheap ideas.
sdcfgy
This is horribly accurate.
I love that phrase.
Why? Letting it burn to the ground will bring about federated git forges faster.
zombot
As if Microslop didn't have enough unfinished cheap ideas already.
Still getting Internal Server Error despite githubstatus claiming “All Systems Operational” and the incident as “Resolved”.

Premature declaration of Resolution is pretty bad user experience, if I’m being honest…

Who needs to wait for the 5xxs return to normal when you can just press end incident and pad your stats!
What's to pad there? Their one nine?
zombot
It can only go up from there.
Could be an honest mistake.
Possible, but highly unlikely in GitHub’s case.
At this point, a mistake like this is indistinguishable from malice.
The original incident was probably an honest mistake as well.
Spent 5 minutes debugging why my builds were failing... then recalled github is not google when it comes to availability... checked status.github.com ... and voila.. there it is.
I'm surprised people still rely so much on GitHub. Anything I want to (have to*) be able to deploy on a moments notice, for hotfixes and what not, was moved away from Github like a year ago, because of the horrible uptime. Did others not get the memo yet?

It's not getting better, if anything it's getting worse. Best time to get off GitHub was yesterday, the second-best day to get off is today. Tangled, Codeberg or self-hosted Forgejo (my approach) or Gitea are all good alternative solutions here.

>> Anything I want to (have to*) be able to deploy on a moments notice, for hotfixes and what not, was moved away from Github like a year ago, because of the horrible uptime. Did others not get the memo yet?

This misses the point in a bad way. You shouldn't have a single deploy funnel to begin with. You should be able to do deploys from multiple places.

Sure, and backups should be replicated over 3 mediums, and you shouldn't deploy bugs into production, and...

Lots of things we should do, only a few we actually have time to fix. If I were choosing between "Migrating away from GitHub" and "Adding another way to deploy in case GitHub is down", I'd take advantage and do the first, because you'll end up having to replicate SCM, build infrastructure and so on anyway, why not do it properly?

antonvs
> This misses the point in a bad way.

This misses the business reality in a bad way. A typical outage with something like Github is not a big deal for most businesses. Sure, some nerds get annoyed, but that's about it. The extra effort in maintaining multiple deploy approaches for a non-trivial system simply doesn't make business sense in most cases.

This is true.

It's also something that's easy to do if you start from the beginning treating your CI/CD system as something that just runs scripts out of your repo.

antonvs
Have you ever worked at a company of any significant size?
viccis
It's free and you can get it up and running pretty much instantaneously with a coding agent. That's it.

They really should be partitioning their infrastructure so that their paying customers are not in the blast radius of outages from the legions of repos running their 700 test vibecoded regression suite every time they change a config file.

I’m sure they’re working on that, but that takes time. It definitely doesn’t happen in the relative “overnight” that this new influx of code spam spun up.
solatic
It's crazy but there's really still so much out there where hosting on GitHub is practically a requirement, with no other solutions on the market. Off the top of my head, stuff I encountered in the last month:

1. Homebrew really, really wants you to host your taps on GitHub. The docs are very GitHub-focused: https://docs.brew.sh/How-to-Create-and-Maintain-a-Tap as are downstream projects like dist: https://axodotdev.github.io/cargo-dist/book/installers/homeb...

2. Terraform and OpenTofu public registries support only GitHub, see e.g. https://developer.hashicorp.com/terraform/registry/modules/p... and https://opentofu.org/docs/language/modules/develop/publish/#...

3. Only GitHub allows you to use free minutes on managed Windows and macOS runners, which are essential for building and signing native Windows and macOS/iOS software. GitLab's are in beta and require a paid plan, Buildkite requires you to pay at least $30/user without any bundled macOS minutes and no hosted Windows runners available. And for small shops, keeping the macOS and Windows runners clean is such a stupid time sink, with the small number of minutes involved it's much more preferable to just start with free managed minutes and plan to pay for managed runner minutes later.

A recent one for me is Anthropic's first party cloud stuff all assumes you're cloning from GitHub.
In those cases, it's more often than not enough to just have a read-only mirror on GitHub, that gets automatically replicated when your true upstream gets updated. Might mean a bit more machinery, but I'll take that over letting some flaky Microsoft service stop me from delivering what I promise my users.
antonvs
> Did others not get the memo yet?

Perhaps we're not so dysfunctional that we regularly need to "deploy on a moment's notice." That's a business smell.

antonvs
We had a similar experience recently with Google: a few team members were trying to diagnose an issue with our systems on Sep 1, for quite a bit longer than 5 minutes. Turns out it was a major Gcloud outage for about 4 hours: https://status.cloud.google.com/incidents/J5ia5t9p3g9Q5Wi7r8...

Google may be more reliable overall, but Gcloud at least tends to be more critical to a business that uses it, so it kind of balances out. I'd much rather have a Github outage than a Gcloud one.

Does 89.99% count as "three nines availability"?
Of course not, guess you meant that as a satire

> https://www.githubstatus.com/ Their status page records the outage anywhere between 6 to 17 minutes across the different features. The detailed status page clearly has recorded time of when the errors were escalated, when the fix was applied and when it was assumed that the system has recovered. Its clearly a lot more than 17 minutes, infact its more than an hour.

Unless there is something more to it, these status pages feel like a blatant lie, wondering how long before some one actually sues them cause they are publishing/advertising wrong SLAs.

Another example

> https://www.githubstatus.com/incidents/0rn90wk115q9

> On September 13, 2026, between 08:43 and 10:44 UTC

The status page records the incident outage duration as 1 hour 12 minutes

So, "git push" is down ... how are github engineers gonna push the fix now?
sudo git push
Make no mistakes.
in vim over ssh of course.. agents can use vim right?
Yeah, but they keep trying to use Emacs commands...

C-x M-c M-unfuck-git

You'd use tramp... What is way better than the vim's solution.
zrail
Real talk I have a M-x unfuck-this-buffer. All it does is toggle the Unicode input mode that I sometimes accidentally trigger but have never figured out how.
edoceo
That can, but cannot exit
rsync and sftp, like god intended
Like the olden days
Nah in the olden days we'd load a tape with patches and the code would read a patch off tape, apply it, then read the next one. lol.
I prefer Dropbox.
Might I ask why you aren't working directly on the prod server? God doesn't use staging or dev, he just codes directly on prod.
The last decade makes me think he needs to modernize his workflow.
I use dev. It’s a cname to prod.
Are you sure? What if we're living in the staging universe?
14u2c
Prodliness is next to Godliness.
sdcfgy
God uses UUCP
xydone
git push --force. No time for the lease either
s3/blob storage
Easy. GitHub is hosted on Sourceforge.
rapind
Based on the comments in this thread, you'd think that Reddit was down.
Complaining that HN is becoming Reddit is imo even worse than posting Reddit-level comments - those are at least sometimes funny
iambenm
"Please don't post comments saying that HN is turning into Reddit." [0]

[0] https://news.ycombinator.com/newsguidelines.html

This isn't that imo... do you think it violated the spirit of that rule?
What do you think parent is trying to say, except that this HN comment section looks like reddit? I'm trying my hardest to apply charitable reading, but even this has limits.
My understanding as well, as if Reddit was down and redditors took to HN to distill their anger. It's not calling a specific comment like this, it's calling most of comments here like this.
Perhaps about the quality of the comments. In my opinion it's not really that bad but reddit-style comments are often cheap shots that are slightly off topic, some of which I see present here
There are many differences IMO. Moderators, for instance, censor a ton on reddit; but the UI is also different. I much preferred old.reddit.com over both the new reddit (which is utter trash) and hackernews (which is simple, but also ... strange and awkward, much more confusing than old.reddit.com - and I hate the fact that after like 5 comments, I get locked out for hours before I can make more comments, that is the worst decision hackernews ever made).
But that's exactly what the HN guidelines are trying to prevent/discourage -- meta discussion about Reddit-quality comments.

It doesn't add to on-topic discussion.

Rapzid
That it looks like reddit is down so a bunch of people used to posting low effort slop over there ended up here posting low effort slop?

This thread isn't "HN" and that comment wasn't saying it was turning into Reddit..

But let's call a spade a spade. I got worried my account would be flagged due to the amount of down voting I was doing.

I don't think length intrinsically denotes quality. One could say that longer sentences show more quality, but I would not even be certain of that either.

The biggest difference I noticed has been between smartphone/tablet users on the one hand, and oldschool desktop computer system users. I belong to the latter group and I think we, as a group, write more, and faster. I'd also like to assume it has a higher quality, but I am not automatically convinced of that either.

Rapzid
Sure, I mean I didn't mention length but to your point:

> Why is GitHub down so often?

Short. Invites conversation. Shows intellectual curiosity.

> It's so over

Reactionary, performative, low effort slop that's out of place.

Unfortunately the latter is on the rise big time and overly prevelant in this thread.

Drupon
Always thought this rule is hilarious because the fact that such comparisons are so overwhelmingly common that they needed to address it with this bitchy little rule (adding one link to an example per word is a classic tell that an internet moderator is spending their highly compensated moderation time Extremely Mad lol).

However, rather than understand the obvious explanation for it, that reddit and HN's karma systems encourage performative comment behavior with people trying to be epic for the peanut gallery, they make the boneheaded decision to just assume it's an "illusion". It's like Dunning-Krueger but for their theory of mind and general emotional intelligence, with the cause being obvious enough that I don't feel I need to name it.

I suspect it'll only become more common as Reddit enshittifies further and further. The platform is basically unusable for anyone, thus its suitability as a 'containment zone' for the redditor phenotype goes down with it.

I hope BlueSky makes a Lemmy-style Reddit alternative, just something to satiate people's social media-ified forum fix.

“People here have opinions that I don’t like.”
sdcfgy
Nah. I can't see anyone blaming it on Trump, Israel, datacentre water having a memory, their parents, their neighbour, spirits or spez.
Why bring trump at all into the discussion, if not just to sh*tpost?
sdcfgy
Everything I listed are symptoms not causes...
user-
I can attest to the outage starting around 11:05, as that’s when my GitHub auth for Tailscale SSH broke. Of course, the next 10 minutes of troubleshooting and checking GitHub’s status page showed “all systems operational,” so I assumed it was something on my end. Then, it started working again.

And while I was typing this comment it just blipped again. Atleast this time I know why.

I feel dumb for setting up tailscale like 7 years ago using the github auth, time to figure out how to just move to email on that account.

> I feel dumb for setting up tailscale like 7 years ago using the github auth, time to figure out how to just move to email on that account.

Unless I've missed something, I don't think you can. I also had to use GitHub auth, as it was the least-bad option of the choices given. I wish they just did email/user + pass + TOTP like regular platforms.

stryan
You can't do just an email address, but if you're running your own domain you can do your own OIDC with it. My tailnet auths against my self-hosted KanIDM server.
"We don't want your passwords."

https://tailscale.com/blog/passkeys

Maybe running your own OpenID service would work?

https://openid.net/developers/how-connect-works/

Too much work though. Easier to just switch to Nebula or Netbird and self-host the full stack.

zrail
There aren't any restrictions anymore on who you can invite to your team. I added a passkey admin account and added other users via email invites. The account owner is still my GitHub account but I pretty much don't use it day to day.
1105 utc? I had (self hosted) runners running at 1336 gmt although when I had a quick look it felt a bit odd - jobs ahead finished but hadn’t updated the pr status
user- OC
11:05 EDT
gwbas1c
What's frustrating is that:

1: The status site says that pull requests are working,

2: A PR I'm working on right now is missing commits,

3: The contact support page errors out.

atsjie
Contact support page is probably just a waste of time even if it did work.
nevir
Definitely not resolved yet.
Yeah can't push or create PRs yet for me.
Embarrassing that a multi-billion dollar corporation that other billion dollar corporations depend on seems to be fine with breaking their clients workflows. Was github always this shoddy or has something changed to cause this many outages?
I mean... they got acquired by Microsoft.
Almost 10 years ago
They used to run their own data centers, and while the unicorn (their error page back in the day) was visible sometimes, it wasn't nearly as bad as things are today.

And of course, Microsoft is saying that none of this downtime is at all related to them moving everything to Azure, and also at the same time they'll fix all this downtime by finishing moving everything to Azure.

I wouldn't hold my breath here.

mjr00
> has something changed to cause this many outages?

AI coding. Which can be interpreted one of two ways:

* The generous way, which is that Github is so overloaded with massive volumes of AI codebases and AI-driven automation that they're hitting a scale they never anticipated; or

* The not-so-generous way, which is that Github itself was one of the first companies to push everyone to AI code as much as possible, which has lead to an eventual breakdown of the stability of the system, as the people responsible for it no longer understand how it actually works, leading to production outages once every few days.

It also coincides with GitHub moving to exclusively Azure infrastructure. Azure is a terrible shit show for reasons that have nothing to do with AI.
Haha I hope this is true, it fits my experience on Azure. 0 machines available. If someone from GitHub wants to chime in.
Azure sucks balls. Its support has no clue whats going on.
It's never the wrong time to link this TFA: https://news.ycombinator.com/item?id=47616242
User error
Rapzid
That's a correlation but I've never seen a proof of causation. The variable can't be isolated because they also started shipping more features.

Also, the uptick in odd issues post acquisition pales in comparison to the scaling issues they've had over the past 1-2 years. As a matter of scale, trying to link this back to Azure doesn't really square. If anything they'd potentially be in a worse spot without having access to the resources of a massive public cloud..

There were internal warnings about making the move from highly customized MySQL bare-metal clusters to cloud infrastructure because of network latency, moving from physical machines to multi-tenant environments and having to rewrite a lot of their existing customization.

The other issue of trying to quietly move terabytes of information while the site is live hasn't helped either. Add in the challenge of keeping highly customized bare-metal databases in sync with a cloud environment isn't easy.

Its this very very very complex migration that's resulted in a lot of the ongoing issues.

https://hostingjournalist.com/news/microsoft-accelerates-git...

The migration underscores Microsoft’s strategy to unify its AI and developer ecosystems under Azure, bolstering performance and reliability for Copilot and related AI workloads. However, not all GitHub employees are confident in the transition. Internal concerns have surfaced about potential service disruptions, particularly given the complexity of moving GitHub’s massive MySQL clusters, which currently run on custom bare-metal infrastructure.

Outages have become more frequent in recent months, a symptom of GitHub’s growing operational strain. Insiders suggest the platform’s infrastructure - originally designed to handle conventional development workloads - is now being pushed to its limits by large-scale AI integrations and surging user activity.

Rapzid
> There were internal warnings about making the move from highly customized MySQL bare-metal clusters to cloud infrastructure because of network latency

I'm sure there were, every plan has detractors. Azure has bare metal offerings; does GitHub have access to them? Vladimir Fedorov, GitHub CTO, cited constrained capacity in their data centers and accelerating Azure migration as critical for them to deal with the increased AI load.

IDK, but again as a matter of degree the big issues they are having seem more strongly correlated with agentic coding explosion. In fact it motivated their accelerated migration plans.

o_m
Give them some slack. Their core product is AI, not version control
ainar-g
> Was github always this shoddy or has something changed to cause this many outages?

The data speaks clearly: https://damrnelson.github.io/github-historical-uptime/

They went woke. Never seen anyone resembling the average dev in their promotion… You can see the correlation in their downtime graph.