Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I think you're misinterpreting the comment you're replying to. They would agree with you that the tiny SRE team described in the article sounds very effective, and likely have a lot to do with why the site is still up and running currently. Work like that should continue. But if 1-3 people can have that degree of impact, what are the other 8000 doing? (Again, this is just me attempting to interpret the point made by the parent, not trying to make one myself.)


The SRE team mentioned in the article is SRE for one component of a complex architecture.

There are probably many such components; I'd imagine SRE alone would be 200+ people

How many of the remaining staff have the knowledge required to keep all of those components running smoothly?


Once you get the automation going the number itself doesn't matter that much.

You might have 200 different apps (hell, we have close to that, only 3 people in ops) but competent team will make sure they deploy in same way and are monitored in same way.

And once you go from "a server" to multiple servers, whether the end number ends up being 20 or 200 isn't that important till you start hitting say switching capacity, and if you're in cloud that's usually not your concern anyway.

Our biggest site (about dozen million users, a bunch of services and caching underneath, few gbits of traffic) took zero actual maintenance for 2022, "it just works", any job was implementing new stuff. It took some time to get to that state but once you do aside from hardware failures it "runs itself"


> Our biggest site (about dozen million users, a bunch of services and caching underneath, few gbits of traffic) took zero actual maintenance for 2022, "it just works", any job was implementing new stuff. It took some time to get to that state but once you do aside from hardware failures it "runs itself"

Nobody is adding changes that blows out the DB? or add some inefficient code that burns CPU much faster?


> Nobody is adding changes that blows out the DB? or add some inefficient code that burns CPU much faster?

not to be flippant but you find that out like 3 environments below production.


It's not 1-3 people. The entire SRE team globally - including the technicians and the engineers with server access - is easily going to be in the hundreds.

The SRE manager is in charge of keeping it all running. He isn't running around the world swapping out servers. He also isn't sitting back with his feet up thinking "All done - now how are my Pokemon doing?"

It's a dynamic process with quality monitoring, budgeting and reports, post-mortems, continual experiments to see if uptime can be improved, and redesigns as hardware and software change.

It's part of the backend, but is only loosely coupled to the content management and delivery system, the ad machine, moderation, marketing, and so on, all of which are going to have similarly complex structures.


So 1-3 people have a big impact, the other 7997 must not be doing anything? I don't think that logic follows.


it doesn't follow. The article posits that "many people think twitter headcount was bloated" then proceeds to describe a (presumably) really efficient work of a small SRE team. These two parts seem completely disconnected from each other - neither one proves, disproves or follows from the other - so it's unclear why the former was mentioned at all.


His personal experience was zero bloat. He was one deep in a critical function for the company. He isn't saying this proves without a doubt that there is no bloat, but he didn't see any in his time there. It's seems like a reasonable addition to the conversation to me.


It's a shame that most the conversation going on here are extrapolated arguments based on this article and another anecdotes. The problem starts when the ones making their points let beliefs on "how things should be" stronger instead of "how they really were".




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: