We've now run infrastructure reviews across companies from five-person startups to engineering orgs with a few hundred people. The technology varies enormously. The underlying issues, surprisingly, don't.
The single most common finding, by a wide margin: nobody has actually tested restoring from backup. Close behind: a deploy process that depends on one specific person being available. Third: monitoring that pages for things nobody acts on, which has quietly trained the team to ignore real alerts along with the noise.
None of these are hard problems, technically. They persist because fixing them never feels as urgent as the feature work competing for the same time — right up until the week they become the most urgent thing in the company. That gap, more than any specific technology choice, is the thing we exist to close.