In the 1%, though, you'll have needs that malloc/free don't fit. Complex object graphs don't really have a single point of ownership or requiring that you free something in all the places that it might be released is too much to handle. In this case you can reach for a garbage collector (including, for example, implementing reference counting). Other times, you may need to make a lot of allocations in a short time where they can all be freed at once. Request processing in a network server is a common example of this: once the request is complete, everything allocated can be dropped, and you generally want minimal latency.
All of this comes with tradeoffs, though. With a GC, you lose predictability and performance changes; sometimes for the better and sometimes for the worse. Depending on the GC, you may not be able to have stable pointers, and you may lose the ability to finalize objects. With an arena, you can't free or reallocate, so you need to scope the arena to a small region of execution (this is where the author went wrong, for example).
Finally, regardless of which approach you use, you're going to need to thread the allocator through the application; probably implement your own datastructures, etc. Depending on how complex your memory management model is, you may need more than one allocator at any given point (e.g., a GC for the persistent data and an arena per connection and per request). If you're implementing your own allocator, you'll also likely have bugs, and allocator bugs tend to be insidious and obnoxious to debug.
If you can avoid going down that route, I recommend it. Sometimes, though, you have enough constraints that you need to brave the jungle.