That said, there's one useful feature that regrettably is poorly supported in most allocator discussion - where some higher alignment is wanted but not at the exact start of the object.
Microsoft has the only major allocator that I'm aware of that supports this; it uses the "align at offset" sense (which simplifies the case of prepending a header to an aligned object). GCC has some minimal support and uses the "align with offset" sense (which IMO makes literally everything else much easier). As a minimal example, if 0x4 is the low nybble of your pointer, Microsoft calls that "align 16 at offset 12", whereas GCC calls it "align 16 with offset 4".
Note that regardless of sense, you can store both the alignment and the offset in a single integer-sized variable by using your compiler's "find highest bit" or "round down to a power of 2" intrinsic. By exclusively using wrapper objects you can in fact hide the internally-chosen sense and support both publicly. This does however remove the possibility of only passing the log of the alignments, popular on BSD allocators.
Limiting `align` to positive `isize` is useful, but further restrictions can reduce the number of overflow checks you often need. I have never found a practical use (mostly: considering huge pages) for more than 3/4 of the bits to be used on alignment - that is, 6 bits (64, a typical cacheline nowadays) for 8-bit pointers, 12 bits (4K, a typical non-huge page nowadays) for 16-bit pointers, 24 bits (16M - suported on many ISAs, just not x86) for 32-bit pointers, and 48 bits (256T - theoretically available on RISC-V) for 64-bit pointers. You'll likely want a +1 when you use this, at least if you implement the logic the way I did.