57 karma · joined December 13, 2023
Not half bad!
Your last paragraph is interesting though, in what way do you think mobile dev is going to change?
I do feel we're heading in a direction where building in-house will become more common than defaulting to 3rd party dependencies—strictly because the opportunity costs have decreased so much. I also wonder how code sharing and open source libraries will change in the future. I can see a world where instead of uploading packages for others to plug into their projects, maintainers will instead upload detailed guides on how to build and customize the library yourself. This approach feels very LLM friendly to me. I think a great example of this is with `lucia-auth`[0] where the maintainer deprecated their library in favour of creating a guide. Their decision didn't have anything to do with LLMs, but I would personally much rather use a guide like this alongside AI (and I have!) rather than relying on a 3rd party dependency whose future is uncertain.
I would say this oversight was a blessing in disguise though, I really do appreciate minimizing dependencies. If I could go back in time knowing what I know now, I still would've gone down the same path.
I think `ts-rest` is a great library, but the lack of maintenance didn't make me feel confident to invest, even if I wasn't using express. Have you ever considered building your own in-house solution? I wouldn't necessarily recommend this if you already have `ts-rest` setup and are happy with it, but rebuilding custom versions of 3rd party dependencies actually feels more feasible nowadays thanks to LLMs. I ended up building a stripped down version of `ts-rest` and am quite happy with it. Having full control/understanding of the internals feels very good and it surprisingly only took a few days. Claude helped immensely and filled a looot of knowledge gaps, namely with complicated Typescript types. I would also watch out for treeshaking and accidental client zod imports if you decide to go down this route.
I'm still a bit in shock that I was even able to do this, but yeah building something in-house is definitely a viable option in 2025.
And yeah, I've been using prometheus' `collectDefaultMetrics()` function so far to see event loop metrics, but it looks like node:perf_hooks might provide a more detailed output... thanks for sharing
I ended up figuring out a fix but it's a little embarrassing... Optimizing certain parts of socket.io helped a little (eg installing bufferutil: https://www.npmjs.com/package/bufferutil), but the biggest performance gain I found was actually going from 2 node.js containers on a single server to just 1! To be exact I was able to go from ~500 concurrent players on a single server to ~3000+. I feel silly because had I been load-testing with 1 container from the start, I would've clearly seen the performance loss when scaling up to 2 containers. Instead I went on a wild goose chase trying to fix things that had nothing to do with the real issue[0].
In the end it seems like the bottleneck was indeed happening at the NIC/OS layer rather than the application layer. Apparently the NIC/OS prefers to deal with a single process screaming `n` packets at it rather than `x` processes screaming `n/x` packets. In fact it seems like the bigger `x` is, the worse performance degrades. Perhaps something to do with context switching, but I'm not 100% sure. Unfortunately given my lacking infra/networking knowledge this wasn't intuitive to me at all - it didn't occur to me that scaling down could actually improve performance!
Overall a frustrating but educational experience. Again, thanks to everyone who helped along the way!
TLDR: premature optimization is the root of all evil
[0] Admittedly AI let me down pretty bad here. So far I've found AI to be an incredible learning and scaffolding tool, but most of my LLM experiences have been in domains I feel comfortable in. This time around though, it was pretty sobering to realize that I had been effectively punked by AI multiple times over. The hallucination trap is very real when working in domains outside your comfort zone, and I think I would've been able to debug more effectively had I relied more on hard metrics.
I could increase this interval, but I'd like to keep it as short as I can afford to to keep that realtime feel (i.e. other players can see what the current turn player is typing).
There is no cross-room communication. I could spawn a process per room but I was trying to address this issue with my current Docker setup where I have multiple `game` containers that run a single node.js process and each process can host multiple rooms.
Not having to use Docker sounds simpler but it's that's where I'm at atm haha.
I agree that the network load feels very small. Maybe it's a socket.io related issue where when many broadcasts are being fired at once, then a shared I/O step gets bottlenecked?
Here's my actual typing broadcast code, I was originally broadcasting from the socket event callback itself but I found performance improved slightly by batching broadcasts per player in a setInterval loop (also note that only 1 player in a given room can be typing at once, so batching broadcasts per room shouldn't address the bottleneck).
/**
* Used to handle very frequent typing events more gracefully to avoid overloading CPU
*/
const TypingUsersMap = new Map<
ConnectionId,
{
socketId: string | null; // doesn't exist for bots
roomId: PublicRoomId;
userId: UserId;
currentInput: string;
}
>();
type ConnectionId = `${UserId}:${PublicRoomId}`;
// ! this should be same as client throttle interval
const TYPING_BROADCAST_INTERVAL = 200;
export let typingBroadcastInterval: NodeJS.Timeout | undefined = undefined;
export const startTypingBroadcastJob = () => {
typingBroadcastInterval = setInterval(() => {
const freshTypingUsersMap = new Map(TypingUsersMap);
TypingUsersMap.clear();
if (freshTypingUsersMap.size === 0) return; // Nothing to do
// Go through each user that has a pending update
for (const [_connectionId, data] of freshTypingUsersMap.entries()) {
const socket = data.socketId
? io.sockets.sockets.get(data.socketId)
: undefined;
// Use the data we stored to perform the broadcast
if (socket) {
// emit to other players
socket
.to(data.roomId)
.volatile.emit(
SOCKET_EVENT_NAMES.USER_TYPING_RES,
data.userId,
data.currentInput
);
} else {
// bots emit to everyone
io.to(data.roomId).volatile.emit(
SOCKET_EVENT_NAMES.USER_TYPING_RES,
data.userId,
data.currentInput
);
}
}
}, TYPING_BROADCAST_INTERVAL);
};
export const stopTypingBroadcastJob = () => {
if (typingBroadcastInterval) {
clearInterval(typingBroadcastInterval);
typingBroadcastInterval = undefined;
}
};
// this is called from the USER_TYPING socket event callback. so effectively every throttled keystroke by the user gets queued.
export const queueTypingEvent = ({
socketId,
roomId,
userId,
currentInput,
}: {
socketId: string | null;
roomId: PublicRoomId;
userId: UserId;
currentInput: string;
}) => {
const connectionId: ConnectionId = `${userId}:${roomId}`;
TypingUsersMap.set(connectionId, {
socketId,
roomId,
userId,
currentInput,
});
}; net.core.wmem_max = 16777216
net.core.rmem_max = 16777216
net.ipv4.tcp_wmem = 4096 65536 16777216
net.ipv4.tcp_rmem = 4096 87380 16777216
Perhaps the reality for low latency multiplayer games is to embrace horizontal scaling and not vertically scaling? Not sure.