Show HN: Herodotus – An IRC bot that logs a channel's activity to JSON
github.com
github.com
I'd much prefer an array with objects as element, each object storing its own timestamp.
In addition, that's just based on my perception; IRC is NOT an ordered protocol, so if multiple people ran this same logger, they'd end up with different stream and different duplicate keys (same thing as events lost I might add).
My log, also, conveniently contains joins, parts, notices, ctcps, and all that other jazz because I used a mature logging library.
var array = [];
client.addListener('message', function (from, to, text) {
var time = Math.floor(new Date() / 1000);
var obj = { nick: from, message: text };
var item = '\"' + time + '\"\:' + JSON.stringify(obj);
array.push(item);
fs.writeFile(file, '{' + array + '}\n', function() {
// console.log("updated");
});
});
This will buffer all the messages in memory (in the variable "array") and rewrite the file on every message. Besides the lack of atomicity mentioned in another comment, that's also really inefficient and will eventually run out of memory. Plus 1000 incoming messages will lead to 1000 sequential rewrites of the same file!The other "Show HN" on the front page right now, for comparison, is Claudia.js. That's not something anyone on HN could trivially write in 5 minutes. It's about 5k lines of javascript. It's already somewhat mature.
For comparison, the project you posted is under 50loc, any of us could trivially write (in fact, piping http://tools.suckless.org/sic/ would be sufficient... or using znc with its "log" module, or using weechat / irssi's builtin logging functions), and doesn't really accomplish something super interesting to HN imo.
I expect everyone who uses IRC in a significant capacity on HN has already solved the most basic level of their logging problem. I solved mine by just running a ZNC bouncer with the log module. The harder level of the logging problem, indexing it and presenting it in a nice UI and solving availability by merging logs from multiple leafs in a netsplit event, I'm not so you can consider commonly solved, but hey, your solution has nothing to say there either.
Others have already mentioned that there are some issues with your code, so I won't touch on them, but I will emphasize that you should probably work on improving code quality and make something both more significant and more interesting prior to doing a Show HN.
In addition, this actually is something people on HN actively couldn't trivially use because it has no configuration (hard coded connection strings) and no docs.
That'd require one log/connection per server of the network, which isn't something ZNC or weechat will do by default.
Have you ever needed that functionality?
For absolute correctness, you'd want to record per-leaf and merge.
However, the last time I needed log-merging functionality was much more boring; some of my log-files were corrupted, and someone else had logfiles that weren't as extensive, thus I needed to munge the two together. The timings were subtly different (because that's how irc works), so I wrote custom code to munge them together, preferring his for the range of corruption.
So yes, I've needed that functionality, though it wasn't actually done due to a netsplit. I can foresee it being useful in the case of a netsplit if your logging is at the ircd level or you run one bot per leaf.
My own IRC log tool [GH: tilpner/ilc] can be used to merge logs, but I rarely use that functionality.
For two log files "a" and "b", in weechats default log format:
ilc sort -f weechat -i <(cat a b) | ilc dedup -f weechat
I've never tested this with logs over 200MB, but sort will read the combined log into memory, which is definitely not optimal.The bar for Show HN doesn't have to be that high. It's ok for people to post projects that they're working on while learning something, for example.
https://github.com/egladman/herodotus/blob/master/server.js#...
How can a developer be that stupid?
file.write is atomic (guaranteed to "work") for PIPE_BUF (posix) octects ~ 16ko at most.
JSON is guaranteed to be unparsable if the file is truncated of the last chars (it is not very resilient).
Hence the write may corrupt your WHOLE log if a non recuperable failure happens or the code is interrupted.
The code DOES not use fixed size allocation ... thus is can crash randomly because of SEGFAULT in the middle of the writing. The history will take cumulative size in memory.
This coding attitude highers the probability of this failure to happen BY DESIGN.
Is it that complex to FIRST write the file, and THEN atomically rename the new file to the old file at worst (resulting only in losing the current session, but not the whole history).
If your log are that precious, why would you not take extra care about protecting them?
This code makes me want to puke.
On the other hand, it is representative of the reason why I am disenchanted by modern coding standards.
Do you live in a state of perpetual nausea? Because there is so much code out there like this...I fear it could overwhelm your life to worry about it.
I think the best we can do is just fix the stuff we want to use and ignore the rest..
But I guess you may lack of context to get it right, as for the rest of your comment. ;)
This comment breaks the HN guidelines so badly that we could put it on a poster of how never to behave here. It doesn't matter how right you are if you express it this abusively.
Since you've done this more than once before, I've banned your account. If you don't want it to be banned, you're welcome to email hn@ycombinator.com and give us reason to believe that your comments will be civil in the future.
All: In addition to https://news.ycombinator.com/newsguidelines.html, please know that HN has extra rules at https://news.ycombinator.com/showhn.html to guide the discussion of new work. The parent comment breaks every one of them.
file.seek is your friend. You should not use any dynamic allocation to make it more robust (yes preallocating fixed memory size with "char circular_ring[N * MAX_SIZE]") Where N and MAX_SIZE are fixed.
Well. It requires a complete rewrite for this code to not create a risk for your users. Your code is like full of hidden landmine by lack of design.
And you know what?
It used to be standard knowledge for introduction to CS for scientific. You know why? Because no users of computers like to lose their precious data made during a very long measurement spilling huge amount of data and costing a lot of resources. (like gold, helium, molybden, electricity, time, wages of qualified operators....)
It seems like coders do not care of their users, like a builder thinking it is reasonable to build 50 store buildings on quicksands because people only judges builder by the look of their creation.
We banned SFjulie1: https://news.ycombinator.com/item?id=11141174. If you violate the rules of this community that badly, you forfeit the right to comment here.
We detached this comment from https://news.ycombinator.com/item?id=11141103 and marked it off-topic.
How do I vouch for a dead comment?