34 karma · joined January 13, 2024
Sounds like a real A player.
It's easy to cherry pick a PDF for marketing purposes and claim you're better. I didn't miss it, I just don't believe marketing announcements at face value. I tried their parser on a PDF with a bit of complex formatting like multiple columns, tables and a couple images and it choked, spitting out one big markdown header with jumbled text. Not impressed.
1. The sign up with email just endlessly redirected, click link in email, ask to sign up with email, put in email, click link in email, etc.
2. Fine, I'll sign in with Google.
3. A PDF parser? Seriously that's what all this fuss is about? There are so many options already out there, PDFBox, iText, Unstructured, PyPDF, PDF.js, PdfMiner not to mention extraction services available from the hyperscalers. Super confused why anyone needs this.
Does that mean failure to adhere to the terms of licenses like ASLv2 or MIT would could only be settled in court if the person who wrote the code actually bring suit and that OSS foundations basically become (more) toothless?
Haha, yeah I think it's the "fanbase" around it (don't want to offend anyone). It's like that old joke, how do you know someone is (a vegan, doing keto, plays guitar)? They'll tell you. There almost seems to be people who want to turn these diets into a competition, oh you're keto? Well, let me tell you how crazy keto I am. Like dude, I'm just trying to not hate myself when I look in the mirror not win the Mr Keto award.
Wow, such a great testament to The Unix Philosophy of building small, modular, focused tools that can be combined together to do all sorts of interesting and more complex tasks. I'm sure no one imagined using these utilities from a helicopter to retrieve rover logs to aid in diagnostics, but here we are. What a cool story.