Show HN: sheet2dict – simple Python XLSX/CSV reader/to dictionary converter
github.com
github.com
One of the things I still don't understand are services like snyk.io, which are supposed to do security analysis. But they penalize a tool like this for not having CoC, Contributing in the GitHub repository, and what is most shocking to me is that they measure Popularity. I understand that if more people are involved in the SW, it is probably safer. But penalizing someone for having few stars on GitHub seems weird to me. Especially when the tool is used by several people / companies and it has over 5,000 downloads.
https://github.com/capitalone/DataProfiler
My question, how do you do header detection? That's a _very_ difficult problem.
https://github.com/h2oai/datatable/blob/385a9b370db32de90135...
Xlsx I know nothing about.
https://github.com/Pytlicek/sheet2dict/blob/main/sheet2dict/...
Thanks for removing some boilerplate from that process for people!
Do you mean "boilerplate" or something else? Because a shell script (e.g. bash with calls to sed, awk, grep, and Perl) is not "boilerplate". It's an implementation. And much more efficient than some of the unstable, high-complexity solutions that claim to be "simple".
TBH, the choice of "solution" doesn't matter much for a small dataset. It is of course overkill to run a glorified REPL just to do some math on a small dataset.
As datasets grow, the introduction of unstable complexity can cause problems.
The tooling that you say is being removed may well be the the fastest and most reliable tools, proven over decades of use.
Is this a better approach?
Isn’t a big point of dataframes providing tools that are more efficient so you don’t have to use Python loops for operations across a body of data?