Tad, a tabular data viewer
tadviewer.com
tadviewer.com
- Sad to see it's an Electron app, but I don't want to start this flamewar all over again
- Mid-sized CSV file opens fairly quick
- If you want to build this into a full-fledged app, be prepared to handle a lot of CSV edge cases, also think about supporting AVRO, Parquet files
- Regarding filtering: I can filter on CONTAINS and '=' but not on NOT-CONTAINS or '<>' I think? This is annoying to filter out empty strings (didn't check in much detail so I might be wrong here)
- Regarding filtering: I have a lot of columns, would be nice if the dropdown would allow me typing some characters to prune the list and quickly find what I'm looking for
- You might think about including some statistics per column, e.g. variance or entropy to allow for exploration in case many columns are present and you quickly want to highlight the more "interesting" ones
Congrats on putting something out! Some competition in this space: https://exploratory.io/, https://github.com/saulpw/visidata and http://www.delimitware.com/ (as others have mentioned).
Then stop mentioning it. Just don't use anything electron and don't mention it. I will use my RStudio and VS Code.
It is, but it's not okay to bring up a common flame war topic and then say "but I don't want to start a flame war". If you don't want to start one, just don't start one. Starting it, and then saying you didn't want to, is disingenuous at best.
The "I don't want to start a flamewar" bit is also a reasonable addendum. It's a terse way of saying, "I don't like this feature, but, people who disagree with me, please don't get up in arms about this, you're free to like this feature if you want."
Frankly, if there's anyone trying to start a flamewar, it's not the original poster; it's people who are edging toward caping up in response to OP's (quite mild) comment.
Is there any substance here?
I think it's fair to talk about this, as long as we go out of our way not to reduce it to flamewar territory. The app is built on Electron, after all.
Maybe avoid mentioning memory usage, since that point has been done to death.
All the JS-the-language criticisms apply, those are probably the biggest at this point. The classic "node.js is cancer" is still somewhat relevant though I never thought it made its point well enough on how forced-asynchronous is at least as bad as the forced OOP and checked exceptions we put up with in e.g. Java again. Callback hell was long a common complaint, until people found Promises / the waterfall structure, but then the complaint is debugging hell. The new await stuff in node 8 sounds nice at least.
For a long time npm didn't get stuff over https by default, I think it does for a year now? The only other criticisms I can think of are just shoddy engineering that turns into drama (I remember seeing something about their colorized output would slow everything down quite a lot) or more taste-level things like how seemingly trivial concepts like leftpad are core libraries (that go through multiple versions because it couldn't have just been done right the first time) and the drama involved when something everyone depends on has an issue.
Disclaimer, I haven't done production node since ~2013. It was actually mostly enjoyable, and I wouldn't mind doing it again, despite on the taste-level I think the whole node ecosystem is just worse than many other options (when you have a choice) for many specific problems. But life is a continuous lesson in Worse is Better.
NPM is quite basic, it pretty much just downloads tarballs recursively from a webserver according to some version spec. The CommonJS module format is fairly loose (and kind of outlived its usefulnes) and so there's no agreed upon package structure and everyone just does whatever they feel like doing, which leads to a huge overall mess. The whole ecosystem is fragmented and very few things are really agreed-upon.
This is about as valuable as "Rails is a webserver that responds to some HTTP verbs". Common.js isn't even a part of npm (the correct capitalization) or node.
I never said it was, it's not part of it per se, but it's the 'blessed' module/package system on Node.js (the correct capitalization). Which is really sad, because CommonJS is retarded (a file has to actually be executed in order to get it's require & export characteristics and they might arbitrarily change in runtime). It's sad that Node.js as well as the JS ecosystem in general is fairly hostile towards ES6 modules - basically any current support of ES6 modules is just translating then / dumbing them down back to CJS. Hopefully this'll change in the future as ES6 gets more adopted and CommonJS will be left to rot (a fate it deserves and should've already met ages ago).
It's a modified dataTable(so search is builtin. I didn't see search in Tad yet), with selectize for categorical variable filters(which is a searchable dropdown list)
Another good to have feature is to show column index. We often need to manipulate the columns in code, a column index is helpful.
Based on column index, you can also select a subset of columns faster with numeric input -- over 100 columns are normal, using checkbox to select is too cumbersome.
+1 for the comment about Electronic/Javascript, also don't want to start a flame war heh.
exploratory.io looks really neat, it's a shame about the pricing and I'm not quite sure if it's a native app once you purchase or if it's an in-browser online-only app, because if it's online only 1) I'd be worried about uploading data to a 3rd party and 2) The internet in Australia is _terrible_ so uploading / downloading larger data sets might be painful.
- I have very sensitive data. Is my data safe?
Any data you import into Exploratory Desktop always stays on your PC and never leave your PC unless you explicitly publish (share) it to Exploratory Cloud (exploratory.io). If you decided to publish the data to Exploratory Cloud for sharing or scheduling, you can share it in a private way so that so that only you and others you have invited can view it. The data is also stored in encrypted. Please take a look at our Privacy Policy for more details.
- Where exactly my data is stored after importing?
All the data you import into Exploratory Desktop is saved as a binary form (R's Rdata format) inside your repository, which is located under '/.exploratory' on your PC.
Qt - The IDE is great. It doesn't look native. Either dynamically linked, or commercially licensed. Runs on Windows, Mac OS X, Linux, Android, iOS, and a range of embedded hardware.
wxWidgets - Native backends, so it always looks like it belongs. Bindings in a ton of languages. Can be both simple or complex, depending on what you need it to do. Runs on Windows, Mac OS X, Linux, and in-progress for Android and iOS.
JavaFX - Java's replacement for Spring. Runs anywhere Java does. Fairly flexible, and easy to use.
Kivy - A Python framework. Mainly aimed at touch-compatibility. Runs on Windows, Linux, macOS, iOS, and Android.
LCL - A Lazarus framework. (Think Pascal). Really easy to use, with great drag'n'drop and the like in the IDE. Runs on Windows, macOS and Linux.
nuklear - Fairly easy to use, with bindings in a lot of languages. Sometimes requires a bit more work getting it to run on some platforms.
This is not all, but the ones I find easy to use (as easy or easier than Electron), and easy to set up and deploy.
Electron doesn't look native, which a lot of consumers don't like. Others don't care. So, depending on your audience, it can be an extra hurdle that some like wxWidgets don't have.
Electron is difficult to get performant. You need to really think and test your perf. Qt, Tk, and a few of the others are much faster, much easier. And if you put in the same effort you need to put in for a fast Electron program, you can end up with blisteringly fast speeds.
Electron bundles are large to download. Some places this doesn't matter, others it does. (Think Australia, Africa and the like where tiny download caps exist). IIRC Qt, the biggest dependency, is about half the size of Electron. Tk and wxWidget are small, nuklear is just a header. It's tiny.
Electron is not good with touchscreens (or "harder to get right"), but more and more people have them. Kivy is great for that usecase.
Electron is "just code". Qt Creator and Lazarus' IDE are phenomenal for putting GUIs together. You'll be surprised how little code you need.
Electron is primarily JavaScript. Some people like that, others don't. If you want an easy language, you can use Python, or Nim. Want more control C or C++. Value both time and control? Go with Java.
Electron has its place.
If you have a really tight time-to-market window, or this is just a tiny personal project, and JavaScript is your "goto" language, then awesome. Use it.
But, the cross-platform GUI world is big. You have a dozen or so mature, stable, proven frameworks.
So if your project takes off, or you have the time to do things right, evaluate if any of these libraries, including Electron, fit your needs.
But if you're going that route, you'll be making your life a lot harder in the cross-platform department.
* It is backed by Qt and rendered by OpenGL, so it performs very well.
* It has very easy syntax, especially automatic variable binding, so there is very little boilerplate for event handingl.
* It supports JavaScript. You can code the logic in either JS or C++, depending on whether you need the performance or not, and it can be easily split between the two.
* They actually support two sets of pre-made controls, one that looks like desktop UI (http://doc.qt.io/qt-5/qtquickcontrols-index.html) and one that looks like Android or iOS (https://doc.qt.io/qt-5/qtquickcontrols2-index.html). Or you can design your own UI from rectangles and other shape. Or combine all three.
It might look a bit dated on some platforms, but it will be lean and fast, unlike Electron.
Enterprises still use Swing and some have touched FX as Swing is in maintenance mode and they don't have a choice.
Some advice on the website: Do some browser sniffing (I know) to display a screenshot of the software on the user's operating system. This immediately answers the question "Does the software work on my computer?" Also, the source code link should be near the Download section. Not everyone is trained that the triangle GitHub icon in the top right means "source code".
Otherwise, nice app. How does it scale to larger CSV files? I have been working on something similar (https://warp.one) for Mac, which streams CSV files (apparently you use a SQLite database behind the scenes as cache?)
Based on prior experiences with CSV, one of the big problems I've seen has been figuring out what text encoding is being used. This problem appears to be more prominent with people outside of the US. It looks like Tad is using fast-csv, which I don't think will properly handle different file encodings. Life would be so much simpler if everyone just used UTF-8.
1. Write your CSV parser with an assumption that the data is ASCII compatible (this means it works with either UTF-8 or Latin-1 out of the box, possibly modulo non-ASCII meta characters). To support additional encodings---such as UTF-16---either the CSV library or the caller must transcode first.
2. Write your CSV parser such that it can work on multiple different encodings. For example, this means looking for `\x2C\x00` when parsing UTF-16LE data instead of just `,`. This introduces implementation complexity, and you'll be unlikely to support the full gamut of encodings that other tools support whose job it is to do that sort of thing.
(2) is kind of weird but probably quite a bit faster than (1), although I can imagine it being useful in very niche circumstances. e.g., "I have a boat load of UTF-16 encoded CSV data and transcoding it to UTF-8 to use this CSV parser isn't worth my time because ______." I can't actually fill in that blank, so solutions in (1) tend to be the way to go.
Now... If you're building a full on CSV tabular viewer, then I might understand why it should handle encoding for you automatically, but when it comes down to it, the viewer is still going to need to choose between (1) and (2). Unless they want to hand roll their own CSV library, I imagine they're just going to pick (1), and when possible, transcode the data first. In that case, it shouldn't really matter whether their underlying CSV parser supports alternative encodings or not.
I think the only other time I've seen something that hacky has been with a date parser that would try to guess the format for you. That pushed me to the firm belief that ISO8601 is the only sensible way to store dates.
> OpenRefine (formerly Google Refine) is a powerful tool for working with messy data: cleaning it; transforming it from one format into another; and extending it with web services and external data.
There are some projects out there using memory mapped files to do fast CSV parsing. Could be a nice way to speed up the memory loading and scroll it in real time. Can't find the link to the library I saw it used in, but it might be an interesting venue to consider. Another library that does it seems to be astropy fast ascii IO module [1].
[1]: http://docs.astropy.org/en/stable/io/ascii/fast_ascii_io.htm...
Not that LibreOffice gets a pass, but it doesn't crash nearly as often for me.
But honestly, if you care about collaborating on text and not its formatting, then I'd suggest hosting an Etherpad instance somewhere :).
Is this supported at all, or do I need to use the copy mechanism? I have a ~4GB file to filter down and analyse, and this almost looks like a nice tool for business users to use to explore the data themselves, but they need to be able to export to excel at some point.
Mainly for the cascaded pivot option.
First remarks : - Requires \n or \r\n end line, not working with \r
- Requires comma as separator (not possible to change it or I don't find how)
- Does not support (for exemple) ANSI encoding
test.csv: UTF-8 Unicode (with BOM) text, with CR line terminators
Apparently, even recent software is not up to date on what the line separator should be.
EDIT: once you install it, a rich README pops up which includes more screenshots and some example datasets. Fun to play around with
Error: /usr/lib/x86_64-linux-gnu/libstdc++.so.6: version `GLIBCXX_3.4.21' not found (required by /tmp/.org.chromium.Chromium.ibqKoR)
$ strings /usr/lib/x86_64-linux-gnu/libstdc++.so.6 | grep
GLIBCXX_
...
GLIBCXX_3.4.19
GLIBCXX_3.4.20
GLIBCXX_DEBUG_MESSAGE_LENGTH(top level comment to hold various suggestions, so upvoting can be used per-suggestion to bubble the best to the top)
i.e. Currently to add a filter you to the bottom, clicking filter, selecting a column name, selecting equals, entering a value.
Instead, right clicking on a cell and selecting "filter > equals" would apply a filter to that column for values matching the selected cell's value. Likewise "filter > contains", "filter > does not equal", "filter > greater than or equal", etc.
These filters would then be appended to the filter at the bottom, so could still be managed there; but just saves some effort when first populating.
- a very important use case for my team in support of business users who are comfortable in Excel which cannot manage large (>1M rows) files. Tad is useful even with 10M+ rows, a user can import, review datafile and filter/pivot to desired subset. The next step for them would be exporting and creating formatted tables or charts for analysis/reporting documents.
Like count, but only counts each distinct value once. Useful for judging data quality (i.e. if you have 600 items with 599 distinct values, chances are there's an invalid duplicate; if you have 1 distinct item chances are that column's not of interest; if you have a few distinct values, you have a potential pivot candidate, etc).
Help mentions INT REAL, and TEXT. Having support for dates would be useful; especially if this enables us to treat dates as multi-part values; e.g. pivot by year & month instead of by the complete value.
I've seen that COUNT is implemented for numeric values; but not for text. There's no reason to limit COUNT to numeric (unlike AVG and SUM).
Definitely still early stage, but that will pass.
I have showed it to a couple people I work with and they have started using it too.
If only there was a tool/service to create such project pages automatically.
Nice work!
Just on that point, Excel allows you to define the field delimiter on import of text files and handles a ton of different formats out of the box, so I doubt this is a point where Tad is better than Excel (though Tad does look promising as a tool to quickly check data files).
It's an Electron app.