Client libraries are better when they have no API?
csvbase.com
csvbase.com
Also, while CSV is a nice/convenient data format for a data analytics use case (like this), it’s certainly not a format I’d choose for an API where clients are likely to be more standard CRUD-ish apps. JSON is great for those, CSVs (with their trickier parsing, “everything is a string” data types, and enforced flatness) are a pain in the ass.
I did think it was interesting to learn about fsspec, didn’t know about that! And this style of client library does seem like a good/convenient one for this specific Python data analysis use case. Enjoyed that part of the article, but had to wade through a fair bit of clickbait style writing to get there.
To understand how awful it is, you just have to ask two questions "How do I represent empty values?" "What do I do when my separator is contained in the data?"
There are answers to these questions, but there's not a standard backing up those answers. And that's what makes CSV is PITA to deal with. It's such a loose "standard" that anything goes.
xml and json are FAR better options even when you just want a table of data.
https://en.wikipedia.org/wiki/CDATA is something far too few people pay attention to when considering the complications of object notation and why markup languages have some escape hatches.
,,
And yeah, dealing with separators is annoying, but pretty much every reader supports quoted values and escapes (not every writer cares up add them, though).
In practice I use csv to process large amount of tabular data (like logs, events, etc) where I care about greppability and performance more than about potential for 0.001% of corrupted data. YMMV of course, if getting it 100% right is important use something else (JSON is not without sins too, consider that JSON numbers are often parsed as floats).
Or is it ,,,,,,,
Or is it tabtabtabtab
Or perhaps it's "NULL,NULL,NULL"
My part of my work is data ingestion and I've seen all these (and more) as answers to the "empty values" question.
I'm not saying that other formats aren't without their problems, they certainly are. However, CSV doesn't just have those problems, it has multiple other problems on top of them.
It's a basic idea with really obvious edge cases addressed in multiple ways depending on who is producing these documents.
Those are rookie questions. What about when the data contains newline/crlf? What if the data contains the quote character and newline?
And why is the file mostly Windows-1252 encoded except some fields that are sometimes, at random, UTF-8 encoded?
The real kicker is when your fellow users are opening the CSV in a spreadsheet with a locale that prefers commas for fractional currency amount and then saving the file back.
Empty string. For example, here are three in a row:
,,
> What do I do when my separator is contained in my values?Use quotes. For example, here is a single comma:
","I'm sorry about the title. As I said below; that was really meant in the spirit of fun.
On the subject of csvbase's content negotiation - yes that is an API. That was covered here some time ago when I wrote about it before: https://news.ycombinator.com/item?id=37526047
The "no API" bit I'm talking about in this article is basically the "trick" (or whatever word you want to use) of avoiding having any user-facing interface and just hooking into stuff that is already there. There is no "API surface" here for the user to learn beyond a url scheme. I think that's nice. And it's mainly what I'm talking about.
> Also, while CSV is a nice/convenient data format for a data analytics use case (like this), it’s certainly not a format I’d choose for an API where clients are likely to be more standard CRUD-ish apps. JSON is great for those, CSVs (with their trickier parsing, “everything is a string” data types, and enforced flatness) are a pain in the ass.
Without wanting to sound too much like a sales pitch: csvbase does offer JSON. Try https://csvbase.com/calpaterson/opcodes-6502.jsonl for JSON lines (or https://csvbase.com/calpaterson/opcodes-6502.json (no 'l') for a paged plain-JSON interface).
I personally think there is no ideal format for this at the moment. JSON is very very large and slow to parse. CSV has well known problems though has massive compatibility and often works well in practice. Parquet is probably closest to the ideal and excellent in many respects but is quite complicated to parse (moreso than CSV? perhaps) and anyway is effectively unstreamable - actually quite annoying for something like csvbase where you really _don't_ want to materialise the dataset while serving it.
> I did think it was interesting to learn about fsspec, didn’t know about that!
Yes it is cool isn't it. Millions of downloads, terabytes of bandwidth of PyPI and no one has heard of it.
There's no need to apologize! I and many others understood exactly what you meant by "no API". Some other people didn't understand the distinction you were drawing and chose to interpret their lack of understanding as you misleading them somehow, but that's on them, not you.
It was a great article that I thoroughly enjoyed! Thanks for sharing!
The python library has no python API, serving instead just as a plugin for fsspec. You then use pandas or anything else the same as you did before, the python library adds no (python) API. i guess technically "use the custom `csvbase://` scheme in the URIs you supply as input, instead of `http`" could be called an "API" if you really want to play gotcha.
I think the point is legit -- the python programmer has to reference nothing specific to the client library here other than the custom URI scheme, to then use remote data from the site with any one of several existing python data libraries, via their own python APIs.
Wikipedia: API stands for Application Programming Interface. In the context of APIs, the word Application refers to any software with a distinct function. Interface can be thought of as a contract of service between two applications.
This is an interface between two applications.
I do like the simplicity of the APIs provided though, but they _are_ APIs
- Where there is an "API key" and a "secret key" when in reality the idiot maintainers could have just concatenated the two and called it a "key". The customer doesn't need to know the abstraction on the server side
- Where one needs to create an "account" and then a "project" before one can even create a "key". WTF is a project
- Where one needs to do a dumb SMS 2FA to even get an API key
- Where one needs to do multiple steps and try/except blocks to get a result in a plain format
I really wish APIs were as just having a free tier that you can just use, and if you want more, just POST $10 worth of Solana to some endpoint and the API gives your IP 10000 more requests. No accounts, no keys, just simple.
I'm not a cryptocurrency fanatic but making APIs easier to access is one really good use case for it. It makes paying for an API as easy as paying for a lemonade at a street stand with cash. No accounts, no billing addresses, no subscriptions, just POST over a nickel along with your API request and get a response, and API owner gets paid.
IP-based requests seems very restrictive. I'd rather post some sort of payment to an endpoint then get some arbitrary secret back that I could use/save/distribute. Maybe you could choose the mode you wanted though, so if you were confident your IP would remain stable and not be shared with anyone else, you could choose that mode.
I just read this article and have very conflicting feelings about it. It is clever, and the nice kind of clever that does not require one to be a mega-brain. On the other hand, it creates invisible and uncontrollable dependencies, such as the one you describe.
Another drawback: something is broken and I want to set a debug point in the code that fetches the CSV data. Unless you know about fsspec it will be hard to follow the breadcrumbs to know how this library injects itself into your code.
But I guess that's the reality of making libraries, even bugs will be relied on, we really don't have any methods for evolving software ecosystems reliably and compatibly.
https://filesystem-spec.readthedocs.io/en/latest/developer.h...
https://setuptools.pypa.io/en/latest/userguide/entry_point.h...
(As a side note here, you might think that "pip install" can only add functionality, and that you can add extra packages locally for your convenience, or keep them around when you switch branches, without having your dev environment diverge from production behavior. But if they register plugins with other packages via the entry point functionality, you might end up coding things that depend on behavior that's not present in other environments! This is especially common in the pytest ecosystem. CI is vital here!)
I apologise for the title. Not intended to mislead, just to be a bit of fun :)
> t's really cool to learn about how fsspec [...]
I'm glad you took something from it. At the last place I worked I spent a lot of time building a library to persist dataframes (intended audience: data analysts) and in retrospect I wish we'd thought of the approach of just having a `csvbase://` url scheme via fsspec.
Currently it doesn't support Parquet which they would have needed though. Still some work to do for csvbase to accept Parquet uploads.
Reading through the article, though, the author isn't advocating for "no APIs," they're advocating for minimalist APIs that use agent detection to "do the right thing."
Chrome tells the API that it can accept HTML, so the server sends data formatted inside of a web page, with an HTML table.
Curl doesn't send that header, so the server sends unformatted CSV. But you could send an Accept header to get the HTML if you wanted.
The benefit of this is that for most use cases, this will "just work."
The downside is, if you want to view the CSV in a browser, or the web page in curl, you need to know (or guess) how the server is deciding what to send you, and take the correct action. An API documented with OpenAPI, while more complicated, explicitly tells you what you can do.
I've long held that software ages far better if you just eliminate the ability of people to provide it input.
That's what the author is saying they don't have. When you import their client library you just have a new URI scheme you can use anywhere that accepts a URI, not layers of extra classes and methods to learn how to interact with.
http://csvbase.example/username/dataset/table
instead of creating a whole protocol handler just for an online service? import pandas as pd; opcodes_6502 = pd.read_csv('https://csvbase.com/calpaterson/opcodes-6502', index_col=0)
But writing back is harder. It's not easy to make pandas do an HTTP PUT and then insert the HTTP basic auth and so on. Plus (not discussed in the article) there is a cache to avoid redownloading the data when it hasn't changed, which is my personal bete noire in "data science" such as it is.But this must be a management vision and effort, if you are always pressed just to deliver something working, you will end up with lots of people reinventing the wheel internally.
The Rack protocol did something similar for http servers in the Ruby world, allowing for a number of middleware. Whole frameworks (Rails, Sinatra, as examples) could be mounted on specific routes.
Sorry, I was the 800th, your post is now outdated!
In every instance I have seen, the free "web API" involves an extra HTTP header(s) that, in lieu of a common one such as "User-Agent", can be used to track, rate-limit and/or selectively block a www user.
The upside of the "web API" idea IMO is the serving of public information in formats other than HTML or PDF. It's great.
But why not just do this without using the extra HTTP header(s), tracking and limitations.
adaptors - someone else's APIs
It sounds like an API and it looks like an API, but don’t let that fool you…
that's still an API
PS: Thanks for all of your work over the years (decades?) on sqlalchemy. I'm sure you're not too surprised to learn that the lions share of code in csvbase is calls to SQLA core.
Really! a bit surprised sure, wasn't really sure what the tool was actually doing to persist data :)
It’s offloaded the work to an adapter.
Please don’t do that.
Edit: Explicit patching would be just fine, like:
from my_project import csvbase_patch
csvbase_patch()
p = pandas.pd(“csvbase://…”)
Then there’s a giant indicator inside the same module that there’s magic happening. If I cmd-f “csvbase”, the magic string in the URL, I’ll stumble across that patch function. Then I can jump to its definition to see what’s happening.This is Python. Hidden monkeypatching isn’t how we do things.
People that write code like this should be forced into bug jail like i was for a year.
But the author is not doing any monkeypatching or changing how pandas works. He is using fsspec [0] to create a filesystem interface that pulls data from his site. fsspec appears to be somewhat standard since pandas, polars, Dask, and other libraries use it.
As soon as I understand that fsspec exists and this library uses it, there is no more magic. I would prefer if the specification of which fsspec to use were not embedded as part of a string, but overall this approach seems pretty reasonable.
The idea is nice! I’m not saying fsspec is bad. It’s not. It’s neat! I just strongly abhor the idea that pip installing a package changes runtime behavior whether that packages is ever even imported.
```
from csvbase import CSVBASE_FS
pd.read_csv('//calpaterson/onion-vox-pops', fsspec=CSVBASE_FS)
```
At 2PM when I’m well rested, caffeinated, and alert, the original code is clever and lovely. At 2AM, I want this kind of easy to understand explicitness.
```
fs = fsspec.implementations.local.LocalFileSystem()
with fs.open('test.csv', mode='r') as fp:
print(pd.read_csv(StringIO(fp.read())))
```But, for simple use cases, what he is doing beats building yet another client library or defining REST or RPC endpoints.
The resource is represented by the URL, no?
What else is needed for this to be called a "REST API"?
The author didn't make this pattern up, it's how fsspec officially recommends implementing backends [0]. Given how widely used fsspec is I think it's fair to say that hidden monkeypatching is how we do things... sometimes.
[0] https://filesystem-spec.readthedocs.io/en/latest/developer.h...
If foo 1.1 takes a dep on Bar 1.2, and baz 2.3 takes a dep on Bar 1.4, both Bar contexts are tested with their respective libraries but forcing Bar to a particular global version can have problems and both versions of Bar are needed to have tested behavior.
Examples mentioned and OP change the global behavior vs proper dependency management.
IIRC (in JDBC) you also used to have to do `Class.forName("name.of.it")` somewhere before trying to do any DB access, to ensure that the static initializers had actually run, but I don't believe it's necessary anymore
(And then of course you have Spring Boot autoconfiguring which is another level of magic up, using automatic subclassing and proxy injection to add things like transaction management. And then you can get into proper classloader hackery)
The "data source name" string when connecting is... basically a JDBC connection string, and some adapters use exactly that iirc, but it's fundamentally an unstructured string that just serves the same purpose. Plugins can use anything they like, and style varies.
"hidden monkeypatching" is essentially how all of Java works.
If so, that’s another reason I’m uninterested in Java. Magic behavior changes based on the prefix of a string passed into a function doesn’t appeal to me at all. The Python equivalent for a database might look like (typed from memory):
from psycopg import connect
conn = connect(host, username, password, dbname)
Where connect() returns a database adapter object with DB-API methods. When something breaks, I can see which DB code is involved. It’s right there in the import. My IDE doesn’t have to parse strings to know which code path subsequent method calls are following.Not fine. Just less horrible.
The original authors of several codebases I've been saddled with over the years would disagree with you.
This is Python and you a peasant. Yes most of the time the snake along the road is a snake, but occasionally it's a Basilisk waiting to destroy you with it's gaze. Knowing this, you the weary traveler, carry inconvenient tools such as mirrors to protect yourself and are well versed in superstition and arcane dark arts. You never let your guard down, because yes there is magic afoot.
"Beautiful is better than ugly."
I think this is ugly. It's visually appealing, but my editor doesn't know that "csvbase" is a magic symbol that selects the code path to follow. That deep ugliness outweighs any skin-deep cleverness.
"Explicit is better than implicit."
That sums it up. This is implicit. `pip install foo` changing runtime behavior of packages that never import foo is as implicit as it gets. From a developer's and a security engineer's point of view, I loathe that code is running simply because it exists.
It's possible to do these things. I contend that it's not Pythonic, and we should not be doing them.
It kind of is, though. This isn’t the first time I’ve seen this; off the top of my head, I believe the Stackdriver libraries do something similar but to the stdlib logging library.
The fact that it’s possible is also a tacit endorsement of doing things this way. Many languages just don’t permit adding to or modifying a namespace.
I would argue monkey-patching is a bad pattern in general, regardless of whether it’s manual or automatic.
IMO, fsspec should have a private package-level variable with a list of these adapters, and expose a ‘fsspec.register_adapter(adapter)’ method to add things to it. I don’t see a need or use in patching here.
It just seems fraught with issues that could be avoided with an adapter registry. Testing seems easier too; I can’t imagine the pain of unit testing a bunch of adapters if they each try to modify the fsspec package directly.
The author is just adding a new backend for a common protocol in the data world. Pandas, Polars, Dask, DuckDB, etc. all support this protocol and the type of people who want to access a dedicated CSV data archive would probably much rather keep their current client API and just add a new connection URI string vs. adding an entire client API for just one data source (or dealing with making requests and passing the data into the dataframe).
There's no need whatsoever for a separate client API. There could be a convention like:
from csvbase import loader as csvloader
df = pd.read_csv(csvloader("calpaterson/onion-vox-pops"))
The user wouldn't have to know anything but what to import to fetch a certain thing, and it's explicit about what's coming from there. There's also less risk of mistyping the URL string ("oops, I just typed cvsbase and accidentally loaded a list of CVS drugstores"), and code completion can tell you that csvloader() fetches things through the csvbase module.Cons:
- It takes 5 seconds longer to type, one time.
Pros:
- You can tell what the code does at a glance.
When I see `pd.read_csv("csvbase://")` during debug, I wonder how pandas knows to speak to csvbase (as the article anticipates). Nothing is imported. Nothing is configured. Things just speak to one another. So, can I also call pd.read_csv("other_csv_server://") like this? When I replace pandas with koalas, will koalas.read_csv("csvbase://") also work? How the wires connect between pandas and csvbase is hidden. Unless you know that the two are obeying some implicit lower layer (the fsspec standard), this becomes a mystery. Mysteries are the last thing you want when debugging.
I don't know which `create_engine()` function you're alluding to. The one I know and have used comes from SQLAlchemy. How it works has always been obvious. I've never seen any mention of fsspec. I looked at its code and it's predictably just a convenient syntax to specify connection information in a single string. The string is simply parsed to extract connection attributes, which are then relayed to the lower DBAPI. There's no mystery involved.
I do take your point - to an extent.
I think the subject you're touching on is basically that of configuration. Configuration is what allows you to change the behaviour of code without making a code change. Configuration is of course very really powerful for good and evil. Someone below mentions ODBC which, yes, 100% is a great example. But the one that sticks in my mind is resolv.conf.
I would contend (:)) that csvbase-client is not doing "fiddly magic at package installation time" but supplying configuration for a url scheme. In much the same way that installing boto3 does for using the `s3://` url scheme.
"Magic" in my book would be monkeypatching the methods of other libs. I would hate to write `csvbase_patch` and have it execute on the import of a specific module. Like you, I prefer to live in a low-magic fantasy world.
This kind of configuration feels very magical in that it's completely behind the scenes. There's nothing in code I can look at that says "this URL should be handed to this adapter". There's no env var that gets slurped in. No .env file that's parsed. There's nothing checked into git that tells me how that URL scheme could possibly work, except the existence of a package in pyproject.toml -- one that's not even imported, just present.
That makes it stateful in a bad way. If `poetry install` (or `pip install -r requirements.txt` or whatever) didn't complete, the module using it will still load and run up until the point that it crashes because that URL doesn't load. If I ran an automated process to prune unused modules from poetry/pip, suddenly behavior changes. I don't like that idea one bit.