It's 2024 and you're not developing the next vim or Postgres. Use bash.
3,497 karma · joined August 13, 2015
tselai.com
It's 2024 and you're not developing the next vim or Postgres. Use bash.
You can run a `select mic(col1, col2) from my_table` between any arbitrary pair of columns to see if there's a strong correlation between the two. Postgres already provides corr(X, Y), but that works only for linear correlations.
The usual disclaimer applies: correlation does not imply causation etc.
For context, there are a few efforts in progress. I'm sure there are more. - https://pgxman.com - https://pgt.dev - pgxn v2: https://github.com/orgs/pgxn/projects/1/views/1
This has been discussed in the past [1], but the Postgres tooling ecosystem has been primarily C-Makefile—mailing list driven, and there used to be a lot of Makefile targets copy-pasting. Whenever major Postgres providers wanted to open source some of their extensions / sub-products. I still feel, however, that a lot of Postgres C-know-how is being slowly forgotten / lost, and I think it will be necessary again soon. Internal things as how Postgres handles varlena, StringInfo, JsonbValue, etc. The core abstractions that make Postgres work.
0: https://github.com/omnigres/omnigres/blob/master/Dockerfile 1: https://redmonk.com/jgovernor/2023/10/10/postgres-the-next-g...
I've found it more helpful to have a PREFIX=$(pwd)/build or even PREFIX=$HOME/.local as an installation destination. And have a .env add PATH=$PREFIX/bin:PATH
I've found it more helpful to have a PREFIX=$(pwd)/build or even PREFIX=$HOME/.local as an installation destination. And have a .env add PATH=$PREFIX/bin:PATH
Now, can jq be used for SQL injection in an SQLite context? That's interesting in theory, but I'd assume any decent driver would validate its input as data and not code.
PRQL and EdgeQL (EdgeDB) are the most interesting ones to watch how they evolve, though.
I've also written a PG extension to make jq available in Postgres [0]
I believe Postgres, in general, will flourish as a host for DSL languages [1].
0: https://github.com/Florents-Tselai/pgJQ 1: https://tselai.com/pgjq-dsl-database.html
https://en.m.wikipedia.org/wiki/John_MacFarlane_(philosopher...
pgJQ embeds the standard jq compiler and brings the much-loved jq lang to Postgres.
It adds a jqprog data type to express jq programs and a jq(jsonb, jqprog) function to execute them on jsonb objects. It works seamlessly with standard jsonb functions, operators, and jsonpath.
And there's this vtable for Parquet extension. https://github.com/cldellow/sqlite-parquet-vtable
But for my use case virtual would be too complicated.
The .warc spec is ideal. I'm not saying we replace it (ref. xkcd: standards).
On top of what uniqueid said, "loading" is much slower and more cumbersome than it sounds. I'm not saying SQLite will replace text (maybe my aphorism sounded too firm). I'm saying that maybe along with the .warc.gz archives at rest, one could have .sql.gz files at rest as well.
In other words: why not move ACID-compliant archives moving around?
Now with the new protocols, dunno maybe its too soon to worry? Then again, maybe its an IPv4 / IPv6 analogy.
> There is a companion format, CDX, which builds indexes from warc files (which in turn are just concatenated records, so rather robust).
Good point. I's planning of combining this fact with the ATTACH option that SQLite has - allowing to query multiple database files [0]
The WARC format is extremely simple and yet so powerful. Most importantly though, there are already pebibytes of already crawled archives.
This is a fairly straightforward mapping of a .warc file to a .sqlite database. The goal is to make such archives SQL-able even in smaller pieces.
The schema I've come up with it's tailored around my requirements, but comment if you can spot any obvious pitfalls.
PS: I do believe that at some point .sqlite will become the defacto standard for such initiatives. Sure, it's not text... but it's pretty close.
There's an "awesome web-scraping" repo with ~5K stars.
A few days ago the maintainer added the Russian flag at the top of README.md
Someone raised this as an issue [0]
"Just a suggestion, having a Russian flag at the top of the page might not be the best idea?" to which the maintainer simply replied
"Why not? I am Russian citizen and I support my country."
Seems trivial, but got me thinking of the implications. E.g. how would a Ukrainian dev feel having his project showcased in this list, under the Russian flag?
[0] https://github.com/lorien/awesome-web-scraping/issues/136
can you elaborate on these patches? Have you worked on any parser-level hooks? I'm building an extension on PG, and have been trying to extend the parser without touching PG's tree, but it looks impossible at this point.
My social circle regularly posts about such issues.
It turns out this makes Facebook think that we're interested in yachting.
Being a data professional myself it makes me frustrated to see such false positives and I'm surprised this special case hasn't been flagged somehow.
[1] https://www.facebook.com/633257566/posts/10159522380347567
[2] https://www.aljazeera.com/news/2021/12/25/16-migrants-dead-i...
Good point and it's a potential trap of tools built from developers for developers.
There's this fine line between finding it useful and being willing to pay for it. Being honest to my self: Do I find this useful ? Yes Makes my web scraping easier? Yes Would I pay for it? Maybe... but the problem is I've never paid for a scraping tool either.
But other people may pay for it. To get to that point though I'll probably need a few more hands on deck, who need salaries and for that you need external funds.
Good point though. Bootstrapping is a viable option too, but fundraising has certain no non-financial benefits that can be appealing.
Hedge funds actually call this "alternative data".
Wasn't aware of the Ticketmaster v. Tickets.com case. Will have to print it along the SCOTUS LinkedIn v. hiQ ruling.
:-)