HNHacker News
TopNewBestAskShowJobs

mgradowski

104 karma · joined December 26, 2018

https://mgradow.ski
submissionscomments
mgradowski··on The Reason So Many Autistic Adults Can't Stay Employed
Hennecke seems to be just one of many "udarniks" [1], common in countries east of the iron curtain.

There isn't really a lineage between that and contemporary capitalist culture IMO.

[1] https://en.wikipedia.org/wiki/Udarnik

mgradowski··on Many hard LeetCode problems are easy constraint problems
Isn't it trivially [1]?
mgradowski··on AI is different
I jokingly alluded to antirez as HN crowd pars pro toto. I agree it doesn't pass as an intellectually honest argument.
mgradowski··on AI is different
Sorry boss, I'm just tired of the debate itself. It assumes a certain level of optimism, while I'm skeptical that meaningfully productive applications of LLMs etc. will be found once hype settles, let alone ones that will reshape society like agriculture or the steam engine did.
mgradowski··on AI is different
I wouldn't trust a taxi driver's predictions about the future of economics and society, why would I trust some database developer's? Actually, I take that back. I might trust the taxi driver.
mgradowski··on Async Queue – One of my favorite programming interview questions
I'm really confused because I had to scroll half the comment section for the word `semaphore`.

This seems to be an interview question about JS esoterica, not concurrent programming.

mgradowski··on Contra Wirecutter on the IKEA air purifier (2022)
UNIX philosophy is alive and well.
mgradowski··on FF4J – Feature Flags for Java
The "4j" suffix literally means "for Java" and is commonly used to indicate that the project is a Java library, e.g. log4j, slf4j, &c.
mgradowski··on Run SQL on CSV, Parquet, JSON, Arrow, Unix Pipes and Google Sheet
What you're really saying is that the database presented in OP is not useful because it only handles DQL.

1. SQL can be thought of as being composed of several smaller lanuages: DDL, DQL, DML, DCL.

2. columnq-cli is only a CLI to a query engine, not a database. As such, it only supports DQL by design.

3. I have the impression that outside of data engineering/DBA, people are rarely taught the distinction between OLTP and OLAP workloads [1]. The latter often utilizes immutable data structures (e.g. columnar storage with column compression), or provides limited DML support, see e.g. the limitations of the DELETE statement in ClickHouse [2], or the list of supported DML statements in Amazon Athena [3]. My point -- as much as this tool is useless for transactional workloads, it is perfectly capable of some analytical workloads.

[1] Opinion, not a fact.

[2] https://clickhouse.com/docs/en/sql-reference/statements/dele...

[3] https://docs.aws.amazon.com/athena/latest/ug/functions-opera...

mgradowski··on Checkmake: Experimental Linter/Analyzer for Makefiles
Excellent name.
mgradowski··on Ask HN: Tools to visualize data in SQL databases?
I'm curious why do you think of DAX as a virtue. My poor SQL-shaped peg brain has never really fit the DAX hole of MS software.

Also, it always struck me as something too complex for the non-technical folks, and not expressive enough for tech-literate analysts/data engineers &c.

mgradowski··on An Infinitely Large Napkin [pdf]
This book exposed me to proper math back in high school, I liked it.
mgradowski··on What would it take to recreate dplyr in Python? (2020)
Huh, even though I would prefer a universal SQL layer, ibis looks quite nice.
mgradowski··on What would it take to recreate dplyr in Python? (2020)
TIL there's an actual 'botched' library with the same name; I actually came up with it independently on a lazy office afternoon :^)
mgradowski··on What would it take to recreate dplyr in Python? (2020)
That's what I do at $dayjob whenever I have to do windowing &c. Figuring out this stuff in Pandas is a waste of time. Before I discovered DuckDB, I would re-learn the API every damn time. I came up with a little utility function, which you can implement yourself :)

``` def sqldf(df: DataFrame, query: str) -> DataFrame: ... ```

mgradowski··on What would it take to recreate dplyr in Python? (2020)
I like interface-only packages in the Julia ecosystem e.g. Tables.jl enables the development of several packages for querying tabular data that work across many concrete implementations; Plots.jl separates the high-level plotting interface from the plotting backend.
mgradowski··on What would it take to recreate dplyr in Python? (2020)
DuckDB and Polars are my bets in the Python data-wrangling space. I grew tired of Pandas' weird-ass API.
mgradowski··on macOS Setup after 15 Years of Linux
I have found this setting to break some apps (e.g. Spotify, my bank's app) though.
mgradowski··on Better SQL JOINs
Neither did I. I tend to avoid RBDMS-specific syntax extensions of questionable benefit.

The only thing I really miss in PG is ASOF JOIN. It's possible to implement it via CROSS JOIN LATERAL, but in my case it resulted in nested loops — coming from ClickHouse it bothers me that there is no way (that I know of) to hint PG to use a specific join strategy.

mgradowski··on Better SQL JOINs
Related — Postgres (other RDBMS's too, probably) has JOIN b USING (column, ...) and NATURAL JOIN. IIRC, both fail when there's ambiguity in column names.
mgradowski··on Fast CSV Processing with SIMD
Nice one!
mgradowski··on Fast CSV Processing with SIMD
Also, you will sleep better at night knowing that your column dtypes are safe from harm, exactly as you stored them. Moving from CSV (or god forbid, .xlsx) has been such a quality of life improvement.

One thing I miss though is how easy it is to inspect .csv and .xlsx. I kinda solved it using [1], but it only works on Windows. More portable recommendations welcome!

[1] https://github.com/mukunku/ParquetViewer

mgradowski··on SSH Tunneling Explained
I use WireGuard with a server in my pantry as a router. Dynamic IP of the server is handled by DuckDNS, and WireGuard gracefully handles client roaming e.g. I can switch from home wifi to mobile internet without interrupting my SSH sessions. Would recommend.
mgradowski··on Ripgrep 13.0
If you like fzf, give skim [1] a go. It's basically faster (anecdotally) fzf, but can also be used as a Rust library. I made a little personal to-do manager using it, together with SQLite for persistence. I also improved my fish shell ergonomics with [2].

[1] https://github.com/lotabout/skim [2] https://github.com/zenofile/fish-skim

mgradowski··on Ask HN: Show me some most beautiful fonts
Source Serif Pro is my Google Docs workhorse .
mgradowski··on Ask HN: Dotfiles Management Tools?
I use GNU Stow with a dotfiles git repo placed in $HOME. Works well.
mgradowski··on Functorio
Dependent types could presumably capture this idea. They are very expressive – see [1] for a perfectly type-safe printf implementation in Idris.

[1]: https://gist.github.com/chrisdone/672efcd784528b7d0b7e17ad9c...

mgradowski··on Imgdiff: Faster than the fastest pixel-by-pixel image difference tool
Nice spelunking! If there is hope of avoiding C++ altogether, then I'll try testing it again on something beefier than my laptop. Thanks for taking the time and effort.
mgradowski··on Imgdiff: Faster than the fastest pixel-by-pixel image difference tool
That's a nice and tidy codebase, real pleasure to read.

> I believe many (most?) opencv & numpy operations release the GIL.

Any idea how I can determine this? I am prototyping a real time machine vision application targeting 2x720p@240fps and I want to avoid writing any C++ for as long as possible.

mgradowski··on Imgdiff: Faster than the fastest pixel-by-pixel image difference tool
Regarding [1], many OpenCV functions in Python support an optional dst argument, similar to how they do in the C++ API. This makes the memory management situation not completely hopeless.

In my experience, the only drawback of the Python API is the fact that it cannot utilize the multithreading module (probably does not release the GIL for long calls).

Page 1 of 2Next →