HNHacker News
TopNewBestAskShowJobs

marsupialtail_2

264 karma · joined October 29, 2022

submissionscomments
marsupialtail_2··on Butterflies flew 2,600 miles across the Atlantic without stopping
Butterfly flies 2600 miles across the ocean just to be caught by a human in a jar to be DNA sequenced ...
marsupialtail_2··on The race to replace Redis
The sincerest form of flattery is when AWS decides to come up with a big consortium to displace you with some open source.

Incidentally the most effectively way to stall a project according to the CIA is to have a huge guiding committee with clearly diverging interests.

Redis will win because it's focused on its users. It's competitors will lose. Like OpenSearch, like OpenCL etc.

marsupialtail_2··on Wordle but with Emojis
code: https://github.com/Vince7778/Emojile
marsupialtail_2··on Show HN: Quickwit – OSS Alternative to Elasticsearch, Splunk, Datadog
No just curious. I understand how your indexing structure based on SSTables could find it challenging to support substring search in general. I think it tradeoff between fast querying and flexible functionality
marsupialtail_2··on Show HN: Quickwit – OSS Alternative to Elasticsearch, Splunk, Datadog
Glad this is getting some love. This is seriously good software. Have you guys supported generic substring search yet? I recall it was not supported as of a few months ago.
marsupialtail_2··on How Query Engines Work
Thanks for the shoutout!
marsupialtail_2··on The Carrot Problem
Yes but there is also the inverse carrot problem. E.g. if the pilots have radar, they are more liable to rely on it and neglect other aspects of flying. Similarly in business, it is simply harder for folks who grew up rich to develop the level of grit that comes natural to the less privileged.

I may sound like a rich apologist, but please believe me when I say it is harder to spend 10 hrs a day cranking on a risky startup if you know you can be clubbing with daddy's money

marsupialtail_2··on Case study: Algorithmic trading with Go
Hi Justin, you might be interested in my blog: https://github.com/marsupialtail/quokka/blob/master/blog/bac... advocating a cloud based approach.

You don't have to use the system I am building, but it's worth thinking about that design.

marsupialtail_2··on The Accidental HFT Firm (2018)
Perhaps I charged too little when I contracted away my 10x random forest inference solution...
marsupialtail_2··on Daft: A High-Performance Distributed Dataframe Library for Multimodal Data
SQL support is very challenging.

I work on Quokka (https://github.com/marsupialtail/quokka). I support Iceberg reads. Recently we are adding SQL support from just parsing the DuckDB logical plan, though that is very challenging as well.

The Python world lacks a standard for a plug and play SQL query optimizer. Apache Calcite is good for the JVM world, but not great if you are trying to cut out the JVM.

marsupialtail_2··on Daft: A High-Performance Distributed Dataframe Library for Multimodal Data
While we are on this topic, the challenge with data lakes for Python based projects like Daft and Quokka (what I work on) is the poor Python support for data lakes like Delta, Iceberg and Hudi. Delta has the best support but its Python API is consistently behind the Java ones. Iceberg doesn't support Python writes. Hudi doesn't support anything Python.

I have users demanding Iceberg writes and Hudi reads/writes. I don't know what to tell them, since I don't have the resources to add a reader/writer myself for those projects.

Hopefully as DuckDB becomes more popular we will see Python bindings for these popular data lake formats this year.

marsupialtail_2··on Daft: A High-Performance Distributed Dataframe Library for Multimodal Data
Hi -- I am the author of Quokka: https://github.com/marsupialtail/quokka, trying to be distributed Polars. I am trying to go for API compatibility, or at least supporting most of the API.

I am not focused on complex data types though.

marsupialtail_2··on Why your dataframe library needs to understand vector embeddings
Open for rebuttals from vector database vendors, especially this one: https://github.com/jdagdelen/hyperDB
marsupialtail_2··on The simple joys of scaling up
You will always be limited by network throughput. Sure that wire is getting bigger but so is your data
marsupialtail_2··on The Inner Workings of Distributed Databases
In case people are interested, I wrote a post about fault tolerance strategies of data systems like Spark and Flink: https://github.com/marsupialtail/quokka/blob/master/blog/fau...

The key difference here is that these systems don't store data, so fault tolerance means recovering within a query instead of not losing data.

marsupialtail_2··on Show HN: Quokka -- Distributed Polars on Ray
I hope in a good way
marsupialtail_2··on Deep Dive into Neon storage engine that enables serverless Postgres
Can you comment on main differences between this and Aurora?
marsupialtail_2··on Show HN: Quokka -- Distributed Polars on Ray
It's more about the API -- currently the API mimics Polars. I actually use DuckDB for a lot of the computation.
marsupialtail_2··on Launch HN: DAGWorks – ML platform for data science teams
would love to collaborate on an integration with pyquokka (https://github.com/marsupialtail/quokka) once I put out a stable release end of this month :-)
marsupialtail_2··on The Lone Developer Problem
If you work on a project long enough eventually you forget about how parts of your project work, and this automatically happens
marsupialtail_2··on CircleCI says hackers stole encryption keys and customers’ source code
If you make everything open source...
marsupialtail_2··on Apache Hudi vs. Delta Lake vs. Apache Iceberg Lakehouse Feature Comparison
me too. Trino for one would be a good start. Adding support for those data lakes is really hard though if you want good performance.
marsupialtail_2··on Apache Hudi vs. Delta Lake vs. Apache Iceberg Lakehouse Feature Comparison
I think the blog post should point out very early that Onehouse is a Hudi company. There are some other recent benchmarks published in CIDR by Databricks that might paint a different picture: https://petereliaskraft.net/res/cidr_lakehouse.pdf
marsupialtail_2··on Writing a Python SQL engine from scratch
OK I'll admit: https://news.ycombinator.com/item?id=34189422 is not a real pure Python SQL engine, this one is.
marsupialtail_2··on Databases in 2022: A Year in Review
Just did.
marsupialtail_2··on Databases in 2022: A Year in Review
I am writing a Python based SQL query engine: https://github.com/marsupialtail/quokka. My personal goal is to get Andy Pavlo to mention it in his year-end blogs.

I agree with many of the points made in the blog by Andy. Writing a distributed database has become way easier due to open source components like Ray, Arrow, Velox, DuckDB, SQLGlot etc.

I personally believe we will see a switch from JVM based technologies to Rust/C based with Python wrapper

marsupialtail_2··on I wrote a SQL engine in Python
Wait this is really cool. How do you parse and optimize SQL?
marsupialtail_2··on I am not a supplier
I sponsor OSS software I depend on and you should too!
marsupialtail_2··on I wrote a SQL engine in Python
Let's make this top-rated comment :-P
marsupialtail_2··on I wrote a SQL engine in Python
That's right. My background is mostly in quantitative finance, where we would use models like linear regression on expert-engineered features based on market data, instead of throwing a deep neural network at raw price data like what some people might imagine.

For push vs. pull, I'd recommend: https://news.ycombinator.com/item?id=27006476.

On single machine, you really should just use Polars. Quokka is faster than Pandas because it can take advantage of multiple cores, but so can Polars -- and it is likely to be faster.

Page 1 of 2Next →