HNHacker News
TopNewBestAskShowJobs

simicd

61 karma · joined April 21, 2022

submissionscomments
simicd··on Microsoft team creates data-storage system that lasts for millennia
At 4.8TB one could add a header section with the full code, instructions how to compile it etc. That would certainly help to reproduce it, assuming civilizations in 10k years still can decypher todays language.
simicd··on The Business of Betting on Catastrophe
It's correct that the number of reinsurers is smaller than that of primary insurers. But the risk born by reinsurers is less correlated, not more. Any given primary insurer has risk clusters (domestic market, line of business, etc.). If a large catastrophe happens in their domestic market they might go bust but what are the chances that it happens simultaneously to all markets globally?

Say you're a primary home insurer in the US. If a hurricane hits you might not have enough capital to rebuild all the homes. A reinsurer which is also covering Europe, Asia, LatAm, etc. is less likely to go bankrupt. The reinsurer can cross-subsidize and use the insurance premiums from other regions to pay out the claims from the US market. All that matters is that on average the loss probabilities and severities are estimated correctly.

And this is just using one line of business as example, reinsurers are covering property, casualty, life and health which add extra layers of diversification.

simicd··on GitHub Copilot: The Agent Awakens
Since it is a fork of VS Code you can install any VS Code extension in Cursor (although manually): https://www.cursor.com/how-to-install-extension
simicd··on Should you ditch Spark for DuckDB or Polars?
From what I understood the article refers to the point that DuckDB doesn't provide its own dataframe API, meaning a way to express SQL queries in Python classes/functions.

The link you shared shows how DuckDB can run SQL queries on a pandas dataframe (e.g. `duckdb.query("<SQL query>")`. The SQL query in this case is a string. A dataframe API would allow you to write it completely in Python. An example for this would be polars dataframes (`df.select(pl.col("...").alias("...")).filter(pl.col("...") > x)`).

Dataframe APIs benefit from autocompletion, error handling, syntax highlighting, etc. that the SQL strings wouldn't. Please let me know if I missed something from the blog post you linked!

simicd··on Sora is here
I suspect so, OpenAI is subject to the EU AI Act [0]. Last time they released the Advanced Voice Mode it also took some time before it became available in the EU. Not sure why UK and Switzerland are delayed as well, they are not in the European Union.

[0] https://openai.com/global-affairs/a-primer-on-the-eu-ai-act/

simicd··on Sora is here
Hmm enough capacity for the rest of the world but not EU:

https://help.openai.com/en/articles/10250692-sora-supported-...

simicd··on Sora is here
And in the UK and Switzerland unfortunately

https://help.openai.com/en/articles/10250692-sora-supported-...

simicd··on Non-elementary group-by aggregations in Polars vs pandas
I'm using both Spark and polars, to me the appeal of polars is additionally it is also much faster and easier to set up.

Spark is great if you have large datasets since you can easily scale as you said. But if the dataset is small-ish (<50 million rows) you hit a lower bound in Spark in terms of how fast the job can run. Even if the job is super simple it take 1-2 minutes. Polars on the other hand is almost instantaneous (< 1 second). Doesn't sound like much but to me makes a huge difference when iterating on solutions.

simicd··on Announcing Polars 1.0 (Blog Post)
Yes only found the announcement [1] that the Polars team and NVIDIA engineers are working on a GPU engine, but other than that no concrete examples. Github issues also don't provide any hints on the status, only one open item [2] where most comments are prior to the announcement.

[1] https://pola.rs/posts/polars-on-gpu/

[2] https://github.com/pola-rs/polars/issues/13111

simicd··on Htmd: A turndown.js inspired HTML-to-Markdown converter for Rust
To add to that, an additional benefit would be you can compile and release it as Python package (Py03/maturin) or compile to WASM so it runs in the browser (with javascript bindings). This makes the code portable while benefiting from Rust's performance/memory safety.
simicd··on DataFusion Comet: Apache Spark Accelerator
In short: Compatible with existing Spark jobs but executing them much faster. Benchmarks in the README file and docs [1] show improvements up to 3x while not even all operations are implemented yet (i.e. if an operation is not available in Comet it falls back to Spark), so there is room for further improvements. Across all TPC-H queries the total speedup is currently 1.5x, the docs state that based on datafusion's standalone performance 2x-4x is a realistic goal [1]

Haven't seen any memory consumption benchmarks but suspect that it's lower than Spark for same jobs since datafusion is designsd from the ground up to be columnar-first.

For companies spending 100s of thousands if not millions on compute this would mean substantial savings with little effort.

[1] https://datafusion.apache.org/comet/contributor-guide/benchm...

simicd··on GQL: A New ISO Standard in Graph Query Language
Yeah totally agree. I think it's great there is an organization that works on establishing standards but paywalling them definitely adds lots of frictions. In an ideal world the org would be fully financed by its members/governments so the standards can freely proliferate. Guess the only ones paying are big corporations who can then claim to be in accordance with a standard.
simicd··on GQL: A New ISO Standard in Graph Query Language
Currently there is very little information available online on GQL, the SIGMOID paper 'Graph Pattern Matching in GQL and SQL/PGQ' linked in the article is a great intro to the topic [1]. The authoritative source is the ISO spec [2] but it is paywalled. In short, GQL ISO is inspired by the following query languages and has the aim to standardize graph querying:

- SPARQL: A query language for RDF data, developed for semantic web by W3C. [3]

- PGQL: Property Graph Query Language developed by Oracle [4]

- Cypher: Neo4j's query language designed for its graph database. [5]

- openCypher: An open-source initiative to make the Cypher available beyond Neo4j database. [6]

- GSQL: TigerGraph's graph query language, SQL-like syntax for graph querying. [7]

- SQL/ PGQ: Another subproject of the SQL standard group introducing graph queries inside SELECT statements. Superseded by GQL (chapter SQL/PGQ Property Graph Query in [8])

[1] https://arxiv.org/abs/2112.06217

[2] https://www.iso.org/standard/76120.html

[3] https://www.w3.org/TR/2013/REC-sparql11-overview-20130321/

[4] https://pgql-lang.org/

[5] https://neo4j.com/docs/cypher-manual/current/introduction/

[6] https://opencypher.org/

[7] https://www.tigergraph.com/gsql/

[8] https://en.m.wikipedia.org/wiki/Graph_Query_Language

simicd··on GQL: A New ISO Standard in Graph Query Language
Agree names are too close, searching GQL on Google returns results related to GraphQL as well.

GraphQL is older, it was introduced in 2015 [1] while the work on the ISO GQL standard officially started in 2019 [2].

[1] https://en.m.wikipedia.org/wiki/GraphQL

[2] https://en.m.wikipedia.org/wiki/Graph_Query_Language

simicd··on What I talk about when I talk about query optimizer (part 1): IR design
Agree, substrait is a really cool project! Related: if you like substrait you might want to check out datafusion too. The project is a query execution engine built on top of Apache Arrow (incl. SQL parser, query planner & optimizer, execution engine, extensible user defined functions, among others) and it implements a substrait provider and consumer: https://github.com/apache/arrow-datafusion/tree/main/datafus...
simicd··on Oxlint – JavaScript linter written in Rust
Have you by any chance used Pyright? If not, I can highly recommend it. The VS Code extension makes writing Python almost as if it's a statically typed language (+ there is a CLI if you want to check types in CI). The docs are claiming that it's 3-5x faster than mypy - I haven't run performance benchmarks myself, all I can say is that for all my code bases it is very fast after the first cold start.

Comparison to mypy: https://github.com/microsoft/pyright/blob/main/docs/mypy-com...