61 karma · joined April 21, 2022
Say you're a primary home insurer in the US. If a hurricane hits you might not have enough capital to rebuild all the homes. A reinsurer which is also covering Europe, Asia, LatAm, etc. is less likely to go bankrupt. The reinsurer can cross-subsidize and use the insurance premiums from other regions to pay out the claims from the US market. All that matters is that on average the loss probabilities and severities are estimated correctly.
And this is just using one line of business as example, reinsurers are covering property, casualty, life and health which add extra layers of diversification.
The link you shared shows how DuckDB can run SQL queries on a pandas dataframe (e.g. `duckdb.query("<SQL query>")`. The SQL query in this case is a string. A dataframe API would allow you to write it completely in Python. An example for this would be polars dataframes (`df.select(pl.col("...").alias("...")).filter(pl.col("...") > x)`).
Dataframe APIs benefit from autocompletion, error handling, syntax highlighting, etc. that the SQL strings wouldn't. Please let me know if I missed something from the blog post you linked!
[0] https://openai.com/global-affairs/a-primer-on-the-eu-ai-act/
https://help.openai.com/en/articles/10250692-sora-supported-...
https://help.openai.com/en/articles/10250692-sora-supported-...
Spark is great if you have large datasets since you can easily scale as you said. But if the dataset is small-ish (<50 million rows) you hit a lower bound in Spark in terms of how fast the job can run. Even if the job is super simple it take 1-2 minutes. Polars on the other hand is almost instantaneous (< 1 second). Doesn't sound like much but to me makes a huge difference when iterating on solutions.
Haven't seen any memory consumption benchmarks but suspect that it's lower than Spark for same jobs since datafusion is designsd from the ground up to be columnar-first.
For companies spending 100s of thousands if not millions on compute this would mean substantial savings with little effort.
[1] https://datafusion.apache.org/comet/contributor-guide/benchm...
- SPARQL: A query language for RDF data, developed for semantic web by W3C. [3]
- PGQL: Property Graph Query Language developed by Oracle [4]
- Cypher: Neo4j's query language designed for its graph database. [5]
- openCypher: An open-source initiative to make the Cypher available beyond Neo4j database. [6]
- GSQL: TigerGraph's graph query language, SQL-like syntax for graph querying. [7]
- SQL/ PGQ: Another subproject of the SQL standard group introducing graph queries inside SELECT statements. Superseded by GQL (chapter SQL/PGQ Property Graph Query in [8])
[1] https://arxiv.org/abs/2112.06217
[2] https://www.iso.org/standard/76120.html
[3] https://www.w3.org/TR/2013/REC-sparql11-overview-20130321/
[5] https://neo4j.com/docs/cypher-manual/current/introduction/
GraphQL is older, it was introduced in 2015 [1] while the work on the ISO GQL standard officially started in 2019 [2].
Comparison to mypy: https://github.com/microsoft/pyright/blob/main/docs/mypy-com...