I stumbled this week on few answers on a topic that I’m following, and although they were really nicely looking, correctly formatted, etc. - they were absolutely wrong - ChatGPT was hallucinating, showing calls to functions that aren’t in specific SDKs, etc. And I’ve noticed that such answers are added to old questions too, adding more noise :-(
Yes from my previous experience there were common language pairs, especially on the short texts - Spanish / Catalan, often German / Dutch, Bulgarian / Ukrainian, Dutch / Afrikaans, and few more…
It includes Python client for Spark Connect, more ANSI SQL compliance, very good improvements in Structured Streaming, new built-in functions, and much more.
it was before, but it was relying on the "normal clusters" with longer startup times, limited scalability, etc. Now it's serverless with fast startup times, scaling to 0, etc.
10 years ago I’ve attended a talk from one of German car producer - they talked about 25+ years of software maintenance. You invest into tools 5 years before release to market, and they should work 20 years after release. So companies need to plan the whole cycle not only code maintenance
I have an Integration of Planet Clojure directly using API (direct posting allows attribution to specific person). So, 8.5k followers will be cut from that news source…
Even for data transformation logic, SQL isn’t the best choice. How would you handle the case when you need to apply the same transformations to few dozens or hundreds columns?
yes, but it's a best thing of the cloud - your cluster doesn't run when you don't need it, plus you can advantages of spot instances, autoscaling, etc.. And you won't do TPC test each every half an hour.
That number should be "The total 3-year price of the entire Priced Configuration must be reported, including: hardware, software, and maintenance charges", so they just took the cost of the hardware used for benchmark, and extended it to 3 years.
If you are using IDE you can look to dbx Python package that has the “dbx sync” command that allows to sync local code with Repos and then you can quickly test your code in notebooks
Another alternative is to split notebooks into “library notebooks” that just define transformations, and “orchestration notebooks” that use code library notebooks to execute a “business logic”.
Although I wasn’t very hard to implement it for a subset of queries like group by, etc. - primary problem was a good OSS sql parser. We used Presto back in 2016th for that
But then you need to push these segments into partitions, and big partitions are really bad, especially for old versions of Cassandra… Although I met customers with partitions of size of 100Gbs…
Yes, 100%. I’m trying to use registration information for cybersecurity stuff, and it’s a mess. Some TLDs just doesn’t provide that information or provide it only to registered accounts or only inside their country. Parsing is a mess. Many have rate limits, like .au has 20 requests/day, .cz - 100 day, but with delay of 3 minutes between requests, …
I read many technical books as ebooks (epub or pdf), but these are primarily about "current tech" that will most probably become obsolete in 2-3 years. But I buy real books for more fundamental technical areas, where the information will be relevant for a long time. Another factor is if the book is math heavy - around deep learning & related topics, although in many cases PDFs are ok.