243 karma · joined October 21, 2019
Distributed ML is tough to train because of very little control over train loop. I personally prefer using single server trainkng even on large datasets, or switch to online learning algos that do train/inference/retrain at the same time.
as for snowflake, I havent heard of people using snowflake to train ML, but sbnowflake is a killer in managed distribited DWH that you dont have to tinker and tune
spark has dataframe API which is similar to pandas api and can be learned in one day, especially if you know python.
same for Airflow and other frameworks, it just a fancy scheduler that anyone can pick up in a couple days.
THis process is not required for your average company where IT is a cost center and CRUD-type apps generator.
https://www.businessinsider.com/china-harvesting-organs-of-u...
fits perfectly the description of Security of complex IT systems, nice way to explain why IT security marketplace is a wild wild west
Anyone successfully implemented actor model framework over kafka?
interested in learning others' experience
yes still SQL is everywhere
2. always hide postgREST behind API gateway/load balancer/waf/ids+ips/rate limiter and you will be more secure from stuff lile this
Probably integrating ML model inference, written in ML.NET could be use case, but we have SQL with ML Server with r/python support now, so.
The problem with CLR is that you need to know and understand sql engine internals in order to write good C# code for CLR integration, otherwise your clr code will be blocking the sql engine
what you probably wanted to mention is table statistics - but they are cleared only if you truncate your table, but then again - once you populate your table - the engine will recalc statistics by itself.
overall RDBMS does a lot of stuff behind the scenes for you, and you should take advantage of it, and instead think about more important atuff - business logic, data modeling, schema evolution, etc
how would you debug HTML, for example? Open it in browser, right? same for sql: run it and see if it works as you wanted
same goes for composability: you can compose SQL same way you can compose HTML, but you still need to understand the big picture
explicit control flow in SQL (cursors, for loops and stuff) is absolutely an antipattern. You need to think in terms of relational algebra and functional programming in order to write clean SQL
or there could be another possibility: You think you have 10x engineer on your team, but actually you don't have them. You just keep employing a lot of 0.1x engineers that are unfit for their position/level/grade
I think the latter view is why companies like amazon/netflix fire a lot of engineers that they think are 0.1x or less than 1x
from https://en.wikipedia.org/wiki/Arctic_methane_emissions : The Permian–Triassic extinction event (the Great Dying) may have been caused by release of methane from clathrates. An estimated 52% of marine genus became extinct, representing 96% of all marine species.
go to a smaller dealer and ask them to buy a car from wholesale auction for you and you cover their costs. that way their margin js transparent and is negotiated upfront. You will be lucky to get 5-10% discount from retail prices, but wholesale prices at Manheim can be 15-30% less than retail, even more if you are willing to buy less than perfect condition car
this tool requires ready mostly clean-ish data to work with. but the #1 problem in data engineering is lack of such data