HNHacker News
TopNewBestAskShowJobs

mremes

11 karma · joined April 20, 2019

submissionscomments
mremes··on Show HN: Iso20022.js – Create payments in 3 lines of code
What XML has to do with bank-specific messages that have to be parsed and processed? It’s just a markup format.
mremes··on KVM Switches
Thanks for sharing! I actually have MX Mechanical keyboard in my setup, which uses Easy-Switch function, but I might need to get multiple dongles. As far as I have understood, in a KVM setup it's one port to connect from a target device to the switch, which is rather a convenience.

Of course, you might run up into a problem that you actually don't want to plug some devices into the switch when switching to another PC (for example some audio equipment)... So in a dream device, there would be switches on-off switches for ports (or groups of ports) and target PC enum switch. Probably some modular approach would work for making groups of ports switchable.

EDIT: https://www.avaccess.com/eu/product-category/kvm-usb-extende... (link to the manufacturer's e-commerce site) seems to have a good repertoire of KVM extenders.

EDIT1: Bluetooth support is still the missing piece in these...

mremes··on KVM Switches
I have been switching my USB peripherals (dongles mostly) and display ports for a while now within my desktop setup, and just now become aware of existence of these. I'm wondering how does this kind of switching work with e.g. Bluetooth.

EDIT: Apparently, based on HN search, many HN users have realized the same within the past few years.

mremes··on Ask HN: How do I improve our data infrastructure?
I disagree with the author of the parent comment in regards of using SQL and using Spark instead. I actually first wrote my "SQL advocation" as a reply to this comment but decided to leave leave this view for what it is and write my own "rant" against complicating "big" data transformations with Spark or EMR (Hadoop Pig) or vendor-locked Spark-instrumentations like AWS Glue.

But I agreed with the parent comment's author about pretty much anything until the third bullet point of the second list. I'd like to get more reasoning behind his SQL hate.

mremes··on Ask HN: How do I improve our data infrastructure?
I'd just go and write out the technical architecture, defining what are the inputs (the raw data) and what are the outputs (matrices for training, testing etc. etc.) on different intervals (usually, data scientists want the previous days' data processed into some format, A/B test results and such) and how are you going to instrument those transformations. It's not just SQL but the DB where that SQL would be run and orchestration (for example with Apache Airflow), and for concrete ETL tasks (nodes in a processing graph) using a combination of open-source modules (usually in Python) and Bash scripts.

It takes time to get experienced in explaining and mapping these things to the domain.

mremes··on Ask HN: How do I improve our data infrastructure?
Don’t worry, I’m having my battles convincing my clients (both business and DS/DEs) that this is a viable paradigm. Here’s a nice-looking z-value recipe by Silota that I just googled up: http://www.silota.com/docs/recipes/sql-z-score.html
mremes··on Ask HN: How do I improve our data infrastructure?
I'll express an advocation for using SQL as a data pipelining language. Firstly, many SQL dialects are multi-platform and provide standardization for transformations. It's a declarative language that doesn't define how computation happens but what.

Where SQL is terrible to write is when one must pivot data. Each column transformation is defined separately (case whens). When the cardinality of a pivoted vector is high, it results in quite a verbose declaration. This problem can be mitigated for example by generating SQL programmatically with templating languages such as Jinja2. Rendering is handled nicely on platforms such as Airflow when running the rendered SQL in cloud (for example on top of Redshift or Presto cluster, BigQuery).

For writing complex transformations, UDFs and cascading subqueries are the way to go. Window functions are useful for scanning subsets of column values (useful for example in vector transformations [doing normalization, regularization etc.])

SQL is also a language with a gentle learning curve which makes it easy to learn for less software-engineering-minded people (BI people and analysts of different departments in a decentralized data science organization). It's established itself as a lingua franca for matrix transformations already for decades.

Data processing is usually done in batches of different intervals as in traditional data science nothing really needs real-time processing for single events. Then Spark shines. But I would rather make a tradeoff of using SQL and Spark side by side when handling real-time processing than losing benefits of using SQL that I listed above.

When data transformations – with some object ontology related to it other than "just maths" – are to be done real-time, then you better start thinking about building an application for that (using your favorite programming languages).

Even with Spark, around 70% of work is done in SparkSQL.