HNHacker News
TopNewBestAskShowJobs

paulasmuth

1,388 karma · joined May 22, 2011

submissionscomments
paulasmuth··on Show HN: Octo.ai, Open source analytics hypervisor
>> What we figured was the biggest problem with any form of machine learning or data based system is that we require high quality historic data. When a company starts the heavy dependance on free/freemium analytics services it means that all that data is usually lost for good. [...]

At the risk of shamelessly promoting our (also open-source) product: We also didn't like the fact that all commercial event analytics tools were completely proprietary and would effectively be closed data silos, so we recently released the open source EventQL event analytics database [1]. Maybe it would be an interesting target to support in octo.ai?

[1] http://eventql.io/

paulasmuth··on Show HN: EventQL – Open-source distributed SQL analytics database in C++11
We do have a csv import util (the API expects JSON) but it's not in the current distribution/release build. I'll add it and update this comment once it's live.

Queries are mostly limited by IO if running on regular hard disks. The number of rows/seconds mainly depends on the number and types of columns that are accessed. For example, if we scan 1.8B rows and only load a single integer column from disk (and the integers are small), we'll only have to load about 1 byte per row from disk (using an idealized model excluding some overheads for illustration purposes). If we want to complete the query in 1.5 seconds that would be a total IO load of 1144MB/s. So (depending on disk speed) around ~15 machines would suffice.

paulasmuth··on Show HN: EventQL – Open-source distributed SQL analytics database in C++11
Yes, it is similar to BigQuery with a couple of differences as you pointed out. The big ones being that it's fully open source and self-hostable.

It's less similar to Apache Drill - Drill is "only" a query engine and doesn't handle the actual data storage. EventQL combines a bigtable-like storage engine (optimized for the analytics use case) with a dremel/bigquery/dremel-like query engine.

paulasmuth··on Module Hub – The Redis Modules Marketplace
Interesting that they chose AGPL as the license for their modules (as opposed to BSD for core redis).

It's good to see such an important project adopt a viral copyleft license. Hope this will help the redis developers to monetize on the huge momentum they've built in the past years.

paulasmuth··on The future of SaaS hosted Git repository pricing
>> The rise of microservices, an approach to development in which you structure your software into smaller individual service-oriented units. Each microservice then runs its own process and they communicate with each other through APIs. This software development approach is known to have four primary benefits: agility, efficiency, resiliency, and revenue

Off-topic, but what?! I understand there is a lot of hype about "microservices" right now but I don't think the upsides of the propagated approach are as clear-cut or uncontroversial as the post makes them out to be.

The article defines "microservices" as the act of splitting up an application into many small, individual programms that talk to each other via RPCs.

How does structuring my code into lots of small programs improve revenue of all things? Or are you talking about the revenue of code repository hosting providers and "microservices" consultants?

Also, this approach surely does not improve resiliency and efficiency per se. It actually does the opposite in my opinion. All the additional RPCs are only adding _more_ points of failure (unreliable IPC) and overhead as compared to straight in-process method calls.

And who's saying that "a microservices architecture" improves a dev teams "agility"? How does one even measure agility? The article is stating this like it's an established fact. In my [anecdotal] experience, spreading an application out over lots of individual repositories and binaries makes it harder to work with, not easier.

>> With such tangible business benefits, it is no wonder why Google, Amazon, and Facebook have been using microservice development practices for over a decade.

The fact is that these are a massive software companies with thousands of engineers. Naturally they run an incredible amount of internal software and services. I think this is often conflated with the idea that splitting up a given application into smaller and smaller parts to produce more individual "services" automatically makes it better somehow.

Don't get me wrong. I'm all for having a clear seperation of concerns and well defined interfaces between individual modules. Doing that has been best practice since I started to code. But the latest "let's split our codebase up into 20 different binaries and repos, because microservices" fad is just bonkers.

paulasmuth··on My name still causes SQL errors
>> You might be interested to know your example doesn't escape that correctly.

It does -- You're right that the double-quote escaping is part of the original SQL standard, while the c-style escaping is an extension to it. However so is a lot of behaviour in modern SQL databases and it doesn't make my example incorrect.

Off the top of my head, here is an incomplete list of databases that implement c-style string escaping: Mysql, Postgres, Vertica, BigQuery, Oracle 9.2

>> some databases implement prepared statements as simple query string interpolation -- Do you have one in mind?

Yes, immediately both mongodb and bigquery come to my mind which are both pretty popular and do not currently support server-side prepared statements. If you google for jdbc drivers for theses databases some will implement 'prepared queries' as string interpolation on the client side.

Random Example: https://github.com/jonathanswenson/starschema-bigquery-jdbc/...

>> how exactly are prepared statements worse?

That's not what I said. My point (and I think we agree here) was that using a correct interpolation routine is just as secure as using a prepared statement with regards to SQL injection vectors.

paulasmuth··on My name still causes SQL errors
>> How are you turning the user supplied string into the SQL string literal string? Is your method for doing so guaranteed to always produce a valid SQL string literal which represents the supplied string?

Yes. This is trivial. Any self-respecting junior programmer should be able to write this routine.

paulasmuth··on My name still causes SQL errors
No, you don't need to unescape it when displaying it.

This SQL literal string:

    'Gijs in \'t Veld' 
Is a representation for these ASCII bytes/content:

    Gijs in 't Veld

In other words. The backslash will not be stored as part of the string into the database. It's just a hint to the SQL parser that the following single quote should be interpreted as a literal single quote char and not as the end-of-string-character.

From a security perspective there is nothing more "proper" about prepared queries than correctly escaped non-prepared queries.

paulasmuth··on My name still causes SQL errors
I can't tell if you're trolling or not, but this is not how string escaping works in SQL or any other programming language that I know.

    >> SELECT 'here is an apostrophe: \'';
    >> returns: here is an apostrophe: '
The backslash is not part of the string but just a hint to the compiler. The literal string '\'' represents a one byte (if ASCII) string containing only a single apostrophe character.

https://en.wikipedia.org/wiki/Escape_character

paulasmuth··on My name still causes SQL errors
This is false. Of course you can use any special characters in a literal string without using prepared statements. You just have to escape them correctly.

Example:

     INSERT INTO USERS (name) VALUES ('Gijs in \'t Veld');
Inserts the following row into the database

     ... | name             | ...
     ========================
     ... | Gijs in 't Veld  | ...

Prepared statements have nothing to do with this and are _in no way_ more resilient to SQL injection vulnerabilities than correctly interpolated non-prepared SQL queries.

(In fact, some databases implement prepared statements as simple query string interpolation in the client/driver. So this might be exactly what's happening when you use 'prepared statements' depending on the db)

paulasmuth··on Show HN: BitKeeper – Enterprise-ready version control, now open-source
The nested repository feature sounds amazing. Dealing with both git submodules and git subtrees has been a huge pain for me.

I'm looking forward to trying this out over the weekend. Is there some kind of util/script to import history from git?

paulasmuth··on Ask HN: Must you have no life, at least early on, to run a successful startup?
Revenue != Income. Gross profit (before paying any taxes) will probably be a fairly small percentage of turnover in manufacturing - you have to pay employees, purchase supplies, etc...

EDIT: just looked it up and most of "industry weeks"'s Top 50 manufacturing companies seem to have profit margins in the 10% range: http://www.industryweek.com/resources/iw50best/2015/48

paulasmuth··on Maybe: run a command, see what it does to your files without actually doing it
Yes, the current 'maybe' implementation should practically break almost any properly written IO code.

At this point, it will only work for the most trivial of demo cases. Still IMO it's a cool demonstration of the linux ptrace facility. And if the author implements the missing sandboxing/emulation layer in a future version and switches to whitelisting instead of blacklisting syscalls I think it could actually run a limited number of programs (forbidding stuff like mmap and network IO).

paulasmuth··on Maybe: run a command, see what it does to your files without actually doing it
While this looks like a nice idea on paper, I would not recommend to use the current implementation of 'maybe' on a system that hosts valuable data.

The tool seems to work by intercepting individual "blacklisted" system calls and then - instead of executing them - returning a nonsense value.

The issue is that this breaks every single POSIX spec and will therefore break any program that does more than a few trivial IO operations and relies on those operations to behave as specified.

So it might work for a simple demo case where a small script only does a single file modification and never checks the results, but for any serious program (think a database, a complex on-disk format or really anything that does network IO) this will lead to corruption and undefined behaviour as system calls will return erroneous success values or invalid file descriptors.

I think to actually make this work one would have to emulate the system calls and make sure everything stays POSIX compliant. Doing this correctly for calls like mmap might get tricky though (and won't be possible from within a python runtime). And even then it isn't obvious how something like network IO would be handled.

paulasmuth··on iPhones 'disabled' if Apple detects third-party repairs
I'm not sure I understand what exactly happened here. Was it previously possible for non-apple engineers to replace the home button or was it not? The guardian's article seems to suggest it was: "Indeed, the phone may have been working perfectly for weeks or months since a repair or being damaged."

If that is the case and it was possible to replace these sensors before, apple's narrative that the "error 53" code was introduced for security reasons doesn't seem to make a lot of sense: If the hardware sensor wasn't designed with secure authorization (e.g. via asymmetric cryptography) in the first place, all they could do now in a software update would be to add some kind of cosmetic device ID check.

However, any such newly introduced check in software could not actually prevent "malicious sensor" attacks but would only add a (possibly trivial) additional step to the attack where you have to spoof the correct device id.

Or maybe my reading of the guardian article is imprecise and replacing the home button has always meant loosing access to at least some security-relevant features?

paulasmuth··on Results of the 2015 Underhanded C Contest
Is there a specific reason for using typedefs rather than (packed) structs? With typedefs you get no real type safety (you can pass a widget_counter_t to a method that expects a foo_t if both typedef to the same thing), do you?
paulasmuth··on BadBarcode: Start a shell by scanning a boarding pass
Agreed on all points.

Still the "airports can be hacked using this ninja barcode trick" spin that they put on the story pushes my buttons. -- The "trick" is absolutely trivial and they haven't even bothered to try and confirm it on _any_ real world device.

This doesn't stop them from framing it in a way that makes it look like a significant discovery of a new vulnerability which "affects the entire barcode scanner-related industries" and then try to insinuate fear of the possible consequences of that "new discovery": "[It's] really a serious problem, not just a bug people could use to get free beer". Even though it's all based on complete speculation in the first place.

In all likelihood, the people building those systems have thought of and closed the attack vector a long time ago (if it was ever there). But of course vice apparently hasn't even asked a single vendor for comment -- maybe the answer would've been "no, it's not a problem in our product".

I find it a typical case of vice reporting. They take something that isn't exactly true/new to begin with and then blow it extremely out of proportion. This creates the illusion that only vice has the hottest, rawest and most uncensored stories abut sex, drugs and crime which nobody else reports about like they do (because they are mostly made up by vice). IMO vice is classic yellow press packaged for the hip and trendy geeks of my/our generation.

</rant>

paulasmuth··on BadBarcode: Start a shell by scanning a boarding pass
I could only speculate how specific applications (e.g. at airports) are built.

However, if some specific vendor was susceptible to this attack, it would just be a stupid, obvious and easily fixable input handling bug in their product.

But the linked article doesn't even demonstrate such an attack against an actual application ("BadBarcode is not a vulnerability of a certain product"). Just a trivial "demo" where they use a virtual keyboard device to enter commands directly into a windows shell and then get excited that it works...

paulasmuth··on BadBarcode: Start a shell by scanning a boarding pass
Am is missing something or is the presented "attack" really trivial? They use a barcode scanner connected as a virtual keyboard to directly enter keystrokes into the windows shell - I imagine the barcode simply reads "CTRL, T, S, O, M, E, C, O, M, M, A, N, D".

It's not surprising that it works to me -- I think it's the intended behaviour.

The article suggests the fact that barcodes can contain arbitrary (non-printable) ascii characters is the discovery of a new vulnerability which can be used to attack a large number of real-world POS/airport check-in systems, but doesn't give a single example or proof for that. This seems like complete speculation on the presence of input handling bugs in systems connected to barcode scanners presented as a fact.

paulasmuth··on On Botnets and Streaming Music Services
A similar thing has happened in the past:

>> "Gracia was selected to represented [sic] Germany in the Eurovision Song Contest with the song "Run & Hide", produced and composed by David Brandes. After the German national pre-selection for the Eurovision Song Contest it was revealed that Brandes had bought thousands of his own CDs to ensure chart placement, a requirement of the ESC"

https://en.wikipedia.org/wiki/Gracia_Baur

paulasmuth··on Ask HN: Who is hiring? (September 2015)
ZBase | Berlin/Amsterdam | full time, freelance, remote

About us: We develop a database/analytics product. Company founded in 2015 by german serial entrepeneurs and ex-Google engineers. We have launched our product in spring and it is already used by some of germany's largest ecommerce sites.

What we are looking for:

A freelancer or a small development firm that can help us with either of these two topics:

  - Ongoing Feature Development in the existing C++ codebase.
    This includes stuff like maintaining and improving our SQL
    engine and mapreduce framework, bringing ML models to 
    production/serving and general backend/API development.

  - Somebody with a ML/Stats background that can help us to
    tune some regression models.
Company is registered in Berlin but we are looking for somebody who wants to work remotely.

If you are interested / would like to learn more please give us a quick ping/one-liner: paul (at) zbase.io

paulasmuth··on Show HN: Serializable Math Expressions Using Cap'n Proto

    >> the advantage of the approach taken in this project (essentially, sending an
    >> abstract syntax tree encoded in capnproto) over just sending a string, is that
    >> there is no parsing-step needed at the receiver, which saves time. (It is a
    >> real-time system)
even though cap'n'proto advertises itself as a "zero copy" serialization scheme that doesn't necessarily need to copy the data when reading it, it still needs to "parse" the incoming message. and the way your interpreter is implemented I think you still need to switch()/recurse over everything in the proto tree. So I think the main difference here would be switching() over an enum vs throwing an extra strncmp/hashmap access in there when loading the expression from the incoming network buffer.

I think cutting that one strncmp per symbol could be a perfomance upside, but I reckon this would be _tiny_ compared to all the networking and evaluation overhead (think low microseconds for a few compares on strings that all fit within the SSO). But might be an interesting thing to benchmark nonetheless, maybe even against another solution that goes all the way and implements a proper micro vm so it can execute "bytecode" directly from the incoming networking buffer without any intermediate at all.

EDIT / side note: just had another look and saw that you are actually storing some symbols as strings in the capnprot messages and then do a lookup in the eval loop, too --- and the symbol lookup is implemented in O(N) by doing a linear scan over the symbol list and comparing the subject with each possible candidate here -- if you care about speed/micro-optimizing I think you can improve this to O(1) or at least logn using a hashmap/multiset -> https://github.com/niekbouman/commelec-api/blob/master/comme...

another thing if you want to microoptmize: you should try to pass stuff by reference instead of copy in your eval loop wherever possible -- removing as many mallocs() as possible from the inner interpreter eval loop should be one of the most low hanging fruits. what stuck out to me was esp. your use of the auto keyword in the eval code -- some thoughts regarding that:

- the "auto" keyword does not automatically take a reference. a plain "auto" may create a copy. use "auto&" or "const auto&" to take a reference

- esp. in the new range based for loops: in the places where you are currently using (for auto item : collection) you are performing a copy of all the items while iterating -- for (const auto& item : collection) will eliminate those copies

EDIT2: Skimmed your paper and saw it included some benchmarks regarding message size. However you are benchmarking your cap'n'proto solution ONLY against MathML, an XML encoded representation (lol).

So I tried one example from your paper -- am I missing something here?

Expression: "P-10+q^2"

- Your capnproto encoding with packing: 82bytes (value from your paper)

- Your capnproto encoding without packing: 280 bytes (value from your paper, actually larger than XML...)

- mathml (xml) without packing: 196 bytes (value from your paper)

- mathml (xml) without packing: 164 bytes (value from your paper)

- string encoding of the same expression ("P-10+q^2"): 9 bytes

[ Hope I am not coming across as too negative, but since you put it in a paper and published it I think it is fair to do some peer review ]

paulasmuth··on Show HN: Serializable Math Expressions Using Cap'n Proto
I have to admit I don't really understand what this is. Looking at the source I can find and a small server programm that sends/receives hardcoded UDP+JSON messages that contain some very specific data (a few numeric parameters, I think you need to understand electrical generators to understand what they mean) but doesn't seem to do anything with the data except to send and receive it and doesn't seem to contain error handling/etc. Also I see a bunch of cap'n'proto schemas to encode a number of mathematical expressions and a small c++ class to execute those expressions.

Since OP is the author of the linked project: How do the serialized expressions and the server program fit together and what would this software be used for? Is it still in early development? Also, why not serialize the expressions as simple strings/s-expressions -- what is the upside of using protobuf for this? doesn't it just add unnecessary complexity and overhead?

paulasmuth··on Tetris as a C++ Template Metaprogram
I realize that. I was joking [more like trying to, it seems] about the fact that parent commented on a technically amazing post with "I could have done this too in X" without providing any further insight, let alone a link to the completed project in X.
paulasmuth··on Creating a low-latency calling network
Yes, this seems like a bunch of work to keep up and running and I agree that most of the meat of their solution is actually in Google's or Amazon's systems running the GeoDNS/push stuff. Hopefully IPv6 will fix it all ™.

However, until then, GCM [1] seems like a really good workaround. And I believe it is actually free of charge and available for both iOS and Android.

[1] https://developers.google.com/cloud-messaging/

paulasmuth··on Creating a low-latency calling network
The load balancing they describe in their article is only used for inbound requests, similar to a traditional HTTP LB setup. If I understand it correctly they don't implement the outgoing message delivery to the called user themselves, but instead use SMS or a Google Service to deliver push notifications.

So the problem of "how do I route request X to user Y" has to be solved by either Google or the Provider that delivers the short message. -- Actually, with their current setup they don't even need to know where a user is located geographically, since they simply choose a proxy server by asking Route53 for a list of close servers (I presume close to the calling user) and then use their connect-then-disconnect-hack to choose the "best" server from that list.

    > Instead we decided to write our own minimal signaling protocol 
    > and use push notifications (at first SMS, then eventually GCM when 
    > it was introduced) in order to initiate calls.
So I imagine you could just stick the address of the switch that was chosen by the caller into the message that gets delivered to the called device when a call is initiated.

Obviously this doesn't solve the hard part of "how not to drop a call when your switch goes down mid-call" though.

paulasmuth··on Tetris as a C++ Template Metaprogram
I was under the impression that asm.js/emscripten are the new hotness these days? Hence not requiring porting everything to JS manually to bring it "to the web" anymore.
paulasmuth··on Tetris as a C++ Template Metaprogram
That's hilarious -- in fact the full quote is even better. This guy hits it right on the nose:

jeff atwood: "any application that can be written in JavaScript, will eventually be written in JavaScript. Writing Photoshop, Word, or Excel in JavaScript makes zero engineering sense, but it's inevitable. It will happen. In fact, it's already happening"

Good foresight considering he said this 2009.

paulasmuth··on Tetris as a C++ Template Metaprogram
Amazing. Maybe we are onto a new kind of Rule 34 here...

"Everything that can be implemented in {c++ templates, javascript} will eventually be implemented in that language"

paulasmuth··on Tetris as a C++ Template Metaprogram
Pics or it didn't happen ;) /jk

On a serious note: modern C++ does have const expression functions, too, albeit not as powerful (yet): http://www.cprogramming.com/c++11/c++11-compile-time-process...

← PreviousPage 2 of 5Next →