HNHacker News
TopNewBestAskShowJobs

bazizbaziz

167 karma · joined May 15, 2015

submissionscomments
bazizbaziz··on The New Three-Tier Application
Workflows/orchestration/reconciliation-loops are basically table stakes for any service that is solving significant problems for customers. You might think you don't need this, but when you start needing to run async jobs in response to customer requests, you will always eventually implement one of the above solutions.

IMO the next big improvement in this space is improving the authoring experience. In short, when it comes to workflows, we are basically still writing assembly code.

Writing workflows today is done in either a totally separate language (StepFunctions), function-level annotations (Temporal, DBOS, etc), or event/reconciliation loops that read state from the DB/queue. In all cases, devs must manually determine when state should be written back to the persistence layer. This adds a level of complexity most devs aren't used to and shouldn't have to reason about.

Personally, I think the ideal here is writing code in any structure the language supports, and having the language runtime automatically persist program state at appropriate times. The runtime should understand when persistence is needed (i.e. which API calls are idempotent and for how long) and commit the intermediate state accordingly.

bazizbaziz··on ARM Mac: Why I'm Worried About Virtualization
One benchmark would be to track down a python/JS/etc based "hello world" demo container. Base one version on Intel and the other on ARM, and measure each versions container build-time and request latency after it is set-up.

If changing the base image is all that's needed and both Dockerfiles otherwise assume ubuntu, this should not take too long.

bazizbaziz··on ARM Mac: Why I'm Worried About Virtualization
This seems like a weird benchmark, reading from /dev/urandom and gzipping random data does not seem like something most folks will want to do. It even appears like /dev/urandom speeds differ greatly on various architectures [0] and there are issues with /dev/random being fundamentally slow due to the entropy pool [1] (but I guess this is why the author uses /dev/urandom).

It would be better to measure something more related to what docker users will actually do, like container build time of a common container, and/or latency of HTTP requests to native/emulated containers running on the some container.

One reason to feel positive about the virtualization issues is that Rosetta 2 provides x86->ARM translation for JITs which an ARM-based QEMU could perhaps integrate into it's own binary translation [2].

[0] https://ianix.com/pub/comparing-dev-random-speed-linux-bsd.h... [1] https://superuser.com/questions/359599/why-is-my-dev-random-... [2] https://developer.apple.com/videos/play/wwdc2020/10686/

bazizbaziz··on Is anyone using the non Pro/Air Macbook for iOS app development?
I use an 2016 Macbook i7/8GB as my daily development system. I love it. It's light and portable, which is important for me.

The main thing to understand about these machines is that the i7 CPUs are about as good as any other CPU on the market (aside from less cores), but they're fanless which means they rely on the case to dissipate heat, and so will thermally throttle during long-running high CPU jobs. They're perfect for short lived jobs, even multi-minute compilations, but will have a hard time getting through repeated long jobs that require long durations of high CPU.In short, great for bursty computations with longer idle times where the laptop has a chance to cool off.

For instance, I build brew packages, full llvm builds, even small ML models, etc, without problems, because the machine starts cold and there is enough time after the job is done for the machine to cool off again. My machine suffers on tasks like Docker+Kubernetes/minikube that run a constantly polling VM in the background that takes 25-100% CPU when running idle.

For instance, someone in this thread mentioned the iOS emulator might be difficult to run. This may not be true, so long as the emulator does not constantly use lots of CPU - if it just uses high CPU in response to input events, it will likely be fine.

bazizbaziz··on Outperforming everything with anything: Python? Sure, why not?
This 'partial interpretation' trick used here has also been used successfully in the database community to accelerate whole queries, etc. Tiark Rompf's group in particular has been pushing this idea to it's limit.

Functional Pearl: A SQL to C Compiler in 500 Lines of Code - https://www.cs.purdue.edu/homes/rompf/papers/rompf-icfp15.pd...

How to Architect a Query Compiler, Revisited - https://www.cs.purdue.edu/homes/rompf/papers/tahboub-sigmod1...

Flare: Optimizing Apache Spark with Native Compilation for Scale-Up Architectures and Medium-Size Data - https://www.usenix.org/conference/osdi18/presentation/essert...

bazizbaziz··on Show HN: IronDB – a resilient key-value store for the browser
Thanks for the response! FWIW I'm overall really interested in this, as I maintain an application that uses localStorage to keep very important data while the app is offline until it can be uploaded.

I recently tested localForage [0] to see if it could be more reliable and get around storage limitations of localStorage. Unfortunately, LocalForage has a very annoying problem where it picks a single storage backend to use, but the one it chooses can switch on page reloads (yes, even on the same device/browser.) and this switching causes data loss as keys stored in the other backends are unaccessible. I'm very interested to see if IronDB can help here! Thanks for working on this.

[0] https://github.com/localForage/localForage

bazizbaziz··on Show HN: IronDB – a resilient key-value store for the browser
> When a value is retrieved via its key, IronDB... > Looks up that key in every store. > Counts each unique, returned value. > Determines the most commonly returned value as the 'correct' value. > Returns this most common correct value > Then IronDB self heals: if any store(s) returned a value different than the determined correct value, or no value at all, the correct value is rewritten to that store. In this way, consensus, reliability, and redundancy is maintained.

I'm not sure any of these properties are ensured the way anyone would want. Why do this? What are the failure modes this protects against? If only one of the backing stores is active when a result is written, and then the others later become active with no data, is blank data returned for the prior result? It seems like recording timestamps would fix this problem nicely on a single system and make this thing overall quite reliable.

bazizbaziz··on When optimising code, never guess, always measure
I didn't mean to say this was a silly thing to do - most modern processors execute instructions out of order on multiple ALUs.

The problem is that the abstraction layer between the python code in question and the processor's instruction stream is so thick that it's hard to say one way or the other that the processor is indeed executing that particular pair of instructions in parallel. It's definitely executing many instructions out of order, but it's unclear (without inspection of the python interpreter and its assembly) what's happening at the machine level.

Looking at the bytecode of the python program at least begins to tells us that the python bytecode of the two versions is fundamentally different, which could account for the performance difference. Although, what exactly makes the material difference is also under debate elsewhere in the thread. :)

bazizbaziz··on When optimising code, never guess, always measure
This comment is excellent. The title of the original post should be: "When optimising code, never guess, always read the bytecode/assembly."

Without actually reading the assembly/bytecode/etc, you end up speculating about silly things like 'the two evaluations and assignments can happen in parallel, and so may happen on different cores.'.

bazizbaziz··on How ProPublica Illinois Uses GNU Make to Load Data
Minor nitpick about their exit code technique [0]: The command checks if the table exists, but it does not appear to re-run if the source file has been updated. Usually with Make you expect it to re-run the database load if the source file has changed.

It's better to use empty targets [1] to track when the file has last been loaded and re-run if the dependency has been changed.

[0] https://github.com/propublica/ilcampaigncash/blob/master/Mak...

[1] https://www.gnu.org/software/make/manual/html_node/Empty-Tar...

bazizbaziz··on Mary Meeker Internet Trends Report 2018
Link to slides off KPCB's site without needing to sign up for an account: http://kpcbweb2.s3.amazonaws.com/files/121/INTERNET_TRENDS_R...

http://www.kpcb.com/internet-trends

bazizbaziz··on How Mapbox Is Winning Over Developers to Challenge Google's Mapping Dominance
Last I checked, MapBox's WebGL based vector tile renderer was a cut above the rest. For simple mapping, Google's stuff may cut it, but MapBox has the ability to draw complex vector polygons over a wide area that render/zoom/scroll fast and provide a nice experience. This is pretty important if you want to provide fine grained analysis of geographic areas.

If you wanted this on GMaps, you were stuck rendering all your vectors to image tiles, hugely increasing the size. MapBox's vector support made this very easy! It may be that Google/free options have caught up in this space, but I haven't re-evaluated in a few years.

I think MapBox could also be a winner in the GIS space as the GIS options have not made the most graceful move to the web.

bazizbaziz··on Why Raspberry Pi Isn't Vulnerable to Spectre or Meltdown
What about the fact that these instructions might get partially executed in the pipeline before the branch gets resolved and the pipeline flushed? If a mis-fetched instruction can reach the LSU stage before the pipeline gets flushed, it might serve as a speculative memory load...
bazizbaziz··on Why Raspberry Pi Isn't Vulnerable to Spectre or Meltdown
> There's no reason to predict a branch if you're not going to execute speculatively.

Not quite. Branch prediction is typically used on non-speculative architectures in order to avoid pipeline bubbles. (You could argue that pipelining is a form of speculation)

Here is the branch prediction documentation for one of the processors they claim is not vulnerable. http://infocenter.arm.com/help/index.jsp?topic=/com.arm.doc....

Whether or not they're vulnerable has more to do with how their pipeline is structured. It's possible for an architecture to be vulnerable if a request to the load store unit can be done within the window between post-branch instruction fetch/exec and a branch resolution. Eyeballing the pipeline diagram from the above docs, it looks like you can maybe get a request to the LSU off before the branch resolves. dramatic music

bazizbaziz··on Vanguard Founder Jack Bogle Says ‘Avoid Bitcoin Like the Plague’
"There is nothing to support bitcoin except the hope that you will sell it to someone for more than you paid for it.”

Genuinely asking, isn't this the plan for many people? How is bitcoin different than other commodities or property in this regard? Would Bogle say the same thing about those investments?

bazizbaziz··on Ask HN: If you could automate anything, what would you automate?
Acquisition and administration of public housing projects. The world needs more publicly housing for people of all economic statuses. It seems incredibly time consuming to build or convert private housing due to vast amounts of paperwork and groups that must be involved to finance and manage these projects. Automation can help to find a business plan, search for funding, help run governance, and do accounting for the on-going operations.
bazizbaziz··on Do we need a third Apache project for columnar data representation?
FWIW it should be possible to vectorize the search across the row store. With 24 byte tuples (assuming no inter-row padding) you can fit 2.6 rows into a 64 byte cache line (a 512 bit simd register). Then it's just a matter of proper bit masking and comparison. Easier said than done, I figure because that remainder is going to be a pain. Another approach is to use gather instructions to load the first column of each row and effectively pack the first column into a register as if it were loaded from a row store and then do a vectorized compare as in the column-store case.

All of that to underscore it's not that one format vectorizes and the other doesn't. The key takeaway here is that with the column store, the compiler can automatically vectorize. This is especially a bonus for JVM based languages because afaik there is no decent way to hand-roll SIMD code without crossing a JNI boundary.

bazizbaziz··on Samsung Made a Bitcoin Mining Rig Out of Old Galaxy S5s
Reminds me of this paper: https://blog.acolyer.org/2017/08/25/towards-deploying-decomm...
bazizbaziz··on The Google Home Mini Is Google’s $49 Answer to the Echo Dot
I notice on the tech spec's page that the Mini also has bluetooth [0]. Would love to know if this means all connected homes will play (the same) audio via bluetooth.

[0] https://store.google.com/product/google_home_mini_specs

bazizbaziz··on How to compare two functions for equivalence, as in (λx.2*x) == (λx.x+x)?
Undecidable in general, but there are two approaches I've seen work based on AST manipulation and comparison:

1. Solvers that search for applicable re-write rules to transform X into Y, such as Cossette for SQL. These may not terminate because undecidability lol http://cosette.cs.washington.edu/

2. Canonicalization of the AST. This is a form of #1 but much more restricted, and the hope is that functions that are equivalent end up canonicalized in the same way. LLVM and GCC do this for a variety of reasons. In the example given, you'd hope that both functions get canonicalized to either the left or right hand side. https://gcc.gnu.org/onlinedocs/gccint/Insn-Canonicalizations... http://llvm.org/docs/MergeFunctions.html

bazizbaziz··on The Age of Nvidia
> "The growing body of Big-data, HPC, and especially machine learning applications don’t need Windows and don’t perform on X86. So 2017 is the year Nvidia slips its leash and breaks free to become a genuinely viable competitive alternative to x86 based enterprise computing in valuable new markets that are unsuited to x86 based solutions."

Google's TPU paper [0] showed the CPUs were relatively competitive in the machine learning space (within 2x of a K80). It's not true that x86 doesn't perform on these workloads.

The existence of the TPU itself threatens Nvidia's dominance in the ML processor space. Google built an ASIC in a short time period that more than rivals a GPU on these tasks. The TPU performance improvements (section 7) make it look very straightforward to get even better performance with a few more years of development effort. With developers moving to higher level libraries, migration between GPU/CPU/TPU becomes painless, so they'll just go with whatever has the lowest TCO. (Google hosted TPUs?)

Aside from machine learning tasks, the author seems to be advocating for the cpu/gpu combinations that AMD is already selling to game console manufacturers. Granted, Nvidia has a piece of this via the Switch. If Microsoft/Qualcomm goes full-on with their ARM-based x86 emulation, then perhaps a future ARM-based Xbox is in the cards driven by an Nvidia chip? /speculation

[0] https://arxiv.org/abs/1704.04760

bazizbaziz··on Building a modern database using LLVM (2013) [pdf]
This is really great work, and awesome that they're developing this as open source. Combined with the results from HyPer folks it sure is starting to look like using LLVM to specialize code on the fly is a good idea for any data processing engine.

Looking more closely at the benchmarking results has me scratching my head, though: their reported 16x performance benefits from codegen for TPCH Q1 has seemingly dropped to 2x when compared to the [REDACTED] database. What's happening?

My guess is that Impala is sort of inefficient in a few places that still need work (which is OK, this is not a criticism of that). I bet that [REDACTED] is quite efficient due to having been in development for least 2x longer than Impala. Maybe even closer to 10x. In which case, getting within 2x is fantastic!

bazizbaziz··on Webhooks do’s and dont’s: what we learned after integrating APIs
How do people in production handle the possibility that your service might miss a webhook notification? If you miss a notification you'll end up with stale data and you won't know it.

Slack has a retry policy for a while but will then just give up. Another webhook provider I've looked at says nothing at all about this sort of thing. How do folks deal with this in production systems?

Seems to me like the best way to address this issue is to use the webhook as a hint that you need to run some other process that guarantees you've got all updates.

bazizbaziz··on Humans 'will bully robot cars', Mercedes chief warns
I think in many ways this could be considered a feature and not a problem. Why does it matter if someone butts in front of you? The passengers in the car can relax and let the computer handle the task. Yes it's important to avoid deadlock while merging during congestion but aggressive drivers will always try to butt in front of likely looking candidates, autonomous or not, and this usually doesn't stop progress of the line in general, just maybe frustrates the non-driver?
bazizbaziz··on The iPhone's new chip should worry Intel
> if Ax series performance starts to exceed corresponding Intel mobile parts by >50% (enough to compensate for the overhead of binary translation for legacy x86 apps).

Another milestone for an Ax laptop would be a good power efficiency story. If Apple can show an Ax based laptop with exceptional battery life under real working conditions that would be quite attractive to many people, even if translated programs experience a 50% drop in performance.

In addition to the Appstore bitcode stuff other comments mention, the Universal Binaries tools that were used for the mass powerpc->intel migration will also help to diminish the need for binary translation in many cases. UBs mean that you don't solely need to rely on Bitcode translation via the iOS/Appstore, either, so these fanciful Ax laptops could run all normal non-appstore apps quite easily.

bazizbaziz··on The iPhone's new chip should worry Intel
One of the reasons to do this would precisely be to free up engineering talent that would otherwise have to be dedicated to developing specialized laptop logic board designs (and other specialized parts). If you can build a relatively flexible internal hardware platform that can handle laptop, tablet or phone use cases, then your production teams can focus on core problems for a specific product while borrowing whatever solutions they need from other teams. Apple already has similar chips and boards running iPads and iPhones, so why not consolidate laptops too?

For some evidence that this could be happening, look at ifixit's macbook teardown[0]. The logic board seems to be approaching the size of an iphone logic board [1]. Someone at Apple has to be asking a what-if question here when looking at this thing and thinking about how they could go all the way.

[0] https://www.ifixit.com/Teardown/Retina+MacBook+2016+Teardown... [1] http://www.cultofmac.com/315469/new-macbook-logic-board-is-o...

bazizbaziz··on RISC instruction sets I have known and disliked
Kinda seems like the compiler just shouldn't allocate r0 for inline assembly on PPC, since it's only valid in special circumstances. Hard to fault the ISA a lot since this is basically the compiler backend author(s) missing a corner case, which is quite easy to do considering the breadth of a compiler backend.
bazizbaziz··on Why “Airbnb for event venues” websites all die/ venue marketplaces' fatal flaws
What sort of things did you guys try to convince people not cut you out? It seems that the marketplace needs very compelling value-add (aside from discovery) in order to get people to not cut it out after discovery. Trying to be the marketplace for existing venues seems quite challenging as most have already figured out their their infrastructure (payment processing, customer support, insurance, reporting, etc).