HNHacker News
TopNewBestAskShowJobs

DevelopingElk

223 karma · joined November 20, 2023

submissionscomments
DevelopingElk··on AWS multiple services outage in us-east-1
My guess is that for IAM it has to do with consistency and security. You don't want regions disagreeing on what operations are authorized. I'm sure the data store could be distributed, but there might be some bad latency tradeoffs.

The other concerns could have to do with the impact of failover to the backup regions.

DevelopingElk··on AWS multiple services outage in us-east-1
Mostly AWS relies on each region being its own isolated copy of each service. It gets tricky when you have globalized services like IAM. AWS tries to keep those to a minimum.
DevelopingElk··on How I bypassed Amazon's Kindle web DRM
It kind of makes sense. It's good enough to stop a non-coder. Anything in the browser can be either broken by a serious coder or has unpleasant tradeoffs.

Amazon would need to drop this feature to seriously lock down their books

DevelopingElk··on Palisades Fire suspect's ChatGPT history to be used as evidence
And Apple did provide all iCloud data they had available.
DevelopingElk··on Show HN: I invented a new generative model and got accepted to ICLR
One more concern I noticed: This generative approach needs not only for each layer to select each output with uniform probability, but also for each layer to select each output with uniform probability regardless of the input.

This is the bad case I am concerned about.

Layer 1 -> (A, B) Layer 2 -> (C, D)

Lets say Layer 1 outputs A and B each with probability 1/2 (perfect split). Now, Layer 2 outputs C when it gets A as an input and D when it gets B as an input. Layer 2 is then outputting each output with probability 1/2, but it is not outputting each output with probability 1/2 when conditioned on the output of layer 1.

If this happens, the claim of exponential increase in diversity each layer breaks down.

It could be that the first-order approximation provided by Split-and-Prune is good enough. My guess though is that the gradient and the split-and-prune are helping each other to keep the outputs reasonably balanced on the datasets you are working on. The split and prune lets the optimization process "tunnel" though regions of the loss landscape that would make it hard to balance the classes.

DevelopingElk··on Show HN: I invented a new generative model and got accepted to ICLR
A thought on why the intermediate L2 losses are important: In the early layers there is little information so the L2 loss will be high and images blurry. In much deeper layers the information from the argmins will dominate and there will be little information left to learn. The L2 losses from the intermediate layers help this by providing a good training signal when there is some information known about the target, but there are still large unknowns.

The model can be thought of as N Discrete Distribution Networks, one of each depth 1 to N, that are stacked on each other and are being trained simultaneously.

DevelopingElk··on Show HN: I invented a new generative model and got accepted to ICLR
First, I think this is really cool. Its great to see novel generative architectures.

Here are my thoughts on the statistics behind this. First, let D be the data sample. Start with the expectation of -Log[P(D)] (standard generative model objective).

We then condition on the model output at step N.

- Expectation of Log[Sum over model outputs at step N{P(D | model output at step N) * P(model output at step N)}]

Now use Jensen's inequality to transform this to

<= - expectation of Sum over model outputs at step N{Log[P(D | model output at step N) * P(model output at step N)]}

Apply Log product to sum rule

= - expectation of Sum over model outputs at step N {Log(P(D | model output at step N)) + Log(P(model output at step N))}

If we assume there is some normally distributed noise we can transform the first term into the standard L2 objective.

= - expectation of Sum over model outputs at step N {L2 distance(D, model output at step N) + Log(P(model output at step N))}

Apply linearity of expectation

= Sum over model outputs at step N [expectation of{L2 distance(D, model output at step N)}] - Sum over model outputs at step N [expectation of {Log(P(model output at step N))}]

and the summations can be replaced with sampling

= expectation of {L2 distance(D model output at step N)} - expectation of {Log(P(model output at step N))}]

Now, focusing on just the - expectation of Log(P(sampled model output at step N)) term.

= - expectation of Log[P(model output at step N)]

and condition on the prior step to get

= - expectation of Log[Sum over possible samples at N-1 of (P(sample output at step N| sample at step N - 1) * P(sample at step N - 1))]

Now, for each P(sample at step T | sample at step T - 1) this is approximately equal to 1/K. This is enforced by the Split-and-Prune operations which try to keep each output sampled at roughly equal frequencies.

So this is approximately equal to

≃ - expectation of Log[Sum over possible samples at N-1 of (1/K * P(possible sample at step N - 1))]

And you get an upper bound by only considering the actual sample.

<= -Log[1/K * expectation of P(actual sample at step N - 1))]

And applying some log rules you get

= Log(K) - expectation of Log[P(sample at step N - 1)]

Now, you have (approximately) expectation of -Log[P(sample at step N)] <= Log(K) - expectation of Log[P(sample at step N - 1)]. You can repeatedly apply this transformation until step 0 to get

(approximately) expectation of -Log[P(sample at step N)] <= N * Log(K) - expectation of Log[P(sample at step 0)]

and WLOG assume that expectation of P(sample at step 0) is 1 to get

expectation of -Log[P(sample at step N)] <= N * Log(K)

Plugging this back into the main objective, we get (assuming the Split-and-Prune is perfect)

expectation of -Log[P(D)] <= expectation of {L2 distance(D, sampled model output at step N)} + N * Log(K)

And this makes sense. You are providing the model with an additional Log_2(K) bits of information every time you perform an argmin operation, so in total you have provided the model with N * Log_2(K) bits for information. However, this is constant so you can ignore it from the gradient based optimizer.

So, given this analysis my conclusions are:

1) The Split-and-Merge is a load-bearing component of the architecture with regards to its statistical correctness. I'm not entirely sure about how this fits with the gradient based optimizer. Is it working with the gradient based optimizer, fighting the gradient based optimizer, or somewhere in the middle? I think the answer to this question will strongly affect this approaches scalability. This will also need a more in-depth analysis to study how deviations from perfect splitting affect the upper bound on loss.

2) With regards to statistical correctness, the L2 distance between the output at step N and D is the only one that is important. The L2 losses in the middle layers can be considered auxiliary losses. Maybe the final L2 loss / L2 losses deeper in the model should be weighted more heavily? In final evaluation the intermediate L2 losses can be ignored.

3) Future possibilities could include some sort of RL to determine the number of samples K and depth N on a dynamic basis. Even a split with K=2 increases NLL loss by Log_2(2) = 1. For many samples after a given depth the increase in loss due to the additional information outweighs the decrease in L2 loss. This also points to another difficulty, it is hard to give fractional information in this Discrete Distribution Network architecture. In contrast, diffusion models and autoregressive models can handle fractional bits. This could be another point of future development.

DevelopingElk··on Intel's Open-Source Strategy Is Changing at Odds with the Ethos of Open-Source
My view on this is that it is sad, but makes sense from a business perspective. Intel is in a really hard place which means making some painful cuts, especially in things with long term benefits. The Linux kernel is mostly funded by big companies. Contributions rise and fall, but it has enough diversity to weather this type of change.
DevelopingElk··on F-Droid and Google’s developer registration decree
I wonder what would happen if F droid signed all software under their keys even though they aren't the developer? Make Google ban them instead of just giving up?
DevelopingElk··on The EU Just Killed ARR
I think that AWS Reserved Instances will still be OK. That is purchasing a product with an expiration date, with the product being compute hours. And its mostly B2B, this article is mostly focusing on B2C. AWS will mostly be in the clear. Services are already billed per call or per monthly usage.
DevelopingElk··on Miles from the ocean, there's diving beneath the streets of Budapest
Much safer. Spelunking can be fairly safe if done with caution and an experienced team. Cave diving causes fatalities even among experts.
DevelopingElk··on FreeBSD Scheduling on Hybrid CPUs
I see the heterogeneous architectures as mostly a plus. If you want the most throughput for a highly parallel workload given a power and silicon budget 100% E cores would be best. If you have some workloads that don't parallelize well then a few P cores are best. Heterogeneous gives possibilities to optimize for both cases. There is another knob to turn, and mistakes can be made, but this should be an overall positive.

My bigger concern with the newer Intel CPUs are the crashes and reliability issues that were reported.

DevelopingElk··on The "high-level CPU" challenge (2008)
For sure. The issue is that many AI workloads require terabytes per second of memory bandwidth and are on the cutting edge of memory technologies. As long as you can get away with little memory usage you can have massive savings, see Bitcoin ASICs.

The great thing with the Von Neumann architecture is it is flexible enough to have all sorts of operations added to it; Including specialized matrix multiplication operations and async memory transfer. So I think its here to stay.

DevelopingElk··on Micron rolls out 276-layer SSD trio for speed, scale, and stability
Not true. High end controllers need DRAM for caching indexes. That's at least a few dollars.

Flash storage is a commodity, we are paying close to the amortized cost of manufactured and sales.

DevelopingElk··on Ageing accelerates around age 50 ― some organs faster than others
It's because age isn't normally distributed. The annual odds of dying go up by a factor of 2 every 8 years. This gives tighter bounds on age than a normal distribution would give.
DevelopingElk··on I solved the century-old mystery of a shipwreck survivor
That, and there were survivors to tell the tale. Ships sinking with all on board lost was a reality. There was always the chance that one who went out to sea might not return. The Titanic's survivors made the story known and memorable.
DevelopingElk··on M.2 SSD Can Self-Destruct by Giving Itself a Burst of Voltage
Flash chips, unlike hard drives, are highly reliant on their controllers. If you hit it with a hammer the data is unrecoverable. No need for a voltage zap.
DevelopingElk··on Astronomers discover 3I/ATLAS – Third interstellar object to visit Solar System
It's 1. A combination of better telescopes and GPU accelerated algorithms for picking out moving objects.
DevelopingElk··on Iran halts cooperation with UN nuclear watchdog
Agreed. Even as someone who thinks the strikes were justified this was an inevitable outcome. Iran was trying to get as far as they could with their enrichment without provoking a military response. Once there was a military response there was no reason to cooperate.

There might still be room for diplomacy and reentering IAEA. But since Iran leaving the IAEA strengthens their bargaining position this is a logical move for them.

DevelopingElk··on Alternative Layout System
Yes. A combination of being hand copied and the text having no punctuation.
DevelopingElk··on Recent CS grad unemployment twice that of Art History grads
I see that although unemployment is higher for CS underemployment is lower than many other majors. So I wouldn't immediately discount CS but the report is saying that CS isn't an immediate high paying job.
DevelopingElk··on Q-learning is not yet scalable
This is correct. I will add that sampling from the distribution thereafter is equivalent to on policy learning.
DevelopingElk··on Q-learning is not yet scalable
I think the article's analysis of overapproximation bias is correct. The issue is that due to the Max operator in the Q learning noise is amplified over timesteps. Some methods to reduce this bias, such as https://arxiv.org/abs/1509.06461 were successful in improving the RL agents performance. Studies have found that this happens even more for the states that the network hasn't visited many times.

An exponential number of states only matters if there is no pattern to them. If there is some structure that the network can learn then it can perform well. This is a strength of deep learning, not a weakness. The trick is getting the right training objective, which the article claims q learning isn't.

I do wonder if MuZero and other model based RL systems are the solution to the author's concerns. MuZero can reanalyze prior trajectories to improve training efficiency. The Monte Carlo Tree Search (MCTS) is a principled way to perform horizon reduction by unrolling the model multiple steps. The max operator in MCTS could cause similar issues but the search progressing deeper counteracts this.

DevelopingElk··on An illustrated guide to Amazon VPCs
Running out of IP addresses within that VPC is a real difficulty for services still using it.
DevelopingElk··on Dimension 126 Contains Twisted Shapes, Mathematicians Prove
Manifolds are generally considered objects of themselves, and it may be difficult to embed then in higher dimensional objects. This is especially the case for tricky manifolds like those with a Kervaire invariant of 1.
DevelopingElk··on New material gives copper superalloy-like strength
The high temperature talked about in the article is close to 800 Celsius. That far exceeds home or even most industrial appliances. The primary use would be in turbines where the combination of strength and heat conductivity can keep the blades from melting and improve efficiency.
DevelopingElk··on Astronomers confirm the existence of a lone black hole
The layman's answer is that since everywhere was very dense there wasn't a gravitational pull to one direction or another since it all cancelled out.
DevelopingElk··on Analytic Combinatorics – A Worked Example
In cryptography it is common to get exact estimates for the number of bit operations required to break a cipher. Asymptomatic analysis is great in some circumstances. It for example explains why quicksort is faster than bubble sort for large inputs. Beyond that you want some empirical results.
DevelopingElk··on JEP draft: Prepare to make final mean final
Yes, from a language design it's ugly and the implementation is convoluted. From a user perspective it's awesome and enables better interfaces.
DevelopingElk··on Airbus to test new engine design for successor to A320
Standard airplanes used for passenger travel have transonic airflow. Some air is forced to be supersonic as it passes over the wing and around the body of the plane.

The Transonic Truss-Braced Wing is still experimental but prototypes have been built and they are seriously considered for the next generation. If I recall correctly where to store the fuel is a bit concern.

← PreviousPage 2 of 3Next →