A relative of mine worked for the UN and interfaced with the UN after they left for a non-profit. Anyone that knows anything about them and also just simply observing what and how they are doing things should have no doubt that it is filled with people that got there by using their connections. And you absolutely constantly run into people that have no business being there other than through nepotism. Btw. I am sure that US staff is less likely to be a total nepo baby, but because the UN "has" to hire from all over the world, most roles are not filled like that.
It will all make sense once you realize who works at the UN, basically nepo babies of all colors and variety, including second cousins of Saudi royalty etc.
what is striking to me is how far reasoning by analogy and generalization can get you. some of the deepest theorems are about relating disparate things by analogy.
It would be great if we got "kernel independent" Nvidia drivers. I have some experience with bare-metal development and it really seems like most of what an operating system provides could be provided in a much better way as a set of libraries that make specific pieces of hardware work, plus a very good "build" system.
This is a terrible idea and direction but it will not stop people from pursuing it and as soon as they have a critical mass of people reviewing each other it will go on for quite a while. Transformers for time series is one of those things that seems to make sense but not really.
I think it is a given that they are aiming for a fully custom training cluster with custom training chips and inference hardware. That would align well with their abilities and actually isn't too hard to pull off for them given that they have very decent processors, GPUs and NPUs already.
That isn’t actually true. I‘ve personally worked on an embedded POWER processor https://github.com/electronicvisions/nux. It is used as a dual core embedded micro-controller in a neuromorphic chip. Each core has just 16kB of memory on chip and 4kB instruction cache. POWER has subsets of the instruction set which have a gcc and llvm target and are well suited for embedded use cases.
Main advantage of RISCV at this point is that it has a larger community. But at the time it didn’t even have a vector instruction set, whereas we could easily modify the POWER one for our purposes.
Personally I am almost certain that the current framing of RL and its relationship to animal behavior is deeply misguided. It proves close to impossible to train animals using this paradigm (not for a lack of trying), i.e. animals such as mice only make any progress when water deprived and under conditions that exploit their natural instincts. Nevertheless they are capable of far more complex natural behaviors. There is a non-zero chance that RL as an explanation of animal behavior is just plain wrong or not applicable.
> “Although the study found inadequate sleep duration was not an issue in brain atrophy in this study, we cannot say there is no association,” she said, noting that a previous CARDIA study showed that shorter sleep was associated with worse white matter integrity, indicating lower cognitive functioning.
That quote seems to directly contradict the headline.
There are tasks, which are implemented as part of the runtime and they appear to plan to integrate libuv in the future. Some of the runtime seems to be fairly easy to hack and have somewhat nice ways of interoperating with both C, C++ and Rust.
It cost 6 billion to build the current collider (in an existing tunnel) and CERN has a budget of 1-2 billion per year. In terms of data they are easily producing 100x of previous colliders.
Compare the number of issues here: https://wiki.archlinux.org/title/Vulkan. The only issue with the Nvidia driver is that there might be another (open-source) driver installed. AMD has several drivers which fail in different situations. The same has been true for OpenGL implementations, the only truly good implementation is the one by Nvidia, this is even well known in the PC game development industry.
I tried AMD exactly once for Linux graphics, it was so unstable that I bought a NVIDIA card within a month and have never looked back. Potentially NVidias approach of staying as far away as possible from the Linux kernel abstractions is much saner than to play ball with them.
It is a very fad driven field. Everyone brands everything. It isn't enough to give things boring titles like, stacked open linear dynamical system with selective observations and learned timestep.
There is nothing theoretical that stops anyone from trying that. It is just a matter of engineering and architecture. It is (in principle) possible to define "spiking transformers" and there is more than one way to do so, but this will not mean necessarily that you will immediately get good performance. Also current generation hardware doesn't as efficiently simulate SNN, so this is another constraint.
Trying to figure that out right now :), there is a tension between an algorithm like this and the biophysical reality of actual brains. My intuition is that whatever the brain does is not so much an approximation of any known algorithm, but rather an analog of such an algorithm. Now that we know how in a spiking neural network gradient computation looks like, we can ask what assumptions are violated in the brain, just like people have done before for the vanilla backpropagation algorithm. There is some recent concrete experimental evidence that points towards a solution.
That you can't backpropagate is a common misconception. There is recent work (I am one of the co-authors), that derives a precise analog of the backpropagation algorithm in spiking neural networks (https://arxiv.org/abs/2009.08378). It computes exact gradients and requires only communication at spike times during the backward pass. The reason this works is because "x is not differentiable" lacks a statement "with respect to y". It turns out that since gradient computation is local, a gradient is defined almost everywhere. The gradient computation is only ill-defined at places where a spike would get added or deleted. This is similar to how in a ReLU network an exact zero input is "non-differentiable".