Scientific Developer working in big data ML/AI.
creator of RPHash and Approximate Bloom Trees. Developer of GPU-Leech Lattice Decoder, Cardinality Shift Clustering, Nokia 6100 Linux LCD driver, as well as other commercial and hobby embedded devices.
CV
https://dblp2.uni-trier.de/pers/hd/c/Carraher:Lee
GH
https://github.com/leecarraher
PhD-Computer Engineering and Computer Science
MS-Computer Science
BS-Computer Engineering/Math
I feel like Jeff hawkins was working on this with HTMs, but they never really took off. I suspect it was hardware adaptability based, memory access patterns for sparse graph training aren't ideal candidates for t/gpus.
I understand charge time is a concern, but it is'nt the theoretical limitation of really any battery technology. From the computing parlance it's embarrassingly parallel. More, smaller cells allows for more current => faster charge times.
theoretical limitations: energy density, number of cycles
economics limitation: material and manufacturing costs.
infrastructure limitations: grids to power these charge rates at scale
Sodium Ion is promising because it drastically lowers the material cost. Cheap batteries can help solve the infrastructure problems as energy reservoirs, but I am more or less not swayed by the fast charge time break through. I can show doubling charge time with some AAs.
Do they mean deterministic k-means, k-means++ ... ? Global optimal k-means is NP-Hard, so linear speedups aren't terribly helpful. It's nice, until you add more input. Standard k-means would be nice, or the k-means++ seed algorithm.
I've considered the move to wired not for quality but for the sad state that Bluetooth pairing headphones has become. Theycan't just be headphones anymore; They require their own app and pairing protocol. They want 19 different touch points and permissions to implement a handful of never used features I get people being frustrated at why they can't just do what copper did for the last century.
probably a lot of economics going on, such as early age vendor lock-in, and new market acquisition loss-leaders, but ultimately it's not cutting edge hardware. So the same reason the laptop you bought 2 years ago is half the cost it is today. Granted, even that is not purely a cost only decision. Stratify any market and see how much you can get each segment to pay, and convince them they are getting the best deal for their money.
are you referring to this paper https://arxiv.org/abs/1501.01711 ? i believe they won best paper at icml or other impact journal. the published paper and algorithm i recall being compact and succinct, something that took less than a day to implement.
i agree, not just the multinomial sampling that causes hallucinations. If that were the case, setting temp to 0 and just argmax over the logits would "solve" hallucinations. while round-off error causes some stochasticity it's unlikely to be the the primary cause, rather it's lossy compression over the layers that causes it.
first compression: You create embeddings that need to differentiate N tokens, JL lemma gives us a bound that modern architectures are well above that. At face value, the embeddings could encode the tokens and provide deterministic discrepancy. But words aren't monolithic , they mean many things and get contextualized by other words. So despite being above jl bound, the model still forces a lossy compression.
next compression: each layer of the transformer blows up the input to KVQ, then compresses it back to the inter-layer dimension.
finally there is the output layer which at 0 temp is deterministic, but it is heavily path dependent on getting to that token. The space of possible paths is combinatorial, so any non-deterministic behavior elsewhere will inflate the likelihood of non-deterministic output, including things like roundoff. heck most models are quantized down to 4 even2 bits these days, which is wild!
for an interesting reversal of the "problem" of the speed of light, IEX is a stock exchange design to combat HFT by adding a physical speed bump by way of 38 miles of fiber optic cable. The general idea being to level the playing field and improve market liquidity using physical communication limits of light.
https://en.wikipedia.org/wiki/IEX
i agree, also add to that, that many python modules are foss projects that are maintained on a limited basis or budget. Refactoring code that may have some unsafe async routines would be costly for an org, and dreadful for recreation.
So you can either have a rich library of modules, or go async and risk something you need not working then having to find a workaround.
Personally, if parallelism is important enough, i use ctypes and openmp. If i need something more portable, i have a few multiprocessing wrappers that implement prange and a few other widgets for shared memory.
what was the position? what are your credentials to fulfill that position? I feel like cover letters, and recommendations are just icing on the cake of core skills and experiences, not the entire cake.
On the surface, cutting less essential resources during a power supply event makes sense, the ranking of essentialness seems problematic. While the decision to stop dumping megawatts of power to train a companies next gen LLM to be used for life saving/sustaining systems makes sense, it's pretty hard to implement in all but the most extreme cases. Hospital vs gpt6 training is an easy decision, but what about deciding between someone who wants to run AC at their unoccupied home vs. cutting power to a multi-day training epoch worth hundreds of thousands of dollars. It all feels very un-capitalistic, which in the US, like it or not, is how many edge cases get resolved. Right now datacenters are just the easy target, but why not Texas' numerous fracking sites, or other less desirable industries. My guess is that an injunction on the constitutionality of this will hold it up in court for a while.
my slightly next gen todo is a notebook on my remarkable. added features are sharing between devices, and since it's eink its a good paper like alternative to sticky-notes. For me beating procrastination can be more important than organizing many subtasks.
FWIW, i only use this for work todos and differentiate todo with calendar(paper calendar and dry erase board for home, outlook for work calendar)
I've used pbzip2 which takes the same parallel blocked compression approach 7zip seems to be taking (using AI's analysis of the changes). Theoretically the compression is less efficient, but i haven't noticed a difference in practice.
It depends on what level of nuts you mean. Some are AGI skeptics about LLMs, theyre probably right, there is likely more breakthroughs required before true AGI. But AGI isn't required to completely disrupt a ton of good, well-paid professions. That is the more worrying scenario. AI is already widening the wealth gap irreparably and with more progress it will only continue.
By defined format I of course don't mean random, or that it doesn't have structure, you have to be able to read it back otherwise what would be the point. I mean when you write out jpeg it has to adhere to a standard otherwise it isnt jpeg. Granted the jpeg format has evolved and has been augmented over the years. But it still is standard across iterations and devices.
Is raw a defined format? Or is it just whatever is most performant for the camera to write to storage, while also allowing for conversation to standard formats. Most are likely similar, running same or similar architectures, but in and of itself there is no raw format spec that must be adhered to. So if one way is more amenable to the system architecture, then that wins.
how much of the work can be ai generated, would a minor human copyrightable addition to the artwork constitute an original work. what would stop someone from generating art and popping a watermark or some imperceivable steganographic addition such that the ai part and human part cannot be disentangled.
maybe it was my prompt, but there seems to be far too much interpretation after the image embedding. In my examples it implicitly started to summarize parts of the text, unfortunately incorrectly. On an invoice with typed lettering it summarized that payments submitted would not post for 2-3 business days, when in reality the text said if you submitted after 2p on a friday, the payment would not post until the following monday. Which is significantly different. I'd be curious if you could ablate those layers in some way, because the one-shot structured text detection recognition was much better than vanilla ocr.
gpus dont implement complex number fp math, you have to bolt it on as extra logic. cufft works because you can recursively predict the imaginary and real component paths in the butterfly network. between layers you have fft->ifft , is this cost memory locality-wise worth it, or is it better to find ways to tamp down n in n^2 self attention by windowing, batching, gating, many other solutions. im not saying this work isn't cool, FNOs are really cool especially for solving PINNs and related continuous problems, are llms continuous problems, does n have to span the entire context window? I'll probably end up experimenting with this as theyve made the code available, but sometimes good theory is good theory, but not necessarily practical.
this seems to follow the similar FNO work by nvidia, and switching to frequency domain is usually in any computer scientist's toolbox at this point, however, I'm curious if this translates to real gains for real architectures. FFT makes use of imaginary numbers to encode the harmonics of the signal, these are generally not amenable to gpu architectures. Would fast walsh hadamard suffice? Sometimes the 'signal mixing' is more important than the harmonics of a compositions of sines. Or do we go further down the rabbit hole of trained transformation and try out wavelets? I am an avid FFT fan, (love fast johnson lindenstrauss transform using the embedded uncertainty principle for RIP), but sometimes real hardware and good theory dont always align (eg there are sub ternary matrix multiplies, but they are rarely used in DL)
From my experience, there has been a noticeable decline in PhD positions available within academia, likely due to tenure and career longevity, and reduced retirement benefits. As a result, many PhDs are forced into the private sector. However, many organizations have removed middle management layers, making merit based advancement less likely, and instead time becomes the dominating factor.
So given the choice between longer tenure or further education, where education is only marginally effective and time is dominant, the clear choice is to start a career as soon as possible. Which is something i wish i would have understood during my studies.
while the cabin may be able to be more spacious (like dirigibles and blimps) the flight my not be more comfortable because a turboprop's ceiling has you flying well within the troposphere where the majority of weather occurs. that may also put wear and tear on the airframe, incurring costs additional to the extended crew hours.
i think the article is confused. i write quite a bit of SIMD code for HPC and c is my goto for low level computing because of the low level memory allocation in c. In fact that tends to be the lion share of work to implement simd operations where memory isn't the bottleneck. if only it were as simple as calling vfmadd132ps on two datatypes.
i certainly used this trick back in my college days, prompted by a similar technique for "seeing" 3d stereoscopes on a computer monitor. I feel like i learned it somewhere on Eric Weisstein's Mathworld because the 3d objects viwer app let you split the image into two stereo images. Unfortunately java applets have been banished from the internet landscape.