HNHacker News
TopNewBestAskShowJobs

simne

748 karma · joined February 15, 2022

submissionscomments
simne··on Genomic study: our capacity for language emerged at least 135k years ago
> total number of humans who ever lived 110B

But why do you think that all 110B individuals have made significant contribution to overall intelligence?

I think, at best case, in every generation exists few significant contributors, but all others are just carriers of genome diversity and nothing more, and all they have done, just disappear as a breath of wind.

So, for 200k years, if one generation 40 years, will be just 5000 generations, ok, lets consider significant 1000 persons from each generation, will be just 5millions, 5 magnitudes less than your estimation, and BTW much closer to known estimations of humanity knowledge and size of datasets used to train largest existing models.

simne··on Genomic study: our capacity for language emerged at least 135k years ago
> humans needed 10 million times more text for cultural evolution

Could you provide source, where exists this 10 million times more text?

- Text which is not stored on some material carrier is just forgotten.

Largest old text system I know is Confucianism. It is huge, but very far from 10 million times 1B.

Other known text systems are all much smaller than Confucianism.

https://en.wikipedia.org/wiki/Confucianism

simne··on Genomic study: our capacity for language emerged at least 135k years ago
LLMs are interest case, but they are nearly flat. Neural system of mammals is structured, as I understand, neocortex consists of about 100 millions structures each ~ 100x400 neurons, something like this. Anyway, just from calculation of human thinking delay appear human NN is just about 500 layers.

Second difference, natural I are feed-forward network, not back propagated as typical AI NN, because natural use some chemical method to "calculate" parameters.

I think, with right structure and some FF method, GAI is already technically achieved.

PS FF NN already exist, but problem that it now using classical method of differential equations which is prohibiting real use, because too much computation need.

simne··on Genomic study indicates our capacity for language emerged 135,000 years ago
There are many cases, when animals (or birds) literally speak and even showing very human like behavior on talking.

But, as I could see, only humans have so significant language culture that lead to great number of manuscripts.

I will even accent - little number of humans write books, but near none other species do this.

simne··on Genomic study: our capacity for language emerged at least 135k years ago
I wonder, if it is possible to detect what exactly changed in brain, as this could be a clue to create GAI.

Even strict list of changed genes could be extremely helpful for AI progress.

simne··on AMD's Strix Halo under the hood
I have seen video. There stated exactly: CPU don't have access to GPU cache, because "they have ran tests, and with this configuration some applications seen two digits speed increase, but nearly none applications they tested shown significant gains with CPU have access to GPU cache".

So, when CPU access GPU memory, CPU just directly access RAM via system bus, but not trying to check GPU cache. And yes, this mean, could be large delay between GPU write cache and data actually delivered to RAM and seen by CPU, but probably, smaller than on discrete GPU on PCIe.

simne··on AMD's Strix Halo under the hood
Plus, main idea of GPU, their "Computing Unit" is not alone.

Mean, in CPU could cut any core and it will work completely separated without other cores.

In GPU, typical, have blocks for example 6x CUs, which have one pipeline for all, and this is how they achieve thousand CUs or more. So, all CUs basically run same program, in some architectures could make limited independent branching with huge penalties on speed, but mostly just one execution path for all CUs.

Very similar to SIMD CPU, even some GPUs was basically SIMD CPUs with extremely wide data bus (or even just VLIW). So, GPU cache sure optimized for such usage, it provide buffer wide enough for all CUs on same time.

simne··on AMD's Strix Halo under the hood
> does the data move from the CPU cache to RAM to the GPU cache

Probably, not. Because it need dedicate channel on hardware level.

- GPU are mostly for streaming applications with large data blocks, so usually, CPU cache architecture is too different from GPU to simply copy (move) data, plus, they are on different chiplets, and dedicate channel means additional interface pins on imposer which are definitely very expensive.

So, when it is possible to make SoC with dedicated channel CPU<->GPU (or between chiplets), but usually it used only on very expensive architectures, like Xeon, or IBM Power, and not used on consumer products.

For example on older AMD products with APU, usually, graphics core have priority over the CPU to access unified RAM, but CPU cache don't have any additions to handle shared with GPU memory.

On latest IBM Power and similarly on Xeon, invented shared L4 cache architecture, where blocks of extremely huge L4 (near to 1Gb per socket on Power, as I remember somewhere about 128Gb on Xeon), could be assigned programmatically to exact core(s) and could give extremely high performance gain for applications running on these cores (usually these things very beneficial for DB or something like zip compressing).

Added: example difference CPU cache to GPU, for CPU usual size of transaction is less than 64bits, may be current 128..256bits but this is not common on consumer hardware (could be on server SoC), just because many consumer applications are not optimal to use large blocks, but for GPU normal to use 256..1024bits bus, so their cache definitely also have 256bits and larger blocks.

simne··on Block Diffusion: Interpolating between autoregressive and diffusion models
Example, as I understand:

(I'm not sure how should look prompt, my guess): Prompt: answer, what word is missing in text query. Query: What is it denoising there?

simne··on Ask HN: Any insider takes on Yann LeCun's push against current architectures?
> words spoken by 110 billion people who ever lived, assuming 1B estimated words per human during their lifetime..comes out to 10 million times the size of GPT-4's training set

This assumption is slightly wrong direction, because not exist human who could consume much more than about 1B words during their lifetime. So humanity could not gain enhancement from just multiply words of one human by 100 billion. I think, correct estimation could be 1B words multiply by 100.

I think, current AI already achieved size need to become AGI, but to finish, probably need to change structure (but I'm not sure about this), and also need some additional multidimensional dataset, not just texts.

I might bet on 3D cinema, and/or on automobile targeting autopilot dataset, or something for real life humanoid robots solving typical human tasks, like fold shirt.

simne··on Ask HN: Any insider takes on Yann LeCun's push against current architectures?
I'm not deep researcher, more like amateur, but could explain some things.

Most problem with current approach, to grow abilities, need to add more neurons, but this is not just energy consuming, but also knowledge consuming, mean, at GPT-4 level all text sources of humanity already exhausted and model become essentially overfitted. So looks like multi-modal models appear not because so good, but because they could learn on additional sources (audio/video).

I seen few approaches to overcome problem of overfitting, but as I understand not exist universal solution.

For example, tried approach to create from current texts some synthetic training data, but this idea is limited by definition.

So, current LLMs appear to hit dead end, and researchers now trying to find exit from this dead end. I believe, nearest years somebody will invent some universal solution (probably, complex of approaches) or suggest another architecture, and progress of AI will continue.

simne··on Block Diffusion: Interpolating between autoregressive and diffusion models
For LLM, standard method to provide to foundation model sentence without one word (random, but already known to model) and ask to fill this word.
simne··on Ask HN: What is the actual cost basis of the stock market?
> individuals that hold

I don't consider they serious, because they could buy penny stocks, which are unregulated and are very frequent target for fraud.

> pensions, mutual funds

Funds are better than individuals, because usually they have some strategy and hiring professionals to control things, but they also prone for mistakes.

simne··on Ask HN: What is the actual cost basis of the stock market?
Berkshire Hathaway is very special type of investor, as from Buffet cite, they only invest as major owner, avoid to be minor. Because as major they have much more access to internal kitchen of company, not just standard for IPO open balance with omitted details.
simne··on Ask HN: What is the actual cost basis of the stock market?
I think, most realistic evaluation is classic VC "divide year sells by bank deposit percents", nothing less and nothing more. So, as example, NASDAQ have about 7.4B sales and divide by 5% will be somewhere about 140B (link below).

I think this approach most realistic, because it use very simple and easy to get input data and don't need sophisticated logic to separate debts.

Unfortunately, I remember statistic S&P from 2008, from all NASDAQ only Microsoft had fully 100% coverage of all debts (so AAA rating), not just 1 year of percents as all others.

For other things exist commodities exchange, etc. Just calculate sum of all exchanges and will got good enough approximation.

And I don't think it is worth to use more deep logic, because will need to spend much more time and resources to gather data, and will achieve basically same, but delayed to few months or more.

Only exclude, it may worth to get somewhere survey of households economy, as they are much more conservative than business, so data from 3-5 years ago will be good enough for first look.

https://www.macrotrends.net/stocks/charts/NDAQ/nasdaq/net-wo...

simne··on Ask HN: EM to Director
I'll describe situation from investors view.

Exists just two approaches to hire director. Each have drawbacks.

1. Grow organically from lowest position (Henry Ford claimed, he use this approach in FMC). Drawbacks - it is slow and such director will be definitely conservative, because he grown in company environment and probably will not want to change anything. Good things are - such way is good check of professional quality, and such director could be more loyal for company.

2. Hire director from outside, farther - better. Drawbacks - high risks to hire person without need knowledge/skills and definitely collective will see him stranger. Good things such person probably will be more open for innovations and will not much depend on company traditions.

Each investor himself decide, what he think more important for company - loyalty or innovation. I usually think, innovation is highest priority, but to be honest, life is complicated, and possible case in which I'll chose more conservative way.

And sure, human behavior is about will, for example, Lee Iacocca grown in FMC from mechanic, but once become president of other company.

simne··on Gemma 3 Technical Report [pdf]
Unfortunately, I could not remember, when median performance of mobile CPU become comparable to business Notebooks.

I think, Apple entered race for speed with iPhone X and iPad 3. For Androids things even worse, looks like median achieved Notebooks speed at about Qualcomm snapdragon 6xx.

simne··on Gemma 3 Technical Report [pdf]
Unfortunately, this is known business model, most known example was Eclipse IDE, which killed all small IDE businesses. Other example, MySQL from Oracle.

Yes, idea, to make basically free something, on which small-medium businesses could survive and grow to something big, so making big death valley between small and big businesses.

Only exception are tiny businesses, living in tiny niches, but for them nearly impossible to overcome gap from tiny to big.

And you should understand, "open models" are in reality open-weight models, as they not disclose sources from which trained, so community cannot remake model from scratch.

Headhunting is sure important, but big business typically are so much finance powerful, so they could just buy talents.

- Headhunting with reputation is really important for small businesses, because they typically very limited in finances.

Medium business typically between small and big, but as I said at beginning, making some strategic things free, create death valley, so it become very hard to be medium.

Reputation is good thing for all, but again, top corporations are powerful non-proportional to size, so in many cases for them is relatively cheap to just maintain neutral reputation, they don't need to spend much to whitening.

simne··on Gemma 3 Technical Report [pdf]
> Is there a way to get something going on hardware a couple years old?

Tensor accelerators are very recent thing, and GPU/WebGPU also recent. RAM was also limited, 4Gb was long time barrier.

So, model should run on CPU and within 4Gb or even 2Gb.

Oh, I forget one important thing - couple years old mobile CPUs was also weak (and btw exception was iphone/ipad).

But, if you have gaming mobile (or iphone), which at that time was comparable to Notebooks, may run something like Llama-2 quantized to 1.8Gb at about 2 tokens per second, not very impressive, but could work.

simne··on Ask HN: Should there be new RPN calculators to replace the TI-84?
I remember. There was relative of MK-61, nearly fully compatible and in theory have extension port:

https://en.wikipedia.org/wiki/Elektronika_MK-52

And sorry for Russian language link, for some reason I have not found English.

https://ru.wikipedia.org/wiki/Электроника_МК-161

https://web.archive.org/web/20140202161045/http://mk.semico....

simne··on Ask HN: Should there be new RPN calculators to replace the TI-84?
I have used Soviet RPN calculator at my school years. It was not compatible with TI or something Western, but as I know ideologically it was very close.

What also interest, in early 2000s, Russia made modern remake, with much improved performance and added digital i/o, and compatible with old programs (even emulated some old glitches), so it could really run tons of programs one could find on old books/magazines.

I could not recommend anybody to buy Russian calculators, just want to say, this is very significant niche (RPN), but hugely depend on compatibility (with old software and old books) and on nostalgia.

https://en.wikipedia.org/wiki/Elektronika_MK-61

simne··on People are just as bad as my LLMs
> But an LLM can't be held accountable.. neither can most employees

Yes and no.

Yes, this is really problem, because at current level of technologies, some thing are inexpensive only if done in large numbers (factor of scale), so for example, just could not exist one person who could be accountable for machine like Boeing-747 (~500 human-years of work per plane).

Unfortunately, modern automobile is considered large system, made from thousands parts, so again, not exist one person to know everything.

And no, Germans said "Ordnung muss sein", which in modern management mean, constant clear organization of the game of the whole team is more important than the success of individual players.

Or, in simple words, right organization, controlled by rules is considered enough reliable to be accountable.

And for example in automobile industry, now normal to consider accountable whole organization.

And for example, Daimler officials few years ago said, Daimler safety systems will use Daimler view on robotic laws - priority will be safety of people inside vehicle. You may know, traditionally used Lem robotic laws, which have totally different view, separated from inside vs outside approach. In civil aviation using approach, to just use simple designs or design with evidence of reliability.

Sure, government regulators could decide something even more original, will see.

Any way, as technology emerge, accountability of machines will be sure subject of many discussions.

simne··on Chasing RFI Waves – Part Seven
In eastern Europe in 80-90s produced small trucks with analog diesel engines, but as I know most enterprises already scraped and most lucky now producing WW or Renault by license.
simne··on Chasing RFI Waves – Part Seven
It is possible, but fortunately still exist countries like Ukraine, and produce old diesel engines for trucks (KRAZ) and for tanks (T-64 line, including T-90, and also used on military trailers like MAZ-537, so when you cannot visually see, you don't know if tank running itself or on trailer).

Main problem, in those countries have not managed to produce small diesel engines for light trucks, and KRAZ engines are something 6L or more (tank diesel are 10L).

I hear they made half-block near-flat modification for armored personnel carrier, and even used it on low-floor city bus, but as I know they was not much market success, so have not produced big series.

simne··on Planes are having their GPS hacked. Could new clocks keep them safe?
I must admit, I agree with nearly all you said. Problem is that Ukraine was weak, and nearly without navy, and civilian ships was unable to resist to Russian navy. And Ukrainian export was blocked, as civilian ships fear to run within range of fire of Russian navy.

When Ukraine got enough weapons to force Russian ships to stay at distance, situation changed dramatically, so export was unblocked.

I think, very similar things happen during WWII.

This is not about only submarines, this is about superiority.

simne··on Planes are having their GPS hacked. Could new clocks keep them safe?
> slower on the surface than most naval vessels of the time

That is point. I'm not agree about most, but will be agree if you say about many.

> they couldn’t outrun pursuing destroyers or corvettes

But problem was, navy have so huge deficit of ships, so some convoys was run without naval support.

Sure, if all convoys was supported with fastest ships with best commands, u-boats will be no problem anymore, and as I understand, once this was happen.

simne··on Planes are having their GPS hacked. Could new clocks keep them safe?
> if the enemy knew its rough location

In these words you hit bull eye.

During WWII, submarines was just very special type of boat. You could check wikipedia about German u-boats - exist about TEN subtypes, from which only latest types have really significant underwater range, but all others was extremely limited in underwater activity.

But, surface ships of that time was even more limited, many could not achieve even half of surface speed of u-boat, so become easy prey.

But if you will try to find some artificial object on sea surface, that is really hard question. Just because sea is huge, so you need to check extremely large space in short time.

Radars are better to spot artificial object on sea surface than visual, just because radar easier to automate. But nothing more. Radar is also have problem of square distance, very similar to visual. So, as it is hard to spot partially surfaced submarine visually, it also hard to spot such sub with radar, because much less part will be on surface, so radar will have much less signal to detect.

Periscope size is nearly undetectable on surface, if it used carefully, just outside detection range of radar.

So, to conclude, Ukraine problem is, we cannot detect partially surfaced submarines on open sea, but they could fire missiles. Fortunately, Russians have very few submarines on Black sea, and after they was hit at harbors, their usage become very limited.

simne··on Planes are having their GPS hacked. Could new clocks keep them safe?
Many small planes don't have INS in typical meaning, but their pilot is INS computer, calculated approximate nav from air data (air speed + weather data + compass or radio compass).
simne··on Planes are having their GPS hacked. Could new clocks keep them safe?
> ballistic missile submarines can’t really use active sonar or surface with any frequency

Detect semi-surfaced submarine at night is really hard, if don't have intelligence data that it will surface on some non-random position.

From experience of Ukrainian war, my country have success with eliminating surface military ships, because have constantly monitoring their moves with satellites, but I cannot remember any case when semi-surfaced submarine was hit.

simne··on Planes are having their GPS hacked. Could new clocks keep them safe?
Unfortunately, these numbers considered state of art for modern classic gyroscopes.

Better are quantum navigation systems, using quantum matter as sensor, but they was too bulky to be used on planes, only last years appear more compact systems, sized like common home fridge.

← PreviousPage 5 of 34Next →