HNHacker News
TopNewBestAskShowJobs

mdda

1,344 karma · joined August 27, 2010

email me : {your.name} at mdda.net my blog : blog.mdda.net (AI and OSS)

Co-organiser of : https://www.meetup.com/Machine-Learning-Singapore/

submissionscomments
mdda··on GPU World
Or : "The young GPU's illustrated human primer"
mdda··on Is Gemini 2.5 good at bounding boxes?
You've got to be careful with PDFs : We can't see how they are rendered internally for the LLM, so there may be differences in how it's treating the margin/gutters/bleeds that we should account for (and cannot).
mdda··on AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms
"Make this better in a loop" is less powerful than using evolution on a population. While it may seem like evolution is just single steps in a loop, something qualitatively different occurs due to the population dynamics - since you get the opportunity for multiple restarts / interpolation (according to an LLM) between examples / and 'novelty' not being instantly rejected.
mdda··on US Administration announces 34% tariffs on China, 20% on EU
and (of course) the company will record the data so that the robots will be able to learn via imitation learning ASAP
mdda··on DeepSeek-R1
I think the "Aha" is that the RL caused it to use an anthropomorphic tone.

One difference from the initial step is that the second time around includes the initial step and the aha comment in the context : It is, after all, just doing LLM token-wise prediction.

OTOH, the RL process means that it has potentially learned the impact of statements that it makes on the success of future generation. This self-direction makes it go somewhat beyond vanilla-LLM pattern mimicry IMHO.

mdda··on Oscilloscope watch ships after 10 years on Kickstarter
"Track your potential while staying current" would be more apt
mdda··on Tarsnap outage postmortem
So you're the one in charge of the unix epoch rollover?
mdda··on AI-generated beer commercial contains joyful monstrosities, goes viral
Did you see the tennis players from Nvidia? https://research.nvidia.com/labs/toronto-ai/vid2player3d/
mdda··on PyTorch 2.0
Google Colab gives you $free GPU (usually a 16Gb T4) preloaded with frameworks, ready to run. Later, you might be tempted by the Pro(+) version, but there's plenty of scope to move up the learning curve before spending any money.
mdda··on OpenAI ChatGPT: Optimizing language models for dialogue
Could you point to any resources online about how to do this? e.g. is this using 8-bit quantisation?
mdda··on Show HN: TensorDock Core GPU Cloud – GPU servers from $0.29/hr
Isn't that for only one quarter of the A100?
mdda··on YOLOv7: Trainable Bag-of-Freebies
To just play with something : https://huggingface.co/spaces/nateraw/yolov6 (There's an images tab, and some samples below).

If you go to the associated code, you'll see that it needs a 'backbone', 'neck' etc. What is a backbone? Questions that arise directly from the code will lead you towards good blog articles, etc. https://huggingface.co/spaces/nateraw/yolov6/blob/main/yolov...

OTOH, you could go and have a look at (for instance) the Stanford vision courses for a more 'theoretical' approach. But the code itself is often solid guide to what's going on (the frameworks used for Deep Learning map well onto what's being discussed in blogs/lectures/papers).

mdda··on No Language Left Behind
"All models are licensed under CC-BY-NC 4.0" :

So, to clarify, does this mean that companies cannot use these models in the course of business, or is it more about selling the translation results directly?

mdda··on How DALL-E 2 Works
Could be... Except their page (should you choose to believe it, of course) specifically addresses the advantages:

"""

"Advantages over Traditional GANs" : Thus, we observe that our model exhibits _better training stability_ and mode coverage.

"Why is Sampling from Denoising Diffusion Models so Slow?" : After training, we generate novel instances by sampling from noise and iteratively denoising it _in a few steps_ using our denoising diffusion GAN generator.

"""

mdda··on How DALL-E 2 Works
Or the two can be combined : https://nvlabs.github.io/denoising-diffusion-gan/index.html
mdda··on 5% of 666 Python repos had comma typo bugs (inc V8, TensorFlow and PyTorch)
The automatic bug report generation tool produces the following:

"Absent comma results in unwatned string concatenation on line 330"

Bug-ception!

mdda··on Singapore: Sovereign City
Also : Zero capital gains tax. Zero estate taxes.
mdda··on I’m leaving London for NYC and taking my tech startup
Failing at a business in the UK: "Told you so! And now everybody knows you're a failure."

Failing at a business in the US: "Every success has a few failures on the journey : If you have another go, it'll prove that you're a fighter!"

Succeeding at a business in the UK: "Who did you screw over to make that money?"

Succeeding at a business in the US: "Awesome! Let me pitch you my idea..."

mdda··on How to train large deep learning models as a startup
The search term you're looking for is "Keyword Spotting" (or "Wake Word Detection") - and that's what's implemented locally for ~embedded devices that sit and wait for something relevant to come along so that they know when to start sending data up to the mothership (or even turn on additional higher-power cores locally).

Here's an example repo that might be interesting (from initial impressions, though there are many more out there) : https://github.com/vineeths96/Spoken-Keyword-Spotting

mdda··on Autonomy founder Mike Lynch can be extradited to US
FWIW, for the International audience, the (excellent) UK series "Spooks" was renamed "MI-5" - it was all about the domestic intelligence services in the UK, and (IMHO) worth looking for on NetFlix...
mdda··on Useful algorithms that are not optimized by Jax, PyTorch, or TensorFlow
"I wanted to use splines for the baseline survival" - isn't this a modelling step that could have been revisited? It seems a somewhat arbitrary choice (there are other ways to ~interpolate that are much more framework friendly) - and it seems that it forced you down a bit of a rabbit-hole.
mdda··on Pfizer is testing a pill that, if successful, could cure Covid-19
Definitely a thing in the UK : http://www.bbc.co.uk/comedy/onlyfools/lingo/
mdda··on Replicating GPT-2 at Home
Look under "FP16 16-bit (Half Precision) Floating Point Calculations" on https://www.microway.com/knowledge-center-articles/compariso...

These raw numbers don't tell the whole story, of course. But IMHO, the convenience of a local 2080Ti outweighs the speed benefits of an _somewhat flaky_ V100 via Colab for day-to-day use (unless memory size is an issue, which you can't really get around).

OTOH, for just trying out stuff / one-offs, Colab is perfect - and bonus points if you score a V100.

mdda··on Transputer
I was a summer intern at the UK company (Perihelion) doing the Atari-based Transputer machines. The word there was that because Atari had invested in the UK company (?) the Atari machine was essential to include as the front-end, even though it didn't really fit the UK designers' idea of what a good front-end machine would be... (nothing against the Atari design, just that its quirks/shortcuts didn't match up with the Transputer cluster backend 'vision').
mdda··on Cook: ‘China Hasn’t Pressured Us’
Where "variant" is Japanese! Much like English is a "variant" of Italian.
mdda··on Meetup.com alternatives
The venue should be available for free if your meetup is attractive to people that could be potential employees of the hosts. And why do you feel the need to give people free food and drinks? Reframe the issue : Everyone who attends should be motivated to come to the event for the content, rather than free food on their way home.
mdda··on Lake found at 11,000 feet in the Alps proves climate change is real
People should probably check the pictures of the water on your cited webpage before they take this as some kind of evidence. The lakes listed seem like mostly geothermally heated, or no longer existing.
mdda··on Missing tattoos: Altered photo lineup by Portland police draws objection
I semi-remember hearing that ear-prints are also somewhat unique (like, say, thumb-prints). Since the CCTV seems to have a decent ear image, couldn't the defense demonstrate that their client's ear is distinctly different in shape?
mdda··on Neural nets typically contain smaller “subnetworks” that can often learn faster
~No. Suppose that 5 of 50 channels in a particular network layer made up the 'lottery ticket' for that layer. The number of 'combos' of 5 channels that were trained at the same time is 50C5 (i.e. ~2million [0]).

Whereas training 5C5 (=1) 10 times only gives you 10 chances to get the right 5 together.

At least, that's one way to think about it.

[0]: : https://www.mathway.com/popular-problems/Finite%20Math/60182...

mdda··on Dandelion Seeds Fly Using ‘Impossible’ Method Never Before Seen in Nature
Perhaps people are wondering how your UNet is going to see the bristles on the other side of the dandelion seed.
Page 1 of 21Next →