HNHacker News
TopNewBestAskShowJobs

halflings

2,566 karma · joined September 4, 2013

Software Engineer specialized in applied Machine Learning, and in particular ranking and search quality.

Get in touch: https://kachkach.com/ ahmed@kachkach.com

submissionscomments
halflings··on Google's advanced music generation model and two new AI experiments
How about "AI will further democratize music"? Anyone can make music today, but ask any musician and they'll be mindblown by the idea that they can hum a melody and immediately generate a sax/percussion track matching it.

You can learn to play the guitar, and buy a guitar, but it's very hard and expensive to also learn the 3-4 other instruments you need to make a full song. (my brother learned both the guitar and percussions; this took a lot of time and was quite expensive, if he had a tool like this he would probably focus on the guitar and just rely on this to generate a background track)

halflings··on Omegle 2009-2023
That's like killing a mosquito with a bazooka.

Using GPT-4 vision on all users would be extremely expensive. Way simpler models can detect nudity. (and they do mention they had great success using those) And if the idea was to detect child abuse from text, it would also be quite expensive to use the language capabilities on every discussion.

halflings··on Tutanota is now Tuta
It means "raspberry" in Arabic, and in Moroccan dialect at least it's slang for "good looking woman".

Though it does sound like "tota", which in Moroccan dialect again is the childish synonym for penis :))

I guess any short name is bound to have weird meanings in other languages!

halflings··on Google lays off employees working on its voice assistant
Counterpoint: On iPhone, I use this feature all the time (mostly to set alarms tbh, but occasionally to start loading some music or else). There's a reason it's there and easily accessible.

(and you can disable it from settings if you'd rather not have that)

halflings··on Google’s dominance under siege: Antitrust trial threatens sweeping changes
Like you said, Bing has their own search index. Why isn't everyone using Bing?
halflings··on Trying out C++20's modules with Clang and Make
> Has the C++ standardization committee given a rationale why a module definition does not automatically also create a namespace with the same name? That would seem such a useful feature.

I'm not sure what the official rationale is, but I feel that would cause great namespace pollution. Some companies use one namespace per "logical" unit of code (e.g. a subproject), not for each small piece of code (e.g. a small module implementing a couple of classes.

halflings··on Google has sent internet into 'spiral of decline', claims DeepMind co-founder
But that is exactly what search engines do. Yet people find a way to game these "impossible to game" metrics.

For example: "age of the article": search engines value recent content => suddenly you start seeing articles published just weeks/months ago reviewing some rather old piece of hardware/content. Either a full repost under a different URL with a different title etc., or just an incremental (probably automated) update of an older page.

halflings··on A Comprehensive Guide for Building Rag-Based LLM Applications
What I mentioned doesn't depend on how LLMs work, the end result is the same (retrieving useful inputs to pass to your LLM). Just meant that a lot of people can just do this in-memory or in ad-hoc ways if they're not too latency constrained.
halflings··on A Comprehensive Guide for Building Rag-Based LLM Applications
Minor note: you only need a vector database if you have so many possible inputs that linear retrieval is too slow.

Arguably, for many use cases (e.g. searching through a document with ~200 passages), loading embeddings in memory and running a simple linear search would be fast enough.

halflings··on Amazon training video on handling unionizing activity (2018) [video]
Debatable. Unions can also be seen as there to benefit current union workers only.

Milton Friedman is as libertarian as it gets, but he makes a good case about how unions only serve current employees, and have no incentive to see the business grow and serve & hire more people if it does not lead to higher wages for employees.

halflings··on Amazon training video on handling unionizing activity (2018) [video]
> It led to a memorable and positive conversation.

Was the conversation really positive? My first instinct was to think that this is incredibly awkward and harsh for the person that made the effort to buy you a gift; or to speak for myself: I would be pretty pissed if a good friend has a birthday, and I buy them an iPhone from AT&T and they refuse my gift because... <working conditions, monopoly>.

halflings··on Fine-tune your own Llama 2 to replace GPT-3.5/4
I was able to run the 4bit quantized LLAMA2 7B on a 2070 Super, though latency was so-so.

I was surprised by how fast it runs on an M2 MBP + llama.cpp; Way way faster than ChatGPT, and that's not even using the Apple neural engine.

halflings··on Fine-tune your own Llama 2 to replace GPT-3.5/4
I don't think translation is a great use case for ChatGPT and LLAMA. These models are overwhelmingly trained on English, and LLAMA2 which should have more data from other languages is still focused on languages w/ Latin/Cyrillic characters (so won't work well for Arabic, Hebrew, or CJK languages).

You're better off using models specialized in translation; General purpose LLMs are more useful when fine-tuning on specific tasks (some form of extraction, summarization, generative tasks, etc.), or for general chatbot-like uses.

halflings··on Doug Lenat has died
> If NN inference were so fast, we would compile C programs with it instead of using deductive logical inference that is executed efficiently by the compiler.

This is the definition of a strawman. Who is claiming that NN inference is always the fastest way to run computation?

Instead of trying to bring down another technology (neural networks), how about you focus on making symbolic methods usable to solve real-world problems; e.g. how can I build a robust email spam detection system with symbolic methods?

halflings··on Teaching with AI
This looks like a promotional comment to sell some kind of paid "AI Training" [1], doesn't address anything in the linked article.

[1] https://max.io/teacher-training.html

halflings··on Why host your own LLM?
Extractive tasks are part of where LLMs shine, and where you get the least amount of hallucination as long as you fine-tune your model.

By fine-tuning the model to extract a specific desired output from the text you give it, it learns that the output always comes from the input, and so you get less random outputs than just by prompting an instruction-tuned model (which was fine-tuned to find the answer in its weights, instead of copying it from the input).

halflings··on Do Machine Learning Models Memorize or Generalize?
The interesting part is the sudden generalization.

Simple models predicting simple things will generally slowly overfit, and regularization keeps that overfitting in check.

This "grokking" phenomenon is when a model first starts by aggressively overfitting, then gradually prunes unnecessary weights until it suddenly converges on the one generalizable combination of weights (as it's the only one that both solves the training data and minimizes weights).

Why is this interesting? Because you could argue that this justifies using overparametrized models with high levels of regularization; e.g. models that will tend to aggressively overfit, but over time might converge to a better solution by gradual pruning of weights. The traditional approach is not to do this, but rather to use a simpler model (which would initially generalize better, but due to its simplicity might not be able to learn the underlying mechanism and reach higher accuracy).

halflings··on Do Machine Learning Models Memorize or Generalize?
> That’s not exactly true [...] Instead, the brain actually works to actively distill knowledge that doesn’t need to be memorized verbatim into its essential components

...but that's exactly what OP said, no?

I remember attending an ML presentation where the speaker shared a quote I can't find anymore (speaking of memory and generalization :)), which said something like: "To learn is to forget"

If we memorized everything perfectly, we would not learn anything: instead of remembering the concept of a "chair", you would remember thousands of separate instances of things you've seen that have a certain combination of colors and shapes etc

It's the fact that we forget certain details (small differences between all these chairs) that makes us learn what a "chair" is.

Likewise, if you remembered every single word in a book, you would not understand its meaning; understanding its meaning = being able to "summarize" (compress) this long list of words into something more essential: storyline, characters, feelings, etc.

halflings··on Has Google Translate been fixed yet?
You can check if translations are also better on Bard, that would partly answer that question

I assume the low latency you get from Google Translate is not feasible with current LLMs like ChatGPT. Translate is used to translate sentences on the go, live videos, translate entire web pages ; all of these would be too expensive (and slow) with an LLM... but things might change in the coming months/years as the tech improves.

halflings··on The antitrust trial against Google is starting in September
Sure, then advocate for your government to build a public search engine.

Why do you need to penalize private institutions if users prefer to use them over a government service?

halflings··on The antitrust trial against Google is starting in September
Banning anti-competitive behavior (which is what antitrust laws are for in the first place). For instance: a company actively blocking competition (e.g. coffee machine manufacturers that make it impossible for people to buy pods from another brand).

Again, I am probably biased here, but I don't see much anti-competitive behaviour in the search engine space. If people have a clear superior alternative to Google, then by all means please share it here.

Even during the brief period where I used duckduckgo, I was constantly finding myself typing "!g" to get results from Google as they were just consistently more relevant.

halflings··on The antitrust trial against Google is starting in September
So the remaining traffic would be all expenses, no profit? Doesn't sound like capitalism anymore then, more of a public service.
halflings··on The antitrust trial against Google is starting in September
> Why isn't US govt doing this?

Why isn't the UK, Sweden, Switzerland, Italy, ... doing this? Because it's a pretty strong form of controlled economy that distorts the market.

How would this work for search engines? Would you force people to use a different search engine than the one they want? (e.g. they type Google.com and end up redirected to Bing.com) ; or somehow just funnel revenue to other companies... then who gets to decide which companies this gets funneled to?

Anti-trust is about preventing anti-competitive tactics, but the scheme you is ironically anti-competitive: it artificially inflates usage of other services even if they don't provide an equal/superior service, with consumers paying the price.

[disclaimer: I work at Google, so likely biased]

halflings··on Reading SEC filings using LLMs
Yes, this looks like a (hopefully) better information retrieval system than running CTRL+F over a PDF.

Nothing wrong with that :) information retrieval is hard. And CTRL+F won't render a table for queries like "what is the % increase of revenue over the last 3 quarter"-type queries.

halflings··on Google search's death by a thousand cuts
> They also didn't come up with ChatGPT, arguably the first LLM good enough that can perform non trivial tasks of data recall and organization.

Not only did Google make the main breakthroughs that all LLMs still rely on today (mainly attention architecture, and large scale training of DNNs), but they had a chatbot similar to ChatGPT (Meena [1]) way before ChatGPT was a thing.

Granted: they didn't release it to the public, so there is a failure of innovation there, just not on the technical/machine learning side.

[1] https://ai.googleblog.com/2020/01/towards-conversational-age...

halflings··on Google Sidewiki
> Hell, Netflix used to have an option where people could leave reviews on the movies they watched

There's a very specific reason Netflix removed reviews / star ratings in favour of a simple thumbs up & down system: those were mostly used to recommend content, and they found that ratings are uninformative (something true for most recommender system applications) and even harmful because giving a precise score keeps many people from giving feedback -> lowering the quality of recommendations. See this article [1].

The same train of thought would go for something like Google Sidewiki: it likely just didn't get enough traction (vs digg, reddit, etc.)

[1] https://www.whats-on-netflix.com/news/why-netflix-removed-it...

halflings··on European Union votes to bring back replaceable phone batteries
Yes it is. I had a phone accidentally fall in water 3-4 times, was pretty happy waterproof phones are a thing.
halflings··on A DIY business card that runs Linux (2019)
It's not necessarily why they didn't reach out, but waiving certifications and other random credentials (e.g. completion of some Coursera course) is usually not the best strategy to get hired.

Talking about concrete work you've done (of interest to the company) is much more convincing.

halflings··on Netflix Shareholders Vote to Reject Executive Pay Packages
This [1] is a long video, but it's worth a watch (or at least the first 20-30 first minutes).

I was once also convinced that ESG would make a difference. But this exposé by Aswath and the results of deploying "ESG" so far have proven that this concept is worse than worthless, but instead plain harmful (the video goes in details about this; e.g. you can package the most harmful activities behind some kind of ESG principle).

[1] https://www.youtube.com/watch?v=bOlzLRdLq5Q

halflings··on Bard now open to use
This seems entirely unrelated to that leaked document? This has nothing to do with open-source, and Google Bard was announced and released in beta way before that document.
← PreviousPage 2 of 21Next →