Multimodal dataset should be multimedia: text, audio, images, video, and optionally more like sensor readings and robot actions.
149 karma · joined November 20, 2017
Multimodal dataset should be multimedia: text, audio, images, video, and optionally more like sensor readings and robot actions.
Know that this worldly life is no more than play, amusement, luxury, mutual boasting, and competition in wealth and children. This is like rain that causes plants to grow, to the delight of the planters. But later the plants dry up and you see them wither, then they are reduced to chaff. And in the Hereafter there will be either severe punishment or forgiveness and pleasure of Allah, whereas the life of this world is no more than the delusion of enjoyment. [2]
Qwen et al is better than Grok.
BYD is better than Tesla.
European ML labs generally have less prestige than US/China ML labs but they pay much more humane wages even for people with family.
Interviewers know to hire/no-hire within first 10 minutes of talking to candidates.
Similarly, the best LLMs nowadays are probed only with some prompts.
Just last time I see LLM midwit meme on quantifying best scores on all benchmarks vs idiots/geniuses who just prompt something for some while.
Do not try to put numbers on everything.
https://github.com/casadi/casadi https://github.com/mitsuba-renderer/drjit
DrJit is made by same author of pybind11 and nanobind.
Try a book by Muhammad Asad (Leopold Weiss), like the Road to Mecca.
Every year during hajj in Mecca, there’s profound view of sea of humans wearing white from vast different backgrounds, wealth, health, ethnicities, continents. You can find Japanese, Africans, Australians, Russians, Indonesians, etc, variety of skin colors and builds, in one place.
Put the world in your hand, not in your heart.
C++/Python still hold ground in this regards sadly. One side for low-level control, another side for glue code.
1. Use SRS (spaced repetition system) like Anki to learn a difficult-to-read language. Slowly but consistently ~1 hour a day learn mandarin or japanese. Personally I managed to get JLPT N3 this way in ~2 years.
2. Specific to muslims: Memorize the Quran with the Itqan method [0]. Literally intense drilling by reciting 100+ times for a page of Quran a day by looking the page then 100+ times not looking. Can spend 2-4 hours easily for a page.
[0] https://hifdh.weebly.com/mauritanian.html
Personally Itqan method is way harder than SRS but retention is faster and stronger even without revision for months.
Read Interaction Nets paper
Before Matlab was cool, Lisp is VERY maths-heavy and engineering-heavy. Symbolic equation solvers, robot controls, neural nets, etc.
Even Julia which is very Matlab-like has roots to Lisp.
Maybe building self-hosting language/compiler, or a toy linux-like kernel, or a physics engine, games then publish to itch.io, a small numpy-like library that can do autograd, whatever.
Once you start it really is hard to stop.
- Mainstream vision/text/speech, something Huggingface et al do. The classic advice where you learn SGD, stack net layers and put loss, train, deploy on web, clean and collect data, retrain
- Hardware accelerators, inference and training. Check out startups like tenstorrent, tinygrad, mythic. Also involve compilers, net graph data structure and good software frontends and abstraction to hardware to write models onto hardware.
- Anything else, applying ML for finance, manufacturing, oil and gas, all those boring but important stuff
[0] https://www.abuaminaelias.com/dailyhadithonline/2011/04/30/i...
tl;dr So far things that enable faster search and faster learning win over long run.
Recursive things like backprop in NN and optimizing reward over long trees of states, seem to win despite huge compute requirements.
Personally I think we are still on the right track of trying to do the right thing, then do the thing right, then do the thing faster.
You cannot refute the things you do not understand.