1,187 karma · joined February 13, 2010
"We hypothesize that this phenomenon emerges predominantly from a misassumption about how these systems are trained. Modern multimodal models are developed on web-scale corpora and are commonly built on top of pretrained large language models, which makes them extraordinarily strong at language modeling, retrieval of statistical regularities, and reconstruction of likely contexts from sparse cues.[48, 25, 24] During the multimodal training, the models are presented with the image, a textual question, and are expected to reconstruct the correct answer. Lacking access to an entire text corpora, a human would intuitively answer the question based on the image in that setup; but we should not infer that this would be the default approach for an AI model. Incentivized to generate the correct next tokens, models might learn to easily ignore the visual information and rely only on their vast prior knowledge, taking the shortest route to the correct answer.[36, 5, 48]"
The crazy thing is that based just on the text in the questions a model was able to "guess" answers:
"When fine-tuned on the public training set of this dataset with images removed (i.e., trained in mirage-mode), our 3-billion-parameter, text-only super-guesser outperformed all frontier multimodal models, including those exceeding hundreds of billions of parameters, on the held-out test benchmark (Figure 3c). It also surpassed human radiologists by more than 10% on average, relying entirely on hidden textual cues in the questions and the structural patterns of the benchmark. In addition, our super-guesser was able to create reasoning traces comparable to, and in some cases indistinguishable from, those of the ground-truth or those generated by frontier multi-modal AI models."
https://open.substack.com/pub/samkriss/p/numb-at-burning-man
The point is to demonstrate "we are not alone in this feeling", that's it...
https://www.thriftbooks.com/w/truck-on-rebuilding-a-worn-out...
[1] https://www.ribbonfarm.com/2009/10/07/the-gervais-principle-...
I guess that’s socialism though, so not gonna happen.
https://www.dir.ca.gov/dlse/callbackandstandbytime.pdf https://www.shrm.org/topics-tools/tools/policies/california-...
This doesn't apply to a lot of people reading, but just a PSA for those in CA where it might.
SF seems to be a lot more in-flux compared to other cities, so if you don't like the scene now just wait a few years and a new one will be along :-)
Maybe it's the other way around? Our minds work this way because that's the rules and we're just emergent from that reality. It's hard to argue that symmetry is not a mathematical ideal first, and a biological approximation to that ideal second.
"If ve ever wanted to be a miner in vis own right -- making and testing vis own conjectures at the coal face, like Gauss and Euler, Riemann and Levi-Civita, deRham and Cartan, Radiya and Blanca -- then Yatima knew there were no shortcuts, no alternatives to exploring the Mines firsthand. Ve couldn't hope to strike out in a fresh direction, a route no one had ever chosen before, without a new take on the old results. Only once ve'd constructed vis own map of the Mines -- idiosyncratically crumpled and stained, adorned and annotated like no one else's -- could ve begin to guess where the next rich vein of undiscovered truths lay buried."
I'm sympathetic to the idea that there is bloat and over-regulation, but most of the laws and regulations on the books are in response to a bad thing happening. It's kind of like a legacy code base - just deleting the repo and starting over from scratch is usually not the best idea, it takes some careful refactoring and judicious tests to move in the right direction.
It's sort of an interesting idea - what are "tests" in the context of the legal/regulatory framework? The constitution? The judiciary?
I learned assembler by typing in listings from magazines and hand dis-assembling and debugging on paper. Your approach seems similar in spirit, but who has the times these days?