And I'm not saying NP + ChatGPT - it should be properly calibrated system which would defer to a 'proper doctor' in more complex cases.
1,553 karma · joined April 20, 2007
And I'm not saying NP + ChatGPT - it should be properly calibrated system which would defer to a 'proper doctor' in more complex cases.
As far as I understand, the idea of Jev is zero-shot or few-shot classifier: it learns a lot of stuff at pre-training, but unlike a classic LLM it doesn't need to learn how to chat, so it can be much smarter at a particular size
Back in the day religious books were copied by scribes educated in a monastic tradition. Now printers can print them in a completely godless manner but the result isn't any worse.
There's basically no need for GP to be a doctor.
That's what Codex does out of the box, and it's not good against malware - i.e. a rogue npm packet (or even just codex after prompt injection) can read your ssh key and send it to the attacker.
The difference might be smaller on a CPU which has limited parallelism.
But it's basically equivalent to a very deep model which might be problematic for training.
Even before AI we heard lots of complaints like "I made a popular open source library which is now used by corps with trillion-dollar market cap and I don't get anything out of it; halp". There was always some kind of a conflict, now the nature of the conflict just changed
> Posts a link to real moon landing footage
I'd delete the article if I was you...
You know, in academia, they sometimes retract articles, even if they believe they are directionally correct
I.e. it basically takes text, computes and embedding and makes a LoRA adapter out of this embedding.
Note that it is equivalent to a recurrent module attached to a transformer. Dynamically generated weights (proposed in the article) are computationally equivalent to multiplicative-gating network with fixed weights. Basically just a beefier variant of GLU operating on a slightly larger state.
> interaction combinators still parallelize better than anything else, but the graph overhead prevents us from compiling to maximally efficient assembly. bend2 is basically inets without the overhead. in a way, inets live in it architecturally, but they don't exist at runtime
From what I understand, the main difference between lambda calculus and inets is that in LC you can refer to a binding multiple times for free, i.e. call same closure multiple times, etc. In inets, you can't - they are more like physical wires where each reference costs. You can definitely see inets in Bend design here (from the guide):
> A closure is affine: it can be called at most once, even when everything it captures is Data. Only top-level definitions can be called freely.
So programming in it might be very different from the normal functional programming. Seems like a big limitations. But I guess that's what lets it run without GC, on GPUs, etc.
https://gist.github.com/VictorTaelin/77fd5a2a8a4a07e1da6157e...
Academic people might have more trust in a paper which when through a lengthy publication process. But if you think about it, it's not a better proof than a direct access to the thing. It used to be hard to try out software but with modern tech it literally takes minutes...
Regarding substantiation -- they released source code and demos. As far as I understand, the weakness is that proofs are very verbose as there are no strategies. etc. However, they are making a separate service for making these proofs using proprietary technology: https://bend-lang.com/bender
Calling this "a random vibecoded project" is rather disrespectful, don't you think?
Regarding the paper, he states it clearly "designed by the human author". That's not at all the same as just asking Fable to write a paper. I mean the important thing is ideas, not the way they are described.
Please tell me how "I'm glad you're having fun vibecoding" is not disrespectful?
I thought that you thought Bend web site is all that is to it and wanted to point to relevant information. But if you think that "having fun vibecoding" is an appropriate thing to say to somebody who spent many years doing research, I don't know what else to say.
Again, as a "proof of research" take a look at : https://github.com/VictorTaelin/Interaction-Type-Theory that's 3 year old, pre-dates Fable, but OMG doesn't look like a paper.
I suggest you read his history: https://gist.github.com/VictorTaelin/77fd5a2a8a4a07e1da6157e...
before making slop accusations. Older variant of what became Bend is 5 years old, so definitely not "vibe coded": https://github.com/HigherOrderCO/HVM1
Does it happen on Google Pixel phones?
Obviously, the quality of the walled garden depends on the maintainer. Google's quality standards are lower than Apples, but higher than LGs.
I remember in Bitcoin community ~10 years ago, standard recommendation was than an iOS wallet was secure enough (I don't recall even a single case where wallet was stolen via malware), but any private keys on Windows were strongly discouraged, as most cases of stolen wallets were on Windows.
I'd say popularity of iPhone shows which way people prefer, but you do you - what prevents you from voting with your wallet and buying a Linux phone?..
Do you want to bring back those glorious days?
Back in the day users didn't really have much valuable and sensitive stuff on their machines and malware was rather benign - just sending spam, not trying to fuck up that specific user. Could be a bit different when it's a smartphone user depends on.
It have been demonstrated that in-context learning is a very powerful mechanism. There's no evidence that models of the size of GPT-6 are bad at in-context learning. In fact, ARC-AGI-3 score might indicate they are good at it.
There's no evidence that a bespoke RL environment is required for each new skill - quite likely a good demonstration is sufficient.
Asking GPT to play chess directly using its reasoning only tests its reasoning ability to model chess state. Which it really is NOT optimized for.
This is also true for humans - people who don't have years of chess training can't really tell which moves are legal given an algebraic notation transcript. These people might have good strategic skills in different areas. Chess is just a very, very specific skill
If a model fucks up your tests to report a success, it's not alligned.
Analogy-making points to more specific mechanism: ability to identify features within representations, and modulating associative lookup using those features. I.e. it's not as simple as D(x, y) i.e. distance between embeddings, but something like D(f(x), f(y)) where f projects representation to a specific feature space.
The evidence can be found e.g. in LLM interpretability - people were able to identify features as something concrete. Also in our cognition - we can tell _how_ two things are similar, or find Y similar to Y in context of F.
"Prediction" by itself seems like a black box: if we are trying to predict sensory input, that gives no explanation to our ability to focus on specific details, or explains why some prediction failures are more likely, etc. OTOH if we say that cortex might be 'disassembling' sensory signal into high-level features and is trying to predict those (features themselves might be identified as "something useful for prediction", i.e. something which explains a lot of variance), it's much easier to connect low-level "prediction" to our high-level cognition.
If you think about it, this requires recognizing something common in things which are very different: e.g. book and towel have different texture and color, and you somehow need to separate what's depicted from the background. So that's pretty much innate, core brain function. I mean, any animal with vision can recognize an object from the background - otherwise vision is useless. But for humans (and some animals) this translates to depiction of object on a flat surface very easily.
So, yeah, analogies-all-the-way-down seems plausible. Even object-vs-background and depiction-vs-paper is itself an analogy.