Copy is all you need
arxiv.org
arxiv.org
So if the implication was that no language model was needed at all and you can just do nearest neighbour on string similarity and patch results together, that implication was clearly wrong.
I think what the paper does show though is that there are methods that can make language models topic-specific without fine-tuning and that yield competitive results even with older models.
Also the fact that evaluating language models is difficult, and we tend to end up with models that game the evaluation benchmarks.
Of course a more explicit approach like this paper is a really good step in that direction by making it easier to trace information provenance. It might still be nontrivial to answer why the model selected this specific piece of information, and why it was composed in this specific way, but it seems trivial to say where the model got the information from. Which is really all we demand from humans too.
Google search once was great too but then ads and SEO killed it.
Quantum computers would have something to say about this, assuming they ever materialize.
Faith and Fate: Limits of Transformers on Compositionality https://arxiv.org/abs/2305.18654
Transformers solve compositional reasoning tasks by reducing multi-step compositional reasoning into linearized subgraph matching without problem-solving skills. They can solve problems when they have reasoning graphs in the memory.
however, the "parenthesis" can be any symbol. even grammatical clauses are one sort of "parenthesis" in the way I'm thinking about them
in fact, my whole idea has got me on a deep dive into the nature of the decimal point (up to which extent is the decimal point representation of numbers and instance of a "fixed point"? I don't know! I cannot understand a fix point just yet; and for me to say I get decimal notation actually means I understand something about p-adic representation; which I'm still working on figuring out)
I thought these models got more 'logical' after training with computer code
Even more than that, I wonder if we could then apply something like this to power some sort of "fact provenance" for the web as a whole, e.g. by populating Wikidata with referenced facts (preferably with extensive human QA).
And systems that allow you to "talk" to a PDF via top results of vector search being added to the prompt are also pretty underwhelming.
[1] https://www.urbandictionary.com/define.php?term=Dickstracted
I once worked with a programmer who, the vast majority of time, would only input text into a text editor via copy and paste.
Think anti-vim. His fingers were locked on mouse and crtl+c/v. It was incredible to watch and his programming speed was very impressive.
Also explains why he was so fast
Just in case you're thinking this: He was not copying large portions of code from stack overflow or anything like that. He was line by line writing code, a few copy and pastes at a time. Often he times would copy and paste single characters to maintain his flow.
In the case where you are copy-pasting out of code you don't really understand. Retyping it gives you time to understand and maybe catch existing bugs in the code you are copying.
My way of doing imperative coding for data science with Python is to write a price of code in Sublime Text, copy and paste to iTerm, run, and get back to the editor. But of course I mapped shift+Enter to do all of that for me. I much prefer this setting to Jupyter Notebooks.
I just alt-tab from editer to terminal, check output, etc, and back. That way I have a bunch of unix text-processing tools (grep, sed, etc...) always available. I'm too reliant on print debuging things as I go along, but it's a deeply ingrained habit.
I get people have different workflows, but not taking advantage of even the minimalist functionality of ones tools I think I will never understand.
Sometimes he would need a portion of a word and he would remember that it was in an email he had open, and he would alt+tab and grab the portion of the word from the email, then alt+tab back to the editor and paste the word portion in.
He would go to extreme lengths to not have to move his hands to the home row on the keyboard.
I will try to implement this with the necessary changes to actually make this work properly, where instead of generating a new answer, it simply highlights the most likely text spans.
(Not sure if the authors have indicated any method for attribution of the original data)
"Set your model temperature as high as possible an generate a completely new and random word"
It acted acted like it understood and generated the word Blazivox. I don't see it on Google at least.
There's some neural / patch blends from 2016 that I always thought were interesting (CNN-MRF) [1], and I think there's a renaissance in those approaches recently (combined with other generators / prompts etc.). You can also argue ViT is "patch based" in a major sense... I am still a big believer in patch + combinations + warping (non-parameteric synthesis) generally, some cool older work from Apple on that in speech land [2].
I go as far as arguing BPE / wordpiece / sentencepiece / tokenizers in general are key for modern approaches (as were word vocab selections in the earlier days of NMT), because they find 'good enough' patches (tokens) for a higher level model to stitch together while still having some creativity / generalization available... but we focus on the model details rather than the importance of the tokenizer (and tokenizer distribution) in publication many times.
[0] http://people.eecs.berkeley.edu/~efros/research/quilting.htm...
COG stands on the line of retrieval-augmented text generation research but takes a radical step forward. Unlike previous work that combines retrieval and generation, in COG, retrieval is generation.
COG shares some ideas with previous work such as replacing the fixed vocabulary with a nonparametric phrase table.
The paper presents experimental results showing the advantages of COG over strong baselines in three experimental settings: standard language modeling (using the WikiText-103 dataset), domain adaptation (using the Law-MT dataset), and an enlarged phrase index (using the En-Wiki dataset).
Despite the promising results, the authors acknowledge that there are some flaws in the COG method. For example, COG may copy a phrase that is incoherent with the previously copied phrase, or it may only copy a part of a complete phrase, leading to inaccurate generation results.
:)
> The dominant sequence transduction models are based on complex recurrent or convolutional neural networks that include an encoder and a decoder. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.
https://proceedings.neurips.cc/paper/2017/file/3f5ee243547de...
Second, this approach seems equivalent to using larger tokens, which means the problems with using tokens instead of letters are just exacerbated