16 karma · joined February 9, 2011
There's a language called Kalamang with only 200 native speakers left. There's a set of grammar books for this language that adds up to ~250K tokens. [1]
They set up a test of in-context learning capabilities at long context - they asked 3 long-context models (GPT 4 Turbo, Claude 2.1, Gemini 1.5) to perform various Kalamang -> English and English -> Kalamang translation tasks. These are done either 0-shot (no prior training data for kgv in the models), half-book (half of the kgv grammar/wordlists - 125k tokens - are fed into the model as part of the prompt), and full-book (the whole 250k tokens are fed into the model). Finally, they had human raters check these translations.
This is a really neat setup, it tests for various things (e.g. did the model really "learn" anything from these massive grammar books) beyond just synthetic memorize-this-phrase-and-regurgitate-it-later tests.
It'd be great to make this and other reasoning-at-long-ctx benchmarks a standard affair for evaluating context extension. I can't tell which of the many context-extension methods (PI, E2 LLM, PoSE, ReRoPE, SelfExtend, ABF, NTK-Aware ABF, NTK-by-parts, Giraffe, YaRN, Entropy ABF, Dynamic YaRN, Dynamic NTK ABF, CoCA, Alibi, FIRE, T5 Rel-Pos, NoPE, etc etc) is really SoTA since they all use different benchmarks, meaningless benchmarks, or drastically different methodologies that there's no fair comparison.
[1] from https://storage.googleapis.com/deepmind-media/gemini/gemini_...
The available resources for Kalamang are: field linguistics documentation10 comprising a ∼500 page reference grammar, a ∼2000-entry bilingual wordlist, and a set of ∼400 additional parallel sentences. In total the available resources for Kalamang add up to around ∼250k tokens.
I think the answer to your question is more situational. For example, the ML community routinely publish supplementary material on Github. While the introductory section of these repositories should be accessible to everyone, they do also go fairly deep with technicals before referencing their respective publication for more details. The PL community also routinely use Github as a distribution platform. For them, it's often much more compact to give the semantics of their language-extension as a system of equations rather than as words. This doesn't absolve them of the obligation of giving readable examples of the specification, but for some of these more technical repositories, it's a nice-to-have as well.
Having said all of this, the ML community has thought up of a rather clever workaround for this already. Github renders ipynb notebooks that has LaTeX in side of its Markdown sections, so many READMEs just reference a Jupyter notebook. This is just a different approach on that problem.
On the other hand, giving the model with more options than it necessarily needs and letting it decide what is important will usually backfire. Rather than learning a few meaningful/functional features, it can just go ahead and completely fit the training data from the very beginning. It will therefore decide that everything is important, because all those extraneous parameters will let it squeeze that last 0.5% out of your training set.
For example, in the subproblem of only one-way streets that we can assume to be composed of one or more segments of straight lines per street, we take a simple street-function $\phi_i(x,y)$ corresponding to street i to be the norm of the projection of <x,y> onto that street. Furthermore we also add into the system some "smoothing-function" to ensure that the overall shape of the final path is doable, for example constraining the distance between successive points. Next, we solve the argmin of the norm equation for each point so that each point is now moved to some linear combination of the streets, and truncate all but the most significant street basis, and rerun until we get to some acceptable tolerance.
Maybe?
According to Web analytics group Net Applications, Apple's Safari Web browser posted an 8.1 percent jump in market share for the month of July
http://marketshare.hitslink.com/report.aspx?qprid=0&qpca... tells me that Safari has a total aggregate of 8.1 percent of the market share in July...