Having models attempt an SVG letter S remains one of my personal/informal LLM benchmarks. They are still pretty bad at it.
1,888 karma · joined September 12, 2014
Working on FOSS and user-friendly alternatives to things like khanacademy, anki, MathAcademy, Alpha School, etc.
Modern, open edtech tooling.
Also http://paritybits.me
Having models attempt an SVG letter S remains one of my personal/informal LLM benchmarks. They are still pretty bad at it.
EG, my own oldest child needed a surgery at birth that would have been logistically impossible even 50 years ago. I'd say that she and I have benefited enormously, despite not being billionaires.
edit: I solemnly swear that the sibling comment with the strikingly similar "impossible 50 years ago" claim is a pure coincidence and that I at least am not a bot campaign. Haha.
Even with LLMs, posts like this don't just fall out of a coconut tree. If you have a set of target benchmarks for your own model, then keeping "the set" of side-by-side comparable models is its own maintenance headache.
Pol here is abbreviated politician.
Practically nobody teaching K-12 has subject-matter masters degrees. It's just not part of the career trajectory. As unusual as a nurse having an M.A. in history or something. Yes, would occur on the margins of people changing course in life, but not the mainline.
Specifically, the question here is about the efficacy of pay-scale bumps for Masters degrees in education. To your point (and my counter-point), teachers get a substantial pay bump* if they hold a M.Ed, but no bump if they hold a masters in their teachable areas.
For persons who can afford it in the moment, taking a one year or two or three year part-time M.Ed. after getting a few years teaching experience (an entrance requirement in most M.Ed. programs) can pay for itself over the next 2-5 years, then is all surplus for the rest of the career.
* - all of the varies a bit by jurisdiction but I think this is "the general case".
I'm about as pro AI-as-a-research--and-writing-assistant and anti AI-witchhunt as they come, but I simply cannot parse what I've quoted here.
Posting slop to arxiv is blatant deception. Posting an article is an attestation that the article is a genuine engagement with the literature. If you're posting things to arxiv that are not sincere engagements with the literature, you are attempting to deceive.
EG: How did Mark Zuckerberg make software five years ago?
He's as capable of opening up an editor as I am, but circumstance had offered him a different interface in terms of human resources. Instead of the editor, he interacts with those humans, who produced the software. This layer between him and the built systems is an abstraction, deterministic or not.
Today, you and I have a broader delegation mandate over many tasks than we did a few years ago.
1. It's amazing how strong the obvious in hindsight is for this research. LLMs have been (rightly) characterized as inscrutable black boxes. If only there were some discipline for learning and extracting semantics from information dense payloads ... !?
2. NLAs seem to be in the ballpark of a safety and interpretability standard that is both enforceable (easy?) and plausibly effective (probably hard to prove definitively, but easy to believe at least partially).
3. NLAs here are trained against the residual stream of a model at some layer (N). It would be interesting to see a sequence of NLAs against a staggered set of layers. There may be a semantically meaningful evolution of 'thought' going from the early to late layers.
4. I would love to see this technique applied against tokens across boundaries of model 'aha!' moments (to what extent is the 'aha' an affectation, or is there actually a sharp turn in the understandings?), and jailbreaks / personality snaps [1].
It'd be quite a coincidence if the training runs discovered an invertible weights>text>weights function that produces text that both "is on topic and intelligible as an inner monologue in context" and also is unrelated to meaning encoded in the activations.
EG, could a misaligned model-in-training optimize toward a residual stream that naively reads as these ones do, but in fact further encodes some more closely held beliefs?
If they are co-trained only on activationWeights->readibleText->activationWeights without visibility into the actual stream of text that the probe-target LLM is processessing, then it seems unlikely that the derived text can both be on-topic and also unrelated to the "actual thoughts" in the activationWeights.
If it's an echo chamber of AI hatred, then I think this makes the case that there are substantial numbers of people in that camp also?
AI as a product is bimodal in terms of the opinions people have of it.
Local subreddits are filled with posts "calling out" usage of AI by local businesses or governments. Consensus is that persons who are found out to be be AI users should be fired or resign, businesses that use it should be boycotted / shamed, etc.
https://www.reddit.com/r/newfoundland/comments/1t3x6q3/aialt...
https://www.reddit.com/r/PEI/comments/1s8rtyn/burger_love_ai...
---
Protests around a data center construction project: https://www.youtube.com/watch?v=11Q9ncOdnDg
Every sniffed out systematic service overcharge can be aggressively undercut by competition.
"Your margin is my opportunity", etc.
> ... If they have your message history and see your mum is dying, they can spike flight tickets for example. And they will know exactly the highest amount you can afford for it.
In my own head, this was the the most far-fetched piece of my own comment!
> Cameras on the aisles as well can enforce that individual tags update while nobody is within 15 feet, etc.
We too are amalgamations of inanimate components - emerged superstructures.
Just cells. Just molecules. Just atoms.
Canada's major grocery chain has migrated entirely to LCD price tagging that can receive OTA updates. There are now no paper price labels in the store.
The same chains have extensive camera coverage on the entrance / exits of the store.
So pricing can be an optimization function as fine grained as persons currently in the store.
Cameras on the aisles as well can enforce that individual tags update while nobody is within 15 feet, etc.
It's hard to even talk or think about without without sounding (and becoming!) conspiratorial. Add a little data from our trusted partners and they can jack specific prices according to urgency - eg, floral bouquets when you're en route to a dance recital.
the article here isn't about the LLM recognizing works that were in the training data. EG, The Old Man and the Sea off the shelf. It's about pegging the author of novel texts, like, say, some letter written by Hemmingway that gets discovered next week and was never before digitized.
There general interest across a variety of disciplines to kick the tires of LLMs with respect to their competence in DOMAIN_X. This is good in general terms, but, especially with larger studies, they tend to be out-of-date by the time of publication, and super out-of-date by the time they hit the media circuit. Out-of-date here in terms of testing against models 1 or 2 or more generations back from SOTA.
The DOMAIN_X experts do have a lot to offer in terms of defining success criteria across domain tasks, but the studies (snapshots in time) could be much more impactful if they were instead packaged as benchmarks (that could track model progress over time, and even steer it).
AI community / industry could probably do some outreach work to streamline or standardize methods for general researchers to produce reusable benchmarks.
The reported variance in Sonnet 4.6's estimates here are actually quite low, and in general terms, not so bad across models. Damn paella.
This does seem like a task well suited to a for-purpose training run against a bunch of labelled data. Is there any reason they wouldn't improve at it?
It was clear (see the linked post from 70 days ago) that the current offering was unsustainable, but I'm a bit taken aback at how sharp the clawback is.
Right now the dashboards show 78 providers online, but someone in-thread here said that they spun one up and got no requests. Surely someone would be willing to beat the posted rate and swallow up the demand?
I expect this is a migration target, but a tactical omission from V1 comms both for legitimate legibility reasons (I can sell x for y is easier to parse than 'I can participate in a marketplace') and slightly illegitimate legibility reasons (obscuring likely future price collapse).
Still - neat project that I hope does well.
[1] Layer Labs, formerly EigenLayer, is company built around a protocol to abstract and recycle economic security guarantees from Ethereum proof of stake.
You - Darkbloom - Operator - Darkbloom - you, vs
You - Provider - you
---
On the censorship point - this is an interesting risk surface for operators. If people are drawn my decentralized model provisioning for its lax censorship, I'm pretty sure they're using it to generate things that I don't want to be liable for.
If anything, I could imagine dumber and stricter brand-safety style censorship on operator machines.