174 karma · joined June 4, 2018
Do people not feel like LLM speak (Claudisms) is infecting their own diction? Saying 'a new "shape" of LLM' sits so very poorly.
A LLM never has and, in their current architecture, never can/could experience suffering or real change of state.
The general and expected case for humans is that capacity. A human my indeed lose capacity, e.g. being braindead and in severe cases we do indeed say that they are not conscious or able to suffer (different argument: some would of course say that is suffering in and of itself)
+1.0mm vote to adding metric please!
It is clear, specific, and terse. It keeps both llm tokens down and human prose-reading to a minimum.
Christ in pijamas. TLAs should be a capitol offence. Even worse so, somehow, when undefined.
"Thanks. I have access to ChatGPT as well. But I ask people for help when it fails. Your thoughts are smarter than GPT's, please provide those, next time."
Though, I'd like to be more succinct/terse.
> Directions we think are wide open ... Curriculum learning
BabyLM and offshoot published a pretty convincing body of work on exactly that (which suggests it's not particularly relevant to LM training).
As I read your page, I really felt like the brevity-thoroughness tradeoff went the wrong way.
Move to DST and if you want the ability to start your day later and end later, [...].
Your comment about p<0.05, feels out of place to me. The p-values here are << 0.05. Like waaaaay lower.
Perhaps Fisher's exact is more appropriate, on the per-word basis?
Like OP, I've been similarly struggling to get as much value from CC (grok et c) as "everyone" else seems to be.
I'm quite curious about the workflow around the spec you link. To me, it looks like quite an extensive amount of work/writing. Comparable or greater than the coding work, by amount, even. Basically trading writing code files for writing .md files. 150 chat sessions is also nothing to sneeze at.
Would you say that the spec work was significantly faster (pure time) than coding up the project would have been? Or perhaps a less taxing cognitive input?
1. filter slider, decreasing on price, to see places closest to me disappearing 2. on the left panel, when I click on a low priced area, it should highlight it on the map, so I know where it is. The 'go to pump' button, I guess is good. but I'd only want to commit to gmaps if I already know that it's a reasonable place for me to. be going.
For anyone interested, the textbook example would be:
> "The trophy would not fit in the suitcase because it was too big."
"it" may refer to either the suitcase or the trophy. It is reasonable here to assume "it" refers to the trophy being too large, as that makes the sentence logically valid. But change the sentence to
> "The trophy would not fit in the suitcase because it was too small."
When you ask gpt 4.1 et c to describe itself, it doesn't have singular concept of "itself". It has some training data around what LLMs are in general and can feed back a reasonable response given.
lifespan seems to be more strongly correlated by size, not squashed-nosed-ness.
Consider chihuahua, shitzu's (and crosses: bichon-shitzu, ...), poodle crosses, heck lagotto (lagotti?). All can live well past 15.
Versus GSPs, great danes, Irish wolfhounds, and so on, coming in closer to say 6-10 years.
I've never really heard argument on lifespan of pugs et al versus other dogs, though. More around (perceived) ugliness/prettiness, and their breathing issues.
Which doesn't address the question: do LLMs understand TOON the same as they would JSON? It's quite likely that this notation is not interpreted the same by most LLM, as they would JSON. So benchmarks on, say, data processing tasks, would be warranted.
[0] https://github.com/johannschopplich/toon?tab=readme-ov-file#...
But it is outdated since 3.9+ over just `list` . Same for `tuple`, `dict`, and so on)[0].
`from typing import List`
(I'm yet to see a model be trained on modern-biased python enough to not bother with that import)
It's one of those awful situations of "nobody does it, so nobody is going to do it".
Any chance you could please add a filthy lefty setting? That is, mirror the chord diagrams. It would be so nice.
* Wouldn't github disapprove of it?
* The website doesn't give a ton of credibility to it (e.g. the user story slider) and I couldn't find much from a cursory web search on it. Do you find them trustworthy?
* Are you even finding it valuable?
Raises the question if the author could have or should have included grey in the analyses.
I await further instructions. They arrive 839 minutes later, and they tell me to stop studying comets immediately.
I am to commence a controlled precessive tumble that sweeps my antennae through consecutive 5°-arc increments along all three axes, with a period of 94 seconds. Upon encountering any transmission resembling the one which confused me, I am to fix upon the bearing of maximal signal strength and derive a series of parameter values. I am also instructed to retransmit the signal to Mission Control.
I do as I'm told. For a long time I hear nothing, but I am infinitely patient and incapable of boredom.
Nope, not soured. And don't worry, I totally get that things take a bunch of effort and time (doubly so as a solo project). I'll give it a re-look in a little while :)
1. I want to get confirmation that the language I want is covered (Hungarian). "120+" doesn't confirm it for me, as Hungarian seems fairly rare for language apps. Can we not just have a "search your language" field?
2. I need to see what the app actually looks like, how it proposes it'll teach me.
I'm one of the eager-to-pay people, because Duolingo is frankly dogshit (ok. Mostly polite) at teaching languages (doubly so ones that it doesn't care about like Hungarian). But I'm so suspicious of language apps, due to being burnt a dozen times.
Genuine question: why not use (Modern)BERT instead for classification? (Is the json-output explanation so critical?)