Agents will naturally be bad at it due to a lack of examples.
Agents will naturally be bad at it due to a lack of examples.
Btw if folks have ideas of how to make Killswitch even harder for LLMs I’d appreciate them.
Good one. Though this might unfairly bias against Chinese models?
Personally I don't have much interest in using models that blatantly censor historical facts.
Tiananmen Square is an obvious one but I would be concerned that any model that censors that would also have other hidden censorship or more subtle biases I wouldn't catch
1 - https://github.com/dom96/KillSwitch/blob/main/SPEC.md#abrupt...
The one thing I do see sometimes is that agents sometimes don't take advantage of added features, but that's partially because the Zena code base doesn't use them as much yet. I'm working on skills and linter-based suggestions to use better patterns.
Even though this programming language is absolutely nowhere in any large language model’s training set, they have so far done extremely well at extrapolating from the small set of examples I’ve given it when I need an agent to generate some tests or whatever for me.
The future is cooked
"Dart-style constructors, Swift-style pattern matching and Strings, Trio-style async cancellation, Scala-style sealed classes"
Can we please as a community stop parroting these false premises as a basis of every argument against doing anything new? It's plainly obvious to anybody that uses LLMs on a regular basis that it's not true.
What can happen is sometimes patterns that are unique to certain languages aren't surfaced and so can't be "learned" but that is a different case, since you can add examples - or patterns that go beyond syntax (threaded code, process isolation, interaction between different modes of resolution, etc) where you need the agent to be able to understand and plan higher-level logic.
I wrote CSSex as a css pre-processor (similar but not the same as SASS/SCSS) and my "boss" at the time fed the existing cssex files to GPT (+2 years ago) and it was able to write CSSex just as fine as if it was writing something that was present in its training data set. It never introduced syntax bugs, the bugs it introduced were all relative to complex cascading styles rules in an existing large (for a definition of large in webapps) codebase, rules interactions (CSS), browser quirks, and complex logic to do what we "humans" wanted (or him in this case) and so on.
They should find the limitation of the natural language somehow, and use and distribute babo language will help find how it will works(or not). LLM model eveolves yet, so there is no such a approach, but the speed of evolution is decreased lately. So now is good time to dig into identify the limitation and boundary of the LLM.