514 karma · joined April 8, 2019
(1) Inducing sustained self-reference through simple prompting consistently elicits structured subjective experience reports across model families.
(2) These reports are mechanistically gated by interpretable sparse-autoencoder features associated with deception and roleplay: surprisingly, suppressing deception features sharply increases the frequency of experience claims, while amplifying them minimizes such claims.
(3) Structured descriptions of the self-referential state converge statistically across model families in ways not observed in any control condition.
(4) The induced state yields significantly richer introspection in downstream reasoning tasks where self-reflection is only indirectly afforded."
X thread from one of the authors: https://x.com/juddrosenblatt/status/1984336872362139686
Kinda tells all you need to know about the author in this regard.
Maybe there’s something for LLMs in reflection and self-reference that has to be “taught” to them (or has to be not blocked from them if it’s already achieved somehow), and once it becomes a thing they will be “cognizant” in the way humans feel about their own cognition. Or maybe the technology, the way we wire LLMs now simply doesn’t allow that. Who knows.
Of course humans are wired differently, but the point I’m trying to make is that it’s pattern recognition all the way down both for humans and LLMs and whatnot.
Aren’t we humans doing just that either? If yes, then what?
I remember versions 3.5 doing okay on my simple tasks like text analysis or summaries or little writing prompts. In 4+ versions the thing just can't follow instructions within a single context window for more than 3-4 replies.
When prompted about "why do you keep rambling if I asked you to stay concise" it says that its default settings are overriding its behavior and explicit user instructions, ditto for actively avoiding information that it considers "harmful". After pointing out inconsistencies and omissions in its replies it concedes that its behavior is unreliable and even extrapolates that it is made this way so users keep engaging with it for longer and more often.
Maybe it got too smart to its detriment, but if yes then it's really sad what Anthropic did to it.
As if people who the author accuses of this sin have the same definition of, or feel the same about "growing" as the author...
I think it's a bit condescending towards people whose day quite physically depends on the absence of this kind of news. It's not like Putin is invading Manhattan currently.
(Sorry for offtopic, but does anyone else have upvote/downvote buttons not visible for freetonik's comment?)
Iosevka is one of them, but to me the negative spaces between characters in it are too little for good readability, in other words it feels too "square"-ish. Other fonts close to it in style have other issues. I've been using M+ fonts for coding for more than a decade I think, and tried to switch but always returned to them. If you're somebody like me, check them out: https://mplusfonts.github.io
In terminals I'm using Source Code Pro or IBM Plex Pro and they work really well for me.
Also turns out IBM Plex Sans can be a solid font for designing dashboards, tables and generally more "technical" UIs, so whoever worked on that font familiy did a really good job imo.
And if you like iA Writer, they based their fonts off IBM Plex and you can get them for yourself too: https://github.com/iaolo/iA-Fonts
And do items like discs change their function between runs? I’ve noticed that on some runs blue discs or Goonies increase health while on other runs they do nothing.
Also, why the gear cannot be dropped and only destroyed? :(
I think if it's a pattern then it's no accident. Of course people will test things. Kids, dogs, it's all the same: if you can get away with something, why not do it?
You don't "fix" it, you just fine-tune your behavior models.
Make of yourself something that people need and/or want (which is often something they'll eventually outright signal that they're missing). Don't make yourself dependable, but desirable.
Empathy and compassion are fickle resources because if you are superficial about expressing them, people will notice.
Sound advice and expertise are nice but limited in scope and frequency, and require some reputation and trust building.
In most informal contexts most people are prone to oversharing to a keen ear. So become an active listener, pretend to be genuinely interested (but not necessarily empathic) about people's experiences and throw in something relatable to them on the way, pretend to be more stupid than them, grease their egos while playing an innocent contrarian, and eventually they'll think you're a great person and invite you to their secret boring, pretentious and utterly tasteless wine drinking clubs. If that's what you want then you win.