HNHacker News
TopNewBestAskShowJobs

fennecfoxy

1,291 karma · joined June 29, 2021

Gay furry. I swear a lot; it's cultural.
submissionscomments
fennecfoxy··on 4,400-Year-Old Tomb of Egyptian Judge Found at Saqqara with Colors on Walls
Well yes, there are only so many attention heads (well whatever magical variant frontier models are using these days) that can attend to the context and so as the context grows attention becomes spread thin.

But with reasoning enabled I find that even with a large context that induces mistakes most of the time a decent model realises and corrects itself before output.

And it's always been known that prompting what it should do is far better than what it shouldn't, since just introducing "DON'T do X" into the prompt means that the tokens for X are present and can be paid attention to in the wrong way.

But even then I've used plenty of "Do X, not Y" recently, especially for tools "Use this for x, don't use this for Y" and models perform like 90% of the time.

I would be interesting to experiment to see performance curves given a restriction on reasoning tokens allowed to n% of context tokens and see if there's some magic number of "reasoning should be at least n tokens for a context of length p" even ignoring the complexity of instructions in the prompt itself.

fennecfoxy··on 4,400-Year-Old Tomb of Egyptian Judge Found at Saqqara with Colors on Walls
Interesting. I agree that it seems that LLMs are wont to fall into recognisable patterns much more readily than humans however I asked a model to evaluate your comment history and it found loads of patterns such as:

“I’m not sure…” / “I don’t…” hedges before disagreement Negation followed by correction/reframing is a recurring structure: essentially “It’s not X, it’s Y”.

And dozens of other examples. I wonder if LLMs really are more liable for it or if it simply seems that way purely because when you use a model it's almost as if you're speaking to the same "person". I think that human language is naturally formed by common agreed meanings of words first, and then phrases.

fennecfoxy··on 4,400-Year-Old Tomb of Egyptian Judge Found at Saqqara with Colors on Walls
For sure. If you write with a casual conversational tone people don't get triggered. But if you are eloquent and use an extended vocabulary thanks to large amounts of reading then that seems to make people think you're using AI.

As someone who loves to read (well, on and off) I wonder what the median vocabulary is and how I and other readers compare. I guess language has always been functional for most people.

fennecfoxy··on Harnessing the Universal Geometry of Embeddings
Ah right, thank you for clarifying.

But I think my point still stands, isn't the geometry information THE information I referred to in the first place? Obviously the vector size gives you the granularity but it's kind of unavoidable to positionally encode information in a latent space...that's literally what they're for?

But yes, it is very cool to know that regardless of exact implementation finding x,y,z representations of some dataset with various relationships (like language) creates similar geometry/clues across all the implementations.

fennecfoxy··on Why AI-generated slopaganda works
Eh people don't care if something is true or false they care about the dopamine hit they get from engaging in tribal behaviours that reinforce their freaky belief systems that often aren't rooted in any kind of reality.

Humans are animals. We are!

fennecfoxy··on Harnessing the Universal Geometry of Embeddings
I'm not as heavy on the maths stuff involved in this as other people commenting appear to be.

But the idea makes sense, of course there is still recoverable data in embeddings, that's the point. Though as I constantly find the more you try to squeeze into an n bit vector the more watered down everything gets.

I suppose a latent space could be encrypted/mapped in some way to resolve that, but how many people are exposing their vectors in the first place?

fennecfoxy··on Ask HN: Fable hacked my piano, can I release the results?
IANAL, especially not an American one.

But if you're worried just pop it on anon GH.

Besides the fact that this sort of protection through obfuscation is dead now anyway. If you can ask an LLM to do it so can I, or anybody else. The only downside is duplication of work/wasted tokens but eh.

AI has already started commoditising software. Hopefully we see more OS' lean into the "safe" layer that runs everything and then temporary/custom interfaces dynamically created by AI on top.

fennecfoxy··on Ask HN: Fable hacked my piano, can I release the results?
Tbf "is it a crime" is hard for even a single lawyer to answer because it depends on: who you are, your skin colour, how rich you are, your sex, whether it's a white collar crime or not, did you commit the crime on behalf of a corpo, etc.

But we like to pretend that the justice system delivers justice evenhandedly I suppose.

fennecfoxy··on New Worlds: We are living in the future of J.G. Ballard or William Gibson
Talking to members of the opposite sex only? How narrow-minded. Cyberpunk/sci-fi often explores futures where sex/gender is less of an issue (especially Ian. Banks' works). Clearly we're not there yet.
fennecfoxy··on Dutch central bank moves 86 tonnes of gold from US citing 'geopolitical unrest'
So America thinks its above the ICC then?

Explains a lot, lmao.

fennecfoxy··on Apple caught off guard by AI demand for Mac Mini and Mac Studio
The Internet has magnified the "my church leader said that gay people are actually demons" effect. Or "if you throw salt over your shoulder you're protected" or "wolves have alphas of the pack".

People don't fact check in the first place, they don't check that facts came from a trusted source, they don't check if information has changed the last time they got it from a trusted source, etc.

Is it just me or is it scary that it seems to me as if the average person just takes everyone's word for most things, as long as the person telling them is a "trusted source" based only on tribal lines?

fennecfoxy··on 'Mad honey' that can stop your heart is being sold online
I don't think that that issue is inherent to women. I think it's inherent to everyone and we get affected on different levels in different contexts.

Any man who's been told to "man up" (most) has experienced the same.

fennecfoxy··on The safest job from AI may be writing
Given what I have seen in recent Hollywood "writing" I would not be surprised if when correctly prompted an LLM can outdo a lot of the disposable, predictable writing in modern shows.

Every story boils down to the same predictable elements, even so-called twists are just repeats of the same old twists. Every story is boy-meets-girl, the friendly character at the start is the real villain, it was all just a dream, the power of friendship, etc.

I just asked GPT5 medium for something along those lines and got (condensed): humanity discovers a phenomenon in that the universe has started compressing causality, things start disappearing because they did not have a large enough influence upon the universe - someone's childhood hobby, a friendship that was tenuous at best and their journey is how to deal with (and ultimately mitigate) this new feature of physics.

fennecfoxy··on The safest job from AI may be writing
I mean a lot of it comes down to:

"Give me a poster for my band playing a gig"

Versus

"Give me a poster for my band playing a gig, it should have x y and z. Use a x' artistic style and include elements of y'. The layout should be z'..."

You ask for the default, you get the default.

fennecfoxy··on Accelerating GPT-5.6 Sol Ultrafast
Oh for sure, but I think generally multiple experts are selected in an MoE pass for a token, so presumably it'd select programming related ones as well as general knowledge/language.

Only problem would be the routing layer works on the previous token as far as I understand so it might need more informational depth than just "a token" to select experts, I suppose in the same way attention works.

I wish I had the GPUs to run those sorts of experiments ha ha.

fennecfoxy··on Accelerating GPT-5.6 Sol Ultrafast
I still personally think that a heavy lean into MoE will be better for that sort of thing. Our brains are subdivided into large parts but I'm sure (and I'm not a brain scientist) that those parts can be subdivided even further into systems that run at various frequencies and latencies depending on what they're used for.

I was thinking about it the other day actually. How our brains evolved structure. I imagine it was purely just down to evolution adding/clustering additional cells around the areas where additional cells were needed. And after long enough a natural brain architecture emerged.

Makes me wonder if we're on the right track with transformer architecture/attention but if it'd be more effective on a larger scale, like MoE with a billion "experts".

fennecfoxy··on Accelerating GPT-5.6 Sol Ultrafast
Fast models is why I was hoping Taalas would get their butts into gear and eventually release a consumer priced card. I'd love to have a pcie card that screams along at 15k t/s even if on a heavily quantized 2026 level model forever.

Faster & cheaper tokens = more reasoning capability and more reasoning = better problem solving as far as I have seen.

fennecfoxy··on AI is removing the middle class of software engineering
I have found this to some degree (even though I work with amazing people day to day).

It's much easier for people to pump out absolute garbage, you know the kind of "just get it done fast" slop that management types cry out for. And then they wonder why everything end ups broken, not being maintained etc.

And I think the contrast between good and bad code output is much more impactful. Someone can pump out 10x the bad code they used to before, never test it never read through it just push push push baby. And then for good code, sure it's increased my output for slop tasks like repetitive unit tests, but a lot of TLC and review is required for good code and I'd say I've had maybe a 2-3x speed up on a lot of things. But not 10x; you only get that when you don't give a fuck.

fennecfoxy··on My phone detects going on a run as “someone snatching my phone and running off”
It's fine. Even re-reading what I'd said it does seem a little extreme, but I still believe the core idea of it. I'm just watching violent criminals in UK get early release while the hapless government and public cannot rub two neurons together to just build more prisons + address problems at the root (income inequality, rising prices, lack of things to do in smaller towns/villages).
fennecfoxy··on Beating GPT-5.6 Sol on retrieval with 100x cheaper open models
Yeah, I always thought the future of this stuff would be hot-pluggable MoE modules or LORAs that are able to be downloaded and applied/used at will like how skills have become a thing.

Like, atm most architectures seem limited by a single context and fixed architecture with no hot loading. Especially for robotics, being able to load/unload various specialised skills on limited mobile hardware will (I hope) definitely become a thing.

fennecfoxy··on New Zealand lost its music media, and what we're building to replace it
When I lived in Auckland (now in the UK) I used to go to gigs at Whammy bar in St. Kevins arcade, the dive bar part is awesome and I think they stop people doing it now but when I was in uni the walls used to be covered by graffiti and scribbles by people.

If you've never been totally recommend. K Road in general is always so amazing (well, like >7 years ago).

fennecfoxy··on New Zealand lost its music media, and what we're building to replace it
Yes, also because I imagine people generally followed the dog damned rules there.
fennecfoxy··on New Zealand lost its music media, and what we're building to replace it
Yep, as a Kiwi living in UK I actually went back to NZ to redo my visa about December 2019, then that was all sorted and I returned about late Jan-early Feb 2020.

I had a stopover in Bangkok and _everyone_ was wearing masks. Not just Asian people, _everyone_. And that's when I was like "oh...this covid thing in the news is gonna be a bit bigger than we thought".

Then I land in the UK after that and all my friends & colleagues either haven't heard of covid or aren't really bothered by it ("it won't come here"). Was back at work for 2 weeks, then the first lockdown hit.

And we had lockdown, after lockdown, after lockdown. I would walk to the supermarket (the one thing we were allowed to do as per the rules) and be walking past parks filled with moron families socialising in the middle of lockdown.

The UK had so many long lockdowns because people here are quite selfish and cannot follow rules - and before people take offence - this is in contrast with NZ culture. There's much more "me me me" and "mine mine mine" culture here in the UK.

fennecfoxy··on My phone detects going on a run as “someone snatching my phone and running off”
No, I'm saying it's only human to retroactively want to do extreme fantastical violence against people who have physically harmed you when you live in a society that a: does not care that you were hurt because you're a man and b: pressures men into fending for themselves, sorting out their own problems and "manning up".

I'm not the type, myself. But damn, no wonder downtrodden men end up snapping - because they most often have nobody to turn to and are forced to exist within a bubble that analyses their weakness and victimhood as much as women are for beauty standards, etc.

fennecfoxy··on Beating GPT-5.6 Sol on retrieval with 100x cheaper open models
I would also like someone to clarify this for me. I (think) I understand what people mean by a neurosymbolic model however common definitions are a bit funky.

If it's neuro (llm/transformer similar) symbolic (symbolic with hand crafted rules) then, imo, that's no different from current tool calling harness implementations and I'd love for someone to explain it further to me if I've misunderstood.

If it's neuro (llm/transformer similar) symbolic (symbolic rules that have also been learned via training) then yeah I can understand how it's a distinct concept.

But every definition I've seen tells me that a standard harness of:

query->llm->[tool call symbol + tool name + tool params] aka symbols governed by logic, including param/arg validation->llm->etc

Already meets the requirements for being "neurosymbolic"...

fennecfoxy··on Beating GPT-5.6 Sol on retrieval with 100x cheaper open models
I'm curious about this too as I've been working with a 1300 page document on procedures for [industry]. Where every single section is primarily about [industry] with minor differences in verbs actions etc for the procedures.

I've found with traditional embeddings that obviously you're getting an average of the content of the chunk even with the semantic awareness magic. And our (or I guess the) core problem seems to be a lack of enough contrast between chunks with makes one-shot pure embedding based RAG extremely difficult and low quality.

fennecfoxy··on Beating GPT-5.6 Sol on retrieval with 100x cheaper open models
It is nice to have a model that can "do it all", though. And surely that's still the end goal? Like how MoE is still somewhat popular in certain areas even after its heyday.

I am wondering if models will end up being some sort of evolution of MoE where it has something internally like the model the author refers to that gets surfaced when it needs to search in some way. I guess it makes sense; our own brains have so many distinct task-specific regions.

fennecfoxy··on Beating GPT-5.6 Sol on retrieval with 100x cheaper open models
By post-train I presume you mean a finetune? Unless that's wrong (please correct me if so).

I haven't looked into model architecture people are working with for this stuff too deeply yet but I presume the core idea is fine-tuning a lightweight reasoning-enabled LLM specifically using search as a metric for training?

fennecfoxy··on Beating GPT-5.6 Sol on retrieval with 100x cheaper open models
Thank you for the link! Super interesting and I appreciate how it was written.

It seems like a lot of the problems I have been running into with RAG on large/complex documents with generally low contrast in the information is not one that has been perfectly solved yet - here I am thinking I'd been a bit behind.

It's just unfortunate that none of the cloud providers are flexible enough to deal with the pace of change. Probably going to have to shove one of those 8b~ models into an instance to use when needed.

It's interesting you mention late interaction (retrieval), I had recently been using ChatGPT as a mirror to throw ideas back at me on this issue and had been musing about how nice it would be to have some sort of hierarchical embeddings that capture a whole chunk, then sentences and then sentence fragments or individual word and it seems that that fits the bill!

fennecfoxy··on Beating GPT-5.6 Sol on retrieval with 100x cheaper open models
Hmmmm, this is a problem I have been facing recently.

Basic embeddings give decent-ish results (in the top say 20 chunks). Basic agentic retrieval gives slightly better results so long as the agent part of it doesn't go down the wrong track.

I like the idea of what's discussed in the link, however atm we are on Bedrock KBs and so locked in to a very basic implementation of RAG, because Amazon doesn't have the foresight to make things flexible enough - including making it an absolute pita to use their hybrid search. But, I guess they "work" reliably.

One of our core issues centers around a 1300 page document all about the same overall topic but with minor various for specific procedures/situations. Typical embeddings waters this down so that each chunk really just represents the common theme and therefore lacks a lot of contrast.

But now that luna's (and others) price has been cut, perhaps I'll start experimenting with giving it free rein to explore the data a little in the same way that I do a web search.

One thing that definitely helped was providing a separate index of each section where I had another model summarise the primary unique topics in each section to act as a guide for the agent. I think either we should be chucking the entire doc at a model (400k tokens...so not really ideal at this time) or improving RAG accuracy. For the latter I think even with embeddings, meaning of words and semantic connections are not enough at all - attention is KV so it is 2 dimensional and once I started getting into it I've kind of realised that 2 dimensions aren't really enough to represent the logic that exists between tokens (i.e. sections of documents that refer to a sequence of actions dependent on some logic that references "variables" from another section, i.e. "if x, y has happened then refer to z sequence). There's much deeper meaning to human language than I think basic embeddings covers.

I think it's becoming clear to me that in the same way that embeddings encode the web of semantic meaning of a chunk of text, I need something similar to a hybrid of the author's model + reranker + super-embeddings that encodes as much of the entire meaning of a text as possible and not just semantic.

Page 1 of 34Next →