HNHacker News
TopNewBestAskShowJobs

heisenburgzero

22 karma · joined July 3, 2016

submissionscomments
heisenburgzero··on Prompt engineering playbook for programmers
In my own experience, if the problem is not solvable by a LLM. No amount of prompt "engineering" will really help. Only way to solve it would be by partially solving it (breaking down to sub-tasks / examples) and let it run its miles.

I'll love to be wrong though. Please share if anyone has a different experience.

heisenburgzero··on Semantic search engine for ArXiv, biorxiv and medrxiv
So did you just combine Title+Abstracts+Authors into a single chunk and embed them or embedded them individually?
heisenburgzero··on Using Ghidra and Python to reverse engineer Ecco the Dolphin
I always wondered where to start learning reverse-engineering. Most people will say learn Assembly first. But from there on, there seems to be not much more concrete information online.

Do people just figured it out by trial & error like common patterns in x86 / arm / arcade platforms slowly?

I can't really find much discussion on details online.

heisenburgzero··on Understand how transformers work by demystifying the math behind them
Thanks. "In-context learning" was the phrase I was looking for.

Also, I found these 2 links pretty good too. 1. http://ai.stanford.edu/blog/understanding-incontext/ 2. http://ai.stanford.edu/blog/in-context-learning/

I'm still not completely convinced. Probably need to dwell on the topic longer.

heisenburgzero··on Understand how transformers work by demystifying the math behind them
I think some earlier NLP applications have something called "Unknown token", which they will replace any unseen word. But for recent implementations, I don't think they are being used anymore.

It still baffles me why such stochastic parrot / next token predictor, will recognize these "Unseen combinations of tokens" and reuse them in response.

heisenburgzero··on Understand how transformers work by demystifying the math behind them
Not completely related. Does anyone know where I can find articles / papers that discuss why transformers, while acting as merely "next token predictor" can handle questions with: 1. Unknown words (or subwords/tokens) that are not seen in the training dataset. Example: Create a table with "sdsfs_ff", "fsdf_value" as columns in pandas. 2. Create examples(unseen in training dataset) and tell the LLM to provide similar output.

I have a feeling it should be a common question, but I just can't find the keyword to search.

PS. If anyone has any links with thoroughly discussion about positional embedding, that would be great. I never got a satisfying answer about the usage of sine / cosine and (multiplication vs addition)

heisenburgzero··on Embeddings: What they are and why they matter
Thanks, that make sense.

I think my comment was not worded properly. I was thinking "geometry properties = linear properties", what I really should say is:

Why does the latent space has geometry properties where we could use functions like cosine similarity to compare?

So when training, the signal will be mapped to latent space that will minimize the error of the objective function as much as possible.

Many applications already use cosine similarity function at the end the network, it would be obvious why they work. I reviewed other cost functions such as Triplet Loss. They use euclidean distances, so I guess it make sense why the geometry properties exist too.

For "and there I guess the point is it's not maximally information dense, so the geometry exists in the redundancy", what does "maximally information dense" means, I still don't quite get it.

heisenburgzero··on Embeddings: What they are and why they matter
Why does the embeddings have linear properties such that you can use functions like cosine similarity to compare? It seems that after the signal going through so many non-linear activation layers, the linear properties should have been broken down / no guarantees.

I wasn't able to find a good answer online.

heisenburgzero··on YOLOv5: State-of-the-art object detection at 140 FPS
This is not the first time something is fishy. Back in the early stages of the repo. They were advertising on the front page that they are achieving similar MAP to the original C++ version. But only to be found out they haven't train it on COCO dataset and test it.
heisenburgzero··on How I Cracked a Keylogger and Ended Up in Someone's Inbox
where did ).exe came from? I thought you need to use VBscript of some sort to download a file from command line.