HNHacker News
TopNewBestAskShowJobs

sdrg822

88 karma · joined May 6, 2019

submissionscomments
sdrg822··on A single firm is behind OpenAI, Anthropic, and Meta hacking scandals
This is incredibly misleading. OpenAI internal systems were pwned, and in all cases, the labs absolutely are responsible for their models.

Yes, vendors are also irresponsible, but this misses the point.

sdrg822··on Agent Mesh for Enterprise Agents
“Model Control Plane” - quite early to already re-use a popular acronym
sdrg822··on LangManus: An Open-Source Manus Agent with LangChain + LangGraph
Congrats on the launch!
sdrg822··on Prompt Caching
+1 it wouldn’t be terribly useful if it were only caching the tokenizer output.
sdrg822··on LangGraph Engineer
Where is the lie in this repo?
sdrg822··on LangGraph Engineer
I think if you actually meet the people you'd realize they're pretty earnest and candid of these things' limitations, though it may not show in tweets and hackernews posts.

This repo, for instance, makes no claims of AGI. It just claims to help bootstrap a starting point for a software project: "it will not attempt to write the logic to fill in the nodes and edges."

sdrg822··on Consistency LLM: converting LLMs to parallel decoders accelerates inference 3.5x
But indexing *is* training. It's just not using end-to-end gradient descent.
sdrg822··on A generalist AI agent for 3D virtual environments
Dang they use Transformer-XL from 2019 haha - didn't realize people still used that / XLNet-like architectures
sdrg822··on Microsoft strikes deal with Mistral in push beyond OpenAI
It is not.
sdrg822··on Show HN: Use natural language to query and visualize 400M tweets
Congrats on the launch!
sdrg822··on Exponentially faster language modelling
Cool. Important note:

""" One may ask whether the conditionality introduced by the use of CMM does not make FFFs incompatible with the processes and hardware already in place for dense matrix multiplication and deep learning more broadly. In short, the answer is “No, it does not, save for some increased caching complexity." """

It's hard to beat the hardware lottery!

sdrg822··on LLaMa running at 5 tokens/second on a Pixel 6
It’s only a matter of time
sdrg822··on Implementing GPTZero from scratch – Reverse engineering GPTZero
Given that GPTZero was an undergrad’s side project it’s not that surprising?
sdrg822··on What a transformer can NOT do
Attention is Turing complete
sdrg822··on Why is Chat GPT so expensive to operate?
For things like BERT where you just want to extract an embedding, the naive way you reach full utilization at inference time is that you :

- run tokenization of inputs on CPU

- sort inputs by length

- batch inputs of similar length and apply padding to make of uniform length

- pass the batches through so a single model can process many inputs in parallel.

For GPT-style decoder models however, this becomes much more challenging because inference requires a forward pass for every token generated. (Stopping criteria also may differ but that’s another tangent).

Every generated token performs attention on every previous token, both the context (or “prompt”) and the previously generated tokens (important for self consistency). this is a quadratic operation in the vanilla case.

Model sizes are large , often spanning multiple machines, and the information for later layers depends on previous ones, meaning inference has to be pipelined.

The naive approach would be to have a single transaction processed exclusively by a single instance of the model. this is expensive! even if each model can be crammed into a single A100 , if you want to run something like Codex or ChatGPT for millions of users with low latency inference, you’d have to have thousands of GPUs preloaded with models, and each transaction would take a highly variable amount of time.

If a model spans multiple machines, you’d achieve a max of 1/n% utilization because each shard has to remain loaded while the others process, and then if you want to do pipeline parallelism like in pipe dream, you’d have to deal with attention caches since you don’t want to have to recompute every previous state each time

sdrg822··on ChatGPT is a ‘code red’ for Google’s search business
Yeah this is purely a risk decision for a prototype not an actual technical limitation.
sdrg822··on Show HN: Explainpaper – Explain jargon in academic papers with GPT-3
Really well done
sdrg822··on Road to Artificial General Intelligence
Which humans?
sdrg822··on The Anglo-Saxon Classroom
Good points!
sdrg822··on The Anglo-Saxon Classroom
Just a slight comment - "&" was indeed called "and" (not "per se"), but in reading the alphabet, it was confusing to say "and and." To clear up confusion, one could say "and per se and," which was smooshed to become "ampersand."

"per se" was used for letters that could also be words "in themselves"

sdrg822··on America’s losing battle against diabetes
Seasoning isn't the problem... Turmeric and cumin added to your pulses aren't going to give you metabolic diseases...
sdrg822··on Zero arrests in 6 months of health care professionals replacing police officers
Agreed, the "but" in the bolded text[1] further confuses, but the article highlights some great progress.

>> STAR, has responded to 748 incidents." >> about 3 percent of calls for DPD service, or over 2,500 incidents, were worthy of the alternative approach

[1] "Chief Pazen is thrilled with the success of STAR, but the time and money it saves will go toward fighting crime, he said."

sdrg822··on Could a fecal transplant one day restore cognitive function in elderly?
> "While it remains to be seen whether transplantation from very young donors can restore cognitive function in aged recipients, the findings demonstrate that age-related shifts in the gut microbiome can alter components of the central nervous system."

Looking forward to future results testing the other direction.

sdrg822··on [dead]
Automated summarization, content moderation/censorship, and advertising.

It could make publishing more profitable, but will it make enable higher quality content? It feels more like a continuation of the race to the bottom we are already engaged in.

sdrg822··on Frugality Is Non-Linear (2019)
Abbreviation for a popular blogger in the financial independence / retire early domain. See: https://www.mrmoneymustache.com
sdrg822··on How GPT3 Works – Visualizations and Animations
While models such as XLNet incorporate recurrence, GPT-{2,3} is mostly just a plain decoder-only transformer model.[1]

[1]https://arxiv.org/abs/2005.14165 [2]https://d4mucfpksywv.cloudfront.net/better-language-models/l...

sdrg822··on A Walk in Hong Kong
Embarrassment may be universal, but the internalization of guilt and shame varies dependent on the culture both in degree and manner. See, e.g.:

https://www.researchgate.net/publication/276464615_A_Cultura...

sdrg822··on A Look at Overnight Stays at US National Parks
Usually backcountry camping requires hiking in to a site, whereas "tent" camping would refer to stays in frontcountry campgrounds that are usually accessible via a car. With exceptions, tent camping would require a reservation, whereas backcountry camping sites often are occupied on a first-come first-served basis.
sdrg822··on “Once-in-a-Hundred Year” Sightings of Bamboo Blossoms Reported in Japan
I also have been wondering whether I am out of the loop on some joke...