HNHacker News
TopNewBestAskShowJobs

sandkoan

128 karma · joined October 11, 2020

https://govindgnana.com , https://blog.govindgnana.com
submissionscomments
sandkoan··on Why Bell Labs Worked
DeepMind.
sandkoan··on Show HN: Sonauto – A more controllable AI music creator
See https://musiccontrolnet.github.io/web/
sandkoan··on World_sim: LLM prompted to act as a sentient CLI universe simulator
sudo !!
sandkoan··on Extreme video compression with prediction using pre-trainded diffusion models
Ilya says this here: https://www.youtube.com/watch?v=AKMuA_TVz3A
sandkoan··on Pretraining data enables narrow selection capabilities in transformer models
For anyone else interested, the paper he's referring to is "The Reversal Curse": https://arxiv.org/abs/2309.12288.
sandkoan··on MemGPT – LLMs with self-editing memory for unbounded context
For anyone who's curious, the paper in question, entitled, "Lost in the Middle: How Language Models Use Long Contexts" (https://arxiv.org/abs/2307.03172)
sandkoan··on Fine-tune your own Llama 2 to replace GPT-3.5/4
This is part of what we're doing at Automorphic. Building shareable, stackable adapters that you can compose like lego bricks.
sandkoan··on Show HN: LLMs can generate valid JSON 100% of the time
This is what we did at Trex (https://github.com/automorphic-ai/trex). The tricky part is doing it quickly and efficiently.
sandkoan··on Llama: Add grammar-based sampling
Also using a similar method: https://github.com/automorphic-ai/trex

Playground: https://automorphic.ai/playground

sandkoan··on TypeChat
Relevant: Built this which generalizes to arbitrary regex patterns / context free grammars with 100% adherence and is model-agnostic — https://news.ycombinator.com/item?id=36750083
sandkoan··on Show HN: Structured output from LLMs without reprompting
Yeah, then it seems we agree. I was just pointing out that it's not necessary to finetune OSS models to behave like OpenAI functions if you're able to do something similar to what we did (no tuning involved!).
sandkoan··on Show HN: Structured output from LLMs without reprompting
Ahh, no, the value of this isn't as much the model as it is the infrastructure enabling structure enforcement.
sandkoan··on Show HN: Structured output from LLMs without reprompting
https://news.ycombinator.com/item?id=36752991
sandkoan··on Show HN: Structured output from LLMs without reprompting
You wouldn't actually want to, because you'd be losing generalizability, and it's a lot of unnecessary work.

I think approach #1 outlined above is the better (more cost- and time-efficient) technique—where a pretrained model already understands JSON (among myriad other formats), and you merely constrain it at text-gen time to valid JSON (or other format).

sandkoan··on Show HN: Structured output from LLMs without reprompting
Ahh, I've been meaning to try FLARE—was it a marked improvement over traditional RAG?
sandkoan··on Show HN: Structured output from LLMs without reprompting
Thanks for the reminder—done!
sandkoan··on Show HN: Structured output from LLMs without reprompting
This is model agnostic, actually—any model on HuggingFace is compatible. So if someone wanted to run this with their own model, they could.
sandkoan··on Show HN: Structured output from LLMs without reprompting
Custom LLM—hence the self-hostability.
sandkoan··on Show HN: Structured output from LLMs without reprompting
Costs add up surprisingly quickly. A quote-colon-space-quote combo alone is four tokens wasted. Now scale that up....
sandkoan··on Show HN: Structured output from LLMs without reprompting
https://news.ycombinator.com/item?id=36753254

Does this help clarify?

sandkoan··on Show HN: Structured output from LLMs without reprompting
The prompt is given to our model as a guiding aid (a suggestion), and the cfg is used to constrain the model to generate only tokens that abide by the schema (an enforcement). That's how we ensure only valid outputs at text generation time.

We also prefill some tokens depending on the set of allowed tokens at a given state, so the model doesn't waste resources trying to predict them.

sandkoan··on Show HN: Structured output from LLMs without reprompting
We have folks playing around with it mostly through the playground / raw HTTP endpoints as opposed to the Python API. And we've got some batch jobs running, which adds further traffic.
sandkoan··on Show HN: Structured output from LLMs without reprompting
Problems with OpenAI:

1) You're wasting GPT tokens on outputting JSON instead of meaningful information.

2) GPT functions won't, with absolute, 100% certainty, return JSON in the schema you want. In 1% to 3% of cases it hallucinates fields, etc.

3) This also allows you to output data in arbitrary non-JSON formats.

4) You can't self-host OpenAI functions.

sandkoan··on Show HN: Structured output from LLMs without reprompting
We enable conforming to arbitrary context free grammars in addition to regex patterns, and have a bunch of speed optimizations, as well.

Though it may not seem too fast right now on account of the hundreds of simultaneous requests we're getting :)

sandkoan··on Show HN: Structured output from LLMs without reprompting
Ahh, that—due to compute limitations we're forced to run a very small model that isn't as capable of converting 5'8" to inches. The larger model is, though.
sandkoan··on Ask HN: How did you migrate off Evernote?
Shoot me an email at govind <at> automorphic <dot> ai
sandkoan··on How to verify your domain on Nostr and Bluesky (for micro.blog users)
Would love an invite as well, if you still have some :)
sandkoan··on Langchain Is Pointless
Wholeheartedly agree—it adds layers of bloat and abstraction between you and the actual ReAct pattern, which can be trivially implemented in maybe 50 sloc.
sandkoan··on Ask HN: How did you migrate off Evernote?
Do you have a copy of the script you used to export all your notes?

I'm currently faced with this existential problem myself.

sandkoan··on FBI obtains Kolektiva.social database while executing warrant on administrator
"RTFM" means "Read The Fucking Manual."
Page 1 of 3Next →