HNHacker News
TopNewBestAskShowJobs

roseway4

1,118 karma · joined May 10, 2016

Founder, Zep. ML, robots, startups.
submissionscomments
roseway4··on Ask HN: Who is hiring? (May 2024)
Zep AI (YC W24) | Developer Advocate | https://www.getzep.com | Full-time | SF Bay / Remote US-only

Apply at: https://jobs.gem.com/zep

Zep is building the long-term memory layer for the LLM application stack. With a fast growing open source community and recently launched cloud service, we're seeking a developer relations engineer to accelerate adoption among developers and help shape our roadmap.

In this highly visible role, you'll engage with the developer community across multiple channels - writing technical content, delivering presentations, creating sample apps/demos, and fielding inquiries. Working directly with our founder + engineering team, you'll help steer our product strategy.

Zep is funded by Y Combinator, Engineering Capital, and angels such as Guillermo Rauch (Vercel).

roseway4··on Show HN: Ellipsis – Automatic pull request reviews
We use Ellipsis across most of our repos. PR titling and summaries are great and definitely reduce the cognitive load for both code contributors and reviewers. Reviews are also a good first line of defense for code quality, catching issues human reviewers haven't caught. Though on occasion it's also a little overly pedantic or confused. I guess not unlike us humans at times ;-)
roseway4··on Show HN: I made a GPU VRAM calculator for transformer-based models
While not as pretty (and mobile-friendly) as the original link, the calculators below support modeling LoRA-based training, alongside full finetuning.

https://huggingface.co/spaces/Vokturz/can-it-run-llm

https://rahulschand.github.io/gpu_poor/

roseway4··on Vim Racer – VI Keyboard Skillz Game
A fun browser-based game for practicing your VI keyboard skillz. Come for the game, enjoy the 80's neon fashion and pixel art font.
roseway4··on We have decided to pause driverless operations across all of our fleets
It is alleged that Cruise neglected to provide the DMV with information regarding an incident that led to a pedestrian being badly injured. Other reports claim they willfully withheld information. It appears to be more than a “software bug.”
roseway4··on Zep: Fast, scalable building blocks for production LLM apps
You could certainly use Postgres for everything. Zep does. But then you'd also have to build the stuff that Zep does, too.
roseway4··on Zep: Fast, scalable building blocks for production LLM apps
Thanks. I definitely want to make a longer video demonstrating Zep end-to-end.
roseway4··on Zep: Fast, scalable building blocks for production LLM apps
Thanks for the advice! Some benchmarks here: https://blog.langchain.dev/zep-x-langchain-slow-chatbots/
roseway4··on Zep: Fast, scalable building blocks for production LLM apps
Zep has Python and TypeScript packages for teams building without LangChain: https://docs.getzep.com/sdk/
roseway4··on Zep: Fast, scalable building blocks for production LLM apps
Zep author here. LangChain is a great framework and has a very broad and active ecosystem. Many of LangChain's core components, such as chat history "memory," history summarization, entity extraction, vector search, and more, don't scale well in production. They operate in memory and often synchronously within the chat loop, resulting in poor user experiences and limiting deployment options.

You can use third-party ecosystem integrations with external SQL databases, vector databases, etc to fix some of these issues. This requires some know-how, infrastructure, and time. Zep's building blocks are turn-key solutions to these challenges and offered in a single service.

The Zep GitHub project, website, and demo video provide a good overview of the project's functionality and how the service solves these issues.

This blog post on the LangChain website offers some benchmark data using Zep vs LangChain's core memory components: https://blog.langchain.dev/zep-x-langchain-slow-chatbots/

roseway4··on LLMs, RAG, and the missing storage layer for AI
You may want to take a look at Zep, an LLM application platform that wraps Postgres, pgvector, embedding models, and more to offer chat memory persistence and document vector search.

The Python and TS SDKs are designed to support drop-in replacements for the bits of LangChain that don’t scale, but nothing stops you accessing Postgres directly.

https://github.com/getzep/zep

Disclosure: I’m the primary author.

roseway4··on Recursively summarizing enables long-term dialogue memory in LLMs
I’m struggling to understand what’s novel here. LLM-based summarization of chat history memory is a well-established technique implemented by many LLM frameworks. Summarizing on every message is, as proposed in the paper, a major performance bottleneck and adds significant latency to the chat loop.

Many implementations utilize a fixed sized buffer, progressively summarizing batches of older memories when they fall out of the buffer. Ideally, this is also done out of band to the chat loop.

I’m an author of Zep[0], an open source long-term memory store, and this is how we implemented summarization.

0: https://github.com/getzep/zep

roseway4··on Pgvector 0.5.0 Feature Highlights and how tos
Jonathan did just that in his prior blog post: https://jkatz05.com/post/postgres/pgvector-hnsw-performance/

pgvector_hnsw outperforms pg_embedding using a query per sec / recall measure across various common embedding widths and a range of dataset sizes.

roseway4··on HNSW Merged into Pgvector
Fantastic news. Postgres is becoming a compelling vector DB for many use cases.
roseway4··on Secret Chinese-run lab in CA illegally stored vials of Covid-19 and other
No. The FresnoBee is a legitimate news organization and the story has been reported on by other news orgs.

https://www.nbcnews.com/news/us-news/officials-believe-fresn...

roseway4··on On-disk HNSW index for Postgres with pg_embedding
As I mentioned elsewhere, there's more to vector database selection than raw performance: Developers leveraging their existing experience with Postgres, existing infrastructure investment (often in managed Postgres on AWS/GCP etc), and a single API into the vector store shared with other parts of their app (their ORM / DB layer).

Many teams can also get away with good performance vs the fastest performance, given smaller index sizes and the other tradeoffs I mentioned.

That said, I can imagine the pgvector folks precaching the new HNSW index support they're working on, as they do with their IVFFLAT index.

* edited for the grammar gremlins

roseway4··on On-disk HNSW index for Postgres with pg_embedding
This is comparing apples and oranges. You've pointed to some technologies from vector database vendors. Neon isn't a vector database vendor. They're a Postgres vendor. The objective of pg_embedding, pg_vector, and others is to offer teams who deploy Postgres the opportunity to use their existing infrastructure for vector search. The added and important benefit here is hybrid search: using existing DB data to filter semantic search results.

What the work done by Neon, the pgvector team, Supabase and others points to is that "speed" isn't the only factor in vector database selection. Developer experience and existing infrastructure investment are too.

roseway4··on Pgvector: Fewer Dimensions Are Better
Kismet this is published today. We've just announced Zep's[0] new Document Vector Store[1] built on pgvector, and the embedding model we ship with is all-MiniLM-L6-v2. Very fast, low-memory, and surprisingly performant.

Our focus with Zep is on LLM app use cases, and in particular, searching over chat histories and documents for RAG apps. You can, however, use Zep to turn any Postgres instance into a vector store with great developer experience. pgvector index configuration and query tuning can be challenging. We've tried to do much of this work for developers.

Also, I'm super excited about the prospect of HNSW support in pgvector, slated for pgvector 0.5.[2]

[0] https://github.com/getzep/zep

[1] https://blog.getzep.com/introducing-the-zep-document-vector-...

[2] https://github.com/pgvector/pgvector/issues/181#issuecomment...

roseway4··on Johannesburg: Video captures moment suspected gas explosion threw buses into air
Johannesburg has municipal natural gas lines. They’re more likely to be the culprit here than sewage.
roseway4··on Space material found on a beach in Western Australia
It's been identified. Reddit FTW. :-)

https://www.reddit.com/r/space/comments/1515q3w/found_on_a_b...

roseway4··on Show HN: Structured output from LLMs without reprompting
When you say “our model”, are you using a custom LLM for completions vs OpenAI or other LLM vendor?
roseway4··on Show HN: Structured output from LLMs without reprompting
Looking at the playground, it appears the few shot examples in the prompt and CFG are duplicative. What is the relationship between the two?

When you say in another comment that using OpenAI functions to output JSON is a waste of tokens, how are you generating the JSON output? And why do your prompts then include few shot examples of JSON objects?

roseway4··on How to Use AI to Do Stuff: An Opinionated Guide
What “better model”? If you’re referring to whatever LLM GPT-4 was before RLHF, by what measure might this model be better? As I pointed out above, what OpenAI is selling is “best” for their customers.
roseway4··on ChatGPT use declines as users complain about ‘dumber’ answers
I’m increasingly seeing ChatGPT as a single LLM application being conflated with an entire domain or industry. And plenty of schadenfreude. Including many comments on this site.

The universe of LLM-driven applications is rapidly expanding, many of them chat-related, others not.

Even if ChatGPT is losing its shine, we’re only at the start of a massive reinvention of user interfaces, creation of new tools for reasoning, semi-autonomous decision making, and far more.

Sure there’s hype. But the reductive saltiness doesn’t add much to the conversation.

roseway4··on How to Use AI to Do Stuff: An Opinionated Guide
Caveat emptor indeed. LLM-as-a-Service vendors know their customer and this customer’s needs. It’s not you. They’re selling an API to companies for whom safety and governance are part of the value proposition.

As others have mentioned, the data retention aspect is specifically /not/ an issue with OpenAI (and other vendors’) APIs.

I don’t really get the saltiness so many on this site have towards LLM vendors. It just sounds entitled.

roseway4··on Could an industrial civilization have predated humans on Earth?
As the Wikipedia article mentions, we’ve already done just this with the JWST…
roseway4··on Show HN: AI companions stack – create and host your own AI companions
I stand corrected. Sometimes my cynicism gets the best of me ;-)
roseway4··on Bard’s latest update: more features, languages and countries
Bard's multimodal abilities are impressive. It's worth noting that PaLM 2 does not offer multimodal capabilities, and your experience here, as you suggested, likely leans on Google's long history of building search across both text and image data.

In a reasoning and creativity shootout, I'm not convinced Palm 2 (and Bard) fairs well against Claude v2 and GPT-3.5.

Is anybody aware of Palm 2 being used in production other than with Bard? I see companies building and deploying with Claude* but have yet to come across Palm 2 users.

* I've fielded several requests to add Claude 2 API support to Zep, an open source memory store I co-authored. https://github.com/getzep/zep

roseway4··on Show HN: AI companions stack – create and host your own AI companions
Well, it is a round-up of their venture investments in the space ;-)

If you're looking to self-host chat memory rather than go all in on Supabase, there's Zep: https://github.com/getzep/zep

Full disclosure: I'm a co-author.

roseway4··on Virus-like transposons cross the species barrier, study shows
tl;dr the presence of exactly the same gene in reproductively isolated species has puzzled scientists for some time. The example cited were two different very cold water fish which both have exactly the same “antifreeze” gene.

This research group identified a type of transposon that infect worms with a virus-like ability to penetrate cells and “infect” hosts. Transposons are “greedy” particles that “transpose” DNA in order to reproduce, often at the expense of their hosts. While doing this, entire genes may be transposed.

The transposon itself likely gained this ability from an encounter with a virus.

Fascinating.

← PreviousPage 2 of 6Next →