HNHacker News
TopNewBestAskShowJobs

marcklingen

78 karma · joined September 6, 2021

co-founder Langfuse

marcklingen.com

Email: hi@

submissionscomments
marcklingen··on ClickHouse acquires Langfuse
Let me know what features are missing in prompt management
marcklingen··on ClickHouse acquires Langfuse
(thank you!)
marcklingen··on Ask HN: Books about people who did hard things
The Innovators, Walter Isaacson

It’s interesting to read how many individuals contributed in all sorts of important ways in the history of computing.

marcklingen··on Show HN: Houseplant – Database Migrations for ClickHouse
This is interesting! Do you plan to add support for distinguishing between blocking and non-blocking/async migrations? This seems to be a common challenge when using Clickhouse and needing to change the ordering of existing tables.

Context: When adding Clickhouse to Langfuse [1] to store all observability data, the migration DX was a big challenge. Some migrations need to copy/reorder data between tables which can take a long time and thus need be performed in the background/async. Many teams run Langfuse self-hosted and reducing downtime of an upgrade is important, thus, we added some custom logic to run these “background migrations” for now [2]. For regular migrations we use golang-migrate, works decently although DX isn’t as good as eg Prisma for Postgres.

[1] OSS LLM Observability, https://github.com/langfuse/langfuse

[2] https://langfuse.com/self-hosting/background-migrations

marcklingen··on Launch HN: Langfuse (YC W23) – OSS Tracing and Workflows to Improve LLM Apps
The Langfuse data model is closely inspired by OpenTelemetry and we plan to add a Collector soonish. Until now, the OTel-semantics for LLMs have not been very stable and exhaustive while LLM capabilities are changing frequently (think prompt caching, realtime, multi-modal).

We are tracking this here including potential tradeoffs: https://github.com/orgs/langfuse/discussions/2509

marcklingen··on Launch HN: Langfuse (YC W23) – OSS Tracing and Workflows to Improve LLM Apps
Thanks! IMO, Discord is good, but GitHub Discussions is the better option for building a growing open-source community. It is indexed and makes it easier to revisit conversations weeks later. Currently we use both but have a strong preference for GitHub Discussions.
marcklingen··on Launch HN: Langfuse (YC W23) – OSS Tracing and Workflows to Improve LLM Apps
All core features are fully open-source and identical to those in Langfuse Cloud, with no limitations on capabilities or scalability (e.g. all v3 infrastructure changes).

We also offer some optional commercial add-on features that can help iterate faster or support very large teams using Langfuse. However, these features are entirely optional and we do our best to be transparent about this across our docs.

marcklingen··on Launch HN: Langfuse (YC W23) – OSS Tracing and Workflows to Improve LLM Apps
Congrats on the Launch!
marcklingen··on Launch HN: Langfuse (YC W23) – OSS Tracing and Workflows to Improve LLM Apps
As you mentioned, this was a significant trade-off. We faced two choices:

(1) Stick with a single Docker container and Postgres. This option is simple to self-host, operate, and iterate on, but it suffers from poor performance at scale, especially for analytical queries that become crucial as the project grows. Additionally, as more features emerged, we needed a queue and benefited from caching and asynchronous processing, which required splitting into a second container and adding Redis. These features would have been blocked when going for this setup.

(2) Switch to a scalable setup with a robust infrastructure that enables us to develop features that interest the majority of our community. We have chosen this path and prioritized templates and Helm charts to simplify self-hosting. Please let us know if you have any questions or feedback as we transition to v3. We aim to make this process as easy as possible.

Regarding OTel, we are considering adding a collector to Langfuse as the OTel semantics are currently developing well. The needs of the Langfuse community are evolving rapidly, and starting with our own instrumentation has allowed us to move quickly while the semantic conventions were not developed. We are tracking this here and would greatly appreciate your feedback, upvotes, or any comments you have on this thread: https://github.com/orgs/langfuse/discussions/2509

marcklingen··on Launch HN: Langfuse (YC W23) – OSS Tracing and Workflows to Improve LLM Apps
Thanks! Really enjoyed working with you maintainers of other projects to help them offer more native LLM observability and evaluation to their users/communities. There is a lot that goes into making the observability/eval part scalable/useful and requirements change on a weekly basis with new advancements. Same applies to other projects and it makes a lot of sense to integrate.

Overview of community integrations: https://langfuse.com/docs/integrations/overview

Packages that depend on Langfuse: https://langfuse.com/faq/all/packages-depending-on-langfuse

marcklingen··on Launch HN: Langfuse (YC W23) – OSS Tracing and Workflows to Improve LLM Apps
Thank you! Working with your team has been great. I love seeing you ship LLM-powered features and appreciate the feedback you have shared along the way.
marcklingen··on Launch HN: Langfuse (YC W23) – OSS Tracing and Workflows to Improve LLM Apps
thank you! if you have any ideas for improvements after having used Langfuse for a while, please contribute them via github discussions: https://langfuse.com/ideas
marcklingen··on Launch HN: Langfuse (YC W23) – OSS Tracing and Workflows to Improve LLM Apps
Thanks for the feedback, glad that you find Langfuse useful!

Can you create an issue with more details on the API performance problems? We monitor strict SLOs on the public API for Langfuse Cloud and are not aware of any ongoing issues, would love to learn more.

marcklingen··on Launch HN: Langfuse (YC W23) – OSS Tracing and Workflows to Improve LLM Apps
appreciate your constructive feedback!

> i wonder if there are new ops solutions for the realtime apis popping up

This is something we have spent quite some time on already, both on designs internally and talking to teams using Langfuse with realtime applications. IMO the usage patterns are still developing and the data capturing/visualization needs across teams is not aligned. What matters: (1) capture streams, (2) for non-text provide timestamped transcript/labels, (3) capture the difference between user-time and api-level-time (e.g. when catching up on a stream after having categorized the input first).

We are excited to build support for this, if you or others have ideas or a wishlist, please add them to this thread: https://github.com/orgs/langfuse/discussions/4757

> retries for instructor like structured outputs mess up the traces, i wonder if they can be tracked and collapsible

Great feedback. Being able to retroactively downrank llm calls to be `debug` level in order to collapse/hide them by default would be interesting. Added thread for this here: https://github.com/orgs/langfuse/discussions/4758

> chatgpt canvas like "drafting" workflows are on the rise (https://www.latent.space/p/inference-fast-and-slow) and again its noisy to see in a chat flow

Can you share an example trace for this or open a thread on github? Would love to understand this in more detail as I have seen different trace-representations of it -- the best yet was a _git diff_ on a wrapper span for every iteration.

> how often do people actually use the feedback tagging and then subsequently finetuning? i always feel guilty that i dont do it yet and wonder when and where i should.

Have not seen finetuning based on user-feedback a lot as the feedback can be noisy and low in frequency (unless there is a very clear feedback loop built into the product). More common workflow that I have seen: identify new problems via user feedback -> review them manually -> create llm-as-a-judge or other automated evals for this problem -> select "good" examples for fine-tuning based on a mix of different evals that currently run on production data -> sanitize the dataset (e.g. remove PII).

Finetuning has been more popular for structured output, sql generation (clear feedback loop / retries at run-time if the output does not work). More teams fine-tune on all the output that has passed this initial run-time gate for model distillation without further quality controls on the training dataset. They usually then run evals on a test dataset in order to verify whether the fine-tuned hits their quality bar.

marcklingen··on Launch HN: Langfuse (YC W23) – OSS Tracing and Workflows to Improve LLM Apps
Thanks for the feedback.

Being unopinionated and API-first has been a core design decision. We want to build the building blocks that everyone needs while acknowledging that most Langfuse users are very sophisticated teams that have a clear idea of what they want to achieve. Over time we will build more abstractions for common workflows to make it easier to get started but new features will always start API-first.

More on this here: https://langfuse.com/why

marcklingen··on Launch HN: Langfuse (YC W23) – OSS Tracing and Workflows to Improve LLM Apps
Thanks AJ, feedback on GitHub/Discord (like yours) has been very helpful to evolve prompt management from a quick addition of the core platform to one of the most-used features -- for which we then actually needed to change a lot of infrastructure to make it reliable and fast (see blog post linked in the original post)
marcklingen··on Launch HN: Langfuse (YC W23) – OSS Tracing and Workflows to Improve LLM Apps
This is a good long-list of projects, although it is not narrowly scoped to tracing/evals/prompt-management: https://github.com/tensorchord/Awesome-LLMOps?tab=readme-ov-...
marcklingen··on Launch HN: Langfuse (YC W23) – OSS Tracing and Workflows to Improve LLM Apps
Thank you for the kind words! Let us know if you have any questions or feedback regarding the self-hosting documentation and experience. We collaborate with many teams that have diverse security needs, including HIPAA, PCI, and on-premises deployments on bare metal without internet access.
marcklingen··on Launch HN: Langfuse (YC W23) – OSS Tracing and Workflows to Improve LLM Apps
If you use an OSS Gateway already, some (e.g. LiteLLM) can natively forward logs to Langfuse: https://docs.litellm.ai/docs/proxy/logging#langfuse

We are looking into adding an Otel Collector as OTel-semantics are maturing around LLMs. For now many features that are key to LLMOPs are difficult to make work with OTel instrumentation as the space is moving quickly. Main thread on this is here: https://github.com/orgs/langfuse/discussions/2509

marcklingen··on Launch HN: Langfuse (YC W23) – OSS Tracing and Workflows to Improve LLM Apps
Positioning/roadmap differs between the different project in the space.

We summarized what we strongly believe in here: https://langfuse.com/why Tldr: open apis, self-hostable, LLM/cloud/model/framework-agnostic, API first, unopinionated building blocks for sophisticated teams, simple yet scalable instrumentation that is incrementally adoptable

Regarding roadmap, this is the near-term view: https://langfuse.com/roadmap

We work closely with the community, and the roadmap can change frequently based on feedback. GitHub Discussions is very active, so feel free to join the conversation if you want to suggest or contribute a feature: https://langfuse.com/ideas

marcklingen··on Launch HN: Langfuse (YC W23) – OSS Tracing and Workflows to Improve LLM Apps
This is an optional instrumentation of the OpenAI SDK which simplifies getting started, tracking token counts, model parameters and streaming latencies.

Langfuse is not in the critical path, this just helps with instrumentation.

You can use the Langfuse Python SDK / Decorator to track any LLM (with some instrumentation code) or use one of the framework integrations.

Here is a fully-featured example using the Amazon Bedrock SDK: https://langfuse.com/docs/integrations/amazon-bedrock

marcklingen··on Ask HN: What's the best way to build multi agent PoCs today
-- langfuse founder/maintainer here

After working with thousands of teams building agents and complex LLM applications, I'd conclude that:

- all frameworks have their strengths and weaknesses, Crew and LangGraph are the popular options if you want to use a framework, both have sdk methods or add-on packages to visualize the agent

- biggest overall strength is to get started quickly

- biggest overall weakness is that complexity is abstracted which can lead to building more complexity than necessary, making it harder to maintain

- many teams build light abstractions and do very well with them

- larger organizations benefit from frameworks as it standardizes how llm applications are built, thus being able to hire and move people between teams while relying on the same abstraction -- without needing to design a generalized abstraction that works for all sorts of use cases

- many use cases do not need an "agent", implementation can be simpler and more predictable

I think you'd enjoy this very recent podcast episode on one way of thinking about building agents with SOTA models: https://www.latent.space/p/claude-sonnet

tldl: many use cases need less abstraction than initially assumed when using SOTA models, focus on tools as they have a lot of leverage on performance/quality

PS: Would love to chat in case you are interested as we currently plan many improvements with regards to agents. I just saw your message on this in the acquired slack and responded there right before getting the HN notification.

marcklingen··on Show HN: Velvet – Store OpenAI requests in your own DB
disclosure: founder/maintainer of Langfuse (OSS LLM application observability)

I believe proxy-based implementations like Velvet are excellent for getting started and solve for the immediate debugging use case; simply changing the base path of the OpenAI SDK makes things really simple (the other solutions mentioned typically require a few more minutes to set up).

At Langfuse (similarly to the other solutions mentioned above), we prioritize asynchronous and batched logging, which is often preferred for its scalability and zero impact on uptime and latency. We have developed numerous integrations (for openai specifically an SDK wrapper), and you can also use our SDKs and Decorators to integrate with any LLM.

> For my client projects, I've been leaning towards open source platforms like Arize so clients have the option of pulling it inhouse if needed. Most often for HIPAA requirements.

I can echo this. We observe many self-hosted deployments in larger enterprises and HIPAA-related companies, thus we made it very simple to self-host Langfuse. Especially when PII is involved, self-hosting makes adopting an LLM observability tool much easier in larger teams.

marcklingen··on Deploy Your Own LLM Observability Platform in 5 Minutes [video]
Langfuse co-founder/maintainer here, while the underlying architecture will change in v3 (thanks for referencing the thread), the docker-composed based self-hosting setup will remain as simple as shown in the video

verdverm, thanks for all your feedback by the way!

marcklingen··on Show HN: Adding Mistral Codestral and GPT-4o to Jupyter Notebooks
Big fan of LiteLLM Proxy and LiteLLM Python SDK to connect to various local models. Might be helpful here as well
marcklingen··on Roast My Docs
co-author here

we put in a lot of effort into our docs and we'd greatly appreciate any criticism or feedback! Langfuse is powerful but the docs should help beginners to quickly get started and then incrementally use more features.

docs are OSS, repo: https://github.com/langfuse/langfuse-docs

built using: https://github.com/shuding/nextra

marcklingen··on Show HN: You don't need to adopt new tools for LLM observability
Fully agree - even as a founder of an ‘LLM observability company’. Observability does not need to be reinvented to get detailed traces/metrics/logs of the LLM part of an application.

LLM Observability usually means: prompts and completions, which model was used, errors and exceptions (rate limits, network errors), as well as metrics (latency, output speed, time to first token when streaming, USD/token and cost breakdowns). All of this is well suited to be captured in the existing observability stack. OpenLLMetry makes this really easy and interoperable - chapeau.

In my view, observability is not the core value that solutions like Baserun, Athina, LangSmith, Parea, Arize, Langfuse (my project) and many others solve for. Developing a useful LLM application requires iterative workflows and tinkering. That's what these solutions help with and augment.

There are specific problems to building an LLM application such as managing/versioning of prompts, running evaluations, blending multiple different evaluation sources, collecting datasets to test/benchmark an application, helping with fine-tuning models on high-quality production completions, debugging root causes of quality/latency/cost issues, ...

Most solutions either replicate logs (LLM I/O) or traces at first, as they are a necessary starting point to then build solutions for the other workflow problems. As the observability piece gets more standardized over time, I can see how integrating with the standard makes a ton of sense. Always happy to chat about this.

marcklingen··on How are generative AI companies monitoring their systems in production?
We struggled with this ourselves while building LLM-based products and then open-sourced our observability/monitoring tool [1]. Many use it to track RAG and agents in production, run custom evals on the production traces (focused on hallucination), and track how metrics are different across releases or customers. Feel free to dm if there is something specific you are looking to solve, happy to help.

[1] https://github.com/langfuse/langfuse

marcklingen··on Show HN: Langfuse – Open-source observability and analytics for LLM apps
Agree – Langfuse stores all prompts/completions, model configuration and metadata. Currently the GET API can be used to use the data for finetuning and we build a wrapper to access a filtered sample via the Python SDK.
marcklingen··on Ask HN: Any tools or frameworks to monitor the usage of OpenAI API keys?
Langfuse founder here

Do the API keys map to end users / customers? If yes reporting via Langfuse (UI & GET API) can solve this.

If that's not why you use multiple API keys, you can add an api key id/hash (to not store the key in langfuse) as metadata and then group/filter by this metadata field.

Dm me on Discord[1] if you have more detailed questions.

[1] https://langfuse.com/discord

Page 1 of 2Next →