421 karma · joined August 9, 2013
https://github.com/tensorzero/tensorzero
(Also, we raised the capital in 2024 and didn't burn most of it.)
The title is misleading unfortunately but that's how social media goes...
I might publish a long-form reflection when the dust settles.
The ~1% figure might be outdated today but it was a best-effort estimate a couple of months ago. TensorZero powered tens of trillions of inference tokens per month. TensorZero is not widely used but it was used by a couple of extreme-scale users.
We are returning the remaining capital to investors.
We started the company two and a half years ago, and raised $7.3m in 2024 (announced only almost a year later). We've spent less than half of this amount.
Earlier this week we came to the difficult decision to wind down the project. The open-source repository remains available on GitHub (Apache 2.0) but won't be actively maintained by the team moving forward.
For example:
- You write a heuristic (regex, code, etc.) that assigns a score to an output
- You make another LLM score the output from your system (aka "LLM-as-a-judge")
- You have an automated system that can verify the generated outputs (e.g. does generated code compile or pass tests?)
People often talk about "LLM evals (evaluations)" which will include a set of evaluators i.e. scoring functions.
We'll make this clearer next time!
```
from openai import OpenAI
# Point the client to the TensorZero Gateway
client = OpenAI(base_url="http://localhost:3000/openai/v1", api_key="not-used")
response = client.chat.completions.create(
# Call any model provider (or TensorZero function)
model="tensorzero::model_name::anthropic::claude-sonnet-4-6",
messages=[
{
"role": "user",
"content": "Share a fun fact about TensorZero.",
}
],
)```
You can layer additional features only as needed (fallbacks, templates, A/B testing, etc).
Thanks!
Good luck!
But broadly speaking, yes, we generate data using a large model, curate the best samples using metrics from the environment, and fine-tune on that data. This isn't a novel technique from an academic perspective; our focus is on applying it to different use cases (e.g. agentic RAG, agentic tool use) and models (OpenAI, Google, Qwen).
Thanks!
We chose a set of tasks with different levels of complexity to see how this approach would scale. For LLMs, the "challenge" with NER is not the task itself but the arbitrariness of the labels in the dataset. I agree it's still much simpler than the other tasks we present (agentic RAG, agentic tool use, maze navigation).
There are definitely strong parallels to model distillation and student-teacher training, with the primary difference being that we don't simply take all the data from the larger model but rather filter the dataset based on metrics from the environment. In the "Does curation even matter?" section, we show that this generally improves the result by a good margin.
We link to Vicuna, which might be the closest reference as prior art: https://lmsys.org/blog/2023-03-30-vicuna/
Thanks!
It's a WIP PR that we plan to merge soon: https://github.com/tensorzero/tensorzero/pull/2273
TensorZero is an open-source stack for industrial-grade LLM applications. It unifies an LLM gateway, observability, optimization, evaluation, and experimentation.
Open Roles:
‣ Back-end Engineering (Rust)
‣ Design Engineering
‣ Developer Relations (DevRel) Engineering
‣ Front-end Engineering (React)
‣ Product Engineering (Full-Stack)
What we offer:
‣ Vast majority of your work → open source
‣ Years of runway
‣ Small and entirely technical team: former Rust compiler maintainer, ML researchers (Stanford, CMU, Oxford, Columbia) with thousands of citations, decacorn CPO
‣ $200-300k base + up to 1% equity + benefits
‣ Onsite (5 days) in New York (Williamsburg, Brooklyn)
More information: https://tensorzero.com/candidate-brief
We plan to continue investigating how it works (+ optimize the models and prompts using TensorZero).
The Gist you shared is a good resource too though!
TensorZero | https://github.com/tensorzero/tensorzero | Staff Front-end / Design Engineer | Remote or Onsite (NYC) | Full-time or Part-time
TensorZero creates a feedback loop for optimizing LLM applications — turning production data into smarter, faster, and cheaper models.
We're looking for a contract / freelance Staff Front-end / Design Engineer with the following skillset:
‣ Must have: expert in TypeScript, React, and web fundamentals
‣ Nice to have: familiar with LLMs, experience with Vite / React Router V7 (RemixJS) / Tailwind
What we offer:
‣ Vast majority of your work → open source
‣ Flexible arrangement: remote or onsite (NYC), full-time or part-time
‣ Small and entirely technical team: former Rust compiler maintainer, ML researchers with 1000's of citations, decacorn CPO
‣ Engagement expected to last a few months
‣ Compensation in line with staff+ experience
Also hiring full-time employees: https://news.ycombinator.com/item?id=43569646
Apply: hello@tensorzero.com