Chronon, Airbnb's ML feature platform, is now open source
medium.com
medium.com
With LLMs, it's more like you buy a large pancake machine that you dump all of your compost into (and you suspect the installers might have hooked up to your sewage line as input too). It triples your electricity bill, it makes bizarre screeching noises as it runs, you haven't seen your cat in a week, but at the end out come some damn fine pancakes.
I apologize. I'm talking about the thing that I was saying was a relief to be not talking about.
FWIW, Chronon does serve context within prompts to personalize LLM responses. It is also used to time-travel new prompts for evaluation.
What does this mean?
The user might say - I need a refund. The bot needs to know contextual information - order details, delivery tracking details etc.
Now you have written a prompt template that needs to be rendered with contextual information. This rendered prompt is what the model will use to decide whether to issue a refund or not.
Before you deploy this prompt to prod, you want to evaluate its performance - instances where it correctly decided to issue or decline a refund.
To evaluate, you can “replay” historical refund requests. The issue is that the information in the context changes with time. You want to instead simulate the value of the context at a historical point in time - or time-travel.
Can you give an example that's not ripe for abuse? This really doesn't sell LLMs as anything useful except insulation from the consequences of bad decisions.
I can see no reason why it would be illegal or inappropriate to use an LLM as part of the initial flow there. In fact I see no reason why it would be illegal for Amazon to simply flip a coin to decide whether to immediately accept your refund. (Appropriateness is another matter!)
I guess you're assuming the LLM would be the only point of contact with no recourse if it rejects you? Which strikes me as very pessimistic, unless you live in a very poorly regulated country.
[0] https://dspy.ai/
Nah, this is nothing new.
We've solved this for ages with "snapshots" or "archives", or fancy indexing strategies, or just a freaking "timestamp" column in your tables.
Standard OLAP solutions right now are really good at "What's the X day sum of this column as of this timestamp", but when every row of your training data has a precise intra-day timestamp that you need windowed aggregations to be accurate as-of, this is a different challenge.
And when you have many people sharing these aggregations, but with potentially different timestamps/timelines, you also want them sharing partial aggregations where possibly for efficiency.
All of this is well beyond the scope that is addressed by standard OLAP data solutions.
Not to mention the fact that the offline computation needs to translate seamlessly to power online serving (i.e. seeding feature values, and combining with streaming realtime aggregations), and the need for online/offline consistency measurement.
That's why a lot of teams don't even bother with this, and basically just log their feature values from online to offline. But this limits what kind of data they can use, and also how quickly they can iterate on new features (need to wait for enough log data to accumulate before you can train).
As long as your OLAP table/projection/materialized view is sorted/clustered by that timestamp, it will be able to efficiently pick only the data in that interval for your query, regardless of the precision you need.
> And when you have many people sharing these aggregations, but with potentially different timestamps/timelines, you also want them sharing partial aggregations where possibly for efficiency.
> All of this is well beyond the scope that is addressed by standard OLAP data solutions.
I think the StarRocks open-source OLAP DB supports this as a query rewrite mechanism that optimizes performance by using data from materialized views. It can build UNION queries to handle date ranges [1]
[1] https://docs.starrocks.io/docs/using_starrocks/query_rewrite...
You need these values accurate as of ~500k timestamps for 10k different page ids, with significant skew for some page ids.
So you have a "left" table with 500k rows, each with a page id and timestamp. Then you have a `page_views` table with many millions/billions/whatever rows that need to be aggregated.
Sure, you could do this with backfill with SQL and fancy window functions. But let's just look at what you would need to do to actually make this work, assuming you wanted it to be serving online with realtime updates (from a page_views kafka topic that is the source of the page views table):
For online serving: 1. Decompose the batch computation to SUM and COUNT and seed the values in your KV store 2. Write the streaming job that does realtime updates to your SUMs/COUNTs. 3. Have an API for fetching and finalizing the AVERAGE value.
For Backfilling: 1. Write your verbose query with windowed aggregations (I encourage you to actually try it). 2. Often you also want a daily front-fill job for scheduled retraining. Now you're also thinking about how to reuse previous values. Maybe you reuse your decomposed SUMs/COUNTs above, but if so you're now orchestrating these pipelines.
For making sure you didn't mess it up: 1. Compare logs of fetched features to backfilled values to make sure that they're temporally consistent.
For sharing: 1. Let's say other ML practitioners are also playing around with this feature, but with a different timelines (i.e. different timestamps). Are they redoing all of the computation? Or are you orchestrating caching and reusing partial windows?
So you can do all that, or you can write a few lines of python in Chronon.
Now let's say you want to add a window. Or say you want to change it so it's aggregated by `user_id` rather than `page_id`. Or say you want to add other aggregations other than AVERAGE. You can redo all of that again, or change a few lines of Python.
Isn’t this just a table with 5bn rows of timestamp, page_type, page_views_t1d, page_views_t7d, page_views_t30d, page_views_t60d, and page_views_t180d? You can even compute this incrementally or in parallel by timestamp and/or page_type.
What’s the magic Chronon is doing?
But even for offline computation, for the same computation logic, the code will be duplicated in lots of places. we have observed the ML practitioners copied sql queries all over. In the end, it is not possible for debugging, feature interpretability and lineage.
Chronon abstracts all those away so that ML practitioners can focus on the core problems they are dealing with, rather than spending time on the ML Ops.
For an extreme use case, one user defined 1000 features with 250 lines of code, which is definitely impossible with SQL queries, not to even mention the extra work to serve those features.
For the offline computations, we will reuse those intermediate results to avoid calculation from the beginning again. So the engine can actually scale sub-linearly.
> So you can do all that, or you can write a few lines of python in Chronon.
It all seems a bit handvwavy here. Will Chronon work as well as the SQL version or be correct? I vote for an LLM tool to help you write those queries. Or is that effectively what Chronon is doing?What does this “last” operation do? There’s definitely a LAST_VALUE() window function in the databases I use. It is available in Postgres, Redshift, EMR, Oracle, MySQL, MSSQL, Bigquery, and certainly others I am not aware of.
Actually, Last is usually called last_k(n), so that you can specify the number of the values in the result array. For example, if the input column is page_view_id and n = 300, it will return the last 300 page_view_id as an array. If a window is used, for example, 7d, it will truncate the results to the past 7d. The LAST_VALUE() seems to return the last value from an ordered set. Hope that helps. Thanks for your interests.
We've also had slowly changing dimensions to solve this type of problem for a decent amount of time for the labels that sit on top of everything, though really these are just fact tables with a similar historical approach.
But please enlighten us on which databases to use so Airbnb (and the rest of us) can stop wasting time.
We've not been developing v2 with ML feature serving in mind so far, but I would love to speak with anyone interested in this use case and figure out where the gaps are.
What now?
To evaluate if a feature is valuable, you could attach the value of the feature to past inferences and retrain a new model to check for improvement in performance.
But this “attach”-ing needs the feature value to be as of the time of the past inference.
For real-time ml systems, it give uou row oriented retrival latencies for features.
Most importantly, it helps modularize your ML system into feature pipelines training pipelines, and inference pipelines. No onolithic ML pipelines.
Data warehouse is more for relatively unstructured or blobby data with moderate read access and capacity for massive files.
OLAP is mostly for feeding streaming and event-driven flows, including but not limited to ML.
another one is the focus on pushing as much compute to the write-side as possible (within Chronon) - specially joins and aggregations.
OLAP databases and even graph databases don't scale well to high read traffic. Even when they do, the latencies are very high.
[1] https://www.starrocks.io/ [2] https://github.com/StarRocks/starrocks [3] https://www.starrocks.io/blog/starrocks-vs-clickhouse-the-qu...
Fundamentally most of the computation needs to happen before the read request is sent.
Tecton, we evaluated, but decided that the time-travel strategy wasn’t scalable for our needs at the time.
A philosophical difference with tecton is that, we believe the compute primitives (aggregation and enrichment) need to be composable. We don’t have a FeatureSet or a TrainingSet for that reason - we instead have GroupBy and Join.
This enables chaining or composition to handle normalization (think 3NF) / star-schema in the warehouse.
Side benefit is that, non ml use-cases are able to leverage functionality within Chronon.
I agree that my statement would be much better if used snowflake schema instead.
OLAP systems are fundamentally designed to scale the read path - former approach. Feature serving needs the latter.
- Since the platform is designed to scale, it would be nice to see scalability benchmarks
- Is the platform compatible with human-in-the-loop workflows? In my experience, those workflows tend to require vastly different needs than fully automated workflows (e.g. online advertising)
re: human-in-the-loop workflows - do you mean labeling?
For our org, those are by far the most complicated to handle. Graph DBs are kind of scaling poorly, while storing state in stream processing jobs is way too large/expensive. Those would also be built on top of API sources, which then lead us to the unfortunate "log & wait" approach for our most important features
In the API itself - you could specify the chain links by specifying the source.
To be precise - a GroupBy(aggregation primitive) can have a Join(enrichment primitive) as a source. To rephrase, you can enrich first and then aggregate and continue this chain indefinitely.
> Graph DBs are kind of scaling poorly
That makes sense. Since you scaling these on the read side it is much much harder than pre-computing on the write side. (That is what Chronon allows you to do)
Online is a bit more involved - you need a month or more to test that your KV store scales against traffic coming from chronon for reads and writes.
Chronon is a full re-write of zipline with 1) a different underlying algorithm for time-travel to address scalability concerns. 2) a different serde and fetching strategy to address latency concerns.
Someone mentioned they wanted to add cadence support.
- inability to back-test new real-time features. People were forced to log-and-wait to create training sets for months. Chronon reduces this to hours or days.
- the difficulty of creating the lambda system (batch pipeline, streaming pipeline, index, serving endpoint) for every feature group. In chronon, you simply set a flag on your feature definition to spin up the lambda system.
https://www.chronon.ai/authoring_features/Source.html#stream...
Yes, exactly! I see there is some kind of support, but is it possible to use the OLTP database as an event source?
For example, say I had a table in my OLTP database, `data.returns`, that had columns `ts`, the event time, and `status`, which can be PENDING or COMPLETED. I'd like to generate point-in-time correct training data, where the feature is the count of completed returns. It seems like all the necessary information to calculate this is there
Admittedly some of the transforms proposed in this article are a little simple & don't represent the full space of feature eng requirements for all large orgs
We have plans to implement this behavior for computing the batch arm of feature serving.
and probably the only link you care about: https://github.com/airbnb/chronon#readme (Apache 2)