HNHacker News
TopNewBestAskShowJobs

nikhilsimha

308 karma · joined January 18, 2014

submissionscomments
nikhilsimha··on ATProto for Distributed Systems Engineers
Shameless plug, we built Chronon to solve a very similar problem. Chronon is now open source (apache 2) and used at Netflix, OpenAI, Airbnb, Stripe etc.

It is used to serve features/context to ML/AI models and rule engines, along with app views.

We are more scale and perf oriented as project given the use-cases and less decentralization oriented.

We also ship with a dataframe / prql like declarative api.

[1] https://chronon.ai

nikhilsimha··on Andy Pavlo joins ClickHouse to establish ClickHouse Labs
Congrats! Andy is incredible!
nikhilsimha··on Benchmarking Opus 5 on SlopCodeBench
the last note about the agent sending emails without approvals is a very common problem!
nikhilsimha··on A way to exclude sensitive files issue still open for OpenAI Codex
Files that codex and any other coding agent has access to, should be opt-in NOT opt-out. I think codex is not the right layer to solve this if you want a sane(one-click) UX. We built our own internal sandboxing-terminal around claude and codex. Where a user-configured base-folder with low-risk code and creds is COPIED into the sandbox BEFORE new session creation. There were many other UX related reasons to build our own terminal. Can share more if anyone is interested.
nikhilsimha··on Nobody ever got fired for using a struct
It is very common to find tables with 1000+ columns in machine learning training sets at e-commerce companies. The largest I have seen had over 10000 columns.
nikhilsimha··on Qwen3.5: Towards Native Multimodal Agents
easily the most memorable comment i have ever seen on hackernews so far. kudos good sir!
nikhilsimha··on Meditation as Wakeful Relaxation: Unclenching Smooth Muscle
usually when a sharp sensation arises in an area, there is a habitual tendency to counteract - unconsciously tense surrounding muscles or antagonistic muscles or switch posture etc.

the idea is to observe with clarity the counteraction and let the sharp sensation arise and pass without the counteraction/resistance.

nikhilsimha··on Some thoughts on autoregressive models
Not saying that our current approaches will lead to intelligence. No one can know.

It could very well be that the internal mechanism of our thought has an auto-regressive reasoning component.

With the full system effectively "combining" short term memory (what just happened) and "pruned" long-term memory (what relevant things i know from the past) and pushing that into a RAW autoregressive reasoning component.

It is also possible that another specialized auto regressive reasoning component is driving the "prune" and "combine" operations. This whole system could be solely represented in the larger network.

The argument that "intelligence cannot be auto-regressive" seems to be without basis to me.

> there is strong evidence that not all thinking is linguistic or sequential.

It is possible that a system wrapping a core auto-regressive reasoner can produce non-sequential thinking - even if you don't allow for weight updates.

nikhilsimha··on F-strings for C++26 proposal [pdf]
just skimmed the proposal, dont see how inline rendered f-strings are more complicated than the alternative.
nikhilsimha··on Feldera Incremental Compute Engine
gotcha! thanks for the clarification
nikhilsimha··on Feldera Incremental Compute Engine
So the writes are O(N) then - to keeps reads at O(1)?
nikhilsimha··on Feldera Incremental Compute Engine
dumb question: how do z-sets or feldera deal with updates to values that were incorporated into the max already?

For example - max over {4, 5} is 5. Now I update the 5 to a 3, so the set becomes {4, 3} with a max of 4. This seems to imply that the z-sets would need to store ALL the values - again, in their internal state.

Also there needs to be some logic somewhere that says that the data structure for updating values in a max aggregation needs to be a heap. Is that all happening somewhere?

nikhilsimha··on Functional languages should be so much better at mutation than they are
i personally like nim’s approach to memory management - implicitly refcounted, but exposes clean manual memory management when needed
nikhilsimha··on Maestro: Netflix's Workflow Orchestrator
great job on open sourcing!
nikhilsimha··on Chinese AI stirs panic at European geoscience society
publicly funded research, but behind paywalls, was scraped to build the chatbot - by “china” not open ai, causes “people” to lose their s**t.

i do think ip infringement is not cool in general - but it doesnt seem right that geo research is private property.

nikhilsimha··on Pg_lakehouse: A DuckDB Alternative in Postgres
duckdb is mit licensed. [1]

datafusion is apache v2 licensed. [2]

pg_lakehouse built on top of data fusion is AGPL v3 + business licensed. [3]

[1] https://github.com/duckdb/duckdb/blob/main/LICENSE

[2] https://github.com/apache/datafusion/blob/main/LICENSE.txt

[3] https://github.com/paradedb/paradedb/blob/dev/LICENSE

Most companies won't touch AGPL v3 license. Maybe not a bad thing, but FYI.

nikhilsimha··on The search for easier safe systems programming
Nim?
nikhilsimha··on Chronon, Airbnb's ML feature platform, is now open source
almost every button click is either powered by a model or guarded by a model.
nikhilsimha··on Chronon, Airbnb's ML feature platform, is now open source
We don't accept hints yet - but we determine what to cache.
nikhilsimha··on Chronon, Airbnb's ML feature platform, is now open source
To be pedantic, even in star schema - the dim tables are denormalized, fact tables are not.

I agree that my statement would be much better if used snowflake schema instead.

nikhilsimha··on Chronon, Airbnb's ML feature platform, is now open source
"Imagine" is the operative word :-)
nikhilsimha··on Chronon, Airbnb's ML feature platform, is now open source
These are reconstruction of features / columns that don’t exist yet.
nikhilsimha··on Chronon, Airbnb's ML feature platform, is now open source
Normalization is overloaded. I was referring to schema normalization (3NF etc) not feature normalization - like standard scaling etc.
nikhilsimha··on Chronon, Airbnb's ML feature platform, is now open source
We do this for training data generation already.

We have plans to implement this behavior for computing the batch arm of feature serving.

nikhilsimha··on Chronon, Airbnb's ML feature platform, is now open source
Snapshots can’t travel back with milliseconds precision or even minute level precision. They are just full dumps at regular fixed intervals in time.
nikhilsimha··on Chronon, Airbnb's ML feature platform, is now open source
Pardon the jargon. But it is a necessary addition to the vocabulary.

To evaluate if a feature is valuable, you could attach the value of the feature to past inferences and retrain a new model to check for improvement in performance.

But this “attach”-ing needs the feature value to be as of the time of the past inference.

nikhilsimha··on Chronon, Airbnb's ML feature platform, is now open source
Imagine you are building a customer support bot for a food delivery app.

The user might say - I need a refund. The bot needs to know contextual information - order details, delivery tracking details etc.

Now you have written a prompt template that needs to be rendered with contextual information. This rendered prompt is what the model will use to decide whether to issue a refund or not.

Before you deploy this prompt to prod, you want to evaluate its performance - instances where it correctly decided to issue or decline a refund.

To evaluate, you can “replay” historical refund requests. The issue is that the information in the context changes with time. You want to instead simulate the value of the context at a historical point in time - or time-travel.

nikhilsimha··on Chronon, Airbnb's ML feature platform, is now open source
We haven’t tried materialize - IIUC materialized is pure kappa. Since we need to correct upstream data errors and forget selective data(GDPR) automatically - we need a lambda system.

Tecton, we evaluated, but decided that the time-travel strategy wasn’t scalable for our needs at the time.

A philosophical difference with tecton is that, we believe the compute primitives (aggregation and enrichment) need to be composable. We don’t have a FeatureSet or a TrainingSet for that reason - we instead have GroupBy and Join.

This enables chaining or composition to handle normalization (think 3NF) / star-schema in the warehouse.

Side benefit is that, non ml use-cases are able to leverage functionality within Chronon.

nikhilsimha··on Chronon, Airbnb's ML feature platform, is now open source
Let’s say you want to compute avg transaction value of a user in the last 90days. You could pull individual transactions and average during the request time - or you could pre compute a partial aggregates and re-aggregate on read.

OLAP systems are fundamentally designed to scale the read path - former approach. Feature serving needs the latter.

nikhilsimha··on Chronon, Airbnb's ML feature platform, is now open source
Two main differences - ability to time travel for training data generation and the ability to push compute to the write side of the view rather than the read side for low latency feature serving.
Page 1 of 4Next →