501 karma · joined May 14, 2017
We plan on adding more model integrations, but it is completely decoupled from the method implementation.
Note that we still need to manage the KV cache in outlines. It’s a small interface change that will be made this week hopefully, but we’ve been focusing on constrained generation so far.
Our method is much more efficient. llama.cpp loops over the entire vocabulary (~50k tokens) at each step to generate the mask. We generate an index at initialization, and building the masks at each step only requires a dictionary lookup (trade speed for memory). Sampling is just as fast as standard sampling.
(1) This feature only requires regex-guided generation. We have a PR for BNF sampling that is about to be merged. (2) ggml loops over the entire vocabulary (~50k tokens) at each step, which introduces a noticeable overhead, and makes it unusable for complex grammars. Our method works by building an index at initialization, and build the masks at each step with a dictionary lookup. Once the index is built, generation is just as fast as standard generation. Doesn't depend on the complexity of the grammar, the size of the LLM or its vocabulary size.
Our method on the other guarantees that the output will follow the specs of the JSON schema. No need to call the LLM several times.
In early stages, for instance, I would look for people who have a shipping fast mentality. In later stage for people who value technical excellence.
Everything else is at best childish and naive, and too often discriminating against certain persons who would actually be a valuable asset. A grown up can work efficiently in a team without going to Thursday beers every week.
I don’t think this schedule would work for everyone but the general idea was: intense work for a short periods of time, take some time to talk with colleagues and go through meetings and then work again, with long breaks in between. That way I could manage 8 hours of productivity without burning out. And yes, sometimes I would get in flow and work for 12 hours straight without eating. But those days were more the exception than the rule. The point is I could have roughly 8 hours most days working this way.
In companies I’ve worked with I’ve always felt babysitted, as though I was unable to discipline myself when not watched all the time by managers. The truth is we don’t all work in the same way, and we are all reasonably interested in our job—-and if we’re not, sitting all day in the office is not going to change that. So why don’t we make room for everyone’s pattern while keeping some team time every day?
I often think that shy people are assholes at first. Now that I’m aware of that I try to test the shyness hypothesis first when I have that feeling.
So exploit it, it’s something humans developed for good reasons. Just be aware of the biais and use it as a good starting point for a behavioral interview.
The other books I read make the field look like a bunch of heuristics that just happen to work.
Business value comes first, and with a better infrastructure we could deliver a lot more value with the same head count.
I don't, which is why I'm asking around. I'm also scheduling chats with people to understand the background and the history to see if it's worth changing anything. I will make a move if that make sense. In the meantime, I'm just gathering information to not make a stupid decision.
> And you're a perm now, in a big corp, doing data science. Relax, you got it made, right?
> So enjoy playing that game, because you are happy to be in a big corp, so the sort of benefits that can bring you is what you want, right?
I don't think this was necessary. I've only worked for startups before, and I was hired in part to see if we can do a better job with the resources we have. Buy in from management is not an issue. I am not asking for life advice.
> Moves from an architecture that is clustered for scale (ie. spark) to one that only scales vertically
I did a quick estimate of the volume, and we won't reach 1Tb before > 5 years. We're not in a line of business where the number of clients can increase dramatically so it's fairly predictable. I don't want to design for imaginary scaling issues.
> Potentially introduces yet more sources of truth for some data.
It is more intended to replace the current mess.
> SQL is terrible language to write transformations in (its a query language, not an ETL pipeline)
Actually this is the point that concerns me the most. The need to transform the data in non-trivial ways. But surely people didn't wait for Spark to do this?
> Unless you can very clearly demonstrate that what you're making is meaningfully better
This is a very good point, and I think I should come up with a quick POC to demonstrate and get buy-in.
> Could you perhaps find better way to orchestrate your spark tasks, eg. with airflow or ADF or AWS Glue or whatever?
I feel that it would just be solving the mess by adding more mess.