LMQL: A query language for programming (large) language models
github.com
github.com
Be sure to play around with our entirely web-based playground IDE here: https://lmql.ai/playground and our example showcase at https://lmql.ai. We are happy to answer any questions that come up.
Re Node.js: We are actively investigating LMQL use from other languages than Python (see also https://github.com/eth-sri/lmql/issues/1). The current interpreter is quite closely tied to a Python environment, meaning to use it from Node.js, you will at least have to host a Python interpreter as a subprocess for now. We are actively looking into gRPC/inter-process communication though, hoping to improve on this a bit in the near future.
Here's the flow I've been picturing while writing code writing models. Feel free to take this idea anyone, would appreciate a heads up if it works and you publish sota before I get around to it :)
Let's give llms ide level info. As they type, if there's a function they're starting to call use a language server to get tooltip docs and put it in a comment just above the line it's writing. Put auto complete help, type docs, etc in their context while they code.
Edit
Second thought, this kind of thing feels very low in compute compared to the LLM calculations, is this the kind of thing that could be passed up to a remote service as a wasm bundle/similar, to control the streaming output?
Indeed, this is a form of LLM prompting that we are also exploring in our preview release channel with something called in-context functions. See https://next.lmql.ai and choose the "In-Context Functions" example in the New Features showcase screen.
With in-context functions, you can provide additional instructions/data to the LLM, that will apply locally only. Once such an in-context function returns, additional instructions/retrieved info is removed and only the LLM-generated end-result remains.
LMQL gives you a concise way to define multi-part prompts and enforce constraint on LLMs. For instance, you can make sure the model always adheres to a specific output format, where parsing of the output is automatically taken care of. Also abstracts a number of things like APIs and local models, tokenisation, optimisation and makes tool integration (e.g. tool function calls during LLM reasoning) much easier.
In practice this saves you a lot of ugly text concatenation and output parsing code, letting you focus on the core logic of your project. Overall, however, you will still use your host language to call LMQL. E.g. we are fully integrated with Python, where LMQL query code simply lives in decorated functions (https://docs.lmql.ai/en/latest/python/python.html).
Per query, the LMQL runtime calls the underlying LM several times, to execute the complete specified (multi-part) query program.
It does not only translate to only one LM prompt, but rather a sequence of prompts, where during generation additional constraining of the LM is applied, to ensure the LM behaves according to the provided template. That's how it support control-flow and external function calls during generation, it actually executes the queries with a proper runtime and only uses LLMs on the backend.