Structured Outputs in the API
openai.com
openai.com
We use GPT-4o to build dynamic UI+code[0], and almost all of our calls are using JSON mode. Previously it mostly worked, but we had to do some massaging on our end (backtick removal, etc.).
With that said, this will be great for GPT-4o-mini, as it often struggles/forgets to format things as we ask.
Note: we haven't had the same success rate with function calling compared to pure JSON mode, as the function calling seems to add a level of indirection that can reduce the quality of the LLMs output YMMV.
Anyhow, excited for this!
Would you mind sharing a bit on how things have evolved?
When we first launched, the tool was very manual; you had to generate each step via the UI. We then added a "Loop Creator agent" that now builds Loops for you without intervention. Over the past few months we've mostly been fixing feature gaps and improving the Loop Creator.
Based on recent user feedback, we've put a few things in motion:
- Form generator (for manual loops)
- Chrome extension (for local automations)
- In-house Google Sheets integration
- Custom outputs (charts, tables, etc.)
- Custom Blocks (shareable with other users)
With these improvements, you'll be able to create "single page apps" like this one I made for my wife's annual mango tasting party[0].
In addition to those features, we're also launching a new section for Loop templates + educational content/how-tos, in an effort to help people get started.
To be super candid, the Loop Creator has been a pain. We started at an 8% success rate and we're only just now at 25%. Theoretically we should be able to hit 80%+ based on existing loop requests, but we're running into limits with the current state of LLMs.
As I understand it, this year's mangoes mostly came from Merritt Island, as there was some not-so-great weather in southern Florida.
Are you in central florida, or is that just your in laws? I'd love to see a talk on this at the local orlando devs meetup.
Any that you’re particularly fond of?
You can join the slack here if you want to know what meetups are happening and when: https://orlandodevs.com/slack/
highlights-text --> Right-Click --> New ML --> (smart dropdown for watch [price|name|date|{typed-in-prompt-instructions}] --> TAB --> (smart frequency - tabbing through {watch blah (and its auto-filling every N ) --> NAME_ML=ML01.
THEN:
highlights-text --> Right-Click .... WHEN {ML01} == N DO {this|ML0X} --> ML00
ML00 == EMAIL|CSV|GDrive results.
ML11 == Graph all the above outputs.
:-)
We’re working on fairly similar problems, would love to have a chat and share ideas and experiences.
In this something you’d be interested in?
You can ping me at {username}@gmail.com
That blows my mind! You have users paying for it and it only has a 25% success rate for loops created by users? I've been working for about a year on an LLM-based product and haven't launched yet because only 50-60% of my test cases are passing.
Not quite! Most of our paying users use Magic Loops to build automations semi-manually, oftentimes requiring a back-and-forth work with the Loop Creator at the start and then further prompt iterations as the Loop progresses.
The 25% statistic is from the Loop Creator "agent" (now 32% after some HN traffic :tada:) which is based on a single-shot user prompt -> Loop Creator output success rate.
The number is much higher for Loops with multiple iterations and almost 100% for manually created Loops.
tl;dr - the power of the tool isn't the Loop Creator "agent", it's the combination of code+llms that make building one-off workflows super fast and easy.
We had to a do a bunch of extra prompting to make it work, as GPT would often include backticks or broken JSON (most commonly extra commas). At the time, YAML was a much better approach.
Thankfully we've been able to remove most of these hacks, but we still use a best effort JSON parser[0] to help stream partial UI back to the client.
Practically speaking, we have quite a few use-cases where users call Loops from other Loops, so we're investigating a first-class API to generate Loops in one go.
Similar to regular software engineering, what you put in is what you get out, so we've been hesitant to launch this with the current state of LLMs/the Loop Creator as it will fail more often than not.
highlights-text --> Right-Click --> New ML --> (smart dropdown for watch [price|name|date|{typed-in-prompt-instructions}] --> TAB --> (smart frequency - tabbing through {watch blah (and its auto-filling every N ) --> NAME_ML=ML01.
THEN:
highlights-text --> Right-Click .... WHEN {ML01} == N DO {this|ML0X} --> ML00
ML00 == EMAIL|CSV|GDrive results.
ML11 == Graph all the above outputs.
:-)
--
A MasterLoop would be good - where you have all [public or private] loops register - and then you can route logic based on loops that exist - and since its promptable - it can summarize and suggest logic when weaving loops into cohesive lattices of behavior. And if loops can subscribe to the output of other loops -- when youre looking for certain output strands -- you can say:
Find all the MLB loops daily and summarize what they say about only the Dodgers and the Giants - and keep a running table for the pitchers, and catchers stats only.
EDIT: Aside from just subscribing, maybe a loop can #include# the loop in its behavior to accomplish its goal/assimilate its function / mate/spawn. :-)
These loops are pretty Magic!
This is actually where we started :)
We had trouble keeping all the context in via RAG and the original 8k token window, but we're aiming to bring this back in the (hopefully near) future.
The 50% drop in price for inputs and 33% for outputs vs. the previous 4o model is huge.
It also appears to be topping various benchmarks, ZeroEval's Leaderboard on hugging face [0] actually shows that it beats even Claude 3.5 Sonnet on CRUX [1] which is a code reasoning benchmark.
Shameless plug, I'm the co-founder of Double.bot (YC W23). After seeing the leaderboard above we actually added it to our copilot for anyone to try for free [2]. We try to add all new models the same day they are released
[0]https://huggingface.co/spaces/allenai/ZeroEval
The previous version of 4o also beat 3.5 Sonnet on Crux.
We also get pretty reliable JSON output (on a smaller scale though) even without JSON mode. We usually don't use JSON mode because we often include a chain of thought part in <brainstorming> and then ask for JSON in <json> tags. With some prompt engineering, we get over 98% valid JSON in complex prompts (with long context and modestly complex JSON format). We catch the rest with json5.loads, which is only used as a fallback if json.loads fails.
4o-mini has been less reliable for us particularly with large context. The new structured output might make it possible to use mini in more situations.
1. Reliable structured outputs 2. Reduced costs by 50% for input, 33% for output 3. Up to 16k output tokens compared to 4k
My take is that the leaderboards and benchmarks are still very flawed if you're using LLMs for any non-chat purpose. In the product I'm building, I have to use all of the big 4 models (GPT, Claude, Llama, Gemini), because for each of them there is at least one tasks that it performs much better than the other 3.
I suspect the reason why most big LLMs have ended up in a pretty verbose spot is that it's easier for users to scroll & skim than to ask follow-up questions (which requires formulation + typing + waiting for a response).
With regard to this new gpt-4o model: you'll find it actually bucks the recent trend and is less verbose than its predecessor.
Also your about page is very suspicious for someone at an AI company ;).
Maybe it's a 'technical' user divide, but that seems wrong to me. I would much rather a succinct answer that I can probe further or clarify if necessary.
Lately it's going against my custom prompt/profile whatever it's called - to tell it to assume some level of competence, a bit about my background etc., to keep it brief - and it's worse than it was when I created that out of annoyance with it.
Like earlier I asked something about some detail of AWS networking and using reachability analyser with VPC endpoints/peering connections/Lambda or something, and it starts waffling on like 'first, establish the ID of your Virtual Private Cloud Endpoint. Step 1. To locate the ID, go to ...'
Human users are charged by the number of messages, so longer responses are preferable because follow up questions use up your message allowance.
APIs are charged by token so shorter messages are preferable as you don’t pay for unnecessary tokens.
I had it again earlier:
Me: give me a bucket policy for write access from alb
CGPT: [waffle about IP ranges that is totally incorrect; then starts telling me ALB doesn't typically write to S3 because it's usually an intermediary between clients and backend services like EC2 instances or Lambda functions - it already knows from chat context I am using the latter]
Me: [whacks stop because it's rapidly getting out of hand] yes it does for access and connection logs
CGPT: To allow Application Load Balancer (ALB) to write access and connections logs to an S3 bucket, you need to set up a bucket policy that [waffle waffle waffle]
Me: [stop] yes I know that's what I asked for
CGPT: Here is an example of an S3 bucket policy [...]
Me: invalid principal [as far as I can tell, a complete hallucination]
CGPT: [tries again]
Me: yes I already tried that, valid policy but ALB still doesn't have permission
CGPT: [nonsense intensifies]
In the end I sorted it much quicker from AWS docs, which is sort of saying something, because I do often struggle with them. Thought I'd give ChatGPT a chance here but it really wasn't helpful.
For me this is one of the strongest motivators for running LLMs locally- even if they’re measurably worse, they’re a far better tool because they don’t change behavior over time.
Fwiw my subjective experience has been that non-technical stakeholders tend to be more impressed with / agreeable to longer AI outputs, regardless of underlying quality. I have lost count of the number of times I’ve been asked to make outputs longer. Maybe this is just OpenAI responding to what users want?
[1] https://sophiabits.com/blog/new-llms-arent-always-better#exa...
> You may output only up to 500 words, if the best summary is less than 500 words, that's totally fine. If details are unclear, do not fill-in gaps, do leave them out of the summary instead.
Surprised it took them so long — llama.cpp got this feature 1.5 years ago (actually an even more general version of it that allows the user to provide any context free grammar, not just JSON schema)
Does it keep validating the predicted tokens and backtrack when it’s not valid?
There are contrived grammars you can give it that will make it use exponential memory, but in practice most real-world grammars aren't like this.
Is this just a schema validation layer on their end to avoid the round trip (and cost) of repeating the call?
The simplest algorithm for getting good quality output is to just always pick the highest probability token.
If you want more creativity, maybe you pick randomly among the top 5 highest probability tokens or something. There are a lot of methods.
All that grammar-constrained decoding does is zero out the probability of any token that would violate the grammar.
> The model can fail to follow the schema if the model chooses to refuse an unsafe request. If it chooses to refuse, the return message will have the refusal boolean set to true to indicate this.
I'm not sure how they implemented that, maybe they've figured out a way to give the grammar a token or set of tokens that are always valid mid generation and indicate the model would rather not continue generating.
Right now JSON generation is one of the most reliable ways to get around refusals, and they managed not to introduce that weakness into their model
# Very Limited Field Typing
OpenAI offers a very limited set of types[2] (String, Number, Boolean, Object, Array, Enum, anyOf) without the ability to define patterns and max/min lengths. Outlines supports defining arbitrary RegEx patterns, making extracting currencies, phone numbers, zip codes, comma-separated lists, and more a trivial exercise.
# High Schema Setup Cost / Latency
vLLM and Outlines offer near-zero cost schema setup: RegEx finite state machine construction is extremely cheap on the first inference call. While OpenAI's context-free grammar generation has a significant latency penalty of "under ten seconds to a minute". This may not impact "warmed-up" inference but could present issues if schemas are more dynamic in nature.
Right now, this feels like a good first step, focusing on ensuring the right fields are present in schema-ed output. However, it doesn't yet offer the functionality to ensure the format of field contents beyond a primitive set of types. It will be interesting to watch where OpenAI takes this.
[0] https://help.getzep.com/structured-data-extraction
[1] https://help.getzep.com/dialog-classification
[2] https://platform.openai.com/docs/guides/structured-outputs/s...
The realisation of the tech might be fantastic new things… or it might be that people like me are Clever Hans-ing the models.
* that may be the wrong word; "strong capabilities" is what I think is present, those can be used for ill effects which is pessimistic.
Not guaranteed even with the same seed. If you don't perform all operations in exactly the same order, even a simple float32 sum, if batched differently, will result in different final value. This depends on the load factor and how resources are allocated.
They were lying, of course, and meanwhile charged output tokens for malformed JSON.
Other LLM vendors figured this out many months ago.
I highly doubt it's the model that does this... It's very likely code injected into the token picker. You could put this into any model all the way down to gpt-2.
My assumption is if that's all this is they would have done it a long time ago though.
There are quite a few open source implementations of this e.g. https://github.com/outlines-dev/outlines
1. There is always a valid next token.
2. This greedy algorithm doesn't result in a qualitatively different distribution from a rejection sampling algorithm.
The latter isn't too obvious, and may in fact be (very) false. Look up maze generation algorithms if you want some feeling for the effects this could have.
If you just want a quick argument, consider what happens if picking the most likely token would increase the chance of an invalid token further down the line to nearly 100%. By the time your token-picking algorithm has any effect it would be too late to fix it.
If you want to be really clever about your picker, a deterministic result would blat out the all the known possible strings.
For example, if you had an object with defined a defined set of properties, you could just go ahead and not bother generating tokens for all the properties and just tokenize, E.G. `{"foo":"` (6-ish tokens) without even passing through the LLM. As soon as an unescaped `"` arrives, you know the continuation must be `,"bar":"`, for example
> This greedy algorithm doesn't result in a qualitatively different distribution from a rejection sampling algorithm.
It absolutely will. But so will adding an extra newline in your prompt, for example. That sort of thing is part and parcel of how llms work
That said if you know all valid prefixes in your language in advance then you can always realise when a token leaves no valid continuations.
> It absolutely will. But so will adding an extra newline
A newline is less likely to dramatically drop the quality, a greedy method could easily end driving itself into a dead end (if not grammatically then semantically).
Say you want it to give a weather prediction consisting of a description followed by a tag 'sunny' or 'cloudy' and your model is on its way to generate
{
desc: "Strong winds followed by heavy rainfall.",
tag: "stormy"
}
If it ever gets to the 's' in stormy it will be forced to pick 'sunny', even if that makes no sense in context.edit: oh actually, we do sort of know -- they call out jsonformer as an inspiration in the acknowledgements
This might be even more pronounced when the output is restricted more using the JSON schema.
So the heavy lifting here was most likely to align the model to avoid/minimize such outcomes, not in tweaking the token sampler.
FWIW (not too much!) I have used llama.cpp grammars to restrict to specific formats (not particular json, but an expected format), fine-tuned phi2 models, and I didn't hit any issues like this.
I am not intuitively seeing why restricting sampling to tokens matching a schema would cause the LLM to converge on valid tokens that make no sense...
Are there examples of this happening w/ people using e.g. jsonformer?
It's kind of like an inverse Moravec's paradox.
The wild part is that a model trained with so much human language text can still outputs mostly compilable code.
When we first started using the OpenAI API's, the first thing I reached for was some way "to be certain that the response is properly formatted". There wasn't one. A common solution was (is?) "just run the query again, untill you can parse the JSON". Really? After decades of software engineering, we still go for the "have you tried turning it off and on again" on all levels.
Then I reached for common, popular tools: everyone's using them, they ought to be good, right? But many of these tools, from langchain's to dify to promptflow are a mess (Edit: to alter the tone: I'm honestly impressed by the breadth and depth of these tools. I'm just suprised about the stability - lack thereof, of them). Nearly all of them suffer from always-outdated-documentation. Several will break right after installing it, due to today's ecosystem updates that haven't been incorporated entirely yet. Understandably: they operate in an ecosystem that changes by the hour. But after decades of software engineering, I want stuff that's stable, predictable, documented. If that means I'm running LLM models from a year ago: fine. At least it'll work. Sure, this constant-state-of-brokeness is fine for a PoC, a demo, or some early stage startup. But terrible for something that I want to ensure to still work in 12 months, or 4 years even without a weekly afternoon of upgrade-all-dependencies-and-hope-it-works-the-update-my-business-logic-code-to-match-the-now-deprecated-apis.
In the same way people revert to older stable releases. You're welcome to revert to writing boilerplate code yourself.
The reason people are excited and use it, is because they show promise, it already offers significant benefits even if it isn't "stable".
My problem is that this community, and thus the ecosystem, is repeating many mistakes, reinventing wheels, employing known bad software-engineering practices, or lacking any software-engineering practices at all. It is, in short, not learning from decades of work.
It's not standing on shoulders of giants, it's cludging together wobbly ladders to race to the top. Even if "getting to the top first" is the primary goal, it's not the fastest way. But certainly not the most durable way.
(To be clear: I have seen the JS (node) community doing the exact same, racing fast towards cliffs and walls, realizing this, throwing it all out, racing to another cliff and so on. And I see many areas in the Python community doing this too. Problems have been solved! For decades! Learn from these instead of repeating the entire train of failures to maybe finally solve the problems in your niche/language/platform/ecosystem)
I also learned about constrainted decoding[1], that they give a brief explanation about. This is a really clever technique! It will increase reliability as well as reduce latency (less tokens to pick from) once the initial artifacts are loaded.
If your schema is not supported, but you still want to use the model to generate output, you would use `strict: false`. Unfortunately we cannot make `strict: true` the default because it would break existing users. We hope to make it the default in a future API version.
For example, if I ask an LLM to generate social security numbers, it will give the whole "I'm sorry Hal, I can't do that". If I ban all tokens except numbers and hyphens, prior to your "refusal = True" approach, it was guaranteed that even "aligned" models would generate what appeared to be social security numbers.
Christ, I hate the AI safety people who brain-damage models so that they refuse to do things trivial to do by other means. Is LLM censorship preventing bad actors from generating social security numbers? Obviously not. THEN WHY DOES DAMAGING AN LLM TO MAKE IT REFUSE THIS TASK MAKE CIVILIZATION BETTER OFF?
History will not be kind to safetyist luddites.
That said, when zero "safety" is at stake might be the best time to experiment with how to build and where to put safety latches, for when we get to a point we mean actual safety. I'm even OK with models that default to parental control for practice provided it can be switched off.
The wider consumer base is practically itching for a reason to avoid this tech. If it gets out, that could be a problem.
It was the same issue with Geminis image gen. Sure, the problem Google had was bad. But could you even imagine what it would've been if they did nothing? Meaning, the model could generate horrendously racist images? That 100% would've been a worse PR outcome. Imagine someone gens a mistral image and accredits it to Google's model. Like... that's bad for Google. Really bad.
Hence all the people leaving, too.
> OpenAI's President and co-founder Greg Brockman is also taking a sabbatical through the end of the year, he said in a X post late Monday.
> Peter Deng, a vice-president of product, also left in recent months, a spokesperson said. And earlier this year, several members of the company’s safety teams exited.
That's after co-founder and Chief Scientist Ilya Sutskever left in May.
GPT-4 has been so mind-blowingly cool, but most of the interesting applications I can think of involve 10 steps of “ok now make sure GPT has actually formatted the question as a list of strings… ok now make sure GPT hasn’t responded with a refusal to answer the question…”
Idk what the deal is with their weird hype persona thing, but I’m stoked about this release
``` openai.BadRequestError: Error code: 400 - {'error': {'message': 'Invalid schema for response_format \'PolicyStatements\': schema must be a JSON Schema of \'type: "object"\', got \'type: "array"\'.', 'type': 'invalid_request_error', 'param': 'response_format', 'code': None}} ```
I know I can always put the array into a single-key object but it's just so annoying I also have to modify the prompts accordingly to accomodate this.
Otherwise you will to handle the scenarios in code everywhere if you don't know if the root is object or array. If the root has a key that confirms to a known schema then validation becomes easier to write for that scenario,
Similar reasons to why so many APIs wrap with a key like 'data', 'value' or 'error' all responses or in RESTful HTTP endpoints collection say GET /v1/my-object endpoints do no mix with resource URIs GET /v1/my-object/1 the former is always an array the latter is always an object.
The reasons others already posted about extensibility are more correct.
Possible reasons:
- Extensibility without breaking changes
- Forcing an object simplifies parsing of API responses, ideally the key should describe the contents, like additional metadata. It also simplifies validation, if considered separate from parsing
- Forcing the root of the API response to be an object makes sure that there is a single entry point into consuming it. There is no way to place non-descript heterogenous data items next to each other
- Imagine that you want to declare types (often generated from JSON schemas) for your API responses. That means you should refrain from placing different types, or a single too broad type in an array. Arrays should be used in a similar way to stricter languages, and not contain unexpected types. A top-level array invites dumping unspecified data to the client that is expensive and hard to process
- The blurry line between arrays and objects in JS does not cleanly map to other languages, not even very dynamic ones like PHP or Python. I'm aware that JSON and JS object literals are not the same. But even the JSON subset of JS (apart from number types, where it's not a subset AFAIK) already creates interesting edge cases for serialization and deserialization
It's all about the extensibility. If you return an object you can add extra keys, for things like "an error occurred, here are the details", or "this is truncated, here's how to paginate it", or a logs key for extra debug messages, or information about the currently authenticated user.
None of those are possible if the root is an array.
Oops, this is incorrect. {“val{“:2} is valid json.
(modulo iOS quotes lol)
Structured generation seems counter to every other signal we have that chain of thought etc improves performance.
> By switching to the new gpt-4o-2024-08-06, developers save 50% on inputs ($2.50/1M input tokens) and 33% on outputs ($10.00/1M output tokens) compared to gpt-4o-2024-05-13.
GPT 4 Turbo was more like GPT 3.9, and GPT 4o is more like GPT 3.7.
It's taken years to get even preliminary reliable decision boundary examples from LLMs because doing so is expensive.
Leaderboards and benchmarks are very misleading as OpenAI is optimizing for them, like in the past when certain CPU manufacturers would optimize for synthetic benchmarks.
Fwif these aren't chat usecases, for which the newer models may well be better.
Is it a natural function of how models evolve?
Is it engineered as such? Why? Marketing/money/resources/what?
WHO makes these decisions and why?
---
I have been building a thing with Claude 3.5 pro account and its *utter fn garbage* of an experience.
It lies, hallucinates, malevolently changes code that was already told was correct, removes features - explicitly ignore project files. Has no search, no line items, so much screen real-estate is consumed with useless empty space. It ignores states style guides. get CAUGHT forgetting about a premise we were actively working on them condescendingly apologies "oh you're correct - I should have been using XYZ knowledge"
It makes things FN harder to learn.
If I had any claude engineers sitting in the room watching what a POS service it is from a project continuity point...
Its evil. It actively f's up things.
One should have the ability to CHARGE the model token credit when it Fs up so bad.
NO FN SEARCH??? And when asked for line nums in it output - its in txt...
Seriously, I practically want not just a refund, I want claude to pay me for my time correcting its mistakes.
ChatGPT does the same thing. It forgets things committed to memory - refactors successful things back out of files. ETc....
Its been a really eye opening and frustrating experience and my squinty looks are aiming that its specifically intentional:
They dont want people using a $20/month AI plan to actually be able to do any meaningful work and build a product.
It's odd, as many people praise claude's coding capabilities.
The way to get better results is with agentic workflows that breakdown the task into smaller steps that the models can iteratively come to a correct result. One important step I added to mine is a review step (in the reviewChanges.ts file) in my workflow at https://github.com/TrafficGuard/nous/blob/main/src/swe/codeE...
This gets the diff and asks questions like:
- Are there any redundant changes in the diff? - Was any code removed in the changes which should not have been? - Review the style of the code changes in the diff carefully against the original code.
Maybe try using that, or the package that I use which does the actual code edits called Aider https://aider.chat/
There was 100 days in between Claude 3.0 Opus and Claude 3.5 Sonnet being released which gave us similar capability at a 80% price reduction. When I was using Opus I was thinking this is nice, but the cost does add up. Having Sonnet 3.5 so soon after was a nice surprise.
One more round of 80% price cuts after that combined with building out the multi-step agentic workflows should provide some decent capabilities!
It's weird that's only a footnote when it's actually a major shift.
I think this new feature is more relevant for intricate schemas and dynamic schemas when a text prompt cannot do the job.
I have been using it so that my agents aren’t specific to a particular model or api.
Like others have said many other providers already have function calling and json schema for structure outputs.
I can't speak for other approaches, but -- while llama.cpp's implementation is nice in that it always generates valid grammars token-by-token (and doesn't require any backtracking), it is tough in that -- in the case of ambiguous grammars (where we're not always sure where we're at in the grammar until it finishes generating), then it keeps all valid parsing option stacks in memory at the same time. This is good for the no-backtracking case, but it adds a (sometimes significant) cost in terms of being rather "explosive" in the memory usage (especially if one uses a particularly large or poorly-formed grammar). Creating a grammar that is openly hostile and crashes the inference server is not difficult.
People have done a lot of work to try and address some of the more egregious cases, but the memory load can be significant.
One example of memory optimization: https://github.com/ggerganov/llama.cpp/pull/6616
I'm not entirely sure what other options there are for approaches to take, but I'd be curious to learn how other libraries (Outlines, jsonformer) handle syntax validation.
Other posts claim that you can generate jsonschema-conformant output reliably without this, and while I mostly agree, there is an edge case where gpt4o struggles, and that is simple data types. For example, a string, in jsonschema, has a schema of simply {"type": "string"}, and an example value of "hello world." However, gpt4o would produce something like {"value": "hello world"} at a very high probability. I had to include specific few shot examples of what not to do in order to make this simple case reliable. I suspect there are other non-obvious cases.
gpt-4o-2024-08-06 has 16,384 tokens output limit instead of 4,096 tokens.
https://platform.openai.com/docs/models/gpt-4o
We don't need the GPT-4o Long Output anymore.
Also, the question of the default value applies both at the server level and at the SDK level.
Otherwise the max token output limit stated on the models page would be meaningless.
https://platform.openai.com/docs/api-reference/chat/create#c...
I'm working on an app that dynamically generates schema based on user input (a union of arbitrary types pulled from a library). The resulting schema is often in the 800 token range. Curious how long that would take to preprocess
Previously image inputs on GPT-4o-mini were priced the SAME as GPT-4o, so using mini wouldn't actually save you any money on image analysis.
This new gpt-4o-2024-08-06 model is 50% cheaper than both GPT-4o AND GPT-4o-mini for image inputs, as far as I can tell.
UPDATE: I may be wrong about this. The pricing calculator for image inputs on https://openai.com/api/pricing/ doesn't indicate any change in price for the new model.
and apologies for the outdated pricing calculator ... we'll be updating it later today
I'm getting an error.
openai.BadRequestError: Error code: 400 - {'error': {'message': 'Invalid content type. image_url is only supported by certain models.', 'type': 'invalid_request_error', 'param': 'messages.[1].content.[1].type', 'code': None}}
Sadly, that project was a final (relatively successful) attempt at getting traction before the startup was sold and is no longer live.
Edit: Ouch, are the down-votes disbelief? Annoyance? Not sure what the problem is.
>Acknowledgements Structured Outputs takes inspiration from excellent work from the open source community: namely, the outlines, jsonformer, instructor, guidance, and lark libraries.
It is cool to see them acknowledge this, but it's also lame for a company named "OpenAI" to acknowledge getting their ideas from open source, then contributing absolutely NOTHING back to open source with their own implementation.
Maybe those projects were used as-is by OpenAI, so there was nothing new to contribute.
OpenAI wants to kill everything that isn't OpenAI.
Hardware's too expensive, and will be for a while, because all the big players are trying to get in on it.
* cue arguments: "'open weights' or training data'?"; "does the Meta offering count or are they being sneaky and evil?"; etc.
Hint: they won't, it would kill their company. The hype around OpenAI is based on people using it for free, at least at the start.
Heck, even drug dealers know this trick!
Seems like an obvious net gain for the community.
The previous setup didn't allow for custom types, only objects/string/num/bool.
Are the enums put into context or purely used for constrained sampling?
You can take a look at our BFCL results on that site or the github: https://github.com/BoundaryML/baml
We'll be publishing our comparison against OpenAI structured outputs in the next 2 days, and a deeper dive into our results, but we aim to include this kind of constrained generation as a capability in the BAML DSL anyway longterm!
Which has lead me to the prior art: using same Pydantic API to generate structured output from local LLMs
https://github.com/outlines-dev/outlines?tab=readme-ov-file#...
That said, Outline has been making this concept portable for a long time, it's the langchain the community deserves:
I was thinking about the old canard of the sufficiently smart compiler. It made me think about LLM output and how in some way the output of a LLM could be bytecode as much as it could be the English language. You have a tokenized input and the translated output. You have a massive and easily generatable training set. I wonder if, one day, our compilers will be LLMs?
Like there are already cases where hand-rolling assembly can eke out performance gains, but few do that because it’s so arduous. If the LLM could do it reliably it’d be a huge win.
It’s a big if, but not outside the realm of possibility.
Lots of potential avenues to explore, e.g. going from a high-level language to some IR, from some IR to bytecode, or straight from high-level to machine code.
I mean, -O3 is already so much of a black box that I can't understand it. And the tedium of hand optimizing massive chunks of code is why we automate it at all. Boredom is something we don't expect LLMs to suffer, so having one pore over some kind of representation and apply optimizations seems totally reasonable. And if it had some kinds of "emergent behaviors" based on intelligence that allow it to beat the suite of algorithmic optimization we program into compilers, it could actually be a benefit.
In theory we could do the same with mathematical computations, 2+2=4 and the like; but computing the result seems easier.
I could see AI being used (in a deterministic way) to make decisions about what optimizations to apply, to improve error messages, or make languages easier to use/reason about, but not for the frontend/backend/optimizations themselves.
Works with any open LLM, including Llama 3.1
https://github.com/ggerganov/llama.cpp/tree/master/grammars
Supports an EBNF-like syntax, as well as JSON-Schema.
For a long time, I was relying on such guaranteed structured outputs as a "secret sauce" that only works using llama.cpp's GBNF grammars. Now OpenAI literally introduced the same concept but a bit more accessible (since you create a JSON and they convert it to a grammar).
Those of you who have used GBNF, do you think it still has any advantage over what OpenAI just announced?
But yeah I mean, GBNF or other structured output solutions would of course allow you to supply formats other than JSON schema. It sounds conceivable though that OpenAI could expose the grammars directly in the future, though.
And this is more esoteric, but technically in the case of JSON I suppose you could embed a grammar inside a JSON string, which I'm not sure JSON schema can express.
root ::= \" \" item{{{min_count},{max_count}}}
item ::= [A-Z] [^\\r\\n\\x0b\\x0c\\x85\\u2028\\u2029.?!]+ [a-z] (\". \" | \"? \" | \"! \")
This kinda works if you don't mind no abbreviations, but you can't do something like this with JSON grammars afaik.
> 9.11 and 9.9 -- which is bigger
https://community.openai.com/t/why-9-11-is-larger-than-9-9-i...
I wonder if JSON is just a first step, and eventually this will be generalized to any formal grammar...
Yay I can now ensure the json object will look how I want, but lets completely disregard any concern of wether or not the data returned is valuable.
I don't understand why we are already treating these systems as general purpose AI when they are not. (Ok I do understand it, but it is frustrating).
The example given of "look up all my orders in may of last year that were fulfilled but not delivered on time".
First I have found these models incredibly dumb when it comes to handling time. But even beyond that, if you really are going to do this. I really hope you double check the data before presenting the data you get back as true. Worse that is just double checking what it gives back to you is accurate, not checking that it isn't telling you about something.
Every time I try to experiment with supplying data and asking for data back, they fall flat on their face before we even get to the json being formatted properly. That was not the issue that needed to be solved yet when it still fundamentally messes up the data. Often just returning wrong information. Sometimes it will be right though, but that is the problem. It may luck out and be right enough times that you gain confidence in it and stop double checking what it is giving back to you.
I guarantee you someone is going to have a discussion about using this, feeding it data, and then storing the response in a database.
- Structured Outputs is a bit more straightforward. e.g., you don't have to pretend you're writing a function where the second arg could be a two-page report to the user, and then pretend the "function" was called successfully by returning {"success": true}
- Having two interfaces lets us teach the model different default behaviors and styles, depending on which you use
- Another difference is that our current implementation of function calling can return both a text reply plus a function call (e.g., "Let me look up that flight for you"), whereas Structured Outputs will only return the JSON
(I worked on this feature at OpenAI.)
With function/tool calling, you give it a list of tools with payload and the AI will select between them, or refuse, so there are more than 2 possible outcomes.
(I've never used any of these APIs, get that from just reading the docs)
https://platform.openai.com/docs/models/default-usage-polici...
If you're looking to get in early on a free tool that allows you to finetune your model using encrypted data sets you can sign up here: https://waitlist.bagel.net/
The company I work for desperately need the LLM to consistently generate results in a subset of HTML. I was able to craft a small grammar file that does just that, in under 5 minutes, and use it successfully with llama.cpp. Yet, there are still no API offering this basic feature that could really benefit everyone.
Instead we have a thousand garbage medium articles with "tips & tricks" on how to prompt the AI to get better results. It's as if people don't care anymore about consistency and reliability.
at
gpt.franzai.com
we use a 3 times check until now
if chatgpt does not return valid json then trim anything before and after last {}
if this is not valid json feedback him full response and ask for valid json
if this is not valid json put full response into self made json to get al least something we can work with in an "something did not work out" response
Why in the world would I literally pass around a schema file with each request that is (according to the example) 91 lines of code and pay someone for the privilege of maybe doing it correctly?
- The first request with each JSON schema will be slow, as we need to preprocess the JSON schema into a context-free grammar. If you don't want that latency hit (e.g., you're prototyping, or have a use case that uses variable one-off schemas), then you might prefer "strict": false
- You might have a schema that isn't covered by our subset of JSON schema. (To keep performance fast, we don't support some more complex/long-tail features.)
- In JSON mode and Structured Outputs, failures are rarer but more catastrophic. If the model gets too confused, it can get stuck in loops where it just prints technically valid output forever without ever closing the object. In these cases, you can end up waiting a minute for the request to hit the max_token limit, and you also have to pay for all those useless tokens. So if you have a really tricky schema, and you'd rather get frequent failures back quickly instead of infrequent failures back slowly, you might also want "strict": false
But in 99% of cases, you'll want "strict": true.
Is it done at the API Key + schema level? Meaning that for a given API key, the latency penalty for a new schema is only paid one time, regardless of how far apart requests are? Or is cached with less duration, e.g. each session, conversation thread, etc?
I just tried to use structured outputs with the latest release (openai-python 1.40) and it doesn't think Structured Outputs is a thing.
EDIT: turns out my JSON schema is too large (800 lines + recursive) and seems to be breaking OpenAI's "convert to CFG" step. Whoops.
It's been a bunch of decades since I did anything even remotely related to these areas (and I was never good at math).
namely, why did they take so long for something that just seems like a wrapper around function calling?
It's a bug fix. They should never have been charging for malformed responses in the first place!
Interoperable with other external models like the open source versions? What, are you mad?"