630 karma · joined March 21, 2015
Email: me@joshdickson.com Website: opennutrition.app Twitter/X: x.com/joshdickson40
I can assure you that you are not overthinking it in terms of figuring that information out. The search experience tries to make it as clear and helpful as possible. If you encounter any situations where it could be more clear, I would love to see them. My contact info is in my bio, or there is a feedback prompt on the site/in-app. Thanks again for checking the project out and your feedback.
> When something doesn't have a reference listed, and just says "sourced from a publicly available first-party datasource", what does that mean?
It depends, and the degree to which it depends is why the citation is ambiguous (although it is true, if imprecise). My goal is to individually cite the individual nutrients but it was simply too costly and time-consuming at the stage of the project at which I did this work.
> what is the process like there for interpreting those values?
Because the degree to which something in the database might be related to those values is so varied, it depends. The reasoning agent had access to those database entires, which is helpful because they tend to contain micronutrient data. It also had access to web data, as well as its own world knowledge, and considers sources in that order. Ultimately it was left up to the agent to decide what the most reasonable fit for each food was, thinking through what an average user likely meant by that entry (e.g. a typical user probably assumes a 'Tomato' is raw), and then to choose the best sources from there. For the chicken salad, it used approximate micronutrient values from the listed references to inform its answer, but adapted the end values for how the dish is described in the description.
> if you had the choice between verified data and fuzzy LLM data, you should go for the human verified data (for now)
Human verification isn't free, and that means it is not available to a lot of people who can't or don't want to pay for something. But if that's something that someone values, I would certainly not diss the human effort!
https://wiki.openfoodfacts.org/ODBL_License
You may disagree with each of those projects as well, but, I am following long-standing licensing in this space. I also have used some OFF data for product naming, and as a result, their terms state I have to maintain their license.
Creating these databases involves a tremendous amount of time and effort, and it would not make sense for me to make this data available to commercial entities to use without attribution. The alternative is not a MIT-licensed dataset, it is no dataset.
Red Beans
- https://www.opennutrition.app/search/red-beans-canned-and-dr...
- https://www.opennutrition.app/search/red-beans-dry-vIh9Ofhcl...
Rice
- https://www.opennutrition.app/search/enriched-white-rice-tlA...
Millions of people use food logging apps to drive behavioral change and help adhere to healthy lifestyles. I believe there is immense societal good in continuing to offer improved tools to accomplish this, especially for free, and that's why I created the project and chose to open source the data.
https://www.opennutrition.app/about#current-state-of-nutriti...
> If I can join the endless queue of feature requests, the ability to scale the portion size and update the nutrition facts would be great
This is all supported in-app if you're in a country with the ability to download it and have iOS (for now). The web product is more of a demo and isn't intended to be used on a day-to-day basis to track your food consumption, but this is a totally reasonable request.
> Also IDK where AI is wrt automated scraping but I've had some success feeding recipes into AI and getting the nutrition facts out. The ability to plop a URL in and get a scraped recipe with a name and nutrition facts would be immense.
I am not doing this for a few reasons, but, you can just screenshot the image of the recipe and use the app to upload that as a meal or recipe and it should parse out the ingredients and portions for you.
Only using the OFF database would be untenable to me as an end user. I think most people do not want to know or care about where the data is coming from, they just want it to be accurate and easy to use. I've listed the usability reasons here for why I can't offer that how I want with only OFF (and that's no dig to OFF, it is a fantastic project, and a primary motivator for this project and its license structure).
Have you asked one of the LLMs used to tell you about the choline content of a food, even ungrounded? They are surprisingly good at reasoning about what kinds of foods tend to contain large amounts of choline because their training datasets will include all kinds of similar data points, even if the single food you're looking for doesn't have it listed explicitly.
https://www.mext.go.jp/en/policy/science_technology/policy/t...
It looks like for unsweetened oat milk:
https://www.opennutrition.app/search/unsweetened-oat-milk-mt...
...it is leaning into a citation from the Australian Nutrient Database (e.g. Oat beverage, fluid, unfortified. Australian Nutrient Database. Public Food Key F006132. ), which is what I instructed it to do if it thought there was an exact match from a governmental database.
It's possible this is a poor general source for oat milk or that's not the beverage intended for the entry to stand for. I'll check it out, thank you for the report.
Background removal lambda if you want to check that out: https://github.com/joshdickson/rembg-lambda
OpenNutrition: https://www.opennutrition.app/search/honey-bunches-of-oats-h...
Via Manufacturer: https://www.honeybunchesofoats.com/product/honey-bunches-of-...
If you wouldn't mind DM'ing me the barcode you're looking at that would be helpful to understand what the nature of the discrepancy is.
Poor phrasing on my end -- yes, absolutely the end data as well as the reasoning, as the reasoning tends to include the final answer.
Maybe I should! Appreciate the feedback.
Right now you'll see that aggregated on some items like this where the reported data is an ensemble of all of the linked resources: https://www.opennutrition.app/search/eggs-eeG7JQCQipwf
Frankly, I just couldn't justify the additional time and monetary expense in doing that if I released this initial version and nobody cared or found it useful. This dataset was also compiled before tools like Claude Citations came out which could make it easier. That is the nature of AI-driven data; I think this is useful now, it is also the worst it will ever be.
1. Generic, non-branded foods
2. Simple prepared foods that ease food entry
3. Restaurant foods
4. Micronutrients beyond those reported by the brand.
OFF is a fantastic project but OpenNutrition is really trying to fit a different niche. OFF does what it does very well; I would never be able to use it to track my food intake.
Appreciate the feedback!
Not really. I do explain in the methodology post how good o1-pro is at the task, but there was a lot of manual effort involved in coming to that conclusion with my own effort to review the LLM's reasoning, and even still, o1-pro is not perfect.