Turn recipe websites into plain text
plainoldrecipe.com
plainoldrecipe.com
If you want to expand that, I recommend parsing JSON+LD and Microformats. Given your parsers folder [2], it looks like you've tried it, but only for specific websites. I would make that generic and check whether the metadata is available on any website. I wrote a blog post on this if you're interested [3].
source: I've built a very similar tool for my cooking app: https://www.mysaffronapp.com/
[1] https://github.com/hhursev/recipe-scrapers
[2] https://github.com/poundifdef/plainoldrecipe/blob/master/par...
You've got a new paying customer :)
I'd been looking for something like your app for a long time.
Ough, your PayPal flow is not working :( fix that and you'll have a paying customer haha
(will autodelete in 30 mins) there's nothing very private but...
The scaling and editing recipe functionality is top notch. Ill probably use this tool now.
So that just requires some more research and testing. Perhaps someone enterprising will read this and make a pull request...
Saffron looks great, I had encountered it before building this for myself. Your blog post is quite illuminating - perhaps the first practical application of LCA that I've seen outside of an interview setting :)
Now I've got tampermonkey watching over my shoulder and backing up everything I look at to a couchdb instance. (Still gotta write some UI and an agent to pull down images, but I've got other irons in the fire at the moment.)
It's a shame that Saffron provides neither on the published recipe pages. If I share a recipe with someone, they might want to import it into a different app.
Anyway, all of you have a lot of neat suggestions! Please do take a look at the “contributing” section of the repo and let me know if you’d like to pitch in!
That's why they explicitly allow some amount of reposting.
The last one is a bit like when a film releases on the same weekend as some blockbuster, it is more likely to go under the radar. The middle condition is mostly luck, unless someone is prepared to manipulate the early upvotes somehow. But the first one is quite easy to time correctly - US-heavy sites tend to have a few time slots where a disproportionately large number of people check it. My guess would be morning, lunch and afternoon slots, and especially slots where different time zones overlap. E.g. an afternoon slot for Europe which overlaps with the Eastern seaboard doing a morning check of HN might work quite well, etc. This can help a fresh post break through the decaying, but highly upvoted older posts that are keeping it from more visibility.
Some days I wish I had bookmarked that link
[0]: https://chanind.github.io/2019/05/07/best-time-to-submit-to-...
Very nice to see your repo and happy to see it being so minimalistic. I use paprikapp[1] to manage my recipes and you can import recipes into paprika using yaml[2].
I used to be a customer of hellofresh (Germany) and had to manually copy paste into Paprika. I recently made a tool to just input a HelloFresh URL and output a yaml file which can be used to import the recipe into paprika into images. [3]. Maybe I can open a PR and add support for HelloFresh recipes.
Good work btw.
(I'm in the process of abandoning LaTeX in favor of a custom Markdown → Pandoc → HTML flow with basically the same layout, though.)
I use an 8.5x11 page with metadata on one side (name, description, time to make, tags, dates made, etc.) and the recipe on the other. On the recipe side I have ingredients at the top of the page in two columns, and then the directions at the bottom in two columns.
We put this page in a plastic sleeve and into a 3-ring binder. When we want to make the recipe, we take it out and tape it to the cupboard with masking tape. That way anyone in the kitchen can easily see it and prep their part.
I have all of my recipes plaintext in an org file with a very simple format...
* Recipe title
** Ingredients
- 1 tsp x
- 2 tbsp y
** Directions
1. Mix together x and y
2. Bake at 350 for 15 minutes
I set up a simple python script using the same package as the OP's website for scraping recipe sites to org format, then I export subsections to latex with a custom class and print to index cards.drives me nuts. I have to rewrite any recipe i want to keep making to actually have all the required instructions.
It was a scientific way of cooking: gather and measure all your ingredients before you start. In a commercial kitchen you still do that: you go to work hours before the doors open to put everything in place (mise en place). That's how you get reliable results.
Even if you don't gather stuff, a good home cook will still scan the list to ensure that they have what they need. Still, It wouldn't be a bad idea to also replicate the measurements in the recipe itself. Perhaps they have the coder's instinct to not duplicate information.
It took me a moment to grok the format, but once I did I found it to be really helpful.
https://coolinfographics.com/blog/2010/4/26/cooking-for-engi...
When I write recipes I group the ingredients by step. So for example I might have a "Sauce Ingredients" block, and the instruction will be "add sauce ingredients to the pan". It makes mise less annoying, I don't have to keep scanning the instructions to figure out how to organize my ingredients.
Keeping ingredients listed separately from instructions dovetails into this practice.
This is assumed in all the time estimates, as well. Time spent getting the ingredients into the bowls isn't counted, so if something calls for 3 chopped onions you have to add the peeling & chopping time to get the real prep time.
One major flaw: it seems like the calories and macros aren’t captured. For bodybuilding and powerlifting types, and other athletics, these are the most important part of a meal.
What would make this a “killer app”, in my view, is if I could request a recipe in JSON instead of just formatted plain text as you do. Then I could use the recipe (and recipe search) in my own home-brewed meal planning program.
As a person with better visual memory for certain kinds of data, having the original page content may have as much meaning as the recipe, for entirely different reasons. Food can be very personal, and recipe books doubly so. A recipe archive can be as personal as we like, or all of that can abstracted away when we don’t need it.
I remember first seeing it on Slashdot way back in the day, before they had user moderation or user meta-moderation.
https://blogs.loc.gov/thesignal/2012/04/a-library-of-congres...
If you save the source document, the code needed to parse your recipe archive is likely to be pretty short. Then you have a corpus to do A/B testing of your recipe parsing code against.
Side note: I feel that moderation and later meta-moderation system on Slashdot was the most transparent, fun, positive moderation system I’ve ever been part of. I wish HN had more than just up and downvotes, for instance. User meta-moderation would help reduce flamewars immensely IMO.
Ctrl+F "print".
Probably 90+% of websites use some tool to format the actual recipe, so the actual recipe pretty consistently includes a link for printing it. Search for the word "print", and it takes you to the part of the page that has the recipe.
(You don't actually have to click the link. It's just a landmark for navigation within the page.)
As an aside: I live in a jurisdiction where recipes cannot be copyrighted, so I have collected all recipes I remotely liked on a web page with only text. All recipes except some of the very last ones "untested, potentially disgusting" headline are in Swedish though: https://koketteriet.se/skrivet/Recept/recept.html
Didn't know that recipes couldn't be copyrighted here in Sweden.
I think I managed to give credit where it is due (in the intro, at least), but about 60% of the recipes are things we veganised ourselves from whatever old family recipes we have.
http://organicpassioncatering.com/
Anyway, I find, as I rapidly approach the big four-oh, I don’t need recipes much anyways, so I have a sleeve-book with some recipes, some pages are just dish names, and others are flavours or ingredients that go together.
I recently went down the rabbit hole of recipe websites while I built a little side project Shopify app for creating recipes on ecommerce stores[1]. Having a plaintext version has been on my backlog to-do list forever, but it seems the vast majority of store owners aren't keen on it. It's been a massive learning experience for me; and I never realized how... bad the recipe website experience was for so many people.
Question: how do you think we, as in us as the people building the current web, can improve the standard recipe display on the web? Obviously removing the 1300 word novel before recipes is a big plus, but what else do you think would improve your day-to-day recipe browsing?
e.g. https://search.google.com/structured-data/testing-tool/u/0/#...
out of interest, did you try using structured data to scale the tool?
Having said that a lot of users of recipe mark-up have problems implementing it properly.
The exception is usually butter, which is measured in grams (and conveniently indicated on the packaging).
I live in Europe and have always weighed exactly 100g on the scale, whether I was asked for 100 grams of butter, sugar or flour.
Are kitchen scales accurate enough for such tiny measurements? 1 teaspoon is 5ml, so if your scale only has 1g of precision you're looking at up to a 10℅ variance (assuming you're measuring water), and that's assuming its accurate at that range. It gets worse when you need to deal with fractional teaspoons.
Yes, and there are lots of recipes that use mass rather than volumetric measurements. For solids, in fact, many sources prefer this because consistency of packing density and other user factors that effect consistency with volumetric measure are far worse than with mass measured with a kitchen scale.
Unless you’re in the USA, in which case a cup is 236.5882365mL exactly.
Apparently it works perfectly fine for household cooking to use units of volume for flour and sugar. Close enough is good enough!
That’s what we commonly use which is easier to use than just weighing everything, but still just as accurate.
It also solves to many issues with ingredients that will just never have equal density, flour variances, packing brown sugar by hand, etc are things of the past.
For a recipe where ingredients go in at different times and require different prep, and measurements are approximate anyway, weighing is more trouble.
I wonder if the ideal country where no-one uses units of volume other than litres and millilitres actually exists.
The really bizarre units of volume are certainly the ones I've seen in Europe - where you get measures marked in "grams of flour" and "grams of sugar". I suppose no-one has ever asked for a hundred grams of sugar of flour, but I'm tempted to every time I see one of them.
Even European recipes still measure small amounts for flavors, leaveners, etc. in volumes. A scale that's precise to sub-gram levels is more expensive. So it's not as if you're not familiar with the concept.
Now that electronic ones are cheap it's clearly the better choice, especially for flour, but that's only in the last decade or two. People are switching over to scales, but the volume measurements are perfectly adequate and they're what our existing cookbooks call for.
I imagine you all have also greatly increased the amount of baking you're doing while staying safe at home.
https://chrome.google.com/webstore/detail/recipe-filter/ahlc...
... but then I noticed Firefox reader view does this. Take a look: https://www.bbcgoodfood.com/recipes/brilliant-banana-loaf
javascript:location.href='https://plainoldrecipe.com/recipe?url='+encodeURIComponent(l...
I could not care less what pivotal moment in your childhood led you to like chocolate cake. I just want to make one.
Everyone I know balks at this, yet it exists. Can some explain why? Where is this tradition coming from? Does this drive page views? I know nothing about the cooking subculture. Is it some sort of normies vs "really into this" thing (and please excuse the phrasing, I could not come up with something better). Because that would mean the page is just not targeted at me. But then again that argument will make sense if the monetization was not through advertisement and view volume.
I didn't realize how curious I was about this.
Edit: just had a thought. Can it be something to do with SEO? Does google like bigger articles with lots of personal text? If this is the reason it would blow my mind. Without knowing it google algo would have made the internet a shittier place, this is very amusing to me.
From the US Copyright Office:
"Mere listings of ingredients as in recipes, formulas, compounds, or prescriptions are not subject to copyright protection. However, when a recipe or formula is accompanied by substantial literary expression in the form of an explanation or directions, or when there is a combination of recipes, as in a cookbook, there may be a basis for copyright protection.”
i find all my recipes on youtube now.
I have the opposite problem trying to get more text on our recipe pages - I am sure some designers would design a recipe e page with no text if we let them.
Of course you don't have structure of the recipe and it does not simply work with a url.
And makes it really easy to reference a week or two later if we want to remake it - just pull it out of the binder.
I do often wonder if many recipe sites ever really think about the user experience. It's almost always going to be on mobile and there's either got to be a simplicity to the presentation (which is what this does brilliantly) OR a really simple way to flick between ingredients and method.
The one that stands out for me is Jamie Oliver (example: https://www.jamieoliver.com/recipes/vegetables-recipes/veggi...) - on mobile there's this super simple tab that flips you between ingredients / method. It's simple, but exactly what you want, especially when you're covered in flour / oil / etc... :-)
As well as importing from websites, you can enter you own recipes. And my wife and I can sync between our phones and laptops so we can follow and edit the same recipes. Very handy.
Next stop: the holy grail of recipe website cleanup, removing the mindless drivel about the author's.. whatever / whoever, their inane anectodes, and generally any trace of their annoying personalities (here's an offending example: https://www.ibreatheimhungry.com/easy-roasted-pork-shoulder-...)
I wonder if any of these websites will find a way to prevent these recipe scrapers from working so that people read the damn ads in their blog texts. Instapaper‘s article extraction gets rejected from time to time and Safari‘s reader mode too, but it’s mostly on major news sites.
It seems you're using the same Python scraper I am: https://github.com/hhursev/recipe-scrapers
Is this aiming to be a domain-specific equivalent of Outline.com and reader-mode?
Otherwise there's a PR open: https://github.com/poundifdef/plainoldrecipe/pull/1