Beautiful Soup
crummy.com
crummy.com
It is fast? no.
But it had a fantastic mission: extracting data from malformed HTML.
Might be less common now but back then (~10+ years ago) it was still rampant. Many if not most parsers would barf on any deviation from the standard, leaving you to hand-roll regex solutions and ugly corner cases.
BS covered a LOT of these cases without forcing you to write terrible code. It mostly just worked, with a reasonable API, and stellar, well-written, example-laden docs.
It is less necessary now. One of the most important parts of the HTML5 standards, IMHO, is that it specifies how to parse HTML that doesn't conform to the standards in a standard way. In principle, every bag of bytes now has a standard-compliant way to parse it that every HTML5 parser should agree on. I don't use this to know how many edge cases the standard and/or implementations have, but it's a lot better than it used to be, and means that every HTML5 parser has many of the capabilities that Beautiful Soup used to (nearly-)uniquely have for parsing messy HTML.
I suspect Beautiful Soup was a non-trivial aspect of how the decision to implement such a spec was decided upon. It proved the idea to be a very valuable one at a time when most languages lacked such a library. Basically, BS won so hard that while it wasn't necessarily directly adopted as a standard, the essence of it certainly was.
Sometimes, finding and using the right library can completely turn around a f'd project.
And BS has been that library for me on at least 2 such projects.
That's because it's often deeply buried under more fashionable abstractions.
from multiprocessing import Pool
def parse(html):
result = []
soup = BeautifulSoup(html, 'html.parser')
for p in soup.select('div > p'):
result.append(p.text)
return result
with Pool(processes=16) as pool:
for texts in pool.imap_unordered(parse, my_html_texts):
for text in texts:
print(text) ==== Total trials: 100000 =====
bs4 lxml total time: 110.9
bs4 html.parser total time: 87.6
bs4 lxml-xml total time: 0.5
bs4 xml total time: 0.5
bs4 html5lib total time: 103.6
pq total time: 8.7
lxml (cssselect) total time: 8.8
lxml (xpath) total time: 5.6
regex total time: 13.8 (doesn't find all p)I copied over only the first two paragraphs of each of his reviews with a link back to the original.
The HTML is a total mess having obviously moved around the web a couple times. So it required a bunch of cleanup. But that wasn't even the hard part. The hard part was getting the correct TMDB ID for the movies because his reviews also have either no useful metadata or metadata that's wrong, like incorrect movie years, misspelled actor names, etc.
I never was able to get API access to letterboxd, but they have a CSV import feature which worked-out well enough.
def clean_text(text):
text = re.sub(r"[\x7f-\x9f]", "", text) # remove control chars
text = re.sub(r"[\xa0\r\t]+", " ", text) # replace with spaces
text = re.sub(r"\n+", "\n", text) # squash runs of newlines
text = re.sub(r"\s+", " ", text) # squash runs of spaces
# Remove newlines unless they appear to be at the end of a sentence
# or if the sentence is shorter than 80 characters.
text = re.sub(r"([^.?!\"\)])\n", r"\1 ", text)
text = re.sub(r"\n([^\n]{,80})\n", r"\1 ", text)
return text.strip()Just in case someone wants a comment overview of what this superbly named library is: web scraping (html parsing) in python
0: https://www.crummy.com/software/BeautifulSoup/bs4/doc/index....
BeautifulSoup(markup, "html.parser")
BeautifulSoup(markup, "lxml")
BeautifulSoup(markup, "lxml-xml")
BeautifulSoup(markup, "xml")
BeautifulSoup(markup, "html5lib")
Looks like lxml w/ xpath is still the fastest with Python 3.10.4 from "Pyquery, lxml, BeautifulSoup comparison" https://gist.github.com/MercuryRising/4061368 ; which is fine for parsing (X)HTML(5) that validates<(EDIT: Is xml/html5 a good format for data serialization? defusedxml ... Simdjson, Apache arrow.js)
==== Total trials: 100000 =====
bs4 lxml total time: 110.9
bs4 html.parser total time: 87.6
bs4 lxml-xml total time: 0.5
bs4 xml total time: 0.5
bs4 html5lib total time: 103.6
pq total time: 8.7
lxml (cssselect) total time: 8.8
lxml (xpath) total time: 5.6
regex total time: 13.8 (doesn't find all p)
bs4 is damn fast with the lxml-xml or xml parsersIt may have been because I learned on the python xml.etree library in base(I moved to lxml because it has the same api but is faster and knows about parent nodes) and had a hard time with the soup api.
But I think it was the way it overloaded the selectors. I did not like the way you could magically find elements. I may have to revisit it and try and figure out why and if I still do not like it.
IIRC, Beautiful Soup doesn't handle javascript, so at least for JS you're forced to use something else.
I'm also looking forward to seeing how people scrape the web once Web Assembly becomes prevalent.
Other than that Playwright is incredible, by far the best browser automation api.
I used it for my https://shot-scraper.datasette.io/ tool, and wrote a bit about CLI-driven scraping using that tool here: https://simonwillison.net/2022/Mar/14/scraping-web-pages-sho...
For these cases it can be useful to do the reverse, and use the BeautifulSoup HTML parser as an alternative parser backend for the lxml package: https://lxml.de/elementsoup.html
Unless you mean the person is a professional restaurant reviewer and you like their opinion.
Oh well, good luck either way.
Wait, which one is the real cyberlurker? ;)
For example, the new Google Play Store website stores the data in AF_initDataCallback calls and can be extracted with re.findall(r"<script nonce=\"\S+\">AF_initDataCallback\((.*?)\);", html_string).
Getting this working in a headless browser driven by Selenium would probably be easier for maintainability.
For example Target is clientside, but has all the data in a `window.FOOBAR = json` variable you can fetch and parse with some substring magic. Much easier than spinning up chromedriver and some package.
c.OnHTML("a[href]", func(e *colly.HTMLElement) {
e.Request.Visit(e.Attr("href"))
})
Sometimes I wish Go idioms included an iterator abstraction, it's easier to understand and less hideous than that functional callback style.Yes: https://github.com/anaskhan96/soup
It works well.
Particularly: On-Topic: Anything that good hackers would find interesting. That includes more than hacking and startups. If you had to reduce it to a sentence, the answer might be: anything that gratifies one's intellectual curiosity.
I'm guessing most of the difficult moderation would be on the newsy more-controversial posts anyways.
If you don't want to participate in this post, it's OK to skip it. I skip dozens of posts a day - the best part is it's more efficient than going to them and putting the effort to whine!
As for not inviting meaningful discussion: there's some good discussion on this post - the very article you claim isn't capable of generating such.
Also, in response to your "low effort posts do not invite meaningful discussion" from a different comment, I don't see how this is any lower effort than every other link only post (i.e. the vast majority)? And there's over 40 comments on this thread now talking about other scrapers, projects you can do with scrapers, better docs, tangential use cases and how to handle them, etc. Seems like a lot of people have a variety of things to say about this, I don't see how that's not "meaningful discussion".
EDIT: I also disagree with requiring a couple of sentences from the submitter. If they have something to say they can say it, otherwise it's fine if they don't try to influence the discussion - it's more interesting to see where the random commenters take something, then trying to chart a course.
One of my favorite things about HN is that it has a lot of both serious oldtimers and high school students. It's actually a place where one can, at least sporadically, get the technical mentorship that many of us longed for, but missed, early in our careers.
Wikipedia submissions are a special case and somewhat different: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que....