HNHacker News
TopNewBestAskShowJobs

aumerle

569 karma · joined March 26, 2016

submissionscomments
aumerle··on Show HN: Browsh – A modern, text-based browser
Rendering images using unicode block symbols does not make them text, it just makes them pixelated. And you can always down sample images to reduce the bandwidth cost and still get a much better rendering than you achieve currently. For example you can convert the images to indexed 256 color compressed PNG and transmit that, which would reduce bandwidth by a factor of ~4x while still giving you much better rendering.

But anyway, it was just a suggestion, it's your project, you should feel free to do what you think is best for it.

aumerle··on Show HN: Browsh – A modern, text-based browser
You might want to consider adding support for displaying proper terminal graphics, instead of using pixelated ones (which I assume are done using unicode block drawing symbols?).

See https://sw.kovidgoyal.net/kitty/graphics-protocol.html

aumerle··on GIF for CLI
See https://sw.kovidgoyal.net/kitty/graphics-protocol.html for a much more comprehensive terminal graphics protocol
aumerle··on Eshell as a main shell
Links are already supported in many terminals, see https://gist.github.com/egmontkob/eb114294efbcd5adb1944c9f3c... and are actually used by the ls program, ls --hyperlinks

The overhead of an HTML DOM vs the typical cell structure used ina terminal is the difference between Jupiter and Mercury.

aumerle··on Eshell as a main shell
The only difference between "rich text" and what you can do in a terminal is changing font families/sizes. And that is not even close to useful enough to justify runnig a terminal on top of a browser.

As for familiar HTML, it is trivial to write a library that accepts "familiar" HTML and converts it to SGR cdes for formatting. I could do it in an afternoon.

And for "awful slow", do the following experiment. Open a large text file in less in your terminal. Then scroll it continuously and monitor CPU usage (of the terminal and X together). Now compare with a real terminal. Think of all the battery life and all the energy you are sacrificing just for the ability to use multiple font sizes and families.

aumerle··on Eshell as a main shell
All these features are already possible in modern terminals. 1) True color support 2) Bold and italic fonts 3) Emojis, Ligatures, unicode in general 4) copy/paste across SSH 5) the ability to split windows into tmux like panes and send text/control the different panes 6) The ability to display images in true color with alpha blending

These are only a small subset of features modern terminals have. There is absolutely no need for terminal 2 or awful slow terminal implementation based on rendering via a DOM.

Many ternminals support all or most of these features, one such: https://github.com/kovidgoyal/kitty

aumerle··on A look at terminal emulators, part 2
That's because alacritty claim's are pure bullshit: https://github.com/jwilm/alacritty/issues/289
aumerle··on A look at terminal emulators, part 1
Simply remap Ctrl+Enter to send some other key combination and bind that in zsh. You can do the remapping either in X or using the terminals native remapping features many terminals have them, for instance kitty.
aumerle··on Introduction to web scraping with Python
http://html5-parser.readthedocs.io/en/latest/#html5_parser.p...
aumerle··on Introduction to web scraping with Python
https://github.com/kovidgoyal/html5-parser
aumerle··on Modern terminal-based text editor
https://github.com/kovidgoyal/kitty

and as for moving the state of terminals forward:

https://github.com/kovidgoyal/kitty/blob/master/protocol-ext...

aumerle··on Show HN: Pdf-Bot, an API/CLI for Generating PDFs Using Headless Chrome
calibre has been able to convert arbitrary HTML files to PDF with Table of Contents with page numbers, links, embedded fonts, arbitrary headers/footers for years, all rendered using WebKit, without a running X server, for years.

ebook-convert file.html file.pdf --pdf-add-toc

aumerle··on Show HN: Fast C-based HTML 5 parsing for Python
Simply pass the argument treebuilder='soup' to the parse function. But note that you wont get as much of a performance boost, because the besutifulsoup tree has to be built in python, not C.
aumerle··on Show HN: Fast C-based HTML 5 parsing for Python
Yeah, I have to deal with what we have in the real world, not in spec land. Yielding blank documents on a <title/> just isn't going to cut it :)
aumerle··on Show HN: Fast C-based HTML 5 parsing for Python
You do know that by default bs4 uses lxml.html to do its parsing, which is also a big blob of C, right?

For example, running the following:

BeautifulSoup('<p>a') /usr/lib/python3.6/site-packages/bs4/__init__.py:181: UserWarning: No parser was explicitly specified, so I'm using the best available HTML parser for this system ("lxml"). This usually isn't a problem, but if you run this code on another system, or in a different virtual environment, it may use a different parser and behave differently.

aumerle··on Show HN: Fast C-based HTML 5 parsing for Python
-Werror is used only when building on the continuous integration servers, not in the released code. That way, you get the best of both worlds. Your code wont break on new compiler releases, but it also wont have neglected warnings.
aumerle··on Show HN: Fast C-based HTML 5 parsing for Python
This has a hard dependency on libxml2. If your language of choice has a libxml2 based tree, it can be ported to that.

Basically your language needs an equivalent of lxml

aumerle··on Show HN: Fast C-based HTML 5 parsing for Python
I use this pattern a lot. For example, in https://github.com/kovidgoyal/kitty where the UI layer is in python and the performance critical code is in C. I find this pattern gives me an optimum in performance + maintainability.
aumerle··on Show HN: Fast C-based HTML 5 parsing for Python
Yes, since I wrote html5-parser for use in calibre (i.e. to parse the HTML in e-books), sigil-gumbo is a more appropriate upstream for me.

I have no objectsion to merging in changes from gumbo-parser that are not in sigil-gumbo over time, but it will need to be done gradually, as there is only so much time I can devote to html5-parser now that it meets the needs of calibre.

aumerle··on Show HN: Fast C-based HTML 5 parsing for Python
https://html5-parser.readthedocs.io/en/latest/#benchmarks
aumerle··on Show HN: Fast C-based HTML 5 parsing for Python
Indeed, I developed html5-parser to eventually replace html5lib in calibre
aumerle··on Show HN: Fast C-based HTML 5 parsing for Python
@nostrademons: Thank you for gumbo, which is an excellent project.
aumerle··on Show HN: Fast C-based HTML 5 parsing for Python
I encounter self closed <title> tags in malformed XHTML files all the time. In fact I encounter them so often, I used to use a dedicated sanitization pass before passing int he html to html5lib, for that reason alone.
aumerle··on Show HN: Fast C-based HTML 5 parsing for Python
I have not tested it, but it is probably slower. The advantage of this is that is follows the HTML 5 parsing spec, so it gives you the same results as a browser when parsing malformed HTML
aumerle··on Show HN: Fast C-based HTML 5 parsing for Python
:) Try parsing the following snippet using HTML 5 parsing algorithm and see what you get:

<html><head><title /></head><body><p>foo

aumerle··on Show HN: Fast C-based HTML 5 parsing for Python
Yes, the lack of maintenance of gumbo-parser was another reason I decided to include a private copy. You are welcome to make your PR against html5-parser's copy of gumbo-parser, and I will review it.
aumerle··on Show HN: Fast C-based HTML 5 parsing for Python
Because it uses a modified version of gumbo to support parsing not-well formed XHTML as well
aumerle··on Show HN: Fast C-based HTML 5 parsing for Python
As permissive as html5lib, which is as permissive as a modern browser, since both are base don the HTML 5 parsing spec.
aumerle··on Show HN: Kitty – A modern, hackable, featureful, OpenGL based terminal emulator
OSX/freebsd support should not be too far away -- see issue 5 in the kitty github for the status of the OS X port.
aumerle··on Show HN: Kitty – A modern, hackable, featureful, OpenGL based terminal emulator
Presumably, if you are going to do network forwarding, your machine will be network conencted, so just install it. kitty is designed to run from a single directory if needed, so you dont even need admin permissions to do that.
← PreviousPage 6 of 7Next →