HNHacker News
TopNewBestAskShowJobs

deusu

258 karma · joined February 28, 2015

submissionscomments
deusu··on Ask HN: Do you use an old or 'unfashionable' programming language?
Pascal - still use it today.

I started with Turbo Pascal 3.0, went through several versions of Borland Pascal (DOS and Windows), several versions of Delphi, and am now using FreePascal. I do most of my work on Linux nowadays. Pascal still gets everything done and after 20+ years I am really fluent in it.

Did a little bit of Javascript with Node.js, also looked at Golang. While I did like some aspects of them, there was always the downside of having to learn a whole new ecosystem. Didn't like the dynamic typing of Javascript. Golang looked better for what I do. But in the end I just wrote a small Pascal library to emulate Golang's channels which I really liked, and that was the end of Golang for me.

My main project in Pascal is a search-engine: https://deusu.org

Sourcecode for that is on GitHub: https://github.com/MichaelSchoebel/DeuSu

So yes, you can (still) do actual stuff in Pascal. Even pretty cutting-edge stuff.

deusu··on Ask HN: Blacklist Forbes?
I don't know what technology these blogs are using, but on the occasions where I looked at the HTML-source, the content was there. It just didn't show with Javascript disabled.

Yes, you can use Javascript to make things prettier. I have no problem with that. But not showing content that is actually there, that's a no-go for me.

deusu··on Ask HN: Blacklist Forbes?
I agree.

But I would go even farther. Block every site where you won't be able to read the content without having Javascript enabled. I have experienced several blogs where you won't see anything without Javascript.

deusu··on Show HN: Advertising-Free Search-Engine
Yes. That's what I intend to do.
deusu··on Show HN: Advertising-Free Search-Engine
The completed index takes up 337gb.

The incoming data to the crawler... not so sure... something on the order of 10-20tb. I didn't really measure this. And I don't keep all the data. There is no "cache" function.

On a 1gbit/s connection it takes about a week to crawl and generate the index.

deusu··on Show HN: Advertising-Free Search-Engine
I usually start a crawl with a list of the top one million websites according to Alexa.com. Unfortunately that list doesn't seem to be available for download anymore, so I have to make due with the latest one I have which is from April 2014. After that URLs get crawled basically in a first-come first-serve order. I simply crawl them in the order I find them.

Interpreting ambiguous words for ranking is beyond what I can do at the moment.

deusu··on Show HN: Advertising-Free Search-Engine
Making the software scale to N servers is one of the reasons behind the planned rewrite.

I'm already considering making "special" indexes for blogs, news etc. like you mentioned. In fact that same suggestion came up in a conversation I had with someone today.

The crawler/indexer is pretty fast as it is. I could crawl about 2 billion pages/month for a cost of about 800€ / 900US$ per month. That's on a 3.4GHz quad-core machine with a 1gbit/s connection and 32gb RAM. I tested that once. Works fine.

I can only imagine that your parser does a lot more than mine does. Crawling is CPU-bound for me. So the more processing the parser has to do, the slower the crawling would be.

deusu··on Show HN: Advertising-Free Search-Engine
Many of the design choices that I made many, many years ago aren't valid anymore. So it won't just be a port to a different language, but a complete redesign and rewrite. Which means that regarding the amount of work, it doesn't really matter which language I use.

JavaScript seems to have the biggest community at the moment. Plus I kinda like Node.js. That's why I'm going that way.

deusu··on Show HN: Advertising-Free Search-Engine
I do the crawling myself. There are currently about 320 million pages in the search-index.
deusu··on Show HN: Advertising-Free Search-Engine
No pagerank. But I use backlink-counts and Alexa.com data. Also position of keywords in url, title and snippet. Length of url and number of elements in the url are also a part.
deusu··on Show HN: Advertising-Free Search-Engine
And why should it? You don't need DeuSu to find itself. You are already on the site. :)
deusu··on Show HN: Advertising-Free Search-Engine
I started writing the software a LONG time ago. As far as I know Solr/Elastic Search didn't exist back then.

Query logs are preserved, but they do NOT contain the IP-adress of the user. So all I know is that someone of the billions of Internet users searched for something embarrassing. :)

deusu··on Show HN: Advertising-Free Search-Engine
The main crawl is several months old. There is a separate news-crawler which checks a few dozen news sites every 10 minutes or so.
deusu··on Show HN: Advertising-Free Search-Engine
That is probably your browser keeping track of what you entered into input-fields that are named "q" for "query".

And no, this data does NOT get transmitted. So DeuSu does not know about what you searched-for on Google.

deusu··on Show HN: Advertising-Free Search-Engine
The rewrite in JavaScript will be mixture of porting and redesign. Some parts of the Pascal-sources are relatively new and can almost be translated as is. But the older the Pascal-source is, the more a redesign is a better idea.

The old parts of the sources are in VERY bad shape. This project started more than 15 years ago on a much, MUCH smaller scale (about 2 million pages compared to 320 million now). Some design-choices simply aren't valid anymore. Also the use of JS has the advantage that it will make the project attractive to a lot more programmers than staying with Pascal could achieve.

The current implementation keeps the search-index as one big thing. I will split that up into many smaller pieces. That way it can run multi-threaded or even on multiple servers. Currently queries are running single-threaded.

Another shortcoming is that building an index involves several manual steps. I want to automate that.

deusu··on Show HN: Advertising-Free Search-Engine
I know of YaCy. It's written in Java which I'm almost completely unfamiliar with.

But I know and follow the YaCy-community. Like myself the lead person of YaCy is located in Germany. We even share the same first name. :)

deusu··on Show HN: Advertising-Free Search-Engine
Given that I'm in the process of rewriting it in JavaScript I'm not sure about that... On the other hand it's probably worth a try. Thanks for the suggestion.
deusu··on Show HN: Advertising-Free Search-Engine
It's open-source: https://github.com/MichaelSchoebel/OpenAcoon

This is currently written in Pascal. I'm in the process of rewriting it JavaScript/Node.js. First part for the rewrite is the ranking. That should be done in about a week or two.

deusu··on Show HN: Advertising-Free Search-Engine
Please keep in mind that this is running on one server. Google has what? Several hundred thousand servers?

Yes, it needs to get better. But it's a good start, I think.

← PreviousPage 2 of 2