Full Time
marginalia.nu
marginalia.nu
Congrats to Viktor and good luck!
Going to go and try your search engine now.
Previous discussion of the search engine a couple of months ago: https://news.ycombinator.com/item?id=35611923 (196 comments)
Many other posts and blog posts over the last couple of years: https://news.ycombinator.com/from?site=marginalia.nu
There are positives with an office but I kinda envy how everyone celebrates the opposite move
I've had a look at the README[0] for the Java sourcecode, but it's highly focused on crawls, database & indexing (understandable for search); would be cool to see a front-end focused write-up.
[0] https://github.com/MarginaliaSearch/MarginaliaSearch/blob/ma...
The search engine is serverside-rendered mustache templates via handlebars[1], via served via spark[2]. It's basically all vanilla Java. I do raw SQL queries instead of ORM, which makes it quite a bit snappier than most Java applications. The sheer size of the database also mandates that basically every query is a primary key lookup. The code is written around that constraint.
Although the search engine is a bit on the slow side since it's routed through cloudflare and I think I'm relatively far away from the closest datacenter so it adds like 100ms to the load times.
Yeah the static stuff being fast is less surprising - it was mainly the search results page that astounded me.
> via served via spark[2]
Had not heard of this Spark (only the other Apache one). Will definitely take a look.
> Although the search engine is a bit on the slow side since it's routed through cloudflare and I think I'm relatively far away from the closest datacenter so it adds like 100ms to the load times.
I've hit the CF loading screen which introduces a big delay, but when I don't see that the loading is really instantaneous.
Same philosophy but just got things right where Spark, being the "first" (in the Java world, using the design inherited by Sinatra[2]) had a few design issues.
Love it. I've seen so many cases where engineers with just basic SQL knowledge (like myself; I'm no JOIN god) can run circles around the queries ORMs generate.
Which is kinda what they are for, that's why they're called "Object-Relational Mappers". Not "Object-Relational Query Generators". Because they suck at the latter.
That anyone with more than 6 months of experience still drag Hibernate up from their chest is just absolutely beyond my comprehension.
That said, the moment you leave the small table simple relations space (by e.g. having a table with a quarter billion rows), then ORM is not a good choice.
Btw, I was half expecting that you quit because of FUTO grants (saw your post on their forums), but I guess it wasn't that. Either ways, rooting for you!
I do think LLMs has the potential to integrate well with search engines and I'd love to see a sort of open source search ecosystem emerge with different projects collaborating to exceed the sum of their parts.
Yeah I was in talks with FUTO and they agreed to help the project out a bit, I won't say more until I have money in hand. It's taking a while but it's not on their end, I just need to sort out some legal stuff first.
Have you considered rewriting it in rust? ;)
In all seriousness, it's great to see something written in a "boring" language like Java, which seems to get a lot of hate in developer circles, hover at the top of HN.
Java really can perform amazingly well, especially if you minimise the use of unneeded libraries and frameworks. Super curious to see how your stack evolves as you get more load.
Best of luck to you on the journey!
Ps there's truly a world of difference between "Spring Boot Developer" and "Software Engineer with Java experience". I suspect a lot of people who hate Java or think it performs badly have only worked with the former group of people.
Java gets a lot of shit, some of it is merited but a lot of it isn't really fair to the language. There's a lot of Java developers that are kinda shit-tier copy-paste developers developing shit-tier copy-paste applications, because the language is so forgiving as to accommodate that, but it's also a competent language that you can do seriously impressive things with.
You can be insanely productive in Java because it's extremely stable and mature. You almost never have to deal with library churn or other upstream changes that urgently needs to be fixed. I can think of exactly two instances in my professional life that's happened, migrating off Java 1.8 and the oh-shit moment of needing to patch log4shell.
My code from late 2000s, with very few modifications to keep up with modern java syntax, writren in bare java talking to postgres in just raw sql runs circles around anything new I've tried or built with for modern web application backend stacks.
IDE support is top class. Thread are just awesome when done right. Static types make code incredibly easy to read and reason about even years later. JVM has been super stable forever. A whole lot of features I need are just baked right into the language, but not obvious at first. Mvn just works. And my reluctance to external libraries actually made me write the logic myself making me much better understand related concerns.
Congratulations and I wish more cool projects picked Java as their language, but I see them use Oracle's ownership as the strawman argument against it. I don't know enough about that.
Personally, I owe a lot to Java.
We've over-abstracted everything with our own abstractions and/or libraries that are as bad as our own stuff and we're doing this in a language and runtime that is pretty much one big premature pessimisation to begin with because we could do significantly more with less. We have no idea what the GC is doing and when, and we care less about that than "using functional programming" or some other equally pointless-in-itself principle that you would be better taking very little from.
Many applications do so many redundant calculations it's almost absurd. Professionally I've seen applications do enormous ORM lookups that fetch thousands of objects, stick them into a hash table by some key, and pick one value and compare against some parameter; and then do this operation in four different places in processing a single request. Gee I wonder why we have 800ms page loads...
Of course sometimes a complex problem is difficult to solve simply, and sometimes being verbose is a good tradeoff to help maintainability.
I tried Marginalia and already get amazing and fast results. This will make the web fun, creative and interesting again.
Just like, I think, fellow countryman proving the world wrong that browsers cannot be created from scratch with Ladybird I think you will succeed also. (At least with search engines the competition gets worse every day.)
Hope to see it continue to grow until the internet goes dark.
A sort of indie internet discovery ecosystem effect is one of the things I've really been hoping to accomplish with Marginalia.
I'm very impressed. Using it I get this old Internet vibe (which someone else also mentioned). Just used it to get some information on a random topic I recently tried to research with Google but failed due to all the SEO crap. It produced several hits of old pages (with the tiny font and the early 2000's graphics and design), but _full_ of information.
Not all the results are good though, it was mostly hit and miss, but the hits were _good_. Will use it from now on.
@marginalia_nu if you're reading this, please know that you're an inspiration and that I crazy appreciate what you're doing.
We need people like you in the world pushing to make interesting things that aren't necessarily profit driven, but instead seek to help add flavor and interest back into the world.
Your search engine is the kind of technology that reminds me of the technology ethos from the 90s and it's so amazing to see you get the chance to actualize it! Don't waste this chance!
No matter what, know that this is the right decision. You have fans, and we're rooting for you!
Thank you for making cool stuff!
Is another support option in progress to replace?
[0] https://www.marginalia.nu/marginalia-search/supporting/patre...
Fixed now.
I like to write browser (puppeteer) tests for user-facing software criteria like "patreon link must work." In the past I've written similar tests for small websites I've created where the purpose is to surface affiliate links for users to click on. My criteria is "from a money standpoint, is this the call to action I want my users to engage in?"
I don't know what type of test this is -- can anyone disambiguate testing terminology for me?
P.S. Browser based testing is brittle but since I often create websites and because I want to really ensure that I'm not 'lying' to myself in tests, using a browser is often the best (albeit slower) choice. These tests usually run in CI and I get notifications if they break.
P.P.S. I wish we had a better mental model for the types of tests than the "testing pyramid." I find the testing pyramid lacking.
I came across this interesting tool for similar tests the other day. It lets you request websites or API’s and then search the return for a string. It’s more for checking uptime, so I dunno if it would be acceptable for this type of test, but it looks like a cool tool.
I have a hunch that every pyramid model is bullshit. It's inherently appealing to present any sequence of things as a pyramid, regardless of whether it makes sense.
What I find quite nice about Marginalia is for discoveries outside the most popular destinations for such topics. For example, looking for a weekend movie but do not want to see all the SEO websites talking about movies. Marginalia surprises you with some unknown websites in the first page :) I use it when I want to be surprised by the results :D
The "temporal freedom" I have in my work (Gad Saad, if you don't know the name). I love being the master of my own day, of my own time. I don't sit in Zoom meetings or have daily standups. I can get up at 5am and work until 11am, and then go hike, play with my dog, get ice cream with my daughters, workout, etc. and then work again from 7pm until midnight or whatever.
Having (almost literally) full control over my daily schedule, week-in, week-out, year after year, is invaluable to me.
One disclosure: a few times a year I do very hard things where I have very little freedom, but they allow me to have lots of freedom the rest of the year.
Not to be a jerk, but I won't be elaborating. And I realize this life isn't for everyone!
I wonder how many great works we'd have built if most weren't trapped elsewhere.
A bit too nervous to pull the trigger just yet
At 50% if you can see an upward trend, ~6 months savings, and have a plan that the time will give you to execute, got for it.
Also if you're going to quit anyway, you might as well ask the company you currently work for if they'll let you go on sabbatical, or part time, or consulting.
That can give you a bit of extra runway/feeling of security.
Dunno what you're doing wrong if you can't land a job with a built-from-scratch internet search engine on your resume.
In general I've had like infrequent but large influx of money from the project, so it's hard to answer. Although I have relatively long runway, no small thanks to nlnet for their generous grant.
On some level it's all a gamble. Either I try to make this work somehow, or I close up shop and keep working as an office drone, because I really can't keep doing both.
My hope is that I'm able to make it work on a wikipedia-like model donation model, maybe supplemented with selling commercial API access (access is free CC-BY-NC-SA). My burn rate is literally my living expenses plus a hundred dollars per month of service costs to I don't have to be spectacularly profitable to sustain flight. ... all that is contingent on making it work quite a lot better than it does now, so I guess I have my work cut out for me.
It's also a weird project, since it's had an almost absurdly positive reaction. For example, many people develop a search engine and get almost lynched on HN for not working exactly like Google or not dealing with some query as expected. Someone found a link to my barely working search engine that didn't properly support multiple-keyword queries and this happens: https://news.ycombinator.com/item?id=28550764
https://www.marginalia.nu/marginalia-search/supporting/patre...
https://www.marginalia.nu/marginalia-search/supporting/buyme...
(the text of the links is correct though)
A book, software, something. I can't quite expense patreon and others may have a similar issue. (Useful "free SaaS" where all there is is the cup of coffee button makes me sad).
My buyer won't even blink if I say that I need a $150 tool: I can just bill it to whatever project it's being used for as long as I get an invoice or a receipt or some kind of documentation. If I say that I found a free tool and I'd like to donate $10 to the author, no one will know how to do that.
It's very hard to read your search results. I've always disliked grid views to represent data. It's very hard to find what you want.
Im not sure. But it looks like you didn't want to copy google and wanted to make something "authentic", same reason why often modern design is unusable.
Every competitor of Google just gave up trying finding a better sexier way. DuckDuckGo, bing etc. Pure copies. A list view, with a good contrasting header is the best way to scan and find the results you want.
If you want to keep it, at least provide a list / grid switcher so users can pick themselves.
Good luck! Happy you get to pursue your passion.
I don't like the basic old school google style list though. It makes very poor use of the screen space. This is primarily a service for desktop users finding desktop content, but I still want something that's accessible to other screen sizes. Really hard to find a good design that works well.
I also really appreciate the desire to use available screen space. It irks me to no end when a site forces a narrow column of info/content and wide empty borders wasting half or more of my screen. Wikipedia recently started doing this and I can't say they're better for it in my opinion.
I do like that feature.
They are not my colors, but the contrast is clear!
Mobile I think the cards are to high. Slightly smaller font, and cutting of after 2-3 sentence a read more link would probably make it easier to sift through your results.
But just my random opinion, good luck!
When a search term yields many results, it's left to the user to the user to search the results for the site that will yield the "best" match for what they're after. It seems like people assume that the better the search engine is, the better it is at predicting what the user is really after by putting it at the top of the listing. But this can be rather difficult when the original search terms are pretty generic and the user is required to scroll and check many results. If there were a way help the user sort the results based on relevant criteria, maybe that would make that search easier. And personally, I like things that give users a little more say in how they get fed information. Allow sort by popularity, frequency of search terms in page, number of pages in site's domain, date of last page edit (no idea if this is possible to get), etc...
Maybe have multiple columns of search results. One column that lists results that match all words in the query, another for only one or two words. Or maybe columns that list results that include the user's query plus likely related topics. Or a set of search refinement tools that can further help the user sort based on any number of criteria, or filter results by specific related terms.
Slightly related, I really like your encyclopedia site. In addition to being incredibly nice to use all on its own, perhaps it (from the 'See Also', 'Further Reading', 'Related articles', etc... sections) could be mined for suggesting additional search terms/info a user could add to their search or filter their results by. For example if I search for Tcl and get a bunch of results, some tools that suggested filtering (or a search instead option?) the results to those that included Tk, expect, and TclX might help me get to what I'm after quicker.
No idea if any of that is practical or would even actually be that useful in practice.
I don't know you personally, but you come across as an earnest lone developer doing something for the passion of it. I think that goes a long way on here, versus someone giving off "portfolio project", "hire me" or "seeking investment" vibes. I've not really found a use case for your engine yet but I am really enjoying seeing your progress.
I have no affiliation but recently came across them from a weekly newsletter (via https://changelog.com/news/48/email).
Marginalia is one of my favorite sites. Wishing you all the best.
https://youtube.com/clip/UgkxgJAzCqKOL5yMg39wmtZi52tw8LAXOEr...
Just today I realized how distracting too many hyperlinks can be. And Wikipedia is full of them! It feels so much easier to read an article without them. Now I just wish Wikipedia had more supporting graphics to help engage readers in a more productive manner.
Beyond the feeling, it's also educational as you learn about your deficients quickly (or, in some cases, too slowly).
I'm wrestling with this now as I'm building my platform and looking to pivot into something that produces revenue.
https://encyclopedia.marginalia.nu/favicon.ico
Otherwise - Great job on the peppy site and breath of fresh air to open the network tab and see 1 html get, and another for the CSS. And the 404 favicon that I guess the browser insists on ;)
I thought this too! So happy someone has tested the assertion.
Good luck! I’ve had Marginalia bookmarked for some time but this story will remind me to try it.
Is there and what is the "peak" amount of optimization feasible in Java for this search engine before one would need to turn to C/Rust/etc to get any more performance out of this on the given hardware?
I recently rewrote a heavy algorithm from Java to Rust, thinking that I'd get faster performance pretty much automatically. It turned out to be significantly slower than my optimized Java algorithm, and I didn't have the experience to tune the Rust version, so I ended up sticking to Java for now.
I'm sure someone who knows how could have tuned the Rust version to get better performance, but native code is not my specialty and the Java version was doing fine.
A warmed up JVM is a lot faster than most people think, especially for a long-running app like a search engine.
It's relying quite heavily on memory mapped I/O and doing some clever things to work around language limitations in how much you can memory map at a time. This permits surprisingly good but not optimal performance.
A bigger drawback is that this type of low level programming in Java is a serious pain in the ass.
Congratulations on cutting loose, always a great feeling.
First I'm hearing of it.
I personally find it hard to put into words, but the old internet and old search engines had this feel to them that you never knew what you were going to get. Each site looked different. Each site had it's own philosophy of content and design. Everybody was winging it. It just felt more personal and interesting. At the risk of hyperbole, now it seems search engines give back mostly SEO blogspam that all looks the same.
Marginalia feels more like the former internet.
And with technical questions, too many results are not correct. It seems that Google search is really going down hill in this regard. I'd like a way to vote down results that are obviously SEO trash, but I'm pretty sure if that were provided, it would be gamed too.
Hoping this turns into something magical. I dearly miss the old web!
Upon closer inspection, I think there's something disconcerting about using this engine in that—I don't know the right words for this but—it filters 100% by all the keywords. So, "rollover" is kind of a strange word, and because you don't have any results with these three words, I see nothing.
I'd prefer to instead see results for "urbit key," given the circumstances. I imagine the algorithms to do this well are complicated though.
Another query that had no results: "multi-band compression maselec"
A query that has surprisingly few results: "qmmf-4." It shows a single forum post from a forum that has probably hundreds of posts matching this query. Why just one?
IOW how queries ought to work. I am perfectly capable of changing the precision of my query to affect recall; I prefer that no "algorithm" ruins that ability.
I did end up getting results with "jit oauth" (quotes not in search), but not great ones.
To be fair, google didn't give me great results for either of those queries either
edit: What I was looking for was related to JWTs, not necessarily oauth, and the actual claim in spec is "JTI" (but I believe a service whose traffic I was inspecting used "jit" instead)
If you're not searching for literal occurrences of "what" and "is", why should they even be in a query?
* this is an instance of an adjunction, which is an important concept in informatics, but I understand that to actually admit the fact would probably be the kiss of death for anything claiming to be a "user" interface.
Lagniappe: marginalia, upon being queried with "precision" and "recall", came up inter alia with http://comonad.com/reader/2009/remodeling-precision/ , which I count as a win for it.
Without a feature like this, I want to have another feature that allows me to search for multiple queries at the same time and interleave the results.
But I also understand the advantages. C'est la vie.
Query understanding is one of the things I'm hoping to address this year. It's very crude right now with some pretty obvious low hanging fruit in the scientific literature that's ripe for implementation.
Going full time is the only way to go for a project you love and want to grow.
What is the business model?
How much visitors does Marginalia have?
EDIT: Also cool project!
It's a simple macroservice architecture.
If you post the domain name (or email me) I can take a closer look tomorrow.