HNHacker News
TopNewBestAskShowJobs

jacobobryant

1,374 karma · joined April 12, 2017

Personal site: https://obryant.dev

Working at https://tyba.ai

Building https://yakread.com and https://biffweb.com

email: hello@obryant.dev

submissionscomments
jacobobryant··on Plan mode is dead
I like writing spec files exclusively by hand and just asking the agent to surface questions about it, which I then clarify by editing the spec file further by hand. It keeps the spec file more manageable than having the agent generate the spec file from your conversation.
jacobobryant··on A Road to Lisp: Which Lisp
as of now neovim works great with clojure, not sure about other lisps. vs code also.
jacobobryant··on Biff.graph: structure your Clojure codebase as a queryable graph
yep, so if it's important for the application you're working on that you always run the minimum number of database queries possible, biff.graph isn't a good fit. Pathom's query planner might work as you've described; I'm not sure.
jacobobryant··on Biff.graph: structure your Clojure codebase as a queryable graph
You could always introduce an explicit caching context by doing something like `(binding [cache (atom {})] ...)` whenever you start using some functions like this. If you were trying to use this approach inside a library then you could wrap the public functions with that. Not sure if that would work for the way you were trying to do it.
jacobobryant··on Biff.graph: structure your Clojure codebase as a queryable graph
ah got it. yeah, in that case you can write a biff -> pathom resolver shim that works for everything. Again though the main thing is just the fact that they have two completely different query engines and aren't guaranteed to give the same results. e.g. off the top of my head I can think of a contrived scenario where biff.graph might not be able to resolve something but Pathom can since it can "look ahead" in the query planning step.

Maybe that kind of situation is fine and the question is really just if there are queries that biff.graph can handle which Pathom can't. If your resolvers are written correctly maybe not? But there have definitely been times with Pathom where I did something wrong that threw off the query planner in ways I didn't expect.

In any case, if I end up wanting to support migrating easily between the two as a core feature, I'd definitely want to e.g. do a bunch of generative tests to find out what kinds of queries end up with different results. Until then, a downside of supporting Pathom resolvers without a shim is that it might give people the false impression that biff.graph is a drop-in replacement for Pathom or vice-versa.

So far though the main target audience I have for biff.graph is people (biff users) who have never even heard of Pathom before, so interchangeability hasn't been a top concern. Though if many people start using biff 2 and then eventually some of the start wanting to migrate to pathom, I'd be down to explore that area.

And haha yeah nice to bump into you again--I think I remembered your username from reddit, assuming it's geokon there.

jacobobryant··on Biff.graph: structure your Clojure codebase as a queryable graph
> If it weren't for those, would Pathom be a drop-in replacement? Or is there different logic?

I could've written biff.graph to work with actual Pathom resolvers. In fact it wouldn't be hard to write a shim that takes Pathom resolvers and returns biff.graph resolvers. Although not all resolvers would work since biff.graph doesn't support everything in EQL (e.g. union queries, attribute parameters).

The query results aren't strictly guaranteed to be the same, so even with a shim I wouldn't recommend dropping biff.graph into a large project that's already using Pathom. And then that's not even getting into all the Pathom features that biff.graph doesn't support at all (lenient mode, plugins, async mode, the graphql adapter...).

But as for the core concepts, yeah I'd say they're pretty close.

> I'm curious in what scenario PathomViz is not giving enough info. I had a lot of trouble getting it working tbh

I had that trouble too heh heh--I tried running it I know at least once but didn't succeed. I don't remember exactly what the issue was... but I probably should figure that out.

Even if I got better at debugging Pathom though, for Biff I would still prefer to have an implementation that's easier for users to understand so that ideally they don't even need extra tools to aid with debugging.

FWIW there is an example here[1] of what the biff.graph error looks like when a nested required attribute can't be resolved. That file also has examples of some additional validation logic I've thrown in, e.g. biff.graph will complain if one resolver declares an attribute as a join and another resolver declares it as a scalar. Sometime for our codebase at work I'll probably write some assertions to do those kinds of checks on our Pathom resolvers.

[1] https://github.com/jacobobryant/biff/blob/v2.x/libs/graph/do...

jacobobryant··on Biff.graph: structure your Clojure codebase as a queryable graph
> From what I hear, the main draw is separating what you want from how you get it, so your calling code can just focus on what it needs. But you can use regular functions to do that. What libraries like Pathom do is leave it open to the caller what shape of data they need.

hmmm... it would be interesting to try an approach where you make heavy use of memoization and then write your functions to take the the minimal set of inputs (e.g. just the primary key for a record). I'm not sure if that's exactly what you had in mind, but here's a strawman example:

  ;; instead of having a resolver with this input
  {:input [:person/age
           :person/name
           {:person/pet [:pet/species
                         :pet/n-legs]}]}
  
  ;; you could have this plain function which calls regular functions to get its
  ;; input, each of which only need a single entity ID for their input
  (defn get-person-stuff [db person-id]
    (let [age         (get-person-age db person-id)
          name        (get-person-name db person-id)
          pet-id      (get-person-pet db person-id)
          pet-species (get-pet-species db pet-id)
          pet-n-legs  (get-pet-n-legs db pet-id)]
      ...))
And you know, I think that would be workable, even though it feels more boilerplatey to me. It would still get you the main benefit of not having to keep track of all the data shapes that are needed by the functions you're calling etc. Some off-the-cuff thoughts:

- with this approach you have a single function for each attribute, so you don't have the situation with pathom/biff.graph where there are multiple resolvers that could be called to get a particular attribute. However note that you could always put an assertion in your codebase that ensures no two resolvers share the same output key, which would then also give you the ability to know exactly what resolvers are being called.

- my example above doesn't include optional inputs, so that's logic you'd also need to write into all your functions: don't fetch the pet data if the pet ID is nil, don't return anything if the person name is nil, etc.

- if you do all that with regular code instead of dependency injection, that does mean you have more code to test, and you have to either supply a test DB (and populate it with everything the functions you're calling need) or mock out the functions. With the dependency injection approach you get plain-old-pure-functions which helps keep your unit tests nice and dumb.

- I like the readability of being able to look at the input / output queries and know exactly what shape of data I'm dealing with.

- There might be performance issues with the memoized functions approach. Pathom and biff.graph both support batch resolvers for example, and I'm not sure if you could do the equivalent as cleanly with the functions approach. And Pathom of course has its additional query planning step which does... stuff.

Going back to your comment, some thoughts:

> But I think letting the caller do subtle query changes that can completely change which resolvers are triggered and how something is fetched is kinda leaky.

This is an area where you might like biff.graph more than Pathom. Since there's no query planning step, the way that biff.graph executes your queries should be fairly predictable. It's basically just doing a depth-first traversal of your query.

(My first bullet point above is relevant too--you can always restrict yourself to having only one resolver per attribute so there's no question of what resolver is getting used.)

> How do you write the perfect resolver for all situations? How do you keep them from accidentally exploding their fetches?

Typically you write resolvers with only one level of joins/nesting and then let the query engine do the rest. so e.g. instead of writing a resolver that returns something like `{:person/pet {:pet/id 1, :pet/toys [{:toy/id 2, ...}, ...]}}`, you would have one resolver that returns `{:person/pet {:pet/id 1}}` and then another resolver that takes a pet ID and returns `{:pet/toys [{:toy/id 2}, ...]}` etc.

So there is a trade-off here in that e.g. you may end up running multiple database queries even though you could've stuffed everything you need into a single database query. That is mitigated by batch resolvers at least so you don't get N+1 query problems.

I've never needed to do this myself yet, but if you do run into any places where the performance isn't good enough, you can always write those bits the regular way (e.g. have a resolver that does a more complex query and returns nested data and/or don't even use pathom/biff.graph for this one bit). i.e. optimize where needed but stick with the default in most places.

> Is it not better to have things be explicit through function calls instead of chasing down disjointed call graphs?

There are pros and cons I think. Sometimes you want to know how an input is being computed and sometimes you want to be able to understand some logic in isolation. In practice I've acclimated quite a bit to the graph structure; I feel like it does a nice job of helping you split your code into the right "chunks".

jacobobryant··on Biff.graph: structure your Clojure codebase as a queryable graph
Thanks for mentioning datajet, I'll be taking a look at that for sure...
jacobobryant··on Biff.graph: structure your Clojure codebase as a queryable graph
> It's very cool you managed to make a mini Pathom - esp in so few lines of code :))

Thanks! The possibility of doing this had been on my mind for a while... and then I finally got around to trying it since all I had to do to get started was say "try making something like pathom but without [...]". I actually have all the prompts and feedback for the initial POC over here[1] since at the time I was using github issues/comments for my LLM-driven-development workflow.

Over the past few weeks as prep for release I went over all the code manually (especially since the whole point of this thing is for the implementation to be easy to understand) and basically rewrote the whole thing, or at least that's what it felt like.

> But the end result looks almost identical? Resolver declarations are a bit reorganized and look a bit cleaner - though you could do that with a wrapper around Pathom. Why not fork Pathom and just make some QOL adjustments?

The main thing I was going for was just to reduce the implementation size; the tweaks I made to e.g. `defresolver` were really just a side thing. To give some more background on the motivations, an issue I've had sometimes with Pathom is figuring out what's going wrong when my queries don't give me the results I'd expect. A few times as part of that I've gone spelunking through the Pathom codebase but still had never built up a complete understanding of how the query planning and execution works, which has meant that my debugging has always been more trial-and-error / black-box than I'd prefer. So I wanted to see "what is the least complex way that I could take an EQL query and figure out the results, even if the way I do it is dumber than the way Pathom does it?"

i.e. I'm trying to minimize the amount of time it takes for someone to read the code and understand exactly what's going on under the hood. Hence layering more code on top of Pathom would only hinder that goal.

[1] https://github.com/jacobobryant/biff.graph/issues?q=is%3Aiss...

jacobobryant··on Biff.graph: structure your Clojure codebase as a queryable graph
Yeah, it's a big eye-opener. I'd like to see if I can figure out an ergonomic way to do it in Python since I do a fair amount of work in that, and passing ORM objects around isn't great.
jacobobryant··on Biff.core: system composition for Clojure web apps
hehe yes. There are plenty of other languages with dominant frameworks etc; I like being in a community of experimenters.
jacobobryant··on Biff.core: system composition for Clojure web apps
If everyone wants to move to biff.core that's fine with me!
jacobobryant··on Biff.core: system composition for Clojure web apps
AI has been working out well for me writing Clojure, both in personal projects and at work. Documentation, not so much... I write all that by hand.

For Biff I've been using AI to generate a rough draft of all the code and then I take a manual pass over things before releasing. Seems to be a good middle ground.

jacobobryant··on Bttf is a command line datetime Swiss army knife
As the author of a different project also named Biff, I do have to warn you that half the comments on your HN posts will be people quoting back to the future--though I haven't decided yet if that's annoying or an engagement hack!

[1] https://github.com/jacobobryant/biff

jacobobryant··on Many Small Queries Are Efficient in SQLite
In some informal benchmarks I wrote using queries + data from a web app I develop, sqlite queries were about 5x faster than postgres.
jacobobryant··on Libre – An anonymous social experiment without likes, followers, or ads
Thanks for the feedback. I've structured Yakread (and its predecessors) as a daily email newsletter because it increases user retention tremendously. It's much less work for users if Yakread can show up in a place they already check regularly (their email inbox) rather than trying to get users right away to build a habit of visiting a new website regularly. The most common approach to this problem for consumer products is to make a mobile app so you can send push notifications; I like email a lot more since it's a bit more decentralized and is/can be less pushy (no pun intended).

But yeah, I wouldn't be opposed to trying out an alternate landing page that shows you article recommendations up front with a signup box somewhere. Could be interesting to see how both approaches perform in an A/B test. Especially if I ever made a concerted effort to get traffic from HN; then structuring the site a bit more like HN would probably be great. Maybe even aggregate comments from bluesky/mastodon? Once I get through the mountain of other TODO items that's been piling up :).

jacobobryant··on Libre – An anonymous social experiment without likes, followers, or ads
I've been working on this kind of thing over the past several years (for a while full time as an attempted entrepreneur, now on the side for the past couple years). The latest iteration is https://yakread.com -- hit "take a look around" and you can see the "home page"/a list of recommendations without signing up. The recommendations are personalized, i.e. the probability you'll see any particular post depends on your individual interactions with past posts, if you've signed up. (it does collaborative filtering with spark mllib). So that may be a bit different from what you had in mind, since your comment sounds more like an unpersonalized system, but with some extra exploration thrown in. However in practice I suspect the biggest thing the collaborative filtering is doing at Yakread's current scale (not much) is learning which items are good/bad in general.

I also do have some methods baked in for doing exploration. "Epsilon greedy" is a common simple approach where x% of the recommendations are purely random. I do a bit more of a linear thing where I rank all the posts by how many times they've been recommended, then I pick a percentage 0 - 100, then I throw out the top x% most popular (previously recommended) items. that also gives you some flexibility to try out different distributions for the x% variable.

The source is at https://github.com/jacobobryant/yakread

jacobobryant··on Structuring large Clojure codebases with Biff
You can do that, it's just slow if there are a lot of results.

Agreed you want to keep data in your main database normalized since it's easier to reason about and avoid bugs/inconsistencies in the data. The inherent trade-off is just that it's more computationally expensive to get the denormalized data.

The idea of materialized views is to get the best of both worlds: your main database stays normalized, and you have a secondary data store (or certain tables/whatever inside your main database, depends on the implementation) that get automatically precomputed from your normalized data. So you can get fast queries without needing to introduce a bunch of logic for maintaining the denormalized data.

The hard part is how do you actually keep those materialized views up to date. e.g. if you're ok with stale data, you can do a daily batch job to update your views. If you want to the materialized views to be always up-to-date then things get harder; the solution described in the article is one attempt at addressing that problem.

jacobobryant··on Structuring large Clojure codebases with Biff
Thanks! It's all from scratch.
jacobobryant··on Structuring large Clojure codebases with Biff
yes, that's part of it.
jacobobryant··on Writing your Clojure tests in EDN files
Agreed, thanks for sharing that post. Good read.
jacobobryant··on Biff – a batteries-included web framework for Clojure
Glad to hear it!
jacobobryant··on Biff – a batteries-included web framework for Clojure
You can use Selmer: https://github.com/yogthos/Selmer
jacobobryant··on Biff – a batteries-included web framework for Clojure
Hey HN. Since this has showed up here maybe a status update would be interesting? This continues to be my main side project--amusingly it's had more traction than any of the startups I tried to build with it. Over the past year I've been working on some experimental features for Biff that are meant to help with medium-to-large codebases[1] (I've been doing this as I rewrite one of my Biff apps from scratch). There haven't been many code releases in that time, so I've got a decently sized backlog of things I'd really like to get to. E.g. XTDB v2 is almost out of beta; once I finish the app rewrite, that's next on my list.

[1] https://biffweb.com/p/structuring-large-codebases/

jacobobryant··on JSX over the Wire
The framework checklist[1] makes me think of Fulcro: https://fulcro.fulcrologic.com/. To a first approximation you could think of it like defining a GraphQL query alongside each of your UI components. When you load data for one component (e.g. a top-level page component), it combines its own query with the queries from its children UI components.

[1] https://overreacted.io/jsx-over-the-wire/#dans-async-ui-fram...

jacobobryant··on Yakread's Ranking Algorithm
It is pretty subjective--I design the algorithm with myself in mind first, but I can definitely see a heavier bias toward exploration/variety being better for some others. Something to experiment with.

Thanks for the sampling ideas! Interesting stuff.

jacobobryant··on Yakread's Ranking Algorithm
Also--I think the pseudo code you have isn't /quite/ correct. If x is greater than p, we don't immediately take a random element from the list; rather we go to the next element and generate a new x and repeat. I.e. with p=0.1, there's a 10% we immediately take the first item, and if we don't do that, then there's a 10% chance we immediately take the second item, etc. we only pick a completely random item as a fallback if we get to the end of the list without picking anything.
jacobobryant··on Yakread's Ranking Algorithm
Yeah that makes sense! Once I'm done with the rewrite I'm planning to go through the app and convert the remaining htmx portions to datastar--should be a decent learning exercise.
jacobobryant··on Yakread's Ranking Algorithm
For the interleaving, yes we want to prefer the item that's been skipped fewer times. I got the wording backwards in the article; I'll fix that.

For shuffling, I was trying to come up with an approach that would recommend the top k items roughly the same amount regardless of how many total items are in the list. E.g. say you have 10 subscriptions that you really like--I want to have those be a reasonable portion of your recommendations whether you've subscribed to 100 other subs or 1000 other subs.

Contrast that to a weighted random shuffle where each subscription's weight is its affinity score and we sample them based on weight without regard to their order in the original list. That approach is much more influenced by the size of the total list, and my experience is the handful of subscriptions that I really liked were always drowned out by all the other "speculative" subscriptions I had accumulated in my account.

The computational complexity ends up being OK because we generally don't actually need to shuffle the whole list. I recommend items in batches of 30, so we just need to get that many items and then we can abort the shuffle. There probably is some more efficient way to implement this though.

During implementation I was mostly thinking of this as "sampling" rather than "shuffling" actually, and just ended up describing it as the latter when I wrote the post.

jacobobryant··on Yakread's Ranking Algorithm
Fun to see this on the front page! I worked on Yakread full time for about 8 months as an attempted startup, after a few years of other recommender system startup ideas. Now it's a side project that I develop on the weekends after my kids fall asleep, aided by caffeine (me, not the kids). I'm in the middle of open-sourcing/rewriting it. Hopefully will be done in a couple months? Then I can finally get back to adding new features. I talked about some potential ones in my previous post: https://obryant.dev/p/rewriting-yakread/

also I guess a link to the actual app wouldn't hurt: https://yakread.com

Page 1 of 10Next →