82 karma · joined August 28, 2026
I was merely thinking of the theoretical aspect of it: performance of opus 5.5 is better than sonnet 5.5 across the board, with the exception of Terminal-Bench. So I was curious why this one stood out. Was it because they focused on it during training? Did sonnet 5.5 had access to more references for this benchmark? But based on my first reading, I concluded that it might just be the safety constraints that made the difference here, and I wanted to share that.
Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.
[1] Section 8.5 of the Sonnet 5.5 System Card
We finally got a nice setup with CloudNativePG + Barman. This allows for point-in-time restores, but there were a lot of lessons to learn along the way.
- The various types of (database) backups (logical, binary, onsite, offsite, snapshots, write-ahead log...) in combination with the various types of data (database, files, cluster configuration...)
- In our earlier approaches, we tried to preserve the old database volume if it was not corrupt, and use that in our restore. This caused so many complications, because you are fighting the recommended approach. So now, when we need to restore, we always restore from backups and the 'live volume' is dropped.
- For a restore, we just spin up a completely new Kubernetes cluster, instead of trying to restore in-cluster. This is a lot easier.
- Many object stores allow for retention periods, which you can put to good use to prevent malicious or accidental removal of backups. HOWEVER, not all of them are really 'locked'. In some services, you can still override the lock with a forced delete; in others, you can still remove the project holding the storage buckets, which will delete the buckets, and so on... so test those things, instead of just blindly depending on a 'retention period' claim.
- We now automatically run a scheduled restore with verifications on a weekly basis. This requirement does shape your environment, so keep that in mind! There is also the question of how you can reliably and automatically verify that the restore restored the latest data (of a live prod environment). Various solutions exist here, but most are not very elegant!
Honestly, this is only worth it if you are already handling sufficient volume. If you are just starting out, then the easier approach is to just go with a hosted database, which will handle backups and point-in-time restores for you.
I'm not completely convinced.
If the Sidero Labs is doing well, then why would they need the acquisition? And if the Sidero Labs is not doing well, then why would Yardi keep the status quo?
We created it as an alternative to the price gouging that happened over the last year at various timesheet SaaS companies. The idea is that good software should not cost 10x as much just because 10x as many people use it. In our opinion, software should also be functionally complete. There should be no arbitrary tiers that only exist to force people into more expensive plans.
The stack is Elixir with Phoenix LiveView, deployed on a Talos Kubernetes cluster. This is different from what we normally build on, and honestly quite fun.
Lately we have been thinking hard about the temporal nature of timesheet software. Many things get quite complicated once you factor in historical versions: the need to retire, lock, change, reopen records or rates, and so on. We have been working hard to polish Abejora to make everything accessible, user friendly, and intuitive.
* They increased the cost of all price tiers.
* They moved features from lower tiers to higher tiers, forcing you to pay more.
* And this all was leveraged via "per seat" pricing. So a modest increase in price quickly becomes a lot, simply because of the multiplicative nature of per-seat pricing.
This per-seat pricing is especially absurd to us. To run the software it makes little difference whether there are 3 users or 6 users, yet the total cost of those 3 additional users was an additional 500 dollars! This got so out of hand that we decided to build our own timesheet software, which we now happily use. I have always seen software as the promise that you only need to build it once, and can reuse and leverage it many times over to reduce your costs. However, this does NOT seem to be the case anymore.
You see the same in streaming services: limits on the amount of devices you can have in your family, moving features into higher cost tiers, and so on...
We try to differentiate in a few ways:
* we try to be as fast and user friendly as possible, because nobody likes filling in timesheets: logical screens that align with the task you need to perform, keyboard support, flexible input...
* democratic pricing. a lot of competitors hide the good stuff behind the expensive subscriptions, and per-seat pricing quickly becomes expensive
* It has grown from our own needs
We use Elixir and Phoenix for our application, which is quite an amazing experience. It's a very different approach from the established "React/JS storefront" with "Traditional backend language REST API" that you see in lots of places. It's also a functional language. It's definitely something new, and something that we would recommend you experience.
We also switched from a simple debian server for hosting with standard CI/CD pipelines to a more exotic setup. We are now running a full kubernetes cluster on Talos Linux, with GitOps CI/CD using Flux. This was also something new for us, and a bit of a challenge to get it working, but now it is amazing to see it in action.
We are always looking for people who can give us feedback on our product. So send us your opinion/experience/recommendations/disappointments, and we will credit your account with 6 free months as a thank you!
There is so much to learn in this space.
There are small things, like dropping a 'jump to main content anchor' before the navigation so you don't have to tab through the top navigation everytime.
But there are also certain pages en layouts which require a lot of thought on a good keyboard navigation flow. Combine that with the need for responsive layouts, our inexperience with accessibility tools, etc; and the required effort quickly adds up.