Pains of building your own billing system
arnon.dk
arnon.dk
However, I will admit, I had a hearty chuckle at this line:
> "why can’t we just dump a file of what we need to bill on S3, and have a CRON job pick it up and collect payment?"
Under no circumstances does my engineer brain think this is a good idea. At all.
But, I will dump one aspect of my engineer-brain thoughts: My favorite "billing architecture" decision is to try to decouple billing as much as possible from credit in a system. For example, if you have a subscription system where the user pays ahead for a given billing period, I prefer to have the entitlement itself just store the expiration date and the details about what entitlements the subscription grants during the time period it is active. The billing system can store the subscription and sync back to the entitlement as-needed. This makes both manual billing by human operators (not to mention debugging and patching around momentary issues) and something like a Stripe integration very easy. You should, of course, be very careful to leave it open for extension in the future, but this seems to be a very nice decision that doesn't, in itself, limit you too much.
Obviously, this is not my original idea, but it's still something I've grown to like a lot, especially after having tried other things less successfully.
But there is probably a way to do it, and even if there isn't, there certainly is some equivalently insane design that is legal.
S3 can also be encrypted at rest. The encryption can be performed client-side or server-side [0]. Nowadays, AWS performs server-side encryption by default, unless you disable it.
[0]: https://docs.aws.amazon.com/AmazonS3/latest/userguide/UsingE...
What threat models does AWS server-side at-rest encryption protect against? Someone breaking into an S3 data center and walking off with a box of hard drives? AWS reselling obsolete hardware on eBay and forgetting to wipe a disk?
And yet what does it not protect against at all? A misconfigured load balancer that can access 169.254.169.254 and dump your AWS access key straight to an attacker, who can then use it to ask AWS to politely decrypt all of your data.
What does it really not protect again? The US.gov issuing a National Security Letter to Bezos threatening jail time if they don't decrypt and hand over the data.
To reach the spirit of the "data must be encrypted at rest" laws, I believe you should encrypt the data using your own keys which are maintained in an HSM that AWS cannot access. This means that you cannot run your entire stack on one cloud, which would be a huge driver of cost and complexity.
A lot of threats. A rogue employee sucking up data. Malware on AWS servers.
Not likely, but certainly possible. Server-side encryption is useful, but if you want to be _certain_ that your data is safe, then client-side encryption does exactly what you're asking for.
Client-side encryption can use a local key, or a key in KMS.
https://docs.aws.amazon.com/amazon-s3-encryption-client/late...
Is there a reason you have such a negative take?
Client-side encryption is a lot better. If the regulations were drafted in a technology and security aware way, maybe they would include something more detailed than "data must be encrypted at rest" which would require such practices. As it stands today, you can be in compliance without actually gaining much in the way of practical security improvement.
> A lot of threats. A rogue employee sucking up data. Malware on AWS servers.
IMO server-side encryption-at-rest does not protect against either of these, although there might be some edge cases where it does. For example, malware on the S3 gateway would see the data decrypted. A rogue employee with very limited access may be stopped by this, but I would guess that most rogue employees would be abusing the support impersonation processes or would deploy evil code that can access the data after decryption.
Client-side encryption would prevent both of these threats as well as the more common ones like losing access to your AWS keys somehow. As long as the method that leaked the AWS key didn't just leak your client-side encryption key as well.
All security evaluations should start with a threat model. If you're trying to protect yourself against APTs or nation states, you have a very different challenge than if you're trying to protect against drive-by ransomware.
Remove old entries, keep a large margin of old entries. Append new records to the file. Use a counter. (Make it 64 bit if you are paranoid.)
On the "receiver" end, keep track of what you last saw. Could be as simple as keeping track of the counter.
Add only records you never saw before.
Process new records.
There are of course nowadays so many better ways, but "real" and stable systems have been built with such duct tape.
Yeah, if that's the model engineer in the "don't build your own billing system" mantra then the article had the opposite effect on me. Now I want to build such a system because this made me feel like I'm significantly more qualified than I previously thought.
Also, with all due respect to those working on billing I have seen that it is generally handed to engineers who are, shall we say, adequate for the task but no more, and unless you're 100% greenfielding you're going to be dealing with their stuff no matter what.
From another cynical point of view, it is a situation where you can fail miserably and bring great pain to the company that is traced directly to you, but there is no success; a functioning billing system is just what is expected. It's all downside and no upside for you personally.
I have a billing system (a component of the larger whole) and I'm pleased to be able to hand it off next month. I've not enjoyed my stay there. Though it was miles better than when I brushed the data lake/data analysis world. If I had to maintain this system for another few years I could and I wouldn't scream if I had to do it again someday. Dealing with the data analysis world beyond just throwing them some data is a "never again".
The data pond, puddle, and lakeside cottage methodology would have spared you that pain.
More seriously, my experience with that world reinforced my view that it's all monkey motion (much effort, negative value).
Otherwise, spot on summary of IT dev work. Which now applies to "web services" too.
I've long been befuddled by the rise of "agile". It's not meant for product dev, right? Well stupid me took forever to realize that most work is actually IT (aka data processing) and not product dev.
It...is?
Very loosely, "product dev" projects have concrete deliverables, deadlines, and high cost of change (updates). So therefore are ameniable to PMI-ish critical path like methodologies. (Avoiding the evergreen unresolvable sophistry wrt "waterfall".)
The original motivation for "agile" was to protect the dev team from terrible customers (clients). Those who cannot or will not commit to requirements, actively resist any kind of formalism, have no clue what "done" even means.
Meaning most IT (data processing) efforts. And now most web stuff.
"Agile" is totally rational in those unwinnable situations. Because the dev team would always have something to show for their suffering. Lessening the likelihood of being the scapegoat come the inevitable failure.
Unfortunately, "agile" became a fad, a belief system, an identity. Curtailing any kind of introspection. Only in the last few years have we been able to have calm, frank discussions about "agile". Mostly because some of its progenators have broken ranks and done mea culpas.
Why? A web based product can iterate, for example. We see the most experimentation in places like Amazon, where they're constantly deploying updates and experiments. Why should such a product have a high cost of change?
Back when we burned CD-ROMs, distributing updates like bug fixes involved non-trivial effort.
Today, libs/packages and apps are somewhere between "agile" and PMI.
FWIW, for web stuff, I'm very interested in "test into prod" methodology. Neither agile or critical path; its something new. First popularized by the gilt.com fella (IIRC). But haven't yet found any teams (to join) that bold.
But you can still evolve APIs; breaking changes can occur much less than normal changes. And even then you can version them and run old and new in parallel if you want. I don't see why that would be slow.
I also don't see why making things in a way that's hard to change suddenly makes something amenable to agile.
I have never advocated for "agile", for any purpose. Ironically, IMHO "agile" is silly make work (ceremony) and too heavy weight. I was just trying to steelman the original justification for "agile" methods.
The PMI based processes my prior teams used had minimal communication and project mgmt overhead and very little drama. I see no reason why IT / web projects shouldn't be similarly managed. It's just that PMI is now like a lost art.
I only mentioned "test into prod" because I'm very curious about its methodology innovations. I think it would have better modeled and explained some of my prior experiences.
For instance, one of my teams would live code and then commit to prod changes while we were talking with our customers (for immediate verification and so forth). Almost zero overhead or ceremony.
We just came up with that system intuitively and couldn't really explain ourselves to others. (We had one senior VP in particular who reacted very negatively to our processes.) If we had the "test into prod" narrative, maybe things would have gone better.
I just wanted to say thank you for giving a robust voice to the frustrations I've felt after building multiple billing systems. This and your linked comment were a pleasure to read, and simply spot on.
Coincidentally, I am also pursuing a career transition toward data analysis. If I understand correctly, you are lamenting "data engineering" in contrast to "ad hoc data analysis deliverables"?
Note that you need to get the above details 100% correct in all cases. They can put you in prison for mistakes (almost always they will just fine the company a lot of money, but they have the right to apply prison and you never know when they will want to make an example of someone)
But there is also a long, long tail of boring problems that you won’t have time to do properly.
To put it a different way: even after you’ve solved the “mathematical” problems of billing, you still need to solve the operational stuff. Like payments integration, GL integration, invoice formatting, tax, reports, etc etc.
I think most engineers underestimate both the difficulty of the “mathematical” problems, and the length of the long, relatively boring tail.
However! The problems are interesting, the challenges are endless, and it’s possible in some orgs to get recognized as the person who is keeping the money flowing.
I do have a sense of frustration that every company seems to start off with a bad design for billing, so I’m always cleaning up after negligence… but it pays the bills.
Sometimes I think I want to build an open-source billing engine. To show off what it looks like if you do it right. But would that be killing the golden goose? Currently there’s endless demand from people who don’t know.
VAT.include(country, type, amount){
if(type == "high" && country == "NL"){
return percentageOf(price, VAT.NL.high) + amount;
}
throw new Error("Error in VAT.include: VAT not found.");
}What I did (and would recommend others do), was a deep dive into Stripes API when designing my schema. It doesn't even cost anything to create an account and run some test charges, refunds, disputes, balance changes, etc. Of course I was modelling a more complicated setup (charges on behalf of customers), but still, useful stuff IMO.
If you have access to an accountant / bookkeeper (relative?) that can help too.
https://arnon.dk/why-you-should-separate-your-billing-from-e...
I think a lot of the issues arise from the difference between payments and billing [0]. When just starting out and signing up your first customers, you primarily care about collecting a few (recurring) payments - and it's really easy to set that up with Stripe (or even just manually invoicing your first customers).
However, over time, more billing requirements gradually sneak in, such as more complex entitlements, multiple plans, grandfathering, and eventually enterprise/high touch sales customers (where the money is!) who need custom billing cycles, terms, and entitlement provisioning. Since billing is never a technical focus, numerous additions and small hacks accumulate over time, taking engineering resources away from the actual product. Eventually, this turns into an unmanageable mess that can significantly slow down the sales process or limit what you can sell.
The complexity of billing is riddled with hidden pitfalls and edge cases, and it has become even more complex now that most plans include many different limits and usage-based components and that most SaaS companies sell globally. Many later-stage companies have teams of 15+ engineers working solely on billing. I fully agree with the author that, unless it's at the core of your product, no organization should build a billing system themselves (Disclaimer: I'm the CTO of Wingback, a SaaS Billing Platform).
[0] https://www.wingback.com/blog/saas-payment-vs-saas-billing
Having worked with SaaS companies using our (Warrant) solution for customer entitlements [0] for the past couple years, this is the approach we arrived at as well (e.g customer stores entitlements in our system and checks against them when needed, adding rules/entitlements as subscriptions are updated/deleted with their payment provider). It makes it easier for companies to work with any (or multiple) payment providers, and there's a clear separation of concerns. Someone shared another blog post by OP about separating your billing and entitlement systems [1] below, but I'll share it here since it's more relevant within the context of this comment thread.
I think the ideal entitlement system is (1) dynamic (i.e. rules stored in a database), (2) can handle one off scenarios (for enterprise customers, etc.), and also has a policy layer built on top (so it supports almost any scenario a developer can throw at it -- e.g. pro plan supports <= 5 seats, growth plan supports X feature up to N times per day, etc). I think it's also a huge benefit to have a UI where non-technical folks can make changes for customers without needing to involve engineering (which was always a drag on engineering in my prior roles as an engineer).
[0] https://warrant.dev/use-cases/pricing-tiers-and-entitlements... [1] https://arnon.dk/why-you-should-separate-your-billing-from-e...
- your billing is just a mechanism to compute the invoices.
- you accumulate billing using a ledger of invoices sent and payments made (and adjustments, etc) - ie, receivables
- you have a set of entitlements, yes
- but in the middle you have a policy mechanism that determines when the entitlements should/should not be applied.
The trick is that entitlements are not just about billing. The policy mechanism is the “glue” between entitlements and everything else. It means you can give an entitlement to the boss of your biggest customer without it necessarily being linked to billing in any way.
It depends on your jurisdiction, the entities involved, the amount of money involved, and probably more. However, it's not possible to know ahead of time, you'd need to go to court and argue why this entity is being unreasonable in demanding the money be returned.
They responded saying they had no case matching my info. I gave up before investing more time because I thought it was unlikely to end well for the amount of effort I'd have to invest.
I started directly depositing money orders into his account. No problems after that. The USA government profited one month of my rental rate. I nearly got evicted when my rent wasn't delivered. No explanation was ever given to me.
I'm glad that a month's rent isn't a lot of money to you but this probably shouldn't inform your opinion on how it is to interact with the state.
[0] https://www.google.com/amp/s/amp.cnn.com/cnn/2021/02/16/busi...
The top comment was talking about money that wasn't even billed. They don't have any actual service or product it was tied back to. Also if they do have agreements with that insurer (just no payments pending), there is going to be a contract overriding default law which probably covers how to handle mistakes.
https://www.reuters.com/markets/us/citigroup-wins-appeal-ove...
heh 20 something years later and he still takes my calls to listen to my crazy business ideas.
And we still want to know why.
You can cash a check today for revenue you accounted for 90 days ago.
You need to reconcile it but it doesn't have to happen at the same time. It's very normal to reconcile months later.
The accountants just go "huh, this one doesn't match up with anything" later on.
That’s not how any of this works lol
How does it work?
In Canada, one of the main players in the space is INTRIA
https://www.cibc.com/en/cibc-websites/intria.html
They have an office full of people processing cheques, cash etc etc. In the old world, all cheques were processed by hand but in the new world, all cheques are processed by a combination of electronic imaging and MICR: people only look at the ones where tech processes fails. Which are a lot of them.
This area of finance is called "Corporate Cash Management" - a bit of misnomer because it is not just cash and is primarily electronic these days. It is essentially about the processes that need to be in place to physically and electronically move money around.
Here are some examples of what the biggest bank in Canada offers in this space.
https://www.rbcroyalbank.com/business/paying-and-receiving/m...
The do not explicitly mention cheques because that space is really well automated these days but it is mostly done really well by technology and most banks are trying to move away to electronic payments. But if you ask for it, you will get it at specially quoted per customer rates.
They certainly have all kind of strange things they must send to all kinds of people.
In terms of business process, it’s better to have one single standard procedure of depositing (not cashing, ha) every check right away. Rather than asking any person handling a check to go through some sort of decision tree on the spot.
If you need to return the money later you can always cut your own check.
Billing systems are of high complexity; I recognise that. However, if Chargebee, Solvimon, Stripe, Recurly, Orb, Metronome, Lago, Togai or anyone else has that body of knowledge, we could instead collect that knowledge in one place.
Indeed, there's no better approach than the one that serves you. If you're a subscription-based SaaS, you have specific solutions for your business. If you're a usage-based API, you have specific solutions to the billing.
But we could have all that knowledge, approaches, paradigms, programming patterns, better and best practices in one place, instead of discouraging the practice. There are also edge cases where a company is not U.S.-based or European, and a billing solution like Stripe wouldn't work, e.g. your company is based in Venezuela, and you can't have a Stripe account. What do you do in that case? You must forcefully build your own billing solution and connect it to the local payment gateways with their arcane SOAP-XML APIs.
--
On a separate note, "building your own billing system" reminds me of the topic of "rolling your own SIEM" with the typical Elastic + Grafana setup.
I don't recommend it, but I understand why it's such a hot path for an IT Security department to do it.
> Billing systems are of high complexity; I recognise that. [...]
They're less complex than whatever your developers are working on. The article tries to paint hard legal requirements as a difficulty, but in practice this means the specs are easy to find and well documented. Parts of the process do change frequently, but those parts are well labeled and well explained.
> Indeed, there's no better approach than the one that serves you. If you're a subscription-based SaaS, you have specific solutions for your business. If you're a usage-based API, you have specific solutions to the billing.
I mean, will your customers allow you to shift the burden of responsibility? Is your income important enough to verify? Can you afford the haircut?
But aren’t we supposed to be automating this stuff? What software do you think the accountants use?
You go to the accountants to ask the questions, but if you’re billing at scale you can’t just push the problem aside. The accountants already have jobs to do and billing is not one of them.
> They're less complex than whatever your developers are working on
I worked on billing systems for a long time. They are definitely more complex than whatever your developers are probably working on. It’s all in the article.
Just take proration. The user joins half way through the month and you have a monthly billing cycle. How much do they get charged?
If you think this is an easy question, you don’t understand billing.
> The article tries to paint hard legal requirements as a difficulty, but in practice this means the specs are easy to find and well documented.
That’s the funniest thing I’ve read here in a while! The legislation is not even kept in one place. It’s often opaque. You need an accountant to explain it to you. It gets worse when you’re working with public companies.
I’m sorry but your replies are the typical “billing is not that hard” of an inexperienced engineer, posted to an excellent article explaining exactly why billing is hard.
If n days of m day month is used, the charge will be n/m * monthly_rate, no? And it would be just another line item in the monthly invoice.
2. Imagine some things are billed on a recurring basis (e.g., every 30 days), and some are every 1st of the month...
3. Assume, for example, that you charge per-user. What if there are 300 users added and 20 removed in one day. Do you refund the remaining time in the month? Is this a credit on their next invoice, a negative item on the current invoice, etc...
There are many more situations that can mess this up even further.
What time zone is the billing performed in? And what time zone is the user in? Is the user part of a group billing scheme that is billed in a different time zone? Or do you bill everyone in the same time zone regardless of where they sign up? You need to know this in order to compute both "n" and "m".
Do you need to break the billing into separate line items, for example, a service component and an entitlement component? What about addons or powerups? Are they all prorated? are you sure?
What about fixed up-front charges? Surely these are not prorated?
Do you charge in arrears or in advance? If you charge in advance, how many days/months in advance? If in arrears, when do you actually raise the charge? You need to know this to compute "m" and "n". How do you remember to pro-rate it? Bonus question: what's the process you use to ensure that arrears charges are always raised? Not a cron job, surely?
Is your billing cycle even monthly? Could it be quarterly? Weekly? Annual?
Do you need to prorate entitlements? What about quantity entitlements like download volume? What about fixed entitlements like per seat licenses?
Do you prorate to the millisecond, or the day? Is there a minimum proration period?
Are there any accounting rules related to proration in the accounting regime of the user being billed?
If you provide a signup bonus or coupon that exceeds the value of the prorated first billing period, how do you apply it to the account?
Ok. That covers the stuff I can think of off the top of my head. There are probably just as many questions that I’ve forgotten.
Ultimately, the problem is not determining how to prorate, since that's a simple ratio. The problem is to work out what "m" and "n" even means, and to compute their values in a way that works within the policies of the business you're billing for. And even a simple business will change requirements over enough time.
Then, of course, there’s dealing with things like refunds and cancellations. Which, in my experience, everyone gets wrong and pushes off to manual processes. And then there's the other 12 things in TFA, each of which has at least as many twists and turns.
(edited to focus only on proration.)
I wouldn’t call billing “difficult” as much as it is “annoying”, but I also have ~18 years experience in it.
I build billing systems in telco for 20+ years, so I don’t think it’s difficult any more either, but with that experience, I’d certainly do it very differently to how most people would approach it.
The features and ideas billing usually needs aren’t “difficult” to write, but if you have someone inexperienced designing the system, your project is inevitably going to run face first in to these “annoyances” and this may present as somewhere between difficult and impossible where if you had an experienced person at the helm from the get go, you won’t be running in to the same issues.
Yeah. I guess the iceburg will present as difficulty in some ways.
For example when presenting an invoice, you end up needing to be able to access just about any part of the system. In telco you need the invoice, line items, past payments, CDRs, entitlement, pre-paid, ...
Without a global data model this becomes really complicated. It's not hard (we settled on GraphQL), but it ends up being big. And it needs to be pretty fast, too.
I could go on.
Just when you think you've learned it all, another crazy requirement pops up.
You're correct when you say people underestimate the complexity.
Even with such an approach, you’ll still miss something, but at least policy mechanisms provide a pattern for refactoring and new features.
I ended up building what we might call a declarative billing platform. You tell it what you want, and it tries to make it happen. Very successful, technically, but sadly I left the industry before it got into serious production.
I say this because building your own Billing System isn't any harder than any other platform problem. The difference is that most companies aren't in the business of building a billing system. They could hire engineers to focus on it, and they could build it just fine (it's not THAT difficult). It would just cost a lot of money and all they would do is recreate something similar to Stripe, Chargebee, and the other billing platforms out there, which have already solved these problems for you. These platforms charge you a relatively low fee (lower than the cost of a team of skilled engineers and PMs) to use it, so it doesn't make sense.
It's not that the problem is debilitatingly difficult. There is just not enough value in building it when off-the-shelf solutions exist. There are some businesses that might benefit from a custom billing solution because they are too unique to use an off-the-shelf solution. These companies do build their own billing systems and do just fine, but maintain a large team of developers dedicated to the billing product. Examples of this might include AWS with their billing, AT&T or Verizon with their billing, and so on. For most companies though, it would cost far more to build and maintain your own billing system, than it might generate in revenue. Especially relative to just paying a 3% finance fee and getting it "for free" with a provider like Stripe or paying a low fee for a wrapper like Chargebee (which charges only $500 for $100,000k in MRR).
So if you generate $100k MRR, would you rather go through the headache of building a billing system and spend $50k a month in a small team of engineers to maintain it, or would you rather just pay $500 for essentially the same thing with no headache? THe choice seems simple to me.
So it's not that Billing Systems are hard... they just don't make sense to build. Any competent engineering team could solve the problems from a billing system, it just isn't worth it.
You need to track affiliate codes back to sales and a user, handle sending payouts to an affiliate at their configured payment provider and if you want to go all out then handle metrics around visitors and present a UI to the affiliate so they can see their conversion rates and payout history.
Fortunately most of that can be built incrementally. As long as you associate a unique code to a user and wire up associating that to a sale with a specific commission % amount everything else can be done manually or skipped. For example I pay affiliates out once per month where I goto Zelle or PayPal and send out the payments. It takes less than 10 minutes. There is no front end for tracking conversions and it's also never been a cause of someone saying they don't want to be an affiliate because of that.
You really want to have "the order" as static as possible, for "the invoice" there is even legal obligation.
Then it goes something like...
I bought 5 and got a discount but I want only 3 and ohh there is a typo in my name and one in my address - sorry about that.
No problem, you credit the invoice, make a new one for the same amount with the name and address correction, then you wait for the items to be returned and checked and credit it again, make a new invoice again. The return shipment stays in limbo for a while, they order 5 more, one arrives broken, they want a refund rather than a replacement. With 2 delivery dates the system doesn't apply the 5+ discount.... and all of a sudden you have tons of transactions and "paperwork" for what you initially imagined to be 2 simple orders. They also forgot their password so they made a new account using a different email address. Looking over the logs a year later it is hard to figure out why they got a discount for just 3 units.
You do need to understand the concepts of invoices, credit, tax periods, pro-rata billing changes and so on... but all of that knowledge can be used to make an informed decision on build vs buy, rather than an automatic reason to outsource.
The only external API you need for software-as-a-service is a credit card processor, two if you're fancy. Sure after the first year you will probably have a bunch of manual work to do and your accountants will tell you the dumb stuff you've done, and you'll learn a lot about accounting :)
(I would still shop around to start a new business today, but with the confidence that building isn't very scary)
The objective as someone working on this kind of billing is simple: offer your sales teams the right options so they can close deals. This involves collaborating with your sales team to figure out what is in the “easy to implement x good enough to close deals” bucket
I can’t stess this enough. Your sales team is closing sales! Stop whining about your perfect billing system. Yes it is shocking but sometimes you have to apply elbow grease to problems involving business relationships
You should try to do things in straightforward ways when possible but really you should have a system that is legible to other teams in your company, and one that is easily to manually tweak even with automation running around it. This is like 80% of why Stripe feels so good to use. Manual tweaks are a fact of life, even if for some businesses they’re rare
Building the system always looks easy, and on day 1, it really is. It also guarantees you'll spend many hours with your Finance Director explaining why the reports you send them are complete garbage; many hours with your Support team explaining why your invoicing failed, why you charged incorrect subscription prices, any many more fun edge cases you never knew existed.
Next up, regulations change that you need to adapt to, or maybe your chosen gateway doesn't support a growing region.
And before you say "just build it better," remember: that's also time. Time not spent on your product, improving your pricing model (oh, you need to build that yourself, too, of course), or any number of things you want to do that actually grow your business instead of standing still.
As someone who has worked on exactly the same billing code and running system as the GP, I can say that you have dramatically underestimated their experience in this area.
> I can say that you have dramatically underestimated their experience in this area.
No, I really haven't.
We added complexity as the business grew, and it was an integrated part of the hosting service that we sold.
I'm not saying it's for every company and every product, but for a bootstrapped British company in 2004 it wasn't even a question, there was nothing to buy.
That experience might colour my opinion that it's easier than it appears, or that you can get away with a lot less than SaaS billing solutions provide.
So that's context I didn't have and it makes a lot of sense for 2004. I wouldn't apply this to 2024 any more than I'd apply my old boss's lessons about dBase to modern database design.
> I truly did spend a lot of time
This is the core of the issue. If I were starting a business in 2024, I'd say "stick Stripe Billing on it" and "spend a lot of time" on things my customers care about instead. The reality of building your own billing system is that you have to learn about all the things we both mentioned above and, as your business grows, start praying that the team you have working on it also learn those things. Otherwise you end up with the problems I mentioned (and I know this because I inherited these problems from my predecessors).
For simple cases where entitlements are limited and the billing is generic, I think it's fine to built it yourself as long as you're aware of the time and cost of doing so vs. buying it. As others have said, time spent on building billing is time NOT spent on building products.
You can ignore everything else, but if you can't calculate your taxes later or end on the bad side of a KYC law, you'll have a really bad time.
Also, keep in mind that just because some system sells for more money than you'd expect to spend and lots of companies that you imagine were savvy use it, it doesn't mean that the system is fit for purpose, and even less that it's fit for the purpose you want to give it. Buying doesn't free you from research.
If it's billing, sure go for it. Being fully able to customize your billing stack is I think a strength, exactly because of all the points the OP is making, plus many other traps. I think it can be genuinely interesting and there's always more to discover, but it's a full time job for a at least a small team for most companies.
If instead your goal was to run a business or build user features, do other things, offloading the billing part is definitely the best (probably the only) option.
Perhaps you're giving Walmart money because they're very good at accepting it, but as a customer there's probably something else you're expecting from the exchange...
1) Month/Quarter close. While this post talks about the account ledger, when you start dealing with a public company that has to report numbers, you have hard-cutoffs and have to make sure everything goes smoothly for month or quarter close.
2) Cash-in-transit accounting: this post assumes your payment processor is 100% correct and nothing goes wrong. When you get big enough, you need to be able to match any money that lands in a bank account with a given invoice/billing entry. You need to be able to detect any invoices that do not have a paired bank statement line. And as others have called out, the reverse can happen where you get paid for something that may not have an invoice item associated with it. Being able to deal with credits on a bank account not associated with a billing invoice can be equally important.
Maybe you can say that these are all accounting departments problem, but they are tightly coupled and the 2 teams need to work together that there are no discrepancies between their books.
Anyone even accounting adjacent to a big public company probably just had a mental brain shudder by mentioning this
Whether it be the actual accountant having to stay up to midnight repeatedly because of close, or someone on tech/business side who needed something from accounting but they just dropped off the face of the planet for N time period because of close, or a family member, or whatever.
I was just coming to hn this morning because I wrote about using FF for entitlements: https://prefab.cloud/blog/modeling-product-entitlements-with... I took some inspiration from another of Arnon's posts about SKU format for the post.
FF don't seem like the perfect place for entitlements, but in my experience they're often the best tool at hand to deal with the challenges. I'd love to hear alternative opinions.
Call me old-school, but I don't get why something like this should be outsourced to a third party.
So your accounts can have N plans. Do each of those plans map directly to a sku that they are paying for?
As you show in your blog (Nice post btw!), while you can have a flag with a numerical value representing a limit, the infrastructure for tracking usage is left to the business to implement. Imagine instead that you emit usage of that lever to an entitlement service and entitlements based on that usage are updated in real-time, even across teams. Also imagine that you have other entitlements that may be dependent on that entitlement that update as well. In addition, as limits are approached or crossed, you can choose to have soft enforcement (I.e. let them continue, but notify sales to reach out) or hard enforcement and display a prompt to upgrade.
In the spirit of OP's link, we work alongside existing billing solutions rather than try to reinvent the wheel there. We're bootstrapped via a previous successful exit and working with early customers, so if anyone's interested in chatting, even to just geek out on this topic, please reach out: trent at planship.io.
Sooo many usage based pricing things out there (ironically with totally non-transparent pricing), but I agree that it doesn't feel like the right solution has been built yet.
Oh man tell me about it. I also think the trend towards metered, pay-as-you-go pricing in general doesn't make sense for many businesses or their customers, but I digress.
Curious if you've thought much on how to distill feature-flag entitlements, service-limit entitlements, role-based entitlements, billing-based entitlements, and so on to a single interface like "Can I [action]?"
Edit: Ha, just saw your reply in a separate thread. Yes, you have thought about it. :) Would be curious to learn more.
1. Some of these are primarily updated in a UI, some are driven from some other system. Flags are customer UI, but Auth is often customer's customer (admins setting RBAC) or derived from SAML Role / etc.
2. There are very different personas / different blast-radiuses for different changes. This means that there should be different interfaces for the creation and administration of the different type of things.
3. Different scales. Dynamic configuration / feature flags are great, but you don't want the rule payload to be GB, which you would have if you tried to store fine grained permissions inside it. Zanzibar is awesome and should be available to more people.
4. It would be best if we could all agree on the names & types of these objects. I think schema files and code-generation are underused.
5. My inclination is that entitlements are just feature flags under the covers, but that there's enough different that a custom UX is warranted.
6. It's easy to add a ton of latency to applications with this stuff. First we spend 50ms getting the user... then we go get the billing for 50ms... then we get the entitlements for 50ms... There are big gains to be had by having this solved holisticaly.
hmu anytime jdwyer at prefab.cloud would love to hear your thoughts.
Will definitely reach out. Thanks!
With multiple data sources reading and writing to SpiceDB [1] (our OSS implementation of Zanzibar), those questions can further be extended into "Can [subject] [action] on [resource]", which allows for supporting not just permissions, but (as you suggested) feature flags, entitlements, role-based access control and even billing-based entitlements (if the billing system's information is supplied in as relationships or dynamically via caveat context [2]).
As a concrete example, feature flags can be represented as a straightforward permission:
definition user {}
definition featureflag {
relation enabled: user
permission is_enabled = enabled
}
They can then be checked directly: check featureflag:somefeature is_enabled user:{currentuserid}
The real power comes into play when different aspects of the system are combined, such as only allowing a feature flag to be enabled if, say, the user also has another permission: definition organization {
relation member: user
}
definition featureflag {
relation enabled: user
relation org: organization
permission is_enabled = enabled & org->member
}
In the above example [3], a feature flag is only enabled for the specific user if they were granted the flag and they are a member of the organization for which the flag was created. While this is somewhat of a constructed example, it demonstrates how combining the models can be used to grant more capabilities.With caveats [2], these kinds of questions can even depend on dynamic data, such as the time of day, whether the user's account balance is positive, or even be random based on some distribution (to enabled, for example, partial enablement of feature flags)
[0]: https://zanzibar.tech/ [1]: https://spicedb.io [2]: https://authzed.com/docs/spicedb/concepts/caveats [3]: https://play.authzed.com/s/eML6cLz9ByAZ/schema
Disclaimer: I'm CTO and a cofounder at AuthZed, and we build SpiceDB
Now... is that the best way? Probably not, for any one of those use cases. But once you've built it, it is very tempting to use them for everything.
I'm really interested in getting stronger opinions here about how to set people up for success. Being the founder of a feature flag company I feel like I'm a rope dealer.
I'm really interesting in providing good authorization primitives as well. This is kinda the whole reason/vision for why Prefab exists. If you have an authorization tool and a separate feature flag tool and a separate billing tool. You're going to be tempted to do some weird un-holy things. If you can use one provider for it, then we can really help when it comes to putting things in the right place. My dream is that you can type user.can_do?(:thing) and get back a response that has checked the authorization, ff & entitlements system and gives you a clear and comprehensible answer.
That's why we decided to offer separate feature entitlements that are tightly coupled to the billing chain and metering as part of Wingback (disclaimer: I'm the CTO). In the end, depending on your plan complexity and how much you have already invested in feature flags, I think both approaches can work well. Having some kind of feature gating in place for your customers will also make your life a lot easier for provisioning customer accounts and being able to offer custom packages.
Perhaps the general case of X is enormously tricky and complex, but in my use case I only need to handle a specific subset of the complexity. Therefore I can build my own solution that only handles the complexity I need, and it will be much simpler than off-the-shelf tools.
I absolutely adopt this stance for X=datetime. My approach to datetime requires two function calls to be provided by the library: convert an epoch time to an ISO formatted time string in a particular TZ, and the inverse. I never touch any other library code; I do all other time manipulation in my own code, in terms of those two functions.
Thing is billing money leaves very little room for error and is regulated.
From the start you'll need to understand all the dos and don't of personal info handling, the cycles and lifetime of the different billing options, cashback thresholds, support idempotency, sane DB transaction management with models that are adapted to the task, invoicing and exposing that to your client etc.
It's not everything at once, but at least half of it just for your first workable implementation.
Many people think it is a good idea to close your eyes and just start walking, resulting in untimely injuries or even death.
1. Countries and even cities have different rules for crossing the street. 2. You may not know this yet but you need to look both ways. 3. Many people do not realize they are color blind, resulting in death and injuries because they cannot tell red from green. 4. You may look both ways, but do you look up as well? In cities people can drop things from windows, in the country birds may fly over head. More and more space junk is falling to earth. 5. Thing may come at you from multiple directions. 6. etc.
Not saying this article is that bad, but I would appreciate an article framed as a how to. I am much more likely to read articles that are not framed negatively.
The article is spot on.
It's the ideal 'how hard can it be?' joke project, except it's all true.
The list of examples is too long to count, but if you approach the problem incorrectly you're soon left with the question 'how much should the customer pay' and then 'why didn't we charge this amount but something else' and when this happens too often finance comes in and asks what did you do because the reports which run the company (including whether to lay off people or not) are not trustworthy.
The issue I have is that while an article that gives people a heads up of what lies ahead would be very useful and helpful this was not that. Instead it was framed as "don't do it". Just not my cup of tea.
Most of my job involved working on their marketing pages, but at some point they had me work on the credit card processing platform, and I've sort drawn a soft line in the sand that I won't do that ever again.
Part of it was just that the code was really messy (it was ColdFusion after all), but a lot of it boiled down to having to deal with the million edge case conditions required to achieve PCI Compliance. Somehow, that company had managed to get a PCI Compliance label, and I have no idea how, because the code was held together with duct tape and prayers; there were hundreds of nested if statements, and if's nested inside else nested inside other ifs.
I'm not saying that it's unnecessarily complicated, but I know I'm not smart enough to deal with it.
[1] I'd like to point out, this was 2012; it was considered kind of outdated even at the time!
itch.io is a good alternative if you sell digital assets and even services as you can verify payments with mail.
If dolibarr didn't have auto handling for it I would have encountered it for the first time in my own program and instantly thrown my hands in the air and just done paper ledgers. I can't even imagine coding around it.
In finance C is a big one! A bad billing system can kill you, and so despite the complexity it might be worth doing just because you now are in control. If you choose to buy a billing system anyway, you still need to understand how it works in enough detail to audit it. A corrupt billing system can hide how it is siphoning your money away and the numbers seem to balance and so you don't realize it is at fault. A bad billing system will not apply some tax it should and when you fail an audit the government will demand you pay - the unexpected bill will kill you. There are many variations on both of the above. You need to audit your systems to ensure they don't happen to you.
Despite the above I generally would suggest you buy a billing system not build your own. However that doesn't absolve you from understanding billing in enough detail to audit yours on a high level. You also need to select one that your independent auditors (which might be too expensive to have but you really want) can audit in whatever detail needed.
I agree you should use an existing system and learn it well, so you can audit.
Services are also mostly made for the average user, and that average user is never you.
That's how it works in Germany.
And I'm not running a shop, just for Google Adsense.
It's just normal procedure. You have to keep all invoices for several years, 1 out of 365 days I do the yearly tax. It's just part of running a business.
* Forward billing vs billing in arrears. We had to treat different customers differently.
* Tiered billing -- e.g., if a customer used > X amount of services, they got a discount.
* Special rules around when billing starts for new customers (e.g., no bill for the first month).
There are probably more I've forgotten. We billed by data usage, mostly, and I used to say that their bill was the integral of the customer's usage graph over a month -- which was true -- but that explanation didn't gain much traction with the accounting folks.Absolute nightmare when we blew up, had to hire 6 people to beef up the billing side of the website and it definitely cost us way more than those 6 people due to mistakes.
I ended up working on parts of it for our customer service portal, lots of band-aides applied in a rush as we scaled.
Sure we could have probably botched implementation of some vendor, but we made every dumb mistake you could make and instead of doing it slowly we did it at scale.
Billing problems in general are a subset of problems one has in banking because banking system basically = billing + interest calculation + payments + accounting + multiple equally complex things and it has to work in real-time in a super-regulated environment while working in perfect sync with other systems you don't control.
If you do everything just right and manage to keep the can rolling down the street for sometime, there will be a moment when regulation changes or your business changes and your architects suddenly resign and you will end up with a mess which requires more engineering dollars than off-the-shelf system would.
Don't build your business on such foundation. It is XXI century already, no need to re-invent stuff which is readily available for peanuts.
Yet, every new neobank comes up with their own shiny corebanking ledger because they can't be bothered to look into somebody else's petrified spaghetti.
The biggest mistake I see companies making is not properly appreciating the fact that there is in fact 4-5 entirely independent business processes going on inside that monolithic concept they refer to as a "billing system", and they REALLY need to be kept seperate.
IMO the fundamental seperation that is absolutely critical to make is distinguishing between your "entitlements system" that tracks what SKUs you offer and which customers have what SKUs (and in what quantity), the "accounting system" that takes input from the entitlements system and tracks the over-time net balanced owed by the customer, and finally the actual "billing system" which periodically takes the balance from the the Accounting system, zeroes the account back to nil, and generates an actual point in time "bill" for the customer to pay.
So many systems don't properly respect these boundaries and suffer the unbelievably painful consequences.
You'll thank me later when you encounter dunning, invoicing terms, etc.
P.s But you'll hate me when you encounter pro-rating...
> Sometimes I see other people fuck up a project over and over, and I say “I could do that better”, and then I get a chance to try, and I discover it was a lot harder than I thought, I realize that those people who tried before are not as stupid as as I believed. That did not happen this time. Moonpig is a really good billing system. It is not that hard to get right. Those other guys really were as stupid as I thought they were.
Very unique circumstances but huge undertaking.
But in all seriousness, we looked at Stripe, thought about "aw man is it worth spending the .5% extra" and then spent about 30 seconds thinking about how much time it would take to roll our own, and it was an absolute no brainer.
(replace Stripe with any of the other similar competitors, not trying to be too biased, but stripe is kind of the mind-share default for SaaS startups, aren't they?)
In some cases, it's "brain dead simple", in a lot of cases it ends up "being much more complex than it seemed".
Related threads: - "Why Stripe doesn't use Stripe Billing": https://news.ycombinator.com/item?id=33191307 - "Stripe's real pricing: a primer" : https://news.ycombinator.com/item?id=33920019
The point is that I’d happily pay someone else to sweat the details than my own team. Our volume is so small that it is a no brainer, especially if you want to sell in multiple countries.
Sounds like an engineer with not enough experience. We all did thought like this once.
This is a solved problem. Don't roll your own queue / job system where you inevitably run into transactional problems. Use something like Airflow.
And mentioning S3 in this sentence has no value. S3 is merely a way to store the data temporarily. But using S3 immediately makes us not consider (more suitable) alternatives.
PCI dss conform cc handling?
What about liability in case of fraud?
Sometimes you had to reimplement whole systems that SAP comes out of the box with. Thanks not to be named fashion brand.
It's a massive shame that simple table stakes features are behind "get a quote". Folks, I'm not even sure this makes good business sense; just charge me 0.5% of revenue for premium and set enterprise up as 'get a quote' with enterprise features or 1M revenue plus.
Everyone can rent out a server. Not everyone can rent out 16.83% of a server and bill it by the second and by the byte.
Really?
Odds are that it's not, and therefore you should farm out that work to someone who _does_ make it their job. There are reasons why companies implement software from Microsoft, Oracle, and SAP, including that it's better to reduce costs on things that don't differentiate you in your market (and it's nice to have someone to "pointedly talk at" when things go wrong).
I have never met an accountant who cared even slightly even about very significant tax compliance issues. Businesses don't want to change their sloppy ways either. They want to make money first and let the accountants clean up the mess later (impossible).
That's not what I'm talking about at all. I'm talking about rounding math. If you always round down, and I always round up, and then we supposedly have 50M identical transactions, what is the likelihood that we are even in the same ballpark? fairly low.
There are even many movie and TV shows involving rounding get rich schemes, including Superman: https://skeptics.stackexchange.com/questions/14925/has-a-pro...
When you are talking about millions or billions of dollars, a rounding error can account for significant sums of money.
When you and your bank are off on how much money you supposedly have by 6 or 8 figures due to rounding, you absolutely better care.
The US has no consistent rounding laws for accounting, but the EU has a consistent rounding rule around the euro. The EU has an entire paper around rounding the euro available here: https://ec.europa.eu/economy_finance/publications/pages/publ... The US basically shrugs and makes no ruling, because if you ever go interact with many different banks, you will find out that they don't all round the same.
Also you might find the wikipedia page interesting: https://en.wikipedia.org/wiki/Rounding There are plenty of rounding options, and sometimes figuring out how an organization you interact with decided to do their rounding can be a fun exercise :)
If you know anything about dealing with money in computers, you know that floating point math and money math are completely incompatible, so you have to do decimal math to deal with money. This is like engineering of financial systems 101. Rounding math is financial systems 201.
The bank’s numbers are authoritative. We don’t know what spot forex rate the bank will use for any given transaction. We don’t know what fees we get charged. There is no such thing as a disagreement with the bank about how much money we have. Our bank accounts were empty a decade ago. Sum all transactions between then and now and you end up with our current balance, to the penny.
There are some edge cases, where for instance your quarterly numbers are reported individually (and therefore rounded) but the annual report uses unrounded quarterly numbers but this is all pretty straightforward.
Unnecessary complexity would lead to rounding footguns, but money in my world is discrete and fixed precision (just like physical).
We have many banks we do business with, one of them rounds differently than the rest. It was fun to figure out how they rounded :) Even more fun altering our software to round differently for transactions involving that bank.
I promise they and you round. As I linked above via Wikipedia, there are many rounding strategies. From what you describe, I'd guess they are Rounding toward zero, i.e. just truncating all digits past 2 decimal places, which is a very reasonable rounding strategy for banks.
If you mean: "Rounds to the nearest value; if the number falls midway it is rounded to the nearest value with an even (zero) least significant bit." Then yes I agree it's commonly used by banks in the USA. I promise though that not every bank in the USA uses it, or they might have subtle differences. Also banks outside the USA might have their own common rounding methods. If you are in the EU, they regulate how to round their EU currency(again, see the link I gave upthread, where I link to the EU standard).
Rounding is complicated.
With that approach, how do you:
1. Make your ERP recognize revenue correctly? An ERP system isn't magically going to know how your revenue should be amortized based on certain events, or what should happen when a partially-amortized order is later upgraded/downgraded to a different package/plan.
2. Ensure that your refund/upgrade/downgrade calculations are consistent between product systems and the ERP system.
Even if you emit all the necessary events (and each contains all the necessary data), there has to be some system that's programmed (or, at least, configured) to apply your revenue recognition rules to the event stream. Given how much difficulty people have even describing how things should work, I don't think you can just abdicate the revenue recognition to another system and just assume whoever configures/programs that system will do it right. Even if you can assume that, it's equivalent to saying "it's complicated so you shouldn't do it yourself; have someone else do it".
The amount of money I've seen startups forking out annually for various billing and platform subscriptions would have been enough for me to build them a better system from scratch in under 6 months. The amount of money that companies waste on rents is ridiculous and it's sad because that money could have gone to skilled engineers instead of hotshot entrepreneurs.
For my own SaaS startup, I built the billing system from scratch, it measures every millisecond of CPU time and every operation used by each account and adds them to the current active bill. At the end of the month (or whenever), I manually trigger a process to close off active bills and the system will automatically display that one as due in the UI and will open up a new one and new usage will be recorded against that new active one.
The trick is to use a function which automatically figures out what bill to use based on a specific flag so you don't need to worry about the implementation details of which bill to record usage against. This function should handle all possible situations; including the few milliseconds of delay which can exist between the old bill being closed and the new bill being opened.
That's where idempotence can be useful as mentioned in the article; that way if the system fails to record usage against a new bill because (for example) that bill was already just created concurrently by a different process then it will retry again later and record the new usage stats against that existing bill.
The trick I use is to have deterministic IDs for bills that are based on the nearest whole unit of time in UTC (rounded up or down); the amount of rounding I use allows me to control the allowable delay between concurrent processes and determines the amount of pending usage data which is kept in-memory within each process before it is flushed to its bill in the database. If the chosen unit of time is a 'whole minute' then it means that it's not possible to generate two bills within 30 seconds of each other by concurrent processes. This is greater than my database update timeout so it effectively guarantees that two processes cannot accidentally create two bills for the same period. If a specific process took longer than 30 seconds to update their usage info, that process would reconnect and either realize that an active bill has now already been created by a different process or, if not, it would try to create it again with a new ID corresponding to the new 'whole minute'.
I have full flexibility to change my billing to any time interval I want. It doesn't even have to be the same for all users.
Anyway, as you can see, it's very simple.
I found the documentation both overly verbose and lacking, so it took me longer than it should have to understand basics like subscription item ids and the 50-100 id types which people have collected into lists. Then there are simple bugs like how the go live button only shows in test mode, with no simple way to copy live mode data back into test environment(s), so developers have to manually export/import stuff that business people set up in their attempt to save time. The most basic features are missing, like a way to copy any object as json from the browser dashboard, forcing us to look down in the logs/events to copy the latest version, often with wildly inconsistent/unpredictable formats.
And Stripe does little to support actual real-world use-cases, for example: subscription events come through webhooks ok, but subscriptions have a state like "active" or "incomplete" (meaning that payment hasn't gone through yet), causing subscriptions to become stateful. Meaning that instead of a user being subscribed (yes/no), the app has to consider additional criteria and edge cases around missing or lapsed payments. And there doesn't appear to be a synchronization mechanism for eventualities like a backend server being down for a day, causing events to be missed. Stripe does a best-effort resend some low number of times, but then the events are lost. It's up to the developer to diff the backend's database with the Stripe subscription list and then sync individual subscriptions manually.
These are all just exactly the kinds of bugs/features I predicted would be there before I used it, which is why they felt like such a slap in the face. I suspect that there are conceptual shortcomings throughout nearly every service provided by Stripe, and I'd be over 50% confident that any ones I predict here would be found upon first use. These are artifacts of favoring agile over waterfall for formal engineering challenges.
Basically what I'm saying is that I consider the Stripe API to be what an MVP payment processing system might have looked like back in the 1990s under older paradigms like SOAP/XML. But nobody has created a wrapper for Stripe yet that works how people expect, more like Venmo/PayPal that maps customer use-cases to CRUD operations.
I thought that maybe Laravel Cashier would do some of the heavy lifting, but it appears to be just a thin wrapper over Stripe, with its own set of oversights.
I know that everyone uses Stripe and apparently loves it, and I applaud its efforts and accomplishments around reducing the friction of payment processing. But I can't help but feel that there is an element of having to drink the kool-aid here. I would still recommend Stripe though, so this rant is directed more at its developers who may not be aware of these issues. Stripe would do well to perform some new user testing like Apple used to do when designing its human interface guidelines, to enumerate the pitfall(s) in each step of Stripe onboarding.
There is no EU directive or law prohibiting this. You don't need any kind of "special data collection" authorization. Most SMBs just print their own invoices with MS Word or some generic software package.
You can have your in-house billing system (but all the other eu crappy regulations apply, eg. GDPR, VATMOSS), what I think you are referring to is becoming a money transmitter, which requires a special license.
https://www.fonoa.com/blog/spain-publishes-draft-ordinance-r...
You can become a financial information service or a financial transaction service.
The information service requires the periodic purchase of a certificate and registration with some very annoying hurdles.
The transaction service requires a certification procedure and a lot of money as well as all of what the information service requires.
The EU is a POS. Lobby-created laws that favor the established, rich companies.
It's an elite club and you're not in it, and they're making sure you won't be part of it. I tried to get information about the information service registration process from the German Bafin, for over 6 months over several mails. Just tell me what to do, I want to open a financial information service.
They have not given me any information at all. The employed lawyer kept evading my questions and responded with things that have no reference or meaning in regards to my questions.
I'm not part of the club and they're keeping me from becoming part of it.
Corrupt.