Spanish legislation as a Git repo
github.com
github.com
The idea: legislation is just patches on patches on patches. Git already solves this. Instead of reading "strike paragraph 3 and replace with...", you get an actual diff.
The repo is the product. Browse any law, git log to see its full reform history, git diff to see exactly what changed.
Built the pipeline in ~4 hours with Claude Code. Source is BOE (Spain's official gazette) consolidated legislation API.
Exploring whether there's a business here — structured legislation API for legaltech/compliance, or just a useful open dataset. Curious what HN would build with this data.
Just thinking how this could maybe used for (automated) research / visualization on the evolution of (spanish - in this case) law
Looking at the commit dates (which seem to be derived from the original publication dates) the history seems quite sparse/incomplete(?) I mean, there have only been 26 commits since 2000.
Yeah, I think everyone is aware. It's just that the last couple dozen commits, to me, looked like commits had been created in chronological order, so that topological order == chronological order.
> I know GitHub prefers sorting with author over commit date, but don't know how topology is handled.
Commits are usually sorted topologically.
Spain is not a country with a Common Law legal system entirely like the US or the UK. They have a civil law system where prior court judgement does not form a strictly binding precedent. Prior judgements can be important, but case law is not really a thing.
Is it not the same in Spain at all?
In other words, you have to hire a lawyer. They really built a great system for themselves, didn't they?
It’s generated by this CLI: https://github.com/se-lex/sfs-processor (md with temporal tags, html or Git commits as output format)
So while this project does track laws, is there any facility to determine which laws from which bodies are relevant to a specific activity in a specific location?
No, cities don't have their own laws, but the autonomous communities do have some influence in some laws and regulations (not all), like the amount of income tax you have to pay and so on. But cities within the autonomous communities don't have their own laws.
Regardless, cities do not have their own "local laws" in the way your comment made it seem. We have national laws, and minor differences in various autonomous communities, since they have some legislative power to control their own industry, commerce, education and some more stuff.
I suspect that this should be qualified by "in the US"
Corps and cities are very similarly structured. Each are charted at the start, with corps getting governed by boards and c-suite types while cities have mayors and city council types. Both file paperwork to exist within the state. Both are subject to state laws, but are allowed to make up regulations specific to them as long as they are within the state's laws.
In the end, it's all just paperwork, at least in the US
I think local government in Spain has at least as much authority as it does in the UK, maybe more, but almost certainly less than it does in the US.
And who knows maybe a way could be found to create smart contracts (smart oracles? smart judges?) and those could lead to instant judgements.
Ed: Nevermind, I missed the "BOE (Spain's official gazette) consolidated legislation API" part. Sending jealous greetings from Germany. We just have a bunch of PDFs in Germany. And the private entity that has been publishing them for decades even claims copyright on them!
As to what can be done with the data, maybe one interesting step could be a graph-database regarding laws which reference other laws or the definitions that they depend on?
Also, in my experience (having built in this space before), regulations aren’t really the issue. Court rulings are, because there’s no open data for them in Spain. And the potential users for a paid product (legal professionals) already know the law; the key players (big law firms) have their own databases of annotated and verified court rulings and other documents.
I think the corollary that comes to mind is that reforms, with their git commits, are incrementally valuable if they refer to other parts of the legislation, previous commits, etc. to give more context as to the intent at the time of the law. So maybe there's a way to distill the legislative process into more PR and commit-oriented work—likely ex post as you did here, but perhaps in the future as part of an actual workflow.
And then maybe I'd pitch the idea to some technologically-inclined local government.
Have you considered embedding semantic hierarchical structure directly in the markdown? Something like https://github.com/wikibonsai/semtree ? It lets you build a navigable tree across markdown files using indented [[wikilinks]] as the organizational spine. Could be a natural fit for legislation that already has an inherent taxonomy (constitutional → organic → ordinary, or by subject area).
Dreaming costs nothing.
Compiling legal data for specific domains and then selling processes that rely on your private compilation is a battle-tested business plan, but there's a lot of manual work involved and the cost of that work becomes a barrier to entry.
Generally speaking, the people who'd like to cross that barrier are both open to ideas and funded well enough to run little experiments.
https://www.irs.gov/privacy-disclosure/tax-code-regulations-...
Verrsioning+search is like feature zero of any law software. In many countries (such as next door Portugal) it's even part of the standard public website provided by the state. Not to diminish your effort but yeah, people have thought of that before x)
note that even in case law systems, legislators can still pretend to be clarifying or reinterpreting things to essentially change law, but at least the supposed principle is that law has to stay coherent and go through all the pertinent checks and balances
in the UK legislation traces back to its origin, with all relevant amendments, see for example:
*Edit*: Woah ! The French crew is here. We are at least 5 quoting a variation of <https://www.legifrance.gouv.fr/> for versioning.
definition qualified_employee_discount
under condition is_property consequence
equals
if employee_discount >=
customer_price \* gross_profit_percentage
then customer_price \* gross_profit_percentage
else employee_discount
Dang, I wish all law was written like this instead of the purposefully obfuscated legalise of (lobbied) legislative lawyers meant to mislead people and slip in loopholes for their interest groups to profit of.Clear legislature is definitely something every person in the world would benefit of - if the the country's administration would want that.
Much harder to hide some impact under the carpet
Many of you asked if the code is shared. It is now: https://github.com/legalize-dev
The pipeline is multi-country by design. France already works as a second country (Légifrance data). Adding a new country means implementing 4 Python interfaces for your national gazette. The rest (markdown, git, web, API) is generic.
I wrote a guide for contributors who want to add their country: https://github.com/legalize-dev/legalize-pipeline/blob/main/...
From this thread alone, people asked about Germany, Portugal, Sweden, Finland, Netherlands, and Brazil. If you know your country's open data source for legislation, a PR is the best way to help.
I'll be honest — I don't know how much time I'll be able to dedicate to this, but I'd love to build something big. Legislation is massive and scaling to more countries takes real infrastructure. I'm setting up an Open Collective to fund hosting and development: https://opencollective.com/legalize
Live site with browsable laws + diffs: https://legalize.dev
Everything is still very precarious, time to time.
Thank you
Whenever a law is about to be changed/removed, run all the tests to make sure no regressions.
If the full compliment of software development practices were applied to legislation and ordinances we would be living in a very different world.
There really, really are.
The legal industry is well aware of that fact - and how many billable hours they stand to lose by making their work more efficient and understandable.
You know how tax prep companies spent over $90m 'lobbying' Congress to ensure that filing your taxes remains difficult and complicated [0]?
Well, lawyers know just as well or better how to butter their bread; and they will pull out every dirty trick they have to scupper attempts to make practising law more transparent or efficient in any way.
0 - https://www.opensecrets.org/news/2023/09/tax-prep-companies-...
For a while I thought about trying to write software that would turn the obscure natural-language diffs in written bills into a readable diff, showing the laws before and after with highlighted changes. But she said they just got the bills as paper printouts which weren't always even up-to-date, so it might not have helped much. Maybe now they're online. And LLMs might make the project easier.
Out of curiosity, like what specifically?
Didn’t DOGE’s failure highlight that it actually wasn’t trivial? I’m skeptical at first glance but open to being proven wrong.
For example, there are thousands of divisions of government out there provisioning largely the same systems in duplicate. E.g. the very local government here has a web portal for the sports venue bookings like pools and tennis courts. They have a waste collection portal. Local tax portal.
Only recently has this been slightly standardized but even those efforts are purely regional. You might get 5 local councils in the city using one SaaS platform, another 5 using another SaaS platform, and another 5 rolling their own. For each function of local government.
Nevermind the fact that a local government in France like this probably has very similar needs to one in Belgium or even the US.
And the worst part is they are terrible at procurement so even when they do consolidate, they're basically getting scammed.
I often think about starting a cost-plus-priced open core project to deal with these issues. Like we build common government functions, and sell it for cost plus 20% markup, with a licence that lets the gov run it themselves if we ever go bust. But then I think procurement is largely a grift game and it might not do well for that reason.
Well what I'm proposing building would be source-available and licensed such that the gov can run it themselves if it ever gets too expensive. The sub-gov entities should really band together for the negotiation though, then they can ask for whatever they want: non-profit vendors, liberal licensing, price agreements. A collective of government buyers form basically a monopsony larger than any individual vendor could ever be.
Yes, I’ve noticed that software like MS Word and Atlassian Confluence now has version control built in
No shade on the author, they made a fun thing. I'm directing my cannons more towards the parent post idea that the world needs software developers for their rare genius to use their beautiful brains to solve problems in ways no actual participant in the system could have ever thought of.
The additude that because you can prompt a LLM to write some python you are also uniquely situated to solve the world's problems is how we built an entire generation of automated solutions worse than what we had before.
Maryland just launched their regs on our platform:
https://regs.maryland.gov (https://github.com/maryland-dsd/law-xml-codified)
Feel free to reach out (email in bio) if you would like your community to publish their official laws on GitHub!
Is the parsing/uploading code shared somewhere else?
Definitely the kind of idea that would have been below my activation energy pre-Claude.
I think this approach should be standard, I have always wondered why the source of truth for these documents is not moved to a repo like git.
For others wondering, while most of the Franco-era laws were nuked in 1978, this does include lots of old laws (ie pre-20th C).
However, the source material starts with a sqashed commit in 1960 :) So no changelog before that. The BOE source though is pretty phenomonal, they've scanned files going back to the 1600s so far.
But getting the entire country's law into git is already an impressive feat.
Not git, but Congress actually does have quite a bit of data digitized. A random example[0] -they even provide XML. The Congress data is going to give you all bills - many of which do not pass, so a different mission than this project.
[0] https://www.congress.gov/bill/118th-congress/house-bill/4818
Git isn't structured for collaborative commits, but community-wide conventions kind of "patches" support for it on top of the git message body, via "Co-Authored-By: name <name@example.com>" which IIRC most platforms support, and the convention itself initially comes from Linux kernel development.
In Brazil we have lexml, a standard to describe the law and their changes over time. It's surprisingly complex.
I understand that Spain was a participant in LexML as well... I gather they've since converted to something else?
You can see how certain articles have the option to check "how that particular article was at each moment in time". That would be way harder to track, but it would be awesome if not only could you "go back in time and see what the law was" but also "how its been evolving".
I'm sure I won't be the only one curious, please enlighten me.
[1]: <https://github.com/EnriqueLop/legalize-es/commit/424cbc96507...>
When someone specifically mentions "built in ~4h with Claude Code" they probably didn't care that much about the outcome quality
You might not have merge conflicts but I imagine you could end up with conflicting guidance from two separate pieces of law (e.g., law A says you must wear green on St. Patrick's day, law B outlaws green pajamas).
Unfortunately Git is based on Unix timestamps so everything pre-1970 has the wrong date.
Well done and great to see items like this and great to see the comments.
When it comes to the law specifically, there's a whole silly setup with transcription companies as well.
https://github.com/se-lex/sfs-processor (processor with Git, md, md with temporal tags, or html as output formats)
Testing may not exhaust all scenarios but it is useful to see where loopholes may exist or whether a bill that sneaks in while you aren't paying attention is unfavorable to your values.
https://github.com/righttoprivacyact/bill/blob/main/tests/te...
It's sort of funny to think about how you'd format Thomas Jefferson's email address so that he could be credited as the author of this article or that amendment. I bet someone's already figured out how to do that though.
Looks like it's been abandoned, though. :(
If you go on particular laws, you can see the previous versions and how it changed. Example: https://www.legifrance.gouv.fr/codes/section_lc/LEGITEXT0000...
Click on "version" then "comparer" buttons and you will see a diff.
With this repo, git log --oneline -- spain/BOE-A-1978-31229.md gives you every reform to the Spanish Constitution in one command. git diff between any two reforms shows you exactly what changed in context. git blame tells you which reform touched which article. These are operations that would take a lawyer hours of cross-referencing, and they're free once the data is structured this way.
The other great thing: you can build tooling on top of it and use it with the CLI.
For instance, the big Lebowski is great and cool.
I'll take a look at data to enrich it :).
Useful for alerts in our concern area, and monitoring proposed legislation iteration and flow through committees to keep ahead.
I can imagine quite a few other more civic interest uses as well!
Hoping to open source some later myself, seems an area ripe for some open civic citizen/hacker projects. Bet some fun startups could be made on top too, gl.
It left out the tables (e.g. under 2.1 Materiales.) and the images (e.g see the very bottom).
The reality is that the official United States Code gives plenty of history for statutes, while the Code of Federal Regulations gives less but still basic history. Both are also provided in XML in bulk, though the former has a modern USLM format and the latter has an archaic schema. Case law is less amenable to git histories.
Using AI to plow through and make sense of all this would put us at risk of people knowing what the USG is getting up to.
[1] https://en.wikipedia.org/wiki/Code_of_Federal_Regulations
[1] https://www.courtlistener.com/help/api/
[2] https://www.courtlistener.com/help/coverage/
It's not a fully stupid idea, many rules can be automated and indeed have already been. The things that courts still have to decide manually are the leftovers that require more human judgment.
Edit: Consider the following words included in law.
“reasonable” “reckless” “due care”
Certain laws, like parts of tax law may be possible to turn into code, like percentages and deadlines, but even those often carry natural language conditions that can't be evaluated so easily. Seriously, try it.
Turns out, ambiguity is an intentional communicative tool.
Ditto with alcohol laws -18 there-. Selling a cyder to a 17 + 11 months guy would have a much smaller fine than a hard liquor to a 14yo.
The reverse it's true, too. 14yo are the minimum age to be legally punished. If you are 13 and barely stole some $20 Steam card -if any- you just got sentenced to spend your formative years in a juvenile center.
But, if you are 13yo gang member and you have a longass list of both petty and hard crimes and the last one has been a bloody crime with serious injuries or homicide... you can be sent as an exception to an adult prison because your mentality and mindset are not the ones from the early teens.
Especially if your body it's really developed for your age and you basically commanded mini-clans as the ones you can see in Ireland, Italy and the like. When you can smack down adults at age 13 and even ilegally drive a car, the Spanish constitution wont save you. Ultimately you must -and can- be trialed as an adult but also be able to finish the mandatory education years until you hit 16. Not easy, of course, but sending these kind of people to juvenile centers just generates more thugs than anything else.
If this is difficult for humans, imagine that for software with exact constraints.
(Many countries' laws are already available online and included in the dataset they're trained on. The project is very cool for humans though.)
P.S: Sadly my PR amendment was repealled
But I'm sure someone at some point might figure it out, you never know :)
It's interesting to wonder what kind of coverage Cycorp managed to achieve internally, and on what domains. Seemingly no job openings at the moment though!
State of Utopia[1] has this manifesto[2]. In our estimation (and we use AI a lot), it is not powerful enough to govern a country yet. We thought it was worth trying anyway.[3] We would like it to be able to handle contexts that are millions of times greater (think more like 1 billion tokens than 1 million tokens), and even so AI governance is a very difficult matter. In addition, once AI governance is achieved, how can you truly trust the governance model not to be corrupted? Transparent government run by AI is an additional point of difficulty. These days, the most difficult unsolved problem is how to introduce voting and users' comments without inviting comment spam and vote rigging. You can watch my latest update here[4] (I'm sorry, it's very quiet), and we welcome your input on all subjects. We have a fully autonomous agent currently running the country, which consists of a Mac Mini and a Claude subscription (plus our own dedicated server in a country that recognizes us, and we have a couple of other embassies by agreement and legal contract). But in practice this government just does whatever I tell it. It's not advanced enough to run a code of laws, which is one of the basic requirements citizens expect of their country. The size of problem space for running a country is larger than models can handle, but many things help.
One of the best hopes we have is with deterministic offline models where we share the pipeline with people ahead of time, so they know exactly how it will work. This could be a trustworthy matter of dispute resolution, if we get the architecture right.
For example, our country could help you sign a contract and in case of dispute, both parties could submit supporting documents and make statements and the offline model they agreed to at the time of signing their contract could adjudicate. This pipeline could be transparent from the start. This won't satisfy everyone, but might provide the minimum standard of having a code of laws that assists with contract enforcement. For now, all you can really do is keep checking our site for updates and leave comments about what direction you'd like the country to take. (For example, you can leave a comment on my latest update on Youtube.)
[2] https://claude.ai/public/artifacts/d6b35b81-0eeb-4e41-9628-5...
[3] https://medium.com/@rviragh/ai-is-not-ready-to-create-utopia...
https://www.legislation.gov.uk/ukpga/1990/18/section/9
Laws being passed are these ludicrous sets of patches:
The main difference is that in Britain the judge decisions become almost-laws, so it's like a repo with too many people with commit right. I think in Spain the judge decisions have less weight and only the legislature has commit permissions .
$ git commit --amend --author="Author Name <author@spanish.gov>" --no-edit
.. with the details for the author of each commit.Then, it would be simply amazing to run gource, sit back, and watch where all the noise is coming from.
Gource:
https://github.com/acaudwell/gource
What gource looks like:
I’ve long wanted to see gource applied in other sociologically-relevant contexts and this’d be a real good one ..