Show HN: how I built the largest open database of Australian law
umarbutler.com
umarbutler.com
In this article, I run through the entire process of how I built my database, from months-long negotiations with governments to reverse engineering ancient web technologies to hacking together a multitude of different solutions for extracting text from documents.
My hope is that the next time someone like me is interested in training an LLM to solve legal problems, they won't have to go down a year-long journey of trying to find the right data!
You can find my database on HuggingFace (https://huggingface.co/datasets/umarbutler/open-australian-l...) and the code used to create it on GitHub (https://github.com/umarbutler/open-australian-legal-corpus-c...).
Just one note - the link in your Github readme to https://umarbutler.com/open-australian-legal-corpus doesn't seem to go anywhere.
For someone interested in using the data (and help out with bugs/issues), where would you suggest starting?
Thanks for the heads up! I've fixed that now.
> For someone interested in using the data (and help out with bugs/issues), where would you suggest starting?
I think the best place to start is by downloading the Corpus (visit https://huggingface.co/datasets/umarbutler/open-australian-l... , and then click "Files and versions" and then "corpus.jsonl"). You can then use my Python library orjsonl to parse the dataset (you'd run, `corpus = orjsonl.load('corpus.jsonl')`). At that point, there's any number of applications you could use the dataset for. You could pretrain a model like BERT, ELECTRA, etc... and share it on HuggingFace. You could connect the dataset to GPT and do RAG over it. Etc...
Fantastic work here! I was griping to my team just last night how painful developing a chunking strategy for Australian Legislation is that while there's (generally) layout consistency within a piece of legislation, that's not true across pieces of legislation... so I can imagine the pain of trying to collate legislation across jurisdictions.
I've reach out via your LinkedIn profile - would be great if there was an opportunity to collaborate.
Absolutely, there's a lack of consistency even within the same jurisdiction and document type. It only gets worse once you want to add multiple jurisdictions and different types of documents. My best strategy so far has been to use recursive chunking where you begin chunking at the largest section of newlines. Ideally though you want some form of semantic chunking where you already know what parts of the document represent Parts, Divisions, Schedules, Sections, Sub-sections, etc...
> I've reach out via your LinkedIn profile - would be great if there was an opportunity to collaborate.
Great! Always happy to connect.
Most of the time I end up having to just take half an hour to manually regex and format plain text.
A particular case I have is where there is a draft bill put out for industry/community consultation. Quickly diffing the releases is the goal but for now usually relies on one (preferably two) subject matter experts to read the whole thing top to bottom to build an understanding. I don't think these would be available via the means you've secured. They are usually hosted on a relevant government entities website as PDFs
One last question/comment, have you considered adding some additional reference info like the federal list of entities?[1]
[1] https://www.finance.gov.au/government/managing-commonwealth-...
It's possible that they're in my database. I have included the as made version of all bills on the Federal Register of Legislation. However, if they haven't had a first reading yet, then probably not.
For processing PDFs, I recommend using `pdfplumber`, which is what I used to build the Corpus. Happy to discuss further if you'd like.
> One last question/comment, have you considered adding some additional reference info like the federal list of entities?
Do mean adding additional metadata? At the moment, I've kept the number of metadata attributes as low as possible. Every attribute added equates to more work to keep it standardised across all the jurisdictions and document types. My plan is to slowly add more attributes as I have time. I'd really like to associate a date with documents but even that is a hurdle. I have to decide what date should be the date of a document (is it the time it was issued, the time it was published, the time it came into force, the time the latest version was issued, etc... and what happens when a document doesn't have a date? should I extract it from its citation? how do I preserve time zone information? etc...).
Yes, additional metadata. Totally understand it adds in a lot of complexity but could help for fine-tuning an LLM.
With regards to dates, not a lawyer, but for Federal I would go with "Start Date", it's always the day following the End Date of the previous comp. The Date of Assent (well the year at least) is in the title, but also the first start date. The registration date can be either before or after the start date depending. [1][2]
The tricky part is when sections have different commencement dates that are detailed in the text. I don't know anywhere that is easily accessible. And, if you think about it, usually the most important information for say businesses being regulated.
I wouldn't worry with timezone per say, it's relative to each particular state.[3] i.e. why polling closes in a federal election at 6pm in each state rather than coordinated with ACT.
[1] Section 12 of the Legislation Act 2003 https://www.legislation.gov.au/Details/C2023C00213
[2] Sections 4 Acts Interpretation Act 1901 https://www.legislation.gov.au/Details/C2023C00213
[3] Sections 37 Acts Interpretation Act 1901 https://www.legislation.gov.au/Details/C2023C00213
(Symantec Endpoint Protection chrome extension)
The Australasian Legal Information Institute is a great resource and yet seems strangely unknown (to the wider public at least.)
Trivia, the only reason I found out about it was when I did some work for an Aus govt agency and found out that they shared their web site with austlii! This was back in the early 2000s.
In terms of completeness, however, AustLII and Jade win out. They seem to have almost everything if not everything. Their data is also much richer than mine. I must give props to AustLII for how they're able to hyperlink terms defined within legislation to their definition. I think they're an invaluable resource for members of the public.
The audience of my database was more so those who want to play around with raw legal data and want to feel secure that they are not breaching any laws in the process. The fact that it is stored in plain text is also beneficial for anyone trying to build ML models that only accept raw text.
Ultimately, however, I’m hoping that in the long term, the Australian Government will see the use in this project and decide to maintain it.
Is it overdue for innovation?
Probably copied / cooperated with each other, which is popular for Commonwealth countries.
One other point, and this is not specific to CanLII (I haven't checked whether this is the case) but I've seen that a lot of legislation databases have poor SEO. In fact, AustLII is usually ranked higher than governments' own websites when searching for laws. I think it's something important to get right because a lot of people just use Google/Bing/Kagi to search websites nowadays rather than using internal search engines.
Because, I just learned about it.
I now have my first two DOIs, one (data) via Dryad and one (code) via Zenodo.
This can be applied to multiple countries around the world. The world of laws at your hands.
It’s an interesting concept
Whats worse is that git is such a perfect solution for legislation.
Probably others, haven't looked.
These types of projects have the potential to influence a nation.
Fwiw and getting formatted text from html did you try
lynx —-dump url >> file.plaintext
There’s nothing at the state level right now. I’ve been considering setting up a statue scraper under the openstates umbrella but it’s a bit of daunting project to start. Lots of yeoman’s work parsing gnarly websites or evading Lexis scraper protections.
even with a database of current laws as they exist right now, the laws to change them primarily come in 2 forms:
1. verbatim additional laws
2. instructions that are essentially diffs to the current law. what words to change, strike out, sections to re-arrange and modify, as well as new lines of code. these have to be spliced in to the prior state of the law
and after we have all that, laws are often following different logic. like logic gates. One set of laws may be using "and" as a set of conditions that must be satisfied all as one, but it also could be using "and" and an "exclusive or", a set of conditions where only one has to be satisfied. but when writing it, those things all flowed grammatically and harmonization of laws wasn't prioritized.
there's a whole lot that can be improved that we don't have the infrastructure to do just yet. someone could do it, but that's the first step.
Notably, it’s thanks to them that in 2020 the Supreme Court ruled Georgia’s legal code, including annotations, is uncopyrightable.