813 karma · joined October 11, 2023
1. They want to continue pretraining a model instead of starting from scratch. But actually people might not know that you can pretty easily reuse model weights even when training with a new tokeniser (I’ve got a blog post on how to do that: https://umarbutler.com/how-to-reuse-model-weights-when-train... ).
2. Because it’s convenient for end users. Tokenising and chunking really large corpora can take a long time and it’s nice that I can use the GPT2 tokeniser and then train a bunch of different models on that data without having to retokenise everything.
Maybe if we’re talking terabytes it might not scale as well but so far in my experience training tokenizers has never been an issue. It’s training models that takes ages.
And if I am wrong, and it does mean to affirm, then doesn’t that by necessity mean that it involves ‘offering an opinion’? Affirmation and/or agreement is not neutral.
Glad to hear it corresponded with your lived experience, it really was surprising to see how the map correlated with my own understandings of the law developed through my degree!
This is the bigger point. In my own university studies, there was a clear segmentation between the common law and statute, although they are certainly interrelated.
It’s also worth noting that the boundaries between cases and legislation were not absolute, there were areas of the cases ‘mainland’ that contained legislation.
My point on the style was that in addition to differences in purposes, they are also textually different, which can indeed bleed into semantics.
You're right — my map does not necessarily represent the underlying semantic structure of Australian law, it is an approximation, one that is biased by the data I used (which as I mentioned, is missing laws and cases from a number of jurisdictions), the embedding model I selected and the dimensionality reduction model I used to project my embeddings into a two-dimensional space, to name a few.
Because I was writing for both legal and data science audiences, I tried to avoid sounding like my inferences are anything more than just inferences but without getting too technical and explaining the inherent limitations of any attempt to semantically map knowledge with today's technology.
I will just say though that, having studied law myself, Australian case law is indeed somewhat of a continuum. A single case may touch on many areas of law and there are no restrictions in terms of subject matter on what precedents a judge may draw upon in reaching a decision, apart from that they are both relevant and binding (or, if they are not binding, are not treated as such).
It was also interesting to observe how the final clusters that developed were uncannily similar to the way in which I was taught law at university. It goes to show, there's a lot of thought put into the design of our legal courses here in Australia. In fact, there are 11 subjects that are mandatory, known as the Priestley 11: https://en.wikipedia.org/wiki/Priestley_11. All of those are reflected on the map, although some have been rolled up into larger categories or divided by other means.
I’d love to see a semantic map of the internet, I’m considering having a crack it as well, but it’d be a monumental task. There is this cool map but it’s quite dated: http://internet-map.net/
I don’t think I came across datamapplot, I’ll have to check it out. I know D3.js would probably be able to solve all my problems but then I’d need to build a lot of it from scratch and I don’t particularly enjoy Javascript (although sometimes you have no choice, I still needed to inject some code to modify the map to enable clicking on data points to open links and not capturing double clicks when zooming out (since there’s so many data points that you’re bound to misclick and end up opening a data point by mistake)).
What if you could take every law, regulation and case in Australia and project them onto a two-dimensional map such that their distance from one another was proportional to their similarity in meaning? What would that look like?
Well, it might look something like what you’ll find in my latest article. I took every law, regulation and case in the Open Australian Legal Corpus, the largest open database of Australian law (which, full disclosure, I authored), and used text embedding and dimensionality reduction models to plot them on a two-dimensional map (excluding noisy documents that I was unable to cluster).
My map represents the first semantic map of Australian law.
Some of the most interesting insights I was able to gain from this endeavour are that:
• Australian case law is more of a continuum than a rigidly defined structure;
• Migration, family and substantive criminal law are the most isolated branches of case law on the map;
• Migration, family and substantive criminal law are the most distant branches of case law from legislation on the map;
• Development law is the closest branch of case law to legislation on the map; and
• The map does not reveal any noticeable distinctions between Australian state and federal law, whether it be in style, principles of interpretation or general jurisprudence.
If you’re interested in learning more about what the map can teach us about Australian law or if you’d like to find out how you can create semantic maps of your own, you can read the full article on my blog, which provides a detailed analysis of my map and also covers the finer details of how I built it, with code examples offered along the way.
is used for tuples, for sets, for frozen sets, for pickles, for bytes and for bytearrays.
I thought it was pretty ingenious but clearly I’m not the only one to think of it.