Four Areas of Legal Ripe for Disruption by Smart Startups
lawtechnologytoday.org
lawtechnologytoday.org
> But even with today’s modern communication tools, both customer experience and lawyer workflow have remained stagnant.
At a large firm, legal practice is unrecognizable compared to even 10-15 years ago. Everything is electronic: filing and docketing, document collection/scanning/OCR, legal research, document management (DMS + version control). Everyone communicates almost exclusively via e-mail, and remote work facilities are ubiquitous.
To the extent that technology is available that's not getting adopted, it's because it's not good enough. Predictive coding can be very helpful, but it also has a fixed setup and training overhead that makes it less efficient for smaller matters. That's why arguably the biggest shift in discovery in the last 15 years hasn't been to automating it, but outsourcing it to contract lawyers.
In the area of research, Westlaw and Lexis still rule because of their completeness and accuracy. If I need a copy of some statute enacted in 1873 I can not only find it, but I can get original scans so I can verify the text is free of OCR errors.
Moreover, things that are easy on the rest of the web are not easy when it comes to legal (or scientific) research. PageRank, for example, works great when everyone searching for "skiing near Tahoe" is looking for the same popular pages. But when you're doing legal research, a lower-court case that directly addresses your issue but isn't widely cited is much more valuable than a highly-cited Supreme Court case that doesn't address your issue. And computers still don't really understand either what issue you're looking for or what issue a case is about. So ancient technology (search for this word near that word) still rules the day.
> There are good reasons for this, as law firms tend to be cost agonistic (since they pass costs directly to their client)
This is oft-stated, but economically fallacious. Price is a function of supply and demand. The client cares about total cost for a particular legal service; she doesn't care about how that cost is broken down. If the client's budget for a matter is $300,000, every dollar that goes to costs is a dollar that doesn't go to the law firm. This is true even if you're billing by the hour, because in the long run, a firm will raise rates until hours x rate = client budget.
If everything is electronic, his firm and the judges they deal with didn't get the message...
But when it came to signing documents, that had to be done on paper. And notarized, which was a simple trip to the UPS store. And the originals mailed around, which was a hassle. The parties involved don't even all live in the same hemisphere, so there were some annoying shipping times.
I work for a very large company with three dozen law offices across the US. Listening to what is considered "normal" in other offices is always interesting. I am, unfortunately, in an office that practices in courts that are paper-dependent, even though there are rules that allow for electronic filing and service in many circumstances, all it takes is one litigant (or the judge!) to say "eh, I'm not comfortable with this, we're going to do this the old fashioned way" and we're done.
I think this is the key. The big revolution in legal research that I'm waiting for is the ability for the computer to understand something like: all cases where a) Claim X was brought as a counterclaim and not as the original claim; b) counter-claimant made Y argument citing to case Z but not case W; d) requested relief was R; and e) court made its decision based on Factor F. Until then, "search for this word near that word (and boolean operators)" is extremely powerful and can get me most of the way pretty quickly.
They're looking for a solution that will let better manage case and customer info internally, and push public info to the public site. The idea is, only enter information once.
Custom development is of course an option, but first I want to see if there are any good available products. Are there any good off-the-shelf or hosted law firm management software packages that provide an API? Thanks for any advice you can spare.
[1] If they don't have a DMS story, they need one!
This is true where you are doing fine-grained research, but there are plenty of lawyers who are constantly delving into new areas of law that are adjacent to or tangential to their primary area of interest or practice. When this happens, it can be phenomenally useful to get up to speed on the area of law by seeing which decisions have the highest PageRank, or PageRank weighted by certain factors, etc.
In this way, one way to think about the startup that implements a PageRank algorithm is that it is not disrupting electronic legal research, but rather disrupting legal textbooks. In my experience, the fastest way to get abreast of a new field is to find the leading textbook. One could imagine a sophisticated database which could altogether remove the need to consult a textbook by painting a picture of the leading cases, key pieces of legislation, which sections are most often referred by which cases in which context, etc.
Automating these aspects of legal practice wouldn't disrupt the way that industry functions. It would just make a lot of paralegals and young lawyers obsolete (and make legal services a lot cheaper).
1. extracting documents from within other documents (attachments out of an email, files out of a zip, embedded excels out of a word doc, images out of a powerpoint, etc)
2. convert all said documents to some kind of standard media format so that the native viewing applications are not needed (all said document types to png, or pdf, or tif)
3. allow full-text searching across all electronic files
With these kinds of tasks available as an automated feature, the real product would just allow a bunch of attorneys to review the documents and apply tags or labels to them. Once they've gone through all the documents, there is generally an output from the system that summarizes their work and provides the relevant documents, notes, etc.
Over the years of writing this kind of software, we've encountered a never-ending amount of complicates with file types, feature requests, etc. The real complexities with this kind of software is making your software work for a large number of customers. Every customer probably has a different idea about what they want this kind of tool to do for them.
That's where I see potential in this market. Ediscovery is a pain-point for law firm clients - especially large corporate clients who are constantly involved in complex litigation. Document review has to happen in order to effectively litigate (gotta find the smoking gun!) but when a bill comes through with hundreds or thousands of attorney-hours devoted to reading your opponent's old emails, ouch. The client hasn't even seen a work product yet.
FYI, Blake also blogged the notes on Peter Thiel's start-up lectures at Stanford.
Which became source material for Zero to One.
In line with Peter Thiel / Palantir's philosophy that the human brain is an amazing machine not to be supplanted by computers, but one that should be used to its fullest, augmented by computers, Judicata's software involves using NLP as much as possible, then feeding or 'striating' that information to lawyers or legally trained people, depending on the complexity of the information extracted, for their confirmation. This is in any case necessary because NLP cannot get close to 100% accuracy for the information they are trying to extract, and you need 100% accuracy in the legal domain (e.g. it would be unacceptable to get the legal claim wrong, c.f. Google search).
One consequence of structured legal texts is improved search. What many don't realise is the degree with which structured search on legal texts will improve legal research. E.g., if I want to find all cases in the last 10 years where the plaintiff claimed breach of duty in an occupiers' liability suit, I simply cannot. To find that batch of cases (accurately) would take me hours. If the legal claim was a structured piece of information, I could just search for it. As an ex-lawyer and ex-legal researcher, the number of hours that could be saved per lawyer per year could easily be in the hundreds, and this is at charge out rates of $300-$1k per hour. This is, similarly to the above comments, in line with Thiel's investment thesis to 'improve something 10-fold' or 'make a quantum advance to cause adoption / change consumer behavior'. I think most people seriously underestimate how significant of an improvement structured search would be.
The other thing that Judicata are flying under the radar about, a little bit, is the ability to use structured legal information for other purposes. High on the list is analytics, which Itai Gurari mentioned at the end of a talk, but merely in passing as if it was inconsequential. I think this is pretty clearly a multi-billion dollar market waiting to be made. If you look at what similar firms are doing in niche areas of law, e.g. Lex Machina, and look at what they are charging, and extrapolate the types of questions you can answer with structured legal information, the potential becomes clear. Again, this is in line with Thiel's investment thesis to 'create a market a dominate it, rather than compete in an existing one'.
The primary difficulty for Judicata or somebody undertaking to do the same thing is that the task is mammoth in just about every respect. As such the optimum strategy is likely to attack a niche jurisdiction and then build out the product. You can't go 'full-lean', because you need at least a semi-complete data set, but you can start 'small'. Hence, Judicata have been working on a niche jurisdiction of law as their first project: Californian Employment Law. While I am not in that jurisdiction (not even in America), that seems to me to be a very reasonable area of law to start with given that most legal claims (I think) are found in California's employment law statute (as opposed to other areas where the legal claims are found in Judge-made common law). Furthermore, there are a ton of neat pieces of information in employment law which you can structure, e.g. in a discrimination case, what factor was the plaintiff allegedly discriminated on - race, age, gender, etc. Finally, uptake would be high among employment lawyers who research at reasonably frequent intervals and have a practical need for more accurate search; compare this to constitutional law for example.
While part of the reason they are operating in stealth mode is simply because it takes so long to build up a semi-complete data set, I think the other part of the reason is because the biggest risk for such a firm is that Lexis, West or Bloomberg will start doing something similar. Imho, it's likely they will eventually but the risk of Judicata catalyzing that process is pretty small.
There are a few other firms operating in this space but with fundamentally different philosophies. My view is that these other firms are simply taking the wrong approach and simply want to release a product and build on it now in the lean tradition. Judicata's product is the type of product where the question is not whether there will be adoption, but rather whether or not you can actually build the product on your budget and in the time frame required. Imho, if Judicata can successfully create what they are planning on, it will flatten their competition. The real question is whether they can.
(And they're welcome.)
The law firm has to pay for the data collection, the disk space to store the data, the transmission of data back and forth (investigator found something good, buys external HD, ships it via the mail), the data analysis, the investigation and reporting, and then sometimes the expert witness.
Sometimes, the law firm does not know what they're looking for, in other words, there is no smoking gun piece of data. Sometimes, the goal is to find something, anything, that would hint, point, or prove a goal.
What this means is that a retainer can either be here is 10hours worth of analysis/investigation to find the email we know was sent that contains this particular text. They do not plan on the analysis and investigation to exceed that and usually the result is we found it/didn't find it and we did it in the time allotted or under the time.
It can also be, "we're looking for evidence that this type of event has occurred". This is where the billed-hours start stacking up. Its hard enough digging through other people's emails, documents, and pictures looking for something, let alone digging without having something in particular to look for.
The point is, I believe that they, the law firms, want to, and need to, bring this in house. This greatly improves the process. But now they need security, real security, because its not their data being stored, it is their client's client's data and so forth. It becomes very sensitive.
They need infrastructure. They need to be able to forensically acquire data. Forensically store data. Forensically analyse data. Forensically share data. The infrastructure needs to be fast, easy, and effecient. We're seeing 6TB hard drives now... shares of much, much larger size. And they do not want to be storing their client's data in someone else's cloud.
Then they either need technicians and investigators or the ability to hire and grant access to their data on their network to the tech/investigator. They need technicians to provide solutions to the inevitable problems run into (i.e. how can I acquire each of the 2 drives in this FusionDrive raid and return to the lab and build the raid on something other than osX?) And they need investigators experienced at honing in on relevant data while digging through vast troves of data.
Pretty shocking when you think the first book on digital evidence didn't come out until 2004. You want a legal field ripe for disruption? It's definitely the eDiscovery space.
Here's the book I was referring to:
http://legalsolutions.thomsonreuters.com/law-products/Treati...
Most contracts are credible because their terms can be enforced by sanctioned violence. "Smart Contracts" are credible because they are enforced by distributed verifiable cryptographically-secure automation.
[1] For example, when the DMV database confirms that a certain car is now registered in my name, you automatically get 4 bitcoin.
Contracts are code. Many of the concepts are similar (defined terms / includes / gotos) so the tools don't need to be changed that much. The challenge is getting others to adopt them.
Usually though, the moment you move into a transaction of even medium complexity, while this might be helpful it can't be ultimately relied on - if you were to track a defined term or a section reference, a change to that term or section reference would impact not only the places where that term or section reference is specifically used, but also where a concept depends on such term or section reference. A basic example of the change of an actor from singular to plural (originally there was one purchaser, now there are two purchasers) - then, every pronoun and verb would need to be changed to plural form. This is why you sometimes find contracts wonkily sticking with a plural defined term when really there is just one entity/person described by the term, or vice versa.
This mostly points to what I always say about the challenges to true legal disruption: common law, statutes and contracts all depend on language subject to interpretation (and in the case of common law, interpretation IS the entire name of the game - see every supreme court case of the past 30 years to see how much interpretation varies), and unless I am missing some major breakthrough, we have not come to a point where language is understood systematically enough to truly "hack" complex legal concepts as currently drafted (aka. using language).
I could, however, imagine a new legal system that depended entirely on data, numbers and systems instead of historical language, but it would first require us to all agree that the system would govern and agree on the rules (or lack thereof) of interpretation, which given the vested interests most of society has in the current system, would be pretty challenging to implement.
That being said, you can see how this spectrum works by comparing the common law system (like US/UK, which relies heavily on interpretation of judicial opinions) and the civil law system (France, Korea, etc., which relies heavily on a more formulaic interpretation of statutes) - way more costly litigation and more politics in common law systems when compared to civil law systems. The trade off is that common law is (arguably) more dynamic (judges can overturn statutes unilaterally - civil rights, etc.), whereas civil law systems require legislatures to make changes to statutes.
Full disclosure: I used to be a lawyer, so I am biased.
Similarly, it would be interesting to try and take a relatively conventional transaction (incorporation documents, convertible notes, Series A) and place the key economic terms in a spreadsheet, and then incorporate the other legal boilerplate (indemnification, securities exemptions, etc.) by reference to a standard T&Cs-type document - I know some incubators already do this by having the key terms as a "fill-in-the-blank field (I think AngelPad does this..). The agreement would then just be a spreadsheet-like form, with the other terms incorporated by reference.
eDiscovery has largely been solved for most corporate environments. There are tools to collect data in a defensible manner, to "process" (i.e., index) it, and to review it. There are even some products that aggregate these functions together, however, it must be well-noted that each of these functions has a different user/customer and occurs at a different timeframe in the discovery process.
Many of the dominant tools do have their warts. But the money that was once in this space--the eDiscovery collection product I wrote sold for a couple million to its first customer--is no longer there. Prices have dropped dramatically and its now a commoditized market. So you'd have to work very hard for very little gain to displace any of the dominant players.
Note that TFA was written by investors in a new eDiscovery startup and TFA seems mostly like latent marketing for them. I don't know anything about them--good luck and all that--but I'm very familiar with the space and I don't envy them.
And they are missing a huge piece--predictive coding and advanced analytics (email threading and near-dup are EXTREMELY common). If you are in NLP, ML and/or IR, the legal industry is probably one of the most exciting places to be. Huge datasets, available annotators and tons of money. It's a red-hot lab of state of the art techniques being tried in the real world instead of on the Reuters, 20-newsgroups, Enron, and other "canned" datasets.
One concern that I have is this "Secure Infrastructure--no installation required". This suggests the cloud. The other is "Upload your data by FTP".
Think about e-discovery for a moment. The toxic material stored is pretty much a superset of all known toxic material--PCI, PII, HIPA, and so on. Many customers require physical control of where this information goes. Zero installation means that it is not a server under your control.
There are many pieces to the EDRM process (http://www.edrm.net/resources/edrm-stages-explained) and it is not clear that DISCO covers all of them. 'tucros3141 in a parallel comment mentions some of them.