Attacking a Pay Wall That Hides Public Court Filings
nytimes.com
nytimes.com
In passing, charging fees is clearly unfair given the role courts play in our society. At the state level, such as in New York state courts, filings are freely accessible.
Governments are sometimes just like people -- they get addicted to revenue streams and have a hard time adapting when those revenues change or go away.
It’s entirely self service.
With fairly high fixed costs and, likely, demand varying a lot as a function of price, picking a break-even price point isn’t really possible up-front.
For example, if they were to charge $1,000,000.— per page, chances are they won’t recover costs, but that doesn’t imply they should ask more.
I think governments should just give up on the idea of recovering costs on this kind of service. That may mean some people will benefit more from it than others, but then, be that so. Alternatively, put a reasonably high cap on the ability to query the system.
The public deserves more than the base line of free access to court records. Give the public access to clean normalized case data like in LexisNexis and Westlaw's outrageously-priced products, including a full database dump for civic hackers to improve.
Aside: I worked on a product which relied on scraping PACER. A foolish integration test on our CI server once racked up a $50,000 PACER bill which we didn't notice until the end of the quarter.
That is not to say that politics and law can’t be data driven. But the data we need in those areas is not data you can derive by analyzing user or government generated databases. We need “hard data” such as how different tax policies impact economic growth. That’s not the sort of data analysis “civic hacking” can give us.
Perhaps we have witnessed different types of citizen contribution. I am not suggesting data science. Many would happily join a volunteer endeavor to clean up our country's case database and build a great UI on top of it which encourages public interest in the judicial system.
> The data is in fact out there
It is not. There is no complete, real-time, or normalized open case data out there, and LexisNexis and Westlaw's products are not priced to be accessible to individual citizens.
What would be the value of that? What problem is that trying to solve that is worth giving up $140 million a year in revenue that’s currently coming almost entirely out of the pockets of lawyers and law firms?
> The data is in fact out there It is not. There is no complete, real-time, or normalized open case data out there, and LexisNexis and Westlaw's products are not priced to be accessible to individual citizens.
I’m not sure what you mean by “normalized” in this context, but Lexis and Westlaw do a ton of additional analysis on top of public cases at their own expense. That information isn’t in PACER. To generate it you’re talking about even more expense. And while the existing databases are not updated in real time, they’re quite complete for federal cases, which is all that PACER covers. And I’ve never seen an interesting application of that data beyond simply allowing you to read the cases.
There are a lot of ways governments generate revenue even if we want to restrict it to the legal industry. Obfuscating public information to the point that there's an oligopoly controlling access is fucked up.
It is a well known phenomenon given things like how quixotic the article 13 upload filter is because the expense locks out new competitors.
[1] https://www.lexisnexis.com/en-us/products/courtlink-for-corp...
[2] https://legal.thomsonreuters.com/en/products/westlaw/dockets
That being said, despite that the overlap between West/LEXIS is very small. The docket tracking is an adjunct service that you use in the unusual situation where you’re keeping tabs on a case you’re not participating in. Its not real time (day after) which is a big thing because the PACER/ECF notification is the official notice that triggers time periods and deadlines. It’s also vastly more expensive than PACER (like $50 per document).
That doesn’t address OP’s point, which is West/Lexis’s market dominance. And that is based on those systems having 200+ years of human annotated and indexed case law. PACER doesn’t have that data.
You mean their client's pockets? Thus making legal assistance more accessible for the wealthy than the poor.
And just because you can't think of a good use of the data doesn't mean there isn't one. I can think of dozens of analyses I would love to run on a large set of cases and filings. You could develop models to predict how likely the language used in a filing is to contribute to a favorable ruling. You could measure the speed with which filings are handled and compare it across jurisdictions. Picking a specific type of case and one where there's a jurisdictional split implicating administrability, you could develop a rough proxy measure for the relative costs of the alternative rules. Etc etc. And that's just the limited subset of what I randomly dreamt up in three seconds that I'm willing to take the time to type out on my phone.
Data analysis like that is possible and even relatively easy with a large, standardized, freely available data set (which BTW is what the person you are replying to meant by normalized). Twitter has been analyzed to death using NLP for that very reason.
The price of legal services is based on supply and demand, not individual lawyers’ cost structures. Moreover, the rules already allow free PACER access for the poor. I’d bet the large bulk of PACER fees are actually coming from lawyers representing corporations and well-off individuals.
You have it exactly backwards here. There is no need to justify stopping the flow of money but for the flow of money to exist in the first place. There is no right to a business model for such a thing can only end in abject insanity of demanding that it rain so you can sell your umbrellas.
>I’m not sure what you mean by “normalized” in this context
[Database normalization](https://en.wikipedia.org/wiki/Database_normalization) means ensuring that the data stays consistent across many sources and avoids redundancy. Issues like "There are two postings one for a Allan Smith in New York Court ABC Room A at 3pm on January 2nd, 1993 and one for Alan Smith in New York Court ABC Room A at 3pm on January 2nd, 1993 - which is correct?"
>but Lexis and Westlaw do a ton of additional analysis on top of public cases at their own expense. That information isn’t in PACER.
Then they have their own valid business model to sell their derivative works to supplement the public data that anyone else can get for free.
No, you’re demanding the government give umbrellas away for free because people have a right to be protected from the rain. The government isn’t stopping you from distributing documents you got on PACER. It’s charging you for access to PACER itself.
> Issues like "There are two postings one for a Allan Smith in New York Court ABC Room A at 3pm on January 2nd, 1993 and one for Alan Smith in New York Court ABC Room A at 3pm on January 2nd, 1993 - which is correct?"
PACER just stores free-form PDF files. There is some metadata, but the substance of your example above would be in the PDF.
> Then they have their own valid business model to sell their derivative works to supplement the public data that anyone else can get for free.
That’s what they do. The original court opinions are free on PACER.
Personally, I have no idea whether it's possible to do meaningful data analysis on PACER data, so I'm not necessarily arguing against your original point. But as a non-lawyer who sometimes gets interested in particular court cases, I find PACER fees obnoxious. Partly because the fees force me to think about whether I actually need a particular piece of information, which is so alien to the normal way I browse anything on the Web. And partly because the fees are the reason I have to actually deal with PACER's 90's-style interface: if all PACER data were free, I'm sure there would be at least one free website mirroring that data with a nicer interface.
Gathering the data of trials and judgement can prove when something isn't working and prove how things are actually working in the field better than even the best rhetorician. The question of 'should we be doing something' is separate from data.
But being able to get the cases together and point out that 'minimum/maximum sentences are constraining judges and juries since in 70% of the cases they note that they would go with lower/higher but they are constrained from it'. It can point out real problems with inequality in execution or corruption. Civic hacking is meant for accountability essentially to show how the system is or isn't working. Which is part of why resistance to it is so worrying.
Don’t confuse my opposition to techno-optimism with opposition to accountability. The issue is not that we shouldn’t have accountability. It’s that there is already a ton of data out there, and civic hackers have had approximately zero impact with it. Because for the most part what you actually need is carefully controlled studies from real institutions, not “civic hacking.” Civic hacking is the techie version of “raising awareness.”
Reforms are sorely needed. Imagine google charging you (and your lawyer) $.10 per search plus $.10 per returned result.
For me (not a lawyer), my PACER usage never reaches the threshold where I get sent a bill. I can go research a case I'm interested in, download the relevant filings, read them over, search for related materials, and not get billed. If I really got into a case or couple of cases, maybe I'd get a bill that is less than an average hardcover book. Meanwhile, I've have spent some multiple of that amount on my own time.
I'd much rather a system like this than an ad-supported social network for legal filings.
My larger beef with PACER is the insanely poor user interface. However, it gets the job done.
Edit to add afterwards: Despite the above, I fully support making PACER taxpayer-funded and free-to-access.
I’ve seen legal professionals not limit their queries, downloading 50 pages at $.10 each, when they’ve already downloaded the document umpteen times and only want to see the first or last page.
That also goes out the window when you want to read a 283-page filing including attachments.
Audio recordings (free) are starting to be made public at the 10th circuit. Looking forward to wider availability of them nationwide.
PACER has a $3 cap per document (no matter how many pages).
March of last year I got billed $12.60 for a 126-page transcript.
Not sure where you're getting your information from.
"Please note that there is a 30-page cap for documents and case-specific reports (e.g., docket report, creditor listing, claims register). You will not be charged more than $3 when you access documents or case-specific reports that are more than 30 pages. However, the 30-page cap does not apply to name search results, lists of cases, or transcripts (when available online)."
1) That PACER charges for people to access "the law." As the article points out, "judicial opinions are free" on PACER. They're also typically published as PDFs on the courts' websites.
2) That the "costs of storing and transmitting data have plunged, approaching zero." That may be true, but the cost of maintaining and upgrading an electronic filing and access system have not approached zero. And there is no free off-the-shelf system that does what PACER does.
3) That PACER is meant for access to public records. PACER is a read-only view of the courts' filing system for lawyers and the court to use. Because lawyers use it, lawyers pay for it. (Individuals representing themselves pro se can use it for free.) It may indeed be desirable to make these legal filings publicly available through some system. But PACER is not that system.
The issue is not really that "PACER should be free," because it's quite reasonable to charge lawyers a fee for a service they use that's provided by the courts. What you really want is something like the various "Open Government" searchable data sites that are out there. PACER is not that system--it's not easy for the public to use, it's search features are primitive, there is no way to do bulk downloads, it's running on the same systems as mission-critical ECF filing functions, etc.
If the mechanism for providing public access to court records is inefficient or poorly-done, let's fix that. Saying "we have to charge for PACER because PACER isn't the right way to do what PACER claims to do" doesn't seem valid to me.
Charge attorneys and other filing groups to use CMECF (the filing side of PACER) and make public access to download documents available for free. Or, make a bulk feed and database dump of documents available to download, like Wikipedia does, and let a group dedicated to open records access take up the charge. I'd rather have the government provide the search service but if it can't or won't, make the data available for free and allow for the opportunity for someone else to take up the charge.
I think in a common law system like ours, opinions must be available to the public as a matter of political morality. I don’t think that same moral justification extends to the other filings in cases, which don’t have the force of law. I think you need some utilitarian justification for incurring the cost of making these more widely available, and I don’t see a compelling one.
Any democracy deserves that, even civil law systems.
The law was constructed in the interest of the public and self representation was deemed reasonable for justice to be always administered when a wrong has been committed and a person wants to seek a verdict.
Well if you're a person who has low finances and in a situation where no lawyer will take your case by interest or financial reasons. You can still file a lawsuit and have the filing fees waived. The problem will be how much it costs to print copies for yourself, the number of defendants and the judge (court). People underestimate how much is needed to print if you're unable to electronically file. Lawyers have access to the online filing system but pro se doesn't necessarily (depending on the state) and has to file in paper with mailing everything out.
Further more, the defendants may be a corporate entity with expensive lawyers on retainers, who have full access to every previous court verdict (similar to the situation) on PACER. They likely have a local database of the downloaded documents and where they can do advance searches. On the other hand the pro se is restricted by PACER in fear of the financial bill when searching and viewing documents to find something similar to his/her own situation.
[0] https://ec.europa.eu/digital-single-market/en/protection-dat...
> However, no other initiatives are anticipated to make the justice system more public. The judge is, for example, not too keen on the idea of publishing all judgments online. “That would require far too much effort to make things anonymous,” he explains. But the court does already publish the most interesting judgments.
But the government here is also not big on paying for things. Building a whole electronic system to disseminate these filings costs a lot of money. Philosophically there is a strong trend in the US that having a right to do something doesn’t mean the government has to pay to help you exercise that right.
Wasn't there an old legal precept that that law needs to be public? (Even Draco saw that as a requirement.) And judicial opinions are part of the law in our system with precedent.
We just need a lawyer to take the time to point out that cost is negligible and the hosting costs would be nothing if they would do something like torrent the legal records :)
The Federal Courts argue that PACER is just "one avenue" of making the information public and that people can access the information for free by going to any Federal Courthouse.
I might buy that argument in 1999, but not in 2019.
In effect it ends up being a parasite on the system akin to corruption or literal highway robbery and it would have been far more sensible to just leave it free and do taxation and budgeting centrally.
It is akin to why bandits were treated with such harshness historically compared to other thieves of almost always being a capital crime. They did damage to the flow of goods and effectively 'broke the roads' because people feared to travel. That was a bad thing for everyone including the people in charge of force of arms so they dealt with the problem using the tool they knew best - violence.
Also, I'd imagine (hope) Google wasn't _the_ source, but instead a better interface/search engine over the data.
Random side note, Bloomberg Law is the only service I am aware of that allows you to do a pretty complete keyword search of PACER (at least for federal district courts, state courts vary). I'm pretty sure Westlaw and Lexis do not do this (also they charge an exorbitant price to do a search, while Bloomberg charges the PACER fee for the first time a document is pulled by anyone plus the seat the license fee). Docket Navigator is a pretty good product too.
The scholar team is small. Anurag[1] thought there was no reason law shouldn't be accessible to normal people too. So we pushed on that direction (I lead an eng team in DC at the time that worked on opening up data that should have been open. We also did election information, etc).
Once PACER/et all turned them down, i'm pretty sure they made some deals but there really wasn't a good and complete source.
Worse, lots of states/etc had locked themselves into exclusive deals and so couldn't give us the data if they wanted to. (They were actually happy to be locked in, it turns out).
They do have some fairly good feeds from paid sources but ...
Anurag is very persistent, but even here, i think he's been focusing on other areas that are more useful to people.
[1] https://www.wired.com/2014/10/the-gentleman-who-made-scholar...
https://blog.archive.org/2017/02/13/internet-archive-offers-...
Mainly because they make hundreds of millions on it, and this would stop them from doing that.
(and since people seem to have crazy views - no, google did not expect to make any money by doing this)