Full text search on 400M US court cases
judyrecords.com
judyrecords.com
in many jurisdictions, sealing a record is the equivalent of destroying it.
your crimes will haunt you forever because the system never forgets, meanwhile they simply go back to business like it never happened
ref https://www.google.com/amp/s/www.nytimes.com/2018/03/18/nyre...
I can see why it might be surprising to find some results when searching. The same data has already been available in many other databases that have existed long before this one and in those described on the info page as well.
A search for "minor consuming" reveals a few hundred thousand cases against children. I'm a little surprised to see that.
I’d imagine there are a fair few people on that list who’s crimes are for things that are now legal.
Why do taxpayers have to pay expensive court proceedings, and offenders have to spend a lot of money for an attorney, and waste a bunch of time.
They don't have to. The usual process for something like a speeding violation in the US is:
1. You get stopped for speeding (or get caught on a speedcam).
2. You receive the ticket in the mail.
3. At this point, you have an option to agree with it and pay the fine OR appear in court and hope they will rule in your favor (which could easily happen if you genuinely believe they were wrong; and half the time, the cop himself will fail to appear in court anyway, so you get the ticket dismissed if it wasn't anything too wild).
You don't have to appear in court (if you choose to accept the ticket and pay the fine) or waste money on attorneys (if you choose to contest the ticket in court). You can literally play it the same way in the US as you just described, by getting the ticket in your mailbox and paying it off (for something like speeding). That's it. Contesting the ticket in court is just another option available to you.
What happened to the parent commenter who mentioned failing to appear in court, they basically didn't pay the ticket they received (aka ignored it) and didn't show up in court to contest it either. That's pretty much it.
There are lots that are amusing, with most names or insults getting a hit. It’s seems the ‘AKA’ field has all sorts of entries.
There are many other public records databases that have similar data, including the federdal judiciary and many state courts across the country.
Some are listed on the info page: https://www.judyrecords.com/info
Sorry, I'm just curious.
It says MySQL 8 and Elasticsearch 7.8. I don't have much experience in elasticsearch, I wanted to know how does elasticsearch makes it faster? Is it like an extension that makes it faster? Or Elasticsearch has its own data store that consumes data from the database and magically makes it faster?
Thanks.
[1]https://www.reddit.com/r/programming/comments/jg4rkv/how_a_s...
Does it work when the view is insanely big? [i am not an expert at DBs, so my vision of a big DB might amuse you, but let’s say I have millions of rows to assemble as a view].
I recently tried a 32M songs dataset [2] and it works great, so I’m on the lookout for larger datasets to benchmark with.
"PACER notwithstanding, CourtListener is the most powerful case law research tool available online — and in many ways is much more powerful."
This is based on CourtListener's 4 million+ written court opinions, which judyrecords has recently integrated. But you're right, CourtListener has more case law research features.
This is pretty neat as I have never seen the records, even though I requested them (out if curiosity) 10 years ago, only to be told they had been destroyed years prior.
Edit: I think I got there eventually - this is something you made?
Is it in a format that could be backed up by a community to protect? Seems like something folks in /r/datahoarder would be interested in backing up.
That's 1024 * 15 * 439,000,000 = 6.7TB roughly.
The cases are all compressed, so I'm not using 6.7TB non-compressed for cases. But there are other request and non-request related records needed too. Just my backups currently.
https://www.reddit.com/r/programming/comments/jg4rkv/comment...
Some sites crash from the page views, and here I have to handle everyone searching 400 million documents too.
Looks like a place I would like to live in.
The one I quoted seems to be some kind of test case?
I am no fan of this at all.
That leaves me with the distinct impression that they're monetizing data about visitors and searches in some horrible way. (Data targeting for mugshot shakedown operations?)
I'm not going near this.
https://www.politico.com/magazine/story/2019/03/20/pacer-cou...
As for major privacy concerns, it is generally the more major crimes that have the larger issue with the victim being known. Knowing that some one was the victim of mischief vandalism is far less a privacy invasion than knowing they were the victim of sexual assault of a child (and even hiding the victim's identity often doesn't do more than hide the name from a passive search).
Then there are the benefits that other posters have raised, such as being useful for knowing past decisions used even in minor trials.
If you look at the info page there is a specific example about how to look up codes of cases that had the same charge.
Being able to see how other offenders are sentenced is useful to make sure people are being treated fairly. Lawyers use this kind of data up to the point of producing analytics from data like that to understand outcomes. Major legal data companies have a large segment of business doing analytics for lawyers handling high and lower level cases.
Here are a few related links: https://cluesearch.org/ https://measuresforjustice.org/
I do have a different cultural background so it's probably natural this feel horrible. Everything about this site would be so illegal in my home country it's almost hilarious in comparison. I'm used to (and fully approve of) a law that you can't keep a list of names in a notebook without a proper reason and everyone's consent, that would already be an illegal register.
"The Court finds that New Century had access to the SQL Data [pg. 536] Structures and that there is enough probative similarity to find that New Century factually copied the SQL Data Structures."
The next question might be to have 'Positive Software' demonstrate that they did not, in fact, take their table schemas from some place else. Like... textbooks? Or... example database schemas from vendors. Or tutorial sites? Or competing products?
There may be something extremely unique about part of their structure, perhaps, but... at the same time, there's often very little variety in how most similar data (crm/sales/lead gen/etc) might be stored to be remotely usable for reportin anyway.
"misappropriation of confidential information". Without seeing the structures in question it may be hard to say, but typically 'confidential info' is qualified with "not elsewhere available"-style clauses.
"... Likewise, the Court finds that there are more than one or a few ways to organize the data structures required for programs such as LoanTrack and LoanForce..."
Yeah, but usually there's only one good way to do stuff. Yes I could just have one row with 940 columns - technically, I could make my program work with that - but it's extremely suboptimal - regardless of whether I've seen anyone else's table structures or not.
Page 1 of 1,763 total cases for: donald j. trump Page 1 of 2,299 total cases for: donald trump
This is a proximity search, to ensure it's actually turning up one of the various permutations of the name (as different court protocols may refer by surname first), rather than documents that just happen to contain each of the terms somewhere.
For fairness, "hillary rodham clinton"~4 turns up 193 cases.
Relevant doc: https://www.judyrecords.com/info (down the page, under "proximity search")
#!/bin/sh
# usage: 1.sh [query] -- perform search
# usage: 1.sh -- process results page 1 to 200
# usage: n=5 1.sh -- process results page 5 to 200
# usage: n=201 1.sh -- quit
# start tmux if not already running then detach
j=https://www.judyrecords.com;
case $# in
1)
tmux set set-remain-on-exit on;
tmux neww links;
tmux send g $j/addSearchJob?search="$@" c-m;
sleep 1.5;
tmux send d;
tmux capturep -p|sed -n /./p;
tmux send g $j/getSearchJobStatus c-m;
sleep 1.5;
tmux send d;
tmux capture -p|sed -n /./p;
;;0)
test $n||n=1;while true;do test $n -le 200||break;
tmux send g $j/getSearchResults?page=$n c-m;
sleep 2;
tmux send Down Down '\'
# small monitor where results page HTML takes 4 spacebar presses to get to bottom;
m=0;while true;do test $m -le 3||break;
# process results -- e.g., print record URLs;
tmux capturep -p|sed -n "/href=..record/{s|.*record.|$j/record/|;s/\"//;p;}";
tmux send Space;
m=$((m+1));done;n=$((n+1));done;
tmux killw;
esac #!/bin/sh
j=https://www.judyrecords.com;
case $# in
1)
tmux new -P -d links;
tmux set set-remain-on-exit on;
tmux send g;
tmux send $j/addSearchJob?search="$@";
tmux send c-m;
sleep 1.7;
tmux send d;
tmux capturep -p|sed -n /./p;
tmux send g;
tmux send $j/getSearchJobStatus;
tmux send c-m;
sleep 1.7;
tmux send d;
tmux capture -p|sed -n /./p;
;;0)
test $n||n=1;while true;do test $n -le 200||break;
tmux send Down
tmux send Down
tmux send g
tmux send $j/getSearchResults?page=$n
tmux send c-m;
sleep 2;
tmux send Down
tmux send Down
tmux send Escape;
tmux send F
tmux send v
tmux send c-u
tmux send 1.htm
tmux send c-m
tmux send o
sed -n "/href=\"\/record/{s,.*record\/,$j/record/,;s,\",,;p;}" 1.htm;
__grepq=$(exec sed -n '/a class=\"goToNextPage/!d;=;q' 1.htm);
test ${#__grepq} -gt 0||break;
n=$((n+1));
done;
tmux killw ;
esacI know this was always public but this makes it too easy for masses to dig through the troves.
Scares me. Next thing I see is some AR glasses that do facial recognition and correlate name -> public records. Could be a nasty blackmail tactic. Some things are close to Black mirror in reality.
> MySQL 8 is used for DB. The seach server uses elasticsearch 7.8.
For reference, I once threw the entirity of open streetmaps at it before it even hit 1.0 to implement a simple reveres geocoding thing. Basically a couple hundred million street segments, some polygons, etc. At the time the geospatial support wasn't great and very new and very CPU intensive. I got away with indexing all of that and running it on a single node cluster with a xeon and 32G of RAM and spinning disk (RAID 1, no SSD). It worked great. Very responsive. Indexing only took about 50 minutes or so. Most of that was my parsing logic. That's not comparable of course, I'd expect this to be faster on the same hardware with a current version of Elasticsearch. They've made a lot of leaps with improving performance, memory usage, cpu usage, disk usage, robustness, etc. in the 7 major versions since then.
When I type a query and press search, would like it if the URL updated with the search in the query string. It would make it easier to share specific queries.
Also, by "case" do you mean "opinions"?
Full disclosure, I've written and contributed to several scrapers for CL, and if there's a large source they're missing I'd like to know.
Note that the CL opinion number you're quoting doesn't include orders from Federal courts that are in the RECAP collection, which accounts for several million additional opinions.
their searches are indexed and have rulings and documents as well.
does this differ from that service?
Do you have any long term plan for the site? I can see this going in a lot of different directions depending on your goals.
Does this have the lower trial court records too?
Hard to read on mobile though
so do I get a court order for each county, the website, the resyndicating source that the website uses or what?
I looked at the reddit page and other people noticed the same thing, the author just said send me the link! Hahaha one by one removal maybe!
Shut it down, enjoy it while it lasts
Should just do a search for expunged or similar terms and remove those entries.
You want secret courts?
Edit: wow, plus family court stuff like a four year custody dispute, kids being adopted, etc
secret courts have cases that are secret from the beginning