Do the police give more tickets at the end of the month to meet quotas?
robert.io
robert.io
1. Do police give more tickets at the end of the month? 2. If they do, is it because of quotas?
The first is a statistical question, the second is much more complex. I can think of a few very good reasons why numbers would spike at the beginning of the month, plummet in the fat middle then spike at the end. None of them really have much to do with statistics directly.
I think this would be compelling if you'd compare stats for years in which "quotas" (whatever they decide to call them) were policy against the stats from years which they weren't.
I didn't actually try to find information on quota policies (specifically for Baltimore and Maryland) until I was writing the analysis section. I was expecting to be able to find a detailed description of policies by state (or something helpful) on Wikipedia, but that didn't really pan out [1]. I wasn't able to find any reliable information elsewhere either. I think the taboo attached to the subject may keep information about policies pretty far under-wraps.
Of course, all that does is prevent them from admitting so publicly. Anecdotal, but a friend of a friend claimed that he goes out some nights and has to write a certain amount of tickets.
Not so anecdotal, a collection of news stories related to Mass ticket quotas:
http://www.dailynewstranscript.com/news/x93766893 http://www.thenewspaper.com/news/04/453.asp http://www.boston.com/news/local/articles/2005/10/22/police_... http://www.thenewspaper.com/news/21/2130.asp http://www.motorists.org/ma/radar.html
Also, when you see enforcements such as "Click It or Ticket," basically the govt provides incentive funds if certain seatbelt tickets quotas are met.
The most useful piece of advice, though, was on how to pull over. Turns out that every time a cop pulls somebody over, he's worried about getting smooshed by traffic or shot by a loon. His advice, which I've followed faithfully, is to pull waaaay over, so the cop can park close to traffic and create a bubble of protected space. Then you turn on the interior lights, roll down the windows, turn off the car, throw the keys on the dash, and hold the steering wheel at the top with both hands.
Authority figures like respect, and this display of considerate compliance has worked wonders for me.
For a given month, you can determine the last day. You can also determine the last weekday.
So just do a series that for each month shows "average weekday citations, non automatic" and "final weekday citations, non automatic" (and maybe "first weekday citations, non automatic").
The right way to solve this problem is to find the exact thing you want - the precise last weekday of the month, the precise first weekday, and a precise weekday average - which is trivial with any decent date/time library. These would be the RIGHT kind of gymnastics. Tallying information by day and then running some kind of regression to approximate last-day traffic is a waste when you can just GET the last-day traffic.
Looking at the last plot, you can't conclude that police give out more tickets at the end of the month, unless you also conclude that police like to give out tickets around the 8th, and don't like to give out tickets around the 11th of each month. Problem is, there are no error bars, if there were, it'd probably show that there is no significant evidence for the hypothesis. The error of the mean of the last plot is probably comparable to the signal itself.
I did make an update with a few, more conventional statistics as well as an adjustment based on day of the week (see "Updates"): http://robert.io/posts/4.html
> as well as for the dates that only happen in a few months out of the year
That actually was compensated for. That's what the "Correcting for more frequent dates" section was about.
If we make the assumption that it doesn't matter, and the quotas are 'rolling', then surely the same could be said about the monthly quota not ending at the actual end of the month?
Also, the 7th, 14th, 21st, and 28th days are all above average.
Another thing they can look at with this data is where they are giving out tickets. If they notice one officer has all his tickets in one location which is known for high accident rates then they know that person is "fishing" rather than actually doing active patroling.
I guess my take-away from that was, cops can give out tickets in short order if people are consistently breaking the law right in front of them. I wouldn't necessarily punish an officer for giving out many tickets in a high-offending area, just based on what I saw that day. Sometimes people in an area need to have the rules reinforced. For every driver that got pulled over, 20 other drivers were witness.
If those people were driving safely, then ticketing them for a technical violation of the law probably isn't helping anything. The point of tickets is to punish the unsafe drivers so that they learn a lesson and have a record that makes it easier to reform or weed out the really bad ones.
So a cop that spends all day writing tickets on technical violations probably should get a talking to from their superiors. They're wasting their time and ticking off solid citizens while actual unsafe behavior gets a pass.
But I think the data is a little more conclusively inconclusive than a mere "I'm not sure I can tell anything from the data".
[1] https://data.baltimorecity.gov/Financial/Parking-Citations-2...
Maybe, there's a possibility of increased ticket rates too due to the extra driving km done by people "going away for the long weekend", and by drivers being less experienced on the routes they're driving (compared to regular commuter traffic who know exactly where the speed/redlight cameras are, and are usually in heavy enough traffic to not be able to exceed the speed limit).
I wonder if there's local (NSW, Australia - for me) tickets-by-day day available to analyse?
maybe a better question would be, for ever hour worked, how many ticks are issued. aka normalizing against the number of patrol men/women active for a day or an hour.
e.g. more police on patrol == more tickets, but does not mean each police man/woman is biased to issue more tickets near the end of the month to meet their quota.
also consider counting backwards. e.g end of the month = 0, 1 day before end of month = 1, 2 days before end of month = 2 etc..
my 2cents.
1. It looks like the total number of tickets in 2009 and 2010 is about 10% that of 2011. I'm guessing that there weren't actually ten times as many tickets given in 2011, so either the data is incomplete (as the author suggested), or there was a typo. If the data is incomplete, I'd suggest normalizing to the 2011 totals; otherwise, the 3-year average doesn't make much sense.
2. The scale of the "normalized" difference graphs (showing "Actual - Expected"). The formula given is
(actual - expected) / total * 1000 = normalized number
If this is the case, then since the scale goes to about +/- 5, the differences are very small (less than 1% away from what you'd expect!). But from eyeballing the data, that doesn't seem right.
In any case, a better scale might be to expect the data to be normally distributed, and scale the differences to # of standard deviations. (See, e.g., http://en.wikipedia.org/wiki/Normal_distribution#Standard_de...)
2) The fact that the normalized numbers were so small was very unintuitive to me at first too, but the important thing to realize is that in that formula, you're dividing the difference, not actual value for the given day, by the total number for the year. When I first ran those numbers I was so confused by the output. I was originally thinking that I'd normalize it by saying "X percent of the total for that year," but since I was working with the differences, and not the actual values, the numbers were too small a fraction.
Either that or I made some huge mistake in my logic...
WRT the use of standard deviations, like I said in the post, I'm not a statistician, so I wasn't really sure what the canonical way of normalizing data was. I pretty much just made one up. Thanks for pointing that out. I'll look into using standard deviation for the next one. :)
Even better than error bars would be box plots of the data corresponding to each day of the month. You can easily produce such plots with R.
The 27th appears in every month, but the time remaining to meet your (possible) quota would be different for different months.
I might do an update with some of the suggestions here, including a graph of tickets / days until end of month. [1]
They'll just task one to follow you until you make a ticketable driving mistake.
I.e. n[i] should be sum over all days d of tickets[d] where d is the i-th day from the end of the month that d occurs in.
Also, re quotas, from my cop buddy (constable) "Of course there are no quotas, that's ridiculous. What they do, is tell you to attend the meeting of the commissioner's court; during which the commissioners bemoan the budget woes and discuss which county employees may have to be laid off if things don't improve, etc. etc." He said that usually inspires the troops, and nobody ever says the word "quota."
the way an actual statistician would try to answer this question would be to first describe a "null" hypothesis, called H0, then talk about the probability of seeing this set results, or an even more extreme set, if H0 were in fact the case.
H0: The police do not change their ticket collection strategies at the end of the month
H1: They try to collect more tickets at the end of the month than the other days of the month
H0 would describe some distribution for the numbers you are seeing. This is where you can put all of your assumptions about how things are - so that you ultimately end up with a model that generates a certain distribution.
Under this model, there's going to be a certain probability of seeing this result, or a more extreme result.
If that probability is less than, say, 5%, then this is saying that you'd only have a 5% chance of seeing these numbers given that there isn't in fact any conspiracy to collect more tickets.
In such a case, you might then "reject" the null hypothesis in favour of H1.
If the probability was higher than 5%, you might say that there is no "significant" evidence in favour of H1.
I wonder if ticket writing happens when there's nothing better to do. It would be interesting to see the correlation between arrests for "real" stuff and traffic tickets. Are there differences in pay rates based on days/times worked? Do more tickets get written when rates are low vs. 2.5x pay on a holiday? Does the PD only staff for necessities when rates are higher or when officers most desire time off?
I'm surprised I haven't read something yet about Stephen Levitt looking at this data. This provides such incredible insight into the impact of incentives and lets us compare the probability that the official incentive policy is perceived as accurate by the officers vs. the likely real but unwritten incentives.
The correct approach would be to search for patterns over time in the dataset (frequency analysis) and see what turns up.
not arbitrary segment the data into week sized blocks because you 'think' there might be a data pattern in there.
I see this sort of 'inspiration based' analysis in web analytics all the time, and it's complete nonsense.
Look for patterns in the data, don't look at the world and try to fit the data to it. You'll end up with stupid and statistically invalid results.
By looking at established segments, at least you know where to look for things that MAY make sense, and then derive hypothesis.
If not, you end with bullcrap analysis we see every single day in literature where researchers find a "link between vegetarians and personality disorder" or things like that, without even probing the fact such links may be coincidental at best.
Your data is now composed of the original set of data [n1, n2, n3... nn] + a NEW DATA SET [s1, s2, s3 ... sn] which is your selection criteria.
This is often masked by the fact that your selection data set is a function f(x) ("we'll just split them up by week and partition by gender"), so it doesn't look like your actually combining two dataset, but you are.
This new set of data S, biases the result of the analysis, and the more complex it is, the closer you are to over-fitting to find the 'golden correlation' you're looking for in your data.
Humans are pattern matching machines. We see patterns in random noise, and hear voices in random tones. It's easy for us to see patterns (or what we think are patterns) in data sets, but its much harder to statistically substantiate those patterns as distinct from randomness.
It's provable, I imagine, that any data pattern can be found in a data set if we have sufficiently complex selection criteria for the samples, and sufficiently random raw data.
dont do this
If you do, at the very least acknowledge that you have selectively modified the original dataset and specifically test your partitioning rules.
Its much, much better to search the dataset for a target pattern, and then investigate the partitioning rules that your search matches against (because you can examine the partitioning rules for complexity, and data-content and easily spot over-fitting).
Maybe on the 28th or so they panic and start issuing tickets. And by the 31st they've made their quota and relax. They don't want to wait to the last day and risk it.
Or maybe these results mean nothing (more likely)!
http://skeptics.stackexchange.com/questions/9578/do-police-o....
In some areas, this has been confirmed.
Could this be true?
The following graphs are based on variation from that expectation.