HNHacker News
TopNewBestAskShowJobs

apwheele

1,899 karma · joined November 10, 2021

Data scientist. Former academic in criminal justice field.

Personal blog at https://andrewpwheeler.com/

Consulting at https://crimede-coder.com/

Large Language Models for Mortals: A Practical Guide for Analysts (book), https://crimede-coder.com/blogposts/2026/LLMsForMortals

submissionscomments
apwheele··on Jev Can't Be Calibrated
While this is true, it is still possible to take the probabilities and make determinations of recall and false positives using conformal sets, https://crimede-coder.com/blogposts/2026/ConfClassification

(Or just use a model to re-calibrate the probabilities, I like the conformal approach though as those rates are what I often care about.)

You just need some labelled data to generate the "corrected" probabilities (or thresholds to meet the specified error rates).

apwheele··on Flock Wants a Closely Surveilled World with No Exit
They are relevant and courts have considered the whole of a persons movements standard in terms of whether querying ALPR data constitutes a search, https://andrewpwheeler.com/2026/08/12/license-plate-reader-s...

Under current case law (Carpenter and recently affirmed in Chatrie) it will definitely be a search, IMO it is just when the sensors become dense enough according to the court (absent states do not make regulation themselves to require a warrant for historical data).

apwheele··on Hallucinated Citations and Accuracy of Papers
I sent it to the moderation email this morning. Will update if I hear back.
apwheele··on Your AGENTS.md file doesn't do anything
This is what the majority of mine look like as well. Just simple instructions for things I need to do repeatedly.

I have not even really needed a formal memory system. If I see an error happen more than once, I just say "hey add a note on this to agents.md". Tends to be verbose but overall works quite well for the projects I am doing.

apwheele··on Gemini Omni 1.1 Flash
Is there a good example someone can show of naturally generating real videos using extending and interpolation from the last frames? (For any of the video generators).

I mean I have not tinkered that much, but trying to even get a video to 30 seconds (I just want a cartoon AI avatar to narrate tutorials) has been incredibly difficult. They drift so easily.

Many AI videos you can tell just stitch short clips together. I just want a continuous scene for like 30 seconds to 2 minutes.

apwheele··on Don't classify, hallucinate!
This is another riff on not embedding a full document, but doing a summarization of the document and embedding the summary for RAG. Nice usecase for high cardinality data!
apwheele··on Flock updates privacy, accountability, security, and transparency safeguards
I did not know about this when I wrote https://andrewpwheeler.com/2026/08/12/license-plate-reader-s... the day prior, but one of the things I did state was deleting data (except in the extreme case, like New Hampshire's 3 minute retention), will not prevent abuse. People just do illegal queries over and over again.

Part of the reason I wrote that post is it is better to retain the data indefinitely and have a warrant standard than continue to reduce the time period in which the data is retained.

apwheele··on License plate reader searches should require a warrant
This is already the current status quo on paper. I bet the majority of the cases IJ identified of mis-use the PDs already had this policy. (That is the same standard when you do a background criminal history check in New York, need to put in a case number, FYI.)

I did not want to go into too many technical details on the post -- I think you could do an AI audit at the point in time of the search query to prevent people putting in junk cases. That said I believe there will always need to be third party audits at a minimum.

Part of the issue is policies on paper are not effectively enforced. So people saying "just have policy X to prevent abuse" is a non-answer.

I wrote the post mainly because people arguing for limited data retention I think is bad for both sides -- it neither protects civil liberties and simultaneously makes it harder for long term criminal investigations.

apwheele··on New Orleans is testing Carbyne’s AI-powered Emergency Call Triage software
I cannot read this article, but this may be partly confused what this function by Carbyne does. If there is a flood of critical systems, it just plays an auto message "If this is about X, thank you and we know about it" if you called from the same area.

I don't remember from prior articles if at that point you just stay on the line, or if you need to press a button. It is not going through an AI voice call tree.

apwheele··on Why Large Language Models Fail at Tabular Prediction
My money is on they used agentic AI coding tools to help set up all the metrics/experiments with traditional models.
apwheele··on That time when I failed the Microsoft interview
My experience like this was for a data scientist position at Facebook in 2022.

So the question was something like, here is a table of predicted probabilities for credit card transactions (hypothetical table, I do not remember the exact quantities), but it looked something like this:

    Amount, Probability, ActualFraud
       $5 ,    0.2     ,  0
     $100 ,    0.5     ,  0
       $3 ,    0.7     ,  1
       $1 ,    0.9     ,  1
And then the question was "what was the threshold used to maximize (precision|recall)?" (I cannot remember which metric they were asking for).

My response was that a proper decision model should take the dollar amount of the transaction into consideration, so the decision should more like be `$*p > threshold`, where this threshold meets your precision and recall for preventing dollars lost instead of the binary yes/no.

The interviewers English was not great, so after all that I was just like "well to answer your exact question, it could be anywhere between `0.5 < p <= 0.7`"

Did not make the second round!

apwheele··on Ask HN: What are you working on? (August 2026)
An app to verify citations in peer review journals, https://veruscite-data.com

that is, to identify when papers have hallucinated references

can see an example of results at https://veruscite-data.com/share/3Pu3dwCfc-ixPgtHutVFXgMU7FY... for a preview. Also can sign up and get two free reviews.

apwheele··on I flagged two research papers for fake authors and both were accepted as orals
It is in alpha (hoping to do Show HN in a few weeks), but for those interested I am working on an application to do this, https://veruscite-data.com/

Most of the folks on HN will be more familiar with genAI tools and can just use the skill Caleb and Isaac provided in this blog post. My tool is just likely more token efficient and has a GUI where you can review the extracted bib and edit it more easily.

apwheele··on Don't ask an LLM for a confidence score
So I agree with the general geist of this, a few counter-examples though:

If you have the raw log-probs, I show how to use conformal inference to set false-positive or recall rates, https://crimede-coder.com/blogposts/2026/ConfClassification

Some of the peer reviewed papers with the older models did show calibration was bad (not close to monotonic). This post with newer models (for one example, classifiying injuries) is not that bad, https://gmcirco.github.io/blog/posts/ai-calibration/calibrat....

I mean it just depends on the application, what level of error you can consider. But the second post shows how to recalibrate the scores as well if you need calibrated probabilities.

apwheele··on Are AI Labs Pelicanmaxxing?
If you look at the raster image ChatGPT generated, that is fine. It is just this example (and other simple SVG icons I have asked for) result in pretty bad SVGs. It just makes me highly suspicious that the LLMs are learning shape primitives and extrapolating to new shapes, vs just having a big dictionary of prior examples and stitching them together.
apwheele··on Are AI labs pelicanmaxxing?
So this is not my experience at all for asking about simple SVG icons for web-pages. Here is one of the examples I have tried for in the past, make a simple cartoon SVG knife for a map icon for a crime map.

https://x.com/CrimeDecoder/status/2080008114615537766

Can see the images for ChatGPT/Claude (Sonnet 5), and Gemini are all quite bad.

Jagged edge of LLMs. How do you explain being able to generate very complicated shapes in the Pelican example but cannot make a much simpler icon without just alluding to it is in the training data?

apwheele··on Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Not an open source, but I discuss it in my book with examples for OpenAI/Anthropic/Gemini, https://crimede-coder.com/blogposts/2026/LLMsForMortals.

All of the models, you need to have a consistent input to get the cache hit. So if you are chatting with a document, and change the system prompt, it will be a cache miss, even if the rest of the items are all the same. If you even pass in the document in not the same order as the prompts, it will be a cache miss. Or if you add tool calls or structured outputs, it will be a cache miss. (Since those generally go at the beginning of the prompt call, not at the end.)

Most of the time when reading documents from URLs directly it will never cache. (Need to typically pass in the bytes directly, or use the provider document store index.)

Gemini has a 4096 minimum token size with the 3 version models before even getting a cache hit. OpenAI it is lower (1024), and is automatic, but only happens in increments of 124. Anthropic can also get cache hits at 1024 tokens, but you need to explicit ask for it (and pay extra).

Caching by default typically lives for 5 minutes since the last cache hit across providers. But some of them you can ask for longer. AWS for Anthropic models can be tricky with multiple endpoint routing, so can get cache misses if it happens to route to a different endpoint.

apwheele··on Kimi K3, and what we can still learn from the pelican benchmark
Just as even a counterpoint to this, I have asked the LLMs to attempt to generate SVG icons for websites. Even though I have requested things much simpler than a Pelican, they have all tended to do quite poorly in my examples.

Because of this, I presume the Pelican has been in the training data for at least a year+.

The models are very useful, I am afraid they have fundamental limitations though generalizing (it is just hard to evaluate effectively). So it will just be whack-a-mole "can your model do X", and there will always be a new X.

apwheele··on Evidence of inconsistencies in evaluation process and selection of winners
I think this is a good meta-lesson for Kaggle. When you have objective metrics to hill-climb towards, AI can do quite well. When you just phone it in and rely on LLM as a Judge, the results are not so great.
apwheele··on GLM 5.2 and the coming AI margin collapse
Agree with the web search point. (I would like for Perplexity to start to offer more models out of the box integrated, like they do now with OpenAI/Gemini models.)
apwheele··on Pruning RAG context down to what the answer actually needs
Processing 500 retrieved chunks here (probably with an additional LLM) will also add a bit of latency (not shown in these graphs!).

So for user facing apps, that scenario is probably not feasible (more like filter 10 chunks). Which as the parent of this comment suggests is fine to add in extra context given the current size of context windows.

apwheele··on The Hitchhiker's Guide to Agentic AI
If I may shamelessly suggest my own book for those looking for a more basic intro to calling the APIs, including a chapter devoted to agentic patterns across the different SDKs (openai, claude, and google), https://crimede-coder.com/blogposts/2026/LLMsForMortals.

Basically I wanted to write a book that did not spend the majority of time on training LLM and architecture, as it is just not relevant to the majority of software engineers (either using agent coding tools to help write code, or using the LLM APIs as backend parts of apps you are building).

This Hitchhikers book has many examples of API use as well, but is very "here is a wall of code". Mine is a bit more gentle progressively building up simple examples -- e.g. I show how to write python code to call a tool in a loop. Then show how you can feed back the error and have the LLM update. Then show how you can do that without manually looping through the agents SDK.

apwheele··on SVGFarm – SVG icons for coding agents
Off topic, but Simon's Pelican on a bike SVG makes me think they are using that is the current models. I ask for much simpler SVG icons periodically (to insert into webpages, map icons, chart icons, etc.) and all of the models tend to do poorly in my experience.

Most recent example was a V to look like a checkmark. Seemingly much simpler than a pelican riding a bike, but Claude/ChatGPT/Google all did quite poorly.

apwheele··on CursorBench 3.1
I read these and think it is just the jagged edge. I do not doubt your personal experience, I have used Composer 2.5 (via Grok and the credits I get with my X premium account) the past month.

I am not building rockets, but have been quite impressed. All the models do dumb things sometimes, it has done the work I have asked it to pretty well though and has done to me some impressive work.

It is fast on Grok, for other models I have worked extensively with I think it is better than gemini 3.1 (3.5 and antigravity for me is worse than the prior gemini cli). And is comparable to Opus 4.6. (Have not used the more recent models in Claude Code.)

apwheele··on The early hiring funnel is now breaking on both ends
An AI written article writing about AI resumes and AI interviewers -- https://www.pangram.com/history/38790b9f-dda1-4ea9-9b6c-3858....
apwheele··on Map Clustering Is Not My Favorite
That is a good point (spiderfy is good if you want to show multiple icons and click on each one specifically).

If you do not like the cluster-marker, you could have an icon just list the total number at that location, and then on hover (or click for tooltip) show the table. And that may honestly be better than the spiderfy example for most applications now that I am thinking about it.

Maybe a better example, ESRI when you click a popup has a series of HTML, https://data-ral.opendata.arcgis.com/datasets/raleigh-police..., so if you click one of the grey circles it contains multiple points, you can basically go through different views (either all the elements or click through to a more detailed view of the individual point).

apwheele··on Map Clustering Is Not My Favorite
Checkout how leaflet does this, it is the "spiderfy" part of markercluster. So you click down into the number bubble, and at the lowest zoom explodes out into multiple point markers.
apwheele··on Map Clustering Is Not My Favorite
The clustered markers in leaflet are jarring (I like them, but when I show maps I make my wife she finds the transitions nauseating).

The default heatmaps for these maps are bad. Heatmaps should use filled contours so the gradations are more easily identified. (Continuous raster maps are blobby.) See the ascii glyph map in this post, https://andrewpwheeler.com/2015/06/12/favorite-maps-and-grap.... I think those should be static for various levels of zooms as well, and not recalibrated when zooming.

Another option (not shown here) is to just use polygons and aggregations, and when zoomed in can turn on that point layer (or just have it appear). Or can just make actual clusters (like DBSCAN).

I have a map I made on my website that shows these (with various interaction tooltips/hover), https://crimede-coder.com/graphs/DurhamHotspots (hotspots of crime in Durham, NC). And an explanation of the cartographic decisions and when to use the different techniques, https://www.youtube.com/watch?v=mBm6sTR08BI

apwheele··on Never talk to the police
Terry stops != traffic stop

You are required to provide a license to operate a vehicle.

apwheele··on Claude Corps
I mean some of the orgs could be like that, but the one near me I looked up (Durham, not affiliated to be clear), is more like an entry level gig to work on various projects (so you work for non-profit, but the orgs they work for may be public or private sector).

They mostly built phone apps oriented to public good projects. (So would just be using Claude Code to build the app itself, they wouldn't be calling Anthropic APIs behind the scenes, at least for those projects.)

Think $85k for an entry level gig subsidized by Anthropic. What is so bad about that?

Page 1 of 7Next →