159 karma · joined August 31, 2021
Also, the ANE only allows some operators to be ran on it right? There's very little transparency/control on what can be offloaded to it and cannot which makes using it difficult.
The worst part of it is as you say we all accept it and no one talks about it.
Is there any recommended reading you'd suggest to look into this more and the impacts of it?
Also I did run a test of FP16 vs FP32 for a large matmul on the Apple GPU and the FP16 calculation was 1.28x faster so it makes sense that they'd go for FP16 as a default.
Also generally I think CoreML isn't the best. The best solution for ORT would probably be to introduce a pure MPS provider (https://github.com/microsoft/onnxruntime/issues/21271), but given they've already bought into CoreML the effort may not be worth the reward for the core team. Which fair enough as it's a pretty mammoth task
You can find more details at my site soon: https://ym2132.github.io
ha, thats probably why I noticed the EyesOff accuracy drops so much at longer ranges, I suppose two models would do better but atm battery drain is a big issue.
I'm not sure if it's important or not, but the app comes from my own problems working in public so I'm happy to continue working on it. I do want to train and deploy an optimised model, something much smaller.
Sounds great, once a POC get's built I'll let you know and can see about the clinical side.
Thanks for the tips! I'll be sure to post something and reach out if I get round to implementing such a model.
Do you remember the cost of Mech Turk? It was something I wanted to use for EyesOff but never could get around the cost aspect.
I need some time to process everything you said, but the EyesOff model has pretty low accuracy at the moment. I'm sure some of these tidbits of info could help to improve the model, although my data is pretty messy in comparison. I had thought of doing more gaze tracking work for my model, but at long ranges it just breaks down completely (in my experience, happy to stand corrected if you're worked on that too).
Regarding the baby screener, I see how this approach could be very useful. If I get the time, I'll look into it a bit more and see what I can come up with. I'll let you know once I get round to it.
Is there any research/papers on this type of autism diagnosis tools for babies?
To your last point, yes I agree. Even the task I setup the model for is relatively easy compared to proper gaze tracking, I just rely on large datasets.
I suppose you could do it in the way you say and then from that gather data to eventually build out another model.
I'll for sure look into this, appreciate the idea sharing!
They create this CNN for exactly this task, autism diagnosis in children. I suppose this model would work for babies too.
Edit: ah I see your point, in the paper they diagnose autism with eye contact, but your point is a task closer to what my model does. It could definelty be augmented for such a task, we’d just need to improve the accuracy. The only issue I see is sourcing training data might be tricky, unless I partner with some institution researching this. If you know of anyone in this field I’d be happy to speak with them.
Any tips on improving accuracy? A lot of it might be due to lack of diverse images + labelling errors as I did it all manually.
Perhaps I have been jaded by the Mac webcam, I agree on most old webcams it wont be great but on newer webcams I have had success.
I did try a calibration approach but it's simply too fragile for in the wild deployment, calibration works great if you only care about one user but when you start looking at other people it doesn't work so well.
Good idea, it may be more fruitful to do that. At least then for the primary user we can be much more certain.
Privacy screens are still useful and I recommend people to use EyesOff and the screen protector. A privacy screen won't stop someone shoulder surfing from directly behind you etc.
There is also better ways to do this sort of task when all you care about is tracking the main user: https://arxiv.org/abs/2504.06237, https://pmc.ncbi.nlm.nih.gov/articles/PMC11019238/
It's a simple (currently macOS) application which aims to target shoulder surfing by using a locally running neural network to detect those looking at your screen.
Built using python the app alerts you to when someone is looking at your screen, using locally running deep learning models + your webcam.
I developed the app when I felt uncomfortable with working in public spaces, I wanted EyesOff to give an extra barrier of security to shoulder surfing scenarios
It uses the webcam and a locally running facial detection model to alert you if it detects someone in frame.
It's FOSS and available @ https://www.eyesoff.app
This is on the roadmap, along with gaze detection.
https://github.com/opencv/opencv_zoo/tree/main/models/face_r...
This model lets you upload a reference face and then matches to those in images. I think this allows for the "approved faces" function you mentioned.
I suppose a difficulty may arise when we run a few models at the same time, however there is probably a lot of room for efficiency on the table and thanks to the small model size we are already in a good place.
- https://arxiv.org/pdf/1912.04958 - StyleGAN2 - https://arxiv.org/pdf/1807.09341 - Causal InfoGAN, written in TinyGrad
Noted I am quite new to writing like this, where exactly was it confusing? I'll definitely take this on board and try to increase the clarity in my writing.