His work has uncovered a number of really tricky bugs in the multicore runtime but what's brilliant is the reports normally come with a minimal reproduction. This makes working out the cause so much easier.
Great work Jan.
5,054 karma · joined July 19, 2007
His work has uncovered a number of really tricky bugs in the multicore runtime but what's brilliant is the reports normally come with a minimal reproduction. This makes working out the cause so much easier.
Great work Jan.
If you want to go further you can export the GeoJSON and then run it through any machine learning pipeline you like.
https://github.com/ucam-eo/geotessera has an image showing our embedding coverage at the moment. Blue areas we have complete coverage for 2024, green areas we cover 2017-2024. We're slowly trying to populate everything 2017-2024 but the constraint is GPU and storage at the moment - each year takes ~20k GPU/200k CPU hours and requires storing and serving 200 terabytes of data. The world is big!
If there is an area you would like prioritised, there's an issue template on the geotessera github repo which we can use to move regions around in the processing queue.
The downside of that approach is that you need to spend valuable labels on learning the spatial feature extraction during training. To fix that we're working on building some pre-trained spatial feature extractors that you should only need to minimally fine-tune.
As I mentioned in one of the other comments, the model is also only pixel-wise. That is, it is not using spatial information for predictions.
For a proper evaluation you would need to be more methodological but as a sanity-check we were very happy with it.
One other thing to point out about the bramble model is that it is pixel-wise. That is each prediction is exclusively only what is within the 10 metre pixel (give or take the georeferencing error).
The easiest way to test is to try out the interactive notebook and drop some labels in known areas.
There is the issue of just how visible truffles are from space though, if they grow under cover. That said, it may still work because you can find habitats that are very likely to have truffles. We've had some promising results looking at fungal biomass.
You are very right on the temporal aspect though, that's what makes the representation so powerful. Crops grow and change colour or scatter patterns in distinct ways.
It's worth pointing out the model and training code is under an Apache2 license and the global embeddings are under a CC-BY-A. We have a python library that makes working with them pretty easy: https://github.com/ucam-eo/geotessera
We're hoping to try it with a few different things for our next field trip, maybe some that are much harder to find than brambles.
When it comes to the satellite images, the model actually used TESSERA (https://arxiv.org/abs/2506.20380) which is a model we trained to produce embeddings for every point on earth that encodes the temporal-spectral properties over a year.
Think of it like a compression of potentially fifty or a hundred observations of a particular point in earth down to a single 128 dimension vector.
Happy to answer any other questions.
$0.15M in / $0.6-0.75M out
edit: Now Cerebras too at 3,815 tps for $0.25M / $0.69M out.
I was looking at: https://arxiv.org/abs/2506.18254 but your approach is even more general.
An advanced extension to this is that there are algorithms which calculate the number of records to skip rather than doing a trial per record. This has a good write-up of them: https://richardstartin.github.io/posts/reservoir-sampling
I've been doing some work with colleagues at Cambridge and Imperial over the last year on using LLMs to improve evidence synthesis, primarily trying to find papers on the effectiveness of certain Conservation interventions. It's becoming clear that you really need to move beyond screening papers only by title and abstract - there's often information buried deep within papers that can only be found with access to full text. My colleague Anil Madhavapeddy has written a bit about our adventures in trying to ingest full-text academic papers: https://anil.recoil.org/notes/uk-national-data-lib
It covers some experiments on weight tying, one of which is actually LoRA and random weights.
There may be different constraints for other runtimes.
(Was a very enjoyable episode though!)
The paper we wrote on retrofitting effect handlers: https://arxiv.org/abs/2104.00250 also has some http benchmarks
Am happy to answer questions people have on this one as well, there's also some other OCaml contributors lurking.
Threads belong to a domain and only one thread can hold the runtime lock for the domain. This is the same behaviour as in OCaml 4.
With OCaml 5 you can have as many domains as you want though (we recommend no more than you have cores though).
In terms of the compiler and runtime development, the OCaml and ML Workshops at ICFP in October have videos that cover some of the experimental work happening: https://watch.ocaml.org/video-channels/ocaml2022/videos and https://www.youtube.com/playlist?list=PLyrlk8Xaylp7f8T7L5SFF...
There's also a compiler development newsletter that's posted on the discuss at regular intervals which details some of the other work happening: https://discuss.ocaml.org/t/ocaml-compiler-development-newsl...
I haven't done a huge amount of investigation but I suspect the cost comes from the extra indirection in the lock-free one.