HNHacker News
TopNewBestAskShowJobs

sadiq

5,054 karma · joined July 19, 2007

Computer Lab @ Cambridge, Co-founded Opsian, OCaml Multicore hacker
submissionscomments
sadiq··on Property-Based Testing of OCaml 5's Runtime System [pdf]
It's worth looking at Jan's trophy cabinet at the bottom of https://github.com/ocaml-multicore/multicoretests/

His work has uncovered a number of really tricky bugs in the multicore runtime but what's brilliant is the reports normally come with a minimal reproduction. This makes working out the cause so much easier.

Great work Jan.

sadiq··on Can a model trained on satellite data really find brambles on the ground?
So the interactive map should do this workflow for you. You place points and it will run the knn classifier over the landscape for you.

If you want to go further you can export the GeoJSON and then run it through any machine learning pipeline you like.

sadiq··on Can a model trained on satellite data really find brambles on the ground?
I would try https://github.com/ucam-eo/tessera-interactive-map , this is relatively easy to get started with and has a nice interface for labeling.

https://github.com/ucam-eo/geotessera has an image showing our embedding coverage at the moment. Blue areas we have complete coverage for 2024, green areas we cover 2017-2024. We're slowly trying to populate everything 2017-2024 but the constraint is GPU and storage at the moment - each year takes ~20k GPU/200k CPU hours and requires storing and serving 200 terabytes of data. The world is big!

If there is an area you would like prioritised, there's an issue template on the geotessera github repo which we can use to move regions around in the processing queue.

sadiq··on Can a model trained on satellite data really find brambles on the ground?
It's possible to use embeddings as input to a convolutional network and then train that using labels. We've done that for at least one of the downstream tasks in the TESSERA paper: https://arxiv.org/abs/2506.20380 to estimate canopy height.

The downside of that approach is that you need to spend valuable labels on learning the spatial feature extraction during training. To fix that we're working on building some pre-trained spatial feature extractors that you should only need to minimally fine-tune.

sadiq··on Can a model trained on satellite data really find brambles on the ground?
I was a lot more optimistic about Gabriel's model than he was. It is essentially a presence-only species distribution model where accuracy depends largely on assumptions around prevalence and which really needs some presence-absence data to calibrate.

As I mentioned in one of the other comments, the model is also only pixel-wise. That is, it is not using spatial information for predictions.

sadiq··on Can a model trained on satellite data really find brambles on the ground?
We did note several places during the trip that didn't contain bramble. The hotspot in the middle of the residential area was also entirely isolated.

For a proper evaluation you would need to be more methodological but as a sanity-check we were very happy with it.

One other thing to point out about the bramble model is that it is pixel-wise. That is each prediction is exclusively only what is within the 10 metre pixel (give or take the georeferencing error).

sadiq··on Can a model trained on satellite data really find brambles on the ground?
It might work. TESSERA's embeddings are at a 10 metre resolution, so it might depend on the size of the features you are looking for. If those features have distinct changes in colour or texture over time or they scatter radar in different ways compared with their surroundings then you should be able to discriminate them.

The easiest way to test is to try out the interactive notebook and drop some labels in known areas.

sadiq··on Can a model trained on satellite data really find brambles on the ground?
If you have some GPS locations of truffles, you could use the notebook Anil mentioned here https://news.ycombinator.com/item?id=45378855 and give it a go.

There is the issue of just how visible truffles are from space though, if they grow under cover. That said, it may still work because you can find habitats that are very likely to have truffles. We've had some promising results looking at fungal biomass.

sadiq··on Can a model trained on satellite data really find brambles on the ground?
That's actually a great idea! I wonder what kind of feature size would be needed though - TESSERA's embeddings are at a 10 metre resolution so for larger structures you might need some kind of spatial aggregation.
sadiq··on Can a model trained on satellite data really find brambles on the ground?
Hyperspectral data is really neat though it's worth pointing out that TESSERA is only trained on multispectral (optical + SAR) data.

You are very right on the temporal aspect though, that's what makes the representation so powerful. Crops grow and change colour or scatter patterns in distinct ways.

It's worth pointing out the model and training code is under an Apache2 license and the global embeddings are under a CC-BY-A. We have a python library that makes working with them pretty easy: https://github.com/ucam-eo/geotessera

sadiq··on Can a model trained on satellite data really find brambles on the ground?
Yes! TESSERA is very new so we're still exploring how well it works for various things.

We're hoping to try it with a few different things for our next field trip, maybe some that are much harder to find than brambles.

sadiq··on Can a model trained on satellite data really find brambles on the ground?
Hi! You can find a bit more about Gabriel's model through some of his posts over the last few weeks: https://gabrielmahler.org/posts/

When it comes to the satellite images, the model actually used TESSERA (https://arxiv.org/abs/2506.20380) which is a model we trained to produce embeddings for every point on earth that encodes the temporal-spectral properties over a year.

Think of it like a compression of potentially fifty or a hundred observations of a particular point in earth down to a single 128 dimension vector.

Happy to answer any other questions.

sadiq··on What's the strongest AI model you can train on a laptop in five minutes?
You might find https://arxiv.org/abs/2401.17377v3 interesting..
sadiq··on Open models by OpenAI
Looks like Groq (at 1k+ tokens/second) and Fireworks are already live on openrouter: https://openrouter.ai/openai/gpt-oss-120b

$0.15M in / $0.6-0.75M out

edit: Now Cerebras too at 3,815 tps for $0.25M / $0.69M out.

sadiq··on Show HN: RULER – Easily apply RL to any agent
Excellent, look forward to giving this a go.

I was looking at: https://arxiv.org/abs/2506.18254 but your approach is even more general.

sadiq··on Reservoir Sampling
This is a really nicely written and illustrated post.

An advanced extension to this is that there are algorithms which calculate the number of records to skip rather than doing a trial per record. This has a good write-up of them: https://richardstartin.github.io/posts/reservoir-sampling

sadiq··on Starting July 1, academic publishers can't paywall NIH-funded research
This is good though it's not clear whether these papers will appear in the PMC Open Access subset (https://pmc.ncbi.nlm.nih.gov/tools/openftlist/) and be bulk downloadable.

I've been doing some work with colleagues at Cambridge and Imperial over the last year on using LLMs to improve evidence synthesis, primarily trying to find papers on the effectiveness of certain Conservation interventions. It's becoming clear that you really need to move beyond screening papers only by title and abstract - there's often information buried deep within papers that can only be found with access to full text. My colleague Anil Madhavapeddy has written a bit about our adventures in trying to ingest full-text academic papers: https://anil.recoil.org/notes/uk-national-data-lib

sadiq··on SeedLM: Compressing LLM Weights into Seeds of Pseudo-Random Generators
There's a related write-up here you might find interesting: https://wandb.ai/learning-at-home/LM_OWT/reports/Parameter-s...

It covers some experiments on weight tying, one of which is actually LoRA and random weights.

sadiq··on OCaml: a Rust developer's first impressions
For 5.0+ you might want to look at https://github.com/ocaml-multicore/eio for how effects can make async much more pleasant
sadiq··on OCaml: a Rust developer's first impressions
There will also be DynArrays in the stdlib from 5.2: https://github.com/ocaml/ocaml/pull/11882
sadiq··on The Garbage Collection Handbook, 2nd Edition
Great that there's a new edition of this coming. This is the definitive reference if you work with garbage collectors, it's also very well written.
sadiq··on OCaml 5.0 Multicore is out
Also worth pointing out there's design constraints on the OCaml 5 GC imposed by some of OCaml's language features (looking at you ephemerons) and C API invariants.

There may be different constraints for other runtimes.

sadiq··on OCaml 5.0 Multicore is out
One thing to point out to anyone listening to that episode is that at the time the plan was to only upstream the multicore GC for 5.0 and then follow up with effects. Instead they both went in to 5.0.

(Was a very enjoyable episode though!)

sadiq··on OCaml 5.0 Multicore is out
We certainly hope so. You may find Thomas Leonard's talk from last year's workshop interesting: https://watch.ocaml.org/videos/watch/74ece0a8-380f-4e2a-bef5...

The paper we wrote on retrofitting effect handlers: https://arxiv.org/abs/2104.00250 also has some http benchmarks

sadiq··on OCaml 5.0 Multicore is out
There's already some discussion on https://news.ycombinator.com/item?id=34013767

Am happy to answer questions people have on this one as well, there's also some other OCaml contributors lurking.

sadiq··on OCaml 5.0 Multicore is out
Just to add to the sibling comment. To maintain backwards compatibility, OCaml 5 has both threads and domains.

Threads belong to a domain and only one thread can hold the runtime lock for the domain. This is the same behaviour as in OCaml 4.

With OCaml 5 you can have as many domains as you want though (we recommend no more than you have cores though).

sadiq··on OCaml 5.0 Multicore is out
There's definitely some work to build atop the new functionality available in 5.0 and make sure there's plenty of good learning material.

In terms of the compiler and runtime development, the OCaml and ML Workshops at ICFP in October have videos that cover some of the experimental work happening: https://watch.ocaml.org/video-channels/ocaml2022/videos and https://www.youtube.com/playlist?list=PLyrlk8Xaylp7f8T7L5SFF...

There's also a compiler development newsletter that's posted on the discuss at regular intervals which details some of the other work happening: https://discuss.ocaml.org/t/ocaml-compiler-development-newsl...

sadiq··on OCaml 5.0 Multicore is out
There are a few OCaml contributors lurking and happy to answer questions if you have them.
sadiq··on On “correct and efficient work-stealing for weak memory models”
There's a high performance Chase-Lev work stealing queue written to exploit the OCaml 5 memory model in the lockfree library: https://github.com/ocaml-multicore/lockfree/blob/main/src/ws...
sadiq··on Static B-Trees: A data structure for faster binary search
So that was in some quick benchmarks against the sequential one in the runtime: https://github.com/ocaml/ocaml/blob/trunk/runtime/skiplist.c

I haven't done a huge amount of investigation but I suspect the cost comes from the extra indirection in the lock-free one.

Page 1 of 9Next →