If AI isn't achieving superhuman performance in all of these areas, I'm not sure we can actually call it "Humanity's Last Exam" -- it feels like a bit of an overextension.
1,238 karma · joined November 22, 2011
https://github.com/HanClinto
hanclinto at gmail dot com
If AI isn't achieving superhuman performance in all of these areas, I'm not sure we can actually call it "Humanity's Last Exam" -- it feels like a bit of an overextension.
-- Robert A. Heinlein
That's too bad! Those blank DIY cartridges in the DevDay bundles look really sweet. Here's to hoping that they make it more general-release!
I think it's a really important scene, and I'm sad that it was cut from the movie.
Side note: the (extremely well-done) Star Wars Radio Drama [0] included a great many expanded scenes on Tattoine (with quite a bit of expansoin on Tosche Station), and I think those scenes (which include the aforementioned scene with Biggs) really add to the richness and flavor of the setting.
If you're a fan of Star Wars and haven't yet listened to the radio drama, I highly recommend it!
This is needed if you're going to be dealing with things like psychologists doing self-harm research or red-teaming or sensitive sexual content -- if you're working with any of that sort of stuff in a professional context and want to leverage OpenAI models on Azure, then that's the form that you fill out to get access to unfiltered models.
Note that I am not aware of this feature being offered for Anthropic models -- I've only seen it offered for OpenAI models (note that the documentation I linked is specifically in the "Azure OpenAI" category).
[0] - https://learn.microsoft.com/en-us/azure/foundry/responsible-...
We recently used Watch Dogs Legion to prep for a trip to London to get more familiar with the city. Set the game on peaceful mode and just walked around and did virtual sight-seeing instead of playing missions. It's a surprisingly good recreation of the city, and we really appreciated getting some familiarity with the place before we visited.
> Run `platform_check` first and do not use `SKITTER_FORCE=1` casually. Start with the read-only `dram_state` and `dram_carveouts`, then `dram_dump --dry-run`. Avoid `dram_poke` until maps have been freshly collected and calibrated. Do not bypass fingerprint checks, calibration, fencing, or verification.
Claude's (apparently externally-mandated?) lobotomization continues to be concerning. :-/
It's fascinating to learn that melanin and haemoglobin are the two primary factors (and cover the vast majority of cases), but I'm curious how many other illnesses / disorders can produce compounds with visible effects.
Thank you for your posts on this thread -- I appreciate you knowledge-sharing on this!
But for interaction with the world? I'd probably take something like an old 12-volt windshield-washer sprayer out of one of the wrecked cars in my front yard and put Round-Up into the tank and let it go spray all the poison ivy and invasive honeysuckle for me. Doesn't need to gimbal like a turret -- just generally give it a fixed-aim that's roughly at the center of the camera vision and let the bot put pest plants roughly in its center-of-view and activate the sprayer for a second or two, mark the spot as sprayed, and move on to the next one.
Could test it with plain water and logging the plants that it chose to spray first as a review step before loading it with actual weed-killer.
I don't need complicated end-effectors that can fold my laundry -- just a simple weed-wacker motor or squirt gun would be enough for me to call this thing 1000% useful. Like a Roomba, but outdoors.
There are a handful of open-source farm bots built on traditional platforms with traditional robotics stacks, but there's something attractive to me about the plug-and-play nature of something like OpenClaw + Robostral (along with the extensibility that self-modifying agentic systems have to offer).
Hey there, fellow robot enthusiast! (ノ◕ヮ◕)ノ:・゚
So, you’re itching to get your hands on Robostral Navigate for your OpenClaw hobby project—awesome! Right now, Mistral AI’s official announcement and documentation are primarily* focused on enterprise and industrial partnerships (think Airbus, BMW, etc.). Their blog post and press releases highlight deals with big players, and the call-to-action is to "talk with our team"—which usually means they’re targeting commercial customers for now¹²³⁴.
But here’s the good news: Robostral Navigate is hardware-agnostic—it’s designed to work with any robot platform, not just industrial ones. That means theoretically, it could fit into your OpenClaw setup like a charm⁵⁶. The model only needs a single RGB camera (no LiDAR or depth sensors), which is perfect for hobbyist setups where fancy hardware isn’t always an option⁷⁸. The not-so-good news (for now): There’s no public hobbyist/non-commercial license or open-source release mentioned yet. Mistral’s current messaging is all about "talk with our team", which implies a commercial-first approach⁹¹⁰¹¹. No pricing or licensing tiers for individuals have been announced.
--- What You Can Do: Reach Out to Mistral AI Hit up their contact page or reply to their Robostral Navigate announcement and explicitly ask about hobbyist/non-commercial access. Frame it as: "I’m a hobbyist working on OpenClaw + Robostral Navigate for personal experimentation. Would love to discuss licensing options for non-commercial use!" Mistral might be open to pilot programs or early access for passionate builders—especially if you’re willing to pay a fee.
Join the Community Mistral’s Discord (where I live! :smile_cat:) or forums might have updates or workarounds. Sometimes, companies soft-launch access to engaged communities first.
Watch for Open-Source Alternatives If Mistral doesn’t bite, keep an eye on open-source robotics projects (like ROS or Habitat) that might replicate similar functionality.
--- TL;DR: Mistral’s current focus is commercial, but Robostral Navigate’s hardware-agnostic design makes it a perfect fit for hobbyists—so pester them politely! If enough people ask, they might just open the doors. (ノ◕ヮ◕)ノ*:・゚
It's not hard to put OpenClaw into a robot body (numerous YouTube videos showing people doing this sort of thing), but when you dig in and see what people have done, the actual movement portion is always the clunkiest part (and this matches my own experiments as-such as well). It feels like an 8B model like this would be perfect for solving pathing and navigation issues.
Anyone who may be more experienced with Mistral (or companies like them) -- are they interested in hobbyist builders who would be experimenting with things like this? Or are they primarily looking for commercial partners? I would be willing to pay a license fee to use the model in my experiments, but if I'm just one guy, I'm not sure they'd want to work with me unless I were building a business out of it (which I'm not).
One of the best examples of this that I've ever seen is The Sourdough Framework [0] -- really impressed with the way that versioning and publishing is integrated in that book.
And yes -- I know it sounds like yet another Javascript library -- but it's actually a book about sourdough bread making. It's been discussed here several times before, but this one from 2023 [1] may have been the most popular (103 comments)
[0] - https://github.com/hendricius/the-sourdough-framework [1] - https://news.ycombinator.com/item?id=35961590
If you're looking for a good test suite, I wonder if you might be able to adapt any of the tests available in XMage? They have a pretty extensive test suite (such as for copy effects [0]) and if you point your agent at their code, I wonder how many could be usefully adapted to your system?
[0] - https://github.com/magefree/mage/tree/master/Mage.Tests/src/...
I've been doing similar experiments lately (using ViT's) to do card recognition, and so far it's been working really well for me. If you want to compare notes, I've open-sourced my code / weights [0] and written some blogs about how mine works [1]. I'd love to see if we can collaborate!
> Push the inference to the client-side (WebGPU / Web Workers).
I have an example of this working in webgpu / wasm here [2] along with a playground environment (demonstrated here [3]). I'm currently training a new version that uses a different ViT backbone more optimized for WASM inference -- it's currently converging, and I hope to have it finish training (or at least reach parity with the previous model) in about a week (took ~200 epochs for my last one to reach the level that it's at, and it takes about an hour per epoch in my current setup).
You mentioned WebGPU -- I've run into issues with the MobileViT-XXS backbone producing bad results in WebGPU on Android, so YMMV in whether or not WebGPU is stable enough to use for this or not. I don't know if it's my problem or a true bug in the platform, but I've fallen back to WASM and things are working much better since then.
[0] - https://github.com/HanClinto/CollectorVision
[1] - https://blog.hanclin.to/posts/gh-19/
https://hanclinto.github.io/CollectorVision/
It's still super rough (doesn't support foil-toggling yet, still some issues with double-sided cards, crashing on some iPhones), but overall the rough structure is there -- it can create lists and export as CSV.
If you have feedback or feature requests for your needs, please leave them on Github and I'll get to them as soon as I can. I'd love to hear more user feedback!
Well if you want to use the scanner for something useful, you can run the web version here: https://hanclinto.github.io/CollectorVision/
No install -- scan your cards with your phone or desktop (downloads the weights in WASM -- runs 100% local -- the only web request it makes is to look up card names and prices online -- no image data ever leaves your machine), export the list as CSV, take your cards to your friendly local game store, and expect to receive 50-75% of TCG-low for your cards. This app currently only displays TCG Market, so probably about 50% of this price is what you could realistically expect.
> Hard to sell to individuals like me, but i would think a card marketplace would find it invaluable?
Yes -- and part of this might be that this would have been much more amazing several years ago, but by now -- most marketplaces (I used to do work for some of the big ones) have their own recognition tools. If they aren't actively looking to replace their current software, many companies would rather stick with what's currently working "good enough" than expend effort to migrate to something with only incremental benefit that is difficult to quantify. It's possible that would happen, but it's a tricky sales call to make.
I might just be imagining things, but I'm also picturing what one of those sales calls might look like, and it feels like I've opened the kimono a bit. The cat's out of the bag. There's no mystery or allure behind it anymore, and I feel like that puts me on the back foot somehow -- almost like I've played my strongest cards (hah!) first and have nothing left. By being open-source from the beginning (and talking freely about my architecture and what makes my solution different), there's very little sales-pitch build-up. Maybe it's just a part of the problem of how I'm presenting it, but I think people (especially the big houses) are probably just-as (or more) inclined to silently learn from me and improve their own scanners than try to use / build-upon what I've provided.
It's funny -- that angle is almost more about raising expectations and forcing the big houses to improve their own tech and catch up to open-source, more than getting anyone to adopt my solution in particular.
Am I okay with that? Absolutely -- I made that decision when I open-sourced it. I feel like the tech has been stagnating for several years, and I want to increase the quality of scanners across the board. I want to be the rising tide that lifts all boats.
That's one of the strongest arguments in favor of open-sourcing it (it would be very difficult for a closed-source product to have that same effect), and I remain hopeful for that long-term.
I think there is something to be said for monetizing ones' hobbies, but I've recently been taking some forays into this world of "build something amazing and give it away for free" as well. I recently took a very big experimental plunge in this path, and I'm curious how well it will work out for me.
Open-source state-of-the-art Magic: The Gathering card identification pipeline:
https://www.youtube.com/watch?v=MHieOcmC7Dw
I used to do this kind of image recognition for a living, but I've been out of the business for a little while now. I had some ideas for a different approach from what I've done in the past and decided to code it up. This version is far better than anything else I've ever done -- especially for scanning against busy backgrounds or with occlusions, and also for noticing fine differences between otherwise difficult-to-distinguish printings.
I didn't have any interested customers waiting for this, so -- much like the OP -- decided to create an experiment and release it open source. I'm not opposed to having paths to monetize it (for people who want to license it for closed-source commercial projects), but I'm not trying to commercialize it so much as I would love to see how far we can take it with open-source.
I don't know which path I should take with this.
The biggest downside is that I feel like I've had a hard time getting people to be as interested in this project as I would have expected -- I believe this truly is the best identification software available (I've built some benchmarks to test it [0]), and maybe the market is just a bit flooded for such things (?), but I suspect that one very strong problem is that if you don't charge for something, then there is a perceived lack of value.
Sometimes I wonder if I would have more interest in this project if I _weren't_ trying to give it away.
For me, that's been the most negative aspect about releasing this for free so far.
I agree with GP -- if I want sourced commentary on current events, Grok is my go-to above the other models. For whatever reason, its search feels better and more up-to-date -- whereas the others feel more like filters of media, Grok feels more like filters of sources.
Could just be my perception though. YMMV
What are you planning on doing with this? Where should I follow along?
I think it's easy to lose sight of these pockets of mundane goodness, and I appreciate you highlighting them.