HNHacker News
TopNewBestAskShowJobs

HenryNdubuaku

1,077 karma · joined October 24, 2021

submissionscomments
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
thanks, we improved on False Negatives this time :)
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
noted, we'd look into this, thanks
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
thanks!, let us know if you ever build it out :)
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
Thanks for testing Needle out! I'd be very interested in hearing more about the finetuning setup to see how we can make both the library's finetuning setup and the model better.
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
Thanks for considering needle. Keep in mind that you can also fine-tune the model to fit your use case more. I think this illustrates the intended deployment pretty well, where both computational resources and compute credits can both be issues for deployment.
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
Thank you fr this. "warm the house" now goes to the thermostat. It's fair that a more deterministic system with just action phrases would be easier to debug/interpret, but I think there is room for both a model that is trained to understand meaning as well as deterministic logic aiding it. To this end, we just started exploring the idea of triggers, and are working towards expanding this even more.
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
I think we might look into creating a baseline like this for our future models
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
Thanks!
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
The model can be sliced and perform the inference using a subset of its layers. The first 4 layers alone are 8 MB, all 20 are 29 MB. Fair point on the copy, tightened it.
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
Thank you!
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
I think that would be very useful for us! The best way to reach us is through the founders@cactuscompute.com email

Thank you!

HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
This is extremely useful feedback for us, thanks! I think the easiest thing here that can be fixed with tool definitions is the number conversions. Additionally, the model tends to work better with fewer tools. We will definitely be focusing on better context usage and followups going forward as well.
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
Thanks for the feedback! Implications and relations are hard for the model to understand (things like go to the living room, then the kitchen, and back), so yes the cleanest use cases involve direct language. Reasoning isn't true reasoning in the way general LLMs do it, it is more like grounding for the model that it generates itself. This can often become nonsensical specifically when the model gets things wrong, providing signal to the confidence.
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
That's a really good point and I think it's not yet clear how well, say, 8-30MB worth of regexs with accompanying algorithmic structure would do on these tasks. I would imagine they do quite well on a well defined task, but it would be much harder to then adapt this set to a new domain. A big part of Needle's promise is how easy it is to finetune. Ultimately I think the two approaches can be more complimentary to each other, rather than choosing only one (see triggers!).
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
Thanks! Certainly giving the model more context on the task it needs to perform would help it. This was actually a part of training that we improved going from Needle 2 to Needle 3
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
As far as I know Jev's architecture isn't public (though I might be mistaken!), but it is a coincidence :)

Needle 3 has been in the making since Needle 2 launched early august, but we are very excited that Jev is bringing more attention to the problem we are trying to solve.

HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
well certainly the environment on the website cannot be a full product, and it isn't claiming to be that. The model, while capable in many dimensions, is also limited by its size. The website is meant to show both the capabilities and the limitations! A real deployment would absolutely need external guardrails, more thoroughly thought out tool sets with better task-specific triggers, perhaps also task-specific finetuning for better confidence grounding. And in my view that's the point of small open models! You can take it and run with it as far as you want.
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
Hey there! If you end up trying out needle on the test suite it would be very useful for us if you could share some failure modes of the model! We are always trying to understand where the model isn't doing good and where we can make it better.

For your question on tool calling, I think you will find that the model is pretty good at simpler tool calls and parallel ones, but can struggle with implied references and multistep reasoning. These are definitely things that can improve with task-specific finetuning but for some things you just have to have a model that is properly sized. That said, we are always trying to improve the model so that it can handle an ever larger set of queries

HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
Hey, thanks for the feedback! I think this is a useful part of a demonstration so I added a 911 tool specifically to demonstrate this capability and the fact that you can guard it with triggers that make it so calling emergency is an unambiguous action given the input. This really shows that constructing the right tool set with the right surrounding setup is a priority when deploying needle.
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
Hey there, yep we found that on Apple devices specifically running on CPU is fast enough that Metal support is not needed. Thanks for flagging this though, and if usecases that would benefit from Metal support come up we will be adding it to the binaries.
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
Hey! Yeah I think for labelling the model would need to have much better world knowledge than its current size allows. Jev really is a very good model, I think it has a very strong place in the upcoming tech stacks. Really good suggestion to make a continuously distilled model, we are going to have to look into that one :)
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
Oh yeah really good use case! Definitely something to finetune the model for so that it gains better task-specific reliability, because rebooting the wrong server could easily be catastrophic.
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
lol well i guess you can turn the whole kitchen into an oven with needle :)

But for real usecases you are able to set explicit minimum and maximum values on the output range of numeric arguments, so that you can avoid situations like these. In this case it was hard for us to do that while keeping a broadly appealing demo since celsius and fahrenheit have different "reasonable" output ranges.

HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
haha i think that's a good demonstration of how external guardrails could help ground tiny models like these to prevent issues from coming up. I wouldn't trust needle to be my autopilot either (:
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
Hey, thanks a lot for this feedback, very useful and actionable for us! Quite a few of these came down to our tool definitions in the playground as well as out triggers. We updated them just now and these should be more reliable. Really this goes to show that needle shines through after putting in the work to make the tool list around it good for your use case. As for the reasoning, yes its main function is really to provide more words/keywords that the model can latch onto when generating the tool call response, since this is a SAN model it needs more grounding in existing context.
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
Hi, thanks for the feedback! And yes absolutely less tools and better tool descriptions make a huge difference for this model.
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
Yes, it was really difficult to compress meaningfully intelligence down to that, and there are many limitations we are aware off and still improving on.
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
100%, data was honestly most of the work, Needle 3 is trained on 360B tokens of structured data and we spend way more time on the generation pipeline than on the model. On Led Zeppelin, Needle doesn't actually need to know it, arguments are copied from the request so it just lifts the name into the artist field. The knowledge went into the engram btw, 70M of the 121M params are n-gram tables, so it can tell artist vs song without an MLP. Also yes, "play their second album" won't work, that needs the world knowledge it doesn't have.
HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
Makes a lot of sense! declare one record with the ~20 keys as fields, cuisine and friends as enums with the common values, and the grammar can't produce a value outside the set; fields with no evidence come back empty, so the output is exactly the diff.

For "which place", query nearby POIs from location in the app and pass the candidate names as an enum field, so Needle picks rather than guesses.

Two caveats: it's text-in, so you need on-device STT first, and opening_hours syntax is the risky bit, so either put the format in the description or capture the raw hours and normalise in code.

HenryNdubuaku··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
presets updated for you now :)
Page 1 of 6Next →