I've spent a few months trying to convince my management on getting a Macbook after fighting with a kernel memory leak that consumes the 128Gb of RAM in my freshly new Dell Pro Max 16 with Windows 11. It seems the best solution my IT gives me is to replace my laptop... which I have already done... three times.
I think this is a good chance to try and port Alyx (ARM64 build) to visionOS and run it on my Vision Pro headset which has been running around unused for the last two years xD
> What modern LLMs do is apply brute-force computation to any domain expressible in language--not just a restricted chess domain. That's the algorithm.
Which means that companies with sufficient computational resources and money will be capable of unlocking problems thousands of times faster and more effective than any individual even when lacking the skills, just by a matter of try and error.
I might be a weirdo, but even the mail client on iOS has serious deficiencies. Just getting an email with a zip attachment where I have to edit / sign a document with a certificate / something, rezip and send back is a hell in iOS, while it simply feels natural in Android.
What's the value of using one of these phones running iOS? iOS is great for media consumption and not so much for productivity... so I don't understand very well the use cases... apart from watching Netflix I guess?
They had other models a while ago, named 'Pulsar', and they were finetunes of Nemotron in collaboration with NVIDIA. I think this will be something similar.
I mean, I think it depends. At home 3 of us we use AI for multiple reasons, from coding apps to asking general questions, and if we would have to pay equivalent subscriptions that would be ~1k a year on AI + submitting all your data to external services. I payed around ~8k on 2 DGX Sparks that, at the moment, serves perfectly fine as a ChatGPT/Claude replacement at home (DS4 Flash peaking at ~170 tokens per sec with 6 concurrent sequences), and even once the technology is obsolete for inference in a few years, I will still have 2 pretty powerful machines for whatever I need + some pretty fast NVME Storage. I don't think its a terribly bad idea.
Not really, I just made a really simple Python library the agent can use and modify as he wishes within a sandboxed environment to explore the SQLite database. When the user prompts the agent, the prompt goes into my coding agent, omp, and starts running Python scripts and throwing SQL sentences until coming up with the answer. It works extremely well in our experience.
Its a read-only SQLite file. And I mean the people using the chatbot knows it uses AI so just like Google they shouldnt pick the first result, but forcing the LLM to mention the sources and not assume acronyms works amazingly well
I have a web user interface connected to a coding agent (OMP) running inside a container, so if any of the tool calls fail or something happens, usually my model recovers autonomously from these situations. The capabilities of models like DS4 Flash are those of frontier models from months ago, so its recovery and autonomous capabilities are quite impressive.
I indexed thousands of documents into a SQLite database with an FTS5 index, plugged it into DeepSeek v4 Flash and got better much better results than any other commercial solutions my company has tried in the past.
The trick was just to let the LLM come up with its own SQL queries for searching... and the results are impressive.
I've been reverse engineering LEON3-FT SPARC v8 BE code, so I wouldn't say it's common :D. When attached to Ghidra through a MCP the things you can do with this are simply crazy.
I am surprised. I've been using DS4 Flash (0731) for weeks now and it works perfectly fine as a replacement for Claude in a large variety of cases. It requires a few more iterations, sure, but it's useful enough to not need a Claude subscription anymore. Among the things I do I've been reverse engineering, writing complex C++ code...
Sure if you use remote AI services, but any companies working on niche markets where they want to protect their IP, or they simply work with sensitive stuff, will rely on local AI instead.
Absolutely agreed, I don't see this as an investment leading to cost-savings in the future, at all.
I see this rather as an investment in improving my capabilities and knowledge of this technology, letting me play with a 'GPU cluster', vLLM and other technologies that otherwise would require me to rent GPUs on the cloud.
I ended up buying 2 DGX Sparks interconnected over QSFP with the intention of getting rid (as much as I could) of any cloud-based AI provider. I'm running DS4 Flash 0731 on it and some OCR models, using Oh My Pi and OpenWebUI as my main ways to interface with the agent... and from someone that has been using Claude for a long time, I can definitely say I don't need it anymore. Not for the stuff I'm doing.
I've been using HarmonyOS for almost a year now (not the AOSP fork, but the actual Hongmeng microkernel flavour)... and this idea of having a iOS-like system with an Android VM running at near native performance works pretty well. Unfortunately is really focused on the chinese market but it feels really polished and promising.
Haven't used IDA much lately, but after looking at the screenshot with that IDA PRO decompiled code in their website I feel like Ghidra is already ahead of them in this area :D
Quite weird that heavy quantization method on a dense model gives better results than slightly quantized MoE models like 35B-A3B from Google.
At this point all the different quantization and 'compression' (look at MPO applied to LLMs...) techniques start feeling a bit like snake oil. It's just gut feeling - or scores on benchmarks models are optimized for - what ends up deciding whether a technique is good enough or not.