You can also go beyond Jev. Qwen 3.5 0.8B is fantastic at basic image classification/question answering (including OCR elements) also. Though rather than looking at logits, I get it to output a structured JSON object and it does simple object classification tasks on a Mac at under 500ms a pop (I forget how far, but I think it's like ~250ms) with good accuracy (depending on task).
I've had similar thoughts! I grew up around a heavy smoker. My dad told me to never smoke and I never have. But I randomly tried a nicotine lozenge in my 40s and found it such a "lightbulb on" moment that I've been using pouches for a few years now and my executive functioning still feels better for it (though not as much now as in the first two years, alas).
I had a similar thing with Xcode the other day. I unticked the "download a code assistance model" thing, clicked to continue and it started downloading it anyway (which, to be fair, you can just X out of).
There might be some practical applications of this sort of idea like in situations where you want/have a heavily restricted vocabulary to build from. You can already do this with LLMs but they can get very... "distressed" if you force logits, whereas this would not.
Come to think of it, I'm now curious if it would do well at building SQL queries, say. (30 minutes later: I tried it, and it can do it reasonably well, but a normal model and linting will outperform it, though the inherent guard rails of a limited vocab are still intriguing.)
FWIW, on my Mac Studio I get ~24-27 tok/s generation between 0-16k context in - that's on the Q6_K GGUF with speculative decoding on. I have spent zero effort optimizing/improving this so far but will be trying the 4 bit MLX next (I've tended to find models drop off somewhat below 6 bit but maybe that isn't the case nowadays).
How's the compute side now, I wonder? Because while the Ultras have impressive memory bandwidth for inference, processing prompts still takes a dog's age on my M3 Ultra. I heard the M5 makes some strides forward in this area, though, and the M7 in particular promises to go a lot further.
More likely double that, even. I think you'd still see many buyers there. You can spend like $16k alone on a RTX 6000 PRO with a mere 96GB of VRAM now..
They seem to be suffering from the supply constraints like everyone else. They phased out the higher capacities on the M3 Ultra Mac Studio a while ago, and if you order a 128GB MBP, say, you're looking at six weeks or more for delivery.
It reminds me conceptually of the idea of using a ST:TNG replicator to just give you another replicator of your own, or asking a stereotypical genie for "infinite wishes". The genie is indeed out of the bottle in many ways.
As well as the headline in/out changes, people heavily using agentic coding tools will want to note the 6x (off peak) and 12x (peak) increase to cache hit pricing on Pro (since cache hit can easily make up 90%+ of input on long sessions).
DeepSeek was hugely underpricing cache hit pricing before and even after this increase they're still cheaper on that metric than every other provider I'm aware of, but it will put an end to those "I used 1 billion tokens and spent $4" reports.
Not to take away from the broader point, but modafinil is lumped in with methylphenidate as a stimulant that "increase[s] the activity of neurotransmitters such as dopamine and norepinephrine". Modafinil does also stimulate the orexin system and is part of why it's so effective for wakefulness and not as a stimulant in the medium/long-term.
No, it’s a licensed black cab thing. London has had mini cabs as well for many decades who didn’t need to study anything and were not even licensed till about 20 years ago. However, their downside is they have to be prebooked and can not be hailed on the street. Apps count as booking and not hailing I guess.
I agree with you to an extent, but you have certainly given me food for thought.
Sticking to LLMs, they seemingly get their intelligence (whatever that really means) from building models rich with knowledge, so you could have a point. But Qwen models seem to be particularly good, even at small model sizes, at maintaining both their own knowledge while acquiescing to and integrating external information in the moment.
Hopefully this boils down to the smaller versions they've teased. In my experience, Qwen models are the closest to the "less knowledge, more intelligence" (yes, the two are hugely correlated!) ideal some tool-dependent tasks need. Even the 3.5 2B can be easily prompted to always lean on tools and not jump to false conclusions (although its actual coding skills are abysmal, as you'd expect).
I'm guessing this is almost entirely about the incredibly low cache read prices. Few have come close to them, nothing has a bigger effect on (a typical) session price, and with the price of RAM right now, they have to be the biggest pain point for them right now? A 10x increase in cache read would be a significant increase, yet would still keep them cheaper than every other provider of their model (at least based on the prices at https://openrouter.ai/deepseek/deepseek-v4-pro#providers)
Shopping with my kids I noticed we kinda have a replacement here in the UK in the shape of HMV (originally a spin off of His Master's Voice). When I was a kid it was mostly a record store, but now it's focused on what I'd call geek/"fandom" tat: posters, figurines, comics, band merch, etc. Whenever I'm dragged into one, it's always full of geek/fandom types hanging out.
Any free hosting solution tends to spiral towards abuse at some point.
Does it work? I'm a paid user, but set up a new free account to see if the deployment stuff worked there too (as I was linking to it from elsewhere) and it said I had to upgrade to a plan that supported it.
I bought one as it was on sale at the same price as the normal iPhone. Pretty good: 8/10. I have no interest in photos, don't need long battery life, and the lighter/thinner the better.
That said, I'd prefer something smaller but thicker with no protrusions/camera bump at all. I hate the wonkiness of all modern iPhones when sat on a table. I think the last totally flat one was the iPhone 5s!
Not a “road” as such, but it’s also quite common in multi storey car parks as well where a ramp is shared by traffic going both up and down to stop unnecessary crossovers at each level.
It's been several years (and the issues are in my office so I could take a look later) but I recall it was a mix. They'd sometimes group big stories together thematically (like when all the billionaires were starting to get into space) but also do longer-term reflections on single, complex stories like the Salisbury Novichok incident.
I don't recall it really surfacing things that didn't make the news cycle, but one thing I failed to mention is they had (have?) really good infographics/visualizations! They have a lot of the more recent ones here https://www.slow-journalism.com/filter/infographics
I was subscribed to this. It's beautifully designed, well written, good paper stock, the works. If the idea appeals to you, it's worth trying.
Ultimately, it didn't work for me, but that's my fault. Despite my best intentions, it turned out I wasn't interested in reading about world affairs beyond the news cycle.
Which, more often than not, causes weird titles I’m more likely to click due to sounding curious and interesting, undoing the entire point of the feature.
I would not have clicked on “How My Images Are Dithered” but this truncated version made it sound more like an interesting bug story.
I wonder if the broad use of AI overviews on Google search results is having an impact. Maybe the numbers make it more profitable to use their compute on several billion searches a day rather than selling API access.
Various devices that I don't want to install the official software or drivers to control. I've done this with several things, but just as an example I have a MIDI MPC pad (the sort of thing samplers/beat makers use) and worked out it had various features not supported officially like controlling the lights on the pads. I also discovered some cheap Chinese lights my daughter owns are controllable over BLE without encryption.