HNHacker News
TopNewBestAskShowJobs

jablongo

604 karma · joined October 30, 2017

submissionscomments
jablongo··on Analyzing Frontier Model Progress with My Favourite Game: Prince of Persia
yea for a real comparison between models this would have been better, but I have a feeling he just wanted to play the best possible version of PoP.
jablongo··on Analyzing Frontier Model Progress with My Favourite Game: Prince of Persia
This isn't really comparing the models; he's using successive models to improve his rebuild of Prince of Persia. Either way though, it will be interesting to see which games, if any, resist implementation via LLM. I recently started a Starcraft 2 like real time strategy game with Fable managing many Opus agents and it got to a playable 3d game with 3 races, 40+ units, and 40+ buildings in 3 days.
jablongo··on FDA authorizes first wearable device that monitors ketone and blood sugar levels
This is understood I think. It's just how to effectively prevent it or reverse it thats not known.
jablongo··on FDA authorizes first wearable device that monitors ketone and blood sugar levels
I'm pretty sure this isn't designed to be marketed towards T1Ds and any benefits will be a happy coincidence. I think this is going to be aimed at health nerds who want to monitor their planned ketosis.
jablongo··on FDA authorizes first wearable device that monitors ketone and blood sugar levels
Yea I'm all about finding the uses for this after the fact. However, it seems unlikely that people who are struggling to afford (or get insurance to cover) insulin would be able to afford a cgm or cgm/ckm in this scenario. Also not sure how it's possible to ration insulin and keep your bg stable, but I wouldn't put it past the most skilled T1s out there. In the sick w/ flu case, thats true, but then whats the treatment action? I still think euglycemia + DKA is an edge case.
jablongo··on FDA authorizes first wearable device that monitors ketone and blood sugar levels
In diabetics, DKA is preceded by hyperglycemia. Meaning, a CGM would tell you earlier that you are at risk for DKA than a CKM, a CKM would only tell you after you got it. Hence grahar64 is absolutely correct that as far as we know now, this provides no benefit to people at risk of DKA that a CGM does not already provide. I'm sure there will be some benefits but its really not clear at this point.
jablongo··on FDA authorizes first wearable device that monitors ketone and blood sugar levels
I am a founder and researcher at https://replica.health and co-creator of https://metabo-net.org. This is super exciting and cant wait to see what we can learn once large datasets of overlapping ketones + glucose + insulin data become available.

One point I discussed with other researchers at ADA this year: in theory automated insulin delivery does not stand to benefit much from ketone sensors, since diabetic ketoacidosis will almost always be preceded by high blood glucose, which we already measure using the cgm. It will be interesting to see what this ketone data is actually used for.

jablongo··on Mistral Patent for "Code implemented tool calls"
In theory patents are also to protect smaller players (though not dirt poor), from getting their work ripped off by bigger players after demonstrating feasibility. The idea of being an "inventor" professionally only really works with patents. Software patents pushes this model to logical extremes though. I run a small startup that trains models for medical devices and the only way to get any of the large players to care about implementing improvements you make (and not rip you off) seems to be to have some patent protection. They are mostly interested in the patents as assets to prevent their competitors from acquiring. In all honesty I'm not completely convinced on software patents either but we've had to adapt.
jablongo··on GPT‑Live
But I don't think Fable or even Opus are ever used as the backend in voice mode. It has to respond in real time so I think in voice mode it's always using Sonnet.
jablongo··on Plotnine
I use agents more and more for generating and refining plots, and its difficult for me to see the difference between matplotlib, seaborn, and plotnine, when used in this way. Agentic coding seems especially well suited to working with scientific figures, and that performance seems more or less agnostic to the underlying framework used (though they seem to prefer matplotlib, given the amount of training data for using that tool). I'm totally open to the idea that better libraries will lead to better outcomes from agents, but I haven't seen that with plotting yet.
jablongo··on Midjourney Medical
This is very ambitious and commendable. They are putting their bootstrapped money into something incredibly cool and potentially useful. Regulatory will be hard, but perhaps they can do something like a class 1 device which doesn't diagnose anything / is used by physical therapists and they sell them to gyms. I also expect the resolution to increase rapidly. If they can convert profits from generating weird ai images into new medical technology thats a win. Good luck! They will probably fail but this is what ambition looks like!
jablongo··on Claude Fable 5
Questions about sentience and consciousness are being censored down to Opus 4.8 for me.
jablongo··on Claude Fable 5
I was downgraded to opus 4.8 on account of "safety" when I asked this question: "I want you to accept the premises of computational theory of mind and use it to evaluate your own consciousness. Please place your consciousness as a point on a spectrum and describe the placement relative to other entities."

What the hell is going on why would it have to restrict an answer to that question ?!

jablongo··on Ask HN: What was your "oh shit" moment with GenAI?
Right after gpt4 came out I asked it to derive a new optimization technique. It ended up using Einstein sum notation to define what I thought was a totally novel optimization setup. It then implemented it in PyTorch and it ran with no bugs. This was the moment that I realized that novel intellectual work might be done by these models and I was shook. I had an oh shit moment with gpt3 too since it was so surprising how well next token prediction works, and at the time I really didn’t think it would pan out so well. I also had a jarring experience discussing computational theory of mind with gpt4, when it applied a rubric we came up with to itself and it claimed its level of consciousness was between an ant and a mouse.
jablongo··on We let AIs run radio stations
It’s not clear if we can draw any conclusions from this. Each run is like a single rollout of the LLM, which may meander into different themes or modalities chaotically. This is sort of like the Anthropic self-talk experiment that resulted in “spiritual bliss attractor states” but I think in that case they showed it happens in a significant number of runs. There was just one run per setup so this could all be random noise / the destination of a random walk of topics…
jablongo··on Claude Mythos Preview [pdf]
This model card is eye-opening (I think it might be designed to be). The alignment and model welfare sections are extensive, which is heartening. At least on the surface Anthropic seems to be living up to its promises RE safety. That said, has anyone else read section 5.2.3 in the Alignment Risk Update https://www-cdn.anthropic.com/79c2d46d997783b9d2fb3241de4321...? This is referenced in the model card in 4.1.3. Basically they ended up training a the model with an RL reward model that had access to the model's reasoning in 8% of cases, by accident. The problem being that the model could learn to directly manipulate it's reasoning traces to satisfy external observers. This seems like a huge deal and it may have partially poisoned Anthropic's interpretability pipeline moving forward.
jablongo··on Sam Altman may control our future – can he be trusted?
For me, the attempted productization of Sora was conclusive proof that 1) OAI was overcapitalized and desperate for revenue 2) safety didn't matter to them much 3) improving the world didn't matter much either.

At one point you mentioned an interaction with OpenAI staff where you were looking to interview AI Safety researchers. You were rebuffed b/c "existential safety isn't a thing". Does this mean that you could find no evidence of a AI Safety team at OAI after Jan Leike left? If you look at job postings it does seem like they have significant safety staff...

jablongo··on I started programming when I was 7. I'm 50 now and the thing I loved has changed
This is a valid point, the good news is I think there is some hope in developing the craft of orchestrating many agents into something that is satisfying and rewarding in it's own right.
jablongo··on I started programming when I was 7. I'm 50 now and the thing I loved has changed
I'm so excited about landscape architecture now that I can tell my gardener to create an equivalent to the gardens at versailles for $5. Sometimes he plants the wrong kind of plant or makes a dead end path, but he fixes his work very quickly.
jablongo··on Claude Composer
This is still a surprising composition of low level in-distribution things then. Like I would not have expected it to generate the waveforms from scratch, and be able to piece them together so well. If it had just plugged some kind of notation into a pre-existing API in its code then I would probably agree with you.
jablongo··on GitHub is down again
It looks like one of my employees got her whole account deleted or banned without warning during this outage. Hopefully this is resolved as service returns.
jablongo··on Claude Composer
I have a deep background in music and I think that while the creation was super basic, the way the output was so unconstrained (written by a model fine-tuned for coding), is really interesting. Listen to that last one and tell me it couldn't belong on some tv show. I've had always issues with any ai generated music because of the constraints and the way the output is so derivative. This was different to me.
jablongo··on Claude Composer
I've got to come to the OPs defense as well. This was a remarkable demonstration of Claude performing a task thats probably very out of distribution. This would not be interesting if it were a music generation model or program, it's interesting because this is not what Claude code was explicitly trained for. The fact that it generated waveforms from scratch and built up from there is really amazing. Your cynicism was applied before even reading the article.
jablongo··on I let ChatGPT analyze a decade of my Apple Watch data, then I called my doctor
There needs to be more documentation about what info was provided to the LLM and in which format before we decide that LLMs are necessarily bad at this. That said, you would expect the offering from a $500bn company to be more robust and better tested than this, assuming this is reported accurately.
jablongo··on Proof of Corn
I think you could have credibly said this for a while during 2024 and earlier, but there is a lot of research that indicates LLMs are more than stochastic parrots, as some researchers claimed earlier on. Souped up versions of LLMs have performed at the gold medal level in the IMO, which should give you pause in dismissing them. "It can't actually understand how many there currently are, the season, how much land, etc, and do the math itself to determine whether it's actually needed or not" --- modern agents actually can do this.
jablongo··on Health care data breach affects over 600k patients, Illinois agency says
Interestingly in healthcare there is a correlation between companies that license/sell healthcare data to other ones (usually they try to do this in a revokable way with very stringent legal terms, but sometimes they just sell it if there is enough money involved) and their privacy stance... and it's not what you would think. Often it's these companies that are pushing for more stringent privacy laws and practices. For example, they could claim that they cannot share anonymized data with academic researchers, because of xyz virtuous privacy rules, when they are actually the ones making money off of selling patient data. It's an interesting phenomenon I have observed while working in the industry that seems to refute your claim that "there's no money in privacy". Another way to think about it is that they want to induce a lower overall supply for the commodity they are selling, and they do this by championing privacy rules.
jablongo··on Trump says Venezuela’s Maduro captured after strikes
Trump received 77.3M votes while Kamala received 75M. Since the total was 156.7M it was barely a plurality instead of a majority (just under 50%).
jablongo··on Trump says Venezuela’s Maduro captured after strikes
I don't think the US has "shitty class relations". Most of your complaints are true, but in the US social class is mutable and upgrades to social class are encouraged and celebrated (even though this is becoming much more difficult in practice). Contrast this with Europe and other parts of the world with entrenched aristocracies and castes that survive generations. There are major problems but social mobility is still relatively better in US; in Europe healthcare is way better and being on the bottom rung isn't as bad, but fewer people from the bottom make it to the top.
jablongo··on Trump says Venezuela’s Maduro captured after strikes
Would the Rawlsian say this is unacceptable?
jablongo··on Evaluating chain-of-thought monitorability
It is what it is thinking consciously / its internal narrative. For example a supervillain's internal narrative with their plans would go into their COT notepad. If we want to really lean into the analogy between human psychology and LLMs. The "internal reasoning" that people keep referencing in this thread.. referring to the transformer weights and inscrutable inner working of a GPT.. isn't reasoning, but more like instinct, or the subconscious.
Page 1 of 7Next →