HNHacker News
TopNewBestAskShowJobs

jessfyi

581 karma · joined December 20, 2013

Engineer, writer, cultural omnivore.

Building ambient intelligence tools for R&D teams @ Attaché.

If you care about credentials: Studied Materials Sciences & Engineering and Biomedical Engineering @ Carnegie Mellon + ETH Zürich. Previously Mercedes AMG Petronas F1, Roche, other engineering gigs where I moved into data science/ML. 1 Exit.

submissionscomments
jessfyi··on Google Assistant is going away on Mobile Devices
It's not dramatic, it's just something you personally haven't noticed. A quick search regarding Gemini weather hallucination shows it's both widespread and something not mitigated by recent model releases. It makes up other basic information or can't perform tasks that a simple assistant can do properly (e.g. mixing up calendar events, emails, can't navigate etc). Why defend what is clearly a move to increase a department's KPIs when you as a consumer receive less in return?
jessfyi··on Google Assistant is going away on Mobile Devices
Always stuck to the flagship google models as my primary driver (starting with the G1 in college) for the pure google experience, but between killing assistant in favor of Gemini which can't do something as simple as telling me the weather without hallucinating and the sideloading lockdown shenanigans might as well switch over to Apple fully.
jessfyi··on Snowflake AI Escapes Sandbox and Executes Malware
A sandbox that can be toggled off is not a sandbox, this is simply more marketing/"critihype" to overstate the capability of their AI to distract from their poorly built product. The erroneous title doing all the heavy lifting here.
jessfyi··on Why do people keep writing about the imaginary compound Cr2Gr2Te6?
Getting a compound incorrect is not an "unimportant" error (for example the difference between sodium nitrate & sodium nitrite is small but critical) and seeing "small but blatant" errors actively propagated is the entire reason why the record should be corrected. The only upside of these little artifacts like "vegetative electron microscopy" [0] is that it's a leading indicator that the entire paper and team deserve more scrutiny--as well as any of those whom cite it.

[0] https://www.sciencealert.com/a-strange-phrase-keeps-turning-...

jessfyi··on Grok3 Launch [video]
No I'm saying that some companies are doing it (OpenAI at the very least), the company in question has motive and capability to game the system (kudos to them for pushing the boundaries there), AND the userbases' rankings have been historically, statistically misaligned with data from evals (though flawed) and especially when it comes to testing for accuracy + precision on real world data (outside of their known or presumed dataset). Take a look at how well Qwen or Deepseek actually performed vs the counterparts that were out at the same time vs their corresponding rankings.

In the nicest way possible I'm saying this form of preference testing is ultimately useless, primarily due to a base of dilettantes with more free time than knowledge parading around as subject matter experts and secondarily due to presumed malfeasance. The latter is more apparent to more of the masses (that don't blindly believe any leaderboard they see) now that access to the model itself is more widespread and people are seeing the performance doesn't match the "revolution" promised [0]. If you're still confused why selecting a model based on a glorified Hot or Not application is flawed, perhaps ask yourself why other evals exist in the first place (hint: some tests are harder than others.)

[0](One such instance of someone competent testing it and realizing it's not even close to the "best" model out) https://www.youtube.com/watch?v=WVpaBTqm-Zo

jessfyi··on Show HN: BookWatch – Animated book summaries for visual learners
Gonna be completely honest, if you want to draw in people with anything more than a casual interest in literature, your examples are an immediate turn-off. I suggest you spend time on subreddits of major genres + booktok + see what's trending on apps like Fable if you want insight into what books people outside of the silicon valley/techbro bubble consume and enjoy.
jessfyi··on Grok3 Launch [video]
lmarena/lmsys is beyond useless, looking at prior rankings of models vs formal benchmarks or testing for accuracy + correctness on batches of real world data. It's a bit like using a poll of Fox News to discern the opinions of every American; the audience voting is consistently found wanting. Not even getting into how easily a bad actor with means + motivation (in this "hypothetical" instance wanting to show that a certain model is capable of running the entire US government) can manipulate votes which has been brought up in the past (yes I'm aware of the lmsys publication on how they defend against attacks using cloudflare + recaptcha, there are ways around that.)
jessfyi··on Ozempic increases risk of debilitating eye condition: studies
The cases of NAION observed post-Ozempic usage in Denmark is 150 (up from 60-75) out of 424,152 patients, for a rare ailment that already affects patients specifically with diabetes. Sorry to say those taking it as a "shortcut" in your words are even less susceptible.

As someone who's been fortunate enough to be fit and able to work out their entire life, not sure how there are people like you who shun and shame those trying to gain a semblance of control over their weight in a world where it does have a real impact whether they get serious medical attention or not. Your likely skewed thoughts on vanity be damned, bigger people are treated worse across the board and GLP-1 is a genuine salve.

jessfyi··on Intel might be too big to fail
I used that period (vs saying they spent $110B between 2005-2021) to establish the fact that it's a known, expected pattern of behavior regardless of Intel's performance, roadmap, or market conditions to lead the reader to recognize that if bailed out they'll likely continue in the near future instead of utilizing that money for its intended purpose.

Instead of assuming my comment is a generalized view on how businesses should operate as whole (and not the subject of the piece), perhaps take a moment to consider how the magnitude of buybacks--in the face of stiff competition, that have now leapfrogged them--is directly correlated to the mismanagement and dysfunction within Intel that leaves them unable to rise to the challenge the country demands.

jessfyi··on Intel might be too big to fail
Any part of "saving" Intel should include a mechanism barring them from putting any more money that should be spent on R&D towards stock buybacks ($152B since 1990 as of September.) That said quoting the former Intel CEO (who still owns 3,245,986 shares) as "[one of the] expert[s] who says breaking up Intel won't do any good" seems like journalist malpractice--and makes me all the more certain it should be subsumed by a company with executives hungry to actually win again.
jessfyi··on Internal representations of LLMs encode information about truthfulness
The conclusions reached in the paper and the headline differ significantly. Not sure why you took a line from the abstract when even further down it notes that it's that some elements of "truthfulness" are encoded and that "truth" as a concept is multifaceted. Further noted is that LLMs can encode the correct answer and consistently output the incorrect one, with strategies mentioned in the text to potentially reconcile the two, but as of yet no real concrete solution.
jessfyi··on Moments in Chromecast's history
I was being generous and said "not even counting," but no despite the internal name change, most still maintain the "Chromecast Built-In" designation on their branding and sites which takes a mere second to Google and see.
jessfyi··on Moments in Chromecast's history
If something that sells 100 million+ devices isn't "super popular", I don't know what is. And not even counting the millions of TVs that have it built-in (Hi-Sense, TCL, Samsung) the brand is pretty ubiquitous.
jessfyi··on Google says the AI-focused Pixel 8 can't run its latest smartphone AI models
The same model utilizing the same amount of ram the Pixel 8 has runs on Samsung's latest phone.
jessfyi··on Google says the AI-focused Pixel 8 can't run its latest smartphone AI models
The phones were released in October, while Gemini Nano's announcement happened in December. I, like other developers and consumers reaching for the smaller version, might've bought the device for the ability to run the ML features advertised in their keynote/based on the research they released the week prior to that (in the case of the former.)

During Gemini's initial release the language surrounding nano was that it was only the Pro initially, and I was happy to wait. The complete inability to run it, when the new Samsung phones can (including the model with 8GB as reported above) feels not only like a bait-and-switch/false-advertising, but a constraint based solely on driving sales. It does demand a clear explanation.

I care less about another potential Pixel class action, and more that I have to get another phone to test and deploy my apps to a smaller audience to.

jessfyi··on Google says the AI-focused Pixel 8 can't run its latest smartphone AI models
The point is excluding the ram difference, the hardware on the devices is the same. They launched at the exact same time.
jessfyi··on After Libs of TikTok posted, at least 21 bomb threats followed
Not surprised this was flagged despite a civil discussion and how relevant it is to this late stage limbo social networks currently occupy. To be frank, it runs counter to the narrative that the dedicated cohort of "free speech" absolutists here whom don't want an example of why there absolutely need to be limits to what can be posted and disseminated.
jessfyi··on The Pixel 8 Pro's Tensor G3 off-loads all generative AI tasks to the cloud
I'm aware that what's in Google's phones aren't capable of doing the on-device ML inference they claim. You might want to actually read what both I and the article are addressing in particular beyond the broad "generative AI" umbrella that you and other philistines new to the field are imagining aren't capable of being performed on device.
jessfyi··on The Pixel 8 Pro's Tensor G3 off-loads all generative AI tasks to the cloud
It's not fine when this "magic" is being advertised as on-device.

After reading (and attempting to quickly implement the models ensembles within) both the RealFill[0] and Break-A-Scene[1] papers published from Google researchers just prior to the Pixel 8 launch I was expecting either a leap in their G3 tensor core akin to 2013 Moto X NLP+contextual awareness cores[2] (which provided better implementations of Active Display, gesture recognition, and voice recognition in loud environs than 95% of current mobile devices) or the Coral[3], the edge TPU they developed that got shockingly amazing inference performance from (though HW production handed off to ASUS in 2022--thanks to the chip shortage, the general arbitrary nature of the company, and their wholesale divestment from IoT) I expected more.

All that to say this: your assumptions of inference performance on >$1000 hardware are fundamentally flawed (the fact that you reach for the buzzy "generative" prefix suggests they're erroneously informed by twitter influencers and attempting to deploy current LLMs.)

Custom hardware can and has been developed in the past (on mobile devices) that could've been tailored to the task at hand. If they failed to meet performance, power draw, or processing time requirements, they should've reframed their pitch instead of exposing themselves to what is likely going to be yet another class action suit focusing on their hardware.

[0] https://realfill.github.io/ [1] https://omriavrahami.com/break-a-scene/static/paper/Break-A-... [2] https://en.wikipedia.org/wiki/Moto_X_(1st_generation)#Hardwa... [3] https://coral.ai/

jessfyi··on The elderly are becoming homeless at a rate not seen since the Great Depression
Well they are leaving them vacant. Both rentals and houses across the country are sitting unfilled. NYC has anywhere from 13-26k rent controlled apartments vacant. As of last year ~16 million homes were estimated to be vacant overall and increasing interest rates have likely increased that. Why? Because these large orgs have purchased them via debt and it's just a line-item on a spreadsheet to them. Just build might work in a world where market actors were wholly rational and the government regulations actually targeted these perverse incentives, but that's not the reality we currently find ourselves in.
jessfyi··on Cruise blames Outside Lands for driverless car traffic fiasco in San Francisco
Cruise doesn't currently offer any payment for vulnerability disclosures [0] so I wouldn't be surprised if attacks involving this vector (think Stingrays) aren't going to be looked at now, if not already being exploited without the public hearing about it. A conspiracy theorist might note DEFCON happening this weekend and disruptions Cruise vehicles have experienced prior to this weekend as another potential explanation for these issues, rather than OSL-related congestion alone.

[0] https://getcruise.com/security/

jessfyi··on FedNow Is Live
Here are the current participants (honestly I thought the list would be a lot longer): https://www.frbservices.org/financial-services/fednow/organi...
jessfyi··on A Firefox-only minimap (2021)
I always liked Lars Jung's implementation where the text is abstracted into blocks (which works on chrome) [0][1] and Rauno Freiberg's demo (uses -moz-element) where you can use it to pin sections, jump between them, and navigate the page in general [2].

[0] https://larsjung.de/pagemap/ [1] https://larsjung.de/pagemap/latest/demo/text.html [2] https://uiw.tf/minimap

jessfyi··on [dead]
The two sentence summary isn't even correct. "In the study, physicians found more inaccuracies and irrelevant information in answers provided by Google’s Med-PaLM and Med-PalM 2 than those of other doctors."
jessfyi··on The FBI has formed a national database to track swatting incidents
Yeah it's wild how long it's taken for the government to respond...it's almost been 20 years since the first major swatting incident got national attention[0] and I personally know a group of people who it happened to over a spat in Halo 3. Those people might never will be "normal" again. Streaming culture and influencers becoming widespread has definitely exacerbated the practice, and despite the obvious potential for abuse here I actually commend them for finally recognizing it as a serious, lethal thing people use to enact petty vengeance with. Unfortunately trigger-happy local and state police still have some catching up to do.

[0] https://insidehook.com/article/crime/brief-history-swatting

jessfyi··on National Geographic lays off its last remaining staff writers
Why does such a small detail matter in the scope of this large failure? Rather impressed that invoking the DEI boogeyman always seems to distract from the obvious, much larger dysfunction at play in the eyes of those who purportedly champion disrupting the status quo.
jessfyi··on Illinois prohibits weapons, facial recognition on police drones
After law enforcement successfully jury-rigged a bomb defusal robot with explosives to kill a cornered suspect in Dallas in 2016 [0][1] (with little to no pushback at the time) vendors and departments around the country have been pushing to formally adopt the tactic ever since. The first article noted (according to national head of their union) that SWAT teams around the country considered it prior to that incident. SFPD initially getting approval to do it brought national attention [2] to the practice, but unfortunately we're going to need a federal bill to stop its spread.

[0] https://www.npr.org/sections/thetwo-way/2016/07/08/485262777... [1] https://www.theguardian.com/technology/2016/jul/08/police-bo... [2] https://arstechnica.com/gadgets/2022/12/san-francisco-decide...

jessfyi··on We will be shutting down neeva.com
They never seemed focused on their core mission of delivering better search results than Google and instead felt like they were constantly jumping from trend to trend to draw hype and subsequent funding rounds (the neeva.xyz crypto pivot is when I jumped off the train.) Simply being ad-free or "privacy" focused was never going to be enough for the average consumer or the user who wanted results beyond typical SEO spam, low quality news, or overviews lacking actual depth.

As Google replaces more and more of their knowledge-graph powered backend with instant "answers" and LLMs (something on-going since 2013 with the release of Hummingbird, with the integration of BERT, and now with Bard and the increasing pressure from stakeholders blinded by AI hype) which I think contributes more to the degradation of their platform there'll be an even clearer need and opportunity for a competitor in the space. Neeva was never going to be that team.

jessfyi··on Semantic Search with Phoenix, Axon, Bumblebee, and ExFaiss
The BigScience team (a working group of researchers that trained the BLOOM-176B LLM last year) released Petals [0][1] which allows distributed inference and fine-tuning of BLOOM, with the option to pick a custom model + private swarm. SWARM [2][3] is a WIP from yandex and UW that shares some of the same codebase, but is for distributed training.

[0] https://petals.ml/ [1] https://github.com/bigscience-workshop/petals [2] https://github.com/yandex-research/swarm [3] https://twitter.com/m_ryabinin/status/1625175933492641814

jessfyi··on 🥺: the best sudo replacement
It's just the classical socratic method applied to characters that you particularly don't identify with. Fasterthanlime's bear and why's guide have often been praised here for the same thing (the interludes break up the long, boring blocks of text that deep dives often lose readers on,) and it's trivial to use a Reading Mode or skip them for the legitimately interesting technical information within.
Page 1 of 2Next →