HNHacker News
TopNewBestAskShowJobs

vipshek

620 karma · joined October 11, 2012

vipshek.com
submissionscomments
vipshek··on OpenAI bots knew about the RubyGems caching vulnerability
I agree with one of the sibling comments that determinism isn't necessary for certifying a product. All engineered products operate under uncertain conditions; we define standards for how those products ought to respond under those conditions and verify them under measurement. Consider robot vacuums, for example.

I also agree that qualitatively, this technology seems different than the others. However, I feel that people tend to overly fixate on their internal stochasticity. Even if LLMs' internal mechanism is nondeterministic, shouldn't we be able to verify their "side effects" aren't harmful? Of course, "harm" is subjective and at this scale, the most effective way to verify behavior is probably some kind of LLM-as-judge...

Anyway, in this case the problems have occurred while actually running the evals themselves, so again, we're in a situation where we can't even confidently test these things and know that they won't cause harm in the outside world.

vipshek··on OpenAI bots knew about the RubyGems caching vulnerability
In the physical world, it seems like when an tool/device/instrument causes harm (or is used to cause harm), we assign blame to either the user of the tool or its creator.

When do we blame the user? When the tool is operating as intended by its creator, and we agree the tool meets certain quality standards and isn't defective.

When do we blame the creator? When the device doesn't meet those quality standards and reasonable use caused harm inadvertently. For example, for consumer devices, certifications like UL/CE are used to define acceptable performance levels and safety standards.

Maybe we need "quality certifications" for AI agents - essentially eval suites that demonstrate those agents won't cause harm under reasonable patterns of usage. Right now, these eval suites are run best-effort by the labs themselves.

The tricky thing is, a lot (all?) of these recent safety incidents have occurred while evaluating these models! This suggests we need much more rigorous standards for how exactly an eval can be run. Perhaps all of them should occur in truly air-gapped environments... though that may run counter to evaluating agents in a realistic way.

Regardless, it feels like the "industry standards" common in, say, electrical engineering and other disciplines are sorely lacking here. Unsurprising given how new these technologies are, but concerning since the blast radius for this technology is likely much larger than other technologies we've encountered in the past, except maybe nuclear technology.

vipshek··on I should have loved biology (2020)
This article is ostensibly about biology but it's really about pedagogy—how traditional education squeezes out the sense of discovery and turns many subjects into rote memorization exercises.

This reminds me of the pedagogical philosophy of Seymour Papert, who was heavily influenced by the ideas of Jean Piaget. Piaget's "genetic epistemology" argues that knowledge and understanding is created by interacting with environments, which traditional educational approaches fail to provide.

Papert combined Piaget's ideas with the emergence of computing to argue that children should be taught subjects in a hands-on, exploratory way - and not just for teaching computing! The idea is that by programming in simplified languages, children can discover ideas in subjects like mathematics and grammar - deriving them as they try to solve problems instead of having them dictated to them.

Suppose that biology was taught in a game-like environment where students were designing cells or organisms in some way. Maybe the game could be structured so that each organelle could be "discovered" by the student as they designed the cell to survive in some environment. Perhaps that'd make the purpose of each cell component more grounded and memorable.

Papert's book Mindstorms covers all this in detail, and I highly recommend reading it. I feel like it's especially important today given fears of how AI will affect childhood education. At best, I hope that computing can be a boon to education instead of a detriment if it's woven into pedagogy thoughtfully.

vipshek··on Delta
This is intriguing. The two relevant features seem to be 1) realtime collaborative multiplayer conversations and 2) conversation-as-document - basically, letting you comment inline in an agent conversation.

For (1), the main value I'd see is in mentoring junior engineers or less technical contributors on a team. If someone puts up a PR with sloppy results, you could actually jump into the thread that produced that PR and see how the results came about, or even coach that contributor on how to do better next time. Also might make it easier to hand off work from one person to another - right now most coding agent sessions are user-local.

On (2), I frequently find myself consuming agents' gigantic text responses and tediously writing 8-bullet-point responses to guide them. It's pretty exhausting. I could see inline comments providing much better ergonomics.

All that being said, Zed has largely fallen out of the conversation for "agentic coding tools", and so this feels like their attempt at creating something like the Cursor Agents Window, Codex, or Claude Code. These two features seem compelling, and I understand they're even compatible with other coding harnesses. But I don't know if there's enough there to have a defensible product; if these features are excellent, others will clone them eventually.

Regardless, would love to give this a shot!

vipshek··on Show HN: Gaussian Splat of a Strawberry
This took me down a rabbit hole that led to this company, doing Gaussian splat videos: https://www.4dv.ai/. Fascinating.
vipshek··on I moved my digital stack to Europe
> This website has been temporarily rate limited

Feels a bit ironic... though this website is hosted on Cloudflare Workers so using an American company anyway?

vipshek··on Ask HN: What are you building that's not AI related?
Very cool! Have you seen https://encore.dev/ ? Haven't used it personally but I saw it on HN last year and have been meaning to try it out.

Seems like your approach is a bit more "batteries-included" but I'd curious for your thoughts on the differences.

vipshek··on Ask HN: What are you building that's not AI related?
https://exploretrees.nyc/

Been tinkering on this Olmsted-inspired map of all the trees in NYC for a long time - just added some more features (seasonal colors and better search/filtering) recently.

If you're in NYC, try finding some cherry blossoms near you: https://exploretrees.nyc/?species=cherry

vipshek··on Your data model is your destiny
I don't have much to say about this post other than to vigorously agree!

As an engineer who's full-stack and has frequently ended up doing product management, I think the main value I provide organizations is the ability to think holistically, from a product's core abstractions (the literal database schema), to how those are surfaced and interacted with by users, to how those are talked about by sales or marketing.

Clear and consistent thinking across these dimensions is what makes some products "mysteriously" outperform others in the long run.

vipshek··on OpenAI to buy AI startup from Jony Ive
I find this perspective bizarre. Though I'm not happy about it all being centralized, the closest thing we have these days to the very niche phpBB forums of the 2000s is various subreddits focused on very specific topics. Scrolling through the front page is slop, sure, but whenever I'm looking for perspectives on a niche topic, searching for "<topic> reddit" is the first thing I do. And I know many people without any connection to the software industry who feel the same way.
vipshek··on AI Blindspots – Blindspots in LLMs I've noticed while AI coding
Perhaps swearing at the LLM actually produces worse results?

Not sure if you’re being figurative, but if what you wrote in your first comment is indicative of the tone with which you prompt the LLM, then I’m not surprised you get terrible results. Swearing at the model doesn’t help it produce better code. The model isn’t going to be intimidated by you or worried about losing their job—which I bet your junior engineers are.

Ultimately, prompting LLMs is simply a matter of writing well. Some people seem to write prompts like flippant Slack messages, expecting the LLM to somehow have a dialogue with you to clarify your poorly-framed, half-assed requirement statements. That’s just not how they work. Specify what you actually want and they can execute on that. Why do you expect the LLM to read your mind and know the shape of nginx logs vs nginx-ingress logs? Why not provide an example in the prompt?

It’s odd—I go out of my way to “treat” the LLMs with respect, and find myself feeling an emotional reaction when others write to them with lots of negativity. Not sure what to make of that.

vipshek··on Microsoft cancels leases for AI data centers, analyst says
I would like to propose a moratorium on these sorts of “AI coding is good” or “AI coding sucks” comments without any further context.

This comment is like saying, “This diet didn’t work for me” without providing any details about your health circumstances. What’s your weight? Age? Level of activity?

In this context: What language are you working in? What frameworks are you using? What’s the nature of your project? How legacy is your codebase? How big is the codebase?

If we all outline these factors plus our experiences with these tools, then perhaps we can collectively learn about the circumstances when they work or don’t work. And then maybe we can make them better for the circumstances where they’re currently weak.

vipshek··on Ask HN: Who is hiring? (January 2025)
Meridian | Founding Engineers (Product, Infra) | NYC, New York (In-person) | https://careers.meridian.tech | Full-time

Meridian develops software to accelerate the next generation of companies building in the physical world across aerospace, defense, automotive, robotics, and more. We automate the administrative work of quality and compliance to help our customers go to market faster, scale their production, and increase their pace of innovation.

Meridian is 3 months old. We’ve already signed paying customers, built and launched our product, and raised an oversubscribed pre-seed round.

For our three first hires, we’re looking for world-class generalist engineers who can ship great product experiences fast while laying the foundations for a platform that will scale to large and complex enterprises in the future. We're offering competitive salaries and above-market equity.

We're building an in-person engineering team that prides itself on shipping excellent products for a user segment (quality engineers in manufacturing) that's been sorely neglected in the past. We ship with speed and quality, own a large product surface area, and are relentlessly customer-focused.

To apply, send us your resume and anything else you’d like to careers@meridian.tech.

vipshek··on NeuralSVG: An Implicit Representation for Text-to-Vector Generation
This is excellent!

I think the utility of generating vectors is far, far greater than all the raster generation that's been a big focus thus far (DALL-E, Midjourney, etc). Those efforts have been incredibly impressive, of course, but raster outputs are so much more difficult to work with. You're forced to "upscale" or "inpaint" the rasters using subsequent generative AI calls to actually iterate towards something useful.

By contrast, generated vectors are inherently scalable and easy to edit. These outputs in particular seem to be low-complexity, with each shape composed of as few points as possible. This is a boon for "human-in-the-loop" editing experiences.

When it comes to generative visuals, creating simplified representations is much harder (and, IMO, more valuable) than creating highly intricate, messy representations.

vipshek··on Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku
Install Cursor (https://cursor.com), go into Cursor Settings and disable everything but Claude, then open Composer (Ctrl/Cmd + I). Paste in your exact command above. I bet it’ll do something pretty close to what you’re looking for.
vipshek··on AlphaCodium outperforms direct prompting of OpenAI's o1 on coding problems
I've completely switched over to Cursor from Copilot. Main benefits:

1. You can configure which LLMs you want to use, whereas Copilot just supports OpenAI models. I just use Claude 3.5 for everything.

2. Chatting with the LLM can produce file edits that you can directly apply to your files. Cursor's experimental "Composer" UI lets you prompt to make changes to multiple files, and then you can apply all the changes with one click. This is way more powerful than just tab-complete or a chat interface. For example, I can prompt something like "Factor out the selected code into a new file" and it does everything properly.

3. Cursor lets you tune what's in LLM context much more precisely. You can @-mention specific files or folders, attach images, etc.

Note I have no affiliation whatsoever with Cursor, I've just really enjoyed using it. If you're interested, I wrote a blog post about my switch to Cursor here: https://www.vipshek.com/blog/cursor. My specific setup tips are at the bottom of that post.

vipshek··on DOJ sues realpage for algorithmic pricing scheme that harms renters
Which state is your town located in, out of curiosity? I'm trying to build a mental rolodex of which states have towns that are development-friendly.
vipshek··on Ask HN: Why are AI generated images so shiny/glossy?
Ah, my mistake. "Meta AI" can generate both text and images, but apparently text prompts are handled by Llama 3.1 while image prompts are handled by Emu. I initially struggled to find the name of the image generation model.
vipshek··on Ask HN: Why are AI generated images so shiny/glossy?
Many AI-generated images you encounter are low-effort creations without much prompt tuning, created using something like DALL-E or Llama 3.1. For whatever reason, the default style of DALL-E, Llama 3.1, and base Stable Diffusion seems to lean towards a glossy "photorealism" that people can instantly tell isn't real. By contrast, Midjourney's style is a bit more painted, like the cover of a fantasy novel.

All that being said, it's very possible to prompt these generators to create images in a particular style. I usually include "flat vector art" in image generation prompts to get something less photorealistic that I've found is closer to the style I want when generating images.

If you really want to go down the rabbit hole, click through the styles on this Stable Diffusion model to see the range that's possible with finetuning (the tags like "Watercolor Anime" above the images): https://civitai.com/models/264290/styles-for-pony-diffusion-...

vipshek··on Tesla Owners Get Only 64% of EPA Range After Just Three Years: Study
Just to clarify, are you saying these are relatively new BEV trucks coming into your shop? If so, do the mileage issues usually boil to battery degradation due to driver behavior as you’re suggesting, or is it because the EPA rated mileage was unrealistic in the first place? Or both?
vipshek··on 4.7 Earthquake in NYC
Dupe: https://news.ycombinator.com/item?id=39942880
vipshek··on On limitations that hide in your blindspot
I was a software intern at Bridgewater in 2013. My impressions from that time are that Bridgewater had an interesting, unique, and very intentional culture, but at some point the firm grew and Dalio started thinking about how to scale that culture. He wrote the Principles book and had everyone read and discuss it; the Dots app was implemented; etc.

The core cultural values seemed reasonable to me, but the efforts to scale the culture felt heavyhanded and seemed like they sometimes backfired. Attempting to quantify someone's "believability" based on subjective data collected on an iPad app was a bit silly and easily gameable. The fact that Dalio and other executives scored highest on basically every facet was an obvious sign that the system was a bit farcical.

Maybe the company managed to solve some of those cultural scaling challenges after 2013? I wasn't around to see what happened after.

vipshek··on Ask HN: How to onboard yourself to a new product/industry in a new job?
In the past year, I've gone from near-zero understanding to pretty deep expertise in the domain of electrical grid interconnection as a product engineer. Here's how I went about it.

I think all the other comments saying "read a lot" and "talk to everyone" are correct first steps, but for me, consuming information has diminishing returns after a short while. After you've reached a point where your brain feels like it's exploding, you should switch your focus to outputting information.

If you're a "write things down" person, then write a synthesized document explaining everything you've learned, and then ask a few trusted coworkers to tear it apart.

If you're a "talk out loud" person, schedule time with coworkers to have a "teachback session" where you give a presentation about everything you've learned. Again, ask them to tear it apart.

It's crucial to build trust with a few coworkers who are willing to critique your output. Get them to rip everything you've created to shreds. Whenever you write or say something that's even slightly off compared to how someone in the industry would say it, make sure you learn about that, and learn how someone in the industry would say it.

This focus on getting the language right - especially the colloquial language of how people actually describe things day-to-day - is important for every role, but I assume it's especially important in marketing, where you need to be able to use the precise language that your customers use.

tl;dr: Read/talk to people at first, but switch to writing/presenting ASAP. Solicit and internalize as much critical feedback as you possibly can.

vipshek··on How the biggest plane would supersize wind energy
From the article:

> Radia estimates the larger turbines could reduce the cost of energy by up to 35% and increase the consistency of power generation by 20% compared with today’s onshore turbines.

Not sure what that translates to in terms of energy output over time.

As for the sibling "Why not airships?" question, the article says:

> Blimps can’t land in windy conditions. Helicopters are more costly than airplanes, and flying with a dangling blade designed to catch wind would prove complex and dangerous.

vipshek··on What Extropic is building
I have no idea about the merits of this approach, but I found this interview with the founders a lot more sensical than the linked article:

https://twitter.com/Extropic_AI/status/1767203839818781085

vipshek··on Bootstrap or VC? [video]
Completely agree there are many other reasons people don't join large, established bureaucratic organizations. My career has involved building software within well-established industries, but not necessarily within the large bureaucracies themselves. There is plenty of room for nimble software companies in those industries, but little of it fits the venture-backed model because it'll likely never achieve a 100x return. That's the sort of thing that gets suffocated by an overemphasis on VC.

I think your point about "inspirational stories" hits at the heart of something important. I believe there are plenty of successful software companies working in niche industries and having great impact, but their stories don't get amplified nearly as much as venture-backed companies do, partially because raising VC is viewed as a stamp of success that makes it easier to get press coverage, hire talented people, etc.

If there's one thing that I think needs to change, it's that we need to tell more inspiring stories about people who've achieved success without needing to raise VC. That'll also reduce pressure on VC to be the one-stop-shop for everyone's business ideas.

(Also, hi Vinay! Not sure you remember me, but I used to hang out on the "codechill" Slack a couple years ago. Hope all's well.)

vipshek··on Bootstrap or VC? [video]
This discussion and the points being made are totally valid and reasonable. But I think there's one thing they completely miss, which is the cultural impact of VC.

Given the massive success of software companies over the past decade or two, hordes of driven, talented, and smart young people have joined VC-backed companies or aspire to do so. It's become culturally accepted that if you're an ambitious person, you should be working at - or better yet, founding - a venture-backed company.

The societal opportunity costs of this phenomenon are significant. I know dozens of young, talented people who are founding companies for no particular reason, simply because it's become the default path for what they're supposed to aspire to.

Is this the "fault" of venture capital? Not really - VCs are just trying to attract talented founders and make their portfolio companies successful.

But by doing so, they've created a negative externality, which is that VC has become the gatekeeper of perceived "success." That means fewer talented, high-energy people working in government institutions or critical industries - precisely where those talented people could have the greatest impact on society at large.

vipshek··on We have reached an agreement in principle for Sam to return to OpenAI as CEO
"Stronger" is ambiguous. If you interpret it as "resilience" then I agree having a single point of failure is usually more brittle. But if you interpret it as "focused", then having a single charismatic leader can be superior.

Concretely, it sounds like this incident brought a lot of internal conflicts to the surface, and they got more-or-less resolved in some way. I can imagine this allows OpenAI to execute with greater focus and velocity going forward, as the internal conflict that was previously causing drag has been resolved.

Whether or not that's "better" or "stronger" is up to individual interpretation.

vipshek··on Show HN: I built a transit travel time map
This is super cool! I particularly like that you render the route along transit lines to reach each destination, which is differentiated from other isochrone maps I've seen. I would love to use this while hunting for apartments in a large city.

I do think there are a number of things you could do to improve the UX here. Hope this doesn't come across as too harsh, but here are some suggestions...

1. Double-click to set origin makes sense, but hover to set a destination is a bit weird, for two reasons: a. There's no way to "lock" a destination by clicking, so I can't pick a route I want to see and fix my view on it. The UI feels jittery as a result. b. In some cases (I'm looking at NYC), loading the route to the hovered destination hangs for several seconds. I don't know if this is just an issue because your server is overloaded, but it's a weird state to be in.

I'd consider: 1) adding a loading state when the route to the destination is being computed, 2) enable clicking a point to lock viewing that route, 3) maybe disabling the hover interaction entirely, though if it were performant, it would be pretty nice to have.

2. The "arrival time" concept is a slightly odd. I'd prefer to just see the amount of time it'll take to get to a destination, rather than an arrival time based on a particular starting time. I don't think anyone's going to use this to plan a specific route at a specific time; instead, they'll use it to explore locations of interest, and explore how long it'll take to get to various destinations. I see you show the trip duration in the bottom-left, but having it be an arrow pointing to a gradient spectrum is much harder to parse than just saying "27 minutes".

3. The fade effect on the side panel is a bit weird. It does draw more attention to the map, but the side panel is still occluding my view, and having it faded just makes the text very hard to read. I'd consider making it un-faded, but adding some way to collapse it.

vipshek··on Beyond introvert vs. extrovert
I did briefly discuss this idea of interactions "draining energy" in the post here: https://www.vipshek.com/blog/interaction#balance-not-batteri....

I understand that the idea of draining vs. charging works for some people, but I always found this concept even more baffling. For me, when I've been in solitude for a while, I feel a restless desire to go out and interact with people. When I'm in that lonely state, I find social interaction "charging" and further solitude "draining."

But when I go interact with people for a long time—especially in a large group where I don't know people well—I feel an urge to retreat to solitude.

So am I charged by solitude and drained by interaction, or the other way around? It feels like the answer is "both"... so what am I supposed to do with all this?

One way you could think about my model is that you have four "batteries", each of which has an optimal "charge level." If you're under that level, you'll feel a desire to be in that state for a while. If you're over it, you'll feel an aversion to being in that state. And it's possible that if you're quite introverted, your optimal level of large-group interaction is close to zero.

If you're somewhere in the middle like me, though, the simple idea that you're "drained" by one thing or the other feels off.

Page 1 of 2Next →