HNHacker News
TopNewBestAskShowJobs

dvt

19,381 karma · joined October 14, 2012

UCLA alum (philosophy, mathematical logic), startup guy, data engineer, CTO, immigrant, amateur kayaker. Find me at https://dvt.name or @davvv on Twitter.
submissionscomments
dvt··on Ask HN: Who wants to be hired? (July 2026)
Location:Los Angeles

Remote: Yes

Relocation: Case-by-case

I'm an engineer and data professional interested in team-building, consulting and architecting data pipelines. At Edmunds.com, I worked on a fairly successful ad-tech product and my team bootstrapped a data pipeline using Spark, Databricks, and microservices built with Java, Python, and Scala.

At ATTN:, I re-built an ETL Kubernetes stack, including data loaders and extractors that handle >10,000 API payload extractions daily. I created SOPs for managing data interoperability with Facebook Marketing, Facebook Graph, Instagram Graph, Google DFP, Salesforce, etc.

More recently, I was the CTO and co-founder of a crypto gaming startup. We raised over $6M and I was in charge of building out a team of over a dozen remote engineers and designers, with a breadth of experience ranging from Citibank, to Goldman Sachs, to Microsoft. I moved on, but retain significant equity and a board seat.

I am also a minority owner of a coffee shop in northern Spain. That I'm a top-tier developer goes without saying. I'm interested in flexing my consulting muscle and can help with best practices, architecture, and hiring.

Would love to connect even if it's just for networking!

Blog: https://ai.dvt.name/whos-david/ (under construction)

GitHub: https://github.com/dvx

Email: [d]@[dvt].[name]

dvt··on HackerRank open sourced its ATS. My resume scored 90/100. Oh wait 74. No – 88
This is a very authoritative answer that should be more nuanced and caveated as implementation-dependent. In some cases, repetition penalties take precedence over sampling; top_k and top_p can also be handled before or after the temperature step. In other cases, `0` is turned into like 1e-10 or some super tiny float value (which can drift if you do any arithmetic with it). Routing, quantization, etc. can also have an effect on sampling. And yes, in some cases, setting temperature to 0 can mean "pure greedy decoding" which makes the decoder about as deterministic as it can get.
dvt··on HackerRank open sourced its ATS. My resume scored 90/100. Oh wait 74. No – 88
An alarming number of people don't understand that LLMs work via purely stochastic processes, so I'm happy to see in-depth pieces like this. I'm looking for a job and maybe this is why it's so hard to get a callback these days: resumes are just dumped in some LLM black hole and no one really knows how it works. The author says:

> temperature 0.1 — low, supposedly nudging the model toward deterministic outputs

This is not correct (and is briefly touched on later in the piece when he sets temperature to 0), temperature is not some kind of "deterministic" switch, but rather it affects the sampling distribution (which becomes more "spiky"—but is still very much a distribution).

dvt··on QSOE: QNX-inspired OS with dual-kernel architecture
This is one of the few AI hills I will die on: not disclosing AI tool usage when you produce a product where the writing is the end result (in this case, the website I'm being sent to) is disrespectful to your readers and users.

    `The quickest way to see QSOE run — no hardware, no -kernel juggling.`

    `Real hardware, real disk.`

    `A working plan, not a contract: milestones may shift as the work reveals what's really next.`
I have no problem with using AI to draft docs, or as an editing tool, or even to help writing (if, e.g., you are not a native speaker) but this is just egregious low-effort slop. If you can't even put the time to write your own documentation (or at least disclose AI tool use), why would I trust you to even test your own sofware?
dvt··on Anonymous GitHub account mass-dropping undisclosed 0-days
Went over a few of these with a pretty keen eye, and they aren't that particularly interesting. The Docker one is just a weird bug, it's not a vulnerability, and certainly not a "0-day" (which is a pretty loaded term and people expect bad stuff to happen).

The nghttp2 nghttpx one is more interesting, and could potentially be used for phishing, but it's very hard to line up properly because the request queue is non-deterministic so basically impossible to target a specific victim (assuming proxy traffic).

The VLC one is just a straight-up crash/bug. And VLC crashes all the time when using weird codecs, so that's nothing new.

Am I missing something here?

dvt··on Overfitted a 900KB Transformer to Compress a 100MB CSV into 7MB
Likely doable with metaparameter tuning (used to work on a team with data scientists that were routinely doing this in various situations). Seems like a cool idea.
dvt··on Show HN: Monolisa v3 – a typeface for developers and creatives
In a world where Fira Code, Hack, JetBrains Mono, and like a zillion others (of equal, if not greater, quality) are offered for free, this is obviously a pure marketing play and it's sad we live in a world where even fucking fonts are so heavily monetized.
dvt··on Big AI labs are hiring philosophers
Usually you need to be well-published/cited in the field, so a minor would likely not qualify. People joke around, but philosophers are some of the smartest people I've ever met, and it's not even particularly close. (I graduated ~10 years ago, so most of them are sadly lawyers or in academia these days, though some are engineers or entrepreneurs.)
dvt··on How to Passive-Aggressively Shame People Who Use LLMs Selfishly
The irony of someone writing a blog post purely to rage-bait (or rage-engage?) against "slop grenades." Slop grenades can also be made by humans you know, and I'd argue this post is at least a flashbang.
dvt··on Wolves are reconquering Europe. Can people learn to live with them?
> There's the killing again.

Do you not understand how ecosystems work? Or how populations of predator/prey are in various (usually) oscillatory states?

> Humans have already stolen so much from nature and the animals

I'm confused, are you saying humans are somehow not part of nature?

dvt··on In praise of memcached
> The problem is that Redis tries very hard to position itself as a persistent data store

What are you talking about? On their website, the top 3 use cases (under the Platform menu) are: caching, streaming, and session management. Literally all of these three are volatile.

dvt··on In praise of memcached
> Considering how complex and error prone this is, I don’t want it in my stack.

Have you ever used Redis before? I've literally never had to manage clustering or had any issues with it, and I've been using Redis for like 15 years (including for games where state had to live in multiple regions and could change on a 30- or 60-tick basis).

dvt··on In praise of memcached
> 1) Wrap your client library so that it's impossible to store anything without an expiry date. You don't want 6-months-old data suddenly coming up in your app!

No need for this client-side complexity, as you should be using `allkeys-lru`. FWIW, should likely be doing this anyway, as (generally speaking) all data stored in Redis is usually regarded as volatile because of what Redis actually is.

dvt··on In praise of memcached
Memcached is meant to be a lightweight memory cache, which makes sense, but contrary to the article's claim that "Redis is brought into a stack as a cache, and it is run with the assumption that people treat it that way"—I have very very rarely experienced this. Redis is brought into a stack because (most importantly!) it's fast and (almost as importantly!) because it's simple. I don't think this article is written by AI, but (and I'm trying to be charitable here), it's just like.. dumb.

> Dealing with memcached downtime is incredibly easy, because client libraries generally ignore connection exceptions. For instance, a simple get will just return the default value (or none) if the server is down.

This is a terrible idea in the context of things that might use Redis. If you use Redis with some kind of complex state (say, a document if you're working on a Notion clone, for instance), wtf even is a "default value"? In fact, I actually also want to know when the thing is down.

> Clustering memcached is wonderful, because memcached actually has no clustering built-in.

Yeah bro, this is yet another one of the reasons people use Redis: it handles consensus and clustering for you. What even is this article? It's a master class in straw-manning architectural decisions: most people use hammers as hammers, but screwdrivers make great hammers too, especially if you also need to screw stuff in! I mean.. technically true?

dvt··on Prompt Injection as Role Confusion
The paper is correct, but I think that anyone that knows anything about LLMs knows this:

> Role tags were a formatting trick that became the security architecture and the cognitive scaffolding of modern LLMs.

LLMs are basically some `f(x) → y` where x and y are strings. That's it. Nothing more to it. If you feed it private x (like secret keys) or do dangerous stuff with y (like running arbitrary non-sandboxed code), that's on you.

Also, roles were never really meant to be a "security architecture," they were just meant to (a) make training/fine-tuning easier, and (b) make conversational LLMs more useful.

dvt··on Apertus – Open Foundation Model for Sovereign AI
> Instead we have a majority of society that wants to see AI fail.

Do you talk to regular people? I work out of coffee shops routinely and literally like 90% of laptops have ChatGPT or Claude open. I was shocked at how many of my friends love the silliest of AI features (like Slack bot summarizing your day or your upcoming meetings), and a lot of decks, proposals, SOW's, etc. are (at least in part) generated with AI these days.

dvt··on Petition against Meta's employee training data collection for ML models
> Indexing by search engines is fully suppressed by robots.txt

Ah yes, the companies that have ignored robots.txt to scrape your website for 20+ years will now not totally, most definitely not ignore (wink wink) polite requests to not use your data for AI training. Also, haven't Meta employees been complicit in getting teenagers addicted to social media and violations of PII until they got caught?

Respect goes out to mathematicians and their Leiden Declaration, which is an actual level-headed approach given the complexities of AI training and usage.

dvt··on The AI Hate Progression
Totally correct, though I do think a contract was signed in the LEGO case you're referencing. So "common law 9/10ths" doesn't even come into play here because people signed stuff.
dvt··on Show HN: Are You in the Weights?
I have a unique last name (maybe that's why), but pretty much nailed it:

    David Titarenco
    Software engineer and open-source contributor

    340 strength · Top 20%

    GPT-5.5 says
    Software engineer and writer known for work
    on developer tools, systems, and programming-
    related articles.

    Claude Opus 4.8 says
    Software engineer and entrepreneur known for
    web/JavaScript development work and contributions
    to open-source projects and tech startup communities.
dvt··on The AI Hate Progression
You're confusing distribution with copyright. Even free flyers or public YouTube videos are copyrighted. If you draw a flyer then I use that flyer as a cover for a book I'm selling, you will likely win in court. (I say likely because it's up to the judge, but there's a lot of case law here.)

> permission you implicitly granted

There is nothing you implicitly grant anyone, except for ownership (which is explicit, because I gave the thing to you).

dvt··on The AI Hate Progression
Funny example, because if you create a flyer, you own the copyright to said flyer :) So if you create a flyer, then if someone else uses that flyer to make money, you can sue them and you will win in court (unless the derived work is transformative, critiques it, yadda yadda). And this is the kind of hair-splitting that can get you into trouble, because I think it's trivial that ChatGPT's training is certainly more transformative than Google's indexing/PageRank, but we're somehow more upset at the latter than we are at the former.
dvt··on The AI Hate Progression
I saw, I'm just re-emphasizing/clarifying.
dvt··on The AI Hate Progression
> Following the same logic, is there a fair difference between pure training and reoffering?

This is the kind of hair-splitting that I was trying to avoid (because, at the end of the day, there is no functional difference, is md5 okay, maybe Markov chains, just a very simple one-layer perceptron?). Once you take someone's copyrighted work and you do anything with it without consent, you're breaking some implicit trust.

However, obviously there's a lot of tension here: free speech. transformed works, copyright owners, profit making, etc., etc. That's why I don't think it's really that important to exactly figure out what consent was broken and when, but rather it's important to be forward-looking and plan for what might come next.

dvt··on The AI Hate Progression
> Your argument boils down to: companies haven't respected copyright or user consent for a long time, so why should we now care?

This is not a very charitable take. My argument is: that cat is out of the bag, so what do we do next? Do you have a plan? Ranting about consent or privacy or copyright won't somehow untrain these models. At least the math people put out the pretty sensible Leiden Declaration[1].

> Don't you see a problem with that? Why should we not care?

This is a lot like saying "do you not see a problem with nuclear weapons? They can wipe out humanity!"—I mean, yeah, sure I do, but ranting about the technology won't somehow put that cat back in the bag, either.

[1] https://leidendeclaration.ai/

dvt··on The AI Hate Progression
Google started indexing copyrighted data without consent in 1999, Yahoo in 1994. Absolute delusion to think that ChatGPT is the one that broke consent.
dvt··on The AI Hate Progression
> At some point, the entire tech industry saw ChatGPT and fell into a collective psychosis and decided that this, this is the next big thing, and that we must pull out the stops to ensure the prophecy is fulfilled of generative AI/LLMs becoming the next big thing.

This is extremely disingenuous, ChatGPT was the fastest-growing consumer product for a reason. Part hype, part usefulness, part novelty. The main problem I have with AI haters (just like AI lovers) is that they can't just be balanced about their takes. It's not that hard to criticize the AI hypetrain without a strawman.

> I'm sure copyrighted material was already being fed into LLMs at this point (I mean, you also had people willingly feeding it in, like the example I gave above) but once the techbros caught on and wanted to accelerate this, suddenly EVERYTHING public facing was fair game to training their models.

This is also just extremely disingenous. For full disclosure, I'm technically a stakeholder here, as I wrote two books which made it into training sets (and part of two class actions), but this cat is way out of the bag. Unless you really want to start splitting hairs about "ingesting" vs "processing" vs "training" vs "transforming," Google and Yahoo and even DDG have been using copyrighted data for a quarter century, if not longer. Folks were bringing this up decades ago, especially record labels that were suing google for copy-pasting lyrics to their main search pages; were you complaining then, too? Because some people were.

> All because investors demanded it, and companies didn't want to be caught with their pants down if these inflated claims of it being the next big thing proved to be true. This is where the hate for me really started, because a lot of these companies forced AI upon you, with no means to opt out. FOMO is a hell of a drug on a corporate scale, ho-ly.

Corporate FOMO is pretty run of the mill, this shouldn't be surprising in the least. I must've done like half a dozen "blockchain hackathons" or "VR demos" back when those technologies were all the rage. I don't really see how it's that big of a big deal, other than being mildly annoying.

> You'll notice a trend here: Consent is just gone. It does not exist when AI enters the room in 90% of cases. Companies just foist it on you and tell you to shut up and like it, or leave.

Consent was gone when Google, Yahoo, etc. started indexing the entire internet. It was doubly gone when Facebook sold PII data to advertisers. It was triply gone when Experian got hacked and the SSN of every taxpayer in the USA was leaked (and no one went to jail lol). Let's stop being dramatic. It just betrays a fundamental misunderstanding of how data collection works.

> Good luck avoiding AI if we buy up all the components for you to build your own computers and devices! Submit everything to the cloud, it's now the only affordable option, suckers!

Blog author and hobbyist photographer discovers free markets.

I don't mean to be dismissive, but this kind of take is boring, uninspired, and (ironically) could've just been written by ChatGPT. Come up with an interesting point or thought-provoking counter-argument and maybe people will take you seriously.

dvt··on ASM Shader Toy
Super cool shader toy! Semi-related, but my favorite demo ever (a tiny 64k Heaven Seven[1]) was fully coded in extremely esoteric assembly and used all kinds of mind-blowing tricks. There's a blog post written about it somewhere by a few Exceed members; real-time raytracing on a CPU in 2000 still blows my mind.

[1] https://www.youtube.com/watch?v=rNqpD3Mg9hY

dvt··on SpaceX to buy Cursor for $60B
Obviously you're trying to be snarky, but I hope you realize that Congress does, in fact, have (fairly broad) statutory authority over executive agencies[1].

[1] https://www.congress.gov/crs-product/R45442

dvt··on SpaceX to buy Cursor for $60B
> the current political direction of the United States

These kinds of comments reek echo-chamber parroting and zero substantive research. As someone that very much enjoys and carefully follows politics, the current political direction points squarely to Republicans getting absolutely pummelled in the midterms, effectively turning Trump's administration into a 2-year lame duck. What are you even talking about?

dvt··on Show HN: Fata – Spaced repetition to fight skill rot from AI coding
> I argue "cleverness" is a learned and honed skill by exposure and exercise.

Even if I were to concede this point, it's certainly not honed by the kind of exercise that OP is advertising. We are deep in Max Howell's "invert a binary tree" territory here.

← PreviousPage 4 of 34Next →