981 karma · joined July 23, 2012
I got 2400 with 10 in a row on first attempt.
Here's the entire decision matrix:
- Is the color, mood, composition, and level of detail exactly like most Nano Banana Pro images? If yes, AI.
- Is the composition/framing utterly perfect? If yes, AI.
- Is the subject ephemeral or random, like a candid photo or something otherwise unlikely to be photographed? If yes, not AI.
- Final decision, if there's any doubt, go entirely off of how centered the subject is.
I think some of this has more to do with the data set than the methodology. You could easily come up with a bunch of examples to fool me specifically, but alas, they didn't.
If we're writing a document that polices the model behavior, we're altering every version of the model+harness we use henceforth. Most things I wrote in a Claude.md because Opus 4.something was crap had roots in Opus, Claude Code, system prompts, and our own bad code we wrote last year. None of these things exist today but the Claude.md file can stick around like it's all still necessary.
I too am morally opposed--I abide by a 100 line (short lines, not paragraphs) limit and edit it rarely.
Most of my annoyances have come and gone, meanwhile I've gotten better at producing 4 good sentences of prompt for most things instead of 4 bad sentences. And we actually have CLI tools that can communicate effectively so instead of telling the model what "gates" to care about, we let the tooling surface area and output express what's important.
I'm hoping that perhaps my personal perspective shades why "instances" comes up, or why the reaction on HN seems to include the wider scope than the article itself covers.
What I hoped to read in the article is how we approach topics like centralization, censorship, moderation, data ownership--and with a technical lens. But I feel like all I got was "here's why instances are the wrong vocabulary" without substantively talking about the part I personally care about and want to marry the technical understanding with. Maybe I just read too shallowly and need to sit with it.
If I have an aesthetic in mind I'll use some screenshots of those sites in the prompt and phrase their inclusion as: "Look at these slightly non-standard designs that work really well for me." So far I've only seen Claude look for through-lines and high-level takeaways--"user likes <design feature> based on the screenshots, so I'll include that"--and screenshots aren't currently a granularity level where it can lift specific details or produce something derivative.
Other than that I try to encourage specific consideration of: type scale, borders and rounding, padding/whitespace, elevation as shadow vs blur, colors. I don't think one needs to pull every customization lever on every project.
I find it funny about meeting requirements when you give them, and making safe choices when you don't give direction. So if you're going to rate the output aesthetics and UX/content, but you don't prompt especially much around the aesthetics, you're only getting the safe assumed defaults. It's good at making bootstrap/tailwind clone designs unless you work that angle. For simple web pages, I've started making this the only focus for initial iteration.
The problem I see is that this is indistinguishable to a reader at a glance.
Distancing the writing from the "AI smell" not only improves the quality by dropping the unnecessary ocean of rhetorical devices, it forces the human to have real weight and agency on what's being said.
I think that act of distancing from raw LLM output through refinement is a huge quality leap. Even if you're only doing the refinement with an LLM, it forces the writing to have more voice and ideas from the author.
I can see the work that went into the analysis here but again, as a casual reader, it's impossible to tell that there were any original ideas here expressed by the author.
This is low-quality--every single day I witness Codex and Claude misunderstand, mislead, and hallucinate responses based on "assumptions" and I have to fact-check them.
If I wanted a statistical analysis and to be the human in the loop, I would ask the LLM myself, and I would definitely NOT read an article that just dumps the LLM output as-is.
The analogous version in black and white is "is this dark grey or light grey?" because that's the one asking you to guess which side of the 50/50 split the color is on.
It would be nice if it showed you 1st, 2nd, semi-finalist, quarter-finalist...
It would also be nice to see progress of some kind, a few minutes in I was wondering if I was near completion or just getting started.
We're building alphaLENS (https://alphalens.io), a suite of tools for professional market research and trading. Public registration opens in days. You can see our company at https://thinkalpha.io.
Stack: Svelte (migrating from React/Redux), MVVM architecture with signal primitives, class-based models with IoC for cross-platform targeting (web + Electron, mobile later). Heavy emphasis on testing and performance.
Looking for: Someone who thinks about problems in detail. You should be comfortable with OOP in JS, understand reactive primitives at a deeper level than hooks, and be able to operate autonomously on a small team.
If this is you, please reach out to me with your resume and a brief intro: jason [at] thinkalpha [dot] io
People want to pretend fundamentals of economics don't exist AND the company has moral obligations to fulfill to consumers. It's laughable.
It's not just nVidia, I've seen other expensive consumer brands getting the same sentiments.
Right, so if you take away the resolver composition (this is graph composition and not route federation), you can do the same things with a similar amount of effort in REST. This is no longer a GraphQL vs REST conversation, it's an acknowledgement that if you don't want any of the benefits you won't get any of the benefits.
My initial goal is to make a functional SillyTavern (AI roleplaying) replacement. SillyTavern builds prompts from a few rigid buckets (character, scenario, lore, system prompt, author's note...), which makes complex setups hard to manage. Content gets duplicated, settings have to be toggled in multiple places, and it’s easy to accidentally carry or modify state across conversations. Over time, it becomes difficult to tell what context is actually in effect.
I’m building an alternative that treats context as small, reusable pieces that can be composed and organized flexibly, rather than locked into fixed categories. Characters, settings, and behaviors can be mixed, reused, or temporarily enabled without duplication or manual cleanup, and edits preserve clear history instead of rewriting the past. The goal is to make managing complex context deliberate and controlled instead of fragile.
Although I’m trying to get the functionality required for roleplaying done first, the app is generic enough for other AI workflows where fine-grained, explicit context control is an improvement over existing chat interfaces. Think: start a new conversation with an assistant and start checking off rules, documents, and instructions to apply to the chat. Regenerate responses with clarifications or additional one-time context layers.