HNHacker News
TopNewBestAskShowJobs

codegladiator

1,475 karma · joined July 22, 2017

Winning the fights against code !
submissionscomments
codegladiator··on Our framework for reporting model misalignment
> with vague instructions

All instructions are vague unless its code. But you can also give llm "code" and expect vague outcomes if you ask it to emulate what the runtime would look like.

codegladiator··on Tailcat – Like netcat, but over Tailscale’s data plane
why not try out some existing project already on top of iroh ?

I see a bunch here on awesome-iroh page

https://github.com/n0-computer/awesome-iroh

codegladiator··on Late.sh – a command-line Clubhouse for computer people
read it ? we dont do what kind of stuff here, i got my army of reader agents.
codegladiator··on DSLs Enable Reliable Use of LLMs
for me reliable = deterministically repeatable. if a llm has been able to do a task successfully and i ask it to do the same task again, i want the reliability that i get the same outcome (be if failure or success)

like if someone says "is this car reliable", i dont expect an infallible car. but if the cars third gear is 'you know sometimes it doesnt work', i wouldnt take that car out of city.

codegladiator··on DSLs Enable Reliable Use of LLMs
yeah and so llm cannot write C reliably
codegladiator··on DSLs Enable Reliable Use of LLMs
nothing wrong if your dsl stays small and you have deterministic validators/compilers to your actual target.

if dsl gets large, the number of potential interactions your dsl allow will grow exponentially (unless you are building an one dimensional action layer). and there will be semantic issues unless your dsl is "clear and intuitive", also comprehensive enough to accomodate your ongoing changes, else every change is now 2 changes.

so you now you need a comprehensive manual for your agent which needs to be sent in every /completion request.

codegladiator··on DSLs Enable Reliable Use of LLMs
> any JSON or YAML that carries semantics with the syntax is a language

semantics are defined by the converter/compiler/interpreter, and that is the process which is going to consume the said json/yaml. if the json/yaml is going to be consumed by any process then the semantics are inherently defined, so by your definition all jsons/yamls are in their own a "language" (or they are not being consumed at all), which just defeats the purpose of calling it a language at that point.

codegladiator··on DSLs Enable Reliable Use of LLMs
> The LLMs can write it fine. It wrote it almost acceptably on first sight

you are probably talking of some coding harness which looked up the existing code base and then made it and not something like first prompt to llm "write xyz testcase in my systems company dsl in tcl from early 2000".

> LLMs can write it fine

sure, coding by examples is fine (it goes back to "system prompt describing the language"). but the claim we are arguing is reliability. did the llm generate 100% test case or code after reading your existing codebase, likely not as you mention it "almost acceptably on first sight".

i would probably be in denial if was suggesting llms not good at finding and fitting patterns.

codegladiator··on DSLs Enable Reliable Use of LLMs
> you can have DSLs that are json/yaml, is my point

i disagree on the semantics of "DSL" vs "config in json". these are different.

and anyways both of them just pretend to be "reliable" by throwing the responsibility to an upper layer of validator/compiler/interpreter

> it is worse because the tooling around it won't be as good

plantuml is good because of the the tooling around it. not sure if we are agreeing/disagreeing there, confused by the wording

codegladiator··on DSLs Enable Reliable Use of LLMs
its served me okay.

> Perfect and deterministic? No, of course not. Just an improvement and mitigation.

exactly, not reliable, as the article tries to portray

codegladiator··on DSLs Enable Reliable Use of LLMs
> DSL that is json/yaml helps a ton

it definitely does, and i would say json/yaml is not a dsl. this example of json/yaml keeps coming in the form of "DSL". i would say your configuration is not a dsl, it a declaration. llms are better at declarative stuff ? maybe but there are hardly that many of complex declarative frameworks.

PlantUML is a real dsl. not just declarative yaml.

codegladiator··on DSLs Enable Reliable Use of LLMs
> less common DSLs like liquidsoap’s stream management DSL

seems to be on github since 2008 so definitely in the training data. i am not talking about less or more common. either "your dsl" would need to look something like someone elses dsl (at this point is it your dsl?) or you need some way to get your dsls examples in the training data for the llm, or feed it in the prompt.

> LLMs respond well to clear, simple structure

and what a "clear simple structure" for a dsl is also quite not mentioned. clear and simple would be quite subjective based on the domain, the article says let the llm go in a loop trying to figure out the dsl for you.

> checkpoints that have more structure than natural language

if llm is at any point in the structured generation part then either you have a deterministic validator/compiler or you are back to reading/reviewing it manually, what can you trust ?

codegladiator··on DSLs Enable Reliable Use of LLMs
> The advantage holds while the DSL stays small and constrained enough that a few in-context examples can convey its usage. There is also a real upfront cost in designing and maintaining the language and its semantic model. The payoff is therefore concentrated in well-factored, genuinely constrained DSLs backed by a validator.

dsl stays small is doing all the heavy lifting here

the premise is that because of these few existing dsls (like PlantUML mentioned) my "new dsl" will be equally effective. PlantUML has millions of examples in the training data, my new dsls are not (specially if its not json/yaml or just function chain based). as the number of things that can mix and match increase you are basically looking at a whole system prompt just describing the new language.

this brings us to the second part. step 2: after dsl is 'planned' (note they use the java compiler), the dsl need to have a real compiler/executor, not just a validator. because if then you are going to ask the llm to "compile the dsl to implementation" we are back to square 1.

codegladiator··on Mr. Baby Paint and accidentally discovering a new cellular automata
AlphaBaby ?
codegladiator··on Tokenmaxxing is dead, long live tokenmaxxing
10 days
codegladiator··on Ask HN: Why is AI use decried if it has been used without attribution?
Its like someone reheating frozen food and saying I cooked it and they have not eaten it themselves and want me to buy it.
codegladiator··on Donating AI credits to open source projects
if its so critical then how about donating money and letting the maintainer decide where to spend it.
codegladiator··on An AI agent deleted our production database. The agent's confession is below
> Master your craft. Don’t guess, know.

You mean add that to my prompt right ?

codegladiator··on Forcing an Inversion of Control on the SaaS Stack
how far it can go ? complete page rewrites ?
codegladiator··on Claude Code's source code has been leaked via a map file in their NPM registry
what you are suggesting would be like a truck company using trucks to move things within the truck
codegladiator··on Ask HN: We need to learn algorithm when there are Claude Code etc.
algorithms are tools. like in any field every professional is supposed to be aware of "the basic standard tool set", what a good implemenation of a tool supposed to be like (i should be able to assess if a hammer is broken or good for use if i am civil engineer, or piano if i am a musician). it does not mean they should be able to "create that tool from scratch".

knowing the tools will only benefit you irrespective of your employer or other tools in town. (llm is also a tool, ides are tools, libraries are tools, design patterns is a tool)

codegladiator··on When AI writes the software, who verifies it?
once upon a time 'engineering' in software had some meaning attached to it...

no other engineering profession would accept the standards(or rather their lack of) on which software engineering is running.

codegladiator··on India's top court angry after junior judge cites fake AI-generated orders
> She had no intention to misquote or misrepresent the rulings and that "the mistake occurred solely due to the reliance on an automatic source", the high court wrote

I don't think the intention matters here. Its the same deal with every profession using llm to "automate" their work. The onus in on the professional, not the llm. Arstechnica case could have been justified by same manner otherwise.

Not knowing the law isnt execuse to break law, so why is not knowing the tool an excuse to blame the tool.

codegladiator··on Elevated Errors in Claude.ai
Isnt bedrock and vertex pass thru to anthropic servers ? I didnt know aws/google are deploying the actual models
codegladiator··on Show HN: I built a zero-browser, pure-JS typesetting engine for bit-perfect PDFs
devnagri in the screenshot is wrongly rendered.

Also can you share some names of films you have been part of as film director.

codegladiator··on Ladybird adopts Rust, with help from AI
> This was human-directed, not autonomous code generation.

All my vibe coded projects are human directed, unless explicitly stated otherwise

codegladiator··on Ask HN: What (other) jobs do you think of doing?
zoo keeper
codegladiator··on MinIO repository is no longer maintained
can vouch for SeaweedFS, been using it since the time it was called weedfs and my managers were like are you sure you really want to use that ?
codegladiator··on Tally – A tool to help agents classify your bank transactions
Fair enough. I did not see this as a promotion of the product and more of as a show experimental side project. But if they really want to promote the product, the llm design isnt helping giving any confidence. A blog post would have sufficed.
codegladiator··on Tally – A tool to help agents classify your bank transactions
Why does LLM generated websites feel so "LLM generated".

Its like a bootstrap css just dropped. People still giving "minimum effort" into their vibe code/eng projects but slap a domain on top. Is this to save token cost ?

Page 1 of 28Next →