HNHacker News
TopNewBestAskShowJobs

sjmaplesec

48 karma · joined July 31, 2018

submissionscomments
sjmaplesec··on Jev is 13.6x faster, 2.7x cheaper than GPT Luna 6
We ran a bunch (2,725) of Tessl verifier tests which are part of our test suite for our internal code base and switched the judge model to Jev and GPT Luna 6

The results:

- Jev is 13.6x faster - Jev is 2.7x cheaper - The models had a 85.9% agreement on verdicts

Jev finished the suite in 32 seconds. GPT Luna 6 took 436.5 seconds.

The interesting part is where the models disagreed, particularly on rules about comment structure and content. We published the approach, what we tested, and a workflow you can reproduce on your own codebase.

sjmaplesec··on Haiku 4.5 + skills outperforms Opus 4.7. 9 models tested with and without skills
Full set of models tested: claude-opus-4-7 claude-opus-4-6 claude-sonnet-4-6 claude-haiku-4-5 gpt-5.4 gpt-5.3-codex gpt-5-codex cursor-composer-2

11 Skills used were here: https://github.com/mcollina/skills

sjmaplesec··on Roast My Skill
Pass in a skill and it'll roast the contents:

"This is not just useless - it is an insult to the very concept of functionality."

sjmaplesec··on Claude Code model comparison: Skill usage
I ran some evals to see which Anthropic models use skills the best, between Opus, Sonnet and Haiku.

I was pretty impressed how good Haiku was with skills at completing various tasks

sjmaplesec··on Googleworkspace/CLI isn't optimized – Test your skills
Link to all the review scans is here - mostly in the 50-70% range https://tessl.io/registry/skills/github/googleworkspace/cli
sjmaplesec··on Googleworkspace/CLI isn't optimized – Test your skills
There's so much more we can do around activation and skills creation. Looking at the eval results, there are even cases where the context makes the agent worse.

Scenario 5, test 1 72% -> 22%

https://tessl.io/eval-runs/019cc02f-bb26-76e0-a7c9-598a7337e...

sjmaplesec··on Agents.md file isn't the problem. Your lack of Evals is
The review eval tests language, activation etc of skills. I guess you could move it all to a skill quick and then run an eval on that if using Tessl. This checks if the way you write the instructions etc are being well understood by the agent
sjmaplesec··on Agents.md file isn't the problem. Your lack of Evals is
An eval is to an LLM as a test is to code.
sjmaplesec··on Agents.md file isn't the problem. Your lack of Evals is
Tessl can generate the evals, both to test anthropic best practices as well as running scenarios with and without the skill to check if it's helping
sjmaplesec··on Agents.md file isn't the problem. Your lack of Evals is
Can add this as a skill or as part of a skill, and so you don't need to keep prompting the same things.
sjmaplesec··on Agents.md file isn't the problem. Your lack of Evals is
No, the context can be human created as much as it could be llm generated. The suggestions are based on Anthropic best practices and allow the agents to activate, and use the skills better, make the text clearer for the agent etc.
sjmaplesec··on Show HN: A package manager for agent skills with built-in evals
This resonates with my experience: we have dozens of internal “playbooks” and prompt snippets floating around, and nobody knows which ones still work after model changes. If you can make “skill quality” visible over time (regressions, drift), that’s valuable. Do you have a CI integration where you can pin a skill version and fail builds if eval scores drop?
sjmaplesec··on Datadog CEO on AI and Observability
This is so true!
sjmaplesec··on Datadog CEO on AI and Observability
I'm pretty sure Olivier Pomel rarely does podcasts, but this was a pretty good one.

Some of my thoughts:

- Customers "lie to themselves" saying they prefer noise to missed issues, when in practice 2 false alarms make them lose faith - and its implications of that to AI.

- The path to agentic adoption comes from finding narrow use-cases they can solve with high confidence, and bringing them to life.

- How foundation models can't understand time-series, and how DataDog's foundation model, Toto (get it?!) tackles it.

sjmaplesec··on AI with IaC – "80% of the value is in codifying 20% of the assumptions"
Page title: Armon Dadgar, Hashicorp co-founder, on AI Native DevOps: Can AI shape the future of Autonomous DevOps workloads?

Link is to an interesting podcast episode about Gen AI being used in infra

sjmaplesec··on SourMint: Malicious code, ad fraud, and data leak in iOS
Totally - Sounds like they started as a legitimate Ad SDK too!
sjmaplesec··on SourMint: Malicious code, ad fraud, and data leak in iOS
Amazing that this has been going for a year now - Let's see how Apple deal with the existing apps on the AppStore.