HNHacker News
TopNewBestAskShowJobs

summarity

5,032 karma · joined September 18, 2016

Product @ GitHub (CodeQL code scanning)

Opinions are my own.

submissionscomments
summarity··on CS240 AI Cheating Retrospective
My comment explicitly did not focus in the use of LLMs and in fact much of the post doesn’t. Leading with the depressing “work alone” slide, it focused on similarity, not AI use. Between students, ruling even “inadvertent” exposure as the same punishable (and literally literally undeniable!) offence. There’s millions of ways to structure a class for learning, this myopic focus is not one of them. Ironically the setup ignores the very basics of human collaboration. AI merely amplifies extremes, it’s a factor but neither the cause nor culprit here.

The title includes “retrospective” but the learning and changes following such a process are absent.

summarity··on CS240 AI Cheating Retrospective
If you remove the mention of AI from this post, the authors process and setup of the notice to students is absolutely mental. The intentions might be good, but damn if this isn’t a comically inept way to set up a kafkaesque nightmare rather than a place for actual learning. The defensive attitude doesn’t help either. Just the sheer energy that went into all the steps mentioned here, not invested in any even superficial attempt to perhaps adapt. That isn’t to “give into AI” or cheating, but if 50% of students break your rules, more rules and litigation about the rules are not the next step to take.
summarity··on America.gov goes crazy on "play Minecraft"
Cynically this is a great way for the unhinged admin to train and harden an AI censorship filter :)

If you’re in need of a writing prompt …

summarity··on Can I opt out of my input or output data being used for training?
“Opt in by default” would mean that it is not enabled by default. Do you mean opt out?
summarity··on New Mac Studio with M5 Max and M5 Ultra
They also cancelled availability of the 512 M3 Ultra months ago in many regions, likely just redirecting memory
summarity··on The AI Situation in Software Development
Keeping a simple log of "accepted/rejected" avenues is basically all that's needed, maybe with a root style guide in the project. I've got several multi-day sessions in 5.6 Sol running without going off the rails in terms of complexity. After a while in this loop it actually starts to remind/berate itself to keep things straight-forward.
summarity··on Racket v9.3
Thank you for Herbie btw, it continues to be one of the most used tools in my personal projects (all 3D math and rendering)
summarity··on Mistral Patent for “Code implemented tool calls”
I agree in general, but can think of at least one counterpoint: https://terathon.com/blog/decade-slug.html

Actually novel implementation is protected, paid the author's bills, and was dedicated to the public domain recently - no massive corp involved.

summarity··on F*: A general-purpose proof-oriented programming language
What? There’s literally a completely interactive book linked right from the home page.
summarity··on Agent Skill to Force Docs in ASD-STE100 Simplified Technical English
"Test B is an alternative to test A."

could mean there's a Test B and a Test A, and they're interchangeable.

Or:

You can run Test B to confirm A works, that is "to test A".

Again the stated goal to clear documentation for non-native or limited-exposure speakers. This doesn't pass that test.

summarity··on Agent Skill to Force Docs in ASD-STE100 Simplified Technical English
Gotta love how the very first example in Issue 9 of the standard is already self-defeating:

> "Test" is an approved noun, but not an approved verb.

> STE: Test B is an alternative to test A.

So much for clear - unless you know the STE specific rule, the sentence is unambiguously ambiguous.

Direct access btw since the official site gates downloads with a Google form: https://www.asd-ste100.org/assets/files/ASD-STE100_ISSUE9.pd...

summarity··on Blender 5.2 LTS
To answer the question at the top, yes there are countless suites that are instantly resumable even after months away. I used to teach Blender (v3) and I think I could find my way around but the overhead isn’t worth it. I’m also though in a privileged position to be able to afford more, smaller commercial tools.

CAD: MoI (cool kids nowadays use Plasticity, but that is taking on similar feature creep - in any case I own and use both)

SubD modelling: nothing has ever beat Silo. And silo still gets updates, but the workflow is the same. Cheetah3D is also in the same boat.

Rendering (I do products, not character/CGI): Maverick. Fastest path tracer in the world, perpetual license, rivals Vray for product viz. if you’re currently using Keyshot - switch now.

Texturing: 3Dcoat or Marmoset Toolbag 5

What most of these have in common is either the workflow is super intuitive, or you have at least 15 years of training material readily accessible that still applies to the software today as it did back then.

Overall it’s a choice between a kitchen sink pipeline (Blender, Rhino) vs a specialised one. If I were to teach again, I would still use Blender in the classroom, and I fully support the project. My brain just works better with smaller, focussed tools.

summarity··on Blender 5.2 LTS
The question is why don’t they integrate cycles? I mean I have a random app for home planning on macOS and even that comes with Cycles as the rendering engine. Can’t be that hard.
summarity··on OpenAI loses trademark dispute at EU court
Key difference between the trademark systems here: in the EU system you don’t get a trademark by trading with a specific name and it then being recognized. It’s the other way around: the name must be unique, not confusing, and highly specific. It’s actually irrelevant whether a product exists or is traded at all.

Having gone through the process and gotten both approvals and rejections, the line is pretty clear.

summarity··on Apple's new SpeechAnalyzer API, benchmarked against Whisper and its predecessor
Vs Voxtral would be a better comparison. No other model, open or closed, has been able to hit such a low AER (Acronym Error Rate ;)) for my meeting transcripts. Seems to understand/infer all the technobabble I use at work. Never have to edit anything. Whisper was catastrophically bad.
summarity··on Kimi K2.7 Code is generally available in GitHub Copilot
It has supported custom, local, any BYOM for quite a while.

I work at GitHub but even then I often use OpenRouter models in the CLI and Copilot App

summarity··on We’re making Bunny DNS free
Their DNS is also scriptable, it’s not just a name server
summarity··on Remembering Planet Source Code: Sharing Code Before GitHub Made It Easy
You can still download VS6 from Microsoft and clone a repo from there, and chances it’ll compile and run are higher than JS project that’s two weeks old
summarity··on Claude Code Routines
If you’re trying this for automating things on GitHub, also take a look at Agentic Workflows: https://github.github.com/gh-aw/

They support much of the same triggers and come with many additional security controls out of the box

summarity··on NimConf 2026: Dates Announced, Registrations Open
It’s growing but not a lot, I have some data here: https://pierretempel.com/p/nim-usage-on-github

Most code I write is still Nim though.

summarity··on Issue: Claude Code is unusable for complex engineering tasks with Feb updates
Not claude code specific, but I've been noticing this on Opus 4.6 models through Copilot and others as well. Whenever the phrase "simplest fix" appears, it's time to pull the emergency break. This has gotten much, much worse over the past few weeks. It will produce completely useless code, knowingly (because up to that phrase the reasoning was correct) breaking things.

Today another thing started happening which are phrases like "I've been burning too many tokens" or "this has taken too many turns". Which ironically takes more tokens of custom instructions to override.

Also claude itself is partially down right now (Arp 6, 6pm CEST): https://status.claude.com/

summarity··on Herbie: Automatically improve imprecise floating point formulas
This is somewhat in line with the approach taken by some softfloat libraries, e.g. https://bigfloat.org/architecture.html
summarity··on Claude Code Found a Linux Vulnerability Hidden for 23 Years
Related work from our security lab:

Stream of vulnerabilities discovered using security agents (23 so far this year): https://securitylab.github.com/ai-agents/

Taskflow harness to run (on your own terms): https://github.blog/security/how-to-scan-for-vulnerabilities...

summarity··on Claude Code Found a Linux Vulnerability Hidden for 23 Years
Already happend: https://arxiv.org/abs/2407.08708
summarity··on Herbie: Automatically improve imprecise floating point formulas
I posted this and it picked up steam over night, so I thought I'd add how I'm using it:

I work on 3D/4D math in F#. As part of the testing strategy for algorithms, I've set up a custom agent with an F# script that instruments Roslyn to find FP and FP-in-loop hotspots across the codebase.

The agent then reasons through the implementation and writes core expressions into an FPCore file next to the existing tests, running several passes, refining the pres based on realistic caller input. This logs Herbie's proposed improvements as output FPCore transformations. The agent then reasons through solutions (which is required, Herbie doesn't know algorithm design intent, see e.g. this for a good case study: https://pavpanchekha.com/blog/herbie-rust.html), and once convinced of a gap, creates additional unit tests and property tests (FsCheck/QuickCheck) to prove impact. Then every once in a while I review a batch to see what's next.

Generally there are multiple types of issues that can be flagged:

a) Expression-level imprecision over realistic input ranges: this is Herbie's core strength. Usually this catches "just copied the textbook formula" instance of naive math. Cancellation, Inf/NaN propagation, etc. The fixes are consistently using fma for accumulation, biggest-factor scaling to prevent Inf, hypot use, etc.

b) Ill-conditioned algorithms. Sometimes the text books lie to you, and the algorithms themselves are unfit for purpose, especially in boundary regions. If there are multiple expressions that have a <60% precision and only a 1 to 2% improvement across seeds, it's a good sign the algo is bad - there's no form that adequately performs on target inputs.

c) Round-off, accumulation errors. This is more a consequence of agent reasoning, but often happens after an apparent "100% -> 100%" pass. The agent is able to, via failing tests, identify parts of an algorithm that can benefit from upgrading the context to e.g. double-word arithmetic for additional precision.

summarity··on It Took Me 30 Years to Solve This VFX Problem – Green Screen Problem [video]
Well sort of, the industry tried to go way beyond that by capturing the entire light field: https://techcrunch.com/2016/04/11/lytro-cinema-is-giving-fil...
summarity··on 100 Jumps
Finally, Desert Golf and Flappy Bird merged into one
summarity··on Thermal Grizzly was scammed twice on raw materials worth €40k
This guys factory is just across the lake from where I live and this is painful to watch. Both Alibaba and the general local industry (metal fabs, train shops, etc) have high degrees of expertise in supply chain verification. You can hire (heck even bribe) experts along the way to reduce fuck ups. The video contained no mention of any audits, any additional paperwork beyond some pictures.

I once had a company that procured very simple electronics (fingerprint readers) from Taiwan and due diligence included travelling there, meeting every single person in the engineering office in person, then touring the contract factory where this would be built, then negotiating shipping and even driver development details.

This took all of one week and the price of a few plane tickets. We didn’t have the cash for professional auditors. In the end we got a product that worked, and even at a lower price (negotiating at a distance is not effective).

summarity··on Turn Dependabot off
No engine can be 100% perfect of course, the original comment is broadly accurate though. CodeQL builds a full semantic database including types and dataflow from source code, then runs queries against that. QL is fundamentally a logic programming language that is only concerned with the satisfiably of the given constraint.

If dataflow is not provably connected from source to sink, an alert is impossible. If a sanitization step interrupts the flow of potentially tainted data, the alert is similarly discarded.

The end-to-end precision of the detection depends on the queries executed, the models of the libraries used in the code (to e.g., recognize the correct sanitizers), and other parameters. All of this is customizable by users.

All that can be overwhelming though, so we aim to provide sane defaults. On GitHub, you can choose between a "Default" and "Extended" suite. Those are tuned for different levels of potential FN/FP based on the precision of the query and severity of the alert.

Severities are calculated based on the weaknesses the query covers, and the real CVE these have caused in prior disclosed vulnerabilities.

QL-language-focused resources for CodeQL: https://codeql.github.com/

summarity··on Turn Dependabot off
Heyo, I'm the Product Director for detection & remediation engines, including CodeQL.

I would love to hear what kind of local experience you're looking for and where CodeQL isn't working well today.

As a general overview:

The CodeQL CLI is developed as an open-source project and can run CodeQL basically anywhere. The engine is free to use for all open-source projects, and free for all security researchers.

The CLI is available as release downloads, in homebrew, and as part of many deployment frameworks: https://github.com/advanced-security/awesome-codeql?tab=read...

Results are stored in standard formats and can be viewed and processed by any SARIF-compatible tool. We provide tools to run CodeQL against thousands of open-source repos for security research.

The repo linked above points to dozens of other useful projects (both from GitHub and the community around CodeQL).

Page 1 of 8Next →