The title includes “retrospective” but the learning and changes following such a process are absent.
5,032 karma · joined September 18, 2016
Opinions are my own.
The title includes “retrospective” but the learning and changes following such a process are absent.
If you’re in need of a writing prompt …
Actually novel implementation is protected, paid the author's bills, and was dedicated to the public domain recently - no massive corp involved.
could mean there's a Test B and a Test A, and they're interchangeable.
Or:
You can run Test B to confirm A works, that is "to test A".
Again the stated goal to clear documentation for non-native or limited-exposure speakers. This doesn't pass that test.
> "Test" is an approved noun, but not an approved verb.
> STE: Test B is an alternative to test A.
So much for clear - unless you know the STE specific rule, the sentence is unambiguously ambiguous.
Direct access btw since the official site gates downloads with a Google form: https://www.asd-ste100.org/assets/files/ASD-STE100_ISSUE9.pd...
CAD: MoI (cool kids nowadays use Plasticity, but that is taking on similar feature creep - in any case I own and use both)
SubD modelling: nothing has ever beat Silo. And silo still gets updates, but the workflow is the same. Cheetah3D is also in the same boat.
Rendering (I do products, not character/CGI): Maverick. Fastest path tracer in the world, perpetual license, rivals Vray for product viz. if you’re currently using Keyshot - switch now.
Texturing: 3Dcoat or Marmoset Toolbag 5
What most of these have in common is either the workflow is super intuitive, or you have at least 15 years of training material readily accessible that still applies to the software today as it did back then.
Overall it’s a choice between a kitchen sink pipeline (Blender, Rhino) vs a specialised one. If I were to teach again, I would still use Blender in the classroom, and I fully support the project. My brain just works better with smaller, focussed tools.
Having gone through the process and gotten both approvals and rejections, the line is pretty clear.
I work at GitHub but even then I often use OpenRouter models in the CLI and Copilot App
They support much of the same triggers and come with many additional security controls out of the box
Most code I write is still Nim though.
Today another thing started happening which are phrases like "I've been burning too many tokens" or "this has taken too many turns". Which ironically takes more tokens of custom instructions to override.
Also claude itself is partially down right now (Arp 6, 6pm CEST): https://status.claude.com/
Stream of vulnerabilities discovered using security agents (23 so far this year): https://securitylab.github.com/ai-agents/
Taskflow harness to run (on your own terms): https://github.blog/security/how-to-scan-for-vulnerabilities...
I work on 3D/4D math in F#. As part of the testing strategy for algorithms, I've set up a custom agent with an F# script that instruments Roslyn to find FP and FP-in-loop hotspots across the codebase.
The agent then reasons through the implementation and writes core expressions into an FPCore file next to the existing tests, running several passes, refining the pres based on realistic caller input. This logs Herbie's proposed improvements as output FPCore transformations. The agent then reasons through solutions (which is required, Herbie doesn't know algorithm design intent, see e.g. this for a good case study: https://pavpanchekha.com/blog/herbie-rust.html), and once convinced of a gap, creates additional unit tests and property tests (FsCheck/QuickCheck) to prove impact. Then every once in a while I review a batch to see what's next.
Generally there are multiple types of issues that can be flagged:
a) Expression-level imprecision over realistic input ranges: this is Herbie's core strength. Usually this catches "just copied the textbook formula" instance of naive math. Cancellation, Inf/NaN propagation, etc. The fixes are consistently using fma for accumulation, biggest-factor scaling to prevent Inf, hypot use, etc.
b) Ill-conditioned algorithms. Sometimes the text books lie to you, and the algorithms themselves are unfit for purpose, especially in boundary regions. If there are multiple expressions that have a <60% precision and only a 1 to 2% improvement across seeds, it's a good sign the algo is bad - there's no form that adequately performs on target inputs.
c) Round-off, accumulation errors. This is more a consequence of agent reasoning, but often happens after an apparent "100% -> 100%" pass. The agent is able to, via failing tests, identify parts of an algorithm that can benefit from upgrading the context to e.g. double-word arithmetic for additional precision.
I once had a company that procured very simple electronics (fingerprint readers) from Taiwan and due diligence included travelling there, meeting every single person in the engineering office in person, then touring the contract factory where this would be built, then negotiating shipping and even driver development details.
This took all of one week and the price of a few plane tickets. We didn’t have the cash for professional auditors. In the end we got a product that worked, and even at a lower price (negotiating at a distance is not effective).
If dataflow is not provably connected from source to sink, an alert is impossible. If a sanitization step interrupts the flow of potentially tainted data, the alert is similarly discarded.
The end-to-end precision of the detection depends on the queries executed, the models of the libraries used in the code (to e.g., recognize the correct sanitizers), and other parameters. All of this is customizable by users.
All that can be overwhelming though, so we aim to provide sane defaults. On GitHub, you can choose between a "Default" and "Extended" suite. Those are tuned for different levels of potential FN/FP based on the precision of the query and severity of the alert.
Severities are calculated based on the weaknesses the query covers, and the real CVE these have caused in prior disclosed vulnerabilities.
QL-language-focused resources for CodeQL: https://codeql.github.com/
I would love to hear what kind of local experience you're looking for and where CodeQL isn't working well today.
As a general overview:
The CodeQL CLI is developed as an open-source project and can run CodeQL basically anywhere. The engine is free to use for all open-source projects, and free for all security researchers.
The CLI is available as release downloads, in homebrew, and as part of many deployment frameworks: https://github.com/advanced-security/awesome-codeql?tab=read...
Results are stored in standard formats and can be viewed and processed by any SARIF-compatible tool. We provide tools to run CodeQL against thousands of open-source repos for security research.
The repo linked above points to dozens of other useful projects (both from GitHub and the community around CodeQL).