HNHacker News
TopNewBestAskShowJobs

_jayhack_

286 karma · joined March 19, 2017

https://github.com/jayhack
submissionscomments
_jayhack_··on A Mathematical Framework for Transformer Circuits (2021)
Shocking to me how little interest the general public has in mechinterp given the alien capabilities demonstrated by LLMs - this and the subsequent transformer-circuits.pub publications will be seen as classic, foundational work in a few years
_jayhack_··on Mistral Patent for “Code implemented tool calls”
Likely from Palmer Luckey, who coined the term 'Chinese instruction manuals'
_jayhack_··on Qwen3.8-Max: A New Bar for Coding and Cowork
Only the ones that beat expectations
_jayhack_··on The Maxwell Conjecture Is False (GPT 5.6 Sol)
formal verifiability e.g. vi Lean
_jayhack_··on Memoirs of Extraordinary Popular Delusions and the Madness of Crowds (1852)
The content on "priming" (significant pillar of the book) has collapsed as part of the reproducibility crisis in psychology. More here: https://replicationindex.com/2017/02/02/reconstruction-of-a-...
_jayhack_··on LLMs are not the black box you were promised
Deep Learning is a Black Box and Here's Why You Should Use Random Forests Because They Are Interpretable

This was the mantra of applied machine learning c. 2010 - 2024 for anyone paying attention. No longer the case.

_jayhack_··on LLMs are not the black box you were promised
For some definitions of better, yes. Chinese is more token efficient for representing fixed text, for example, although this does not always lead to better performance on downstream tasks.
_jayhack_··on LLMs are not the black box you were promised
As the author - this was adapted from a thread posted on X in March (linked in article). AI did the adaptation, I wrote the original article. It seems like it inserted grammatically correct hyphens, otherwise the copy is mine.
_jayhack_··on LLMs are not the black box you were promised
Hello, I am the author - this is not an LLM-generated article, I wrote this by hand and had an LLM adapt it from a thread on X. You can see the original thread here: https://x.com/mathemagic1an/status/2035850046735098065

> the fact that language models have human-interpretable representations and neurons has been known since BERT... Circuits research also does not come from Anthropic... The article does not claim Anthropic invented the field, rather that they have had important contributions to it. This is intended as an overview into a specific set of ideas that are working for mechanistic interpretability. Not a formal literature review.

_jayhack_··on Launch HN: Freestyle – Sandboxes for Coding Agents
Would love to understand how you compare to other providers like Modal, Daytona, Blaxel, E2B and Vercel. I think most other agent builders will have the same question. Can you provide a feature/performance comparison matrix to make this easier?
_jayhack_··on Cameras and Lenses (2020)
Great article. For another fantastic explainer on optics, see 3Blue1Brown's video on refraction: https://www.youtube.com/watch?v=KTzGBJPuJwM
_jayhack_··on Codemaps: Understand Code, Before You Vibe It
> static analysis tools that produce flowcharts and diagrams like this have existed since antiquity, and I'm not seeing any new real innovation other than "letting the LLM produce it".

Inherent limitation of static analysis-only visualization tools is lack of flexibility/judgement on what should and should not be surfaced in the final visualization.

The produced visualizations look like machine code themselves. Advantage of having LLMs produce code visualizations is the judgement/common sense on the resolution things should be presented at, so they are intuitive and useful.

_jayhack_··on Meta Superintelligence Labs' first paper is about RAG
Vector embedding is not an invention of the last decade. Featurization in ML goes back to the 60s - even deep learning-based featurization is decades old at a minimum. Like everything else in ML this became much more useful with data and compute scale
_jayhack_··on Game over for pure LLMs. Even Rich Sutton has gotten off the bus
Gary Marcus has been taking victory laps on this since mid-2023, nothing to see here. Patently obvious to all that there will be additional innovations on top of LLMs such as test-time compute, which nonetheless are structured around LLMs and complementary
_jayhack_··on Show HN: Vibe Kanban – Kanban board to manage your AI coding agents
Very cool and interesting project. Ideas like this are a threat to traditionally-conceived project management platforms like Linear; that being said, Linear and others (Monday, ClickUp, etc.) are pushing aggressively into UX built for human/AI collaboration. I guess the question is how quickly they can execute and how many novel features are required to properly bring AI into the human project workspace
_jayhack_··on Measuring the Impact of AI on Experienced Open-Source Developer Productivity
This does not take into account the fact that experienced developers working with AI have shifted into roles of management and triage, working on several tasks simultaneously.

Would be interesting (and in fact necessary to derive conclusions from this study) to see aggregate number of tasks completed per developer with AI augmentation. That is, if time per task has gone up by 20% but we clear 2x as many tasks, that is a pretty important caveat to the results published here

_jayhack_··on Show HN: AI for Building Design, Planning, and Permitting
It looks like you used AI-generated videos for customer testimonials. This should be illegal
_jayhack_··on Type-constrained code generation with language models
Also worth checking out MultiLSPy, effectively a python wrapper around multiple LSPs: https://github.com/microsoft/multilspy

Used in multiple similar publications, including "Guiding Language Models of Code with Global Context using Monitors" (https://arxiv.org/abs/2306.10763), which uses static analysis beyond the type system to filter out e.g. invalid variable names, invalid control flow etc.

_jayhack_··on Launch HN: mrge.io (YC X25) – Cursor for code review
If you are looking for an alternative that can also chat with you in Slack, create PRs, edit/create/search tickets and Linear, search the web and more, check out codegen.com
_jayhack_··on Fintech founder charged with fraud; AI app found to be humans in the Philippines
Related article from mid-pandemic: https://www.theinformation.com/articles/shaky-tech-and-cash-...

A friend asked me to do diligence on this company circa 2021 given my personal background in ML. The founder was adamant they had a "100% checkout success rate" based on AI, which was clearly false. He also had 2 other startups he was running concurrently (?)

Live and learn!

_jayhack_··on Show HN: Codemcp – Claude Code for Claude Pro subscribers – ditch API bills
Codegen is an MCP client that you can trigger via Slack: codegen.com
_jayhack_··on Show HN: Codemcp – Claude Code for Claude Pro subscribers – ditch API bills
This is true and is essentially a form of arbitrage. Anthropic is eating the cost of your elevated queries with their $20 flat fee subscription.

The "famously huge API token costs" you are referring to is Cline passing the Anthropic API cost through to you with no markup. You even input your own API token.

_jayhack_··on Show HN: Codegen – OSS Python Library for Advanced Code Manipulation
Codemodder has extensive Java support, which Codegen does not support at the moment. Otherwise, my understanding of Codemodder is that it is focused on AST-level syntactical modifications. Codegen computes a richer graph datastructure, and this can be used for sophisticated modifications that depend on inheritance hierarchies, function usages, cross-file references and more.

Codemodder is written in Java, whereas you can write Codegen in a jupyter notebook or anywhere you can run Python.

_jayhack_··on Operator research preview
Unfortunately a lot of the things we want agents to interact with don't expose neat APIs. Computer use and, eventually, physical locomotion are necessary for unlocking agent interactivity with the real world.
_jayhack_··on Can AI do maths yet? Thoughts from a mathematician
If you think the purpose of pure math is to provide employment and entertainment to mathematicians, this is a dark day.

If you believe the purpose of pure math is to shed light on patterns in nature, pave the way for the sciences, etc., this is fantastic news.

_jayhack_··on Refactoring Python with Tree-sitter and Jedi
Good call, thank you
_jayhack_··on Refactoring Python with Tree-sitter and Jedi
We do advanced static analysis to provide programmatic access to the type system, etc., based on tree-sitter and in-house tech.

This enables APIs such as `function.call_sites`, `symbol.usages`, `class.parent_classes`, and more!

_jayhack_··on Refactoring Python with Tree-sitter and Jedi
Interesting refactor!

This is trivial with codegen.com. Syntax below:

  # Iterate through all files in the codebase
  for file in codebase.files:
      # Check for functions with the pytest.fixture decorator
      for function in file.functions:
          if any(d.name == "fixture" for d in function.decorators):
              # Rename the 'db' parameter to 'database'
              db_param = function.get_parameter("db")
              if db_param:
                  db_param.set_name("database")
                  # Log the modification
                  print(f"Modified {function.name}")
Live example: https://www.codegen.sh/codemod/4697/public/diff
_jayhack_··on Show HN: ts-remove-unused – Remove unused code from your TypeScript project
This is simple and configurable with Codegen - see the results on `renovate` here: https://www.codegen.sh/codemod/4553/public/diff

Most transformations like this are not possible with pure static analysis and require some domain knowledge (or repo-specific knowledge) in order to pull off correctly. This is because some code gets "used" in ways that are not apparent i the code.

Enjoy!

_jayhack_··on The Criminal Charges Against Aaron Swartz Were Fair and Reasonable (2013)
Part 2: https://volokh.com/2013/01/16/the-criminal-charges-against-a...
Page 1 of 2Next →