HNHacker News
TopNewBestAskShowJobs

ashater

47 karma · joined September 22, 2016

submissionscomments
ashater··on What if you could stop your AI agent before it makes a mistake?
Do you want to monitor what an Al agent is about to do before it acts?

In our new paper, Beyond the Black Box: Interpretability of Agentic Al Tool Use, we explore how mechanistic interpretability can help surface signals around tool-use decisions, missed calls, unnecessary calls, and higher-risk actions.

ashater··on TinyLoRA – Learning to Reason in 13 Parameters
Likely reasoning is part of the original model. It is well known that it is not possible to get a 1bn parameter model to reason, even with RL.
ashater··on Beyond the Black Box: Interpretability of LLMs in Finance
We want to do both. In finance, highly regulated industry, understanding how models work is critical. In addition, mech interp will allow us to understand which current or new architectures could work better for financial applications.
ashater··on Beyond the Black Box: Interpretability of LLMs in Finance
Thank you for reading. One of the main reasons we've written the paper is to help with model validation of LLM usage in our highly regulated industry. We are also engaging with regulators.

The industry at the moment is mostly using closed sourced vendor models that are very hard to validate or interpret. We are pushing to move onto models, with open source weights and where we can apply our interpretability methods.

Current validation approaches are still very behavioral in nature and we want move it into mechanistic interpretation world.

ashater··on Beyond the Black Box: Interpretability of LLMs in Finance
Our paper provides evidence of features in Finance but I would suggest reading seminal papers from Anthropic https://www.anthropic.com/news/golden-gate-claude and https://transformer-circuits.pub/2024/scaling-monosemanticit...

Monosemantic behavior is key in our research.

ashater··on Beyond the Black Box: Interpretability of LLMs in Finance
Thank you. Agreed, we are exploring different ways to apply these interpretability methods to a wide range of transformer based methods, not just decoder based generative applications.
ashater··on Beyond the Black Box: Interpretability of LLMs in Finance
Paper introduces AI explainability methods, mechanistic interpretation, and novel Finance-specific use cases. Using Sparse Autoencoders, we zoom into LLM internals and highlight Finance-related features. We provide examples of using interpretability methods to enhance sentiment scoring, detect model bias, and improve trading applications.
ashater··on Beyond the Black Box: Interpretability of LLMs in Finance
Our paper introduces AI explainability methods, mechanistic interpretation, and novel Finance-specific use cases. Using Sparse Autoencoders, we zoom into LLM internals and highlight Finance-related features. We provide examples of using interpretability methods to enhance sentiment scoring, detect model bias, and improve trading applications.
ashater··on City simulator I made in Scratch
My younger daughter, who is pretty good at making games in Scratch is not that interested in jumping into text/code based programming. I do think Scratch makes things a lot easier and text based programming is not thar appealing to kids. I will try to start her with Pygame but even that might make it seem very arcane and not very visual.
ashater··on The Next Great Leap in AI Is Behind Schedule and Crazy Expensive
Another piece of evidence that LLMs are plateauing
ashater··on Interview with Terence Tao in Barcelona
Terrence Tao has a very healthy view of what AI can and cannot achieve shorter term. It is refreshing to hear more grounded views from a top mathematician who is clearly well versed in the topic.
ashater··on Learning containers from the bottom up
Good article, steps one level below container managers like Docker or k8s. Obviously not the indepth of how Linux kernel manages container processes but a good write-up.
ashater··on FPGA Developer Tutorials
pynq is a good starting board. What's really good is the software environment connected to the fpga logic via Python ecosystem. There are lots of examples via Jupyter notebooks.
ashater··on Towards an Untrepreneurial Economy?
how do I modify the link?
ashater··on Towards an Untrepreneurial Economy?
The paper describes the rise of “tech entrepreneurship as a lifestyle” and how more entrepreneurship seems primarily motivated by people wanting to “be entrepreneurs” rather than by those with potentially valuable ideas that are likely to lead to economically gainful, productive activity. The result: fewer high-growth firms because the types of firms entering the market simply aren’t capable of attaining success as these types of firms could in the past, the researchers say.