7,331 karma · joined August 3, 2012
Recent public projects include:
- Contextify: https://contextify.sh (A tool to assist CLI-based agentic programming workflows)
– FileKitty: https://github.com/banagale/FileKitty (a prompt engineering utility)
– Chief of Staff: https://chiefofstaffhq.com (a pre-generative AI text-to-speech SaaS)
I also very occasionally write at: https://banagale.com
Feel free to say hello, rob @ the domain above.
This is site meta though, see footer for contact methods to get direct answers on stuff like this.
Wasn’t there a dive on Freddy Pharkus up on FP recently?
These kind of folks still exist and still create. There is a lot more noise to sort through these days.
Except they were not for past few years as they misfired on the attempt to compete with vscode. That had a big impact on pycharm, which seemed starved for resources for so long. The company eventually declared a year of Django, but even that failed to really make an impact.
Arguably, Jetbrains had first insight into AI based code completion via rapid rise of the TabNine plugin but missed that opportunity also.
"Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5"
and
"We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5."
and
"In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one."
I realize it is corporate communications but "most common areas of feedback" and is a bit sterile. If the company wants authenticity and trust its easy to say that they found it hard to follow. And that it did not meet a quality bar they generally expect from their releases.
If this is not true, that it Opus 5 output was generally acceptable and we might see something like that again, that is an important consideration for potential customers or investors.
I don't remember seeing any RW content, but there was some John Lennon stuff. He was mostly in good taste, but I remember seeing something and being, uncomfortable with it.
I don't really see why Bob Barker should be any less protected, then John Lennon or Robin Williams. None of these people were saints.
But there is or at least used to be deference provided to certain people or establishments.
I haven't seen this video that was so upsetting, but perhaps it's not the reanimation so much as the content. Some of it is so obviously satire it's not to be taken seriously.
But, I remember somebody used Bruce Lee as a vessel to spread the gospel on Sora. I had to frown at that, as it was not how he would have behaved.
Contextify (https://contextify.sh)
It is an application for Linux, macOS and Windows that files your AI chats into a database your AI can read using a skill or MCP.
It backs up the valuable content of your Claude Code and Codex conversations (JSONL transcripts) and turns them into a complete history you don’t have to manage.
Use Contextify during normal flow of CLI AI conversations:
1. Resume unfinished work → "where did we leave off on that?"
2. Recover the intent/scope → "what was the actual goal of this whole effort?"
3. Verify it got done → "did we ever finish that, and which session proves it?"
4. Recall a fix → "how did we fix this the last time it broke?"
Just add "use total recall" or invoke the skill directly via /total-recall or $total-recall.You can self-host and keep all data local or use the source available Contextify Cloud sync server so your history on all machines in your cluster or desktop and laptop are always backed up and immediately available for use.
I have not been a Windows user since the days of the Hackintosh. But urging here and from users to cover the platform pushed me into it and I ended up really enjoying the challenge.
Right now I’m working on an expanded version of the windows app and support for integrating a 3rd party local LLM.
Also, IIRC there was reporting that Apple considered doing the compute for Vision Pro on a device at the hip. Instead of what they ended up with which was just a battery.
I could see this coming back around again where a future iPhone is all of the brains or a lot of the brains for AR glasses.
Contextify indexes your entire local Claude Code and Codex sessions into a local SQLite/FTS index as they happen. This lets you aim your agents at your real-time and complete historical AI-coding (or other project type) history.
It works purely locally but has an option for hosted or self-hosted cloud sync to read it on another machine, along with a hosted option if you want turn-key.
HN and some existing users requested a Windows client. In some cases it is because their self-hosted setups include Mac, Linux and Windows boxes that are all generating CLI AI transcripts (sometimes headlessly)
This product started as a macOS app; 1.8.0 brings it to Windows (x64 and ARM64) and extends the Linux CLI. It is free for personal use and has a source-available self-hosted option. I have commercial licenses available, and a few very much appreciated customers so far.
On Windows it's a background agent plus a CLI, with no history-browsing UI yet.
Lots of interesting development was needed to pull this off (including a shared Swift core!) in a way where it could be updated as easily as prior to the release. I'll share some of that in a comment below.
If you use Windows and would be open to giving this a go, I'd appreciate any feedback here or via rob@contextify.sh.
I have a bunch of skill tooling that build out iterm2 windows that cover worktrees specified in a project-config.yml. It even includes tab color.
More recently I added a skill to recover all of these windows in their prior state including resuming the conversation the agent was running. This is to recover from the need to restart from an OS update or something bad happened.
All of my skills are custom to my workflows except Contextify (more on that at the end) This extends to how I distribute them across multiple development machines.
They live in my `cli-ai-setup` repo alongside agent settings, git worktree tooling, iTerm workspace restoration (important, machines have to restart and crash sometimes), code review scripts, and machine setup guides. I use git to carry changes between machines and a setup script to symlink the skills into a shared directory that Claude Code and Codex both use.
I have a `skills-and-settings` skill specifically for deciding where new skills belong and how to make them available (project, global, application). I also have a custom `skill-create` skill that turns sessios into new skills or updates existing ones.
Importantly, I also have entire custom applications I have not yet made open source that my cli-ai-stack relies on. I do expect to distribute these so they live in their own repo and are symlinked or installed in as appropriate.
For maintenance, I've largely handled this manually and organically. When a skill is not performing, I'll use the context of the situation as the ~1 shot or pull in more examples for the ai:
This skill seems to not be performing as expected on [something].
This has happened a couple of times now use /total-recall to find similar recent situations for example [something I remember]"
Recommend updates to the skill and upon approval commit and push them...etc.
My other machines watch this repo and the symlink structure means that the updates are carried into live cli-ai sessions almost immediately.This past week I was exploring the automatic skill improvement behavior described in the Anthropic blog guest post with their partner org. I'd previously build a "dreaming" skill that works okay and think there may be some value yet to plumb there.
For skill creation, I have a skill that reads the official skill docs for both Claude Code and Codex. This way the skills are built to handle both platforms particularities. I automatically pull those docs into local Markdown daily, so it has a regularly refreshed reference for what each tool supports.
As mentioned above, I have built Contextify (https://contextify.sh) which provides a sql database of all of my Claude Code an Codex session transcripts across all of my development machines. The skill for this (/total-recall) is the most important skill I have and I use it constantly.
I also have found what I believe is a bad training bias in the design of release related CI workflows toward proof of release artifact provenance.
Both major frontier models love provenance programming in CI, so much that they will spin endlessly trying to solve basic CI functionality at the same time as ensuring SHA's match up across lengthy (often already complex) cross-system pipelines.
I had thought some of my durable context was causing this, and sought to strip anything that might be triggering this behavior.
But then I come upon some more work in release workflows comes up, and boom its back! I couldn't believe it, I called the AI out on it and it agreed it had been told specifically not to do this but was doing it anyway. It did kindly stop and remove the commit(s).
Somewhere, something was oversampled in training because the AI will try their damndest to build this stuff. The worst of it is that it can often involve lengthy, sometimes resource-heavy CI runs so the validation of this unnecessary stuff can have very long feedback loops.
And, sometimes you actually need the provenance. In this case, I've had success forcing the AI to split the work up into functional capability completely devoid of artifact ~chain of custody and get that right before attempting any kind of provenance work.
Bit of a rabbit hole on this, but the above cost me a lot of burned tokens so hopefully helps someone...or some AI.
The car started from a low-poly DeLorean glTF. I had the model rebuild the geometry into inline Canvas code rather than just dropping a model into three.js.
The flight itself is procedural: position/rotation, hover, pitch, wobble, flames, camera, etc. are all driven by code and can be scrubbed/debugged.
That link has a link to an interactive page where you can adjust the car’s path and when it reaches 88mph if you want to see.
It is not mobile friendly atm, though.
I would like to apply textures to the people and scenery in the gauntlet scene. Just haven’t had the time or tokens, as getting motion and pathing and camera angle right seemed the mvp.
I posted this previously and people wanted more info so here is how I built it, with a deeper dive into some film scene recreation I worked on focusing on a scene from Apocalypto:
https://banagale.com/cinematic-canvas-ai-film-animation.htm
I’d love to improve the look of these things.
On the Fable 5.2 eval summary, Opus 5 only beats Fable on SWE-bench multilingual and multimodal.
I primarily use the models via interactive sessions enhanced with custom tools and skill. For that Opus 5's benchmark superiority has not materialized into greater productivity and frankly has been quite a let down.
The outputs are too often unreadable even after adding recommended prompts. There is an ongoing problem with the heron_brook system prompt affecting orchestration. [1]
I've used Opus 4.8 since the second week Opus 5 was released.
Over this time, Fable 5 has been reliably fantastic. Both in planning and direct execution on complex changes across code and infra.
I'm a bit surprised that there doesn't (seem) to be a section discussing ~performance across different modalities. This system card and blog post too-often default to an API-based use case when the gander primarily experience Anthropic's models via interactive sessions.
I understand waiting to comment until Opus 5.1 is available and handles these problems, though I am hopeful that Anthropic will confront the elephant in the room on Opus 5's failure to delivery great interactive sessions and the widespread negative feedback on the release.
It would show the org is paying attention, taking steps to balance model evals between interactive and API use. Also, some empathy for customers that wasted time trying to make opus 5 work for them.
Or Codex from the ChatGPT app.
If you have supplied tooling on your system, you have access to all of this and more.
I think the idea is most people don’t have a machine up and available, nor maintain skills for interacting with their core services?
One thing that keeps me from adopting codex more deeply is the architecture around mobile access.
Claude Code makes this trivial /rc and you are done.
Codex requires you run the desktop app and additional authentication requirements. The result of this has been codex is almost always relegated to fleet worker rather than orchestrator.
I wonder if the emphasis on this work feature has something to do w the persistent hurdles to remote control if codex sessions.
Why bother creating hundreds of (presumably paid) accounts and the tooling to create and consume the data if not of tangible value to their core mission?
https://www.anthropic.com/news/detecting-and-preventing-dist...
For example, having young kids is not a good time to have a nice teak dining room table. (Or new wood floors for that matter) (or really anything nice you don’t want to hide in a locked box or room)
Also, pets, like a cat is wonderful. But without great diligence a more lower mid-tier couch will get clawed up the same as an ikea special. At least it may be more comfortable in the meantime.
I suspect wealth may be having your cake and eating it too—-you have fancy furniture that kids or pets wreck and you file that under dc.
All of it screams for automated materials reprocessing. One can only hope that the collision of robotics and AI will reduce the amount of landfill we create.
Or that the man design was done for so long by a close collaborator of Larry’s. And that the man base has been frequently spectacular both on fire and not.
The historical perspective and old photos are great but I’m not much for the conclusion. I would be willing to bet the title and byline are editorialized by the publisher.
You can still go out there with not much, be new and participate and just smile at some famous or not person also on their way back from the Portos.
There is a lot of money involved, no doubt and there is a lot of bullshit too, like art cars that do fundraisers and take volunteer contributions only to remain largely private vessels for a particular crew, a few celebs etc.
And there’s a lot to not like about the org or how art grants are handled, it’s a big system it has the exploits and nepotism that might infect anything scaled and offering power.
But if you go and do not look closely and just give in small ways none of that would likely be visible to you.
You could have the time of your life. Or not. Which I think suggests the opportunity for adventure and chaos still remains at that thing in the desert.
Rainbow Gathering is barter-based.
So, the premise is a bit of a self-own.
That said, the solution may still have merit regardless of the model that sent them looking.
Whether the distillation has constituted "attacks" or has or will meet the bar of "stealing" IP is not super interesting to me, though.