HNHacker News
TopNewBestAskShowJobs

alexsmirnov

69 karma · joined May 21, 2011

submissionscomments
alexsmirnov··on GLM-5.3 and the spread of advanced cyber capabilities
This is exactly the problem that Anthropic hides in their article. Security capabilities needed mostly not for the attackers ( but they need it, of course ) but for users and developers to protect their own code and systems. I do use glm ( from openrouter and abliteration.ai ), and kimi model for security reviews on pull requests. Claude rejected even to edit instruction files. I did port CapitalOne vulnhunt project into skills, and Claude refused even to edit them, not talking about execution.
alexsmirnov··on Show HN: PaperMono, e-ink fridge magnet shopping list with mobile web page
KISS ! I do use whiteboard + camera. Much simpler and reliable.
alexsmirnov··on How much oil-market buffer is left?
I do have SUV ( Pathfinder ) for camping, skiing, kayaking, and other long trips. And 2 seats electric Smart for local commute - groceries, school drop off/pickup, gym visits etc. This is 95% of my family car use. The 90 miles charge is about $2, I do it 2 times a week. Yes, this is CA
alexsmirnov··on Ask HN: How do you manage skills files?
I did create special repo that has all skills/agents/scripts and MCP configuration. All agents ( currently 3 in use: claude code, opencode, and pi ) packed in docker container, with artifacts required by project technologies, and task at hand ( planning, code review, documentation management, web design, ... ). Nothing but a small file with list of technologies committed to project.

For evaluation, there is command to record session, repository commit, and observed problems that stored in special database, so each session can be reproduced. Developer commit reports, I do analysis, refine and evaluate system.

alexsmirnov··on The growing divide between AI hype and software engineering reality
In my feeling, it is decades old. With LLMs and agentic coding, it's deja vu of some 200x working with with inexperienced offshore teams. The same misunderstanding problems, the same corners cut, the same attempts to present bullshit as a "production ready", and the same "Yes, Sir, you are absolutely right" answer to criticue.
alexsmirnov··on "Opus 5 is a really bad model"
"The single worst issue: the system prompt never explicitly tells it to ignore auto-injected files like CLAUDE.md or *.rule.md." - Yes, it is! Right after the content of CLAUDE.md, Claude code inserts:

<system-reminder> IMPORTANT: this context may or may not be relevant to your tasks. You should not respond to this context unless it is highly relevant to your task. </system-reminder>

alexsmirnov··on UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
I did run both kimi-k3 and glm-5.2 with Capital One vulnhunt [1]. No rejections, they did find the same problems on my project that I used for testing. gpt-5.6, gemini pro, and opus all rejected to follow. gpt even declined to edit skill files.

[1] https://www.capitalone.com/tech/open-source/announcing-vulnh...

alexsmirnov··on Microsoft deleted my account and OneDrive
Shit happens. I had my iMac and Time machine hard drives died 2 days apart...
alexsmirnov··on Write code like a human will maintain it
The problem of duplicated code described in the article, goes in a different way in reality: AI does not update 4 places in the same way, but implement them a slightly different. I found diverged business logic all the time. For example, file upload dialog for a document with the same meaning: in one place, it accepts pdf only. Another allows to upload pdf or docx. The last accepts pdf, doc, docx, and txt.
alexsmirnov··on The short leash AI coding method for beating Fable
> This happens but far less often than it used to, and the case for full autonomous agents is getting stronger, not weaker.

This is that I do not see. My journey, just couple weeks ago, Claude Code + Opus 4.8. The task was not too complicated, 4 new API endpoint plus events streamed from client by websocket.

1. Multiply iterations on API definitions, refine request/response models, database schema, whole flow. A lot of corrections, removing contradictions, manual changes in document. Opus went of rails all the time. 500+ lines final document

2. API Integration tests. Once again, back and forth. AI was unable to create tests directly from document, so 2 iterations: Create placeholders with Given-When-Than comments, review an correct by hand. Second iteration was to implement tests. A lot of mistakes corrected after review.

3. Implementation. CC got api document, working tests ( modifications blocked by hook ), 6+ "best practices" skills ( most promptly ignored ), "rubber duck" and "code simplifier" agents, pre cooked scipts to run tests, linter, and check for compilation errors. Plan + execution + review, multiply corrections on the way. Feature implemented, all tests passed.

4. Code review. At average, found one issue per 20 lines of code. Not count code style, things like: Use in memory semaphore in kubernetes service (deployment described in CLAUDE.md ), 8 database calls to update the same record during a single request. One column at a time! Read-modify-save without transaction. Mistakes in business logic, failure recovery, authorization.

The result: almost one workweek, $100+ in tokens, and one thought: did it worth the effort ? P.S. I have a team of 2 developers. Just got PR to review from one of them. 80% slop.

alexsmirnov··on Monetization Gateway: Charge for any resource behind Cloudflare via x402
If I pay for vending machine by corporate card on a business trip, it looks more like B2B
alexsmirnov··on The Radiation Exposure Lie
From the first hand: I was in the air base service regiment, Ovruch ( 30 km from reactor ), 1 year in service at the time of disaster.

The average numbers have a little sense. By our measurements, it was't an "even" distribution, but some "hot spots" that had 10x time radiation level than surrounding territory.

The article focuses on cancer, but for me and my buddies, the worst was an impact on immune system. In the six months after disaster, the healthy boy before, I got: pneumonia, chicken pox ( and 2 more with me ), furunculosis ( the whole company got it as well ), and endless flu/fevers. The weak immune continued for about 10 years.

alexsmirnov··on Where is the AI jobs crisis?
I more like my high school math teacher explanation: "Put your one hand on dry ice, another in the boiling water. On average, you feel warm and cozy"
alexsmirnov··on Using Git's rerere feature to escape recurring conflict hell
We usually squash feature branches before merge. To squash before rebase, I use git reset --soft $(git merge-base develop HEAD) && git commit && git rebase develop - you have to resolve final conflicts only
alexsmirnov··on OpenAI backs Illinois bill that would limit when AI labs can be held liable
Much longer than that, and was available way before an internet. I graduated STEM high school in St. Petersburg in 1981, and I had several classmates who were big funs of chemistry. That they were able to create from textbooks, school lab ingredients, and understanding:

WWI era poison gas, tear gas, potassium cyanide, and bunch of explosives like acetone peroxide.

LLMs have all of that knowledge in training data

alexsmirnov··on Qwen3.6-Plus: Towards real world agents
Exactly.

I did create my own MCP with custom agents that combine several tools into a single one. For example, all WebSearch, WebFetch, Context7 exposed as a single "web research" tool, backed by the cheapest model that passes evaluation. The same for a codebase research

Use it with both Claude and Opencode saves a lot of time and tokens.

alexsmirnov··on Study: 'Security Fatigue' May Weaken Digital Defenses
Almost instantly, compared to my experience working for a big health care provider... I waited 6 moths for IT department to allow me install development tools on work laptop.

And while security rules created enormous roadblocks for work, whey also left enough holes to be exploited. Before getting required permissions, I managed to create dual boot with linux and share files between 'approved' and 'illegal' systems

alexsmirnov··on Ask HN: AI productivity gains – do you fire devs or build better products?
> you start off checking every diff like a hawk, expecting it to break things, but honestly, soon you see it's not necessary most of the time.

I see it's necessary ALL the time. The AI generated code can be used as scaffolding, but it's newer get close to real production quality. The expierence from small startup with team of 5 developers. I do review and approve all PRs, and none ever able to pass AI code review from the first iteration.

alexsmirnov··on Claudetop – htop for Claude Code sessions (see your AI spend in real-time)
The calculation misses subagents tokens, that can be a significant differences. Better to parse session jsonl files ~/.claude/projects/<project>/<session id>.jsonl and ~/.claude/projects/<project>/<session id>/subagents/<agent id>.jsonl
alexsmirnov··on Code Review for Claude Code
This mostly matches my own estimates for pr-review command that I use. But it's pretty sophisticated: 6 specialized agents, best practices skills, CVE database, bunch of scripts. To reduce cost, most of agents use cheap open source models.
alexsmirnov··on When AI writes the software, who verifies it?
Actually, they extremely bad at that. All training data contains cod + tests, even if tests where created first. So far, all models that I tried failed to implement tests for interfaces, without access to actual code.
alexsmirnov··on What Claude Code chooses
Considering how little data needed to poison llm https://www.anthropic.com/research/small-samples-poison , this is a way to replace SEO by llm product placement:

1. create several hundreds github repos with projects that use your product ( may be clones or AI generated )

2. create website with similar instructions, connect to hundred domains

3. generate reddit, facebook, X posts, wikipedia pages with the same information

Wait half a year ? until scrappers collect it and use to train new models

Profit...

alexsmirnov··on Why Developers Keep Choosing Claude over Every Other AI
I do in similar way, connect claude code to litellm router that dispatches model requests to different providers: bedrock, openai, gemini, openrouter and ollama for opensource models. I have special slash command and script that collect information about session, project and observed problems to evaluation dataset. I can re-evaluate prompts and find models that do a job in particular agent faster/cheaper, or use automated prompt optimization to eliminate problems.
alexsmirnov··on My AI Adoption Journey
For me, AI is the best for code research and review

Since some team members started using AI without care, I did create bunch of agents/skills/commands and custom scripts for claude code. For each PR, it collects changes by git log/diff, read PR data and spin bunch of specialized agents to check code style, architecture, security, performance, and bugs. Each agent armed with necessary requirement documents, including security compliance files. False positives are rare, but it still misses some problems. No PR with ai generated code passes it. If AI did not find any problems, I do manual review.

alexsmirnov··on LNAI – Define AI coding tool configs once, sync to Claude, Cursor, Codex, etc.
I did create and actively use a similar tool, but with different purpose: configure AI tools for each team member to use the same code style and architecture guides across projects. It includes: - build docker images for claude code and opencode dev containers. - creates custom MCP server that works as a proxy and combines several tools into a single one ( for example, web search, fetch, and context7 tools exposed as a single "web_research" that invokes custom code to answer question ) - copy code style, documentation, and best practice rules for technologies used in our projects - deploys a bunch of helper scripts useful for development - configure agents, skills, hooks, and commands to use those rules. Configuration changed per "mode" : documentation, onboarding, code review, and web development all have different settings. - run AI tools in docker container with limited permissions - feedback tool to generate session report, that is used for automatic evaluation and prompt optimization.

This came out of necessity, as active using of AI assistants in uncontrollable way significantly degraded code quality. The goal is to enforce the same development workflow across team This is internal tool. If someone interesting, I can create a public repo from it

alexsmirnov··on Ask HN: Do you also "hoard" notes/links but struggle to turn them into actions?
I do use Obsidian on pair with Claude code and git.

I organize notes by tags, folders, and links from tree of "map of content" notes. Those documented as rules for AI. All notes came to "Inbox" folder, and from time to time I run special script that checks inbox, formats notes, tags them, and put in the most appropriate place. "git diff" to check results and fix mistakes, reset if it went wrong.

As notes organized by the limited number of well defined rules, they became easy to search and navigate by AI. Claude Code easily finds requested notes, working as advanced search engine, and they became a starting point for "deep research" : find relevant notes, follow links, detect gaps, search internet. Repeat until reach required confidence level.

The most advanced workflow so far is combination of TRIZ (Theory of Inventive Problem Solving) + First Principles Framework. Former generates ideas and hypotheses, later validates them and converge on final answer.

alexsmirnov··on The Code-Only Agent
This was implemented far ago, at least by huggingface "smolagents". https://huggingface.co/docs/smolagents/index . I did use them, with evaluations. For the most cases, modern models tool call outperforms code agent. They just trained to use tools, not a code
alexsmirnov··on Don't fall into the anti-AI hype
This is exact the impression that I got. Every question or task given to LLM returns pretty reasonable, but flawed result. For the coding, those are hard to spot but dangerous mistakes. They all look good and perfectly reasonable, but just wrong. Anthropic compared Claude Code to a "slot machine", and I fell that AI coding now is something close to gambling addiction. As small wins keep gambler to make more bets, so correct results from AI keep developers to use it: "I see it made correct solution, let's try again!" At a startup CTO, I review most of the pull requests from team members, and team uses AI tools actively. The overall picture strongly confirms your second conclusion.
alexsmirnov··on Why users cannot create Issues directly
This is not about understanding the message, but switching user mental activity. I go myself in the similar situations many times. One example: I tried to pay my bills in online bank application, but got into error. After several attempts, I did read message and it say "Header size exceed..." . It give me clue that app probably put too much history into cookies. Clear browser data, log in again, and all got works.

Even when error message was clearly understandable for my expertise, it took surprisingly long tome to switch from one mental activity - "Pay bills", to another - "Investigate technical problem". And you have to throw away all short memory to switch into another task. So all rumors about "stupid" users is direct consequence from how human mind works.

alexsmirnov··on The Gorman Paradox: Where Are All the AI-Generated Apps?
A lot of discussions around vibe AI coding flaws: awful architecture, performance problems, security holes, lack of maintainability, bugs, and low code quality. All correct, but none of those is matter if:

- you create small utility that covers only features needed only for you. As many researches show that any individual uses only less than 20% of software functionality, your tool covers only 10-20% that matters for you

- it only runs locally, on user computer or phone, and never has more than one customer. Performance, security, compliances do not matter

- the code lies next to application, and small enough to fix any bug instantly, in a single AI agent run

- as a single user, you don't care about design, UX, or marketing. Do the job is only matter

It means, majority of vibe coded applications run under radar, used only by a few individuals. I can see it myself: I have a bunch of vibe code utilities that never intended for a broad auditory . And, many of my friend and customers, mention the same: "I vibe coded utility that does ... for me". This means a big consequences for software development: the area for commercial development shrinks, nothing that can be replaced by the small local utility has a market value.

Page 1 of 2Next →