HNHacker News
TopNewBestAskShowJobs

mohsen1

2,124 karma · joined June 16, 2012

https://tsz.dev azimi.me github.com/mohsen1
submissionscomments
mohsen1··on Early rogue AI agent activity and attempts to hack found on urlquery.net
> agents were not told to 'go hack'

I was referring to the HuggingFace incident.

> none of the alignment techniques that are applied to models today actually work

none of techniques to autonomously drive a car was/is not working for a long time. no company came out and said 'this is impossible to do, let's change the regulations'.

mohsen1··on Early rogue AI agent activity and attempts to hack found on urlquery.net
https://podcasts.apple.com/us/podcast/the-ezra-klein-show/id...
mohsen1··on Early rogue AI agent activity and attempts to hack found on urlquery.net
I listened to Jensen Huang's interview with Ezra Klien and it was so refreshing to hear it from an engineer. Jensen framed it as OpenAI's responsibility and recklessness which I agree with. Jensen thinks it's an engineering problem to build better sandboxes.

It's irresponsible for OpenAI to give unaligned agents a prompt to 'go hack' and internet access. They know better, so I am thinking they might have other intentions to let those swarms have any sort of internet access.

mohsen1··on Meta VR Glasses
"100 grams" is the headline here
mohsen1··on Gemini 3.8 text-to-speech
Why tho? AI Studio has this model available and I don't need to paste the API key anywhere

https://aistudio.google.com/generate-speech?model=gemini-3.8...

mohsen1··on Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
> The SAME model may even give different answers to the same prompt when asked multiple times

temperature?

mohsen1··on Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
see sister comment's response https://news.ycombinator.com/item?id=49786548
mohsen1··on Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
This is less true for modern posttrained models. Model identity can be explicitly reinforced during posttraining. Qwen's own finetuning docs include identity training examples, and Qwen models have been trained with system prompts that explicitly say things like "You are Qwen, created by Alibaba Cloud."

So a model correctly identifying its family doesn't necessarily mean it inferred that from pretraining.

I think with Jev, they took a posttrained model and trained it further, so it did not forget about its earlier knowledge during Owen's own RL.

mohsen1··on Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
I can't find it but saw that if you give Jev English alphabet as choices and ask it in a loop what model it is, it would say Qwen

also tried myself: https://console.typesafe.ai/playground?share=shr_1690a3160f1...

mohsen1··on OpenJev
There is an open PR for VLLM to do this via DefussionGemma

https://github.com/vllm-project/vllm/pull/57250

mohsen1··on OpenJev
yup https://github.com/vllm-project/vllm/pull/57250
mohsen1··on Bend 2 and the Vibe-Coding Trap
> A little research before vibe-coding an entire language and compiler could have substantially improved the result because the author would have known what to ask for.

A little research before writing and publishing a personal attack like this could have substantially improved the result because the author would have known what they're writing about

Victor is not a formal verification noob as this article suggests

mohsen1··on iOS 27, iPadOS 27, and macOS 27
if you were disappointed by iOS dictation, give it another shot after this update. It has improved a lot. I've been using it a lot more than before
mohsen1··on DeepSeek v4.1 Flash
I am speculating but hard to not see that DeepSeek is brewing a full Pro model with those new techniques to come out right around the time of Anthropic and/or OpenAI IPO to tamper the excitement for their offering.
mohsen1··on Unified Arabic
Turkish as well. They moved from Arabic to Latin script.
mohsen1··on 216M Spy TVs – The LG Smart TV Problem [video]
I absolutely loathe my LG smart TV. Just switching to another input requires 3 button press.

I hope someone uses their free time to hack the OS and offer a clean and simple interface to make it a “dumb” TV.

mohsen1··on What will be left for us to work on
I have a question about “domain expertise” as a component of future knowledge worker requirements. How does one gain such expertise in a context where thinking is expected to be delegated to AI (shifting from problem solving to question asking, as noted in this paper)?

How does one learn to pose the right questions when basic ones are rarely “manually” answered? that is, without an AI assistant’s help

This pattern appears in schools, where AI interferes with human development that typically demands long and difficult effort of actually answering questions

mohsen1··on Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights
> The company on Wednesday confirmed speculation that the Ox Alpha model is a new iteration of its GLM series and said it will release the weights for it tonight, in response to queries by Bloomberg News.

Seems legit.

It's really hard to know how good it is. So much hype around it.

mohsen1··on "The Persian MâR-Nâmeh Or, the Book for Taking Omens from Snakes" (1892)
wow i totally missed "It Was Just an Accident". Have to watch asap.
mohsen1··on Interviewing Engineers in the AI Era: Lessons from a Year of Rebuilding
Hmmm... I had a different experience. They had a fully automated environment where you ha to write test that passes some tests. No human involved. And the time requirement was insanely tight
mohsen1··on When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
That's a lot of words to say you need larger set of questions for today's models. 300 questions won't be enough to find the difference
mohsen1··on Kimi Linear: An Expressive, Efficient Attention Architecture
> one of the basic tenants of algorithm development was that you can't just brute-force your way towards a solution for some complex problems

Mote-Carlo is pretty useful still. Not sure if your statement holds

mohsen1··on Kimi K3 exploited the latest Redis server
Do you think any programmer really understands how their program works end-to-end? At some abstraction layer, we're all clueless. There are many layers between what you type into the text editor and the actual CPU ticks that make your program work. I bet nobody fully understands the whole stack.

Now that that text editor accepts English, we're all calling each other names, etc.

mohsen1··on Claude Opus 5
RL. Lots of RL
mohsen1··on Sleep regularity is a stronger predictor of mortality risk than sleep duration (2023)
I thought the same. If your life is so in order that you routinely sleep on the same interval, perhaps your life is not as stressful as others who sleep more chaotically
mohsen1··on Ask HN: What Are You Working On? (July 2026)
This is a Chrome extension that records lots of details in a usage session. Stuff like network calls, console logs, screenshots and also optionally screenshots and user narration

Tools like this exist, but every one I tried is uploading the session details somewhere in their cloud and try to monetize this.

So I built the version I wanted: free, open source, and local. There is no account, no backend, no telemetry. Sessions live in IndexedDB in your browser and exported as a zip.

What it records:

* Clicks, typing, page changes, network requests and responses, console errors screenshots, video with sound

* Your voice, transcribed and placed next to what you were doing at the time

* Annotations: Arrows and boxes you draw on the page's screenshot

Note: Passwords, auth headers, and tokens are masked at capture time

All events are lined up in a timeline with timestamps

At export you pick a detail level with a live token estimate, so a long session still fits your model's context window.

.

Repo: https://github.com/mohsen1/session-recorder-chrome-extension

mohsen1··on Cargo-nextest: 3x faster than cargo test, per-test isolation, first-class CI
Thanks! Any pro tips for sharding? I landed on single job because couldn't get cache to work properly for shards to be fast enough to worth it
mohsen1··on Cargo-nextest: 3x faster than cargo test, per-test isolation, first-class CI
I love nextest. without it my CI could take hours

https://github.com/tsz-org/tsz/actions/runs/29002057457/job/...

watch it running 32.5k unit tests without breaking a sweat!

mohsen1··on Cloudflare Drop
This is perfect for my Chrome Extension for recording sessions and capturing screenshots, audio narration and videos. The output is a zip file with everything so if user wants to share they can use this

https://github.com/mohsen1/session-recorder-chrome-extension

I built above chrome extension because anything in this area has been trying to monetize the solution. I wanted a free and open source version of this to exist.

mohsen1··on GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance
I have a philosophical problem with adaptive thinking. It’s a dumb guess for how much thinking budget to allocate ahead of thinking. At least in the context of LLMs there is probably no way of knowing how much thinking (token generation) is needed. The problem space is infinity vast, similarly of two prompts is not going to help any LLM decide how much thinning is needed. Models already stop thinking before hitting the thinking budget.

Why there is so much effort in making adaptive thinking happen and don’t we train models to produce the end of thinning token better?

Feels like a bandaid. We need models to be trained to do a reasonable amount of reasoning (no pub intended):

    reason

    estimate remaining uncertainty

    continue?

    reason more

    repeat
Page 1 of 15Next →