HNHacker News
TopNewBestAskShowJobs

sbinnee

293 karma · joined March 17, 2022

ML and AI researcher

Based in Seoul, Korea

https://sbinnee.github.io/

submissionscomments
sbinnee··on Show HN: GridTravel- A community based travel app for users to share routes
I love the idea. Unfortunately, not available in my country. I hope it to be successful and see this app in my country in the future.
sbinnee··on Build iterative repair loops with Codex
The term "repair" immediately caught my eyes.

> Repair: apply focused edits to a copied artifact using the findings and the latest validation feedback.

I am used to seeing "fix" if it's broken or "optimization" otherwise. I was making a guess maybe this is their strategy to make it more approachable for non-techie. But I have no idea why.

FYI, Anthropic recently introduced "outcomes"[1]. There are all chasing evals.

[1] https://platform.claude.com/docs/en/managed-agents/define-ou...

sbinnee··on AI is making me dumb
I've reserved a time span next month to learn TypeScript. I don't intend to rule out AI entirely during that process. My plan is to read a book from cover to cover, and then write code. I am pretty sure I heard of this method from Mitchell Hashimoto on some podcast. I am exited and scared at the same time, because like OP I've spent a lot of time prompt-coding.
sbinnee··on Codex is now in the ChatGPT mobile app
I don't like this direction. For accessibility aspect, sure it is good. But Codex is a coding product. I am increasingly concerned of lack of reviewing practice. I doubt that a mobile app is good for reviewing code changes.

> Stay connected to active work from anywhere

... (and anytime because it's on your phone). No thanks.

sbinnee··on A few words on DS4
It is a big thing for sure to have a competitive local agentic model. I've replaced gemini 3 flash preview with DeepSeek v4 flash for all of my personal use cases. Starting from chat app, language learning, and even hobby coding. For coding, I couldn't get decent results no matter which sota latest models I used before. It's not close to Opus or Codex models. It's a flash model and makes mistakes here and there (I just saw `from opentele while import trace`, new Python syntax!)

But I found its tool calling is reliable than other oss models I tried. I assume that it attributes to interleaved thinking. Its reasoning effort is adjusted automatically by queries. I enjoy reading these reasoning traces from open models because you can't see them from proprietary models.

I would love to try DS4 so bad. Well, I don't have a machine for it. I will just stick to openrouter. I wish I can run a competitive oss model on 32GB machine in 3 years.

sbinnee··on Scorched Earth 2000 is back
OMG. One of my favorite games. It was fun to explore all the weapons and utilities with my brother.
sbinnee··on MacBook Neo Deep Dive: Benchmarks, Wafer Economics, and the 8GB Gamble
12gb bump soon? I don’t see that happening. It’s Apple.
sbinnee··on Googlebook
As a linux desktop user, I would buy it only if I can wipe ChromeOS out of it.
sbinnee··on Googlebook
I guess it's just mobile chip and everything AI related connects directly to google services through internet.
sbinnee··on I returned to AWS, and was reminded why I left
I also tried. Only service I use is s3 for personal backup. I pay around 15 cents per month.
sbinnee··on Agents need control flow, not more prompts
I have been telling this to my team that 1000 lines of instructions are deemed to fail no matter how great of instruction following capability of a model. I have been reviewing hundreds of line changes daily basis for about a month. I couldn’t help becoming a prayer.
sbinnee··on DeepClaude – Claude Code agent loop with DeepSeek V4 Pro
After some time replacing gemini 3 flash preview with deepseek v4 flash for a chat model, the biggest difference is the auto reasoning effort. Gemini flash is super fast and perfect for a chat model. But when I need some thought experiments with a handful of constraints, it struggles a bit and I switch to sonnet. But with deepseek v4 flash, it can do long complex reasoning and it gets things often right. Generating a lot of reasoning tokens means that it takes a lot of time of course. But I am happy to find a cheaper model and excited to try something other than gemini flash. Gemini flash has been so good that I was locked on it for a while.
sbinnee··on Ghostty is leaving GitHub
GitHub has become a place where you seek people’s attention. There are other places you can freely host your projects. GitLab was always available. I just haven’t logged in for I don’t know how long. An open source project is essentially a show window to the internet by a lonely developer. Ghostty has already established a great community. It’s already on display on a skyscraper. The project is mature enough that it needs a dedicated discussion forum or something like that. I am excited to see where it will find home and how it will evolve.
sbinnee··on Is my blue your blue?
In Korean, we have an adjective "푸르다". It is somewhere between blue and green. You can say trees are that, oceans are that. It also means unripe.

Yeah, so to me, tortoise is definitely blue.

Edit: typo tortoise -> turquoise

sbinnee··on Quarkdown – Markdown with Superpowers
> (academia, hating everything modern, will also hate you if you use typst)

I chuckled. I'd love to try out typst when the time comes. But for writing a journal paper, it's still going to be latex.

sbinnee··on DeepSeek v4
Price is appealing to me. I have been using gemini 3 flash mainly for chat. I may give it a try.

input: $0.14/$0.28 (whereas gemini $0.5/$3)

Does anyone know why output prices have such a big gap?

sbinnee··on Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model
I don’t think I ever heard you said excellent for the pelican test. It looks excellent indeed!

The trend went to MoE model for some times and this time around is dense model again. I wonder if closed models are also following this trend: MoE for faster ones and dense for pro model.

sbinnee··on Claude Opus 4.7
The comment section is already long, but I knew that I could found comments about "hmm" that I started noticing. Yes, it is so irritating to me too. Also, one additional thing I noticed was that verbose information has been more and more being obfuscated. I run CC with --verbose option for months, and I can see verbose mode is not verbose anymore. I wish I can do -vvv maximum verbose mode.
sbinnee··on Ask HN: Who is using OpenClaw?
I was going to try hermes agent after hearing OpenClaw constantly breaks and this hermes buzz is a better one. I t all sounds a lot of maintenance work.
sbinnee··on Guide.world: A compendium of travel guides
It looked cool, and I thought that it might be a new community where articles belong to this site. But when I clicked two articles, Seoul and Singapore, both were behind paywalls. So it seems it's just an aggregation of internet articles it seems?
sbinnee··on Make tmux pretty and usable (2024)
Though I also customize my tmux setup, the best way to use tmux is just to learn and remember the basics. Once you change the prefix bind or any other basic binds, you will have hard time on a new machine.

Btw, you can place tmux config at ~/.config/tmux/tmux.cong. No reason to clutter home dir.

sbinnee··on WiiFin – Jellyfin Client for Nintendo Wii
I love this kind of project. I am pretty sure the developer had a Wii console sitting around somewhere and thought about how to make it useful again. Wait, I have a PS2 sitting around somewhere…
sbinnee··on GitHub Stacked PRs
Is this going to be a part of triage task? If so, it makes sense. Whether a human developer or an AI made a big PR, AI goes review it and if necessary makes stacked PRs. I don’t see any human contributors using this feature to be honest because it’s an extra work and they should have found a better way to suggest a large PR.
sbinnee··on European AI. A playbook to own it
I lived and studied in France. So it's only natural for me to try out Mistral's every major updates. I had the same sentiment with you, that their models were just a bit lagging. But on the other hand, I understand their shift. Their value is not in SOTA coding, math, or puzzle solving performance in my opinion. They will catch up. To me it makes sense that they focus on something else, how to scale through talents and propagate their models with policies in Europe.
sbinnee··on South Korea introduces universal basic mobile data access
Crazy even to me, a Korean. I just woke up and saw this news on HN. Over the years I watched the price of data was going down drastically in Korea. I had always complained that data in France was much cheaper like 30gb for 10 eur. Then when I came back after around the end of pandemic, the price of data in Korea was actually quite cheap.

Do I know why? No idea. The article alluded fast AI adoption but even senior Korean citizens are all addicted to youtube videos. Soon they will start using AI. Young people are already heavily using AI for everything. So I don’t think it’s for AI adoption.

The recent hacking incidence was a big one, true. But the price had been going down even before.

sbinnee··on Muse Spark: Scaling towards personal superintelligence
It is fair to think so because that is what everyone is doing. But being Meta and considering Llama, if MSL is going to keep releasing models and wants to join back the AI war, they may actually open weights just to get more attention. Once they establish a sizable community, they can start guarding their frontier models.
sbinnee··on Muse Spark: Scaling towards personal superintelligence
> but you can try it out today on meta.ai (Facebook or Instagram login required).

I guess I will have to wait. I hope at least soon it will be available on Openrouter. Overall, I am really excited to try it out.

sbinnee··on People Love to Work Hard
People have different priorities, different purposes, and different passions. Maybe a success is as simple as to build a team that has some overlaps of these things.
sbinnee··on The Claude Code Source Leak: fake tools, frustration regexes, undercover mode
I had a good laugh. I am too polite but I do remember using wth a few times in the past week. haha
sbinnee··on Gonon: Building a Clock with No Numerals
I wonder if this article got inspired by the recent movie Project Hail Mary?

Another thing related to this subject is the frequency in watch making. I own a manual winding watch that I wear everyday. It is certainly an engineering marvel. These watches are ticking by the hair spring and its frequencies are targeted to 2.5Hz to 4Hz (5 times per second, or 8 times per second). I don't know the rationale behind these numbers. I guess that they must have been a combination of engineering constraints and finding a good balance to keep every second accurate.

← PreviousPage 3 of 8Next →