HNHacker News
TopNewBestAskShowJobs

b89kim

12 karma · joined March 8, 2026

submissionscomments
b89kim··on Navier-Stokes – Tristan Buckmaster [pdf]
OpenAI addressed the dispute in official and denied direct usage of data. However, they acknowledged the possibility that session data was used for model improvement.

> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .

b89kim··on Navier-Stokes – Tristan Buckmaster [pdf]
Sebastien denied most accusations and apologized only for inappropriate wording, providing his perspective.

- He insisted that OpenAI initiated research based solely on rumors and never accessed their Codex sessions.

- He mistakenly believed Tristan and Levent were solving the same problem in Anthropic.

- proposed two option. (1) Tristan becoming the lead author to revise OpenAI’s work, or (2) OpenAI providing internal model to support and bridge their research.

- Sebastien insisted there was no intention to alter authorship. He was simply uncomfortable sharing OpenAI’s work,model with an Anthropic researcher. Additionally, He believed Levent’s credit seemed limited as their work focused on Euler.

- complained that negotiations with Tristan and Levent were difficult

[1] https://xcancel.com/SebastienBubeck/status/20973794116915163...

b89kim··on Navier-Stokes – Tristan Buckmaster [pdf]
If OpenAI's proposal is true, it strongly implies they accessed Tristan's private session logs. OpenAI's 'Terms of Use' allow using Codex session logs for model improvement. However, using private user sessions to develop research would still be highly controversial.

Separately, Terence Tao noted there is a low probability OpenAI actually solved the general regularity problem.

b89kim··on Navier-Stokes – Tristan Buckmaster [pdf]
- Tristan and his co-author (Harvard/Anthropic) developed a theoretical framework and validated it using Codex and Claude.

- OpenAI did related research around similar timeframe.

- Tristan claimed OpenAI offered a proposal that included dropping the Anthropic-affiliated co-author.

- Sebastian (a prominent OpenAI researcher involved) denied these claims.

- Tristan have no concrete evidence that OpenAI accessed their session.

- OpenAI's theory may hold up, but it will require long-term validation to confirm.

b89kim··on Reasons robotics is hard
Main problem of humnoid is price. Price starts from 20,000$ and we need jetson thor(On-device) or 5090 for learning capabilities. Their ability per cost is not greater than human. The other problem is hard to predict/simulate real world contact physics. Thus, most of company use VLA for task with contact. Even if VLA is advancing rapidly, it's still immature to be deployed in practical.
b89kim··on Ask HN: Who is using FPGA for ML inference?
You could build small ML model like talos-v2. But larger model/LLM requires much more engineering cost than you expected. Optimizing HDL has too many control knob/param to solve by RL.
b89kim··on Bonsai 27B: A 27B-Class model that runs on a phone
I had no failure mode in 3.6 9B thinking with llama.cpp. After release, there were updates for both model and llama.cpp.
b89kim··on Building a robotics research setup that lives next to my desk
If you're using depth, you're better off starting with a diffusion policy (DP). We benchmarked ACT, DP, pi0,pi05 on the same task, ACT underperformed in most cases.

There is already plenty of research around multimodal diffusion policies. While DP typically doesn't require pre-training, you can boost data size by depth estimation model+Open data.

b89kim··on Building a robotics research setup that lives next to my desk
Adding a depth channel rarely yields a massive performance gain, likely due to data scarcity and the fact that modern VLAs are good at guessing distance directly from RGB. I have used multiple RGB-D cameras, but it is hard to get stable images without jitter. Depth can still be useful for high-level reasoning. PI also uses bounding-box or segmentation data from PI-05 for that.

PI smartly combined discretized tokens with flow-matching for efficient training, and it works well in most cases. Still, end-effector representation may be better for teleop with devices like a SpaceMouse, VR, or VibeTracker. PI-07 also supports EEF, but I am not sure how much data is needed to fine-tune PI-05 for that.

I'd suggest starting with the default pi05 model. Data strategy is probably more important than model improvements. Since VLA performance is highly dependent on the data/action distribution and it's easy to modify. After that, you can add high-level reasoning like PI05. I visited a Chinese VLA company that already adopted the PI-05 approach, and it works quite well in practice.

b89kim··on Building a robotics research setup that lives next to my desk
- A single arm is sufficient for validating basic Pick/Place tasks, but more complex scenarios require Bi-arm

- Calibration is not required for VLA models.

- RGB or Stereo RGB inputs are sufficient for ACT, DP, and PI0/PI05.

- ROS2 is not strictly required, but it can be useful for sharing/co-developing codes. For instance, the Stanford team built a custom framework for diffusion policy instead. I also developed similar framework because ROS2 is not optimized for bi-manual manipulation or VLA workloads.

b89kim··on Building a robotics research setup that lives next to my desk
I could confirm 50-100 demonstrations are enough for fine-tuning pi0/pi05. I did research with aloha and humanoid. It works from 20~40ep(5~10min) but success rate would be 70~80%. Pi0 tech paper suggests to use over 1~4 hours of data. I could get 95% success rate for pick&place with 1 hour of humanoid. Anyway, required hours for good SR depend on generality of data. Long Horizon task over 5 min is not working as paper because PI removed high level(subtask) reasoning part in released pi05.
b89kim··on Raspberry Pi 5 – 16GB RAM
The Raspberry Pi is a single-board computer with native support for UART, SPI, I2C, CSI, and more. There's a large ecosystem of HATs, sensors, and peripherals built specifically for it. Most mini PCs rely on USB for peripherals, which isn't ideal for embedded use cases. Additionally, mini PCs tend to be discontinued within 2–3 years, whereas the Pi has much longer lifecycle over a decade. It's closer to an ARM development board, and those alternatives aren't cheap.

There are plenty of Pi clone boards at lower prices, but they have smaller communities and less documentation. When you hit an unexpected problem, it can be hard to find solutions or get support.

b89kim··on Pyodide: a Python distribution based on WebAssembly
ChatGPT's Canvas uses Pyodide for sandboxing, but it's not designed for coding agents. Node.js environment is usually better for agents. Pyodide restricts server-side functionality, and fetching external URLs often needs proxying due to sandbox. By the way, pyodide is still good option for interactive visualizer or deploying small webapps require data processing.
b89kim··on How to run Qwen 3.5 locally
I’ve been testing these on other tasks—IK, Kalman filters, and UI/DB boilerplate. Qwen3.5 is multimodal and specialized for js/webdev or agentic coding. It’s not surprising MoE model have some limitations in specific area. I understand most LLM have limited ability in mathematical/physical reasoning. And I don't think these tasks represent general performance. I'm just sharing personal experiences for those curious.
b89kim··on How to run Qwen 3.5 locally
I’ve been benchmarking GGUF quants for Python tasks under some hardware configs.

  - 4090 : 27b-q4_k_m
  - A100: 27b-q6_k
  - 3*A100: 122b-a10b-q6_k_L
Using the Qwen team's "thinking" presets, I found that non-agentic coding performance doesn't feel significant leap over unquantized GPT-OSS-120B. It shows some hallucination and repetition for mujoco codes with default presence penalty. 27b-q4_k_m with 4090 generates 30~35 tok/s in good quality.