> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .
12 karma · joined March 8, 2026
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .
- He insisted that OpenAI initiated research based solely on rumors and never accessed their Codex sessions.
- He mistakenly believed Tristan and Levent were solving the same problem in Anthropic.
- proposed two option. (1) Tristan becoming the lead author to revise OpenAI’s work, or (2) OpenAI providing internal model to support and bridge their research.
- Sebastien insisted there was no intention to alter authorship. He was simply uncomfortable sharing OpenAI’s work,model with an Anthropic researcher. Additionally, He believed Levent’s credit seemed limited as their work focused on Euler.
- complained that negotiations with Tristan and Levent were difficult
[1] https://xcancel.com/SebastienBubeck/status/20973794116915163...
Separately, Terence Tao noted there is a low probability OpenAI actually solved the general regularity problem.
- OpenAI did related research around similar timeframe.
- Tristan claimed OpenAI offered a proposal that included dropping the Anthropic-affiliated co-author.
- Sebastian (a prominent OpenAI researcher involved) denied these claims.
- Tristan have no concrete evidence that OpenAI accessed their session.
- OpenAI's theory may hold up, but it will require long-term validation to confirm.
There is already plenty of research around multimodal diffusion policies. While DP typically doesn't require pre-training, you can boost data size by depth estimation model+Open data.
PI smartly combined discretized tokens with flow-matching for efficient training, and it works well in most cases. Still, end-effector representation may be better for teleop with devices like a SpaceMouse, VR, or VibeTracker. PI-07 also supports EEF, but I am not sure how much data is needed to fine-tune PI-05 for that.
I'd suggest starting with the default pi05 model. Data strategy is probably more important than model improvements. Since VLA performance is highly dependent on the data/action distribution and it's easy to modify. After that, you can add high-level reasoning like PI05. I visited a Chinese VLA company that already adopted the PI-05 approach, and it works quite well in practice.
- Calibration is not required for VLA models.
- RGB or Stereo RGB inputs are sufficient for ACT, DP, and PI0/PI05.
- ROS2 is not strictly required, but it can be useful for sharing/co-developing codes. For instance, the Stanford team built a custom framework for diffusion policy instead. I also developed similar framework because ROS2 is not optimized for bi-manual manipulation or VLA workloads.
There are plenty of Pi clone boards at lower prices, but they have smaller communities and less documentation. When you hit an unexpected problem, it can be hard to find solutions or get support.
- 4090 : 27b-q4_k_m
- A100: 27b-q6_k
- 3*A100: 122b-a10b-q6_k_L
Using the Qwen team's "thinking" presets, I found that non-agentic coding performance doesn't feel significant leap over unquantized GPT-OSS-120B. It shows some hallucination and repetition for mujoco codes with default presence penalty. 27b-q4_k_m with 4090 generates 30~35 tok/s in good quality.