388 karma · joined January 8, 2023
> GPT-5.6-Sol is retiring. This conversation will automatically switch to GPT-6-Sol
I don't recall OAI retiring a model so early lol. Similar arch?
> Every day CI runs about 20..80 million tests in 600 commits and 300 pull requests
> Last year, ClickHouse spent 360 years of machine time for CI
I am no user of CH so I can't talk about their product. But we are talking about a company with 686 employees as per their LinkedIn, where ClickHouse is clearly the core of their business. Considering all of this, is 50+ commits a day that much?
[1] https://presentations.clickhouse.com/2026-openhouse-sf/great...
I even forgot about this. God damn it, I enjoyed the SE so much because of this.
Edit: Is there that many demand for a foldable iPhone that justifies Apple making such engineering investment?
Stadlmann improved it from 246 to 240, OpenAI later claimed 186 I think?
Maybe someone can help clarify? I am no expert at all, but I can't help but see similarities.
I feel like P/D dissaggregation will be the next big one for providers, as prefill tends to be compute bound while decode mem bound which I guess each will have a different type of node
> We make money via compute credits + Enterprise Hub + HF Pro subs [...]
I guess this plus custom inference deployments, external inference providers, partnerships with the big cloud AWS, Azure, etc.
It will include Pre-training and Mid-training. ETA 3 months.
Live wandb dashboard: https://wandb.ai/marin-community/marin_moe/reports/535B-A23B...
GitHub issue tracking progress: https://github.com/marin-community/marin/issues/8435
23T tokens (not deduped) directly in S3 for anyone to see: s3://marin-us-east-02a/marin/datakit/store_4d2e363d
Intermediate checkpoints also in S3
This! It's literally in their best interest for open-weights models to succeed
as founding members is crazy !
Their base models and architecture has quickly become the go-to for local inference and fine-tuning, even when they introduced some tricky things like GDN, so many people use it, that it was matter of days/weeks until lots of OSS frameworks adopted it.
If those numbers translate well to its general capabilities, with the great caching DeepSeek has, I feel like this model will get tons of usage.
Wonder if this has something to do due with space constraints. If the study was done in a controlled nest, it must be space bounded one way or another. Dynamics might change when in real-world?
I'd like to read the paper to skim over the methodology but it's not open-access :(
[1] https://www.uni-wuerzburg.de/fileadmin/uniwue/2026/0702Ameis...
Since C++ has no HTTP client in its std lib, I really had no other choice but to use curl. Same with OpenSSL. It'd be quite naïve of me to re-implement the whole HTTP stack and SHA256 from scratch =)