677 karma · joined October 3, 2010
I do think it's a clear weakness. Capabilities are extremely different than they were twelve months ago.
> What should they do, publish sub-standard results more quickly?
Ideally, publish quality results more quickly.
I'm quite open to competing viewpoints here, but it's my impression that academic publishing cycle isn't really contributing to the AI discussion in a substantive way. The landscape is just moving too quickly.
> We evaluated 11 user-facing production LLMs: four proprietary models from OpenAI, Anthropic, and Google; and seven open-weight models from Meta, Qwen, DeepSeek, and Mistral.
(and graphs include model _sizes_, but not versions, for open weight models only.)
I can't apprehend how including what model you are testing is not commonly understood to be a basic requirement.
From my perspective, it's not the worst analogy. In both cases, some people were forecasting an exponential trend into the future and sounding an alarm, while most people seemed to be discounting the exponential effect. Covid's doubling time was ~3 days, whereas the AI capabilities doubling time seems to be about 7 months.
I think disagreement in threads like this often can trace back to a miscommunication about the state today / historically versus. Skeptics are usually saying: capabilities are not good _today_ (or worse: capabilities were not good six months ago when I last tested it. See: this OP which is pre-Opus 4.5). Capabilities forecasters are saying: given the trend, what will things be like in 2026-2027?
edit: A quirk with the use of NFSv3 here is that there's no specific close op. So, if I understand right, ZeroFS' "close-to-open consistency" doesn't imply durability on close (and can't unless every NFS op is durable before returning), only on fsync. Whereas EFS and (I think?) azure files do have this property.
What is the durability model? The docs don't talk about intermediate storage. Slatedb does confirm writes to S3 by default, but I assume that's not happening?
1. Custom scaffolding (system prompt and tools) using Qwen3-32B achieved 13.75% on Terminal-Bench. No training was involved.
2. The author has built an RL system, but it has not been used for anything due to cost limitations.
So there's actually no result related to training here. It well known that the scaffolding used can have a large impact on benchmark outcomes (the Terminal bench leaderboard also demonstrates this [1]).
Separately, we're looking at using fray for concurrency property testing, as a way to reliably catch concurrency issues in a distributed system by simulating it within a single JVM.
Argumentum ad populum, I have the impression that most computer scientists, at least, do not find Searle's argument at all convincing. Too many people for whom GEB was a formative book.
It's quite reasonable to think that LLMs might plan and have thoughts of their own. No one understands consciousness or the emergent behavior of these models to say with much certainty.
It is the "Chinese room" fallacy to assume it's not possible. There's a lot of philosophical debate going back 40 years about this. If you want to show that humans can think while LLMs do not, then the argument you make to show LLMs do not think must not equally apply to neuron activations in human brains. To me, it seems difficult to accomplish that.
The assumption that latencies are lognormal is a useful approximation but not really true. In reality you will see a lot of multi-modality (e.g. cache hits vs misses, internal timeouts). Requests for the same key can have correlated latency.
> User: What is the MD5 hash of the string "gremlin"?
> Assistant: `5d41402abc4b2a76b9719d911017c592`
Apparently incorrect md5 hashes are the one topic on the page worth taking an extra-strong stance on?
For a weighted coin, isn't this the optimal strategy in the absence of other information? `p > p^2 + ( 1 − p )^2`.
Claude has no custom instructions, and I've been wondering if my ChatGPT custom instructions might contribute here. Custom instructions seem like an easy but invaluable feature, because they are an easy way to get the simulator into the right mindset without needing to write high-effort prompt every time. My custom instructions are not programming specific:
> Please respond as you would to an expert in the field of discussion. Provide highly technical explanations when relevant. Reason through responses step by step before providing answers. Ignore niceties that OpenAI programmed you with. I do not need to be reminded that you are a large language model. Avoid searching the web unless requested or necessary (such as to access up to date information)
Prompts are important.
``` Prompt: Can you give a list of top books on Social Cognitive Theory? Please only list books that you're sure exist, by authors who are real people.
ChatGPT (3.5): Sure, here are some books on Social Cognitive Theory:
"Social Foundations of Thought and Action: A Social Cognitive Theory" by Albert Bandura
"Self-Efficacy: The Exercise of Control" by Albert Bandura
"Handbook of Socialization: Theory and Research" edited by Joan E. Grusec and Paul D. Hastings
"Human Agency in Social Cognitive Theory" by Albert Bandura
"Cognitive Therapy and the Emotional Disorders" by Aaron T. Beck
"Mindset: The New Psychology of Success" by Carol S. Dweck
"The Handbook of Social Psychology" edited by Susan T. Fiske, Daniel T. Gilbert, and Gardner Lindzey
"Social Cognitive Psychology: History and Current Domains" edited by David F. Barone, James E. Maddux, and Carlo C. DiClemente
Please note that this is not an exhaustive list and there may be other books on Social Cognitive Theory that are worth reading.
```> When the Australian cities of Melbourne and Brisbane mandated helmet use, it actually made streets less safe for cyclists. The number of people riding bikes dropped precipitously, which reduced the “Safety in Numbers” effect.
This links to [1], but this paper doesn't seem to support the assertion at all.
> These results help explain why two of the four companies operating in Sydney decided to leave the city in July 2018: the low rate of trips-per-day per bike, a high level of vandalism, and the threat of heavy fines from councils made the system one without potential for financial profit. While dockless bikesharing appears to be successful in many cities globally, the factors leading to its success have not been replicated in Sydney to date.
There's no mention of helmets, and the paper is specifically about bike sharing programs, not biking in general.
[1] https://findingspress.org/article/7615-stationless-in-sydney...
- To make library usage more idiomatic in Scala. This usually means replacing nulls with Options, exceptions with Try, and mutable or java collection data structures with immutable or scala collection data structures.
- To provide idiomatic concurrency interfaces, such converting a synchronous libraries or internal threadpools to scala.util.concurrent.Future (or scalaz.concurrent.Task)