1,336 karma · joined December 11, 2020
I think so for a few reasons:
1. The Reuters article does explicitly say the model is compatible with domestic chips for inference, without mentioning training. I agree that the Reuters passage is a bit confusing, but I think they mean it was developed to be compatible with Ascends (and other chips) for inference, after it had been trained.
2. The z.ai blog post says it's compatible with Ascends for inference, without mentioning training, consistent with the Reuters report https://z.ai/blog/glm-5
3. When z.ai trained a small image model on Ascends, they made a big fuss about it. If they had trained GLM-5 with Ascends, they likely would've shouted it from the rooftops.
4. Ascends just aren't that good
Also, you can definitely train a model on one chip and then support inference on other chips; the official z.ai blog post says GLM-5 supports "deploying GLM-5 on non-NVIDIA chips, including Huawei Ascend, Moore Threads, Cambricon, Kunlun Chip, MetaX, Enflame, and Hygon" -- many different domestic chips. Note "deploying".
I've only seen information suggesting that you can run inference with Ascends, which is obviously a very different thing. The source you link also just says: "The latest model was developed using domestically manufactured chips for inference, including Huawei's flagship Ascend chip and products from leading industry players such as Moore Threads, Cambricon and Kunlunxin, according to the statement."
from collections import defaultdict
def find_cheryls_birthday(possible_dates):
# Parse the dates into month and day
dates = [date.split() for date in possible_dates]
months = [month for month, day in dates]
days = [day for month, day in dates]
# Step 1: Albert knows the month and says he doesn't know the birthday
# and that Bernard doesn't know either. This implies the month has no unique days.
month_counts = defaultdict(int)
day_counts = defaultdict(int)
for month, day in dates:
month_counts[month] += 1
day_counts[day] += 1
# Months with all days appearing more than once
possible_months = [month for month in month_counts if all(day_counts[day] > 1 for m, day in dates if m == month)]
filtered_dates = [date for date in dates if date[0] in possible_months]
# Step 2: Bernard knows the day and now knows the birthday
# This means the day is unique in the filtered dates
filtered_days = defaultdict(int)
for month, day in filtered_dates:
filtered_days[day] += 1
possible_days = [day for day in filtered_days if filtered_days[day] == 1]
filtered_dates = [date for date in filtered_dates if date[1] in possible_days]
# Step 3: Albert now knows the birthday, so the month must be unique in remaining dates
possible_months = defaultdict(int)
for month, day in filtered_dates:
possible_months[month] += 1
final_dates = [date for date in filtered_dates if possible_months[date[0]] == 1]
# Convert back to original format
return ' '.join(final_dates[0]) if final_dates else "No unique solution found."
# Example usage:
possible_dates = [
"May 15", "May 16", "May 19",
"June 17", "June 18",
"July 14", "July 16",
"August 14", "August 15", "August 17"
]
birthday = find_cheryls_birthday(possible_dates)
print(f"Cheryl's Birthday is on {birthday}.")ETA: I agree that you did not say this was happening in your original comment, but it seems to me your comment implied that these companies were actually influencing major decisions (since that's the topic of the OP).
an AEI interview https://www.aei.org/workforce-development/the-future-of-remo... which seems pretty balanced overall (and doesn't take a prescriptive position)
an AEI piece https://www.aei.org/research-products/report/the-trade-offs-... which seems pretty balanced too
a Heritage piece (from early in Covid) https://www.heritage.org/jobs-and-labor/report/labor-policy-... that seems mostly bullish on remote work (but mostly focuses on other issues, like labor rights)
a McKinsey report (also from fairly early in Covid) https://www.mckinsey.com/featured-insights/future-of-work/wh... which is mostly descriptive and also seems pretty balanced
a Cato piece https://www.cato.org/commentary/remote-work-here-stay-mostly... which argues in favor of remote work
OP was written by the person who co-founded GiveWell[1] to make charitable giving more effective, and who while running Open Philanthropy oversaw lots of grants to things like innovation policy[2], scientific research[3], and land use reform[4].
Anyway, more broadly I think you present a false dilemma. You can both prepare for tail risks and also make important marginal and efficiency improvements.
[1] https://www.givewell.org/ [2] https://www.openphilanthropy.org/focus/innovation-policy/ [3] https://www.openphilanthropy.org/focus/scientific-research/ [4] https://www.openphilanthropy.org/focus/land-use-reform/
[1] https://www.lesswrong.com/posts/GpSzShaaf8po4rcmA/qapr-5-gro...
I've reproduced dwaltrib's results using World Bank data on 251 countries, and I get a Pearson's r of 0.82 and a p value of 5.6e-61 (!). I.e. a strong correlation, with high confidence. It makes sense too -- larger countries generally have more people, and more people generally generate more economic activity.
Code if you want to try yourself:
import pandas as pd
gdp = pd.read_csv("~/Downloads/API_NY.GDP.MKTP.CD_DS2_en_csv_v2_5551501.csv").set_index("Country Name")
land_area = pd.read_csv("~/Downloads/API_AG.LND.TOTL.K2_DS2_en_csv_v2_5552158.csv").set_index("Country Name")
gdp["GDP"] = gdp["2020"]
gdp["Land"] = land_area["2020"]
gdp = gdp.dropna(subset=["GDP", "Land"])
from scipy import stats
print(stats.pearsonr(gdp.Land, gdp.GDP))
#+RESULTS: : PearsonRResult(statistic=0.8151313879150333, pvalue=5.621180589722219e-61)
There absolutely is a correlation between land mass and nominal GDP.
The 1e35 FLOP number is meant as a conservative upper bound and comes from here: https://www.lesswrong.com/s/5Eg2urmQjA4ZNcezy/p/rzqACeBGycZt...
The major fabs all have roadmaps for approaching 1 nm, and there are other advances that could allow you to keep going either if transistor size scaling stops (e.g., vertical scaling). (That said, I definitely don't think it's a given that HW price-performance keeps doubling at the same rate 10+ years.)
It’s completely true that LLMs are trained on next-token prediction (although some, like ChatGPT, are then additionally trained using reinforcement learning with human feedback). It’s also completely true that this fact profoundly influences the texts they generate. So I don’t think it’s unreasonable to call LLMs autocomplete engines or to emphasise next-token prediction. But I think it’s subtly misleading:
- Though LLMs were trained to optimise success on next-token prediction, that is not necessarily what they do. We don’t know what it is they do. The training process reinforces behaviours/heuristics in the model that tend to cause it to make better next-token predictions on in-distribution data. This does not mean that those behaviours/heuristics are fundamentally “about” optimising next-token prediction, especially when the model encounters out-of-distribution data.
- The usual example here is human evolution. Humans were shaped by a process that optimised for reproductive fitness. This gave us a bundle of drives such as family kinship, prestige and sexual pleasure – drives that aren’t fundamentally about optimising for reproduction, which becomes evident as we enter a new environment – one with contraceptives, say.
- Optimising for a task for which intelligence is useful encourages the optimised thing to become more intelligent. Sam Altman gave expression to this last week when he wrote, “Language models just being programmed to try to predict the next word is true, but it’s not the dunk some people think it is. Animals, including us, are just programmed to try to survive and reproduce, and yet amazingly complex and beautiful stuff comes from it.”
- The forms of intelligence that are useful in doing next-token prediction are different from those that are useful in human reproduction, but I think there’s a considerable overlap, as (1) some fundamental abilities, for example using and applying concepts, just seem very broadly useful and (2) the data LLMs are trained on are written by humans, for humans and often about humans and things that matter to us.
I think the "they're just autocomplete" take also hides other properties of LLMs, like them seeming to (as mentioned in another comment, and in the post) contain and use world models, and being able to learn general algorithms.
Beyond that, many intellectual capabilities are also useful for successful bullshitting, so even if we can confidently say that that's all they do (in some sense), that doesn't mean they don't also do something-like-human-reasoning etc.
- https://ourworldindata.org/grapher/excess-deaths-cumulative-...
This is pretty unfair IMO. Aaronson wrote (emphasis mine):
> In this piece Chomsky, THE INTELLECTUAL GODFATHER GOD OF an effort that failed for 60 years to build machines that can converse in ordinary language, condemns the effort that succeeded.
He never wrote that Chomsky himself was involved in this efforts, only that he was influential to it (which may or may not be true, but Marcus never argues that point).
- LLMs are sometimes said to be “just” shallow pattern matchers, “just” massive look-up tables or “just” autocomplete engines. These comparisons amount to a form of (methodological) reductionism. While there’s some truth to them, I think they smuggle in corollaries that are either false or at least not obviously true.
- For example, they seem to imply that what LLMs are doing amounts merely to rote memorisation and/or clever parlour tricks, and that they cannot generalise to out-of-distribution data. In fact, there’s empirical evidence that suggests that LLMs can learn general algorithms and can contain and use representations of the world similar to those we use.
- They also seem to suggest that LLMs merely optimise for success on next-token prediction. It’s true that LLMs are (mostly) trained on next-token prediction, and it’s true that this profoundly shapes their output, but we don’t know whether this is how they actually function. We also don’t know what sorts of advanced capabilities can or cannot arise when you train on next-token prediction.
- So there’s reason to be cautious when thinking about LLMs. In particular, I think, caution should be exercised (1) when making predictions about what LLMs will or will not in future be capable of and (2) when assuming that such-and-such a thing must or cannot possibly happen inside an LLM.
https://www.erichgrunewald.com/posts/against-llm-reductionis...
about "the future is distant and we can't know anything about it", see e.g. this[2] post on the effective altruism forum.
> Possible misconception: “Trying to influence the far future is pointless because it is impossible to forecast that far.”
> My response: “Considering far future effects doesn’t necessarily require predicting what will happen in the far future.”
> [...]
> Nuclear war could feasibly happen tomorrow. Climate change is an ongoing phenomenon and catastrophic climate change could happen within decades. In The Precipice, Toby Ord places the probability of an existential catastrophe occuring within the next 100 years at 1 in 6, which is concerningly high.
> Ord is not in the business of forecasting events beyond a 100-year time horizon, nor does he have to be. These existential threats affect the far future on account of the persistence of their effects if they occur, but not on account of the fact that they might happen in the far future. Therefore whilst it is true that a claim has to be made about the far future, namely that we are unlikely to ever properly recover from existential catastrophes, this claim seems less strong than a claim that some particular event will happen in the far future.
i recommend op engage with some longer longtermist writings and try to understand the view from the inside. another poster mentioned toby ord's the precipice, but there is also plenty of material online. for example, i recommend this interview with carl shulman about extinction risk.[3]
[1] https://slatestarcodex.com/2017/04/07/yes-we-have-noticed-th...
[2] https://forum.effectivealtruism.org/posts/ocmEFL2uDSMzvwL8P/...
[3] https://80000hours.org/podcast/episodes/carl-shulman-common-...
- https://forum.effectivealtruism.org/posts/fd96FtLFACeAshqJP/...
- https://forum.effectivealtruism.org/posts/ErKzbKWnQMwvzRX4m/...
or my interview with lucia coulter, one of the co-founders:
- https://www.erichgrunewald.com/posts/interview-with-lucia-co...