HNHacker News
TopNewBestAskShowJobs

lllllm

241 karma · joined February 18, 2023

submissionscomments
lllllm··on Ask HN: Who is hiring? (October 2025)
Swiss AI Initiative | https://www.swiss-ai.org/ | Hybrid/ONSITE (in Europe)

We are a young team, and the creators of the Apertus LLM, the currently leading open-data open-weights AI model.

Join us to work on cutting edge LLM training in the open. We do pretraining, alignment, reasoning, multilinguality and multimodality - all at the intersection of engineering and research.

This is a joint team between ETH Zurich and EPFL in Lausanne, running on the Alps supercomputer (one of the largest public institution GPU cluster). Visa sponsoring possible, work language is English.

https://careers.epfl.ch/job/Lausanne-AI-Research-Engineers-S...

lllllm··on Apertus 70B: Truly Open - Swiss LLM by ETH, EPFL and CSCS
yes this seems a good way to go. for example you can already find many quantized versions under https://huggingface.co/models?search=apertus%20mlx and elsewhere
lllllm··on Apertus 70B: Truly Open - Swiss LLM by ETH, EPFL and CSCS
thank you!
lllllm··on Apertus 70B: Truly Open - Swiss LLM by ETH, EPFL and CSCS
We hear you, nevertheless this is one of the very few open-weights and open-data LLMs, and the license is still very permissive (compare for example to Llama). Personally of course I'd like to remove the additional click, but the universities also have a say in this.
lllllm··on Apertus 70B: Truly Open - Swiss LLM by ETH, EPFL and CSCS
The pretraining (so 99% of training) is fully global, in over 1000 languages without special weighting. The posttraining (See section 4 of the paper) had also as many languages as we could get, and did upweight some languages. The posttraining can easily be customized to any other target languages
lllllm··on Apertus 70B: Truly Open - Swiss LLM by ETH, EPFL and CSCS
common crawl anyway respects the CCbot opt-out every time they do a crawl.

we went a step further because back in old ages (2013 is our oldest training data) LLMs did not exist, so website owners opting out today of AI crawlers might like the option to also remove their past contents.

arguments can be made either way but we tried to remain on the cautious side at this point.

we also wrote a paper on how this additional removal affects downstream performance of the LLM https://arxiv.org/abs/2504.06219 (it does so surprisingly little)

lllllm··on Apertus 70B: Truly Open - Swiss LLM by ETH, EPFL and CSCS
martin here from the apertus team, happy to answer any questions if i can.

the full collection of models is here: https://huggingface.co/collections/swiss-ai/apertus-llm-68b6...

PS: you can run this locally on your mac with this one-liner:

pip install mlx-lm

mlx_lm.generate --model mlx-community/Apertus-8B-Instruct-2509-8bit --prompt "who are you?"

lllllm··on Apertus 70B: Truly Open - Swiss LLM by ETH, EPFL and CSCS
we compared to GPT-OSS-20B, Llama 4, Qwen 3, among many others. Which models do you think are missing, among open weights and fully-open models?

Note that we have a specific focus on multilinguality (over 1000 languages supported), not only on english

lllllm··on Apertus 70B: Truly Open - Swiss LLM by ETH, EPFL and CSCS
we didn't have time to write one yet, but there is the tech report which has a lot of details already
lllllm··on Apertus 70B: Truly Open - Swiss LLM by ETH, EPFL and CSCS
posttraining codebase is here: https://github.com/swiss-ai/posttraining
lllllm··on Apertus 70B: Truly Open - Swiss LLM by ETH, EPFL and CSCS
we released 81 intermediate checkpoints of the whole pretraining phase, and the code and data to reproduce. so full audit is surely possible - still it would depend on what you consider 'practical' here.
lllllm··on Apertus 70B: Truly Open - Swiss LLM by ETH, EPFL and CSCS
benchmarks: we provide plenty in the over 100 page tech report here https://github.com/swiss-ai/apertus-tech-report/blob/main/Ap...

quantizations: available now in MLX https://github.com/ml-explore/mlx-lm (gguf coming soon, not trivial due to new architecture)

model sizes: still many good dense models today lie in the range between our small and large chosen sizes

lllllm··on ETH Zurich and EPFL to release a LLM developed on public infrastructure
this is what this paper tries to answer: https://arxiv.org/abs/2504.06219 the quality gap is surprisingly small between compliant and not
lllllm··on ETH Zurich and EPFL to release a LLM developed on public infrastructure
absolutely! i've sent you a linkedin message last week. but here seems to work much better, thanks a lot!
lllllm··on ETH Zurich and EPFL to release a LLM developed on public infrastructure
we kept all 1800+ (script/language) pairs, not only the quality filtered ones. the question if a mix of quality filtered and not languages impacts the mixing is still an open question. preliminary research (Section 4.2.7 of https://arxiv.org/abs/2502.10361 ) indicates that quality filtering can mitigate the curse of multilinguality to some degree, so facilitate cross-lingual generalization, but it has to be seen how strong this effect is on larger scale
lllllm··on ETH Zurich and EPFL to release a LLM developed on public infrastructure
no. the main source is fineweb2, but with additional filtering for compliance, toxicity removal, and quality filters such as fineweb2-hq
lllllm··on ETH Zurich and EPFL to release a LLM developed on public infrastructure
Yes this is an interesting question. In our arxiv paper [1] we did study this for news articles, and also removed duplicates of articles (decontamination). We did not observe an impact on the downstream accuracy of the LLM, in the case of news data.

[1] https://arxiv.org/abs/2504.06219

lllllm··on ETH Zurich and EPFL to release a LLM developed on public infrastructure
No, the model has nothing do to with Llama. We are using our own architecture, and training from scratch. Llama also does not have open training data, and is non-compliant, in contrast to this model.

Source: I'm part of the training team

lllllm··on Planet squeezed in between two stars
animation of it: https://youtu.be/ewg36czOOiI?si=moL9g9Xz2-vVClZX
lllllm··on Text Is All You Need
it takes quadratically more time the larger your context is.
lllllm··on Text Is All You Need
The current systems like chatGPT actually have just such two parts. One is the raw LLM as you describe. The second one is another network acting as a filter on top of the first one. To be more precise, that second part is the process of finetuning with Reinforcement Learning from Human Feedback (RLHF). It trains a reward model to say if the first one was good or bad. Currently it's done very similarly to standard supervised learning (with human labelling) to say if the first model behaved good or bad, aligned or not with 'our' values.

Anyway, while I remain sceptical about the roles of these in-flesh hemispheres, the artificial chatGPT-like systems indeed do have such left and right parts