Kolibri: A Sovereign Open-Weight Model
aleph-alpha.com
additional paper: https://tej.as/blog/aleph-alpha-kolibri
aleph-alpha.com
additional paper: https://tej.as/blog/aleph-alpha-kolibri
One of humanity’s biggest problems here is we don’t know how to do moderation.
We have two modes. One is a brick taped to the accelerator and damn all consequences, driven by national pride or corporate greed or egos. The other is a brick taped to the brake driven by histrionic doomers and anti-everything pessimists.
The extremes are loud and fit in a tweet. Nuance is quiet and contemplative and usually requires an essay or a book. It’s also dynamic. Nuanced positions evolve over time as new things are learned. Extremes tend to be fixed and rigid. All this, I think, gives them higher memetic fitness in the discourse.
I don’t think this is new. Look at nuclear power, a largely pre-Internet example. You had pro nukes who minimized and hand waved away any risk and anti nukes that wanted it utterly outlawed. Nobody said “hey this is a great zero carbon source of energy but we really need to think it through carefully and manage it well.” Or if they did they were drowned out by the loud screaming extremes.
That's what they call this and I think it's a pretty clear definition these days to people in the industry.
It’s seemed crazy to me that anyone thought these could stay closed or even SHOULD be closed source.
Open source LLMs "democratize" access to the "intelligence booster" that is AI. But while that has several benefits, it also has several downsides.
Humanity has the serious problem of being underdeveloped in the "spiritual" department. Ethics is often considered some sort of lifestyle choice, but it's actually the difference between order and chaos in a society.
Everybody being able to do anything means somebody will be able to do something you don't like. At an arbitrary scale.
Wonder what safeguards Kolibri uses to prevent this? Or if they even can
Miscreants with even half a braincell will find themselves in illegal drugs business over terrorism.
Political and religious fanatics with even a half a braincell figure out terrorism will only damage their cause.
People who end up as terrorists have already failed the two above tests.
Given enough smarts, you can simulate anything. Quantum AI is an active goal.
I think we should stay in reality for now. They write good rust and bad emails.
But the applicability and the ROI of LLMs for other use cases than coding is a lot squishier. Also correspondingly the risks are unspecific.
As for what to do, ethical disclosure of vulnerabilities provided a good framework for disclosing software vulnerabilities discovered with the assistance of LLMs. What is going to be novel and calls for our spiritual development in other domains?
What do average people currently want? They're taught, the most important thing was being rich. So they will ask their AI to make them rich. Most real life ways to get there are "sketchy" to say the least, usually downright unethical and anti-social, but US society turns a blind eye when the "Wolf of Wall Street" comes out on top and the schemes don't easily fit into average people's abilities of moral judgement.
-> Large parts of US society suddenly engaging in all kinds of "semi-legal/hyper-illegal" fraud schemes, at the expense of already saturated environmental and societal resilience. Guaranteed collapse.
Or, let's get rid of those pesky neighbors/wrong-colored people/annoying opinions? Again, "legal" is a pretty squishy concept and only really applies when you don't have the legal expertise to get around it. Now you can.
Or, look at the basics: what is "real"? You only "know" because you trust certain people and institutions. Generative AI can help with that /s.
It's not only about "building weapons of mass destruction". It's about doing the same shit as usual, but a thousand times faster/amplified. Look up poly-/metacrisis for starters. Going faster with AI when there's a wall in front of you isn't the best idea.
In other words, ambitious frauds have already explored all of the angles and bought all the ads. At worst, LLMs will create a few more successful but less ingenious frauds.
"Ambitious frauds" haven't "explored all the angles".
You imply "LLMs" to be and stay less intelligent than humans, in particular yourself. You're mistaken.
..or maybe the commenter did look at superhuman persuasion long enough to believe it would be best to channel those ever the same fear fantasies from the LLM through their account to the reader.
On a more serious note, just look at the doom premises here: "Large parts of US society suddenly going criminal" is from the movie "The Purge", I think. It is fiction.
The idea that generative AI takes away our ability to find out reality. ... I don't know. People write about that a lot, but it still seems very far fetched.
Maybe through some terminally online overconsumption, like with social media? I wouldn't know.
With new AI capabilities we will have to adjust, I am sure. Media, science, education and law are changing very visibly right now. Those p(doom) narrations just seem to be pre-IPO hype though.
It is just so so dangerous. That is why they want to go public and only want to care about optimizing for the next quarter ...right before breaking into AGI. /s
It's very clear to that any specifics couldn't capture the risks, because the capabilities, including the risks, are one level higher than any specific techonolgy. It is the process of advancing technology itself, in accelerating speed, that poses the risk.
While being unable to judge whether you should in the first place.
Here you are claiming that AI products are going to reach AGI or RSI in the foreseeable future. Of course you can't "capture the risks" with specifics because those are inherently unspecific futures. It's a bit like saying when we invent antigravity all hell will break loose.
I would believe those future risks more if there were a progression of risks. What other than finding vulns has those characteristics?
edit: why is the parent rationale sensible for AI and not nuclear technology?
Nuclear weapons are a wholly unique threat to mankind.
While "some chatbot" isn't the problem, general intelligence superior to humans absolutely is.
AI allows anybody to enact essentially anything. And your "level a major city" is just a small task really. The problem there is your lack of imagination, not the actual impossibility of that task.
Thats the annoying part about these conversations. People try to smuggle in the conclusion of "And now I wave a magic wand that does literally anything, by magic" when discussing stuff that everyone can use and see right now that clearly isn't that.
No it doesn't. Where is your evidence for this?
> While "some chatbot" isn't the problem, general intelligence superior to humans absolutely is.
This doesn't exist. Why are you pretending it does?
>AI is different in that it enables technological development in ways quite unlike the wheel in a general sense.
The wheel already enabled almost every tech. Without the wheel nothing you see around yourself would be possible.
There's a massive barrier between knowing how to build a nuke and actually building it.
Also, isn't proliferation the basis for MAD? Well if we believe in that, naturally it means that the world would be safest if every individual had their own nuke ;)
But the alternative just seems… so much worse to me?
A select few groups gating access to the ability to do everything seems like neo-fuedalism in the making.
And to be fair even the gating that we do have (daybreak, CVP, etc.) is already being circumvented via keys being stolen and sold on the dark web.
Clearly not a "better" scenario. The real problem though seems people feigning helplessness? You can't leave society "to its own". You are part of it and go where it goes. So better start steering.
When access to AI gives you abilities you cannot use responsibly, you shouldn't have access to that. Just like you shouldn't be allowed to drive a car or fly a plane or command a rocket without proper guardrails, safeguards, prerequisites, etc.
"General" intelligence isn't present in humans, why does it need to be in AI?
I don't think this is realistically a problem at all. It just makes certain types of research cheaper and less time-consuming. And, again, this is also a problem with the american services.
What do you mean "cannot"? As in you are granted abilities that have no responsible use?
You live in a curated world and rarely or never encounter such things. Precisely because your environment is curated that way.
Look at how you can't buy WMDs. They have no responsible use for you.
As is shown to us by filthy rich people every day.
Or did you mean the burglar in the fawellas?
Not arguing that we shouldn't strive to do far far better here, better is by our own imagining. There is no development scale. There are no aliens or prior non human civilizations to compare against. So we're not underdeveloped. We are as we are. For all we know we're at peak capacity and humanity will never be better.
We have a lot of details in the tech report if you want to go deeper.
I’m one of the authors of model2vec, and working on training classifiers for this. I think model2vec could be better, but I haven’t had the opportunity to try this at scale. So if you did, knowing about it would be helpful!
Why expose yourself to this liability?
(This comment was originally posted to https://news.ycombinator.com/item?id=49943034, but we're merging the threads.)
Bravo team!
I felt like a learnt a lot of details about the entire process. Questions arose during my reading, and searching for the answers led to more learning.
No marketing BS, lots of actual information. And their level of openness is really neat!
We as many other’s were curious to try and benchmark it.
On that note, as a small gesture of support, we’ve hosted and made Kolibri-1 free for anyone to try for the next few days.
No GPU. No setup. Just try it. tesseracted.com/kolibri-1-chat/
No, I'm not gonna do that, I asked you to do that Mr Kolibri.
Also feedback on that interface: It's very annoying while answering. It almost immediately shows a list of sources, which on my screen fill up all the space and then when it starts answering it keeps those in view but also scrolls down the tiny part of actual text its outputting but I can't scroll up to start reading from the top, coz it keeps scrolling. I have to wait until it's completely done generating its output.
Uses Brave search, the idea is to test how well the Model can decided when to leverage search or not.
What I couldn't tell but maybe you can tell us: it also complained that it couldn't read the full text as something was cut off. Is that because of the tool you gave it, of brave itself or is it the model?
https://aleph-alpha.com/en/blog/bounding-hallucinations-merl...
Gemini: Here's a list of links to check
https://minecraft.wiki/w/Minecraft_Support_Virtual_Agent#Sys...
(for context, the Minecraft Support chatbot, aka Merl, had a meme because she kept saying "I don't know" to questions like "How to craft a diamond pickaxe".
What is the canonical interpretation of the song ‘Glass Flowers’ by Armand Uso?
(A made-up song and name)
And received: "Glass Flowers" by Armand Uso, a track from the 1978 album The Art of Falling in Love, is generally understood as a melodic reflection on the fragility and impermanence of love. [...]
I saw similar responses for other questions. Qwen3.8-27B correctly refused without web search.That said, I couldn't find any evaluations in the report targeting that specifically. They evaluate on MC for hallucination and look about comparable to Qwen3.5 35B-A3B there.
disclaimer: I‘m part of the training team, happy to answer any questions
- Are there non-LLM approaches to the above task with the goal of achieving a non-hallucinatory agent?
It seems like a very suitable size for local AI models on reasonably high end consumer devices, given it's low active parameter count and a mixed 8bit/4bit quant would fit easily inside 64GB of memory.
And that is a good thing - no need to hide it. Given the growing cost of keeping up, these few non-US, non-Chinese companies really need to do more sharing of efforts and costs.
Canada too is very much in need of sovereign AI options, but funding that on its own would be pretty much a waste of money. Would love to see this new German Canadian company cooperate with Mistral too, or maybe one of the Korean AI companies.
The costs of "keeping up" aren't really growing, on the contrary.
Also, once the Cohere takeover is complete will they still be able to use this "sovereign" claim despite being 90% owned and 100% operated out of Toronto?
The main problems with big corp AI are due to control of access in the first place and control of what they output.
When you make your industry reliant on such choke points, you render yourself the opposite of "sovereign" for sure. Having multiple independent suppliers at least mediates that.
What's so "not-sovereign" for an open-weights American or an open-weight-open-training-process Chinese LLM?
Do we also need sovereign Linux (maybe), sovereign Postgres (most likely not), sovereign Python (def not)?
Given your other comments I can see why it was banned.
I’m most familiar with Canada, where sovereign is usually just an excuse to overpay someone connected for an inferior product with no strategic value.
Also, it's not like it's fixed in time, they are continually building more powerful models as well. Simply because you're in second place doesn't mean you quit the journey. Though, I'm sure the US has a huge vested interest of convincing people to "just quit and submit."
No. Thanks.
Qwen models have solid performance on benchmarks and we're transparent about this in the report. We're just happy to share a European alternative in the small model space, where I don't think we can afford to be fully dependent on China. We also put a lot of effort into German language quality things that don't show up in evals at all (style, Grammar, German reasoning, etc).
Sovereignty is imho mostly about choice and control over your data. Cohere is no different on that front and personally I'm quite excited about what we will build together post-merger.
Right now one could run an open model for most government applications and it would be good enough, you just cannot trust any of these.
So having a sovereign controlled model audit the first one would basically act like a “trust adapter”.
If the second model is cheap and fast enough, there is a business model.
You don’t even need to audit all the intermediate steps, just tool calls and end results.
Since no individual buyer cares that much about geostrategic issues, it's almost impossible to get some random European company to think about buying anything other than 'whatever US or China' are making.
There has to be very concerted push to make a difference.
Legal mandates around sovereignty (be very careful here) could make a difference.
But there will also be market reaction: US/China companies will bend on some level and provide things like 100% EU hosted and even comply with some data source stuff.
The only real path is to be more competitive as a continent.
You clearly haven’t worked in enterprise in the EU. Location of data processor is the first thing they check. That’s why every major cloud has regions with different offerings, not just for HA and redundancy
Yes - security concerns are now very real, but do not fall for this idea that anyone really cares about structural concerns.
Tons of EU companies are still selling crap to Russia, happy to look the other way wile the bear devours a neighbour, as long as profits are there.
What has happened is that Trump has given a face to the reality, and so companies are adjusting on some level.
But 1) it will be nominal 2) Trump will be gone and the impetus will fade 3) the conglomerates will react 4) modest legislation will mean ...
The can will get kicked down the road.
The invasion of E. Europe by Russia has not even caused defence spending or the size of Armies to change, doctrine is barely changing.
Nothing will make those beheamoths change other than other forms of structural concern.
The 'Riet of the Right' - partly due to collapse of Auto / China (among many other things) might cause some change. But the change won't happen until well after the damage has been done.
Brexit should have been an opportunity for institutional reform of the EU, but no - they blamed it all on populism.
Certain governing entities in Europe will try to move away from Microsoft, they will be pulled back.
There is hope, and certain champions can rise and causes pieces to collapse.
If SUSE had any true entrepreneurial whereiwthal - they would create a true consumer / prosumer / enterprise-user friendly variation and brand their flavour of Unix - and make sure that all Euopean governments us it exclusively, which would cascade into widespread use.
There are variations of insta, youtube, netflix that should all be Europe based that could 'theoretically happen' but it needs some structural impetus.
Things usually don't change. Usually there needs to be a collapse and re-order of the system for that to happne.
The people in charge just want to keep their jobs, their very high salaries, and protect their retirement, and will be happy to 'sell out' whatever other imperatives along the way.
Ending with a positive not - I would say current conditions mean the change is now 'plausible, it not likely' whereas before it was 'not very plausible'. So there's a candle, a bit of wind, but not a lot of tinder or dry wood.
AI models act as a force multiplier for intelligence, in particular for generating information according to someone's wishes.
I.e. deepfakes and social media mis-/disinformation campaigns are a thing and having powerful AI allows you to do those at scales that can overwhelm society's resilience.
In general, even if you have "aligned" AI: aligned with whom or what?
Whom are you comfortable with lording as a some demi-god over you, dictating what to believe?
Those laws would not exist if European champions were leading the world.
"are you comfortable with lording as a some demi-god over you, dictating what to believe?"
Yes, Europe handed over all of the decisions about everything to foreign powers, now they have to enact regulations to try to constrain it.
Zuck et. al. make the investments decisions for Europe, by virtue of you all giving him the money and power to do that.
Stop giving him the money/power, then this regulation won't exist
This is why each country needs to educate and maintain their own researchers across as broad a spectrum of disciplines as possible. Also why there needs to be healthy locally financed (through taxes and grants in addition to consumer spending) ecosystem of reporters and media, so that social media (and LLMs) can be easily discarded as sources of social truth and information like other entertainment products in favour of the trusted ones.
As for technical truths, I don't think using informal social media like the stackexchanges or LLMs to explain how camera lenses or checksums work yields materially worse outcomes than consulting Wikipedia or Knuth. For most things, where it doesn't matter, the blind copy/pasters will prevail. Where it does matter, the organization of the work itself has to set out with establishing accountability that encourages the appropriate amount of diligence anyways. And also the point above about having some local expertise.
There is no such thing. Countries spend gazillions on Academia, the Academics do waht they want.
There are no real government places or jobs for that kind of thing. It comes at great expense and vague outcome. Or it gets entangled in bureaucracy.
What you're hinting at is a form of 'utopian governance' - like - it perfectly makes sense on paper, and it's actually rational. But it's completely unworkable in the context of how governments and organizations actually work.
It could on work in theory, not particularly well in practice.
Or are we worried that open models send secret telemetry?
I get that doesn't invalidate the real "point" of the model, but...
Qwen3.8 27B scored notably higher in most of the provided benchmarks, including the German-specific ones. The only "downside" is that inference is much more costly and slow, since it's a dense model.
Qwen3.8 Flash-Next appears to usually "benchmark higher" than 27B, while remaining fast.
I'm sure I could dig up the equivalent benchmarks for Flash and do the comparison myself, but as far as inference goes, it's messy. Consider that Qwen3.5 35B-A3B scores higher than Qwen3.6 on some of the German-specific benchmarks.
So it seems superficially plausible that Qwen3.8 Flash-Next might not be "27B but faster" in the ways that are important for this model. Or it could just "be superior" in all ways.
Either way, I don't think an LLM has to be "the best" at anything to be worthwhile, necessarily. And I kind of distrust benchmarks on top of that, so...
People follow the latest frontier lab models with great attention and migrate to the next big model on their subscriptions. Meanwhile these local models have quietly gotten REALLY good. It is not even an exaggeration. It has happened in the last couple of months.
"Local model you can run at 40 t/s on a gaming machine that is better than Opus 4.6" is way less exciting than "OpenAI IS DOING CRIME!!! OpenAI SOLVED NAVIER STOKES. DARIO SAYS GLM 5.3 BAD! SLOW DOWN THE FRONTIER!".
(edit: also... totally ignore that 27B dense column over there where Qwen 3.8 27B beats Kolibri on nearly every single benchmark. Why would I choose to run this model?)
What is the upper bound on the value of more intelligence applied to your problem domain?
Reiterating what I said yesterday, I think what most people did miss here is that coding is probably not the main goal of this model. It can also code, but the real target is the business processes that aren't code.
Think for example a support department, where new cases (in german!) should be automatically processed (or automatically draft processed) based on existing cases, without pumping all that data through a cloud api.
Or being a general sparring thing for non-IT workers.
__
If you want to code, you already operate in english, so there's the existing and better coding models for you.
But if you're some regular worker somewhere, you might just want to ask the thing to retrieve data on this specific project done three years ago and go through the shared email folder relating to it.
__
You can see that that might be the target, given that they've explicitly focused on getting the model to say "I don't know" instead of just making stuff up. That is exactly the capability you need for these kinds of use-cases.
You can also see that in how it is optimized for speed and a fairly small KV footprint. The idea likely being that a business can pay like a few hundred euros once per worker and then have this kind of local capability for a team of 30.
Its speed is pretty good on the other hand with only 3B active parameters I am getting around 170 tkn/s on fp8
Yeah. It’s a German model.
Aleph Alpha Kolibri: How the sovereign German LLM works - https://news.ycombinator.com/item?id=49943034 - Oct 2026 (162 comments)
Good to see Europe adding toe what Mistral is doing. +100
We have a separate repo for BF16: Kolibri-1-BF16. Will still be tough to put it on a Mac though :(
Disclaimer: I'm part of the team that trained Kolirbi
> trained it on infrastructure in Germany
Made me chuckle. About the second worst place in Europe to train a model if you really care about environmental constraints and energy requirements.
https://app.electricitymaps.com/
Good initiative on the sovereignty front though.
More specifically, I asked the model about what I should put on my contact page, which can cost you in the order of 500 € in Germany if you don't write the right magic words.
Kolibri incorrectly referenced the "Telemediengesetz" ("telecommunication act"), which has been superseded by the "Digitale-Dienste-Gesetz" (DDG, "digital services act") since 2024. The model knows about the DDG, but does not reference it unless specifically instructed to do so.
If anyone of the developers reads this, you can fix this by introducing a recency bias during training. You can even control it by conditioning the model on a date provided with the system prompt or first prompt, so you can travel in time.
I’ll be curious to see which it becomes in the next year - nba “very narrow record”, or a sign you can hang with the big boys.
This thing is worse than a Qwen3.8 27B.
Sovereign works when talking about building a commodity supply or something, not in literally the world’s most competitive and fast moving field.
Those seeking sovereign capability would be better off aiming to be best at something, even something much narrower than an all round LLM. Or just fast following and making something that matches leading performance, which is close to what the Chinese labs do currently.
I think it's not simple to directly compare a 27B dense model with an MoE model like ours. As we know, dense models need all params active for every token. Whereas, MoE models (especially sparse ones like Kolibri) fewer active parameters and correspondingly less compute per token.
Among the MoE models we compared against in our tech report and model card, though, Kolibri performs very well in our evaluation, including against models with 12B active parameters. It also best model in the group we tested within that range of active params.
So, I think it's fairer to see this as a trade-off. Kolibri needs less compute per token but more memory, while Qwen3.8 27B needs far less memory and more compute per token. In the report, both are actually on the quality-vs-serving-cost Pareto frontier among the models we evaluated, just at different points.
> intellectual-property safety
My suspicion is that you simply can't build an even slightly competitive model without liberally stealing your training data, in 2026, as much as I'd like it to be otherwise. You can get to the point that I suspect most of the frontier labs are at, where you've laundered the initially stolen data through the creation of huge amounts of derivative synthetic data, but still. Anyone who isn't comfortable stealing their training data is bringing a knife to a gun fight, and is going to die a noble but inevitable death.
There's also the ability to distill other models, which is also not illegal (though I'm sure they like to come after whomever for TOS violations, but thats a civil matter).
And, of course, the obligatory copying-isn't-theft observation. A recent supreme court judgment put it well.
> Since the statutorily defined property rights of a copyright holder have a character distinct from the possessory interest of the owner of simple “goods, wares, [or] merchandise,” interference with copyright does not easily equate with theft, conversion, or fraud. The infringer of a copyright does not assume physical control over the copyright, nor wholly deprive its owner of its use. Infringement implicates a more complex set of property interests than does run-of-the-mill theft, conversion, or fraud.
Folks are pretty smart here, I think we can handle these nuances, even if we don't agree about whether they are good.
Edit: reading through the full text of their post, it looks like they are using common crawl, which is likely just as much of a copyright infringement as Anna's Archive -- it's not like published works have a unique claim to copyright. I think this strengthens your point, though: I was expecting to see scans as training data, but it doesn't appear to be the case.
Copyright is a government mandated monopoly that was only granted in order to advance the arts and science. Any interpretation that runs contrary to that is bollocks being used by the religiously or financially motivated to serve their own petty interests to the detriment of societies.
I think it’s a safe assumption that they’re leaning into “sovereign” because performance is bad.
There is a proliferation of sovereign models under development specifically to address data sovereignty, and a loss of performance is absolutely acceptable over the risk that a once ally will turn adversarial, or a foreign business stops serving what has become critical infrastructure.
Unfortunately, Qwen3.6 35B A3B isn't really a useful coding model. You'd probably want Qwen3.8 27B at a minimum, which requires at least 32GB of VRAM (not system RAM) to run semi-comfortably.
So this isn't going to be a competitive model for hobbyists, and you'd have to be a bit desperate to use it for coding. But if you work in a regulated industry and don't mind paying for a bit of extra hardware, it isn't catastrophically bad, either. Probably would work fine for information extraction or as a "classifier" like Jev. (Almost any GGUF model can be turned into a classifier using llama-server. See pi.dev codemode for sample code.)
So they're not a real contender yet, but they look like they're probably at least minimally credible.
Thus training on 'clean' data is like trying to unscramble an egg.
I feel I messed up your quip =/ I'm new here, go ez. Not looking for excuses to hate on China either.
For what it's worth, in my language we don't have a word for copyright either. We have the concept, though, we just call it literally Creators Rights זכויות יוצרים and the borders of what is and what isn't covered broadly map to the familiar concepts of IP.
A real sovereign effort could invest heavily in this, whatever people accuse China of “stealing” I’m sure they are also generating tons of their own data and are probably the primary sovereign doing so outside the US labs.
You're right that especially big models benefit from training on copyrighted material in terms of world knowledge (especially from books). However, in the small model space imho agentic capabilities where the model looks up knowledge on the fly are much more important. That's what we focused on quite a bit during training. Personally, I also don't think stealing stuff is okay.
Regulate large cloud services and proprietary software - yes! But not on the basis of "Intellectual Property".
The "legal" issues here are very very complex and we should not passively wait for or accept corrupt court rulings, international trade agreements, proposed laws, or worst of all propaganda that pushes a parochial and craven view on this.
"an LLM" -- does that mean they are effectively learning from that LLM the German encyclopedic style? makes me wonder which LLM and how that is really sovereign.
> built for sovereign mission-critical work in regulated areas including public administration, industrials and aerospace
Cool to be doing more independent model lineages, but I sure hope no one actually uses this as part of any aerospace engineering...
Model Training as Code - https://news.ycombinator.com/item?id=48673450 - June 2026 (24 comments)
How is a 27B dense model bigger than a 78B MoE?
I wonder how far we are from this. How far are we from LLM's Debian moment?
If they can follow Qwen3.8-Flash-Next, the could draft off the huge reduction in VRAM requirements.
Congrats to the team.
What matters is to have good sovereign models in 5, 10, 20 years time.
The next thought is culture-specific workflows, requirements etc that are best trained into weights by people that actually understand those needs.
Society can't depend on the whims of the private sector.
Yeah we can, we do it all the time. Nothing in society works without relying on private actors to keep doing what they have always done. Unless you go full soviet command economy, which in fact works worse.
But I don't see why this needs to be done on a state-by-state level, a sovereign EU model seems to be much more realistic in terms of funding.
The annoying part of optical matrix multiplication is the conversion from electronic to optic and back.
Why not use electron-optics for matrix multiplication and use dynodes for amplification.
Not all needs are going to be met by using or finetuning general purpose models, so ability to build your own SOTA ML models (LLMs or not) is also important.
I mean, given the fuel prices it would be very nice, alongside more charging infrastructure in the EU. :(
Not to detract from the overall point, though to be honest, it's probably worthwhile to do improvements and refinements with the current technology, even if the future holds something vastly different. Both cause of gaining expertise and also maybe an improvement or two along the way, that might carry forward.
Like I haven't seen many steam locomotives around and for me the difference between saturated steam engines and superheated ones isn't very material in regards to transportation, but the latter is used in modern turbines.
Coal is burned and converted to electricity in coal power plants and the electricity is used to move the electric locomotives around.
Even with distribution losses, the thermal efficiency of coal power plant -> electric network -> electric train is much higher then the old steam locomotives (Only about 5% of the potential energy produced by a steam locomotive’s boiler is translated to the wheels in the form of actual driving power.)
And who are you to be asking such questions?
Someone that builds things instead of merely "managing".
Someone pissed by this exact style of default github pfp business-person ruining this country. If they haven't left for SV already, in which case I hope that they will stay there.
> Aleph Alpha is just a sad joke by now.
And this isn't ad-hominem for lack of factual critic. The personal stuff _is_ the reason. How else do you call out someone driven by ego without mentioning them and their ego?
The guy is only ego. There is no substance to attack.
this is, because it can mean so much stuff that you can't really tell you are calling out their "ego" or just trolling.
Until you can at least open yourself to the possibility that there are better more inviting ways to talk to people online, you'd think this is just the natural way people talk online, but I'd say it's not. It's just an aesthetic choice.
If you truly think that I am the problem deep down in this nested subthread that violates the spirit and letter of the rules of the platform, I have zero faith in the calibration of your perception.
Serious projection issues, and anger to boot. Maybe take a walk outside, that was a very silly comment to get so worked up about.
AI is here to stay - will be around for hundreds/millions of years. Whether company A is a few years ahead of company B is irrelevant. In 10 years time company A, who started the industry, may be gone completely - also irrelevant.
VW EV sales are up in South America, Canada and Europe. It's true that their total global numbers are down, but that's mainly because they lost like 30% in China. Which obviously sucks considering how big the Chinese market is, but then, the Chinese market is it's own sort of thing.
Don't get me wrong, I think your point is valid for a lot of traditional European car brands. Mercedes is certainly one of them which has lost it's "our engines are nicer" brand, but I suspect VW is one of the brands that will do just fine. Especially with the id polo coming out next year.
Mercedes is doing something interesting with its level 3 DRIVE PILOT. As far as I know, it's the only autonomous driving system/auto pilot on the market that actually takes full legal responsibility for the vehicle and its actions. If you put a new Mercedes in self driving and it crashes, Mercedes takes full liability.
Yes, it currently only goes up to 100 kph and it will buzz you to take over if it can't guarantee prefect safety anymore. But this approach is really the only one I think is interesting. All this self driving shit is worthless to me if I'm still liable in the end. So I respect Mercedes for putting their money where their mouth is.
But anyway. I think the main oversight when dismissing this is that not every use-case is coding a SV-style startup app. That market is quite saturated, so it would make sense to create something locally for the use-cases currently underserved by LLMs.
We will probably learn more about what this can really do once quants become available that can be run by people without an SV salary (and the biases that come with that).
Then what does it good at? Sending faxes?
Next day, the books would arrive. This was maybe 1996, but I thought it was pretty cool.
Maybe you can relate to seeing webcams or Facetime when it came out or something.
Being able to see a signed document or a picture someone took 5000km away quickly was mind boggling in its age.
They were never common. They were mostly business tools, and consumers never wanted them. Sometimes you'd be required to send someone a fax, and you'd have to go to an office store. (And pay rather a lot; dollars per page.)
I hope you get to give one a try some day, for kicks, but the thrill will wear off fast.
Are you sure about that? I remember we had one at home when I grew up, so I assumed they were very common at the time. (And it's not like we had this crazy business driven family use of it.)
> I never saw one in a TV show
Now that you mention it, I think I only remember one from the simpsons.
But based on [1] my mom really was an outlier for getting a fax machine. Mabye we had enough contact with beurocracy to make it work it. :D Even in 1995 only 1.5 million out of 36 million (4,1%) households had one. (Although I remember it from ~2000, so too bad the data doesn't go there.)
For the US with 14 million out of 98 million households (14%) fax machines were more popular for ordinary households than in germany. (But home-use was never the issue in the fax debate anyway, just fun facts :) )
[1] https://www.oecd.org/content/dam/oecd/en/publications/report...
I even experienced the opposite case: Some feature was first rejected since everyone agreed that it would break some behavior. After talking to several people again, I discovered that the use case which required this behavior did not exist anymore and had been officially phased out years before. Maybe, some RAG LLM could have told me, based on company documents, that there is some edge case that can be removed making the way free for the new feature.
This means that it’s always on time, it uses acronyms for everything and when there’s a decision to be made, it sets up a committee.
Tell that to Deutch Bahn.
We also talk more about German pre-training data in the tech report: https://aleph-alpha.com/downloads/tech-report.pdf section 2.3.2.2.
Disclaimer: I am part of the team that trained Kolibri
Of course, it's impossible to know for sure what was LLM processed or not, but this post did get classified that way. That's why the software flagged it.
That explains why they're framing it around sovereign deployment for public administration / aerospace rather than chasing general English benchmarks. The Pareto-frontier claim on throughput vs quality (Figure 1, 8xB200 evals) is also refreshingly honest — serving cost matters a lot for regulated on-prem use.
Would love to see more detail on how the synthetic German data was validated for quality, and how MergeMix data mixing affected German vs English trade-offs. Apache 2.0 open weights is a big plus here.