HNHacker News
TopNewBestAskShowJobs

molf

1,834 karma · joined March 22, 2012

https://voormedia.com

https://zxcv.art

submissionscomments
molf··on Mistral raises 1.7B€, partners with ASML
ASML CEO: Mistral investment not aimed at strategic autonomy for Europe

"In the long run, all AI models will be similar. It's about how you use the models in a well-protected environment. We will never allow our data and that of our customers to leave ASML. So a partner must be willing to work with us and adapt its model to our needs. Not only did Mistral want to do that, it is also their business model."

https://fd.nl/bedrijfsleven/1569378/asml-ceo-strategische-au...

--

Full article translated:

“A good reason to collaborate.” That's how ASML's CEO described his company's remarkable €1.3 billion investment in French AI company Mistral on Wednesday. Since the investment was leaked by Reuters on Sunday, there has been much speculation about ASML's reasons for investing in the European challenger to giants such as OpenAI and Anthropic. Analysts and commentators pointed to the geopolitical implications or the strong French link between the companies. But according to ASML CEO Christophe Fouquet, the reason was purely business. “Sovereignty has never been the goal.”

Mistral AI is a start-up founded in 2023 that specializes in building large language models. The French CEO of ASML and Mistral CEO Arthur Mensch met at an AI summit in Paris earlier this year and decided to work together to use Mistral's models to further improve ASML's chip machines.

Surprising investment

Each ASML machine generates approximately 1 terabyte of data per day. “Our machines are very complex,” Fouquet explains in an interview with the FD. "We have highly advanced control systems on our machines to enable them to operate very quickly and with great accuracy. The amount of data our machines generate gives us the opportunity to use AI. With the current software and machine learning models, we are limited in what we can do with the data and how quickly we can adjust the machine,“ says the CEO. ”AI is the next step in making better use of all that data."

ASML has invested in other companies in the past, such as German lens manufacturer Zeiss and Eindhoven-based photonics company Smart Photonics, but those were either suppliers or potential customers. Mistral is neither.

Running AI models in-house

According to the ASML CEO, the Dutch company's investment in Mistral stems from the conviction that both companies can create value together. If Mistral becomes more valuable as a result of the collaboration, ASML can benefit from that.

ASML is the main investor in a new €1.7 billion financing round for Mistral. This makes Mistral an important AI player in Europe, but small compared to its American rivals. OpenAI raised $40 billion in its latest round alone. Anthropic, the company behind the Claude program, which is popular among programmers, just closed a $13 billion round.

“European sovereignty was not the goal”

According to Fouquet, the reason for the collaboration lies primarily in the way Mistral develops its AI models. “In the long run, all AI models will be similar. It's about how you use the models in a well-protected environment,” says Fouquet. “We will never allow our data and that of our customers to leave ASML. So a partner must be willing to work with us and adapt its model to our needs. Not only did Mistral want to do that, it is also their business model.”

According to Fouquet, the collaboration is not motivated by a desire for greater European sovereignty. “That was not the goal. But if it contributes to that, we are happy,” says Fouquet.

ASML supports EU initiatives to strengthen the chip sector in Europe, but always maintains a politically neutral stance in the geopolitical struggle between the United States, China, and the European Union. This is understandable, as the company has major customers in all regions, such as TSMC in Taiwan, SK Hynix in South Korea, SMIC in China, and Intel in the US.

“Two birds with one stone”

Although ASML itself does not play the European card, some analysts and politicians do see such a motive for the collaboration with Mistral. “Thousands of large companies worldwide make extensive use of AI in their product development by using the services of OpenAI, Meta, Microsoft, Google, Mistral, without investing in these companies,” writes investment bank Jefferies in a commentary. “We also do not believe that ASML needed an investment in an AI company to benefit from AI models in its lithography products. In our view, the investment stems primarily from geopolitical motives to support and develop a European AI company and ecosystem,” the bank states.

Wouter Huygen, CEO of AI consultancy Rewire, also sees a clear link to European sovereignty. “ASML is known for taking internal technology development very far. It is therefore quite understandable that ASML is taking this step: access to and influence on the development of a strategic technology. Plus European sovereignty. That's two birds with one stone.”

molf··on Mistral raises 1.7B€, partners with ASML
Not only are the CEO + COO French, they recently hired Le Maire, French ex-minister of Finance as a strategic advisor. ASML has also been rumoured to exit the Netherlands and relocate to France.

It is definitely a political move.

molf··on Google can keep its Chrome browser but will be barred from exclusive contracts
If there were a culture of always including the original source, or journalists massively advocating to include the original source, then surely the CMS would cater to it. I think it's safe to draw the conclusion that most journalists don't care about it.
molf··on Do Things That Don't Scale (2013)
I feel this is an insanely distorted take.

How about extreme and utter irrelevance (such as after building a thing nobody wants)?

Or how about this, arguably the most common: slightly successful; nobody hates it but nobody loves it either. Something people feel mildly positive about, but there is zero “hype” and also no “moat” and nobody cares enough to hate it.

molf··on The bitter lesson is coming for tokenization
The key insight is that you can represent different features by vectors that aren't exactly perpendicular, just nearly perpendicular (for example between 85 and 95 degrees apart). If you tolerate such noise then the number of vectors you can fit grows exponentially relative to the number of dimensions.

12288 dimensions (GPT3 size) can fit more than 40 billion nearly perpendicular vectors.

[1]: https://www.3blue1brown.com/lessons/mlp#superposition

molf··on “Don’t mock what you don't own” in 5 minutes (2022)
But ultimately it suggests this test; which only tests an empty loop?

  def test_empty_drc():
      drc = Mock(
          spec_set=DockerRegistryClient,
          get_repos=lambda: []
      )

      assert {} == get_repos_w_tags_drc(drc)
Maybe it's just a poor example to make the point. I personally think it's the wrong point to make. I would argue: don't mock anything _at all_ – unless you absolutely have to. And if you have to mock, by all means mock code you don't own, as far _down_ the stack as possible. And only mock your own code if it significantly reduces the amount of test code you have to write and maintain.

I would not write the test from the article in the way presented. I would capture the actual HTTP responses and replay those in my tests. It is a completely different approach.

molf··on “Don’t mock what you don't own” in 5 minutes (2022)
I’m not sure this is good advice. I prefer to test as much of the stack as possible. The most common mistake I see these days is people testing too much in isolation, which leads to a false sense of safety.

If you care about being alerted when your dependencies break, writing only the kind of tests described in the article is risky. You’ve removed those dependencies from your test suite. If a minor library update changes `.json()` to `.parse(format="json")`, and you assumed they followed semver but they didn’t: you’ll find out after deployment.

Ah, but you use static typing? Great! That’ll catch some API changes. But if you discover an API changed without warning (because you thought nobody would ever do that) you’re on your own again. I suggest using a nice HTTP recording/replay library for your tests so you can adapt easily (without making live HTTP calls in your tests, which would be way too flaky, even if feasible).

I stopped worrying long ago about what is or isn’t “real” unit testing. I test as much of the software stack as I can. If a test covers too many abstraction layers at once, I split it into lower- and higher-level cases. These days, I prefer fewer “poorly” factored tests that cover many real layers of the code over countless razor-thin unit tests that only check whether a loop was implemented correctly. While risking that the whole system doesn’t work together. Because by the time you get to write your system/integration/whatever tests, you’re already exhausted from writing and refactoring all those near-pointless micro-tests.

molf··on Selfish reasons for building accessible UIs
Not just in dev tools; that mess is also in your source code...
molf··on In case of emergency, break glass
The points about visual hierarchy are spot on, in particular on macOS. I think Apple has two realistic paths forward to resolve this mess:

1. Double down on the aesthetic and gradually redesign apps to improve the hierarchy. That would mean adapting UX across countless apps to serve the new look.

2. Tone down the glass effects and shadows drastically. Preserve existing app layouts without compromising usability as much. We'll be left with shimmering buttons and panels, a bit more blurred transparency than in the 'current' design language.

My guess is they will end up choosing option 2, simply because it’s cheaper.

molf··on How we’re responding to The NYT’s data demands in order to protect user privacy
It would help tremendously if OpenAI would make it possible to apply for zero data retention (ZDR). For many business needs there is no reason to store or log any request at all.

In theory it is possible to apply (it's mentioned on multiple locations in the documentation), but in practice requests are just being ignored. I get that approval needs to be given, and that there are barriers to entry. But it seems to me they mention zero-data retention only for marketing purposes.

We have applied multiple times and have yet to receive ANY response. Reading through the forums this seems very common.

molf··on A maths proof that is only true in Japan
Totally get your point, but math is still a human creation. The symbols, language, and frameworks we use are cultural, and disagreement over proofs like this one shows math depends on shared understanding, not just objective truth.
molf··on A South Korean grand master on the art of the perfect soy sauce
This is the first time I hear about keeping soy sauce in the fridge. Is this common?
molf··on Fast Allocations in Ruby 3.5
C itself is fast; it's calls to C from Ruby that are slow. [1]

Crossing the Ruby -> C boundary means that a JIT compiler cannot optimize the code as much; because it cannot alter or inline the C code methods. Counterintuitively this means that rewriting (certain?) built-in methods in Ruby leads to performance gains when using YJIT. [2]

[1]: https://railsatscale.com/2023-08-29-ruby-outperforms-c/ [2]: https://jpcamara.com/2024/12/01/speeding-up-ruby.html

molf··on LLM function calls don't scale; code orchestration is simpler, more effective
Of course it would hallucinate. It would just pick arbitrary/wrong values.
molf··on AI Horseless Carriages
Agree 100%.

We built a very niche business around data extraction & classification of a particular type of documents. We did not have access to a lot of sample data. Traditional ML/AI failed spectacularly.

LLMs have made this super easy and the product is very successful thanks to it. Customers love it. It is definitely transformative for them.

molf··on UML diagram for the DDD example in Evans' book
This seems like a business problem more than a design issue. Systems need to evolve alongside the business they support. Starting out with a simple design and evolving it over time to something more nuanced is a feature. Your colleague was right, and you were also right; except for the part where all nuances of the ideal solution need to be present on day 1.

The clients you have on day one are often very different from the ones you’ll have a few years in. Even if they’re the same organisations, their business, expectations, and tolerance for complexity likely have changed. And the volume of historical data can also be a factor.

A pattern I’ve seen repeatedly in practice: 1. A new system that addresses an urgent need feels refreshing, especially if it’s simple. 2. Over time (1, 3, 10 years? depending on industry), edge cases and gaps start appearing. Workarounds begin to pile up for scenarios the original system wasn’t built to handle. 3. Existing customers start expecting these workarounds to be replaced with proper solutions. Meanwhile, new customers (no longer the early adopter type) have less patience for rough edges.

The result is increasing complexity. If that complexity is handled well, the business scales and can support growing product demands.

If not… I'm sure many around here have experiences where that leads (to borrow Tolstoy: “All happy families are alike; each unhappy family is unhappy in its own way.”).

At the same time a market niche may open for a competitor that uses a simpler approach; goto step 1.

The flip side, and this is key: capturing all nuances on day 1 will cause complexity issues that most businesses at this stage are not equipped to handle yet. And this is why I believe it is mostly a business problem.

molf··on How to protect your phone and data privacy at the US border
"It didn't happen to me, therefore it is not a problem."
molf··on Coding Isn't Programming
I think he argues that more thought should go into explicitly designing the behaviour of a software system ("What?"), independently from its implementation ("How?"). Explicitly designed program behaviour is an abstraction (an algorithm or a set of rules) separate from its implementation (code written in a specific language).

Even if a user cannot precisely explain what a program needs to do, programmers should still explicitly design its behaviour. The benefit, he argues, is that having this explicit design enables you to verify whether the implementation is actually correct. Real-world programs without such designs, he jokes, are by definition bug-free, because without a design you can't determine if certain behaviour is intentional or a bug.

Although I have no experience with TLA+ (which he designed for this purpose in the context of concurrency), this advice does ring true. I have mentored countless programmers, and I've observed that many (junior) programmers see program behaviour as an unintentional consequence of their code, rather than as deliberate choices made before or during programming. They often do not worry about "corner cases", accepting whichever behaviour emerges from their implementation.

Lamport says: no, all behaviour must be intentional. Furthermore, he argues that if you properly design the intended program behaviour, your implementation becomes much simpler. I fully agree!

molf··on The lottery of the snakebite antivenom industry
> The first snakebite antivenom was made in the mid-1890s and the method has changed little since: snakes are “milked” for their venom, which is injected into horses or sheep and the antibodies that their immune systems then produce are extracted via the animals’ blood.

This is expensive, and many people develop allergies to the antibodies.

A cool practical application of DeepMind's AlphaFold is that it allows one to "design" proteins because it is now possible to predict how they will fold. [0]

This means it is possible to create synthetic protein antibodies that neutralise snake venom, specifically designed for humans [1].

There is a recent Veritasium video [2] that contains a great explanation. The whole video is worth a watch!

[0]: https://ui.adsabs.harvard.edu/abs/2023Natur.620.1089W/abstra...

[1]: https://www.nature.com/articles/s41586-024-08393-x

[2]: https://www.youtube.com/watch?v=P_fHJIYENdI&t=20m36s

molf··on Show HN: App that simulates a software engineer's daily job
Is it satire because many of the buttons and links don't work so you can't actually finish anything?

If it’s serious, I doubt it provides much real value. Developer experience is built on accomplishing actual tasks, not the surrounding administration and ceremony. Aspiring devs should probably work on personal/toy projects or open source.

That said, it looks great—well done!

molf··on Anthropic achieves ISO 42001 certification for responsible AI
You are referring to Article 5 1.f?

"1 The following AI practices shall be prohibited: (...)

"f) the placing on the market, the putting into service for this specific purpose, or the use of AI systems to infer emotions of a natural person in the areas of workplace and education institutions, except where the use of the AI system is intended to be put in place or into the market for medical or safety reasons"

See recital 44 for a rationale. [1] I don't think this is "hilarious". Seems a very reasonable, thoughtful restriction; which does not prevent usage for personal use or research purposes. What exactly is the problem with such legislation?

[1]: https://artificialintelligenceact.eu/recital/44/

molf··on Don’t look down on print debugging
To be fair though: printing also will likely impact timing and can change concurrent behaviour as well.

Still agree that print debugging is more useful in such situations (and I prefer it in general).

molf··on Postzegelcode
In the case of sloppy writing, the letter will be ejected from the sorting machine and reviewed by a human (same process as for illegible addresses).

If a human cannot find a match then the sender will receive a request to pay for the missing postage (presuming the code is invalid). If the sender is unknown the recipient will receive a payment request (which can be appealed).

Speaking from experience you would normally want the letter to arrive at its destination so I take care to write the code very clearly. I imagine this is true for most people.

molf··on Postzegelcode
For all mail except the first, the same thing that happens if you send a letter with no postage: the sender will receive a request to pay after delivery. If the sender is unknown the recipient will receive a request to pay (which can be appealed).
molf··on JavaScript-Is-Weird as a compressor
Using `xz -9 -e` results in pretty reasonable compression, in comparison:

     19648  dommy-2.0.js
     28092  dommy-2.0.weird.js.xz
    115023  lodash-4.17.15.js
    138808  lodash-4.17.15.weird.js.xz
      7114  modernizr-custom.js
     11148  modernizr-custom.weird.js.xz
molf··on Show HN: uFuzzy.js – A tiny, efficient fuzzy search that doesn't suck
Many search engines default to returning hits when at least 1 keyword match. Not returning any results just because not every single keyword could be matched leads to a poor user experience. However, good relevance scoring is key here. Hits that match most keywords should be somewhere at the top. The reasoning is that a user will try to find the result they are looking for in the top hits.

In the case of MiniSearch you can change this default and configure it to only return hits if all keywords match [0].

Searching for "super ma" seems like an autocomplete query, for which most users have subtly different expectations than regular search. MiniSearch has a separate method for that, essentially baking in different default settings. [1]

[0] https://lucaong.github.io/minisearch/classes/_minisearch_.mi...

[1] https://lucaong.github.io/minisearch/classes/_minisearch_.mi...

Edit:

> s/everything/anything

No need for the sneer.

molf··on Show HN: uFuzzy.js – A tiny, efficient fuzzy search that doesn't suck
A blunt answer to your question is: I didn't take the time to read everything.

Sorry for missing that feature. I stand by my overall point though. There is more to search than finding matches. Handling real world typing errors is one (try searching for "mario avdentures" or "mario adventutes").

Ordering results by relevance is not trivial either. And it does not appear this library intends to tackle that problem completely. Relevance scoring should take into account the term frequency and the document/field length, as well as average length.

I don't mean to discount your project though. There are many distinct use cases related to searching and filtering, and power to you for solving it in a way that works for you (and I'm sure for others as well). I just wanted to share my experience in exploring the space of full text search libraries that handle long form text and typos gracefully.

molf··on Show HN: uFuzzy.js – A tiny, efficient fuzzy search that doesn't suck
I have spent many days searching for the best JavaScript implementation of full text search that handles typos (substitutions) well. Implementing a good indexing algorithm for this is not easy. In particular if you are indexing large amounts of text (documents) instead of short strings.

I settled on MiniSearch. [0] It is fast & small enough and fairly feature complete.

Afterwards I made a few contributions to improve performance and implement a better scoring algorithm. So I'm probably a bit biased now. Take my recommendation with a grain of salt.

Personally I think that OP's library does not perform searches, fuzzy or otherwise. It's much more similar to 'grep'. Try searching for "mario adventures". It won't actually find the most obvious results, because the order of the keywords in the search string must match the order of the keywords in the indexed text.

[0]: https://github.com/lucaong/minisearch

[1]: https://leeoniya.github.io/uFuzzy/demos/compare.html?libs=uF...

molf··on Ask HN: Pitfalls in jumping ship to a client company?
You're correct in your intuition that this will affect the business relationship. Would your potential new employer prefer you over the business relationship with your current employer? If no, this is a risk for you which could potentially lose you both job opportunities. In any case you need to discuss this with your potential new employer as soon as you apply.

Some managers are empathic if you are open and honest and you might try disclosing your plans to your current employer somewhat early during the hiring process. This might it less likely for the business relationship with the client (your potential new employer) to sour. Of course, there is the possibility it backfires. Depends on the relationship between you and them.

molf··on Give nothing, expect nothing: Gitlab latest punching bag for entitled users
If you have more than 5 users, the free tier is not an option.

We liked and paid for Gitlab, but the price hike from $4 to $19 per user is absurd. Forcing an annual subscription on us was the motivation we needed to migrate away.

← PreviousPage 2 of 9Next →