HNHacker News
TopNewBestAskShowJobs

BenoitP

2,852 karma · joined August 8, 2013

https://benoit.paris

http://explicable.ai

benoit.paris.753@gmail.com

https://www.linkedin.com/in/benoitparis/

Kismet: 3b638cc74046887b5af69269b5c351d53acf4b702871f1c13691369db2917550

submissionscomments
BenoitP··on Calyx: Intermediate Language for Hardware Accelerators
From a glance: it seems like it is more specific for accelerators. Chisel deals in the general way, but Calyx seems to have primitives for wiring systolic arrays for example. If you plan to do an OoO CPU, choose Chisel. If you want to do a DSP or TPU, choose Calyx.
BenoitP··on Measuring energy usage: regular code vs. SIMD code
Please look at these number with the following grain of salt, when optimizing programs. They are trumped by cache usage patterns [1]:

    Loading (fig IV)
    
    Location    Energy (pJ = pico Joules)
    L1          64   pJ/Byte
    L2          121  pJ/Byte
    L3          254  pJ/Byte
    RAM         1250 pJ/Byte
    
    Adding integers with SIMD (fig VI)
    428 pJ/op, for 8 Byte/op; this means: 
    53 pJ/byte

So it takes actually 20 times to fetch data from RAM than to add it to something! And most often this is also the source of latency. Generally energy is linear to the distance signal has to travel, and that's the same for latency.

That's why successful data structures are sized a tiny bit under L1/L2 sizes! (BTree chunks, ring buffers).

If you've been following hardware, it's all about putting RAM closer to compute at the moment with chiplets at the moment.

[1] https://tu-dresden.de/zih/forschung/ressourcen/dateien/abges...

BenoitP··on AMD Athlon K7 Easter egg has a revolver and map of Texas etched onto the chip
Chips was a mistake
BenoitP··on AMD Athlon K7 Easter egg has a revolver and map of Texas etched onto the chip
Fun factoid: Chips need to have patterns, even where there is nothing to print! This is because you want the acid etchant to be consumed the same way everywhere. So you choose a random repeating pattern that has the same density as your other signals density; it will serves no purpose other than consuming nitric acid. But you have freedom in choosing the exact shape.

Source: a friend of mine working at a chip company 20 years ago. They considered disparaging the competition billions of time in tiny 250nm letters. It occurred right after a patent dispute where it was highly likely they had reverse-engineered part of their design. A disdainful message only their competitor would read, to the tune of "we know you're cheating, and fart in your general direction". It was mostly a joke though, they did not go through with it.

----

AMD's one is only in one corner of the chip, so I don't think it applies here.

----

EDIT: this was a long time ago. Seems like filling has evolved quite a lot [1]. Timing, signal integrity, and probably capacitance too are so constrained nowadays you can no longer write a personalized message to your competitors.

[1] https://semiengineering.com/knowledge_centers/materials/fill...

BenoitP··on Show HN: Rank a random subject every day
Apple is #1, but a lot of other fruits are way less known. This is the equivalent of every not-well-thought map visualization being a population density map.
BenoitP··on Ask HN: What have you built with LLMs?
I'd buy that. I'd buy that for interview preparation as well. Maybe 5$ per hour, up to 15$. I wouldn't buy a subscription, only actual consumption of the service.

Please consider putting it in online.

BenoitP··on More misdrilled holes on 737 MAX in latest setback
I've read that he was the token engineer in the see of accountants, MBAs and consultants.

Cultural change goes much deeper than changing the CEO. Boeing might not be salvageable on the cheap side. Decapitating the company at levels 1-4 may not even be enough. It may require acknowledging that changing the flight envelope and software-patching it was a mistake.

This is a major can of worms on itself. This means a plane redesign, 6-8 years delay, and probably another round of WTO-unfriendly subsidy.

BenoitP··on Why Don't We Teach People How to Parent?
Just like tech goes through cycles and everything old is new again (Corba to XML to JSON to gRPC), maybe morals are too on a cycle. In tech the typical time constant is for the young gen to burn and churn. In morals it is a few generations (as grandparents influence it)

Maybe the cycle will bring back social shame again, which to me is at a historical low. And we'll again have manners manuals with content copied from 1924. Politeness, saying please and thank you, waiting for adults to finish before interrupting, doing house chores, no shouting, etc.

BenoitP··on Ask HN: Freelancer? Seeking freelancer? (February 2024)
SEEKING WORK | Paris, France | Remote-open

---------------------------

Machine learning engineer, specialized in Explainable AI / ML. 13 years professional coding experience, 24 years personal.

Recent highlights:

* Seamless integration of hand-crafted rules with ML-build rules

* Intuitive, visual data/signal explorer: http://explicable.ai

* Implementation in Spark/Scala of treeinterpreter, currently used in production

* Participation to the FICO-Google Explainable Machine Learning Challenge

---------------------------

Tech: SHAP, RuleFit, Random Forest, Word2Vec, PCA, t-SNE, LSH, Scikit-Learn, Spark, Flink, Weka, Databricks, BigQuery, Hive, Postgres, MySQL, Oracle, AWS, Linux, Maven, Git, Java, Scala, Python, CAML, Elm, Javascript, Typescript, React, Spring, Primefaces, d3.js, Three.js

Résumé/CV: https://www.linkedin.com/in/benoitparis/

Github: https://github.com/benoitparis/

Email: benoit.paris.753@gmail.com

Blog: https://benoit.paris

BenoitP··on Ask HN: Who wants to be hired? (February 2024)
Machine learning engineer, specialized in Explainable AI / ML. 13 years professional coding experience, 24 years personal.

Highlights:

* Seamless integration of hand-crafted rules with ML-build rules

* Intuitive, visual data/signal explorer: http://explicable.ai

* Implementation in Spark/Scala of treeinterpreter, currently used in production

* Participation to the FICO-Google Explainable Machine Learning Challenge

Location: Paris, France

Remote: open to it

Willing to relocate: for the right job, yes

Technologies: SHAP, RuleFit, Random Forest, Word2Vec, PCA, t-SNE, LSH, Scikit-Learn, Spark, Flink, Weka, Databricks, BigQuery, Hive, Postgres, MySQL, Oracle, AWS, Linux, Maven, Git, Java, Scala, Python, CAML, Elm, Javascript, Typescript, React, Spring, Primefaces, d3.js, Three.js

Résumé/CV: https://www.linkedin.com/in/benoitparis/

Github: https://github.com/benoitparis/

Email: benoit.paris.753@gmail.com

Blog: https://benoit.paris

BenoitP··on Ask HN: Could we just re-invent original Google?
TL;DR: Goodhart's law rots everything it touches

SEO is in the structure of the internet now. Original Google was great because there was no incentive yet to buy a domain and blogspam it. Google getting shitty is just a natural instance of Goodhart's law, applied to domains and content.

Now, Google originally was based on PageRank; which based itself on every domain being a unit of authority. These have been compromised and drowned by SEO, but the concept remains valid and we could choose people as units of authority. For example PageRank on scientific papers accurately reproduce Nobel Prizes attributions. A person publishing papers is a solid enough foundation for this unit of authority.

It remains to be organized though. And if we take people as units of authority, it means they'd have to 'cite' or vote for each other. This has social consequences and might not be doable. Are you ready to refuse to cite your boss when he/she ask you to do so? Maybe if the vote is secret and delayed by 5 years?

BenoitP··on Useful Models: The Barbell Strategy
Seems like the technical risk budget in startup. As a startup is an exercice in exploration vs risk management that has to be successful; it should make a big bet in one dimension, and be really conservative on the rest.
BenoitP··on Reasons to avoid static type checking in Python
Yeah the formulation is quite biased, but to some extent there is some truth in the statements:

> Your codebase is old, large and has been working fine without static type checking for years. While Python’s type system is designed to allow gradual adoption of static type checking, the total cost of adding type annotations to a large extant codebase can be prohibitive.

That's a convoluted and quite dishonest way of warning that you're trapped if you don't enforce types from the start; but it does have the merit of being on the list.

BenoitP··on Ask HN: Does (or why does) anyone use MapReduce anymore?
The concept is quite alive, and the fancy deep learning have it: jax.lax.map, jax.lax.reduce.

It's going to stay because it is useful:

Any operation that you can express with an associative behavior is automatically parallelizeable. And both in Spark and Torch/Jax this means scalable to a cluster, with the code going to the data. This is the unfair advantage of solving bigger problems.

If you were talking about the Hadoop ecosystem, then yes Spark pretty much nailed it and is dominant (no need to have another implementation)

BenoitP··on Amazon Fined $35M in France over 'Overly Intrusive' Surveillance
The total revenues of Amazon's activities in France were over €9 bn in 2022. So we're at about 0.3%. A warning shot IMHO
BenoitP··on Amazon Fined $35M in France over 'Overly Intrusive' Surveillance
Indeed, but courts in France have the habit of always firing warning shots. Non-compliance can trigger another court decision with forceful consequences.

So far GDPR has only yielded symbolic amounts. But the legislative cap is at profit-destroying levels.

BenoitP··on Sam Altman Says AI Using Too Much Energy Will Require Breakthrough Energy Source
Facebook is not alone, and there is growth. Also cooling is to be taken account of.

And third: renewables need to be associated with its backup like hydro/step or ... batteries which cost a lot. Gas can't be taken in as it's not CO2-free. All that unless training and inference happen when there's the corresponding wind and sun shining. And I'm not seeing that happening right now.

BenoitP··on Jazelle DBX: Allow ARM processors to execute Java bytecode in hardware
> A few programmers will remain as the LLMs high priests.

That's interesting.

It's controversial to say that in 2024, but not all opinions have the same value. Some are great, but some are plain dumb. The current corporate right opinion is to praise LLMs as end-all be-all. I've been asked to advise a private banking family office wanting to get into LLMs. For advising their clients' financial decisions. I politely declined. Can there be a worse use case? LLMs are parrots with the brain size of the internet. With thoughts of random origin mixed together randomly. It produces wonderful form, but abysmal analysis.

IMHO as LLMs will begin to be indistinguishable from real users (and internet dogs), there's going to be a resurging need to trace origin to a human; and maybe to also rank their opinions as well. My money is on some form of distributed social proof designating the high priests.

BenoitP··on Jazelle DBX: Allow ARM processors to execute Java bytecode in hardware
I'm still not decided on AOT vs JIT being the endgame.

In theory JIT should be higher performance, because it benefits from statistics taken at actual runtime. Given a smart enough compiler. But as a piece of code matures and gets more stable, the envelope of executions is better known and programmers can encode that at compile-time. That's the tradeoff taken by Rust: ask for more proofs from the programmers, and Rust is continuing to pick up speed.

That's also what the Leyden project / condensers [1] is about, if I understand correctly. Pick up proofs and guarantees as early as possible and transform the program. For example by constant-propagating a configuration file taken up during build-time.

Something I've pondered over the years: a programmer's job is not to produce code. It is to produce proofs and guarantees (yet another digression/rant: generating code was never a problem. Before LLMs we could copy-paste code from StackOverflow just fine)

In the end it's only about marginal improvements though. These could be superseded by changes of paradigm like RAM getting some compute capabilities; or programs being split into a myriad of specialized instructions. For example filters, rules and parsing going inside the network card; SQL projections and filters going into the SSD controller; or matrix-multiplication going into integrated GPU/TPU/etc just like now.

[1] https://openjdk.org/projects/leyden/notes/03-toward-condense...

BenoitP··on Jazelle DBX: Allow ARM processors to execute Java bytecode in hardware
The gains seem to not have been high enough to sustain that project. Nowadays CPUs plan, fuse and reorder so much of micro-code that lower-level languages can sort of be considered virtual as well.

But Java and similar languages extract more freedom-of-operation from the programmer to the runtime: no memory address shenanigans, richer types, and to some extent immutability and sealed chunks of code. All these could be picked up and turned into more performance by the hardware; with some help from the compiler. Sort of like SQL being a 4th-gen language, letting the runtime collect statistics and chose the best course of execution (if you squint at it in the dark with colored glasses)

More recent work about this is to be found on the RISC-V J extension [1], still to be formalized and picked up by the industry. Three features could help dynamic languages:

* Pointer masking: you can fit a lot in the unused higher bits of an address. Some GCs use them to annotate memory (refered-to/visited/unvisited/etc.), but you have to mask them. A hardware assisted mask could help a lot.

* Memory tagging: Helps with security, helps with bounds-checking

* More control over instruction caches

It is sort of stale at the moment, and if you track down the people working on it they've been reassigned to the AI-accelerator craze. But it's going to come back, as Moore's law continues to end and Java's TCO will again be at the top of the bean-counter's stack.

[1] https://github.com/riscv/riscv-j-extension

BenoitP··on Backing up the power grid with green methanol
I'll add that German plans up to now were to build turbines that could take both gas and H2. Even though these have very different combustions. H2 was brought in to power the green transition. The turbines were to be developed.

Now they're actually planning them and cutting cost, surprise surprise, the turbines are gas only.

BenoitP··on Serious incident to the 777-300ER on 5 April 2022 at CDG [pdf]
Sounds like my coworker telling a marketing campaign operator he has yet again let Excel interpret phone numbers as integers. And that he should learn to properly wield Excel. The downstream system is a call center made of humans. They lose a few seconds inputting the phone number manually though. No lead is lost. I don't want to speculate too much, but I think we might be in the clear. The client is not losing too much money. But if he complains once, we'll have to implement more checks to validate the Excel file. Then I'll have to buy my coworker breakfast, so that he doesn't bitch the whole implementation time. Such is life.
BenoitP··on Ask HN: OCR for 100 year old (German) handwritten cursive script?
I worked in the same space as a company that does this with ML (and charges for it), using some form of Recurrent Neural Network IIRC. Maybe LSTMs?

They had a contract to index historical French archives composed of handwritten latin documents in elasticsearch.

Depending of the historical relevance of your documents (read: some academic funds), they may be able to help. Doesn't hurt to contact them:

https://teklia.com/

BenoitP··on Solar will supply almost all growth in U.S. electricity generation through 2025
Cost figures are very hard to come by as:

* Financing and returns are spread over 60 years for nuclear, which makes it hugely dependent on the interest rate. Which we don't know in advance, and also is at risk of irrational public perception (Leading in democracies to early closure of perfectly fine plants. That affects ROI hugely).

* Nuclear too profits from scale: you have to maintain an army of competent engineers. But once done, you can copy-paste plants cheaply. And raw fuel is less than 10 of the operating cost.

* LCOEs (financially Levelized Costs) are all over the place and highly dependent on plugging holes in production

* Chinese solar panels are state-sponsored to an unspecified amount

* Chinese solar panels are produced using coal electricity, and with lots fossil heat in the process. Models for pricing it in a CO2-free (or levelized with nuclear) are inexistent.

* Even calculating the EROEI (Energy Return Over Energy Invested) is difficult as you need to define an envelope. For example: do you need to account for producing gasoline from solar energy to power the mining truck necessary for your materials. How about powering it with batteries? Which you would need more mining.

* Battery storage needs are of humongous size. What's marginally fine now will probably be humongously expensive at scale. If done on lithium, we need about 180 years of mining for 28 days storage. And that's without accounting for other uses of lithium.

* Having cheap or expensive energy from one source may help you get more of another.

TLDR: At scale, it's a system with 150 variables dependent on each other, including geopolitical game theory hidden variables; and of course the climate crisis affecting the whole thing. But maybe we can trust physics: nuclear forces >> electric forces.

BenoitP··on GE Vernova announces order of 674 wind turbines, providing 2.4 GW of power
From wikipedia [1]

> Typical capacity factors of current wind farms are between 25 and 45%

Offshore being about 10% higher than on land.

Of note is that it is in unconstrained conditions, defined by the weather; when it is needed or when a flexible production can reduce its output, or when flexible consumption can absorb it. As renewables' share increase, this is becoming a problem in some parts of the world: electricity prices can even go negative at times.

[1] https://en.wikipedia.org/wiki/Capacity_factor

BenoitP··on Power over fiber
> efficiently converts optical power to electrical power

Damn, I thought it was just a copper pair going along the fibers. I hope they provide some thick safety glasses with that.

I don't know if it'd come out as laser light, but 600 mW puts it into class 4 (eyes are damaged at a few 100m distance):

https://www.lasersafetyfacts.com/resources/FAA---visible-las...

BenoitP··on Ask HN: Freelancer? Seeking freelancer? (January 2024)
SEEKING WORK | Paris, France | Remote-open

---------------------------

Machine learning engineer, specialized in Explainable AI / ML. 13 years professional coding experience, 24 years personal.

Recent highlights:

* Seamless integration of hand-crafted rules with ML-build rules

* Intuitive, visual data/signal explorer: http://explicable.ai

* Implementation in Spark/Scala of treeinterpreter, currently used in production

* Participation to the FICO-Google Explainable Machine Learning Challenge

---------------------------

Tech: SHAP, RuleFit, Random Forest, Word2Vec, PCA, t-SNE, LSH, Scikit-Learn, Spark, Flink, Weka, Databricks, BigQuery, Hive, Postgres, MySQL, Oracle, AWS, Linux, Maven, Git, Java, Scala, Python, CAML, Elm, Javascript, Typescript, React, Spring, Primefaces, d3.js, Three.js

Résumé/CV: https://www.linkedin.com/in/benoitparis/

Github: https://github.com/benoitparis/

Email: benoit.paris.753@gmail.com

Blog: https://benoit.paris

BenoitP··on Ask HN: Who wants to be hired? (January 2024)
Machine learning engineer, specialized in Explainable AI / ML. 13 years professional coding experience, 24 years personal.

Highlights:

* Seamless integration of hand-crafted rules with ML-build rules

* Intuitive, visual data/signal explorer: http://explicable.ai

* Implementation in Spark/Scala of treeinterpreter, currently used in production

* Participation to the FICO-Google Explainable Machine Learning Challenge

Location: Paris, France

Remote: open to it

Willing to relocate: for the right job, yes

Technologies: SHAP, RuleFit, Random Forest, Word2Vec, PCA, t-SNE, LSH, Scikit-Learn, Spark, Flink, Weka, Databricks, BigQuery, Hive, Postgres, MySQL, Oracle, AWS, Linux, Maven, Git, Java, Scala, Python, CAML, Elm, Javascript, Typescript, React, Spring, Primefaces, d3.js, Three.js

Résumé/CV: https://www.linkedin.com/in/benoitparis/

Github: https://github.com/benoitparis/

Email: benoit.paris.753@gmail.com

Blog: https://benoit.paris

BenoitP··on Do you really need foreign keys?
My rule of thumb has been: enable them strictly in DEV and INT environments, disable in PROD. They can catch schema discrepancies, but can impede ingestion rates.

Also some referential errors are sort of ok in PROD, as long as it's only about not dropping user data; which can be dealt with later on (INT gets reset with PROD user data from a backup each week, it also helps in the restore plan, fk are enabled, errors are caught, then data gets pruned heavily)

If referential integrity is a business-level bug, then of course we should enable them.

BenoitP··on 'e/acc' – Silicon Valley's favorite obscure theory about progress at all costs
I believe it comes at the exact time -and in opposition to- of the masses coming to terms with the climate change challenge.

We have to transition away from fossil fuels, but they are so intertwined with our lives it also means strong degrowth (count the times transporting was necessary to your activity today, or the number of times you touched a plastic or steel item, or just consider half your nitrogen came from an industrial Haber process column)

SV's philosophy is based on growth, and this is anathema. IMHO this is a half-conscious attempt to impose an alternative, or to shed future potential obligations.

It feels like we're in a Greek tragedy: everyone plays the role one was made to honestly play, and that makes the catastrophe inevitable.

← PreviousPage 5 of 27Next →