2,852 karma · joined August 8, 2013
http://explicable.ai
benoit.paris.753@gmail.com
https://www.linkedin.com/in/benoitparis/
Kismet: 3b638cc74046887b5af69269b5c351d53acf4b702871f1c13691369db2917550
Loading (fig IV)
Location Energy (pJ = pico Joules)
L1 64 pJ/Byte
L2 121 pJ/Byte
L3 254 pJ/Byte
RAM 1250 pJ/Byte
Adding integers with SIMD (fig VI)
428 pJ/op, for 8 Byte/op; this means:
53 pJ/byte
So it takes actually 20 times to fetch data from RAM than to add it to something! And most often this is also the source of latency. Generally energy is linear to the distance signal has to travel, and that's the same for latency.That's why successful data structures are sized a tiny bit under L1/L2 sizes! (BTree chunks, ring buffers).
If you've been following hardware, it's all about putting RAM closer to compute at the moment with chiplets at the moment.
[1] https://tu-dresden.de/zih/forschung/ressourcen/dateien/abges...
Source: a friend of mine working at a chip company 20 years ago. They considered disparaging the competition billions of time in tiny 250nm letters. It occurred right after a patent dispute where it was highly likely they had reverse-engineered part of their design. A disdainful message only their competitor would read, to the tune of "we know you're cheating, and fart in your general direction". It was mostly a joke though, they did not go through with it.
----
AMD's one is only in one corner of the chip, so I don't think it applies here.
----
EDIT: this was a long time ago. Seems like filling has evolved quite a lot [1]. Timing, signal integrity, and probably capacitance too are so constrained nowadays you can no longer write a personalized message to your competitors.
[1] https://semiengineering.com/knowledge_centers/materials/fill...
Please consider putting it in online.
Cultural change goes much deeper than changing the CEO. Boeing might not be salvageable on the cheap side. Decapitating the company at levels 1-4 may not even be enough. It may require acknowledging that changing the flight envelope and software-patching it was a mistake.
This is a major can of worms on itself. This means a plane redesign, 6-8 years delay, and probably another round of WTO-unfriendly subsidy.
Maybe the cycle will bring back social shame again, which to me is at a historical low. And we'll again have manners manuals with content copied from 1924. Politeness, saying please and thank you, waiting for adults to finish before interrupting, doing house chores, no shouting, etc.
---------------------------
Machine learning engineer, specialized in Explainable AI / ML. 13 years professional coding experience, 24 years personal.
Recent highlights:
* Seamless integration of hand-crafted rules with ML-build rules
* Intuitive, visual data/signal explorer: http://explicable.ai
* Implementation in Spark/Scala of treeinterpreter, currently used in production
* Participation to the FICO-Google Explainable Machine Learning Challenge
---------------------------
Tech: SHAP, RuleFit, Random Forest, Word2Vec, PCA, t-SNE, LSH, Scikit-Learn, Spark, Flink, Weka, Databricks, BigQuery, Hive, Postgres, MySQL, Oracle, AWS, Linux, Maven, Git, Java, Scala, Python, CAML, Elm, Javascript, Typescript, React, Spring, Primefaces, d3.js, Three.js
Résumé/CV: https://www.linkedin.com/in/benoitparis/
Github: https://github.com/benoitparis/
Email: benoit.paris.753@gmail.com
Blog: https://benoit.paris
Highlights:
* Seamless integration of hand-crafted rules with ML-build rules
* Intuitive, visual data/signal explorer: http://explicable.ai
* Implementation in Spark/Scala of treeinterpreter, currently used in production
* Participation to the FICO-Google Explainable Machine Learning Challenge
Location: Paris, France
Remote: open to it
Willing to relocate: for the right job, yes
Technologies: SHAP, RuleFit, Random Forest, Word2Vec, PCA, t-SNE, LSH, Scikit-Learn, Spark, Flink, Weka, Databricks, BigQuery, Hive, Postgres, MySQL, Oracle, AWS, Linux, Maven, Git, Java, Scala, Python, CAML, Elm, Javascript, Typescript, React, Spring, Primefaces, d3.js, Three.js
Résumé/CV: https://www.linkedin.com/in/benoitparis/
Github: https://github.com/benoitparis/
Email: benoit.paris.753@gmail.com
Blog: https://benoit.paris
SEO is in the structure of the internet now. Original Google was great because there was no incentive yet to buy a domain and blogspam it. Google getting shitty is just a natural instance of Goodhart's law, applied to domains and content.
Now, Google originally was based on PageRank; which based itself on every domain being a unit of authority. These have been compromised and drowned by SEO, but the concept remains valid and we could choose people as units of authority. For example PageRank on scientific papers accurately reproduce Nobel Prizes attributions. A person publishing papers is a solid enough foundation for this unit of authority.
It remains to be organized though. And if we take people as units of authority, it means they'd have to 'cite' or vote for each other. This has social consequences and might not be doable. Are you ready to refuse to cite your boss when he/she ask you to do so? Maybe if the vote is secret and delayed by 5 years?
> Your codebase is old, large and has been working fine without static type checking for years. While Python’s type system is designed to allow gradual adoption of static type checking, the total cost of adding type annotations to a large extant codebase can be prohibitive.
That's a convoluted and quite dishonest way of warning that you're trapped if you don't enforce types from the start; but it does have the merit of being on the list.
It's going to stay because it is useful:
Any operation that you can express with an associative behavior is automatically parallelizeable. And both in Spark and Torch/Jax this means scalable to a cluster, with the code going to the data. This is the unfair advantage of solving bigger problems.
If you were talking about the Hadoop ecosystem, then yes Spark pretty much nailed it and is dominant (no need to have another implementation)
So far GDPR has only yielded symbolic amounts. But the legislative cap is at profit-destroying levels.
And third: renewables need to be associated with its backup like hydro/step or ... batteries which cost a lot. Gas can't be taken in as it's not CO2-free. All that unless training and inference happen when there's the corresponding wind and sun shining. And I'm not seeing that happening right now.
That's interesting.
It's controversial to say that in 2024, but not all opinions have the same value. Some are great, but some are plain dumb. The current corporate right opinion is to praise LLMs as end-all be-all. I've been asked to advise a private banking family office wanting to get into LLMs. For advising their clients' financial decisions. I politely declined. Can there be a worse use case? LLMs are parrots with the brain size of the internet. With thoughts of random origin mixed together randomly. It produces wonderful form, but abysmal analysis.
IMHO as LLMs will begin to be indistinguishable from real users (and internet dogs), there's going to be a resurging need to trace origin to a human; and maybe to also rank their opinions as well. My money is on some form of distributed social proof designating the high priests.
In theory JIT should be higher performance, because it benefits from statistics taken at actual runtime. Given a smart enough compiler. But as a piece of code matures and gets more stable, the envelope of executions is better known and programmers can encode that at compile-time. That's the tradeoff taken by Rust: ask for more proofs from the programmers, and Rust is continuing to pick up speed.
That's also what the Leyden project / condensers [1] is about, if I understand correctly. Pick up proofs and guarantees as early as possible and transform the program. For example by constant-propagating a configuration file taken up during build-time.
Something I've pondered over the years: a programmer's job is not to produce code. It is to produce proofs and guarantees (yet another digression/rant: generating code was never a problem. Before LLMs we could copy-paste code from StackOverflow just fine)
In the end it's only about marginal improvements though. These could be superseded by changes of paradigm like RAM getting some compute capabilities; or programs being split into a myriad of specialized instructions. For example filters, rules and parsing going inside the network card; SQL projections and filters going into the SSD controller; or matrix-multiplication going into integrated GPU/TPU/etc just like now.
[1] https://openjdk.org/projects/leyden/notes/03-toward-condense...
But Java and similar languages extract more freedom-of-operation from the programmer to the runtime: no memory address shenanigans, richer types, and to some extent immutability and sealed chunks of code. All these could be picked up and turned into more performance by the hardware; with some help from the compiler. Sort of like SQL being a 4th-gen language, letting the runtime collect statistics and chose the best course of execution (if you squint at it in the dark with colored glasses)
More recent work about this is to be found on the RISC-V J extension [1], still to be formalized and picked up by the industry. Three features could help dynamic languages:
* Pointer masking: you can fit a lot in the unused higher bits of an address. Some GCs use them to annotate memory (refered-to/visited/unvisited/etc.), but you have to mask them. A hardware assisted mask could help a lot.
* Memory tagging: Helps with security, helps with bounds-checking
* More control over instruction caches
It is sort of stale at the moment, and if you track down the people working on it they've been reassigned to the AI-accelerator craze. But it's going to come back, as Moore's law continues to end and Java's TCO will again be at the top of the bean-counter's stack.
Now they're actually planning them and cutting cost, surprise surprise, the turbines are gas only.
They had a contract to index historical French archives composed of handwritten latin documents in elasticsearch.
Depending of the historical relevance of your documents (read: some academic funds), they may be able to help. Doesn't hurt to contact them:
* Financing and returns are spread over 60 years for nuclear, which makes it hugely dependent on the interest rate. Which we don't know in advance, and also is at risk of irrational public perception (Leading in democracies to early closure of perfectly fine plants. That affects ROI hugely).
* Nuclear too profits from scale: you have to maintain an army of competent engineers. But once done, you can copy-paste plants cheaply. And raw fuel is less than 10 of the operating cost.
* LCOEs (financially Levelized Costs) are all over the place and highly dependent on plugging holes in production
* Chinese solar panels are state-sponsored to an unspecified amount
* Chinese solar panels are produced using coal electricity, and with lots fossil heat in the process. Models for pricing it in a CO2-free (or levelized with nuclear) are inexistent.
* Even calculating the EROEI (Energy Return Over Energy Invested) is difficult as you need to define an envelope. For example: do you need to account for producing gasoline from solar energy to power the mining truck necessary for your materials. How about powering it with batteries? Which you would need more mining.
* Battery storage needs are of humongous size. What's marginally fine now will probably be humongously expensive at scale. If done on lithium, we need about 180 years of mining for 28 days storage. And that's without accounting for other uses of lithium.
* Having cheap or expensive energy from one source may help you get more of another.
TLDR: At scale, it's a system with 150 variables dependent on each other, including geopolitical game theory hidden variables; and of course the climate crisis affecting the whole thing. But maybe we can trust physics: nuclear forces >> electric forces.
> Typical capacity factors of current wind farms are between 25 and 45%
Offshore being about 10% higher than on land.
Of note is that it is in unconstrained conditions, defined by the weather; when it is needed or when a flexible production can reduce its output, or when flexible consumption can absorb it. As renewables' share increase, this is becoming a problem in some parts of the world: electricity prices can even go negative at times.
Damn, I thought it was just a copper pair going along the fibers. I hope they provide some thick safety glasses with that.
I don't know if it'd come out as laser light, but 600 mW puts it into class 4 (eyes are damaged at a few 100m distance):
https://www.lasersafetyfacts.com/resources/FAA---visible-las...
---------------------------
Machine learning engineer, specialized in Explainable AI / ML. 13 years professional coding experience, 24 years personal.
Recent highlights:
* Seamless integration of hand-crafted rules with ML-build rules
* Intuitive, visual data/signal explorer: http://explicable.ai
* Implementation in Spark/Scala of treeinterpreter, currently used in production
* Participation to the FICO-Google Explainable Machine Learning Challenge
---------------------------
Tech: SHAP, RuleFit, Random Forest, Word2Vec, PCA, t-SNE, LSH, Scikit-Learn, Spark, Flink, Weka, Databricks, BigQuery, Hive, Postgres, MySQL, Oracle, AWS, Linux, Maven, Git, Java, Scala, Python, CAML, Elm, Javascript, Typescript, React, Spring, Primefaces, d3.js, Three.js
Résumé/CV: https://www.linkedin.com/in/benoitparis/
Github: https://github.com/benoitparis/
Email: benoit.paris.753@gmail.com
Blog: https://benoit.paris
Highlights:
* Seamless integration of hand-crafted rules with ML-build rules
* Intuitive, visual data/signal explorer: http://explicable.ai
* Implementation in Spark/Scala of treeinterpreter, currently used in production
* Participation to the FICO-Google Explainable Machine Learning Challenge
Location: Paris, France
Remote: open to it
Willing to relocate: for the right job, yes
Technologies: SHAP, RuleFit, Random Forest, Word2Vec, PCA, t-SNE, LSH, Scikit-Learn, Spark, Flink, Weka, Databricks, BigQuery, Hive, Postgres, MySQL, Oracle, AWS, Linux, Maven, Git, Java, Scala, Python, CAML, Elm, Javascript, Typescript, React, Spring, Primefaces, d3.js, Three.js
Résumé/CV: https://www.linkedin.com/in/benoitparis/
Github: https://github.com/benoitparis/
Email: benoit.paris.753@gmail.com
Blog: https://benoit.paris
Also some referential errors are sort of ok in PROD, as long as it's only about not dropping user data; which can be dealt with later on (INT gets reset with PROD user data from a backup each week, it also helps in the restore plan, fk are enabled, errors are caught, then data gets pruned heavily)
If referential integrity is a business-level bug, then of course we should enable them.
We have to transition away from fossil fuels, but they are so intertwined with our lives it also means strong degrowth (count the times transporting was necessary to your activity today, or the number of times you touched a plastic or steel item, or just consider half your nitrogen came from an industrial Haber process column)
SV's philosophy is based on growth, and this is anathema. IMHO this is a half-conscious attempt to impose an alternative, or to shed future potential obligations.
It feels like we're in a Greek tragedy: everyone plays the role one was made to honestly play, and that makes the catastrophe inevitable.