HNHacker News
TopNewBestAskShowJobs

BenoitP

2,852 karma · joined August 8, 2013

https://benoit.paris

http://explicable.ai

benoit.paris.753@gmail.com

https://www.linkedin.com/in/benoitparis/

Kismet: 3b638cc74046887b5af69269b5c351d53acf4b702871f1c13691369db2917550

submissionscomments
BenoitP··on Ask HN: How Can I Make My Front End React to Database Changes in Real-Time?
I did something akin to that with the following monstrosity of DevOps:

- set up Kafka

- set up Flink

- CDC on your DB, into Kafka

- have your server maintain a websocket to clients, read Kafka from it, and dispatch messages according to a field from Kafka

- put your reaction logic into Flink, insert into the Kafka topic read by the server

I even went to opening JIRA tickets in Apache Flink, so that you may choose the Kafa topic to insert to dynamically. In case you have multiple servers to dispatch messages to.

It was very clunky, and a pain to operate. However it was short-lived and is no longer a concern of mine. What I got away from that: adding a time to data serving is just multiplying the complexity of data bugs you can have.

BenoitP··on Ten years of improvements in PostgreSQL's optimizer
It describes all the way the SQL could be executed, then choses the faster plan. For example: if you're looking for the user row with user_id xx, do you read the full table then filter it (you have to look at all the rows)? Or do you use the dedicated data structure to do so (an index will enable to do it in the logarithm of the number of rows)?

A lot more can be done: choosing the join order, choosing the join strategies, pushing the filter predicates at the source, etc. That's the vast topic of SQL optimization.

BenoitP··on What makes concurrency so hard?
According to the framework the author uses, how would you describe these in terms of state space management (apart from immutability, which the author dealt with)?
BenoitP··on Stacksort (2013)
This semi-old internet artifact, to be viewed in light of today's LLM training data. Once again XKCD was not much off the mark.
BenoitP··on Ask HN: How to Transition from Software Engineer to AI/ML Engineer
I did a similar transition as a freelance, it was not a simple task. I started 7 years ago.

First, I got lucky to be doing data close to where a business case allowed it. Then had to fight my client so that I could _improve_ his best selling marketing campaign. I had to resort to implement explainable ML to convince him the back box was taking the same decisions, and only then adding long tail signals. At the time SHAP did not exist, but its precursor did (interpretableTree by Saabas, I ported it from Python to Scala/Spark because Python scared the client)

This was one-shot. Back to data integration after that.

I participated in an ML challenge, nights and weekends. And developed ideas from there, and personal techniques that I showcased here: http://explicable.ai . Although I didn't make any money from it (I tried, but while people like nice pixels, they don't need them; and I interest mostly engineers, who won't/can't buy it but will definitely want to do the same for themselves)

But it did act as a great portfolio piece. And landed me my first full ML project. It was a very small fixed time contract and ended recently, but I had a lot of fun doing it.

Anyway, this doesn't garantee you the greenest grass. For example, I'm currently looking for a salaried position for personal reasons and it's still tough. Btw if anyone wants to hire an MLE in Paris or remote worldwide, here is my resume: https://benoit.paris/CV_Benoit_Paris_EN.pdf

Now for your question: MOOCs can help a lot, and it's definitely something to put on your CV. Lots of good resources on YouTube as well. Karpathy's stuff is awesome for intuition (see for example LSTMs https://karpathy.github.io/2015/05/21/rnn-effectiveness/). You may also want to widen your focus to breakthroughs in other domains, reading papers and code as they come (UMAP, Nerfs/GaussianSplats). https://paperswithcode.com/ is awesome for that. Also be sure to continue doing projects you can show a prospective employer in parallel.

TL;DR: Keep at it, but make sure you produce something, even if it's just nice pixels.

BenoitP··on Visualizing Attention, a Transformer's Heart [video]
3Blue1Brown is really a godsend. I had been trying and failed to get an intuition for the self-attention mechanisms, and also why you want to stack you transformer layers. I went through a lot of material (Key-Query-Value interpretations, SVD of these, etc.) but his comment in chapter 5 video [1] made it click for me: "But you should think of the primary goal of this network that it (the word embedding) flows through as being to enable each one of those vectors to soak up a meaning that's much more rich and specific than what mere individual words could represent".

Here is my intuition:

Transformers are merely a Convnet for token sequences. In the sense that information is first local and each data means something wrt its immediate neighbors. You extract data locally, and then you dezoom to and redo exactly the same at a higher level. Convnet have been shown to reproduce edge detecting kernels first (vertical edge), then assemble them into basic shapes (vertical+oblique: corner), then bigger shapes (something round), then link shapes together (round+round separated by a horizontal shape with rectangles on top), and then it's a car! [2]

Transformers do the same. They assemble meaning among tokens ("ju" + "mps" : a conjugated verb), then assemble these at a higher level ("cat jumps" -> noun+verb), then it's a sentence "the cat jumps over the wall", then a paragraph, then a chapter, then a book. All differentiable so any deviation from something coherent can tweak its knobs higher up the chain at any level of detail.

Convnets just relate to 2D data topology, Transformers to long 1D sequences. This begs the question of 3D data btw: will Nerfs/Gaussian-Splats have their hierarchical-dezoom moment? (have they already done so?)

So, why haven't LLMs appeared before? RNN and LSTM modelled sequences accurately earlier, but these are not "BigData". They are small, by design, and must forget the long tail of residual meaning. Transformers are a hardware-friendly way of keeping all the joint-interactions across tokens in large windows. Because sometimes text needs it: you mention a cat at the beginning, and 3 chapters later you can refer to it without saying "cat". Same for code: "import java.sql.Date" will have a pretty important meaning 5000 tokens down the line.

LLMs are a BigData instance all over again. Dumb algorithms that scale better do better than complex ones, just out of the sheer volume of ingested examples. That's why -to my knowledge- you still have logistic regression at the core of placing ads. You don't model 1 person clicking one thing, you just let 100 teach you how they do it. You just crush the problem with data. It's also been called "The Bitter Lesson" [3]. And LLMs, as Moore's law progressed, were the first structure to ingest a sizeable portion of the internet.

[1] https://youtu.be/wjZofJX0v4M?si=Zdz4sQAvch5B-QA1&t=1177

[2] https://www.cs.cmu.edu/~epxing/Class/10708-19/notes/lecture-...

[3] http://www.incompleteideas.net/IncIdeas/BitterLesson.html

BenoitP··on Tesla Cancels Mass-Market $25,000 Car, Musk Says This Is a Lie
Elon probably said on both occurrences that robotaxis and the 25k$ cars were priority #1.

Anyway it doesn't hurt -even if this is a lie- to say the cheap cars are coming. It sucks financial derisking off of competitors. To competition, Tesla doing a cheap car means no margin for a decade, if ever. Sometimes a little lie can go a long way for moat building.

BenoitP··on Intel discloses $7B operating loss for chip-making unit
Yup. This is more of a signal on 1) experience with new EUV processes and 2) cost of labor. That's not taking into account the expensive EUV amortization.

A hard truth I've heard in CEO circles is that western middle classes are falling down in terms of purchasing-power; And they will continue to do so until they join Chinese middle classes levels of revenue.

BenoitP··on Microsoft posted on the FFMPEG tracker that their issue is 'high priority'
Hi, this is Elon Musk himself. Please refrain from outing publicly my relaxing bug-fixing hobby. Thank you.
BenoitP··on Ask HN: Freelancer? Seeking freelancer? (April 2024)
SEEKING WORK | Paris, France | Remote-open

---------------------------

Machine learning engineer, specialized in Explainable AI / ML. 13 years professional coding experience, 24 years personal.

Recent highlights:

* Seamless integration of hand-crafted rules with ML-build rules

* Intuitive, visual data/signal explorer: http://explicable.ai

* Implementation in Spark/Scala of treeinterpreter, currently used in production

* Participation to the FICO-Google Explainable Machine Learning Challenge

---------------------------

Tech: SHAP, RuleFit, Random Forest, Word2Vec, PCA, t-SNE, LSH, Scikit-Learn, Spark, Flink, Weka, Databricks, BigQuery, Hive, Postgres, MySQL, Oracle, AWS, Linux, Maven, Git, Java, Scala, Python, CAML, Elm, Javascript, Typescript, React, Spring, Primefaces, d3.js, Three.js

Résumé/CV: https://www.linkedin.com/in/benoitparis/

Github: https://github.com/benoitparis/

Email: benoit.paris.753@gmail.com

Blog: https://benoit.paris

BenoitP··on Ask HN: Who wants to be hired? (April 2024)
Machine learning engineer, specialized in Explainable AI / ML. 13 years professional coding experience, 24 years personal.

Highlights:

* Seamless integration of hand-crafted rules with ML-build rules

* Intuitive, visual data/signal explorer: http://explicable.ai

* Implementation in Spark/Scala of treeinterpreter, currently used in production

* Participation to the FICO-Google Explainable Machine Learning Challenge

Location: Paris, France

Remote: open to it

Willing to relocate: for the right job, yes

Technologies: SHAP, RuleFit, Random Forest, Word2Vec, PCA, t-SNE, LSH, Scikit-Learn, Spark, Flink, Weka, Databricks, BigQuery, Hive, Postgres, MySQL, Oracle, AWS, Linux, Maven, Git, Java, Scala, Python, CAML, Elm, Javascript, Typescript, React, Spring, Primefaces, d3.js, Three.js

Résumé/CV: https://www.linkedin.com/in/benoitparis/

Github: https://github.com/benoitparis/

Email: benoit.paris.753@gmail.com

Blog: https://benoit.paris

BenoitP··on DBOS Operating System
Damn, Stonebreaker, Zaharia. Seems like the big guns are out for the new data model in town.

This tech feels like fancy RDBMS triggers backed (and made unique) by a distributed log. Looks very much like the same tech as Flink's statefun and Restate.dev.

Just like SSD gave us LSM trees and RocksDB, it seems like this is the higher level abstraction that Kafka enables. It's going to be interesting to see how it pans out.

BenoitP··on Starship's Third Flight Test [video]
I guess some risk analysis can be made with the plane's and Starship's cross-section.

Let say they are cubes of 30m each. The expected area where they might both be present to be a 5km square. That's a 0.000025 chance of collision at most; and I suppose the plane is away from the center of the expected Starship Gaussian. I'd personally ride in that plane and risk that, even to just to get a glimpse of the reentry.

BenoitP··on Starship's Third Flight Test [video]
FR24 shows a Dassault Falcon 900EX doing circles around the expected Starship re-entry location (Perth to Perth flight plan):

https://www.flightradar24.com/MXJ/345b8f09

BenoitP··on Starship's Third Flight Test [video]
> 26000 km/h

> 27,900 kph

Seems like they want to test the limit. Same speeds as LEO, but guaranteed to come down.

BenoitP··on Starship's Third Flight Test [video]
It is, how to reproduce:

Twitter stream page, F12, network tab, look for m3u8 file, right click, copy url, open in VLC

BenoitP··on Starship's Third Flight Test [video]
Updated, Thanks!
BenoitP··on Starship's Third Flight Test [video]
When: 8:25 AM CT

Launch window: 7:00 AM CT - 8:50 AM CT

--- Updates:

(future)T+40: Starship relight and entry

T+12: Elevator music engaged, please stay tuned for T+40

T+11: Payload door testing

T+8: Upper stage SECO, nominal orbit insertion

T+7: (mine) KSP moment for booster reentry, instabilities. Signal cut off because of exhaust conducts electricity and absorbs RF. Status unknown

T-11: Still no blockers. Watching winds, may have hold at T-40s.

T-30: Broadcast started

T-60: (SpaceX Twitter) The Starship team is go for prop load but keeping an eye on winds, now targeting 8:25 a.m. CT for liftoff

T-65: (SpaceX Twitter) Shifting T-0 a few more minutes to give boats time to clear the keep out area, now targeting 8:10 a.m. CT

T-65: (SpaceX Twitter) New liftoff time is 8:02 a.m. CT, team is clearing a few boats from the keep out area in the Gulf of Mexico

T-45: No blockers

T-90: (SpaceX Twitter) Weather is 70% favorable for today’s third integrated flight test of Starship. The live webcast will begin ~30 minutes before liftoff

---- Streams:

High Quality VLC: Open VLC, Media, Open Network Stream, paste following, Play:

(higher quality) https://prod-ec-us-west-2.video.pscp.tv/Transcoding/v1/hls/g...

(lower latency) https://prod-ec-us-west-2.video.pscp.tv/Transcoding/v1/hls/g...

NASASpaceflight: https://www.youtube.com/watch?v=RrxCYzixV3s

Spaceflight Now: https://www.youtube.com/watch?v=EfnkZFtHPmM

Everyday Astronaut: https://www.youtube.com/watch?v=ixZpBOxMopc

LabPadre Space: https://www.youtube.com/watch?v=LMyXho_YCK8

(FR) Techniques Spatiales: https://www.youtube.com/watch?v=BRXfWLVMEQ8

---- Mission profile:

https://www.spacex.com/launches/mission/?missionId=starship-...

---- Links:

https://twitter.com/SpaceX

https://twitter.com/elonmusk

https://old.reddit.com/r/spacex/comments/1bb8scf/rspacex_int...

BenoitP··on Marcel Grossmann and his contribution to the general theory of relativity
Well the comment by pnin made it about GR, when the parent by boringuser2 was about the general contributions to relativity, both SR and GR. I believe "general theory of relativity" to be different than "general relativity" here, the former encompassing both SR and GR. Maybe that's where our misunderstanding comes from. Also, if you believe that GR has absolutely nothing to do with SR, then no Lorentz has no relevance to the genesis of General Relativity.

Anyway I don't believe we're having a discussion worth having here.

BenoitP··on Marcel Grossmann and his contribution to the general theory of relativity
Which is the source for General Relativity, which suffered too from an attribution dispute with Hilbert's works.

My point being: science doesn't happen in a vacuum, and is not totally ordered

BenoitP··on Marcel Grossmann and his contribution to the general theory of relativity
> Neither Poincare nor Lorentz are relevant to the genesis of General Relativity

Well, that's just plain wrong. From the horse's mouth:

> As we know, this is connected with the relativity of the concepts of "simultaneity" and "shape of moving bodies." To fill this gap, I introduced the principle of the constancy of the velocity of light, which I borrowed from H. A. Lorentz's theory of the stationary luminiferous ether, and which, like the principle of relativity, contains a physical assumption that seemed to be justified only by the relevant experiments

More from here:

https://en.wikipedia.org/wiki/Relativity_priority_dispute

BenoitP··on Marcel Grossmann and his contribution to the general theory of relativity
Relativity was ripe for discovery. And a lot of other scientists came very close:

https://en.wikipedia.org/wiki/Relativity_priority_dispute

But that's just how science work, and no ill will would have been employed by any participant; just as their fanclub want to pit them against one another.

BenoitP··on Marcel Grossmann and his contribution to the general theory of relativity
Plagiarized is too strong of a word. Poincaré too based his work on Lorentz's. And both with Einstein he sort of derived E = m c * 2, independently and earlier. But Einstein's publication was more complete.

Science is not totally ordered, the same invention can occur at two different places from the same shoulders of the same giant. Science is just partially ordered.

BenoitP··on Biden proposes 30% tax on crypto mining
> It's absurd to tax something that can so easily migrate to a different country

It creates an incentive in that country to tax it. If you have followed EU's Carbon Border Adjustment Mechanism, it is exactly how the EU intends to export its legislation. It says: your imports are going to be taxed on their CO2 emissions, unless you implement a tax of your own on their CO2 emissions. The exporting country now has every incentive to tax it and remove the tariff.

BenoitP··on Bitcoin rally pushes BlackRock ETF over $10B in record time
Uncorrelated is the key word IMHO. And I wouldn't say it is liquid, but that's a good thing here IMHO too.

A lot of asset management is for pensions. This means that long term stability is quite important, and when designing portfolios negative or neutral correlations are highly desired (square deviation sigma goes down with weighted sum of correlation coefficients here [1]).

One one hand all companies rely on energy, electricity and gas and oil. To transform matter, to transport goods. Despite being a few percent GDP, energy is a multiplier/enabler of economic activity. Electricity's price is also being defined at the margin by peaker gas plants. And gas price is also highly linked to oil.

On the other hand, Bitcoins being mostly 'hodled' their price would be greatly defined by the block rewards. Which cost of is function of electrical scarcity. The highest the electricity price, the more it cost to produce the same amount of Bitcoin as per the consensus hash rate adjustments.

It's not perfect analysis, and I have not crunched the numbers; but I would not be surprised for Bitcoin to have negative or neutral correlations to a lot of assets because of energy.

[1] https://en.wikipedia.org/wiki/Modern_portfolio_theory

BenoitP··on The new croissant taking Paris by storm
No

The half-life of Tiktok trends is about a few days. I assure you this diabetes hazard nothingburger is not being talked about in Paris; and is going to join the countless other 'croissant hamburgers' 'inventions' shortly. Why is this on HN?

BenoitP··on This week, xAI will open source Grok
Also a company building an AI accelerator [1]

I guess Gro(q|k) is the new Spar(k|c)[1][2][3][4][5][6][7][8]

----

EDIT: I originally thought xAI was referring to XAI [9]

[1] https://groq.com/

[2] https://spark.apache.org/

[3] https://en.wikipedia.org/wiki/SPARC

[4] https://en.wikipedia.org/wiki/SPARC_(tokamak)

[5] https://en.wikipedia.org/wiki/Adobe_Express

[6] https://sparkjava.com/

[7] https://en.wikipedia.org/wiki/Spark

[8] https://en.wikipedia.org/wiki/SPARC_(disambiguation)

[9] https://www.darpa.mil/program/explainable-artificial-intelli...

BenoitP··on 'We had to educate Oracle about our contract,' CIO says after Big Red audit
> Earlier last year, it made changes to the Oracle Java SE subscription model, basing it on a per-employee metric many said would increase costs for users.

Would running Java under Corretto 21 change this?

BenoitP··on Ask HN: Who wants to be hired? (March 2024)
Machine learning engineer, specialized in Explainable AI / ML. 13 years professional coding experience, 24 years personal.

Highlights:

* Seamless integration of hand-crafted rules with ML-build rules

* Intuitive, visual data/signal explorer: http://explicable.ai

* Implementation in Spark/Scala of treeinterpreter, currently used in production

* Participation to the FICO-Google Explainable Machine Learning Challenge

Location: Paris, France

Remote: open to it

Willing to relocate: for the right job, yes

Technologies: SHAP, RuleFit, Random Forest, Word2Vec, PCA, t-SNE, LSH, Scikit-Learn, Spark, Flink, Weka, Databricks, BigQuery, Hive, Postgres, MySQL, Oracle, AWS, Linux, Maven, Git, Java, Scala, Python, CAML, Elm, Javascript, Typescript, React, Spring, Primefaces, d3.js, Three.js

Résumé/CV: https://www.linkedin.com/in/benoitparis/

Github: https://github.com/benoitparis/

Email: benoit.paris.753@gmail.com

Blog: https://benoit.paris

BenoitP··on Ask HN: Freelancer? Seeking freelancer? (March 2024)
SEEKING WORK | Paris, France | Remote-open

---------------------------

Machine learning engineer, specialized in Explainable AI / ML. 13 years professional coding experience, 24 years personal.

Recent highlights:

* Seamless integration of hand-crafted rules with ML-build rules

* Intuitive, visual data/signal explorer: http://explicable.ai

* Implementation in Spark/Scala of treeinterpreter, currently used in production

* Participation to the FICO-Google Explainable Machine Learning Challenge

---------------------------

Tech: SHAP, RuleFit, Random Forest, Word2Vec, PCA, t-SNE, LSH, Scikit-Learn, Spark, Flink, Weka, Databricks, BigQuery, Hive, Postgres, MySQL, Oracle, AWS, Linux, Maven, Git, Java, Scala, Python, CAML, Elm, Javascript, Typescript, React, Spring, Primefaces, d3.js, Three.js

Résumé/CV: https://www.linkedin.com/in/benoitparis/

Github: https://github.com/benoitparis/

Email: benoit.paris.753@gmail.com

Blog: https://benoit.paris

← PreviousPage 4 of 27Next →