HNHacker News
TopNewBestAskShowJobs

fforflo

3,497 karma · joined August 13, 2015

Data & AI Tools - Postgres - Python

tselai.com

submissionscomments
fforflo··on Why we no longer use LangChain for building our AI agents
I'll bite:

"More artists than engineers": yes and no. I've been working with Pandas and Scikit-learn since 2012, and I haven't even put any "LLM/AI" keywords on my LinkedIn/CV, although I've worked on relevant projects.

I remember collaborating back then with PhD in ML, and at the end of the day, we'd both end up using sklearn or NLTK, and I'd usually be "faster and better" because I could write software faster and better.

The problem is that the only "LLM guy." I could trust with such a description, someone who has co-authored a substantial paper or has hands-on training experience in real big shops.

Everyone else should stand somewhere between artist and engineer: i.e., the LLM work is still greatly artisanal. We'll need something like scikit-learn, but I doubt it will be LangChain or any other tools I see now. You can see their source code and literally watch in the commit history when they discover things an experienced software engineer would do in the first pass. I'm not belittling their business model! I'm focusing solely on the software. I don't think they their investors are naive or anything. And I bet that in 1-2 years, there'll be many "migration projects" being commissioned to move things away from LangChain, and people would have a hard time explaining to management why that 6-month project ended up reducing 5K LOC to 500 LOC.

For the foreseeable future though, I think most projects will have to rely on great software engineers with experience with different LLMs and a solid understanding of how these models work.

It's like the various "databricks certifications" I see around. They may help for some job opportunities but I've never met a great engineer who had one. They're mostly junior ones or experienced code-monkeys (to continue the analogy)

fforflo··on Why we no longer use LangChain for building our AI agents
Ah didn't know that. IIRC I first heard this analogy with regards to Java Spring framework, which had the "longest java class name" somewhere in its JavaDocs. It should have been something like 150+ chars long. You know... AbstractFactoryTemplate... type of thing.
fforflo··on Why we no longer use LangChain for building our AI agents
LLM frameworks like LangChain are causing a java-fication or Python .

Do you want a banana? You should first create the universe and the jungle and use dependency injection to provide every tree one at a time, then create the monkey that will grab and eat the banana.

fforflo··on Koheesio: Nike's Python-based framework to build advanced data-pipelines
It's a library written by some devs who thought it might be useful to others, too, and/or are proud enough to share their work. It's not that they'll publish it in an SEC filing.

For every n-th solution in the market, n-1 existing ones could have been used, but they weren't for (many times) good reasons.

And looking at their Makefile and pyproject.toml I can see that they knew what they wanted.

fforflo··on Python's many command-line utilities
if you want complex subcommands and a truly fluent CLI interface, go with a click. https://click.palletsprojects.com/

else: argparse is more than enough https://docs.python.org/3/library/argparse.html

fforflo··on Python's many command-line utilities
My favorite and probably most useful? `python3 -m venv ./venv` Honestly, forget about conda, poetry, virtualenvwrapper etc and just use that one.
fforflo··on Amber: Programming language compiled to Bash
I don't share the "oh my, why bash, why not English then?" sentiment in the comments.

I've done a bunch of DSLs (CLI with some elaborate syntax really) that compile to bash as a target. It just works for me. Bash is always available; the constructs you can use are there (coreutils are mostly enough for the primitives and xargs for parallelization) It has been great so far for basic cases. Where things get complex (as other's have said) is handling failures, errors and intrinsic cases. That's when you're reminded why people didn't stuck with bash in the first place and we got other scripting languages.

fforflo··on Amber: Programming language compiled to Bash
I had done a SQL-like ->Bash "language," and the reason was that I wanted to have access to CLI tools as-is, in my case, curl programs. Had I chosen Python as a target, I'd have to use requests/urllib, which would be much more verbose.

The same applies if one needs awk, sed, etc... The constructs are just there.

Also the fact that you can just pipe to bash, makes development much easier.

fforflo··on The Time I Convinced the CTO Not to Outsource Our Developers
I agree with your points, but as a side note, I've noticed that French companies have never been too excited about offshoring and remote work in particular.

I've been freelancing remote/hybrid in Europe for years now, and on the recruitment side, I recall always turning down recruiter calls primarily because the company was adamant about 100% onsite. Maybe freelancing is not that popular, either. Is this the case?

fforflo··on SEQUEL: A Structured English Query Language (1974)
Exactly.

Especially around UPDATE [0] and MERGE [1].

A lot of "data engineering" complexity and data pipelines are necessary because of the lack of more-than-superficial knowledge of SQL.

Inexperienced developers talk about DAGs of execution and whatnot, without ever realizing they describe atomic (ACID) operations an RDBMS can do out of the box.

[0] https://www.postgresql.org/docs/16/sql-update.html

[1] https://www.postgresql.org/docs/16/sql-merge.html

fforflo··on Experience with SQLite as a Store of Files and Images
A different approach to backups is an important aspect here. Also, writing-heavy things can get trickier with concurrency and locks. BLOBs usually have a more archive nature. In a scenario where I had a 5-10MB pdf document (BLOB) for each 4KB JSON document, I had to store these in separate db files. When I want to use them within the same query, I would use ATTACH https://www.sqlite.org/lang_attach.html
fforflo··on Dotfiles: Unofficial Guide to Dotfiles on GitHub
I guess you're referring to this ? https://github.com/aspiers/stow/issues/65
fforflo··on [dead]
> There’s so many projects that just seem to call out to the OpenAI api even though they say something silly like “99% local.”

That used to be the case for lots of business-facing products that wanted to capitalize on the hype quickly a year ago but I think the dust has settled (?)

For developer tools, however things like llamafile are pretty much the standard. Not to mention the pain of maintaining multiple keys, different response formats etc.

fforflo··on [dead]
It's the other way around actually: standard jq but with AI capabilities.
fforflo··on [dead]
Not all LLM models are remote. You can do just fine with local ones.
fforflo··on Data Science at the Command Line, 2nd Edition (2021)
Yes, that's extremely important. I've had great success, but replacing Airflow, Luigi, and friends with a cron-ed Makefile target refreshing some database tables (usually Materialized views).

I've then used this tool to visualize the execution graph. https://github.com/lindenb/makefile2graph

The result looks like this: https://tselai.com/data/graph.png

The convenient thing is that each node in the execution graph is in a different environment. Some are shell scripts, a some are Python scripts while others are SQL queries.

fforflo··on Hacking on PostgreSQL Is Hard
That's an ongoing discussion in the Postgres community [0], primarily because Postgres is 100% written in C.

As for the future, I don't know. Not many companies can maintain full forks. I've noticed, however, from my professional experience that businesses are willing to hire freelancers to code specific things. (e.g., I've had a few Postgres gigs to write custom Postgres extensions. 50% C and 50% SQL/PgSQL). But yeah, I guess lots of C projects will become like COBOL ones.

0: https://redmonk.com/jgovernor/2023/10/10/postgres-the-next-g...

fforflo··on Hacking on PostgreSQL Is Hard
Development-wise they're completely different. Extensions are much easier / more isolated. Developing a pg extension is all about compiling some C code into a .so object and dynamically using a templated Makefile to load it into the running Postgres process.

If you want some boilerplate to get you started I've been using this cookie-cutter template to bootstrap lots of extensions https://github.com/Florents-Tselai/cookiecutter-postgres-ext...

Demo walkthrough: https://www.youtube.com/watch?v=zVxY3ZmE5bU

Contributing to core Postgres, it's an entirely different story.

fforflo··on NASA's Voyager 1 Resumes Sending Engineering Updates to Earth
That's what happens when you're freed from "SEO-optimized content". It's also a culture thing I'd probably put under military philosophy. I've worked with ex-military engineers, and you can tell from how they communicate. Writing technical reports and memos is a skill.
fforflo··on Show HN: What Are You Working On?
Unix-inspired so it's entirely text oriented.
fforflo··on Show HN: What Are You Working On?
I'm working on a small Unix-inspired DSL for LLM pipelines.
fforflo··on EU science advisers back call for a 'CERN for AI' to aid research
Unfortunately, EU research money nowadays is free money for private companies. How this works: a 3-year research project with a consortium of academia+industry of 2-3 universities (1M/each) and 4-5 private companies(50k-200k).

With this money, universities do fund PhD students but with a stipend/salary of an average software engineer. In contrast, the "industry partners" do the bare minimum work a couple of days before the quarterly review. And they use 20% of that money to sub-contract their work packages to freelancers (who, btw are paid better than the PhD students) and the other 80% to cover normal OpEx.

And you know... the "work packages" can be something like 5K-10K for a bootstrap/hugo website.

Edit: yes, I over-simplify things a lot but that's the general modus operandi.

fforflo··on Why SQLite Performance Tuning Made Bencher 1200x Faster
Not really. In this context they're a hint to the optimizer to evaluate the CTE result store=materialize it somewhere and use that snapshot when the CTE is queried subsequently.

The alternative means the CTE is copy-pasted essentially and them evaluated.

The closest one can get to materialized views in SQLite is probably with CREATE TABLE AS. Drop and recreate.

fforflo··on An electric new era for Atlas
Do they list Sora as a potential competitor?
fforflo··on An electric new era for Atlas
I have zero ties to the industry. Am I right to assume there's a lot of DoD-driven echo chamber? Material being produced for the big clients and contracts ?
fforflo··on An electric new era for Atlas
What's the best way/resource to get an honest/pragmatic view of where things stand with the "robots market" in general and how much and fast things are really progressing?

I remember seeing prototypes from Toshiba when I was 10 (20 years ago), and every few months, there is a company releasing an "amazing video." its mother company then spins it off like there's no adequate progress, and so on.

fforflo··on Show HN: PostgreSQL index advisor
The convenient thing about this is that it's written in vanilla Pl/PgSQL. It can be tempting to copy the `index_advisor(text)`function in a session and start hard-coding stuff and heuristics :D .

Most meaningful extensions need to be compiled, installed, created dropped.

fforflo··on The Archeologists of Athens (2023)
I just returned from our daily walk around the streets the article mentioned and opened HN.

> Beneath the seductive surface of the present...

Unfortunately, this surface has become less and less seductive, and history is being buried deeper and deeper. Deep below overpriced coffee shops, AirBnBs, and "luxury apartments". There's no continuity whatsoever. Historians of the future will probably split Athenian history into classical, Roman, and touristic eras.

fforflo··on Why doesn't Facebook use Git? [video]
Actually I knew that and I did spot the hallucination, but I was fine with it as long as it answered the core of the question.
fforflo··on Why doesn't Facebook use Git? [video]
20 mins video for this? The battle for attention has reached new levels (and I pay YouTube premium)

I asked ChatGPT:

- Why doesn't Facebook use git? In 5 lines

- Facebook doesn't use Git primarily because of the scale of their codebase and the number of developers working on it. Git, while powerful, can struggle with extremely large repositories and high volumes of concurrent updates. Instead, Facebook has developed its own version control system, called Mercurial, which is optimized for their specific needs, including handling large codebases and providing faster performance for their workflow.

Sounds legit, Lmk if it got this right

← PreviousPage 3 of 5Next →