HNHacker News
TopNewBestAskShowJobs

cl42

3,147 karma · joined June 28, 2012

Reach out to me on Twitter: @wojciech
submissionscomments
cl42··on Scores for Adults Are Dropping on Tests of Basic Skills
I wonder about this regularly and think back to discussions on this topic back in 2020-2021. Are there more recent studies or research on this topic you'd recommend?
cl42··on Scores for adults are dropping on tests of basic skills
This is very interesting. The study it's reporting on has all the numerical data here, which speaks volumes: https://www.oecd.org/en/publications/2024/12/do-adults-have-...
cl42··on Ask HN: What blogs, researchers, 'thinkers' do you follow?
Fantastic list, thank you!
cl42··on Ask HN: What blogs, researchers, 'thinkers' do you follow?
I go back and forth on paying for it. I tend to subscribe for a few months, cancel, and restart.

I imagine there's not many such analysts because the quality of writing is so high but... I wish there were more! :)

cl42··on Ask HN: What blogs, researchers, 'thinkers' do you follow?
What a great list. Thank you!
cl42··on Machiavelli and the Emergence of the Private Study
This is a really interesting question because I think the definition of a home office has changed quite a bit.

In North America, there was a period in the 80s and 90s where the Desktop PC was very much a shared device. You'd have it sitting somewhere like the living room or basement, maybe near the TV, you'd have a landline phone next to it, etc.

I think a lot of families had those, but it's very different from the idea of a "home office" where you have a separate/isolated work room.

cl42··on Machiavelli and the Emergence of the Private Study
When I was in undergrad, I worked with Barry Wellman (one of the early proponents of Social Network Analysis). One of his research projects back in the 1970s-90s was interviewing people about their home offices, and they'd have them take photos of their desk, computer setup, etc. Really cool stuff to see how people decide to focus. I wonder how much it's changed?

This is a nice article going in the other chronological direction!

cl42··on Notes on Anthropic's Computer Use Ability
I really, really like this new product/API offering. Still crashes quite a bit for me and obviously makes mistakes, but shows what's possible.

For the folks who are more savvy on the Docker / Linux front...

1. Did Anthropic have to write its own "control" for the mouse and keyboard? I've tried using `xdotool` and related things in the past and they were very unreliable.

2. I don't want to dismiss the power and innovation going into this model, but...

(a) Why didn't Adept or someone else focused on RPA build this?

(b) How much of this is standard image recognition and fine-tuning a vision model to a screen, versus something more fundamental?

cl42··on Show HN: FinetuneDB – AI fine-tuning platform to create custom LLMs
Main requirement is to programmatically send my chat logs. Not a big deal though, thanks!
cl42··on Show HN: FinetuneDB – AI fine-tuning platform to create custom LLMs
Was looking for a solution like this for a few weeks, and started coding my own yesterday. Thank you for launching! Excited to give it a shot.

Question: when do you expect to release your Python SDK?

cl42··on Ask HN: How have you integrated LLMs in your development workflow?
100%. I usually use it for prototyping new features.

Two examples...

1. Recently wanted to build a Chrome plugin and never built one before. Used o1-preview to build it all for me.

2. Wanted to build a visualization of the world with color-coded maps using D3. Again, hadn't used D3 much in the past... Claude basically wrote all the code for me and then I just had to make edits to fit my site/template.

cl42··on Ask HN: Former gifted children with hard lives, how did you turn out?
Those affected by this post might appreciate this book: "The Drama of the Gifted Child: The Search for the True Self" (https://www.amazon.ca/Drama-Gifted-Child-Search-Third/dp/046...)
cl42··on Ask HN: Must-Read Books for Startups?
I like this book, but am curious if you've seen any third-party case studies written about using her framework?

I tried to use it a while back but found it a bit bureaucratic for a startup. Felt better-placed for large orgs with distinct product management teams.

cl42··on Ask HN: What are you working on (August 2024)?
Thanks for asking! Yes, but we're actively addressing them.

We do a few things under the hood to make hallucinations significantly less likely. First, we make sure every single statement made by the LLM has a fact ID associated with it... Then we've fine-tuned "verification" LLMs that review all statements to make sure that assertions being made are backed up by facts, and that the facts are actually aligned with the assertion.

It's still possible for the LLM to hallucinate in this process, but the likelihood is much lower.

cl42··on Ask HN: What are you working on (August 2024)?
I've been passionate about superforecasting and AI for a long time, so decided to try and use LLMs to structure data about the world -- commodities, politics, etc. It's now a startup we're trying to get off the ground: https://emergingtrajectories.com/

Most of my days are spent reading the news and working on LLMs, which has been a blast. As an example, here's a dashboard that tracks major supply and demand shocks to various commodities around the world: https://emergingtrajectories.com/c/commodities

cl42··on Ask HN: How many users does the average YC founder cold call/talk to?
Went through S12, and also talk to a lot of YC founders. A few thoughts:

1. Quality over quantity. "Talk to users" works best when you have a well-defined market or ICP. Talking to 10 users in your ultra-specific niche is way better than talking to 100 users across multiple niches.

2. Your script matters a lot. Asking leading questions will get you results you can't trust. I think The Mom Test (https://www.amazon.com/Mom-Test-customers-business-everyone/...) is a great intro on this topic.

3. Talk to enough users that you start being able to predict their answers. If you are running interviews and still getting new/surprising answers to your questions, then it means you either haven't spoken to enough people or have a poorly defined ICP... If you need #s, I generally find that after 10 interviews in a very focused ICP, you should start seeing patterns.

Finally, there is an exception to every rule. Your specific market might need more interviews, or you might have such a good insight that you skip formal interviewing all together.

cl42··on The Pitfalls of Defining Hallucination [pdf]
I didn't realize the above was a temporary URL. Here is the Arxiv link: https://arxiv.org/abs/2401.07897
cl42··on macOS Sequoia adds weekly permission promptfor screenshot, screen recording apps
> Having to click through these confirmation nags every week, for every such utility you use, is not a little thing at all. It’s the sort of thing companies do when decisions like this are made by people looking to cover their asses, not make insanely great products.

... or alternatively, when agreeing to using such an app is such a huge privacy nightmare that it might just be safer (for the user) to ask the user to opt-in every week, especially if the company which runs the OS is known for promoting a privacy-friendly brand.

cl42··on Open Source Farming Robot
This is incredibly helpful, thank you!

If you have resources I can read or learn from about all this, please share them. You've clearly got wisdom in this space!

cl42··on Open Source Farming Robot
A lot of people are criticizing this product. Does anyone know what "best in class" small-scale farming or gardening projects are? Very curious! Also any community recommendations would be great.
cl42··on Ask HN: Will peer to peer services overtake centralised corporations?
From a historical (and academic) perspective, most decentralized services seem to tend towards monopolization.

Telephones, radio broadcasting, and the Internet are all examples of once-decentralized, democratizing forces that were eventually centralized from a corporate control perspective.

I don't know if this will change in the future; it'd require either a very active legal agenda or incredibly engaged citizens/consumers.

"The Master Switch"[1] is a great book on the above.

[1] https://www.amazon.com/Master-Switch-Rise-Information-Empire...

cl42··on Solving the out-of-context chunk problem for RAG
100% agree with you. I've built a # of RAG systems and find that simple Q&A-style use cases actually do fine with traditional chunking approaches.

... and then you have situations where people ask complex questions with multiple logical steps, or knowledge gathering requirements, and using some sort of hierarchical RAG strategy works better.

I think a lot of solutions (including this post) abstract to building knowledge graphs of some sort... But knowledge graphs still require an ontology associated to the problem you're solving and will fail outside of those domains.

cl42··on PyCon 2024 Videos Released
Some interesting videos/talks:

1. https://www.youtube.com/watch?v=LauNJNKECOM --> reviews GT, a library for generating beautiful tables (https://posit.co/blog/introducing-great-tables-for-python-v0...).

2. https://www.youtube.com/watch?v=gRS8uu3GGpk --> how to communicate ideas with diagrams. Goes into PyFlow, which I wasn't aware of before. Neat!

3. https://www.youtube.com/watch?v=zHm-f9E7aIY --> JupyRest, deploying web services via notebooks.

cl42··on Ask HN: How do you manage files and backups as an individual?
One more here: Ask HN: How do you do personal backups in 2023? https://news.ycombinator.com/item?id=38576809
cl42··on Investors Pour $27.1B into A.I. Startups, Defying a Downturn
I really dislike this sort of headline because it misses a core part of the analysis between seed-stage, versus late-stage startups.

NYT says the $27.1B represents ~50% of investments in startups ($56B in total).

But the $27.1B includes $1B for CoreWeave, $1B for Scale AI, and $6B for xAI. Elon Musk's company is an outlier IMHO, and the other two might as well be public companies soon. Certainly late stage enough not to be bellwethers for early stage VC investing.

Remove the $8B above, and AI startups got about 34% of all startup funding. Is that a lot? I don't know, but all of a sudden it sounds more reasonable.

cl42··on The Operational Wargame Series: The best game not in stores now (2021)
If you're curious, here's a 165-page Taiwan war game run across multiple scenarios and events: https://www.csis.org/analysis/first-battle-next-war-wargamin...
cl42··on The Operational Wargame Series: The best game not in stores now (2021)
I believe in this case that's very much the goal -- less about who wins, and more about the options/tactics debated to inform actual military battle prep.
cl42··on Ask HN: Is anyone building automated long-term investing software?
I'm using LLMs to basically build "junior analysts" that monitor very niche types of companies -- think, junior mining companies, or very specific commodities futures... A lot of these spaces have tons of terrible companies and there's a lot of noise, so if you use a framework that is concrete enough, you can have LLM agents do various types of research for you, fill in the framework, and sift through the noise for you.

Case in point, my framework for mining companies is here: https://emergingtrajectories.com/a/pub/mining_company_risk_f... You can see the scores here: https://emergingtrajectories.com/c/copper_mining_companies

"Long term" -- we'll see, I expect to hold positions for 12-24 months.

For those interested, my work above is influenced by two important books: "You Can Be a Stock Market Genius Even if You're Not Too Smart" by Joel Greenblatt and "Superforecasting: The Art and Science of Prediction" by Philip Tetlock. The idea from Joel's writing is to look for less liquid or less popular asset classes (or ones that structurally can't be invested in by the pros who are smarter/better-resourced than you), and Tetlock really drills process and research for long-term forecasting.

cl42··on Researchers describe how to tell if ChatGPT is confabulating
Thank you! This is so helpful.

It's also interesting to see what temperature value they use (1.0, 0.1 in some cases?)... I have a feeling using the actual raw probability estimates (if available) would provide a lot of information without having to rerun the LLM or sample quite as heavily.

cl42··on Ask HN: How do you get people to try out your product?
1. "Show, don't tell." If you have to explain in detail your value prop and rationally convince a potential user that they will be better off using your product, then you might be wrong about how good the product is. If you're truly solving a problem for them, you should be able to show them quickly how you automate a process or generate a helpful result.

2. Ask for feedback, and specifically, try to see if the problem you are solving is truly relevant to them. Most users have dozens of "problems" they have every day, and they choose NOT to solve most of those problems because they have better things to do. Are you sure you're not solving a problem so low on their list of priorities that they simply don't care?

3. Go to where your users are -- conferences, events, web forums, whatever. If you validated #1 and #2, then showing them or presenting to them will get them excited and will get you users.

I find reaching out to people to get feedback via LinkedIn (I do enterprise sales) is a great way to validate a problem, and only then worry about scaling or getting them to try.

Books that might be helpful: [1] "The Mom Test", for interviewing users, and [2] "Competing Against Luck", which introduces 'jobs to be done' and talks about how your biggest competition isn't another product, but users deciding to do NOTHING.

← PreviousPage 3 of 18Next →