https://www.nature.com/articles/s41562-025-02259-6
https://www.theguardian.com/money/2019/feb/19/four-day-week-...
1,399 karma · joined February 9, 2025
https://vivis.dev
https://findsubstack.com
https://pythonkoans.substack.com
https://www.nature.com/articles/s41562-025-02259-6
https://www.theguardian.com/money/2019/feb/19/four-day-week-...
Can we start tagging titles in HN with [AI-generated] or something?
I know some people have no problem with it, but it might help others (like me) to steer clear
> uv add pydantic --bounds major
So not really sure what he's complaining about
This was new. I'm surprised that a model specifically designed for security research and gated to professionals is refusing legitimate requests
Gemma4 edge models were promised to be great for agentic use, but have been really disappointing in all my tests. They fail at the most basic tool use scenarios.
Have you run any tool-use benchmarks for Needle, or do you plan to? Would be great if you could add results to the repo if so.
This means you don't have to muck around with supplying the right documentation for each version of each dependency, or worry about hallucinated interfaces (at least with the latest models).
In the past you'd have to dig through a foreign codebase manually to figure out why a documented interface for a dependency is not working as expected, but frontier models automate that quite well.
> We find that weaker models’ degradation originates primarily from content deletion, while frontier models’ degradation is attributable to corruption of content.
I think we largely already knew this. This is why we fudge around with harnesses and temperature etc.
Prompting via text alone is a really bad way to generate images. Ideally you want Canny Control to draw an outline of the image with elements in the exact locations where you want them. It's why comfyui is so great.
The ability to edit images and specify regions in the image for the prompt is a step in the right directions though. ChatGPT and Gemini have this.
Yes, you can do image-> text on existing styles, but something always gets lost in translation.
Midjourney probably has the best baseline, and --sref is a really easy way to differentiate
Check out this paper - https://arxiv.org/abs/2506.13405
- Gemini Nano-1: 46% MMLU, 1.8B
- Gemini Nano-2: 56% MMLU, 3.25B
- Gemma4 E2B: 60.0% MMLU, 2.3B
- Gemma4 E4B: 69.4% MMLU, 4.5B
Sources:
- https://huggingface.co/google/gemma-4-E2B-it
- https://android-developers.googleblog.com/2024/10/gemini-nan...
Would love to know exactly what the latest process is to keep slop out of training data.
I realize that most researchers use AI to assist with writing, but when the topic of your paper is "cognitive surrender", I struggle to take any content in there seriously.
This was exactly the reason why GPT-2 was restricted for general release in 2019.
Check out section 4 - https://cdn.openai.com/GPT_2_August_Report.pdf
It's a newsfeed constructed from 130k substack RSS feeds but limited to the past 24h.
Its helping me discover writers other than just what the algorithm gives me.
A newsfeed for Substack posts from the past 24h. Its helping me discover writers other than just what the algorithm gives me.
I'm building this mostly to scratch my own itch.
A newsfeed for Substack posts from the past 24h. Its helping me discover writers other than just what the algorithm gives me.
They know how to run a good marketing campaign.
Just google "taste is the new moat"
Doesn't deserve to be on the front page.