HNHacker News
TopNewBestAskShowJobs

letitgo12345

416 karma · joined December 26, 2014

submissionscomments
letitgo12345··on AI researchers debate how close we are to recursive self-improvement
Models that are strong enough to improve themselves without a human (I.e. ai researcher) in the loop
letitgo12345··on US seeks cheaper hunter-killer drones after Iran destroys $1B worth of Reapers
Well there's long term impact but yea that doesn't create enough political pressure to make process efficient
letitgo12345··on Domain expertise has always been the real moat
Long codex sessions lead to a lot of cached token hits, esp when you resume them after a few hours.
letitgo12345··on Small models also found the vulnerabilities that Mythos found
Can't you execute the bug to see if the vulnerability is real? So you have a perfect filter. Maybe Mythos decided w/o executing but we don't know that.
letitgo12345··on After outages, Amazon to make senior engineers sign off on AI-assisted changes
Worth noting that this is when they used Amazon's own AI product, not when using Claude Code or Codex.
letitgo12345··on [dead]
Seems the same tbh
letitgo12345··on Gemini Embedding: Powering RAG and context engineering
LLMs can use search engines as a tool. One possibility is Google embeds the search query through these embeddings and does retrieval using them and then the retrieved result is pasted into the model's chain of thought (which..unless they have an external memory module in their model, is basically the model's only working memory).
letitgo12345··on AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms
Most straightforward would be to ask the model to generate different evaluation metrics (which they already seem to do) and use each one as one of the dimensions
letitgo12345··on AI cracks superbug problem in two days that took scientists years
Or the humans did think of it and were actively proceeding to test that hypothesis
letitgo12345··on Humanity's Last Exam
I think the idea for this is anything that can be set in a literal exam for humans. So anything that would take the best human in that topic in the world say more than an hour to complete is out.

Also IIRC 42% of the questions are math related, not memorization of knowledge.

letitgo12345··on The AI Bubble Is Bursting
I'm the real world, judges I know are using it to do case summaries that used to take weeks, Goldman is using it to do 95% of IPO filings work and I personally am using O1 pro to write a ton of code.

AI's biggest use cases are for doing actual work, not necessarily replacing regular interactions with your mobile or entertainment devices

letitgo12345··on AlphaProof's Greatest Hits
Think more is made of this asterix than necessary. Quite possible adding 10x more GPUs would have allowed it to solve it in the time limit.
letitgo12345··on Notes on Guyana
They found tons of oil
letitgo12345··on Mira Exits OpenAI
Maybe it is but it's not the only company that is
letitgo12345··on Backlash over Amazon's return to office comes as workers demand higher wages
It's an excellent way of getting your best people who have options to quit while the worst ones who don't are forced to stick around
letitgo12345··on [dead]
This is a site setup by ppl unaffiliated with OAI it seems and has wrong claims -- ex O1 doesn't solve 83% of IMO problems -- it solves 83% of AIME problems which are significantly easier.
letitgo12345··on AlphaProteo generates novel proteins for biology and health research
One question is how specific the binding is -- what's the level of off-target effects, etc.
letitgo12345··on Ilya Sutskever's SSI Inc raises $1B
This might be the largest seed round in history (note that 1B is the cash raised, not the valuation). You think that's an indication of the hype dissipating?
letitgo12345··on Former Google CEO Eric Schmidt's Leaked Stanford Talk
Tensorflow losing has nothing to do with Google getting bored -- it's vice versa.

Tensorflow is a symbolic framework, which is less intuitive to work with for most people than the Pytorch. Not to mention the errors Tensorflow generates are more annoying to debug (again more an issue with the fact that it's symbolic than any lack of effort on part of Google)

Google tried to fix it by introducing an eager mode in Tensorflow but by then it was too late.

letitgo12345··on The AI Scientist: Towards Automated Open-Ended Scientific Discovery
Feels like the next generation of models could truly start replacing lower level ML and software engineers
letitgo12345··on "Jeff Bezos and Amazon tried to imprison my husband"
Looks like her argument is that code of conduct is not legally enforceable and that Amazon itself has argued it cannot be used by employees to sue Amazon. Hard to feel sorry for Amazon here for me despite Amazon seemingly being morally in the right in this case
letitgo12345··on Claude 3.5 Sonnet
Sounds like an excuse tbh. Esp when other companies are pushing ahead beyond OAI and open source is close to rivaling them
letitgo12345··on Gemini AI
GPT-4 is also rumored to have consumed 5x less compute to train
letitgo12345··on OpenAI's Murati Aims to Re-Hire Altman, Brockman After Exits
Or even possible. They won't have twitter api access that they did until recently. And they certainly won't have the gigantic dataset ChatGPT is collecting.
letitgo12345··on OpenAI board in discussions with Sam Altman to return as CEO
Why then has no one come close to replicating GPT-4 after 8 months of it being around?
letitgo12345··on OpenAI board in discussions with Sam Altman to return as CEO
Otoh Ilya wasn't a main contributor for GPT-4 as per the list of contributions. gdb was.
letitgo12345··on OpenAI offers $10M pay packages to poach Google researchers
Publish high impact papers or develop a social circle of influential folks and impress them
letitgo12345··on OpenAI Data Partnerships
And that's not counting the data advantage China already has with the large scale surveillance it does.

US regulators act like they control the world

letitgo12345··on GPT-4-turbo preliminary benchmark results on code-editing
They have tons of usage data by now to figure out which queries to devote model capacity to
letitgo12345··on Phind Model beats GPT-4 at coding, with GPT-3.5 speed and 16k context
Just prompt it to implement the function
Page 1 of 7Next →