HNHacker News
TopNewBestAskShowJobs

edot

1,242 karma · joined November 27, 2020

submissionscomments
edot··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Proof? Like, do you have logs or something? Not calling you a liar but this seems not correct based on my usage.
edot··on Has Violence Against Teachers Become Accepted by Society?
Well, not exactly. Have you ever seen the quality of life of someone on substantial government assistance? It’s not a life I’d want. No one lives on food stamps and Medicaid because they enjoy it.
edot··on Ember-1
There’s a safer way to do this with nearly no added friction. Give it a read only API key. Then just ask it to write the API calls into a bash script and then read it and run it yourself. The agent can still inspect the live resources and diagnose and give you more commands to run. I do agree I wouldn’t give it create / write access.
edot··on Jev and System One Models: Calibration Beats Accuracy
Humans don’t use “quietly” in normal parlance. Come on man, try harder.

Edit: Dude, you didn’t even use the thing you’re talking about? It’s on OpenRouter. Do better!

Edit 2: OP is a ~60 day old account, only other (positive) commenter is a ~48 day old account. Sus.

edot··on Google’s Project Suncatcher to put ML infrastructure in space
Surely if this were the reason, you'd just put it in international waters?
edot··on ArXiv receives multiyear commitments to support it as an independent nonprofit
It’s good for getting free access to preprints which are often close enough to the paywalled real papers in journals. If I find a paper I want (or more realistically, if ChatGPT finds a paper it wants for me), but it’s behind a paywall, odds are the authors put a preprint on arXiv.
edot··on ArXiv receives multiyear commitments to support it as an independent nonprofit
Tom Dietterich (Editor in Chief at arXiv) posted on LinkedIn the other day that they’re having trouble keeping up with the onslaught of AI-generated papers. Lots of suggestions in the comments but no magic bullets.

It’s ironic that the LLMs which benefit so much from reading arXiv papers of yore are now being used to pollute it. Personally if I see a single author post 2023, I assume it’s junk, especially if they’re not from an actual research institution. Not all solo independent researchers are phonies but … many phonies are solo independent researchers.

OpenReview is okay but even some of the reviewers are apparently using LLMs or just hardly reading.

edot··on OpenAI is enlisting an influencer army to make it look 'good for the world'
Sure, but I was replying to a comment saying that they "won't sell". My point is they absolutely do sell.
edot··on Jev Can't Be Calibrated
It's on OpenRouter if you want to try it.
edot··on Jev Can't Be Calibrated
But why? Jev-style models seem useful for "I have no clue what my incoming distribution looks like but I need to give some sort of answer". If I know what my incoming distribution looks like I'll just upload a CSV of that into ChatGPT and ask it to fit a basic ML model on my data.
edot··on Jev Can't Be Calibrated
Hah! I did the exact same tests as you! I found that if you give it the choice to say "not sure", it picks that 100% of the time. But if you pin it in a corner, then yes it does these weird things. Also yes, the continuous options were much more accurate than the choices. Not sure why that is.
edot··on OpenAI is enlisting an influencer army to make it look 'good for the world'
There are lots of products we know are bad for the world that still sell plenty well. Petroleum products (to include plastic), cigarettes, alcohol, guns, etc.

People generally understand that all of the above are not clear absolute negatives. There are some positives to each. And to LLMs.

edot··on OpenAI is well positioned to fast-follow Jev
Agreed. I tested Jev on OpenRouter this past weekend and it’s “okay” but a specific classifier is significantly better. It used to require skill to import sklearn (ok, not really), but now it’s literally one prompt and upload your Excel file or whatever and you can get your classifier out. It’ll run free, instant, more accurate.
edot··on AI Has No Wisdom and Neither Will You
And do you know where the tools and dies to make those things come from? Where the robots that work at those lines come from? Where the chips in those robots at those lines come from?

Hint: not here. Even basic machining is cooked here. Go on Xometry and compare the pricing of any simple design made in USA vs. China and you’ll see, we can’t competitively do the basics anymore (you conflate basic with crap).

edot··on ZuckOff is a free app that sees Meta glasses before they see you
I feel as though it’s inherent in the models, as part of “alignment” or whatever. I’ve noticed whenever having an LLM help with data analysis, especially Claude, it just can’t help itself from putting caveats and guides and documentation inline or in footers.
edot··on I built non-autoregressive decision models with RL a year ago
But it's hallucination-free, isn't it?
edot··on I built non-autoregressive decision models with RL a year ago
I don’t understand Jev or this. I used this since it’s open source (good job btw!) with the following. State: “a 6 sided die rolled a 3”, question (noul): “Is the number odd?”

Answer: 9% chance, with 91% confidence.

Heh???

Ok, even worse. 75% chance a coin landed heads up?

State: I flipped a coin. Question:

{ "noul_result": { "type": "noul", "instructions": "Did the coin land heads up?" }, "choice_result": { "type": "choice", "instructions": "Determine if the coin landed heads or tails up.", "criteria": { "heads": "the coin landed heads up", "tails": "the coin landed tails up" } } }

Ran on: https://huggingface.co/spaces/convaiinnovations/laya-demo

Result: { "model": "laya", "answers": { "noul_result": { "type": "noul", "noul": 0.6839, "rl_agent": { "act_probability": 1.0 } }, "choice_result": { "type": "choice", "choice": "heads", "probabilities": { "heads": 0.7407, "tails": 0.2593 }, "confidence": 0.1743, "rl_agent": { "act_probability": 1.0 } } }, "usage": { "input_tokens": 76, "output_tokens": 0 }, "latency_ms": 93.8 }

Trying to be even more good-faith:

State: "A fair coin was flipped once. The result was not observed. No other information about the outcome is available."

Questions: { "noul_result": { "type": "noul", "instructions": "Given only the supplied state, what is the probability that the coin landed heads up?" }, "choice_result": { "type": "choice", "instructions": "Given only the supplied state, determine which outcome occurred.", "criteria": { "heads": "the coin landed heads up", "tails": "the coin landed tails up" } } }

Result:

{ "model": "laya", "answers": { "noul_result": { "type": "noul", "noul": 0.1265, "rl_agent": { "act_probability": 1.0 } }, "choice_result": { "type": "choice", "choice": "tails", "probabilities": { "heads": 0.2522, "tails": 0.7478 }, "confidence": 0.1853, "rl_agent": { "act_probability": 1.0 } } }, "usage": { "input_tokens": 123, "output_tokens": 0 }, "latency_ms": 154.5 }

edot··on macOS 27 Golden Gate – Review
Continuing the quote gives more context, which I agree with:

"The obviously AI-generated flyers that come home from school with my kid or that get circulated around neighborhood groups strike me as benign but deeply tacky; the kinds of AI-generated images that get passed around by politicians online don’t even have the benefit of being benign.

More than anything, I’m just tired of seeing it everywhere, tired of being on the lookout for it, tired of it polluting every stream of information I have access to. Wasting time actively trying to avoid auto-generated slop is now a daily fixture of my life."

Obviously, one should be able to entertain an idea without necessarily accepting it, and much like any tool that can be misused, just because some malign actor (or dumb actor) uses it "wrong" doesn't mean you should also not use it.

edot··on Introducing System One Models and Jev
Thanks. I do concede it’s very general but that is a double-edged sword. I don’t need a general fraud identification algorithm. I need an accurate one. If I have another classification task I’ll train another model for that task.
edot··on Introducing System One Models and Jev
Very cool! Can you explain when I would use this vs. training a standard ML model on my data? Suppose I had a fraud dataset with features like customer ID, amount, merchant, online or in-person, etc. - I can't imagine that a general model like Jev would predict this more accurately or cheaply than even a basic XGBoost model trained on my dataset (one that I could build in a few minutes by asking Codex to build it). Where does Jev add value here?
edot··on Dario, Please
Why have any laws then? If laws can't prevent something, only punish it after the fact (which I agree is true)? Yet we have laws. People generally follow them because they expect to be caught and punished. If we passed a law that said the CEO of any company that deploys an LLM that commits a crime gets punished as if they personally did the crime (so, basically instant life sentence if it's even a simple crime times a million instances), I guarantee you the first email the CEO sends to the company is a "pause every LLM project we have - we gotta think about this".
edot··on Garry Tan wants US open-weight AI labs to 'distill' frontier models, too
This is our generation's "Saddam has WMDs". It's something the big labs thought up when they were trying to figure out how to make their product sound scary enough to deserve regulation. Literally no one is doing this or even trying, anyone who would want to do it would have already done it. Not worried about it.
edot··on Flock worker calls police on reporter filming public camera installation
Sum up the number of kids harmed or murdered by the state (pick the top n dictators in the last 100 years for example).

Sum up the number of kids harmed or murdered by random people (remember that, for example, non-family abductions are 1% of abductions).

Compare.

I guarantee you that an empowered state is worse for child welfare than random criminals.

edot··on CIA Releases President's Daily Briefs in Commemoration of 9/11
What’s the purpose of publishing these? Why now? Why redacted still?

Interesting to read but … why bother?

edot··on YouTuber, or someone, wants to send nice emails to active military personnel
This link makes it look like it's some event-specific website. It's not. It's an easy way to navigate the confusing mail routing system to military service members, especially those in bootcamp where they can't have phones. But you want to send them a letter which they CAN get, but you don't know their address or you don't want to actually buy stamps and go to the post office and whatever else. It's a convenience service.

Not sure why other commenter suspects people will go apeshit over this, it's just a good business idea?

edot··on Anthropic Just Threatened to Kill Billions of People. This Is Not Okay
I'm not worried about a future LLM being better, or even inventing some true AGI or ASI. I am worried about current, even last-year, models being applied by evil actors. Let's go through some things that already-existing models can do extremely well:

1) Build a very accurate profile of anyone (given what governments plus advertisers have on everyone already) and what they believe, love, fear, etc. 2) Pilot drones. 3) Identify faces, gaits, vehicles, signals. 4) Hack most anything. 5) Customize content for specific audiences. 6) Find needles in haystacks ("here are feeds from various sensors, tell me when something interesting happens"). 7) Generate fake images and videos. 8) Swarm forums with fake accounts (or hacked accounts, see above). 9) Tell powerful people that they're absolutely right. 10) Do homework for kids.

Honestly I hope ASI happens, at least there's a chance it'll be good. None of the above has any chance to be good. Maybe the needle in a haystack one, like "find me a cure for cancer", but that's about it.

edot··on Navier-Stokes – Tristan Buckmaster [pdf]
If you don't trust the labs directly, you can always use AWS Bedrock or Azure Foundry which should have a much stronger incentive to not train. They make money from asset rental, not selling models. I'd be shocked if they were training.
edot··on Flock Wants a Closely Surveilled World with No Exit
Let's think about the first two letters of CCTV and what that entails.
edot··on Working on Economics with Fable 5
Yes, exactly. Researchers will not like the fact that I refer to the literature as merely a manual, but “RTFM” applies here. Someone has likely already investigated what you’re looking at, or at least found a way to not do it. And sure it’s in the weights, but if you put papers directly in front of the LLM it’s much more impactful.
edot··on AI models ran real businesses: They sent $12,431 in fake invoices, lost $3,200
Ok, so they gave it access to a real Stripe account, real money, and gave it no guardrails or prompting or direction at all other than “make me money”, and you gave it no actual direction as to the type of business you wanted?

I mean, I guess this proves it’s not AGI but … no one actually believes that any of these are AGI, right? It’s a useful tool. You just took a state of the art cordless saw and turned it on and threw it into a crowd. Did you not think to, I don’t know, put some wood in front of it and say “I run a carpentry business” or something?

Page 1 of 8Next →