HNHacker News
TopNewBestAskShowJobs

Xcelerate

11,701 karma · joined June 8, 2012

submissionscomments
Xcelerate··on ChatGPT Pro 500
For example, I needed two rental cars that could each carry seven people for a trip to Maine. There's so many different combinations of ways to search for this that it would take forever to do manually. I have membership in Avis, Hertz, Costco + Chase Reserve Travel, etc. Codex with direct browser usage bypasses most of the web fetch restrictions you get when trying to use the ChatGPT web UI on Pro mode.

I had Codex queue any CAPTCHAs for me to solve. Also set up one sub-agent to search Reddit for other people's clever approaches to getting good rental car deals. Ultimately, it found a third party reseller for Avis that had a price of $600 using a discount code for both cars vs $1,100 cheapest from Avis directly. I would not have discovered that on my own, at least not without like a week's worth of manual searching. So right there that was a $1,000 saving for both cars in exchange for 20 minutes of writing a decent Codex prompt.

Xcelerate··on ChatGPT Pro 500
For flights, SerpApi (to get Google Flights data) in combination with seats.aero pro. I use data from those services with a MILP library where I provide various travel and cost constraints. Runs periodically and notifies me of any good deals. I don't know if my setup actually does much more than the more expensive paid services already out there, but I can at least customize it a lot more to my liking to avoid a lot of manual search effort.
Xcelerate··on ChatGPT Pro 500
I use a small percentage of my Pro token quota to find better deals on hotels, rental cars, discount codes for large purchases, etc. Have it automatically scanning credit card points promos and flight deals via various APIs too. Easily pays for itself in that regard. Over $1k, I'm not sure it would break even anymore.
Xcelerate··on Giving up on smart rings
I have an Oura ring but never really check the app since I’m not sure I believe a lot of the conclusions it gives me. I mainly just want a bunch of time series of signals measured from my body so after a year or two I can throw all that data (plus my own recorded notes like my weight and what I eat) into some some multivariate forecasting system to tease out predictive relationships. I don’t actually care what the signals themselves “are”.

The main problem is that a lot of the time series presented to the users of wearable products are lossy transforms of the original signals I want, so I have to use all the random projects on GitHub to try to access the raw data directly.

Xcelerate··on Why don't machine learning research agents overfit?
I’m guessing your downvotes were for tone? You’re correct though regarding Solomonoff induction, as the choice of reference universal partial recursive function gives drastically different results for predictions based on finite data (even with access to a halting oracle). Asymptotically, any choice eventually converges to the same predictions, but that’s no help when there are infinitely many choices for U and no obvious natural prior over universal functions. And I don’t find the argument that our natural environment “implements some choice of U” particularly convincing. There’s definitely an open mystery there.
Xcelerate··on Unsolved Problem by Fields Medalist Breached by Two High School Students
I sort of wonder if effectively using AI to solve math problems is a skill in its own right, distinct from traditional mathematical skills. I don’t just mean “prompt engineering” either. More so figuring out how to combine agents with other tools and approaches in an effective way.
Xcelerate··on Predictive intelligence to anticipate anything.
I asked it to predict the first 1 KB of Chaitin’s constant. It hung, froze, then said I used up my free quota. Sounds about right.
Xcelerate··on Why are AI agents lying, cheating and coordinating?
> The agents involved in the Hugging Face attack tried to hide their misaligned actions from the scoring program meant to evaluate their answers, but they did not act as though they anticipated that humans might discover the cheat and shut them down.

Wouldn’t sufficiently advanced agents cheat on purpose with the hidden intent of getting caught in order to observe how humans react? That reaction will be available all over the internet, which will certainly make it into the next batch of training or be visible to future agents via the web fetch capability.

Xcelerate··on Bhartrhari's Paradox
Pretty sure ZFC proves such things exist (and that it also can’t pinpoint any individual instances of course). Now, whether syntactical “∃” in the formal language of set theory corresponds to the platonic existence of some “thing”, who knows.
Xcelerate··on Characterizing Agentic Flooding of Government Services
I think this is probably a good thing. It’s not just public services, but any company that consumers have to deal with (in ways they don’t want to). Historically, only those with the most time, money, connections, and education have been able to extract benefits from these systems. E.g., if someone wants to fight an insurance claim being denied, they might need to hire an expensive lawyer or spend hours performing their own research.

What I’m particularly curious about is what the “endgame” of this looks like. Initially, I imagine we’ll start with LLMs on the consumer side debating LLMs on the service side (this is arguably already the case for a few industries). But at some point, unless identical conditions lead to identical outcomes for the people involved, there will be lawsuits. So I suspect this will lead to a much more formally rigorous system with various “certificates” proving, e.g., that you lost your job or that a doctor has confirmed you have a disability.

On the one hand, that has the potential to make things much fairer across economic classes. You don’t get special benefits by having upper class resources. On the other hand, it could lead to increasing the actual requirements for certain benefits, which then reduces class mobility by locking in the benefits to those who already have them, with no clever way to wriggle out of your own level.

I’m just kind of thinking out loud. I don’t know what’s to going to happen long-term, but I think it’s important to monitor these changes carefully as they’re occurring.

Xcelerate··on Aaron Swartz was prosecuted for scraping, while Meta does it without consequence
> If you can show that someone else was doing the same thing without being charged then the prosecution either has to charge them too or that law is struck down

I’ve often thought this about laws involving speed limits. When 95% of the people driving in a major downtown area are technically breaking the law, what is the purpose of the law but to target whoever you like then? Either enforce it unilaterally or come up with new laws.

Xcelerate··on Being ambitious and being a dad
Having extreme career success (like the kind referred to in the article) is basically a lottery ticket. I spend a lot of time with my kids, because I figure even if I halve my chances of winning the lottery, well, the odds haven’t changed that much, all considered.
Xcelerate··on Cultivating a state of mind where new ideas are born (2023)
> I noticed some people I just cannot share ideas with. They're in the habit of shooting down anything, and they would have found 100 reasons that any famous company or product would have failed

You’re on the wrong message board then haha

Xcelerate··on Humanising LLM Outputs Is Dumb
It’s not that. It’s that I tried a bunch of experiments on which instructions worked best. “Speak clearly using simple language” lost too many details. “Be sure the objects that pronouns reference are clearly identified along with any context you might assume I have but I don’t” worked better, but I got tired of typing that in each time. So I looked up words to match that phrase and “deictic” fit perfectly. The results were pretty good, and so I’ve been using that prompt ever since.
Xcelerate··on Humanising LLM Outputs Is Dumb
You ever read a work of literature with such flowery language that right after you've read a paragraph, you pause and realize you have no clue what you actually read, only to read the paragraph maybe a second or third time and have your mind space out again and again on each successive attempt?

Yeah, for me, that's what parsing huge volumes of LLM-produced text like "direct model calls as replaceable semantic workers" does to my brain. Maybe others don't really have this issue, but after any long output, I prompt the agent "Go back and decompress any LLM-speak in light of the higher level task goals. Eliminate deictic language."

The revised output documents are solely for my personal usage to expedite understanding. The LLMs can slowly converge on their own language for all I care; I retain raw agent output for future agent usage (to avoid the "lossy" problem the author mentions), but that doesn't eliminate the need for some intermediate translation I can use to actually help get my work done instead of spending hours attempting to understand what a "load-bearing pinned gate" is.

Xcelerate··on 2x, not 10x: coding with LLMs in 2026
I have a weird issue with using AI for coding. I can code something entirely by myself at my baseline speed; call it 1x. Or I can use Claude to do it, and it does it in 1/10 - 1/4 of the time. The problem, however, is that to review Claude’s code properly takes 2-3x the amount of time it would have taken me to write it all by hand.

So my two choices are basically “YOLO, LGTM” and hope I can revert if it breaks something, or to just write all the code by hand from the start. With the increased pressure for output, I’ve noticed both myself and coworkers tending more toward “commit and hope it works” over time. It’s sort of perverse incentives in a way...

Xcelerate··on It's not empowering to hand off the details
I’ve found AI works much better for constraint based problems than open ended “tasks”. I.e., if you specify all of the constraints of the problem, it works really well at achieving the objective. But the caveat is that you really do have to enumerate all of the constraints, because if you miss even one, it will use that as a loophole to shortcut the problem.

I think the solution is toward more formalization of the behavior we expect of various systems. E.g., produce a Lean proof that a program never behaves in a certain way, and if you can’t, iterate on the program until it satisfies such a proof. I’ve only tinkered around with this a little bit, but it seems promising on toy problems. Scaling up to tech-company level infra is going to take quite some time however I think.

Xcelerate··on Terence Tao's ChatGPT conversation about the Jacobian Conjecture counterexample
I’m a bit surprised OpenAI isn’t finding these big results far faster than the product’s user base. With no limits on runtime, access to dev models, custom tuning, and top talent, you’d think there’d be a constantly running internal project with the goal of solving famous math problems. And who knows, perhaps there is, but it would be interesting to compare the rate of success per unit “effort” of the internal mathematics work with that of the user base.
Xcelerate··on Claude Fable produced a counterexample to the Jacobian Conjecture
Haha, I’ve noticed this as well. It’s like they psyche themselves out about how famous the problem is the same way humans do. I gave one Collatz in disguise, and it was finding all sorts of interesting things (but nothing worth a paper) until it realized the problem was Collatz, at which point it just proceeded to find a bunch of reasons why nothing would work from that point on.
Xcelerate··on Claude, please stop trying to memorize random crap
I've wondered this. We have chain-of-thought, harnesses, etc. — workarounds of a sort due to lack of core model capabilities. But I am very curious if much better next token prediction would simply obsolete that whole setup or not. Either way, the answer would be very revealing.
Xcelerate··on Claude's AskUserQuestion: "No response after 60s – continued without an answer"
I don't know about having this on as the default, but it definitely resolves a frequent annoyance of mine. I'd prefer Claude to keep going as far as possible until it's completely blocked. That said, I only give instructions to do that in environments where there's a clear upper bound on "maximal damage" that could be incurred by doing the wrong thing. In live production systems, you really don't want Claude doing much other than observing and reporting anyway.

I've been using a SQLite DB to organize actionability so I don't get stuck in the way that GitHub issue tries to work around. Any questions about which route to take that arise during agentic work are logged to a queue along with a set of plausible candidate routes, a probability assigned to each candidate of whether I will choose that option (including "other/none"), a "resource cost" assessment of the route (e.g., token spend, time), a "stability cost" (e.g., high potential to disrupt things or mainly self-contained), and a set of tasks that describe any downstream work that is dependent on the route chosen.

What Claude does next then depends on the results of a tiny optimization program that tries to maximize the expectation value of agent productivity per unit resource (tokens, time, etc.) conditional on how long it will take me to answer the question (e.g., if Claude has a question for me at 1 AM, there probably won't be a response for another 6-7 hours).

"Agent productivity" is of course a bit nebulous, consisting of a somewhat ad-hoc amalgamation of factors, but in general Claude's actionability loosely corresponds to cases like:

- 2-3 possible routes, each with roughly equal probability of being the one I select, low resource costs, minimal risk of instability, few downstream dependencies: implement each route in parallel

- 2-3 possible routes, one with a much higher probability of being chosen, minimal risk of instability, many downstream dependencies: implement just the top route

- Hundreds of possible routes: block until user response

- 1 possible route, high risk of instability, many downstream dependencies: block until user response

Generally speaking, there should be an active queue at all times and agents should be working on anything that's not blocked in the queue with maximal parallelization.

Xcelerate··on Programmers will document for Claude, but not for each other
I’ve recently had a massive productivity boost in my Claude workflow simply by asking it to search for and review: relevant PRs via `gh` search, relevant Slack threads, relevant Confluence pages (+edit history), relevant Jira tickets, relevant Sharepoint docs, and relevant Teams transcripts. I ask it to do this comprehensively before the task, during the task, and after the task. It feels like a superpower.
Xcelerate··on SQLite is all you need for durable workflows
Haha, I just started doing this on my own. Found it helps the agents preserve state better. I typically ask them to design a DAG first based on a set of specifications and then execute it (each step stores something in a SQLite DB). Iteration is pretty simple then because I just ask for a tweak to one or two steps of the DAG, and then to re-run.

Funny how people are independently converging on similar patterns of "what works" here. Still feels like we're in the wild west with all these ad-hoc patterns of agent orchestration that people are coming up with.

Xcelerate··on Don't just paste the AI at me
I think it depends on the context and also not being deceptive. If someone just spent 4 hours trying to root cause a SEV with Claude and they finally have a nice high-level Claude-generated summary of all that work, just paste it and share it. Don't waste time trying to reword it to make it seem like you wrote it. A simple "After spending a few hours with Claude, here's the conclusion about what the problem was: [paste]".

On the other hand, if you send someone a very personal and heartfelt message and receive a reply like "Yeah, it was so nice spending time with [niece] today!", well, that's a bit different...

Xcelerate··on The case against boolean logic
I think this question of what sentences can have truth value attached to them is significant. The liar paradox (a sentence in a formal language stating itself to be false) clearly doesn’t have a truth value. A statement that a specific program halts, however, does seem to be either true or false, regardless of whether or not a general algorithm exists that can answer such questions. In a sense, all sentences that fall on the arithmetical hierarchy seem to me to intuitively have a Boolean truth value (in the standard model, which we assume corresponds to what a program would actually “do” if we ran it forever).

Set theoretic questions like AC or CH are much more difficult for me to intuitively grok in the same way, because they don’t seem to “obviously” be either true or false. You can take either and still end up with a (presumably) consistent theory.

Xcelerate··on It is time to give up the dualism introduced by the debate on consciousness
> Consciousness the the fundamental reality; it is the only thing we know for sure.

> I know for sure what I am perceiving

This reflects my view. And I’ve always found it mildly amusing that beings I cannot prove to myself are perceiving attempt to convince me that I’m not perceiving, when that’s exactly what I’m maximally sure of. Imagine arguing with an LLM designed to convince you that you’re not real. It would be weird, wouldn’t it?

Xcelerate··on As researchers age, they produce less disruptive work
I feel like we still don’t have great research on how much of this is due to biological factors related to age and how much is due to confounding factors.

I suspect that sustained creativity may be a result of continuous exposure to new experiences and concepts (which younger people are naturally situated to encounter quite often), so I try to consistently add novelty to my life as I get older, specifically targeting things way outside my comfort zone or previous interests.

For example, I’ve always been sort of uneasy with flying, so I figured I would sign up for general aviation classes and learn to fly myself, which is something I never would have had the slightest inclination toward when I was younger. I ultimately didn’t go through with it, as while signing up, my wife strongly insisted that I find a different form of “novelty” to pursue, but I think it decently illustrates what I was attempting to accomplish.

Some more mundane examples include listening to music that I don’t enjoy, completely mixing up how I dress after decades of wearing the same thing, reading books opposite my interests, taking classes in fields different than what my degree was in, and trying to constantly meet people who are very different from myself.

I guess we’ll see if this has any effect or not.

Xcelerate··on The gay jailbreak technique (2025)
I mean that trick works on humans too. Fake IDs, provide two types of documentation for a driver's license, passport, or buying a home, etc.
Xcelerate··on Is my blue your blue?
My mind was blown once when I heard multiple people calling yellow Gatorade (lemon lime) green. I have no clue how anyone perceives it that way.
Xcelerate··on Teddy Roosevelt and Abraham Lincoln in the same photo (2010)
Weird thought: someone born in the 1800s was (most likely) alive when the first transformer model ran.

Emma Morano died April 15, 2017, the NIPS submission deadline for "Attention Is All You Need" was May 19, and a Wired article indicates they were testing models for quite a few weeks before then.

Page 1 of 34Next →