HNHacker News
TopNewBestAskShowJobs

fxwin

602 karma · joined December 2, 2024

https://fxwin.net/
submissionscomments
fxwin··on CS240 AI Cheating Retrospective
Universities teach one of the most influential and ubiquitous programming languages, crazy
fxwin··on Solving Factorio Quality
Would even recommend peaceful mode, biters to me always felt like the least fun part of the game, especially while still learning how things worked. It's like you have a coworker come over and interrupt whatever you were doing, except instead of doing small talk they close half your tabs, clear your browser data so you are logged out of every site, and steal your coffee mug
fxwin··on Show HN: Koi.rest – watch some fish and regain your balance
> I suppose that puts me in that group of programmers that are focused on the end-product, agnostic as to how the thing was made. I have truly never enjoyed the plumbing part of the job—just want clean tap water.

FWIW I think most people don't belong to either group exclusively. I often care mostly about the outcome (e.g. at work or when i make small tools for myself), but occasionally i care mostly about the process (e.g. when learning a new programming language or solving problems akin to advent of code)

fxwin··on Show HN: Koi.rest – watch some fish and regain your balance
I think "tools for personal use" implies both that their utility is more important than craftsmanship, and that they aren't public.
fxwin··on Contrastive Language Models
is that how jev works too? how do these differences make those approaches more "instinctive and emotional"? I know they're different on a technical level, I'm asking which characteristic difference necessitates the use of this loaded term from pop psych
fxwin··on Contrastive Language Models
so whats the difference between this and a non-reasoning LLM, or just any generic classifier method that necessitates a new term? there is nothing more intuitive or automatic about jev or this than any of the other currently used AI models
fxwin··on Contrastive Language Models
You're describing external aspects but to me, "instinct" says much more about internal processes than the properties you mentioned, and I haven't seen anything that tells me how these models draw on anything similar to these internal processes to generate their outputs (at least not more than generic LLMs do)
fxwin··on Contrastive Language Models
i'm well aware of the origin and meaning of the term

> "System 1" is fast, instinctive and emotional

this implies more than just "fast", which is precisely why i don't like its present usage

fxwin··on Contrastive Language Models
I really hope that "System One" won't stick around as a new buzzword simply meaning "fast".
fxwin··on Strands Harness
if anything, using oh-my-pi opened that can more than using pi (or better: both!) would have
fxwin··on Strands Harness
it's less about the interfaces and more that it feels disingenuous to compare token usage to a pi setup that injects a bunch of stuff into context rather than the base version. it makes me think they tested it with pi, found they couldnt meaningfully beat its token usage + pass rate, and omitted the comparison. i'd love to be proven wrong, but this seems like the easy explanation
fxwin··on Strands Harness
why would you include oh-my-pi in the comparison but not vanilla pi?
fxwin··on Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
for some reason this is really funny to me. it's like the "black museum" black mirror episode where a consciousness in a toy animal can only communicate using very primitive predefined responses
fxwin··on I built non-autoregressive decision models with RL a year ago
> not to mention the egregious target leakage

I was curious about this so I skimmed the paper [0]:

> SalesRLAgent achieved 96.7% accuracy, outperforming the best commercial alternative by 23.7 percentage points and the best LLM approach by 34.7 percentage points.

For a fuzzy natural language task like this, this magnitude of improvement should already set off alarm bells (Though i admit I'm not even sure what accuracy is even measured here, and the paper doesn't help either). Also, "best LLM" here refers to GPT-4 (at the time of upload, the public already had access to GPT-o3 and). I would have loved to contextualize the performance by looking at model size, but the paper is frustratingly devoid of detail in that regard:

> The core of SalesRLAgent is a reinforcement learning architecture consisting of: • A state encoder network that processes Azure OpenAI embeddings and features • A policy network that estimates conversion probability based on the current state • A value network that estimates the expected cumulative reward • A meta-learning module that assesses prediction confi dence

Also:

> Beyond technical metrics, we evaluated SalesRLAgent in real-world sales environments through A/B testing. [...] After 90 days across 217 representatives and 12,433 con versations, we observed: • 43.2% increase in conversion rate for the test group

This would be a pretty huge result but the fact that this is just shoved into a single paragrpah with no further discussion on methodology, baselines and setup makes me very suspicious.

[0] https://arxiv.org/abs/2503.23303

fxwin··on How GLM built its own inference infrastructure
It's presumptuous for them to assume that a reader of their blog is familiar with their product?

Also I feel like the obvious way to read the very first sentence is that GLM is a language model

> As we develop GLM, the model sometimes exhibits capabilities that surprise us

fxwin··on The k-server conjecture is true
That's fair, I had assumed they hadn't read the article at all ;) If they have, it would have been nice to know which parts of the problem definition the paper describes as "simple" they struggled with, otherwise I default to "I haven't read the article and would like a summary for a technically inclined layman"
fxwin··on The k-server conjecture is true
I take ELI5 in hackernews comments to mean: "Explain like i'm someone with a vaguely technical background but no knowledge in this particular field", not "I'm a literal five year old". I think it's probably close to impossible to explain this adequately to an actual five year old while staying true to the essence of the paper.
fxwin··on The k-server conjecture is true
I should have added "for a technical paper". They are generally not written for five year olds, and "metric space" is a term I've seen introduced anywhere between semesters 1 and 3 in most technical Bachelor's degrees.
fxwin··on The k-server conjecture is true
I feel like the paper itself does a fairly good job:

> The [k-server] problem’s definition is simple: There are k servers located at points of a metric space. At each time step, a request arrives at a point of the metric space. An online algorithm must serve the request immediately by moving a server to the requested location, without knowledge of future requests. The goal is to minimize the total distance traveled by servers.

> The k-server conjecture states that a deterministic online algorithm can achieve competitive ratio k on every metric space.

I only had to look up what "competitive" means in this context, and wikipedia [0] had this to say about it:

> An algorithm is competitive if its competitive ratio—the ratio between its performance and the offline algorithm's performance—is bounded.

The ratio by which this performance is bounded for a k-competitive algorithm is k (plus some constant) [1]. We can consider the analogy of k support technicians ("servers) located in different (physical) locations ("in metric space"): The conjecture/theorem states that in any metric space (Not necessarily two- or three-dimensional), there exists an online algorithm that results in travelled distances of no more than roughly k times that of the optimal distance if all requests were known in advance.

[0] https://en.wikipedia.org/wiki/Competitive_analysis_(online_a...

[1] https://www14.in.tum.de/personen/albers/papers/brics.pdf Section 1.1

fxwin··on Fable 5.1 Solves the Cyphral Distich, a 370-year-old cipher
its not quite the correct method though, the commenter suggests using the first row as a page index, and the second row as the word index, but the actual solution was using both rows as word indices within the 32 paragraphs ("Proquiritations") paired to the 32 numbers in each row
fxwin··on Fable 5.1 Solves the Cyphral Distich, a 370-year-old cipher
I found this german blog: https://scienceblogs.de/klausis-krypto-kolumne/2014/11/17/we...

which links to: https://archive.org/details/s9notesqueries03londuoft/page/12...

which is in reference to the original proquiritations here: https://archive.org/details/worksofsirthomas00mait/page/416/...

i had also never heard of this before today and wonder if people had even seriously tried to decipher this at all?

fxwin··on Mathematicians want proof OpenAI didn't use their work
Because in academia, your name and your body of work is a big deal that can open or close doors, and assigning credit is an important part of it. This isn't a new thing/exclusive to the AI era either.
fxwin··on No Man's Sky Cosmos
I haven't played it myself, but surely an article written one month after the game's release doesn't do it justice, considering there have been multiple major updates attempting to address the gripes players initially had with the game in the 10 years since then (And review trends on steam indicate that these attempts were successful)
fxwin··on Understanding Computer Memory Architecture and SSD Internals
False positives (Pangram labels something slop that isn't) I'm not that worried about, since the worthwile articles generally circle back to me some way through the communities I'm involved in, so I like to err on the side of not instantly reading something anyways.

For false negatives, that depends on the type of article at hand. If an error leads to me spending 2 minutes reading an article that ended up being not worth reading I'd be fine with an order of magnitude more than if false positives happen on articles that take 30 minutes to go through. Maybe on the order of 1-3% overall?

fxwin··on Understanding Computer Memory Architecture and SSD Internals
Unfortunately, given the sheer magnitude of text on the internet, we need heuristics to gauge whether something is worth reading or not. For authors/blogs that I don't yet know, LLM-generated text is a clear anti-signal for me, and Pangram has been very useful in that regard.
fxwin··on Why getting your hands dirty is good for you
> "I have a bucket of mud, in which I stick my hands for 15 minutes every morning.

Not quite but close enough :D https://www.youtube.com/watch?v=DWzWwx8T3nE

fxwin··on Hackers have withdrawn ~4k BTC (~$320M) from the Liquid Federation wallet
In a rug pull, the thing you own (usually some kind of digital asset) goes down in value, leaving you with less money than you started with, whereas theft leaves you no longer possessing the asset itself.
fxwin··on Hackers have withdrawn ~4k BTC (~$320M) from the Liquid Federation wallet
how is it a rug pull? if it's an inside job, doesn't that just make it theft?
fxwin··on Poetry book that Anthropic tried to censor
"censor" feels like the wrong word here
fxwin··on Terpstra Keyboard
> I'm just going off what regular typing keyboards go for.

Are there any regular typing keyboards with velocity sensitive keys? I genuinely don't know, it's just the first thing I'd compare this to

Page 1 of 7Next →