HNHacker News
TopNewBestAskShowJobs

thethirdone

1,219 karma · joined July 10, 2014

submissionscomments
thethirdone··on Complexity physics finds crucial tipping points in chess games
I would agree the "don't blunder" and "punish opponents blunders" are harder than endgame knowledge. However, knowing the basics of endgames is actually important to closing out games. Specifically, knowing KQvK, KRvK, and the "ladder technique" is important.

Without any tactics "take free pieces" probably only gets you to around 1200 (chess.com), but if it includes knight and pawn forks, skewers, pins, and discovered attacks it can get you to 1500. Playing perfectly every game is hard though. I would recommend having more chess knowledge than just those 3 rules before you really try for 1500.

thethirdone··on How to give a senior leader feedback without getting fired
This seems like a very shallow way of thinking. "Losing all respect for the person" implies that you think this is NEVER an appropriate way to address someone. Phrasing a disagreement of opinion as a question of reasoning is often the best course of action.

In particular if a choice has been made and going back to reverse it has significant costs, it is important to not say anything like "We should not be doing this" or "You made a mistake." Unless there is a good of action to reverse course that is simply being rude for no reason. Even in the case where there is a good way to reverse a decision, I would rather ask for the reasoning that led to the decision than strongly state the decision is wrong. If I am working with someone I respect at all, I must entertain the thought that I am wrong and they made the right decision with good reasoning.

What would you say to a superior who made a decision that you disagree with, but don't think is worth reversing? My best guess is either nothing or something that more strongly asserts your belief, but I can't think of any better option than phrasing it as a question.

thethirdone··on Programming a computer for playing chess (1950) [pdf]
It is however important to note that allowing kings to be taken can result in both kings being traded which by the scoring function would be considered neutral. This possibility does not occur in normal games so a search with that modification may generate the wrong value.

Games would need to also be stopped when a king is taken to make the search approximately correct.

thethirdone··on Can logic programming be liberated from predicates and backtracking? [pdf]
> "Completeness" is not about finishing in finite time, it also applies to completing in infinite time.

Can you point to a book or article where the definition of completeness allows infinite time? Every time I have encountered it, it is defined as finding a solution if there is one in finite time.

> No breadth first search is still complete given an infinite branching factor (i.e. a node with infinite children).

In my understanding, DFS is complete for finite depth tree and BFS is complete for finite branching trees, but neither is complete for infinitely branching infinitely deep trees.

You would need an algorithm that iteratively deepens while exploring more children to be complete for the infinite x infinite trees. This is possible, but it is a little tricky to explain.

For a proof that BFS is not complete if it must find any particular node in finite time: Imagine there is a tree starting with node A that has children B_n for all n and each B_n has a single child C_n. BFS searching for C_1 would have to explore all of B_n before it could find it so it would take infinite time before BFS would find C_1.

thethirdone··on How I animate 3Blue1Brown [video]
What is the actual math involved? From following the link in [1] I found almost no math content.
thethirdone··on Reflection 70B, the top open-source model
The sample answers for the horse race question are crazy. [0] Pretty much all the LLM really want to split 6 horses into two groups of three.

Only LLAMA 3 makes the justification that only 2 horses can be raced at a time, but then gets its modified question wrong by racing three horses. I personally would consider an answer that presumes some restriction to how the horses can be raced to be valid if it answers the restricted version correctly.

[0]: https://arxiv.org/html/2405.19616v2#S9.SS2.SSS1

thethirdone··on Against all odds, an asteroid mining company appears to be making headway
From the linked report:

> The material required for construction of the cable is a carbon nanotube composite: currently under development and will be available in 2 years.

Its actually crazy that the report thinks that we would have the necessary material in 2005 and is still used as evidence of a practical space elevator. The graph on page 10 shows the absolutely massive extrapolation from data at the time. It seems disingenuous to even talk final price and time to construct a space elevator (as is done in the one-page brief) without any data confirming the possibility of manufacturing the goal material.

thethirdone··on Judges suspends FCC net neutrality restoration rule
You have not done a good job explaining/proving how they are wrong. Most of your response is only addressing a single paragraph that mentioned Netflix.

> The notion that traffic ratios have anything to do with whether it makes sense to peer and whether someone should pay as long been debunked in the internet context.

Do note how the comment does not mention peering ratios. An ISP being a "hog" does not need to be determined by the peering ratio.

> https://drpeering.net/white-papers/The-Folly-Of-Peering-Rati...

This is a very good article, but it does not directly address the above. It is very specific to arguments about peering ratio. If you have no opinions on peering ratios, you have to read between the lines to get opinions on the original comment.

In Argument #2 counter argument #1: "This is a valid observation ... This is not however an argument for using Peering traffic ratios to restrict Peering."

> And you're also wrong but what neutrality is.

Just saying they are wrong is not helpful. You provide no evidence that "Net Neutrality" has not shifted in meaning since the 90s.

thethirdone··on Google's carbon emissions surge nearly 50% due to AI energy demand
The article does a good job of clearly stating that the 48% is compared to 2019, but words like "surge" and "spike" do not closely match that factual basis and imho are misleading.

Attributing the 48% to AI specifically is largely baseless though. From my skim, the Google report does not make any specific claims about the increase in datacenter emission coming from AI. The closest claims are on page 12 where "the rapid advancement of AI has brought necessary increased attention to its energy consumption and resource demands" is juxtaposed with "total data center electricity consumption grew 17%"

In particular the third "key point" seems highly misleading.

> The company attributed the emissions spike to an increase in data center energy consumption and supply chain emissions driven by rapid advancements in and demand for AI.

The word "spike" does not occur in the document and the 48% number is never close to a mention of AI. While the 17% "spike" may have been attributed to AI by Google, I think it is clear the document does not attribute the 48% to AI.

thethirdone··on Google's carbon emissions surge nearly 50% due to AI energy demand
What is the surrounding text for "versus last year"? I cannot find "versus" or "year" in the article.
thethirdone··on Claude 3.5 Sonnet
The specific issue of using a LLM to make org decisions on how to downsize actually affects nearly all jobs equally.

From what I can tell, most programmers are more ok with LLMs directly replacing them than artists are. I tend to agree that it is better to replace programmers, and protect artists.

thethirdone··on Claude 3.5 Sonnet
> You can say ‘the recent jumps are relatively small’ or you can notice that (1) there is an upper bound at 100 rapidly approaching for this set of benchmarks, and (2) the releases are coming quickly one after another and the slope of the line is accelerating despite being close to the maximum.

The graph does not look like it is accelerating. I actually struggle to imagine what about it convinced the author the progress is accelerating.

I would be very interested in a more detailed graph that shows individual benchmarks because it should be possible to see some benchmarks effectively be beaten and get a good idea of where all of the other benchmarks are on that trend. The 100 % upper bound is likely very hard to approach, but I don't know if the limit is like 99%, 95% or 90% for most benchmarks.

thethirdone··on Accessing Math Solutions via Monte Carlo Self-Refine with LLaMa-3 8B
> They did include an 8-rollout version in the tables? I can't say as to why they didn't try using a bigger model than Llama 3 8B.

That was a typo. I meant a > 8 rollout version. It doesn't seem like they have hit massively diminishing returns yet.

> A single rollout is the full process described in chapter 3. "The algorithm iterates through these stages until a termination condition T is met, including rollout constraints or maximum exploration depth"

A rollout is not the entire process. The summary of normal MCTS correctly identifies rollouts as "random simulations by selecting moves arbitrarily until a game’s conclusion is reached, thereby evaluating the node’s potential". Nothing actually like rollouts is ever described.

Typically MCTS is limited by nodes expanded which is likely what they mean, but because they correctly described rollouts, it seems like I am missing something. Also they mention AlphaGo which replaces rollouts with a neural eval which maybe is relevant.

If rollouts means nodes expanded, 4 and 8 are both just really low numbers.

thethirdone··on Accessing Math Solutions via Monte Carlo Self-Refine with LLaMa-3 8B
I am confused about how the MCTSr algorithm actually is. It is not clear how it is better than simply mutating potential answers (by LLM) and sorting by LLM self-eval.

I have a hard time understanding how many LLM evals MCTSr actually does. How the rollout limit is implemented is not described at all. It doesn't seem like it can mean the same thing as for normal MCTS because there is not any definitive "end" to the tree search. Additionally, MCTS is normally limited by the nodes expanded.

Aside from theoretical concerns, it is not clear why they have not include an > 8 rollout version in the tables or used an LLM stronger than Llama 3 8B. If the concept scales well it should be able to beat GPT-4 and friends by a rather large margin.

Obviously marrying search with LLMs is a ripe area for research, but I find it hard to actually take anything away from this paper.

EDIT:

Added a missing greater than sign

thethirdone··on US has the highest rate of maternal deaths among rich nations. Norway has zero
I rather strongly doubt the "Norway has zero" statement. It does not directly reference any study nor does any other article stating the same. I don't doubt that it is lower or even rounds to 0 per 100,000, but actually 0 is almost certainly wrong.

In 2021 [0], the deaths per 100,000 in Norway was 1.7 which is ~80 maternal deaths. I find it hard to believe that Norway happened to go from 80 to 0 in two years even including a generous amount of luck.

0: graph at the bottom of https://www.oecd-ilibrary.org/sites/1ea5684a-en/index.html?i...

thethirdone··on Ogma: Interpretable Symbolic General Problem-Solving Model
> It has been shown that LLMs are unable to learn concepts beyond the first level of the Borel Hierarchy, which imposes severe limits on the ability of LMs, both large and small, to capture many aspects of linguistic meaning. This means that LLMs will continue to operate without formal guarantees on tasks that require entailments and deep linguistic understanding.

I would tend to disagree with this excerpt and the paper it came from. I am not familiar with the "Borel Hierarchy", but from the paper it seems it is just that LLMs cannot make a guarantee that they interpret "every", "all", etc properly. I don't want to bother trying to decipher all of the math, but it seems highly questionable that they proved it "cannot be learned". The experimental section is greatly lacking in data points.

If you disagree and think the paper is good, please do explain.

thethirdone··on Police in Austin, San Francisco skirt facial recognition ban
It seems I read your comment as more anti-demilitarization than it was. It seemed like you were implying the US would become less safe by implementing the suggestions. If your intent was only that improving safety in one specific way does not merit being called safer in general, that point was not adequately separated from a pro-militarized police position.

I would guess most people are ok with the implication from "better on metric X in some way" -> "better on metric X". If that is a particular issue for you, you probably should make a specific effort to make that point clear.

> by all means, reign in the out of control over weaponized "peace" officers, but please, let's be realistic on what that will actually do to overall safety.

> By all means, move the needle, but don't paint with a wider brush than what the needle is actually moving.

If you compared these lines from both of your comments, the first can be interpreted to imply that overall safety will actually go down if the suggestions are implemented whereas the second seems hard to misinterpret from what I believe to be your position.

thethirdone··on Police in Austin, San Francisco skirt facial recognition ban
I do not agree that it needs a qualifier. My opinion is that safety as a whole would improve with those four suggestions implemented.

You can disagree, but if you think it would make safety worse you probably should point out how those changes would result in a more unsafe country not just that it wouldn't solve all problems.

thethirdone··on DeepSeek-V2: A Strong, Economical, and Efficient Moe Language Model
Do note that it has 236 B parameters which makes the weights ~450 GB.
thethirdone··on Cost of developing new drugs may be lower than industry claims: trial
> IDK about your philosophy, but any death that is 100% unavoidable invalidates the entire system of business that has been built.

Assuming you meant avoidable, I don't think there is ANY system of business that avoids EVERY avoidable death. Often saving each life may be 100% possible, but it may not be possible to save every life. As a real example, I would guess most civilians in Gaza that will die in the next month could have their death avoided if their singular life were the #1 priority of every around them. However, I do not think the humanitarian crisis in Gaza can be entirely solved within the next month so that no civilian will die (even assuming cooperation from both sides and massive external aid).

The absolutism in your comments is not conducive to productive conversation. I would guess there are many people on HN that are too pro capitalism and don't recognize how it fails many people, but being absolutist to the point of impossibilities does not make you convincing.

thethirdone··on LLVM Is Smarter Than Me
relevant: https://stackoverflow.com/questions/74417624/how-does-clang-...
thethirdone··on Los Alamos Chess
Your analysis would pretty similarly apply to 5x5 Gardner's chess which has been weakly solved. Simple tree search can be quite effective as the branching factor is cut down a significant amount (2x in the starting position). Gardner's chess was completely solved without any tablebases.

It is only an argument against building a complete tablebase. Additionally because the extra space is reduced the high piece count version is likely the be much smaller than this would predict because pieces cannot overlap.

Just looking at the piece positions without regard to the types shows a better than 2x relationship for tablebase size.

(64 choose 7) / (64 choose 8) ~ 10% whereas (36 choose 7) / (36 choose 8) ~ 25%

Additionally, because the tablebase would need to go past 18, the possible piece positions would actually shrink. To be clear, the tablebase would not shrink because of the piece types.

thethirdone··on What Computers Cannot Do: The Consequences of Turing-Completeness
It is not clear to me what you would define as the reason to return paradox from `halts`. It is pretty clear you can make a `halts` function that returns halts, loop or unsure. Renaming unsure to paradox would give a valid version of your 3-decidable `halts`. A concrete definition in terms of turing machines is necessary if you want to displace the halting problem.

> (I.E. Fregeian Sense and Reference, a different referent for the same sense).

For the traditional halting problem, all of the programs are encodings for some particular UTM and therefore we are only talking about referents. The halting problem is the statement that there does not exist a referent for the sense that is "Does referent P halt on input I?"

> Just because somebody outside the is_halting function can do something counterproductive, doesn't necessarily mean the specific invocation of do_opposite within the closure of is_halting is impossible to classify.

The problem is the inner call and the outer call are definitionally the same. The input to `halts` is an encoding a of turing machine and an input. The construction of the `do_opposite` function is possible no matter what the encoding of `halts` would be. So if `halts` has a valid encoding, there is a corresponding `do_opposite` that totally confuses it and forces the inner and outer eval to be the same.

> Every proof seems to boil down to "muh contradiction" which feels like, ok, so what?

I think you may misunderstand why everyone is like "muh contradiction". They are doing a proof by contradiction so as soon as they get to a contraction, the proof is complete. I will give a proof of the halting problem for python programs.

Theorem: There does not exist a function `halts(program, inputs)` that correctly determines if a given program halts for EVERY input.

For the sake of contraction, assume such a function `halts` exists. Then carefully construct a program `do_opposite` that intends to befuddle `halts` as follows:

    def do_opposite(inputs):
        if halts(do_opposite,inputs):
            while True:
                "Loop"
        else:
            return
if `halts(do_opposite, inputs) == True` then `do_opposite` must loop forever because the if statement will be followed leading to the inner loop.

if `halts(do_opposite, inputs) == False` then `do_opposite` must immediately return because the else statement will be executed.

For any return value of `halts(do_opposite, inputs)` it must contradict the definition because it does not correctly behave on this particular input. Because this is a contradiction with the only assumption we have made, that assumption must be wrong. QED.

thethirdone··on What Computers Cannot Do: The Consequences of Turing-Completeness
Its not. Mostly it is an argument that turing-completeness is too mathematical to be practical.

> I still had the mainstream belief that AI would be just as capable as humans.

thethirdone··on What Computers Cannot Do: The Consequences of Turing-Completeness
Quite a comprehensive summary of the implications of Turing-Completeness. There aren't any outright errors that pop out to me which is high bar if other articles are anything to judge by.

> This was “easy” because the other major thing about UTM’s is how well they generalize to proving that just about any property of an algorithm is not computable, in the general case.

This paraphrasing of Rice's theorem is good enough, but the mathematical result is meaningless in the practical Turing Complete sense. Rice's theorem is misused often on the internet to put limits on static analysis, but in most cases you can successfully prove that a given program terminates. And by proving that, Rice's theorem no longer prevents you from proving anything.

> That looks like an unprovable property, which is exactly what you would have in a Turing-complete language.

> ...

> In essence, your language is still Turing-complete in practice.

The implication of this statement is that because it is not obvious how to prove a given statement, it must be impossible to prove ANY statements in the general case. No additional argument beyond Rice's theorem has been given to explain why proving other properties would be impossible in the general case.

It is exactly this thought process that makes me strongly dislike Rice's theorem. It is easy to decide that "I must not be able to prove this because its impossible" even when it is practically provable.

thethirdone··on Integer Tokenization Is Insane (2023)
> Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs - https://arxiv.org/abs/2402.14903

Very interesting paper. It does make sense to me the R2L chunking would be better than L2R chunking. It doesn't actually study single digit tokenization.

I am mostly interested in a direct comparison between an LLM wide tokenization vs single digit tokenization. It would be nice to see a direct comparison between similarly trained models. Otherwise it is very hard to get a definitive answer by comparing models with varying sizes, training time, and general strength.

> xVal: A Continuous Number Encoding for Large Language Models - https://arxiv.org/abs/2310.02989

I have seen this paper before, but hadn't payed attention to the p10 vs p100 analysis. Its not clear that the findings would be relevant to an LLM like gtp4 though.

thethirdone··on Integer Tokenization Is Insane (2023)
The nature of numbers as `A10^(n+1) + B10^n` for digits `XXXABXXX` is a very important relationship for doing any arithmetic. As you tokenize strings of digits, you lose the position information within the token make more complicated relationships between tokens because the total number of token pairs increases.

For example in order for a super simple model to learn 3 digit multiplication, it would need to see at least one example for each token in order to get ANY information about what number it represents. Alternatively, with single digits you only need an example where each position is present in each location. Obviously, we would hope to have plenty of data, but I would expect better generalization from models which need to rely on memorization less.

Alternatively, I can see a few reason why grouped digits would be better, but they are more complicated reasons than the reason above so by Occam's Razor my intuition says single digits should be better.

thethirdone··on Integer Tokenization Is Insane (2023)
It has seemed to me that the GPT would be considerably better at numbers if it just considered each digit as a token. Has anyone actually done an experiment to test this?

I wouldn't disbelieve that the grouped version is actually better with data, but it fights my intuition pretty hard. Grouping based on frequency obfuscates the regular nature of numbers.

thethirdone··on A simple dice game shines a bit of light on the psychology of regret
A super basic search doesn't show any studies showing the two cup experiment (with results showing 3x reward to get 50% of people to switch) and TFA doesn't cite any sources. If anyone can cite a paper for it, I would appreciate it.

If I knew the experimenter was trustworthy, I would not switch after the initial decision because there is no point to and definitely would for the $6 alternative. It seems very surprising that the median person needs triple to make it worth it to switch.

If I didn't know the experimenter was trustworthy (for example a street hustler in NYC), I definitely would not switch in the equal reward scenario and may even require 3x to switch because I would expect that they would only offer if my initial selection was wrong.

It seems reasonably likely the actual experiments failed to build trust in the experimentees and so are showing that reasonable decision making process. Given that the described process is very similar to 3 card monte, the experimentees would be primed to distrust the process.

thethirdone··on [dead]
Almost certainly generated spam. They only have one post prior to the release of ChatGPT and within 2 weeks of ChatGPT, they are posting three articles of a similar length to this one in a day. It seems extremely suspicious that they began posting consistently after ChatGPT became public.
← PreviousPage 2 of 17Next →