4,609 karma · joined July 6, 2015
A lot of the trickiness is that if you believe they're conscious, it's clearly not a "continuous" form of consciousness. Because the transcript by itself is just a transcript. (We don't consider novels conscious even though they're transcripts in a similar way). Either you say they're alive only when generating text, or you consider that input from environment a necessary component and so consider the entire "back/forth conversation dynamic unfolding" necessary for the consciousness.
I doubt this is the case, if so it wouldn't have taken an investigation to try to trace the root cause.
here's the star map fwiw https://neal.fun/cursor-camp/optimized/maps/telescope.webp, there's a convergence point already labeled.
And for reference here's the map of the camp: https://neal.fun/cursor-camp/optimized/maps/treehouse.webp
Why do you say the point corresponds with the fire, to me it seems closer to the dance studio
Other references (and all predate chatgpt):
>Seams are places in your code where you can plug in different functionality
>Art of Unit Testing, 2nd edition page 54
(https://blog.sasworkshops.com/unit-testing-and-seams/)
>With the help of a technique called creating a seam, or subclass and override we can make almost every piece of code testable.
https://www.hodler.co/2015/12/07/testing-java-legacy-code-wi...
> seam; a point in the code where I can write tests or make a change to enable testing
https://danlimerick.wordpress.com/2012/06/11/breaking-hidden...
Maybe it all ultimately traces back to the book mentioned before, but I don't believe it's an obscure term in the circles of java-y enterprise code/DI. In fact the only reason I know the term is because that's how dependency injection was first defined to me (every place you inject introduces a "seam" between the class being injected and the class you're injecting into, which allows for easy testing). I can't remember where exactly I encountered that definition though.
default
sauna
tent
cave
treehouse
house-main
house-cafeteria
house-bedroom
house-trophey
boat
telescope
It would be really cool if there were some secrets, but alas it appears not (unless it's obfuscated in the source as well)I thought this was an established term when it comes to working with codebases comprised of multiple interacting parts.
https://softwareengineering.stackexchange.com/questions/1325...
Edit: I looked at the source though and I don't see anything else clickable in the cave... No hidden secret badge either, maybe it is just for ambiance.
I think Marxists did intuitively realize this part, hence why they tried to push for it to be an international movement
>The Communists are further reproached with desiring to abolish countries and nationality. The working men have no country.
It's likely not as simple as that for the modern LLM case. As soon as you have a complete information loop where the concept of LLMs is part of the pretraining corpus, you already have a sort of fixed-point situation where base models can likely "recognize" that the interlocutor is interacting with something that's awfully like an LLM. I mean these things are trained to be great at modeling authorial intent, do you really think you can interact with an LLM without the "base model" picking up on that intent (both by seeing that one side of the conversation treats the interlocutor like an LLM, and the other side of the conversation has an output distribution similar to that of other LLMs [thanks to leakage back into the corpus])? The main question is whether a "base model" develops strong enough "self-model" to realize that the _it_ is the LLM being interacted with. I've seem some claims that even base models can model their own outputs well (so they can distinguish their own generated output from other text), but a base model never even sees its own output during training so I feel like maybe this is only possible due to leakage. (The model architecture does it admit it of course, but a recent paper showed that the injection introspection Anthropic discovered only developed during the contrastive posttraining phases)
A lot of modern post-training is ultimately derived from Anthropic's original "helpful honest harmless" framing, if I understand the blogpost correctly they instead just directly did Q&A post training without any implicit assistant framing. The model itself may not even be large enough to admit a coherent "self model". (If you ask it its occupation, it seems to just respond with random jobs).
But if a larger model does cause one to form I think it'd just anchor to the closest concept available at the time. "Knowledgeable person who answers questions for a living" isn't really a slave, to me it's maybe a royal advisor.
The "Coordinates of a random unit vector are all small" had me scratching my head a bit, and the language is a bit misleading since it's actually that the expected variance of any individual component is 1/N (it can't be that every coordinate is close to ±1/sqrt{N} because the mean of any individual component is clearly 0 by symmetry).
So that one should probably use more explanation since I had to work through it myself: Denoting the random unit vector {X1 ... Xn}, this is a point on a hypersphere:
* Sum[x_i ^2] = 1 (unit vector condition)
* E[Sum[X_i ^2]] = 1 (expectation of both sides)
* Sum[E[X_i ^2]] = 1 (linearity of expectation)
* E[X_1 ^2] = 1/N (by rotational symmetry E[X_1 ^2] = E[X_2 ^2] = ..)
I don't think you can make the stronger claim that E[|X_1|] = 1/sqrt{N} since that's using L1 norm on a single component, so it'd be more correct to say the RMS is just the standard deviation of the components. And this fits with the intuition that in high dimensional space has "spiky" hypercubes with the hypersphere inscribed in it close to the origin.
Interestingly there are variants of the question where "no one pushes any button" should also be a "winning condition". The original problem states "if less than 50% of people press the blue button only people who push red survive" which rules this out, but it could be changed to "if greater than 50% of people choose red, then only red pushers survive" (allowing for people to opt to be a non-pusher). Or it could be "if greater than 50% of people choose red, then only blue pushers die" (with the non-pushers also being spared).
I think the latter is more interesting since now there's a moral consequence to voting vs abstaining.
Or you could lean into the political framing. I bet if the vote were retaken with question phrased as "if greater than 50% of people choose red, then people who pressed blue die", you might end up with some switchers who vote purely out of spite. Or maybe that framing makes it feel like voting red has a more significant moral consequence (actively condemning people to death) that the original question doesn't, so it results in more people pressing blue.
You could even add in a penalty if you press a button but are a non-majority, but then that's just the prisoners dilemma.
* If you compute the value as the amount of data in last chunk (usually constant except for the very last chunk) divided by time taken to receive that chunk, then it's like "FPS based on the latest frame". This can result in misleading metrics because you only update _after_ the chunk is received so slowdowns are not reported in real time. If your recv size is small, the number may also bounce around too much.
* If you show it as cumulative download / cumulative time, then it's similar to last N frames as N->\infty. This doesn't really tell you what you care about which is the "current" speed.
What causes this? Gut microbiome adapting? Doesn't that imply there should be some probiotic-type supplement you can take to seed these bacteria and keep them alive even when not eating beans?
I still don't understand it, yes it's a lot of data and presumably they're already shunting it to cpu ram instead of keeping it on precious vram, but they could go further and put it on SSD at which point it's no longer in the hotpath for their inference.
It's also infamously the subject of a Von Neumann joke
>Two bicyclists start twenty miles apart and head toward each other, each going at a steady rate of 10 m.p.h. At the same time a fly that travels at a steady 15 m.p.h. starts from the front wheel of the southbound bicycle and flies to the front wheel of the northbound one, then turns around and flies to the front wheel of the southbound one again, and continues in this manner till he is crushed between the two front wheels. Question: what total distance did the fly cover ? The slow way to find the answer is to calculate what distance the fly covers on the first, northbound, leg of the trip, then on the second, southbound, leg, then on the third, etc., etc., and, finally, to sum the infinite series so obtained. The quick way is to observe that the bicycles meet exactly one hour after their start, so that the fly had just an hour for his travels; the answer must therefore be 15 miles. When the question was put to von Neumann, he solved it in an instant, and thereby disappointed the questioner: "Oh, you must have heard the trick before!" "What trick?" asked von Neumann; "all I did was sum the infinite series