365 karma · joined June 9, 2023
I would say that the "more secure way" is to just use ComfyUI without installing any obscure nodes from unknown developers. You can do pretty much anything using just the default nodes and the big node packs.
The author of the repo is claiming that their repo is hacked, but this is an obvious lie, because their very first GitHub commit is the one where they push the malware. Nobody would hack an empty GitHub account.
I don't know if the author of the repo is lying when they say that Nullbulge is behind the attack (perhaps the author is part of Nullbulge, perhaps not).
Where's the evidence that there's any significant use of NSFW AI by women?
My library solves the following problem: how to tokenize text in a way that is compatible with llama3.
If you don't have any particular constraint (as in "tokenize text in a way that is compatible to model X"), then you can just write your own tokenization that tokenizes the text however you want. It doesn't really make sense to use a complicated tokenization scheme from some LLM model if you don't need to be compatible with that model.
If you really want each word to be its own token, you can easily do that by just splitting on whitespace and punctuation (though that will lead to a huge vocabulary).
The input string "What" (without trailing space) tokenizes into 1 token. The input string "What " tokenizes into 2 tokens. In theory, one might have a tokenizer that would simply tokenize "What " into a single token, but the actual tokenizers we have will tokenize that into at least 2 tokens.
> Curious then why this is called "LLaMA 3 tokenizer" what does it have to do with llama3?
When you input text into any of the LLaMA 3 models, the first step in the process is tokenizing your input. This library is called "LLaMA 3 tokenizer", because it produces the same tokenization as the official LLaMA 3 repo.
When I said that different models use different tokenization schemes, I am talking in comparison to other models, such as LLaMA 1, or GPT-4. Different models use different tokenizers, so the same text is tokenized into different tokens depending on if you're using GPT-4 or LLaMA 3 or what not.
When you enter the word "what", the 3 tokens were: start-of-string token, the token "what", and end-of-string token. I made a change now to hide the special start-of-string and end-of-string tokens so that the visualization is a bit simplified.
Adding a space to input changes the tokenization of the input. Sometimes the resulting token count is the same (if the space is merged into some other text), sometimes the resulting token count increases by one (if the space does not get merged).
That part of the tokenizer is working correctly.
> Also occasionally a space appears as a capital G (in Chrome)
Fixed, thanks for reporting! This is a fork of my earlier tokenizer for LLaMA 1 and the demo visualizer had special handling for tokens 0-256 in LLaMA 1. This LLaMA 3 tokenizer doesn't have same special tokens, so some tokens would be visualized in a weird way (like that G thing you reported). I removed that special handling now and it fixed the visualization issue.
> Question: Is there a special ruleset that llama3 follows that other LMs don't as far as what qualifies as a token?
Different models use different tokenization schemes. Most models use some kind of variant of Byte Pair Encoding, trained with their data (the tokenizer itself is also trained, not only the language model).
The artist, sure. But once the artist goes through all that work to set it up, will they get access to food which they need in order to continue living? No, they won't. Because the consumers will prefer products that are easy to pay for.
> No—I am talking broadly about artists selling their work through crypto, and some of it does include $5 comics, like Sloth Zine:
aaand the link you gave is an NFT collectible...
The price point was not the main point here - sure an NFT doesn't have to be expensive it can also be cheap. That was not the point. My point was that the audience for ponzi-like NFT games is different from the audience for more "traditional" consumers of art. If you are, for example, an artist who currently makes comic books that people will pay for with Visa and Mastercard, you will not be able to provide food on the table by switching into crypto payments.
You're conflating collectibles, which are purchased for speculative purposes, to the kinds of art which are purchased for consumption. Yes, you can use crypto to build ponzi-like games for digital assets, everybody knows that. The question here is can an artist charge like $5 for a comic book or $10 for a monthly subscription. And the answer to that question is no, not to an extent where they would be able to afford food, which they need in order to continue living.
> You don’t need tons of paperwork to set up on and off ramps.
I personally needed a passport, a driver's license, and a recent electricity bill. I also needed a mobile phone with video camera, plus a banking connection which is crypto friendly.
Sure! Acquiring the required paperwork and banking connections to get on crypto on/off ramps is difficult. That's one reason.
Anyway, see how you retreated your argument to "which includes payers"? So it seems that you actually agree with me that artists wouldn't be able to set up crypto payments to provide a livelihood for themselves, due to reasons which are outside their control.
The fact that you disagree with the explanation due to ideological purity reasons is not grounds for downvoting it.
You're living in complete fantasy land. There are many reasons people find it hard to make payments in crypto, other than "just don't want to".
Again, I'm reiterating the point that artists need to eat food to live, and artists would be happy to set up payments in crypto, if that would lead to them gaining access to food. But it wouldn't, so they won't.
> You can look at the code. It's like 50 lines. It is logically trivial
Ah yes, the thing I always love to do before purchasing a naughty comic book: code reviews! I just love to read the code that verifies that my $2 purchase of a comic book. I love sweating nervously and verifying that clicking this button will not have unforeseen consequences, such as one where I lose my house because I missed something in code review. wHy dOEsNT evERyONe jUSt uSE crYPtO
Literally no artist anywhere has ever received an actual payment from an actual customer with this method (I'm not counting the artist themselves making a test payment).
No, if Gumroad bans you, you can't just set up your own crypto smart contracts and make money that way. This is some academic crypto anarchist fantasy.
I didn't use a Trie. The main data structures used are Linked List and Priority Queue.
The tokenization begins by transforming each character to a token. After that merges are applied in priority order as long as merges are possible.
The current state of tokenization is represented as a linked list where each node (token) points to the previous node and next node.
Nodes of the Linked List are added to a Priority Queue _if_ a merge to their corresponding "next node" is possible according to vocabulary (we use a Hash Map to perform these checks fast). Even though we are technically adding nodes of a linked list into the Priority Queue, you can think of the Priority Queue as holding "potential merges" in the order of their priority (priority according to the trained tokenizer merge data).
When we poll a node from the priority queue, we first check that the merge is still possible (because the situation may have changed between the time the node was added to the priority queue, and the time when it was polled from the queue). If the merge is possible, then this is guaranteed to be the most highest-priority merge, so we apply the merge. At this point we need to do several things: we need to mark the previous nodes of the merge as "deleted", we need to create a new node representing the merged token, we need to update the pointers of the adjacent nodes in the linked list, and we need to check if new merges have become possible, and if they have, we need to add the corresponding nodes to the priority queue.