In Slashdot's scoring system, comments are scored from -1 to 5, based on factors like insightfulness, informativeness, and whether they're interesting or funny. Here are the hypothetical scores for the comments from Hacker News:
owlninja (42 minutes ago): This user is asking about the applicability of LLMs for text categorization in large datasets. It's a straightforward query, indicating a need for information. Score: 3 (Interesting) - As it opens up a discussion on a relevant and technical topic.
causalmodels (18 minutes ago): This comment provides a direct solution with a resource (zod-gpt). Score: 4 (Informative) - It not only addresses the query but also provides a specific tool to help achieve the goal, adding value to the discussion.
quickthrower2 (41 minutes ago): This comment suggests an approach for creating buckets using a sample and an LLM. Score: 3 (Interesting) - It proposes a practical method, contributing constructively to the original query.
These scores are subjective and would depend on the perspectives of the individual moderators or the community's voting on Slashdot.
https://medium.com/gft-engineering/using-text-embeddings-and...
You can reduce sentences to vectors and then create similarity scores to build a graph over the corpus. If you choose to create clusters then use a llm to summarise them to create labels.
It works quite well for dynamic classification in my purposes but I’m sure there is a better way.
Then the main challenge just becomes prompt design, which can sometimes be nebulous for NLP annotation.