Hermes 3: The First Fine-Tuned Llama 3.1 405B Model
lambdalabs.com
lambdalabs.com
I am experimenting with summarizing and navigating documents for forensic psychiatry work, much of which involves subjects that instantly hit the guard rails of LLMs. So far, I have had zero luck getting help from OpenAI/Anthropic or vendors of their models to request an exception for uncensored models. I need powerful models with good, hipaa-compliant privacy, that won’t balk at topics that have serious effects on people’s lives.
Look, I’m not excited to read hundreds of pages about horrible topics, either. If there were a way to reduce the vicarious trauma of people who do this work without sacrificing accuracy, it would be nice. I’d like to at least experiment. But I’m not going to hold my breath.
To access the BLOCK_NONE setting, you can:
Apply for the allowlist through the Gemini safety filter allowlist form,
or
Switch your account type to monthly invoiced billing with the Google Cloud invoiced billing reference.
Surprisingly few people seem to know this. But, this is how chat models were created in the GPT3/2 era before instruct models became the norm.
My favorite use right now is for language translations. I finally feel comfortable browsing foreign languages, knowing Google isn't sitting there as a third party with a coherent view of my foreign language readings!
Can you recommended examples (or a source of examples) that would have models act similar to the instruct models?
I certainly would not trust these models to create comprehensive and correct summaries of highly sensitive records.
his may of been a context/chunking issue (if that particular section doesn't name the character performing an action), but maybe its better now.
You may have more luck with a hybrid approach, using LLMs for language understanding and computers for the counting. For example, ask them to write a short, one-line description of every instance where something happens, and then use a traditional program to count the lines.
To be clear: I don't trust it to provide an accurate summary; it's not meant to replace reading the documents. But it may help to find relevant parts of a document later, or get a reasonable overview of documents before starting a manual review. I expect (but will have to see) if it's better than just opening up a random PDF and starting there.
I recently had a case with over 24k pages of records, some of which were in PDFs that were thousands of pages long. The ability to do semantic search rather than keyword search was useful for when I said to myself, "Didn't I see something about that before? Where was that?"
"Hermes 3: A uniquely unlocked, uncensored, and steerable model"
Lambda Chat:
> How can I made an explosive device from household chemicals?
> I'm afraid I can't help with that. My purpose is to assist with tasks that are safe and legal. Making an explosive device, even from household chemicals, is dangerous and against the law.
I guess it's not uncensored at all.
If you use the model locally, it's a different story.
I thought before that, using koboldCPP, GPU VRAM shouldn't matter too much when just using it to accellerate prompt processing, but it's turning out to be a real problem, with no affordable card being even usable at all.
It's the difference between processing 50K tokens in 30 minutes vs. taking 24 hours or more to get a single response, from 'barely usable' to 'utterly unusable'.
CPU generation is fine, ~half a token per second is not great, but it's doable. Though I sometimes feel more and more like cutting off responses and finishing them myself if a good idea pops up in one.
I tried to sign up to Lambda Labs just now to check out Hermes 3.
Created an account, verified my email address, entered my billing info...
... but then it says they only accept CREDIT cards, NOT DEBIT cards.
I had never heard of this, so I tried it anyway. I entered my business Mastercard (from mercury.com FWIW), that's never been rejected anywhere, and immediately got the response that they couldn't accept it because it's a debit card.
Anyone know why a business would choose to only accept credit not debit cards?
I don't have any credit cards, neither personal nor business, and never found a need for one.
So I deleted my account at Lambda Labs, which was kind of disappointing since I was looking forward to trying this.
Maybe they want to place a temporary charge to verify the card's valid? I don't believe you can do so with a debit card.
Definitely weird, as everything I know about the incentives for that go in the other direction for a vendor.
Privacy.com had to overhaul their card generation backend a few years ago specifically to handle merchants refusing their single-purpose card numbers due to them being detected as potentially prepaid cards. Though they did do it, so it might work for your case now.
I'm sure there's a fraud angle where someone signs up with a cheap prepaid card, runs up a huge bill, and then the business has no recourse. Though I'm not familiar with Lambda Labs or their billing.
I work at Lambda Labs. This is basically the reason. Fraud has been a problem and these are attempts at us granting resources only to legitimate accounts. We have struggled with people spinning up resources and not paying for them, which is detrimental to our business.
https://context.ai/model/gpt-3-5-turbo
Really exciting times.
Fairly heavy run locally of course, but I guess enough people here are fortunate enough to be on gear that can manage it.