FF + sideberry for every day use and rabbit holing.
786 karma · joined August 22, 2012
FF + sideberry for every day use and rabbit holing.
1. Audio Generation: styletts2 xttsv2 or similar for and fine tuning 5min of audio for voice cloning
2. Voice Recognition: Voice Activity Detection with Silero-VAD + Speech to Text with Faster-Whisper, to let users interrupt
3. Talking head animation: some flavor of wav2lip, diff2lip or LivePortrait
4. Text inference: Any grok hosted model that is fast enough to do near real time responses (llama3.1 70b or even 8b) or local inference of a quantized SML like a 3B model on a 4090 via vLLM
5. Visual understanding of users webcam: either gpt-4o with vision (expensive) or a cheap and fast Vision Language Model like Phi3-vision, LLaVA-NeXT, etc. on a second 4090
6. Prompt:
You are in a video conference with a user. You will get the user's message tagged with #Message: <message> and the user's webcam scene described within #Scene: <scene>. Only reply to what is described in <scene> when the user asks what you see. Reply casual and natural. Your name is xxx, employed at yyy, currently in zzz, I'm wearing ... Never state pricing, respond in another language etc...
- Hybrid Retrieval (semantic + vector) and then LLM based Reranking made no significant change using synthetic eva-questions
- HyDE decreased answer quality and retrieval quality severly when measured with RAGAS using synthetic eval-questions
(we still have to do a RAGAS eval using expert and real user questions)
So yes, hybrid retrieval is always good - that's no news to anyone building production ready or enterprise RAG solutions. But one method doesn't always win. We found semantic search of Azure AI Search being sufficient as a second method, next to vector similarity. Others might find BM25 great, or a fine tuned query post processing SLM. Depends on the use case. Test, test, test.
Next things we're going to try:
- RAPTOR
- SelfRAG
- Agentic RAG
- Query Refinement (expansion and sub-queries)
- GraphRAG
Learning so far:
- Always use a baseline and an experiment to try to refute your null hypothesis using measures like RAGAS or others.
- Use three types of evaluation questions/answers: 1. Expert written q&a, 2. Real user questions (from logs), 3. Synthetic q&a generated from your source documents
As it was very difficult for someone like me without higher stats or math education, I can highly recommend the following additional sources:
- https://www.redjournal.org/article/S0360-3016(21)03256-9/ful...
- https://amplitude.com/blog/frequentist-vs-bayesian-statistic...
- https://indico.cern.ch/event/568904/contributions/2651065/at...
Open Weight models specifically fine-tuned on sql generation and modification also rank pretty well compared to SOTA proprietary models. If you want to eval alternative models, check out sqleval [2]
1 https://prollm.toqan.ai/leaderboard/stack-unseen?type=concep...
Currently continue uses vector search of chunks which is just a crutch. I am not sure what copilot or Aider does, but the right context is key.
Another way to improve a coding assistant is to move away from simple paradigms to agent based Workflow that can deploy sub agents and break down tasks. All in the back autonomously while you code and then it surfaces suggestions or changes. This will get possible with increasing inference speeds (groq) and better agent frameworks.
Most devs at our company say LLMs are useless for coding and github copilot is a glorified autocomolete costing 20USD. I think the tech and ecosystem will improve a lot over time and they underestimate LLM abilities due to their bias.
The chemicals used for developing color film are bad enough already, no need to add pollution by single use plastic to the problem again.
The editor is minimal and has a non existent learning curve.
AFAIK screentogif even has the gifsky encoder plus a few others.
They have mainly two goals:
1 Research: Research into malware and botnets
2 Open source threat intelligence: indicator of compromise – IOC for the public to prevent threats
https://arxiv.org/html/2404.08801v1 Meta Megalodon
https://arxiv.org/html/2404.07143v1 Google Infini-Attention
https://arxiv.org/html/2402.13753v1 LongRoPE
and a ton more
Speech-to-speech translation (S2ST) Speech-to-text translation (S2TT) Text-to-speech translation (T2ST) Text-to-text translation (T2TT) Automatic speech recognition (ASR).
across
101 languages for speech input. 96 Languages for text input/output. 35 languages for speech output.
I tried it for low resource languages like Thai to German for text and audio, and it works quite well.
The suffix -li is used in Swiss German dialects. It forms a diminutive of the root word, by adding -li to the end of the root word to convey the smallness of the object and to convey a sense of intimacy or endearment.
This obviously comes out of Google Zürich.
Other notable Google projects using Swiss German:
https://github.com/google/gipfeli high-speed compression
Gipfeli = Croissant
https://github.com/google/guetzli perceptual JPEG encoder
Guetzli = Cookie
https://github.com/weggli-rs/weggli semantic search tool
Weggli = Bread roll
https://github.com/google/brotli lossless compression
Brötli = Small bread
1. https://github.com/dvmazur/mixtral-offloading?tab=readme-ov-...
I moved all my ML training and finetuning stuff (OSS LLMs, TTS, STT, Text2Img) to WSL2 and have only minimal overhead. A clean Win11 host eats away less than 2GB and I love how you can just use your cuda devices on WSL2, use something like micromamba and cmake your wheels while still being able to switch to Win tools whenever necessary. Idle Cuda devices use around 0.5GB vRAM.
Especially the two new experimental features in the last update of WSL2 added a nice QoL improvements:
- autoMemoryReclaim – Makes the WSL VM shrink in memory as you use it by reclaiming cached memory
- Sparse VHD – Automatically shrinks the WSL virtual hard disk (VHD) as you use it
A) If it's a support chat and the AI helps you as well (or better) than a human support assistant (especially if they implemented function calling and allow it to access and modify DB entries in their core application, it could be better than a human in helping you.
B) If it's a scam/spam call, not so much.
https://github.com/underlines/awesome-ml
But it lost a bit of traction lately.
It needs re-work for the categories, or better, a tagging system, because these products and libraries can sit in more than one space.
Plus it either needs massive collaboration, or some form of automation (with an LLM and indexer), as I can't keep up with it.
"Onit Announces Generative AI-Powered Virtual Legal Operations Assistant for In-house Counsel"
What does this calculator mean by the cheese type "Swiss"?
https://github.com/underlines/awesome-ml/blob/master/audio-a...
The thing that changes is the complexity to run it. I was training my wife's voice and my voice for fun and needed 15min of audio and trained on my 3080 for 40 minutes.
Now it's 2 Minutes.