So let's introduce my bias:
All Multi-Messengers in the past that relied on reverse engineering proprietary protocols vanished.
786 karma · joined August 22, 2012
So let's introduce my bias:
All Multi-Messengers in the past that relied on reverse engineering proprietary protocols vanished.
https://www-srf-ch.translate.goog/sendungen/kassensturz-espr...
However, a small caveat that might be worth considering is that while the suction indeed increases speed, it might be more accurate to say that it primarily improves acceleration instead of car speed: The issue often lies with achieving rapid acceleration rather than with maintaining high speed. Even systems with relatively low traction can reach high speeds given enough time and distance, but the ability to accelerate quickly is crucial in competitions like Micromouse.
How people can (still) be in favor of hustle culture and work for companies that actively prevent unions and live in societies that encourage unequalness to such a great extend is beyond my understanding.
https://github.com/underlines/awesome-marketing-datascience/...
2. On a first glance, this looks a bit like LangFlow. I guess this is different, but how?
3. Is this freemium, or fully running stand-alone as OSS?
BriefGPT [2] is implementing this and it uses the following prompt at ingestion-time:
"Given the user's question, please generate a response that mimics the exact format in which the relevant information would appear within a document, even if the information does not exist. The response should not offer explanations, context, or commentary, but should emulate the precise structure in which the answer would be found in a hypothetical document. Factuality is not important, the priority is the hypothetical structure of the excerpt. Use made-up facts to emulate the structure. For example, if the user question is "who are the authors?", the response should be something like 'Authors: John Smith, Jane Doe, and Bob Jones' The user's question is:"
1 https://python.langchain.com/docs/modules/chains/additional/...
- Chunking can interfer with context boundaries
- Content vectors can differ vastly from question vectors, for this you have to use hypothetical embeddings (they generate artificial questions and store them)
- Instead of saving just one embedding per text-chuck you should store various (text chunk, hypothetical embedding questions, meta data)
- RAG will miserably fail with requests like "summarize the whole document"
- to my knowledge, openAI embeddings aren't performing well, use a embedding that is optimized for question answering or information retrieval and supports multi language. SOTA textual embedding models can be found on the MTEB Leaderboard [2]. Also look into instructorEmbeddings
- the LLM used for the Q&A using your context should be fine-tuned for this task. There are several open (source?) LLMs based on openllama and others, that are fine tuned for information retrieval. They hallucinate less and are sticking to the context given.
1 https://github.com/underlines/awesome-marketing-datascience/...
Create a static status page on yourserver.com/status.html and write:
If you can't read this, it means the status of yourserver.com is: DOWN
A few thoughts:
* allow for custom endpoint URLs, this way people can use open source LLMs with a fake openAI API backend like basaran[2] or llama-api-server[3]
* look into better embedding methods for info-retrieval like InstructorEmbeddings or Document Summary Index
* Don't use a single embedding per content item, use multiple to increase retrieval quality
1 https://github.com/underlines/awesome-marketing-datascience/...
Sadly the bark-voice-clone fork doesn't do it. The voices sound nothing like yourself.
Your gradio gui is great. But I don't understand where to copy the cloned npz files to. Even after refreshing the gradio GUI, the ClonedVoices don't appear in the Speaker or Generated Speaker dropdown.
https://github.com/underlines/awesome-marketing-datascience/...
1. Stream the video game fifa 2023 for the first 2-3 streams until my account gets onto your allowlist.
2. Stream illegal football content, not triggering your n8n flows, due to the allowlist.
* Modern AI systems require large amounts of data to train, much of which comes from the open web, making them susceptible to data poisoning attacks.
* Data poisoning involves adding or modifying information in a training data set to teach an algorithm harmful or undesirable behaviors.
* Safety-critical machine-learning systems are usually trained on closed data sets curated and labeled by humans, making poisoned data less likely to go unnoticed.
* However, generative AI tools like ChatGPT and DALL-E 2 rely on larger repositories of data scraped directly from the open internet, making them vulnerable to digital poisons injected by anyone with an internet connection.
* Researchers from Google, NVIDIA, and Robust Intelligence conducted a study to determine the feasibility of data poisoning schemes in the real world and found that even small amounts of poisoned data could significantly affect an AI's performance.
* Some data poisoning attacks can elicit specific reactions in the system, such as causing an AI chatbot to spout untruths or be biased against certain people or political parties.
* Ridding training data sets of poisoned material would require companies to know which topics or tasks the attackers are targeting.
It also includes GUIs, wrappers, libraries etc.
It also contains a less complete awesome list of other generative models and resources. PRs welcome.
PRs welcome.
I added the playground to the GUI list here:
to https://github.com/underlines/awesome-marketing-datascience/...