HNHacker News
TopNewBestAskShowJobs

underlines

786 karma · joined August 22, 2012

submissionscomments
underlines··on Show HN: Beeper Mini – iMessage client for Android
Gives me Trillian vibes.

So let's introduce my bias:

All Multi-Messengers in the past that relied on reverse engineering proprietary protocols vanished.

underlines··on Hacking my filter coffee machine
My Italian mother worked in the coffee industry for years and she taught me just one thing: Filter coffee tastes awful my son. Only Germans, your granny and Americans drink it. An she is right :D /s
underlines··on Shein Files for U.S. IPO
EU and the country I live in try to battle super fast fashion:

https://www-srf-ch.translate.goog/sendungen/kassensturz-espr...

underlines··on Shein reportedly seeks $90B valuation in IPO
Unfortunately in German. Use the auto translated subs: https://www.youtube.com/watch?v=2Go4Npf1hYU
underlines··on Volute Shaped Blower [video]
There are impostor channels using similar logos and fake builds using cheap labor and concrete. They do this mostly in India in order to gain views and make ad revenue from it. Maybe you were talking about such a channel?
underlines··on Microsoft Security Bulletin MS98-010 – Critical (1998)
Back Orifice was nice, but then came Sub7 :D
underlines··on New world record with an electric racing car: From 0 to 100 in 0.956 seconds
radicalbyte, you're absolutely right that the use of fans in Micromouse increases the traction, and therefore the speed at which the maze is solved. Suction allows for impressive performances.

However, a small caveat that might be worth considering is that while the suction indeed increases speed, it might be more accurate to say that it primarily improves acceleration instead of car speed: The issue often lies with achieving rapid acceleration rather than with maintaining high speed. Even systems with relatively low traction can reach high speeds given enough time and distance, but the ability to accelerate quickly is crucial in competitions like Micromouse.

underlines··on Show HN: Cosmic Media – Search millions of stock photos and videos
Hitting all the APIs on every keystroke. Who released this?
underlines··on Europa the Tech Third World
Switzerland has great workers rights plus salaries slightly below the median (in tech), but way higher for most other professions. Overall much better distributed wealth. Here you're easily getting 120k+ for regular tech jobs in your 30ies.

How people can (still) be in favor of hustle culture and work for companies that actively prevent unions and live in societies that encourage unequalness to such a great extend is beyond my understanding.

underlines··on Arc Browser 1.0
Binaries aren't available for Windows. They show a "Join Windows waitlist" button that leads to an Email form.
underlines··on Show HN: Floneum, a graph editor for local AI workflows
Thanks for your clarifications. I added it to my awesome list:

https://github.com/underlines/awesome-marketing-datascience/...

underlines··on Show HN: Floneum, a graph editor for local AI workflows
1. This looks great and I love how it supports llama.cpp, what about GPTQ, exLlama or c_transformers support?

2. On a first glance, this looks a bit like LangFlow. I guess this is different, but how?

3. Is this freemium, or fully running stand-alone as OSS?

underlines··on AI for AWS Documentation
The inverse idea of Hypothetical Embeddings, HyDE [1] "HyDE is an embedding technique that takes queries, generates a hypothetical answer, and then embeds that generated document and uses that as the final example."

BriefGPT [2] is implementing this and it uses the following prompt at ingestion-time:

"Given the user's question, please generate a response that mimics the exact format in which the relevant information would appear within a document, even if the information does not exist. The response should not offer explanations, context, or commentary, but should emulate the precise structure in which the answer would be found in a hypothetical document. Factuality is not important, the priority is the hypothetical structure of the excerpt. Use made-up facts to emulate the structure. For example, if the user question is "who are the authors?", the response should be something like 'Authors: John Smith, Jane Doe, and Bob Jones' The user's question is:"

1 https://python.langchain.com/docs/modules/chains/additional/...

2 https://github.com/e-johnstonn/BriefGPT

underlines··on AI for AWS Documentation
historical repo name, it's really not that anymore, besides a very old list of marketing stuff that i rarely update. I should rename the repo, but I hesitate :)
underlines··on AI for AWS Documentation
RAG is very difficult to do right. I am experimenting with various RAG projects from [1]. The main problems are:

- Chunking can interfer with context boundaries

- Content vectors can differ vastly from question vectors, for this you have to use hypothetical embeddings (they generate artificial questions and store them)

- Instead of saving just one embedding per text-chuck you should store various (text chunk, hypothetical embedding questions, meta data)

- RAG will miserably fail with requests like "summarize the whole document"

- to my knowledge, openAI embeddings aren't performing well, use a embedding that is optimized for question answering or information retrieval and supports multi language. SOTA textual embedding models can be found on the MTEB Leaderboard [2]. Also look into instructorEmbeddings

- the LLM used for the Q&A using your context should be fine-tuned for this task. There are several open (source?) LLMs based on openllama and others, that are fine tuned for information retrieval. They hallucinate less and are sticking to the context given.

1 https://github.com/underlines/awesome-marketing-datascience/...

2 https://github.com/embeddings-benchmark/mteb

underlines··on Databricks Strikes $1.3B Deal for Generative AI Startup MosaicML
unfortunately not openllama-33b yet
underlines··on Databricks Strikes $1.3B Deal for Generative AI Startup MosaicML
Coincidentally this comes after MosaicML released the best open source commercially usable LLMs on huggingface: mpt-30b, the first open source LLM with 8k context length that can be extended even further with ALiBi and has been trained on a whopping 1 trillion tokens vs. 300 billion for Pythia and OpenLLaMA, and 800 billion for StableLM.
underlines··on MusicGen: Simple and controllable music generation
I generated a 1:20 sample using your prompt "four-on-the-floor downtempo progressive track with soft pads, no vocals" using the audiocraft-webui fork, which allows for longer generation by overlapping generations.

https://sndup.net/njs2/

underlines··on Run and create custom ChatGPT-like bots with OpenChat
added to the list :) thanks!
underlines··on Stack Overflow Is Down
It makes perfect sense:

Create a static status page on yourserver.com/status.html and write:

If you can't read this, it means the status of yourserver.com is: DOWN

underlines··on Run and create custom ChatGPT-like bots with OpenChat
Disclaimer: I am curating LLM-tools on github [1]

A few thoughts:

* allow for custom endpoint URLs, this way people can use open source LLMs with a fake openAI API backend like basaran[2] or llama-api-server[3]

* look into better embedding methods for info-retrieval like InstructorEmbeddings or Document Summary Index

* Don't use a single embedding per content item, use multiple to increase retrieval quality

1 https://github.com/underlines/awesome-marketing-datascience/...

2 https://github.com/hyperonym/basaran

3 https://github.com/iaalm/llama-api-server

underlines··on Bark: A transformer based text to audio system
I wish there was an easy way to fine tune bark, so we could truly clone our voice for bark inference.

Sadly the bark-voice-clone fork doesn't do it. The voices sound nothing like yourself.

Your gradio gui is great. But I don't understand where to copy the cloned npz files to. Even after refreshing the gradio GUI, the ClonedVoices don't appear in the Speaker or Generated Speaker dropdown.

underlines··on Ask HN: What's your favorite GPT powered tool?
There are so many useful tools, that I keep an Awesome list up to date with openAI API, as well as open LLM tools. Especially the up to date list of open LLM models might be of interest to some, in case someone wants to be independent of OpenAI:

https://github.com/underlines/awesome-marketing-datascience/...

underlines··on Lactic Acid Bacteria as Markers for the Authentication of Swiss Cheeses
To be honest, I don't know anybody in Europe who buys American cheese :D
underlines··on Lactic Acid Bacteria as Markers for the Authentication of Swiss Cheeses
lol. I did my Chemical Laboratory Assistant Apprenticeship at exactly this place and my uncle just retired as a Chemist 5 months ago, working at Agroscope Liebefeld for over 30 years.
underlines··on We glued together content moderation to stop soccer pirates
Anyone who hosts an illegal stream website (which usually makes them thousands of $$ per day), would immediately try the following:

1. Stream the video game fifa 2023 for the first 2-3 streams until my account gets onto your allowlist.

2. Stream illegal football content, not triggering your n8n flows, due to the allowlist.

underlines··on It doesn’t take much to make machine-learning algorithms go awry
summarizing the article's important points with vicuna-7b:

* Modern AI systems require large amounts of data to train, much of which comes from the open web, making them susceptible to data poisoning attacks.

* Data poisoning involves adding or modifying information in a training data set to teach an algorithm harmful or undesirable behaviors.

* Safety-critical machine-learning systems are usually trained on closed data sets curated and labeled by humans, making poisoned data less likely to go unnoticed.

* However, generative AI tools like ChatGPT and DALL-E 2 rely on larger repositories of data scraped directly from the open internet, making them vulnerable to digital poisons injected by anyone with an internet connection.

* Researchers from Google, NVIDIA, and Robust Intelligence conducted a study to determine the feasibility of data poisoning schemes in the real world and found that even small amounts of poisoned data could significantly affect an AI's performance.

* Some data poisoning attacks can elicit specific reactions in the system, such as causing an AI chatbot to spout untruths or be biased against certain people or political parties.

* Ridding training data sets of poisoned material would require companies to know which topics or tasks the attackers are targeting.

underlines··on SafeGPT: New tool to detect LLMs' hallucinations, biases and privacy issues
you mean huggingGPT?
underlines··on List of LLMs based on llama, alpaca, vicuna and it's descendants
I try to keep tab on all publicly available llama descendants like alpaca, vicuna, gpt4all etc. Including it's authors, download links, model specifications and data-sets.

It also includes GUIs, wrappers, libraries etc.

It also contains a less complete awesome list of other generative models and resources. PRs welcome.

PRs welcome.

underlines··on An LLM playground you can run on your laptop
Awesome! Does it support safetensors, new ggml format, triton or cuda model files?

I added the playground to the GUI list here:

to https://github.com/underlines/awesome-marketing-datascience/...

← PreviousPage 4 of 9Next →