HNHacker News
TopNewBestAskShowJobs

sumo43

216 karma · joined September 9, 2023

Interested in LLMs that browse the web natively. you can reach me at artem@lmresearch.net
submissionscomments
sumo43··on Tongyi DeepResearch – open-source 30B MoE Model that rivals OpenAI DeepResearch
Try running this using their harness https://huggingface.co/flashresearch/FlashResearch-4B-Thinki...
sumo43··on Tongyi DeepResearch – open-source 30B MoE Model that rivals OpenAI DeepResearch
I made a 4B Qwen3 distill of this model (and a synthetic dataset created with it) a while back. Both can be found here: https://huggingface.co/flashresearch
sumo43··on Π0: Our First Generalist Policy
I think the fine tuned policies are still very brittle, but I agree that this is super promising. It's also one of the most open (the model is still closed) research blogposts we've seen from any private embodied AI lab
sumo43··on Octo: An Open-Source Generalist Robot Policy
seems like an improvement on the aloha approach? You still need to finetune it on roughly the same amount of OOD examples. Contrast this with google's approach over 2023, which was training large vision-language models with the goal of generalizing on OOD.
sumo43··on Ask HN: Who wants to be hired? (May 2024)
Location: US

Remote: Yes

Willing to relocate: Yes (US)

Technologies: Python, PyTorch, HuggingFace, C++

Résumé/CV: https://drive.google.com/file/d/1qY-m1tKz4_QpHgxaGryC2vk-DGs...

Email: sumo43@proton.me

ML Engineer & Research Scientist. Ex AI grant startup, hedge fund. I've previously worked on inference for LLMs and vision models & have experience with data curation and multinode training. Looking for summer internships or part time positions

sumo43··on DBRX: A new open LLM
Maybe true for instruct, but pretraining datasets do not usually contain GPT-4 outputs. So the base model does not rely on GPT-4 in any way.
sumo43··on Ask HN: Looking for a project to volunteer on? (February 2024)
SEEKING VOLUNTEERS: open source self-play training for language models

we are a small team associated with EleutherAI. looking to push the frontier of open source language models through self-play. so far we have implemented SPIN. compute included.

email tyoma9k@gmail.com

edit: formatting

sumo43··on Run LLMs at home, BitTorrent‑style
For training you would need more memory. As for the pooling, Theoretically yes but wouldn't latency play as much, if not a greater part in the response time here? Imagine a tensor-parallel gather where the other nodes are in different parts of the country.

Here I'm assuming that Petal uses a large number of small, heterogenous nodes like consumer gpus. It might as well be something much simpler.

sumo43··on Run LLMs at home, BitTorrent‑style
Cool service. It's worth noting that, with quantization/QLORA, models as big as llama2-70b can be run on consumer hardware (2xRTX 3090) at acceptable speeds (~20t/s) using frameworks like llama.cpp. Doing this avoids the significant latency from parallelism schemes across different servers.

p.s. from experience instruct-finetuning falcon180b, it's not worth using over llama2-70b as it's significantly undertrained.

sumo43··on Ask HN: Who here is working on DARPA AICC challenge?
Hello, I'm planning to participate in this challenge. I have experience training/prompting and building products from LLMs, I've also participated in a few CTFs.

sumo43@proton.me