HNHacker News
TopNewBestAskShowJobs

dimitry12

54 karma · joined August 26, 2009

https://twitter.com/spring_stream/
submissionscomments
dimitry12··on Bitwarden CLI compromised in ongoing Checkmarx supply chain campaign
From Bitwarden official statement: https://community.bitwarden.com/t/bitwarden-statement-on-che...

"a malicious package that was briefly distributed"

"investigation found no evidence that end user vault data was accessed or at risk"

"The issue affected the npm distribution mechanism for the CLI during that limited window, not the integrity of the legitimate Bitwarden CLI codebase or stored vault data."

"Users who did not download the package from npm during that window were not affected."

Downplaying so hard it's disgusting. Bitwarden failed and became a vector of attack. A vendor who is responsible for all my passwords. What a joke. All trust lost: by the incident and comms-style. Time to move before they make an even bigger mistake.

dimitry12··on Show HN: Smol machines – subsecond coldstart, portable virtual machines
https://github.com/earendil-works/gondolin is another project addressing a similar use-case.
dimitry12··on Self-hosted disposable sandboxes with exe.dev-like UX
(I am not the author of vmtree)

Would be pretty magical if when you need a sandbox, you just SSH into it and it's already there, right? exe.dev popularized this UX, but I self-host a lot of things already, so I wanted something similar, but on my own server, which has plenty of spare RAM and CPU cores.

I tasked Gemini DeepResearch with finding what I can use to glue together something like exe.dev. In a typical Gemini Deep Research fashion, it came back with a very obscure recommendation claiming everyone is using it. Repository had one fork and six stars at the time of search.

While obscure, Gemini IMO found a gem: https://github.com/kkovacs/vmtree

It's a collection of a few short bash scripts which automatically provision sandboxes using ssh as a trigger. It has other nice touches such as: controlling how sandboxes get cleaned up, and provisioning subdomains with or without HTTP authentication to expose services running inside the sandbox.

I love it. Mine is deployed entirely inside the VM (HTTPS and SSH DNAT'ed), and doesn't interfere with other VMs on my server.

dimitry12··on Paradagm – cool spreadsheet UX for interacting with LLMs
Now I want open-source self-hosted BYOM version of this
dimitry12··on Paradagm – cool spreadsheet UX for interacting with LLMs
Saw this today and instantly liked the UX. This is not the first attempt to cross spreadsheets and LLMs, but I like the conceptual simplicity here and how clearly it packages "multi-chat + one extra dimension" pattern.

I imagine it works by:

- treating non-enrich columns as inputs

- running a prompt contained in the "description" of the enrich-column against web-enabled LLM

- populate the answer

- repeat for every enrich-column (each column corresponds to a different prompts against the same "subject")

- repeat for each row/subject

I run flows like this in Python almost every day and they were able to capture it perfectly in their UI I think.

I see they realized that "cold-email marketing" is a killer use-case and built-in templated send-email feature as well.

dimitry12··on Tell HN: Beware confidentiality agreements that act as lifetime non competes
Thank you!
dimitry12··on Tell HN: Beware confidentiality agreements that act as lifetime non competes
Thank you! Not my states but seems spot on and I can extract keywords from there.
dimitry12··on Tell HN: Beware confidentiality agreements that act as lifetime non competes
What are the keywords for finding a lawyer who can advise on non-competes?

Asking because it turned out nearly impossible to find a local lawyer to advise on a dispute couple months ago - with 9 out of 10 telling me they only do divorces or real estate or immigration. I was literally calling one by one from a list based on what I believe were relevant search criteria on State Bar website.

dimitry12··on Launch HN: Vassar Robotics (YC X25) – $219 robot arm that learns new skills
If you bridge recorded trajectories with LVLM, then cameras are necessary visual input for LLM to decide which sub-tasks need to be performed to accomplish long-horizon task, and sub-tasks correspond to pre-recorded ("blind") trajectories which are replayed.

If you go beyond pre-recorded "blind" trajectories into more robust task-policies (which you would have to train from many demonstrations) then cameras become necessary to execute the sub-task.

dimitry12··on Launch HN: Vassar Robotics (YC X25) – $219 robot arm that learns new skills
https://github.com/TheRobotStudio/SO-ARM100/tree/main/Simula... I hope applies to this first gen of the product.
dimitry12··on Launch HN: Vassar Robotics (YC X25) – $219 robot arm that learns new skills
Thank you for confirming! Love how simple yet magical your demos look, the elegance of bridging LLM-driven long-horizon planning with the arm.
dimitry12··on Launch HN: Vassar Robotics (YC X25) – $219 robot arm that learns new skills
SO-ARM101 has a leader-arm, which is the arm with same exact dimensions and same servos - but used to read/record the trajectory. You move it with your own hand and teleoperate the follower-arm in real-time. Follower-arm is visible in the demo videos.

If you fully control the environment: exact positions of arm-base and all objects which it interacts with - you can just replay the trajectory on the follower-arm. No ML necessary.

You can use LLM to decide which trajectories to replay and in which order based on long-horizon instruction.

dimitry12··on Launch HN: Vassar Robotics (YC X25) – $219 robot arm that learns new skills
Do I understand correctly that chess-moving demo decomposes into:

- you recorded precise arm-movement using leader-arm - for each combination of source- and target- receptacles/board-positions (looking at the shim visible in the video, which I assume ensures the exact relative position of the arm and chess-board);

- the recorded trajectories are then exposed as MCP-based functions?

Bought the kit. Thank you for the great price! Are table-clamps included?

dimitry12··on Show HN: Whenish – Plan Group Events in iMessages
> Whenish is an iMessage app

Where can I read more about using iMessage as a medium for generic multi-player collaboration? Or if you can just share the right keywords, I will appreciate that!

dimitry12··on Launch HN: Magic Patterns (YC W23) – AI Design and Prototyping for Product Teams
I can't find the "Coming from Hackernews?" button. Where should I look for it?
dimitry12··on R1 Computer Use
No content, no code. "Roadmap" and "Training pipeline" in README are summaries of Section 2.3 of "DeepSeek-R1"-paper.

Sad.

dimitry12··on Open source inference time compute example from HuggingFace
"1B solver + 8B verifier + search" beating 0-shot 70B is nice, agree.

"1B solver + 8B verifier + search" beating 1B-0-shot or 1B-majority as baselines isn't illustrative imo. In other words, by using larger verifier, HF's replication fails to establish a "fair" baseline. Still an awesome blog and release/repository from HF's group - I love it!

dimitry12··on Open source inference time compute example from HuggingFace
"Solver" is `meta-llama/Llama-3.2-1B-Instruct` (1B model, and they use 3B for another experiment), and verifier is `RLHFlow/Llama3.1-8B-PRM-Deepseek-Data`.

See https://github.com/huggingface/search-and-learn/blob/b3375f8... and https://github.com/huggingface/search-and-learn/blob/b3375f8...

In the original paper, they use PaLM 2-S* as "solver" and its fine-tune as "verifier".

dimitry12··on Open source inference time compute example from HuggingFace
From a practical standpoint, scaling test-time compute does enable datacenter-scale performance on the edge. I can not feasibly run 70B on my iphone, but I can run 3B even if takes a lot of time for it to produce a solution comparable to 70B's 0-shot.

I think it *is* an unlock.

dimitry12··on Open source inference time compute example from HuggingFace
To spend more compute at inference time, at least two simple approaches are readily available:

1) make model output a full solution, step-by-step, then induce it to revise the solution - repeat this as many times as you have token-budget for. You can do this via prompting alone (see Reflexion for example), or you can fine-tune the model to do that. The paper explores fine-tuning of the base model to turn it into self-revision model.

2) sample step-by-step (one "thought"-sentence per line) solutions from the model, and do it at non-zero temperature to be able to sample multiple next-steps. Then use verifier model to choose between next-step candidates and prefer to continue the rollout of the more promising branches of "thoughts". There are many many methods of exploring such tree when you can score intermediate nodes (beam search is an almost 50 years old algorithm!).

dimitry12··on Open source inference time compute example from HuggingFace
I believe this is a valid point: HF's replication indeed uses larger off-the-shelf model as a verifier.

In contrast, in the original paper, verifier is a fine-tune of the exact same base model which is used to sample step-by-step solutions (="solver").

dimitry12··on Open source inference time compute example from HuggingFace
In this paper and HF's replication the model used to produce solutions to MATH problems is off-the-shelf. It is induced to produce step-by-step CoT-style solutions by few-shot ICL prompts or by instructions.

Yes, the search process (beam-search of best-of-N) does produce verbose traces because there is branching involved when sampling "thoughts" from base model. These branched traces (including incomplete "abandoned" branches) can be shown to the user or hidden, if the approach is deployed as-is.

dimitry12··on Open source inference time compute example from HuggingFace
Verifier is trained with soft values of reward-to-go for each solution-prefix, obtained from monte-carlo rollouts of step-by-step solutions sampled from the "base" model.

In other words: 1) sample step-by-step solutions from "base" model; 2) do it at non-zero temperature so that you can get multiple continuation from each solution-prefix; 3) use MATH-labels to decide if full solution (leaf/terminal node in MC rolloout) has reward `1` or `0`; 4) roll up these rewards to calculate reward-to-go for each intermediate step.

Yes, verifier trained in this manner can be used to score solution-prefixes (as a process verifier) or a full-solution (as an outcome verifier).

In the original paper (https://arxiv.org/abs/2408.03314) they fine-tune a fresh verifier. HF's replication uses an off-the-shelf verifier based on another paper: https://arxiv.org/abs/2312.08935

dimitry12··on Show HN: Llama 3.2 Interpretability with Sparse Autoencoders
Curious about that too. There are plenty of forks left, for example: https://github.com/plastic-labs/llama3_interpretability_sae (no affiliation)
dimitry12··on Model Context Protocol
Looking at https://github.com/modelcontextprotocol/python-sdk?tab=readm... it's clear that there must be a decision connecting, for example, `tools` returned by the MCP server and `call_tool` executed by the host.

In case of Claude Desktop App, I assume the decision which MCP-server's tool to use based on the end-user's query is done by Claude LLM using something like ReAct loop. Are the prompts and LLM-generated tokens involved inside "Protocol Handshake"-phase available for review?

dimitry12··on Show HN: Openpanel – An open-source alternative to Mixpanel
Awesome!
dimitry12··on Show HN: Openpanel – An open-source alternative to Mixpanel
Looks great as a self-host alternative if/when you make self-hosting feasible.
dimitry12··on “Yes” means “no”: The language of VCs
Can you please expand on the topic of "learn the marketing side if only by doing it semi-professionally for a client"?

I mean, one side of this spectrum is doing affiliate marketing or direct-sales/MLM. Other point on this spectrum might be for an engineer to go get hired as a social media "manager" (lots of "jobs" like this on Upwork).

What possibilities do you have in mind?

dimitry12··on ChatGPT releases “Professional Plan” for $42/mo
Can anyone who has the visible "Upgrade plan" option share a link to it? I wonder if it's only disabled in UI and we can still upgrade.
dimitry12··on Show HN: TensorDock Core GPU Cloud – GPU servers from $0.29/hr
Lambda Labs has (slow, low IOPS) cloud filesystem to persist data between instances. Attached storage does not persist but is high bandwidth and high IOPS, which is a necessity if training small-medium sized models.
Page 1 of 3Next →