HNHacker News
TopNewBestAskShowJobs

tgtweak

4,336 karma · joined January 18, 2016

submissionscomments
tgtweak··on A 10 year old Xeon is all you need (for 26B-A4B MTP Drafters without GPU)
It may work - depending on your ram speeds it might not even be that much slower.
tgtweak··on Mythos Finds a Curl Vulnerability
I feel like, if it was a codebase without using any security analysis tools, there would have been some more significant findings - perhaps they can re-run it on an 18 month old commit and see how many it found that were subsequenty found and fixed?

Anyway, I think the case that frontier and next-gen models will get increasingly adept at finding vulnerabilities and that those on the receiving end of those vulnerabilities need to be on top of it.

tgtweak··on Cloudflare to cut about 20% of its workforce
>Cloudflare expects second-quarter revenue of $664 million to $665 million, just under analysts' estimate of $665.3 million

Is this considered below expectations on wallstreet... enough to merit an 18% stock cut?

tgtweak··on The map that keeps Burning Man honest
You can definitely add some telemetry to this that records and analyzes realtime location to "map" the litter, even when using a device like this. The conveyor actually seems very well suited to an external camera that records and analyzes the mess to a degree that should be suitable for the purpose of "recording" litter types and concentrations based on the location, without resorting to manual sweep/dust bins which actually sounds pretty insane at this scale.
tgtweak··on Motherboard sales 'collapse' amid unprecedented shortages fueled by AI
When RAM and an SSD cost more than an entire system used to it's not surprising to see this.
tgtweak··on Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model
Depends entirely on quantization. Q6_K with max context length (262144) is ~40GB of VRAM.

Q8 with the same context wouldn't fit in 48GB of VRAM, it did with 128k of context.

tgtweak··on Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model
I've been using it in a few harnesses (FP8 quant, max context length) and it does seem to get tripped up by tool use, often repeating the same tool when it failed previously - that's usually not a great sign for long-term context and multi-step reasoning. It is excellent at one-shotting though and might be most useful as a sub-agent for a stronger frontier coordinator.
tgtweak··on Windows 9x Subsystem for Linux
If you run it in qemu, all good.
tgtweak··on A new spam policy for “back button hijacking”
Was honestly thinking "yeah nice Google, now do it for Android" since the worst offenders are apps (looking at you, Tiktok)
tgtweak··on All elementary functions from a single binary operator
There is a huge market for "its faster" at the cost of efficiency, but I don't think your claim that an EML hardware block would be inherently less inefficient than the same workload running on a GPU. If you think it would be, back it up with some numbers.

A 10-stage EML pipeline would be about the size of an avx-512 instruction block on a modern CPU, in the realm of ~0.1mm2 on a 5nm process node (collectively including the FMA units behind it), at it's entirety about 1% of the CPU die. None of this suggests that even a ~500 wide 10-stage EML pipeline would be consuming anywhere near the power of a modern datacenter GPU (which wastes a lot of it's energy moving things from memory to ALU to shader core...).

Not sure if you're arguing from a hypothetical position or practical one but you seem to be narrowing your argument to "well for simple math it's less efficient" but that's not the argument being made at all.

tgtweak··on All elementary functions from a single binary operator
For basic arithmetic, this is not required nor would it be faster, as it is not likely advantageous for bulk static transcendal functions. Where this becomes interesting is when combining them OR when chaining them where today they must come back out to the main process for reconfiguration and then re-issued.

Practical terms: Jacobian (heavily used in weather and combustion simulation): The transcendental calls, mostly exp(-E_a/RT), are the actual clock-cycle bottleneck. The GPU's SFU computes one exp2 at a time per SM. The ALU then has to convert it (exp(x) = exp2(x × log2(e))), multiply by the pre-exponential factor, and accumulate partial derivatives. It's a long serial chain for each reaction rate.

The core of this is the Arrhenius rate, (A × T^n × exp(-E_a/(R×T))), which involves an exponentiation, a division, a multiplication, and an exponential. On a GPU, that's multiple SFU calls chained with ALU ops. In an EML tree, the whole expression compiles to a single tree that flows through the pipeline in one pass.

GPU (PreJacGPU) is currently the state of the art for speed on these simulations - a moderate width 8-depth EML machine could process a very complex Jacobian as fast as the gpu can evaluate one exp(). Even on a sub-optimal 250mhz FPGA, an entire 50x50 Jacobian would be about 3.5 microseconds vs 50 microseconds PER Jacobian on an A100.

If you put that same logic path into an ASIC, you'd be about 20x the fPGA's speed - in the nanoseconds per round. And this is not like you're building one function into an ASIC it's general purpose. You just feed it a compiled tree configuration and run your data through it.

For anything like linear algebra math, which is also used here, you'd delegate that to the dedicated math functions on the processor - it wouldn't make sense to do those in this.

tgtweak··on All elementary functions from a single binary operator
I actually don't think this is true -

Traditional processors, even highly dedicated ones like TMUs in gpus, still require being preconfigured substantially in order to switch between sin/cos/exp2/log2 function calls, whereas a silicon implementation of an 8-layer EML machine could do that by passing a single config byte along with the inputs. If you had a 512-wide pipeline of EML logic blocks in modern silicon (say 5nm), you could get around 1 trillion elementary function evaluations per second on 2.5ghz chip. Compare this with a 96 core zen5 server CPU with AVX-512 which can do about 50-100 billion scalar-equivalent evaluations per second across all cores only for one specific unchanging function.

Take the fastest current math processors: TMUs on a modern gpu: it can calculate sin OR cos OR exp2 OR log2 in 1 cycle per shader unit... but that is ONLY for those elementary functions and ONLY if they don't change - changing the function being called incurs a huge cycle hit, and chaining the calculations also incurs latency hits. An EML coprocessor could do arcsinh(x² + ln(y)) in the same hardware block, with the same latency as a modern cpu can do a single FMA instruction.

tgtweak··on All elementary functions from a single binary operator
You could also make an analog EML circuit in theory, using electrical primitives that have been around since the 60s. You could build a simple EML evaluator on a breadboard. Things like trig functions would be hard to reproduce, but you could technically evaluate output in electrical realtime (the time it takes the electrical signal to travel though these 8-10 analog amplifier stages).
tgtweak··on All elementary functions from a single binary operator
Yes actually, it is very regular which usually lends itself to silicon implementations - the paper event talks about this briefly.

I think the bigger question is whether it will be more energy-optimal or silicon density-optimal than math libraries that are currently baked into these processors (FPUs).

There are also some edge cases "exp(exp(x))" and infinities that seem to result in something akin to "division by zero" where you need more than standard floating-point representations to compute - but these edge cases seem like compiler workarounds vs silicon issues.

tgtweak··on All elementary functions from a single binary operator
This paper seems to suggest that a chip with 10 pipeline stages of EML units could evaluate any elementary function (table 4) in a single pass.

I'm curious how this would compare to the dedicated sse or xmx instructions currently inside most processor's instruction sets.

Lastly, you could also create 5-depth or 6-depth EML tree in hardware (fpga most likely) and use it in lieu of the rust implementation to discover weight-optimal eml formulas for input functions much quicker, those could then feed into a "compiler" that would allow it to run on a similar-scale interpreter on the same silicon.

In simple terms: you can imagine an EML co-processor sitting alongside a CPUs standard math coprocessor(s): XMX, SSE, AMX would do the multiplication/tile math they're optimized for, and would then call the EML coprocessor to do exp,sin,log calls which are processed by reconfiguring the EML trees internally to process those at single-cycle speed instead of relaying them back to the main CPU to do that math in generalized instructions - likely something that takes many cycles to achieve.

tgtweak··on All elementary functions from a single binary operator
This could have some interesting hardware implications as well - it suggests that a large dedicated silicon instruction set could accelerate any mathematical algorithm provided it can be mapped to this primitive. It also suggests a compiler/translation layer should be possible as well as some novel visualization methods for functions and methods.
tgtweak··on Who is Satoshi Nakamoto? My quest to unmask Bitcoin's creator
Has Back not produced any c++ code from his thesis or days in University? That would be more useful for satoshi-profiling than his written prose, I would think.
tgtweak··on GLM-5.1: Towards Long-Horizon Tasks
Share the harness for that browser linux OS task :)
tgtweak··on 15 Years of Forking
Also, the fact waterfox has almost zero telemetry out of the box and doesn't even have a bonafide number for active users as a result is a good hint at how sincere the "anti-tracking" ethos is applied. Most open source applications have some form of telemetry, even "anonymous" telemetry - but this does really stand out.
tgtweak··on 15 Years of Forking
Adless monetization is a very difficult challenge - something I've worked on many times over the years (compute-monetization, shopping-commission monetization, payment interchange monetization) and it's always been very difficult to compete with ads. I think the waterfox approach of "ads if you're OK with it" opt-in and sane defaults is the better one, but it's very difficult to make ends meat compared to competitors offering full monetization on by default when you're only getting it (and getting less per search) if users opt-in.

Compute/resource monetization is the one after all these years that has done the best at replacing ads as a means of monetization for users, and it requires a very intelligent scheduling system + ethical ecosystem to work (most have just tried running crypto miners that cost users more electricity than they earn).

tgtweak··on Copilot edited an ad into my PR
"Save time by changing your default browser to edge and enabling onedrive"

"just tips bro"

tgtweak··on I wanted to build vertical SaaS for pest control, so I took a technician job
I know about 4 friends that have left their parent company, built a killer product that same company didn't think to build or didn't believe in, only to get acquired by that same company after a few years... some have done it multiple times.

I think this falls in exactly that situation. You see how janky these national companies are doing things, plot out a disruptive course, then disrupt them in a particular region so that you can extrapolate how much that will hurt at national scale and force a buyout that's way beyond the multiple you bought those small operators for.

tgtweak··on Ripgrep is faster than grep, ag, git grep, ucg, pt, sift (2016)
codex is basically a ripgrep wrapper at this point :)
tgtweak··on Drugwars for the TI-82/83/83 Calculators (2011)
Ludes in this version... definitely 1984
tgtweak··on Astral to Join OpenAI
I mean you pirouetted onto the AI hype train before running out of working capital - I guess that's doing great financially by some definitions.
tgtweak··on Astral to Join OpenAI
To raise $4m seed from AAA partners usually requires connections + track record/credability of the founders - looks like they have that here since they raised 3 rounds with zero revenue.
tgtweak··on Astral to Join OpenAI
So I don't see how the acquisition is collateral - it's an acquihire plain and simple, if anything else it would be supply chain insurance as they clearly use a lot of these tools downstream. As you noted the licensing is extremely permissive on the tools so there appears to be very little EV there for an acquirer outside of the human capital building the tools or building out monetized features.

I'm not too plugged into venture cap on opensource/free tooling space but raising 3 rounds and growing your burn rate to $3M/yr in 24 months without revenue feels like a decently risky bag for those investors and staff without a revenue path or exit. I'd be curious to see if OpenAI went hunting for this or if it was placed in their lap by one of the investors.

OpenAI has infamously been offering huge compensation packages to acquire talent, this would be a relative deal if they got it at even a modest valuation. As noted, codex uses a lot of the tooling that this team built here and previously, OpenAI's realization that competitors that do one thing better than them (like claude with coding before codex) can open the door to getting disrupted if they lapse - lots of people I know are moving to claude for non-coding workflows because of it's reputation and relatively mature/advanced client tools.

tgtweak··on Astral to Join OpenAI
Are you implying that the revenue multiple on this acquisition is lower than openAIs and that they'd be making money by acquiring and folding into their valuation multiple? I think that's not the case and I would wager non existent.

This was an acquihire (the author of ripgrep, rg, which codex uses nearly exclusively for file operations, is part of the team at Astral).

So, 99% acquihire , 1% other financial trickery. I don't even know if Astral has any revenue or sells anything, candidly.

tgtweak··on Astral to Join OpenAI
Amusing that the best python tools are written entirely in rust.
tgtweak··on Nvidia NemoClaw
I think the more useful tool would be an LLM prompt proxy/firewall that puts meaningful boundaries in place to prevent both exfiltration of sensitive data and instructions that can be destructive. Using the same context loop for your conversational/coding workflow makes the task at hand and the security of that task very hard to differentiate.

Sending POST?DEL requests? risky. Sending context back to a cloud LLM with credentials and private information? risky. Running RM commands or commands that can remove things? risky, running scripts that have commands in them that can remove things? risky.

I don't know how we've landed on 4 options for controls and are happy with this: "ask me for everything", "allow read only", "allow writes" and "allow everything".

Seems like what we need is more granular and context-aware controls rather than yet another box to put openclaw in with zero additional changes.

← PreviousPage 2 of 34Next →