1,647 karma · joined March 15, 2011
Founder at Flybrix, Dir. DSP at General Radar, Founding Engineer at hCaptcha, Zigfu YC S11
amirhirsch.com
I can’t believe our Google account setup is different from any other startup in SF. Anyone have success with this? Do they even have a bot at Google that tracks this attrition?
I think you only wanted clarification of CTF (Capture the Flag) and not AI (Artificial Intelligence) and not GPT-4 (Generative Pre-Trained Transformer version 4) and not CLI (Command Line Interface) and not MCP (Model Context Protocol) and not LLM (Large Language Model)
Quoting TFA (The Fucking Article): “just adapt bro”
lol at the BGP example
I've always felt that it was these kind of folks that caused the 2008 financial crisis
Previous-co could never get argo billing to match argo analytics, and with no support from CF over months we backed away from CF completely in fear that scale-up would present new surprise unknown/untraceable costs
Previous-previous-co is probably the largest user of web worker
Iron dome has nothing to do with that systems.
Good lunch spot for a nudnik
For the first several rounds (when every tree value is in use) Combine the stage 5 XOR with the subsequent round’s tree XORs. You can determine even/odd in hash stage 5 starting with a ^ (a>>16) without Xoring the constant, then you can only need one XOR, this saves you a ton of XORs
Create separate instruction bundles for the first round, rounds 1-5 (combining hash stages 5 XOR with next round tree XORs) and 6-9 (not every tree node is used anymore), round 10 round 11-14 and round 15 and combine them.
you can use add_imm in parallel to load consts. stage 0 you have to do load the tree first and the vals, by later stages when everything is in scratch, you could use 12 scalar XORs and 6 vector XORs on scratch. once you vload vals, you can start to do XORs but can only advance so much at a time, so I’m starting to work on getting hash stages moving to different rounds faster to hide the initial vloads and get to the heavy load section sooner and spread the load pain.
Let me put down my thought process: You have to start to think of designing a 6-slot x8-len vector pipeline doing 48 hashes in parallel first which needs at least 10 steps —- if you convert three stages to multiply adds and do parallel XORs for the other three) —- the problem with 10 cycle hashing is you need to cram 96 scalar xors along side your vector pipeline, so that will use all 12 ALUs for 8 of those cycles. Leaving you only 24 more scalar ops per hash cycle which isn’t enough for the 48 tree value xors..
so you must use at least 11 steps per hash, with 96 xors (including the tree value xor) done in the scalar alus using 8 steps, and giving 3*12 Alu ops per hash cycle. You need 12 more ops per hash to do odd/even, so you must be 12 stages, and just do all of the hash ops in valu, 4 cycles of 12 alus doing modulo, 8 cycles x 12 alus free
With 12 steps and 48 parallel you’re absolute minimum could be 4096/48 x 12 = 1,024 cycles, since stage 10 can be optimized (you don’t need the odd/even modulo cycle, and can use some of those extra scalar cycles to pre-xor the constant can save you ~10 cycles. 1024 gonna be real hard, but I can imagine shenanigans to get it down to 1014, sub-1000 possible by throwing more xor to the scalar alus.
BROADCAST LOAD SCHEDULE
======================================================================
Round | Unique | Load Strategy
------|--------|------------------------------------------
0 | 1 | 1 broadcast → all 256 items
1 | 2 | 2 broadcasts → groups
2 | 4 | 4 broadcasts → groups
3 | 8 | 8 broadcasts → groups
4 | 16 | 16 broadcasts → groups
5 | 32 | 32 broadcasts → groups
6 | 63 | 63 loads (sparse, use indirection)
7 | 108 | 108 loads (sparse, use indirection)
8 | 159 | 159 loads (sparse, use indirection)
9 | 191 | 191 loads (sparse, use indirection)
10 | 224 | 224 loads (sparse, use indirection)
11 | 1 | 1 broadcast → all 256 items
12 | 2 | 2 broadcasts → groups
13 | 4 | 4 broadcasts → groups
14 | 8 | 8 broadcasts → groups
15 | 16 | 16 broadcasts → groups
Total loads with grouping: 839Total loads naive: 4096
Load reduction: 4.9x
I think I'm going to get sub 900 since i just realized i can in-parallel compute whether stage 5 of the hash is odd just by looking at bits 16 and 0 of stage 4 with less delay.....
The vendor tools are still a barrier to the high-end FPGA's hardened IP
Constraint propagation from SICP is a great reference here:
I needed this info, thanks for putting it up. Can this really be an issue for every data center?