Tl;dr: if you blast in/out the right 5GHz binary digital signal from a SERDES block on an FPGA connected to a short wire you can talk to your phone!
2,321 karma · joined February 16, 2009
[ my public key: https://keybase.io/benzn; my proof: https://keybase.io/benzn/sigs/ijTuK7E61Do30-K2liLR8gjVuthKb_5Z0_z0m4kwRKM ]
Tl;dr: if you blast in/out the right 5GHz binary digital signal from a SERDES block on an FPGA connected to a short wire you can talk to your phone!
I would guess that unlike fiber -- rarely saturated in a consumer context -- people have ~always wanted more than what cable can provide and thus the operators needed to be strategic about the allocation of bandwidth between DL and UL, hence the asymmetry.
[1] https://colab.research.google.com/github/newhouseb/handcraft...
If you look at the training and test set info:
> We report results on two protocols: (1) Same layout: We train on the training set in all 16 spatial layouts, and test on remaining frames. Following [31], we randomly select 80% of the samples to be our training set, and the rest to be our testing set. The training and testing samples are different in the person’s location and pose, but share the same person’s identities and background. This is a reasonable assumption since the WiFi device is usually installed in a fixed location. (2) Different layout: We train on 15 spatial layouts and test on 1 unseen spatial layout. The unseen layout is in the classroom scenarios.
Depending on how they selected various frames -- let's just say it was random -- the model could have learned something to effect of "this RF pattern is most similar to these two other readings I'm familiar with" (from the surrounding frames) and can therefore just interpolate between the resulting poses associated with those RF patterns (that the model has compressed/memorized into trained weights).
If you look at the meshes between the image ground truth and the paper's results, you'll see that they are strikingly similar. I find this also suspect because WiFi-band RF interacts a lot more with water than with clothes and so you would expect the outline/mesh to get the "meat-bag" parts of you correct but not be able to guess the contours of baggy clothes. That is... unless it has memorized them from the training set.
The mostly widely accessible form of this is probably BNF grammar biasing in llama.cpp: https://github.com/ggerganov/llama.cpp/blob/master/grammars/...
> let's say we had a grammar that had a key "healthy" with values "very_unhealthy" or "moderately_healthy." For broccoli, the LLM might intend to say "very_healthy" and choose "very" but then be pigeonholed into saying "very_unhealthy" because it's the only valid completion according to the grammar.
That said, you can use beam search to more or less solve this problem by evaluating the joint probability of all tokens in each branch of the grammar and picking the one with the highest probability (you might need some more nuance for free-form strings where the LLM can do whatever it wants and be "valid").
1. Fancy token selection w/in batches (read: beam search) is probably fairly hard to implement at scale without a significant loss in GPU utilization. Normally you can batch up a bunch of parallel generations and just push them all through the LLM at once because every generated token (of similar prompt size + some padding perhaps) takes a predictable time. If you stick a parser in between every token that can take variable time then your batch is slowed by the most complex grammar of the bunch.
2. OpenAI appears to work under the thesis articulated in the Bitter Lesson [i] that more compute (either via fine-tuning or bigger models) is the least foolish way to achieve improved capabilities hence their approach of function-calling just being... a fine tuned model.
[i] http://www.incompleteideas.net/IncIdeas/BitterLesson.html
That said, you might want to do something like (backtracking) beam-search which uses various heuristics to simultaneously explore multiple different paths because the semantic information may not be front-loaded, i.e. let's say we had a grammar that had a key "healthy" with values "very_unhealthy" or "moderately_healthy." For broccoli, the LLM might intend to say "very_healthy" and choose "very" but then be pigeonholed into saying "very_unhealthy" because it's the only valid completion according to the grammar.
That said, there are a lot of shortcuts you can take to make this fairly efficient thanks to the autoregressive nature of (most modern) LLMs. You only need to regenerate / recompute from where you want to backtrack from.
1. Modify the output token probabilities to fit any arbitrary use case
2. Perhaps do trigger some sort of backtracking / beam-search
(I'm not Grant but we've chatted on twitter and built similar things)
It was clear (our) GC was not used to this because they were constantly telling us that things were happening when they very clearly weren't (thanks to the live video we had from to the site).
AidKit runs the largest guaranteed income programs in the country. We replace convoluted workflows of glued together spreadsheets, bank portals and human tedium with one unified platform to allow organizations that do good, to do so faster and more effectively. AidKit is in many ways a technological petri dish to explore how we can bring the leverage of a bespoke software team to those without a software background through domain modeling and low-code tooling.
We're looking for an Engineering Manager to help AidKit's engineering organization grow from a tight knit team of folks always jumping in to do what it takes to one where everyone has the space and support needed to do their most leveraged work. Full post here: https://aidkit-inc.rippling-ats.com/job/637223/engineering-m...
We're bootstrapped and profitable.
If you're interested, please reach out to me at ben@aidkit.org!
[1] https://github.com/newhouseb/clownfish#so-how-do-i-use-this-...
The model can choose to call a function; if so, the content will be a stringified JSON object adhering to your custom schema (note: the model may generate invalid JSON or hallucinate parameters).
So sadly, it is just fine tuning. There's no hard biasing applied :(. You were so close, but so far OpenAI![1] https://github.com/newhouseb/clownfish
[2] https://platform.openai.com/docs/guides/gpt/function-calling
So... we left and switched to Smarty and haven't looked back since. We spend tens of thousands on geocoding annually. Really mind-boggling behavior from GCP.
Anecdotally in my own interpretability work (without layer norm), my models also learn rotations fairly frequently. I attributed this to the way I was doing positional embeddings (as rotations), but perhaps there's more to it.
Namely: once you include layer normalization your model is more or less forced to find ways to represent absolute quantities in a way that won't be normalized away and a great way to achieve this is to... store things as rotations of a unit tensor! With that as your primitive, it's fairly natural to rotate around a circle to compute modular addition.
I'd be curious to explore if a different algorithm is learned if one were to stop normalizing at various points. I wouldn't be surprised if a large hurdle to mechanistic interpretability turns out to be that the models have learned complicated rotations in non-obvious coordinate spaces that are tricky to identify after the fact.
When I initially started implementing this I was hung up on similar concerns. For example in GPT2/PotatoGPT the MLP player is 4x the width of the residual stream. I went down a rabbit hole of addition and multiplication in Typescript types (the type system is Turing complete, so it's technically possible!) and after crashing my TS language server a bunch I switched tacticts.
Where I ended up was to use symbolic equivalence, which turned out to be more ergonomic anyway, i.e.
type Multiply<A extends number, B extends number> =
number & { label: `${A} * ${B}` }
const Multiply = <A extends number, B extends number>(a: A, b: B) =>
a * b as Multiply<A, B>;
such that tensor([
params.EmbeddingDimensions, // This is a literal with known size
Multiply(4, params.EmbeddingDimensions)] as const)
is inferred as Tensor<readonly [768, Multiply<4, 768>]>
Notably, switching to a more symbolic approach makes it easier for type checking dimensions that can change at runtime, so something like: tensor([Var(tokens.length, 'Sequence Length'),
Multiply<4, Var(tokens.length, 'Sequence Length')>])
infers as Tensor<readonly [
Var<'Sequence Length'>,
Multiply<4, Var<'Sequence Length'>>]>
And you'll get all the same correctness constraints that you would if these were known dimensions.The downside to this approach is that typescript won't know that Multiply<4, Var<'A'>> is equivalent to Multiply<Var<'A'>, 4> but in practice I haven't found this to be a problem.
Finally, on more complicated operators/functions that compose dimensions from different variables Typescript is also very capable, albeit not the most ergonomic. You can check my code for matrix multiplication and Seb's writeup for another example of a zip function).
// An empty 3x4 matrix
const tensorA = tensor([3, 4])
// An empty 4x5 matrix
const tensorB = tensor([4, 5])
const good = multiplyMatrix(tensorA, tensorB);
^
Inferred type is Tensor<readonly [3, 5]>
const bad = multiplyMatrix(tensorB, tensorA);
^^^^^^^
Argument of type 'Tensor<readonly [4, 5]>' is not
assignable to parameter of type '[never, "Differing
types", 3 | 5]'.(2345)
I prototyped this for PotatoGPT [1] and some kind stranger on the internet wrote up a more extensive take [2]. You can play with an early version on the Typescript playground here [3] (uses a twitter shortlink for brevity)[1] https://github.com/newhouseb/potatogpt