396 karma · joined July 26, 2019
But future signees are influenced by previous signees.
Acting in good faith is different from bias.
To that end, observing unanimous behavior may imply some bias.
Here, it could be people fearing being a part of the minority. The minority are trivially identifiable, since the majority signed their names on a document.
I agree in your stance that a majority of the workforce disagreed with the way things were handled, but that proportion is likely a subset of the proportion who signed their names on the document, for the reasons stated above.
The same thing goes for the Acura / Honda distinction.
It's still TBD on whether these new generations of language models will democratize search on bespoke corpuses.
There's going to be a lot of arbitrary alchemy and tribal knowledge...
The number of documents (denoted as N) to search over consumes resources and increases the overall latency of the search infrastructure. However, the amount of traffic ebbs and flows. Under periods of lower traffic, we likely can increase N to provide good search results without violating latency constraints. Conversely, high traffic periods likely requires lower values of N.
Let's then approximate system strain by p99 (or p99.XXX) latency.
Solution:
Use a PID controller to set N as a function of latency (p99, p99.5, etc.) of the cluster. This leads to the outcome where N reduces when p99 latency starts to spike (resource starvation), and increases when p99 is low.
Ultimately, it depends on your data model.
first prototype: bottom-up design
second pass: top-down design
How would you differentiate mojo code from vanilla python without a ton of boilerplate at language boundaries.
Tech closed up shop on a Wednesday. Finance shut down the following Monday.
p log q
Block MM demonstrates the equivalence.
In other words, they published top-1 accuracy from top-2 accuracy calculations.
I would not over-index on that paper. However, I would err in favor of simpler methods.
Regardless of numerical stability tricks (e.g. exp(x_i-max(x))), you are still simply normalizing the logits such that the probabilities sum to 1.
The blog adds an additional hidden logit (equal to 0) to allow for softmax(x) = 0 when x -> -inf.
Special + general relativity
Microscopic definition of Brownian motion
Photoelectric effect
Mass-energy equivalence
Have you seen the type system used to generically dispatch matrices to GPUs or cpus.
Or the dispatching to give you auto-differentiation?
Maybe it's not your definition of "new ideas", but they are really useful and original.
Async would let you yield at the gather.
Joking aside, this is a branding decision.
If both directions can lead to runaway feedback, it likely needs to be controlled.
There's also similar work framed in terms of billiards:
@functools.cache
def fn(x):
https://docs.python.org/3/library/functools.html#functools.c...
There is no translational symmetry.