441 karma · joined November 29, 2017
A 2026 EE Times article [1] refers to "compensation" and "calibration" techniques.
[1] https://www.eetimes.com/mythic-rises-from-the-ashes-with-125...
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amazonbot/0.1; +https://developer.amazon.com/support/amazonbot) Chrome/119.0.6045.214 Safari/537.36"
"x-forwarded-for":"44.210.204.255" "x-real-ip":"44.210.204.255"
This is a bit outside my area of expertise, so I don't know how reliable these x-forwarded-for and x-real-ip are.
At least, it claimed to be AmazonBot…
> Starting from version 6.6 of the Linux kernel, [CFS] was replaced by the EEVDF scheduler.[citation needed]
> We’re not raising emotionally intelligent kids. We’re raising kids to navigate human unpredictability as if it’s a design flaw. Because when you grow up with a machine that always gets you, messy human behavior feels broken. We’re not preparing kids to handle people.
I don’t think there’s anything wrong with escaping into fantasy in the right time and place, but young kids (and even well-adjusted adults) can have problems self-moderating and letting fantasy substitute for engaging with reality.
nvcc -arch=sm_86 prospero.cu -o prospero
cuobjdump -sass prospero | grep -E 'FFMA|FADD|FMUL|FMNMX|MUFU\.RSQ' | wc -lFor mine, I paste in a video or playlist URL and it downloads the video and creates a lower resolution transcoded version suitable for streaming to my phone. It also extracts an audio-only version in case that’s more appropriate.
https://arstechnica.com/tech-policy/2025/02/youtube-briefly-...
(PS: Ars Technica is a bit sluggish for me this evening. Not sure why.)
“The Meha API utilizes it's home-grown” -> “its”
Also, I got a relay access denied error when I tried to email you at info@meha.ai
The simplest and highest-impact tests are all the edge cases - if an input matrix/vector/scalar is 0/1/-1/NaN, that usually tells you a lot about what the outputs should be.
It can be difficult to determine sensible numerical limit for error in the algorithms. The simplest example is a dot product - summing floats is not associative, so doing it in parallel is not bitwise the same as serial. For dot in particular it's relatively easy to come up with an error bound, but for anything more complicated it takes a particular expertise that is not always available. This has been a work in progress, and sometimes (usually) we just picked a magic tolerance out of thin air that seems to work.
Solvers are tested using analytical solutions and by inverting them, e.g. if we're solving Ax = y, for x, then Ax should come out "close" to the original y (see error tolerance discussion above).
One of the most surprising things to me is that the suite has identified many bugs in vendor math libraries (OpenBLAS, MKL, cuSparse, rocSparse, etc.) - a major component of what we do is wrap up these vendor libraries in a common interface so our users don't have to do any work when they switch supercomputers, so in practice we test them all pretty thoroughly as well. Maybe I can let OpenBLAS off the hook due to the wide variety of systems they support, but I expected the other vendors would do a better job since they're better-resourced.
For this reason we find regression tests to be useful as well.
* is the treatment of existing work semi-thorough (even experts don’t know everything) and fair?
* are the claims novel w.r.t the existing work? If not, provide a reference to someone who has already done it.
* can you understand the experiments?
* do the experiments and their results lead to the conclusions claimed as novel?
* does the writing inhibit understanding of the technical content?
No peer review I have ever seen or done would catch anything but the most egregious bug of this nature.
https://dl.acm.org/doi/pdf/10.1145/3624062.3624203
Table 6
http://impact.crhc.illinois.edu/shared/Thesis/dissertation-h...
1. Information sensitivity. Even ignoring classified information, there are quite a few things we can't even put into a Google search. It's definitely a no-go for this to end up in a training dataset.
2. "Hallucinations"
Making LLM available through some infrastructure that is already approved for sensitive information will definitely help with the first point, and allow us to experiment with more areas where it might be helpful. I presume this would come along with guarantees about the interactions not being used for training.
It might be even better if some company would sell an appliance we could install on-prem with similar non-training guarantees. Then we could leverage these new tools for very sensitive information, which could be a great help.
They sent me back a simple letter a couple weeks later saying they'd resolved the issue the way I requested and provided me the updated documentation.
I'll definitely consider this approach next time something like this happens.
I understood the comment to mean the posters are not posting about ChatGPT, but using ChatGPT to generate vapid (?, or inauthentic?) posts to farm engagement where there would otherwise be silence.