HNHacker News
TopNewBestAskShowJobs

andrewgross

96 karma · joined March 23, 2012

submissionscomments
andrewgross··on Mistral Small 3
Super easy to get started, but lacking for larger datasets where you want to understand a bit more about predictions. You generally lose things like prediction probability (though this can be recovered if you chop the head off and just assign output logits to classes instead of tokens), repeatability across experiments, and the ability to tune the model by changing the data. You can still do fine tuning, though itll be more expensive and painfaul than a BERT model.

Still, you can go from 0 to ~mostly~ clean data in a few prompts and iterations, vs potentially a few hours with a fine tuning pipeline for BERT. They can actually work well in tandem to bootstrap some training data and then use them together to refine your classification.

andrewgross··on The impact of competition and DeepSeek on Nvidia
Got it. I’ll review the paper again for that portion. However, it still sounds like the end result is not VRAM savings but efficiently and speed improvements.
andrewgross··on The impact of competition and DeepSeek on Nvidia
Ahh got it, thanks for the pointer. I am surprised there is enough correlation there to allow an entire GPU to be specialized. I'll have to dig in to the paper again.
andrewgross··on The impact of competition and DeepSeek on Nvidia
Is there a concept of an expert that persists across layers? I thought each layer was essentially independent in terms of the "experts". I suppose you could look at what part of each layer was most likely to trigger together and segregate those by GPU though.

I could be very wrong on how experts work across layers though, I have only done a naive reading on it so far.

andrewgross··on The impact of competition and DeepSeek on Nvidia
> The beauty of the MOE model approach is that you can decompose the big model into a collection of smaller models that each know different, non-overlapping (at least fully) pieces of knowledge.

I was under the impression that this was not how MoE models work. They are not a collection of independent models, but instead a way of routing to a subset of active parameters at each layer. There is no "expert" that is loaded or unloaded per question. All of the weights are loaded in VRAM, its just a matter of which are actually loaded to the registers for calculation. As far as I could tell from the Deepseek v3/v2 papers, their MoE approach follows this instead of being an explicit collection of experts. If thats the case, theres no VRAM saving to be had using an MOE nor an ability to extract the weights of the expert to run locally (aside from distillation or similar).

If there is someone more versed on the construction of MoE architectures I would love some help understanding what I missed here.

andrewgross··on How much memory bandwidth do large Amazon instances offer?
4 x 32GB for now. I need to investigate manual OC as EXPO doesn't work with all 4 slots populated. Another option is to try the 192GB 4x48GB Corsair kits at 5200.
andrewgross··on How much memory bandwidth do large Amazon instances offer?
7950x3d w/ 128GB at stock timings (~3200MT/s?). Showed a high base but no increase with threads, need to investigate what is happening.

  1 54.7
  2 50.6
  3 49.6
  4 48.2
  5 47.9
  6 47.4
  7 47.1
  8 46.6
  9 46.5
  10 46.2
  11 46.1
  12 45.9
  13 45.8
  14 45.7
  15 45.7
  16 45.7
  17 45.7
  18 45.8
  19 45.9
  20 45.8
  21 45.8
  22 45.6
  23 45.6
  24 45.5
  25 45.5
  26 45.5
  27 45.5
  28 45.4
  29 45.4
  30 45.4
  31 45.4
  32 45.4
andrewgross··on Show HN: PyEQS – query Elasticsearch like a Django Queryset
Are you referring to the drive to use more of a languages' built in operators to build queries, instead of my approach? I think the Mozilla library uses an approach like that.

http://elasticutils.readthedocs.org/en/latest/

andrewgross··on Show HN: PyEQS – query Elasticsearch like a Django Queryset
It is a modified version of what we use at Yipit. I stripped out mostly things related to working with Django/Flask apps. It comes in very handy when making sure we maintain style and push fewer broken commits. Feel free to re-use.
andrewgross··on Show HN: PyEQS – query Elasticsearch like a Django Queryset
Good point, I will have to swap out some testing libs but should be possible.
andrewgross··on Show HN: PyEQS – query Elasticsearch like a Django Queryset
Thanks for sharing. Nice to see an alternative approach for defining these sorts of things.
andrewgross··on Show HN: PyEQS – query Elasticsearch like a Django Queryset
Fixed. Completely missed that.
andrewgross··on Show HN: PyEQS – query Elasticsearch like a Django Queryset
Thanks, checking this out.
andrewgross··on Show HN: PyEQS – query Elasticsearch like a Django Queryset
Agreed. I had a rough time when I was first learning how to manually build queries as there are few examples of complex queries.
andrewgross··on Show HN: PyEQS – query Elasticsearch like a Django Queryset
Thanks for sharing this, didn't know about it. Always nice to see how other people tackle the same issue.
andrewgross··on Show HN: PyEQS – query Elasticsearch like a Django Queryset
We used to use Haystack, but found it a bit too opinionated for us once we wanted to do some custom stuff. It is a bit more faithful to the Django queryset API, something we had to abandon to let us use more of the complex Elasticsearch query features.
andrewgross··on Show HN: NSString + Ruby
A friend of mine made NSArray extensions with a Ruby syntax a while back. Certainly useful if you find these helpful.

https://github.com/nsantorello/nsarray-extensions

andrewgross··on Ask HN: Can someone explain why this line exists?
Awesome, thanks for the explanation
andrewgross··on Ask HN: Can someone explain why this line exists?
Thanks, was reviewing it with a coworker and crept up on the same conclusion. Seems a bit hard to test and definitely alarming at first glance.
andrewgross··on How Yipit deploys from Github with multiple private repos
Not a bad idea since it would make easier to pull changes outside of the automation without needing to specify the GIT_SSH variable.
andrewgross··on How Yipit deploys from Github with multiple private repos
I know where you are coming from with compiled languages. A big reason for creating this flow was for working with repositories that are not our primary code repository. We have several internal tools and forks that don't have a complex deployment process. We wanted the simplicity of being able to push to master and knowing that it will be out in ~30 minutes. Our primary code base is rolled out with Fabric instead of Chef (for now) but Fabric will still need to execute a pull from Github at some point during the rollout.
andrewgross··on Amazon AWS Throttling Non-Elasticache Memcache Traffic
Switch memcache to UDP for testing and see what it looks like, curious if TCP connection startups are causing issues.