HNHacker News
TopNewBestAskShowJobs

bananaheel

15 karma · joined November 14, 2018

submissionscomments
bananaheel··on Mamba outperforms transformers "everywhere we tried"
Still chewing but my takeaways thus far:

SSM models strengths are with continuous data like audio and video, they struggle with discrete data like text/ DNA. This newest architecture uses selective attention to try to address the weaknesses around discrete data with some loss in performance in continuous tasks, empirically shown here with audio. The empirical exploration was limited to smaller size models, the performance as larger scales is yet to be explored in practice.

I found my deepest understanding of the selection mechanism came from struggling with the discretization in section 2, followed by the deeper explanations of the variables involved in 3.5.2. This video gives excellent background to SSMs, along with a detailed walk through of the paper itself.[a]

I am still coming to understand S4, SSMs in general but the video suggested this annotated explainer that has been helping a lot[b].

I would also point out section 3.1 and it’s discussion of the tradeoffs between compression and effectiveness as particularly interesting.

I do wonder how many different GPUs / hardware architectures will be able to execute the optimizations that are described as critical. I think the nature of the optimizations is the part of the paper I understand least well.

The paper taken at face value looks very exciting. The promise of a very large context window with 5x throughput for inference would be huge if it proves to scale well. I do wonder if it will make sense to train SSMs without this selection mechanism for specific continuous use cases where it seems to perform better or if other architectures will prove to better serve those cases.

----

[a] https://www.youtube.com/watch?v=ouF-H35atOY

[b] https://srush.github.io/annotated-s4/

bananaheel··on Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Still chewing but my takeaways thus far:

SSM models strengths are with continuous data like audio and video, they struggle with discrete data like text/ DNA. This newest architecture uses selective attention to try to address the weaknesses around discrete data with some loss in performance in continuous tasks, empirically shown here with audio. The empirical exploration was limited to smaller size models, the performance as larger scales is yet to be explored in practice.

I found my deepest understanding of the selection mechanism came from struggling with the discretization in section 2, followed by the deeper explanations of the variables involved in 3.5.2. This video gives excellent background to SSMs, along with a detailed walk through of the paper itself.[a]

I am still coming to understand S4, SSMs in general but the video suggested this annotated explainer that has been helping a lot[b].

I would also point out section 3.1 and it’s discussion of the tradeoffs between compression and effectiveness as particularly interesting.

I do wonder how many different GPUs / hardware architectures will be able to execute the optimizations that are described as critical. I think the nature of the optimizations is the part of the paper I understand least well.

The paper taken at face value looks very exciting. The promise of a very large context window with 5x throughput for inference would be huge if it proves to scale well. I do wonder if it will make sense to train SSMs without this selection mechanism for specific continuous use cases where it seems to perform better or if other architectures will prove to better serve those cases.

----

[a] https://www.youtube.com/watch?v=ouF-H35atOY

[b] https://srush.github.io/annotated-s4/

bananaheel··on Amazon permanently shuts down Prime Pantry
This is more of a capturing of current state than a full blown documentary. I still learned a lot about amazon’s logistics.

https://youtube.com/watch?v=2qanMpnYsjk

bananaheel··on del.icio.us is now an unconfigured Apache install
Your experience mirrors my own. While I had hopes that pinboard would stay stable, perhaps even slowly add features over time I’ve seen the opposite. The lack of status page has left me thinking that perhaps I should ping the maintainer to ask if some downtime is planned. Alas their twitter account doesn’t appear to be business related.

I’m on the lookout for a replacement but not many competitors can match the open data format.

bananaheel··on Ask HN: How do I choose the right resource to learn CS fundamentals?
I found CS50’s schedule very difficult to keep up with while working full-time. I fell behind at points but I didn’t find much if any penalty in taking a little longer to finish the class. I was motivated and putting the time in to absorb the material to the point where I was satisfied.
bananaheel··on Ask HN: How do I choose the right resource to learn CS fundamentals?
Harvard’s CS50 [1] is where I started in my self learning journey. I found it very difficult but it’s given me a really good base to build upon.

CS50 hits the sweet spot of excellent online material, large online community and fun.

1: https://www.edx.org/course/cs50s-introduction-to-computer-sc...

bananaheel··on Gmail really wants me to say yes
I stopped using gmail because of this feature.