75 karma · joined October 11, 2011
It addresses key challenges in multimodal AI development:
- Managing diverse inputs - Scaling Generative AI apps - Ensuring extensibility
Built on Ray for seamless scaling, Aana offers a unified framework for multiple data types, easy integration with popular ML frameworks, and a modular architecture.
Base: https://huggingface.co/mobiuslabsgmbh/Mixtral-8x7B-v0.1-hf-a...
Instruct: https://huggingface.co/mobiuslabsgmbh/Mixtral-8x7B-Instruct-...
Shout-out to Artem Eliseev and Denis Mazur for suggesting this idea ( https://github.com/mobiusml/hqq/issues/2 )
The 2-bit version can run on a 24GB Titan RTX.
In terms of perplexity scores on the wikitext2 dataset, the results are as follows: Mixtral: 26GB / 3.79 Llama2-70B: 26.37GB / 4.13
- Blog: https://mobiusml.github.io/hqq_blog/ - Code: https://github.com/mobiusml/hqq - Models: https://huggingface.co/mobiuslabsgmbh/
No data calibration needed, extremely fast , works on both language and vision models!
* Why does it matter? Quantization significantly reduces GPU memory requirements but degrades the quality of the models. Having faster and more accurate quantization methods is extremely valuable for the ML community.
* Approach: Sparsity-based error formulation between the original weights and their dequantized version. We used a Half-Quadratic solver to derive a closed-form solution that is 100x faster than backprop via Pytorch's Autograd.
* Quantization speed: ~ 1 minute for Llama2-13B ~ 4 minutes for LLama2-70B (over 50x faster than GPTQ)
* Findings: - Larger models quantized to 3/2-bit outperform smaller full-precision models with similar or lower memory requirements. - Successful 2-bit quantization requires a lower group-size (e.g., 32 or 16) and compression of both the zero-point and the scaling factor for lower memory usage.
While we acknowledge our view might be slightly biased, we genuinely believe that our work will significantly benefit the open-source software (OSS) machine learning community. Code and model are in Apache permissive license.
It is not distilling the model, it is reducing the model weights on the fly and uses LoRA for training/fine-tuning. After the training phase, we explain how to merge the LoRA weights with the pruned weights to achieve faster inference speed
In a nutshell, we've managed to reduce the model's parameter count by up to 50%, double the training speed, and increase inference speed by 1.25 times.
For those interested in the technical details or looking to replicate our results, the code is openly available for community use and contributions
A personal anecdote is that a few years back the automatic door sensors in my university did not work on my skin tone.
Publications ( number, when, where, citations) is the primary currency/value in which one is judged within the peers in academia, and reputation outside the immediate academic community has a much lower weight. Whereas for online market places solid revune is the first priority and then comes reputation ( which is a means for the higher reveune). In academia it is the reverse, with reputation ( in a small clique) being the primary motivator, and funding being the means to gather it.
A PhD system trains you to think about unsolved problems in an given domain deeply with a larger time runway. The end goal is not a tangible product that reaches millions of people, but rather a set of ideas that can take a crack at the unsolved problems in your field in a novel way. A good work should inspire others in the field, and eventually a larger audience to pick them up and expand and build on top of it. To give a small example, a majority of the fundamentals of machine learning was charted out by many, many PhD works over the last 40 years. Implementing a linear classifier is 2 lines of code in 2018, but many Bothans died to bring us this information :-) .
The goals of industry are more immediate. Expect for a privileged few research labs in industry, your work is expected to be monetized, and rightly so. The goal is for you, if you run the business, else your management team to first figure out a problem of high relevance and monetary value. Build products/solutions for that problem, that can be used by someone who is less versed/ambivalent of your technical solutions. Efficacy of solving that particular problem often defines the merit of your contribution.
The fundamental of choosing the PhD or industry should be taking stock of what kind of contribution you want to make as an individual. If it is a few set of ideas to science, which on a later date might become something fundamental in our understanding of the world, then PhD is a good path. If it is a set of contributions towards a product/solution that eases the pain of many users then go into the industry first.
The work is not at all contradictory to Adorno, especially in the sense that it is explicitly trying to as non-reductionist as possible, and assuming notion of aesthetics is a dynamic entity .
There is a finite pattern in the dataset; more interesting, it has its interesting share of subtleties ( for example, as opposed to a image classification problems), and the technological question is whether we can capture these.
But there is another interesting data question. For our work, we curated our training set with the help of expert curators. But the dataset itself is a metamorphising entity; i.e. it is subject to revision ( it is a continuous process for us at the moment), but more interestingly it is a chance for open debate between our curators. In some sense, technology allow to codify and challenge our notion of aesthetics ( especially with the evolution in our training sets) at a given point of time.
In this case, please trying to take shortcuts towards their goal of improving the amount of publications and grant approvals. There is reward in the system for this kind of behaviour. You being a reviewer/jury position unfortunately do not have the luxury of a filter.
Is there a way to early catch this , by looking at past trends ?
Slight detour is that this is one of my rationale for spending time reviewing papers for journals/conferences. In average, only 10 to 20% of papers I review really stands out or appeal to me, which is correlated to acceptance rate of a top journal/conference.
Hardness/fun factor in puzzles is a matter of personal taste. Hence the ability to personalize is very interesting. Not everyone like to solve 10000+ piece puzzles , nor color line adherence, but people invest time and effort in solving them.
But you are right, some of the puzzles can be super-hard ( for example, the Seurat puzzle ) that we used to joke between ourself to name our paper "taking the fun out of puzzles".
Personally, what was fascinating for me is the shape of the puzzle curve it produced. Most of the common puzzles are grid based (i.e. four neighbours - up , down, left, down ). But in this scheme, there can be strange neighborhood pieces, with even stranger shapes.
One being an advise I got from one of my PhD advisors: All creative tasks might appear that it requires enormous amount of courage and effort. But usually it is more like a kitchen sink heaped with a lot of unwashed dishes. Chances are that once you wash one dish, you will end up cleaning the full lot; and you often get a strange form of pleasure while you are performing the task.
The other one is this essay http://www-rohan.sdsu.edu/~psargent/Mills_Intell_Craft.pdf on intellectual craftsmanship by Wright Mills. I do now a days actively collect memories of pure immersion and pleasure I experienced while my craft got exposed and exploited to its potential. The thought of me improving as a craftsman, coupled with these memories is a powerful self motivator to me. The shit feeling I gets when I waste my time is another reference. One of the potent lessons was also that craft can be improved only by dedicating time ( which is pleasurable); and by disassociating the end result and fears. The toughest part is to replay this logic while I find myself slipping into vortex of non productivity, but that is something I can work on and probably in my control.
While the pen falling from a desk do point out to the incompleteness ( non-Godel sense) of the standard model, which is widely accepted ( http://home.web.cern.ch/about/physics/standard-model : last paragraph ), it does not falsify it. Science is full of open holes, and no one knows ( my bet is against) that it will be completely patched up; but it is the best form of reasoning we have in understanding things, and its ongoing goal is to seek explanations that with the least amount of uncertainty possible.
Good, grief. No!!! It is a way of managing uncertainty and saying something with a precision that is available at a given point of time.
> How many discoveries are being overridden by new discoveries coming from future ?
This is beauty/and USP of science. Every scientific proof is always open for scrutiny and revision in light of new data or discovery ( tenants of falsifiability kick in here). That is, it tries hard NOT to be dogmatic by being provisional. For example, science says that we are confident Higgs Boson exists "accounting for one-in-a-million chance on the contrary" ( 5-sigma).
Let me flip your argument on the converse; success rate at which we could make ground breaking theories [ like evolution, theory of relativity , uncertainty principle ] ( which is standing the test of time for extended period of time) using the scientific method is sheer staggering and amazing. The methodology has accelerated our progress and understanding by leaps and bounds which no alternate system has managed to do so, so far!
> The amount of data accessible to the people in the past is a lot more when compared to current.
I lost you completely here. Can you please elaborate and the rest of the paragraph. ( My belief: If you take 20 random guesses; one of them turned out to be true; it is more likely to be a coincidence than a mystical insight. If on the contrary, the Monte Carlo filter I routinely simulate might just be the most insightfully entity I have encountered ).
> The division between religion/science is very small
Epistemologically they are apples and oranges! Falsifiability is not applicable to religion nor is it is provisional and routinely advocates absolute (and imho dogmatic) reasoning!
The way we approach and do science has evolved drastically ( http://en.wikipedia.org/wiki/History_of_scientific_method ). For example empirical falsifiability which is one of the primary tenant of modern science is less than 100 years old, but forms an essential part on how we do science now a days.
Parent comment's point being; while we may be trying to understand the same principle/phenomena, not only the data available to thinkers that time was very sparse compared to the present; but also the level of rigour applied was of significantly lower standards. While there might be scattered scientific truth in the vedas ( or any other religious document) ; it is insolent to believe that it is good reference manual for scientific knowledge.
The fact that Switzerland is a really rich country, with the wealth distributed not so unfairly, makes this experiment even more interesting ( compared to the communist style takeover which happened in past in then poor nations including Russia or China).
To summarize, one of the things that makes capitalism work is competition; and money necessarily is not the only thing ( and might not be the primary thing) that we compete for!
I remember getting my first PC at the age of 10, and a feeling of empowerment it brought me. Be it writing my first program that did sometime substantial, or playing Prince of Persia ( and clearing those levels ); what was substantial was the feeling that I am in control of this machine, and I can add/mould things as I want it! This sense of empowerment ( and later the ability to make use of the technology ) was possible only due to the latent fun quotient associated whilst using it (and it is best unspoiled by lack of supervision/or other's deciding how I should use the technology ).
This is very pertinent in an Indian context, where most act for "empowerment of poor" is coupled with the assumption that "poor are incapable of making their own decision".
To bring out the real magic out of techniques like deep learning ( http://en.wikipedia.org/wiki/Deep_learning ), availability of large training sets and the infrastructure required to crunch them are a pre-requisite. Once you have that, it is turning out to be a different ball game all together http://deeplearning.net/2012/12/13/googles-large-scale-deep-.... It turns out that groups like google research are the ones at present which have access to such dataset and infrastructure.
I also predict the reverse shift to happen within few years, once the interesting fundamental research problems has been tackled such people might move back to universities. If that happens, that is indeed a healthy process of academica and industry supplementing each other.