MaMMUT: A simple vision-encoder text-decoder architecture for multimodal tasks
ai.googleblog.com
ai.googleblog.com
https://arxiv.org/abs/2301.12597
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
Were you expecting a response?
They're both multi-modal models for zero-shot answering.
I'm not sure what the purpose of this comment is. If no experts see this, I won't get an answer. You don't need to reply with something completely useless that adds nothing to the topic.
HN is not an AIML discussion forum. Putting 10% more work in the question could yield 400% better responses.
Linking to what BLIP2 is a good start
https://huggingface.co/docs/transformers/main/en/model_doc/b...
This reply serves exactly no purpose.
Have you looked at YOLOv8 vs BLIP2?
For my use they're kind of able to be used interchangeably in this case, depending on which can detect objects in the image and generate a caption or image/text match (BLIP2) or detect via object detection or segmentation in SAM.
0: https://lucidrains.github.io
1: https://github.com/lucidrains
2: https://paperswithcode.com/paper/coca-contrastive-captioners...
If it's supposed to stay secret, what's the point of "here's instructions for how to reproduce our big secret"?
Presumably the societal purpose of papers is to share knowledge, and the individual purpose is to take credit and win prestige.
It seems like the first purpose would be better served by also publishing code etc, and the second purpose wouldn't be harmed by it?
Look at GPT-3+, OpenAI gets fame and fortune while people struggle to reproduce their last-gen models.
They just want to get onto the next research instead of taking time to publish a clean open-source implementation, which can (and will) be done by somebody else anyways.
You are free to treat a paper from Google like a paper from any random stranger, sure. But it doesn’t change the fact that many ML researchers I know in my R1 University (and many other top universities) always mention how much insights they get from these papers even when they couldn’t always release the code/models.
Yes, its describing an architecture.
> No interactive product of any kind? The degree of Google’s inability to actually ship anything even now is totally mindblowing.
Yes, Google is comically bad at shipping AI products (mostly, from their description, for “safety” reasons).
OTOH, they are very good at putting out papers that other people turn into products, so this kind of thing isn’t without value.
Google (and other large industry labs) have the budget and resources to experiment.
They should be encouraged and thanked for publishing what works.
It's not super hard to reimplement a published paper. But running multiple experiments is hard for individuals or smaller companies.
Fixed.
Are you aware of other cases?
I don't bother to read most Google papers unless someone tells me that they're doing something astounding. Just because I know I don't have access to their models, their code or their data. So what's the point?
As a community we need to stop accepting and stop citing papers like these.
There is no science without replicability, and it is literally impossible to replicate this work. It's not worth the paper it's printed on.
It's fine if Google wants to play with its toys at home. But we should stop pretending this is research of any value.
OpenAI is killing it with ChatGPT, a publicly accessible product with research papers that have been reproduced. Facebook released huge LLMs for free. Stability released successful image models for free. etc.
Meanwhile Google has ... Tensorflow? Dying. TPUs? Only used by Google themselves or when they're free. Bard? A joke compared to ChatGPT. Imagen? Never released.
Remember when Google said they were going to have an AI call your hair salon and make appointments for you? Yeah...
But hey, at least they've got that golden mountain of PII they harvested from everyone that's been oh so valuable in building new, market defining products... It's not like small companies are running circles around them using publicly available hardware and publicly available data...
And they've still got search, a that product just keeps getting better and more useful by the day...
I think that ran into social problems, not technical ones.
https://www.theverge.com/2019/5/9/18538194/google-duplex-ai-...