Patent for attention-based sequence transduction neural networks (2019)
patents.google.com
patents.google.com
2019-10-22 - Application granted
Should probably put a (2019) in the title. Furthermore considering how fast the ML space moves, the fact that google hasn't used this to create a model significantly better than competitors seems to show that the patented architecture did not perform better than others.
Because it seems pretty clear that the criticism is that Google is an advertising company that uses search as leadgen, and is struggling to ship AI products.
One has to ask: was that to the benefit of society?
Google may make good use of this tech. But it would be better for all of us if everyone did and didn't have to pay them a fee for the privilege.
This is reverse logic because then Google would not have published the tech and would have kept it a trade secret.
In this case, google might not have funded the development of transformers which could dramatically reduce costs of everything for humanity.
Same goes for drug development.
That said, I think there’s a question around software patents and how long they should last. Perhaps it should just be to recoup costs, plus some multiple. I’m not sure.
That doesn't seem likely to me. And if it were true that Amazon was so vulnerable that checkout flow click parity would wipe them out, I hardly think the law should intercede to save them - that seems like a scenario where they had a bad business model.
I went to a talk done by Amazon's longest serving IP lawyer and he explicitly called out that Amazon has only gone to court (or asserted its patent, can't remember the phrase) over a kindle related patent.
Provisional Application No. 62/510,256, filed on May 23, 2017
There are however other patents and applications: Neural machine translation with latent tree attention
US20180300317A1 James BRADBURY
I'm not an expert so I can't read the spec and claims as to relevance with authority. However, it still wouldn't count directly as prior art as it was published in Oct 2018. Patents are now first to file and not first to invent (to match the rest of the world).edit: They do also site non-patent prior are regarding attention. Whether this covers self attention I'm not clear and of course their claims have been reviewed by the examiner in light of the art so they're presumed valid until re-exam.
Luong et al. "Effective approaches to attention based neural machine translation," arXiv 1508.04025v2, Sep. 20, 2015, 11 pages.Luong et al. "Effective approaches to attention based neural machine translation," arXiv 1508.04025v2, Sep. 20, 2015, 11 pages.There is an investor's grace period, as long as your own public prior art disclosure is the earliest, you get 1 year to file even under first to file. First to file refers to people filing for undisclosed inventions, disclosure acts as prior art against anyone else filing, along with the grace period for delaying your filing.
Often "wow the company detailing internals at this conference is so generous!" is actually them getting an extra year on the of patent expiration date by taking advantage of the grace period.
Whole areas will be off limits and dominated by a few companies — the patent system is a system for the olden times.
The patent system unfairly enforced incumbent advantage and needs to be reformed.
Given the recent determinations, it almost sounds like since it won't be human made, it might not be protectable? Not to mention the exceptions generally made for modifications to the IP to improve it.
AFAIK Courts generally have determined it must be "transformative" for fair use, and for patents - but like a substantive change. You can't just add a pixel in a corner and call it good.
https://www.justia.com/intellectual-property/patents/types-o...
Slightly different (was for getting a patent for ai work) - https://cdn.arstechnica.net/wp-content/uploads/2023/02/AI-CO...
If the code's algorithm is patented, you cannot get around the patent by using different function and variable names and perturbing the code organization.
I bet you're correct and we wont be able to use ML models to strip code of its patents the way we use them to strip code of its copyright protection. But I wouldn't have expected the copyright stripping trick to work either, yet here we are.
You can disguise borrowed code such that, in the first place, nobody will suspect that it's derived from anything, and even then, they won't have evidence that it's derived.
Changing a few names and superficial organization details, on a large and significant work, will likely not fool anyone.
If it's a small module or just one function, quite probably easily so. You don't even have to understand how the code works; if you just know the paper from where the original author got the code, you can just say you implemented the detailed requirements there and validated it on the available test vectors.
To my understanding patents do allow for either "improvements" or anything that achieves the same result as long as it's not the same solution as the patent?
They would certainly be a lot trickier then copywrite/IP but I think LLM's would still be able to generate possible solutions? One thing I'm thinking of for example is medication analogs - there's common substitutes you can make that achieve the same or better results that you can make.
To my understanding redbull actually did this (without ai) to modafinil with this patent - https://patents.google.com/patent/US20210380545A1/en
EDIT: Modafinil might have expired but it looks like redbull filed their patent before the expiration.
You can look at their supplied diagrams and general summary to confirm.
This patent specifically covers ONLY transformers in which there is an encoder and a decoder.
Claim 1 of the patent contains the following:
"...the sequence transduction neural network comprising: an encoder neural network configured to receive the input sequence and generate a respective encoded representation of each of the network inputs in the input sequence... and a decoder neural network configured to receive the encoded representations and generate the output sequence."
Claims 29 and 30 (the only other independent claims) also specify an encoder and a decoder. So long as your transformer network does not make use of an encoder in combination with a decoder, this patent does not apply to you.
"1. A method of generating an output sequence comprising a plurality of output tokens from an input sequence comprising a plurality of input tokens, the method comprising, at each of a plurality of generation time steps: generating a combined sequence for the generation time step that includes the input sequence followed by the output tokens that have already been generated as of the generation time step; processing the combined sequence using a self-attention decoder neural network, wherein the self-attention decoder neural network comprises a plurality of neural network layers that include a plurality of masked self-attention neural network layers, and wherein the self-attention decoder neural network is configured to process the combined sequence through the plurality of neural network layers to generate a time step output that defines a score distribution over a set of possible output tokens; and selecting, using the time step output, an output token from the set of possible output tokens as the next output token in the output sequence."
Attention based models don't necessarily need to be sequence to sequence. They can be classifiers, decoder only, etc. Attention is just one tool in the ML architecture toolkit.
It's obviously math.
This whole "computer implemented invention" workaround is a complete sham.
To think that the EU wastes billions annually on this broken institution, while completely failing to properly fund startups is simply infuriating.
- cornered the market on AI researchers
- has the most researchers
- the best researchers
- has a head start on the GPU issue with TPUs
- has patents on the models everyone else uses
- already has distribution
- developed the models everyone else uses
- has models years ahead of everyone else (e.g video ones)
They sure are likely to be a filure, they're just late to the party. Keep in mind Apple is always late to the party.
While I would agree with you the CEO of Google doesn't know what he's doing, Google has the best of everything to succeed in the race. It is incredible they are doing this all openly and sharing their research with everyone while doing it.
The complaint with OpenAI is they are going too fast (the 6 months letter..) with Google they were going at a responsible speed - which is where in a game theory fight with someone else going faster, would leave them in the position they are today in terms of perception, but not the loser.