Show HN: Unbabel API – Human Corrected Machine Translation
blog.unbabel.com
blog.unbabel.com
It would be great if the feedback from the human workers could get reintegrated into the translation models in an online fashion so that they get better over time. I realize they're probably outsourcing their machine translation, but that would be a terrific fully integrated pipeline.
We would love a good provider to enrich video with translations (for when we find a good one that offers machine transcription.)
Not yet, but could you send us an example?
Unfortunately SpeakerText doesn't offer non-post-processed prices, Koemei integrated it into their own product and VoiceBase didn't offer post-processing on request, which we would need for integration into our product.
Which format will become mainstream probably depends on HTML5 adoption, which is detailed here http://www.3playmedia.com/how-it-works/how-to-guides/html5-v... Currently WebVTT seems to be in the lead.
Those formats don't accommodate for timestamps per spoken word though, which would be possible with machine transcription and which I would pay a premium for.
The students themselves translate those documents as part of the course, and once a consensus is reached (a given amount of students came up with the same translation) the document is returned to the submitter, who pays for its translation.
Almost like how recaptcha works :-)
Seems a really good idea, however I looked at:
http://news.unbabel.co/fr/fobo-est-presente-a-san-francisco-...
Vs.
http://techcrunch.com/2014/01/10/fobo/
and the translation, by _5_ people, is really poor: Just look at the first line :
>Maintenant, vous sauriez déjà que Craigslist n'est pas bon comme un lieu pour vendre votre produits.
This translation is unpleasant to read and has some mistakes. Google translate gave me a better translation.
I guess the problem comes from the fact that translators are not native in the target language. When you use the product, can you request native language translators ?
For me, having a native translator seems a must have, I hope your community will grow enough for it to be possible!
"Per word" of the source language, or the target language? Sum of both? What about languages which have a different concept of "words" in written text (e.g. Chinese, Turkish, ...).
And by the way... "cent" of which currency? :)
Edit: I just saw that the list of supported languages does not contain languages with "exotic" types of word boundaries (yet).
That's to say that the service is super interesting, but I guess the final user still has some manual editing to do if he/she wants to use this kind of translations in a professional environment. Given the price, it's still a great deal.
One last thing: the table in your home page shows that the translation from English to Italian is not available (Italy's listed only under the "from" column and not in the "to" row, if I'm reading that correctly).
Good luck with the product :)
As it stands, we have an awesome community of volunteer translators who take care of most our needs, but sometimes in times of high demand we could use help getting through a large batch of loans to translate. That said, when we tested out some external services for leveraging machine translation and translation memories, what we found were a few problems that keep us from being able to leverage external solutions
1) Our volunteers don't like "post-editing", meaning what you are doing here of a human manually fixing up a machine translation. Since you are paying your translators, I imagine they don't mind though.
2) It seems the majority of companies are focused on English -> Foreign Language, whereas the vast majority of our translation needs are Foreign Language -> English, and this proved decisive with most of the software being geared in such a way.
3) Our partners are often in remote areas with not always the highest level of education, and often are writing in a language that is their second language (say perhaps French in a Senegal where the person's native language is Wolof), so the grammar of the French is not going to be great to begin with. This throws off the machine translation and makes it nearly impossible to develop a translation memory that is segmented in the right way to actually produce usable translation suggestions.
4) We need to review the text for policy guidelines (say for instance a partner puts in directions to a business by accident in a region where our borrowers are anonymized for safety reasons). But if we send a translation out to a service like yours and then just have the English back, and then need to report an issue in it back to the partner, the reviewer who would just know English would not be able to communicate back to the partner the issue and identify it in the original language version.
Anyways, just some food for thought in what we've had trouble with in the space of trying to help us get our lenders connected to our borrowers by providing them accurate translations of the borrowers' stories.
One thing that we could be useful for is to actually help your reviewers to communicate with the partner, that is to help in translating the communication itself.
In any case, Kiva is an amazing organization, so obviously we would love to find out how we can help in any way.
On a side note, just something that I find all the time. The unbabel blog does not link anywhere to the main site and I had to type unbabel.com manually to go there. Isn't this something that you guys care about in terms of traffic source ?
Thanks for the heads up about the blog. Corrected at the end of the post.
(nvm, I think my idea is differentiated enough that I still might make it some day)
- I don't think that starting with machine translation to maintain coherence in style is a good idea, while AI is still in it's infancy. Things like sentiment detection, NLP, etc are still too infant in my opinion -- this is baked into the premise of the idea as a whole... We still NEED humans to write good translations - it seems unreasonable to start at the assumption that you will get high-quality output from the imperfect machine process that you are trying to improve (if that makes sense).
I think at best, you will START with bad style, at worst, people will essentially re-translate the chunks to make more sense anyway, and you're left with the hodge-podge.
- If your method WAS suited for medium/long-form, I would suggest adding another tier of worker-bee: the proof-reader. Allow worker bees to apply/become proof readers, and create multiple proofs for large documents. These workers would have qualified for longer-form proofing and possibly editing. A possible increase to the relative pay of the proof-readers (as they are even more closely linked to your revenue and customer satisfaction, and are doing more work to boot), and providing multiple or a combined proof to the customer (up-charge for this) would be a great addition to what you already offer. This will probably do wonders for quality control, and will remove the problem above (I think, to the extent humanly possible). This also gives the people who work with you chance for improvement, chance to build a personal brand, and a chance to take pride in their work (and maybe even build personal/business relationships/trust that benefit the company).
- Why not play in all the vertical space that you guys are in? Part of my version of this service dictates a flat rate for a certain length, and a CLEAR indication that that kind of service is for people with small blurbs to translate. Some companies only need to translate small blurbs (disconnected paragraphs, tag lines, etc), and could benefit immensely and constantly (if you make a brochure for your company, or even an earnings report, etc, you would need this service EVERY month/year, for example). I don't think you would have to make too many structural changes to accommodate such a group of potential customers.
- I have not operated a system like this at scale, so all my suggestions are largely baseless (keep that in mind please)
Oh and my idea was to rid the world of "Engrish", especially at the corporate level.
We have been thinking about having the editor position, we are experimenting with the concept and how it fits with out current workflow.
Anyway, thank you for your ideas, really cool.
If you want to offer that for 1 cent per word you are going to get exactly what you are paying for. 90% of qualified translators are bad enough that I would never let them translate anything for me. I cannot imagine the remaining 10% will work for 1 ct/word.
Also what if you want to translate multiple snippets of text, and keep them somehow consistent? (for example some .po files from a project, translated one entry at a time).
We have an endpoint for bulk translations. We are working on a way to submit XLIFF and PO files directly. Part of the consistency is achieved by the first step of MT. We are working on keeping consistency between editors by propagating their changes.
https://s3.amazonaws.com/unbabel-assets-production/img/chart...
Italian - Portuguese has a "n/a" style dot, but Portuguese - Portuguese translation says "Can take some time".
Let's say I wanted to Unbabel something from German to English. How long is 'Can take some time?'
Also, how long does it take to Unbabel something given 'Regular service' conditions?
Last question, how long before 'Unbabel' catches on as a verb?
Also payment in BC or similar would be great.