Sad to see all the non Chinese open source models being at least one generation behind.
244 karma · joined November 30, 2015
Sad to see all the non Chinese open source models being at least one generation behind.
Interestingly they explicitly mention AI generated voices, does that mean voices generated by traditional TTS engines are fine?
My goal was to help volunteers that were in the field in Nepal communicate in English -> Nepali and back. Even though this was somewhat effective, there was still a communication gap because most people in Nepal in remote parts could not even read in Nepali.
I looked around for solutions but couldn't find any Nepali Text To Speech solutions. The builder brain in me fired up and I decided to build a Nepali Text To Speech engine using some of the groundwork that was laid by Madan Puraskar Pustakalaya (Big Library in Nepal) which they had abandoned halfway.
I spend all night hacking along to build a web app that let the volunteers paste translated text and have it spoken. The result was https://nepalispeech.com/ and the first iteration of this was built in just 13 ish hours.
I hope the people that got affected by the earthquake are in a better situation now.
I left this feedback a bunch of times with different people in the org and I really hope they scrap their janky "frameworks" and dev tooling. They should just move to more standard open source tools options that evolve and get better quicker than barely maintained internal tooling.
Open Source tools also have bigger communities and resources online to debug and solve issues, compared to mediocre documentation from internal wiki. If absolutely needed, make thin wrappers above open source tools. Some other teams within Amazon had the luxury of using better tools, but I bet a lot of people are in the shoes I was in.
I had to switch from OpenAI's models to GPT-J because OpenAI's policies were restrictive. How did you get around that? My guess is that since you are only outputting Emoji's it might be allowed.
Now imagine a machine learning model doing this instead of humans, it could infer so much more information about how you think and how your brain works based off of that data and do even more harm. Ever thought about buying something and an ad for that casually pops up in your feed? It is going to be similar to that, but instead of ads for a product it could be some misinformation that could cause paranoia (just an example).
I will gladly take a system that can drive itself confidently under 40mph in city, and use basic lane assist (and maybe lane changes and exit ramps like in Teslas) for highways. Cant wait!
Tailwind let me quickly create any UI I want, and Daisy helps me reduce the amount of time needed to style basic elements like inputs and all.
Firebase lets me get a low latency real time data source for my application, that can scale infinitely. It also handles some other parts of building an application that normally take a lot of time: authentication, storage management etc. And the pricing is really really cheap once you consider how much it costs to create an infrastructure that scales as well as firebase does, unless you model your data wrong and end up using a lot of db read/write cycles unnecessarily.
The backend API piece is not missing, you can use either Next.js API or firebase functions for backend piece, I use those for things like stripe billing backend etc.
This stack is enough for most projects out there, and when its not enough its flexible enough that you can integrate it with other things. And that timeline I mentioned was for a production ready MVP.
I am not completely sure if its the codec Apple uses that is making the difference or if it is because the source is lossless, but bass sounds punchier and instrument separation is much better with Apple Music compared to Spotify's best quality.
I was really looking forward to Spotify's lossless that was announced (I think it's called Spotify HQ?), but I haven't heard about that in a long time.
Edit: If you plug in your Airpod Max with a hardwire cable, the difference is even more noticeable, even with iPhone's internal DAC.
Also want to clarify that the differences and not so pronounced that you will be able to tell right away, but if you listen to Apple Music for a while and try to listen to the same song on Spotify, it will not sound the same. It's one of those cases where you don't know what you don't know till you know it once.
"CSS modules still seem to solve every problem Tailwind solves, and better."
- Not necessarily true. Unlike css modules, tailwind removes the whole "think about a name for your class" mindset, reducing friction from the development process. It also unifies some base level design decisions like spacing and colors, which developers would have to rely on "best practices" otherwise, which don't necessarily get strictly enforced.
"I'm not sure how Tailwind works, but any CSS that's built at runtime and JS and inserted into the DOM dynamically should be avoided, and is an example of favoring developer experience over end user experience."
- You are right, looks like you are not sure how Tailwind works. Tailwind does not build anything at runtime, it all happens at build time. Tailwind will compile only the things you need (using the new JIT mode) into a css stylesheet which is sent to the frontend. Not much different from how sass or scss works.
"CSS modules let you use the full power and control of vanilla CSS, without having to worry about styles bleeding across components."
- Tailwind does not stop you from using vanilla css, but in most cases you do not need to. As per their website, you can think of it as an API to use parts of CSS, instead of CSS replacement. I think you are confusing Tailwind as a replacement for something like CSS Modules, but those two are completely unrelated. You can still use CSS Modules while using Tailwind. Think of it as an api to your design sytem just like you could think of an ORM as an API to your database.