94 karma · joined August 8, 2017
ml@nullpilot.de
Looks like I have some experiments to run soon.
This supposes that the music is the end goal, and the very point of my comment is that it doesn't always have to be, and in those cases "just don't do it" also means not doing whatever comes after.
Just as you state below, this doesn't replace creating music for the creation's sake. I don't believe it will, or should. It merely replaces having nothing at all, or having the 100,000th video with the same upbeat stock sound.
We would be better off if the other 99.9% didn't have worry about making ends meet, than if we do whatever it takes to keep the status quo of the 0.1% intact. That does not only go for artists.
By the same logic synthesizers shouldn't have been invented that allowed people to make advanced sounds without tediously learning an instrument first, consumers should remain priced out of microphones and editing software, etc.
Like I said, I am not trying to feign ignorance on the drawbacks of the tech which is very real and far from negligible. I am not a tech bro AI maximalist. I just do believe that hyperbole will not put the djinn back into the bottle, and pretending like there isn't a real market between nothing and paying or being a composer isn't adding anything to the conversation.
[1] https://developer.mozilla.org/en-US/docs/Web/CSS/Nesting_sel...
What I am thinking doesn't have to compete with the whole Office or Google Docs suite, or HubSpot. There is imo enough value in being an economic all-in-one email solution for "the rest of us". Point DNS there, get email, CRM, maybe even a small transactional mail quota and some marketing email features.
I have two projects I would onboard yesterday if this existed, and I have spent more than a hot minute thinking about going for it myself, just to scratch my own itch. The main reason so far that I haven't is that Migadu + Resend is enough for my needs at this time, and I can do without CRM for now.
What I think is outside their target market, but what would be a great next step for projects with slightly more needs would be a Google Workspace/Outlook sort of service with a similar approach. Simple CRM features, shared contact management, calendars, signatures, etc. Roughly what Proton is doing without the Silicon Valley pricing, or maybe even Hubspot.
As operators of app stores, Google and Apple have no incentive to remove shovelware, as they earn a healthy share of every buck made by their creators, at almost no cost.
Similarly, ad networks have no incentives to remove harmful ads, as long as nobody complains enough to cut into profits.
And above all, social media platforms have little to no incentive to remove emotionally manipulative content. It's effectively a symbiotic relationship between the content and "the algorithm".
How do we change the incentives without resorting to censorship? I think the government bodies have in part failed to address these things, because the product to be regulated is information. It's always a slippery slope towards censorship, which makes for a perfect target for lobby work.
Google pretty much only for image search and maps, and even those have better alternatives now.
I'm not sure I can follow. This isn't specific to MP4 as far as I can tell. MP4 is what I cared about, because it's specific to my use case, but it wasn't the source of my woes. If my target had been a more adaptive or streaming friendly format, the problem would have still been to get there at all. Getting raw, code-generated bitmaps into the pipeline was the tricky part I did not find a straightforward solution for. As far as I am able to tell, settling on a different format would have left me in the exact same problem space in that regard.
The need to convert my raw bitmap from rgba to yuv420 among other things (and figuring that out first) was an implementation detail that came with the stack I chose. My surprise lies only in the fact that this was the best option I could come up with, and a simpler solution like I described (that isn't using ffmpeg-cli, manually or via spawning a process from code) wasn't readily available.
> You don't need to know this any more.
To get to the point where an encoder could take over, pick a profile, and take care of the rest was the tricky part that required me to learn what these terms meant in the first place. If you have any suggestions of how I could have gone about this in a simpler way, I would be more than happy to learn more.
Looking into it, and working through it, part of my experience was a lack of resources at the level of abstraction that I was trying to work in. It felt like I was missing something, with video editors that power billion dollar industries on one end, directly embedding ffmpeg libs into your project and doing things in a way that requires full understanding of all the parts and how they fit together on the other end, and little to nothing in-between.
Putting a glorified powerpoint in an mp4 to distribute doesn't feel to me like it is the kind of task where the prerequisite knowledge includes what the difference between yuv420 and yuv422 is or what Annex B or AVC are.
My initial expectation was that there has to be some in-between solution. Before I set out, what I had thought would happen is that I `npm install` some module and then just create frames with node-canvas, stream them into this lib and get an mp4 out the other end that I can send to disk or S3 as I please.* Worrying about the nitty gritty details like how efficient it is, many frames it buffers, or how optimized the output is, would come later.
Going through this whole thing, I now wonder how Instagram/TikTok/Telegram and co. handle the initial rendering of their video stories/reels, because I doubt it's anywhere close to the process I ended up with.
* That's roughly how my setup works now, just not in JS. I'm sure it could be another 10x faster at least, if done differently, but for now it works and lets me continue with what I was trying to do in the first place.
The first solution suggested when googling around is to just create all the frames, save them to disk, and then let ffmpeg do its thing from there. I would have just gone with that for a one-off task, but it's a pretty bad solution if the video is long, or high res, or both. Plus, what I really wanted was to build something more "scalable/flexible".
Maybe I didn't know the right keywords to search for, but there really didn't seem to be many options for creating frames, piping them straight to an encoder, and writing just the final video file to disk. The only one I found that seemed like it could maybe do it the way I had in mind was VidGear[1] (Python). I had figured that with the popularity of streaming, and video in general on the web, there would be so much more tooling for these sorts of things.
I ended up digging way deeper into this than I had intended, and built myself something on top of Membrane[2] (Elixir)
[1] https://abhitronix.github.io/vidgear/ [2] https://membrane.stream/
Some examples of things I made use of:
Colors > Levels. Mostly to get noise to go across the whole range from black to white, not just some in-between. Multiple iterations of reducing the output levels and then stretching those out to the full range again can create interesting artifacts.
Colors > Curves. Same as above, and also final processing. I used a fixed set of gradients (UIGradients) and usually changed them up.
Colors > Map > Gradient Map. Combining the above with gradient mapping, or combining it with noise can result in interesting artifacts as well.
Colors > Desaturate > Color To Gray. Adds grain, texture and contrast around edges. Lots you can do with that, especially by distorting the result.
Filters > Distorts > Kaleidoscope, Filters > Map > Panorama Projection, Filter > Map > Paper Tile. With the exception of the galaxy one, all the ones in the album are somehow the result of those.
I did most of these in 2019 when I was deeply burned out, and clicking around for hours was all I was able to do. It really was kind of therapeutic.
So yeah, I am really glad GIMP exists.
Thank you to the creators and maintainers.
25h+ spacey ambient, very relaxed, almost but not entirely without words
https://open.spotify.com/playlist/0tLvg6WcRMc8bn8McRfruK?si=...
12h+ downtempo, no vocals
https://open.spotify.com/playlist/77rLJZoIo7VYjrHWk2gKc7?si=...
7h+ melodic post rock, no vocals
https://open.spotify.com/playlist/6DWXAfI1IPtzIVTXcBd1bU?si=...
Minuum does that to make sense of what you type on the tiny board. It gets better over time and I've ported my data over from the last phone. But, given how much has happened in ML over the last few years, I have no doubt it could be done much better now. I enjoy it on a per-language level, but my experience with mixing dictionaries has always been awful and more of a hindrance than anything else. I eventually gave up and switched to just swiping through languages. That has been second nature since, with basically no mental overhead.
I personally wouldn't expect mixed-dictionary as a feature, and if it's there it would have to be pretty much flawless to convince me to give up that control. Then again, this might have been a problem that particularly affects Minuum, and with no overlap between the space that letters occupy, maybe it's much more manageable.
Either way, I'm excited for what you're doing and appreciate the opportunity to geek out about this topic. Thanks for reading!