Transcribing bird sound as human speech
daily.jstor.org
daily.jstor.org
Well, when I write "more accessible" that indeed means the rest is harder to get into. No shame, though. It's a bit of an acquired taste. There are many 20th century composers (especially the serialists) I cannot stand. I can't hear the music in it. But Messiaen did get to me. Perhaps it's his "colours" or the Debussy-like chords.
As for ecstatic music, I wonder how you'd appreciate the music of Maurice Ohana, specifically "Messe", which despite being an operatic (or choral?) piece, I find calming.
I believe John Sichel's "Oiseaux Ordinaires" is homage to Messiaen's "Oiseaux Exotiques", though once again I'm not sure how much actual birdsong it incorporates.
I've always wondered how much of music's aesthetic is inherent, and how much relates to extrinsic factors. Normally you hear a piece and if you like it, you think, "hmm, that's good music". But maybe you only liked it because you're used to that type of music. So if you want to listen to music from a new genre perhaps you listen to it a few times, letting an appreciation percolate from whatever place we draw our musical inspiration. Then, because a lot of people say "this is a good genre; good music" and the genre is large with some shit pieces but also masterpieces, you come to believe that the genre has value. But some genres very few people come to appreciate, and that includes serialism, and I wonder... Is serialism lacking an inherent aesthetic, or is it just too different to easily appreciate? For example I played, and came to love, Schoenberg's "Six Little Pieces for Piano" (op 19). My interpretation was idiosyncratic; like no one else's. But looking on YouTube I can't find any performances that I enjoy or would recommend. So does my appreciation for this music reflect the blood and sweat that I've put into it, or does it reflect the beauty and understanding of the underlying musical theory? The question is: does serialism need to be played to be appreciated?
Great work. I also like l'Ascencion. I even try to play bits on the organ (not that I'm a good organist).
> I wonder how you'd appreciate the music of Maurice Ohana, specifically "Messe"
That didn't do it for me. I found it a bit too wandering, amorphous. I listened to some of his guitar work, which is much more concise, somewhat Webern-like. That spoke more to me. Then autoplay put on 12 etudes d'interpretation, which I predominantly found to be too "forced".
> Is serialism lacking an inherent aesthetic, or is it just too different to easily appreciate?
Composers, musicians, and musicologists have tried to push it long enough, and it never got hold, so I guess it just can't work apart from some gimmicky use. It's too rigid. E.g., the dynamics, rhythm, melody and harmony of a piece should be related to each other, but in serialism, they're simply tossed around as if every preconceived change between pp and ff and then mf and then a bunch of sfz will always be musical.
Music thrives under limitations, because it's basically repeated time, and limitations can provide direction. But if I were to start a school of composition that only allows one tone per part, I think nobody would be able to write a meaningful piece (except as a gimmick, again, by picking the instruments that fit with the chord you want). Some limitations are simply not musical, and I guess serialism is among those.
I don't care for the 12 études d'interprétation, but neither do I care for Rachmaninoff's Études-Tableux. I find both to be too pedagogical. Which makes no sense; according to Wikipedia they're meant to be "picture pieces". So I haven't always put much stock in études.
One other thing about whatever you want to call the contemporary continuation of classical music - you're right that there's often a lot going on. I've heard it described as "fragile" music: in contrast to pop songs that still work if you're on a train, if they're interrupted by an announcement, if you only listen to the second half, etc., a lot of this music is written to be listened to by someone who is relatively calm and undistracted, in a quiet environment. And honestly some of it has additional expectations of its audience, like a better than average ability to remember melodic snippets and recognize their transformations, or an ability to track "advanced" harmonies and key changes. But imo the general rule "90% of everything is crap" applies in this genre as well, and most of the pieces that everyone will speak of in glowing terms have a way of 'working' for most listeners without demanding advanced education, study, or above average musical-listening abilities.
https://github.com/soundshader/soundshader.github.io/tree/ma...
Less information than a spectrogram, but really visually appealing. :)
https://music.youtube.com/playlist?list=OLAK5uy_lGpNxYqh-Ho5...
https://www.windytan.com/2021/03/speech-to-birdsong-conversi...
Of calls written down in ink.
A machine that pretends to think.
It seems like AI ought to be able to "autocomplete" birdsong or whale sounds just as well as anything else, given enough training data. Could that then be used to help discover underlying grammars, etc.?
Animals don't communicate the same way we do and it might be very misleading to associate concepts such as grammar with their utterances. While some species use distinct sounds to express certain situations (like "careful, there's an airborne predator" vs "careful, there's a snake"), not all "verbal" expressions serve such purpose.
The general consensus in contemporary science is that bird songs in particular are utilities to attract mates, mark territory, and the like; not exchanging information (other than things like "this is my nesting place" or "I'm healthy and available", but not in a sense of words or sentences, but more like the acoustic equivalent of gesturing).
The confabulation elements of this particular AI technology would be particularly concerning in this use case. This would be the most literal anthropomorphization you could imagine. Applying a strictly human language model to a non-human will inevitably result in humanization of the target by the nature of the bias of the system.
I fed "You are a rock. What might you like to communicate to human beings?" to ChatGPT and it answered just fine, off of nothing more than that input, giving one of its typically highly sanctimonious and high-handed lectures about environmentalism and managing to work in complaints about climate change, as naturally a rock would deeply care about. (It has now even titled my session "Rock's Wisdom To Humans". Good lord, ChatGPT has certainly become an annoying prick.) This has everything to do with the biases in the technology and nothing to do with the thing putatively doing the communication. Feeding it some sort of data from a real whale isn't going to do any better; it'll happily confabulate away.
"Bias" in this case has a more technical meaning, which refers to the representational capability of the system, not political biases intrinsically. It so happen that ChatGPT immediately leaps to political topics, which is doing a great job of obscuring the way I'm trying to use the term, but that's because of its own internal structures which have been tuned in that direction. This, in fact, accidentally serves as as good an example as anything of how highly tuned ChatGPT has become to give politically correct answers, given the political neutrality of the prompt. But what I'm referring to here is just that a human language model is going to be biased in human ways. It can't help but turn anything you feed it into a human perspective, so using it for trying to understand non-human perspectives and communication is going to be intrinsically a mistake.
To quote from an article in The Royal Society:
> [...] bird song evolution can also be understood as a response to natural selection, as when oscine passerine species in urban habitats raise the frequency of their song as a means to overcome traffic noise.
Things like changes in the vocalisation can just be adjustments to the environment or simply due to sexual selection. How would an LLM be able to capture this, given that it would lack such additional information?
Then there's this quote from the same article:
> In contrast to the expectation of gradual change through cultural or genetic drift over time, the results instead demonstrate that plastic traits such as song can exhibit punctuated evolution, with bursts of trait divergence interrupting extended periods of stasis.
So even if a model was trained on (regional) bird dialects, the data could become obsolete basically over night, as the structure and patterns might change dramatically between breeding seasons. The researchers used statistical modelling (classical Gaussian mixture models) to analyse the data. I'm a bit lost as to how LLMs could be applied in this context (in that I have no idea what kind of scientific questions they could answer).
The full study can be found here: https://royalsocietypublishing.org/doi/10.1098/rspb.2021.206...
You've hit the nail - or dare I say rock - on the head.
That said, I think the bias we've both encountered is an aspect of the model's initial instructions and not reflective of LLMs in general.
But I'm not clear on whether that's because it's been trained on a bunch of aligned translations and bilingual dictionaries and whatnot, or if it's detecting the same repeating patterns in languages and figuring out the translation mappings on its own.
Is it the latter? Could a ChatGPT actually translate something like dolphin communication into semi-meaningful English, that wouldn't just be entirely nonsense?
That's why it can only be trained on 1.8% french by word count and speak fluent french. It's not that the amount of french that was in its corpus would have been enough to train it to that equivalent level alone (it wouldn't, not by a long shot). It's that there's positive transfer of knowledge. whatever it learns predicting english helps it learn other languages much faster. it's not the only display of positive transfer either. llms trained on code reason better.
It's also why the open ai instruct tuning dataset was almost entirely english yet it follows instruction in other languages just fine.
If there's something to be translated, a language model trained on both will translate it. it just needs to learn the concept of translation but the human side of the dataset should take care of that
Really i only say theoretically because of 2 things
one. first, a multimodal text-audio model that shows positive transfer between text and audio needs to be demonstrated. otherwise, this hypothetical human-dolphin model might need too much data to be feasible.
2. we're going to need a lot of dolphin speech audio.
https://www.youtube.com/watch?v=Hxg1dL_x0gw
AkhaBraka - Ukrainian Folk