How to create minimal music with code in any programming language
zserge.com
zserge.com
> For example, here’s a familiar riff “E B D E D B A B”. It uses only 4 notes - D and E from one octave and A + B from the other, lower octave. If we look at the note frequency table, the pitches of those notes in octaves 5 and 4 would be 659.2Hz (E), 587.3Hz (D), 440Hz (A) and 493.8Hz (B).
Instead of equal tempered pitches (which are generated by repeatedly multiplying by the 12th root of 2), use pitches that form whole number ratios with respect to the root of your key.
A decent 12-note chromatic scale would be something like 1:1, 16:15, 9:8, 6:5, 5:4, 4:3, 45:32, 3:2, 8:5, 5:3, 15:8, and 2:1. So, for instance if your root is C=256 hz, you'd have 256 hz for C, (25616)/15 for C#, (2569)/8 for D, and so on. This is a just intonation (JI) scale. An advantage of JI is that the musical intervals blend with each other better, whereas the advantage of equal temperament (ET) is that you can use any note as the root and change keys on the fly without causing problems. If you're doing electronic music, though, there's a lot less reason to stick with ET: you can always just multiply your whole tuning table by a constant if you want to switch key.
JI scales can be made with any ratios: the above maps well to most conventional music, but there's no rule that you have to limit yourself to 12 steps per octave. Using the harmonic series is another option, or adding in ratios that have a 7 in them.
Someone in the comments pointed out that the analysis was done assuming even temperament, and the singer was actually totally nailing just intonation.
It's one of those anecdotes that has stuck with me since as a reminder to try to not get tunnel vision on a specific metric.
It sounds much more "in tune" to the ear, because of the nice ratios.
It really wouldn't surprise me that raw recordings of many singers are in fact more just than equal temperament.
That seems more plausible than that people naturally sing in an artificial tuning based on repeated multiplication of the twelvth root of two, even if that's the musical context that people are most exposed to these days. Those whole-number ratios are just too good to resist, even if a singer isn't fully aware of what they're doing.
All in all, I think it's safe to say that string players and singers prefer slightly flatter thirds in harmonies than 12-EDO -- which is what the above amounts to -- and deal with commas accordingly.
My wife and 1 daughter both sing in acapella groups, but their various different groups over the years have never done anything in JI, although my wife also sings Balkan songs which are whole other story of course.
Maybe a portion of university-trained musicians might know it from their music history classes (or maybe they're Adam Neely viewers), but I believe most professional and amateur musicians have no idea about just intonation or equal temperament.
In practice, 5/4 major thirds are at the very low end of acceptible major thirds, with the Pythagorean 81/64 at the top. I find slightly sharper major thirds to sounds better generally [for western common practice harmony].
An interesting system, which I wish I've had more time experimenting with before suggesting it, is 55EDO, using a system where each whole tone is broken into nine equal parts; a five part interval is a major semitone and a four part interval a minor semitone. The major scale is made from whole tones and two major semitones. A C# is slightly flatter than a Db in this system, and you only have to worry about enharmonics (old meaning: futzing with the small distances between similar pitches) in complicated chord progressions.
Not suggesting not to try JI of course, I'd recommend experimenting with it (including the suggested 7-limit JI as well)! Scales in the end can only be judged in the context of what you want to do with them -- they certainly all have tradeoffs.
I ask because I'm an electronic musician working in a DAW (Ableton mostly) and am trying to find the best workflow to start exploring these concepts. Ideally, there would be an interface for switching between tuning systems that's as easy as drawing in a time-signature change, but micro-tuning seems to be a low priority for most DAWs. The only one I know to even begin integrating alternate tuning systems is Logic, and those settings are tucked away deep in the preferences.
Does anyone here have a resource for starting to familiarize oneself with how to integrate alternate tunings into their composition? The simpler the better, I suppose. As soon as one sets their foot into the world of alternate tunings you're usually flooded with 40+ tunings. As someone who's worked in Equal Temperament their whole life I would like to pick ONE system and really drill down on it until I get a grip on how to integrate these tunings into programs that are built for ET.
Personally I have experimented with alternate tuning using SuperCollider. It definitely doesn't qualify as simple, but the upside of code-as-music is that if you can imagine it and express it, you can do it. I recently made a track that switches tuning systems periodically by responding to a particular MIDI CC. I have to imagine that those sorts of shenanigans would be pretty hard without some code.
To live a bit more in Ableton-land, you might be able to find a Max device that does something similar-- maybe a MIDI insert that converts MIDI notes to note + pitch bend?
The only role for a DAW in this comes if the DAW labels MIDI data with note names. If you're not using 12TET, these labels will likely be wrong. But since MIDI itself doesn't inherently use 12TET, just 128 notes numbered from 0 to 127, any DAW can still be used to work in non-12TET tunings and scales.
> NoteOnEvent Struct Reference
> Detailed Description
> Note-on event specific data.
> Pitch uses the twelve-tone equal temperament tuning (12-TET).
https://steinbergmedia.github.io/vst3_doc/vstinterfaces/stru...
You can, of course, implement a synth that doesn't follow the spec, if you like.
(This documentation is for VST3, which does not use MIDI. VST2 has been deprecated for years.)
You missed the tuning field in the struct, which effectively makes it use 12TET only as a reference, not as the only tuning standard.
From what you say, it seems like the VST2 spec sets a standard scale to fill in this gap, which is interesting. It certainly would help with interoperability.
Anyway, I learned something from the clarifications arising from each of you correcting each other :-) (Though it's probably not worth continuing discussing whether there's any difference at all between "deals in 12-TET" and "uses 12-TET only as a reference" since you both seem to mean the same thing by these by now.)
A standard workaround if you want to do microtonal MIDI is the one note per channel trick: each simultaneously-sounding voice is on its own channel, which can be independently pitch-bent.
For instance, if you want to play a note 10 cents above middle C, you pick a MIDI channel you're not using right now, send a pitch bend command to raise it by 10 cents, then send a note-on for note 60.
This gives you 16-voice polyphony (because MIDI has 16 channels). The downsides are that this only works on multitimbral synths (i.e. it works with most "romplers" but very few analog synths), and you need a convenient way to set all the channels parameters the same.
The one note per channel trick can also enable individual note volume swells by using channel volume. One note per channel has been standardized recently in MIDI as MPE.
There's also MIDI 2.0 now, which has per-note pitch bend among other features. Time will tell if it actually gets widely adopted, and whether synths actually implement the full spec or if they leave out features that only a small minority of their users care about.
His idea was actually pretty cool: you start out with an F, C, and G major chord tuned just like in the ratios a couple comments up, which gives you an entire C major scale, and then there are a handful of accidentals. The syntonic comma (+) is what turns A into the fifth above D rather than a third above F, # applied to the top note of a major third turns it into a minor third (defined by the division of a major triad into a major and minor third), 7 applied to a Bb lowers it to coincide with the 7th harmonic of C, 11 applied to an F slightly raises it to be the 11th harmonic of C, and so on. Each accidental also has an inverse operation. With just these accidentals above, the system is able to uniquely represent every ratio whose prime factorization contains numbers up to 11.
Unfortunately MIDI is not so good for non-12-EDO. Some instruments support SYSEX messages with per-note-number alterations. There's also a program out there, forgot the name, which takes a MIDI channel and splits it across multiple channels to apply independent pitch bends per note when there's polyphony. I hope MIDI 2 will help...
I don't know much about real DAWs, though there are some Adam Neely videos exploring microtonality and alternative tuning systems where he pulls up some plugin -- maybe that will help you find something.
Having seen a video of Jacob Collier's projects, it seems like he might do microtonality using only pitch correction. I think if you're familiar with your commas -- there aren't many after all -- you just learn how many cents you need.
Edit: I thought to look up Sevish, a microtonal musician who seems to actually care about making good music, and he has some resources on his site: https://sevish.com/music-resources/
Perhaps you mean alt-tuner? http://www.tallkite.com/alt-tuner.html
Well this is a fun rabbit hole!
On the violin, it's common to adjust intonation to different concerns: to bring out multiple voices, to put emphasis on the melodic line, or to nail the chords.
I don't know how hard it would be to do this digitally though. I imagine it would be quite tedious to work on a fretless digital music software.
Here's a visualization I like, that uses 22 notes: https://www.youtube.com/watch?v=jA1C9VFqJKo
Speaking of large-number EDOs, some people I know locally are doing guitar conversions to 41-EDO. The trick that seems like it absolutely shouldn't work but it does is to omit half the frets, so that only half of the notes are available on each string.
The Well-Tuned Piano [1] also has a really nice lattice, though the Wikipedia page doesn't reflect this yet. (The B7+ should just be an A. I think scholars have taken La Monte Young too literally when he said everything was an overtone of an extremely low Eb.) Here, just drew a picture: https://imgur.com/S5OndL6.png (light gray is the piano key it's mapped to). It's illuminating taking the "chords" in Kyle Gann's paper and seeing them on this lattice, which shows major and minor septimal triads in each triangle.
However, not only must one adjust the tuning each time you want to switch keys, just intonation actually doesn't work for all intervals even within a single key: https://music.stackexchange.com/a/22017. This is something I often don't see mentioned.
If you're programming the music entirely in software, you could avoid this by basically forgetting about the notion of notes/pitch classes/scales/keys entirely and focus solely on the frequencies and ratios. I wouldn't be surprised if a lot of the best music of the distant future is partly created in such a way (by humans, AIs, and/or both working together). But if you're working in a DAW, the best option may be to consider a few intervals off-limits, if you want everything to be pure.
The probelm is solved in JI by re-defining the key of C major to include both 10:9 and 9:8. So, instead of one D, you have two slightly different Ds.
Equal temperament just uses a note almost half way in between to play the role of both notes, but that's just an approximation. (This double-role can be useful for certain chord progressions that work out in equal temperament even though mathematically they shouldn't.)
JI definitely becomes a lot more useful and interesting if you allow for more than 12 notes per octave. I'd consider 16 to be about a reasonable minimum for conventional 5-limit music, but having more opens up more possibilities.
In the context of the original article which seemed to be geared more to just implementing some sort of basic sequencer or picking random notes from a scale, I think picking notes from a JI scale rather than 12-tone equal temperament is an easy way to do something different from what everyone else does with a good chance of having a result that sounds good. More serious music composition or performance in JI generally requires more awareness of the various quirks of JI versus equal temperament. I do wish tooling was better; in some ways, MIDI kind of locked everyone into 12-tone equal temperament, and getting out of that rut is going to require some kind of large change. (MPE and Midi-2.0 are possibilities; time will tell if they pan out.)
https://www.youtube.com/watch?v=6NlI4No3s0M
Exploring the lattice, JI and ET:
https://www.youtube.com/watch?v=I49bj-X7fH0
https://www.youtube.com/watch?v=M5OJgsHXSmY
> the Tonnetz (German: tone-network) is a conceptual lattice diagram representing tonal space first described by Leonhard Euler in 1739.
If you want to go a little deeper than the article, here's a few more notes:
The way the author generates the sawtooth and square waves is "analytically". That means treating them like a function of time and simply calculating the waveform position each point in time. As you can see, it's really simple. Unfortunately, it also introduces a ton of aliasing and will sound pretty harsh, especially at higher pitches. If you're familiar with nearest-neighbor sampling in image resizing, think about how nasty and chunky the resulting image looks when you scale it down using nearest neighbor. Analytic waveforms do the equivalent thing for audio.
Fixing that problem is surprisingly hard. It's a lot like texture rendering in games where there's a bunch of filtering, tricks, and hacks you can do to make things look smooth without burning a ton of CPU cycles.
---
The clever trick the author uses to simulate strings is: https://en.wikipedia.org/wiki/Karplus%E2%80%93Strong_string_...
---
The low-pass filter they show is a simple first-order 6 dB/octave digital low-pass filter. Filtering is fundamental to electronic music. The majority of synthesizers you hear in electronic music use "subtractive synthesis" which means to start with the simple sawtooth and square waves introduced early in the article and use filters to tame their harsh overtones. I find the math behind filter design, especially the moving filters used for synths, really difficult, but also interesting.
https://www.izotope.com/en/learn/digital-audio-basics-sample...
`tail -f output.txt | aplay`
I know this is a lame takeaway, but I'm just so happy to learn about this.
aplay was the missing link.
What's the common API for video across OSes?
What's the common API for "put this string in a message dialog on the screen" across OSes?
And meanwhile there's JACK, RtAudio, PortAudio and several other crossplatform APIs, most of which have wrappers in non-C languages (despite the problems with realtime audio that this can lead to).
And I disagree that it's deadly simple... Even sound synthesis (where it's "just math") is very hard, because you have to watch out for aliasing. Same if you're changing pitch or doing anything other than playing back a WAV.
Also, it's hard-realtime, which limits the programming languages you can use. Then, if you're not using the OS audio mixer (lots of different reasons not to), you gotta do the mixing yourself, which is another can of worms.
By the way, the article you linked is really good, thanks for that. Anyone else reading this discussion will surely be interested in that article too.
Sorenson is low key one of the most innovative minds in the scene. Extempore is such a blast to work with.
http://debu.gs/entries/making-music-with-computers-two-uncon...
My only contribution is to point out that the late Hudak was very interested in making music with software and spent a significant portion of The Haskell School of Expression, https://books.google.com/books/about/The_Haskell_School_of_E... on it.
However, this is false:
“ that’s why CD music had a sample rate of 22000 Hz. Modern sound cards however tend to use sampling rates twice as high - 44100 Hz or 48000 Hz”
CDs are/were 44100 hz, they were never 22000, because you need two samples per wave and because you need slight over sampling and to make the conversion from 48k in to bad math to discourage DAT machine as a lossless copy format for the masses.
What do you mean by this?
But regardless, you need to slightly over sample when doing A/D conversion because you need to do a very strong low pass filter before you hit the sample rate to avoid harmonic distortion. So you need the filter to be some amount higher than the target maximum reproduced frequency, because the filter is in the analog domain and itself introduces audible artifacts.
So you must sample at a higher frequency than you intend target. 48k makes sense because you have more ‘extra’ frequencies to work with, so it’s easier to get the filtering out of the audible space.
> Recall that we couldn’t just throw it on a computer and resample it back then.
Of course you could. How else do you think CDs were mastered and produced? The CD was developed with the idea of using it on computers, and for computer storage. And if, in practice, someone wanted to copy it, they'd just record the CD to tape using analog out (not quite lossless, but nobody cared back then), or copy CD/CD with no resampling required.
The 44.1k is a mostly historical accident, and it goes back to the Red Book CD audio spec being reused from an earlier project for digital audio that Sony was working on, the U-matic, where 44.1k was basically how much audio would fit evenly into some number of lines in the tape's storage.
I don’t think we are aiming at the same concept. I was there, as perhaps you were too.
Average consumers did not posses, could not afford, technology to duplicate CDs easily or to convert a CD to DAT.
I can’t recall all this after so much time has passed, but the gist is that somehow the record industry influenced the deployment of technology to reduce lossless duplication ability for consumers.
It wasn’t until home computers became so fast and ubiquitous that this changed.