Nick Rothwell Hello, and welcome to the Sound On Sound People and Music Industry Podcast channel with me, Nick Rothwell In this episode, I talk to Gérard Assayag, who is an IRCAM research director, head of the IRCAM Music Representation team, founded in 1992, and former head of the IRCAM research lab STMS, a joint lab between IRCAM, CNRS and the Sorbonne University. His research interests range from machine learning and computational musicology to computer-assisted composition and interaction. So hi, Gérard. Thank-you for joining us on the Sound On Sound podcast. I assume you're in Paris at the moment, is that right? Gérard Assayag I'm in Paris. I'm at home in Paris. Nick Okay. So yes, it'd be nice to get some background really I guess. So background on yourself, like an origin story, your education, your kind of skills and interests. Could you kind of fill us in on that? Gérard Sure, sure. So I've been working at IRCAM for many years. I got there in the early '80s. I was still a student. I was studying computer science after studying and practicing music and a few other things. And I went there to assist a friend, a composer who needed some computer insight. And I began to develop tools for the composer that were kind of successful and then finally I never left this house, IRCAM. So there I created a team called Music Representations, who was initially specialised in computer-assisted composition. So we designed a few very popular tools, like Patchwork and OpenMusic, that are still used by many composers and musicologists around the world and that deal with the symbolic and structural aspects of music. So these tools are offline. These tools are not designed for the stage and for live interaction, but mostly for conceiving and designing musical ideas and musical structures. And more recently, I moved into the realm of improvised interaction with machines, which is a totally different field, a totally different issue, but still related to the idea of handling symbolic structures in music. So the originality of our approach is that we're not only dealing with sound processing, like the people who are designing for sound effects for instance. But we're trying to deal in the very time of the stage, in the real time of the stage, we're trying to deal with the internal structures of music, trying to understand, have the machine understand what's going on at the structural and symbolic level while the musician is playing, in order to be, so that's what we call machine listening, in order to be able to react in a significant way. But I guess we're gonna talk a lot about this. Nick Yeah. I'm gonna pull back a bit and just for the small number of listeners we might have who don't know much about IRCAM, could you also give some introduction to IRCAM and what else is going on there and how what you do fits into that. Gérard Yeah. Actually, IRCAM is a big thing. It might be the biggest facility in the world for music research, where researchers, engineers and artists are gathered together in order to work in the same place and exchange ideas and interact on a daily basis. And IRCAM is both a research and a music production house. So it has its own concert hall, which we can say a word of later because it's a very original and unique place in the world. And this institute, because IRCAM means Institute for Coordination of Acoustic and Music, was created by the famous conductor and composer Pierre Boulez in the '70s. Pierre Boulez was, at this time he was a conductor in the US, in New York and several other places and he was called back by the French government at this time. They were creating the Centre Pompidou the modern art museum and they wanted to have a kind of a department for music. And they called Pierre Boulez to create this department and organise it. And actually Boulez did much more because he created an independent institute that is related to Centre Pompidou, but it also quite independent and is actually in a different building than the Centre Pompidou. And that's still around today. Nick So IRCAM started in what, 1970 something? Gérard Yeah. It was '77 I think, the foundation of IRCAM. I think the project of Centre Pompidou and the idea of IRCAM was born in the early '70s. And IRCAM was opened in '77, I guess. Nick Okay. So you've been there a while, but your background is more on the computer science side. Is that right? I was a I was a, a musician. I studied, music and musicology, and I went into computer science driven by music actually. Because I wanted to be able to model music and have the machine be able to get some music what they call now machine musicianship. That there is have some intelligence of music and be able to understand music at the, at an intelligent level and be able to under-analyse and generate music at that level. But actually, my first my first exploration into what, what's called now artificial intelligence, but, at this time, we were just trying to get the machine to do, funny and fancy things, was not in music, was in poetry. So I was still illiterate in, in computer science. I was, I was, studying, music and, but, in, in the early '80s, and we had these microcomputers coming into the game, so it was a time where Apple was, bringing the Apple II. Microsoft was bringing the PC, and there was a, a company called Commodore who, brought a computer called the PET, PET Commodore- I remember that. Which was the... You remember that? Yes, I do. You remember that? We had so... Yeah, so I didn't have a computer, but a friend the, a friend's father had a company in, Paris, on the Champs-Élysées in Paris, a business, and they had a PET Commodore for the for the business, and they would let us use the office at night when nobody was there in order to get the machine to the machine and, and write our own program. So I wrote this program for generating poetry, for generating poems, that was based on Noam Chomsky's and Hall phonology Generative phonology. They had, they'd written a book called Generative Phonology, drawing on Chomsky generative theory, but applied to the phonology, that is the sound, the sonic level of, of language. And we used the, the model, the formal model described in the books in order to write a program that would be able to generate sequences of phonemes that would have some nice rules of euphony and try to bring the meaning, not by a semantic model but just by refining the sonic properties of, of the language. Because if you, if you constrain, the sonic properties of, what your, you, you generate to be in the real, realm of what the language produces, you're also increasing the probability of having a meaning. And so we generated a kind of a collection of poems a collection of 13 poems which had this very, this very interesting, sonic structure. So, as you see, it was not very close from music because sound was the important thing at this time. So after that, I decided that I had to study seriously computer science in order to be able to understand what I was doing actually. And then I went into computer music. Okay. So this makes sense with some of the notes I have here which... Well, you, you say in, your writings that you're working symbolically with music. So I guess if you started with the symbol- symbology of poems, then that kind of makes sense. Yes. Yes. The language. Yeah. The language structure. Yep. Okay, so you moved into IRCAM from that. So what were your, your kind of first experiments on the music side then? You know, what was your, what was your path through that in terms of the technology and the outcomes? So, when, when I get to, IRCAM with this, composer of France, my first action was to design a system, that was able to handle music notation, because at this time they didn't have that at IRCAM. IRCAM was created with the kind of the utopia of create, of creating sound, of creating unheard sound and they, they became very proficient in that. But I thought that we should not, let the classical, the traditional approach to music, which is based on, music notation. we, we, we shouldn't let this down because composer was were actually using that by hand in order to write their pieces. Composer at IRCAM were, were writing hybrid, what you call hybrid pieces. That is pieces for instrument, for classical instrument with scores, but also an electronic part that would play along with the musician in a passive way or in an interactive way. But they were still writing the score in a traditional way, and I thought that the, this visual and symbolic representation of music should also be modeled inside the computer, so the composer could handle the whole spectrum of the music representation. This is why a few years later the team I created was called Music Representation. That is from the idea, which can be represented by schematics, to the score, which is a visual and symbolic representation. To the sound and to, to the signal, which is a mathematical representation of the sound, and to the acoustics, to the physical acoustics, which is the way sounds actually unfold in the physical space. I thought that we should be able to model, to have a vision, an integrated and, and coherent, vision and model of the whole spectrum. so, so I brought this, music notation tool, and the, the, and then the musician were able to use a computer to design their own scores, but not only write their score, just say write down their own idea, but also write computer program that would generate the scores. be able to create an idea as a mode, as a general model, and then explore all the possibilities through the scores that were generated, all the possibilities that could unfold from this, this idea. And this was totally new at IRCAM. So this is what I broke. And, and I was still, at this time, I didn't have a, a job at IRCAM. I didn't have really a position. But the, the, but the institute was very open. They would let me a, a seat in a, in an office with a desk, and I had access to the computer, and I was programming this. And then it was actually very successful. And after a while, they proposed me to stay for good in this institute and I created this team. The team was created in 1992. Mm-hmm. Okay. So in the early days, this was non-real time. Yes, you, yes, you were rendering out scores. Exactly. At that time- This was absolutely non-real time. Mm-hmm. So we were the non-real time team at IRCAM, which was which was kind of specific, because all the rest at IRCAM was obsessed by building new machines that would be able to handle the signal processing in real time. Okay. That, that was, that was IRCAM's signature. So we were both marginal in the mainstream of IRCAM, but also bringing something very useful and very original, so we had our place. Okay, so at some stage You, well, became more interested in this idea of co-creativity. Was that happening at the time, or was that kind of a later thought? So that's much more recent because this idea of co-creativity was born ... so we had ... There, there was a first wave in the, when was ... Yeah, it was born in the 2000s. There was a first waves in starting in 2004, where we begin to build software for improvised interaction. We were not talking about creativity and co-creative at this time, but just trying to explore this idea of having software that could, actually answer it to musicians. And the first iteration of this software was called OMax. And, the, the name OMax came from Open Music, which was a non, non-real-time, compositional software that we that we gave the nickname OM, like Open Music. And Max and you know Max. Yeah. Max is the software that everybody use all over the place for real-time interaction. Max, by the way, was designed by Miller Puckette at IRCAM in the, in the late '80s, early, '90s, and then, was commercialized by Cycling '74 and became a, a worldwide, success. So a lot of people were using Max for interaction, and a lot of people were using Open Music for composition. We brought them together and called the system OMax. So we had Open Music computing the kind of high-level symbolic structural things that came into play in the listening and interacting process. And Max was computing all the real-time signal processing and interactive, part of the system. So Because even in real-time interaction, and even in human musician and human improviser, there are processes that are not fully real time. That is, a musician could listen to the environment, get a picture of what's happening, get an idea of the thing he's gonna do, and then maybe do it, but it might happen a few seconds later. Sometimes it's instant, some sometimes it's half a second later, and sometimes it could be one minute later because the idea has kind of, matured, inside the brain. So there are many things that, that can be done by a system which is not fully real time but has the power, the computing the language, the computing language power for processing high-level ideas, including, of course, the use of AI, techniques. So that was done by OpenMusic, and OpenMusic was reacting with Max, and as soon as OpenMusic has something ready, it would call Max and say, "Hey, I, I have this. Would you like to play it?" And in the other direction, Max would tell OpenMusic, "Hey, I just got this from the environment. You think it's interesting, and could you elaborate something from it?" So we had this OMax system, which was kind of successful too and was the first iteration and then had many, many other software as descendants until the software that we recently developed and that we use a lot now, which is called Somax for some obscure reason that I can explain later. And meanwhile, between OMax and Somax, there was this idea of co-creativity which came. At the same time, I got a very important grant from the European Union called an ERC grant. ERC stands for European Research Council. So they give research grants at three levels. At the junior, for junior researcher, for middle researcher, and for senior. And I had the, I had the senior grant because it was quite recently. It was five years ago and that was, I was not young anymore. So I got the senior grant, which was a great news because a senior grant gives a lot of money. It was, like, 2.5 million euros. And they, they give this grant on behalf of a personal, research. It's not, it's not a grant for a lab or a consortium of lab and, and industry, but it's on the head of a researcher in order to acknowledge a life trajectory in research and also some new ideas that the, that the applicant is, is bringing. So the project was called REACH. So the project was awarded in 2020 and unfolded, between 2021 and 2025. It just finished, last December. It was called REACH for Reaching Co-Creativity in Cyber Human Musicianship. Quite a complicated name, but it gives a nice acronym and, and this is where I, I fully developed my idea of, co-creativity. So at the time I was, writing the grant the application, people were talking a lot about machine creativity, artificial creativity, computer or machine creativity. And it was really, really developing well in the universities. You could see chairs of, artificial crevit- creativity inside the music technology, and music, depart- masters in, for programs and, departments. And I was wondering if this was really the right way to approach the m- the musical AI in the sense that it's difficult to use the word creat- creativity in, in its full meaning when it's applied to a complex system like a, a, a computer or an AI model that is not a subject in the philosophical sense, that is not a subject in the sense that it has a, a self-consciousness and a, a reflexive action upon its own action. Maybe it will be coming, but it was not the case. It's not the case yet, and it was not the case at this time. And I thought that maybe that was what that what they call in philosophy, Oh, I forgot the name. Oh, oh, sorry, an aporia. An, an aporia, you know, in philosophy, it's a kind of it's, a, a concept that bears its own contradiction. So it's not totally, it's not totally consistent. And also I was watching how music AI was unfolding with all the systems that are now totally proficient, like Suno and all these generative representation machine learning models that are that Totally based on the idea to replicate what's already existing in a very fancy and brilliant way. But they, they are trained on a lot of existing music, and basically they can produce more of this existing music. And of course, this was not at all the concept at IRCAM, where we're always trying to create new music, new things, new, new ideas. And this is where I had this idea that the salvation was in the very concept of interaction. That is, that it was in the very interaction process between the musicians and the machine that the new ideas and the creative, ideas would emerge and unfold in a way that is consistent with the way it happens in complex systems. So you know that in complex system, there are emergent, phenomena this is the, this is true in, for instance, in the living organism in, in biology. So you have, you have properties, for instance, you, in the cell, the, in the human, in the, human and animal cells and vegetal cells. You have pro-properties that are hi-highly complex, and that you cannot totally, fully explain by understanding the constituent processes that contribute to this property and be-because the property is an emergent property. So I thought that this was exactly what was happening in music when you had sophisticated systems like AI systems that were sophisticated enough to, to react continuously to a musician playing, and musician, which are obviously by themselves very complex systems that also listen and learn from artificial systems and change the way they behave depending on what they the system produce. So you have this double the, this crossfeedbacks loops because both systems, generate, produce, and learn at the same time and change the way they're producing and learning depending on what the other is doing. So obviously, you have an exponentiation phenomenon, and at some point, you have new shapes, new sonic shapes that appear that are interesting, that are st-stimulating enough for the musician and also for the machine, but that cannot be fully explained musically if, if you listen separately to the streams produced by the independent sources. And this is, and this is exactly what I call co-creativity. I guess there's also some kind of neuroscience here as well. For example, musicians who get into certain flow states when they're creating. Flow state. Yeah. So, so in that... Yeah, it's very interesting to mention the flow state because in, in that case it's a particular kind of flow state that is shared. Ah, okay. Because, because usually the flow the flow state in psychology is linked to a single subject, as the subject enters the flow state. But in this case, it's actually the system, created by the interaction between the different subjects that, that create a kind of shared, distributed, or joint action based, flow state. And it's very interesting because I didn't, I didn't think of it. It's an... I think it's an idea that would be worth exploring, yeah, the, the distributed flow states. Yeah. Well, I guess also musicians who play, you know- In a band, when there's more than one musician on stage, they go into their Exactly. Exactly ... playing space, and they collaborate perhaps without realizing quite what they're doing. And that was the model. That, that, that was the idea. That was the model, I think. Mm-hmm. The band, yeah. So to clarify here, you talk, we talk about audio analysis, but you're also, your core interest here is musical representation, so do you still have underlying this notions Yeah generating scores and working with those scores? Yeah. The, yeah, this idea. So it's not exactly a score in the sense that in the, in the unfolding of the live experiment, you don't... If it's improvised, you don't need scores, although some musician do use now dynamic score that are that can be generated on the fly, so you can absolutely integrate all this. But we, we do use symbolic structures underneath the model. Mm-hmm. So what the musician are playing is first analysed audio-wise by signal processing, but then transformed into symbolic units and symbolic structures, and the system interact at the symbolic level before transforming back the symbolic, its own symbolic proposals into sounds. So, and all this happens very quickly. So people think that it's all done in audio but at the core of the system, we, we, we still have the symbolic structure that we could absolutely translate into scores. And actually, we have some, some composer now who are interested in both composition and improvisation, who are really interested in having these scores generated on the fly, and generate proposal for the musician on stage using iPads or any, display. So can you actually explain what these symbolic systems and scores are like? Can you kind of explain that in, in simple terms? We're not just talking about dots on a, on a, a stave and a musical score here. We're talking, is it gestural? Is it some kind of symbolic description of the sounds or the way they're performed or what? Yeah, so for now it's a series of, of descriptors of the music and the sounds, such as the pitch, the chromas, the MFCC's, which are the Mel-frequency Cepstral Coefficient, which is an audio, descriptors, and other audio descriptors, and that was the, that was the first layer. And, and the, the, the audio stream from the musician was segmented using all these descriptors into a discrete collections of, units. And this is where we get into the symbolic, the symbolic domain. But now we're more and more interested into the gesture, so we developed-- we first developed, r-real-time a, a new machine listening system which can detect and recognise, classify and recognise in real time the instrumental playing techniques in, in-inc-including the classical playing techniques or, or, or the advanced playing, techniques, such as, you know, tremolo, staccato, vibrato, multiphonics, flutter-tongue, et cetera. So every instrument has his catalog of, playing techniques. So now we have this, this machine list-listening model based on deep learning, on machine learning, that can do the job in real time and inform, Somax the AI agent, in real time of the, what the musician is, doing. So we have now a layer that is above the basic descriptors, su-su-such as, such as pitch chromas, MFCCs, et cetera that, that gives a hint at the intention of the musician, because if the musician is switching from a flat sound to a multiphonic, that's the system, the system is aware of the, this intention. And now we are aggregating this information into higher level gesture by integrating in longer time spans this information, in order to bring the idea of a musical gesture. And then it could be bigger and become a, a compositional, gesture that, that would be structurally above the level of the sound that the musician can produce. So we're moving more and more toward the idea of gesture, and the gesture can move to toward the idea of shape, of, musical shape. And, so we're, we're integrating, at, at a higher and higher level the musical information and bring this to the system to help its decision process, and also anticipate. So integrate the music that ha-has been played in the past on a longer time span and create anticipation, in the future also over a longer time span. Yeah, 'cause I guess the big challenge with AIs and LLMs at the moment is they don't have a good way of analysing time. They can't kind of form a shape of a piece of music over a long time. Yeah. There, there are two drawbacks with the LLM. So, which the first one is that they use language prompts, where in, in a dialogic process, which is not the, the, which is not the time, and they do not integrate time in the sense of the musical time that is actually unfolding in the live, experience. So this is what we're trying to attain. And I used to say that we have a conception of musical AI where the prompt is music. So we don't, so, so we don't want to, to... because when, when you're in the live ex-ex-experience, of course, musician do not c-communicate by issuing prompts and request using language, but the request is music, and it's not-- It's sometimes in a dialogic way, but it's usually not in a dialogic way. Music is a polyphonic process, so things happen simultaneously. So if you're trying to design an AI system for which the prompt is music, and the prompt is coming continuously, and the answer is generated continuously and in parallel to this prompt coming, then you get into trouble because nobody's doing this in the industry because they, they are not interested in that, which is nice for us, because otherwise you cannot compete, I mean, with, you know. So for now, the industry doesn't have its eye on this because the big thing in, in the industry is to be able to generate offline commercial music that you'll be able to use in a, in a commercial, situation. And it's fine, so lets them do that and l-leave us this, niche where we are, we are working. But it will be coming. It will be coming, and I know that Google has already issued its first interactive model for music, where they pretend that m-musician can play, you know, directly in real time with the model, and the model will play along with the musician, which, in a way is, the same kind of stuff. So, so it's coming, and I heard what they're doing. It's not, it's not yet fantastic musically, but it will be coming. So it's a race, in a way. Okay, but it sounds like your particular edge here is the fact that you have this symbolic representation at the centre of this. So that means you can separate your analysis from your response. It also means that symbolic representation can be passed between artists. It can be manipulated in other ways. So I guess that gives Yes ... more flexibility. Yep. In a way, in a way. Okay. We'll see. So okay, we're kind of coming to the end here, but something I, I caught on one of the websites, which I just want to quickly ask you about, slight change of topic, is you had a, a conference series or an event called Improtech. Was that last year? Oh, yeah. That's very important. Improtech is very important. We created Improtech in 2012. It's a kind of festival. We call it a workshop festival because it's a mixture of concert and workshops, lectures, demos, et cetera. And it was, dedicated in the first place to improvisation with machines. So we had concert with great musicians and always, machines coming into the game. So for instance in New York, we invited George Lewis who played with Jerry Allen, great jazz person, with his System Voyager. So we tried to identify people who were proficient in this, activity of bringing together musicians and machines. We had Vijay Iyer with Steve Lehman and Bernard Lubat, who is a great French, musician and oh we, and David Wetzel performed with Roscoe Mitchell. Oh, wow. So from the... Yeah, so from the very beginning, we had the idea of inviting really the greatest improviser you could think of, who were interested in this experiment of improvising with, with machines and also having, workshops and lectures around it. And then we, we did it again, but, it was very irregular. It was not a yearly or a biannual, festival. We just did it, again, when we had the opportunity and the support and, and, people wanting to host it and help in organising. So it was New York, 2012. Philadelphia, 2017. Athens, Greece, 2019. Then there was the, the COVID break, and the next one was 2023 in a little village in France because there is a there is a great, improvisation festival in this village in Uzès with Bernard Lubat. And then it was, and then it accelerated. That was two, 2023. We had Tokyo, 2024, and IRCAM in 2025 was the last. And always with a, a very high level of, musicality and a very important musician coming and a particular form of concert, very long, like two hours, two hours and a half, three hours, and many, many, many short pieces. so, so it's a very improbable festival because it's very expensive because we invite everybody, the musicians, the lecturers, the workshops. So we have to have a very good support for that. So the last one, I could organise it using my grant, my ERC Reach European, European grant, and the next, the next one, we have no idea because we need this kind of, of support in order to be able to do it. So the interesting in Improtech is this is really the best place to, to see the most, proficient and the most recent music AI systems dedicated to stage, to improvisation and live music. Mm-hmm. And are you also or artists making recordings using this, this, machinery as well? Oh yeah, everything is recorded. We, we have always recorded the concert with a very high quality, and also shoot, high quality, videos, and it's all, it's, and it's all online. Super. And you can, you can find it. There's a YouTube channel for the ERC RICH project and everything's there. Excellent. Well, this has been super interesting. Thank you so much for your time. Yeah, I mean, I'm really looking forward to seeing, what you get up to and how your work progresses over the next few years. As I say, it just sounds super fascinating that you've made this stuff work. It sounds like it's been a massive challenge, but it sounds like there's been great fruitful outcomes as well. Thank you, Nick Thank you for listening, and be sure to check out the show notes page for this episode where you'll find further information along with web links and details of all the other episodes. Oh, and just before you go, let me point you to the soundonsound.com/podcasts website page where you can explore what's playing on our other channels. This has been a Project Cassiope production by me, Nick Rothwell, for Sound on Sound.