Collective Craft: How Artists Collaborate to Train AI-based Audio Synthesis Model for Music
This project studies how artists train AI-based audio synthesis models for musical performance and installations. Rather than treating training as a hidden technical step, it shows how artists and collectives make it into a situated, collaborative and craft-like practice.
Overview
Recent advances in generative AI have opened new possibilities for neural audio synthesis in music. Yet most research and tools focus on composition interfaces, live performance control or model architectures, while the act of training models remains underexplored from the perspective of artists.
This project reframes training as part of the artistic process. Data curation, model optimization and evaluation are not only technical operations: they are moments where aesthetic intentions, embodied musical knowledge and technical constraints are negotiated within a collective.
Research questions
The study asks how artists collectively make sense of opaque neural audio synthesis training processes, and how artistic intentions are communicated, negotiated and inscribed into trained models.
It also examines how collaboration is organized when expertise is distributed across artists, researchers, engineers, performers, instrument designers and other members of a creative project.
Method
We conducted semi-structured interviews with 13 artists across 8 musical projects using custom-trained audio synthesis models for performance or installations. The projects involved models such as RAVE, DDSP, WaveNet, autoregressive systems and music-to-latent approaches, with interaction modalities including voice, gesture, instruments and controllers.
The interviews focused on data collection, model training and collaboration. We then conducted a thematic analysis to understand how collectives communicate during training, how they build practical knowledge, and how they evaluate model behavior through musical practice.
Findings
First, collectives develop their own modes of communication to bridge the gap between machine learning concepts and embodied musical knowledge. Artists create shared vocabularies, notations and listening practices to make models musically meaningful and responsive to artistic intent.
Second, artists develop situated know-how through iterative training. Rather than preparing a perfect dataset in advance, they often train quickly, listen, adjust the dataset, retrain and repeat. This makes dataset construction a form of sculpting guided by aesthetic goals.
Third, musical engagement becomes a primary way to evaluate and steer models. Improvisation, collective listening, jamming and performance allow artists to test what the system can support musically and where its expressive boundaries lie.
Contribution
The project challenges narratives that frame AI music tools as automatically democratizing creative practice. Training remains complex and often requires technical mediation, but artistic decision-making still guides how models are shaped, judged and used.
By foregrounding model training as a site of interaction and collaboration, the study argues for more interactive, collaborative and musically grounded tools for AI-based audio synthesis.