Lady Napalm
18 Followers

The Data-Driven DJ (Part #1)

3 days ago
Subscribe to get closer to your favourite creators

Our heroine—a mild-mannered data scientist by day—is by night Lady Napalm, a DJ of the ill-est repute. She needs a faster way to design set lists and is consequently designing a graph database to facilitate rapid song selection.

The design and construction of this graph database will be discussed in this and the next few posts on the subject, after which our heroine will discuss the development of AI agents that produce candidate set lists using the information contained within this database.

The Challenge

Our heroine regularly performs four- to five-hour DJ sets, divided into sections by tempo and vibe based on the mood she wants to project during various phases of an event (e.g., a dramatic beginning, an energetic first half, a chill second half, and then a mind-blowing, upbeat finale). Music selection is further driven by event theme, for example disco for a “Studio 54”-themed party or psychedelic trance music for an “Alice In Wonderland”-themed event.

Each track in her song library is annotated by tempo information, length in minutes, key information, an algorithmically generated estimate of the song’s “energy”, music genre, and sometimes a Likert scale-based rating of how much our heroine likes the song. Artist and album information is indicated as well.

The challenge is to synthesize all this disparate information into a compelling set list within a reasonable time frame. Our heroine’s current method for assembling a five-hour set can take over half a day--longer if a significant portion of the songs included within the set are new to her track collection.

But our heroine has a day job and therefore needs a faster approach. To begin creating this faster approach, she’ll organize her track data into a graph database, for reasons elucidated below. But first, we must consider the DJ set list design constraints she must consider:

Set Design Constraints

Constraints apply during DJ set planning:

Harmonic Mixing

First, transitions between songs must be harmonically congruent to avoid dissonance. When beginning a transition out of a song written in a specific musical key, there are only a limited number of keys that the next song can be written in for that transition to sound good.

For example, if a DJ begins a transition out of a song written in C-minor, to ensure the mix into to the next song sounds congruent the next song must (generally) be written in one of C-minor, G-minor, F-minor, or Eb-major. (The reason for this is based on western music theory and will not be discussed here. See Harmonic Mixing on Wikipedia for the gritty details).

So when assembling set lists, our heroine must consider how songs mixed in series chain together with respect to key. Following the remainder of the introductory material regarding why our heroine needs a better way to design set lists, this post will conclude by demonstrating how to use graph database infrastructure to quickly address these harmonic mixing requirements during set list construction.

Note that the most sophisticated harmonic mixing techniques also consider songs’ algorithmic “energy” estimate, but we’ll address how that attribute impacts set planning and harmonic mixing decisions within a future post.

Publishing Rules

Our heroine records her performances and posts them on Mixcloud afterwards for the public to enjoy. To ensure copyright compliance, Mixcloud requires that uploaded sets follow specific rules per their agreements with major record labels:

A set may not include more than four songs by a given artist.

No more than three of a given artist’s songs in a set may originate from the same album.

An artist cannot be played more than three times consecutively.

No more than two songs from the same album may be played consecutively.

An exception exists when the recording is made by a DJ who included tracks that she also composed. In that case the DJ may include more than four of her original compositions but still must obey the above-stated rules for other content included in the performance recording.

Don’t Bore the Audience

A set that lasts only an hour can get away with mixing within a single genre at a single tempo. But because our heroine plays four- to five-hour sets, she must include more variety to keep the audience engaged. Keys and energy levels must also vary significantly to achieve this.

Finally, she doesn’t want to play the same song twice by mistake!

Track Data Infrastructure

To address the need for quicker set list design given the constraints discussed above, our heroine will begin by improving her song data infrastructure:

The Current System

Our heroine currently uses rekordbox to store track data and prepare set lists. This proprietary software facilitates track tempo estimation, track key estimation, and the marking up of track transition points. It presents her track library as stacked dataset, only one that can’t be filtered in complex ways.

Alternative Infrastructure Ideas

Relational Database

rekordbox offers an XML export of its track data, which can easily be reformatted for insert into a relational database. Doing so would allow far more complex queries of the track data than rekordbox allows, which would make it easier to identify candidate tracks by (genre, tempo, key) combinations using SQL.

However, one can best think of a DJ set as a “path” through the space of all possible available tracks...

...and while relational database are good for many things, they are not good at tracing pathways through the data they contain. Using SQL for this purpose would require a series of self-joins, one for each number of songs that one wants to include in the set list.

But graph databases are built precisely for performing this task in an intuitive and mathematically elegant manner.

Graph Database

Graph databases systems organize information into mathematical graphs, a pattern that our heroine believes best reflects her own natural way of thinking. Nodes express objects, each akin to an individual row in a stacked dataset. Edges between the nodes indicate relationships between these objects.

One may annotate nodes and edges (henceforth called “relationships”) with attributes. For example, a node representing a music track might contain attributes such as song name, tempo, song length, etc. A separate node representing an album of tracks might contain relationships connecting that node to each track node with a (directional) “HAS TRACK” attribute.

The image above provides an example visual representation for all the Depeche Mode remixes in our heroine's track collection, as stored within her graph database. As we will demonstrate over the next few posts, one can write queries against this data to trace “paths” (i.e., setlists) through the network, subject to the constraints described above.

Here is the Cypher code used to extract the data shown in the image above:

MATCH p=(genre:MUSIC_GENRE)<-[:TRACK_HAS_GENRE]-(track:MUSIC_TRACK)-[rtal:TRACK_HAS_ALBUM]->(album:MUSIC_ALBUM)-[aha:ALBUM_HAS_ARTIST]->(artist:MUSIC_ARTIST {name: 'Depeche Mode'}), q = (track)-[:TRACK_HAS_KEY]->(key:MUSIC_KEY) RETURN DISTINCT p, q;

While the output displayed above is graphical, for visual consumption by a human brain, we can also obtain the results in tabular form as required for final setlist construction.

Next: Our heroine will encode the harmonic mixing rules described above into her graph database to ensure their enforcement during track list selection.

Comments