AI Is Learning Arabic Music. It Is Also Flattening It.

Generative platforms are absorbing Arabic music faster than anyone can audit it — rounding quarter tones into Western scales, defaulting to two dialects, and settling ownership questions in courtrooms where no Arab rights holder has a seat.

by

Ghurba Team

20 min read

20 min read

Saint Levant Is Building a Career Out of Refusing to Choose

Stability AI put it plainly: most AI audio models have never heard a maqam.

A Voice Returns, and a Legal Vacuum Opens

In May 2023, the Egyptian composer Amr Mostafa posted a 32-second teaser of a new song carrying the voice of Umm Kulthum, generated with artificial intelligence, nearly half a century after her death. Her family filed a complaint with Egypt's prosecutor general. Alam El Fan, which holds rights to much of her catalogue, objected publicly. Within weeks the dispute was settled and the complaint withdrawn, resolving the individual case and none of the questions underneath it. Egyptian law had, and still has, no specific provision governing the synthetic resurrection of a voice.

The episode is usually told as a story about one composer's judgement. It is more useful as a preview: the technologies now absorbing Arabic music operate in a space where the region's laws, institutions and rights holders are largely absent. What the machines learn, and what they flatten, is being decided elsewhere.

A condenser microphone lit against a dark background

What the Models Learned From

The two most prominent generative music platforms, Suno and Udio, have never published their training data. When the major labels sued both companies in June 2024, the platforms' court filings acknowledged training on copyrighted recordings while arguing the practice was fair use. There is consequently no public accounting of how much Arabic music sits inside these systems, from which countries, in which dialects, or under what claimed authority. That opacity is the baseline fact of this subject, and every claim about representation has to live with it.

What can be measured is the research field the commercial tools grew out of, and there the numbers are stark. A 2024 study titled "Missing Melodies" found Global North music outnumbers Global South music by roughly six to one in the datasets used to train and evaluate music-AI systems, a ratio that worsens slightly in commercial applications. The same study notes that the dominant symbolic formats underlying many models cannot express notes outside the Western twelve-tone scale at all.

The Quarter-Tone Problem

That last limitation is not an implementation detail. Arabic music is organised around maqamat, melodic frameworks whose identity depends on intervals the Western scale does not contain. A system that quantises pitch to twelve equal divisions cannot state the difference between maqam rast and a major scale; it can only round it away. Researchers studying global inequalities in AI-generated music describe exactly this failure mode: models default to the tuning, rhythm and timbre of their training majority, and reproduce other traditions as ornament rather than structure.

The industry has started conceding the point. Stability AI, promoting a maqam fine-tuning project in late 2025, put it plainly: most AI audio models have never heard a maqam. The admission matters more than the product. If the flagship systems of a supposedly global technology require special remedial training to handle the organising principle of Arabic music, then what they produce under the label "Arabic" or "Middle Eastern" is, by default, an approximation built from the outside: strings, a darbuka loop, a vocal melisma, the postcard rather than the place.

Two Dialects Standing In for Twenty

The flattening repeats in language. Research on Arabic language models consistently finds coverage skewed toward Modern Standard Arabic and the highest-resource dialects, with dialectal Arabic significantly underrepresented. A 2026 evaluation framework, AraRegBias, tested outputs across six dialect groups and found statistically significant dialect-dependent distortions, with lower-resource varieties, Iraqi Arabic in particular, faring worst. These are text studies, not music studies, because no equivalent benchmark yet exists for song. But lyric generation runs on precisely these models, and songwriting in Arabic is dialect writing: a Sudanese song is not an Egyptian song with different spelling.

The likely outcome is not that AI produces no Arabic lyrics. It is that it produces plausible Egyptian and Levantine lines on demand, weaker Gulf material, and near-nothing convincing in Iraqi, Sudanese, Yemeni or Amazigh-inflected Maghrebi registers, and that commercial pressure then treats the two well-served dialects as "Arabic" in general. The diversity loss would arrive not as censorship but as convenience.

An audio waveform rendered in white on a black screen

The Flood Is Already Here

While the representation question waits for data, the volume question has answered itself. Deezer, the only major platform publishing detection figures, said in July 2026 that fully AI-generated tracks now exceed half of all daily uploads, around 90,000 tracks a day, up from ten per cent at the start of 2025. The same disclosures note that AI tracks draw only one to three per cent of streams, and that a large majority of the streams they do draw were flagged as fraudulent. Synthetic catalogue, in other words, functions mostly as royalty-pool dilution, a dynamic Ghurba has examined before in the context of streaming fraud.

For Arabic repertoire the exposure is specific. The region's royalty pools are already thin, as our reporting on where MENA's music money goes details, and its trade-press scrutiny thinner. A flood of cheap synthetic "Arabic-style" functional music, laid over playlists in a market where editorial oversight is scarce, competes directly with working musicians for the same pooled pennies. No platform currently publishes region-level AI-content data, so the scale of this in Arabic markets is unknown, which is itself the point.

Settlements Without the Region

The ownership questions are now being resolved, quickly, in rooms with a short guest list. In October 2025 Universal settled its case against Udio and announced a licensed AI platform trained on authorised music. Warner followed with Udio in November 2025, then became the first major to settle with Suno. Sony remains in litigation. The American Federation of Musicians has sued both majors, alleging member recordings were licensed to the AI companies without compensation or credit.

Whatever these deals become, note who negotiated them: American and European corporations, on behalf of catalogues they control, in United States courts. No Arab collecting society sat at that table, in most of the region because none exists with the standing to sit there. No MENA culture ministry, no regional label consortium, no Arabic-repertoire body has a public position on AI training licences. The composer Hani Shenouda's warning during the Umm Kulthum dispute, that the region's intellectual property law urgently needs updating for AI, remains, three years on, a warning.

The Case Against Alarm

The honest complication is that the region is not standing outside this technology looking in. Amr Mostafa is not a Silicon Valley executive; he is one of Egypt's most successful hitmakers, and he reached for the tool immediately. Regional producers use AI stem separation, mastering and vocal tools daily. Arabic-centric AI efforts, from Gulf-funded language models to maqam fine-tuning projects, demonstrate that the flattening is a dataset and incentive problem, not a property of the technology. For an independent artist in Algiers or Amman, AI production tools genuinely lower costs that once required studios and sessions players they could not afford.

It is also true that "AI will destroy the tradition" has been said before, about recording, about radio, about the synthesiser, and the tradition absorbed them all. Cheap tools have historically widened who gets to make Arabic music. The precedent deserves weight.

What the precedent does not cover is ownership. The synthesiser did not contain a copy of every recording made before it, and its manufacturer did not collect rent on the output. The economic structure of generative AI, in which CISAC projects 24 per cent of music creators' revenues at risk by 2028, is what distinguishes this technology from its predecessors, and that projection was built on markets with functioning royalty systems. The region mostly lacks them.

What Cannot Be Known Yet

Nobody outside the AI companies knows how much Arabic music their models contain, or from where. No benchmark measures dialect or maqam fidelity in generated music. No platform reports AI-content figures for Arabic-language markets. Egypt's promised intellectual property reform has produced no AI-specific statute. And there is no public inventory of which Arabic catalogues, if any, are included in the licensed training sets now being negotiated. Each of these is a decision someone will make; none of them has been made in public.

Approximation by Default

Generative systems are not hostile to Arabic music. They are indifferent to it, which in a dataset economy amounts to the same thing. Left to defaults, the models will keep producing an "Arabic" that is a rounding of the real thing: two dialects for twenty, twelve tones for twenty-four, one region for many. The alternative, models that learn the music's actual grammar under licences its owners actually signed, is technically available and institutionally unclaimed. Flattening is not the technology's destiny. It is just its default setting, and defaults win until someone with standing changes them.

Key Facts

  • Suno and Udio have not disclosed training data; in 2024 court filings both acknowledged training on copyrighted recordings, arguing fair use.

  • Global North music outnumbers Global South music roughly 6:1 in music-AI research datasets ("Missing Melodies," arXiv, 2024).

  • Dominant symbolic music formats cannot express notes outside the Western 12-tone scale, excluding Arabic quarter tones.

  • Research on Arabic language models finds persistent skew toward MSA and high-resource dialects, with Iraqi Arabic among the worst served.

  • Deezer reported in July 2026 that fully AI-generated tracks exceed 50% of daily uploads (~90,000/day), up from 10% in January 2025, while drawing only 1–3% of streams.

  • UMG and Warner settled with Udio (Oct–Nov 2025) and Warner with Suno (Nov 2025); Sony remains in litigation; the AFM has sued both majors over the deals.

  • CISAC projects 24% of music creators' revenues at risk from generative AI by 2028 (€10bn cumulative).

  • In 2023, Umm Kulthum's family filed and later withdrew a complaint over an AI rendition of her voice; Egypt still has no AI-specific voice or likeness statute.

Published

This is an industry analysis based on public data, court records and published research. Ghurba Records has no commercial relationship with any AI company, platform or label named in this piece.

Sources

Primary and institutional sources

  • "Universal Music Group and Udio Announce Udio's First Strategic Agreements for New Licensed AI Music Creation Platform," PR Newswire, October 2025. Supports settlement and licensed-platform claims. prnewswire.com

  • "Global economic study shows human creators' future at risk from generative AI," CISAC, December 2024. Supports the 24%/€10bn projections. cisac.org

  • "AI Music Tops 50% of Daily Uploads on Deezer," Deezer Newsroom, July 2026. Supports upload, stream-share and fraud figures. newsroom-deezer.com

Academic research

  • "Missing Melodies: AI Music Generation and its 'Nearly' Complete Omission of the Global South," arXiv:2412.04100, December 2024. Supports dataset-ratio and 12-tone limitation claims. arxiv.org

  • "Bias beyond Borders: Global Inequalities in AI-Generated Music," arXiv:2510.01963, October 2025. Supports cultural-bias failure modes. arxiv.org

  • "The Landscape of Arabic Large Language Models," arXiv:2506.01340 / Communications of the ACM. Supports MSA and dialect-coverage claims. arxiv.org

  • "AraRegBias: Evaluating Dialectal and Stereotypical Bias in Arabic Large Language Models," Research Square / SSRN, 2026. Supports dialect-bias findings. researchsquare.com

Original reporting

  • "Egypt: AI-generated Umm Kulthum voice sparks legal debate," Middle East Eye, 2023. Supports the complaint and legal-vacuum account. middleeasteye.net

  • "Egyptian composer ends dispute with Umm Kulthum's family over AI rendition of her voice," The National, May 25, 2023. Supports settlement details. thenationalnews.com

  • "US musicians' union files amended lawsuit against Universal and Warner over Suno and Udio AI deals," Music Business Worldwide, 2026. Supports AFM litigation and settlement timeline. musicbusinessworldwide.com

  • "Launch, Train, Settle: How Suno And Udio's Licensing Deals Made Copyright Infringement Profitable," Forbes, December 18, 2025. Critical context on the settlements. forbes.com

Background reading

  • "How do Arab singers, musicians see rise of artificial intelligence in music?" Daily News Egypt, July 12, 2023. Regional practitioner perspectives. dailynewsegypt.com

All sources accessed August 1, 2026.

Explore Topics

Icon

0%

Explore Topics

Icon

0%