SignalMatrix · Pre-seed 2026
We turn live TV and radio into quality speech data for the big languages AI still can't hear.
One pipeline. Two products.
Live TV and radio, captured as it airs.
Quality data is now the bottleneck.
“We've achieved peak data and there'll be no more.”
Ilya Sutskever, OpenAI co-founder · NeurIPS, December 2024
$14.3B
Meta paid for 49% of Scale AI, a data company, in June 2025.
$1.2B
Surge AI's estimated 2024 revenue, selling human data to frontier labs.
10 months
David AI went from a $5M seed to a $50M Series B, selling audio data alone.
Sources: The Verge / NeurIPS 2024; Meta–Scale deal reporting, June 2025; Forbes via Sacra (Surge, press estimate); David AI and Bloomberg, 2025.
Hours of transcribed training audio in OpenAI's Whisper, the reference speech model.
Sources: Whisper (arXiv 2212.04356); Ethnologue. Bars on a log scale. Kurdish, Oromo, Igbo, Zulu, Tigrinya and Wolof: also 0 hours.
Not online. On the radio and TV people actually use.
59%
of Africans get news from radio at least a few times a week, across 38 countries. Radio is still the continent's top news source.
Real speech
Talk shows and call-in radio are spontaneous, multi-speaker and dialect-switching. Exactly the speech clean benchmarks miss.
71–99%
of the time, Whisper large-v3 writes Pashto audio in the wrong script. Models fail where real speech lives.
Sources: Afrobarometer Dispatch AD1173 (Round 10, 2024/2025); Pashto ASR benchmark, arXiv 2604.04598 (2026).
Nobody has all three.
Thousands of channels, 24/7, with scheduling and failure recovery.
Who has it today: media monitors, English-first.
Multi-speaker, code-switching, and cheap enough per hour to run at scale.
Who has it today: labs, trained on clean audio.
Cross-references, entity linking and provenance for every segment.
Who has it today: nobody.
How one broadcast segment becomes data.
Live TV and radio captured as it airs, with station, country and timestamp on every segment.
→ recording + metadataSpeech-to-text tuned for dialects, with speaker, entity and topic tags.
→ time-coded segmentsEvery segment quality-checked, cross-referenced against independent sources and linked into a knowledge graph.
→ claims with sourcesIllustrative output, and the signals pulled from it.
الحكومة قررت رفع أسعار الوقود ابتداءً من الشهر الجاي.
“The government decided to raise fuel prices starting next month.”
Five breakthroughs SignalMatrix has built, where off-the-shelf models fall short.
−60%
CostCheaper per hour than Whisper-class transcription.
5,000+
ScaleChannels our ingestion pipeline is built to carry.
40+
ReachLanguages our transcription stack already handles.
8
Real conversationSpeakers separated at once. Single-speaker models break.
Vision
Beyond audioOn-screen figures tie each claim to who said it.
585,996 hours of real speech, and counting.
Live counts from the production warehouse.
Live today, deepest in Arabic, Farsi, Dari and Kurdish.
~2,000
TV channels
27
Countries
8
Languages live
Shaded: countries with recorded broadcast, from the production warehouse. Outlined: next-round countries.
Revenue now, scale later.
Verified, dialect-tagged broadcast speech, delivered as it airs.
Real-world speech for the languages models can't hear yet.
Who pays for real-world speech.
$2.8B spent by investment managers on alternative data in 2025, up 17% in a year.
Billions spent each year by AI labs on training data, rising as public data runs out.
Sovereign AI programs in MENA, Africa and South Asia, with almost no native speech data.
Serviceable market: non-English broadcast speech across 50 languages. Estimate to add.
Source: Neudata, “The state of the alternative data market in 2026,” February 2026.
Design partners working with the current corpus today.
Investors just funded this category. We own the half nobody's taking.
Every hour we collect makes the next one cheaper.
Past broadcasts are gone once they air. Competitors can't get them back.
Broadcaster licenses take time to sign.
Dialect-level ASR and data you can trust are both hard. This team has done each before.
Every new language is a new corpus on the same pipeline, not a rebuild.
8
Live todayArabic and its major dialects, Farsi, Dari, Kurdish and four more.
20
This roundPashto, Urdu, Hausa, Swahili and more high-speaker, low-data languages. First paying contracts.
50
NextAcross Africa and Asia: Amharic, Yoruba, Somali, Bengali and beyond.
Pre-seed on a SAFE, room to $4M. About 24 months of runway.
$3M
Use of fundsVision
The default supplier of real-world speech for every language AI can't hear yet.
50
languages
5,000+
channels
24/7
captured and verified
The world's speech disappears as it airs. Help us keep it.