SignalMatrix · Pre-seed 2026

The internet is silent.
The airwaves aren't.

We turn live TV and radio into quality speech data for the big languages AI still can't hear.

Hausa is 0.003% of the web 94M people speak it on air every day Heard · Transcribed · Verified
02

Broadcast in. AI-ready data out.

One pipeline. Two products.

2,000 channels

Live TV and radio, captured as it airs.

→

SignalMatrix

1 · Heard 2 · Transcribed 3 · Verified
→
Revenue now

Intelligence feed

Largest prize

Training data

Every segment carries station, time, speaker, dialect and sources.
03

AI ran out of internet.

Quality data is now the bottleneck.

Ilya Sutskever

“We've achieved peak data and there'll be no more.”

Ilya Sutskever, OpenAI co-founder · NeurIPS, December 2024

$14.3B

Meta paid for 49% of Scale AI, a data company, in June 2025.

$1.2B

Surge AI's estimated 2024 revenue, selling human data to frontier labs.

10 months

David AI went from a $5M seed to a $50M Series B, selling audio data alone.

Sources: The Verge / NeurIPS 2024; Meta–Scale deal reporting, June 2025; Forbes via Sacra (Surge, press estimate); David AI and Bloomberg, 2025.

04

Big languages. Almost no data.

Hours of transcribed training audio in OpenAI's Whisper, the reference speech model.

LanguageSpeakersWhisper hoursPer 1M speakers
English1.5B438,218292
Arabic335M7392.2
Persian82M240.3
Hindi611M120.02
Bengali274M1.30.005
Hausa94M00
Pashto21M00
Amharic78M00
Error rates halve with every 16× more data. A data problem, not a model problem.

Sources: Whisper (arXiv 2212.04356); Ethnologue. Bars on a log scale. Kurdish, Oromo, Igbo, Zulu, Tigrinya and Wolof: also 0 hours.

05

These languages live on air.

Not online. On the radio and TV people actually use.

Where

59%

of Africans get news from radio at least a few times a week, across 38 countries. Radio is still the continent's top news source.

What

Real speech

Talk shows and call-in radio are spontaneous, multi-speaker and dialect-switching. Exactly the speech clean benchmarks miss.

Consequence

71–99%

of the time, Whisper large-v3 writes Pashto audio in the wrong script. Models fail where real speech lives.

Sources: Afrobarometer Dispatch AD1173 (Round 10, 2024/2025); Pashto ASR benchmark, arXiv 2604.04598 (2026).

06

Three things have to be true at once.

Nobody has all three.

Capture at scale

Thousands of channels, 24/7, with scheduling and failure recovery.

Who has it today: media monitors, English-first.

Dialect-level ASR

Multi-speaker, code-switching, and cheap enough per hour to run at scale.

Who has it today: labs, trained on clean audio.

Verification

Cross-references, entity linking and provenance for every segment.

Who has it today: nobody.

We built all three.
07

Heard. Transcribed. Verified.

How one broadcast segment becomes data.

1

Heard

Live TV and radio captured as it airs, with station, country and timestamp on every segment.

→ recording + metadata
2

Transcribed

Speech-to-text tuned for dialects, with speaker, entity and topic tags.

→ time-coded segments
3

Verified

Every segment quality-checked, cross-referenced against independent sources and linked into a knowledge graph.

→ claims with sources
08

What one segment looks like.

Illustrative output, and the signals pulled from it.

Illustrative output · one broadcast segment

الحكومة قررت رفع أسعار الوقود ابتداءً من الشهر الجاي.

“The government decided to raise fuel prices starting next month.”

Egyptian Arabic Speaker 2 · host Entity: government Topic: fuel prices Cross-referenced: 2 sources
Signals extracted from every segment
  • Speakers, up to 8 separated at once
  • Entities and topics, linked in a graph
  • Claims, each tied to a speaker
  • Sentiment, events, contradictions
  • Cross-references to other sources
  • On-screen figures, via vision
09

The hard engineering is already done.

Five breakthroughs SignalMatrix has built, where off-the-shelf models fall short.

−60%

Cost

Cheaper per hour than Whisper-class transcription.

5,000+

Scale

Channels our ingestion pipeline is built to carry.

40+

Reach

Languages our transcription stack already handles.

8

Real conversation

Speakers separated at once. Single-speaker models break.

Vision

Beyond audio

On-screen figures tie each claim to who said it.

What's left to 50 languages is licensing and data, not R&D.
10

What we've built.

585,996 hours of real speech, and counting.

Captured585,996 h780,505 recordings · 40.9 TB archived · 66 years of airtime
Transcribed88,524 hPrioritised by language demand
Segments100,384Cut, labelled and time-coded
Claims169,423Each tied to a speaker · 190,548 extractions
Cross-references85,763To 7,054 independent sources
Entities42,314In the knowledge graph

Live counts from the production warehouse.

11

2,000 channels. 27 countries. 8 languages.

Live today, deepest in Arabic, Farsi, Dari and Kurdish.

TanzaniaKenyaPakistanUganda United States of AmericaSomaliaRussiaFranceMaliMauritaniaNigerNigeriaCameroonGabonIsraelJordanUnited Arab EmiratesQatarIraqIndiaAfghanistanIranGermanyTurkeyChinaUnited KingdomSaudi ArabiaEgyptEthiopiaDjibouti Bahrain
Broadcast captured today Next languages this round

~2,000

TV channels

27

Countries

8

Languages live

Shaded: countries with recorded broadcast, from the production warehouse. Outlined: next-round countries.

12

One asset. Two buyers.

Revenue now, scale later.

First revenue

Intelligence feed

Verified, dialect-tagged broadcast speech, delivered as it airs.

  • Risk and narrative intelligence platforms
  • Hedge funds tracking emerging markets
  • Newsrooms and media groups
Largest prize

Training and evaluation data

Real-world speech for the languages models can't hear yet.

  • Frontier AI labs
  • Voice AI companies
  • Sovereign AI programs
Training data is sold only through a rights-cleared track.
13

Three markets. One dataset.

Who pays for real-world speech.

1
First revenue

$2.8B spent by investment managers on alternative data in 2025, up 17% in a year.

2
Largest prize

Billions spent each year by AI labs on training data, rising as public data runs out.

3
Highest need

Sovereign AI programs in MENA, Africa and South Asia, with almost no native speech data.

Serviceable market: non-English broadcast speech across 50 languages. Estimate to add.

Source: Neudata, “The state of the alternative data market in 2026,” February 2026.

14

Both markets are already testing the data.

Design partners working with the current corpus today.

Training data market

A voice-AI company

  • Testing dialect-tagged broadcast speech
  • As training and evaluation data
Intelligence feed market

A risk-intelligence platform

  • Testing the verified broadcast feed
  • For narrative tracking
One product, two buyers, both engaged before the round.
15

Only SignalMatrix checks all four boxes.

Investors just funded this category. We own the half nobody's taking.

Non-English depth
Real-world broadcast
Verified
Rights-cleared data
Speech-data vendors
?Varies
✕Studio or scripted
✕No
✓Yes
Media monitors
✕English-first
✓Yes
✕No
✕No
Cast Insights$4.5M pre-seed, Jul 2026
✕US focus
✓Yes
✕No
?Not stated
SignalMatrixPre-seed 2026
✓Dialect-level
✓Yes
✓Two-source
✓Yes
We make the rest of the world's speech usable, and trustworthy.
16

The loop that compounds.

Every hour we collect makes the next one cheaper.

More hours captured
Better dialect models
Lower cost per hour
More languages viable
More broadcaster licenses

Can't be recaptured

Past broadcasts are gone once they air. Competitors can't get them back.

Slow to copy

Broadcaster licenses take time to sign.

17

Built by the people who built Arabic AI.

Dialect-level ASR and data you can trust are both hard. This team has done each before.

Abdelrahman Mansour
Founder and CEO

Abdelrahman Mansour

  • 15+ years building Arabic-language media and information platforms.
  • Co-founder of Matsadaash, a leading Arabic information platform.
  • Grew funding for information initiatives from $0 to $8M.
Abu Bakr Soliman
Co-founder and CTO

Abu Bakr Soliman

  • Built SILMA 1.0, an early Arabic LLM that beat Gulf sovereign models on the Open Arabic LLM Leaderboard.
  • Author of AraVec, among the most-cited Arabic NLP work.
  • Led AI at Unifonic, including a leading MENA social listening platform.
18

The road to 50 languages.

Every new language is a new corpus on the same pipeline, not a rebuild.

8

Live today

Arabic and its major dialects, Farsi, Dari, Kurdish and four more.

20

This round

Pashto, Urdu, Hausa, Swahili and more high-speaker, low-data languages. First paying contracts.

50

Next

Across Africa and Asia: Amharic, Yoruba, Somali, Bengali and beyond.

We count languages, not dialects.
19

$3M for 20 languages and first revenue.

Pre-seed on a SAFE, room to $4M. About 24 months of runway.

$3M

Use of funds
  • 40% capture, new languages and broadcaster rights
  • 30% machine learning and data quality
  • 20% go-to-market and partnerships
  • 10% infrastructure and reserve
What it buys

By the end of the round

  • 20 languages live in production
  • First paying contracts across feeds and data
  • Rights-cleared track signed with first broadcasters

Vision

Every language AI can hear.

The default supplier of real-world speech for every language AI can't hear yet.

50

languages

5,000+

channels

24/7

captured and verified

The world's speech disappears as it airs. Help us keep it.