Modulate Raises $25 Million to Read Emotion and Fake Voices

Voice Models
Modulate raised $25 million on September 28. Future Ventures led the round. The company wants machines to hear emotion and tell a real voice from a fake one, not only write down the words.
By Shashi Bellamkonda · September 29, 2026
$25M
new round, Future Ventures led
$60M
total funding Modulate reports
10M+
hours of audio a month
600M+
hours processed in total
Models that transcribe speech are now very good at turning languages into written words. Understanding emotion, and catching a fake voice, still has to be done well. I think Modulate is on the right track.

Modulate is a Boston company started in 2017 by Carter Huffman and Mike Pappas, who met as physics students at MIT. On September 28 it announced $25 million from Future Ventures, with Hyperplane and Lakestar in the round. Modulate says total funding is now $60 million. TechCrunch cited PitchBook putting earlier funding at $41 million and the last valuation at $170 million (Modulate, 2026; Mehta, 2026).

They first built voice skins for games. That turned into ToxMod, which listens for abuse in titles such as Call of Duty and Rec Room. The product they sell now is Velma. It listens to the raw call, not only the written transcript, and looks for emotion, tone, intent, and whether the voice was generated by a machine (Modulate, 2026).

Huffman said voice problems "can't be solved from a transcript" (Huffman, 2026).

Healthcare, games, and sales calls need the same listening work

Call centers will use this. Hospitals can use it when a caller claims to be a doctor. Game and social platforms can use it when abuse or grooming hides inside ordinary chat. Companies can use it to watch whether a voice agent is frustrating a customer while the call is still open.

I also think about sales. After a live call, a manager often guesses from memory whether the other side still wants the deal. Emotion and intent on that recording are a better input than a recap written an hour later.

Pitch, pace, and a cloned voice do not survive if you only keep the text. That is the layer Modulate is funding.

Velma uses more than 100 small audio models

Modulate does not send every second of audio through one giant model. It runs more than 100 smaller models together and calls that stack an Ensemble Listening Model. The company says this is up to twice as accurate on real hits as using a large language model for the same jobs, with seven times fewer false alarms, and far less compute. Those numbers are Modulate's claims. Ask how they tested them (Modulate, 2026).

In July its transcription work ranked first on Hugging Face's Open ASR Leaderboard. Velma Deepfake Detect, which shipped in March, ranks first on Hugging Face's Speech Deepfake Arena, with 98.9 percent accuracy on that public set. Batch transcription is $0.03 an hour. Some writeups put deepfake checks at $0.25 an hour. A public test is useful. It is not the same as your own call recordings (Modulate, 2026; Riley, 2026).

The company says it now runs more than 10 million hours of audio a month and has passed 600 million hours in total. TechCrunch says the team is about 40 to 45 people and plans to hire about 10 more. They are also working on versions that stay on a customer's own machines (Modulate, 2026; Mehta, 2026).

Someone has to hear the person on the line

If you build a voice agent you usually buy a speaking voice from one company and a reasoning model from another. You still need something that hears the human. Modulate is trying to be that piece so every team does not train its own emotion and fake-voice models.

The new money goes to research, engineering, developer tools, and partners. What they published this week is the volume, the two Hugging Face ranks, and the transcription price. The sales-call idea is mine. It is not in the press release as a product.

If you already record calls

Take one queue that has a transcript and still needs a person to judge mood or a fake. Run Velma on that queue. Write down the misses. That sheet tells you whether this layer is worth buying.

Sources: Modulate. "Modulate Raises $25M to Scale Its Lead in Frontier Audio-Native AI." 28 Sept. 2026, https://www.modulate.ai/press-releases/modulate-raises-25m-to-scale-its-lead-in-frontier-audio-native-ai. Carter Huffman and Mike Pappas. "We Started by Asking What Machines Could Hear. Now We're Building What Comes Next." Modulate, 28 Sept. 2026, https://www.modulate.ai/blog/we-started-by-asking-what-machines-could-hear-now-were-building-what-comes-next. Ivan Mehta. "Modulate raises $25M for its voice models and analysis suite." TechCrunch, 28 Sept. 2026, https://techcrunch.com/2026/09/28/modulate-raises-25m-for-its-voice-models-and-analysis-suite/. Duncan Riley. "Voice AI startup Modulate raises $25M to bring audio-native models to more developers." SiliconANGLE, 28 Sept. 2026, https://siliconangle.com/2026/09/28/voice-ai-startup-modulate-raises-25m-to-bring-audio-native-models-to-more-developers/.

Disclaimer: This blog reflects my personal views only. Content does not represent the views of my employer, Info-Tech Research Group. AI tools may have been used for brevity, structure, or research support. Please independently verify any information before relying on it.