Mr. Ühss, what does AudioStack do, and what problem in audio production are you aiming to solve?
 

AudioStack is an agentic audio production platform. As an audio infrastructure provider, we help companies not only digitize their audio production processes but also bring them into the era of Gen AI and agentic technology.
We solve the problem that traditional audio production is very slow, labor-intensive, and expensive. As part of their transformation, media companies, advertising agencies, and publishers are increasingly relying on AI-native processes in audio production. Specifically, this includes audio content such as audio spots, video soundtracks, podcasts, and audio articles, which can be produced fully automatically and much faster using our technology.
What used to take days or weeks now takes minutes or seconds—and is fully automated at the push of a button.
The platform offers various AI workflows and agents such as copywriters, AI voice actors, and producers, as well as the ability to personalize content or automate it based on data.

 

What was the idea behind founding AudioStack?
 

When we started with the idea almost 10 years ago, our vision was to make audio simpler and more personalized, but also more accessible and scalable. Social media feeds are highly personalized and algorithm-driven—but audio, as a medium, is very static and always the same. Why is that, actually? Why can’t audio be as relevant and personalized as my social media feed, we asked ourselves back then.
What was already possible in social media, display, and video—the rapid creation and personalization of content and advertising—was still lagging far behind in audio. We wanted to be the Stripe or Twilio of audio—the infrastructure that empowers companies to produce any kind of spoken audio professionally and automatically at scale. Today, our customers produce thousands of hours of audio content at the push of a button with a quality that was unimaginable just a few years ago. Today’s voice models are extremely expressive and are evolving very rapidly, enabling entirely new use cases.

 

What training data does an audio AI model need, and what are the biggest challenges?
 

A high-quality audio AI model requires three things: high-quality, studio-grade voice recordings; very precise annotations regarding emotion, speaking style, and language; and a wide variety of voices, dialects, and contexts. The biggest challenge lies not in the quantity, but in the quality and legal compliance. Many of today’s AI voices are created using data models rather than actual recordings. Quality improvements therefore stem primarily from increasingly sophisticated model training methods.

 

How can you recognize high-quality AI audio results?
 

The best indicator of quality is quite simple: the listener doesn’t realize it’s AI. High-quality AI voices have natural pauses, emotional variation, correct intonation—even with proper nouns, brand names, and numbers—and they match the tone of the context. Poor AI voices sound monotonous, “too clean,” or have incorrect intonation. The ultimate test in the market is always a blind comparison: if listeners can’t tell the difference, the quality is where it needs to be—and in most languages today, that’s actually almost always the case.

 

How is AI audio changing the way brands communicate with their target audiences?
 

AI audio finally makes audio dynamic. Today, brands can produce ads that vary depending on the time of day, weather, location, station, or target audience—without having to go back into the studio every time. This completely changes the logic: Instead of a single ad that runs the same everywhere, hundreds or thousands of variations are created, each of which is more relevant and contextually appropriate.
The most fundamental shift is the transition from campaign-based logic to always-on communication in audio. Until now, audio has been a medium for large, infrequent productions. With AI, brands can now communicate continuously, situationally, and reactively. Audio is thus evolving from a broadcast channel into a medium for dialogue.

 

In your opinion, will AI replace or complement human voice actors in the long term?
 

Quite clearly: complement. Human voice actors remain irreplaceable when it comes to brand identity, sophisticated creative productions, audiobooks, or character voices. AI takes over where scaling, speed, or multilingualism are required—that is, in use cases that would be impossible to realize with traditional production. The market is getting bigger, not smaller.

 

Where do you personally draw the line between efficient automation and authentic storytelling?
 

To be honest, that line is constantly shifting. A few years ago, AI mainly handled functional content. Today, I hear AI-generated ads on TV, Spotify, and the radio that are fully effective on an emotional level. As a general rule: use automation where scaling, language, or speed are key. Humans should handle the central creative idea and the main emotional message. A strong creative idea remains human. The replication, localization, and distribution of this idea can and should be automated. As a tech company, we have a strong creative team that develops creative and innovative ideas. There is always a need for an “orchestrator” who creatively guides and refines the central idea. With AI as a tool, ideas can be developed more quickly and rapidly drafted in various design scenarios.

 

Where do the greatest opportunities lie for companies in the field of AI and audio?
 

For a long time, audio was the last major medium without true data integration. With AI-powered audio, brands are now able for the first time to combine audio creation with audio data—that is, to produce content in a data-driven way and test which voice, tone, and call-to-action work best with which target audience. This is a new dimension of performance that simply wasn’t possible before.
The most exciting opportunities, however, lie beyond advertising: audio accompaniment for apps, voice-controlled brand experiences, personalized audio messages to customers, and audio versions of newsletters or reports. We are only at the beginning of what becomes possible when audio is as easy to produce as text.
 

How do you address ethical issues surrounding synthetic voices and potential deepfake audio applications?
 

Synthetic voices are a technology like any other, which is why they are offered by major tech companies such as Google, IBM, Amazon, and others. I believe we must therefore distinguish between two things: legitimate synthetic voices used in professional business contexts and abusive deepfakes, which often occur in private settings.
Artificial voices are a valuable tool that we want to help shape responsibly. Deepfake applications or identity theft are clearly prohibited in our Terms of Service and technically prevented. We do not clone voices without the explicit consent of the person in question, and we work exclusively with business customers in transparent B2B contexts. Trust is the most important asset in our industry, and we consistently protect it.

 

What developments can we expect to see over the next 3–5 years?
 

Three developments will shape the market: First, AI voices will become indistinguishable from human ones—this is already a reality today in many fields, sectors, and languages. What we expect next are local dialects and accents, as well as the ability to configure these voices ourselves, as we already know from Gen AI tools with prompting.
Second, real-time audio personalization will be used much more extensively by brands, and audio will come closer to the personalization capabilities of visual and text formats. 
Third, audio production will shift from a manual studio process to an automated API process, supported by agent-based workflows and embedded within CMS, ad-tech stacks, and marketing platforms.
Much of this will take a conversational form—that is, through natural interaction between humans and AI—and increasingly, agent-to-agent interaction, meaning with little or no human input.

In 3 to 5 years, we will no longer just listen to linear audio content, but will interact with audio.

 

Will AI develop its own distinctive “voices”?
 

Marketing will certainly move in this direction, even if the development of these voices is initiated and orchestrated by us humans. There are already very extensive voice marketplaces where all kinds of voice variations can be generated and further customized.
Brands will develop their own identities. A brand voice that belongs exclusively to that brand, sounds consistent worldwide, and conveys the same personality across all touchpoints.
Just as we have iconic human voice actors today, we will have iconic AI voices in a few years. Some will be hybrid forms, others completely synthetic. The exciting thing about this is that an AI voice can become an international, multilingual brand signature that sounds consistent worldwide, in every language, for decades to come. This is not possible with human voice actors.