Voice, video and the enterprise: notes from LocWorld and NAB
Voice and speech are the hot space of the moment. On the end of AI anxiety in big enterprises, the murky law of synthesised voices, and why LLMs are getting worse at chess.
I attended two conferences in October — LocWorld, a leading enterprise-localisation event, and NAB East, part of the broadcast world. The single dominating topic in both, with huge convergence on how it'll be used: you guessed it, AI.
Transforming voices and frames
If you're an '80s kid like me, you might remember Scotty time-travelling in Star Trek IV and talking into a computer mouse. Well, that future is here. Meta's voice-enabled AI products point to a not-too-distant world where conversational interfaces are entirely ubiquitous — so it's unsurprising that voice and speech have become the hot topic of the moment, with genuine evolutionary development at every layer. Bessemer's recent AI Voice Roadmap is well worth a read.
Voice in enterprise
The rise of AI voice over the last 18 months hasn't been a straight, steep line. Enterprises have spent a couple of years figuring out how to use AI both practically and with a sense of good governance. We're coming to the end of a period of AI anxiety in many large organisations — which means they're now turning to deploy, and deploy rapidly.
At IBC, Brightcove unveiled integrated AI tools at both the foundational and platform levels, and they're not the only major video platform signalling a shift toward smarter creation and delivery — Vimeo's AI-powered video translation preserves a speaker's voice and nuance across languages. Both saw share-price upticks after Q3 earnings. For many large organisations, simply deploying a foundational model like ElevenLabs isn't a straightforward path to AI 'just working'; that's where intermediary distribution and application platforms come in. SaaS platforms are becoming embedded-AI platforms, and those left behind will be just that — history.
Upstream, the production toolbox is changing too: Adobe released AI-powered auto-captioning. But scaling captions and voice across languages, with approval workflows and brand control, requires enterprise-grade platforms — so large enterprises are still figuring out where in the lifecycle AI is best applied.
Foundational voice models: safe and ethical use
How do native voice models prevent misuse? As my colleague put to ElevenLabs: who or what stops me uploading George Clooney's voice and synthesising a script? The law here is surprisingly unprotective, and turns on where the voice was obtained — you might upload Clooney's voice, but you presumably got it from a film or interview, which is itself protected intellectual property.
A voice cannot be copyrighted. According to Midler v. Ford Motor Co., a voice is as distinctive and personal as a face. While a recording of a voice may be copyrighted, a voice itself may not — though performances by actors, including vocal performances, are protected. A key question is where the voice artist's content was obtained from.
Platforms are on the front foot — legal fuzziness is bad for buyer confidence. As ElevenLabs' GTM director for media and entertainment told me, they algorithmically check a voice as it's uploaded against a bank of protected voices and won't process a match; for everything else they run a validation, basically a Voice CAPTCHA. Robust input–output verification will become as normal in voice interfaces as the biometric fingerprint is now.
Limits of understanding
Changing tack: a colleague shared an amusing finding that LLMs are really bad at chess — and getting worse, not better. Despite remarkable outputs, generative AI still shows volatility and a lack of coherent world-understanding, failing on simple repeated tasks with minor context changes.
I was surprised how quickly performance deteriorated as soon as we added a detour. If we close just 1 percent of the possible streets, accuracy immediately plummets from nearly 100 percent to 67 percent.
Recent MIT research made the point vividly: adding detours to a map of New York caused navigation models to fail, and the maps they'd internally generated turned out to be an imagined city of impossible streets and random flyovers. A reminder that AI's intelligence is still fundamentally synthetic — a real challenge for fields that demand contextual awareness. That's all for this month. Beam me up, Scotty.