Kakao has significantly upgraded its proprietary artificial intelligence (AI) model, Kanana-o, with advanced voice generation capabilities. This enhancement allows the AI to produce speech that accurately reflects user-specified nuances such as speaking style, emotion, intonation, and even regional accents, all dictated through natural language commands. Kakao claims that internal evaluations show Kanana-o outperforms OpenAI’s voice models in key areas.
Kanana-o’s Advanced Voice Synthesis Features
Announced on the 4th, Kakao’s refined Kanana-o model demonstrates a sophisticated ability to interpret and apply user instructions to voice output. For instance, users can instruct the AI to read text in specific ways, such as “read it quickly,” “read it in a sad tone,” or “read it with a Gyeongsang Province accent.” The AI is designed to faithfully render these directives into the generated speech.
This capability opens up a wide range of potential applications. Kanana-o could be employed for tasks requiring specific vocal characteristics, such as narrating sports broadcasts with an energetic commentator’s voice, delivering news with a professional anchor’s tone, or reading children’s stories with distinct character voices. Furthermore, the model supports complex instructions that combine multiple parameters simultaneously, allowing users to specify a desired speaking style, speed, and emotional tone all in a single command.
Performance Benchmarks and Technical Innovations
Kakao has presented performance data from an internal evaluation using the “InstructTTSEval” benchmark for Korean, which measures the ability of text-to-speech (TTS) systems to follow instructions. In this benchmark, Kanana-o achieved a score of 94.50. For comparison, OpenAI’s “GPT-4o mini TTS” scored 91.10, and Google’s “Gemini 1.5 Flash preview TTS” scored 95.38.
Beyond instruction following, Kakao has also focused on improving the efficiency of the voice generation process itself. To achieve this, the company integrated its self-developed “LM-SPT” (Language Model – Speech Tokenizer) technology into Kanana-o. LM-SPT is a technique that compresses audio data into a smaller number of tokens, thereby reducing the amount of data the AI model needs to process.
Tokens serve as the fundamental units by which AI models process audio or text. By reducing the token count, the processing load is lessened, potentially leading to faster generation times and lower computational costs.
LM-SPT’s Impact on Audio Quality
Kakao reports that LM-SPT has demonstrated superior performance in its own internal tests. These tests evaluate various aspects of synthesized speech, including comprehension and generation of Korean and English audio, the naturalness of the synthesized voice, and the similarity of the generated voice to the original speaker’s characteristics. The results indicate that the integration of LM-SPT has led to noticeable improvements in these areas for Kanana-o.
Future Applications and Vision
Noh Byung-seok, a leader at Kakao’s Unified Foundation Model division, highlighted the core focus of the Kanana-o upgrade. “We have concentrated on enabling users to express the desired speaking style, emotion, and intonation according to natural language instructions,” Noh stated. He further elaborated on the company’s strategic outlook, saying, “We plan to apply Kanana-o to various services in the future.”
This advancement in AI-driven voice synthesis signifies Kakao’s commitment to developing more nuanced and human-like AI interactions. The ability to control speech characteristics with such precision could lead to more engaging and personalized user experiences across a multitude of digital platforms and services.
The potential applications are vast, ranging from enhanced accessibility tools for individuals with communication challenges to more immersive virtual assistants and sophisticated content creation aids. As AI technology continues to evolve, features like those introduced in Kanana-o are likely to become increasingly integral to how we interact with technology.
