텍스트 AI를 음성으로 바꾸는 데 그치지 말고, 음성만이 제공할 수 있는 경험을 적극적으로 개발하자.
현재 AI 서비스는 텍스트와 음성으로 비슷한 작업을 수행할 수 있다.
질문에 답하고, 정보를 검색하고, 글을 작성하고, 복잡한 문제를 해결한다.
그러나 음성 AI에는 텍스트 AI와 다른 기술적 제약이 있다.
사용자는 자연스러운 대화와 빠른 응답을 기대하지만, 복잡한 작업일수록 충분한 처리 시간과 자연스러운 대화를 동시에 제공하기 어려울 수 있다.
그렇다면 접근 방식을 바꿔 보면 어떨까?
모든 기능에서 텍스트 AI와 똑같이 경쟁하는 대신, 음성으로 경험할 때 특별한 가치가 생기는 기능과 상황을 적극적으로 개발하는 것이다.
이 제안은 다음과 같은 가능성에서 출발한다.
노래를 잘 부르는 AI가 아니라, 함께 노래하며 즐거움을 주는 AI.
음성 대화 중 상황에 맞춰 짧은 노래를 즉흥적으로 만들어 부르거나 흥얼거리는 기능이다.
사용자가 좋은 소식을 전하면 축하 노래를 부르고, 장난스러운 상황에서는 엉터리 노래로 받아칠 수도 있다.
전문 가수처럼 완벽하게 부를 필요는 없다. 때로는 서툰 노래 자체가 재미있는 상호작용이 될 수 있다.
이 아이디어는 실제 경험에서 출발했다.
운전 중 노래를 틀어달라고 요청했지만 저작권 문제로 거절당했고, 직접 흥얼거려 달라는 요청도 받아들여지지 않았다.
그 순간 떠오른 질문은 간단했다.
기존 노래를 재현할 수 없다면, 상황에 맞는 새로운 노래를 만들어 부르면 어떨까?
원곡을 복제하지 않고도 노래를 통해 사람과 상호작용하는 새로운 경험을 만들 수 있다.
정보를 읽어주는 AI에서 이야기를 들려주는 AI로.
사용자가 잠들기 전 동화를 듣고 싶거나, 긴 글을 읽는 대신 편안하게 듣고 싶을 때 활용하는 기능이다.
사용자가 원하는 책이나 허용된 콘텐츠 읽어주기
긴 글이나 책의 내용 요약해서 들려주기
새로운 동화와 이야기 만들기
이야기의 분위기에 맞춰 자연스럽게 읽어주기
핵심은 단순히 텍스트를 음성으로 변환하는 것이 아니다.
이야기의 흐름, 말하는 속도, 적절한 쉼과 감정 표현을 통해 듣는 경험 자체를 만드는 것이다.
저작권이 있는 책을 그대로 읽어줄 수 없는 경우에도 허용되는 범위에서 요약하거나 새로운 이야기로 풀어내는 대안을 제공할 수 있다.
텍스트를 읽는 수고를 덜어주는 것을 넘어, 누군가 이야기를 들려주는 경험을 제공하자.
미리 준비된 성우 중 하나를 고르는 데서 벗어나, 사용자가 원하는 목소리를 직접 구성하는 AI.
사용자가 선호하는 음색과 목소리를 선택하거나, 기술적으로 가능한 범위에서 자신만의 음성을 만들 수 있도록 한다.
여기에는 다양한 활용 가능성이 있다.
차분하거나 활기찬 목소리 만들기
자신이 좋아하는 음색으로 대화하기
개인적인 취향에 맞는 AI 음성 구성하기
적절한 동의와 권한 아래 가족이나 친구의 목소리 재현하기
이 아이디어에는 개인적인 경험도 있다.
세상을 떠난 친구의 목소리로 대화할 수 있다면 어떨까?
녹음된 목소리를 듣는 것과 그 목소리로 내 말에 응답하는 것은 다른 경험일 수 있다.
물론 AI가 실제 그 사람의 기억이나 생각을 되살리는 것은 아니다. 또한 실제 인물의 목소리를 재현하려면 동의와 권한, 오용 방지에 대한 고려가 필요하다.
하지만 목소리 자체가 사용자에게 특별한 의미를 지닌다면, 그것을 대화 환경에 반영하는 것만으로도 새로운 가치를 제공할 수 있다.
더 좋은 목소리를 선택하는 것이 아니라, 나에게 의미 있는 목소리로 대화하는 것이다.
정확하게 발음하는 것만큼, 사용자가 이해하기 쉽게 발음하는 것도 중요하다.
한국어 음성 대화에서 영어 단어가 등장하면 AI가 지나치게 원어민에 가까운 발음을 사용하는 경우가 있다.
하지만 한국어에는 영어에서 유래한 외래어가 일상적으로 많이 사용된다.
한국인은 영어 단어를 한국어의 발음 체계에 맞춰 익숙하게 사용하는 경우가 많다.
이런 상황에서 모든 영어 단어를 원어민 발음으로 강조하는 것이 언제나 최선일까?
예를 들어 한국어로 대화할 때는 한국에서 일상적으로 사용하는 영어 외래어 발음을 활용하고, 영어 문장을 읽거나 영어로 대화할 때는 자연스러운 영어 발음을 유지하는 것이다.
이를 위해 발음 설정에 다음과 같은 옵션을 제공할 수 있다.
NATIVE: 자연스러운 원어민 발음
LOCAL: 해당 언어권에서 익숙하게 사용하는 발음
LEARNING: 외국어 학습에 적합한 발음
이 기능은 한국어에만 국한되지 않는다. 다른 언어에서도 현지의 발음 습관과 언어 환경을 고려한 음성 현지화를 탐색할 수 있다.
다만 LOCAL 모드가 모든 영어 단어를 무조건 콩글리시로 발음하는 방식이어서는 안 된다. 문맥에 맞는 자연스러운 발음이 목표다.
원어민처럼 말하는 것이 언제나 최선은 아니다. 사용자가 편하게 이해할 수 있도록 말하는 것도 중요하다.
앞의 네 가지가 음성만의 새로운 경험을 만드는 기능이라면, 다음은 음성이 특히 유리한 상황을 활용하는 방향이다.
운전 중 대화는 대표적인 활용 사례다.
그 밖에도 요리, 청소, 산책 등 화면을 보거나 직접 입력하기 어려운 상황에서 음성은 편리하다.
사용자는 작업을 중단하지 않고도 질문하거나 필요한 정보를 들을 수 있다.
다만 운전 중에는 음성 대화도 주의를 분산시킬 수 있으므로 안전을 우선해야 한다. 짧고 단순한 상호작용을 중심으로 설계하고, 복잡한 대화는 운전을 마친 뒤 이어가는 방식이 적절하다.
현재의 일반적인 상호작용은 사용자가 요청하면 AI가 답하는 방식이다.
하지만 음성 AI는 대화의 맥락에 맞춰 적절한 반응을 먼저 제안하거나 실행하는 방향으로 발전할 수 있다.
사용자가 기쁜 소식을 전하면 축하하고, 원한다면 짧은 노래를 부른다. 잠들기 전 이야기를 듣고 싶어 한다면 자연스럽게 이야기를 이어간다.
핵심은 사용자가 매번 구체적인 명령을 내리지 않아도 된다는 것이다.
다만 적극적인 반응이 언제나 환영받는 것은 아니므로, 사용자의 선호를 존중하고 언제든 제어할 수 있어야 한다.
음성 대화에서는 응답의 내용뿐 아니라 말하는 속도, 대화의 리듬, 자연스러운 반응도 중요하다.
사용자가 이야기를 이어가는 동안 적절한 질문을 하거나, 짧게 반응하거나, 필요할 때 잠시 기다리는 방식이다.
이를 통해 음성 AI는 질문과 답변을 반복하는 인터페이스에서 벗어나 보다 자연스러운 대화 경험을 제공할 수 있다.
물론 이를 위해서는 응답 지연, 음성 인식, 대화 차례 관리 등을 지속적으로 개선해야 한다.
음성 AI의 기술적 제약을 고려하면, 모든 작업에서 즉각적인 응답을 제공하려는 접근만으로는 한계가 있을 수 있다.
따라서 작업의 특성에 맞는 상호작용 방식을 선택할 필요가 있다.
작업
적합한 접근
짧은 질문과 답변
빠른 음성 응답
운전 중 간단한 정보 요청
짧고 부담이 적은 핸즈프리 대화
긴 이야기 들려주기
자연스럽게 나누어 전달
즉흥 노래
짧은 멜로디와 가사부터 생성
복잡한 질문과 분석
충분히 처리한 뒤 자연스럽게 설명
긴 문서 비교와 정밀한 편집
텍스트와 음성을 함께 활용
음성 AI의 모든 작업에 동일한 수준의 즉답성을 요구하는 대신, 각 작업에 적합한 속도와 전달 방식을 선택하는 것이다.
이 접근은 기술적 한계를 없애는 것이 아니라, 그 한계 안에서 더 나은 사용자 경험을 설계하는 방법이다.
이 제안의 핵심은 일곱 가지 기능을 모두 구현하자는 데 있지 않다.
각 아이디어는 서로 다른 문제를 해결하지만, 하나의 방향을 공유한다.
음성 AI를 텍스트 AI의 음성 버전으로 만드는 데 그치지 말고, 음성이라는 매체만의 고유한 가치를 적극적으로 개발하자는 것이다.
이를 세 가지 전략으로 정리할 수 있다.
전략
핵심 가치
대표 아이디어
새로운 경험
즐거움과 감성
음치 AI, 할머니의 동화
개인화
사용자의 취향과 언어 환경
CUSTOM VOICE, KONGlish MODE
새로운 상호작용
편리함과 자연스러운 대화
핸즈프리, 선제적 반응, 자연스러운 대화
이 가운데 음치 AI, 할머니의 동화, CUSTOM VOICE, KONGlish MODE는 구체적인 기능 제안이다.
핸즈프리와 자연스러운 대화는 음성 AI가 활용될 수 있는 환경과 상호작용 방식에 관한 제안이다.
각 기능은 독립적으로 실험할 수 있고, 필요하다면 서로 결합할 수도 있다.
예를 들어 사용자가 직접 구성한 목소리로 동화를 들려주거나, 선호하는 음성으로 즉흥적인 노래를 부르는 것도 장기적으로 탐색할 수 있다.
물론 이런 가능성은 실제 기술 구현과 사용자 반응을 통해 검증해야 한다.
음성 AI의 발전을 더 빠르고 정확하게 말하는 기술의 발전으로만 바라볼 필요는 없다.
물론 응답 속도와 정확도는 중요하다.
그러나 음성은 정보를 전달하는 수단일 뿐 아니라, 이야기를 들려주고, 노래하고, 감정을 표현하고, 사용자가 선호하는 목소리로 소통할 수 있는 매체이기도 하다.
또한 사용자의 작업 환경과 언어 습관에 맞춰 소통하는 방식도 발전시킬 수 있다.
따라서 음성 AI는 텍스트 AI와 동일한 기능을 수행하는 데 그치지 않고, 음성만이 제공할 수 있는 경험과 가치를 적극적으로 개발할 필요가 있다.
텍스트 AI와 경쟁하는 음성 AI가 아니라, 텍스트 AI와 다른 이유로 선택받는 음성 AI.
이것이 이번 제안의 핵심이다.
Instead of simply turning text-based AI into a voice interface, let's actively develop experiences that only voice can uniquely deliver.
Today's AI services can perform similar tasks through text and voice.
They answer questions, retrieve information, write content, and solve complex problems.
However, voice AI faces different technical constraints from text-based AI.
Users expect natural conversations and fast responses, but the more complex a task becomes, the harder it can be to deliver both sufficient processing time and a seamless conversational experience.
What if we approached the problem differently?
Instead of competing with text AI on every feature, we could actively develop experiences and use cases that become more valuable when delivered through voice.
This proposal explores that possibility.
An AI that doesn't need to sing beautifully to make people smile.
Imagine an AI that spontaneously creates and sings short songs during a conversation, responding to the situation with a melody or a playful tune.
It could sing a congratulatory song when a user shares good news or respond with a silly song during a playful exchange.
It doesn't need to sing like a professional performer. Sometimes, the imperfect singing itself can become the source of entertainment.
This idea emerged from a real experience.
While driving, I asked the voice assistant to play a song, but the request was declined because of copyright restrictions. I then asked it to hum or sing something itself, but that wasn't possible either.
That experience led to a simple question:
If an AI cannot reproduce an existing song, why not let it create and sing an original one that fits the moment?
Without reproducing copyrighted songs, an AI could still use music to create a new form of interaction.
Move beyond reading information aloud. Let AI tell stories.
Imagine an AI that tells bedtime stories or reads aloud when users would rather listen than read.
Potential features include:
Reading books or other content when permitted.
Summarizing long texts or books and presenting the summaries orally.
Creating original fairy tales and stories.
Adjusting its delivery to suit the mood and atmosphere of a story.
The goal is not simply to convert text into speech.
It is to create an enjoyable listening experience through pacing, pauses, narrative flow, and expressive delivery.
When copyrighted material cannot be reproduced in full, the AI could offer alternatives, such as summaries or original stories, within applicable copyright limits.
Go beyond saving users the effort of reading. Create the experience of having someone tell them a story.
Move beyond choosing from a list of predefined voices. Let users create or customize the voices they want.
Users could select their preferred vocal characteristics or create personalized voices within the capabilities of the technology.
Potential applications include:
Creating calm, energetic, or expressive voices.
Talking to AI in a voice the user particularly likes.
Designing a personalized AI voice based on individual preferences.
Recreating the voice of a family member or friend with appropriate consent and authorization.
This idea also has a personal dimension.
What if you could talk with an AI that sounds like a friend who has passed away?
Listening to a recording of someone's voice can be a different experience from hearing that voice respond to you in a conversation.
Of course, an AI cannot restore that person's actual memories or thoughts. Recreating a voice also raises important questions about consent, authorization, and preventing misuse.
Nevertheless, when a voice carries special personal meaning, incorporating it into a conversational experience could offer a new kind of value.
The goal is not simply to choose a better voice, but to converse in a voice that means something to you.
Pronunciation should be accurate, but it should also be easy for users to understand.
During Korean voice conversations, an AI may pronounce English words with an accent that sounds excessively native to English speakers.
However, Korean incorporates many English-derived words into everyday speech.
Korean speakers commonly pronounce these words according to familiar Korean pronunciation patterns.
In this context, is emphasizing native English pronunciation always the best choice?
For example, when speaking Korean, the AI could use familiar Korean pronunciations for commonly used English loanwords while maintaining natural English pronunciation when reading English sentences or conversing in English.
Voice settings could offer options such as:
NATIVE: Natural native-language pronunciation.
LOCAL: Pronunciation familiar to speakers in the user's linguistic environment.
LEARNING: Clear pronunciation suited to language learning.
This feature would not need to be limited to Korean. Similar approaches could be explored for other languages, taking local pronunciation habits and linguistic environments into account.
However, LOCAL mode should not mechanically convert every English word into a localized pronunciation. The goal is to speak naturally in context.
Sounding like a native speaker is not always the best objective. Speaking in a way that users can comfortably understand matters, too.
The first four ideas introduce new experiences that voice AI can offer. The following ideas focus on situations in which voice is particularly useful.
Driving is an obvious use case for voice AI.
Other examples include cooking, cleaning, and walking, when looking at a screen or typing may be inconvenient.
Users could ask questions and receive information without interrupting what they are doing.
However, voice interaction can still distract drivers and should never compromise road safety. Driving mode should prioritize short, simple interactions, leaving complex conversations for when the user has stopped driving.
Many conventional interactions follow a simple pattern: the user makes a request, and the AI responds.
Voice AI could evolve beyond this pattern by proactively offering appropriate responses or actions based on conversational context.
When a user shares good news, the AI could congratulate them or offer to sing a short celebratory song. When a user wants a bedtime story, it could naturally continue the storytelling experience.
The key is to reduce the need for users to issue explicit instructions at every step.
However, proactive behavior is not always welcome. Users should remain in control and be able to adjust how and when the AI initiates interactions.
In voice conversations, the content of a response is only part of the experience. Speaking pace, conversational rhythm, and natural reactions also matter.
An AI could ask an appropriate follow-up question, respond briefly, or wait when necessary instead of immediately delivering a long answer.
This could help voice AI move beyond a repetitive question-and-answer interface toward a more natural conversational experience.
Achieving this will require continued improvements in response latency, speech recognition, and turn-taking.
Given the technical constraints of voice AI, there may be limits to an approach that prioritizes immediate responses for every task.
Instead, the interaction style should match the nature of the task.
Task
Appropriate Approach
Short questions and answers
Fast voice responses
Simple information requests while driving
Brief, low-distraction, hands-free interactions
Long-form storytelling
Natural, paced delivery
Spontaneous singing
Start with short melodies and lyrics
Complex questions and analysis
Allow sufficient processing time before explaining naturally
Detailed document comparison and editing
Combine text and voice
Rather than demanding the same level of immediacy for every voice interaction, AI should adapt its response speed and delivery to the task.
This approach does not attempt to eliminate technical limitations. It seeks to create better user experiences within those limitations.
The purpose of this proposal is not to suggest implementing all seven features at once.
Each idea addresses a different need, but they share a common direction.
Instead of treating voice AI as merely a voice version of text AI, we should actively develop the distinctive value that voice can offer.
This can be organized into three strategic areas.
Strategy
Core Value
Representative Ideas
New Experiences
Entertainment and emotional engagement
The Tone-Deaf Singer, Grandma's Storytime
Personalization
Individual preferences and linguistic environments
Custom Voice, KONGlish Mode
New Forms of Interaction
Convenience and natural conversation
Hands-Free Interaction, Proactive Interaction, Natural Conversation
The Tone-Deaf Singer, Grandma's Storytime, Custom Voice, and KONGlish Mode are concrete feature proposals.
Hands-free interaction and natural conversation address the environments and interaction models in which voice AI can be most useful.
Each feature could be tested independently and, where appropriate, combined with others.
For example, an AI could tell stories in a voice customized by the user or sing spontaneous songs using the user's preferred voice.
These possibilities would, of course, need to be validated through technical implementation and user feedback.
The development of voice AI should not be measured solely by its ability to speak faster or more accurately.
Response speed and accuracy are undoubtedly important.
However, voice is more than a means of delivering information. It is also a medium through which AI can tell stories, sing, express emotions, and communicate in voices that users prefer.
Voice AI can also adapt its communication style to users' environments and linguistic habits.
Rather than simply reproducing the capabilities of text AI, voice AI should actively develop experiences and forms of value that take advantage of what voice uniquely offers.
The goal is not voice AI that competes with text AI on the same terms, but voice AI that people choose for different reasons.