Breaking the Boundaries of Interaction: Creating Human-Like Live Co-Performances with Multiple AI Characters
26-08-26
Author: Bull, BD & MKT Director | Editor: Helen, Deputy Marketing Manager
Over the past period, Ubitus’ AI VTuber UbiChan has not only reached major viewership milestones and accumulated more than 10,000 hours of companionship, but has also, through the joint efforts of the UbiChan team, begun experimenting with two-person and even three-or-more-character AI interactions during live streams.
The goal is to challenge the limitations of today’s LLM frameworks and enable UbiChan to interact naturally in the same virtual room with pet cats such as Little Lemon and Xiao Hua, as well as online viewers.

Multi-character co-performance: characters talk to each other and question one another

Multi-character co-performance: viewers may even join the scene, creating a live performance with multiple participants
When many companies imagine an AI brand ambassador, they often assume that connecting the latest large language model, or LLM, will naturally make everything work and attract a crowd of fans. In reality, however, this often requires many carefully designed interaction segments.
When the scene shifts to “multiple pure AI characters chatting on the same screen,” if the system simply lets several chatbots take turns reading scripts, the result becomes extremely stiff and mechanical.
Why Does Multi-Character Chat Easily Feel Like “Taking Turns Reading Scripts”?
Most dialogue systems follow a simple pattern: one input, one output. But in real human group conversations, the basic unit is not a complete sentence, but an “event.”
Someone may want to correct another person. Someone may understand the joke and laugh. Someone may jump in to answer first. Or someone may be completely at a loss.
If the system only selects the next speaker after one person finishes speaking, the characters cannot produce authentic reactions while others are talking.
To solve this problem, the UbiChan team treats the entire live-streaming system as a real-time improvisational performance system. The natural feeling we pursue is not about letting multiple characters endlessly take turns speaking, but about giving them states, reactions, and even allowing a controlled degree of overlap, interruption, and collective silence.
Hybrid Architecture: A Single Core Director and Event-Driven Interaction
Under the limitations of the current LLM framework, where each input generally produces one output, we do not recommend directly running four independent LLMs that act separately when creating natural co-performance among multiple AI characters.
Instead, the UbiChan team designed a hybrid architecture consisting of a single core director, multi-character states, event-driven interaction, and interruptible audio scheduling.
- Single Core Director Agent: After summarizing and organizing viewers’ comments, the core model does not merely ask, “Who should answer next?” Instead, it simultaneously simulates the internal intentions, impulse intensity, and emotions of all characters on stage, such as UbiChan, Little Lemon, and Xiao Hua.
- Event-driven interaction: Interaction is no longer based on fixed turns. Instead, it is triggered by events such as semantics, timing, or performance actions. Even when characters are not speaking, they remain in a “living” state and are always ready to react.
Interruptions, Quick Answers, and “Meaningful Silence”
In human-like interaction, imperfect micro-reactions are often more realistic than long speeches. Therefore, we introduced the following designs into the system:
- Mutual interruption: Triggered by content and character state. When a major mistake occurs or the main speaker talks for too long, another character can step in at a clause boundary, allowing 300 to 800 milliseconds of audio overlap.
- Low-cost reactions: Short voice clips and animations such as laughter, agreement, and sighs are added. These reactions are triggered instantly by rules and can greatly reduce the perceived latency of LLM and TTS processes.
- Dramatic silence: When encountering a bad joke or an awkward question, 1.5 to 3 seconds of collective silence, paired with actions or facial expressions, becomes part of the performance rather than a sign that the program has crashed.
During character co-performance, dialogue can be interrupted naturally.
Turning Complexity into a Smooth Audiovisual Experience
To turn these intentions into reality, the timing control of audio and animation is critical. Each character is treated as an independent audio channel, supporting streaming output, instant cancellation, fade-out, and continuation from clause boundaries.
Visually, we use a single scene to manage all characters, allowing interaction among characters within the same space to become the core of the shared-stage experience.
These seemingly behind-the-scenes technical details are the results of the UbiChan team’s careful and continuous refinement. From the companionship of a single VTuber to the challenge of smooth multi-character co-performance, this is an ongoing evolution.
Admittedly, with the current Core Director Agent technology, there are still many immature aspects. When large amounts of short-term, mid-term, and long-term memory, as well as RAG integration, are introduced, the response speed is not always as immediate as we might expect.
However, perhaps in the near future, when large language models evolve to the next stage, multi-AI-character co-performance may once again find new design methods and achieve another breakthrough.
No matter what, we deeply understand that supporting a live stream with a true “show experience” requires extremely low-latency computing power and multitasking collaboration behind the scenes. This is also the original intention behind Ubitus’ continued optimization of GPU cloud infrastructure and the launch of the UbiOne platform:
to keep technical complexity backstage, while allowing the soul of interaction to shine on stage.
About Ubitus
As a member of the NVIDIA Connect program, Ubitus leverages NVIDIA’s support and cutting-edge GPU technology to accelerate AI innovation. The company delivers advanced AI solutions, including UbiGPT (a large language model), UbiONE (an AI-powered avatar creation platform), and UbiArt (an image generation tool), providing customized solutions to meet the diverse needs of various industries.
As a cloud gaming pioneer, Ubitus enables Nintendo and other game companies to establish cloud gaming services and supports the global streaming of multimedia content, including interactive and virtual reality experiences.
Contact
TEL : +886-2-2717-6123 (Taipei)
+81-3-6435-3295 (Tokyo)
Media contact: pr@ubitus.ai
Business inquiry: contact@ubitus.ai
Website:www.ubitus.ai