PRODUCT DESIGN · GEN AI · UX

Wix Legends: How we built a Video-First AI Digital Presence Platform


Wix Legends is a new, cutting edge initiative that introduces agentic AI personas to the Wix platform, enabling users to create, customize and deploy a live digital avatar that engages visitors in real time across synchronized video, voice, and chat. This next evolution of online presence allows creators and businesses to scale their personal brand 24/7 while fully tailoring their Legend's look, behavior, and knowledge base to effortlessly automate knowledge sharing, lead capture, FAQs and more.

Building the next evolution of online presence

Wix Legends marks a shift into a fully interactive and personal experience where your online presence isn't just a portfolio or a storefront - it’s a live, autonomous extension of your brand. It’s a huge opportunity to change how creators and businesses actually connect with people on the web.



As part of the core Wix Legends team, we were given total freedom to explore a completely unmapped territory for Wix, and I was incredibly excited to be part of this project from the ground up.

With no pre-existing Wix frameworks to follow, it was an incredible opportunity to help shape the product strategy, influence the tech stack, and design a futuristic conversational system completely from scratch as a standalone Wix offering.

Research & initial ideation

In early 2025, the team began market research and exploration of existing technologies that were available to us. We benchmarked the emerging AI landscape as well as evaluated voice and audio APIs like ElevenLabs and Hume to ground the platform. As we became more familiar with the market, we began to recognize opportunities and possible differentiators for our own solution. 


We recognized that most existing products remained within their specific domains: video avatars, voice cloning, or basic persona “knowledge”. It became clear that a single holistic solution that put them all together was missing - this became our big differentiator and north star. With Wix’s established reach and understanding of the potential user base that would adopt this product, we were excited to get started. 




Fully embracing conversational UI

Our ultimate goal was to give users the tools to build an interactive video avatar - a new, natural way for people to interact with AI in a more natural “face-to-face” conversation. Yet, as we began to map out the onboarding we realized that we were instinctively designing a system that didn’t adhere to these same ideas. Then it hit us: why use traditional patterns to build something that is meant to feel human?


“Her” (2013), Warner Bros. Pictures

We instinctively threw out the traditional blueprint we had been using and moved to our own video and voice technology to power the creation and management flows, making the system just as conversational as the end product itself. Besides being a unique experience in today’s landscape, we would be teaching the user how it feels to interact with an AI avatar as they were creating one for themselves. We knew that this approach was ambitious and planned to measure its performance as we developed it further.

Ditching the abstract circles and humanizing the OS

We now understood that our system needed to speak and interact with the user while walking them through the experience, yet we still didn’t know how it should look, sound or feel.

Examining how today's AI landscape handles voice conversation shows a clear pattern of abstract shapes and pulsating colors to represent AI during conversation, which to us felt cold and robotic and not the approach we wanted to take.



To break away from these sterile, faceless interfaces and create an interaction that felt distinctly human, we created "Atlas" - our own custom digital interviewer and OS designed to bring true face-to-face connection to the conversational experience. The creation of Atlas had various iterations that took us through many combinations of different voices, visual looks and most importantly communication style and tone-of-voice.

In parallel, we continued to incorporate Atlas into the system itself - giving her increasingly more capabilities such as extracting information automatically and leading the user through flows and answering questions.

 

Because replacing traditional flows and patterns with a more conversational system was highly experimental, we ran extensive testing and usability to optimize Atlas’s voice, pacing and demeanor so she felt natural rather than rigid and robotic.

Acting as the system’s “Onboarding” guide, Atlas had to juggle multiple skills simultaneously while keeping the user moving through the main goal we set for her:

  • Introduce the platform to the user and set expectations
  • Extract the required information from the user
  • Answer general & specific questions about the current screen
  • Facilitate a natural, personal & conversational experience


Atlas extracting information through conversation during the onboarding.

Our design risk paid off! Usability sessions and onboarding funnel data confirmed a major reduction in friction and dropoff rates compared to the original straightforward baseline. Users expressed that it was a real “wow” moment and felt unlike anything that they’ve experienced in the past. 

Atlas guides users seamlessly through the entire setup flow before transitioning into the main workspace where she remains as a persistent assistant capable of modifying settings and capturing data through casual dialogue.

Designing frictionless identity cloning

Early stages of the product depended heavily on existing technologies for video cloning. The initial version required reading out scripts and sitting still through silent calibration segments which led to high failure rates, frustration and significant drop off. We quickly understood that the experience needed to go in a different direction - reducing friction and bringing the user to a “wow” moment as quickly as possible.


First iteration of video cloning, requiring reading a script and remaining silent for 1 min.

Instead of continuing to battle with the technology, we came up with  an ingenious idea that helped us solve these issues and completely eliminiate the aforementioned friction points. We developed an internal process that could seamlessly create your video avatar from just a single uploaded image.

Updated video cloning flow using just a single image upload for the full creation process.

Behind the scenes, our upgraded flow takes the uploaded image and generates a high-quality and “clone-optimized” image (centering, lighting etc.) that is then automatically sent for video generation, resulting in a 100% video cloning success rate. We later added the ability for users to adjust and customize the image using generative prompts, as well as generate intent-based suggestions according to their goal (i.e. a bicycle mechanic would get a related image suggestion).

Generating your own custom Legend “style” based on your uploaded image.

For voice cloning, Atlas would record the user seamlessly and organically during the onboarding conversation with no need to read any script. Of course, fallbacks were put in place in case the user didn’t talk (or has no mic) during onboarding.

At the end of the video cloning process, we generate a short preview of your Legend before it’s sent for full generation. This gives users the chance to see and hear their Legend and continue making changes if they want.


Preview of the user’s video clone (including voice) that was generated from a single uploaded image.

Adapting to support Custom Personas

Testing showed that many users still preferred brand characters over personal clones of themselves, so we quickly shifted our focus to accommodate creating custom personas. This required a substantial overhaul of the onboarding flow and system architecture.


A selection of a few pre-made personas that we allow users to use for their Legend.

The original onboarding and workspace were redesigned to support both options - creating yourself, or choosing from pre-made stock avatars. We created a variety of pre-made video and voice pairings, while also enabling users to generate entirely custom personas from scratch. 


Voice customization - choose an existing voice or create your own.

This ultimately gave business owners full control over exactly what their Legend looks and sounds like - flexibility that proved crucial for user satisfaction and adoption. 

A flexible workspace experience

What does a conversational workspace look like? This was one of the biggest questions we had to tackle for users who finish the onboarding and want to manage their Legend.

We wanted to keep the great conversational experience from the onboarding inside the workspace as well, so Atlas is always 1 click away and ready to perform any action the user needs - from changing a setting, editing information or adding knowledge.


Atlas updating settings via conversation with the user.

Inner workspace pages remain accessible from the side menu at any time, allowing users to see and manually adjust all their information for full control. Since this is a completely new approach to interacting with a system, we are continuously monitoring usage to optimize the experience within the workspace.


Atlas within the workspace alongside inner-pages.

Extracting, collecting and managing knowledge

Creating your Legend’s look & feel is super fun and exciting, but we knew that adding knowledge quickly and easily was key to the success of the product, as well as showing users the actual value their Legend brings.

We split our approach to 2 different methods of knowledge building:  conversational and manual.


Atlas extracting knowledge from user conversation.

Conversational: Keeping true to our product ideology, we offer an interactive and conversational way to add knowledge through Atlas. By naturally talking to Atlas from within the workspace, she automatically knows when to extract information and where best to put it. Additionally, Atlas has the ability to change settings like custom instructions, behavior settings or bio information upon request.


Adding knowledge manually via the Knowledge page.

Manual: We realized that many users have pre-existing knowledge already on hand, so we created an engine for extracting knowledge from uploaded links, files, text and images. Users simply drop in their existing assets while the system processes the materials in seconds, enabling the Legend to immediately give accurate answers to visitor questions.

Creating an improvement loop

Because the Legend speaks on the user’s behalf, giving users the ability to analyze and improve its performance was crucial for us to include. How do users know that their Legend is being helpful and bringing value to visitors?

Legend owners have the ability to view conversations, see what the main conversation points were and filter by intent.



Additionally, its possible to dive into a specific conversation to see what users are interested in and how the Legend answers.



While its possible to manually view visitor conversation summaries to learn about satisfaction levels, we knew that we needed a dedicated system for improving the Legend’s answers. 

The “Knowledge Gap” feature dynamically summarizes user transcripts and highlights specific knowledge gaps where the Legend lacked information. Our system weighs every user message against its relevancy and whether there is enough knowledge to answer the question. If it detects a relevant missing subject, a knowledge gap is exposed to the owner as well as example questions that triggered the gap itself.

This allows Legend owners to proactively feed the system new data based on real user needs, which in turn brings more value to the users talking to the Legend. This extremely powerful feedback loop gives the Legend owner a true feeling that the Legend represents them in the best way possible.

The End-User Experience: Talking To a Legend

Putting in the effort of creating a Legend, customizing its behavior and adding knowledge culminates with sharing your legend with your audience. We focused on giving visitors a great low-latency video experience and high conversational value.


Experiencing a video call with a Wix Legend.

Visitors can choose whether they want a video call, voice call or just a text chat (both on mobile and desktop). Besides replying in voice and text, the Legend can expose interactable widgets such as video embeds, links, files and rich-text if it recognizes the need. These widgets are a powerful differentiator for our product and bring immense value to the end user, as well as allowing the Legend owner to share more with their visitors.

Widgets exposed by the Legend during conversation open the side panel - serving as both a history and chat.

What the future holds

As vibe-coding and LLMs become the new norm for retrieving information and executing tasks, its clear that the next generation of  products will look entirely different than what we know today. Making these interactions feel like realistic and natural conversational experiences remains a huge and exciting challenge for us.

This is why the Wix Legends team is continuing its focus on pushing the envelope further, exposing powerful integrations and adding more robust features so that your Legend can do more for you, and your users. 

All projects, photos & images © Copyright 2026, Roy Sherizly