Multi-modal chat: voice, video, and AI avatar interaction on the website
Taylor Jennings
Today Chili Piper's website chat is text-only. Customers want the on-site experience to be multi-modal: a visitor lands, and can choose to keep typing, switch to an audio call, or jump to a video call, with an AI avatar that can answer questions face-to-face rather than as a text bot.
Requested capability: extend the chat widget so a visitor can move between text, voice, and video in a single session, and so an AI avatar can conduct the conversation.
Why it matters: voice and video are becoming a core way people expect to interact with websites, and a human-like avatar raises engagement and trust at the top of the funnel. There's also a qualification angle, an AI avatar handles early conversations without the "humans are incentivized to close, not help" problem, so bots take first contact and humans get looped in only when it counts. This is squarely on the 2026 website-conversion narrative.
Note: this bundles two separable pieces, (a) an AI avatar persona and (b) live audio/video escalation from chat. Keeping them together for now given early/single-customer demand; split if either gains independent traction.