Architectural Overview

To eliminate incoming call friction while maintaining full booking continuity with enterprise scheduling platforms (such as Zenoti REST APIs), the voice channel is routed into a voice-capable AI framework. Our implementations leverage two core architectural layers:

  • Telephony Layer: Telephony providers (Twilio, Telnyx, or Vapi) provision primary business phone numbers and stream call audio over WebSockets or Webhooks.
  • Integration Middleware Layer: Middleware server exposing tool endpoints mapped directly to enterprise Center and Booking API endpoints.

01. Google Gemini Direct Solution Native Multimodal Voice

An ultra-low latency architecture engineered for direct speech-to-speech interaction without intermediate transcription delays.

Architecture Stack

Twilio Media Streams → Custom Node.js/Python WebSocket Server → Gemini Multimodal Live API

How It Works

1. Audio Streaming

Inbound calls stream raw PCM audio directly to the Gemini Multimodal Live API over a persistent WebSocket session.

2. Native Execution

Gemini handles speech recognition, conversational logic, and voice synthesis natively within a single model loop, maintaining low latency (~300–500 ms).

3. Real-Time Tool Execution

Gemini uses Function Calling during the live conversation to query booking systems for open time slots or provider availability and executes booking payloads.

Key Advantages

  • Ultra-low conversational latency (~300–500 ms).
  • Natural voice interruptions and barge-in support.
  • Simplified technical stack requiring no separate STT or TTS vendor subscriptions.

02. OpenClaw + Claude Solution Pipeline Integration

A highly customizable pipeline architecture designed for complex scheduling rules, multi-step agentic workflows, and persistent memory across visits.

Architecture Stack

Twilio Voice Webhook → STT Engine (Deepgram) → OpenClaw (Claude Model) → TTS Engine (ElevenLabs / Cartesia)

How It Works

1. Real-Time Transcription

Incoming caller audio is transcribed in real time via Deepgram streaming.

2. Agentic Reasoning

OpenClaw processes text streams using Claude (e.g., Claude 3.5 Haiku for speed or Claude 3.7 Sonnet for complex requests) while maintaining persistent session context across interactions.

3. Tool Execution

OpenClaw triggers tool executions directly against REST APIs or through custom Model Context Protocol (MCP) servers.

4. Speech Synthesis

Response strings are converted back to natural audio via ElevenLabs or Cartesia and played back to the caller seamlessly.

Key Advantages

  • Superior instruction-following for multi-step service scheduling rules.
  • High precision in handling complex appointment modifications and edge cases.
  • Persistent caller memory and history tracking across recurring visits.

REST API Function Mapping

Standardized tool mappings connecting conversational AI agents directly to enterprise booking backends:

Interaction API Endpoint AI Tool Function
Identify Client GET /v1/guests?phone={number} lookup_guest_by_phone()
Check Availability GET /v1/centers/{center_id}/slots get_barber_slots()
Hold Appointment POST /v1/bookings reserve_booking_slot()
Confirm & Text POST /v1/bookings/{id}/confirm confirm_and_send_sms()

Account Setup & Infrastructure Provisioning

Deployment Prerequisites Checklist

  • Cloud Platform Account: Configure cloud console project, enable Vertex AI / Developer APIs, and activate Multimodal Live API access for bidirectional WebSocket streaming.
  • Telephony Provisioning: Provision dedicated telephony accounts, import primary business numbers, and enable Webhook/Media Streams streaming to middleware endpoints.
  • API Credentials: Obtain enterprise API keys, Center IDs, and OAuth tokens to grant guest lookup and booking capabilities.
  • Middleware Infrastructure: Provision serverless container environments (e.g., Cloud Run) to execute persistent WebSocket bridges.