Documentation

Wirevox AI Documentation

Everything you need to build, deploy, and scale hyper-realistic AI phone agents. Explore our guides, API references, and core concepts.

Section 1 — Getting Started

The complete guide to signing up, setting up your workspace, and making your first AI voice call with Wirevox.


1.1 Platform Overview

What Is Wirevox?

Wirevox is an enterprise-grade AI voice platform that lets you build, deploy, and manage intelligent phone agents. These agents can answer inbound calls 24/7, book appointments, route callers to humans, respond to FAQs using your own business data, and integrate with your existing CRM and calendar tools — all without writing a single line of code.

Who Is It For?

Wirevox is built for:

  • Small & mid-size businesses that want a 24/7 AI receptionist to answer calls, qualify leads, and book appointments.
  • Agencies & resellers that build and manage AI agents on behalf of their own clients, using Wirevox's white-label agency layer.
  • Enterprise teams that need multi-agent architectures with CRM integrations, knowledge bases, and strict tenant isolation.
  • Developers & SaaS builders that want to embed conversational AI telephony into their products.

Core Value Proposition

Capability What It Means
Modular Voice Pipeline Real-time audio streaming bridging Telnyx telephony with Deepgram/Gladia (STT), OpenAI/Gemini/Claude (LLMs), and ElevenLabs/Deepgram/OpenAI (TTS). Sub-500ms response latency.
RAG Knowledge Base Upload PDFs, DOCX, CSVs, or crawl your website — your agent becomes an expert on your business data.
Functions & Tool Calling Your agent can book appointments, send SMS, transfer calls, fire webhooks, and sync with CRMs — all mid-conversation.
Multi-Agent Architecture Build specialized agents (receptionist, sales, support) and route callers between them.
Agency / White-Label Resell AI agents under your own brand with custom domains, plans, and subaccount management.
9 Native Integrations Google Calendar, Outlook Calendar, Clio, Clio Grow, Filevine, OpenDental, Jobber, Email (Mailgun/SMTP).

1.2 Creating an Account

Wirevox supports two sign-up methods: email + password and Google OAuth.

Option A: Sign Up with Email

  1. Navigate to the sign-up page at /signup.
  2. You'll see the heading "Create account" with the subtitle "Start deploying voice agents in minutes."
  3. Fill in the following required fields:
    • First name — e.g., Jane
    • Last name — e.g., Doe
    • Email — e.g., you@company.com
    • Password — must meet all four requirements (see below)
  4. Click "Create account".
Password Requirements

As you type your password, a real-time strength meter appears with four progress bars. Your password must satisfy all four criteria to turn the bar fully green:

Requirement Rule
Length At least 8 characters
Lowercase Contains at least one lowercase letter (a–z)
Uppercase Contains at least one uppercase letter (A–Z)
Number Contains at least one digit (0–9)

The bar color changes dynamically as you satisfy more criteria:

  • 1 met → Red
  • 2 met → Amber
  • 3 met → Blue
  • 4 met → Green ✓
Email Verification

After submitting, you're redirected to a "Check your inbox" screen. Wirevox sends a verification link to your email address.

  • The verification email is sent via Supabase Auth.
  • A resend button appears with a 60-second cooldown timer. After the cooldown, you can click "Resend Email" to request a new link.
  • If you entered the wrong email, click "Change email" to go back to the sign-up form.

Once you click the verification link in your inbox, your account is activated and you're signed in automatically.

Option B: Sign Up with Google

  1. Navigate to /login or /signup.
  2. Click the "Continue with Google" button at the top of the form.
  3. Google's account picker opens (Wirevox uses prompt: 'select_account' so you can choose which Google account to use).
  4. After authorizing, you're redirected back to Wirevox.
  5. On first Google sign-in, you'll see the "Welcome to Wirevox!" confirmation page where:
    • Your first and last name are pre-filled from your Google profile.
    • Your email is shown (read-only, greyed out).
    • Confirm your name and click "Create Account" to proceed.

Note: If you originally signed up with Google and later try to use email/password login, Wirevox shows a helpful error message: "It looks like you registered this email using Google. Please click 'Continue with Google' instead."

By signing up (either method), you agree to Wirevox's Terms of Service and Privacy Policy. This notice appears at the bottom of both the sign-in and sign-up forms.


1.3 Workspace Setup

After your account is created and email is verified, the Onboarding Guard checks your profile. If your onboarding_complete flag is false, you're routed to the Workspace Setup page before you can access the dashboard.

Workspace Setup Form

The workspace setup screen has the heading "Set up your workspace" with the subtitle "Tell us a bit about your business to get started." It collects four fields in a 2×2 grid:

Field Required? Default Description
Workspace Name ✅ Yes The name of your business or organization (e.g., "SunPath Energy"). This becomes the name displayed in the sidebar and workspace switcher.
Website Link Optional Your business website URL. Used for context and can be crawled later for your knowledge base.
Industry Optional Searchable dropdown with 36 industry categories (see list below). Helps Wirevox tailor prompt templates and recommendations.
Timezone Optional America/New_York Searchable timezone selector. Used by agents for scheduling-related functions.

Click "Create workspace" to save. Your profile is marked onboarding_complete: true and you're redirected to the dashboard at /overview.

Supported Industries

The industry dropdown includes a search filter and the following categories:

Accounting & Tax Services Agriculture & Farming Architecture & Planning
Arts & Entertainment Automotive (Dealerships, Repair, etc) Beauty & Personal Care
Biotechnology Construction & Contracting Consulting Services
Dental E-Commerce & Retail Education & E-Learning
Energy & Utilities Financial Services & Wealth Management Fitness, Wellness & Sports
Food & Beverage Government & Public Sector Healthcare & Medical
Home Services (Plumbing, HVAC, Roofing) Hospitality, Travel & Tourism Human Resources & Staffing
Insurance IT Services & Cybersecurity Legal Services
Logistics & Supply Chain Manufacturing Marketing, Advertising & PR
Media & Publishing Non-Profit & NGO Real Estate
Real Estate Investing Software & SaaS Telecommunications
Transportation & Delivery Veterinary Services Other

1.4 Onboarding Survey

Shortly after your first login, a floating modal appears over the dashboard asking four quick questions. This survey can be skipped at any time via the "Skip" link in the top-right corner.

The survey is designed to help the Wirevox team understand its user base and tailor the product experience. It consists of 4 multiple-choice questions shown one at a time with a progress bar:

Question 1 — Where did you hear about us?

  • Google / Search
  • LinkedIn
  • Twitter / X
  • Friend / Colleague
  • YouTube / Podcast
  • Other

Question 2 — What is your company size?

  • Just me
  • 2 – 10
  • 11 – 50
  • 51 – 200
  • 200+

Question 3 — Current monthly spend on voice / phone tasks?

  • $0 (Handling it myself)
  • $1 – $1,000 / mo
  • $1,000 – $5,000 / mo
  • $5,000 – $10,000 / mo
  • $10,000+ / mo

Question 4 — What is your main goal with Wirevox?

  • Automate inbound support
  • Qualify outbound leads
  • Handle appointment booking
  • Just exploring

Clicking any option immediately advances to the next question (no submit button). After the final question, your answers are saved to your profile and the modal closes. If you skip the survey, it is marked as survey_completed: true with { skipped: true } and will not appear again.


1.5 Dashboard Tour

Once onboarding is complete, you land in the main dashboard. The interface is a two-panel layout: a collapsible sidebar on the left and the main content area on the right.

The sidebar contains your workspace context and navigation links. It has two states: expanded (default) and collapsed (auto-collapses when you enter settings, logs, knowledge, functions, playground, or agency pages).

Workspace Section (Top)

At the top of the sidebar is the Workspace Switcher — showing your workspace name and logo. Clicking it opens a dropdown to switch between workspaces or create a new one.

Below the workspace switcher, a credit balance indicator shows your remaining credits (included_credits + purchased_credits).

Icon Label Route Description
Vox /w/:id/vox AI agent builder assistant (coming soon)
🤖 Agents /w/:id/agents View, create, and configure your AI agents
🧪 Playground /w/:id/playground Test voices and agent configurations
📊 Analytics /w/:id/analytics Dashboard with call volume, usage trends, and active call indicators
📋 Logs /w/:id/logs Call history, transcripts, and chat logs
👥 Contacts /w/:id/contacts Leads and contacts extracted from calls
📚 Knowledge /w/:id/knowledge Manage your RAG knowledge base (documents, websites, text)
🔌 Functions /w/:id/functions Create and manage custom functions (webhooks, built-in tools)
📞 Numbers /w/:id/numbers Buy and manage phone numbers
🏢 Agency /w/:id/agency Agency dashboard (only visible on agency-tier workspaces)
⚙️ Settings /w/:id/settings Workspace preferences, billing, team, integrations, and more
Bottom Section

The bottom of the sidebar shows your profile avatar (initials), your display name (or truncated email), and a dropdown menu with:

  • Account Settings/account (profile, auth, system preferences, danger zone)
  • Sign Out

Mobile Navigation

On mobile devices, the sidebar is replaced by a bottom navigation bar with five tabs: Home (analytics), Agents, Logs, Plans (billing), and More.


1.6 Your First Agent — Step by Step

This walkthrough takes you from zero to a working AI voice agent in under 5 minutes.

Step 1: Open the Agents Page

Click "Agents" in the sidebar. If you have no agents yet, you'll see an empty state with a "+ New Agent" button.

Step 2: Select Agent Type

Click "+ New Agent". A modal appears with three agent type cards:

Agent Type Status Description
Inbound ✅ Available (marked "NEW") Answers inbound calls, handles customer FAQs, books appointments, and routes complex issues to humans 24/7.
Outbound Agent 🔜 Coming Soon Runs proactive voice campaigns to qualify leads and book meetings on autopilot.
AI Chat Agent 🔜 Coming Soon Resolves visitor queries and captures leads directly from your website in real-time.

Click the Inbound card to proceed.

Step 3: Choose a Prompt Template

After selecting "Inbound," the Template Selector opens. This is where you pick a starting configuration for your agent. You can:

  • Filter by industry using the horizontal category scroller.
  • Filter by language using the dropdown (English, Spanish, French, German, Italian, Portuguese, Dutch).

Available templates include:

Template Industry Language
Blank (start from scratch)
General Receptionist General English
Dental Office Dental English
Dental Office (French) Dental French
Dental + Google Calendar Dental English
Legal Intake Legal English
Real Estate Agent Real Estate English
Real Estate (French) Real Estate French
Home Services Home Services English
Med Spa / Aesthetics Healthcare English
Insurance Agency Insurance English
Auto Dealership Automotive English
Salon & Barbershop Beauty English
Salon (French) Beauty French
Restaurant Food & Beverage English
Restaurant (French) Food & Beverage French
Hotel Front Desk Hospitality English
Hotel (French) Hospitality French
Fitness Studio Fitness English
E-Commerce Support E-Commerce English
IT Helpdesk IT Services English
Mortgage Broker Financial English
Childcare Center Education English
Education / Tutoring Education English
Travel Agency Travel English
Event Venue Events English
Veterinary Clinic Veterinary English

Each template pre-fills:

  • Agent name — e.g., "Sarah" for dental, "James" for legal
  • System prompt — a complete, industry-specific prompt with guardrails
  • Voice — a matching TTS voice (voice ID + provider)
  • Template variables — placeholders like {{business_name}}, {{business_hours}}

Select a template and click "Create Agent" to generate your agent.

Step 4: Configure Your Agent

After creation, you're taken to the Agent Settings page — a comprehensive configuration panel where you can customize everything about your agent. Key sections:

  1. System Prompt — The AI's personality, instructions, and behavior rules. You can edit the template-generated prompt or write your own.
  2. AI Model — Choose from GPT-4o, GPT-4.1, GPT-5.x, Claude 3.5/4.5/4.6 Sonnet, Gemini Flash, and more.
  3. Voice — Pick from 50+ voices across ElevenLabs, Deepgram Aura-2, and OpenAI TTS providers. Preview each voice before selecting.
  4. STT (Speech-to-Text) — Choose between Deepgram Nova-3 (mono or flux) and Gladia.
  5. Language — Select the primary language and optionally enable bilingual mode.
  6. Call Settings — Configure max call duration (1–30 minutes), greeting message, and interruption sensitivity.
  7. Functions — Attach built-in tools (SMS, transfer, booking) or custom webhooks.
  8. Knowledge — Link knowledge base sources to give your agent context about your business.

Step 5: Make a Test Call

Every agent has a built-in Test Call panel accessible from the agent settings page. This lets you talk to your agent directly from your browser — no phone number required.

  1. Click the phone icon or "Test" button in the agent settings header.
  2. The test panel opens with a visual "Liquid Aura" orb animation.
  3. Click the call button to connect. Your browser microphone activates and streams audio to the backend via WebSocket.
  4. Talk to your agent. The AI responds in real-time with the voice and behavior you configured.
  5. Click "Hang Up" to end the test call.

Note: Test calls consume credits just like real calls. The cost depends on your agent's AI stack (STT + LLM + TTS combination).

Step 6: Buy a Phone Number and Go Live

Once you're happy with your agent's behavior:

  1. Go to Numbers in the sidebar.
  2. Search for a phone number by area code or country.
  3. Purchase the number (pricing depends on your plan — each plan includes a certain number of free phone numbers).
  4. Assign the number to your agent.
  5. Your agent is now live — callers to that number will be answered by your AI agent 24/7.

1.7 Glossary

Key terms you'll encounter throughout the Wirevox platform:

Term Definition
Agent An AI-powered phone assistant that you create, configure, and deploy. Each agent has its own system prompt, voice, AI model, and phone number.
Workspace An isolated environment containing your agents, phone numbers, knowledge bases, functions, and billing. You can create multiple workspaces.
Credits Abstract units of usage. Every call minute, SMS, and function execution is measured in credits. The dollar value of a credit depends on your subscription plan.
Included Credits Credits that come with your monthly subscription. These reset every billing cycle.
Purchased Credits One-time credit top-ups bought separately. These never expire and are never reset, even if you cancel your plan.
STT (Speech-to-Text) The service that converts the caller's spoken audio into text. Wirevox supports Deepgram Nova-3 and Gladia.
LLM (Large Language Model) The AI brain that processes the caller's words and generates intelligent responses. Wirevox supports OpenAI (GPT), Anthropic (Claude), and Google (Gemini) models.
TTS (Text-to-Speech) The service that converts the AI's text response back into natural-sounding speech. Wirevox supports ElevenLabs, Deepgram Aura-2, and OpenAI TTS.
Voice Pipeline The real-time audio processing chain: Phone Audio → STT → LLM → TTS → Phone Audio. This is the "modular pipeline" that powers every call.
Realtime Model An alternative pipeline mode using speech-to-speech models (OpenAI Realtime or Gemini Live) that skip the separate STT/LLM/TTS stages.
RAG (Retrieval-Augmented Generation) A technique where the AI searches your uploaded documents/websites for relevant information and injects it into its context before answering. This is what powers the Knowledge Base.
Knowledge Base A collection of documents, website pages, and text sources that your agent can search during calls to answer questions accurately. Uses vector embeddings stored in PostgreSQL (pgvector).
Functions Actions your agent can perform mid-conversation, like booking an appointment, sending an SMS, transferring to a human, or firing a webhook to Zapier/Make.com. Powered by OpenAI Function Calling.
Barge-In When a caller interrupts the AI while it's speaking. Wirevox uses Voice Activity Detection (VAD) to instantly stop the AI's speech and listen to the caller.
Backchanneling Filler words the AI uses while processing (e.g., "Mhm", "I see", "Un moment" in French). These are language-specific and configurable.
Carryover Billing Wirevox's per-second billing system. Short calls accumulate seconds across multiple calls for the same agent, and you're only charged when a full minute is reached. This prevents overcharging on brief calls.
System Prompt The set of instructions that define your agent's personality, behavior, knowledge, and constraints. This is what makes each agent unique.
Template Variables Placeholders in your system prompt (e.g., {{business_name}}, {{caller_name}}) that are dynamically replaced with real data at call time.
Concurrent Calls The number of simultaneous phone calls your account can handle at once. This limit depends on your subscription plan and any purchased add-ons.
Post-Call Processing Actions that happen automatically after a call ends — like generating a summary, extracting lead information, or firing a webhook.
Agency A reseller account that manages AI agents on behalf of multiple client businesses (subaccounts), with white-label branding.
Subaccount A client workspace created and managed by an agency. Subaccounts have their own agents, calls, and usage — fully isolated from other subaccounts.
Telnyx The telephony provider that powers Wirevox's phone number provisioning, call routing, and media streaming.
Supabase The backend infrastructure provider used for the database (PostgreSQL), authentication, file storage, and real-time subscriptions.

Next: Section 2 — Core Concepts


Section 2 — Core Concepts

A deep dive into the fundamental building blocks of the Wirevox platform: workspaces, agents, the voice pipeline, realtime models, the credit system, and multi-agent architecture.


2.1 Workspaces

A workspace is the top-level organizational unit in Wirevox. Everything — agents, phone numbers, knowledge bases, functions, integrations, billing, and team members — lives inside a workspace. Workspaces are fully isolated from each other.

What a Workspace Contains

Resource Description
Agents AI voice assistants you build and deploy
Phone Numbers Purchased from Telnyx, assigned to agents
Knowledge Bases Documents, websites, and text sources for RAG
Functions Custom webhooks and built-in tools (SMS, transfer, booking)
Integrations OAuth connections to CRMs and calendars
Team Members Users you've invited with specific roles
Billing Subscription plan, credits, usage tracking
Call History Logs of all inbound/outbound calls and leads

Multi-Workspace Support

Users can own and belong to multiple workspaces. Common use cases:

  • Multiple businesses — A user who owns a dental practice and a law firm can create separate workspaces for each with different agents, numbers, and billing.
  • Separate environments — A development workspace and a production workspace.
  • Team access — A user might be the owner of their own workspace and a team member in someone else's.

Workspace Types

The system recognizes three workspace access types, determined by the tenantContext:

Access Type Description
personal A standard workspace owned directly by the user. This is the default.
agency A workspace on the Agency tier that can create and manage subaccounts.
subaccount A workspace created by an agency for one of their clients. The subaccount user sees a white-labeled dashboard controlled by the agency.

Workspace Switcher

The Workspace Switcher appears at the top of the sidebar and lets you:

  1. View all workspaces you have access to in a scrollable dropdown list.
  2. Switch between workspaces by clicking on one — the URL updates to /w/:workspaceId/... and all data reloads.
  3. Create a new workspace via the "+ New workspace" button at the bottom of the dropdown. This opens a modal with the same fields as the initial setup: Workspace Name (required), Website Link, Industry, and Timezone.

Each workspace gets a unique gradient avatar (one of 10 curated gradients like "Aurora Blue", "Emerald Glow", "Midnight AI", etc.) derived from its ID hash. If you upload a custom logo, it replaces the gradient.

URL Structure

Every workspace-scoped route is prefixed with the workspace ID:

/w/{workspaceId}/agents
/w/{workspaceId}/agents/{agentId}
/w/{workspaceId}/knowledge
/w/{workspaceId}/settings/billing
/w/{workspaceId}/agency/overview

The active workspace ID is persisted in localStorage under the key wirevox_active_workspace, so it survives page refreshes and new tabs.

Workspace Resolution Priority

When you load the app, the system determines which workspace to activate using this priority:

  1. URL parameter — If the URL contains /w/:workspaceId, that workspace is used (if the user has access).
  2. localStorage — Falls back to the last-used workspace stored in wirevox_active_workspace.
  3. First workspace — If neither is available, the user's first workspace is selected, with a preference for subaccount → agency → personal.

Workspace Settings

Each workspace has its own settings accessible at /w/:id/settings:

Setting Page Description
Preferences Workspace name, timezone, language
Integrations OAuth connections (Google Calendar, Outlook, Clio, etc.)
Billing Subscription plan, credit balance, top-ups, invoices
Team Invite members, manage roles
Usage Credit usage breakdown, call analytics
Notifications Email notification preferences
Danger Zone Delete workspace

2.2 Agents

An agent is an AI-powered voice assistant that answers phone calls on your behalf. Each agent is a fully self-contained configuration: its own personality (system prompt), voice, AI model, phone number, knowledge, and tools.

Agent Types

Wirevox supports three agent types, though only Inbound is currently available:

Type Status Direction Description
Inbound ✅ Live Receives calls Answers incoming calls to a phone number. Use cases: receptionist, support, scheduling, intake, FAQ.
Outbound 🔜 Coming Soon Makes calls Proactively dials a list of contacts for sales, lead qualification, appointment reminders.
AI Chat 🔜 Coming Soon Text-based Resolves visitor queries via a website chat widget.

Agent Configuration

Every agent has a comprehensive settings page with the following configurable properties:

Property Description Default
Name Display name shown in the dashboard From template
System Prompt The AI's instructions, personality, and behavior rules From template
AI Model The LLM that powers the agent's intelligence gpt-4o-mini
Pipeline Mode modular (STT→LLM→TTS) or realtime (speech-to-speech) modular
Voice The TTS voice used to speak to callers aura-asteria-en (Deepgram)
STT Provider Speech-to-text engine: auto, nova-3, flux, or gladia auto
Primary Language The agent's primary spoken language en (English)
Secondary Language Optional bilingual mode None
Greeting Message A custom first sentence spoken when the call connects None (AI generates naturally)
Max Call Duration Hard time limit for calls (1–30 min, system max 40 min) 40 min (system default)
Interruption Sensitivity Number of words required to trigger barge-in 3 words
Silence Timeout Seconds of silence before the AI prompts "Are you still there?" 15 seconds
Hangup Timeout Seconds of total silence (no user speech) before auto-hangup 30 seconds
Patience Level How long the AI waits after the user stops speaking before responding: low (600ms), medium (2.6s), high (4.6s) low
Knowledge Base Linked knowledge sources for RAG-powered answers None
Functions Attached tools (SMS, transfer, booking, webhooks) None
Phone Number The Telnyx number assigned to receive calls None (must be purchased)

Agent Lifecycle

  1. Created — Agent is generated from a template (or blank). It exists in a draft state.
  2. Configured — User customizes the prompt, voice, model, and attaches functions/knowledge.
  3. Published — Agent configuration is "locked" as a published version. Version history is maintained so you can roll back.
  4. Live — A phone number is assigned. Calls to that number are routed to this agent in real-time.

Dynamic Prompt Compilation

Before every call, Wirevox's compileSystemPrompt function transforms the agent's raw system prompt into a fully contextualized prompt by:

  1. Injecting custom variables — Replacing {{business_name}}, {{business_hours}}, and any user-defined {{variable_name}} placeholders with their configured values.
  2. Injecting real-time context — Prepending today's date, current local time (in the agent's timezone), and the caller's phone number (formatted as XXX-XXX-XXXX).
  3. Injecting knowledge base instructions — If a knowledge base is linked, its description/instructions are appended.
  4. Overriding unavailable tools — If the prompt references a function (via [[Function: tool_name]]) that is not connected, a critical system override is injected telling the AI to apologize and offer to take details instead of hallucinating the action.

Example of the injected real-time context block:

[REAL-TIME CONTEXT]
TODAY'S DATE: Wednesday, July 16, 2026 (2026-07-16)
CURRENT LOCAL TIME: 12:35 AM
Use this date and time as the reference for all relative dates and business hours reasoning.

CALLER PHONE (ALREADY KNOWN): 416-555-1234. This is a confirmed fact — do NOT ask for it...
--------------------------------

2.3 The Voice Pipeline (Modular Mode)

The Voice Pipeline is the real-time audio processing engine that powers every phone call. In its default modular mode, it chains together four independent services:

┌──────────┐     ┌──────────┐     ┌──────────┐     ┌──────────┐
│  Phone   │────▶│   STT    │────▶│   LLM    │────▶│   TTS    │
│  Audio   │     │ (Speech  │     │ (AI      │     │ (Text    │
│ (Telnyx) │◀────│  to Text)│     │  Brain)  │     │  to      │
│          │     │          │     │          │     │  Speech) │
└──────────┘     └──────────┘     └──────────┘     └──────────┘
     ▲                                                   │
     └───────────────────────────────────────────────────┘
                    Audio sent back to caller

How a Call Flows

  1. Telnyx receives an inbound call and opens a WebSocket media stream to the Wirevox backend at wss://.../media-stream/:agent/:call.
  2. Raw μ-law audio (8kHz, base64-encoded) flows from the phone into the server.
  3. The STT provider (Deepgram or Gladia) receives the audio stream via its own WebSocket and returns real-time transcripts (both interim and final).
  4. The transcript is accumulated, filtered through barge-in logic (see below), and when the user finishes speaking, the text is flushed to the LLM.
  5. The LLM (OpenAI, Gemini, or Claude) generates a streaming response, sentence by sentence.
  6. Each sentence is sent to the TTS provider (Deepgram, OpenAI, or ElevenLabs), which synthesizes speech audio.
  7. The TTS audio is streamed back through the Telnyx WebSocket to the caller's phone in real-time.

Supported Providers

Speech-to-Text (STT)
Provider Model Use Case Credits/min
Deepgram Nova-3 (monolingual) Single-language agents, lowest latency 2
Deepgram Nova-3 Flux (multilingual) Bilingual agents, automatic language detection 2
Gladia Default Alternative STT with different accuracy characteristics 3

The auto STT setting selects Flux for bilingual agents and Nova-3 for monolingual agents.

Large Language Models (LLM)
Provider Models Credits/min
OpenAI GPT-4o Mini, GPT-4o, GPT-4.1 (Mini/Nano), GPT-5 (Mini/Nano), GPT-5.1, GPT-5.2, GPT-5.4 (Mini/Nano) 1–5
Anthropic Claude 3.5 Sonnet, Claude 4.5 Sonnet, Claude 4.6 Sonnet 3–5
Google Gemini 1.5 Flash, Gemini 2.0 Flash, Gemini 2.5 Flash 1
Text-to-Speech (TTS)
Provider Voices Credits/min
Deepgram Aura-2 voices (Asteria, Orion, Luna, Stella, Athena, Hera, etc.) 2
OpenAI Alloy, Echo, Fable, Onyx, Nova, Shimmer 2
ElevenLabs Premium ultra-realistic voices (Rachel, Drew, etc.) 10

TTS provider is auto-detected from the selected voice ID — Wirevox inspects the voice name and routes to the correct provider.

Barge-In (Interruption Detection)

Barge-in is how Wirevox detects when a caller interrupts the AI mid-sentence. The system uses a sophisticated multi-layered approach:

  1. Word count threshold — By default, the caller must speak 3 or more words (configurable via interruptionWords) for an interruption to trigger. This prevents cutting off the AI on single-word reactions.

  2. Backchannel filtering — Short conversational reactions ("yeah", "ok", "mhm", "sure", "right", "thanks") are classified as backchannels and are ignored — they don't interrupt the AI. The system maintains separate backchannel word lists for English and French.

  3. Filler word grace — If the first word of a transcript is a known filler (e.g., "um", "uh"), the interruption threshold is increased by +1 word. This prevents cutting off on "um, yes" but still triggers on "um, actually I have a question."

  4. Confidence scaling — The STT confidence score is checked against dynamic thresholds:

    • At exactly the word limit: requires 75% confidence
    • Above the word limit: requires 60% confidence (length implies intent)
    • Below the word limit: requires 85% confidence
  5. Echo cancellation — A fuzzy scoring algorithm (Jaccard similarity + subsequence matching) compares the caller's transcript against the AI's recent words. If the score exceeds 0.55 and the transcript is longer than 5 characters, it's classified as echo/feedback and suppressed.

When a valid barge-in is detected, the pipeline:

  • Immediately aborts the current LLM generation
  • Stops all queued TTS synthesis
  • Clears the audio buffer
  • Marks the AI as no longer speaking

Silence & Hangup Timers

Timer Default Behavior
Silence Timer 15 seconds After 15s of silence (no user speech, AI not speaking), the AI prompts: "Are you still there?" This fires once per silence period.
Hangup Timer 30 seconds Checked every 5 seconds. If the user hasn't spoken for 30s and the AI isn't speaking, the call is automatically hung up.
Max Duration Timer 40 minutes Hard cutoff. When reached, the call is severed regardless of state. Users can configure a lower limit (1–30 minutes).

Patience Levels

The patience level controls how long the AI waits after the user stops speaking before it begins generating a response. This handles cases where the user is thinking or pausing mid-sentence.

Level Delay Best For
low 600ms Fast-paced conversations (default)
medium 2,600ms Users who pause while thinking
high 4,600ms Complex questions where callers take time to formulate

During the patience delay, RAG knowledge base search runs in parallel so it's ready when the LLM starts.

Bilingual Mode

When an agent has a secondary language configured, the pipeline activates bilingual mode:

  1. The system prompt is augmented with an operational directive telling the AI to start in the primary language but seamlessly switch if the caller speaks the secondary language.
  2. The AI is instructed to prefix every response with a 2-letter language tag: [EN] for English, [FR] for French, etc.
  3. The pipeline parses these tags and routes audio to the correct TTS voice — using primaryVoiceId for the primary language and secondaryVoiceId for the secondary.
  4. STT is set to Deepgram Flux (multilingual) or Gladia for automatic language detection.
  5. The backchannel word list is extended with language-specific words (e.g., French: "oui", "d'accord", "merci", "parfait").

TTS Processing

Before sending text to the TTS engine, the pipeline performs critical cleanup:

  • Strips markdown formatting: **bold**, *italic*, numbered lists (1. ), bullet dashes, headings (###), and inline code backticks
  • Strips bracketed narrations like [Checking availability...] but preserves bilingual language tags like [EN]
  • Collapses multiple spaces left by stripping

Knowledge Base Integration (RAG)

During a call, when the user says something substantive (≥2 meaningful words after removing conversational filler), the pipeline:

  1. Cleans the query — Strips conversational filler ("hi, can you tell me about...") while protecting compound phrases ("right now", "how much").
  2. Searches the vector database — Queries pgvector for the top 3 chunks above the similarity threshold (default: 0.25).
  3. Stores in a sliding buffer — Results are kept in a FIFO buffer of size 3, so follow-up questions can reference previous context.
  4. Injects into the LLM — When runLLM() fires, the KB buffer contents are appended to the system prompt as a [KNOWLEDGE BASE] block.
  5. Circuit breaker — If the KB search takes longer than 1,500ms, it's abandoned and the LLM proceeds without context to prevent dead air.

2.4 Realtime Models (Speech-to-Speech)

In addition to the modular pipeline, Wirevox supports Realtime Models — a fundamentally different architecture where a single AI model handles speech input and speech output directly, without separate STT/LLM/TTS stages.

┌──────────┐          ┌───────────────────┐          ┌──────────┐
│  Phone   │─────────▶│  Realtime Model   │─────────▶│  Phone   │
│  Audio   │          │  (Speech-in,      │          │  Audio   │
│ (Telnyx) │◀─────────│   Speech-out)     │◀─────────│ (Telnyx) │
└──────────┘          └───────────────────┘          └──────────┘

Supported Realtime Models

Model Provider Audio Format Credit Cost/min
GPT Realtime OpenAI PCM 24kHz 20 (platform) + 20 (model) = 24 total
Gemini Live Google PCM 16kHz 10 (platform) + 10 (model) = 14 total

How Realtime Mode Works

  1. The RealtimePipeline opens a persistent WebSocket directly to the model provider (OpenAI or Google).
  2. Inbound phone audio (μ-law 8kHz from Telnyx) is converted to PCM and resampled to the model's native rate (24kHz for OpenAI, 16kHz for Gemini).
  3. Audio is streamed continuously to the model. The model performs its own voice activity detection (VAD) — no separate STT is needed.
  4. The model generates speech audio responses directly, which are resampled back to 8kHz μ-law and streamed to the caller.

Key Differences vs. Modular Pipeline

Aspect Modular Pipeline Realtime Pipeline
Architecture STT → LLM → TTS (3 services) Single model (1 service)
Latency ~400–600ms (sum of 3 hops) ~200–300ms (single hop)
Voice selection 50+ voices across 3 TTS providers Limited to model's built-in voices
VAD / Barge-in Custom multi-layer barge-in logic Model's native VAD (server-side)
Tool calling Full support Full support
Knowledge Base Sliding buffer with query cleaning Tool-based (search_knowledge_base function)
Cost Variable (4 + STT + LLM + TTS credits) Fixed (platform + model fee)
Bilingual Full support with voice switching Limited (model-dependent)

Realtime Pipeline — System Prompt

The realtime pipeline appends a critical directive to the system prompt:

"You are running in a realtime voice environment. If you need to use a tool to look up information, you MUST immediately say a short filler phrase (like 'Let me check that for you' or 'One moment please') BEFORE calling the tool, so the user is not met with dead air."

This is necessary because in realtime mode, there's no separate TTS pipeline to generate filler — the model must produce it natively.

OpenAI Realtime Configuration

The OpenAI adapter connects to wss://api.openai.com/v1/realtime and configures:

  • Input transcription: Whisper-1 model
  • Turn detection: Server-side VAD with threshold 0.5, prefix padding 300ms, silence duration 800ms
  • Output modality: Audio only
  • Tool choice: Auto

Gemini Live Configuration

The Gemini adapter connects to wss://generativelanguage.googleapis.com/ws/... using the gemini-2.5-flash-native-audio-latest model with:

  • Response modality: Audio
  • Voice: Aoede (prebuilt)
  • Full tool support via functionDeclarations

2.5 Credits System

Credits are the universal currency of the Wirevox platform. Every billable event — call minutes, SMS messages, function executions — is measured in credits. Credits are abstract units; their real-world dollar value varies by subscription plan.

Credit Pricing Per Plan

Plan Cost Per Credit Example: 10 credits
PAYG (Pay-As-You-Go) $0.020 $0.20
Starter $0.015 $0.15
Growth $0.0125 $0.125
Agency $0.008 $0.08

Two Credit Pools

Every workspace has two independent credit pools:

Pool Source Reset Behavior
Included Credits Monthly subscription Reset to plan allowance every billing cycle
Purchased Credits One-time top-ups via Stripe Never expire. Never reset. Survive plan changes and cancellations.

Deduction Waterfall

When credits are consumed, they are deducted in this strict order:

1. Included Credits  ──▶  Deduct first
2. Purchased Credits ──▶  Deduct only if included credits are exhausted
3. Both exhausted    ──▶  Call continues, billing logs a warning

Special Case: Unlimited Tier

If included_credits = -1, the workspace is on an unlimited plan. The billing RPC skips all balance checks and returns remaining: -1.

Cost-Per-Minute Formula

Every agent has a different cost per minute based on its AI stack:

Cost Per Minute = Platform Base + STT + LLM + TTS

For realtime pipeline agents, the formula simplifies to:

Cost Per Minute = Platform Base + Realtime Model Fee
Example Calculations
Agent Stack Platform STT LLM TTS Total
Deepgram STT + GPT-4o Mini + Deepgram TTS 4 2 1 2 9 cr/min
Deepgram STT + GPT-4o + ElevenLabs TTS 4 2 3 10 19 cr/min
Gladia STT + Claude 4.6 Sonnet + ElevenLabs TTS 4 3 5 10 22 cr/min
OpenAI Realtime (speech-to-speech) 4 20 24 cr/min
Gemini Live (speech-to-speech) 4 10 14 cr/min

Per-Second Carryover Billing

Wirevox bills per-minute but tracks duration per-second with agent-level carryover. This prevents charging a full minute for a 5-second call.

How it works:

total_seconds  = agent.carryover_seconds + new_call_duration
minutes_to_bill = floor(total_seconds / 60)
new_carryover  = total_seconds % 60

Walk-through (agent at 9 cr/min):

Call Duration Carryover Before Total Billed Minutes Credits Carryover After
Call 1 13s 0 13 0 0 13
Call 2 15s 13 28 0 0 28
Call 3 28s 28 56 0 0 56
Call 4 22s 56 78 1 9 18

Note: Carryover seconds are not rate-stamped. If you change an agent's AI model between calls, the leftover seconds are billed at the new rate. Maximum discrepancy: 59 seconds × rate delta.

Event Costs (Flat Fees)

These are charged per-use, not per-minute:

Event Credits Description
send_sms_confirmation 2 Sending an SMS during or after a call
transfer_call 20 Transferring the caller to a human via PSTN
Inbound SMS 2 Receiving an SMS on a Wirevox number
Outbound SMS 2 Sending an SMS from a Wirevox number
All other tools 0 Free (booking, webhooks, end call, etc.)

Dynamic Pricing

All credit costs are loaded from the Supabase credit_costs table with a 5-minute cache. A Supabase Realtime channel (cost-registry-changes) listens for any changes to the table and instantly invalidates the cache — so pricing updates propagate to all running servers without a deploy.

If the database is unreachable, hardcoded fallback defaults are used.


2.6 Multi-Agent Architecture

Wirevox supports architectures where multiple specialized agents work together to handle different types of caller needs.

Agent-to-Agent Transfer (Cold Transfer via PSTN)

Currently, Wirevox supports Cold Transfers over PSTN, which is the industry standard:

  1. The caller asks to speak to a specialist (e.g., "Can I talk to your tech support?").
  2. The AI agent triggers the transfer_call function.
  3. The AI says goodbye and disconnects its audio streams (STT, LLM, TTS all shut down instantly — $0 AI cost after transfer).
  4. The platform dials the target phone number via Telnyx, bridging the original caller to the new destination.
  5. Cost: 20 credits for the transfer + telephony charges for the outbound leg.

Transfer Types (Current & Future)

Transfer Type Status Description AI Cost After Transfer
Cold Transfer (PSTN) ✅ Supported AI blindly dials a 10-digit phone number $0 — AI shuts down
Warm Transfer (PSTN) ❌ Not yet AI puts caller on hold, briefs the human, then bridges High — AI stays active until drop-off
Cold Transfer (SIP) ❌ Not yet AI transfers over internet (VoIP) $0 — AI shuts down
Warm Transfer (SIP) ❌ Not yet AI briefs human over SIP before bridging Low telephony, high AI
Stateful AI Handoff ❌ Not yet System silently swaps the system prompt from "Receptionist" to "Tech Support" mid-call Same as a normal call — no extra cost

Stateful Handoff (The Modern Way)

The platform has infrastructure for a Stateful AI Handoff — instead of transferring the phone call, the system swaps the AI's system prompt mid-call. The HANDOFF_CACHE (a 3-minute TTL in-memory cache) stores context state for the transition:

  • The caller stays on the same phone connection
  • No extra telephony legs
  • No extra AI model fees
  • The new "agent" inherits the full conversation history

This is the most cost-effective multi-agent approach and avoids the "triple telephony charge" problem of AI-to-AI PSTN transfers.

Designing a Multi-Agent System

A typical multi-agent setup might look like:

                    ┌─────────────────────┐
    Inbound Call ──▶│  Receptionist Agent  │
                    │  (General Intake)    │
                    └────────┬────────────┘
                             │
              ┌──────────────┼──────────────┐
              ▼              ▼              ▼
     ┌────────────┐  ┌────────────┐  ┌────────────┐
     │   Sales    │  │  Support   │  │  Scheduling │
     │   Agent    │  │   Agent    │  │   Agent     │
     └────────────┘  └────────────┘  └────────────┘

Each agent has its own:

  • System prompt tailored to its specialization
  • Functions — Sales has CRM webhooks, Support has ticketing, Scheduling has calendar integrations
  • Knowledge base — Each can reference different document sets
  • Phone number — Or share routing via the receptionist's transfer function

Concurrent Call Limits

Each subscription plan includes a base number of concurrent call slots — the maximum number of simultaneous phone calls the workspace can handle:

max_concurrent = plan.base_concurrent_calls + workspace.extra_concurrent_calls

Additional slots can be purchased as Stripe add-ons:

Plan Add-On Price
Starter ~$15/mo per extra slot
Growth ~$13/mo per extra slot
Agency ~$5/mo per extra slot

When a plan upgrade occurs, existing concurrent add-on subscriptions are automatically migrated to the new tier's pricing.


Previous: Section 1 — Getting Started Next: Section 3 — Agent Configuration


Section 3 — Agent Configuration

A comprehensive reference for every setting available when configuring a Wirevox AI agent, from system prompts to voice selection, model tuning, and publishing.


3.1 System Prompt

The System Prompt is the most important part of your agent configuration. It defines who your agent is, how it behaves, what it knows, and what it can do. Think of it as writing detailed instructions for a new hire — the more specific and well-structured your prompt, the better your agent performs.

The Prompt Editor

The agent settings page features a full-screen Prompt Editor — a syntax-highlighted textarea with two special features:

  1. Variable highlighting — Any {{variable_name}} in your prompt is rendered in blue with a blue background. These are template variables that get replaced with real values at call time.
  2. Function highlighting — Any [[Function: action_name]] reference is rendered in green with a green background. These tell the AI when and how to use attached tools.

Insert Toolbar

Above the editor, an Insert toolbar provides quick insertion of:

  • {{variable}} — Opens a dropdown of your defined custom variables. Click one to insert it at the cursor position.
  • [[Function: action]] — Opens a dropdown of all functions attached to the agent. Click one to insert a function reference.

Autocomplete

As you type [[ in the prompt editor, an autocomplete popup appears showing matching function names. Use arrow keys to navigate and Enter to select, or Escape to dismiss. This works like IDE autocompletion — it filters as you type after the [[Function: prefix.

Prompt Warnings

The editor performs live validation and shows warnings for:

Warning Type Condition Meaning
Undefined variable {{name}} used in prompt but no variable named name is defined The placeholder won't be replaced — it'll appear literally in the AI's instructions
Missing function [[Function: tool]] referenced but tool is not attached to this agent At call time, the AI will receive a critical override telling it this tool is unavailable
Disconnected integration Function references a tool from a disconnected integration (e.g., Google Calendar OAuth expired) The tool exists but can't execute — the AI will apologize and offer alternatives

Warnings appear as an amber badge in the editor toolbar. Click it to see details.

Writing Effective Prompts

Best practices for phone AI prompts:

  1. Write conversationally — The AI speaks on a phone call. Write instructions the way you'd brief a human receptionist, not the way you'd write an essay.
  2. Define the persona — Give the agent a name, a tone (friendly, professional, energetic), and a role ("You are Sarah, the front desk receptionist at Maple Dental").
  3. Set explicit boundaries — What the agent should NOT do is as important as what it should do. Example: "You must NEVER provide medical advice. Always direct health questions to the dentist."
  4. Structure with sections — Use clear headings or labeled sections: [IDENTITY], [BUSINESS HOURS], [BOOKING RULES], [ESCALATION POLICY].
  5. Reference functions explicitly — Tell the AI when to use tools: "When a caller wants to book an appointment, use [[Function: check_availability]] to find open slots."
  6. Use variables for data — Don't hardcode business hours, addresses, or policies into the prompt. Use {{business_hours}}, {{address}}, etc. so you can change them without editing the prompt.

How Variables Are Compiled

At call time, the compileSystemPrompt function processes the raw prompt in this order:

1. Custom Variables:    {{business_hours}}  →  "Mon-Fri 9am-5pm"
2. Legacy Variables:    {{business_name}}   →  "SunPath Energy"
3. Real-Time Context:   Prepend date, time, caller phone number
4. Knowledge Base:      Append KB instructions (if linked)
5. Tool Overrides:      Replace [[Function: X]] with error message (if X is disconnected)

Variable Naming Rules

Variable names must contain only:

  • Letters (a–z, A–Z)
  • Numbers (0–9)
  • Underscores (_)

Any other characters are stripped when you type a variable name in the Variables panel.


3.2 AI Model Selection

The AI model is the "brain" of your agent — it processes the caller's words and generates intelligent responses. Wirevox supports models from three providers.

Pipeline Mode Toggle

Before selecting a model, choose your Pipeline Mode using the toggle at the top of the configuration sidebar:

Mode Description
Modular (default) Separate STT → LLM → TTS pipeline. Maximum voice selection, full bilingual support, fine-grained control.
Realtime Single speech-to-speech model. Lower latency, limited voice options, simpler architecture.

The model dropdown changes based on your selected pipeline mode.

Modular Mode — LLM Options

Model Provider Credits/min Best For
GPT-4o Mini OpenAI 1 Default choice. Fast, cost-effective, great for most use cases.
GPT-4o OpenAI 3 Better reasoning for complex conversations.
GPT-4.1 OpenAI 3 Latest GPT-4 series with improved instruction following.
GPT-4.1 Mini OpenAI 1 Lighter variant of GPT-4.1.
GPT-4.1 Nano OpenAI 1 Ultra-lightweight, fastest response times.
GPT-5 Mini OpenAI 2 Next-gen mini model with enhanced capabilities.
GPT-5 Nano OpenAI 1 Ultra-fast next-gen nano model.
GPT-5.1 OpenAI 4 Advanced reasoning and nuanced conversations.
GPT-5.2 OpenAI 4 Improved GPT-5.1 with better context handling.
GPT-5.4 OpenAI 5 Flagship model. Best intelligence, highest cost.
GPT-5.4 Mini OpenAI 2 Balanced next-gen model.
GPT-5.4 Nano OpenAI 1 Cost-effective next-gen option.
Claude 3.5 Sonnet Anthropic 3 Coming Soon. Strong at nuanced, empathetic conversations.
Claude 4.5 Sonnet Anthropic 4 Coming Soon. Enhanced reasoning and tool use.
Claude 4.6 Sonnet Anthropic 5 Coming Soon. Anthropic's latest flagship.
Gemini 1.5 Flash Google 1 Ultra-fast, great for simple FAQ agents.
Gemini 2.0 Flash Google 1 Improved Flash with better tool calling.
Gemini 2.5 Flash Google 1 Latest Flash with enhanced multilingual support.

Recommendation: Start with GPT-4o Mini (1 credit/min). It handles 90% of use cases well. Upgrade to GPT-4o or GPT-5.x only if your agent needs complex multi-step reasoning or very nuanced conversations.

Realtime Mode — Model Options

Model Provider Credits/min Description
OpenAI Realtime OpenAI 20 Speech-to-speech via GPT Realtime API. Uses Whisper for input transcription, server-side VAD.
Gemini Live Google 10 Speech-to-speech via Gemini's native audio model (gemini-2.5-flash-native-audio-latest).

When you switch between realtime models, the voice automatically adjusts — OpenAI Realtime defaults to "Alloy", Gemini Live defaults to "Puck".

Model Switching Behavior

When you change pipeline modes:

  • Modular → Realtime: The model switches to the last-used realtime model (or gpt-4o-realtime-preview default). The voice switches to the last-used realtime voice (or alloy).
  • Realtime → Modular: The model restores to the last-used modular model (or gpt-4o-mini). The voice restores to the last-used modular voice (or aura-asteria-en).

Each mode maintains its own independent model + voice state, so switching back and forth doesn't lose your configuration.


3.3 Voice Selection

The voice determines how your agent sounds to callers. Wirevox provides a comprehensive voice catalog with voices from three providers (modular mode) and two realtime providers.

Modular Mode Voices

Deepgram Aura-2 (2 credits/min)

The default TTS provider. Low latency, natural-sounding, cost-effective.

English Voices:

Voice ID Name Gender
aura-asteria-en Asteria ⭐ Female
aura-orion-en Orion Male
aura-luna-en Luna Female
aura-perseus-en Perseus Male
aura-angus-en Angus Male
aura-2-amalthea-en Amalthea Female
aura-2-theia-en Theia Female
aura-2-vesta-en Vesta Female
aura-2-andromeda-en Andromeda Female
aura-2-zeus-en Zeus Male
aura-2-saturn-en Saturn Male
aura-2-phoebe-en Phoebe Female
aura-2-pandora-en Pandora Female
aura-2-orpheus-en Orpheus Male
aura-2-iris-en Iris Female
aura-2-harmonia-en Harmonia Female
aura-2-cordelia-en Cordelia Female
aura-2-cora-en Cora Female
aura-2-arcas-en Arcas Male

French Voices:

Voice ID Name Gender
aura-2-agathe-fr Agathe Female
aura-2-hector-fr Hector Male
OpenAI TTS (2 credits/min)

High-quality voices with consistent tone. Same cost as Deepgram.

Voice ID Name Gender
shimmer Shimmer Female
alloy Alloy Neutral
echo Echo Male
fable Fable Neutral
onyx Onyx Male
nova Nova Female
marin Marin Female
cedar Cedar Male
ElevenLabs (10 credits/min)

Premium ultra-realistic voices. 5× the cost of Deepgram/OpenAI but significantly more natural.

Voice ID Name Gender
eleven-21m00Tcm4TlvDq8ikWAM Rachel Female
eleven-pNInz6obpgDQGcFmaJgB Adam Male
eleven-EXAVITQu4vr4xnSDxMaL Bella Female
eleven-MF3mGyEYCl7XYWbV9V6O Elli Female
eleven-ErXwobaYiN019PkySvjV Antoni Male
eleven-aMSt68OGf4xUZAnLpTU8 Juniper Female
eleven-hFskf6X0TFndppvQxiEF Finch Male
eleven-g6xIsTj2HwM6VR4iXFCw Jessica Female
eleven-TxGEqnHWrfWFTfGW9XjX Josh Male
eleven-VR6AewLTigWG4xSOukaG Arnold Male
eleven-nPczCjzI2devNBz1zQrb Brian Male

Realtime Mode Voices

OpenAI Realtime Voices
Voice ID Name Gender
alloy Alloy ⭐ Neutral
echo Echo Male
shimmer Shimmer Female
verse Verse Neutral
ballad Ballad Male
coral Coral Female
sage Sage Female
ash Ash Neutral
marin Marin Female
cedar Cedar Male
Gemini Live Voices
Voice ID Name Gender
Puck Puck ⭐ Neutral
Charon Charon Male
Kore Kore Female
Fenrir Fenrir Male
Aoede Aoede Female

Voice Preview

Every voice in the catalog can be previewed before selection. Click the play button next to any voice to hear a sample audio clip loaded from /assets/voices/{voiceId}.mp3. Only one preview plays at a time — clicking a new voice stops the previous one.

Auto-Detection of TTS Provider

The pipeline auto-detects which TTS provider to use based on the voice ID:

Voice ID Pattern Detected Provider
Starts with eleven- or known ElevenLabs names ElevenLabs
One of alloy, echo, fable, onyx, nova, shimmer OpenAI
Everything else (including aura-*) Deepgram

You never need to manually set the TTS provider — selecting a voice is sufficient.


3.4 STT Configuration

Speech-to-Text (STT) converts the caller's spoken audio into text that the LLM can process. This setting only applies in Modular pipeline mode (Realtime models handle their own speech recognition).

Available STT Providers

Provider Setting Value Credits/min Description
Deepgram Nova-3 nova-3 2 Monolingual. Lowest latency, best accuracy for single-language agents.
Deepgram Nova-3 Flux flux 2 Multilingual. Automatic language detection for bilingual agents.
Gladia gladia 3 Alternative provider with different accuracy characteristics.
Auto auto 2 Smart default — selects flux for bilingual agents, nova-3 for monolingual.

Recommendation: Leave this on auto unless you have a specific reason to change it. Auto makes the right choice for your language configuration.

Patience Level (Endpointing)

The Patience Level controls how long the STT waits after the user stops speaking before treating the utterance as complete and sending it to the LLM. This is critical for natural conversation flow.

Level Delay Description
Low (default) 600ms Snappy responses. Best for simple Q&A or fast-paced conversations.
Medium 2,600ms Gives the user time to pause mid-thought. Good for complex topics.
High 4,600ms Very patient. Best for callers who think out loud or give long answers.

During the patience delay, the system starts the knowledge base search in parallel so context is ready the moment the LLM fires.


3.5 Language & Bilingual Mode

Supported Languages

Wirevox currently supports voices in these languages:

Code Language Native Voices Available
en English 19 Deepgram + 8 OpenAI + 11 ElevenLabs
fr French 2 Deepgram (Agathe, Hector)

Additional languages (Spanish, German, Italian, Portuguese, Dutch) are supported at the LLM level but do not yet have dedicated Deepgram voices. You can use OpenAI or ElevenLabs voices for these languages.

Enabling Bilingual Mode

  1. Set your Primary Language (e.g., English).
  2. Set your Secondary Language (e.g., French).
  3. The STT automatically switches to Flux (multilingual) mode.
  4. Choose a Primary Voice for the primary language and a Secondary Voice for the secondary language.

How Bilingual Mode Works at Runtime

When bilingual mode is active, the pipeline injects an Operational Directive into the system prompt:

[OPERATIONAL DIRECTIVE]
Primary Language: English
Secondary Language: French
Rule: You MUST initiate the conversation and formulate all internal thoughts
natively in the Primary Language. However, you are fully bilingual — if the
user speaks the Secondary Language, you must seamlessly transition and
continue the conversation in that language.

CRITICAL BILINGUAL RULE: You MUST prefix every single response with a 2-letter
language tag indicating the language you are speaking.
Use [EN] for English and [FR] for French.
Example: "[FR] Bonjour, comment puis-je vous aider?"

The TTS engine then parses these tags and routes audio to the correct voice.

Backchannel Words by Language

The barge-in system uses language-specific backchannel word lists:

English backchannels: yes, yeah, yep, ok, okay, right, sure, gotcha, alright, perfect, great, good, thanks, etc.

French backchannels (added when French is primary or secondary): oui, ouais, d'accord, entendu, exactement, absolument, parfait, super, bien, bon, voilà, merci, non, etc.

French fillers: euh, hein, bah, ben, bof, mouais, hmm, ah, oh.


3.6 Call Settings

The Call Settings section in the configuration sidebar controls how the agent handles the mechanics of phone calls.

Ring Duration

Setting Range Default Description
Ring Duration 0–10 seconds 0s Simulates a phone ringing before the AI picks up. Set to 0 for instant pickup. Useful for making the experience feel more natural to callers who expect a few rings.

Greeting Message

The greeting message is the first thing your agent says when it picks up the call.

  • Custom greeting: You define exactly what the agent says. Example: "Thank you for calling Maple Dental. How can I help you today?"
  • No greeting (default): The AI generates a natural opening based on its system prompt.

Custom greetings bypass the LLM entirely — they're sent directly to TTS, saving latency on the first response.

Barge-In Sensitivity (Interruption Words)

Controls how many words a caller must say to trigger an interruption when the AI is speaking.

Value Behavior
1 word Very sensitive — even "yes" interrupts the AI
2 words Moderate — short phrases interrupt
3 words (default) Balanced — natural conversation fragments trigger, single-word reactions don't
5+ words Low sensitivity — only substantial speech interrupts

Filler Words

A customizable list of words that the pipeline treats as filler — they don't trigger the AI to respond and don't count as interruptions when the AI is speaking.

Default fillers:

yeah, uh, um, mhmm, ah, hm, like, right, mhm, uhhuh, uh-huh,
hmm, oh, umm, uhh, huh, mmm, aha, aah, ooh, mm, ugh, eh, er

You can add or remove fillers from the configuration sidebar. New fillers are entered one at a time and added to the list.

Silence Timeout

Setting Range Default
Silence Timeout 5,000–60,000ms 15,000ms (15s)

How long the agent waits in silence before prompting "Are you still there?" This fires once per silence period. After the prompt, the hangup timer continues independently.

Hangup Timeout

Setting Range Default
Hangup Timeout 10,000–120,000ms 30,000ms (30s)

Total seconds of silence (no user speech, AI not speaking) before the call is automatically disconnected. This timer is checked every 5 seconds and resets whenever the user speaks or the AI finishes speaking.

Max Call Duration

Setting Range Default System Max
Max Call Duration 1–30 minutes Not set (uses system default) 40 minutes

A hard time limit for calls. When reached, the call is severed regardless of state. This protects against runaway calls and unexpected credit consumption.

RAG Threshold

Setting Range Default
RAG Threshold 0.0–1.0 0.25

The minimum vector similarity score required for a knowledge base chunk to be included in the AI's context. Lower values return more (potentially less relevant) results. Higher values return only highly relevant matches.


3.7 Post-Call Processing

After every call ends, Wirevox performs several automatic post-processing steps.

Call Summary Generation

The LLM generates a structured summary of the call, including:

  • Purpose — Why the caller called
  • Key topics discussed — Main points of the conversation
  • Actions taken — Any tools that were used (bookings, SMS, transfers)
  • Outcome — How the call concluded (resolved, transferred, hung up)
  • Follow-up needed — Whether human action is required

Lead Extraction

From every call, the system extracts structured lead data:

  • Name — Caller's full name (if provided)
  • Phone number — The caller's phone number (from caller ID or stated)
  • Email — If the caller provided one
  • Sentiment — Positive, neutral, or negative
  • Custom fields — Any additional data captured by functions

Webhook Delivery

If post-call webhooks are configured (via custom functions), the system fires HTTP POST requests with the call data to your specified endpoints. This enables:

  • CRM synchronization
  • Slack/Teams notifications
  • Zapier/Make.com automations
  • Custom analytics pipelines

3.8 Publishing & Versioning

Wirevox uses a draft/published model for agent configuration. This means you can freely experiment with changes without affecting live callers.

How It Works

┌─────────┐                ┌─────────────┐                ┌──────────┐
│  Draft   │───Publish───▶ │  Published   │◀──Incoming───│  Callers  │
│ (Your    │               │  (What       │    Calls      │          │
│  edits)  │◀──Discard────│  callers     │               │          │
│          │               │  experience) │               │          │
└─────────┘                └─────────────┘                └──────────┘

The Draft

Every change you make in the agent settings page is saved to the draft version. Changes are auto-saved with a 2-second debounce — you'll see a save status indicator in the header:

Status Indicator
saving "Saving..." text appears
saved "Saved at 3:45 PM" confirmation
error Error state (usually network issues)

If you navigate away or close the browser while changes are pending, a beforeunload handler fires a final save via keepalive: true to ensure nothing is lost.

Publishing

When your draft is ready, click "Publish" in the header. This opens a Publish Modal where you can:

  1. Review what changed since the last publish
  2. Add an optional version note describing your changes
  3. Confirm the publish

Publishing creates a new version snapshot and makes the draft the new live configuration. All incoming calls immediately use the published version.

The header shows an amber dot with "Unpublished changes" when your draft differs from the published version.

Version History

Click "History" in the agent settings menu to open the Version History Modal. This shows:

  • A chronological list of all published versions
  • Each version's note (if provided) and timestamp
  • The ability to restore any previous version

Restoring a version replaces your current draft with the selected historical version.

Discarding Changes

If you want to abandon your draft and revert to the last published version, click "Discard Changes" in the more menu. This restores the published version as your current draft.

Agent Status

Each agent has a status that controls whether it actively answers calls:

Status Behavior
Active Agent answers incoming calls normally
Paused Agent exists but doesn't answer calls. Calls to its number will ring unanswered.

You can toggle status from the agent settings header. An agent must have a phone number assigned to be activated — attempting to activate without a number shows an error: "Assign a number to the agent to start it."

Other Agent Actions

Available from the more menu (⋮) in the agent header:

Action Description
Duplicate Agent Creates a copy of the current draft as a new agent with all settings preserved.
Delete Agent Permanently removes the agent, its draft, and all published versions. This action cannot be undone.

Cost Display

The agent header shows a live cost estimate in credits per minute. Hovering over it reveals a tooltip with the full cost breakdown:

Cost Breakdown
──────────────────────
  Platform        4
  STT (Listening) 2
  LLM (Thinking)  1
  TTS (Speaking)  2
──────────────────────
  Total          9 cr/min
  ≈ $0.135 / min

If the draft cost differs from the published cost (e.g., you changed the model), the tooltip shows both with a "(Draft)" label. The dollar estimate uses your plan's overage rate for the conversion.


3.9 Configuration Sidebar Sections

The right-side Configuration Sidebar (340px wide, collapsible) organizes all non-prompt settings into accordion sections:

Pipeline Mode (Always Visible)

A toggle between Modular and Realtime at the top of the sidebar. Changes the available options in all sections below.

General Section

  • Timezone — The agent's local timezone (34 options from Adelaide to UTC). Used for real-time context injection and scheduling functions.
  • Model — LLM selection (see Section 3.2). Changes based on pipeline mode.

Variables Section

  • Define {{name}}value pairs
  • Each variable has a name field (alphanumeric + underscore only) and a multi-line value textarea
  • Add variables with "+ Add Variable" button
  • Remove with the ✕ button
  • Variables are highlighted in blue in the prompt editor

Call Settings Section

  • Ring Duration (slider: 0–10s)
  • Additional call mechanics settings

Voice & Language Section

  • Primary Language selector
  • Secondary Language selector (optional, enables bilingual mode)
  • Voice Selector — visual voice picker with preview buttons
  • STT Provider selector
  • Patience Level selector

Knowledge Section

  • Link/unlink knowledge bases via the Knowledge Selector Modal
  • Shows linked KB name and document count
  • RAG Threshold slider

Functions Section

Functions are managed in a separate Functions Panel (slide-out from the right), organized into three stages:

Stage When It Runs Examples
Pre-call Before the call connects (IVR-style routing) Language detection, caller identification
During call While the AI is in conversation Book appointment, send SMS, transfer call, check availability, create contact
Post-call After the call ends Send email summary, sync to CRM, fire webhook

Each function shows its name, category, and provider (if it's an integration). Functions can be added from the Action Catalog which shows all available built-in actions and custom webhooks.


Previous: Section 2 — Core Concepts Next: Section 4 — Phone Numbers & Telephony


Section 4 — Phone Numbers & Telephony

Everything about acquiring, managing, and routing phone numbers on Wirevox — from searching and purchasing through Telnyx, to the complete inbound call lifecycle, transfers, SMS, and billing.


4.1 Buying Phone Numbers

Phone numbers are the bridge between the real world and your AI agents. Wirevox provisions numbers through Telnyx, a carrier-grade telephony provider.

Prerequisites

  • You must be on a paid plan (Starter, Growth, or Agency). Free-tier users cannot purchase phone numbers.
  • The Phone Numbers page (/w/:id/numbers) shows an "Unlock Telephony Features" screen for free users with a prompt to upgrade.

The Numbers Page

The Numbers page has two tabs:

Tab Description
My Numbers Your purchased number inventory with agent assignments
Buy New Numbers Search and purchase new numbers from Telnyx

Searching for Numbers

On the "Buy New Numbers" tab, you can filter available numbers by:

Filter Options Description
Country US, GB, CA, AU, NZ, DE, FR, IE, NL The country where the number is registered
Type local, toll-free Local numbers have an area code; toll-free numbers (e.g., 1-800) are nationwide
Search Query Area code or contains Smart parsing: 3-digit input = area code search, 4+ digits = "contains" search
Max Price Any, $2, $5, $10 Filter by monthly Telnyx cost
Sort Price ascending/descending Order results by monthly fee

The search auto-fires when you switch to the search tab or change country/type filters. The query input has a 500ms debounce. Results are paginated (20 per page) from up to 200 results returned by Telnyx.

Purchasing Flow

  1. Click "Buy" next to a number in the search results.
  2. A confirmation modal appears showing:
    • The phone number (formatted as +1 (XXX) XXX-XXXX)
    • Number type and region
    • Whether it's free (within your plan's free number allowance) or $5/mo (extra number)
  3. Click "Confirm Purchase".
  4. The backend:
    • Verifies your subscription and free number quota
    • If it's a paid number: adds a $5/mo line item to your Stripe subscription (proration_behavior: 'always_invoice' — charged immediately)
    • Calls telephony.buyNumber() to provision the number on Telnyx
    • Saves the number to the database with status active
  5. You're redirected to the "My Numbers" tab where the new number appears.

Free vs. Paid Numbers

Each plan includes a set number of free phone numbers:

Plan Free Numbers Included
Free 0
PAYG 0
Starter 1
Growth 3
Agency 5

Numbers beyond the free allowance cost $5/month each, added as a subscription line item on Stripe.

Free Number Abuse Prevention

A safeguard prevents users from infinitely cycling free numbers:

If you release a free number (moving it to grace period) and then try to buy a new one, the new number is treated as paid ($5/mo) even if you're under the free limit. You must either reclaim the graced number or pay for a new one.

This prevents the exploit of releasing a free number, buying a new free one, and repeating to effectively get unlimited numbers.


4.2 Assigning Numbers to Agents

Once purchased, numbers need to be assigned to agents to start routing calls.

Assignment Rules

  1. One-to-one mapping — Each number can be assigned to exactly one agent, and each agent can have at most one number.
  2. Published agents only — You cannot assign a number to an agent that has never been published. The system validates publishedVersionId exists.
  3. Same workspace — The agent must belong to the same workspace as the number.

Assignment Flow

  1. On the "My Numbers" tab, each number shows an agent dropdown.
  2. Select an agent from the list. Only unpublished or already-assigned agents are excluded.
  3. If the number is currently assigned to a different active agent, a confirmation modal warns you:

    "This number is currently routing calls to [Agent Name]. Reassigning will stop that agent."

  4. Confirming the reassignment:
    • Pauses the previous agent (sets status to paused)
    • Assigns the number to the new agent

Unassigning

Select "Unassigned" in the agent dropdown to remove the agent mapping. The number remains in your inventory but doesn't route calls.


4.3 Inbound Call Flow

When a caller dials your Wirevox number, a precise sequence of events orchestrates the entire call:

Step-by-Step Lifecycle

┌──────────────┐        ┌──────────┐        ┌──────────────┐
│   Caller     │──1───▶ │  Telnyx  │──2───▶ │  Wirevox     │
│   dials      │        │  (PSTN)  │        │  Backend     │
│   your #     │        │          │        │              │
└──────────────┘        └──────────┘        └──────┬───────┘
                                                   │
                            3. Answer call + open  │
                               WebSocket stream    │
                                                   ▼
                                            ┌──────────────┐
                                            │  Media       │
                                            │  Stream      │
                                            │  Handler     │
                                            └──────┬───────┘
                                                   │
                              4. Load published    │
                                 agent config      │
                              5. Compile prompt    │
                              6. Assemble tools    │
                              7. Create pipeline   │
                                                   ▼
                                            ┌──────────────┐
                                            │  Pipeline    │
                                            │  (Voice or   │
                                            │   Realtime)  │
                                            └──────┬───────┘
                                                   │
                              8. STT → LLM → TTS  │
                              9. Audio back to     │
                                 caller            │
                                                   ▼
                                            ┌──────────────┐
                                            │  Call Ends   │
                                            │  (hangup /   │
                                            │   timeout)   │
                                            └──────┬───────┘
                                                   │
                             10. Save transcript   │
                             11. Post-call         │
                                 analysis          │
                             12. Fire webhooks     │

Detailed Breakdown

1. Telnyx Webhook — Telnyx receives the call and sends a webhook to the Wirevox backend. The webhook payload contains the caller's number, the dialed number, and a call_control_id for managing the call.

2. Agent Resolution — The backend looks up which agent is assigned to the dialed number.

3. Answer + Media Stream — The backend answers the call via the Telnyx API and opens a WebSocket media stream at:

wss://{host}/media-stream/{agentId}/{callId}

4. Load Published Config — The handler loads the agent's published version (not the draft). This ensures that edits you're making in the dashboard don't affect live calls.

// Real calls MUST run on the published snapshot, not the live draft.
if (agent && agent.publishedVersionId) {
  const pubVer = await db.getAgentVersion(agent.publishedVersionId);
  agent.agentConfig = pubVer.config;
}

5. Compile System Prompt — The raw prompt is processed through compileSystemPrompt() with:

  • Custom variables replaced
  • Real-time context injected (date, time, caller phone)
  • Knowledge base instructions appended
  • Unavailable tool overrides injected

6. Assemble & Filter Tools — All attached functions are assembled into OpenAI tool schemas. Tools whose required integrations are disconnected are removed. A strict prompt-driven enforcement filter ensures only tools referenced in the prompt (via [[Function: name]]) are exposed, plus always-enabled system tools (end_call, search_knowledge, system_infolist).

7. Create Pipeline — Based on pipelineMode, either a VoicePipeline (modular) or RealtimePipeline (realtime) is instantiated.

8. Audio Processing — Telnyx sends μ-law 8kHz audio frames as base64. The pipeline processes them through STT → LLM → TTS.

9. Audio Response — TTS audio is sent back to Telnyx via the WebSocket as base64 media events.

10–12. Call Ends — When the WebSocket closes:

  • The transcript record is saved to the database
  • runPostCallAnalysis() kicks off background processing (summary, lead extraction)
  • Post-call webhooks fire if configured

Barge-In Audio Flushing

When the user interrupts (barge-in), the pipeline sends a clear event to Telnyx:

{ "event": "clear" }

This flushes Telnyx's audio playback buffer. Without this, audio already sent to Telnyx continues playing for 1–3 seconds after interruption, making the AI seem unresponsive.

Welcome Greeting Timing

After the media stream connects and STT is ready, there's a 300ms stabilization delay before the greeting plays. This ensures the bidirectional audio socket is fully established — without it, the first 100–200ms of the greeting gets clipped.

pipeline.welcomeTimer = setTimeout(() => {
  pipeline.triggerWelcome(defaultGreeting);
}, 300);

If the user speaks before the greeting fires, the cancelGeneration() method clears this timer.


4.4 Outbound Calls

🔜 Coming Soon — Outbound calling is under development.

Outbound agents will support:

  • Campaign creation — Define a list of contacts to call
  • CSV upload — Import phone numbers and names in bulk
  • Concurrent dispatch — Call multiple leads simultaneously
  • Answering Machine Detection (AMD) — Detect voicemail and handle appropriately
  • Call dispositions — Track outcomes (answered, no answer, voicemail, busy)

The backend already has partial support for outbound agents in the media stream handler — outbound calls inject the lead's name into the greeting:

"Hi {leadName}, this is {agentName}. Do you have a quick minute?"

4.5 Call Transfers

Call transfers let your AI agent connect the caller to a human when needed — for escalation, specialist routing, or emergency situations.

How Cold Transfer Works

  1. The AI decides to transfer (based on the caller's request or prompt instructions).
  2. The AI calls the transfer_call function with a target phone number.
  3. The tool handler:
    • Returns { disconnect: true } to the pipeline
    • The pipeline immediately shuts down all AI services (STT, LLM, TTS stop — $0 AI cost from this point)
    • The WebSocket closes
  4. Before disconnecting, the tool initiates a Telnyx transfer — bridging the original caller to the target number via PSTN.
  5. The caller hears ringing and connects to the human.

Transfer Cost

Component Credits
Transfer function execution 20 credits (flat)
Telnyx telephony (outbound leg) Standard Telnyx rates
AI cost after transfer $0 — pipeline shut down

Important Constraint

  • Transfer targets must be 10-digit PSTN phone numbers (or E.164 format).
  • The AI cannot transfer to a SIP endpoint, another AI agent (via the platform), or an internal extension — only real phone numbers.

4.6 SMS

Wirevox supports both sending and receiving SMS on your phone numbers.

Sending SMS (During Call)

The send_sms_confirmation function lets your AI agent send a text message to the caller during a conversation. Common use cases:

  • Sending appointment confirmation details
  • Sharing a link or address
  • Confirming a booking reference number

Cost: 2 credits per outbound SMS.

Inbound SMS

When someone texts your Wirevox number:

  • The SMS is received via a Telnyx webhook
  • Cost: 2 credits per inbound SMS

SMS Requirements

  • The phone number must be SMS-enabled (most Telnyx numbers are)
  • The workspace must have sufficient credits

4.7 Number Billing

Free Numbers

Each plan includes free numbers (see 4.1). Free numbers have isFree: true and monthlyCost: 0 in the database.

Extra Numbers ($5/mo)

Numbers beyond the free quota cost $5/month each. This is implemented as a Stripe subscription line item:

  • When you buy an extra number, a quantity: 1 item is added (or quantity incremented) on the STRIPE_PRODUCT_NUMBER_{TIER} product
  • proration_behavior: 'always_invoice' means you're charged immediately for the remainder of the billing period
  • When you release an extra number, the quantity is decremented with proration_behavior: 'none'no refund for the current period, charges simply stop next cycle

Tier-Specific Products

Each plan tier has its own Stripe product for extra numbers:

Env Variable Plan
STRIPE_PRODUCT_NUMBER_STARTER Starter
STRIPE_PRODUCT_NUMBER_GROWTH Growth
STRIPE_PRODUCT_NUMBER_AGENCY Agency

When a user changes plans, the existing extra number subscription items are migrated to the new tier's product.

Number Quota API

The /api/numbers/quota endpoint returns:

{
  "tier": "starter",
  "freeLimit": 1,
  "freeUsed": 1,
  "extraCount": 2,
  "extraMonthlyCost": 10
}

This powers the UI quota display showing "1/1 free numbers used, 2 extra ($10/mo)".


4.8 Grace Period (30-Day Number Reservation)

When you release a number, it doesn't get permanently deleted immediately. Instead, it enters a 30-day grace period.

How Grace Works

Active Number ──Release──▶ Grace Period (30 days) ──Expires──▶ Permanently Released
                               │                                  (Telnyx deletes)
                               │
                            Reclaim ──▶ Active Number (reactivated)

During Grace Period

  • The number is reserved on Telnyx (no one else can buy it)
  • It does not route calls — callers will get no answer
  • It is unassigned from any agent
  • Stripe billing for extra numbers is stopped (no charges during grace)
  • A blue/amber banner appears on the Numbers page showing grace numbers with:
    • Days remaining (calculated as ceil((graceUntil - now) / 86400000))
    • "↩ Reclaim" button — reactivates the number
    • "Release Now" button — permanently releases immediately

Reclaiming

Click "Reclaim" to reactivate a graced number:

  • Requires an active subscription
  • If the number was free and the free quota isn't full, it's reclaimed as free
  • If it's a paid reclaim, a new Stripe line item is added
  • Status changes back to active, graceUntil is cleared

Permanent Release

Click "Release Now" (or wait 30 days for automatic expiry):

  • telephony.releaseNumber() removes the number from Telnyx
  • db.deleteNumber() permanently removes the database record
  • The number becomes available for anyone to purchase on Telnyx

Automatic Cleanup

A cron job (POST /api/numbers/cleanup-grace) runs periodically to find and release expired grace numbers. It's protected by an x-cron-secret header so only the cron scheduler can invoke it.


4.9 Telnyx Configuration

Wirevox uses Telnyx as its telephony provider for all phone operations.

Core Integration Points

Component Telnyx Feature Description
Number Provisioning Number Orders API Searching available numbers and purchasing them
Call Control Call Control API Answering, hanging up, and transferring calls
Media Streaming WebSocket Media Streaming Real-time audio I/O during calls
SMS Messaging API Sending and receiving text messages

WebSocket Media Stream

When a call is answered, Telnyx opens a WebSocket to:

wss://{your-host}/media-stream/{agentId}/{callId}

The WebSocket exchanges JSON messages:

Event Direction Description
connected / start / stream_started Telnyx → Server Stream initialization
media Telnyx → Server Audio frame: { media: { payload: "base64..." } }
media Server → Telnyx Response audio: { event: "media", media: { payload: "base64..." } }
clear Server → Telnyx Flush playback buffer (used during barge-in)
stop Telnyx → Server Stream ended (call disconnected)

Audio Format

Direction Encoding Sample Rate Bit Depth
Telnyx → Server μ-law (mulaw) 8,000 Hz 8-bit
Server → Telnyx μ-law (mulaw) 8,000 Hz 8-bit

For Realtime pipelines, audio is resampled:

  • OpenAI Realtime: 8kHz μ-law → 24kHz PCM (input), 24kHz PCM → 8kHz μ-law (output)
  • Gemini Live: 8kHz μ-law → 16kHz PCM (input), 16kHz PCM → 8kHz μ-law (output)

Local Development (ngrok)

For local development, Telnyx webhooks need a public URL. Use ngrok to tunnel:

ngrok http 3001

Then configure the Telnyx webhook URL to point to your ngrok URL.


Previous: Section 3 — Agent Configuration Next: Section 5 — Knowledge Base (RAG)


Section 5 — Knowledge Base (RAG)

Complete technical reference for Wirevox's Retrieval-Augmented Generation system — from creating knowledge bases and ingesting documents to the embedding pipeline, vector search, and how context is injected into live phone calls.


5.1 Overview

The Knowledge Base gives your AI agent access to real information — your business policies, FAQs, product catalogs, service menus, team bios, troubleshooting guides, and any other factual data it needs to answer caller questions accurately.

Without a knowledge base, the AI only knows what's in its system prompt. With a knowledge base, the AI can dynamically search and retrieve relevant information during a call, then use it to give precise, grounded answers.

Architecture Summary

┌─────────────────────────────────────────────────────────────┐
│                    INGESTION PIPELINE                        │
│                                                             │
│  Source (PDF/URL/Text/DOCX/CSV/XLSX)                         │
│       │                                                      │
│       ▼                                                      │
│  Document Processor  ──extract text──▶  Raw Text             │
│       │                                                      │
│       ▼                                                      │
│  LLM Pre-Processing  ──organize──▶  Clean, Structured Text   │
│       │                                                      │
│       ▼                                                      │
│  Text Chunker  ──split──▶  1500-char overlapping chunks      │
│       │                                                      │
│       ▼                                                      │
│  OpenAI Embeddings  ──vectorize──▶  1536-dim vectors         │
│       │                                                      │
│       ▼                                                      │
│  Supabase (pgvector)  ──store──▶  knowledge_chunks table     │
└─────────────────────────────────────────────────────────────┘

┌─────────────────────────────────────────────────────────────┐
│                    RETRIEVAL (AT CALL TIME)                   │
│                                                             │
│  Caller Speech  ──STT──▶  "Who are your hygienists?"         │
│       │                                                      │
│       ▼                                                      │
│  Query Cleaning  ──strip filler──▶  "hygienists"             │
│       │                                                      │
│       ▼                                                      │
│  OpenAI Embedding  ──vectorize──▶  query vector              │
│       │                                                      │
│       ▼                                                      │
│  pgvector cosine similarity  ──search──▶  top 3 chunks       │
│       │                                                      │
│       ▼                                                      │
│  Inject into system prompt  ──▶  LLM answers with context    │
└─────────────────────────────────────────────────────────────┘

Two Retrieval Modes

Mode Trigger Description
Passive RAG Every user utterance (automatic) The pipeline automatically searches the KB when the user speaks. Results are injected into the system prompt silently.
Active Tool LLM decides to call search_knowledge The AI proactively searches the KB when it decides it needs information. Works like any other function call.

Both modes can operate simultaneously — passive RAG fires automatically, and the active tool is available as a fallback for deeper or more targeted queries.


5.2 Creating a Knowledge Base

UI Structure

The Knowledge page (/w/:id/knowledge) has a sidebar + detail layout:

  • Left sidebar (280px) — Lists all knowledge bases in the workspace. Click "+" to create a new one.
  • Right panel — Shows sources, chunks, and management actions for the selected KB.

Creating a New KB

Click "+" in the sidebar to start creating. Fill in:

Field Required Description
Name A short identifier for the KB (e.g., "Dental Clinic FAQ")
Description What this KB contains (helps you organize multiple KBs)
Instructions Special instructions appended to the system prompt when this KB is linked to an agent (e.g., "Always cite your source when referencing this knowledge base")

Once created, the KB appears in the sidebar and you can start adding sources.

Editing & Deleting

  • Click the more menu (⋮) to rename, edit the description/instructions, or delete the entire KB.
  • Deleting a KB removes all its chunks and unlinks it from any agents using it.

5.3 Adding Sources

Sources are the raw materials that get processed into searchable knowledge. There are four ways to add information:

File Upload

Supported formats: PDF, TXT, DOCX, CSV, XLSX

  • Click "Add Item" → "Upload File" or drag and drop a file directly onto the source list.
  • Files are limited to 10 MB per upload.
  • Total KB size is limited to 50 MB.
  • If a file with the same name already exists, you'll see a duplicate confirmation asking whether to re-process and overwrite.
How Each Format Is Processed
Format Library Processing
PDF pdf-parse Extracts raw text from all pages. Image-based PDFs fail with an error.
TXT Native Read as UTF-8. Title derived from filename.
DOCX mammoth Extracts raw text (formatting stripped).
CSV csv-parse Parses with column headers. Each row becomes Key: Value pairs separated by ---.
XLSX xlsx Reads all sheets. Each sheet prefixed with Sheet: {name}. Rows formatted as Key: Value pairs.

URL Crawling

Two crawl modes available:

Single URL Mode
  1. Click "Add Item" → "Import from URL".
  2. Paste a URL and click "Import".
  3. The system fetches the page, extracts visible text using Cheerio, and ingests it.
Multi-Page Mode
  1. Switch to "Multi-page" mode.
  2. Paste your root URL and click "Scan".
  3. The scanner discovers up to 100 internal links by:
    • Parsing all <a> tags on the page for same-domain URLs
    • Falling back to sitemap.xml if fewer than 2 links are found
  4. Select which pages to import (auto-selects first 20).
  5. Click "Import Selected" — crawling runs in the background.
URL Text Extraction

The crawler uses a priority-based extraction strategy:

1. Try content selectors:  main, article, [role="main"], .content, #content, .post-content, .entry-content
2. Fallback to:           body text (with nav/footer/header stripped)
3. If < 50 chars:          Jina Reader API fallback (https://r.jina.ai/{url})
4. If still < 50 chars:    Error — "page might be heavily JavaScript-dependent"

Elements removed before extraction: script, style, nav, footer, header, aside, iframe, noscript, svg, form, [role="navigation"], [role="banner"], [role="complementary"].

Background Processing

Multi-page crawls are processed sequentially in the background — the API returns immediately and the UI polls every 3 seconds to update source statuses. Each URL goes through:

  1. Insert a placeholder chunk with status: 'processing'
  2. Fetch and extract text
  3. Run through the full ingest pipeline
  4. Delete the placeholder on success, or update to status: 'error' on failure
Duplicate URL Handling

If you try to import a URL that already exists as a source, a confirmation dialog asks whether to re-crawl and overwrite. Re-crawling deletes existing chunks for that URL first, then re-processes.

Manual Text Document

  1. Click "Add Item" → "Write Document".
  2. Enter a name and paste/type your content (minimum 200 characters).
  3. The text is processed through the full ingest pipeline.

Retry Failed Crawls

If a URL crawl fails (e.g., timeout, 403 error), the source shows an error badge with the error message and a "Retry" button to re-attempt.


5.4 The Ingest Pipeline

Every piece of content goes through the same 4-step pipeline before it becomes searchable:

Step 1: LLM Pre-Processing

Before chunking, the raw text is sent through GPT-4o Mini for intelligent cleaning and reorganization:

Input:  Messy scraped web text with nav elements, repeated headers, "Click Here" CTAs
Output: Clean, organized sections with headers like "## Hygienists" and "## Office Hours"

Rules the LLM follows:

  • Group related information together
  • Make each section self-contained
  • Preserve ALL factual details (names, qualifications, dates, prices, phones, addresses)
  • Remove web-specific noise ("click here", "fill out the form below", "call us to book")
  • Convert CTAs into factual information suitable for a voice AI
  • Use clear section headers
  • Never add information not in the original

Skip condition: Texts shorter than 500 characters skip LLM pre-processing (not worth the API cost).

Fallback: If the LLM call fails, the raw text is used as-is.

Step 2: Text Chunking

The processed text is split into overlapping chunks using a smart boundary-aware algorithm:

Parameter Value Description
Chunk Size 1,500 characters Target size for each chunk
Chunk Overlap 100 characters Overlap between consecutive chunks for context continuity

Boundary detection priority (for chunk end):

  1. Paragraph break (\n\n) — within the last 20% of the chunk
  2. Sentence end (. / ? / ! / .\n) — searched backward
  3. Forward sentence completion — if backward search fails, look forward to finish the current sentence
  4. Space — final fallback

Overlap start snapping: The start of each overlapping chunk snaps to the nearest sentence or paragraph boundary to avoid starting mid-sentence.

If the entire text is shorter than 1,500 characters, it becomes a single chunk.

Step 3: Embedding

Each chunk is converted into a 1,536-dimensional vector using OpenAI's text-embedding-3-small model:

  • Chunks are batched in groups of 100 for API efficiency
  • Each batch call to OpenAI returns an array of embedding vectors
  • The result is an array of { text, embedding } objects

Step 4: Storage

Chunks are stored in the knowledge_chunks table in Supabase (PostgreSQL with pgvector):

Column Type Description
id UUID Primary key
knowledge_base_id UUID FK to the knowledge base
source_type text pdf, url, text, docx, csv, xlsx
source_name text Original filename or URL
content text The chunk text
embedding vector(1536) The 1,536-dim embedding vector
chunk_index integer Position of this chunk in the source
metadata jsonb Additional info (title, status, error message)
created_at timestamptz When the chunk was created

Deduplication: Before re-ingesting, all existing chunks for the same source_name are deleted, preventing duplicate entries.


When a query comes in (either from passive RAG or the active tool), the system performs a cosine similarity search using pgvector.

Search Flow

Query Text → embedQuery() → 1,536-dim vector → db.searchSimilarChunks() → top K results

Search Parameters

Parameter Default Description
topK 3 Maximum number of chunks to return
threshold 0.25 (active tool) / configurable via RAG Threshold slider (passive) Minimum cosine similarity score

Search Variants

Function Scope Use Case
searchKnowledge() Entire KB General search — used by passive RAG and the global search_knowledge tool
searchKnowledgeBySources() Specific sources only Named Knowledge Functions — scoped to user-selected source documents

Result Format

Each result contains:

{
  "id": "chunk-uuid",
  "content": "The chunk text...",
  "sourceName": "services.pdf",
  "similarity": 0.82
}

Weak Result Detection

The system logs warnings for searches where the top result has a similarity score below 0.40:

[Pipeline] ⚠ KB WEAK RESULT — top score 0.312 for raw query: "do you accept bitcoin?"

This helps identify knowledge gaps — topics callers ask about that aren't covered in the KB.


5.6 Passive RAG (Automatic Retrieval)

Passive RAG runs automatically during every Modular pipeline call. The user doesn't configure it — it just works when a KB is linked to the agent.

How It Fires

  1. The user finishes speaking (STT produces a transcript).
  2. The pipeline checks if the utterance is meaningful — it strips common conversational words (yes, no, yeah, okay, please, thanks, etc.) and requires at least 2 meaningful words remaining.
  3. If meaningful, _retrieveKBContext() fires in parallel with the patience delay timer.
  4. A 1,500ms circuit breaker ensures RAG never blocks the LLM. If the search takes longer, it's abandoned and the LLM proceeds without context.

Query Cleaning

Before searching, the pipeline cleans the caller's speech to improve embedding quality:

Filler stripping patterns:

1. Greetings:    "hi", "hey", "hello", "good morning"  (removed from start)
2. Filler words: "can you", "please", "just", "like", "actually", "um", "uh"
3. Request forms: "tell me", "I want to know", "I'm wondering"
4. Articles:     "at the", "in the", "of the"

Protected phrases (not destroyed by filler stripping):

"right now", "right away", "no longer", "no refund", "not yet",
"how much", "how many", "how long", "at least", "at most"

Example:

Input:  "hi can you tell me who are the available hygienists right now"
Output: "available hygienists right now"

If cleaning removes too much (result < 5 characters), the original text is used.

Context Buffer (Sliding FIFO Window)

Retrieved KB context is stored in a sliding buffer that persists across conversation turns:

Property Value Description
_kbContextBuffer Array Stores { query, context, timestamp } entries
_kbBufferMaxSize 3 Maximum entries before FIFO eviction

This means the AI has access to relevant KB context from the current and previous 2 turns, enabling follow-up questions like:

Caller: "What are Dr. Smith's hours?" → KB retrieves Dr. Smith info
Caller: "And what's her specialty?" → Dr. Smith info still in buffer

System Prompt Injection

When runLLM() fires, the KB buffer is injected into the system prompt:

[KNOWLEDGE BASE — Use this information to answer the user's question accurately.
If you cannot find the answer in this context, politely state that you do not
have that information on file.]

{chunk 1 text}

{chunk 2 text}

{chunk 3 text}

The injection is non-destructive — it appends to a clone of the system message, not the original.


5.7 Active Knowledge Tool

The active tool (search_knowledge) lets the LLM proactively search the KB during a call. Unlike passive RAG (which fires on every utterance), this is a function the AI explicitly invokes when it decides it needs information.

Tool Definition

{
  "name": "search_knowledge",
  "description": "Search the knowledge base for information relevant to the current conversation...",
  "parameters": {
    "query": {
      "type": "string",
      "description": "A specific search query. Be descriptive..."
    },
    "category": {
      "type": "string",
      "description": "Optional category to narrow results.",
      "enum": ["url", "file", "text"]
    }
  }
}

Per-Call Cache

The active tool includes a per-call result cache to avoid redundant embedding searches:

  • Cache key: normalized (lowercased, trimmed) query string
  • Cache lifetime: duration of a single phone call
  • On cache hit: returns cached results instantly (no API call)

Response Format

Results found:

{
  "results": [
    { "content": "Dr. Smith specializes in...", "source": "team.pdf", "relevance": 0.87 },
    { "content": "Office hours are Mon-Fri...", "source": "https://example.com/hours", "relevance": 0.72 }
  ]
}

No results:

{
  "results": [],
  "message": "No relevant information found in the knowledge base for this query. You may need to answer based on your general knowledge or let the customer know you don't have that specific information."
}

When Is the Tool Available?

The search_knowledge tool is added to the agent's tool list only if a knowledge base is linked (knowledgeBaseId is set). It is part of the ALWAYS_ENABLED_TOOLS whitelist, meaning it's always available even if not explicitly referenced in the system prompt via [[Function: search_knowledge]].


5.8 Named Knowledge Functions

Named Knowledge Functions are user-created, scoped search tools that only search specific sources within a KB. They provide more targeted retrieval than the global search_knowledge tool.

Use Case

Imagine your KB has 50 sources covering your entire business. You want the AI to search only your troubleshooting guide when a caller reports a technical issue, not your pricing page or team bios. A Named Knowledge Function like "Router Troubleshoot" scoped to troubleshooting_guide.pdf achieves this.

How They Work

  1. You create a function in the Functions panel and select "Knowledge" as the category.
  2. Choose which specific sources (files/URLs) the function should search.
  3. The function gets a sanitized tool name: "Router Troubleshoot""router_troubleshoot".
  4. At call time, buildNamedKnowledgeToolHandler() creates a handler that uses searchKnowledgeBySources() — filtering results to only the selected source names.
Feature search_knowledge (Global) Named Function (Scoped)
Scope Entire knowledge base Selected sources only
Auto-enabled Yes (always available if KB linked) Must be attached to agent and referenced in prompt
Cache Per-call Per-call
topK 3 3

5.9 Chunk Management

The UI provides granular control over individual chunks within each source.

Viewing Chunks

Click a source in the source list to open the Chunk Viewer Modal. This shows:

  • All chunks in order (chunk_index)
  • Each chunk's text content
  • Created timestamp

Editing Chunks

Click on a chunk to enter edit mode:

  1. Modify the chunk text in a textarea.
  2. Click "Save".
  3. The backend calls updateChunkContent() which:
    • Re-embeds the new text via embedQuery()
    • Updates both the content and embedding columns in the database
    • This ensures the new text is correctly searchable

Deleting Chunks

Click the delete button (🗑️) on any chunk:

  1. A confirmation modal appears.
  2. The chunk is deleted from the database.
  3. The source's chunk count updates.
  4. If it was the last chunk in a source, the chunk modal closes automatically.

Bulk Operations

Select multiple sources using checkboxes:

Action Description
Bulk Delete Delete all selected sources and their chunks
Bulk Update Re-crawl all selected URL sources (fetches fresh content)

Bulk Update is only available for URL sources — it re-crawls each URL in the background and replaces existing chunks with fresh content.

A search bar at the top of the source list filters sources by name in real-time.


5.10 Linking a KB to an Agent

To make a knowledge base available to an agent:

  1. Open the agent's settings page.
  2. In the Configuration Sidebar, open the Knowledge section.
  3. Click the Knowledge Selector to choose a KB from your workspace.
  4. Set the RAG Threshold (0.0–1.0, default 0.25) to control minimum relevance.

When linked:

  • Passive RAG activates automatically for every call.
  • The search_knowledge tool becomes available to the AI.
  • The KB's instructions (if any) are appended to the system prompt.

Unlinking

Click "Unlink" in the Knowledge section. This removes the KB association — passive RAG stops, the search tool is removed, and the instructions are no longer appended.


5.11 Limits & Performance

Size Limits

Limit Value
KB total size 50 MB
Per-file upload 10 MB
URL scan limit 100 internal links
RAG circuit breaker 1,500ms timeout

Performance Characteristics

Operation Typical Latency
File upload + ingest (small PDF) 3–8 seconds
URL crawl (single page) 5–15 seconds
Vector search (query → results) 100–300ms
LLM pre-processing (per source) 2–5 seconds
Embedding (100 chunks batch) 1–2 seconds

Embedding Model

Property Value
Model text-embedding-3-small
Provider OpenAI
Dimensions 1,536
Batch Size 100 chunks per API call

Database

Property Value
Database Supabase PostgreSQL
Extension pgvector
Index Type Cosine similarity
Table knowledge_chunks

Previous: Section 4 — Phone Numbers & Telephony Next: Section 6 — Functions & Integrations


Section 6 — Functions & Integrations

Technical architecture of Wirevox's tool-calling pipeline — spanning core tools, CRM/calendar integrations, the dynamic Tool Assembler, the orphan tool filter, runtime normalization, and the shared Post-Execution Layer (Executor).


6.1 Overview

Functions (also referred to as tools) allow your AI agent to take action in the real world. Rather than just speaking, the agent can book appointments, lookup patient records, log CRM leads, send email/SMS notifications, and transfer calls to humans.

Wirevox utilizes a declarative, three-layer tool execution model:

┌─────────────────────────────────────────────────────────────┐
│ 1. DEFINITION & ASSEMBLY (Compile Time)                     │
│                                                             │
│  Workspace Integrations Connected                           │
│        │                                                      │
│        ▼                                                      │
│  Orphan Filter (filterOrphanTools.js)                       │
│  Strips tools for disconnected integrations                   │
│        │                                                      │
│        ▼                                                      │
│  Tool Assembler (assembleTools.js)                          │
│  Generates OpenAI schemas, custom fields & prompt overrides  │
└─────────────────────────────────────────────────────────────┘
                               │
                               ▼
┌─────────────────────────────────────────────────────────────┐
│ 2. NORMALIZATION & DISPATCH (Call Runtime)                  │
│                                                             │
│  LLM invokes tool (e.g., jobber_create_lead)                │
│        │                                                      │
│        ▼                                                      │
│  Universal Dispatcher (inboundHandler.js)                   │
│  - Cleans spoken phone numbers/emails to standard digits     │
│  - Rejects incomplete emails with instructions to AI        │
│  - Generates warm handoff tokens for transfers               │
└─────────────────────────────────────────────────────────────┘
                               │
                               ▼
┌─────────────────────────────────────────────────────────────┐
│ 3. CORE EXECUTOR & POST-PROCESSING (Shared Layer)           │
│                                                             │
│  Shared Executor (executor.js)                              │
│  - Executes thin adapter API calls                          │
│  - Inserts local leads (duplicate check & failover path)    │
│  - Sets call outcomes (priority check: booked > message)    │
│  - Fires destination-specific webhooks                       │
│  - Deducts billing credits (e.g., SMS, transfer tools)      │
└─────────────────────────────────────────────────────────────┘

6.2 Tool Categories & Function Catalog

Wirevox organizes actions into 8 distinct function categories managed via the Functions page (/w/:id/functions):

Category UI Label Target Tool Schema / Custom Behavior
integrations Integrations CRM/Calendar tools Generated dynamically from integration adapters
call_transfer Call Transfer transfer_call Core tool. Redirects caller via Telnyx
sms_confirmation SMS Confirmation send_sms_confirmation Sends SMS via Telnyx. Supports templates
email_confirmation Email Confirmation send_email_confirmation Sends emails. Supports templates
gather_information Gather Information system_infolist Augments lead fields with custom properties
knowledge Knowledge Named scoped search Generates a scoped search tool for selected docs
workflow Workflow Named graph SOP Injects a conversational tree state-machine prompt
take_message Take a message take_message Core tool. Collects caller message details

6.3 Core Tools (Always Available)

Core tools are built into the platform and do not require external API configurations.

transfer_call

Transfers the caller to another number.

  • Parameters: reason (string)
  • Behavior: Shuts down the AI pipeline to freeze billing and commands Telnyx to bridge the call.

system_infolist

Saves caller details directly to the local database as a lead.

  • Parameters: name (string, required), phone (string, required), email (string, optional), reason (string, optional), plus dynamic custom fields.

send_sms_confirmation

Sends a text confirmation to the caller.

  • Parameters: toPhone (string, required), message (string, required)

send_email_confirmation

Sends an email confirmation to the caller.

  • Parameters: toEmail (string, required), subject (string, required), body (string, required)

take_message

Logs a message for the business owner.

  • Parameters: callerName (string, required), callerPhone (string, required), message (string, required), recipientName (string, optional), callerEmail (string, optional), urgency (enum: normal, urgent)

6.4 Integrations Catalog

Integration tools are loaded from Connected Workspace Providers.

Integrations Catalog (INTEGRATION_TOOLS)
├── Google Calendar (google_calendar)
│   ├── google_check_availability
│   ├── google_book_appointment
│   ├── google_lookup_appointment
│   ├── google_cancel_appointment
│   └── google_reschedule_appointment
├── Outlook Calendar (outlook_calendar)
│   ├── outlook_check_availability
│   ├── outlook_book_appointment
│   ├── outlook_lookup_appointment
│   ├── outlook_cancel_appointment
│   └── outlook_reschedule_appointment
├── Jobber CRM (jobber)
│   ├── jobber_lookup_client
│   ├── jobber_create_lead
│   └── jobber_create_request
├── Clio Manage (clio)
│   └── clio_create_intake
├── Clio Grow (clio_grow)
│   └── clio_grow_create_lead
├── Filevine (filevine)
│   └── filevine_create_intake
├── Open Dental (open_dental)
│   ├── open_dental_lookup_patient
│   ├── open_dental_create_patient
│   ├── open_dental_check_availability
│   └── open_dental_book_appointment
└── Email (email)
    └── email_send_summary

6.5 The Tool Assembler (assembleTools)

When a call initializes, the assembleTools function maps attached agent functions into OpenAI-compatible tool definitions.

Category Mapping Details

  1. Named Knowledge & Workflows: Generates tool definitions named dynamically matching the sanitized function name (lowercase, non-alphanumeric replaced by _). Configs (_knowledgeConfig and _workflowConfig) are attached directly to the tool definition to guide the dispatcher.
  2. Integrations: Resolves enabled tools from the integration config (parsing the enabledTools object or tools array) and pushes their definitions to the schema pool.
  3. Template Overrides:
    • For SMS Confirmation and Email Confirmation, if a custom message template is defined, the tool's description is dynamically overwritten:

      "Send an SMS text confirmation. The user has provided a template: "[template]". You MUST use this exact template for your 'message' parameter, but replace any bracketed variables (like [date], [time]) with the actual values..."

  4. Core Category Mapping: Maps standard Call Transfer, Take a Message, SMS, and Email functions to their hardcoded core parameter schemas.

Dynamic system_infolist Augmentation

If a gather_information category function is attached to the agent:

  • The assembler reads custom field schemas defined in config.fields (or legacy config.customFields).
  • It dynamically appends these custom properties to the system_infolist parameter properties block.
  • For example, if you configure a field pet_breed in the UI, the LLM will see pet_breed as a valid argument on the system_infolist tool definition during call execution.

6.6 The Orphan Tool Filter (filterOrphanTools)

To prevent the LLM from attempting to execute integrations that are configured but disconnected in the workspace, the system uses filterOrphanTools:

  1. It parses all enabled tools for an agent.
  2. It collects unique providers needed (e.g., google_calendar, jobber).
  3. It query-checks the database for the active workspace to verify connection states.
  4. Any tool whose integration provider is disconnected is stripped from the active list.
  5. This acts as a safeguard, ensuring the LLM is never presented with invalid tool definitions.

6.7 Normalization & Dispatch (inboundHandler)

The receptionist handler (buildInboundToolHandler) acts as the universal dispatcher.

Normalization Helpers

Voice transcripts often spell out numbers or emails. Before dispatching arguments to any tool, the handler normalizes contact details:

// Convert words to digits
"five five five one two three four" ➔ "5551234"
  • Phone Numbers: Normalizes strings by converting word numerals to digits, stripping non-numeric characters, and ensuring E.164 formatting:
    • 10 digits ➔ prepends +1 (North America)
    • 11 digits starting with 1 ➔ prepends +
  • Emails: Normalizes spelled-out emails by replacing spoken symbols:
    • " at "@
    • " dot ".
    • " underscore "_
    • " dash " / " hyphen "-
    • All spaces are collapsed.

Invalid Email Blocking

If the normalized email does not contain both a @ and a ., the dispatcher blocks execution and returns an error message:

{
  "error": "Email address appears incomplete — it is missing the @ symbol or domain. Ask the caller to repeat their full email address before retrying."
}

This instructs the LLM to verify and request the email again, preventing trash entries.

Warm Agent Transfers & Handoff Tokens

When transferring calls to another agent (destinationType === 'agent'):

  1. The dispatcher looks up the target agent's phone number.
  2. If transferType === 'warm', the dispatcher generates a Handoff Token:
    • It captures the last 15 lines of the transcript.
    • It saves this metadata to the database via db.createHandoff().
    • It injects the token into custom headers sent to Telnyx: X-Wirevox-Handoff: <token>.
  3. When the receiving agent picks up the call, the backend resolves this token and pre-loads the transcript context, allowing the target agent to resume the conversation seamlessly.

6.8 Shared Post-Execution Layer (Executor)

All tool executions route through the shared executeTool function in executor.js. This centralizes post-execution logic and database synchronization.

1. Duplicate Lead Prevention & Fallback

  • Lead Insertion: If a tool's manifest defines createsLead: true, the executor constructs a lead record and saves it locally. To prevent multiple tools (e.g., booking an appointment and taking a message) from creating duplicate lead entries, the executor sets ctx._leadCreatedThisCall = true and skips subsequent lead creation.
  • Failover Storage: If an external API call fails (e.g., Google Calendar returns 500), the executor catches the error, inserts a local lead record with a [FAILED] prefix in the notes, and responds to the LLM with a graceful fallback message so the AI can continue speaking.

2. Outcome Priorities

The call's outcome column is updated based on tool execution. To avoid downgrading the status during a call, the executor enforces an OUTCOME_PRIORITY hierarchy:

booked (6) > rescheduled (5) > cancelled (4) > intake/transferred (3) > message_taken (2) > inquiry (1) > other (0)

A lower-priority outcome will never overwrite a higher-priority one (e.g., if a booking succeeds, a subsequent call transfer will not overwrite the call's outcome from booked to transferred).

3. Dynamic Webhook Routing

If the tool's manifest specifies a webhookEvent, the executor looks up the webhook URL configured for the parent function and fires the webhook asynchronously.

Augmenting system_infolist Webhooks

For system_infolist, the executor reads the gather_information configuration:

  • It only includes arguments where destWebhook: true is configured in the field metadata (or falls back to all arguments for legacy structures).
  • If a webhook URL is configured, it fires the information.gathered event with the filtered payload.

4. Billing & Duration Freezing

  • Billing Deductions: For billable actions, the executor calls db.deductCredits using dynamic cost registries (e.g., SMS sending, transfers).
  • Handoff Duration Freeze: For call transfers, the executor updates the call record's transferred_at timestamp. This freezes AI call duration billing, ensuring the user is only billed for AI duration up to the moment they are connected to a human.

Previous: Section 5 — Knowledge Base (RAG) Next: Section 7 — Post-Call Processing & Webhooks


Section 7 — Post-Call Processing & Webhooks

What happens after the caller hangs up — background transcript analysis, outcome classification, lead synchronization, email summaries, and webhook delivery.


7.1 Overview

When a call ends (the WebSocket closes), Wirevox doesn't just save the transcript and move on. A background post-call pipeline runs automatically to enrich every call with structured data:

Call Ends (WebSocket closes)
       │
       ├─ 1. Save transcript to database
       │
       ├─ 2. runPostCallAnalysis() ───▶ Background
       │      │
       │      ├─ a. Classify outcome (GPT-4o Mini)
       │      ├─ b. Generate summary
       │      ├─ c. Extract caller details
       │      ├─ d. Update call record
       │      ├─ e. Sync lead record
       │      └─ f. Fire post-call hooks (email summary)
       │
       └─ 3. Fire call.ended webhook (if configured)

Trigger Point

The pipeline fires from [mediaStream.js](file:///c:/Users/iysal/.gemini/antigravity/scratch/Wirevox App/backend/src/routes/telephony/mediaStream.js) in the ws.on('close') handler:

ws.on('close', async () => {
  pipeline.end();
  await db.updateCall(callId, {
    status: 'ended',
    endedAt: new Date().toISOString(),
    transcript: transcriptRecord,
  });

  // KICK OFF BACKGROUND ANALYSIS
  runPostCallAnalysis(callId, transcriptRecord, agent);

  // Fire call.ended webhook
  if (postCallWebhook) {
    fireWebhook(postCallWebhook, 'call.ended', {
      callId, agentName, agentType, transcript, messageCount
    });
  }
});

The analysis runs asynchronously — it doesn't block the WebSocket close. Even if analysis fails, the call is still marked as ended.


7.2 LLM-Powered Transcript Analysis

The core of post-call processing is a single call to GPT-4o Mini with a structured JSON response.

Input

The full transcript is formatted as:

AGENT: Hi, this is Dr. Smith's office. How can I help you?
USER: I'd like to schedule a cleaning appointment.
AGENT: Of course! I can help with that...

System Prompt

The LLM is given explicit instructions to:

  1. Classify the outcome — choose from a strict enum of valid outcomes
  2. Write a summary — concise 1-2 sentences
  3. Extract caller details — name, phone, email, reason for calling

Output Format

The response uses response_format: { type: 'json_object' } to guarantee structured output:

{
  "outcome": "booked",
  "summary": "Caller scheduled a dental cleaning for next Tuesday at 2pm with Dr. Smith.",
  "callerName": "Sarah Johnson",
  "callerPhone": "+15551234567",
  "callerEmail": "sarah.j@email.com",
  "reason": "Schedule a dental cleaning"
}

Valid Outcomes

Outcome Description
booked An appointment was successfully scheduled
transferred The call was transferred to a human staff member
message_taken The caller left a message for the team
inquiry The caller asked questions but didn't book or leave a message
cancelled An existing appointment was cancelled
rescheduled An existing appointment was rescheduled to a new time
intake Caller info was submitted to a CRM (Clio, Jobber, Filevine, etc.)
other Wrong number, disconnected early, or any other outcome

Outcome Priority Respect

If an outcome was already set during the call by a tool execution (e.g., google_book_appointment set booked), the post-call analysis skips re-classification and only runs enrichment (summary, caller details). This prevents the LLM from accidentally downgrading a confirmed booking to an "inquiry."

const skipClassification = existingOutcome
  && existingOutcome !== 'inquiry'
  && existingOutcome !== 'other';

Only inquiry and other outcomes are eligible for reclassification.


7.3 Call Record Update

After analysis, the call record in the calls table is updated with:

Field Source Description
outcome Tool or LLM classification Final call outcome
notes LLM-generated summary 1-2 sentence call summary
endedAt Server timestamp When the call ended
leadName LLM extraction Caller's name (if detected)
leadPhone LLM extraction or caller ID Caller's phone number

Phone Number Backfill

If the LLM couldn't extract a phone number from the transcript (e.g., the caller never stated it), the system backfills from:

  1. call.callerNumber — the Telnyx caller ID
  2. call.leadPhone — any phone captured during the call by tools

This ensures email summaries always include the phone number, even on quick inquiry calls where the caller didn't provide it verbally.


7.4 Lead Synchronization

After updating the call record, the system syncs data to the leads table (the Contacts page in the dashboard).

Two Scenarios

Scenario A: Lead Already Exists

If a tool created a lead during the call (e.g., google_book_appointment, system_infolist), the system finds the most recent lead for this agent (created within the last 5 minutes) and enriches it:

Update Description
notes Overwritten with the LLM-generated rich summary
status Updated to match the call outcome
name Backfilled if the tool didn't capture it
phone Backfilled if the tool didn't capture it
email Backfilled if the tool didn't capture it
Scenario B: No Lead Exists

If no tool created a lead during the call (e.g., a quick inquiry where the caller just asked a question), the system creates a new lead from the post-call analysis. This ensures every caller appears in the Contacts page, even if no tools were triggered.

await db.insertLeads(agent.id, [{
  name: analysis.callerName || 'Unknown Caller',
  phone: analysis.callerPhone || '',
  email: analysis.callerEmail || '',
  notes: analysis.summary || '',
  status: STATUS_MAP[outcome] || 'new',
}]);

Lead Status Mapping

Call Outcome Lead Status
booked booked
rescheduled booked
intake intake
message_taken message_taken
transferred transferred
cancelled cancelled
inquiry new
other new

7.5 Email Summary Hook

If the agent has the email_send_summary integration tool enabled, a professional HTML email is sent to the business owner after every call.

Trigger Logic

The post-call processor checks if:

  1. The agent's published function snapshot includes email_send_summary as an enabled tool
  2. The workspace has an email integration configured (with a target email address)

If both conditions are met, the email fires.

Email Content

The email is a professional HTML template with:

Section Content
Header Dark gradient banner: "Call Summary — Handled by {Agent Name}"
Table Caller Name, Phone, Email, Reason, Appointment details, Notes
Footer "Sent securely by Wirevox Receptionist"

Email Data Sources

Field Source
callerName LLM extraction from transcript
callerPhone LLM extraction or caller ID backfill
callerEmail LLM extraction
reason LLM extraction
notes LLM-generated summary
appointmentDate/Time From booking tool result (if applicable)

Important: Post-Call Only

The email_send_summary tool has postCallOnly: true in its tool definition. This means:

  • It is NOT fed to the live AI during a call — the LLM never sees or tries to call it
  • It fires deterministically after the call ends, via the post-call processor
  • This ensures the email always contains the complete call summary, not a partial one

7.6 Webhooks

Wirevox fires webhook events at two levels: per-tool (during the call) and per-call (after the call ends).

Webhook Architecture

export async function fireWebhook(webhookUrl, event, payload) {
  const body = {
    event,           // e.g. "call.ended"
    timestamp,       // ISO 8601
    data: payload,   // Event-specific data
  };

  // Fire-and-forget with 10-second timeout
  await fetch(webhookUrl, {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify(body),
  });
}

Key Properties

Property Value
Method POST
Content-Type application/json
Timeout 10 seconds (AbortController)
Failure behavior Non-fatal — logged and swallowed
Retry None (fire-and-forget)

Webhook Events

Per-Tool Events (fired during the call by the Executor)
Event Trigger
appointment.booked google_book_appointment or outlook_book_appointment or open_dental_book_appointment succeeds
appointment.cancelled google_cancel_appointment or outlook_cancel_appointment succeeds
appointment.rescheduled google_reschedule_appointment or outlook_reschedule_appointment succeeds
call.transferred transfer_call tool fires
message.taken take_message tool succeeds
information.gathered system_infolist tool succeeds (with Gather Information function attached)
clio.intake_created clio_create_intake succeeds
clio_grow.lead_created clio_grow_create_lead succeeds
jobber.lead_created jobber_create_lead succeeds
jobber.request_created jobber_create_request succeeds
filevine.intake_created filevine_create_intake succeeds
email.summary_sent email_send_summary completes
Per-Call Events (fired after the call ends)
Event Trigger Payload
call.ended WebSocket closes callId, agentName, agentType, transcript, messageCount

Webhook URL Configuration

Webhook URLs are configured at two levels:

  1. Per-Function Webhook — Set in the function's configuration (config.webhookUrl). Per-tool events route to this URL.
  2. Post-Call Webhook — Set in the agent's configuration (agentConfig.postCallWebhook). The call.ended event routes to this URL.

Compatible Platforms

Webhooks work with any HTTP endpoint that accepts POST requests:

  • Zapier — Use a "Webhooks by Zapier" trigger
  • Make (Integromat) — Use a "Webhooks" module
  • n8n — Use a "Webhook" trigger node
  • Custom endpoints — Any server that accepts JSON POST requests

Webhook Payload Structure

Every webhook follows the same envelope format:

{
  "event": "appointment.booked",
  "timestamp": "2026-07-16T05:27:00.000Z",
  "data": {
    "callerName": "Sarah Johnson",
    "callerPhone": "+15551234567",
    "service": "Dental Cleaning",
    "preferredDate": "2026-07-23",
    "preferredTime": "14:00",
    "agentName": "Front Desk AI"
  }
}

7.7 Error Handling & Resilience

Analysis Failure

If the GPT-4o Mini analysis call fails (API error, timeout, etc.), the call is still saved with a fallback:

await db.updateCall(callId, {
  outcome: 'other',
  notes: 'Post-call analysis failed.',
  endedAt: new Date().toISOString(),
});

This prevents calls from remaining in a "zombie" state with no end marker.

Email Hook Failure

Email summary failures are non-fatal. If the SMTP send fails:

  • The error is logged
  • The rest of the post-call flow continues
  • The call record and lead sync are unaffected

Webhook Failure

Webhook delivery failures are completely non-fatal and non-blocking:

  • 10-second timeout via AbortController
  • Errors are logged but swallowed
  • No retries — fire-and-forget
  • Never affects call flow or post-call processing

7.8 Processing Timeline

Here's the typical timing for the entire post-call pipeline:

Step Timing Description
Transcript Save ~50ms Immediate DB write on WebSocket close
LLM Analysis ~1-3s GPT-4o Mini processes transcript
Call Update ~50ms Update call record with outcome + summary
Lead Sync ~50-100ms Create or enrich lead record
Email Summary ~1-2s Build HTML + SMTP delivery
Webhook Delivery ~100-500ms POST to configured URLs

Total: ~2-6 seconds after call ends. All operations run in the background — the caller has already hung up and the WebSocket is closed.


Previous: Section 6 — Functions & Integrations Next: Section 8 — Billing & Credits


Section 8 — Integrations Setup

Technical reference for how Wirevox connects to third-party CRMs and Calendars — covering OAuth flows, static API keys, credential storage, and the cleanup lifecycle.


8.1 Overview

The Integrations page (/w/:id/integrations) allows users to connect external systems to their workspace. Once an integration is connected, its corresponding tools become available to all agents within that workspace.

Wirevox supports two primary authentication patterns for external systems:

  1. OAuth 2.0 (Google Calendar, Outlook Calendar, Jobber, Clio Manage, Clio Grow)
  2. Static API Keys / PATs (Filevine, Open Dental)

Database Storage

All integration credentials are stored in the integrations table:

Column Description
workspace_id Links the integration to a specific workspace
provider The integration ID (e.g., google_calendar, jobber)
access_token Short-lived OAuth access token
refresh_token Long-lived OAuth refresh token (OR static API key)
token_expiry Timestamp when the access_token expires
email Captured user email (used for Google/Outlook/Email integrations)

[CODE SMELL FLAG] The refresh_token Column

Currently, Wirevox uses the refresh_token column to store static credentials for non-OAuth integrations (like Filevine PATs and Open Dental Customer Keys) because the table lacks a dedicated api_key or credentials JSONB column.

While functional, this is a known architectural code smell. Future refactoring should introduce a dedicated JSONB credentials column to separate OAuth tokens from static API keys.


8.2 OAuth 2.0 Flow

For OAuth-based integrations (Google, Outlook, Jobber, Clio, Clio Grow), Wirevox implements a standard Authorization Code flow.

1. Generating the Auth URL

When a user clicks "Connect" in the UI, the frontend calls GET /api/integrations/{provider}/auth.

The backend generates the OAuth consent URL. Crucially, it passes the userId and workspaceId as a JSON-stringified state parameter. This ensures the system knows which workspace to attach the credentials to when the callback returns.

// Example: Google Calendar
const state = JSON.stringify({ u: req.userId, w: workspaceId });
const url = oauth2Client.generateAuthUrl({
  access_type: 'offline',
  prompt: 'select_account consent',
  scope: [...scopes],
  state,
});

2. The Callback Handler

When the user approves the connection, the provider redirects to /api/integrations/{provider}/callback?code=...&state=...

The callback route:

  1. Parses the state parameter to recover the userId and workspaceId.
  2. Exchanges the code for an access_token and refresh_token.
  3. Calls db.upsertIntegration to store the tokens in the database.
  4. Redirects the user back to the frontend Integrations page with a success flag.

3. Token Refreshing

When an AI agent executes a tool (e.g., booking an appointment), the integration adapter checks the token_expiry. If the token is expired (or within a 5-minute buffer), the backend automatically uses the refresh_token to fetch a new access_token, updates the database, and then proceeds with the API call.


8.3 Specific Provider Configurations

Google Calendar

  • Auth Type: OAuth 2.0
  • Scopes: calendar.events, calendar.readonly, userinfo.email
  • Settings: access_type=offline, prompt=select_account consent (forces a refresh token to be issued)
  • Extra: Captures the user's email address to display in the UI. Currently defaults to the primary calendar.

Outlook Calendar

  • Auth Type: OAuth 2.0
  • Scopes: Calendars.ReadWrite, User.Read, offline_access
  • Extra: Captures the user's email address.

Jobber (Field Services CRM)

  • Auth Type: OAuth 2.0
  • Scopes: clients, requests
  • Auth Type: OAuth 2.0
  • Scopes: Standard CRM scopes.
  • Auth Type: Static API Key (Personal Access Token)
  • Flow: The user manually inputs a PAT, Client ID, and Client Secret in the UI.
  • Validation: The backend immediately tests the credentials by attempting a token exchange with the Filevine API. If successful, the combined credentials payload is stored in the refresh_token column.

Open Dental

  • Auth Type: Static API Key (Customer Key)
  • Flow: The user inputs an Open Dental Developer Customer Key in the UI.
  • Validation: The backend tests the key by calling the Open Dental API to fetch providers. If successful, the key is stored in the refresh_token column.

Email Notifications

  • Auth Type: None
  • Flow: The user simply inputs a target email address in the UI.
  • Storage: Stored in the email column. Used by the Post-Call processing pipeline to send call summaries via Wirevox's internal SMTP server.

8.4 The Cleanup Lifecycle (cleanupToolsForProvider)

When a user deletes an integration from their workspace, it's not enough to simply delete the API keys.

If agents in the workspace are actively using tools from that provider (e.g., google_book_appointment), disconnecting the integration would leave those tools broken and "orphaned" on the agent.

To handle this, routes/integrations.js implements a cleanup function:

const PROVIDER_TOOL_PREFIXES = {
  google_calendar: 'google_',
  jobber: 'jobber_',
  // ...
};

async function cleanupToolsForProvider(userId, workspaceId, provider) {
  const prefix = PROVIDER_TOOL_PREFIXES[provider];
  
  // 1. Get all agents in the workspace
  const agents = await db.getAgents(userId, workspaceId);
  
  for (const agent of agents) {
    const tools = await db.getAgentTools(agent.id);
    
    // 2. Find any tools starting with the provider prefix
    for (const tool of tools) {
      if (tool.toolName.startsWith(prefix)) {
        // 3. Hard delete the tool from the agent
        await db.deleteAgentTool(agent.id, tool.toolName);
      }
    }
  }
}

When DELETE /api/integrations/{provider} is called:

  1. The backend attempts to revoke the token with the provider (if OAuth).
  2. Deletes the row from the integrations table.
  3. Calls cleanupToolsForProvider to silently strip the associated tools from all agents in the workspace.

This ensures the system state remains consistent and agents do not attempt to use tools they no longer have credentials for.

(Note: There is also an filterOrphanTools.js layer that acts as a runtime safeguard against orphan tools, but this cleanup script handles the persistent database cleanup).


Previous: Section 7 — Post-Call Processing & Webhooks Next: Section 9 — Billing & Credits


Section 9 — Billing & Credits

Technical architecture of Wirevox's billing system — covering abstract credits, the real-time dynamic Cost Registry, Stripe integration, and the carryover billing mechanism.


9.1 Overview: The Credit System

Wirevox does not bill directly in minutes or API requests. Instead, it uses Credits as an abstract internal unit of measurement.

This abstraction solves a critical problem: Not all AI calls cost the same. An agent using gpt-4o with elevenlabs TTS is vastly more expensive to operate than an agent using gpt-4o-mini with deepgram TTS.

By abstracting usage into Credits:

  1. Users have a single, unified balance.
  2. The platform can charge dynamically based on the exact AI stack the user configures for their agent.
  3. The platform can bill flat fees for discrete actions (e.g., sending an SMS, or transferring a call) against the same balance.

The dollar value of a single Credit depends on the user's Subscription Plan (their "overage rate").


9.2 The Cost Registry (costRegistry.js)

The actual cost of an action is calculated at runtime by the costRegistry.js module.

Dynamic Database Configuration

To allow pricing to be adjusted without redeploying code, the system loads credit values from a Supabase table (credit_costs). The backend subscribes to Supabase Realtime (postgres_changes) to instantly invalidate its local cache whenever a cost is updated in the database.

If the database is unreachable, the system falls back to a hardcoded DEFAULT_COSTS map.

Per-Minute Calculation

The cost of 1 minute of call time is calculated as: Total Cost = Platform Base + STT Cost + LLM Cost + TTS Cost

Example Default Values:

  • Platform Base: 4 credits (Covers Telnyx telephony and basic infrastructure)
  • STT (Deepgram): 2 credits
  • LLM (GPT-4o-mini): 1 credit
  • TTS (Deepgram): 2 credits
  • Total: 9 credits per minute

If a user upgrades their agent to use ElevenLabs TTS (10 credits) and GPT-4o (3 credits), their per-minute cost dynamically scales to 19 credits per minute.

Flat Action Costs

The registry also defines flat costs for tool executions. These are billed deterministically when the Executor processes a tool:

  • func_sms: 2 credits (Billed per SMS sent)
  • func_transfer: 20 credits (Billed once when a call is successfully transferred to a human)

9.3 Per-Second Billing & Carryover

Wirevox bills callers for exact AI duration.

To prevent charging users for fractional minutes while still maintaining integer credit balances, the system uses a Carryover Buffer (billing_carryover_seconds on the agent record).

How it Works:

  1. A call lasts for 1 minute and 15 seconds.
  2. The user is immediately billed for 1 minute (e.g., 9 credits).
  3. The remaining 15 seconds are added to the agent's carryover buffer.
  4. On the next call, if the call lasts 50 seconds, the system adds the 15 seconds from the buffer (Total = 65 seconds).
  5. The user is billed for 1 minute (9 credits), and the buffer now holds 5 seconds.

Design Decision: Un-stamped Carryover

Carryover seconds are stored as raw durations, not rate-stamped values. If a user changes their agent's AI stack (e.g., swapping to a more expensive TTS model), any accrued seconds in the buffer will be billed at the new rate when they eventually cross the 60-second threshold.

This is an accepted architectural simplification because manual AI model reconfiguration is rare, and the maximum discrepancy is limited to 59 seconds' delta.


9.4 Stripe Integration & Tiers

Billing is handled via Stripe, implemented in routes/stripe.js. The platform supports a hybrid SaaS model with fixed monthly tiers and dynamic pay-as-you-go top-ups.

Subscription Tiers

Tier Monthly Price Description
PAYG $0 Pay-As-You-Go. User only pays for manual top-ups.
Starter $99 Includes a baseline amount of monthly recurring credits and 1 concurrent call slot.
Growth/Standard $299 Higher volume, lower per-credit overage rate.
Agency $999 White-label capabilities, sub-account management, lowest overage rate.

Add-on Products (Concurrent Calls)

To prevent abuse, the platform restricts how many AI calls can happen simultaneously per workspace. Users can purchase Extra Concurrent Call Add-ons via Stripe. The price of the add-on scales inversely with the subscription tier:

  • Starter Add-on: $15/mo per extra line
  • Growth Add-on: $13/mo per extra line
  • Agency Add-on: $5/mo per extra line

Upgrades with Proration Simulation

Wirevox provides a seamless upgrade path. When a user previews an upgrade in the UI, routes/stripe.js calls Stripe's Invoice Preview API.

This calculates the exact prorated amount the user will be charged today (crediting them for unused time on their current plan, and factoring in any existing add-ons) before they commit to the change.

Dynamic Top-Ups

If a user runs out of credits mid-month, they can trigger a manual Top-Up. The cost of a top-up depends on their plan's overage rate (overageRateCreditsCents).

  • The backend accepts the desired credit amount (minimum 50).
  • Calculates the dollar value: Amount * overageRateCreditsCents.
  • Generates a Stripe Checkout session for a one-time payment.

Previous: Section 8 — Integrations Setup