Section 1 — Getting Started
The complete guide to signing up, setting up your workspace, and making your first AI voice call with Wirevox.
1.1 Platform Overview
What Is Wirevox?
Wirevox is an enterprise-grade AI voice platform that lets you build, deploy, and manage intelligent phone agents. These agents can answer inbound calls 24/7, book appointments, route callers to humans, respond to FAQs using your own business data, and integrate with your existing CRM and calendar tools — all without writing a single line of code.
Who Is It For?
Wirevox is built for:
- Small & mid-size businesses that want a 24/7 AI receptionist to answer calls, qualify leads, and book appointments.
- Agencies & resellers that build and manage AI agents on behalf of their own clients, using Wirevox's white-label agency layer.
- Enterprise teams that need multi-agent architectures with CRM integrations, knowledge bases, and strict tenant isolation.
- Developers & SaaS builders that want to embed conversational AI telephony into their products.
Core Value Proposition
| Capability | What It Means |
|---|---|
| Modular Voice Pipeline | Real-time audio streaming bridging Telnyx telephony with Deepgram/Gladia (STT), OpenAI/Gemini/Claude (LLMs), and ElevenLabs/Deepgram/OpenAI (TTS). Sub-500ms response latency. |
| RAG Knowledge Base | Upload PDFs, DOCX, CSVs, or crawl your website — your agent becomes an expert on your business data. |
| Functions & Tool Calling | Your agent can book appointments, send SMS, transfer calls, fire webhooks, and sync with CRMs — all mid-conversation. |
| Multi-Agent Architecture | Build specialized agents (receptionist, sales, support) and route callers between them. |
| Agency / White-Label | Resell AI agents under your own brand with custom domains, plans, and subaccount management. |
| 9 Native Integrations | Google Calendar, Outlook Calendar, Clio, Clio Grow, Filevine, OpenDental, Jobber, Email (Mailgun/SMTP). |
1.2 Creating an Account
Wirevox supports two sign-up methods: email + password and Google OAuth.
Option A: Sign Up with Email
- Navigate to the sign-up page at
/signup. - You'll see the heading "Create account" with the subtitle "Start deploying voice agents in minutes."
- Fill in the following required fields:
- First name — e.g.,
Jane - Last name — e.g.,
Doe - Email — e.g.,
you@company.com - Password — must meet all four requirements (see below)
- First name — e.g.,
- Click "Create account".
Password Requirements
As you type your password, a real-time strength meter appears with four progress bars. Your password must satisfy all four criteria to turn the bar fully green:
| Requirement | Rule |
|---|---|
| Length | At least 8 characters |
| Lowercase | Contains at least one lowercase letter (a–z) |
| Uppercase | Contains at least one uppercase letter (A–Z) |
| Number | Contains at least one digit (0–9) |
The bar color changes dynamically as you satisfy more criteria:
- 1 met → Red
- 2 met → Amber
- 3 met → Blue
- 4 met → Green ✓
Email Verification
After submitting, you're redirected to a "Check your inbox" screen. Wirevox sends a verification link to your email address.
- The verification email is sent via Supabase Auth.
- A resend button appears with a 60-second cooldown timer. After the cooldown, you can click "Resend Email" to request a new link.
- If you entered the wrong email, click "Change email" to go back to the sign-up form.
Once you click the verification link in your inbox, your account is activated and you're signed in automatically.
Option B: Sign Up with Google
- Navigate to
/loginor/signup. - Click the "Continue with Google" button at the top of the form.
- Google's account picker opens (Wirevox uses
prompt: 'select_account'so you can choose which Google account to use). - After authorizing, you're redirected back to Wirevox.
- On first Google sign-in, you'll see the "Welcome to Wirevox!" confirmation page where:
- Your first and last name are pre-filled from your Google profile.
- Your email is shown (read-only, greyed out).
- Confirm your name and click "Create Account" to proceed.
Note: If you originally signed up with Google and later try to use email/password login, Wirevox shows a helpful error message: "It looks like you registered this email using Google. Please click 'Continue with Google' instead."
Legal Agreement
By signing up (either method), you agree to Wirevox's Terms of Service and Privacy Policy. This notice appears at the bottom of both the sign-in and sign-up forms.
1.3 Workspace Setup
After your account is created and email is verified, the Onboarding Guard checks your profile. If your onboarding_complete flag is false, you're routed to the Workspace Setup page before you can access the dashboard.
Workspace Setup Form
The workspace setup screen has the heading "Set up your workspace" with the subtitle "Tell us a bit about your business to get started." It collects four fields in a 2×2 grid:
| Field | Required? | Default | Description |
|---|---|---|---|
| Workspace Name | ✅ Yes | — | The name of your business or organization (e.g., "SunPath Energy"). This becomes the name displayed in the sidebar and workspace switcher. |
| Website Link | Optional | — | Your business website URL. Used for context and can be crawled later for your knowledge base. |
| Industry | Optional | — | Searchable dropdown with 36 industry categories (see list below). Helps Wirevox tailor prompt templates and recommendations. |
| Timezone | Optional | America/New_York |
Searchable timezone selector. Used by agents for scheduling-related functions. |
Click "Create workspace" to save. Your profile is marked onboarding_complete: true and you're redirected to the dashboard at /overview.
Supported Industries
The industry dropdown includes a search filter and the following categories:
| Accounting & Tax Services | Agriculture & Farming | Architecture & Planning |
| Arts & Entertainment | Automotive (Dealerships, Repair, etc) | Beauty & Personal Care |
| Biotechnology | Construction & Contracting | Consulting Services |
| Dental | E-Commerce & Retail | Education & E-Learning |
| Energy & Utilities | Financial Services & Wealth Management | Fitness, Wellness & Sports |
| Food & Beverage | Government & Public Sector | Healthcare & Medical |
| Home Services (Plumbing, HVAC, Roofing) | Hospitality, Travel & Tourism | Human Resources & Staffing |
| Insurance | IT Services & Cybersecurity | Legal Services |
| Logistics & Supply Chain | Manufacturing | Marketing, Advertising & PR |
| Media & Publishing | Non-Profit & NGO | Real Estate |
| Real Estate Investing | Software & SaaS | Telecommunications |
| Transportation & Delivery | Veterinary Services | Other |
1.4 Onboarding Survey
Shortly after your first login, a floating modal appears over the dashboard asking four quick questions. This survey can be skipped at any time via the "Skip" link in the top-right corner.
The survey is designed to help the Wirevox team understand its user base and tailor the product experience. It consists of 4 multiple-choice questions shown one at a time with a progress bar:
Question 1 — Where did you hear about us?
- Google / Search
- Twitter / X
- Friend / Colleague
- YouTube / Podcast
- Other
Question 2 — What is your company size?
- Just me
- 2 – 10
- 11 – 50
- 51 – 200
- 200+
Question 3 — Current monthly spend on voice / phone tasks?
- $0 (Handling it myself)
- $1 – $1,000 / mo
- $1,000 – $5,000 / mo
- $5,000 – $10,000 / mo
- $10,000+ / mo
Question 4 — What is your main goal with Wirevox?
- Automate inbound support
- Qualify outbound leads
- Handle appointment booking
- Just exploring
Clicking any option immediately advances to the next question (no submit button). After the final question, your answers are saved to your profile and the modal closes. If you skip the survey, it is marked as survey_completed: true with { skipped: true } and will not appear again.
1.5 Dashboard Tour
Once onboarding is complete, you land in the main dashboard. The interface is a two-panel layout: a collapsible sidebar on the left and the main content area on the right.
Sidebar Navigation
The sidebar contains your workspace context and navigation links. It has two states: expanded (default) and collapsed (auto-collapses when you enter settings, logs, knowledge, functions, playground, or agency pages).
Workspace Section (Top)
At the top of the sidebar is the Workspace Switcher — showing your workspace name and logo. Clicking it opens a dropdown to switch between workspaces or create a new one.
Below the workspace switcher, a credit balance indicator shows your remaining credits (included_credits + purchased_credits).
Navigation Items
| Icon | Label | Route | Description |
|---|---|---|---|
| ✨ | Vox | /w/:id/vox |
AI agent builder assistant (coming soon) |
| 🤖 | Agents | /w/:id/agents |
View, create, and configure your AI agents |
| 🧪 | Playground | /w/:id/playground |
Test voices and agent configurations |
| 📊 | Analytics | /w/:id/analytics |
Dashboard with call volume, usage trends, and active call indicators |
| 📋 | Logs | /w/:id/logs |
Call history, transcripts, and chat logs |
| 👥 | Contacts | /w/:id/contacts |
Leads and contacts extracted from calls |
| 📚 | Knowledge | /w/:id/knowledge |
Manage your RAG knowledge base (documents, websites, text) |
| 🔌 | Functions | /w/:id/functions |
Create and manage custom functions (webhooks, built-in tools) |
| 📞 | Numbers | /w/:id/numbers |
Buy and manage phone numbers |
| 🏢 | Agency | /w/:id/agency |
Agency dashboard (only visible on agency-tier workspaces) |
| ⚙️ | Settings | /w/:id/settings |
Workspace preferences, billing, team, integrations, and more |
Bottom Section
The bottom of the sidebar shows your profile avatar (initials), your display name (or truncated email), and a dropdown menu with:
- Account Settings →
/account(profile, auth, system preferences, danger zone) - Sign Out
Mobile Navigation
On mobile devices, the sidebar is replaced by a bottom navigation bar with five tabs: Home (analytics), Agents, Logs, Plans (billing), and More.
1.6 Your First Agent — Step by Step
This walkthrough takes you from zero to a working AI voice agent in under 5 minutes.
Step 1: Open the Agents Page
Click "Agents" in the sidebar. If you have no agents yet, you'll see an empty state with a "+ New Agent" button.
Step 2: Select Agent Type
Click "+ New Agent". A modal appears with three agent type cards:
| Agent Type | Status | Description |
|---|---|---|
| Inbound | ✅ Available (marked "NEW") | Answers inbound calls, handles customer FAQs, books appointments, and routes complex issues to humans 24/7. |
| Outbound Agent | 🔜 Coming Soon | Runs proactive voice campaigns to qualify leads and book meetings on autopilot. |
| AI Chat Agent | 🔜 Coming Soon | Resolves visitor queries and captures leads directly from your website in real-time. |
Click the Inbound card to proceed.
Step 3: Choose a Prompt Template
After selecting "Inbound," the Template Selector opens. This is where you pick a starting configuration for your agent. You can:
- Filter by industry using the horizontal category scroller.
- Filter by language using the dropdown (English, Spanish, French, German, Italian, Portuguese, Dutch).
Available templates include:
| Template | Industry | Language |
|---|---|---|
| Blank (start from scratch) | — | — |
| General Receptionist | General | English |
| Dental Office | Dental | English |
| Dental Office (French) | Dental | French |
| Dental + Google Calendar | Dental | English |
| Legal Intake | Legal | English |
| Real Estate Agent | Real Estate | English |
| Real Estate (French) | Real Estate | French |
| Home Services | Home Services | English |
| Med Spa / Aesthetics | Healthcare | English |
| Insurance Agency | Insurance | English |
| Auto Dealership | Automotive | English |
| Salon & Barbershop | Beauty | English |
| Salon (French) | Beauty | French |
| Restaurant | Food & Beverage | English |
| Restaurant (French) | Food & Beverage | French |
| Hotel Front Desk | Hospitality | English |
| Hotel (French) | Hospitality | French |
| Fitness Studio | Fitness | English |
| E-Commerce Support | E-Commerce | English |
| IT Helpdesk | IT Services | English |
| Mortgage Broker | Financial | English |
| Childcare Center | Education | English |
| Education / Tutoring | Education | English |
| Travel Agency | Travel | English |
| Event Venue | Events | English |
| Veterinary Clinic | Veterinary | English |
Each template pre-fills:
- Agent name — e.g., "Sarah" for dental, "James" for legal
- System prompt — a complete, industry-specific prompt with guardrails
- Voice — a matching TTS voice (voice ID + provider)
- Template variables — placeholders like
{{business_name}},{{business_hours}}
Select a template and click "Create Agent" to generate your agent.
Step 4: Configure Your Agent
After creation, you're taken to the Agent Settings page — a comprehensive configuration panel where you can customize everything about your agent. Key sections:
- System Prompt — The AI's personality, instructions, and behavior rules. You can edit the template-generated prompt or write your own.
- AI Model — Choose from GPT-4o, GPT-4.1, GPT-5.x, Claude 3.5/4.5/4.6 Sonnet, Gemini Flash, and more.
- Voice — Pick from 50+ voices across ElevenLabs, Deepgram Aura-2, and OpenAI TTS providers. Preview each voice before selecting.
- STT (Speech-to-Text) — Choose between Deepgram Nova-3 (mono or flux) and Gladia.
- Language — Select the primary language and optionally enable bilingual mode.
- Call Settings — Configure max call duration (1–30 minutes), greeting message, and interruption sensitivity.
- Functions — Attach built-in tools (SMS, transfer, booking) or custom webhooks.
- Knowledge — Link knowledge base sources to give your agent context about your business.
Step 5: Make a Test Call
Every agent has a built-in Test Call panel accessible from the agent settings page. This lets you talk to your agent directly from your browser — no phone number required.
- Click the phone icon or "Test" button in the agent settings header.
- The test panel opens with a visual "Liquid Aura" orb animation.
- Click the call button to connect. Your browser microphone activates and streams audio to the backend via WebSocket.
- Talk to your agent. The AI responds in real-time with the voice and behavior you configured.
- Click "Hang Up" to end the test call.
Note: Test calls consume credits just like real calls. The cost depends on your agent's AI stack (STT + LLM + TTS combination).
Step 6: Buy a Phone Number and Go Live
Once you're happy with your agent's behavior:
- Go to Numbers in the sidebar.
- Search for a phone number by area code or country.
- Purchase the number (pricing depends on your plan — each plan includes a certain number of free phone numbers).
- Assign the number to your agent.
- Your agent is now live — callers to that number will be answered by your AI agent 24/7.
1.7 Glossary
Key terms you'll encounter throughout the Wirevox platform:
| Term | Definition |
|---|---|
| Agent | An AI-powered phone assistant that you create, configure, and deploy. Each agent has its own system prompt, voice, AI model, and phone number. |
| Workspace | An isolated environment containing your agents, phone numbers, knowledge bases, functions, and billing. You can create multiple workspaces. |
| Credits | Abstract units of usage. Every call minute, SMS, and function execution is measured in credits. The dollar value of a credit depends on your subscription plan. |
| Included Credits | Credits that come with your monthly subscription. These reset every billing cycle. |
| Purchased Credits | One-time credit top-ups bought separately. These never expire and are never reset, even if you cancel your plan. |
| STT (Speech-to-Text) | The service that converts the caller's spoken audio into text. Wirevox supports Deepgram Nova-3 and Gladia. |
| LLM (Large Language Model) | The AI brain that processes the caller's words and generates intelligent responses. Wirevox supports OpenAI (GPT), Anthropic (Claude), and Google (Gemini) models. |
| TTS (Text-to-Speech) | The service that converts the AI's text response back into natural-sounding speech. Wirevox supports ElevenLabs, Deepgram Aura-2, and OpenAI TTS. |
| Voice Pipeline | The real-time audio processing chain: Phone Audio → STT → LLM → TTS → Phone Audio. This is the "modular pipeline" that powers every call. |
| Realtime Model | An alternative pipeline mode using speech-to-speech models (OpenAI Realtime or Gemini Live) that skip the separate STT/LLM/TTS stages. |
| RAG (Retrieval-Augmented Generation) | A technique where the AI searches your uploaded documents/websites for relevant information and injects it into its context before answering. This is what powers the Knowledge Base. |
| Knowledge Base | A collection of documents, website pages, and text sources that your agent can search during calls to answer questions accurately. Uses vector embeddings stored in PostgreSQL (pgvector). |
| Functions | Actions your agent can perform mid-conversation, like booking an appointment, sending an SMS, transferring to a human, or firing a webhook to Zapier/Make.com. Powered by OpenAI Function Calling. |
| Barge-In | When a caller interrupts the AI while it's speaking. Wirevox uses Voice Activity Detection (VAD) to instantly stop the AI's speech and listen to the caller. |
| Backchanneling | Filler words the AI uses while processing (e.g., "Mhm", "I see", "Un moment" in French). These are language-specific and configurable. |
| Carryover Billing | Wirevox's per-second billing system. Short calls accumulate seconds across multiple calls for the same agent, and you're only charged when a full minute is reached. This prevents overcharging on brief calls. |
| System Prompt | The set of instructions that define your agent's personality, behavior, knowledge, and constraints. This is what makes each agent unique. |
| Template Variables | Placeholders in your system prompt (e.g., {{business_name}}, {{caller_name}}) that are dynamically replaced with real data at call time. |
| Concurrent Calls | The number of simultaneous phone calls your account can handle at once. This limit depends on your subscription plan and any purchased add-ons. |
| Post-Call Processing | Actions that happen automatically after a call ends — like generating a summary, extracting lead information, or firing a webhook. |
| Agency | A reseller account that manages AI agents on behalf of multiple client businesses (subaccounts), with white-label branding. |
| Subaccount | A client workspace created and managed by an agency. Subaccounts have their own agents, calls, and usage — fully isolated from other subaccounts. |
| Telnyx | The telephony provider that powers Wirevox's phone number provisioning, call routing, and media streaming. |
| Supabase | The backend infrastructure provider used for the database (PostgreSQL), authentication, file storage, and real-time subscriptions. |
Section 2 — Core Concepts
A deep dive into the fundamental building blocks of the Wirevox platform: workspaces, agents, the voice pipeline, realtime models, the credit system, and multi-agent architecture.
2.1 Workspaces
A workspace is the top-level organizational unit in Wirevox. Everything — agents, phone numbers, knowledge bases, functions, integrations, billing, and team members — lives inside a workspace. Workspaces are fully isolated from each other.
What a Workspace Contains
| Resource | Description |
|---|---|
| Agents | AI voice assistants you build and deploy |
| Phone Numbers | Purchased from Telnyx, assigned to agents |
| Knowledge Bases | Documents, websites, and text sources for RAG |
| Functions | Custom webhooks and built-in tools (SMS, transfer, booking) |
| Integrations | OAuth connections to CRMs and calendars |
| Team Members | Users you've invited with specific roles |
| Billing | Subscription plan, credits, usage tracking |
| Call History | Logs of all inbound/outbound calls and leads |
Multi-Workspace Support
Users can own and belong to multiple workspaces. Common use cases:
- Multiple businesses — A user who owns a dental practice and a law firm can create separate workspaces for each with different agents, numbers, and billing.
- Separate environments — A development workspace and a production workspace.
- Team access — A user might be the owner of their own workspace and a team member in someone else's.
Workspace Types
The system recognizes three workspace access types, determined by the tenantContext:
| Access Type | Description |
|---|---|
personal |
A standard workspace owned directly by the user. This is the default. |
agency |
A workspace on the Agency tier that can create and manage subaccounts. |
subaccount |
A workspace created by an agency for one of their clients. The subaccount user sees a white-labeled dashboard controlled by the agency. |
Workspace Switcher
The Workspace Switcher appears at the top of the sidebar and lets you:
- View all workspaces you have access to in a scrollable dropdown list.
- Switch between workspaces by clicking on one — the URL updates to
/w/:workspaceId/...and all data reloads. - Create a new workspace via the "+ New workspace" button at the bottom of the dropdown. This opens a modal with the same fields as the initial setup: Workspace Name (required), Website Link, Industry, and Timezone.
Each workspace gets a unique gradient avatar (one of 10 curated gradients like "Aurora Blue", "Emerald Glow", "Midnight AI", etc.) derived from its ID hash. If you upload a custom logo, it replaces the gradient.
URL Structure
Every workspace-scoped route is prefixed with the workspace ID:
/w/{workspaceId}/agents
/w/{workspaceId}/agents/{agentId}
/w/{workspaceId}/knowledge
/w/{workspaceId}/settings/billing
/w/{workspaceId}/agency/overview
The active workspace ID is persisted in localStorage under the key wirevox_active_workspace, so it survives page refreshes and new tabs.
Workspace Resolution Priority
When you load the app, the system determines which workspace to activate using this priority:
- URL parameter — If the URL contains
/w/:workspaceId, that workspace is used (if the user has access). - localStorage — Falls back to the last-used workspace stored in
wirevox_active_workspace. - First workspace — If neither is available, the user's first workspace is selected, with a preference for subaccount → agency → personal.
Workspace Settings
Each workspace has its own settings accessible at /w/:id/settings:
| Setting Page | Description |
|---|---|
| Preferences | Workspace name, timezone, language |
| Integrations | OAuth connections (Google Calendar, Outlook, Clio, etc.) |
| Billing | Subscription plan, credit balance, top-ups, invoices |
| Team | Invite members, manage roles |
| Usage | Credit usage breakdown, call analytics |
| Notifications | Email notification preferences |
| Danger Zone | Delete workspace |
2.2 Agents
An agent is an AI-powered voice assistant that answers phone calls on your behalf. Each agent is a fully self-contained configuration: its own personality (system prompt), voice, AI model, phone number, knowledge, and tools.
Agent Types
Wirevox supports three agent types, though only Inbound is currently available:
| Type | Status | Direction | Description |
|---|---|---|---|
| Inbound | ✅ Live | Receives calls | Answers incoming calls to a phone number. Use cases: receptionist, support, scheduling, intake, FAQ. |
| Outbound | 🔜 Coming Soon | Makes calls | Proactively dials a list of contacts for sales, lead qualification, appointment reminders. |
| AI Chat | 🔜 Coming Soon | Text-based | Resolves visitor queries via a website chat widget. |
Agent Configuration
Every agent has a comprehensive settings page with the following configurable properties:
| Property | Description | Default |
|---|---|---|
| Name | Display name shown in the dashboard | From template |
| System Prompt | The AI's instructions, personality, and behavior rules | From template |
| AI Model | The LLM that powers the agent's intelligence | gpt-4o-mini |
| Pipeline Mode | modular (STT→LLM→TTS) or realtime (speech-to-speech) |
modular |
| Voice | The TTS voice used to speak to callers | aura-asteria-en (Deepgram) |
| STT Provider | Speech-to-text engine: auto, nova-3, flux, or gladia |
auto |
| Primary Language | The agent's primary spoken language | en (English) |
| Secondary Language | Optional bilingual mode | None |
| Greeting Message | A custom first sentence spoken when the call connects | None (AI generates naturally) |
| Max Call Duration | Hard time limit for calls (1–30 min, system max 40 min) | 40 min (system default) |
| Interruption Sensitivity | Number of words required to trigger barge-in | 3 words |
| Silence Timeout | Seconds of silence before the AI prompts "Are you still there?" | 15 seconds |
| Hangup Timeout | Seconds of total silence (no user speech) before auto-hangup | 30 seconds |
| Patience Level | How long the AI waits after the user stops speaking before responding: low (600ms), medium (2.6s), high (4.6s) |
low |
| Knowledge Base | Linked knowledge sources for RAG-powered answers | None |
| Functions | Attached tools (SMS, transfer, booking, webhooks) | None |
| Phone Number | The Telnyx number assigned to receive calls | None (must be purchased) |
Agent Lifecycle
- Created — Agent is generated from a template (or blank). It exists in a draft state.
- Configured — User customizes the prompt, voice, model, and attaches functions/knowledge.
- Published — Agent configuration is "locked" as a published version. Version history is maintained so you can roll back.
- Live — A phone number is assigned. Calls to that number are routed to this agent in real-time.
Dynamic Prompt Compilation
Before every call, Wirevox's compileSystemPrompt function transforms the agent's raw system prompt into a fully contextualized prompt by:
- Injecting custom variables — Replacing
{{business_name}},{{business_hours}}, and any user-defined{{variable_name}}placeholders with their configured values. - Injecting real-time context — Prepending today's date, current local time (in the agent's timezone), and the caller's phone number (formatted as
XXX-XXX-XXXX). - Injecting knowledge base instructions — If a knowledge base is linked, its description/instructions are appended.
- Overriding unavailable tools — If the prompt references a function (via
[[Function: tool_name]]) that is not connected, a critical system override is injected telling the AI to apologize and offer to take details instead of hallucinating the action.
Example of the injected real-time context block:
[REAL-TIME CONTEXT]
TODAY'S DATE: Wednesday, July 16, 2026 (2026-07-16)
CURRENT LOCAL TIME: 12:35 AM
Use this date and time as the reference for all relative dates and business hours reasoning.
CALLER PHONE (ALREADY KNOWN): 416-555-1234. This is a confirmed fact — do NOT ask for it...
--------------------------------
2.3 The Voice Pipeline (Modular Mode)
The Voice Pipeline is the real-time audio processing engine that powers every phone call. In its default modular mode, it chains together four independent services:
┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐
│ Phone │────▶│ STT │────▶│ LLM │────▶│ TTS │
│ Audio │ │ (Speech │ │ (AI │ │ (Text │
│ (Telnyx) │◀────│ to Text)│ │ Brain) │ │ to │
│ │ │ │ │ │ │ Speech) │
└──────────┘ └──────────┘ └──────────┘ └──────────┘
▲ │
└───────────────────────────────────────────────────┘
Audio sent back to caller
How a Call Flows
- Telnyx receives an inbound call and opens a WebSocket media stream to the Wirevox backend at
wss://.../media-stream/:agent/:call. - Raw μ-law audio (8kHz, base64-encoded) flows from the phone into the server.
- The STT provider (Deepgram or Gladia) receives the audio stream via its own WebSocket and returns real-time transcripts (both interim and final).
- The transcript is accumulated, filtered through barge-in logic (see below), and when the user finishes speaking, the text is flushed to the LLM.
- The LLM (OpenAI, Gemini, or Claude) generates a streaming response, sentence by sentence.
- Each sentence is sent to the TTS provider (Deepgram, OpenAI, or ElevenLabs), which synthesizes speech audio.
- The TTS audio is streamed back through the Telnyx WebSocket to the caller's phone in real-time.
Supported Providers
Speech-to-Text (STT)
| Provider | Model | Use Case | Credits/min |
|---|---|---|---|
| Deepgram | Nova-3 (monolingual) | Single-language agents, lowest latency | 2 |
| Deepgram | Nova-3 Flux (multilingual) | Bilingual agents, automatic language detection | 2 |
| Gladia | Default | Alternative STT with different accuracy characteristics | 3 |
The auto STT setting selects Flux for bilingual agents and Nova-3 for monolingual agents.
Large Language Models (LLM)
| Provider | Models | Credits/min |
|---|---|---|
| OpenAI | GPT-4o Mini, GPT-4o, GPT-4.1 (Mini/Nano), GPT-5 (Mini/Nano), GPT-5.1, GPT-5.2, GPT-5.4 (Mini/Nano) | 1–5 |
| Anthropic | Claude 3.5 Sonnet, Claude 4.5 Sonnet, Claude 4.6 Sonnet | 3–5 |
| Gemini 1.5 Flash, Gemini 2.0 Flash, Gemini 2.5 Flash | 1 |
Text-to-Speech (TTS)
| Provider | Voices | Credits/min |
|---|---|---|
| Deepgram | Aura-2 voices (Asteria, Orion, Luna, Stella, Athena, Hera, etc.) | 2 |
| OpenAI | Alloy, Echo, Fable, Onyx, Nova, Shimmer | 2 |
| ElevenLabs | Premium ultra-realistic voices (Rachel, Drew, etc.) | 10 |
TTS provider is auto-detected from the selected voice ID — Wirevox inspects the voice name and routes to the correct provider.
Barge-In (Interruption Detection)
Barge-in is how Wirevox detects when a caller interrupts the AI mid-sentence. The system uses a sophisticated multi-layered approach:
Word count threshold — By default, the caller must speak 3 or more words (configurable via
interruptionWords) for an interruption to trigger. This prevents cutting off the AI on single-word reactions.Backchannel filtering — Short conversational reactions ("yeah", "ok", "mhm", "sure", "right", "thanks") are classified as backchannels and are ignored — they don't interrupt the AI. The system maintains separate backchannel word lists for English and French.
Filler word grace — If the first word of a transcript is a known filler (e.g., "um", "uh"), the interruption threshold is increased by +1 word. This prevents cutting off on "um, yes" but still triggers on "um, actually I have a question."
Confidence scaling — The STT confidence score is checked against dynamic thresholds:
- At exactly the word limit: requires 75% confidence
- Above the word limit: requires 60% confidence (length implies intent)
- Below the word limit: requires 85% confidence
Echo cancellation — A fuzzy scoring algorithm (Jaccard similarity + subsequence matching) compares the caller's transcript against the AI's recent words. If the score exceeds 0.55 and the transcript is longer than 5 characters, it's classified as echo/feedback and suppressed.
When a valid barge-in is detected, the pipeline:
- Immediately aborts the current LLM generation
- Stops all queued TTS synthesis
- Clears the audio buffer
- Marks the AI as no longer speaking
Silence & Hangup Timers
| Timer | Default | Behavior |
|---|---|---|
| Silence Timer | 15 seconds | After 15s of silence (no user speech, AI not speaking), the AI prompts: "Are you still there?" This fires once per silence period. |
| Hangup Timer | 30 seconds | Checked every 5 seconds. If the user hasn't spoken for 30s and the AI isn't speaking, the call is automatically hung up. |
| Max Duration Timer | 40 minutes | Hard cutoff. When reached, the call is severed regardless of state. Users can configure a lower limit (1–30 minutes). |
Patience Levels
The patience level controls how long the AI waits after the user stops speaking before it begins generating a response. This handles cases where the user is thinking or pausing mid-sentence.
| Level | Delay | Best For |
|---|---|---|
low |
600ms | Fast-paced conversations (default) |
medium |
2,600ms | Users who pause while thinking |
high |
4,600ms | Complex questions where callers take time to formulate |
During the patience delay, RAG knowledge base search runs in parallel so it's ready when the LLM starts.
Bilingual Mode
When an agent has a secondary language configured, the pipeline activates bilingual mode:
- The system prompt is augmented with an operational directive telling the AI to start in the primary language but seamlessly switch if the caller speaks the secondary language.
- The AI is instructed to prefix every response with a 2-letter language tag:
[EN]for English,[FR]for French, etc. - The pipeline parses these tags and routes audio to the correct TTS voice — using
primaryVoiceIdfor the primary language andsecondaryVoiceIdfor the secondary. - STT is set to Deepgram Flux (multilingual) or Gladia for automatic language detection.
- The backchannel word list is extended with language-specific words (e.g., French: "oui", "d'accord", "merci", "parfait").
TTS Processing
Before sending text to the TTS engine, the pipeline performs critical cleanup:
- Strips markdown formatting:
**bold**,*italic*, numbered lists (1.), bullet dashes, headings (###), and inline code backticks - Strips bracketed narrations like
[Checking availability...]but preserves bilingual language tags like[EN] - Collapses multiple spaces left by stripping
Knowledge Base Integration (RAG)
During a call, when the user says something substantive (≥2 meaningful words after removing conversational filler), the pipeline:
- Cleans the query — Strips conversational filler ("hi, can you tell me about...") while protecting compound phrases ("right now", "how much").
- Searches the vector database — Queries
pgvectorfor the top 3 chunks above the similarity threshold (default: 0.25). - Stores in a sliding buffer — Results are kept in a FIFO buffer of size 3, so follow-up questions can reference previous context.
- Injects into the LLM — When
runLLM()fires, the KB buffer contents are appended to the system prompt as a[KNOWLEDGE BASE]block. - Circuit breaker — If the KB search takes longer than 1,500ms, it's abandoned and the LLM proceeds without context to prevent dead air.
2.4 Realtime Models (Speech-to-Speech)
In addition to the modular pipeline, Wirevox supports Realtime Models — a fundamentally different architecture where a single AI model handles speech input and speech output directly, without separate STT/LLM/TTS stages.
┌──────────┐ ┌───────────────────┐ ┌──────────┐
│ Phone │─────────▶│ Realtime Model │─────────▶│ Phone │
│ Audio │ │ (Speech-in, │ │ Audio │
│ (Telnyx) │◀─────────│ Speech-out) │◀─────────│ (Telnyx) │
└──────────┘ └───────────────────┘ └──────────┘
Supported Realtime Models
| Model | Provider | Audio Format | Credit Cost/min |
|---|---|---|---|
| GPT Realtime | OpenAI | PCM 24kHz | 20 (platform) + 20 (model) = 24 total |
| Gemini Live | PCM 16kHz | 10 (platform) + 10 (model) = 14 total |
How Realtime Mode Works
- The
RealtimePipelineopens a persistent WebSocket directly to the model provider (OpenAI or Google). - Inbound phone audio (μ-law 8kHz from Telnyx) is converted to PCM and resampled to the model's native rate (24kHz for OpenAI, 16kHz for Gemini).
- Audio is streamed continuously to the model. The model performs its own voice activity detection (VAD) — no separate STT is needed.
- The model generates speech audio responses directly, which are resampled back to 8kHz μ-law and streamed to the caller.
Key Differences vs. Modular Pipeline
| Aspect | Modular Pipeline | Realtime Pipeline |
|---|---|---|
| Architecture | STT → LLM → TTS (3 services) | Single model (1 service) |
| Latency | ~400–600ms (sum of 3 hops) | ~200–300ms (single hop) |
| Voice selection | 50+ voices across 3 TTS providers | Limited to model's built-in voices |
| VAD / Barge-in | Custom multi-layer barge-in logic | Model's native VAD (server-side) |
| Tool calling | Full support | Full support |
| Knowledge Base | Sliding buffer with query cleaning | Tool-based (search_knowledge_base function) |
| Cost | Variable (4 + STT + LLM + TTS credits) | Fixed (platform + model fee) |
| Bilingual | Full support with voice switching | Limited (model-dependent) |
Realtime Pipeline — System Prompt
The realtime pipeline appends a critical directive to the system prompt:
"You are running in a realtime voice environment. If you need to use a tool to look up information, you MUST immediately say a short filler phrase (like 'Let me check that for you' or 'One moment please') BEFORE calling the tool, so the user is not met with dead air."
This is necessary because in realtime mode, there's no separate TTS pipeline to generate filler — the model must produce it natively.
OpenAI Realtime Configuration
The OpenAI adapter connects to wss://api.openai.com/v1/realtime and configures:
- Input transcription: Whisper-1 model
- Turn detection: Server-side VAD with threshold 0.5, prefix padding 300ms, silence duration 800ms
- Output modality: Audio only
- Tool choice: Auto
Gemini Live Configuration
The Gemini adapter connects to wss://generativelanguage.googleapis.com/ws/... using the gemini-2.5-flash-native-audio-latest model with:
- Response modality: Audio
- Voice: Aoede (prebuilt)
- Full tool support via
functionDeclarations
2.5 Credits System
Credits are the universal currency of the Wirevox platform. Every billable event — call minutes, SMS messages, function executions — is measured in credits. Credits are abstract units; their real-world dollar value varies by subscription plan.
Credit Pricing Per Plan
| Plan | Cost Per Credit | Example: 10 credits |
|---|---|---|
| PAYG (Pay-As-You-Go) | $0.020 | $0.20 |
| Starter | $0.015 | $0.15 |
| Growth | $0.0125 | $0.125 |
| Agency | $0.008 | $0.08 |
Two Credit Pools
Every workspace has two independent credit pools:
| Pool | Source | Reset Behavior |
|---|---|---|
| Included Credits | Monthly subscription | Reset to plan allowance every billing cycle |
| Purchased Credits | One-time top-ups via Stripe | Never expire. Never reset. Survive plan changes and cancellations. |
Deduction Waterfall
When credits are consumed, they are deducted in this strict order:
1. Included Credits ──▶ Deduct first
2. Purchased Credits ──▶ Deduct only if included credits are exhausted
3. Both exhausted ──▶ Call continues, billing logs a warning
Special Case: Unlimited Tier
If included_credits = -1, the workspace is on an unlimited plan. The billing RPC skips all balance checks and returns remaining: -1.
Cost-Per-Minute Formula
Every agent has a different cost per minute based on its AI stack:
Cost Per Minute = Platform Base + STT + LLM + TTS
For realtime pipeline agents, the formula simplifies to:
Cost Per Minute = Platform Base + Realtime Model Fee
Example Calculations
| Agent Stack | Platform | STT | LLM | TTS | Total |
|---|---|---|---|---|---|
| Deepgram STT + GPT-4o Mini + Deepgram TTS | 4 | 2 | 1 | 2 | 9 cr/min |
| Deepgram STT + GPT-4o + ElevenLabs TTS | 4 | 2 | 3 | 10 | 19 cr/min |
| Gladia STT + Claude 4.6 Sonnet + ElevenLabs TTS | 4 | 3 | 5 | 10 | 22 cr/min |
| OpenAI Realtime (speech-to-speech) | 4 | — | 20 | — | 24 cr/min |
| Gemini Live (speech-to-speech) | 4 | — | 10 | — | 14 cr/min |
Per-Second Carryover Billing
Wirevox bills per-minute but tracks duration per-second with agent-level carryover. This prevents charging a full minute for a 5-second call.
How it works:
total_seconds = agent.carryover_seconds + new_call_duration
minutes_to_bill = floor(total_seconds / 60)
new_carryover = total_seconds % 60
Walk-through (agent at 9 cr/min):
| Call | Duration | Carryover Before | Total | Billed Minutes | Credits | Carryover After |
|---|---|---|---|---|---|---|
| Call 1 | 13s | 0 | 13 | 0 | 0 | 13 |
| Call 2 | 15s | 13 | 28 | 0 | 0 | 28 |
| Call 3 | 28s | 28 | 56 | 0 | 0 | 56 |
| Call 4 | 22s | 56 | 78 | 1 | 9 | 18 |
Note: Carryover seconds are not rate-stamped. If you change an agent's AI model between calls, the leftover seconds are billed at the new rate. Maximum discrepancy: 59 seconds × rate delta.
Event Costs (Flat Fees)
These are charged per-use, not per-minute:
| Event | Credits | Description |
|---|---|---|
send_sms_confirmation |
2 | Sending an SMS during or after a call |
transfer_call |
20 | Transferring the caller to a human via PSTN |
| Inbound SMS | 2 | Receiving an SMS on a Wirevox number |
| Outbound SMS | 2 | Sending an SMS from a Wirevox number |
| All other tools | 0 | Free (booking, webhooks, end call, etc.) |
Dynamic Pricing
All credit costs are loaded from the Supabase credit_costs table with a 5-minute cache. A Supabase Realtime channel (cost-registry-changes) listens for any changes to the table and instantly invalidates the cache — so pricing updates propagate to all running servers without a deploy.
If the database is unreachable, hardcoded fallback defaults are used.
2.6 Multi-Agent Architecture
Wirevox supports architectures where multiple specialized agents work together to handle different types of caller needs.
Agent-to-Agent Transfer (Cold Transfer via PSTN)
Currently, Wirevox supports Cold Transfers over PSTN, which is the industry standard:
- The caller asks to speak to a specialist (e.g., "Can I talk to your tech support?").
- The AI agent triggers the
transfer_callfunction. - The AI says goodbye and disconnects its audio streams (STT, LLM, TTS all shut down instantly — $0 AI cost after transfer).
- The platform dials the target phone number via Telnyx, bridging the original caller to the new destination.
- Cost: 20 credits for the transfer + telephony charges for the outbound leg.
Transfer Types (Current & Future)
| Transfer Type | Status | Description | AI Cost After Transfer |
|---|---|---|---|
| Cold Transfer (PSTN) | ✅ Supported | AI blindly dials a 10-digit phone number | $0 — AI shuts down |
| Warm Transfer (PSTN) | ❌ Not yet | AI puts caller on hold, briefs the human, then bridges | High — AI stays active until drop-off |
| Cold Transfer (SIP) | ❌ Not yet | AI transfers over internet (VoIP) | $0 — AI shuts down |
| Warm Transfer (SIP) | ❌ Not yet | AI briefs human over SIP before bridging | Low telephony, high AI |
| Stateful AI Handoff | ❌ Not yet | System silently swaps the system prompt from "Receptionist" to "Tech Support" mid-call | Same as a normal call — no extra cost |
Stateful Handoff (The Modern Way)
The platform has infrastructure for a Stateful AI Handoff — instead of transferring the phone call, the system swaps the AI's system prompt mid-call. The HANDOFF_CACHE (a 3-minute TTL in-memory cache) stores context state for the transition:
- The caller stays on the same phone connection
- No extra telephony legs
- No extra AI model fees
- The new "agent" inherits the full conversation history
This is the most cost-effective multi-agent approach and avoids the "triple telephony charge" problem of AI-to-AI PSTN transfers.
Designing a Multi-Agent System
A typical multi-agent setup might look like:
┌─────────────────────┐
Inbound Call ──▶│ Receptionist Agent │
│ (General Intake) │
└────────┬────────────┘
│
┌──────────────┼──────────────┐
▼ ▼ ▼
┌────────────┐ ┌────────────┐ ┌────────────┐
│ Sales │ │ Support │ │ Scheduling │
│ Agent │ │ Agent │ │ Agent │
└────────────┘ └────────────┘ └────────────┘
Each agent has its own:
- System prompt tailored to its specialization
- Functions — Sales has CRM webhooks, Support has ticketing, Scheduling has calendar integrations
- Knowledge base — Each can reference different document sets
- Phone number — Or share routing via the receptionist's transfer function
Concurrent Call Limits
Each subscription plan includes a base number of concurrent call slots — the maximum number of simultaneous phone calls the workspace can handle:
max_concurrent = plan.base_concurrent_calls + workspace.extra_concurrent_calls
Additional slots can be purchased as Stripe add-ons:
| Plan | Add-On Price |
|---|---|
| Starter | ~$15/mo per extra slot |
| Growth | ~$13/mo per extra slot |
| Agency | ~$5/mo per extra slot |
When a plan upgrade occurs, existing concurrent add-on subscriptions are automatically migrated to the new tier's pricing.
Previous: Section 1 — Getting Started Next: Section 3 — Agent Configuration
Section 3 — Agent Configuration
A comprehensive reference for every setting available when configuring a Wirevox AI agent, from system prompts to voice selection, model tuning, and publishing.
3.1 System Prompt
The System Prompt is the most important part of your agent configuration. It defines who your agent is, how it behaves, what it knows, and what it can do. Think of it as writing detailed instructions for a new hire — the more specific and well-structured your prompt, the better your agent performs.
The Prompt Editor
The agent settings page features a full-screen Prompt Editor — a syntax-highlighted textarea with two special features:
- Variable highlighting — Any
{{variable_name}}in your prompt is rendered in blue with a blue background. These are template variables that get replaced with real values at call time. - Function highlighting — Any
[[Function: action_name]]reference is rendered in green with a green background. These tell the AI when and how to use attached tools.
Insert Toolbar
Above the editor, an Insert toolbar provides quick insertion of:
{{variable}}— Opens a dropdown of your defined custom variables. Click one to insert it at the cursor position.[[Function: action]]— Opens a dropdown of all functions attached to the agent. Click one to insert a function reference.
Autocomplete
As you type [[ in the prompt editor, an autocomplete popup appears showing matching function names. Use arrow keys to navigate and Enter to select, or Escape to dismiss. This works like IDE autocompletion — it filters as you type after the [[Function: prefix.
Prompt Warnings
The editor performs live validation and shows warnings for:
| Warning Type | Condition | Meaning |
|---|---|---|
| Undefined variable | {{name}} used in prompt but no variable named name is defined |
The placeholder won't be replaced — it'll appear literally in the AI's instructions |
| Missing function | [[Function: tool]] referenced but tool is not attached to this agent |
At call time, the AI will receive a critical override telling it this tool is unavailable |
| Disconnected integration | Function references a tool from a disconnected integration (e.g., Google Calendar OAuth expired) | The tool exists but can't execute — the AI will apologize and offer alternatives |
Warnings appear as an amber badge in the editor toolbar. Click it to see details.
Writing Effective Prompts
Best practices for phone AI prompts:
- Write conversationally — The AI speaks on a phone call. Write instructions the way you'd brief a human receptionist, not the way you'd write an essay.
- Define the persona — Give the agent a name, a tone (friendly, professional, energetic), and a role ("You are Sarah, the front desk receptionist at Maple Dental").
- Set explicit boundaries — What the agent should NOT do is as important as what it should do. Example: "You must NEVER provide medical advice. Always direct health questions to the dentist."
- Structure with sections — Use clear headings or labeled sections:
[IDENTITY],[BUSINESS HOURS],[BOOKING RULES],[ESCALATION POLICY]. - Reference functions explicitly — Tell the AI when to use tools: "When a caller wants to book an appointment, use [[Function: check_availability]] to find open slots."
- Use variables for data — Don't hardcode business hours, addresses, or policies into the prompt. Use
{{business_hours}},{{address}}, etc. so you can change them without editing the prompt.
How Variables Are Compiled
At call time, the compileSystemPrompt function processes the raw prompt in this order:
1. Custom Variables: {{business_hours}} → "Mon-Fri 9am-5pm"
2. Legacy Variables: {{business_name}} → "SunPath Energy"
3. Real-Time Context: Prepend date, time, caller phone number
4. Knowledge Base: Append KB instructions (if linked)
5. Tool Overrides: Replace [[Function: X]] with error message (if X is disconnected)
Variable Naming Rules
Variable names must contain only:
- Letters (a–z, A–Z)
- Numbers (0–9)
- Underscores (
_)
Any other characters are stripped when you type a variable name in the Variables panel.
3.2 AI Model Selection
The AI model is the "brain" of your agent — it processes the caller's words and generates intelligent responses. Wirevox supports models from three providers.
Pipeline Mode Toggle
Before selecting a model, choose your Pipeline Mode using the toggle at the top of the configuration sidebar:
| Mode | Description |
|---|---|
| Modular (default) | Separate STT → LLM → TTS pipeline. Maximum voice selection, full bilingual support, fine-grained control. |
| Realtime | Single speech-to-speech model. Lower latency, limited voice options, simpler architecture. |
The model dropdown changes based on your selected pipeline mode.
Modular Mode — LLM Options
| Model | Provider | Credits/min | Best For |
|---|---|---|---|
| GPT-4o Mini ⭐ | OpenAI | 1 | Default choice. Fast, cost-effective, great for most use cases. |
| GPT-4o | OpenAI | 3 | Better reasoning for complex conversations. |
| GPT-4.1 | OpenAI | 3 | Latest GPT-4 series with improved instruction following. |
| GPT-4.1 Mini | OpenAI | 1 | Lighter variant of GPT-4.1. |
| GPT-4.1 Nano | OpenAI | 1 | Ultra-lightweight, fastest response times. |
| GPT-5 Mini | OpenAI | 2 | Next-gen mini model with enhanced capabilities. |
| GPT-5 Nano | OpenAI | 1 | Ultra-fast next-gen nano model. |
| GPT-5.1 | OpenAI | 4 | Advanced reasoning and nuanced conversations. |
| GPT-5.2 | OpenAI | 4 | Improved GPT-5.1 with better context handling. |
| GPT-5.4 | OpenAI | 5 | Flagship model. Best intelligence, highest cost. |
| GPT-5.4 Mini | OpenAI | 2 | Balanced next-gen model. |
| GPT-5.4 Nano | OpenAI | 1 | Cost-effective next-gen option. |
| Claude 3.5 Sonnet | Anthropic | 3 | Coming Soon. Strong at nuanced, empathetic conversations. |
| Claude 4.5 Sonnet | Anthropic | 4 | Coming Soon. Enhanced reasoning and tool use. |
| Claude 4.6 Sonnet | Anthropic | 5 | Coming Soon. Anthropic's latest flagship. |
| Gemini 1.5 Flash | 1 | Ultra-fast, great for simple FAQ agents. | |
| Gemini 2.0 Flash | 1 | Improved Flash with better tool calling. | |
| Gemini 2.5 Flash | 1 | Latest Flash with enhanced multilingual support. |
Recommendation: Start with GPT-4o Mini (1 credit/min). It handles 90% of use cases well. Upgrade to GPT-4o or GPT-5.x only if your agent needs complex multi-step reasoning or very nuanced conversations.
Realtime Mode — Model Options
| Model | Provider | Credits/min | Description |
|---|---|---|---|
| OpenAI Realtime | OpenAI | 20 | Speech-to-speech via GPT Realtime API. Uses Whisper for input transcription, server-side VAD. |
| Gemini Live | 10 | Speech-to-speech via Gemini's native audio model (gemini-2.5-flash-native-audio-latest). |
When you switch between realtime models, the voice automatically adjusts — OpenAI Realtime defaults to "Alloy", Gemini Live defaults to "Puck".
Model Switching Behavior
When you change pipeline modes:
- Modular → Realtime: The model switches to the last-used realtime model (or
gpt-4o-realtime-previewdefault). The voice switches to the last-used realtime voice (oralloy). - Realtime → Modular: The model restores to the last-used modular model (or
gpt-4o-mini). The voice restores to the last-used modular voice (oraura-asteria-en).
Each mode maintains its own independent model + voice state, so switching back and forth doesn't lose your configuration.
3.3 Voice Selection
The voice determines how your agent sounds to callers. Wirevox provides a comprehensive voice catalog with voices from three providers (modular mode) and two realtime providers.
Modular Mode Voices
Deepgram Aura-2 (2 credits/min)
The default TTS provider. Low latency, natural-sounding, cost-effective.
English Voices:
| Voice ID | Name | Gender |
|---|---|---|
aura-asteria-en |
Asteria ⭐ | Female |
aura-orion-en |
Orion | Male |
aura-luna-en |
Luna | Female |
aura-perseus-en |
Perseus | Male |
aura-angus-en |
Angus | Male |
aura-2-amalthea-en |
Amalthea | Female |
aura-2-theia-en |
Theia | Female |
aura-2-vesta-en |
Vesta | Female |
aura-2-andromeda-en |
Andromeda | Female |
aura-2-zeus-en |
Zeus | Male |
aura-2-saturn-en |
Saturn | Male |
aura-2-phoebe-en |
Phoebe | Female |
aura-2-pandora-en |
Pandora | Female |
aura-2-orpheus-en |
Orpheus | Male |
aura-2-iris-en |
Iris | Female |
aura-2-harmonia-en |
Harmonia | Female |
aura-2-cordelia-en |
Cordelia | Female |
aura-2-cora-en |
Cora | Female |
aura-2-arcas-en |
Arcas | Male |
French Voices:
| Voice ID | Name | Gender |
|---|---|---|
aura-2-agathe-fr |
Agathe | Female |
aura-2-hector-fr |
Hector | Male |
OpenAI TTS (2 credits/min)
High-quality voices with consistent tone. Same cost as Deepgram.
| Voice ID | Name | Gender |
|---|---|---|
shimmer |
Shimmer | Female |
alloy |
Alloy | Neutral |
echo |
Echo | Male |
fable |
Fable | Neutral |
onyx |
Onyx | Male |
nova |
Nova | Female |
marin |
Marin | Female |
cedar |
Cedar | Male |
ElevenLabs (10 credits/min)
Premium ultra-realistic voices. 5× the cost of Deepgram/OpenAI but significantly more natural.
| Voice ID | Name | Gender |
|---|---|---|
eleven-21m00Tcm4TlvDq8ikWAM |
Rachel | Female |
eleven-pNInz6obpgDQGcFmaJgB |
Adam | Male |
eleven-EXAVITQu4vr4xnSDxMaL |
Bella | Female |
eleven-MF3mGyEYCl7XYWbV9V6O |
Elli | Female |
eleven-ErXwobaYiN019PkySvjV |
Antoni | Male |
eleven-aMSt68OGf4xUZAnLpTU8 |
Juniper | Female |
eleven-hFskf6X0TFndppvQxiEF |
Finch | Male |
eleven-g6xIsTj2HwM6VR4iXFCw |
Jessica | Female |
eleven-TxGEqnHWrfWFTfGW9XjX |
Josh | Male |
eleven-VR6AewLTigWG4xSOukaG |
Arnold | Male |
eleven-nPczCjzI2devNBz1zQrb |
Brian | Male |
Realtime Mode Voices
OpenAI Realtime Voices
| Voice ID | Name | Gender |
|---|---|---|
alloy |
Alloy ⭐ | Neutral |
echo |
Echo | Male |
shimmer |
Shimmer | Female |
verse |
Verse | Neutral |
ballad |
Ballad | Male |
coral |
Coral | Female |
sage |
Sage | Female |
ash |
Ash | Neutral |
marin |
Marin | Female |
cedar |
Cedar | Male |
Gemini Live Voices
| Voice ID | Name | Gender |
|---|---|---|
Puck |
Puck ⭐ | Neutral |
Charon |
Charon | Male |
Kore |
Kore | Female |
Fenrir |
Fenrir | Male |
Aoede |
Aoede | Female |
Voice Preview
Every voice in the catalog can be previewed before selection. Click the play button next to any voice to hear a sample audio clip loaded from /assets/voices/{voiceId}.mp3. Only one preview plays at a time — clicking a new voice stops the previous one.
Auto-Detection of TTS Provider
The pipeline auto-detects which TTS provider to use based on the voice ID:
| Voice ID Pattern | Detected Provider |
|---|---|
Starts with eleven- or known ElevenLabs names |
ElevenLabs |
One of alloy, echo, fable, onyx, nova, shimmer |
OpenAI |
Everything else (including aura-*) |
Deepgram |
You never need to manually set the TTS provider — selecting a voice is sufficient.
3.4 STT Configuration
Speech-to-Text (STT) converts the caller's spoken audio into text that the LLM can process. This setting only applies in Modular pipeline mode (Realtime models handle their own speech recognition).
Available STT Providers
| Provider | Setting Value | Credits/min | Description |
|---|---|---|---|
| Deepgram Nova-3 | nova-3 |
2 | Monolingual. Lowest latency, best accuracy for single-language agents. |
| Deepgram Nova-3 Flux | flux |
2 | Multilingual. Automatic language detection for bilingual agents. |
| Gladia | gladia |
3 | Alternative provider with different accuracy characteristics. |
| Auto | auto |
2 | Smart default — selects flux for bilingual agents, nova-3 for monolingual. |
Recommendation: Leave this on
autounless you have a specific reason to change it. Auto makes the right choice for your language configuration.
Patience Level (Endpointing)
The Patience Level controls how long the STT waits after the user stops speaking before treating the utterance as complete and sending it to the LLM. This is critical for natural conversation flow.
| Level | Delay | Description |
|---|---|---|
| Low (default) | 600ms | Snappy responses. Best for simple Q&A or fast-paced conversations. |
| Medium | 2,600ms | Gives the user time to pause mid-thought. Good for complex topics. |
| High | 4,600ms | Very patient. Best for callers who think out loud or give long answers. |
During the patience delay, the system starts the knowledge base search in parallel so context is ready the moment the LLM fires.
3.5 Language & Bilingual Mode
Supported Languages
Wirevox currently supports voices in these languages:
| Code | Language | Native Voices Available |
|---|---|---|
en |
English | 19 Deepgram + 8 OpenAI + 11 ElevenLabs |
fr |
French | 2 Deepgram (Agathe, Hector) |
Additional languages (Spanish, German, Italian, Portuguese, Dutch) are supported at the LLM level but do not yet have dedicated Deepgram voices. You can use OpenAI or ElevenLabs voices for these languages.
Enabling Bilingual Mode
- Set your Primary Language (e.g., English).
- Set your Secondary Language (e.g., French).
- The STT automatically switches to Flux (multilingual) mode.
- Choose a Primary Voice for the primary language and a Secondary Voice for the secondary language.
How Bilingual Mode Works at Runtime
When bilingual mode is active, the pipeline injects an Operational Directive into the system prompt:
[OPERATIONAL DIRECTIVE]
Primary Language: English
Secondary Language: French
Rule: You MUST initiate the conversation and formulate all internal thoughts
natively in the Primary Language. However, you are fully bilingual — if the
user speaks the Secondary Language, you must seamlessly transition and
continue the conversation in that language.
CRITICAL BILINGUAL RULE: You MUST prefix every single response with a 2-letter
language tag indicating the language you are speaking.
Use [EN] for English and [FR] for French.
Example: "[FR] Bonjour, comment puis-je vous aider?"
The TTS engine then parses these tags and routes audio to the correct voice.
Backchannel Words by Language
The barge-in system uses language-specific backchannel word lists:
English backchannels: yes, yeah, yep, ok, okay, right, sure, gotcha, alright, perfect, great, good, thanks, etc.
French backchannels (added when French is primary or secondary): oui, ouais, d'accord, entendu, exactement, absolument, parfait, super, bien, bon, voilà, merci, non, etc.
French fillers: euh, hein, bah, ben, bof, mouais, hmm, ah, oh.
3.6 Call Settings
The Call Settings section in the configuration sidebar controls how the agent handles the mechanics of phone calls.
Ring Duration
| Setting | Range | Default | Description |
|---|---|---|---|
| Ring Duration | 0–10 seconds | 0s | Simulates a phone ringing before the AI picks up. Set to 0 for instant pickup. Useful for making the experience feel more natural to callers who expect a few rings. |
Greeting Message
The greeting message is the first thing your agent says when it picks up the call.
- Custom greeting: You define exactly what the agent says. Example: "Thank you for calling Maple Dental. How can I help you today?"
- No greeting (default): The AI generates a natural opening based on its system prompt.
Custom greetings bypass the LLM entirely — they're sent directly to TTS, saving latency on the first response.
Barge-In Sensitivity (Interruption Words)
Controls how many words a caller must say to trigger an interruption when the AI is speaking.
| Value | Behavior |
|---|---|
| 1 word | Very sensitive — even "yes" interrupts the AI |
| 2 words | Moderate — short phrases interrupt |
| 3 words (default) | Balanced — natural conversation fragments trigger, single-word reactions don't |
| 5+ words | Low sensitivity — only substantial speech interrupts |
Filler Words
A customizable list of words that the pipeline treats as filler — they don't trigger the AI to respond and don't count as interruptions when the AI is speaking.
Default fillers:
yeah, uh, um, mhmm, ah, hm, like, right, mhm, uhhuh, uh-huh,
hmm, oh, umm, uhh, huh, mmm, aha, aah, ooh, mm, ugh, eh, er
You can add or remove fillers from the configuration sidebar. New fillers are entered one at a time and added to the list.
Silence Timeout
| Setting | Range | Default |
|---|---|---|
| Silence Timeout | 5,000–60,000ms | 15,000ms (15s) |
How long the agent waits in silence before prompting "Are you still there?" This fires once per silence period. After the prompt, the hangup timer continues independently.
Hangup Timeout
| Setting | Range | Default |
|---|---|---|
| Hangup Timeout | 10,000–120,000ms | 30,000ms (30s) |
Total seconds of silence (no user speech, AI not speaking) before the call is automatically disconnected. This timer is checked every 5 seconds and resets whenever the user speaks or the AI finishes speaking.
Max Call Duration
| Setting | Range | Default | System Max |
|---|---|---|---|
| Max Call Duration | 1–30 minutes | Not set (uses system default) | 40 minutes |
A hard time limit for calls. When reached, the call is severed regardless of state. This protects against runaway calls and unexpected credit consumption.
RAG Threshold
| Setting | Range | Default |
|---|---|---|
| RAG Threshold | 0.0–1.0 | 0.25 |
The minimum vector similarity score required for a knowledge base chunk to be included in the AI's context. Lower values return more (potentially less relevant) results. Higher values return only highly relevant matches.
3.7 Post-Call Processing
After every call ends, Wirevox performs several automatic post-processing steps.
Call Summary Generation
The LLM generates a structured summary of the call, including:
- Purpose — Why the caller called
- Key topics discussed — Main points of the conversation
- Actions taken — Any tools that were used (bookings, SMS, transfers)
- Outcome — How the call concluded (resolved, transferred, hung up)
- Follow-up needed — Whether human action is required
Lead Extraction
From every call, the system extracts structured lead data:
- Name — Caller's full name (if provided)
- Phone number — The caller's phone number (from caller ID or stated)
- Email — If the caller provided one
- Sentiment — Positive, neutral, or negative
- Custom fields — Any additional data captured by functions
Webhook Delivery
If post-call webhooks are configured (via custom functions), the system fires HTTP POST requests with the call data to your specified endpoints. This enables:
- CRM synchronization
- Slack/Teams notifications
- Zapier/Make.com automations
- Custom analytics pipelines
3.8 Publishing & Versioning
Wirevox uses a draft/published model for agent configuration. This means you can freely experiment with changes without affecting live callers.
How It Works
┌─────────┐ ┌─────────────┐ ┌──────────┐
│ Draft │───Publish───▶ │ Published │◀──Incoming───│ Callers │
│ (Your │ │ (What │ Calls │ │
│ edits) │◀──Discard────│ callers │ │ │
│ │ │ experience) │ │ │
└─────────┘ └─────────────┘ └──────────┘
The Draft
Every change you make in the agent settings page is saved to the draft version. Changes are auto-saved with a 2-second debounce — you'll see a save status indicator in the header:
| Status | Indicator |
|---|---|
saving |
"Saving..." text appears |
saved |
"Saved at 3:45 PM" confirmation |
error |
Error state (usually network issues) |
If you navigate away or close the browser while changes are pending, a beforeunload handler fires a final save via keepalive: true to ensure nothing is lost.
Publishing
When your draft is ready, click "Publish" in the header. This opens a Publish Modal where you can:
- Review what changed since the last publish
- Add an optional version note describing your changes
- Confirm the publish
Publishing creates a new version snapshot and makes the draft the new live configuration. All incoming calls immediately use the published version.
The header shows an amber dot with "Unpublished changes" when your draft differs from the published version.
Version History
Click "History" in the agent settings menu to open the Version History Modal. This shows:
- A chronological list of all published versions
- Each version's note (if provided) and timestamp
- The ability to restore any previous version
Restoring a version replaces your current draft with the selected historical version.
Discarding Changes
If you want to abandon your draft and revert to the last published version, click "Discard Changes" in the more menu. This restores the published version as your current draft.
Agent Status
Each agent has a status that controls whether it actively answers calls:
| Status | Behavior |
|---|---|
| Active | Agent answers incoming calls normally |
| Paused | Agent exists but doesn't answer calls. Calls to its number will ring unanswered. |
You can toggle status from the agent settings header. An agent must have a phone number assigned to be activated — attempting to activate without a number shows an error: "Assign a number to the agent to start it."
Other Agent Actions
Available from the more menu (⋮) in the agent header:
| Action | Description |
|---|---|
| Duplicate Agent | Creates a copy of the current draft as a new agent with all settings preserved. |
| Delete Agent | Permanently removes the agent, its draft, and all published versions. This action cannot be undone. |
Cost Display
The agent header shows a live cost estimate in credits per minute. Hovering over it reveals a tooltip with the full cost breakdown:
Cost Breakdown
──────────────────────
Platform 4
STT (Listening) 2
LLM (Thinking) 1
TTS (Speaking) 2
──────────────────────
Total 9 cr/min
≈ $0.135 / min
If the draft cost differs from the published cost (e.g., you changed the model), the tooltip shows both with a "(Draft)" label. The dollar estimate uses your plan's overage rate for the conversion.
3.9 Configuration Sidebar Sections
The right-side Configuration Sidebar (340px wide, collapsible) organizes all non-prompt settings into accordion sections:
Pipeline Mode (Always Visible)
A toggle between Modular and Realtime at the top of the sidebar. Changes the available options in all sections below.
General Section
- Timezone — The agent's local timezone (34 options from Adelaide to UTC). Used for real-time context injection and scheduling functions.
- Model — LLM selection (see Section 3.2). Changes based on pipeline mode.
Variables Section
- Define
{{name}}→valuepairs - Each variable has a name field (alphanumeric + underscore only) and a multi-line value textarea
- Add variables with "+ Add Variable" button
- Remove with the ✕ button
- Variables are highlighted in blue in the prompt editor
Call Settings Section
- Ring Duration (slider: 0–10s)
- Additional call mechanics settings
Voice & Language Section
- Primary Language selector
- Secondary Language selector (optional, enables bilingual mode)
- Voice Selector — visual voice picker with preview buttons
- STT Provider selector
- Patience Level selector
Knowledge Section
- Link/unlink knowledge bases via the Knowledge Selector Modal
- Shows linked KB name and document count
- RAG Threshold slider
Functions Section
Functions are managed in a separate Functions Panel (slide-out from the right), organized into three stages:
| Stage | When It Runs | Examples |
|---|---|---|
| Pre-call | Before the call connects (IVR-style routing) | Language detection, caller identification |
| During call | While the AI is in conversation | Book appointment, send SMS, transfer call, check availability, create contact |
| Post-call | After the call ends | Send email summary, sync to CRM, fire webhook |
Each function shows its name, category, and provider (if it's an integration). Functions can be added from the Action Catalog which shows all available built-in actions and custom webhooks.
Previous: Section 2 — Core Concepts Next: Section 4 — Phone Numbers & Telephony
Section 4 — Phone Numbers & Telephony
Everything about acquiring, managing, and routing phone numbers on Wirevox — from searching and purchasing through Telnyx, to the complete inbound call lifecycle, transfers, SMS, and billing.
4.1 Buying Phone Numbers
Phone numbers are the bridge between the real world and your AI agents. Wirevox provisions numbers through Telnyx, a carrier-grade telephony provider.
Prerequisites
- You must be on a paid plan (Starter, Growth, or Agency). Free-tier users cannot purchase phone numbers.
- The Phone Numbers page (
/w/:id/numbers) shows an "Unlock Telephony Features" screen for free users with a prompt to upgrade.
The Numbers Page
The Numbers page has two tabs:
| Tab | Description |
|---|---|
| My Numbers | Your purchased number inventory with agent assignments |
| Buy New Numbers | Search and purchase new numbers from Telnyx |
Searching for Numbers
On the "Buy New Numbers" tab, you can filter available numbers by:
| Filter | Options | Description |
|---|---|---|
| Country | US, GB, CA, AU, NZ, DE, FR, IE, NL | The country where the number is registered |
| Type | local, toll-free |
Local numbers have an area code; toll-free numbers (e.g., 1-800) are nationwide |
| Search Query | Area code or contains | Smart parsing: 3-digit input = area code search, 4+ digits = "contains" search |
| Max Price | Any, $2, $5, $10 | Filter by monthly Telnyx cost |
| Sort | Price ascending/descending | Order results by monthly fee |
The search auto-fires when you switch to the search tab or change country/type filters. The query input has a 500ms debounce. Results are paginated (20 per page) from up to 200 results returned by Telnyx.
Purchasing Flow
- Click "Buy" next to a number in the search results.
- A confirmation modal appears showing:
- The phone number (formatted as
+1 (XXX) XXX-XXXX) - Number type and region
- Whether it's free (within your plan's free number allowance) or $5/mo (extra number)
- The phone number (formatted as
- Click "Confirm Purchase".
- The backend:
- Verifies your subscription and free number quota
- If it's a paid number: adds a
$5/moline item to your Stripe subscription (proration_behavior: 'always_invoice'— charged immediately) - Calls
telephony.buyNumber()to provision the number on Telnyx - Saves the number to the database with status
active
- You're redirected to the "My Numbers" tab where the new number appears.
Free vs. Paid Numbers
Each plan includes a set number of free phone numbers:
| Plan | Free Numbers Included |
|---|---|
| Free | 0 |
| PAYG | 0 |
| Starter | 1 |
| Growth | 3 |
| Agency | 5 |
Numbers beyond the free allowance cost $5/month each, added as a subscription line item on Stripe.
Free Number Abuse Prevention
A safeguard prevents users from infinitely cycling free numbers:
If you release a free number (moving it to grace period) and then try to buy a new one, the new number is treated as paid ($5/mo) even if you're under the free limit. You must either reclaim the graced number or pay for a new one.
This prevents the exploit of releasing a free number, buying a new free one, and repeating to effectively get unlimited numbers.
4.2 Assigning Numbers to Agents
Once purchased, numbers need to be assigned to agents to start routing calls.
Assignment Rules
- One-to-one mapping — Each number can be assigned to exactly one agent, and each agent can have at most one number.
- Published agents only — You cannot assign a number to an agent that has never been published. The system validates
publishedVersionIdexists. - Same workspace — The agent must belong to the same workspace as the number.
Assignment Flow
- On the "My Numbers" tab, each number shows an agent dropdown.
- Select an agent from the list. Only unpublished or already-assigned agents are excluded.
- If the number is currently assigned to a different active agent, a confirmation modal warns you:
"This number is currently routing calls to [Agent Name]. Reassigning will stop that agent."
- Confirming the reassignment:
- Pauses the previous agent (sets status to
paused) - Assigns the number to the new agent
- Pauses the previous agent (sets status to
Unassigning
Select "Unassigned" in the agent dropdown to remove the agent mapping. The number remains in your inventory but doesn't route calls.
4.3 Inbound Call Flow
When a caller dials your Wirevox number, a precise sequence of events orchestrates the entire call:
Step-by-Step Lifecycle
┌──────────────┐ ┌──────────┐ ┌──────────────┐
│ Caller │──1───▶ │ Telnyx │──2───▶ │ Wirevox │
│ dials │ │ (PSTN) │ │ Backend │
│ your # │ │ │ │ │
└──────────────┘ └──────────┘ └──────┬───────┘
│
3. Answer call + open │
WebSocket stream │
▼
┌──────────────┐
│ Media │
│ Stream │
│ Handler │
└──────┬───────┘
│
4. Load published │
agent config │
5. Compile prompt │
6. Assemble tools │
7. Create pipeline │
▼
┌──────────────┐
│ Pipeline │
│ (Voice or │
│ Realtime) │
└──────┬───────┘
│
8. STT → LLM → TTS │
9. Audio back to │
caller │
▼
┌──────────────┐
│ Call Ends │
│ (hangup / │
│ timeout) │
└──────┬───────┘
│
10. Save transcript │
11. Post-call │
analysis │
12. Fire webhooks │
Detailed Breakdown
1. Telnyx Webhook — Telnyx receives the call and sends a webhook to the Wirevox backend. The webhook payload contains the caller's number, the dialed number, and a call_control_id for managing the call.
2. Agent Resolution — The backend looks up which agent is assigned to the dialed number.
3. Answer + Media Stream — The backend answers the call via the Telnyx API and opens a WebSocket media stream at:
wss://{host}/media-stream/{agentId}/{callId}
4. Load Published Config — The handler loads the agent's published version (not the draft). This ensures that edits you're making in the dashboard don't affect live calls.
// Real calls MUST run on the published snapshot, not the live draft.
if (agent && agent.publishedVersionId) {
const pubVer = await db.getAgentVersion(agent.publishedVersionId);
agent.agentConfig = pubVer.config;
}
5. Compile System Prompt — The raw prompt is processed through compileSystemPrompt() with:
- Custom variables replaced
- Real-time context injected (date, time, caller phone)
- Knowledge base instructions appended
- Unavailable tool overrides injected
6. Assemble & Filter Tools — All attached functions are assembled into OpenAI tool schemas. Tools whose required integrations are disconnected are removed. A strict prompt-driven enforcement filter ensures only tools referenced in the prompt (via [[Function: name]]) are exposed, plus always-enabled system tools (end_call, search_knowledge, system_infolist).
7. Create Pipeline — Based on pipelineMode, either a VoicePipeline (modular) or RealtimePipeline (realtime) is instantiated.
8. Audio Processing — Telnyx sends μ-law 8kHz audio frames as base64. The pipeline processes them through STT → LLM → TTS.
9. Audio Response — TTS audio is sent back to Telnyx via the WebSocket as base64 media events.
10–12. Call Ends — When the WebSocket closes:
- The transcript record is saved to the database
runPostCallAnalysis()kicks off background processing (summary, lead extraction)- Post-call webhooks fire if configured
Barge-In Audio Flushing
When the user interrupts (barge-in), the pipeline sends a clear event to Telnyx:
{ "event": "clear" }
This flushes Telnyx's audio playback buffer. Without this, audio already sent to Telnyx continues playing for 1–3 seconds after interruption, making the AI seem unresponsive.
Welcome Greeting Timing
After the media stream connects and STT is ready, there's a 300ms stabilization delay before the greeting plays. This ensures the bidirectional audio socket is fully established — without it, the first 100–200ms of the greeting gets clipped.
pipeline.welcomeTimer = setTimeout(() => {
pipeline.triggerWelcome(defaultGreeting);
}, 300);
If the user speaks before the greeting fires, the cancelGeneration() method clears this timer.
4.4 Outbound Calls
🔜 Coming Soon — Outbound calling is under development.
Outbound agents will support:
- Campaign creation — Define a list of contacts to call
- CSV upload — Import phone numbers and names in bulk
- Concurrent dispatch — Call multiple leads simultaneously
- Answering Machine Detection (AMD) — Detect voicemail and handle appropriately
- Call dispositions — Track outcomes (answered, no answer, voicemail, busy)
The backend already has partial support for outbound agents in the media stream handler — outbound calls inject the lead's name into the greeting:
"Hi {leadName}, this is {agentName}. Do you have a quick minute?"
4.5 Call Transfers
Call transfers let your AI agent connect the caller to a human when needed — for escalation, specialist routing, or emergency situations.
How Cold Transfer Works
- The AI decides to transfer (based on the caller's request or prompt instructions).
- The AI calls the
transfer_callfunction with a target phone number. - The tool handler:
- Returns
{ disconnect: true }to the pipeline - The pipeline immediately shuts down all AI services (STT, LLM, TTS stop — $0 AI cost from this point)
- The WebSocket closes
- Returns
- Before disconnecting, the tool initiates a Telnyx transfer — bridging the original caller to the target number via PSTN.
- The caller hears ringing and connects to the human.
Transfer Cost
| Component | Credits |
|---|---|
| Transfer function execution | 20 credits (flat) |
| Telnyx telephony (outbound leg) | Standard Telnyx rates |
| AI cost after transfer | $0 — pipeline shut down |
Important Constraint
- Transfer targets must be 10-digit PSTN phone numbers (or E.164 format).
- The AI cannot transfer to a SIP endpoint, another AI agent (via the platform), or an internal extension — only real phone numbers.
4.6 SMS
Wirevox supports both sending and receiving SMS on your phone numbers.
Sending SMS (During Call)
The send_sms_confirmation function lets your AI agent send a text message to the caller during a conversation. Common use cases:
- Sending appointment confirmation details
- Sharing a link or address
- Confirming a booking reference number
Cost: 2 credits per outbound SMS.
Inbound SMS
When someone texts your Wirevox number:
- The SMS is received via a Telnyx webhook
- Cost: 2 credits per inbound SMS
SMS Requirements
- The phone number must be SMS-enabled (most Telnyx numbers are)
- The workspace must have sufficient credits
4.7 Number Billing
Free Numbers
Each plan includes free numbers (see 4.1). Free numbers have isFree: true and monthlyCost: 0 in the database.
Extra Numbers ($5/mo)
Numbers beyond the free quota cost $5/month each. This is implemented as a Stripe subscription line item:
- When you buy an extra number, a
quantity: 1item is added (or quantity incremented) on theSTRIPE_PRODUCT_NUMBER_{TIER}product proration_behavior: 'always_invoice'means you're charged immediately for the remainder of the billing period- When you release an extra number, the quantity is decremented with
proration_behavior: 'none'— no refund for the current period, charges simply stop next cycle
Tier-Specific Products
Each plan tier has its own Stripe product for extra numbers:
| Env Variable | Plan |
|---|---|
STRIPE_PRODUCT_NUMBER_STARTER |
Starter |
STRIPE_PRODUCT_NUMBER_GROWTH |
Growth |
STRIPE_PRODUCT_NUMBER_AGENCY |
Agency |
When a user changes plans, the existing extra number subscription items are migrated to the new tier's product.
Number Quota API
The /api/numbers/quota endpoint returns:
{
"tier": "starter",
"freeLimit": 1,
"freeUsed": 1,
"extraCount": 2,
"extraMonthlyCost": 10
}
This powers the UI quota display showing "1/1 free numbers used, 2 extra ($10/mo)".
4.8 Grace Period (30-Day Number Reservation)
When you release a number, it doesn't get permanently deleted immediately. Instead, it enters a 30-day grace period.
How Grace Works
Active Number ──Release──▶ Grace Period (30 days) ──Expires──▶ Permanently Released
│ (Telnyx deletes)
│
Reclaim ──▶ Active Number (reactivated)
During Grace Period
- The number is reserved on Telnyx (no one else can buy it)
- It does not route calls — callers will get no answer
- It is unassigned from any agent
- Stripe billing for extra numbers is stopped (no charges during grace)
- A blue/amber banner appears on the Numbers page showing grace numbers with:
- Days remaining (calculated as
ceil((graceUntil - now) / 86400000)) - "↩ Reclaim" button — reactivates the number
- "Release Now" button — permanently releases immediately
- Days remaining (calculated as
Reclaiming
Click "Reclaim" to reactivate a graced number:
- Requires an active subscription
- If the number was free and the free quota isn't full, it's reclaimed as free
- If it's a paid reclaim, a new Stripe line item is added
- Status changes back to
active,graceUntilis cleared
Permanent Release
Click "Release Now" (or wait 30 days for automatic expiry):
telephony.releaseNumber()removes the number from Telnyxdb.deleteNumber()permanently removes the database record- The number becomes available for anyone to purchase on Telnyx
Automatic Cleanup
A cron job (POST /api/numbers/cleanup-grace) runs periodically to find and release expired grace numbers. It's protected by an x-cron-secret header so only the cron scheduler can invoke it.
4.9 Telnyx Configuration
Wirevox uses Telnyx as its telephony provider for all phone operations.
Core Integration Points
| Component | Telnyx Feature | Description |
|---|---|---|
| Number Provisioning | Number Orders API | Searching available numbers and purchasing them |
| Call Control | Call Control API | Answering, hanging up, and transferring calls |
| Media Streaming | WebSocket Media Streaming | Real-time audio I/O during calls |
| SMS | Messaging API | Sending and receiving text messages |
WebSocket Media Stream
When a call is answered, Telnyx opens a WebSocket to:
wss://{your-host}/media-stream/{agentId}/{callId}
The WebSocket exchanges JSON messages:
| Event | Direction | Description |
|---|---|---|
connected / start / stream_started |
Telnyx → Server | Stream initialization |
media |
Telnyx → Server | Audio frame: { media: { payload: "base64..." } } |
media |
Server → Telnyx | Response audio: { event: "media", media: { payload: "base64..." } } |
clear |
Server → Telnyx | Flush playback buffer (used during barge-in) |
stop |
Telnyx → Server | Stream ended (call disconnected) |
Audio Format
| Direction | Encoding | Sample Rate | Bit Depth |
|---|---|---|---|
| Telnyx → Server | μ-law (mulaw) | 8,000 Hz | 8-bit |
| Server → Telnyx | μ-law (mulaw) | 8,000 Hz | 8-bit |
For Realtime pipelines, audio is resampled:
- OpenAI Realtime: 8kHz μ-law → 24kHz PCM (input), 24kHz PCM → 8kHz μ-law (output)
- Gemini Live: 8kHz μ-law → 16kHz PCM (input), 16kHz PCM → 8kHz μ-law (output)
Local Development (ngrok)
For local development, Telnyx webhooks need a public URL. Use ngrok to tunnel:
ngrok http 3001
Then configure the Telnyx webhook URL to point to your ngrok URL.
Previous: Section 3 — Agent Configuration Next: Section 5 — Knowledge Base (RAG)
Section 5 — Knowledge Base (RAG)
Complete technical reference for Wirevox's Retrieval-Augmented Generation system — from creating knowledge bases and ingesting documents to the embedding pipeline, vector search, and how context is injected into live phone calls.
5.1 Overview
The Knowledge Base gives your AI agent access to real information — your business policies, FAQs, product catalogs, service menus, team bios, troubleshooting guides, and any other factual data it needs to answer caller questions accurately.
Without a knowledge base, the AI only knows what's in its system prompt. With a knowledge base, the AI can dynamically search and retrieve relevant information during a call, then use it to give precise, grounded answers.
Architecture Summary
┌─────────────────────────────────────────────────────────────┐
│ INGESTION PIPELINE │
│ │
│ Source (PDF/URL/Text/DOCX/CSV/XLSX) │
│ │ │
│ ▼ │
│ Document Processor ──extract text──▶ Raw Text │
│ │ │
│ ▼ │
│ LLM Pre-Processing ──organize──▶ Clean, Structured Text │
│ │ │
│ ▼ │
│ Text Chunker ──split──▶ 1500-char overlapping chunks │
│ │ │
│ ▼ │
│ OpenAI Embeddings ──vectorize──▶ 1536-dim vectors │
│ │ │
│ ▼ │
│ Supabase (pgvector) ──store──▶ knowledge_chunks table │
└─────────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────┐
│ RETRIEVAL (AT CALL TIME) │
│ │
│ Caller Speech ──STT──▶ "Who are your hygienists?" │
│ │ │
│ ▼ │
│ Query Cleaning ──strip filler──▶ "hygienists" │
│ │ │
│ ▼ │
│ OpenAI Embedding ──vectorize──▶ query vector │
│ │ │
│ ▼ │
│ pgvector cosine similarity ──search──▶ top 3 chunks │
│ │ │
│ ▼ │
│ Inject into system prompt ──▶ LLM answers with context │
└─────────────────────────────────────────────────────────────┘
Two Retrieval Modes
| Mode | Trigger | Description |
|---|---|---|
| Passive RAG | Every user utterance (automatic) | The pipeline automatically searches the KB when the user speaks. Results are injected into the system prompt silently. |
| Active Tool | LLM decides to call search_knowledge |
The AI proactively searches the KB when it decides it needs information. Works like any other function call. |
Both modes can operate simultaneously — passive RAG fires automatically, and the active tool is available as a fallback for deeper or more targeted queries.
5.2 Creating a Knowledge Base
UI Structure
The Knowledge page (/w/:id/knowledge) has a sidebar + detail layout:
- Left sidebar (280px) — Lists all knowledge bases in the workspace. Click "+" to create a new one.
- Right panel — Shows sources, chunks, and management actions for the selected KB.
Creating a New KB
Click "+" in the sidebar to start creating. Fill in:
| Field | Required | Description |
|---|---|---|
| Name | ✅ | A short identifier for the KB (e.g., "Dental Clinic FAQ") |
| Description | ❌ | What this KB contains (helps you organize multiple KBs) |
| Instructions | ❌ | Special instructions appended to the system prompt when this KB is linked to an agent (e.g., "Always cite your source when referencing this knowledge base") |
Once created, the KB appears in the sidebar and you can start adding sources.
Editing & Deleting
- Click the more menu (⋮) to rename, edit the description/instructions, or delete the entire KB.
- Deleting a KB removes all its chunks and unlinks it from any agents using it.
5.3 Adding Sources
Sources are the raw materials that get processed into searchable knowledge. There are four ways to add information:
File Upload
Supported formats: PDF, TXT, DOCX, CSV, XLSX
- Click "Add Item" → "Upload File" or drag and drop a file directly onto the source list.
- Files are limited to 10 MB per upload.
- Total KB size is limited to 50 MB.
- If a file with the same name already exists, you'll see a duplicate confirmation asking whether to re-process and overwrite.
How Each Format Is Processed
| Format | Library | Processing |
|---|---|---|
pdf-parse |
Extracts raw text from all pages. Image-based PDFs fail with an error. | |
| TXT | Native | Read as UTF-8. Title derived from filename. |
| DOCX | mammoth |
Extracts raw text (formatting stripped). |
| CSV | csv-parse |
Parses with column headers. Each row becomes Key: Value pairs separated by ---. |
| XLSX | xlsx |
Reads all sheets. Each sheet prefixed with Sheet: {name}. Rows formatted as Key: Value pairs. |
URL Crawling
Two crawl modes available:
Single URL Mode
- Click "Add Item" → "Import from URL".
- Paste a URL and click "Import".
- The system fetches the page, extracts visible text using Cheerio, and ingests it.
Multi-Page Mode
- Switch to "Multi-page" mode.
- Paste your root URL and click "Scan".
- The scanner discovers up to 100 internal links by:
- Parsing all
<a>tags on the page for same-domain URLs - Falling back to sitemap.xml if fewer than 2 links are found
- Parsing all
- Select which pages to import (auto-selects first 20).
- Click "Import Selected" — crawling runs in the background.
URL Text Extraction
The crawler uses a priority-based extraction strategy:
1. Try content selectors: main, article, [role="main"], .content, #content, .post-content, .entry-content
2. Fallback to: body text (with nav/footer/header stripped)
3. If < 50 chars: Jina Reader API fallback (https://r.jina.ai/{url})
4. If still < 50 chars: Error — "page might be heavily JavaScript-dependent"
Elements removed before extraction: script, style, nav, footer, header, aside, iframe, noscript, svg, form, [role="navigation"], [role="banner"], [role="complementary"].
Background Processing
Multi-page crawls are processed sequentially in the background — the API returns immediately and the UI polls every 3 seconds to update source statuses. Each URL goes through:
- Insert a placeholder chunk with
status: 'processing' - Fetch and extract text
- Run through the full ingest pipeline
- Delete the placeholder on success, or update to
status: 'error'on failure
Duplicate URL Handling
If you try to import a URL that already exists as a source, a confirmation dialog asks whether to re-crawl and overwrite. Re-crawling deletes existing chunks for that URL first, then re-processes.
Manual Text Document
- Click "Add Item" → "Write Document".
- Enter a name and paste/type your content (minimum 200 characters).
- The text is processed through the full ingest pipeline.
Retry Failed Crawls
If a URL crawl fails (e.g., timeout, 403 error), the source shows an error badge with the error message and a "Retry" button to re-attempt.
5.4 The Ingest Pipeline
Every piece of content goes through the same 4-step pipeline before it becomes searchable:
Step 1: LLM Pre-Processing
Before chunking, the raw text is sent through GPT-4o Mini for intelligent cleaning and reorganization:
Input: Messy scraped web text with nav elements, repeated headers, "Click Here" CTAs
Output: Clean, organized sections with headers like "## Hygienists" and "## Office Hours"
Rules the LLM follows:
- Group related information together
- Make each section self-contained
- Preserve ALL factual details (names, qualifications, dates, prices, phones, addresses)
- Remove web-specific noise ("click here", "fill out the form below", "call us to book")
- Convert CTAs into factual information suitable for a voice AI
- Use clear section headers
- Never add information not in the original
Skip condition: Texts shorter than 500 characters skip LLM pre-processing (not worth the API cost).
Fallback: If the LLM call fails, the raw text is used as-is.
Step 2: Text Chunking
The processed text is split into overlapping chunks using a smart boundary-aware algorithm:
| Parameter | Value | Description |
|---|---|---|
| Chunk Size | 1,500 characters | Target size for each chunk |
| Chunk Overlap | 100 characters | Overlap between consecutive chunks for context continuity |
Boundary detection priority (for chunk end):
- Paragraph break (
\n\n) — within the last 20% of the chunk - Sentence end (
./?/!/.\n) — searched backward - Forward sentence completion — if backward search fails, look forward to finish the current sentence
- Space — final fallback
Overlap start snapping: The start of each overlapping chunk snaps to the nearest sentence or paragraph boundary to avoid starting mid-sentence.
If the entire text is shorter than 1,500 characters, it becomes a single chunk.
Step 3: Embedding
Each chunk is converted into a 1,536-dimensional vector using OpenAI's text-embedding-3-small model:
- Chunks are batched in groups of 100 for API efficiency
- Each batch call to OpenAI returns an array of embedding vectors
- The result is an array of
{ text, embedding }objects
Step 4: Storage
Chunks are stored in the knowledge_chunks table in Supabase (PostgreSQL with pgvector):
| Column | Type | Description |
|---|---|---|
id |
UUID | Primary key |
knowledge_base_id |
UUID | FK to the knowledge base |
source_type |
text | pdf, url, text, docx, csv, xlsx |
source_name |
text | Original filename or URL |
content |
text | The chunk text |
embedding |
vector(1536) | The 1,536-dim embedding vector |
chunk_index |
integer | Position of this chunk in the source |
metadata |
jsonb | Additional info (title, status, error message) |
created_at |
timestamptz | When the chunk was created |
Deduplication: Before re-ingesting, all existing chunks for the same source_name are deleted, preventing duplicate entries.
5.5 Vector Search
When a query comes in (either from passive RAG or the active tool), the system performs a cosine similarity search using pgvector.
Search Flow
Query Text → embedQuery() → 1,536-dim vector → db.searchSimilarChunks() → top K results
Search Parameters
| Parameter | Default | Description |
|---|---|---|
topK |
3 | Maximum number of chunks to return |
threshold |
0.25 (active tool) / configurable via RAG Threshold slider (passive) | Minimum cosine similarity score |
Search Variants
| Function | Scope | Use Case |
|---|---|---|
searchKnowledge() |
Entire KB | General search — used by passive RAG and the global search_knowledge tool |
searchKnowledgeBySources() |
Specific sources only | Named Knowledge Functions — scoped to user-selected source documents |
Result Format
Each result contains:
{
"id": "chunk-uuid",
"content": "The chunk text...",
"sourceName": "services.pdf",
"similarity": 0.82
}
Weak Result Detection
The system logs warnings for searches where the top result has a similarity score below 0.40:
[Pipeline] ⚠ KB WEAK RESULT — top score 0.312 for raw query: "do you accept bitcoin?"
This helps identify knowledge gaps — topics callers ask about that aren't covered in the KB.
5.6 Passive RAG (Automatic Retrieval)
Passive RAG runs automatically during every Modular pipeline call. The user doesn't configure it — it just works when a KB is linked to the agent.
How It Fires
- The user finishes speaking (STT produces a transcript).
- The pipeline checks if the utterance is meaningful — it strips common conversational words (
yes,no,yeah,okay,please,thanks, etc.) and requires at least 2 meaningful words remaining. - If meaningful,
_retrieveKBContext()fires in parallel with the patience delay timer. - A 1,500ms circuit breaker ensures RAG never blocks the LLM. If the search takes longer, it's abandoned and the LLM proceeds without context.
Query Cleaning
Before searching, the pipeline cleans the caller's speech to improve embedding quality:
Filler stripping patterns:
1. Greetings: "hi", "hey", "hello", "good morning" (removed from start)
2. Filler words: "can you", "please", "just", "like", "actually", "um", "uh"
3. Request forms: "tell me", "I want to know", "I'm wondering"
4. Articles: "at the", "in the", "of the"
Protected phrases (not destroyed by filler stripping):
"right now", "right away", "no longer", "no refund", "not yet",
"how much", "how many", "how long", "at least", "at most"
Example:
Input: "hi can you tell me who are the available hygienists right now"
Output: "available hygienists right now"
If cleaning removes too much (result < 5 characters), the original text is used.
Context Buffer (Sliding FIFO Window)
Retrieved KB context is stored in a sliding buffer that persists across conversation turns:
| Property | Value | Description |
|---|---|---|
_kbContextBuffer |
Array | Stores { query, context, timestamp } entries |
_kbBufferMaxSize |
3 | Maximum entries before FIFO eviction |
This means the AI has access to relevant KB context from the current and previous 2 turns, enabling follow-up questions like:
Caller: "What are Dr. Smith's hours?" → KB retrieves Dr. Smith info
Caller: "And what's her specialty?" → Dr. Smith info still in buffer
System Prompt Injection
When runLLM() fires, the KB buffer is injected into the system prompt:
[KNOWLEDGE BASE — Use this information to answer the user's question accurately.
If you cannot find the answer in this context, politely state that you do not
have that information on file.]
{chunk 1 text}
{chunk 2 text}
{chunk 3 text}
The injection is non-destructive — it appends to a clone of the system message, not the original.
5.7 Active Knowledge Tool
The active tool (search_knowledge) lets the LLM proactively search the KB during a call. Unlike passive RAG (which fires on every utterance), this is a function the AI explicitly invokes when it decides it needs information.
Tool Definition
{
"name": "search_knowledge",
"description": "Search the knowledge base for information relevant to the current conversation...",
"parameters": {
"query": {
"type": "string",
"description": "A specific search query. Be descriptive..."
},
"category": {
"type": "string",
"description": "Optional category to narrow results.",
"enum": ["url", "file", "text"]
}
}
}
Per-Call Cache
The active tool includes a per-call result cache to avoid redundant embedding searches:
- Cache key: normalized (lowercased, trimmed) query string
- Cache lifetime: duration of a single phone call
- On cache hit: returns cached results instantly (no API call)
Response Format
Results found:
{
"results": [
{ "content": "Dr. Smith specializes in...", "source": "team.pdf", "relevance": 0.87 },
{ "content": "Office hours are Mon-Fri...", "source": "https://example.com/hours", "relevance": 0.72 }
]
}
No results:
{
"results": [],
"message": "No relevant information found in the knowledge base for this query. You may need to answer based on your general knowledge or let the customer know you don't have that specific information."
}
When Is the Tool Available?
The search_knowledge tool is added to the agent's tool list only if a knowledge base is linked (knowledgeBaseId is set). It is part of the ALWAYS_ENABLED_TOOLS whitelist, meaning it's always available even if not explicitly referenced in the system prompt via [[Function: search_knowledge]].
5.8 Named Knowledge Functions
Named Knowledge Functions are user-created, scoped search tools that only search specific sources within a KB. They provide more targeted retrieval than the global search_knowledge tool.
Use Case
Imagine your KB has 50 sources covering your entire business. You want the AI to search only your troubleshooting guide when a caller reports a technical issue, not your pricing page or team bios. A Named Knowledge Function like "Router Troubleshoot" scoped to troubleshooting_guide.pdf achieves this.
How They Work
- You create a function in the Functions panel and select "Knowledge" as the category.
- Choose which specific sources (files/URLs) the function should search.
- The function gets a sanitized tool name:
"Router Troubleshoot"→"router_troubleshoot". - At call time,
buildNamedKnowledgeToolHandler()creates a handler that usessearchKnowledgeBySources()— filtering results to only the selected source names.
Scoped Search vs. Global Search
| Feature | search_knowledge (Global) |
Named Function (Scoped) |
|---|---|---|
| Scope | Entire knowledge base | Selected sources only |
| Auto-enabled | Yes (always available if KB linked) | Must be attached to agent and referenced in prompt |
| Cache | Per-call | Per-call |
| topK | 3 | 3 |
5.9 Chunk Management
The UI provides granular control over individual chunks within each source.
Viewing Chunks
Click a source in the source list to open the Chunk Viewer Modal. This shows:
- All chunks in order (
chunk_index) - Each chunk's text content
- Created timestamp
Editing Chunks
Click on a chunk to enter edit mode:
- Modify the chunk text in a textarea.
- Click "Save".
- The backend calls
updateChunkContent()which:- Re-embeds the new text via
embedQuery() - Updates both the
contentandembeddingcolumns in the database - This ensures the new text is correctly searchable
- Re-embeds the new text via
Deleting Chunks
Click the delete button (🗑️) on any chunk:
- A confirmation modal appears.
- The chunk is deleted from the database.
- The source's chunk count updates.
- If it was the last chunk in a source, the chunk modal closes automatically.
Bulk Operations
Select multiple sources using checkboxes:
| Action | Description |
|---|---|
| Bulk Delete | Delete all selected sources and their chunks |
| Bulk Update | Re-crawl all selected URL sources (fetches fresh content) |
Bulk Update is only available for URL sources — it re-crawls each URL in the background and replaces existing chunks with fresh content.
Source Search
A search bar at the top of the source list filters sources by name in real-time.
5.10 Linking a KB to an Agent
To make a knowledge base available to an agent:
- Open the agent's settings page.
- In the Configuration Sidebar, open the Knowledge section.
- Click the Knowledge Selector to choose a KB from your workspace.
- Set the RAG Threshold (0.0–1.0, default 0.25) to control minimum relevance.
When linked:
- Passive RAG activates automatically for every call.
- The
search_knowledgetool becomes available to the AI. - The KB's instructions (if any) are appended to the system prompt.
Unlinking
Click "Unlink" in the Knowledge section. This removes the KB association — passive RAG stops, the search tool is removed, and the instructions are no longer appended.
5.11 Limits & Performance
Size Limits
| Limit | Value |
|---|---|
| KB total size | 50 MB |
| Per-file upload | 10 MB |
| URL scan limit | 100 internal links |
| RAG circuit breaker | 1,500ms timeout |
Performance Characteristics
| Operation | Typical Latency |
|---|---|
| File upload + ingest (small PDF) | 3–8 seconds |
| URL crawl (single page) | 5–15 seconds |
| Vector search (query → results) | 100–300ms |
| LLM pre-processing (per source) | 2–5 seconds |
| Embedding (100 chunks batch) | 1–2 seconds |
Embedding Model
| Property | Value |
|---|---|
| Model | text-embedding-3-small |
| Provider | OpenAI |
| Dimensions | 1,536 |
| Batch Size | 100 chunks per API call |
Database
| Property | Value |
|---|---|
| Database | Supabase PostgreSQL |
| Extension | pgvector |
| Index Type | Cosine similarity |
| Table | knowledge_chunks |
Previous: Section 4 — Phone Numbers & Telephony Next: Section 6 — Functions & Integrations
Section 6 — Functions & Integrations
Technical architecture of Wirevox's tool-calling pipeline — spanning core tools, CRM/calendar integrations, the dynamic Tool Assembler, the orphan tool filter, runtime normalization, and the shared Post-Execution Layer (Executor).
6.1 Overview
Functions (also referred to as tools) allow your AI agent to take action in the real world. Rather than just speaking, the agent can book appointments, lookup patient records, log CRM leads, send email/SMS notifications, and transfer calls to humans.
Wirevox utilizes a declarative, three-layer tool execution model:
┌─────────────────────────────────────────────────────────────┐
│ 1. DEFINITION & ASSEMBLY (Compile Time) │
│ │
│ Workspace Integrations Connected │
│ │ │
│ ▼ │
│ Orphan Filter (filterOrphanTools.js) │
│ Strips tools for disconnected integrations │
│ │ │
│ ▼ │
│ Tool Assembler (assembleTools.js) │
│ Generates OpenAI schemas, custom fields & prompt overrides │
└─────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ 2. NORMALIZATION & DISPATCH (Call Runtime) │
│ │
│ LLM invokes tool (e.g., jobber_create_lead) │
│ │ │
│ ▼ │
│ Universal Dispatcher (inboundHandler.js) │
│ - Cleans spoken phone numbers/emails to standard digits │
│ - Rejects incomplete emails with instructions to AI │
│ - Generates warm handoff tokens for transfers │
└─────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ 3. CORE EXECUTOR & POST-PROCESSING (Shared Layer) │
│ │
│ Shared Executor (executor.js) │
│ - Executes thin adapter API calls │
│ - Inserts local leads (duplicate check & failover path) │
│ - Sets call outcomes (priority check: booked > message) │
│ - Fires destination-specific webhooks │
│ - Deducts billing credits (e.g., SMS, transfer tools) │
└─────────────────────────────────────────────────────────────┘
6.2 Tool Categories & Function Catalog
Wirevox organizes actions into 8 distinct function categories managed via the Functions page (/w/:id/functions):
| Category | UI Label | Target Tool | Schema / Custom Behavior |
|---|---|---|---|
integrations |
Integrations | CRM/Calendar tools | Generated dynamically from integration adapters |
call_transfer |
Call Transfer | transfer_call |
Core tool. Redirects caller via Telnyx |
sms_confirmation |
SMS Confirmation | send_sms_confirmation |
Sends SMS via Telnyx. Supports templates |
email_confirmation |
Email Confirmation | send_email_confirmation |
Sends emails. Supports templates |
gather_information |
Gather Information | system_infolist |
Augments lead fields with custom properties |
knowledge |
Knowledge | Named scoped search | Generates a scoped search tool for selected docs |
workflow |
Workflow | Named graph SOP | Injects a conversational tree state-machine prompt |
take_message |
Take a message | take_message |
Core tool. Collects caller message details |
6.3 Core Tools (Always Available)
Core tools are built into the platform and do not require external API configurations.
transfer_call
Transfers the caller to another number.
- Parameters:
reason(string) - Behavior: Shuts down the AI pipeline to freeze billing and commands Telnyx to bridge the call.
system_infolist
Saves caller details directly to the local database as a lead.
- Parameters:
name(string, required),phone(string, required),email(string, optional),reason(string, optional), plus dynamic custom fields.
send_sms_confirmation
Sends a text confirmation to the caller.
- Parameters:
toPhone(string, required),message(string, required)
send_email_confirmation
Sends an email confirmation to the caller.
- Parameters:
toEmail(string, required),subject(string, required),body(string, required)
take_message
Logs a message for the business owner.
- Parameters:
callerName(string, required),callerPhone(string, required),message(string, required),recipientName(string, optional),callerEmail(string, optional),urgency(enum:normal,urgent)
6.4 Integrations Catalog
Integration tools are loaded from Connected Workspace Providers.
Integrations Catalog (INTEGRATION_TOOLS)
├── Google Calendar (google_calendar)
│ ├── google_check_availability
│ ├── google_book_appointment
│ ├── google_lookup_appointment
│ ├── google_cancel_appointment
│ └── google_reschedule_appointment
├── Outlook Calendar (outlook_calendar)
│ ├── outlook_check_availability
│ ├── outlook_book_appointment
│ ├── outlook_lookup_appointment
│ ├── outlook_cancel_appointment
│ └── outlook_reschedule_appointment
├── Jobber CRM (jobber)
│ ├── jobber_lookup_client
│ ├── jobber_create_lead
│ └── jobber_create_request
├── Clio Manage (clio)
│ └── clio_create_intake
├── Clio Grow (clio_grow)
│ └── clio_grow_create_lead
├── Filevine (filevine)
│ └── filevine_create_intake
├── Open Dental (open_dental)
│ ├── open_dental_lookup_patient
│ ├── open_dental_create_patient
│ ├── open_dental_check_availability
│ └── open_dental_book_appointment
└── Email (email)
└── email_send_summary
6.5 The Tool Assembler (assembleTools)
When a call initializes, the assembleTools function maps attached agent functions into OpenAI-compatible tool definitions.
Category Mapping Details
- Named Knowledge & Workflows: Generates tool definitions named dynamically matching the sanitized function name (lowercase, non-alphanumeric replaced by
_). Configs (_knowledgeConfigand_workflowConfig) are attached directly to the tool definition to guide the dispatcher. - Integrations: Resolves enabled tools from the integration config (parsing the
enabledToolsobject ortoolsarray) and pushes their definitions to the schema pool. - Template Overrides:
- For
SMS ConfirmationandEmail Confirmation, if a custom message template is defined, the tool's description is dynamically overwritten:"Send an SMS text confirmation. The user has provided a template: "[template]". You MUST use this exact template for your 'message' parameter, but replace any bracketed variables (like [date], [time]) with the actual values..."
- For
- Core Category Mapping: Maps standard Call Transfer, Take a Message, SMS, and Email functions to their hardcoded core parameter schemas.
Dynamic system_infolist Augmentation
If a gather_information category function is attached to the agent:
- The assembler reads custom field schemas defined in
config.fields(or legacyconfig.customFields). - It dynamically appends these custom properties to the
system_infolistparameter properties block. - For example, if you configure a field
pet_breedin the UI, the LLM will seepet_breedas a valid argument on thesystem_infolisttool definition during call execution.
6.6 The Orphan Tool Filter (filterOrphanTools)
To prevent the LLM from attempting to execute integrations that are configured but disconnected in the workspace, the system uses filterOrphanTools:
- It parses all enabled tools for an agent.
- It collects unique providers needed (e.g.,
google_calendar,jobber). - It query-checks the database for the active workspace to verify connection states.
- Any tool whose integration provider is disconnected is stripped from the active list.
- This acts as a safeguard, ensuring the LLM is never presented with invalid tool definitions.
6.7 Normalization & Dispatch (inboundHandler)
The receptionist handler (buildInboundToolHandler) acts as the universal dispatcher.
Normalization Helpers
Voice transcripts often spell out numbers or emails. Before dispatching arguments to any tool, the handler normalizes contact details:
// Convert words to digits
"five five five one two three four" ➔ "5551234"
- Phone Numbers: Normalizes strings by converting word numerals to digits, stripping non-numeric characters, and ensuring E.164 formatting:
- 10 digits ➔ prepends
+1(North America) - 11 digits starting with 1 ➔ prepends
+
- 10 digits ➔ prepends
- Emails: Normalizes spelled-out emails by replacing spoken symbols:
" at "➔@" dot "➔." underscore "➔_" dash "/" hyphen "➔-- All spaces are collapsed.
Invalid Email Blocking
If the normalized email does not contain both a @ and a ., the dispatcher blocks execution and returns an error message:
{
"error": "Email address appears incomplete — it is missing the @ symbol or domain. Ask the caller to repeat their full email address before retrying."
}
This instructs the LLM to verify and request the email again, preventing trash entries.
Warm Agent Transfers & Handoff Tokens
When transferring calls to another agent (destinationType === 'agent'):
- The dispatcher looks up the target agent's phone number.
- If
transferType === 'warm', the dispatcher generates a Handoff Token:- It captures the last 15 lines of the transcript.
- It saves this metadata to the database via
db.createHandoff(). - It injects the token into custom headers sent to Telnyx:
X-Wirevox-Handoff: <token>.
- When the receiving agent picks up the call, the backend resolves this token and pre-loads the transcript context, allowing the target agent to resume the conversation seamlessly.
6.8 Shared Post-Execution Layer (Executor)
All tool executions route through the shared executeTool function in executor.js. This centralizes post-execution logic and database synchronization.
1. Duplicate Lead Prevention & Fallback
- Lead Insertion: If a tool's manifest defines
createsLead: true, the executor constructs a lead record and saves it locally. To prevent multiple tools (e.g., booking an appointment and taking a message) from creating duplicate lead entries, the executor setsctx._leadCreatedThisCall = trueand skips subsequent lead creation. - Failover Storage: If an external API call fails (e.g., Google Calendar returns 500), the executor catches the error, inserts a local lead record with a
[FAILED]prefix in the notes, and responds to the LLM with a graceful fallback message so the AI can continue speaking.
2. Outcome Priorities
The call's outcome column is updated based on tool execution. To avoid downgrading the status during a call, the executor enforces an OUTCOME_PRIORITY hierarchy:
booked (6) > rescheduled (5) > cancelled (4) > intake/transferred (3) > message_taken (2) > inquiry (1) > other (0)
A lower-priority outcome will never overwrite a higher-priority one (e.g., if a booking succeeds, a subsequent call transfer will not overwrite the call's outcome from booked to transferred).
3. Dynamic Webhook Routing
If the tool's manifest specifies a webhookEvent, the executor looks up the webhook URL configured for the parent function and fires the webhook asynchronously.
Augmenting system_infolist Webhooks
For system_infolist, the executor reads the gather_information configuration:
- It only includes arguments where
destWebhook: trueis configured in the field metadata (or falls back to all arguments for legacy structures). - If a webhook URL is configured, it fires the
information.gatheredevent with the filtered payload.
4. Billing & Duration Freezing
- Billing Deductions: For billable actions, the executor calls
db.deductCreditsusing dynamic cost registries (e.g., SMS sending, transfers). - Handoff Duration Freeze: For call transfers, the executor updates the call record's
transferred_attimestamp. This freezes AI call duration billing, ensuring the user is only billed for AI duration up to the moment they are connected to a human.
Previous: Section 5 — Knowledge Base (RAG) Next: Section 7 — Post-Call Processing & Webhooks
Section 7 — Post-Call Processing & Webhooks
What happens after the caller hangs up — background transcript analysis, outcome classification, lead synchronization, email summaries, and webhook delivery.
7.1 Overview
When a call ends (the WebSocket closes), Wirevox doesn't just save the transcript and move on. A background post-call pipeline runs automatically to enrich every call with structured data:
Call Ends (WebSocket closes)
│
├─ 1. Save transcript to database
│
├─ 2. runPostCallAnalysis() ───▶ Background
│ │
│ ├─ a. Classify outcome (GPT-4o Mini)
│ ├─ b. Generate summary
│ ├─ c. Extract caller details
│ ├─ d. Update call record
│ ├─ e. Sync lead record
│ └─ f. Fire post-call hooks (email summary)
│
└─ 3. Fire call.ended webhook (if configured)
Trigger Point
The pipeline fires from [mediaStream.js](file:///c:/Users/iysal/.gemini/antigravity/scratch/Wirevox App/backend/src/routes/telephony/mediaStream.js) in the ws.on('close') handler:
ws.on('close', async () => {
pipeline.end();
await db.updateCall(callId, {
status: 'ended',
endedAt: new Date().toISOString(),
transcript: transcriptRecord,
});
// KICK OFF BACKGROUND ANALYSIS
runPostCallAnalysis(callId, transcriptRecord, agent);
// Fire call.ended webhook
if (postCallWebhook) {
fireWebhook(postCallWebhook, 'call.ended', {
callId, agentName, agentType, transcript, messageCount
});
}
});
The analysis runs asynchronously — it doesn't block the WebSocket close. Even if analysis fails, the call is still marked as ended.
7.2 LLM-Powered Transcript Analysis
The core of post-call processing is a single call to GPT-4o Mini with a structured JSON response.
Input
The full transcript is formatted as:
AGENT: Hi, this is Dr. Smith's office. How can I help you?
USER: I'd like to schedule a cleaning appointment.
AGENT: Of course! I can help with that...
System Prompt
The LLM is given explicit instructions to:
- Classify the outcome — choose from a strict enum of valid outcomes
- Write a summary — concise 1-2 sentences
- Extract caller details — name, phone, email, reason for calling
Output Format
The response uses response_format: { type: 'json_object' } to guarantee structured output:
{
"outcome": "booked",
"summary": "Caller scheduled a dental cleaning for next Tuesday at 2pm with Dr. Smith.",
"callerName": "Sarah Johnson",
"callerPhone": "+15551234567",
"callerEmail": "sarah.j@email.com",
"reason": "Schedule a dental cleaning"
}
Valid Outcomes
| Outcome | Description |
|---|---|
booked |
An appointment was successfully scheduled |
transferred |
The call was transferred to a human staff member |
message_taken |
The caller left a message for the team |
inquiry |
The caller asked questions but didn't book or leave a message |
cancelled |
An existing appointment was cancelled |
rescheduled |
An existing appointment was rescheduled to a new time |
intake |
Caller info was submitted to a CRM (Clio, Jobber, Filevine, etc.) |
other |
Wrong number, disconnected early, or any other outcome |
Outcome Priority Respect
If an outcome was already set during the call by a tool execution (e.g., google_book_appointment set booked), the post-call analysis skips re-classification and only runs enrichment (summary, caller details). This prevents the LLM from accidentally downgrading a confirmed booking to an "inquiry."
const skipClassification = existingOutcome
&& existingOutcome !== 'inquiry'
&& existingOutcome !== 'other';
Only inquiry and other outcomes are eligible for reclassification.
7.3 Call Record Update
After analysis, the call record in the calls table is updated with:
| Field | Source | Description |
|---|---|---|
outcome |
Tool or LLM classification | Final call outcome |
notes |
LLM-generated summary | 1-2 sentence call summary |
endedAt |
Server timestamp | When the call ended |
leadName |
LLM extraction | Caller's name (if detected) |
leadPhone |
LLM extraction or caller ID | Caller's phone number |
Phone Number Backfill
If the LLM couldn't extract a phone number from the transcript (e.g., the caller never stated it), the system backfills from:
call.callerNumber— the Telnyx caller IDcall.leadPhone— any phone captured during the call by tools
This ensures email summaries always include the phone number, even on quick inquiry calls where the caller didn't provide it verbally.
7.4 Lead Synchronization
After updating the call record, the system syncs data to the leads table (the Contacts page in the dashboard).
Two Scenarios
Scenario A: Lead Already Exists
If a tool created a lead during the call (e.g., google_book_appointment, system_infolist), the system finds the most recent lead for this agent (created within the last 5 minutes) and enriches it:
| Update | Description |
|---|---|
notes |
Overwritten with the LLM-generated rich summary |
status |
Updated to match the call outcome |
name |
Backfilled if the tool didn't capture it |
phone |
Backfilled if the tool didn't capture it |
email |
Backfilled if the tool didn't capture it |
Scenario B: No Lead Exists
If no tool created a lead during the call (e.g., a quick inquiry where the caller just asked a question), the system creates a new lead from the post-call analysis. This ensures every caller appears in the Contacts page, even if no tools were triggered.
await db.insertLeads(agent.id, [{
name: analysis.callerName || 'Unknown Caller',
phone: analysis.callerPhone || '',
email: analysis.callerEmail || '',
notes: analysis.summary || '',
status: STATUS_MAP[outcome] || 'new',
}]);
Lead Status Mapping
| Call Outcome | Lead Status |
|---|---|
booked |
booked |
rescheduled |
booked |
intake |
intake |
message_taken |
message_taken |
transferred |
transferred |
cancelled |
cancelled |
inquiry |
new |
other |
new |
7.5 Email Summary Hook
If the agent has the email_send_summary integration tool enabled, a professional HTML email is sent to the business owner after every call.
Trigger Logic
The post-call processor checks if:
- The agent's published function snapshot includes
email_send_summaryas an enabled tool - The workspace has an email integration configured (with a target email address)
If both conditions are met, the email fires.
Email Content
The email is a professional HTML template with:
| Section | Content |
|---|---|
| Header | Dark gradient banner: "Call Summary — Handled by {Agent Name}" |
| Table | Caller Name, Phone, Email, Reason, Appointment details, Notes |
| Footer | "Sent securely by Wirevox Receptionist" |
Email Data Sources
| Field | Source |
|---|---|
callerName |
LLM extraction from transcript |
callerPhone |
LLM extraction or caller ID backfill |
callerEmail |
LLM extraction |
reason |
LLM extraction |
notes |
LLM-generated summary |
appointmentDate/Time |
From booking tool result (if applicable) |
Important: Post-Call Only
The email_send_summary tool has postCallOnly: true in its tool definition. This means:
- It is NOT fed to the live AI during a call — the LLM never sees or tries to call it
- It fires deterministically after the call ends, via the post-call processor
- This ensures the email always contains the complete call summary, not a partial one
7.6 Webhooks
Wirevox fires webhook events at two levels: per-tool (during the call) and per-call (after the call ends).
Webhook Architecture
export async function fireWebhook(webhookUrl, event, payload) {
const body = {
event, // e.g. "call.ended"
timestamp, // ISO 8601
data: payload, // Event-specific data
};
// Fire-and-forget with 10-second timeout
await fetch(webhookUrl, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify(body),
});
}
Key Properties
| Property | Value |
|---|---|
| Method | POST |
| Content-Type | application/json |
| Timeout | 10 seconds (AbortController) |
| Failure behavior | Non-fatal — logged and swallowed |
| Retry | None (fire-and-forget) |
Webhook Events
Per-Tool Events (fired during the call by the Executor)
| Event | Trigger |
|---|---|
appointment.booked |
google_book_appointment or outlook_book_appointment or open_dental_book_appointment succeeds |
appointment.cancelled |
google_cancel_appointment or outlook_cancel_appointment succeeds |
appointment.rescheduled |
google_reschedule_appointment or outlook_reschedule_appointment succeeds |
call.transferred |
transfer_call tool fires |
message.taken |
take_message tool succeeds |
information.gathered |
system_infolist tool succeeds (with Gather Information function attached) |
clio.intake_created |
clio_create_intake succeeds |
clio_grow.lead_created |
clio_grow_create_lead succeeds |
jobber.lead_created |
jobber_create_lead succeeds |
jobber.request_created |
jobber_create_request succeeds |
filevine.intake_created |
filevine_create_intake succeeds |
email.summary_sent |
email_send_summary completes |
Per-Call Events (fired after the call ends)
| Event | Trigger | Payload |
|---|---|---|
call.ended |
WebSocket closes | callId, agentName, agentType, transcript, messageCount |
Webhook URL Configuration
Webhook URLs are configured at two levels:
- Per-Function Webhook — Set in the function's configuration (
config.webhookUrl). Per-tool events route to this URL. - Post-Call Webhook — Set in the agent's configuration (
agentConfig.postCallWebhook). Thecall.endedevent routes to this URL.
Compatible Platforms
Webhooks work with any HTTP endpoint that accepts POST requests:
- Zapier — Use a "Webhooks by Zapier" trigger
- Make (Integromat) — Use a "Webhooks" module
- n8n — Use a "Webhook" trigger node
- Custom endpoints — Any server that accepts JSON POST requests
Webhook Payload Structure
Every webhook follows the same envelope format:
{
"event": "appointment.booked",
"timestamp": "2026-07-16T05:27:00.000Z",
"data": {
"callerName": "Sarah Johnson",
"callerPhone": "+15551234567",
"service": "Dental Cleaning",
"preferredDate": "2026-07-23",
"preferredTime": "14:00",
"agentName": "Front Desk AI"
}
}
7.7 Error Handling & Resilience
Analysis Failure
If the GPT-4o Mini analysis call fails (API error, timeout, etc.), the call is still saved with a fallback:
await db.updateCall(callId, {
outcome: 'other',
notes: 'Post-call analysis failed.',
endedAt: new Date().toISOString(),
});
This prevents calls from remaining in a "zombie" state with no end marker.
Email Hook Failure
Email summary failures are non-fatal. If the SMTP send fails:
- The error is logged
- The rest of the post-call flow continues
- The call record and lead sync are unaffected
Webhook Failure
Webhook delivery failures are completely non-fatal and non-blocking:
- 10-second timeout via AbortController
- Errors are logged but swallowed
- No retries — fire-and-forget
- Never affects call flow or post-call processing
7.8 Processing Timeline
Here's the typical timing for the entire post-call pipeline:
| Step | Timing | Description |
|---|---|---|
| Transcript Save | ~50ms | Immediate DB write on WebSocket close |
| LLM Analysis | ~1-3s | GPT-4o Mini processes transcript |
| Call Update | ~50ms | Update call record with outcome + summary |
| Lead Sync | ~50-100ms | Create or enrich lead record |
| Email Summary | ~1-2s | Build HTML + SMTP delivery |
| Webhook Delivery | ~100-500ms | POST to configured URLs |
Total: ~2-6 seconds after call ends. All operations run in the background — the caller has already hung up and the WebSocket is closed.
Previous: Section 6 — Functions & Integrations Next: Section 8 — Billing & Credits
Section 8 — Integrations Setup
Technical reference for how Wirevox connects to third-party CRMs and Calendars — covering OAuth flows, static API keys, credential storage, and the cleanup lifecycle.
8.1 Overview
The Integrations page (/w/:id/integrations) allows users to connect external systems to their workspace. Once an integration is connected, its corresponding tools become available to all agents within that workspace.
Wirevox supports two primary authentication patterns for external systems:
- OAuth 2.0 (Google Calendar, Outlook Calendar, Jobber, Clio Manage, Clio Grow)
- Static API Keys / PATs (Filevine, Open Dental)
Database Storage
All integration credentials are stored in the integrations table:
| Column | Description |
|---|---|
workspace_id |
Links the integration to a specific workspace |
provider |
The integration ID (e.g., google_calendar, jobber) |
access_token |
Short-lived OAuth access token |
refresh_token |
Long-lived OAuth refresh token (OR static API key) |
token_expiry |
Timestamp when the access_token expires |
email |
Captured user email (used for Google/Outlook/Email integrations) |
[CODE SMELL FLAG] The refresh_token Column
Currently, Wirevox uses the refresh_token column to store static credentials for non-OAuth integrations (like Filevine PATs and Open Dental Customer Keys) because the table lacks a dedicated api_key or credentials JSONB column.
While functional, this is a known architectural code smell. Future refactoring should introduce a dedicated JSONB credentials column to separate OAuth tokens from static API keys.
8.2 OAuth 2.0 Flow
For OAuth-based integrations (Google, Outlook, Jobber, Clio, Clio Grow), Wirevox implements a standard Authorization Code flow.
1. Generating the Auth URL
When a user clicks "Connect" in the UI, the frontend calls GET /api/integrations/{provider}/auth.
The backend generates the OAuth consent URL. Crucially, it passes the userId and workspaceId as a JSON-stringified state parameter. This ensures the system knows which workspace to attach the credentials to when the callback returns.
// Example: Google Calendar
const state = JSON.stringify({ u: req.userId, w: workspaceId });
const url = oauth2Client.generateAuthUrl({
access_type: 'offline',
prompt: 'select_account consent',
scope: [...scopes],
state,
});
2. The Callback Handler
When the user approves the connection, the provider redirects to /api/integrations/{provider}/callback?code=...&state=...
The callback route:
- Parses the
stateparameter to recover theuserIdandworkspaceId. - Exchanges the
codefor anaccess_tokenandrefresh_token. - Calls
db.upsertIntegrationto store the tokens in the database. - Redirects the user back to the frontend Integrations page with a success flag.
3. Token Refreshing
When an AI agent executes a tool (e.g., booking an appointment), the integration adapter checks the token_expiry.
If the token is expired (or within a 5-minute buffer), the backend automatically uses the refresh_token to fetch a new access_token, updates the database, and then proceeds with the API call.
8.3 Specific Provider Configurations
Google Calendar
- Auth Type: OAuth 2.0
- Scopes:
calendar.events,calendar.readonly,userinfo.email - Settings:
access_type=offline,prompt=select_account consent(forces a refresh token to be issued) - Extra: Captures the user's email address to display in the UI. Currently defaults to the
primarycalendar.
Outlook Calendar
- Auth Type: OAuth 2.0
- Scopes:
Calendars.ReadWrite,User.Read,offline_access - Extra: Captures the user's email address.
Jobber (Field Services CRM)
- Auth Type: OAuth 2.0
- Scopes:
clients,requests
Clio Manage & Clio Grow (Legal)
- Auth Type: OAuth 2.0
- Scopes: Standard CRM scopes.
Filevine (Legal)
- Auth Type: Static API Key (Personal Access Token)
- Flow: The user manually inputs a
PAT,Client ID, andClient Secretin the UI. - Validation: The backend immediately tests the credentials by attempting a token exchange with the Filevine API. If successful, the combined credentials payload is stored in the
refresh_tokencolumn.
Open Dental
- Auth Type: Static API Key (Customer Key)
- Flow: The user inputs an Open Dental Developer
Customer Keyin the UI. - Validation: The backend tests the key by calling the Open Dental API to fetch providers. If successful, the key is stored in the
refresh_tokencolumn.
Email Notifications
- Auth Type: None
- Flow: The user simply inputs a target email address in the UI.
- Storage: Stored in the
emailcolumn. Used by the Post-Call processing pipeline to send call summaries via Wirevox's internal SMTP server.
8.4 The Cleanup Lifecycle (cleanupToolsForProvider)
When a user deletes an integration from their workspace, it's not enough to simply delete the API keys.
If agents in the workspace are actively using tools from that provider (e.g., google_book_appointment), disconnecting the integration would leave those tools broken and "orphaned" on the agent.
To handle this, routes/integrations.js implements a cleanup function:
const PROVIDER_TOOL_PREFIXES = {
google_calendar: 'google_',
jobber: 'jobber_',
// ...
};
async function cleanupToolsForProvider(userId, workspaceId, provider) {
const prefix = PROVIDER_TOOL_PREFIXES[provider];
// 1. Get all agents in the workspace
const agents = await db.getAgents(userId, workspaceId);
for (const agent of agents) {
const tools = await db.getAgentTools(agent.id);
// 2. Find any tools starting with the provider prefix
for (const tool of tools) {
if (tool.toolName.startsWith(prefix)) {
// 3. Hard delete the tool from the agent
await db.deleteAgentTool(agent.id, tool.toolName);
}
}
}
}
When DELETE /api/integrations/{provider} is called:
- The backend attempts to revoke the token with the provider (if OAuth).
- Deletes the row from the
integrationstable. - Calls
cleanupToolsForProviderto silently strip the associated tools from all agents in the workspace.
This ensures the system state remains consistent and agents do not attempt to use tools they no longer have credentials for.
(Note: There is also an filterOrphanTools.js layer that acts as a runtime safeguard against orphan tools, but this cleanup script handles the persistent database cleanup).
Previous: Section 7 — Post-Call Processing & Webhooks Next: Section 9 — Billing & Credits
Section 9 — Billing & Credits
Technical architecture of Wirevox's billing system — covering abstract credits, the real-time dynamic Cost Registry, Stripe integration, and the carryover billing mechanism.
9.1 Overview: The Credit System
Wirevox does not bill directly in minutes or API requests. Instead, it uses Credits as an abstract internal unit of measurement.
This abstraction solves a critical problem: Not all AI calls cost the same. An agent using gpt-4o with elevenlabs TTS is vastly more expensive to operate than an agent using gpt-4o-mini with deepgram TTS.
By abstracting usage into Credits:
- Users have a single, unified balance.
- The platform can charge dynamically based on the exact AI stack the user configures for their agent.
- The platform can bill flat fees for discrete actions (e.g., sending an SMS, or transferring a call) against the same balance.
The dollar value of a single Credit depends on the user's Subscription Plan (their "overage rate").
9.2 The Cost Registry (costRegistry.js)
The actual cost of an action is calculated at runtime by the costRegistry.js module.
Dynamic Database Configuration
To allow pricing to be adjusted without redeploying code, the system loads credit values from a Supabase table (credit_costs). The backend subscribes to Supabase Realtime (postgres_changes) to instantly invalidate its local cache whenever a cost is updated in the database.
If the database is unreachable, the system falls back to a hardcoded DEFAULT_COSTS map.
Per-Minute Calculation
The cost of 1 minute of call time is calculated as:
Total Cost = Platform Base + STT Cost + LLM Cost + TTS Cost
Example Default Values:
- Platform Base: 4 credits (Covers Telnyx telephony and basic infrastructure)
- STT (Deepgram): 2 credits
- LLM (GPT-4o-mini): 1 credit
- TTS (Deepgram): 2 credits
- Total: 9 credits per minute
If a user upgrades their agent to use ElevenLabs TTS (10 credits) and GPT-4o (3 credits), their per-minute cost dynamically scales to 19 credits per minute.
Flat Action Costs
The registry also defines flat costs for tool executions. These are billed deterministically when the Executor processes a tool:
func_sms: 2 credits (Billed per SMS sent)func_transfer: 20 credits (Billed once when a call is successfully transferred to a human)
9.3 Per-Second Billing & Carryover
Wirevox bills callers for exact AI duration.
To prevent charging users for fractional minutes while still maintaining integer credit balances, the system uses a Carryover Buffer (billing_carryover_seconds on the agent record).
How it Works:
- A call lasts for 1 minute and 15 seconds.
- The user is immediately billed for 1 minute (e.g., 9 credits).
- The remaining 15 seconds are added to the agent's carryover buffer.
- On the next call, if the call lasts 50 seconds, the system adds the 15 seconds from the buffer (Total = 65 seconds).
- The user is billed for 1 minute (9 credits), and the buffer now holds 5 seconds.
Design Decision: Un-stamped Carryover
Carryover seconds are stored as raw durations, not rate-stamped values. If a user changes their agent's AI stack (e.g., swapping to a more expensive TTS model), any accrued seconds in the buffer will be billed at the new rate when they eventually cross the 60-second threshold.
This is an accepted architectural simplification because manual AI model reconfiguration is rare, and the maximum discrepancy is limited to 59 seconds' delta.
9.4 Stripe Integration & Tiers
Billing is handled via Stripe, implemented in routes/stripe.js. The platform supports a hybrid SaaS model with fixed monthly tiers and dynamic pay-as-you-go top-ups.
Subscription Tiers
| Tier | Monthly Price | Description |
|---|---|---|
| PAYG | $0 | Pay-As-You-Go. User only pays for manual top-ups. |
| Starter | $99 | Includes a baseline amount of monthly recurring credits and 1 concurrent call slot. |
| Growth/Standard | $299 | Higher volume, lower per-credit overage rate. |
| Agency | $999 | White-label capabilities, sub-account management, lowest overage rate. |
Add-on Products (Concurrent Calls)
To prevent abuse, the platform restricts how many AI calls can happen simultaneously per workspace. Users can purchase Extra Concurrent Call Add-ons via Stripe. The price of the add-on scales inversely with the subscription tier:
- Starter Add-on: $15/mo per extra line
- Growth Add-on: $13/mo per extra line
- Agency Add-on: $5/mo per extra line
Upgrades with Proration Simulation
Wirevox provides a seamless upgrade path. When a user previews an upgrade in the UI, routes/stripe.js calls Stripe's Invoice Preview API.
This calculates the exact prorated amount the user will be charged today (crediting them for unused time on their current plan, and factoring in any existing add-ons) before they commit to the change.
Dynamic Top-Ups
If a user runs out of credits mid-month, they can trigger a manual Top-Up. The cost of a top-up depends on their plan's overage rate (overageRateCreditsCents).
- The backend accepts the desired credit amount (minimum 50).
- Calculates the dollar value:
Amount * overageRateCreditsCents. - Generates a Stripe Checkout session for a one-time payment.
Previous: Section 8 — Integrations Setup