Introduction
Building an interactive AI Voice Assistant no longer requires expensive API infrastructure or complex Python backend frameworks like FastAPI or Django. In this tutorial, we will construct a lightweight, high-performance Web-based AI Voice Assistant using PHP (Backend), JavaScript (Frontend), Google Gemini 2.5 Flash-Lite API, and Microsoft Edge-TTS.
This setup supports real-time voice input, live Web Search grounding via Gemini, full audio output playback, dynamic relative time-stamps, and multi-language adaptation (English, Hindi, Urdu, Chinese, Burmese, etc.).
Prerequisites & Server Requirements
Before deploying the application, ensure your Linux VPS or Web Server meets the following specifications:
-
PHP 7.4+ or 8.x with
curlenabled. -
Python 3.x and
pipinstalled. -
Microsoft Edge-TTS CLI tool installed on your server:
pip install edge-tts
Project File Structure
Organize your application files in your root web server directory as follows:
/voice-assistant/
│
├── index.html # Frontend Chat Interface & Audio Recorder
├── api.php # Backend API Processor (Gemini & Edge-TTS)
├── audio/ # Directory where generated MP3 files are stored
└── uploads/ # Temporary directory for incoming WebM voice files
1. Backend Script (api.php)
This PHP script handles incoming user audio/text, passes it to the Gemini 2.5 API (with Google Search grounding enabled), and converts the output text into speech using Edge-TTS.
Security Note: Replace
your_api_keywith your actual Google Gemini API key.PHP
<?php
// CORS & JSON Header
header(“Access-Control-Allow-Origin: *”);
header(“Access-Control-Allow-Headers: Content-Type”);
header(“Content-Type: application/json; charset=UTF-8″);// Preflight options request handling
if ($_SERVER[‘REQUEST_METHOD’] === ‘OPTIONS’) {
http_response_code(200);
exit();
}// CONFIGURATION SETUP
$apiKey = ‘your_api_key’; // Replace with your Gemini API Key$userMessage = ”;
$audioFilePath = null;
$audioMimeType = null;// 1. Handle incoming voice recording
if (isset($_FILES[‘audio’]) && $_FILES[‘audio’][‘error’] === UPLOAD_ERR_OK) {
$uploadDir = __DIR__ . ‘/uploads’;
if (!is_dir($uploadDir)) {
mkdir($uploadDir, 0777, true);
}
$audioFilePath = $uploadDir . ‘/input_’ . time() . ‘.webm’;
move_uploaded_file($_FILES[‘audio’][‘tmp_name’], $audioFilePath);
$audioMimeType = ‘audio/webm’;
}// 2. Handle incoming text query
if (isset($_POST[‘message’])) {
$userMessage = trim($_POST[‘message’]);
} else {
$rawInput = file_get_contents(‘php://input’);
$inputData = json_decode($rawInput, true);
if (isset($inputData[‘message’])) {
$userMessage = trim($inputData[‘message’]);
}
}if (empty($userMessage) && !$audioFilePath) {
echo json_encode([‘error’ => ‘No message or audio provided’]);
exit();
}// STEP 1: Process request with Gemini 2.5 Flash Lite API
$url = “https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash-lite:generateContent?key=” . $apiKey;$parts = [];
if ($audioFilePath && file_exists($audioFilePath)) {
$base64Audio = base64_encode(file_get_contents($audioFilePath));
$parts[] = [
“inline_data” => [
“mime_type” => $audioMimeType,
“data” => $base64Audio
]
];
}$promptText = !empty($userMessage) ? $userMessage : “Please listen to this audio and respond naturally.”;
$parts[] = [“text” => $promptText];$payload = [
“system_instruction” => [
“parts” => [
[“text” => “You are a friendly AI Voice Assistant. Respond naturally and concisely. Use Google Search for accurate real-time information.”]
]
],
“contents” => [
[
“parts” => $parts
]
],
“tools” => [
[“google_search” => new stdClass()] // Enable Google Search Grounding
]
];$ch = curl_init($url);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
curl_setopt($ch, CURLOPT_POST, true);
curl_setopt($ch, CURLOPT_HTTPHEADER, [‘Content-Type: application/json’]);
curl_setopt($ch, CURLOPT_POSTFIELDS, json_encode($payload));
curl_setopt($ch, CURLOPT_SSL_VERIFYPEER, false);$response = curl_exec($ch);
curl_close($ch);$responseData = json_decode($response, true);
if (isset($responseData[‘error’])) {
echo json_encode([‘reply’ => ‘API Error: ‘ . $responseData[‘error’][‘message’]]);
exit();
}$replyText = $responseData[‘candidates’][0][‘content’][‘parts’][0][‘text’] ?? ‘Sorry, I could not generate a response.’;
// STEP 2: Generate Audio via Edge-TTS
$audioUrl = null;
$execDebug = [];if (!empty($replyText)) {
$cleanText = preg_replace(‘/[*#\_`\-]/u’, ‘ ‘, $replyText);$audioDir = __DIR__ . ‘/audio’;
if (!is_dir($audioDir)) {
mkdir($audioDir, 0777, true);
}$fileName = ‘voice_’ . time() . ‘_’ . rand(1000, 9999) . ‘.mp3’;
$filePath = $audioDir . ‘/’ . $fileName;$edgeTtsPath = ‘/usr/local/bin/edge-tts’;
// Change voice parameter here for different languages
$voice = “en-US-AriaNeural”; // Default: English (US)$command = “{$edgeTtsPath} –voice ” . escapeshellarg($voice) . ” –text ” . escapeshellarg($cleanText) . ” –write-media ” . escapeshellarg($filePath) . ” 2>&1″;
exec($command, $execDebug, $returnStatus);
if (file_exists($filePath)) {
$audioUrl = ‘audio/’ . $fileName;
}
}// STEP 3: Return JSON response to client
echo json_encode([
‘reply’ => $replyText,
‘audio_url’ => $audioUrl,
‘debug_exec’ => $execDebug
]);
?>
2. Frontend Interface (index.html)
The frontend contains dynamic audio controls, interactive speech recording capabilities, relative timestamps (1m ago), and an audio toggle.
<!DOCTYPE html>
<html lang=”en”>
<head>
<meta charset=”UTF-8″>
<meta name=”viewport” content=”width=device-width, initial-scale=1.0″>
<title>Voice AI Assistant</title>
<style>
* { box-sizing: border-box; }
body { font-family: ‘Segoe UI’, Tahoma, Geneva, Verdana, sans-serif; max-width: 650px; margin: 30px auto; padding: 20px; background-color: #f4f7f6; }.header-container { display: flex; justify-content: space-between; align-items: center; margin-bottom: 15px; }
h2 { margin: 0; color: #333; font-size: 22px; }.auto-play-toggle { display: flex; align-items: center; gap: 8px; font-size: 14px; color: #555; background: #e9ecef; padding: 6px 12px; border-radius: 20px; cursor: pointer; }
.auto-play-toggle input { cursor: pointer; }.chat-box { border: 1px solid #ddd; height: 420px; overflow-y: auto; padding: 15px; margin-bottom: 15px; border-radius: 10px; background: #fff; box-shadow: 0 2px 5px rgba(0,0,0,0.05); }
.msg-wrapper { display: flex; flex-direction: column; margin: 12px 0; }
.msg-wrapper.user { align-items: flex-end; }
.msg-wrapper.bot { align-items: flex-start; }.msg-row { display: flex; align-items: center; gap: 8px; max-width: 85%; }
.msg { padding: 10px 14px; border-radius: 12px; line-height: 1.5; word-wrap: break-word; font-size: 15px; }
.user .msg { background: #007bff; color: white; border-bottom-right-radius: 2px; }
.bot .msg { background: #e9ecef; color: #333; border-bottom-left-radius: 2px; }.speaker-btn { background: #e0e0e0; border: none; border-radius: 50%; width: 32px; height: 32px; display: flex; align-items: center; justify-content: center; cursor: pointer; font-size: 14px; transition: 0.2s; }
.speaker-btn:hover { background: #ccc; }.time-stamp { font-size: 11px; color: #888; margin-top: 4px; padding: 0 4px; }
.input-group { display: flex; gap: 10px; }
input[type=”text”] { flex: 1; padding: 12px; border-radius: 6px; border: 1px solid #ccc; font-size: 16px; }
button { padding: 12px 18px; border-radius: 6px; background: #007bff; color: white; border: none; font-size: 16px; cursor: pointer; transition: 0.2s; }
button:hover { background: #0056b3; }
button:disabled { background: #aaa; cursor: not-allowed; }.mic-btn { background: #28a745; }
.mic-btn.recording { background: #dc3545; animation: pulse 1s infinite; }
@keyframes pulse { 0% { opacity: 1; } 50% { opacity: 0.5; } 100% { opacity: 1; } }
</style>
</head>
<body><div class=”header-container”>
<h2>AI Voice Assistant</h2>
<label class=”auto-play-toggle”>
<input type=”checkbox” id=”autoPlayToggle” checked>
<span>Auto Voice 🔊</span>
</label>
</div><div class=”chat-box” id=”chatBox”></div>
<div class=”input-group”>
<input type=”text” id=”userInput” placeholder=”Ask a question…” onkeydown=”if(event.key===’Enter’) sendMessage()”>
<button id=”micBtn” class=”mic-btn” onclick=”toggleRecording()”>🎙️</button>
<button id=”sendBtn” onclick=”sendMessage()”>Send</button>
</div><script>
let currentAudio = null;
let mediaRecorder = null;
let audioChunks = [];
let isRecording = false;function timeAgo(date) {
const seconds = Math.floor((new Date() – date) / 1000);
if (seconds < 10) return “Just now”;
if (seconds < 60) return `${seconds}s ago`;
const minutes = Math.floor(seconds / 60);
if (minutes < 60) return `${minutes}m ago`;
const hours = Math.floor(minutes / 60);
if (hours < 24) return `${hours}h ago`;
return date.toLocaleTimeString([], { hour: ‘2-digit’, minute: ‘2-digit’ });
}setInterval(() => {
document.querySelectorAll(‘.time-stamp’).forEach(el => {
const timestamp = el.getAttribute(‘data-time’);
if (timestamp) {
el.innerText = timeAgo(new Date(parseInt(timestamp)));
}
});
}, 10000);async function sendMessage() {
const inputField = document.getElementById(‘userInput’);
const message = inputField.value.trim();
if (!message) return;appendMessage(message, ‘user’);
inputField.value = ”;const formData = new FormData();
formData.append(‘message’, message);await sendRequest(formData);
}async function toggleRecording() {
const micBtn = document.getElementById(‘micBtn’);if (!isRecording) {
try {
const stream = await navigator.mediaDevices.getUserMedia({ audio: true });
mediaRecorder = new MediaRecorder(stream);
audioChunks = [];mediaRecorder.ondataavailable = e => audioChunks.push(e.data);
mediaRecorder.onstop = async () => {
const audioBlob = new Blob(audioChunks, { type: ‘audio/webm’ });
appendMessage(‘🎤 Sent Voice Message’, ‘user’);const formData = new FormData();
formData.append(‘audio’, audioBlob, ‘voice.webm’);await sendRequest(formData);
};mediaRecorder.start();
isRecording = true;
micBtn.classList.add(‘recording’);
micBtn.innerText = ‘⏹️ Stop’;
} catch (err) {
alert(“Microphone access denied: ” + err);
}
} else {
mediaRecorder.stop();
isRecording = false;
micBtn.classList.remove(‘recording’);
micBtn.innerText = ‘🎙️’;
}
}async function sendRequest(formData) {
const sendBtn = document.getElementById(‘sendBtn’);
const micBtn = document.getElementById(‘micBtn’);
const autoPlay = document.getElementById(‘autoPlayToggle’).checked;sendBtn.disabled = true;
micBtn.disabled = true;try {
const res = await fetch(‘api.php’, {
method: ‘POST’,
body: formData
});const data = await res.json();
const reply = data.reply || ‘No response received.’;appendMessage(reply, ‘bot’, data.audio_url);
if (autoPlay && data.audio_url) {
playAudio(data.audio_url);
}} catch (err) {
console.error(“Fetch Error:”, err);
appendMessage(‘Server Connection Error.’, ‘bot’);
} finally {
sendBtn.disabled = false;
micBtn.disabled = false;
}
}function appendMessage(text, type, audioUrl = null) {
const chatBox = document.getElementById(‘chatBox’);
const now = new Date();
const timeString = timeAgo(now);const wrapper = document.createElement(‘div’);
wrapper.className = `msg-wrapper ${type}`;const row = document.createElement(‘div’);
row.className = ‘msg-row’;const msgDiv = document.createElement(‘div’);
msgDiv.className = ‘msg’;
msgDiv.innerText = text;row.appendChild(msgDiv);
if (type === ‘bot’ && audioUrl) {
const speakerBtn = document.createElement(‘button’);
speakerBtn.className = ‘speaker-btn’;
speakerBtn.innerHTML = ‘🔊’;
speakerBtn.title = ‘Play/Pause Audio’;
speakerBtn.onclick = () => playAudio(audioUrl);
row.appendChild(speakerBtn);
}const timeDiv = document.createElement(‘div’);
timeDiv.className = ‘time-stamp’;
timeDiv.setAttribute(‘data-time’, now.getTime());
timeDiv.innerText = timeString;wrapper.appendChild(row);
wrapper.appendChild(timeDiv);chatBox.appendChild(wrapper);
chatBox.scrollTop = chatBox.scrollHeight;
}function playAudio(url) {
if (currentAudio && !currentAudio.paused && currentAudio.src.endsWith(url)) {
currentAudio.pause();
return;
}if (currentAudio) {
currentAudio.pause();
currentAudio.currentTime = 0;
}currentAudio = new Audio(url);
currentAudio.play().catch(e => {
console.log(“Autoplay blocked or audio error:”, e);
});
}
</script>
</body>
</html>
How to Change Languages (Multi-Language Support)
To switch the spoken AI voice to another language, update the $voice variable in api.php:
Target Language Neural Voice Parameter ($voice)
English (US) Hindi (India) Urdu (Pakistan) Chinese (Mandarin) Burmese (Myanmar) |