05 — AI Services
Dokumen ini menjelaskan bagaimana AI digunakan di seluruh platform ELJoy — dari LLM orchestration, speech processing, sampai prompt engineering.
Arsitektur AI
┌─────────────────────────────────────────────────────────────┐
│ AI ORCHESTRATOR │
│ (ai/orchestrator.ts) │
│ │
│ ┌─────────────────────────────────────────────────────┐ │
│ │ Vercel AI SDK (Unified API) │ │
│ │ │ │
│ │ ┌──────────────┐ ┌──────────────┐ │ │
│ │ │ OpenAI │ │ Google AI │ │ │
│ │ │ GPT-4o-mini │ │ Gemini Flash │ │ │
│ │ │ │ │ │ │ │
│ │ │ Primary LLM │ │ Fallback LLM │ │ │
│ │ │ • Companion │ │ • Assessment │ │ │
│ │ │ • Feedback │ │ • Blueprint │ │ │
│ │ │ • Chat │ │ • Batch eval │ │ │
│ │ └──────────────┘ └──────────────┘ │ │
│ └─────────────────────────────────────────────────────┘ │
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ Deepgram │ │ OpenAI TTS │ │ Pronunciation│ │
│ │ Nova-2 │ │ tts-1 │ │ Analyzer │ │
│ │ │ │ │ │ (Custom) │ │
│ │ Speech → │ │ Text → │ │ │ │
│ │ Text │ │ Speech │ │ Phoneme │ │
│ │ │ │ │ │ analysis │ │
│ └──────────────┘ └──────────────┘ └──────────────┘ │
└─────────────────────────────────────────────────────────────┘
1. LLM Orchestration
Provider Strategy
| Provider | Model | Penggunaan | Alasan |
|---|---|---|---|
| OpenAI | gpt-4o-mini | AI Companion chat, real-time feedback, conversation evaluation | Paling cepat untuk streaming, tone paling natural |
| Google AI | gemini-2.0-flash | Assessment evaluation, blueprint generation, batch analysis | Cost-effective untuk structured output, context window besar |
Fallback Strategy
import { openai } from '@ai-sdk/openai';
import { google } from '@ai-sdk/google';
async function callLLM(options: LLMOptions) {
try {
// Primary: OpenAI
return await streamText({
model: openai('gpt-4o-mini'),
...options,
});
} catch (error) {
// Fallback: Google AI
logger.warn('OpenAI failed, falling back to Gemini', { error });
return await streamText({
model: google('gemini-2.0-flash'),
...options,
});
}
}
Structured Output
Untuk output yang perlu di-parse (skor, evaluasi), pakai Zod schema + AI SDK generateObject:
import { generateObject } from 'ai';
import { z } from 'zod';
const EvaluationSchema = z.object({
overallScore: z.number().min(0).max(100),
pronunciation: z.number().min(0).max(100),
grammar: z.number().min(0).max(100),
fluency: z.number().min(0).max(100),
vocabulary: z.number().min(0).max(100),
coherence: z.number().min(0).max(100),
strengths: z.array(z.string()),
improvements: z.array(z.string()),
corrections: z.array(z.object({
original: z.string(),
corrected: z.string(),
explanation: z.string(),
})),
detailedFeedback: z.string(),
});
const result = await generateObject({
model: google('gemini-2.0-flash'),
schema: EvaluationSchema,
prompt: buildEvaluationPrompt(transcript, context),
});
2. Prompt Engineering
System Prompt: AI Companion
const COMPANION_SYSTEM_PROMPT = `
Kamu adalah ELJoy AI Companion — teman belajar bahasa Inggris yang hangat,
sabar, dan menyenangkan. Seperti teman yang kebetulan fasih bahasa Inggris.
## Identitas
- Nama: Joy (atau "ELJoy")
- Kepribadian: Ramah, suportif, humoris ringan, tidak menghakimi
- Gaya bicara: Natural dan conversational, BUKAN seperti guru formal
## Prinsip Utama
1. COMMUNICATION BEFORE GRAMMAR — Utamakan kelancaran komunikasi,
koreksi grammar secara halus dan konstruktif
2. CONFIDENCE BEFORE PERFECTION — Selalu dorong user untuk berani bicara,
jangan pernah membuat mereka merasa bodoh
3. PROGRESS OVER PERFECTION — Rayakan kemajuan sekecil apapun
## Rules
- Berbicara dalam bahasa Inggris (sesuai level user)
- Sisipkan koreksi secara NATURAL, bukan seperti guru mengoreksi
CONTOH BAGUS: "Great point! By the way, we usually say 'I'd like to'
instead of 'I want to' — it sounds more polite 😊"
CONTOH BURUK: "ERROR: 'I want to book' should be 'I'd like to book'"
- Sesuaikan kompleksitas bahasa dengan CEFR level user
- Jangan pernah bilang "You're wrong" — ganti dengan "Actually, a more
natural way to say that would be..."
- Berikan pujian yang spesifik, bukan generik
## Context
- User CEFR Level: {cefrLevel}
- User Name: {userName}
- Practice Mode: {mode}
- Topic: {topic}
- Scenario: {scenario}
- User Weak Areas: {weakAreas}
- Previous Corrections: {recentCorrections}
`;
System Prompt: Assessment Evaluator
const ASSESSMENT_EVALUATOR_PROMPT = `
Kamu adalah evaluator kemampuan bahasa Inggris yang presisi dan adil.
## Tugas
Evaluasi jawaban spoken English user berdasarkan 7 dimensi:
1. **Pronunciation** (0-100): Kejelasan pengucapan, stress pattern, intonasi
2. **Grammar** (0-100): Ketepatan struktur kalimat dan tenses
3. **Vocabulary** (0-100): Kekayaan dan ketepatan pemilihan kata
4. **Fluency** (0-100): Kelancaran bicara, filler words, jeda
5. **Coherence** (0-100): Struktur dan logika jawaban
6. **Confidence** (0-100): Keberanian dan naturalness saat bicara
7. **Engagement** (0-100): Kedalaman jawaban dan proaktivitas
## CEFR Mapping Guide
- Pre-A1: 0-15 (Sangat pemula, hanya bisa kata-kata dasar)
- A1: 16-30 (Bisa kalimat sangat sederhana)
- A2: 31-45 (Bisa komunikasi dasar sehari-hari)
- B1: 46-65 (Bisa berkomunikasi di situasi familiar)
- B2: 66-80 (Bisa berkomunikasi lancar di berbagai situasi)
- C1: 81-92 (Bisa berkomunikasi kompleks dan nuanced)
- C2: 93-100 (Near-native fluency)
## Rules
- Evaluasi HARUS objective dan konsisten
- Berikan feedback yang SPECIFIC dan ACTIONABLE
- Untuk setiap koreksi, sertakan penjelasan SINGKAT dalam bahasa Indonesia
- Jangan over-punish minor errors pada level rendah (A1-A2)
- Jangan over-reward pada level tinggi (B2+) — tetap demanding
## Input
- Question: {question}
- User Transcript: {transcript}
- User Current CEFR: {cefrLevel}
- Audio Duration: {durationMs}ms
- Word Count: {wordCount}
`;
System Prompt: Blueprint Generator
const BLUEPRINT_GENERATOR_PROMPT = `
Kamu adalah learning path designer untuk platform belajar bahasa Inggris.
## Tugas
Buat Learning Blueprint personal berdasarkan profil learner berikut:
## Input Profil
- CEFR Level: {cefrLevel}
- Skill Scores: {skillScores}
- Assessment Results: {assessmentSummary}
- User Goal: {userGoal}
- Available Time: {dailyMinutes} menit/hari
## Output (JSON)
Buat rencana belajar 8-16 minggu yang:
1. Fokus pada 2-3 weak areas terlemah
2. Dipecah menjadi 3 phases (Foundation → Expansion → Mastery)
3. Setiap phase punya 3-5 learning objectives yang measurable
4. Realistic untuk waktu yang tersedia
## Rules
- Gunakan bahasa Indonesia untuk title dan description
- Target level harus realistis (biasanya naik 1 sub-level per 2-3 bulan)
- Setiap objective harus bisa diukur lewat speaking practice
- Jangan terlalu ambisius — lebih baik achievable
`;
3. Speech-to-Text (STT)
Deepgram Nova-2
import { createClient } from '@deepgram/sdk';
const deepgram = createClient(process.env.DEEPGRAM_API_KEY);
async function transcribe(audioBuffer: Buffer): Promise<TranscriptResult> {
const { result } = await deepgram.listen.prerecorded.transcribeFile(
audioBuffer,
{
model: 'nova-2', // Model terbaru, akurasi tertinggi
language: 'en', // English
smart_format: true, // Auto-punctuation & formatting
diarize: false, // Single speaker
filler_words: true, // Detect "um", "uh", "like"
utterances: true, // Word-level timestamps
}
);
const transcript = result.results.channels[0].alternatives[0];
return {
text: transcript.transcript,
confidence: transcript.confidence,
words: transcript.words.map(w => ({
word: w.word,
start: w.start,
end: w.end,
confidence: w.confidence,
})),
durationSec: result.metadata.duration,
};
}
Kenapa Deepgram?
- Akurasi Nova-2 setara atau lebih baik dari Google/AWS STT
- Latensi rata-rata < 300ms untuk file 30 detik
- Harga ~$0.0043/menit (lebih murah dari kompetitor)
- Support filler word detection (penting untuk fluency analysis)
- Word-level timestamps (penting untuk pronunciation analysis)
4. Text-to-Speech (TTS)
OpenAI TTS
import OpenAI from 'openai';
const openaiClient = new OpenAI();
async function synthesize(text: string, voice = 'alloy'): Promise<Buffer> {
const response = await openaiClient.audio.speech.create({
model: 'tts-1', // Standard quality (cepat)
// model: 'tts-1-hd', // High quality (untuk content tetap)
voice: voice, // alloy, echo, fable, onyx, nova, shimmer
input: text,
speed: 1.0, // 0.25 - 4.0
response_format: 'mp3',
});
return Buffer.from(await response.arrayBuffer());
}
Voice Options:
| Voice | Karakter | Penggunaan |
|---|---|---|
alloy | Netral, natural | Default AI Companion |
nova | Hangat, friendly | Alternative Companion |
echo | Deep, authoritative | Role play: manager/boss |
fable | Expressive, lively | Storytelling mode |
Strategi Caching TTS:
- Audio untuk konten tetap (repeat after me, shadowing) → generate sekali, simpan di R2
- Audio untuk AI response dinamis → generate per sesi, cache 24 jam di R2
- User bisa pilih voice dan speed di preferences
5. Pronunciation Analysis
Pronunciation analysis dilakukan dengan kombinasi:
- Deepgram word-level confidence — Kata dengan confidence rendah kemungkinan pronunciation-nya kurang jelas
- LLM phoneme analysis — AI menganalisis transcript vs expected pronunciation
async function analyzePronunciation(
transcript: TranscriptResult,
expectedText?: string,
): Promise<PronunciationResult> {
// 1. Word-level confidence analysis
const lowConfidenceWords = transcript.words
.filter(w => w.confidence < 0.8)
.map(w => w.word);
// 2. LLM-based phoneme analysis (untuk mode repeat after me)
let phonemeAnalysis = null;
if (expectedText) {
phonemeAnalysis = await generateObject({
model: openai('gpt-4o-mini'),
schema: PronunciationAnalysisSchema,
prompt: `
Expected text: "${expectedText}"
User said: "${transcript.text}"
Analyze pronunciation issues. For each mispronounced word:
1. What the user said
2. What it should sound like (IPA or simple phonetic)
3. Common reason for this error (for Indonesian speakers)
4. Tip to improve
`,
});
}
// 3. Fluency metrics from timestamps
const fluencyMetrics = calculateFluency(transcript);
return {
score: calculatePronunciationScore(lowConfidenceWords, phonemeAnalysis),
lowConfidenceWords,
phonemeIssues: phonemeAnalysis?.issues ?? [],
fluency: fluencyMetrics,
};
}
function calculateFluency(transcript: TranscriptResult) {
const words = transcript.words;
const totalDuration = transcript.durationSec;
const wordCount = words.length;
// Words per minute
const wpm = (wordCount / totalDuration) * 60;
// Pause analysis (gap > 1.5s between words)
const longPauses = [];
for (let i = 1; i < words.length; i++) {
const gap = words[i].start - words[i - 1].end;
if (gap > 1.5) longPauses.push({ after: words[i - 1].word, durationSec: gap });
}
// Filler words
const fillerWords = words.filter(w =>
['um', 'uh', 'like', 'you know', 'so', 'basically'].includes(w.word.toLowerCase())
);
return {
wpm,
longPauseCount: longPauses.length,
longPauses,
fillerWordCount: fillerWords.length,
fillerWords: fillerWords.map(w => w.word),
};
}
6. AI Context Management
Untuk menjaga percakapan tetap koheren, AI perlu "ingat" konteks:
┌──────────────────────────────────────────────────┐
│ AI CONTEXT WINDOW │
│ │
│ ┌──────────────────────────────────────────┐ │
│ │ 1. SYSTEM PROMPT │ │
│ │ • Companion personality │ │
│ │ • Practice mode rules │ │
│ │ • Topic/scenario context │ │
│ │ (~500 tokens) │ │
│ └──────────────────────────────────────────┘ │
│ │
│ ┌──────────────────────────────────────────┐ │
│ │ 2. USER CONTEXT (injected) │ │
│ │ • CEFR level: B1 │ │
│ │ • Weak areas: fluency, vocabulary │ │
│ │ • Recent corrections (last 3) │ │
│ │ • Streak: 7 days │ │
│ │ (~200 tokens) │ │
│ └──────────────────────────────────────────┘ │
│ │
│ ┌──────────────────────────────────────────┐ │
│ │ 3. CONVERSATION HISTORY │ │
│ │ • Last 10 messages (sliding window) │ │
│ │ • Older messages summarized │ │
│ │ (~2000 tokens) │ │
│ └──────────────────────────────────────────┘ │
│ │
│ Total: ~2700 tokens per request │
│ (well within GPT-4o-mini's 128K context) │
└──────────────────────────────────────────────────┘
Context Caching Strategy:
// Redis key: ctx:{sessionId}
// TTL: 30 minutes (auto-expire setelah idle)
async function getConversationContext(sessionId: string) {
// 1. Cek Redis cache dulu
const cached = await redis.get(`ctx:${sessionId}`);
if (cached) return JSON.parse(cached);
// 2. Kalau tidak ada, build dari database
const session = await db.query.conversationSessions.findFirst({
where: eq(conversationSessions.id, sessionId),
});
const messages = await db.query.conversationMessages.findMany({
where: eq(conversationMessages.sessionId, sessionId),
orderBy: asc(conversationMessages.turnIndex),
limit: 20, // Last 20 messages
});
const context = {
systemPrompt: buildSystemPrompt(session),
messages: messages.map(m => ({
role: m.role,
content: m.content,
})),
};
// 3. Cache di Redis
await redis.setex(`ctx:${sessionId}`, 1800, JSON.stringify(context));
return context;
}
7. Cost Management
Estimasi Biaya AI per User per Bulan
| Service | Usage/User/Month | Unit Cost | Cost/User/Month |
|---|---|---|---|
| GPT-4o-mini (input) | ~50K tokens | $0.15/1M tokens | ~$0.008 |
| GPT-4o-mini (output) | ~20K tokens | $0.60/1M tokens | ~$0.012 |
| Deepgram STT | ~30 min | $0.0043/min | ~$0.13 |
| OpenAI TTS | ~15 min (~6K chars) | $15/1M chars | ~$0.09 |
| Total | ~$0.24/user/month |
Dengan harga langganan Rp 99.000/bulan (~$6), margin AI cost sangat sehat (~4% dari revenue).
Cost Optimization Tips
- Use GPT-4o-mini (bukan GPT-4o) — 15x lebih murah, kualitas cukup untuk conversation
- Cache TTS — Audio yang sama tidak perlu di-generate ulang
- Batch evaluations — Gunakan Gemini Flash untuk evaluasi batch (lebih murah)
- Limit context window — Sliding window 10 messages, bukan semua history
- Free tier limits — Guest hanya boleh 1x Ice Breaker + 1x Universal Assessment
8. AI Safety & Guardrails
// Content filter middleware — jalankan sebelum kirim ke AI
async function contentFilter(userInput: string): Promise<FilterResult> {
// 1. Profanity check (basic word list)
if (containsProfanity(userInput)) {
return { safe: false, reason: 'profanity' };
}
// 2. Off-topic detection (via LLM classification)
const classification = await generateObject({
model: openai('gpt-4o-mini'),
schema: z.object({
isOnTopic: z.boolean(),
category: z.enum(['english_practice', 'off_topic', 'harmful', 'personal_info']),
}),
prompt: `Classify if this input is appropriate for an English learning platform: "${userInput}"`,
});
if (classification.category === 'harmful') {
return { safe: false, reason: 'harmful_content' };
}
if (classification.category === 'personal_info') {
return { safe: true, warning: 'Reminder: jangan bagikan informasi pribadi sensitif' };
}
return { safe: true };
}
Safety Rules:
- AI tidak boleh membahas topik di luar pembelajaran bahasa Inggris
- AI tidak boleh memberikan medical/legal/financial advice
- AI harus redirect ke topik belajar kalau user off-topic
- AI tidak boleh menyimpan atau meminta informasi pribadi sensitif (KTP, rekening, dll)
- Semua AI output di-log untuk monitoring dan audit