7.6 · Sistem Evaluasi per Mode Latihan
Prinsip
Mengevaluasi Shadowing sangat berbeda dari mengevaluasi Debate. Setiap mode memiliki dimensi dan bobot evaluasi yang berbeda - dan framework harus cukup fleksibel untuk menangani ini.
Matriks Evaluasi per Mode
| Mode | Fluency | Accuracy | Vocabulary | Coherence | Interaction | Pronunciation | Task Completion |
|---|---|---|---|---|---|---|---|
| Repeat After Me | ◐ | ◔ | - | - | - | ● | ● |
| Shadowing | ● | - | - | - | - | ◐ | ◐ |
| Role Play | ◐ | ◐ | ◐ | ◐ | ● | ◔ | ● |
| Daily Conversation | ● | ◔ | ◐ | ◐ | ● | ◔ | - |
| Picture Description | ◐ | ◐ | ● | ● | - | ◔ | ◐ |
| Storytelling | ◐ | ◐ | ◐ | ● | ◐ | ◔ | ◐ |
| Opinion Giving | ◐ | ◐ | ● | ● | ◐ | ◔ | ◐ |
| Problem Solving | ◐ | ◐ | ◐ | ● | ● | ◔ | ◐ |
| Discussion | ◐ | ◐ | ◐ | ◐ | ● | ◔ | ◐ |
| Debate | ◐ | ◐ | ● | ● | ● | ◔ | ◐ |
| Create Conversation | ◐ | ◐ | ◐ | ◐ | ● | ◔ | ◐ |
| Create Speech | ● | ◐ | ● | ● | - | ◐ | ● |
Legenda: ● Tinggi | ◐ Sedang | ◔ Rendah | - Tidak dievaluasi
Konfigurasi Evaluasi
Setiap mode didefinisikan sebagai konfigurasi:
{
"mode": "role_play",
"evaluation": {
"dimensions": {
"fluency": { "weight": 0.15, "enabled": true },
"accuracy": { "weight": 0.15, "enabled": true },
"vocabulary": { "weight": 0.15, "enabled": true },
"coherence": { "weight": 0.10, "enabled": true },
"interaction": { "weight": 0.20, "enabled": true },
"pronunciation": { "weight": 0.05, "enabled": true },
"task_completion": { "weight": 0.20, "enabled": true }
},
"passing_threshold": 0.6,
"mastery_threshold": 0.8
}
}
Threshold
| Threshold | Nilai | Artinya |
|---|---|---|
| Below threshold | < 0.4 | Perlu banyak perbaikan |
| Developing | 0.4 - 0.59 | Di jalur yang benar, butuh latihan lagi |
| Competent | 0.6 - 0.79 | Cukup baik, bisa lanjut |
| Mastery | ≥ 0.8 | Menguasai - evidence of mastery bisa dikumpulkan |
Catatan: Threshold ini adalah data internal - pengguna TIDAK melihat skor numerik. Pengguna melihat feedback naratif.
Acceptance Criteria
- Setiap mode memiliki konfigurasi evaluasi yang terdefinisi
- Dimensi evaluasi bisa di-enable/disable per mode
- Bobot setiap dimensi bisa diatur per mode
- Threshold mastery bisa dikonfigurasi (tidak hard-coded)
- Hasil evaluasi tersimpan sebagai data untuk Evidence of Mastery
- Pengguna TIDAK melihat skor numerik - hanya feedback naratif