Reference standard
We compare model outputs with a clinically meaningful endpoint or validated assessment appropriate to the intended use.
Four principles guide every validation:
We compare model outputs with a clinically meaningful endpoint or validated assessment appropriate to the intended use.
We evaluate data reflecting the intended population, languages, devices and realistic recording conditions.
We test on held-out, external or prospective data that did not influence model development.
We document sensitivity, specificity, AUROC or calibration as appropriate, together with uncertainty, subgroup, language, device and noise checks, and known limitations.
The table reports sensitivity and specificity for eight intended use cases, alongside their reference measures, cohort sizes and evaluated languages. Cohorts were balanced by class (healthy control versus pathological), age and sex, plus weight and smoking habits where applicable.
| Condition | Reference measure | Sensitivity | Specificity | Participants | Languages |
|---|---|---|---|---|---|
| ConditionStress | Reference measureSalivary cortisol | Sensitivity98% | Specificity92% | Participants400 | LanguagesFrench, Italian, English, Spanish, German, Portuguese and Russian |
| ConditionAnxiety | Reference measureGAD-7 | Sensitivity61% | Specificity79% | Participants1,132 | LanguagesFrench, Italian, English, Spanish, German, Portuguese, Chinese and Russian |
| ConditionDepression | Reference measurePHQ-9 | Sensitivity77% | Specificity83% | Participants1,933 | LanguagesFrench, Italian, Chinese, English, Spanish, German, Portuguese, Malay and Russian |
| ConditionParkinson’s disease | Reference measureUPDRS | Sensitivity98% | Specificity96% | Participants1,621 | LanguagesFrench, Slovak, English, Italian, Spanish, German, Arabic and Turkish |
| ConditionAlzheimer’s disease | Reference measureMMSE | Sensitivity97% | Specificity92% | Participants2,488 | LanguagesEnglish, Greek, French, Slovak, Chinese, German and Italian |
| ConditionMild cognitive impairment | Reference measureMoCA | Sensitivity81% | Specificity76% | Participants2,566 | LanguagesSlovak, Mandarin Chinese, English, German and Italian |
| ConditionRespiratory function | Reference measureCOPD, COVID-19 or asthma diagnosis | Sensitivity72% | Specificity82% | Participants30,312 | LanguagesSpanish, English and French |
| ConditionType 2 diabetes | Reference measureHbA1c levels | Sensitivity73% | Specificity70% | Participants1,440 | LanguagesSpanish, English and French |
Supported voice biomarkers can be integrated through the Virtuosis API.
Gervaise L et al. Voice biomarkers for multi-condition health screening using a short natural speech sample: a real-world validation. International Vocal Biomarkers Conference (IVBC). 2026.
Gervaise L et al. Speech-based biomarkers for scalable detection of Alzheimer’s disease, mild cognitive impairment, and Parkinson’s disease. European Journal of Neurology. 2026.
De Kalbermatten C, Gervaise L, Qorraj M. Family voice in rare diseases (FAVOUR-RD). European Conference on Rare Diseases (ECRD). 2026.
Gervaise L. Evaluation of voice-based screening for depression across controlled and real-world speech datasets. Voice AI Symposium. Frontiers. 2026.
Mekki M, Gervaise L. Voice-based depression detection: a bio-inspired and language-invariant model. Applied Machine Learning Days (AMLD). 2026.
Sensitivity is the proportion of participants with the reference condition whom the model correctly identifies; specificity is the proportion without the condition whom it correctly identifies. These values apply to the evaluated cohorts, reference measures and recording conditions shown here and should not be assumed to transfer unchanged to every population, device or intended use.
No. Each condition is evaluated against its own reference measure, cohort, decision threshold and validation design. Sensitivity and specificity therefore describe performance within that specific evaluation; they should not be used to rank conditions or infer that one condition is easier to identify than another.
Virtuosis models have been trained using recordings from many diverse languages, so we expect them to achieve similar performance in languages not listed in the table. We nevertheless monitor performance when deploying in a new language. Our validation datasets also include additional languages represented by smaller groups, typically 10–30 participants. To avoid overinterpreting these smaller samples, the table lists only the main languages evaluated for each condition.