How well do we distinguish AI-generated voices versus human-generated voices? The context of phone scams
DOI:
https://doi.org/10.32735/S0718-22012026000624196Keywords:
Auditory recognition by non-expert listeners, artificial intelligence (AI), voice cloningAbstract
Current technology allows for artificial voice generation. There are numerous software programs available on the web for free that allow reproducing voices from a simple sample of an original speaker. These applications have already been used to conduct telephone scams. We evaluated the performance of non-expert listeners in distinguishing familiar human voices versus voices generated by artificial intelligence (AI) under real-life conditions.
Downloads
References
Arik, S., Chen, J., Peng, K., Ping, W., Zhou, Y. (2018). Neural Voice Cloning with a Few Samples. En S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, & R. Garnett (Eds.), Advances in Neural Information Processing Systems (Vol.31). CurranAssociates,Inc. https://proceedings.neurips.cc/paper_files/paper/2018/file/4559912e7a94a9c32b09d894f2bc3c82-Paper.pdf
Bali, A.S., Basu, N., Weber, P., Rosas-Aguilar, C., Edmond, G., Martire, K.A., Morrison, G.S. (2024). Speaker identification in courtroom contexts - Part III: Groups of collaborating listeners compared to forensic voice comparison based on automatic-speaker-recognition technology. Forensic Science International, 360, 112048. https://doi.org/10.1016/j.forsciint.2024.112048
Chen, T., Kumar, A., Nagarsheth, P., Sivaraman, G., Khoury, E. (2020). Generalization of Audio Deepfake Detection, in Proc. Odyssey 2020 The Speaker and Language Recognition Workshop, 2020, pp. 132-137.
Edmond, G., San Roque, M. (2009). Quasi-justice: ad hoc expertise and identification evidence. Criminal Law J., 33, pp. 8-33.
Edmond, G., Martire, K., San Roque, M. (2011). Unsound law: issues with (expert’) voice comparison evidence. Melbourne Univ. Law Rev., 35, pp. 52-112.
Geradts, Z., Franke, K. (Eds.), Artificial Intelligence (AI) in Forensic Sciences (pp. 3-20). Wiley.
Hughes, V., Rhodes, R. (2018). Questions, propositions and assessing different levels of evidence: Forensic voice comparison in practice. Science & Justice, 58(4), 250-257. https://doi.org/10.1016/J.SCIJUS.2018.03.007
IIElevenLabs (Aplicación de página web). [Software de audio]. [Fecha de último acceso:12/01/25]. https://elevenlabs.io/about
Laub, C.E., Wylie, L.E., Bornstein, B.H. (2013). Can the courts tell an ear from an eye? Legal approaches to voice identification evidence. Law Psychol. Rev. 37, pp. 119-158.
LeCun, Y., Bengio, Y., Hinton, G. (2015). Deep learning. Nature, 521, pp. 436-444. [Review]
Morrison, G:S., Enzinger, E., Zhang, C. (2018). Forensic speech science, in: I. Freckelton, H. Selby (Eds.), Expert Evidence (Ch. 99), Thomson Reuters, Sydney, Australia.
Moustafa, N. (2022). Digital Forensics in the Era of Artificial Intelligence (1st ed.). CRC Press. https://doi.org/10.1201/9781003278962
Morrison G.S., Enzinger E., Hughes V., Jessen M., Meuwly D., Neumann C., Planting S., Thompson W.C., van der Vloed D., Ypma R.J.F., Zhang C., Anonymous A., Anonymous B. (2021). Consensus on validation of forensic voice comparison. Science & Justice, 61, 229-309.https://doi.org/10.1016/j.scijus.2021.02.002
Mcuba, M., Singh, A., Adeyemi Ikuesan, R., Venter, H. (2023). The Effect of Deep Learning Methods on Deepfake Audio Detection for Digital Investigation, Procedia Computer Science, Volume 219, pp. 211-219, ISSN 1877-0509, https://doi.org/10.1016/j.procs.2023.01.283.(https://www.sciencedirect.com/science/article/pii/S1877050923002910)
Napolitano, D. (2020). The Cultural Origins of Voice Cloning. xCoAx. https://www.researchgate.net/publication/342924151
Ormerod, O. (2001). Sounds familiar? Voice identification evidence. Crim. Law Rev. (10), pp. 595-622.
Robson, J. (2018). ‘Lend me your ears’: an analysis of how voice identification evidence. is treated in four neighbouring criminal justice systems. Int. J. Evid. Proof 22, pp. 218-238. https://doi.org/10.1177/1365712718782989.
Rosas, C., Sommerhoff, J., Morrison, G.S. (2019). A method for calculating the strength. of evidence associated with an earwitness’s claimed recognition of a familiar speaker. Science & Justice, 59, pp. 585-596
Rosas, C., Sommerhoff, J., Pacheco, J., Sáez, C. (2020). Yo lo reconocería por su voz… El caso de Emilio Berkhoff. Alpha (Osorno), (51), 137-160. https://dx.doi.org/10.32735/s0718-2201202000051851
Rose, P. (2002). Forensic Speaker Identification, Taylor and Francis, London UK.
Saini, K., Sonone, S.S., Sankhla, M.S., Kumar, N. (Eds.). (2024). Artificial Intelligence in Forensic Science: An Emerging Technology in Criminal Investigation Systems (1st ed.). CRC Press. https://doi.org/10.4324/9781003287810
Saleema, A., Thampi, S. M. (2018). Voice Biometrics: The Promising Future of Authentication in the Internet of Things. In Handbook of Research on Cloud and Fog Computing Infrastructures for Data Science, IGI Global, pp. 360-389.
Simonite, T. (2022). A Zelensky Deepfake Was Quickly Defeated. The Next One Might Not Be. Wired.
Solan, L.M., Tiersma, P.M. (2003). Hearing voices: speaker identification in court. Hastings Law J. 54, pp. 373-435.
Sherrin, C. (2016). Earwitness evidence: the reliability of voice identifications. Osgoode Hall Law J. 52, pp. 819-862. https://digitalcommons.osgoode.yorku.ca/ohlj/ vol52/iss3/3.
Suguna, S.K., Dhivya, M., Paiva, S. (Eds.). (2021). Artificial Intelligence (AI): Recent Trends and Applications (1st ed.). CRC Press. https://doi.org/10.1201/9781003005629
Yarmey, A.D., Yarmey, A.L., Yarmey, M.J., Parliament, L. (2001). Commonsense beliefs and the identification of familiar voices. Applied Cognitive Psychology, 15, pp. 283-299. http://dx.doi.org/10.1002/acp.702
Yarmey, A.D. (2007). The psychology of speaker identification and earwitness memory, in: R.C.L. Lindsay, D.F. Ross, J.D. Read, M.P. Toglia (Eds.), The Handbook of Eyewitness Psychology, Memory for People, vol. II, Lawrence Erlbaum, Mahwah NJ, pp. 101-136. https://doi.org/10.4324/9781315805535.ch5.
Ypma, R.J.F., Ramos, D., Meuwly, D. (2023). AI-based forensic evaluation in court: The desirability of explanation and the necessity of validation. In Geradts, Z., Franke, K. (Eds.), Artificial Intelligence (AI) in Forensic Sciences (pp. 3-20). Wiley.
Zawali, B., Ikuesan, R.A., Kebande, V. R., Furnell, S., A-Dhaqm, A. (2021). Realising a Push Button Modality for Video-Based Forensics. Infrastructures, vol. 6, no. 4, p. 54.
Zhao, L., Chen, F. (2020). Research on voice cloning with a few samples. Proceedings - 2020 International Conference on Computer Network, Electronic and Automation, ICCNEA 2020, 323-328. https://doi.org/10.1109/ICCNEA50255.2020.00073
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Claudia Rosas, Mateo Castro, Jorge Guzmán

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
