How well do we distinguish AI-generated voices versus human-generated voices? The context of phone scams

Authors

  • Claudia Rosas Universidad Austral de Chile (Chile) https://orcid.org/0000-0002-8544-7965
  • Mateo Castro Universidad Austral de Chile (Chile)
  • Jorge Guzmán Policía de Investigaciones de Chile (Chile)

DOI:

https://doi.org/10.32735/S0718-22012026000624196

Keywords:

Auditory recognition by non-expert listeners, artificial intelligence (AI), voice cloning

Abstract

Current technology allows for artificial voice generation. There are numerous software programs available on the web for free that allow reproducing voices from a simple sample of an original speaker. These applications have already been used to conduct telephone scams. We evaluated the performance of non-expert listeners in distinguishing familiar human voices versus voices generated by artificial intelligence (AI) under real-life conditions.

Downloads

Download data is not yet available.

References

Arik, S., Chen, J., Peng, K., Ping, W., Zhou, Y. (2018). Neural Voice Cloning with a Few Samples. En S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, & R. Garnett (Eds.), Advances in Neural Information Processing Systems (Vol.31). CurranAssociates,Inc. https://proceedings.neurips.cc/paper_files/paper/2018/file/4559912e7a94a9c32b09d894f2bc3c82-Paper.pdf

Bali, A.S., Basu, N., Weber, P., Rosas-Aguilar, C., Edmond, G., Martire, K.A., Morrison, G.S. (2024). Speaker identification in courtroom contexts - Part III: Groups of collaborating listeners compared to forensic voice comparison based on automatic-speaker-recognition technology. Forensic Science International, 360, 112048. https://doi.org/10.1016/j.forsciint.2024.112048

Chen, T., Kumar, A., Nagarsheth, P., Sivaraman, G., Khoury, E. (2020). Generalization of Audio Deepfake Detection, in Proc. Odyssey 2020 The Speaker and Language Recognition Workshop, 2020, pp. 132-137.

Edmond, G., San Roque, M. (2009). Quasi-justice: ad hoc expertise and identification evidence. Criminal Law J., 33, pp. 8-33.

Edmond, G., Martire, K., San Roque, M. (2011). Unsound law: issues with (expert’) voice comparison evidence. Melbourne Univ. Law Rev., 35, pp. 52-112.

Geradts, Z., Franke, K. (Eds.), Artificial Intelligence (AI) in Forensic Sciences (pp. 3-20). Wiley.

Hughes, V., Rhodes, R. (2018). Questions, propositions and assessing different levels of evidence: Forensic voice comparison in practice. Science & Justice, 58(4), 250-257. https://doi.org/10.1016/J.SCIJUS.2018.03.007

IIElevenLabs (Aplicación de página web). [Software de audio]. [Fecha de último acceso:12/01/25]. https://elevenlabs.io/about

Laub, C.E., Wylie, L.E., Bornstein, B.H. (2013). Can the courts tell an ear from an eye? Legal approaches to voice identification evidence. Law Psychol. Rev. 37, pp. 119-158.

LeCun, Y., Bengio, Y., Hinton, G. (2015). Deep learning. Nature, 521, pp. 436-444. [Review]

Morrison, G:S., Enzinger, E., Zhang, C. (2018). Forensic speech science, in: I. Freckelton, H. Selby (Eds.), Expert Evidence (Ch. 99), Thomson Reuters, Sydney, Australia.

Moustafa, N. (2022). Digital Forensics in the Era of Artificial Intelligence (1st ed.). CRC Press. https://doi.org/10.1201/9781003278962

Morrison G.S., Enzinger E., Hughes V., Jessen M., Meuwly D., Neumann C., Planting S., Thompson W.C., van der Vloed D., Ypma R.J.F., Zhang C., Anonymous A., Anonymous B. (2021). Consensus on validation of forensic voice comparison. Science & Justice, 61, 229-309.https://doi.org/10.1016/j.scijus.2021.02.002

Mcuba, M., Singh, A., Adeyemi Ikuesan, R., Venter, H. (2023). The Effect of Deep Learning Methods on Deepfake Audio Detection for Digital Investigation, Procedia Computer Science, Volume 219, pp. 211-219, ISSN 1877-0509, https://doi.org/10.1016/j.procs.2023.01.283.(https://www.sciencedirect.com/science/article/pii/S1877050923002910)

Napolitano, D. (2020). The Cultural Origins of Voice Cloning. xCoAx. https://www.researchgate.net/publication/342924151

Ormerod, O. (2001). Sounds familiar? Voice identification evidence. Crim. Law Rev. (10), pp. 595-622.

Robson, J. (2018). ‘Lend me your ears’: an analysis of how voice identification evidence. is treated in four neighbouring criminal justice systems. Int. J. Evid. Proof 22, pp. 218-238. https://doi.org/10.1177/1365712718782989.

Rosas, C., Sommerhoff, J., Morrison, G.S. (2019). A method for calculating the strength. of evidence associated with an earwitness’s claimed recognition of a familiar speaker. Science & Justice, 59, pp. 585-596

Rosas, C., Sommerhoff, J., Pacheco, J., Sáez, C. (2020). Yo lo reconocería por su voz… El caso de Emilio Berkhoff. Alpha (Osorno), (51), 137-160. https://dx.doi.org/10.32735/s0718-2201202000051851

Rose, P. (2002). Forensic Speaker Identification, Taylor and Francis, London UK.

Saini, K., Sonone, S.S., Sankhla, M.S., Kumar, N. (Eds.). (2024). Artificial Intelligence in Forensic Science: An Emerging Technology in Criminal Investigation Systems (1st ed.). CRC Press. https://doi.org/10.4324/9781003287810

Saleema, A., Thampi, S. M. (2018). Voice Biometrics: The Promising Future of Authentication in the Internet of Things. In Handbook of Research on Cloud and Fog Computing Infrastructures for Data Science, IGI Global, pp. 360-389.

Simonite, T. (2022). A Zelensky Deepfake Was Quickly Defeated. The Next One Might Not Be. Wired.

Solan, L.M., Tiersma, P.M. (2003). Hearing voices: speaker identification in court. Hastings Law J. 54, pp. 373-435.

Sherrin, C. (2016). Earwitness evidence: the reliability of voice identifications. Osgoode Hall Law J. 52, pp. 819-862. https://digitalcommons.osgoode.yorku.ca/ohlj/ vol52/iss3/3.

Suguna, S.K., Dhivya, M., Paiva, S. (Eds.). (2021). Artificial Intelligence (AI): Recent Trends and Applications (1st ed.). CRC Press. https://doi.org/10.1201/9781003005629

Yarmey, A.D., Yarmey, A.L., Yarmey, M.J., Parliament, L. (2001). Commonsense beliefs and the identification of familiar voices. Applied Cognitive Psychology, 15, pp. 283-299. http://dx.doi.org/10.1002/acp.702

Yarmey, A.D. (2007). The psychology of speaker identification and earwitness memory, in: R.C.L. Lindsay, D.F. Ross, J.D. Read, M.P. Toglia (Eds.), The Handbook of Eyewitness Psychology, Memory for People, vol. II, Lawrence Erlbaum, Mahwah NJ, pp. 101-136. https://doi.org/10.4324/9781315805535.ch5.

Ypma, R.J.F., Ramos, D., Meuwly, D. (2023). AI-based forensic evaluation in court: The desirability of explanation and the necessity of validation. In Geradts, Z., Franke, K. (Eds.), Artificial Intelligence (AI) in Forensic Sciences (pp. 3-20). Wiley.

Zawali, B., Ikuesan, R.A., Kebande, V. R., Furnell, S., A-Dhaqm, A. (2021). Realising a Push Button Modality for Video-Based Forensics. Infrastructures, vol. 6, no. 4, p. 54.

Zhao, L., Chen, F. (2020). Research on voice cloning with a few samples. Proceedings - 2020 International Conference on Computer Network, Electronic and Automation, ICCNEA 2020, 323-328. https://doi.org/10.1109/ICCNEA50255.2020.00073

Downloads

Published

2026-06-01

Issue

Section

Articles

How to Cite

How well do we distinguish AI-generated voices versus human-generated voices? The context of phone scams. (2026). ALPHA. Journal of Arts, Literature and Philosophy, 1(62), 209-225. https://doi.org/10.32735/S0718-22012026000624196