How synthetic voice fraud became the fastest-growing cybercrime, and what it takes to fight back.
January 2024. A finance employee in Hong Kong sat through a video call with his company’s CFO and several colleagues. Everyone looked right. Everyone sounded right. He followed their instructions and wired $25 million to five bank accounts.
Every person on that call was generated. The CFO, the colleagues, the voices. None of them existed.
That story got coverage. What did not make the news: it happens every day now, at smaller scale, in dozens of countries.
Europol’s 2025 threat assessment flagged AI-generated voice as a “critical enabler” of social engineering attacks. The UK’s National Crime Agency was blunt about it in late 2024: synthetic media is no longer a future threat, it is operational.
Three seconds. That is all it takes.
Tools like ElevenLabs, Resemble AI, and a dozen open-source alternatives can clone a human voice from a few seconds of audio. Not a rough approximation. A replica that carries your intonation, your rhythm, the way you hesitate before finishing a thought.
Casual listeners cannot tell the difference. In controlled tests, trained phoneticians struggled too.
A 2023 University College London study asked participants to identify AI-generated speech across multiple languages. Detection accuracy sat around 73%. One out of four fakes slipped through, and those were people who knew they were being tested. In a real phone call with background noise and emotional pressure, forget it.
The most disturbing version of this is not a wire transfer. It is the parent who picks up the phone and hears a child sobbing, then a stranger demanding money before anything bad happens. The voice was cloned from a short clip posted to social media. No child was ever in danger. By the time the parent verifies that, they may already have paid. Investigators call this virtual kidnapping, and the cloned-voice version is spreading because it removes the one thing the old scam lacked: a voice that sounds exactly like the person you love.
This is not a lab problem
In 2019, criminals cloned a CEO’s voice to authorize a $243,000 wire transfer. The employee who took the call said it sounded exactly like his boss. Same slight German accent, same cadence. He had no reason to question it.
The technique did not stay in 2019. In early 2025, fraudsters impersonated Italy’s Defense Minister, Guido Crosetto, using a synthetic version of his voice to call prominent business leaders and request urgent money transfers tied to a fabricated government operation. At least one transfer of close to one million euros went through before the fraud was caught. The targets were experienced people. They recognized the minister’s voice, and that was the whole point.
By 2025, the FBI was tracking cases where scammers cloned the voices of family members to extract ransom payments from panicked parents. The voice says “Mom, I have been in an accident.” It sounds real because acoustically it nearly is.
Why traditional defenses fail
Most organizations still use voice as an authentication factor. “Please say your passphrase.” Banks, insurers, government agencies. The assumption was always that a voice is hard to fake. That stopped being true around 2022. The infrastructure has not caught up.
The banks themselves now admit it. In a 2025 BioCatch survey, 91% of financial institutions said they were reconsidering voice verification for high-value transactions because of cloning risk. Several large cloud providers have pulled back from offering voice biometrics as a standalone authentication method. The reason is visible in the attack volume: Pindrop measured a 1,633% jump in deepfake voice fraud against contact centers in the first quarter of 2025 alone.
Voiceprint matching maps a speaker’s vocal characteristics against a stored template. It was built for a world where the threat was a human mimic. Neural synthesis breaks that model entirely. The clone does not approximate the voiceprint, it replicates it. The match scores return high because the system was never designed to ask the right question: is this a real speaker, or a synthesis model?
Standard anti-spoofing caught the first generation of synthetic voices without much trouble. Early text-to-speech left obvious artifacts. Metallic resonance, unnatural pauses, discontinuities you could hear if you paid attention. A basic classifier flagged them reliably. That era ended. The latest neural vocoders produce speech that is smooth, temporally consistent, perceptually natural. The artifacts did not disappear. They migrated to domains that human ears cannot access.
The arms race
The ASVspoof challenge series, run by the International Speech Communication Association since 2015, is the main benchmark for voice anti-spoofing research. Each edition tells the same story: attack quality improves faster than detection quality. ASVspoof 5 showed that the best synthesis systems fool baseline detectors over 90% of the time. The top detection systems still catch most of them, but the gap narrows every cycle.
Why is voice different from image or video deepfakes? A fake photo can be examined. You zoom in, check pixel-level inconsistencies, run it through analysis tools on your own schedule. A voice arrives during a phone call, and you decide whether to trust it in real time. There is no “zoom and enhance” for audio in a live conversation. By the time doubt sets in, the transfer already went through.
The cost of producing a convincing clone has collapsed. In 2020 it required serious compute and expertise. Today a teenager with a laptop and a free trial can do it before dinner. The barrier fell to zero. The damage potential stayed in the millions.
What holds up
The research community landed on something by 2024: no single detection approach survives the full range of synthesis methods. A detector trained on one family of synthesis artifacts misses outputs from another. Too many attack surfaces for one classifier.
What holds up, consistently, is ensemble analysis. Multiple detection methods running side by side, each covering blind spots the others miss. The ASVspoof results confirm this every year. Ensembles beat single models across every metric.
But ensembles are not enough if they only examine one layer of the signal. A voice recording carries information across multiple dimensions. A sophisticated deepfake might reproduce one aspect cleanly and reveal itself through another. You need systems that look at all of it at once.
There is a second problem that gets less attention: traceability. In an insurance investigation or a legal proceeding, “our AI flagged it” is not enough. You need to show what was detected, where in the signal, and on what scientific basis. Black-box detection is a research tool. For regulated industries, every finding needs a citation trail.
What we built at ORAVYS
I founded ORAVYS in Israel in part because I could not find a tool I would have trusted myself. I will be direct about what we are and what we are not. We are not a lie detector. We are not a surveillance tool. We are a voice intelligence platform that produces multi-dimensional analysis of audio recordings.
When a recording comes in, we do not run one model and produce a score. We run thousands of independent analysis engines across multiple scientific paradigms. Each engine targets a different aspect of what makes a voice authentic or synthetic. The output is not binary. It is a forensic profile that shows where anomalies exist, what kind they are, and what the published literature says about them.
We built a citation engine into the platform. Every metric in an ORAVYS report links back to peer-reviewed research. When the system flags an anomaly, the report explains which studies established that feature as a deepfake indicator. A fraud investigator, a judge, an insurance adjuster can trace every finding to its scientific foundation.
Privacy is not negotiable here. GDPR-compliant from day one. Recordings processed and deleted by default. No voiceprints stored, no cross-session speaker linking, no retraining on client data without explicit opt-in. We support Global Privacy Control. All processing stays within EU infrastructure.
I want to be honest about the limits. The synthesis side is moving faster than I expected when I started. The next generation of vocoders is starting to generate outputs that are harder to distinguish from authentic speech. Detection is going to keep moving. We do not always stay ahead.
Who needs this
Insurance fraud costs European insurers an estimated 13 billion euros per year, and voice-based claims keep growing. An adjuster reviewing a recorded statement needs to know: is this person genuinely distressed, or performing? Is this actually the policyholder, or a synthetic voice filing a claim? We give them the tools to answer those questions with documented evidence.
Legal proceedings involve audio recordings more than ever. Depositions, witness statements, recorded threats, contested phone calls. The question of audio authenticity used to come up once in a while. Now it comes up constantly. Attorneys need analysis that withstands Daubert scrutiny, with clear methodology and traceable scientific backing.
Corporate security teams face voice phishing at executive level. The CEO fraud scenario that used to require a talented impersonator now requires three seconds of audio scraped from a conference talk.
Call centers process millions of hours of voice data. Screening for emotional distress, cognitive load, authenticity markers at that volume requires automation that does not sacrifice accuracy.
What comes next
The deepfake voice problem will get worse before it gets better. Synthesis models keep improving. Open-source releases mean every advance becomes available to researchers and criminals at the same time. Regulation is slow. The EU AI Act classifies emotion recognition as high-risk and imposes transparency requirements, with enforcement of its transparency obligations starting 2 August 2026. That helps. But regulation alone will not solve an arms race that moves at the speed of a GitHub push.
What matters now is building detection infrastructure that is rigorous enough to stay ahead, transparent enough to hold up in court, and ethical enough to deserve the trust it requires. Voice analysis should never be deployed without consent. Results should always be presented as probabilistic assessments, not certainties. The humans making the decisions stay in the loop.
The voice is the most personal form of communication we have. That it can now be copied and weaponized at scale is a real problem. The science to fight back exists, and the question is whether we build the systems to apply it fast enough.
I think we can. That is what I am working on.
Eliot Cohen Bacrie is the Founder & CEO of ORAVYS, an Israel-based voice intelligence company building forensic-grade audio analysis for security, insurance, legal, and clinical applications.
Website: https://oravys.com
ORAVYS Group: https://oravysgroup.com Anti-Deepfake: https://anti-deepfake.com Voice Forensic: https://voiceforensic.com Claims Detect: https://claimsdetect.com Voice DD: https://voicedd.com Voice Intuition: https://voiceintuition.com YouTube: https://youtube.com/@oravys
References
- ASVspoof Consortium. (2024). ASVspoof 5: Crowdsourced Speech Data, Deepfakes, and Adversarial Attacks at Scale. ASVspoof 2024 Workshop, co-located with Interspeech 2024.
- BioCatch. (2025). 2025 AI, Fraud, and Financial Crime Survey.
- Europol. (2025). EU Serious and Organised Crime Threat Assessment (EU-SOCTA) 2025.
- FBI Internet Crime Complaint Center. (2025). 2024 Annual Report.
- Groh, M., Epstein, Z., Firestone, C., & Picard, R. (2022). Deepfake detection by human crowds, machines, and machine-informed crowds. Proceedings of the National Academy of Sciences, 119(1).
- Mai, K. T., Bray, S., Davies, T., & Griffin, L. D. (2023). Warning: Humans cannot reliably detect speech deepfakes. PLoS ONE, 18(8), e0285333.
- MarketsandMarkets. (2025). Deepfake AI Market, Global Forecast to 2031.
- Muller, N. M., Czempin, P., Dieckmann, F., Froghyar, A., & Bottinger, K. (2022). Does audio deepfake detection generalize? Proc. Interspeech 2022, 2783–2787.
- Pindrop. (2025). 2025 Voice Intelligence and Security Report.
- Scherer, K. R. (2003). Vocal communication of emotion: A review of research paradigms. Speech Communication, 40(1–2), 227–256.
- Cummins, N., Scherer, S., Krajewski, J., Schnieder, S., Epps, J., & Quatieri, T. F. (2015). A review of depression and suicide risk assessment using speech analysis. Speech Communication, 71, 10–49.
- Tsanas, A., Little, M. A., McSharry, P. E., Spielman, J., & Ramig, L. O. (2012). Novel speech signal processing algorithms for high-accuracy classification of Parkinson’s disease. IEEE Transactions on Biomedical Engineering, 59(5), 1264–1271.
- UK National Crime Agency. (2024). National Strategic Assessment of Serious and Organised Crime 2024.
- Yi, J., Wang, C., Tao, J., et al. (2023). Audio Deepfake Detection: A Survey. arXiv:2308.14970.