Biden, Drake, Swift, Johansson, Trump. None of them could stop it.
I build voice detection systems for a living. Over the past three years I have written more than 3,000 analysis engines that examine a recording across many dimensions. I thought I understood the problem.
Last winter someone sent me two short audio clips and asked me to tell which one was real. I could not. Same intonation, same micro-hesitations on the same vowels. One was a real speaker. The other was a neural vocoder. My ears, trained on thousands of hours of voice analysis, did not catch it.
That moment shifted how I think about this work. It stopped being a technical problem and started being an operational one.
Here are five cases that show why.
The phone call
On January 21, 2024, thousands of New Hampshire voters picked up the phone and heard Joe Biden. The voice was unmistakable. “What a bunch of malarkey,” he said, telling them to stay home on primary day. A political consultant named Steve Kramer had commissioned the call using AI. The voice generation itself cost about a dollar on ElevenLabs. The FCC fined Kramer $6 million and settled with Lingo Telecom, the carrier, for $1 million (FCC, 2024).
Frankly, I do not think the fine deters anyone serious. Six million sounds heavy in the abstract. In practice, Kramer reached somewhere between 5,000 and 25,000 people (estimates from Nomorobo call-tracking data; Kramer himself claims the lower end). The damage-to-cost ratio is significant, and any consultant willing to take the risk once will take it again with a burner LLC.
What bothers me more is what the call proved about infrastructure. There is no layer in the telecom stack that checks whether a voice is human. None at all. The FCC declared AI-generated robocalls illegal on February 8, 2024, under the Telephone Consumer Protection Act. It was necessary, and it was already too late.
The Secret Service protects the president’s body. Nobody protects his voice.
The song
This one is short, because the fact does most of the work.
In April 2023, a track called “Heart on My Sleeve” appeared on Spotify, TikTok, and YouTube. It sounded exactly like Drake and The Weeknd. Billboard reported over 600,000 streams on Spotify before Universal Music Group had it pulled on April 17, 2023.
Neither artist had recorded a single note.
An anonymous producer going by Ghostwriter977 had trained an AI on their voices and published the result. The Recording Academy rewrote its Grammy eligibility rules in June 2023: no human authorship, no eligibility. And the music industry walked into a question it had never needed to answer. If anyone can produce a Drake song, what is a Drake song worth?
I do not know the answer, honestly. The labels are still litigating it.
The scam nobody noticed
Taylor Swift sold Le Creuset cookware in January 2024. Or rather, that is what millions of users on Facebook, Instagram, and TikTok came to believe. Deepfake ads showed her endorsing the products at a steep discount. Her voice. Her likeness. A link to buy. Le Creuset itself had to issue a statement disavowing any partnership (New York Times, January 9, 2024; CBS News, 2024).
The products did not exist. Swift had no idea the campaign was running.
What makes this case different is the business model behind it. The Biden call was political sabotage. The Drake track was a stunt. The Swift ads were a business, with margins. Someone was generating revenue, at scale, with a stolen voice as the sales engine. MrBeast posted about the same thing happening to him on X in October 2023. So did Tom Hanks. Oprah too.
Celebrity trust gets built over decades. It can be borrowed in a few seconds of audio and turned against the celebrity’s own audience, in a weekend.
The minister
The four cases above involve famous voices used against the public. The next one shows what happens when a famous voice gets used against other people’s money.
In early 2025, Italian business leaders started getting phone calls from what sounded like Guido Crosetto, the Italian Defense Minister. The voice was right. The story was urgent and official: an Italian journalist had been kidnapped abroad, the government needed a discreet ransom advanced, the funds would be reimbursed by the Bank of Italy. At least one businessman wired close to a million euros to a foreign account before the fraud came apart. Crosetto himself went public to warn that his voice was being used. Italian prosecutors opened an investigation (reported by Italian and international press, February 2025).
Notice the shift. Nobody here wanted to embarrass the minister. The minister was the credential. His recognizable voice was the thing that switched off the victim’s caution, so the actual target, a private wire transfer, would go through. The voice was not the prize. It was the key.
That is the pattern that scales. A trusted voice opens a door that a stranger never could.
The company that did not take no for an answer
In May 2024, OpenAI launched a ChatGPT voice called “Sky.” Within hours, people noticed. It sounded like Scarlett Johansson. Not vaguely. The intonation, the warmth, the pacing.
Johansson’s statement on May 20 was precise: Sam Altman had personally asked her to voice the system. She declined. Two days before the demo, her agent received a second request. Before they could respond, the system was already out there. Altman tweeted a single word that day, “her.” Johansson called the situation “shocking” and retained legal counsel. OpenAI paused the voice on May 19 (Washington Post, May 20, 2024).
This case unsettles me more than the others. A scammer cloning a celebrity is a crime, full stop. An $80 billion company cloning an actress after she explicitly said no is something different. It is a signal about where the boundary of consent sits inside the AI industry, which is to say, nowhere clearly defined.
US federal law does not protect vocal likeness specifically. Tennessee passed the ELVIS Act in March 2024, the most targeted state law so far. The EU AI Act requires disclosure when AI imitates real people (Article 50), with enforcement starting August 2026.
If Johansson, with her lawyers, her platform, and her name, struggled to protect her voice from a Fortune 500 company, I genuinely do not know what an ordinary person is supposed to do.
The most cloneable voice on Earth
Donald Trump.
Decades of television, radio, rallies, courtrooms, podcasts. Thousands of hours in broadcast quality. A vocal signature so distinctive that a synthesis model can fit it on a laptop in an afternoon.
He is the easiest voice cloning target that exists. That, on its own, is a national security problem.
AI-generated audio of Trump making fabricated statements circulated on social media throughout 2024. Deepfake audio and video promoting cryptocurrency scams appeared on YouTube and Telegram, documented by Bitdefender, Elliptic, and several other cybersecurity firms. The FBI’s IC3 reported $4.57 billion in investment fraud losses in 2023, with AI-generated content as a growing vector (FBI IC3, 2024). I cannot tell you what fraction of that traces specifically to deepfake Trump audio. The attribution data does not exist yet, and I am skeptical it ever will at the granularity people want.
In September 2023, the NSA, FBI, and CISA published a joint cybersecurity information sheet on deepfake threats to organizations. That tells you something about how seriously the intelligence community is taking the problem.
Three seconds
All these cases follow a pattern I have spent three years studying. A few seconds of reference audio. A synthesis model that reproduces a speaker’s vocal character. A generator that produces new speech in the target voice, often with so few detectable artifacts that the obvious tells from 2021 are simply gone.
ElevenLabs, Resemble AI, OpenVoice, Coqui TTS. These are not prototypes. A University College London study found that listeners correctly identified synthetic speech only about 73% of the time (Mai et al., 2023, PLOS ONE). For a binary question, that is barely better than a coin flip. For high-quality clones with more reference audio, accuracy dropped further.
A cloning subscription costs $5 to $22 per month. The open-source versions are free. The compute fits on a laptop.
Here is the part that should worry you more than any celebrity case. You do not need to be famous to be cloned. You need to have spoken in public, anywhere, for a few seconds. A voicemail greeting. A wedding speech someone posted. A reel. The same technology that faked Biden runs the “grandparent scam,” where a senior gets a call in a grandchild’s exact voice, panicked, asking for money now. It runs the virtual kidnapping calls, where a parent hears their own child crying and a stranger naming a ransom. The child is fine and at school. The voice came off a public clip. These calls do not make headlines individually. They happen daily.
I want to be honest about something. Detection is not a solved problem either. My own systems hit very high accuracy on known conditions, but research has shown that certain codec pipelines not seen during training can drop performance to near-chance levels on out-of-distribution data. This is an arms race. The attackers are currently cheaper and faster than the defenders. I am not sure that gap closes any time soon.
What I think has to change
Deepfake voice fraud increased 1,300% in 2024 according to Pindrop’s Voice Intelligence Report, from about one incident per month to seven per day across their monitored networks. The same firm reported deepfake vishing up more than 1,600% in the first quarter of 2025. Deloitte projects generative AI fraud losses hitting $40 billion by 2027. The curve is not flattening.
Detection needs to move to the point of reception. Not after the call, not after the wire transfer, not after the election. At the moment the voice enters the system. That is what I am building at ORAVYS, a voice intelligence platform that runs thousands of forensic engines on a single recording and returns a detailed assessment. Detection alone is reactive by nature, though. It tells you a voice is fake after someone already tried to use it.
The generators keep getting better. Every month, the synthesis models improve. They adapt. They learn to bypass whatever detection catches them today. Detection is a race that never finishes, because the other side never stops running.
The fundamental shift is to stop asking “is this fake?” and start asking “is this proven real?”
We watermark images. We sign documents. We fingerprint people at borders. Voices, the most natural and trusted form of human communication, have no proof of origin. No signature. No chain of custody. There is the gap.
I built VoiceSign to start closing it. It is an open-source voice watermarking system that lets a speaker embed an inaudible cryptographic signature into their voice. If you receive a call from someone who has signed their voice, you can verify it really is them and not a clone. If the signature is missing, you know to be careful. Think of it as HTTPS for the human voice.
The AI can improve as much as it wants. It can produce a clone that sounds indistinguishable from the original. It will never be able to generate the cryptographic signature of the real speaker. The signature is not in the voice. It is in the math. And the math does not care how good the synthesis gets.
VoiceSign only works if the speaker opts in, though. It does not protect people who have never heard of it, which is most people. The real solution, the one I keep circling back to, is structural. A global, secured voice registry. A place where any person can deposit their authentic voiceprint, encrypted and legally recognized. Not a database that anyone can query to clone you, the opposite: a reference point that makes unauthorized cloning traceable and prosecutable.
If your voice is registered, any clone can be compared against the original and flagged. If a call comes in claiming to be you, the system checks against your registered record. The voice ends up tethered to its owner the way a title deed tethers property to a person.
This does not exist yet. It probably should.
Biden, Drake, Swift, Crosetto, Johansson, Trump. Some of the most recognizable voices on the planet, plus a defense minister whose voice was used to drain a stranger’s bank account. All compromised. None of them could prevent it, because the system was built to trust voices, not to verify them.
If their voices are not safe, I would not assume yours is either. And the grandparent scam already proves it does not check whether you are famous.
Eliot Cohen Bacrie is the founder of ORAVYS, the voice intelligence platform. Thousands of forensic engines. Based in Israel.
oravys.com
References
- Federal Communications Commission. (2024). Declaratory Ruling: AI-Generated Voices in Robocalls. February 8, 2024.
- Federal Communications Commission. (2024). Forfeiture Order against Steve Kramer ($6M) and Lingo Telecom settlement ($1M). FCC-24–84.
- Billboard. (2023). Universal Music Takes Down AI Drake/Weeknd Track. April 18, 2023.
- Recording Academy. (2023). Grammy Awards: AI-Generated Music Eligibility Guidelines. June 16, 2023.
- New York Times. (2024). No, That’s Not Taylor Swift Peddling Le Creuset Cookware. January 9, 2024.
- CBS News. (2024). Celebrity Deepfake Scam Ads on Social Media. 2024.
- Reuters / ANSA. (2025). Italian Defense Minister Crosetto warns of AI voice scam targeting business leaders. February 2025.
- Washington Post. (2024). Scarlett Johansson Says OpenAI Copied Her Voice After She Said No. May 20, 2024.
- Tennessee General Assembly. (2024). ELVIS Act. Signed March 21, 2024.
- European Union. (2024). AI Act, Article 50.
- NSA, FBI, CISA. (2023). Contextualizing Deepfake Threats to Organizations. September 12, 2023.
- FBI Internet Crime Complaint Center. (2024). 2023 Internet Crime Report. $4.57 billion in investment fraud losses.
- Mai, K.T., Bray, S., Davies, T., Griffin, L.D. (2023). Warning: Humans cannot reliably detect speech deepfakes. PLOS ONE, 18(8), e0285333.
- Pindrop. (2025). Voice Intelligence and Security Report. 1,300% increase in deepfake fraud in 2024; deepfake vishing up over 1,600% in Q1 2025.
- Deloitte Center for Financial Services. (2024). Fighting Fraud in the Age of Generative AI.
- Bitdefender. (2024). Crypto Scams Using Deepfake Political Figures on YouTube.
- Elliptic. (2024). AI Political Deepfake Scams Targeting Crypto Users.