Eliot Cohen Bacrie

The Science of Voice: When It Stopped Being Proof

For most of human history, hearing someone was as good as seeing them. That assumption just expired.

A voice used to carry two guarantees at once. It told you who was speaking, and it told you they were there, in that moment, saying that thing. We built a lot on top of those two guarantees without ever writing them down. A phone call from your bank. A voicemail from your boss authorizing a payment. A recording played in a courtroom. A father’s voice on the line saying he needs help. We trusted the sound itself, because for a hundred years the sound could not be faked convincingly enough to matter.

That is the part that changed. Not slowly. In about two years a voice stopped being a reliable signal of either identity or presence, and most of the systems that depend on it have not caught up.

What we were actually relying on

Think about how much of ordinary life runs on the assumption that a voice is hard to forge. Call centers ask you to “say a passphrase” and treat the voice as the lock. Newsrooms air audio clips as evidence of what a public figure said. Courts admit recordings. Families pick up the phone at two in the morning and act on what they hear. None of these were designed as security systems. They are habits, built on a property of the physical world that held true until recently. A specific person’s voice was effectively unique, and you could not produce it on demand without that person.

Strip that property away and the habits keep running anyway, because nobody told them to stop. That gap, between a defense that no longer works and the systems that still rely on it, is where the damage is happening right now.

This is already costing real money and real trust

These are not projections. They already happened.

In early 2024, a finance employee at the engineering firm Arup in Hong Kong joined a video call with people he believed were his colleagues, including the chief financial officer. They were all synthetic. He authorized transfers totaling about 25.6 million dollars. Everything he saw and heard on that call looked and sounded right. That was the problem.

In January 2024, ahead of the New Hampshire primary, thousands of voters received a robocall using a synthetic version of President Biden’s voice telling them not to vote. One phone call, cloned, broadcast at scale, aimed at an election.

And the version that lands closest to home: the grandparent scam. A senior gets a call, hears what sounds exactly like a grandchild in trouble, panicked, asking for money fast. The voice is generated. A few seconds of audio scraped from social media is enough to build it. This one does not need a nation-state or a corporate target. It needs a phone number and someone who loves the person on the other end.

The common thread is not the technology. It is that in every case, the victim did the reasonable thing. They trusted a voice. That used to be safe.

Why detection has to be neutral, and separate

Once you accept that a voice can be generated, you need a way to ask a plain question and get an honest answer. Was this recording produced by a person, or by a machine. Not “does it sound real,” because it will sound real. The whole point of the technology is that it sounds real.

That answer has to come from somewhere with no stake in the outcome. If the company that generated the audio is also the one certifying whether audio is genuine, you do not have a check, you have a marketing department. If the bank that accepts a voice as authorization is also the only judge of whether that voice was authentic, the incentive runs the wrong way. The role that is missing is a neutral arbiter. An independent referee whose only job is to look at a piece of audio and say, with calibrated confidence, real or generated, and to stand behind that the way a forensic lab stands behind a fingerprint analysis.

That is the role I work on. Not building voices. Judging them. The distinction matters more than it sounds, because the value is entirely in the independence.

The deadline is on the calendar

Regulators have started to treat this as the structural problem it is. The EU AI Act’s transparency rules for synthetic media take effect on the 2nd of August 2026. Generated audio and video will carry disclosure obligations. That date does two things at once. It signals that “we did not know it was fake” is about to stop being an acceptable answer for institutions, and it creates demand for someone who can actually make the determination at scale, on the record.

I am not neutral about whether this matters, so take the framing for what it is. But the incidents are real, the regulatory clock is real, and the underlying shift is not reversible. Voices will keep getting easier to generate, not harder.

Where this leaves us

The old world had a quiet, free safety feature built in. You could usually trust your own ears. We are not getting that back. What replaces it is not better ears, it is infrastructure: a way to verify that a recording is what it claims to be, supplied by someone with no reason to lie about the answer.

For most of history the voice was its own proof. From here, proof has to be added on top. That is the work, and the next few years decide how badly it hurts before it gets built.

Eliot Cohen Bacrie. Founder, ORAVYS, Inc.

Sources: Arup deepfake video-call fraud, Hong Kong (Jan 2024, approx. 25.6M USD). Biden robocall, New Hampshire primary (Jan 2024). Grandparent / family-emergency voice scams (widely documented, FTC and FBI advisories). EU AI Act transparency obligations for synthetic media (effective 2 Aug 2026).