Deepfakes are now a well-established security problem. According to a Gartner survey, 35% of organizations say they have been hit by deepfake-based attacks. Signicat instead found that, in the financial sector, fraud attempts using artificial intelligence increased by 2,137% over three years.
The techniques are becoming more and more sophisticated, but according to multimedia forensics specialists the next big threat could be represented by synthetic audio, destined to have significant consequences: from the management of interviews and journalistic statements to fraud, up to impersonation and false requests for help.
Synthetic audio is the alarm of the future
We talked about it with Giulia Boato, full professor at the Department of Information Engineering and Science at the University of Trento. Among the leading Italian experts in multimedia forensics, she is co-founder of TrueBees, a deep-tech startup that has developed a technology capable of reconstructing the entire chain of manipulation of multimedia content from creation to dissemination on social media, and which in the coming months will also launch a specific ‘detection’ system that looks at the new frontier of voice cloning.
Creating a credible voice with today’s technologies takes very little time. What is the path to identifying synthetic audio? What are the difficulties and timing?
Today, creating a credible voice requires little audio material and very accessible tools. To identify it, however, the work starts from the best possible file, ideally from the original one, and different levels are analysed. It is possible to analyze the acoustic signal, the spectrogram, and therefore the frequencies, the pauses, the breathing, the background noise, the coherence of the background noise with the voice, therefore between the voice and the environment, any cuts, compressions.
The difficulty lies in the fact that the files arrive dirty, recorded by the phone, forwarded via WhatsApp, compressed, noisy, perhaps with music or other voices in the background. An initial screening can be done in a few minutes, but a reliable forensic evaluation, especially in a journalistic or legal context, requires more time, accuracy, because it must explain not only the result, therefore what inconsistencies are found, but also how solid that result is and what uncertainties remain.
Does voice cloning hide greater pitfalls than video deepfakes?
The voice has a particular emotional power. When we hear the voice of a child, a parent, a colleague or an authoritative figure, the brain tends to recognize it before even thinking about it. It is an intimate, familiar signal and for this reason it can lower your defenses.
The main risk is impersonation: a person who seems like our child and asks us for help, a fake manager who orders a bank transfer, a politician or a journalist who is attributed with words he never said. Voice cloning fraud works because it combines technology and emotional pressure. The urgency, the fear, the trust.
With TrueBees you have developed a technology capable of reconstructing the entire chain of manipulations of a multimedia content. Are there clues or anomalies that allow us to understand if a voice is human or manipulated?
The analysis on the reconstruction of the manipulation chain, and therefore also on the diffusion of social media, has been done and is available for images and partially for videos, because it is much more complex to analyze a video signal. We are working on the voice now, we currently have an analyzer in the works, a synthetic voice detector, but not yet the possibility of analyzing voices shared on social media. This second step is ongoing.
As regards discrimination, the manipulated voice certainly presents anomalies in the rhythm, intonation, pauses, breathing, and in the way in which emotions are expressed. Sometimes the voice seems credible but is not consistent with the environment. Maybe the background noise changes, the audio level jumps or there are traces of cuts and recoding. The point is that no single detail is enough to tell if an audio is fake. We work in a broader approach and therefore we not only try to look at the content, but we try to analyze its history.
What is the margin of error? Are there any videos or audios that are more difficult to analyze?
The margin of error is not a single number that applies to all cases. It depends on the quality of the file, the duration, the noise, the type of manipulation that has been done, the compression and also whether there is any comparison material available. A long, clean audio, close to the original version, is very different to analyze than a voice message of a few seconds forwarded via Whatsapp by many people, and then recompressed.
The most difficult cases are the short, very compressed, noisy ones, with overlapping voices, music, cuts, filtering. Video is also difficult to analyze if there is low resolution, recompression, social media passages that can erase traces that we analyze. A system is reliable not only if it says yes or no, therefore true or false: what we are trying to do is understand when the evidence is enough to give an answer and when it is not sufficient and the artificial intelligence must also be able to raise its hand and say I don’t know.
Does the human eye still have any use? Are there details a person can see with the naked eye to tell if a video is a deepfake?
The human eye still has a usefulness as an early warning, not as a final judge. It can make us notice small details, lips that are not perfectly synchronized with the voice, unnatural movements, slightly deformed hands, inconsistent shadows, backgrounds that seem unstable or too stable and too regular, sudden changes in quality, expressions that seem less than credible and not very human.
The problem is that these signals are becoming less and less evident, the technologies they generate are improving very quickly and synthetic contents today are precisely built to overcome the visual controls of our eyes. What we have studied about the perception of synthetic faces is that people can indeed very easily make mistakes and remain very sure of their incorrect evaluation. So the human eye serves to suspect or can serve to suspect, but it is absolutely necessary to certify with an instrument and a technical analysis.
A politician, a famous person, but anyone in general, caught on video doing something compromising, can justify themselves by saying “it’s a deep fake”…
It is a very real risk not only for the false to appear true, but also for the true to be dismissed as false. And a person caught in a compromising situation can say it is a deepfake and create enough doubt to even confuse public opinion. In the future what we also hope for, and for which we are working, is to have tools that are also accessible to citizens and relationships. But we must not imagine them as magic wands.
A detector can help carry out an initial check and in sensitive cases we need more robust tools, capable of explaining precisely why we arrive at a certain result. What traces have been found and with what confidence. The direction is to combine detection, provenance and traceability. So not only asking is it true or false, but understanding where the content comes from, how it was produced and what transformations it has potentially undergone.
Deepfakes, audio manipulation, cloning of the human voice. Being a journalist is becoming a “risky” profession. How does verification change in this intricate scenario?
The journalist can no longer limit himself to asking himself if an image, a video, a voice seem true, but there is a call to reconstruct the context, therefore who provided that content, where it appeared the first time, if it can be reconstructed, if there is an original version, what steps the file has undergone, and perhaps even if there are other sources that confirm something with respect to that media.
This does not mean that the journalist must become a forensic engineer, but perhaps for us it means that the editorial offices can equip themselves with protocols or methods, tools, obviously new skills shared with forensic analysts. So, for example, always keep the original files, avoid working only on forwarded content, document every step you take on the file and also declare when a check is not, let’s say, conclusive. The verification becomes less instinctive, more procedural perhaps, but it is good news, because in a very fragile ecosystem the method becomes the public responsibility of us who operate and who respond responsibly to the news.
How are AI scams evolving?
Scams via artificial intelligence are certainly becoming more scalable, credible and more personalized, because previously many frauds were recognizable by linguistic errors, crude images and very robotic voices, while today a scammer can create well-written messages, realistic profiles, false and very credible documents and images, as well as voices very similar to those of real people. The most worrying cases are certainly fake family emergencies, fake managers asking for bank transfers, manipulated video calls, impersonations of authorities or professionals.
The defense is not to become distrustful of everything, but to create simple rules and above all to create awareness in people that there is a need to have a tool to verify and that one must not act under pressure, taking the necessary time, and perhaps agreeing on words or methods of safety and exchange within the family or in the company. The new rule is not to stop trusting, but to learn to make trust verifiable.
Dossier is the exclusive subscription investigative section of The Vermilion. If you want to support our journalistic work and subscribe, visit our showcase by clicking here.