Products

Solutions

Resources

About Us

Deepfake Researchers: "We Can No Longer Automatically Trust Audio"

Kamil Malinka and Anton Firc


Cloning a voice now takes a few seconds of recording and a cheap subscription. Researchers Kamil Malinka and Anton Firc from Brno University of Technology explain in an interview why deepfakes work even when they're far from perfect, and why every company should "break" its own employees before a real attacker does. 


Kamil Malinka and Anton Firc are researchers at Brno University of Technology, and members of the Security@FIT research group, which focuses on information technology security. Their research focuses primarily on the security aspects of artificial intelligence, such as the risks associated with the misuse of deepfakes. As part of their teaching activities, they offer cybersecurity courses in which they systematically integrate new trends and technologies, including AI.


Deepfakes are a significant threat to businesses and government institutions. What are the most common attacks today? 

Kamil Malinka: On an organizational level, it's identity thefts. They often result in high-value fraudulent wire transfers or the leakage of sensitive data. For the banking sector, a successful attack on voice biometrics can be particularly severe — leading not only to direct financial losses but also to significant reputational damage.

Government bodies, on the other hand, are more likely to be targeted by deepfakes designed to sway public opinion, interfere with elections, or, in the worst-case scenarios, serve as weapons of hybrid warfare. 


Convincing deepfakes once required costly hardware and programming skills. How easy is it to create a high-quality voice or video clone now?

Anton Firc: Technical expertise is no longer a requirement for creating deepfakes. All it takes is a monthly subscription of a few dozen dollars for a commercial speech synthesizer, and the user can handle everything through a user-friendly web interface. In fact, to generate a convincing clone, a few seconds of sample audio or a single photograph is often more than enough. 

Kamil Malinka: Beyond paid services, there is a vast amount of open-source models, which attackers can fine-tune to fit their specific needs. What's more, our experience shows that an attack doesn't even require flawless, high-quality deepfakes to be successful. Context and social engineering often do the rest. 


Can attackers now easily modify a voice or face in real time during a live video call without the other party noticing? 

Kamil Malinka: Real-time facial modification such as face-swapping has actually been feasible on standard, off-the-shelf hardware for several years now. We routinely use this technology ourselves during penetration testing of facial biometrics systems and in our educational workshops. 

Anton Firc: When it comes to voice, the situation is a bit different. While we can generate extremely high-quality voice deepfakes using Text-to-Speech (TTS) methods, these aren't entirely practical for live, interactive conversations. An attacker would either need pre-recorded responses for various scenarios or would have to artificially delay the conversation to buy time for synthesizing an answer. 

Kamil Malinka: Because of this limitation, attackers typically rely on a hybrid strategy: they use a pre-rendered line to introduce a new person into the conversation (for instance, CEO introducing "our new legal counsel" and setting the context). After that initial hook, the impersonator simply steps in and continues the scenario using their natural voice without further modification. 


Kamil Malinka and Anton Firc


Some employees have transferred millions after receiving a fake voice command from their "CEO". How can companies defend themselves against such social engineering?

Anton Firc: The standard best practice is to implement multi-factor authentication. In other words, never authorize transactions based solely on a voice call — no matter how convenient it may seem. Always require an additional verification factor. While voice can remain a secondary part of identity verification, it should never be the single point of trust, especially for sensitive operations or high-value financial transfers. 

Kamil Malinka: That secondary factor could be, for example, a cryptographically signed email containing the full transaction details, or requiring the use of authenticated video channels where the identity of the other party can be properly validated. Some companies are even moving toward phasing out voice-based authentication altogether. Crucially, continuous internal training and employee education are key to raising awareness about these attack vectors and reducing the likelihood of staff falling victim to social engineering tactics. 


For governments, deepfakes may be an even greater threat. Are they already being used in cyber warfare or to destabilize public administration?

Kamil Malinka: Deepfakes have already become a routine component of hybrid warfare and disinformation operations. They are a powerful weapon used to sway public opinion, undermine trust in public institutions, and create chaos among the general population during critical moments. 

Anton Firc: A well-known example includes attempts to interfere in elections in Slovakia. We also witnessed an attempt to influence enemy military actions during the opening days of the invasion of Ukraine. In high-stakes situations like these, the danger lies not only in the deepfake content itself, but also in the speed at which it spreads and the strategic timing of its release. 


Human ears can’t spot deepfakes anymore and every institution is now vulnerable to these types of attacks. Download Phonexia's free whitepaper for cross-industry blueprints to build your deepfake defenses.


Generative AI developers say their guardrails prevent cloning the voices of celebrities or politicians. How hard is it for criminals to bypass them?

Kamil Malinka: On a positive note, many companies began introducing these protective measures voluntarily, without needing strict legislative mandates. Most major providers of speech synthesis platforms, for instance, check audio inputs against known voices of globally prominent public figures. 

Anton Firc: On the other hand, the effectiveness of these defenses varies widely. Without giving away step-by-step instructions, our internal security testing of selected commercial services revealed that we could usually find workable ways to bypass these guardrails with relative ease. 


When organizations and government institutions want to actively defend against deepfake threats, what steps should they take? 

Kamil Malinka: The key starting point is conducting a comprehensive risk assessment that specifically accounts for this emerging threat. Deepfakes manifest in many different forms across various attack vectors, so relying on a generic, one-size-fits-all approach doesn't work.  

For instance, a bank utilizing voice biometrics faces entirely different threats and defenses than a manufacturing company protecting its trade secrets, or an online casino executing age-verification checks during customer onboarding. 

Anton Firc: However, it's crucial to emphasize that having a security policy on paper is simply not enough. Organizations need to safely "break" their own people, processes, and technologies before a real-world attacker does. 

In practice, this means actively testing whether employees fall for a deepfake call from a spoofed executive, whether the helpdesk grants unauthorized access over the phone, or whether a synthetic identity bypasses digital onboarding. These kinds of deepfake penetration tests quickly reveal critical vulnerabilities that standard paper-based risk assessments often miss. 


Kamil Malinka and Anton Firc


Mandatory digital watermarking for AI-generated audio and video is widely discussed. Is it a viable long-term solution, or can attackers easily bypass it?

Anton Firc: It is definitely a step in the right direction to improve the overall situation, but watermarking will not be a silver bullet on its own. Cybersecurity in this space will have to rely on a defense-in-depth strategy combining multiple layers of protection. Various valuable initiatives that focus on addressing these exact standards are already emerging — such as the C2PA (Coalition for Content Provenance and Authenticity). 

Kamil Malinka: I remain somewhat skeptical, however. Ensuring that a watermark cannot be stripped away is technically formidable — whether through secondary AI processing, injecting noise, re-recording ambient audio, or simply via lossy compression formats. 

Anton Firc: Furthermore, freely accessible open-source models present another significant challenge. Enforcing mandatory watermarking or guardrails on open-source frameworks where users have full control over the code will be extremely difficult to achieve.


How do you see deepfakes evolving in the coming years? Will they become an even greater threat? 

Anton Firc: We have been tracking the quality trajectory of deepfakes for several years now. From the perspective of an average user, we reached a tipping point about two years ago where generated deepfakes became virtually indistinguishable from real media. In the near future, telling authentic audio or video apart from synthetic media will only become harder. Voice is particularly vulnerable because it often operates without visual context, and people naturally default to trusting what they hear in everyday communication. 

Kamil Malinka: Speech generators, synthesizers, and voice modifiers are technologically converging and increasingly rely on shared underlying architectures. Put simply, the neural network at the end of the process won't care whether its input is original human speech, a modified voice, or raw text — it will produce a highly convincing audio output regardless. That is why we must prepare for an era where we can no longer automatically trust audio or video at face value. Verifying the source and origin of content will become essential. 


You spend your days researching deepfakes and cybercrime. Do you ever feel paranoid, and how do you keep a healthy trust in the digital world?

Kamil Malinka: The key is not to panic. In cybersecurity, we always operate on the principle that absolute security simply doesn't exist. Deepfakes aren't introducing an entirely new paradigm; they are just another "new kid on the block" expanding an already broad arsenal of cyber threats. At their core, they are simply another sophisticated tool designed to trick people. That's why we need to view deepfakes as a natural evolution of our broader security reality, rather than an isolated threat. 

Start Detecting Deepfakes Today


Just like the researchers say, countering deepfakes takes a dual strategy. When organizations combine disciplined human safeguards with advanced, forensic-grade AI detection tools like the Phonexia Speech Platform, they can take back control of their data, safeguard their operations, and confidently establish a reliable baseline of truth in a world where authenticity can't be assumed.


Stay Close to Phonexia's Innovation

Stay Close to

Phonexia's Innovation

Join our newsletter for exclusive product news, events, case studies,

and breakthroughs in voice biometrics and speech recognition.

Join our newsletter for exclusive

product news, events, case studies,

and breakthroughs in voice biometrics

and speech recognition.

By subscribing, you agree to our Privacy Policy. You can unsubscribe anytime.

By subscribing, you agree to our Privacy Policy.

You can unsubscribe anytime.