HimalCyberX
AI Security

AI-Powered Phishing and Deepfakes: How the Attacks Actually Work

Deepfake fraud hit $1.1 billion in US losses in 2025 alone. Here's the real attack mechanics behind cases like Arup — and what actually stops them.

HimalCyberX Research9 min read
AI-Powered Phishing and Deepfakes: How the Attacks Actually Work

AI-Powered Phishing and Deepfakes: How the Attacks Actually Work

Deepfake-related fraud losses in the United States reached an estimated $1.1 billion in 2025, roughly triple the $360 million reported the year before. This isn't a slow-building trend — it's a step change caused by a specific technical shift: the tools needed to convincingly impersonate a real person's voice and face have gone from requiring specialized skill and expensive equipment to being available as a monthly subscription that costs less than a streaming service.

This article breaks down how these attacks are actually built and executed, using real documented cases, and then walks through the controls that meaningfully reduce risk.

The Technical Pipeline Behind a Voice Clone

Modern voice cloning doesn't require hours of recorded speech. Commercial and open-source voice synthesis tools can produce a working clone from as little as three seconds of source audio, achieving roughly 85% accuracy against the target's real voice in independent testing. Attackers don't need to record their target directly — a few seconds pulled from a company earnings call, a conference talk on YouTube, a podcast appearance, or even a LinkedIn video clip is enough raw material. The entire acquisition and cloning pipeline can be assembled with consumer-accessible tools costing on the order of $30 a month for API access to a voice synthesis service.

Video deepfakes follow a parallel path: real-time face-swap and video-generation tools can now drive a live, interactive video call rather than just producing a static pre-rendered clip, replicating facial movement, expression, and lip-sync closely enough to hold a real-time conversation and respond naturally to questions.

Case Study: The Arup Deepfake Video Call ($25.6 Million)

The case that put this threat on the map happened in January 2024. A finance employee at UK-based engineering firm Arup, working from the company's Hong Kong office, received what appeared to be a message from the company's UK-based CFO describing an urgent, confidential transaction. Following standard security practice, the employee didn't act on the message alone — they requested a video call to verify it, exactly what security awareness training recommends.

That's precisely the step the attackers had anticipated. The employee joined a video call where every other participant — the CFO and several senior colleagues — was an AI-generated deepfake, built from publicly available footage of the real executives. Believing the verification had succeeded, the employee authorized 15 separate wire transfers in a single day, totaling HK$200 million (approximately US$25.6 million). The funds were never recovered. Arup confirmed the incident and reported it to Hong Kong police.

The case is instructive precisely because the victim did the "right" thing by industry-standard training — verifying an unusual request via video call — and the verification channel itself was the thing compromised. This is the core lesson that shapes effective defense: verification only works if the channel used to verify is one the attacker cannot also fabricate.

This Pattern Is Now Routine, Not Exceptional

Arup was the case that made headlines, but it's no longer an outlier. Newer reporting has documented additional incidents at a comparable or greater scale, including a Fortune 500 case in early 2026 involving losses of roughly $28 million. According to the FBI's IC3 2025 annual report, business email compromise losses that included a confirmed deepfake audio or video component rose 312% year over year, with average losses from AI-augmented BEC exceeding $4.1 million per incident, compared to roughly $1.3 million for traditional text-only BEC. Separately, one industry study found that 85% of surveyed companies experienced at least one deepfake-related security incident in the preceding twelve months, and vishing (voice phishing) attacks using cloned voices surged 1,633% in the first quarter of 2025 alone.

The threat has also expanded beyond financial fraud. In May 2025, the FBI's IC3 issued a public warning that malicious actors were using AI-generated voice messages to impersonate senior US government officials, targeting both the officials themselves and their contacts to extract sensitive information or redirect communications — demonstrating that the same underlying technique scales from stealing money to conducting espionage without any change to the core method.

Why Human Judgment Fails Here

It's tempting to believe that with enough training, people can learn to spot a fake. The data says otherwise, consistently, across multiple independent studies. A meta-analysis covering 56 separate studies found that people correctly identify high-quality deepfake video only about 24.5% of the time — worse than random chance. Separately, industry survey data has found that roughly 70% of people admit they cannot reliably tell whether a voice they're hearing is real or synthetically cloned. Detection performance in these studies doesn't meaningfully improve with warning or motivation; the synthetic media is simply good enough that the normal cues people rely on to judge authenticity no longer function. Any security strategy that depends on employees "just paying closer attention" is building on a foundation the research doesn't support.

How Real Liveness Detection Technology Works

Because human detection fails, organizations handling high-stakes identity verification increasingly rely on dedicated liveness detection technology rather than a reviewer's judgment. This isn't a single technique — modern systems typically combine several methods:

  • Passive liveness detection analyzes a single image or short video clip (like a selfie) for digital signals invisible to the human eye — compression artifacts, inconsistent skin texture under different lighting, unnatural eye reflections, and other statistical fingerprints left behind by face-swap algorithms or diffusion-based generation models. Because it requires no special action from the user, it has much higher completion rates in practice — one documented case saw completion rates rise from about 60% with an active, prompt-based liveness check to over 95% after switching to passive detection.

  • Active/challenge-based liveness detection asks the user to perform an action (turning their head, following an on-screen prompt) or uses a controlled illumination sequence during capture to confirm the response is happening in real time, rather than being a replay or an injected pre-recorded clip.

  • Voice liveness detection, used in call center and phone-based verification, analyzes content-agnostic vocal features — pitch, rhythm, timbre, unnatural pauses or tonal inconsistencies — during a live conversation to flag synthetic speech in near real time, rather than only after the call has ended.

Vendor-reported accuracy for detecting known deepfake generation techniques under laboratory conditions frequently exceeds 99%. However, multiple independent analyses caution that this accuracy consistently declines once systems face compressed, real-world audio and video. Detection systems trained on limited demographic data can show uneven accuracy across different accents and racial groups — meaning liveness detection reduces risk substantially but should not be treated as an infallible technical fix on its own.

Control 1: Out-of-Band Verification That Can't Be Faked by the Same Attack

The single highest-leverage defense, based directly on how the Arup case succeeded despite "verification," is ensuring the verification channel is genuinely independent of the channel the request arrived through. If a request arrives by video call, verify it through a pre-registered phone number dialed from a separate system — not a number provided in the same message or call. Practical implementations include pre-agreed code words for sensitive transactions known only through prior, separately-verified communication, and callback procedures that only use internally registered contact information, never a number or address supplied as part of the request itself.

Control 2: Mandatory Delay and Dual-Approval for High-Value Transfers

Time-delay requirements on unusual or high-value transactions — even a short mandatory holding period — give a second reviewer or an automated fraud system a chance to flag an anomaly before funds actually leave the organization. Combined with dual-approval requirements for transfers above a defined threshold, this directly targets the pattern seen in the Arup case, where a single authorized individual approved 15 transfers in one day without an independent second check.

Control 3: Deploying Liveness and Detection Technology at Points of Real Risk

For organizations with meaningful exposure — frequent large wire transfers, executive-level communications, customer-facing call centers — deploying dedicated liveness detection or real-time deepfake detection technology at those specific high-risk points is now a mainstream, commercially available control, not a research curiosity. The right deployment model depends on the channel: passive liveness detection for identity-verification workflows, and real-time voice/audio analysis for call center and phone-based verification scenarios.

Control 4: Behavior-Based Training, Not Detection-Based Training

Because the underlying generation technology keeps improving, training people to spot specific visual or audio "tells" has a short shelf life — whatever tell exists today may be fixed in the next model update. More durable training focuses on a consistent behavioral habit: treat any unusual request involving money, credentials, or sensitive data as requiring out-of-band verification by default, regardless of how convincing or urgent the request seems or how senior or familiar the requester appears to be. This reframes the goal from "detect the fake" to "always verify through a separate channel," which remains effective even as the underlying deepfake technology improves.

Control 5: An Incident Response Playbook Specific to This Threat

Every organization handling financial transactions or sensitive communications should have a documented, rehearsed process for a suspected deepfake-driven social engineering attempt: who to notify immediately, how to freeze a pending transaction before it clears, and how to preserve evidence — call recordings, message logs, timestamps — for later investigation. Regulatory context is shifting quickly here too: the EU AI Act's Article 50 provisions impose transparency and digital watermarking obligations on providers of AI-generated synthetic media, and NIS2 now explicitly interprets voice-based social engineering as a vector organizations must have controls against, meaning documented response procedures increasingly aren't just good practice but a compliance expectation in some jurisdictions.

Conclusion

The Arup case is often retold as a story about how convincing the technology has become, but the more useful lesson is structural: the fraud succeeded because the verification method itself could be faked by the same attack it was meant to catch. Effective defense doesn't depend on getting better at spotting fakes — it depends on building verification processes and technical controls that don't rely on human perception in the first place.

Refrences

Share Article

Newsletter

Stay Ahead of the Threat

Weekly cybersecurity intelligence, research and practical security guides.

No spam. Unsubscribe anytime. Privacy Policy