
The dawn of the vishing!
How AI-driven conversations and synthetic voices could scale vishing, and why trusted phone interactions may need stronger verification on both sides.
How voice phishing builds trust
"Hello, this is Alice from Bank ABC. As part of our due diligence process, and as mandated by the Central Bank, we must validate and update your personal details if anything changes. First, we need to validate your contact details. Please provide your registered mobile number and email address."
"Hello, this is Mark from IT. As part of our ongoing systems monitoring, I can see you haven't changed your password in over six months. Under our security policy, you must change your password every 6 months to protect our systems. For your convenience, I have just sent you an SMS with a link where you can change your password. Have a good day."
Both examples are what we call voice phishing (Vishing). Both examples fall under the so-called Tech-Support scam, or call center scam. A typical pattern of this type of scam is that the victim is called, and the criminal poses as a technical support person, customer service representative, or government employee (police, IRS, etc.) to gain the victim's trust.

According to the most recent FBI IC3 Report, this type of fraud caused losses of more than 1B USD in the U.S. alone in 2022. Another shocking figure is the number of victims, which more than doubled in just the last 3 years, and almost half of all victims are reported to be over 60 years old, whereas they bear 70% of overall losses.
The current situation is bad, so how much worse can it get? Much, much, worse! Let me explain.
The current operating model is quite simple. A fraudster rents an open-space office somewhere in Southeast Asia, most commonly in India, with a huge English-speaking population. In a country where the youth unemployment is 20%+, it is not hard to find several tens of call center "employees," especially if you incentivize them with a proper bonus scheme. Equip them with a computer, phone, and a chair, and you are all set.

As easy as this might sound, it requires some real effort, especially the part handling the employees and staying under the radar of the police. It is also difficult to scale quickly as the actual "work" is done by employees who need to be hired, at least a bit trained, and retained. As this business grows, the general public becomes more aware due to country-wide awareness campaigns. After a while, you might attract the attention of vigilantes and/or social media streamers (e.g., Scammer Payback and many others) who start to disturb your business and cut into your revenue. Still, as the report shows above, this business is growing fast, and despite efforts by global communities and the US, along with the Indian government, there are no signs of slowing down.
From call centres to AI-supported voice scams
But what if we could improve on this business model by leveraging the newest advancements in AI? Let's see whether we can improve the model above a bit. The first thing to focus on is the hard-to-scale part: the call center agent. Could we replace the agent with a Large Language Model? Such a model would need to converse with a potential victim, understand the responses, and lead the conversation to a desired pre-defined outcome. Do we have such a capability today? Well, looking at ChatGPT-4, it would seem we do. Or, if your expectations are high, we are very close. Actually, with a general model like ChatGPT4, we have a far more knowledgeable "call center agent" than one we could hire, not to mention its ability to adapt to a new use case or scenario seamlessly.

OK, we have an agent who can interact with potential victims via prompts, but we need to translate the text into a call conversation. So, can we transform the ChatGPT responses into natural language? Yes, we do! We have several natural Language Processing (NLP) AI text-to-speech models that can generate custom voices on the fly (see the screenshot of customization options from PlayHT above). We can define the language, tone, gender, and other attributes so the generated voice sounds as human-like as possible (say bye-bye to undesired accents). We apply the same approach to voice-to-text translation of human replies, which are fed back to the LLM agent for processing.

The above diagram depicts this new and improved setup, where the human call center agents are replaced by AI models capable of interacting with the potential victims and leading the conversation into the desired end-state - be it credential harvesting or an actual direct financial loss via enforced funds transfer or other defined outcome.
Such a setup would be scalable, elastic, easily moved from one physical location to another, and easily deployed or shut down as needed. The setup wouldn't require a physical location or substantial human involvement and could be easily provided "As A Service" from a "safe jurisdiction."
Now, after going through the above scenario, I hope you understand the logic behind this blog's title. If the above steps are sufficiently fine-tuned and polished, we will see a massive increase in overall vishing losses in the months and years to come.
Verifying who is on the other end
The LLM model might become so authentic and well-trained on even customer-specific data that we will start questioning every phone conversation we have. Today, we as customers have to confirm our identity (e.g., via MFA, tokens, etc.) to assure companies that it is genuinely us interacting with them; in the future, both sides will need to authenticate. Organizations/companies will need a way to confirm to their customers that it is indeed them trying to reach out.
We are indeed living in exciting times, but please be vigilant and spread awareness about the risks to all around you - your family, friends, and colleagues!
References & Further Reading
[1] 2022 Internet Crime Report
2022 reporting year. Page 16 distinguishes tech/customer-support fraud from government impersonation; the combined loss total is not tech-support fraud alone.
[2] Internet Crime Complaint Center annual reports
Official annual-report archive, including the 2018-2022 editions relevant to the historical chart.