Ai

New AI Model Decodes Keystrokes from Zoom Audio with High Accuracy

New AI Model Decodes Keystrokes from Zoom Audio with High Accuracy

Compiled by the editorial desk with reference to the research paper and comments provided to The Guardian.

In a recent study, researchers from the University of Surrey in the UK demonstrated that an AI model can identify which keys are being pressed on a laptop keyboard with up to 95% accuracy, using only audio recordings from a nearby smartphone or a Zoom call. The findings, which have not yet undergone peer review, underscore a growing cybersecurity threat in an era where microphones are embedded in countless personal devices.

The team trained a neural network classifier by pressing 36 keys on a 2021 MacBook Pro, each key 25 times, while recording the sounds through an iPhone's microphone and through Zoom's audio capture. After training, the model achieved a 95% accuracy rate on the iPhone recordings and 93% on the Zoom audio. According to the paper, this performance surpasses earlier keystroke-detection systems, marking a significant advancement in acoustic side-channel attacks.

Co-author Ehsan Toreini, a lecturer in software security at the University of Surrey, told The Guardian that he expects such attacks to become more accurate over time, noting the widespread presence of smart devices with microphones in homes. The researchers argue that the combination of ubiquitous microphones and modern deep learning techniques creates a dangerous environment where sensitive information, including passwords, could be intercepted by malicious actors.

While the study used a specific laptop model, the authors suggest that most keyboards share similar acoustic characteristics, meaning a single trained model could potentially be effective against a wide range of devices. This would lower the barrier for attackers, who might otherwise need to develop separate models for different keyboard types.

Mitigating the Risk

The researchers point out that the AI model struggles to detect when the shift key is pressed, so incorporating capital letters into passwords could reduce vulnerability. They also recommend varying typing styles, using fingerprint or facial recognition when available, and enabling two-step authentication as general security measures.

For those concerned about privacy during video calls, the simplest advice may be to avoid typing sensitive information while on camera. As the research indicates, the sound of keystrokes alone can be enough for an AI to reconstruct what is being typed.

The study adds to a growing body of work on acoustic side-channel attacks, which exploit sound emissions to extract information. While the technique is not new, the high accuracy achieved with consumer-grade equipment and widely used communication platforms like Zoom raises fresh concerns about digital privacy.

As AI continues to advance, the researchers caution that such attacks will likely become more sophisticated. They call for increased awareness and the development of countermeasures to protect against this emerging threat.