Detection of site phishing using neural network-enabled site image analysis leveraging few-shot learning
Abstract
Website phishing detection is enabled using a siamese neural network. One twin receives a query image associated with a website page. The other twin receives a subset of a set of reference website images together with positive (phishing) examples that were used to train the networks, the subset of reference website images having been determined by applying an identifier associated with a brand of interest. The operation of applying the identifier significantly reduces the relevant search space for the inferencing task. If the inferencing determines a sufficient likelihood that the website page is a phishing page, control signaling is generated to control a system to take a given mitigation action.
Claims
exact text as granted — not AI-modified1 . A method of protecting an online system from a phishing attack, comprising:
training a neural network using a set of training data comprising screenshots of reference website pages together with a set of associated positive phishing examples; during training, and in lieu of utilizing a validation set against which a level of generalization of the neural network is inspected during the course of training to identify over-fitting, monitoring a fit quality of the neural network with respect to the training data using a gradient disparity metric; following training of the neural network:
receiving a query screenshot associated with a website page; and
applying the query screenshot to a first instance of the neural network, and applying at least some of the screenshots of the reference website pages to a second instance of the neural network, thereby generating an output that indicates a likelihood that the website page is a phishing page.
2 . The method as described in claim 1 , wherein the first instance and the second instance together comprise a siamese neural network.
3 . The method as described in claim 1 , wherein the output is generated in a timescale measured in seconds from receipt of the query screenshot.
4 . The method as described in claim 1 , wherein training implements a Probably Approximately Correct (PAC)-Bayesian learning framework.
5 . The method as described in claim 1 , wherein the neural network is trained while subjected to a triplet loss function.
6 . The method as described in claim 5 , wherein the triplet loss function trains the neural network to closely embed features of a first class while maximizing a distance between embeddings of one or more second classes distinct from the first class.
7 . The method as described in claim 1 , wherein the screenshots of reference website pages together with the set of associated positive phishing examples comprise an entirety of the training data available to the online system.
8 . The method as described in claim 1 , wherein the online system is accessed through a content delivery network (CDN).
9 . The method as described in claim 1 , wherein, responsive to the output indicating a likelihood that the website page is a phishing page, generating signaling information to cause a control system to take a given action with respect to the website page.
10 . The method as described in claim 1 , wherein the at least some of the screenshots of the reference website pages applied to the second instance of the neural network are determined by applying an identifier associated with a brand of interest.
11 . A computer program product comprising a non-transitory computer readable medium, the computer readable medium comprising computer program code configured to execute in one or more hardware processors to protect an online system from a phishing attack, the computer program code configured to:
train a neural network using a set of training data comprising screenshots of reference website pages together with a set of associated positive phishing examples; during training, and in lieu of utilizing a validation set against which a level of generalization of the neural network is inspected during the course of training to identify over-fitting, monitor a fit quality of the neural network with respect to the training data using a gradient disparity metric; following training of the neural network:
receive a query screenshot associated with a website page; and
apply the query screenshot to a first instance of the neural network, and apply at least some of the screenshots of the reference website pages to a second instance of the neural network, thereby generating an output that indicates a likelihood that the website page is a phishing page.Join the waitlist — get patent alerts
Track US2025294055A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.