US2025373627A1PendingUtilityA1
Vulnerabilities and Protections in Large Language Models
Est. expiryMay 30, 2044(~17.8 yrs left)· nominal 20-yr term from priority
H04L 63/1416H04L 63/145
50
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Large Language Model (LLM) security includes monitoring an LLM; detecting an attack on the LLM and defining an attack type of a plurality of attack types based on the monitoring, providing a notification of the attack; and causing a defense to the attack based on the attack type. Advantageously, the security can be configured to be executed between a user outside of the LLM. Further, the security can be configured to defend against multi-turn attacks.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for large language model security comprising steps of:
inline monitoring a Large Language Model (LLM); detecting an attack on the LLM and defining an attack type of a plurality of attack types based on the monitoring; providing a notification of the attack; and causing a defense to the attack based on the attack type.
2 . The method of claim 1 , wherein the monitoring includes monitoring a user input to the LLM.
3 . The method of claim 1 , wherein the defense includes blocking the user input to the LLM.
4 . The method of claim 1 , wherein the plurality of attack types includes prompt hacking and adversarial attack.
5 . The method of claim 4 , wherein prompt hacking is one of a prompt injection and a jailbreaking attack.
6 . The method of claim 4 , wherein the adversarial attack is one of a backdoor attack and a data poisoning attack.
7 . The method of claim 1 , wherein causing the defense includes causing any of a prevention-based defense and a detection-based defense.
8 . The method of claim 1 , wherein any of the monitoring, detecting, providing a notification, and causing the defense is performed by an intermediate system before a query reaches the large language model.
9 . The method of claim 1 , wherein the defense includes one of removing, altering, and redesigning an output.
10 . The method of claim 1 , wherein the detection includes one of response-based detection and prompt-based detection.
11 . The method of claim 1 , wherein the defense includes one of system-mode self-reminder prompts, smooth LLM, black-box defense, and pretrained language model defense.
12 . A non-transitory computer-readable medium comprising instructions that, when executed, cause one or more processors to perform steps of:
inline monitoring a large language model (LLM); detecting an attack defining an attack type of a plurality of attack types on the large language model based on the monitoring; providing a notification of the attack; and causing a defense to the attack based on a type of the attack type.
13 . The non-transitory computer-readable medium of claim 12 , wherein the defense includes blocking a user input to the LLM.
14 . The non-transitory computer-readable medium of claim 12 , wherein the plurality of attack types includes prompt hacking and adversarial attack.
15 . The non-transitory computer-readable medium of claim 12 , wherein causing the defense comprising causing any of prevention-based defense and detection-based defense.
16 . The non-transitory computer-readable medium of claim 12 , wherein any of the monitoring, the detecting, providing a notification, and causing the defense is performed by an intermediate system before a query reaches the large language model.
17 . The non-transitory computer-readable medium of claim 12 , wherein the defense includes one of removing, altering, and redesigning an output.
18 . The non-transitory computer-readable medium of claim 12 , wherein the detection includes one of response-based detection and prompt-based detection.
19 . The non-transitory computer-readable medium of claim 12 , wherein the defense includes one of system-mode self-reminder prompts, smooth LLM, black-box defense, and pretrained language model defense.
20 . The non-transitory computer-readable medium of claim 12 , wherein the defense includes a three step clustering based defense comprising representation learning, clustering, and filtering.Join the waitlist — get patent alerts
Track US2025373627A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.