US12482164B2ActiveUtilityA1

Method for protecting anonymity from media-based interactions

Assignee: GoBeyondPriority: Apr 27, 2022Filed: Apr 27, 2023Granted: Nov 25, 2025
Est. expiryApr 27, 2042(~15.8 yrs left)· nominal 20-yr term from priority
G10L 21/003G10L 21/043G10L 2021/0135G10L 15/05G06F 21/6254G10L 13/033G06T 13/40
28
PatentIndex Score
0
Cited by
12
References
20
Claims

Abstract

Disclosed herein are a computing device, method, and computer-readable medium embodiments for altering video and/or audio data during a media-based interaction to protect anonymity of one or more users in the media-based interaction for the purpose of removing bias from the media-based interaction. In some aspects, video and audio data including facial information and speech data of a user may be streamed during the media-based interaction between client devices. The video and audio data may be altered in real-time during the media-based interaction to change the visual and/or audio aspects of one or more users during the media-based interaction. This alterations to the video and audio data may be based on identifying visual and audio features from a list of one or more visual and audio identifiers associated with age, sex, and/or gender. The alteration results in a new virtual representation presented during the interaction that anonymizes the user.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method for anonymizing a user during a media-based interaction, wherein the media-based interaction comprises video data and audio data, the method comprising:
 receiving the video data, wherein the video data comprises facial information of the user;   receiving the audio data, wherein the audio data comprises speech data associated with the video data;   accessing one or more identifiers, wherein each of the one or more identifiers is associated with one or more bias categories, the one or more identifiers comprising a mapping, wherein each of the one or more identifiers is mapped to a set of one or more synonyms that are neutral with respect to at least one of the one or more bias categories;   applying a natural language filter to the audio data to generate altered speech of the user based on the one or more identifiers, wherein the natural language filter is configured to:
 detect a set of one or more words in the audio data; 
 match a word in the set of one or more words with an identifier in the one or more identifiers; and 
 replace, in the audio data, the word with a mapped synonym in the set of one or more synonyms; 
   identifying one or more facial expressions from the facial information of the user;   generating, based on the one or more facial expressions, a neutral facial representation of the user, wherein the neutral facial representation of the user anonymizes the user;   generating a virtual representation of the user, wherein the virtual representation comprises:
 a visual component that includes the neutral facial representation of the user; and 
 an audio component comprising the altered speech of the user; and 
   transmitting, during the media-based interaction, the virtual representation to one or more computing devices over an electronic network.   
     
     
         2 . The method of  claim 1 , wherein the one or more bias categories comprise age, race, and gender. 
     
     
         3 . The method of  claim 2 , wherein each of the one or more synonyms has a confidence score corresponding to a meaning of the identifier and a meaning of each of the one or more synonyms. 
     
     
         4 . The method of  claim 1 , further comprising applying a bias score to the audio data, wherein the bias score corresponds to a number of identifiers associated with one or more of the one or more bias categories identified in the audio data. 
     
     
         5 . The method of  claim 1 , wherein the visual component of the visual representation is based on at least one of the one or more bias categories corresponding to the user. 
     
     
         6 . The method of  claim 1 , applying the natural language filter further comprises replacing a frequency of the audio data with one or more neutral frequencies, wherein the one or more frequencies are different from a voice frequency associated with the user. 
     
     
         7 . The method of  claim 1 , applying the natural language filter further comprises replacing a tempo of the audio data with one or more neutral tempos, wherein the one or more tempos are different from the tempo of the user's voice. 
     
     
         8 . The method of  claim 1 , wherein the list of identifiers and their associated mappings include at least one customizable feature comprising adding an identifier to the list of identifiers. 
     
     
         9 . A computing device for anonymizing a user during a media-based interaction, wherein the media-based interaction comprises video data and audio data, the computing device comprising:
 a processor, wherein the processor further comprises a processing unit; and   a memory, wherein the memory contains instructions stored thereon that when executed by the processor cause the computing device to:
 receive the video data, wherein the video data comprises facial information of the user; 
 receive the audio data, wherein the audio data comprises speech data associated with the video data; 
 access one or more identifiers, wherein each of the one or more identifiers are associated with one or more bias categories, the one or more identifiers comprising a mapping, wherein each of the one or more identifiers is mapped to a set of one or more synonyms that are neutral with respect to at least one of the one or more bias categories; 
 apply a natural language filter to the audio data to generate altered speech of the user based on the one or more identifiers, wherein the natural language filter is configured to:
 detect a set of one or more words in the audio data; 
 match a word in the set of one or more words with an identifier in the one or more identifiers; and 
 replace, in the audio data, the word with a mapped synonym in the set of one or more synonyms; 
 
 identify one or more facial expressions from the facial information of the user; 
 generate, based on the one or more facial expressions, a neutral facial representation of the user, wherein the neutral facial representation of the user anonymizes the user; 
 generate a virtual representation of the user, wherein the virtual representation comprises: 
 a visual component that includes the neutral facial visual representation of the user; and 
 an audio component comprising the altered speech of the user; and 
   transmit, during the media-based interaction, the virtual representation to one or more computing devices over an electronic network.   
     
     
         10 . The computing device of  claim 9 , wherein the one or more bias categories comprise age, race, and gender. 
     
     
         11 . The computing device of  claim 10 , wherein each of the one or more synonyms has a confidence score corresponding to a meaning of the identifier, a meaning of each of the one or more synonyms, and a meaning of a plurality of words surrounding the identifier. 
     
     
         12 . The computing device of  claim 9 , wherein the memory contains further instructions stored thereon that when executed by the processor cause the computing device to:
 apply a bias score to the audio data, wherein the bias score corresponds to a number of identifiers associated with one or more of the one or more bias categories identified in the audio data.   
     
     
         13 . The computing device of  claim 9 , wherein the visual component of the visual representation is based on at least one of the one or more bias categories corresponding to the user. 
     
     
         14 . The computing device of  claim 9 , wherein the natural language filter further comprises replacing a frequency of the audio data with one or more neutral frequencies, wherein the one or more frequencies are different from a voice frequency associated with the user. 
     
     
         15 . The computing device of  claim 9 , wherein the natural language filter further comprises replacing a tempo of the audio data with one or more neutral tempos, wherein the one or more tempos do not match the tempo of the user's voice. 
     
     
         16 . The computing device of  claim 9 , wherein the list of identifiers and their associated mappings include at least one customizable feature comprising adding an identifier to the list of identifiers. 
     
     
         17 . A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations for anonymizing interaction during a media-based interaction, wherein the media-based interaction comprises video data and audio data, the operations comprising:
 receiving the video data, wherein the video data comprises facial information of the user;   receiving the audio data, wherein the audio data comprises speech data associated with the video data;   accessing one or more identifiers, wherein each of the one or more identifiers are associated with at least one or more bias categories, the one or more identifiers comprising a mapping, wherein each of the one or more identifiers is mapped to a set of one or more synonyms that are neutral with respect to the at least one of the one or more bias categories;   applying a natural language filter to the audio data to generate altered speech of the user based on the one or more identifiers, wherein the natural language filter is configured to:
 detect a set of one or more words in the audio data; 
 match a word in the set of one or more words with an identifier in the one or more identifiers; and 
 replace, in the audio data, the word with a mapped synonym in the set of one or more synonyms; 
   identifying one or more facial expressions from the facial information of the user;   generating, based on the one or more facial expressions, a neutral facial representation of the user, wherein the neutral facial representation of the user anonymizes the user;   generating a virtual representation of the user, wherein the virtual representation comprises:
 a visual component that includes the neutral facial representation of the user; and 
 an audio component comprising the altered speech of the user; and 
   transmitting, during the media-based interaction, the virtual representation to one or more computing devices over an electronic network.   
     
     
         18 . The non-transitory computer-readable medium of  claim 17 , wherein the one or more bias categories comprise age, race, and gender. 
     
     
         19 . The non-transitory computer-readable medium of  claim 18 , wherein each of the one or more synonyms has a confidence score corresponding to a meaning of the identifier, a meaning of each of the one or more synonyms, and a meaning of a plurality of words surrounding the identifier. 
     
     
         20 . The non-transitory computer-readable medium of  claim 17 , further comprising applying a bias score to the audio data, wherein the bias score corresponds to a number of identifiers associated with one or more of the one or more bias categories identified in the audio data.

Join the waitlist — get patent alerts

Track US12482164B2 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.