US2021366471A1PendingUtilityA1

Method and system for processing audio communications over a network

Assignee: TENCENT TECH SHENZHEN CO LTDPriority: Nov 3, 2017Filed: Aug 4, 2021Published: Nov 25, 2021
Est. expiryNov 3, 2037(~11.3 yrs left)· nominal 20-yr term from priority
G10L 15/1822G06F 9/454G06F 40/205G10L 15/00G06F 40/263G06F 40/279H04W 24/06G06F 40/58G10L 15/005G10L 15/26H04L 65/601
53
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method of processing audio communications over a social networking platform, comprising: at a server: receiving a first audio transmission from a second client device in a source language distinct from a default language associated with the first client device; obtaining current user language attributes for the first client device, which are indicative of a current language used for the audio and/or video communication session at the first client device; when the current user language attributes suggest a target language currently used for the audio and/or video communication session at the first client device is distinct from the default language: obtaining a translation of the first audio transmission from the source language into the target language; and sending, to the first client device, the translation of the first audio transmission in the target language to be presented to a user at the first client device.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method of processing audio communications over a social networking platform, the method comprising:
 at a sever that has one or more processors and memory, wherein, through the server, a first client device has established an audio and/or video communication session with a second client device over the social networking platform:
 receiving a first audio transmission from the second client device, wherein the first audio transmission is provided by the second client device in a source language that is distinct from a default language associated with the first client device; 
 obtaining one or more current user language attributes for the first client device, wherein the one or more current user language attributes are indicative of a current language that is used for the audio and/or video communication session at the first client device; 
 in accordance with a determination that the one or more current user language attributes suggest a target language that is currently used for the audio and/or video communication session at the first client device is distinct from the default language associated with the first client device:
 obtaining a translation of the first audio transmission from the source language into the target language; and 
 sending, to the first client device, the translation of the first audio transmission in the target language, wherein the translation is presented to a user at the first client device. 
 
   
     
     
         2 . The method of  claim 1 , wherein the obtaining the one or more current user language attributes and suggesting the target language that is currently used for the audio and/or video communication session at the first client device further comprises:
 receiving, from the first client device, facial features of the current user and a current geolocation of the first client device;   determining a relationship between the facial features of the current user and the current geolocation of the first client device; and   suggesting the target language according to a determination that the relationship meets predefined criteria.   
     
     
         3 . The method of  claim 1 , wherein the obtaining the one or more current user language attributes and suggesting the target language that is currently used for the audio and/or video communication session at the first client device further comprises:
 receiving, from the first client device, an audio message that has been received locally at the first client device;   analyzing linguistic characteristics of the audio message received locally at the first client device; and   suggesting the target language that is currently used for the audio and/or video communication session at the first client device in accordance with a result of analyzing the linguistic characteristics of the audio message.   
     
     
         4 . The method of  claim 1 , further comprising:
 obtaining vocal characteristics of a voice in the first audio transmission; and   according to the vocal characteristics of the voice in the first audio transmission, generating a simulated first audio transmission that includes the translation of the first audio transmission spoken in the target language in accordance with the vocal characteristics of the voice of the first audio transmission.   
     
     
         5 . The method of  claim 4 , wherein the sending, to the first client device, the translation of the first audio transmission in the target language to a user at the first client device includes:
 sending, to the first client device, a textual representation of the translation of the first audio transmission in the target language to the user at the first client device; and   sending, to the first client device, the simulated first audio transmission that is generated in accordance with the vocal characteristics of the voice in the first audio transmission.   
     
     
         6 . The method of  claim 1 , wherein the receiving a first audio transmission from the second client device further comprises:
 receiving two or more audio packets of the first audio transmission from the second client device, wherein the two or more audio packets have been sent from the second client device sequentially according to respective timestamps of the two or more audio packets, and wherein each respective timestamp is indicative of a start time of a corresponding audio paragraph identified in the first audio transmission.   
     
     
         7 . The method of  claim 6 , wherein the obtaining the translation of the first audio transmission from the source language into the target language and sending the translation of the first audio transmission in the target language to the first client device further comprise:
 obtaining respective translations of the two or more audio packets from the source language into the target language sequentially according to the respective timestamps of the two or more audio packets; and   sending a first translation of at least one of the two or more audio packets to the first client device after the first translation is completed and before translation of at least another one of the two or more audio packets is completed.   
     
     
         8 . The method of  claim 6 , further comprising:
 receiving a first video transmission while receiving the first audio transmission from the first client device, wherein the first video transmission is marked with the same set of timestamps as the two or more audio packets; and   sending the first video transmission and the respective translations of the two or more audio packets in the first audio transmission with the same set of timestamps to the first client device such that the first client device synchronously present the respective translations of the two or more audio packets of the first audio transmission and the first video transmission according to the same set of timestamps.   
     
     
         9 . A computer server through which a first client device has established an audio and/or video communication session with a second client device over a social networking platform, the computer server comprising:
 one or more processors;   memory; and   one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for:
 receiving a first audio transmission from the second client device, wherein the first audio transmission is provided by the second client device in a source language that is distinct from a default language associated with the first client device; 
 obtaining one or more current user language attributes for the first client device, wherein the one or more current user language attributes are indicative of a current language that is used for the audio and/or video communication session at the first client device; 
 in accordance with a determination that the one or more current user language attributes suggest a target language that is currently used for the audio and/or video communication session at the first client device is distinct from the default language associated with the first client device:
 obtaining a translation of the first audio transmission from the source language into the target language; and 
 sending, to the first client device, the translation of the first audio transmission in the target language, wherein the translation is presented to a user at the first client device. 
 
   
     
     
         10 . The computer server of  claim 9 , wherein the obtaining the one or more current user language attributes and suggesting the target language that is currently used for the audio and/or video communication session at the first client device further comprises:
 receiving, from the first client device, facial features of the current user and a current geolocation of the first client device;   determining a relationship between the facial features of the current user and the current geolocation of the first client device; and   suggesting the target language according to a determination that the relationship meets predefined criteria.   
     
     
         11 . The computer server of  claim 9 , wherein the obtaining the one or more current user language attributes and suggesting the target language that is currently used for the audio and/or video communication session at the first client device further comprises:
 receiving, from the first client device, an audio message that has been received locally at the first client device;   analyzing linguistic characteristics of the audio message received locally at the first client device; and   suggesting the target language that is currently used for the audio and/or video communication session at the first client device in accordance with a result of analyzing the linguistic characteristics of the audio message.   
     
     
         12 . The computer server of  claim 9 , wherein the one or more programs further include instructions for:
 obtaining vocal characteristics of a voice in the first audio transmission; and   according to the vocal characteristics of the voice in the first audio transmission, generating a simulated first audio transmission that includes the translation of the first audio transmission spoken in the target language in accordance with the vocal characteristics of the voice of the first audio transmission.   
     
     
         13 . The computer server of  claim 12 , wherein the sending, to the first client device, the translation of the first audio transmission in the target language to a user at the first client device includes:
 sending, to the first client device, a textual representation of the translation of the first audio transmission in the target language to the user at the first client device; and   sending, to the first client device, the simulated first audio transmission that is generated in accordance with the vocal characteristics of the voice in the first audio transmission.   
     
     
         14 . The computer server of  claim 9 , wherein the receiving a first audio transmission from the second client device further comprises:
 receiving two or more audio packets of the first audio transmission from the second client device, wherein the two or more audio packets have been sent from the second client device sequentially according to respective timestamps of the two or more audio packets, and wherein each respective timestamp is indicative of a start time of a corresponding audio paragraph identified in the first audio transmission.   
     
     
         15 . The computer server of  claim 14 , wherein the obtaining the translation of the first audio transmission from the source language into the target language and sending the translation of the first audio transmission in the target language to the first client device further comprise:
 obtaining respective translations of the two or more audio packets from the source language into the target language sequentially according to the respective timestamps of the two or more audio packets; and   sending a first translation of at least one of the two or more audio packets to the first client device after the first translation is completed and before translation of at least another one of the two or more audio packets is completed.   
     
     
         16 . The computer server of  claim 14 , wherein the one or more programs further include instructions for:
 receiving a first video transmission while receiving the first audio transmission from the first client device, wherein the first video transmission is marked with the same set of timestamps as the two or more audio packets; and   sending the first video transmission and the respective translations of the two or more audio packets in the first audio transmission with the same set of timestamps to the first client device such that the first client device synchronously present the respective translations of the two or more audio packets of the first audio transmission and the first video transmission according to the same set of timestamps.   
     
     
         17 . A non-transitory computer readable storage medium storing one or more programs, the one or more programs, when executed by a computer server through which a first client device has established an audio and/or video communication session with a second client device over a social networking platform, cause the computer server to perform operations comprising:
 receiving a first audio transmission from the second client device, wherein the first audio transmission is provided by the second client device in a source language that is distinct from a default language associated with the first client device;   obtaining one or more current user language attributes for the first client device, wherein the one or more current user language attributes are indicative of a current language that is used for the audio and/or video communication session at the first client device;   in accordance with a determination that the one or more current user language attributes suggest a target language that is currently used for the audio and/or video communication session at the first client device is distinct from the default language associated with the first client device:
 obtaining a translation of the first audio transmission from the source language into the target language; and 
 sending, to the first client device, the translation of the first audio transmission in the target language, wherein the translation is presented to a user at the first client device. 
   
     
     
         18 . The non-transitory computer readable storage medium of  claim 17 , wherein the obtaining the one or more current user language attributes and suggesting the target language that is currently used for the audio and/or video communication session at the first client device further comprises:
 receiving, from the first client device, facial features of the current user and a current geolocation of the first client device;   determining a relationship between the facial features of the current user and the current geolocation of the first client device; and   suggesting the target language according to a determination that the relationship meets predefined criteria.   
     
     
         19 . The non-transitory computer readable storage medium of  claim 17 , wherein the obtaining the one or more current user language attributes and suggesting the target language that is currently used for the audio and/or video communication session at the first client device further comprises:
 receiving, from the first client device, an audio message that has been received locally at the first client device;   analyzing linguistic characteristics of the audio message received locally at the first client device; and   suggesting the target language that is currently used for the audio and/or video communication session at the first client device in accordance with a result of analyzing the linguistic characteristics of the audio message.   
     
     
         20 . The non-transitory computer readable storage medium of  claim 17 , wherein the one or more programs further include instructions for:
 obtaining vocal characteristics of a voice in the first audio transmission; and   according to the vocal characteristics of the voice in the first audio transmission, generating a simulated first audio transmission that includes the translation of the first audio transmission spoken in the target language in accordance with the vocal characteristics of the voice of the first audio transmission.

Join the waitlist — get patent alerts

Track US2021366471A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.