Method and system for processing audio communications over a network
Abstract
A method of processing audio communications over a social networking platform, comprising: at a server: receiving a first audio transmission from a second client device in a source language distinct from a default language associated with the first client device; obtaining current user language attributes for the first client device, which are indicative of a current language used for the audio and/or video communication session at the first client device; when the current user language attributes suggest a target language currently used for the audio and/or video communication session at the first client device is distinct from the default language: obtaining a translation of the first audio transmission from the source language into the target language; and sending, to the first client device, the translation of the first audio transmission in the target language to be presented to a user at the first client device.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method of processing audio communications over a social networking platform, the method comprising:
at a sever that has one or more processors and memory, wherein, through the server, a first client device has established an audio and/or video communication session with a second client device over the social networking platform:
receiving a first audio transmission from the second client device, wherein the first audio transmission is provided by the second client device in a source language that is distinct from a default language associated with the first client device;
obtaining one or more current user language attributes for the first client device, wherein the one or more current user language attributes are indicative of a current language that is used for the audio and/or video communication session at the first client device;
in accordance with a determination that the one or more current user language attributes suggest a target language that is currently used for the audio and/or video communication session at the first client device is distinct from the default language associated with the first client device:
obtaining a translation of the first audio transmission from the source language into the target language; and
sending, to the first client device, the translation of the first audio transmission in the target language, wherein the translation is presented to a user at the first client device.
2 . The method of claim 1 , wherein the obtaining the one or more current user language attributes and suggesting the target language that is currently used for the audio and/or video communication session at the first client device further comprises:
receiving, from the first client device, facial features of the current user and a current geolocation of the first client device; determining a relationship between the facial features of the current user and the current geolocation of the first client device; and suggesting the target language according to a determination that the relationship meets predefined criteria.
3 . The method of claim 1 , wherein the obtaining the one or more current user language attributes and suggesting the target language that is currently used for the audio and/or video communication session at the first client device further comprises:
receiving, from the first client device, an audio message that has been received locally at the first client device; analyzing linguistic characteristics of the audio message received locally at the first client device; and suggesting the target language that is currently used for the audio and/or video communication session at the first client device in accordance with a result of analyzing the linguistic characteristics of the audio message.
4 . The method of claim 1 , further comprising:
obtaining vocal characteristics of a voice in the first audio transmission; and according to the vocal characteristics of the voice in the first audio transmission, generating a simulated first audio transmission that includes the translation of the first audio transmission spoken in the target language in accordance with the vocal characteristics of the voice of the first audio transmission.
5 . The method of claim 4 , wherein the sending, to the first client device, the translation of the first audio transmission in the target language to a user at the first client device includes:
sending, to the first client device, a textual representation of the translation of the first audio transmission in the target language to the user at the first client device; and sending, to the first client device, the simulated first audio transmission that is generated in accordance with the vocal characteristics of the voice in the first audio transmission.
6 . The method of claim 1 , wherein the receiving a first audio transmission from the second client device further comprises:
receiving two or more audio packets of the first audio transmission from the second client device, wherein the two or more audio packets have been sent from the second client device sequentially according to respective timestamps of the two or more audio packets, and wherein each respective timestamp is indicative of a start time of a corresponding audio paragraph identified in the first audio transmission.
7 . The method of claim 6 , wherein the obtaining the translation of the first audio transmission from the source language into the target language and sending the translation of the first audio transmission in the target language to the first client device further comprise:
obtaining respective translations of the two or more audio packets from the source language into the target language sequentially according to the respective timestamps of the two or more audio packets; and sending a first translation of at least one of the two or more audio packets to the first client device after the first translation is completed and before translation of at least another one of the two or more audio packets is completed.
8 . The method of claim 6 , further comprising:
receiving a first video transmission while receiving the first audio transmission from the first client device, wherein the first video transmission is marked with the same set of timestamps as the two or more audio packets; and sending the first video transmission and the respective translations of the two or more audio packets in the first audio transmission with the same set of timestamps to the first client device such that the first client device synchronously present the respective translations of the two or more audio packets of the first audio transmission and the first video transmission according to the same set of timestamps.
9 . A computer server through which a first client device has established an audio and/or video communication session with a second client device over a social networking platform, the computer server comprising:
one or more processors; memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for:
receiving a first audio transmission from the second client device, wherein the first audio transmission is provided by the second client device in a source language that is distinct from a default language associated with the first client device;
obtaining one or more current user language attributes for the first client device, wherein the one or more current user language attributes are indicative of a current language that is used for the audio and/or video communication session at the first client device;
in accordance with a determination that the one or more current user language attributes suggest a target language that is currently used for the audio and/or video communication session at the first client device is distinct from the default language associated with the first client device:
obtaining a translation of the first audio transmission from the source language into the target language; and
sending, to the first client device, the translation of the first audio transmission in the target language, wherein the translation is presented to a user at the first client device.
10 . The computer server of claim 9 , wherein the obtaining the one or more current user language attributes and suggesting the target language that is currently used for the audio and/or video communication session at the first client device further comprises:
receiving, from the first client device, facial features of the current user and a current geolocation of the first client device; determining a relationship between the facial features of the current user and the current geolocation of the first client device; and suggesting the target language according to a determination that the relationship meets predefined criteria.
11 . The computer server of claim 9 , wherein the obtaining the one or more current user language attributes and suggesting the target language that is currently used for the audio and/or video communication session at the first client device further comprises:
receiving, from the first client device, an audio message that has been received locally at the first client device; analyzing linguistic characteristics of the audio message received locally at the first client device; and suggesting the target language that is currently used for the audio and/or video communication session at the first client device in accordance with a result of analyzing the linguistic characteristics of the audio message.
12 . The computer server of claim 9 , wherein the one or more programs further include instructions for:
obtaining vocal characteristics of a voice in the first audio transmission; and according to the vocal characteristics of the voice in the first audio transmission, generating a simulated first audio transmission that includes the translation of the first audio transmission spoken in the target language in accordance with the vocal characteristics of the voice of the first audio transmission.
13 . The computer server of claim 12 , wherein the sending, to the first client device, the translation of the first audio transmission in the target language to a user at the first client device includes:
sending, to the first client device, a textual representation of the translation of the first audio transmission in the target language to the user at the first client device; and sending, to the first client device, the simulated first audio transmission that is generated in accordance with the vocal characteristics of the voice in the first audio transmission.
14 . The computer server of claim 9 , wherein the receiving a first audio transmission from the second client device further comprises:
receiving two or more audio packets of the first audio transmission from the second client device, wherein the two or more audio packets have been sent from the second client device sequentially according to respective timestamps of the two or more audio packets, and wherein each respective timestamp is indicative of a start time of a corresponding audio paragraph identified in the first audio transmission.
15 . The computer server of claim 14 , wherein the obtaining the translation of the first audio transmission from the source language into the target language and sending the translation of the first audio transmission in the target language to the first client device further comprise:
obtaining respective translations of the two or more audio packets from the source language into the target language sequentially according to the respective timestamps of the two or more audio packets; and sending a first translation of at least one of the two or more audio packets to the first client device after the first translation is completed and before translation of at least another one of the two or more audio packets is completed.
16 . The computer server of claim 14 , wherein the one or more programs further include instructions for:
receiving a first video transmission while receiving the first audio transmission from the first client device, wherein the first video transmission is marked with the same set of timestamps as the two or more audio packets; and sending the first video transmission and the respective translations of the two or more audio packets in the first audio transmission with the same set of timestamps to the first client device such that the first client device synchronously present the respective translations of the two or more audio packets of the first audio transmission and the first video transmission according to the same set of timestamps.
17 . A non-transitory computer readable storage medium storing one or more programs, the one or more programs, when executed by a computer server through which a first client device has established an audio and/or video communication session with a second client device over a social networking platform, cause the computer server to perform operations comprising:
receiving a first audio transmission from the second client device, wherein the first audio transmission is provided by the second client device in a source language that is distinct from a default language associated with the first client device; obtaining one or more current user language attributes for the first client device, wherein the one or more current user language attributes are indicative of a current language that is used for the audio and/or video communication session at the first client device; in accordance with a determination that the one or more current user language attributes suggest a target language that is currently used for the audio and/or video communication session at the first client device is distinct from the default language associated with the first client device:
obtaining a translation of the first audio transmission from the source language into the target language; and
sending, to the first client device, the translation of the first audio transmission in the target language, wherein the translation is presented to a user at the first client device.
18 . The non-transitory computer readable storage medium of claim 17 , wherein the obtaining the one or more current user language attributes and suggesting the target language that is currently used for the audio and/or video communication session at the first client device further comprises:
receiving, from the first client device, facial features of the current user and a current geolocation of the first client device; determining a relationship between the facial features of the current user and the current geolocation of the first client device; and suggesting the target language according to a determination that the relationship meets predefined criteria.
19 . The non-transitory computer readable storage medium of claim 17 , wherein the obtaining the one or more current user language attributes and suggesting the target language that is currently used for the audio and/or video communication session at the first client device further comprises:
receiving, from the first client device, an audio message that has been received locally at the first client device; analyzing linguistic characteristics of the audio message received locally at the first client device; and suggesting the target language that is currently used for the audio and/or video communication session at the first client device in accordance with a result of analyzing the linguistic characteristics of the audio message.
20 . The non-transitory computer readable storage medium of claim 17 , wherein the one or more programs further include instructions for:
obtaining vocal characteristics of a voice in the first audio transmission; and according to the vocal characteristics of the voice in the first audio transmission, generating a simulated first audio transmission that includes the translation of the first audio transmission spoken in the target language in accordance with the vocal characteristics of the voice of the first audio transmission.Join the waitlist — get patent alerts
Track US2021366471A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.