US2023017401A1PendingUtilityA1

Speech Activity Detection Using Dual Sensory Based Learning

Assignee: BLACKBERRY LTDPriority: Dec 4, 2020Filed: Sep 16, 2022Published: Jan 19, 2023
Est. expiryDec 4, 2040(~14.4 yrs left)· nominal 20-yr term from priority
G06V 40/165G06V 10/751G06V 20/40H04N 7/15H04M 3/568G10L 25/78H04M 3/569
67
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A dual sensory input speech detection method includes receiving, at a first time, a first video image input of a conference participant of the video conference and a first audio input of the conference participant; communicating the first video image input to the video conference; identifying the first video image input as a first facial image of the conference participant; determining, based on the first facial image, the first video image input indicates the conference participant is in a speaking state; identifying the first audio input as a first speech sound; determining, while in the speaking state, the first speech sound originates from the conference participant; and communicating the first audio input to an audio output for the video conference.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method comprising:
 monitoring video including one or more images of a user;   determining whether the one or more images indicate lip or mouth movement associated with the user; and   adjusting audio associated with the user based on the one or more images of the user.   
     
     
         2 . The method of  claim 1 , further comprising transmitting the audio based on the one or more images indicating lip or mouth movement associated with the user. 
     
     
         3 . The method of  claim 1 , wherein adjusting the audio comprises unmuting the audio. 
     
     
         4 . The method of  claim 1 , wherein adjusting the audio comprises muting the audio. 
     
     
         5 . The method of  claim 1 , wherein adjusting the audio comprises adjusting gain of the audio. 
     
     
         6 . The method of  claim 1 , wherein adjusting the audio comprises attenuating the audio. 
     
     
         7 . The method of  claim 1 , further comprising providing the user one or more options for overriding the adjusting the audio. 
     
     
         8 . The method of  claim 1 , wherein the video and audio are received at a communication device. 
     
     
         9 . The method of  claim 8 , wherein the video is received at a camera integrated with the communication device and the audio is received at a microphone integrated with the communication device. 
     
     
         10 . The method of  claim 1 , wherein the method further comprises:
 as a result of determining that the one or more images do not indicate lip or mouth movement:
 not transmitting the audio; and 
 transmitting the video. 
   
     
     
         11 . The method of  claim 1 , wherein the user is a participant in a video conference. 
     
     
         12 . The method of  claim 1 , wherein the method comprises a method of authentication. 
     
     
         13 . A system, comprising:
 a memory storing instructions;   a processor coupled to the memory and configured to execute the instructions to cause the system to:   monitor video including one or more images of a user;   determine whether the one or more images indicate lip or mouth movement associated with the user; and   adjust audio based on the one or more images associated with the user.   
     
     
         14 . The system of  claim 13 , wherein the processor is further configured to execute instructions to cause the system to transmit the audio based on the one or more images indicating lip or mouth movement associated with the user. 
     
     
         15 . The system of  claim 13 , wherein adjusting the audio comprises unmuting the audio. 
     
     
         16 . The system of  claim 13 , wherein adjusting the audio comprises muting the audio. 
     
     
         17 . The system of  claim 13  wherein adjusting the audio comprises adjusting gain of an audio capture device. 
     
     
         18 . The system of  claim 13 , wherein adjusting the audio comprises attenuating the audio. 
     
     
         19 . The system of  claim 13 , further comprising providing the user one or more options for overriding the adjusting the audio. 
     
     
         20 . The system of  claim 13 , wherein the video and audio are received at the system. 
     
     
         21 . The system of  claim 20 , wherein the system further comprises a camera and a microphone coupled to the processor, and wherein the video is received at the camera and the audio is received at the microphone. 
     
     
         22 . The system of  claim 13 , wherein the processor is further configured, as a result of determining that the one or more images do not indicate lip or mouth movement, to:
 not transmit the audio; and   transmit the video.   
     
     
         23 . The system of  claim 13 , wherein the user is a participant in a video conference. 
     
     
         24 . The system of  claim 13 , wherein the processor further configured for authentication. 
     
     
         25 . A non-transitory computer readable media storing instructions that when executed by a processor cause the processor to:
 monitor video including one or more images of a user;   determine whether the one or more images indicate lip or mouth movement associated with the user; and   adjust audio based on the one or more images associated with the user.   
     
     
         26 . The non-transitory computer readable media of  claim 25 , wherein the instructions when executed by the processor cause the processor to transmit the audio based on the one or more images indicating lip or mouth movement associated with the user. 
     
     
         27 . The non-transitory computer readable media of  claim 25 , wherein adjusting the audio comprises unmuting the audio. 
     
     
         28 . The non-transitory computer readable media of  claim 25 , wherein adjusting the audio comprises muting the audio. 
     
     
         29 . The non-transitory computer readable media of  claim 25 , wherein adjusting the audio comprises adjusting gain of an audio capture device. 
     
     
         30 . The non-transitory computer readable media of  claim 25 , wherein adjusting the audio comprises attenuating the audio. 
     
     
         31 . The non-transitory computer readable media of  claim 25 , wherein the instructions when executed by the processor cause the processor to provide the user one or more options for overriding the adjusting the audio. 
     
     
         32 . The non-transitory computer readable media of  claim 25 , wherein the video and audio are received at a system comprising the processor. 
     
     
         33 . The non-transitory computer readable media of  claim 25 , wherein the instructions when executed by the processor cause the processor to receive the video via a camera and receive the audio via a microphone. 
     
     
         34 . The non-transitory computer readable media of  claim 25 , wherein the instructions when executed by the processor cause the processor, as a result of determining that the one or more images do not indicate lip or mouth movement, to:
 not transmit the audio; and   transmit the video.   
     
     
         35 . The non-transitory computer readable media of  claim 25 , wherein the user is a participant in a video conference. 
     
     
         36 . The non-transitory computer readable media of  claim 25 , wherein the instructions when executed by the processor cause the processor to be configured for authentication.

Join the waitlist — get patent alerts

Track US2023017401A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.