Speech recognition and translation method and translation apparatus
Abstract
There are a speech recognition and translation method as well as a translation apparatus. The method includes: entering a speech recognition state in response to the translation button being pressed, and collecting a voice of a user through the sound collecting device; importing the collected voice into each of a plurality of speech recognition engines through the processor to obtain a confidence of the voices corresponding to a plurality of different candidate languages, and determining a source language used by the user based on the confidence and a preset determination rule; exiting the speech recognition state in response to the translation button being released in the speech recognition state, and converting the voice of the source language to a target voice of a preset language through the processor; and playing the target voice through the sound playback device.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A speech recognition and translation method for a translation apparatus, wherein the translation apparatus comprises a processor, a sound collecting device electrically coupled to the processor, and a sound playback device electrically coupled to the processor; wherein the translation apparatus is further provided with a translation button; wherein the method comprises:
entering a speech recognition state in response to the translation button being pressed, and collecting a voice of a user through the sound collecting device; importing the collected voice into each of a plurality of speech recognition engines through the processor to obtain a confidence of the voices corresponding to a plurality of different candidate languages, and determining a source language used by the user based on the confidence and a preset determination rule, wherein each of the plurality of speech recognition engines corresponds to each of the plurality of different candidate languages; exiting the speech recognition state in response to the translation button being released in the speech recognition state, and converting the voice of the source language to a target voice of a preset language through the processor; and playing the target voice through the sound playback device.
2 . The method of claim 1 , wherein the step of determining the source language used by the user based on the confidence and the preset determination rule comprises:
determining a language in the candidate languages with the highest confidence as the source language used by the user.
3 . The method of claim 1 , wherein the step of importing the collected voice into each of the plurality of speech recognition engines through the processor to obtain the confidence of the voices corresponding to the plurality of different candidate languages, and determining the source language used by the user based on the confidence and the preset determination rule comprises:
importing each of the voices into each of the plurality of speech recognition engines through the processor to obtain a plurality of first texts corresponding to each of the candidate languages and a plurality of the confidence; filtering the candidate languages to obtain a plurality of first languages, wherein a value of the confidence of the first language is greater than a first preset value, and a difference between the values of the confidences of any two adjacent first languages is less than a second preset value; determining whether an amount of one or more second language included in the first language is 1, wherein the first text corresponding to the second language conforms to a text rule of the one or more second languages; determining the second language as the source language, if the amount of the second language is 1; taking a third language in each of the one or more second languages as the source language, wherein in all the one or more second languages, a syntax of the first text corresponding to the third language has the highest matchingness with a syntactic rule of the third language.
4 . The method of claim 3 , wherein the step of converting the voice of the source language to the target voice of the preset language comprises:
translating the first text corresponding to the source language into a second text of the preset language; and converting the second text to the target voice through a speech synthesis system.
5 . The method of claim 1 , wherein the translation apparatus further comprises a wireless signal transceiving device electrically coupled to the processor, wherein the step of importing the collected voice into each of the plurality of speech recognition engines through the processor to obtain the confidence of the voices corresponding to the plurality of different candidate languages comprises:
importing the voice to a client corresponding to each of the plurality of speech recognition engines through the processor; transmitting the voice to a corresponding server in a form of streaming media in real time and receiving the confidence returned by each of the servers through the wireless signal transceiving device by each client; stopping the transmission of the voice in response to detecting a packet loss, a network speed being less than a preset speed, or a disconnection rate being greater than a preset frequency; and transmitting all the voices collected in the speech recognition state in a form of file to the corresponding server and receiving the confidence returned by each server through the wireless signal transceiving device by each client in response to detecting the translation button being released in the speech recognition state, or recognizing the voice by calling a local database through the client to obtain the confidence.
6 . The method of claim 3 , wherein the translation apparatus further comprises a touch screen electrically coupled to the processor, wherein the method further comprises:
importing each of the voices to the plurality of the speech recognition engines through the processor to obtain a word probability list corresponding to each of the candidate languages; displaying the first text corresponding to the source language on the touch screen after the source language is recognized; and switching the first word in the first text displayed on the touch screen pointed by a click of the user to a second word, in response to detecting the click on the touch screen.
7 . The method of claim 1 , wherein the translation apparatus is provided with a motion sensor electrically coupled to the processor, wherein the method further comprises:
setting a first action and a second action of the user detected through the motion sensor as a first preset action and a second preset action, respectively; entering the speech recognition state in response to detecting the user having performed the first preset action through the motion sensor, and exiting the speech recognition state in response to detecting the user having performed the second preset action through the motion sensor.
8 . A translation apparatus, wherein the apparatus comprises:
a recording module configured to enter a speech recognition state in response to a translation button being pressed, and collecting a voice of a user through a sound collecting device; a voice recognizing module configured to import the collected voice into each of a plurality of speech recognition engines to obtain a confidence of the voices corresponding to a plurality of different candidate languages, and determining a source language used by the user based on the confidence and a preset determination rule, wherein each of the plurality of speech recognition engines corresponds to each of the plurality of different candidate languages; a voice converting module configured to exit the speech recognition state in response to the translation button being released in the speech recognition state, and converting the voice of the source language to a target voice of a preset language; and a playback module configured to play the target voice through the sound playback device.
9 . A translation apparatus, wherein the apparatus comprises:
an equipment body; a recording hole, a display screen, and a button disposed on a body of the equipment body; a processor, a storage, a sound collecting device, a sound playback device, and a communication module disposed inside the equipment body; the display screen, the button, the storage, the sound collecting device, the sound playback device, and the communication module are electrically coupled to the processor; the storage stores a computer program executable on the processor, and the following steps are performed when the processor executes the computer program: entering a speech recognition state in response to the translation button being pressed, and collecting a voice of a user through the sound collecting device; importing the collected voice into each of a plurality of speech recognition engines to obtain a confidence of the voices corresponding to the plurality of different candidate languages, and determining a source language used by the user based on the confidence and a preset determination rule, wherein each of the plurality of speech recognition engines corresponds to each of the plurality of different candidate languages; exiting the speech recognition state in response to the translation button being released in the speech recognition state, and converting the voice of the source language to a target voice of a preset language; and playing the target voice through the sound playback device.
10 . The apparatus of claim 9 , wherein:
a bottom of the equipment body is provided with a speaker window; inside the equipment body is provided with a battery and a motion sensor both electrically coupled to the processor, and an audio signal amplifying circuit electrically coupled to the sound collecting device; and the display screen is a touch screen.Join the waitlist — get patent alerts
Track US2021365641A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.