US2018218735A1PendingUtilityA1

Speech recognition involving a mobile device

Assignee: APPLE INCPriority: Dec 11, 2008Filed: Mar 27, 2018Published: Aug 2, 2018
Est. expiryDec 11, 2028(~2.4 yrs left)· nominal 20-yr term from priority
G10L 15/26G10L 15/30G10L 2015/0631G10L 2015/025
48
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A system and method of speech recognition involving a mobile device. Speech input is received ( 202 ) on a mobile device ( 102 ) and converted ( 204 ) to a set of phonetic symbols. Data relating to the phonetic symbols is transferred ( 206 ) from the mobile device over a communications network ( 104 ) to a remote processing device ( 106 ) where it is used ( 208 ) to identity at least one matching data item from a set of data items ( 114 ). Data relating to the at least one matching data item is transferred ( 210 ) from the remote processing device to the mobile device and presented ( 214 ) thereon.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . An electronic device, comprising:
 one or more processors;   a memory; and   one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for:
 receiving speech input; 
 converting the speech input to a set of phonetic symbols on the electronic device; and 
 transferring data relating to the phonetic symbols to a remote processing device over a communications network; 
 receiving data relating to at least one matching data item from a remote processing device; and 
 presenting data relating to the at least one received matching data item. 
   
     
     
         2 . The electronic device of  claim 1 , wherein presenting data relating to the at least one received matching data item comprises:
 displaying an orthographic representation of the at least one received matching data item.   
     
     
         3 . The electronic device of  claim 1 , wherein presenting data relating to the at least one received matching data item comprises:
 outputting data corresponding to a map coordinate of a location represented by a received matching data item.   
     
     
         4 . The electronic device of  claim 1 , wherein presenting data relating to the at least one received matching data item comprises:
 outputting data corresponding to an identification code for a media item represented by the received matching data item.   
     
     
         5 . The electronic device of  claim 1 , wherein presenting data relating to the at least one received matching data item comprises:
 generating a spoken form of the at least one received matching data item.   
     
     
         6 . The electronic device of  claim 1 , wherein a number of the matching data items received from the remote processing device corresponds to a lower of: a maximum number of data items to be displayed, or a number of data items arranged in order of decreasing posterior probability corresponding to the phonetic symbols down to a predetermined threshold value. 
     
     
         7 . The electronic device of  claim 6 , wherein the posterior probability is based on a match score, wherein the match score is generated by matching phonetic symbols to phonetic reference forms corresponding to the matching data items, and wherein the match score is normalized. 
     
     
         8 . The electronic device of  claim 1 , wherein the set of phonetic symbols comprises:
 a sequence of phonetic symbols, a lattice of phonetic symbols, or a combination thereof   
     
     
         9 . The electronic device of  claim 1 , wherein the one or more programs further include instructions for:
 storing data representing the speech input; and   performing a rescoring process using the data relating to the at least one received matching data item and the stored speech input data.   
     
     
         10 . The electronic device of  claim 9 , wherein the rescoring process further comprises producing a network including data representing phonetic specifications of the received data relating to at least one matching data item. 
     
     
         11 . The electronic device of  claim 9 , wherein the rescoring process further comprises:
 generating, for each received matching data item to be rescored, sequences or lattices of acoustic hidden Markov models (HMMs), the HMMs representing phonetic units corresponding to a phonetic specification of reference pronunciations, wherein the reference pronunciations correspond to each received matching data item to be rescored.   
     
     
         12 . The electronic device of  claim 11 , wherein the rescoring process further comprises:
 receiving compressed data based on shared common elements in the phonetic specification to reduce an amount of data received at the electronic device for the rescoring process.   
     
     
         13 . The electronic device of  claim 11 , wherein the rescoring process further comprises:
 receiving data specifying a general-purpose sub-grammar representing a set of alternative number-word sequences for the rescoring process; and   determining, using the sub-grammar, a most likely one of the number-word sequences to be presented.   
     
     
         14 . The electronic device of  claim 11 , wherein the rescoring process further comprises:
 receiving phonetic specifications corresponding to the received matching data items, the specifications associated with indices;   selecting one or more phonetic specifications to be presented;   transferring indices corresponding to each of the selected phonetic specifications to the remote processing device;   receiving, from the remote processing device, data representing a full description of data items corresponding to the transferred indices; and   presenting the data representing the full description of data items.   
     
     
         15 . The electronic device of  claim 11 , wherein the rescoring process further comprises:
 receiving, from the remote processing device, a predetermined number of matching data items corresponding to best matches of the phonetic symbols and data items in a set of data items.   
     
     
         16 . The electronic device of  claim 10 , wherein each of the sequences or lattices is matched against a sequence of frames of spectrum parameters corresponding to the speech input data stored in the electronic device to produce a match score. 
     
     
         17 . The electronic device of  claim 16 , wherein the matching of the sequences or lattices involves at least one of a Viterbi time alignment and a full forward probability method. 
     
     
         18 . The electronic device of  claim 1 , wherein the mobile device comprises a mobile telephone. 
     
     
         19 . A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device, cause the electronic device to:
 receive speech input;   convert the speech input to a set of phonetic symbols on the electronic device; and   transfer data relating to the phonetic symbols to a remote processing device over a communications network.   
     
     
         20 . A method, comprising:
 at an electronic device having one or more processors:
 receiving speech input; 
 converting the speech input to a set of phonetic symbols on the electronic device; and 
 transferring data relating to the phonetic symbols to a remote processing device over a communications network.

Join the waitlist — get patent alerts

Track US2018218735A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.