US2006026206A1PendingUtilityA1

Telephony-data application interface apparatus and method for multi-modal access to data applications

Assignee: LOGHMANI MASOUDPriority: Oct 7, 1998Filed: Jun 3, 2005Published: Feb 2, 2006
Est. expiryOct 7, 2018(expired)· nominal 20-yr term from priority
G06Q 30/06G06Q 30/0625G06Q 30/0633
52
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A voice-optimized database system and method of audio vector valuation (AVV) provides means to search a database in response to a spoken query. Audio vectors (AVs) are assigned to phonemes in the names or phrases constituting searchable items in a voice-optimized database and in spoken queries. Multiple AVs can be stored which correspond to different pronunciations of the same searchable items to allow for less precision on the part telephone callers when stating their queries. A distance calculation is performed between the audio vectors of database items and spoken queries to produce search results. Existing databases can be enhanced with AVV. Several alternate samples of a spoken query are generated by analyzing the phonemic string of the spoken query to create similar, alternate phonemic strings. The phonemic string of the spoken query and the alternate phonemic strings are converted to text and used to search the database.

Claims

exact text as granted — not AI-modified
1 . A system for allowing users to browse and purchase items available via the Internet comprising: 
 a telephony-Internet interface device for connecting users to the Internet, at least one of said users communicating with said telephony-Internet interface device via a graphical interface device comprising a display and a user input interface, and an audio interface device comprising at least one of an analog line, a digital line, a wireless line, and a Voice over Internet Protocol (VOIP) line, respectively, in one of a simultaneous and sequential manner;    a memory device corresponding to at least one online source available to said users via the Internet for storing records for said items; and    a processing device connected to said memory device and said telephony-Internet interface device for allowing said users to request selected said items using at least one of electronic queries and spoken queries generated via said graphical interface device and said audio interface device, respectively, said telephony-Internet interface device being converting said spoken queries to electronic commands for transmission to said processing device, said processing device receiving from said memory device markup language-type pages comprising data relating to selected ones of said records in response to at least one of said electronic queries and said electronic commands, said markup language-type pages comprising audio tags and non-audio tags, said telephony-Internet interface parsing said markup language-type pages to obtain at least a portion of said data therefrom using said audio tags and to generate an audio message corresponding to said at least a portion of said data for playback to said users.    
   
   
       2 . A system as claimed in  claim 1 , further comprising: 
 a voice-enabled electronic shopping cart connected to said telephony-Internet interface device and online shops via the Internet and operable to process a transaction between said user and said online shops during which said user uses spoken queries only to browse said online shops and select different said items for placement in said voice-enabled electronic shopping cart.    
   
   
       3 . A system as claimed in  claim 1 , wherein said voice-enabled electronic shopping cart tracks and stores transaction data relating to the user's selection, and one of purchasing at least one of said items in said voice-enabled electronic shopping cart and terminating said transaction by one of plurality of termination events comprising hanging up a user telephone employed by said at least one of said users to communicate with said telephony-Internet interface, terminating connection to one of said digital line, said VOIP line and said wireless line, and interrupting said transaction via one of said processing device and said voice-enabled shopping cart.  
   
   
       4 . A system as claimed in  claim 3 , wherein said telephony-Internet interface device provides said user with an audio message which provides a menu of options corresponding to respective said online shops, to receive a response signal from said user on said at least one of said analog line, said digital line, said VOIP line and said wireless line indicating a selection of one of said options, and processing said response signal to generate an hypertext transfer protocol request to access the selected one of said online shops.  
   
   
       5 . A system as claimed in  claim 2 , wherein said voice-enabled electronic shopping cart stores account information relating to account holders comprising account identification codes and transaction data, said transaction data comprising item identification codes for said items in said voice-enabled shopping cart, and said user is one of said account holders, said voice-enabled electronic shopping cart stores said transaction data for said transaction, to maintain said transaction data following one of said plurality of termination events, and to recall said transaction data during a subsequent transaction involving at least one of said processing device and said voice-enabled electronic shopping cart and at least one of said user and another said user having access to the corresponding one of said account identification codes.  
   
   
       6 . A system as claimed in  claim 2 , wherein said voice-enabled electronic shopping cart comprises an audio interface directives module for generating tags and inserting said tags in said markup language-type pages prior to transmission to said telephony-Internet interface device, said tags being operable to facilitate parsing of said markup language-type pages by said telephony-Internet interface device.  
   
   
       7 . A system as claimed in  claim 6 , wherein at least one of said online shops employs scripts to facilitate modification of search results pages generated in response to said queries with tags and to assist said telephony-Internet interface device to interpret the tags for audio interaction with said user.  
   
   
       8 . A system as claimed in  claim 1 , wherein said telephony-Internet interface device performs a plurality of operations comprising converting said spoken queries to text, performing speech recognition operations on said spoken queries and converting text received from said processing device to speech.  
   
   
       9 . A system as claimed in  claim 1 , wherein said telephony-Internet interface device comprises: 
 a telephony interface module for connecting to said online shops via the Internet and to said user via at least one of an analog line, a digital line and a wireless line, said user connecting to said telephony interface module via at least one of a user computer, a user telephone and a telecommunications device connected to said at least one of an analog line, a digital line, a Voice over Internet Protocol line and a wireless line, said user computer comprising a microphone for inputting said spoken queries, said telephony interface module being operable to process a call from one of said user telephone, said user computer and said telecommunications device, to convert said spoken queries to commands, and to convert information received via the Internet in response to said commands into an audio format;    an audio vector valuation module;    a data presentation module connected to said telephony interface module; and    an Internet interface module connected to said data presentation module and operable to connect to the Internet;    wherein said data presentation module converts said spoken queries to said electronic commands for transmission to said processing device, said processing device retrieving selected ones of said records from said memory device in response to said queries and generating and transmitting said markup language-type pages to said telephony-Internet interface device comprising said data relating to said selected ones of said records, said data presentation module parsing said markup language-type pages received from said processing device to obtain at least a portion of said data therefrom and generating an audio message corresponding to said at least a portion of said data for playback to said user via said telephony interface module;    wherein said audio vector valuation module is operable to convert said spoken queries received from said user to phonemes for said electronic command transmitted to said processing device, said phonemes being used to generate audio vectors to facilitate retrieval of said records via said processing device.    
   
   
       10 . A system as claimed in  claim 9 , wherein said data presentation module converts said spoken queries received from said user to at least one of phonemes and data for said electronic command transmitted to said processing device.  
   
   
       11 . A system as claimed in  claim 1 , further comprising: 
 a voice-enabled electronic shopping cart connected to said telephony-Internet device and online shops via the Internet for processing a transaction between said user and said online shops during which said user uses spoken queries only to browse said online shops and select different said items for placement in said voice-enabled electronic shopping cart, said voice-enabled electronic shopping cart tracking and storing transaction data relating to the user's selection, and one of purchasing at least one of said items in said voice-enabled electronic shopping cart and terminating said transaction by one of plurality of termination events comprising hanging up a user telephone, terminating connection to one of said digital line and said wireless line, and interrupting said transaction via one of said processing device and said voice-enabled shopping cart, said voice-enabled electronic shopping cart comprising an audio interface directives module for generating audio tags and inserting said audio tags among non-audio tags in said markup language-type pages prior to transmission to said telephony-Internet interface device, said telephony-Internet interface device parsing said markup language-type pages and generate audio interaction with said user based on said audio tags.    
   
   
       12 . A system as claimed in  claim 1  wherein said processing device establishes a single session during which to receive from said memory device markup language-type pages comprising data relating to selected ones of said records in response to at least one of said electronic queries and said electronic commands from a corresponding one of said users who requested it, said telephony-Internet interface parsing said markup language-type pages to obtain at least a portion of said data therefrom using said audio tags, generating an audio message corresponding to said at least a portion of said data for playback to said corresponding one of said users, and generates said markup language-type pages concurrently with said audio message for selective use by said corresponding one of said users.  
   
   
       13 . An interface for connecting telephony users to access one or more online information sources using speech, text or graphics comprising: 
 a telephony interface module configured for access by said users via audio interface devices;    a data presentation module connected to said telephony interface module; and    an online interface module connected to said data presentation module and configured for access by said users via graphical interface devices comprising display devices and user input interfaces;    wherein said telephony interface module performs speech recognition and conversion operations between speech and text with respect to multiple users to generate commands, said data presentation module converts said commands into at least one online communication protocol, and said online interface module connects to said online information sources and retrieve information therefrom in response to at least one of said commands and requests generated via said graphical interface devices, said data presentation module parsing markup language-type pages provided by said online information sources to extract selected information, said markup language-type pages comprising audio tags and non-audio tags, said selected information being provided to said users on said audio interface devices via audio messaging using said audio tags and to said users on said graphical interface devices via at least one of text and graphics.    
   
   
       14 . An interface as claimed in  claim 13 , wherein said data presentation module is configured to establish first and second sessions for at least one of said users to communicate with said interface via a computer and a telephone, and said online interface module is configured to establish a single session with which to retrieve said information requested by said user and to provide said information to said data presentation module for playback via at least one of said computer and said telephone.  
   
   
       15 . An interface as claimed in  claim 13 , 
 wherein said telephony interface module processes a call from one of a user telephone and a user computer, performs at least one of speech recognition and conversion operations between speech and text on spoken queries to generate electronic commands, and manages multiple connections to different users, said data presentation module converts said electronic commands to HTTP-type electronic commands that comprise at least one of text and said phonemes and data, receives one of said markup language-type pages in response to said electronic commands, and manages processing of said electronic commands from a plurality of users, and said online interface module manages multiple connections to different web sites, said telephony interface module, said data presentation module and said Internet interface module being configured to allow establishment of, respectively, said connections to different users, said spoken queries from a plurality of users, and said connections to different web sites independently of each other, and to relate selected ones of said connections to different users to said connections to different web sites for said processing of spoken queries corresponding to said selected ones of said different users.    
   
   
       16 . A voice-enabled electronic shopping cart system for allowing users to browse and purchase items available via online shops on the Internet using spoken queries and electronic queries comprising: 
 a voice-enabled electronic shopping cart;    wherein at least one of said online shops comprises at least one memory device for storing records for said items available via said online shop and a processing device, said processing device being connected to said memory device and said voice-enabled electronic shopping cart, said processing device being operable to retrieve selected ones of said records from said memory device in response to said spoken queries and said electronic queries and to generate and transmit markup language-type pages to said users comprising data relating to said selected ones of said records and being formatted to allow generation of audio messages for said user corresponding to said data.    
   
   
       17 . A voice-enabled electronic shopping cart system as claimed in  claim 16 , further comprising a telephony-Internet interface device for connecting to said online shops via the Internet and to said users via at least one of an analog line, a digital line and a wireless line, said users accessing said system via at least one of a user computer, a user telephone and a telecommunications device connected to said at least one of an analog line, a digital line, a Voice over Internet Protocol (VOIP) line and a wireless line, said user computer comprising a microphone for inputting said spoken queries, 
 said telephony-Internet interface device converting said spoken queries to an electronic command for transmission to said processing device,    said processing device generating and transmitting markup language-type pages to said telephony Internet interface device comprising data relating to said selected ones of said records, said telephony-Internet interface device parsing said markup language-type pages to obtain at least a portion of said data therefrom and generating an audio message corresponding to said at least a portion of said data for playback to said user,    said telephone-Internet interface device being operable via the conversion of said spoken queries and parsing of markup language-type pages to track and store transaction data relating to use of spoken queries to browse said online shops, select different said items for placement in said voice-enabled electronic shopping cart, and purchase at least one of said items in said voice-enabled electronic shopping cart.    
   
   
       18 . A voice-enabled shopping cart system as claimed in  claim 16 , further comprising an audio interface directives module for generating tags and inserting said tags in said markup language-type pages, said tags being operable to facilitate parsing of said markup language-type pages by said telephony-Internet interface device.  
   
   
       19 . A method of allowing a user to interact with data applications using spoken queries comprising the steps of: 
 receiving a spoken query from said user;    converting said spoken query to text;    generating an electronic request using at least one electronic command comprising said text;    providing said electronic request to a processing device connected to said at least one of said data applications, said processing device being operable to search a memory device associated with said data application in response to said electronic request to locate selected information available via said data application, said processing device being operable to generate search results comprising at least one markup language-type page having information relating to said selected information;    providing audio tags singly or in combination with non-audio tags in said at least one markup language-type page identifying said information;    receiving said search results;    parsing said search results to obtain a portion of said information identified by said audio tags in said at least one markup language-type page; and    generating at least one audio message to indicate said portion of said information to said user.    
   
   
       20 . A method as claimed in  claim 19 , wherein said receiving step comprises the steps of: 
 generating a pre-recorded audio message instructing said user to select one of said data applications using one of a dual-tone multiple frequency tone and a voice command; and    generating a hypertext transfer protocol request to access a selected one of said data applications in accordance with said one of a dual-tone multiple frequency tone and a voice command.    
   
   
       21 . A method as claimed in  claim 19 , wherein said generating step for generating search results comprises the steps of: 
 receiving a home page from said selected data application; and    providing said text to at least one of said home page and another page generated by said selected data application for processing said spoken query.    
   
   
       22 . A method as claimed in  claim 19 , further comprising the step of configuring said online shop to provide tags in said search identifying said portion of said information.  
   
   
       23 . A method as claimed in  claim 19 , wherein said parsing step comprises the step of extracting said portion of said information using at least one of said tags.  
   
   
       24 . A method as claimed in  claim 19 , wherein said parsing step comprises the steps of: 
 extracting said portion of said information using scripts corresponding to respective said data applications to indicate where said portion of said information is located in markup language-type pages with only non-audio tags received therefrom; and    generating an audio interaction.    
   
   
       25 . A method as claimed in  claim 19 , wherein said generating step for generating said at least one audio message comprises the steps of: 
 converting said portion of said information into speech; and    combining said speech with a pre-recorded audio message.    
   
   
       26 . A method as claimed in  claim 19 , further comprising the steps of: 
 tracking and storing transaction data relating to the selection of said items in a voice-enabled electronic shopping cart; and    purchasing at least one of said items in said voice-enabled electronic shopping cart.    
   
   
       27 . A method as claimed in  claim 19 , wherein said converting step comprises the step of converting said spoken query to phonemes, and said providing step comprises the step of searching said database using audio vectors based on said phonemes.  
   
   
       28 . A method as claimed in  claim 19 , further comprising the step of receiving an audio command, said audio command being specified by said information identified by said audio tags.  
   
   
       29 . A method as claimed in  claim 19 , further comprising the steps of: 
 said processing device establishing a single session with which to retrieve said selected information and to provide said selected information to said user; and 
 receiving said search results during said single session to provide said search results to said user using both said audio tags and said non-audio tags.  
   
   
   
       30 . A method as claimed in  claim 29 , further comprising the step of generating said at least one markup language-type page concurrently with said audio message for selective use by said user.  
   
   
       31 . A method as claimed in  claim 29 , further comprising the step of 
 generating said at least one markup language-type page concurrently with said audio message for selective use by said user.    
   
   
       32 . A telephony-data application interface device for allowing a user to interact with data applications using spoken queries comprising: 
 a telephony interface module for receiving spoken queries from users;    a data presentation module connected to said telephony interface module for generating an electronic request using at least one electronic command in response to said spoken query; and    a data application interface module for connecting said data presentation module to said data applications, said electronic request being provided to a processing device connected to said at least one of said data applications that is operable to search a memory device associated with said data application in response to said electronic request to locate selected information, and to generate search results comprising at least one markup language-type page having said selected information and audio tags in combination with non-audio tags to identify said selected information;    wherein said data presentation module is configured to receive said search results, to parse said search results to obtain a portion of said information identified by said audio tags in said at least one markup language-type page, and to generate at least one audio message to indicate said portion of said information to said user.    
   
   
       33 . A telephone-data application interface device as claimed in  claim 32 , further comprising an Internet interface module connected to said data presentation module and operable to connect to the Internet and manage multiple connections to different web sites, wherein said telephony interface module is configured to process a call from one of a user telephone and a user computer and being operable to manage multiple connections to different users, said data presentation module is configured to process commands from said telephony interface module corresponding to spoken queries from a plurality of users, and said Internet interface module is operable to manage multiple connections to different data applications, said telephony interface module, said data presentation module and said Internet interface module being configured to allow establishment of, respectively, said connections to different users, said commands from a plurality of users, and said connections to different data applications independently of each other, and to relate selected ones of said connections to different users to said connections to different data applications for processing of said commands corresponding to said selected ones of said different users.  
   
   
       34 . A telephone-data application interface device as claimed in  claim 33 , wherein said data presentation module is configured to establish first and second sessions for at least one of said users to communicate with said interface via a computer and a telephone, and said data application interface module is configured to establish a single session with which to retrieve said information requested by said user and to provide said information to said data presentation module for playback via at least one of said computer and said telephone.  
   
   
       35 . A voice-enabled shopping cart for allowing a user to interact with data applications using spoken queries comprising: 
 a transaction module for communicating with a user establishing a session with at least one of a computer and a telephone to browse at least a selected one of the data applications, the user having a telephony-Internet interface to communicate with the selected data application via telephone;    a communications module for connecting to the internet and establishing communication between the selected data application and the transaction module; and    a monitoring module for monitoring each user transaction and user selections during a browsing session and allowing a user to terminate the session prior to purchase of any user selections and call the data application at another time to purchase at least some of the user selections.    
   
   
       36 . A voice-enabled electronic shopping system for allowing a user to browse and purchase items available via online shops on the Internet using spoken queries comprising: 
 a Voice over Internet Protocol (VOIP) module for communicating with a user who has accessed the internet with a telecommunications device equipped with a microphone for input of said spoken queries; and    a data presentation module coupled to the VOIP module and the internet and operable to convert said spoken queries to an electronic command for transmission to a selected online shop via the internet;    wherein the online shop stores records for said items available via the online shop in a memory device, retrieves selected ones of said records from said memory device in response to said spoken queries, and generates and transmits markup language-type pages to said data presentation module that comprise data relating to said selected ones of said records, said data presentation module parses said markup language-type pages received from the online shop to obtain at least a portion of said data therefrom and to generate an audio message corresponding to said at least a portion of said data for playback to said user via said VOIP module;    wherein the online shop also stores for each of said records at least one item phoneme vector for parsing of said memory device by the online shop in response to said spoken queries, each of said records having at least one searchable field comprising item data that is characterized by phonemes, said phonemes being assigned respective values, said phonemes having similar pronunciation being assigned similar values, each said phoneme vector comprising said values corresponding to said phonemes in said at least one searchable field; and    wherein said data presentation module converts said spoken queries received from said user to phonemes for said electronic command transmitted to the online shop, said phonemes being used to generate spoken query phoneme vectors to facilitate retrieval of said records via the online shop by comparing said spoken query phoneme vector with at least a portion of said item phoneme vector corresponding to each of said items.    
   
   
       37 . A voice-enabled electronic shopping system as claimed in  claim 36 , further comprising an audio vector valuation module coupled to the data presentation module, the audio vector valuation module generating alternate pronunciations of said spoken queries for which phonemes are also determined by the data presentation module.  
   
   
       38 . A voice-enabled electronic shopping system as claimed in  claim 36 , further comprising telephony interface module coupled to the data presentation module to receive spoken queries via at least one telephone.  
   
   
       39 . A voice-enabled electronic shopping system as claimed in  claim 38 , further comprising an internet interface module coupled to the data presentation module to connect to the internet.

Join the waitlist — get patent alerts

Track US2006026206A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.