System and Method for Enhancing Locative Response Abilities of Autonomous and Semi-Autonomous Agents
Abstract
A computer system and method according to the present invention can receive multi-modal inputs such as natural language, gesture, text, sketch and other inputs in order to simplify and improve locative question answering in virtual worlds, among other tasks. The components of an agent as provided in accordance with one embodiment of the present invention can include one or more sensors, actuators, and cognition elements, such as interpreters, executive function elements, working memory, long term memory and reasoners for responses to locative queries, for example. Further, the present invention provides, in part, a locative question answering algorithm, along with the command structure, vocabulary, and the dialog that an agent is designed to support in accordance with various embodiments of the present invention.
Claims
exact text as granted — not AI-modified1 . A system for manipulating virtually displayed objects in a virtual environment, comprising:
at least one input device adapted to receive at least one of speech, gesture, text and touchscreen inputs; and a computer processor adapted to execute a program stored in a computer memory, the program being operable to provide instructions to the computer processor including:
receiving user input via the at least one input device, wherein the user input comprises a query regarding at least one target object in the virtual environment;
interfacing with the virtual environment;
sensing the at least one object in the virtual environment; and
deriving an optimum utterance in response to the query.
2 . The system of claim 1 wherein the optimum utterance is a minimum cost response.
3 . The system of claim 1 wherein interfacing with the virtual environment and sensing the at least one object in the virtual environment are conducted by a virtual agent.
4 . The system of claim 3 wherein the query is a locative query for a target object in the virtual environment, and wherein deriving an optimum utterance includes selecting at least one candidate landmark object from a group of available candidate landmark objects based upon the at least one candidate landmark object being on a potential path between the virtual agent and the target object.
5 . The system of claim 3 wherein the query is a locative query for a target object in the virtual environment, and wherein deriving an optimum utterance includes selecting at least one candidate landmark object from a group of available candidate landmark objects based upon the at least one candidate landmark object being larger in size than the target object.
6 . The system of claim 3 wherein deriving an optimum utterance includes deriving a set of object pairs in the virtual environment and computing at least one inter-object relation associated with each object pair.
7 . The system of claim 6 wherein the at least one inter-object relation is one of visibility, distance, spatial relation and background contrast level.
8 . The system of claim 3 wherein deriving an optimum utterance includes deriving a cost computation for each candidate landmark object from a group of available candidate landmark objects.
9 . The system of claim 8 wherein deriving a cost computation for each candidate landmark object is limited by a predetermined maximum number of landmarks admissible in the utterance.
10 . The system of claim 8 wherein deriving a cost computation for each candidate landmark object includes constructing a cost tree with a plurality of nodes, with each of the plurality of nodes representing a candidate landmark object, wherein the nodes are arranged into at least one path, and wherein the at least one path is ordered by descending volume of the nodes.
11 . The system of claim 10 wherein deriving a cost computation further includes deriving a cost associated with each node of the at least one path.
12 . The system of claim 11 wherein deriving a cost associated with each node includes determining a visual area scan cost.
13 . The system of claim 12 wherein determining the visual area scan cost includes determining a distance scan factor based upon the distance between two candidate landmark objects and inversely proportional to the volume of one of the two candidate landmark objects.
14 . The system of claim 12 wherein determining the visual area scan cost includes determining a relational scan difficulty factor.
15 . The system of claim 11 wherein deriving a cost associated with each node includes determining a target recognition difficulty cost.
16 . The system of claim 15 wherein determining a target recognition difficulty cost based upon the user's familiarity with the candidate landmark objects.
15 . The system of claim 11 wherein deriving a cost associated with each node includes determining a target background contrast cost.
16 . The system of claim 11 wherein deriving a cost associated with each node includes determining a description length penalty factor that is proportional to the number of candidate landmarks in the determined utterance response.
17 . The system of claim 11 wherein deriving a cost associated with each node includes determining a parent node cost.
18 . A computer-implemented method, comprising:
receiving, by a computer engine, user input that comprises a query regarding at least one target object in the virtual environment; interfacing, via the computer engine, with the virtual environment; sensing, via a virtual agent associated with the computer engine, the at least one target object in the virtual environment; and deriving an optimum utterance in response to the query.
19 . The method of claim 18 , wherein interfacing with the virtual environment and sensing the at least one object in the virtual environment are conducted by a virtual agent, and wherein deriving an optimum utterance in response to the query includes:
deriving a set of object pairs in the virtual environment and computing at least one inter-object relation associated with each object pair; deriving a cost computation for each candidate landmark object from a group of available candidate landmark objects, wherein deriving a cost computation for each candidate landmark object includes constructing a cost tree with a plurality of nodes, with each of the plurality of nodes representing a candidate landmark object, and deriving a cost associated with each node, wherein the cost associated with each node is determined by determining a visual area scan cost, a relational scan difficulty factor, a target recognition difficulty cost, a target background contrast cost, a description length penalty factor and a parent node cost.
20 . The method of claim 19 wherein the cost associated with each node is calculated through the sum of the visual area scan cost, the relational scan difficulty factor, the target recognition difficulty cost, the target background contrast cost, the description length penalty factor and the parent node cost.
21 . The method of claim 19 wherein the nodes are arranged in one or more paths as part of the cost tree, and wherein the path associated with the lowest cost is used to generate the optimum utterance.
22 . A system for deriving a response to a locative query in a virtual environment, comprising:
at least one input device adapted to receive at least one of speech, gesture, text and touchscreen inputs; and a computer processor adapted to execute a program stored in a computer memory, the program being operable to provide instructions to the computer processor including:
receiving user input via the at least one input device, wherein the user input comprises a query regarding the location of at least one target object in the virtual environment;
interfacing with the virtual environment via a virtual agent;
sensing the at least one object in the virtual environment by the agent; and
deriving an optimum utterance in response to the query based upon a determined cost computation for each candidate landmark object from a group of available candidate landmark objects.Join the waitlist — get patent alerts
Track US2012306741A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.