US2023042836A1PendingUtilityA1

Resolving natural language ambiguities with respect to a simulated reality setting

Assignee: APPLE INCPriority: Sep 24, 2019Filed: Oct 19, 2022Published: Feb 9, 2023
Est. expirySep 24, 2039(~13.2 yrs left)· nominal 20-yr term from priority
G06F 3/017G10L 15/22G06F 3/013G10L 15/26G06N 20/00G06F 3/167G06T 19/006G10L 2015/228G06F 3/011G06F 2203/0381G06V 20/20G06V 40/18
65
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The present disclosure relates to resolving natural language ambiguities with respect to a simulated reality setting. In an exemplary embodiment, a simulated reality setting having one or more virtual objects is displayed. A stream of gaze events is generated from the simulated reality setting and a stream of gaze data. A speech input is received within a time period and a domain is determined based on a text representation of the speech input. Based on the time period and a plurality of event times for the stream of gaze events, one or more gaze events are identified from the stream of gaze events. The identified one or more gaze events is used to determine a parameter value for an unresolved parameter of the domain. A set of tasks representing a user intent for the speech input is determined based on the parameter value and the set of tasks is performed.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic system with a display and one or more images sensors, the one or more programs including instructions for:
 generating, based on a stream of gaze data and a displayed setting, a stream of gaze events corresponding to a plurality of event times;   receiving speech input within a time period;   identifying, based on the time period and the stream of gaze events, an application of a plurality of displayed applications within the displayed setting; and   causing the application to close.   
     
     
         2 . The non-transitory computer-readable storage medium of  claim 1 , wherein the displayed setting is a simulated reality setting. 
     
     
         3 . The non-transitory computer-readable storage medium of  claim 1 , wherein the stream of gaze data is determined based on image data from the one or more image sensors. 
     
     
         4 . The non-transitory computer-readable storage medium of  claim 1 , wherein identifying the application includes:
 determining a domain based on a text representation of the speech input; and   identifying an application corresponding to the determined domain and a gaze event of the stream of gaze events.   
     
     
         5 . The non-transitory computer-readable storage medium of  claim 4 , wherein the application is a parameter of a task to close the application. 
     
     
         6 . The non-transitory computer-readable storage medium of  claim 4 , wherein identifying the application corresponding to the determined domain and the gaze event of the stream of gaze events includes determining a first application that corresponds to the gaze event during the time period. 
     
     
         7 . The non-transitory computer-readable storage medium of  claim 1 , wherein the speech input includes a deictic expression, and wherein the application corresponds to the deictic expression. 
     
     
         8 . The non-transitory computer-readable storage medium of  claim 1 , wherein each gaze event in the stream of gaze events occurs at a respective event time of the plurality of event times and represents user gaze fixation on a respective gazed object of the plurality of gazed objects. 
     
     
         9 . The non-transitory computer-readable storage medium of  claim 8 , wherein the plurality of gazed objects corresponds to a plurality of applications. 
     
     
         10 . The non-transitory computer-readable storage medium of  claim 8 , wherein each gaze event in the stream of gaze events is identified from the stream of gaze data based on a determination that a duration of the user gaze fixation on the respective gazed object satisfies a threshold duration. 
     
     
         11 . The non-transitory computer-readable storage medium of  claim 1 , wherein generating the stream of gaze events includes determining respective durations of gaze fixations on a plurality of gazed objects, and wherein the one or more gaze events are identified based on the respective durations of gaze fixations on the plurality of gazed objects. 
     
     
         12 . The non-transitory computer-readable storage medium of  claim 1 , wherein one or more gaze events are identified based on a determination, from the plurality of event times, that the one or more gaze events occurred closest to the time period relative to other gazed events in the stream of gaze events. 
     
     
         13 . The non-transitory computer-readable storage medium of  claim 1 , wherein the speech input includes an ambiguous expression corresponding to the application, further comprising:
 determining a reference time at which the ambiguous expression was spoken, wherein one or more gaze events are identified based on a determination that the one or more gaze events each occurred within a threshold time interval from the reference time.   
     
     
         14 . The non-transitory computer-readable storage medium of  claim 1 , the one or more programs further including instructions for:
 detecting a gesture event based on an image data from the one or more image sensors of the electronic system; and   identifying one or more objects to which the gesture event is directed, wherein the one or more objects are identified within a field of view of a user, and wherein the one or more objects corresponds to an application.   
     
     
         15 . The non-transitory computer-readable storage medium of  claim 1 , wherein the plurality of gazed objects further includes one or more physical objects of a physical setting. 
     
     
         16 . An electronic system, comprising:
 a display;   one or more images sensors;   one or more processors; and   memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for:
 generating, based on a stream of gaze data and a displayed setting, a stream of gaze events corresponding to a plurality of event times; 
 receiving speech input within a time period; 
 identifying, based on the time period and the stream of gaze events, an application of a plurality of displayed applications within the displayed setting; and 
 causing the application to close. 
   
     
     
         17 . A method, performed by an electronic system having one or more processors, memory, a display, and one or more image sensors, the method comprising:
 generating, based on a stream of gaze data and a displayed setting, a stream of gaze events corresponding to a plurality of event times;   receiving speech input within a time period;   identifying, based on the time period and the stream of gaze events, an application of a plurality of displayed applications within the displayed setting; and   causing the application to close.

Join the waitlist — get patent alerts

Track US2023042836A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.