US2025218126A1PendingUtilityA1
Multi-modal query and response architecture for medical procedures
Assignee: INTUITIVE SURGICAL OPERATIONSPriority: Dec 29, 2023Filed: Dec 26, 2024Published: Jul 3, 2025
Est. expiryDec 29, 2043(~17.4 yrs left)· nominal 20-yr term from priority
G06F 40/30G06F 40/279G16H 50/20G16H 70/20G16H 50/70G06V 20/47G06V 10/82G16H 20/40G16H 30/40G06F 40/40G06V 2201/03G06T 2210/41G06T 2210/56G06T 2219/004G06V 20/46G06V 20/41G06T 19/00
46
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Aspects of this technical solution can receive a text prompt from a user, generate, based at least in part on a plurality of sets of data associated with at least one medical procedure, an output corresponding to the text prompt, where each of the plurality of sets of data has a different modality, where the plurality of sets of data comprises depth data, and provide the output for display.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A system, comprising:
one or more processors, coupled with memory, to: receive a text prompt from a user; generate, based at least in part on a plurality of sets of data associated with at least one medical procedure, an output corresponding to the text prompt, wherein each of the plurality of sets of data has a different modality, wherein the plurality of sets of data comprises depth data; and provide the output for display.
2 . The system of claim 1 , wherein the output includes a text response and a visual annotation of video data, the text response includes a number responsive to the text prompt, and the plurality of sets of data comprises the video data.
3 . The system of claim 1 , wherein the output includes a text response and a portion of video data within a time period, the text response and the time period each responsive to the text prompt, and the plurality of sets of data comprises the video data.
4 . The system of claim 1 , wherein the output includes a text response and an image, the text response includes at least a partial description of the image and is responsive to the text prompt.
5 . The system of claim 4 , wherein
the partial description of the image includes a description of at least one of an absolute location or a relative location within the medical environment of a first object in the plurality of sets of data, the relative location of the first object relative to a second object in the medical environment; or the partial description of the image includes a description of a layout of one or more objects in the plurality of sets of data, the image corresponds to a 3D reconstruction of a medical environment based on point cloud data, the medical environment corresponds to the medical procedure, and the plurality of set of data includes the point cloud data.
6 . The system of claim 1 , wherein the output includes a text response and a plurality of portions of video data, the text response includes a number responsive to the text prompt, a number of the plurality of portions of the video data corresponds to the number responsive to the text prompt, and the plurality of sets of data comprises the video data.
7 . The system of claim 6 , wherein
the number corresponds to a number of people involved in a task in a medical environment corresponding to the medical procedure, and the plurality of portions of the video data depicts one or more of the number of people involved in the task; the number corresponds to a number of instruments used in a task in a medical environment corresponding to the medical procedure, and the plurality of portions of the video data depicts one or more of the number of instruments used in the task; or the number corresponds to a number of events by a person involved in a task in a medical environment corresponding to the medical procedure, and the plurality of portions of the video data depicts one or more of the events.
8 . The system of claim 1 , wherein the output includes a text response and a data visualization output, the text response includes at least a partial description of the data visualization output.
9 . The system of claim 1 , the processors to:
determine a modality that is responsive to the text prompt, wherein the modality includes at least one of video data, an annotation of video data, an image, and a data visualization.
10 . The system of claim 1 , the processors to:
extract one or more features from one of more of the plurality of sets of data; and generate one or more fused features each including one or more of the features each having the different modality; and generate the output based on one or more of the fused features.
11 . The system of claim 1 , the processors to:
select, according to a determination that the plurality of sets of data includes video data, a neural network configured to extract features from the video data; and extract, by the neural network, the one or more features from the video data.
12 . A system, comprising:
one or more processors, coupled with memory, to: receive multi-modal data comprising video data, analytics data, and metadata for one or more medical procedures each to update one or more models; generate, using a first model configured to detect image features, a first feature that identifies an object in the video data for a plurality of medical procedures, wherein the video includes at least one of medical staff, a patient, a robotic system or instrument, or medical environment; generate, using a second model, a second feature that identifies features in a text prompt; generate, by a third model and based on the first feature and the second feature, an output responsive to the input text prompt, wherein the output comprises at least one of text or media content; determine, based on the first feature and the second feature, a loss with respect to the output; and update at least one of the first model, the second model, and the third model based on the loss.
13 . The system of claim 12 , wherein the output includes a text response and a visual annotation of the video data, and the text response includes a number responsive to the text prompt.
14 . The system of claim 12 , wherein the output includes a text response and a portion of the video data within a time period, the text response and the time period each responsive to the text prompt.
15 . The system of claim 12 , wherein the output includes a text response and an image, the text response includes at least a partial description of the image and is responsive to the text prompt, and the media content comprises the image.
16 . The system of claim 15 , wherein the partial description of the image includes a description of a location of an object in the multi-modal data.
17 . The system of claim 15 , wherein the partial description of the image includes a description of a layout of one or more objects in the plurality of sets of data, and the image is based on the metadata.
18 . The system of claim 12 , wherein the output includes a text response and a plurality of portions of the video data, the text response includes a number responsive to the text prompt, and a number of the plurality of portions of the video data corresponds to the number responsive to the text prompt.
19 . The system of claim 12 , wherein the output includes a text response and a data visualization output, and the text response includes at least a partial description of the data visualization output.
20 . A system, comprising:
one or more processors, coupled with memory, to: receive a text prompt from a user; determine a text prompt feature for the text prompt; identify, based at least in part on the text prompt feature, a plurality of sets of data associated with at least one medical procedure, wherein each of the plurality of sets of data has a different modality, wherein at least one set of the plurality of sets of data comprises depth data; process the plurality of sets of data based on the text prompt feature; and generate, based on the processing of the plurality of sets of data, an output responsive to the text prompt feature.Join the waitlist — get patent alerts
Track US2025218126A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.