Efficient Hybrid Generative AI via Context Filtering/Focused Attention
Abstract
Various embodiments include systems and methods for performing efficient hybrid AI processing. A computing system may be configured to receive multimodal data, determine user intent based on information available to the processor, generate filtered input data by performing context filtering on the multimodal data based on the determined user intent, generate data segments by segmenting the filtered input data based on the determined user intent, and convert the data segments into tokens representing attributes of the data segments. The computing system may assign a priority to each of the tokens based on their relevance to the determined user intent, generate an enhanced prompt based on the assigned token priorities, send the enhanced prompt to an artificial intelligence (AI) model, receive inference results from the AI model, generate a final output based on the received inference results and locally processed data, and present the final output to a user.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computing device, comprising:
at least one memory; and at least one processor coupled to the at least one memory and configured to:
receive multimodal data;
determine user intent based on information available to the at least one processor;
generate filtered input data by performing context filtering on the multimodal data based on the determined user intent to generate filtered input data;
generate filtered data segments by segmenting the filtered input data based on the determined user intent;
convert the filtered data segments into tokens representing attributes of the data segments;
assign a priority to each of the tokens based on their relevance to the determined user intent;
generate an enhanced prompt based on the assigned token priorities;
send the enhanced prompt to an artificial intelligence (AI) model;
receive inference results from the AI model;
generate a final output based on the received inference results and locally processed data; and
present the final output to a user.
2 . The computing device of claim 1 , wherein the at least one processor is configured to:
assign a priority to each of the tokens based on their relevance to the determined user intent by generating a bitmap indicating importance of each token, the generated bitmap including at least one of:
a hard bitmap that includes binary values; or
a soft bitmap that includes a range of values; and
generate the enhanced prompt based on the assigned token priorities by selecting tokens for transmission based on the generated bitmap and a dynamically updated threshold value for token transmission.
3 . The computing device of claim 2 , wherein the at least one processor is further configured to adjust the dynamically updated threshold value based on at least one of:
battery life; network bandwidth; computational resources; or communication costs.
4 . The computing device of claim 1 , wherein the at least one processor is configured to receive the multimodal data by receiving at least two or more of:
visual data; auditory data; textual data; or sensor data.
5 . The computing device of claim 1 , wherein the at least one processor is configured to generate the filtered data segments by segmenting the filtered input data based on the determined user intent by:
generating bounding boxes around specific objects of interest within visual data based on the determined user intent.
6 . The computing device of claim 1 , wherein:
the at least one processor is further configured to compress context data to reduce a data size of the context data in response to determining that a large volume of the context data is relevant to the determined user intent; and the at least one processor is configured to generate the enhanced prompt based on the assigned token priorities by generating the enhanced prompt based on the assigned token priorities and the compressed context data.
7 . The computing device of claim 1 , wherein the at least one processor is configured to generate the final output by integrating the inference results with locally collected context information and user profile information.
8 . The computing device of claim 1 , wherein the at least one processor is configured to present the final output to the user includes at least one of:
displaying information on an electronic display of the end-user device; providing audio feedback; or performing a responsive action.
9 . The computing device of claim 1 , wherein the at least one processor is further configured to:
monitor user interactions with the end-user device to collect attention-based metrics and feedback data; and update user profile information or context information based on the collected attention-based metrics and feedback data.
10 . The computing device of claim 9 , further comprising adjusting operations of the end-user device based on the updated user profile information or the updated context information.
11 . The computing device of claim 1 , wherein the at least one processor is configured to determine the user intent based on the information available to the processor by deriving the user intent from sensory data obtained from one or more input devices.
12 . The computing device of claim 11 , wherein the at least one processor is configured to derive the user intent from the sensory data obtained from one or more input devices by deriving the user intent from gaze detection data obtained from augmented reality (AR) glasses worn by the user.
13 . The computing device of claim 1 , wherein the at least one processor is configured to send the enhanced prompt to the AI model and receive the inference results from the AI model by sending the enhanced prompt to a cloud-based AI model and receiving the inference results from the cloud-based AI model.
14 . The computing device of claim 1 , wherein the at least one processor is configured to send the enhanced prompt to the AI model and receive the inference results from the AI model by sending the enhanced prompt to a local AI model and receiving the inference results from the local AI model.
15 . A method performed by a processor of an end-user computing device of applying multimodal data to an artificial intelligence (AI) model, the method comprising:
receiving multimodal data; determining user intent based on information available to the processor; generating filtered input data by performing context filtering on the multimodal data based on the determined user intent to generate filtered input data; generating filtered data segments by segmenting the filtered input data based on the determined user intent; converting the filtered data segments into tokens representing attributes of the data segments; assigning a priority to each of the tokens based on their relevance to the determined user intent; generating an enhanced prompt based on the assigned token priorities; sending the enhanced prompt to an AI model; receiving inference results from the AI model; generating a final output based on the received inference results and locally processed data; and presenting the final output to a user.
16 . The method of claim 15 , wherein:
assigning a priority to each of the tokens based on their relevance to the determined user intent comprises generating a bitmap indicating importance of each token, the generated bitmap including at least one of:
a hard bitmap that includes binary values; or
a soft bitmap that includes a range of values; and
generating the enhanced prompt based on the assigned token priorities comprises selecting tokens for transmission based on the generated bitmap and a dynamically updated threshold value for token transmission.
17 . The method of claim 15 , wherein generating the filtered data segments by segmenting the filtered input data based on the determined user intent comprises:
generating bounding boxes around specific objects of interest within visual data based on the determined user intent.
18 . The method of claim 15 , wherein generating the final output comprises integrating the inference results with locally collected context information and user profile information.
19 . The method of claim 15 , wherein sending the enhanced prompt to the AI model and receiving the inference results from the AI model comprise at least one or more of:
sending the enhanced prompt to a cloud-based AI model and receiving the inference results from the cloud-based AI model; or sending the enhanced prompt to a local AI model and receiving the inference results from the local AI model.
20 . A non-transitory processor-readable medium having stored thereon processor-readable instructions configured to cause a processor of a computing device to perform operations comprising:
receiving multimodal data; determining user intent based on information available to the processor; generating filtered input data by performing context filtering on the multimodal data based on the determined user intent to generate filtered input data; generating filtered data segments by segmenting the filtered input data based on the determined user intent; converting the filtered data segments into tokens representing attributes of the data segments; assigning a priority to each of the tokens based on their relevance to the determined user intent; generating an enhanced prompt based on the assigned token priorities; sending the enhanced prompt to a local or remote artificial intelligence (AI) model; receiving inference results from the AI model; generating a final output based on the received inference results and locally processed data; and presenting the final output to a user.Join the waitlist — get patent alerts
Track US2026099672A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.